AI systems are prone to make mistakes, only we tend to call them hallucinations. It appears that the term hallucination was first used in 2000 in a paper called “Proceedings: Fourth IEEE International Conference on Automatic Face and Gesture Recognition,” where it carried positive meanings in computer vision, rather than the negative ones we see today.


In February 2023, Google’s Bard AI Google found out how its own Bard generative AI could produce errors in the program’s first public demo, where Bard stated that the James Webb Space Telescope “took the very first pictures of a planet outside of our own solar system,” when the first such photo was taken 16 years before the JWST was launched.
Once the error became known, Google’s stock price lost as much as 7.7%, or $100 billion, in the next day of trading.
AI hallucinations occur when large language models (LLMs) or computer vision systems generate false or misleading information, sometimes called “confabulations.” These inaccuracies often stem from insufficient or biased training data and can range from minor errors to complete fabrications. LLMs, like those powering ChatGPT and similar chatbots, use statistics to create seemingly correct language, which can make their hallucinations appear plausible and lead to poor decision-making.
AI hallucinations can arise from various sources, including poor data quality, flawed generation methods, and unclear input from users. These hallucinations can manifest as sentence contradictions, prompt contradictions, factual inaccuracies, or irrelevant and random information.
Several real-world examples illustrate this issue. As well as Google’s Gemini chatbot made a false claim about the James Webb Space Telescope, while Meta’s Galactica LLM generated inaccurate and biased scientific information. Similarly, OpenAI’s ChatGPT has faced defamation lawsuits and has been known to produce nonsensical or contradictory responses. Newspapers such as the Guardian have highlighted how sources cited as coming from that newspaper simply do not exist.
AI hallucinations undermine user trust, can lead to the perception of AI systems having human-like qualities (generative anthropomorphism), and can potentially spread misinformation or be weaponised for cyberattacks. These are hard to fix because LLMs are often “black boxes”, and AI infrastructure is expensive. Detecting hallucinations requires careful fact-checking, which can be difficult. However, users can ask the model to self-evaluate or familiarise themselves with the model’s data sources to aid in this process