Hallucinations in Large Language Models: Challenges and Mitigation Approaches
Abstract
Large language models - like GPT-3.5, GPT-4, LLaMA, or PaLM - have changed artificial intelligence fast; they write smooth, contextsensitive text now. They power things such as chatbots, school helpers, writing aids, legal review tools, even health advice apps. Still, these systems face a big flaw: hallucination - that’s when they say false stuff, make up facts, or give illogical answers but sound sure about it. Because of this, people can’t always rely on them, particularly where mistakes matter most - say, medicine, classrooms, courts. This paper looks into what hallucinations are, along with different kinds you might see. It digs into why they happen - like skewed data, guesswork in predictions, or missing real-world ties. A range of fixes gets reviewed, including pulling facts during output (RAG), learning from user feedback (RLHF), checking claims before sharing, and adjusting models for specific topics. In the end, it’s clear we won’t wipe them out entirely, though mixing methods helps a lot - tapping outside info, tweaking prompts smartly, using people in the loop, plus solid testing standards.
References
Z. Ji, N. Lee, R. Frieske, et al., “Survey of Hallucination in Natural Language Generation,” ACM Computing Surveys, no. 12, pp. 1–38, 2023.
E. M. Bender, T. Gebru, A. McMillan-Major, S. Shmitchell, “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” In Proceedings of ACM FAccT, pp. 610–623, 2021.
Y. Bang, S. Cahyawijaya, N. Lee, et al., “A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity,” arXiv preprint arXiv: 2302. 04023, 2023.
S. Lin, J. Hilton, O. Evans, “TruthfulQA: Measuring How Models Mimic Human Falsehoods,” In Findings of ACL, pp. 3214–3229, 2022.
Refbacks
- There are currently no refbacks.