In the rapidly evolving world of conversational AI, especially in voice agents, the promise accent and noise testing of Retrieval-Augmented Generation (RAG) is alluring. However, as companies like Suprmind, Air Canada, and giants like OpenAI explore voice AI solutions, new failure modes emerge — notably, the phenomenon I term hallucination on hallucination.
This post unpacks this layered problem, focusing on seven failure points unique to voice agents, the limits of RAG technology, knowledge base (KB) hygiene, and the indispensability of live tools for customer-specific facts. We’ll also examine strategies such as high-precision entity confirmation and readback to mitigate error compounding.
Understanding Hallucination in RAG
Hallucination in generative AI refers to when models produce plausible-sounding but factually incorrect or nonsensical information. RAG, introduced to reduce hallucination by grounding generation in retrieved documents, itself is not immune to errors.
But what happens when retrieval steps produce flawed inputs, and the generator compounds these errors by hallucinating further? I call this hallucination on hallucination — a critical failure mode especially pernicious in real-world voice agents, where users expect precise, context-sensitive answers.
Seven Failure Points in Modern Voice Agents
Based on extensive experience shipping voice AI systems with telcos and retail giants, and analyzing thousands of real call snippets (e.g., “B three one seven two”), here are seven key failure points where hallucination on hallucination can arise:
Flawed Speech-to-Text (STT) transcription: Mishearing or mis-transcribing key data leads to a corrupted input for downstream retrieval. Imperfect Retrieval Premise: RAG’s retrieved documents may be outdated, irrelevant, or contradictory due to poor KB hygiene. Ambiguous Query Expansion: Automatic query reformulations can widen the retrieval scope incorrectly, pulling in noise. Generation Overreach: The generative model attempts to bridge knowledge gaps, sometimes hallucinating plausible-sounding facts. Compounded Error Feedback Loops: Misleading generations influence user inputs, reinforcing earlier mistakes. Text-to-Speech (TTS) Synthesis Errors: Mispronunciations or unnatural emphasis can confuse users, leading to misunderstanding. Insufficient Entity Confirmation: Lack of explicit readback of critical data leads to silent failures and trust erosion.Table: Summary of Failure Points and Impact on Hallucination
Failure Point Description Effect on Hallucination Speech-to-Text Errors Misheard words, especially numerics and entities. Initial data corruption, feeding false premises. Retrieval Premise Failure Outdated or irrelevant KB docs retrieved. Generative model hallucinates on wrong context. Query Expansion Errors Overbroad or misdirected queries to the KB. Noise retrieval, further compounding confusion. Generation Overreach Model fabricates plausible answers beyond evidence. Incorrect outputs presented as fact. Error Feedback Loops User repeats or accepts wrong data. Compounds hallucinations over session. Text-to-Speech Issues Mispronunciations affecting understanding. User misinterpretation and interactive errors. Entity Confirmation Lapses No explicit readback or verification. Errors go unnoticed, reducing trust.RAG Limits and the Importance of Knowledge Base Hygiene
The root of hallucination on hallucination often lies in a flawed retrieval premise. RAG models rely heavily on the premise that retrieved documents are both relevant and accurate. However, knowledge bases in customer service environments are frequently plagued by:
- Stale information after policy or product changes. Contradictory documents without clear versioning. Incomplete metadata that hinders precision ranking.
Dirty knowledge bases set the stage for RAG to retrieve problematic inputs. When the generator then extrapolates from these ambiguous or incorrect documents, hallucination compounds. This is especially damaging in voice agents, where users often receive answers without visual transcript to verify in real-time.
Suprmind’s approach to combating this hinges on vigilant KB hygiene, continuous document curation, and sophisticated metadata tagging to maintain freshness and context. Without these, hallucination risk remains unacceptably high.
Live Tools as the Source of Truth for Customer-Specific Facts
This is where companies like Air Canada have invested heavily. General KB documents can only go so far in delivering personalized, up-to-the-minute information for flight status, booking changes, or https://bizzmarkblog.com/my-callers-claim-another-agent-promised-a-discount-how-should-the-bot-respond/ loyalty program data. Live customer-specific systems accessed through APIs serve as the source of truth.
Integrating RAG with live tools requires:

- Robust API call orchestration within the conversational flow. Fail-fast mechanisms when real-time data is unavailable. Fallback paths that explain limitations clearly.
Ensuring that the voice agent references live tools rather than hallucinating up-to-date customer facts is essential for trust and user satisfaction.
High-Precision Entity Confirmation and Readback
OpenAI’s evolving voice agent iterations underscore that high-precision entity confirmation and explicit readback remain among the best guardrails. Simply relying on prompt engineering to keep hallucination in check doesn’t suffice.
Key techniques include:
Explicit readback of critical entities such as booking numbers, addresses, or account IDs. For example, confirming “I have your seat as 12C; is that correct?” Multiple-channel confirmation: combining STT confidence scores, retrieval verification, and live tool confirmation before delivering responses. Threshold-based dialog strategies: if confidence dips below a threshold, escalate to a human or ask clarifying questions.These practices close the error loops that otherwise feed hallucination on hallucination in RAG-based voice agents.
The ACL 2025 Paper and Future Research Directions
The upcoming ACL 2025 conference features a notable paper examining error compounding in voice-based RAG systems. Its main contributions include:

- Quantitative metrics identifying points where retrieval errors trigger generator hallucination cascades. Proposals for multi-stage verification pipelines incorporating live tools and entity confirmation. Novel dataset release capturing real telephony audio with annotated hallucination events.
This research validates many lessons that implementations at Suprmind and Air Canada have already discovered through production experience. Emphasizing metrics that measure factual correctness over mere tone or fluency is key.
Summary and Best Practices
Hallucination on hallucination in RAG-based voice agents is a multifaceted problem rooted in upstream errors, flawed retrieval premises, and compounded generative overreach. To minimize its occurrence, companies must:
- Maintain knowledge base hygiene rigorously. Leverage live tools as sources of truth for customer-specific facts. Implement high-precision entity confirmation and readback steps. Monitor and control failure points along the entire pipeline — from STT through TTS. Adopt evidence-based metrics focusing on truthfulness and factuality, not just conversational tone.
As OpenAI, Suprmind, and Air Canada demonstrate, success in voice AI depends on engineering across these layers — not just smarter prompts.
What is the source of truth for this analysis?
This post draws upon my 12 years leading contact center AI implementations, ship records of IVR-to-voice AI migrations, real telephony audio evaluation suites, public tech disclosures from Suprmind, Air Canada customer service innovations, OpenAI voice product iterations, and recent ACL 2025 research insights.
Stay tuned for more deep dives into pitfall avoidance in voice agents and best practices for robust conversational AI.