AI Hallucination Guardrails: When AI Is Right, But Risky
On a recent episode of Evolving with Jessica Nguyen, Eric Dodson Greenberg, general counsel of Cox Media Group, shared a story from a healthcare fraud trial. The lawyer, someone Eric knows, ran everything through AI, from real-time transcripts to depositions. At one point the AI recommended arguing that the defendant sales reps had not understood the law. The lawyer declined, reasoning that ignorance of the law is a weak defense that would let the other side argue the reps did whatever it took to close a sale. The jury hung, and the four jurors who sided with the defense each said afterward that the reps simply hadn't known they were doing anything wrong.
Eric called the story "vivid and ambiguous." The AI had found a real narrative in the facts, yet nobody in the conversation was confident the lawyer had erred by leaving it out. Most discussions of AI hallucinations miss this kind of ambiguity. Guardrails are built to catch outputs that are false, but for an in-house legal team, a correct output applied without business context can be just as costly. Understanding where guardrails help, and where they stop, starts with the failure they were designed for.
What Are AI Hallucinations
An AI hallucination is an output that is fluent and coherent but is fabricated, inaccurate, or unsupported by the source material. Hallucinations range from a subtly wrong date or clause reference to case law, statutes, and quotations that do not exist. Large language models generate text by predicting likely sequences of words from their training data rather than by retrieving verified facts, so a fluent answer and a true one can look identical. Even the most capable generative AI systems confidently confabulate, because better predictions do not change the underlying mechanism.
The risk is sharpest in legal research, where one invented authority can undermine an entire brief. A 2024 Stanford study found that LexisNexis's Lexis+ AI and Thomson Reuters' Westlaw AI-Assisted Research hallucinated between 17% and 33% of the time, even with retrieval built into both products.
What Are AI Guardrails
AI guardrails are the enterprise response to that problem. They are the controls, policies, checks, and automated validations that constrain how an artificial intelligence system behaves before, during, and after it generates an output. Guardrails sit between raw model capability and production use, determining what an LLM can see, what it may do, and what must be verified before an answer reaches a lawyer. For legal teams, guardrails separate a general-purpose chatbot like ChatGPT or Copilot from a system that can be trusted with contract language.
Types of AI Guardrails
Guardrails fall into three categories, defined by when they operate in the AI workflow.
Input Guardrails and Prompt Engineering
Input guardrails shape what goes into the model. Prompt templates standardize how questions are asked, role constraints limit the model to a defined task, and topic boundaries keep out queries the system was not built to answer. The most important input guardrail is context injection, which supplies the governing contract or applicable policy so the model answers from the organization's own documents instead of its training data alone.
Output Guardrails and Validation Checks
Output guardrails review what the model produces before anyone relies on it. They validate claims against source documents, confirm that legal citations resolve to real authorities, flag low-confidence responses, and screen for confidential content. Audit trails belong in this layer as well, recording which sources informed an answer so a reviewer can verify it.
Self-Correction Loops and Feedback Mechanisms
More advanced systems add iterative validation. Retrieval-augmented generation (RAG) grounds an answer in retrieved documents, and a verification step then checks the draft against those sources and revises unsupported statements before delivery. Corrections from human reviewers can also feed back into the system, improving the output on the next matter.
When AI Is Right But Still Risky
The guardrails above test whether an output is accurate and permissible, but none tests whether it is appropriate for the business. That gap is why general-purpose legal AI tools fall short for enterprise legal teams. An output can be accurate, well supported, and compliant while still creating risk, because the system lacks the business context, policies, and precedent behind the right answer for a specific deal. Eric argues that human judgment has to take over here. "It's not treating the output as a conclusion to be adopted," he said. "It's an insight to be metabolized by our judgment and our experience."
Correct Language That Violates Internal Playbooks
An AI drafting tool can produce a limitation of liability clause that is legally sound and market standard while conflicting with the position the company has held for years. Without access to the organization's legal playbook, the model cannot know the clause concedes a point the team has deliberately refused to give up. A reviewer who assumes the draft reflects house positions may never catch it.
Accurate Advice Without Business Context
An AI system can correctly recommend holding firm on a standard indemnity without knowing that the counterparty is a strategic partner or that a prior negotiation nearly collapsed over the same clause. Both facts change the right answer, yet neither appears in the contract.
Compliant Outputs That Create Unintended Exposure
Some outputs are fully compliant while creating commercial risk. One legal team found its AI tool generating more than 30 redlines on agreements with key strategic partners, each legally defensible and collectively damaging to relationships the business depended on. Overly aggressive terms slow procurement, strain vendor relationships, and push the business to bypass legal.
Related: Learn how GCs navigate decisions outside their legal expertise.
Enterprise Impact of AI Hallucinations
Whether an output is false or only missing context, it exposes a legal team to three kinds of consequence, and how well the team limits them decides whether enterprise AI adoption succeeds.
Legal and Compliance Exposure
The most visible consequences so far have come from the courts. In Mata v. Avianca, a federal judge imposed Rule 11 sanctions on lawyers who submitted a brief citing cases ChatGPT had invented. As of September 2026, researcher Damien Charlotin's database tracks more than 2,000 cases worldwide involving hallucinated content in legal filings. Many judges have since issued standing orders requiring lawyers to disclose their AI use. ABA Formal Opinion 512 also makes clear that the duty of competence extends to understanding the limits of generative AI tools. In-house teams face the same exposure in less public forms, such as hallucinated contract language that comes up in a dispute or incorrect regulatory guidance that leads to an audit failure.
Reputational and Client Trust Risks
When an AI-generated error reaches a counterparty, a regulator, or the board, the damage extends beyond the error itself. Business stakeholders who see legal send out flawed work question everything else legal produces, and a counterparty who catches a mistake gains leverage. For a department working to be seen as a strategic partner, one visible error can set that effort back by months.
Operational Bottlenecks from Manual Review
Without guardrails the team trusts, lawyers have to review every AI output line by line, doing the original work plus the verification and eliminating the efficiency gains that justified the investment. It is a common reason AI adoption stalls in legal departments.
Why Guardrails Are Not Enough on Their Own
Better guardrails address much of that exposure, but they have limits. Validation checks catch fabricated citations and unsupported claims, yet they cannot supply context the system never had or decide whether a correct answer serves the business. Jessica Nguyen made the same point in her conversation with Eric, noting that legal is "in the business of getting things right" and that judgment and accountability remain the lawyer's role.
Effective AI deployment therefore pairs guardrails with a strong knowledge foundation, thoughtful system design, human oversight where judgment matters, and clear governance. Teams that treat this as a maturity journey, learning where AI needs a human in the loop, build trust faster than teams expecting one implementation to settle the question.
How In-House Legal Teams Build Trustworthy AI Systems
Context turns a right-but-risky AI system into a reliable one. Guardrails work only as well as the information they can check against. In most legal departments, however, that information is scattered across a CLM, email, Slack, shared drives, and the memories of senior lawyers. A system that cannot see the company's playbooks, its history with a counterparty, or the business priorities behind a request will keep producing answers that are accurate but incomplete.
Trustworthy AI for legal requires bringing that institutional knowledge into one system where guardrails can operate on it. Sandstone connects every person, company, and document tied to a piece of legal work, so agents and lawyers start from full context — that's what we mean by Legal Relationship Management. When a request arrives through intake, Sandstone automatically attaches the relevant playbook positions, counterparty history, prior redlines, and business context before an AI agent drafts a response or routes the work. Agents handle first-pass work grounded in the company's own positions, and judgment calls go to the lawyers best placed to make them.
Learn how Sandstone enables in-house legal departments with AI.
FAQs About AI Hallucination Guardrails
How do AI guardrails differ from AI alignment?
Guardrails are operational controls that constrain an AI system's behavior in production, such as input restrictions, output validation, and review workflows. Alignment refers to the deeper technical work, done largely during model training, of ensuring AI systems pursue the goals their developers intend. Legal teams control their guardrails directly, while alignment is mostly the responsibility of model providers.
Can AI guardrails eliminate hallucinations entirely?
No guardrail system can guarantee zero hallucinations. Layered controls combining input validation, retrieval grounding, output checks, and human oversight can still reduce them to a level acceptable for enterprise use, and they work best when the system is grounded in the organization's own documents and precedent.
Works Cited
- American Bar Association Standing Committee on Ethics and Professional Responsibility. Formal Opinion 512: Generative Artificial Intelligence Tools. July 29, 2024.
- Charlotin, Damien. AI Hallucination Cases database. Accessed September 27, 2026. https://www.damiencharlotin.com/hallucinations/
- Magesh, Varun, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Stanford RegLab and Stanford HAI, 2024. https://dho.stanford.edu/wp-content/uploads/Legal_RAG_Hallucinations.pdf
- Mata v. Avianca, Inc., No. 22-cv-1461 (PKC) (S.D.N.Y. June 22, 2023).
- Nguyen, Jessica, host. "Eric Dodson Greenberg on Relationships and the Future of Law." Evolving with Jessica Nguyen. Sandstone.
