Evidence · Incident

When hallucinations reach the official record

A pattern has formed across the last few months of AI incident reporting, and it has nothing to do with attackers. Organisations with mature review processes have put AI-fabricated content into records that are supposed to be authoritative — and discovered it only when someone outside the organisation checked.

Four cases, one shape

  • Sullivan & Cromwell. The Wall Street firm told a federal court that a filing contained errors resulting from AI hallucinations, including fabricated case citations, as reported by The New York Times and The Guardian in April 2026.
  • A sanctioned appellate brief. Bloomberg Law reported a U.S. Court of Appeals sanctioned a Maryland lawyer US$1,000 for a brief drafted with generative AI that cited nonexistent cases.
  • KPMG. The firm withdrew a published report after multiple case studies were found to be inaccurate and apparently AI-generated, according to PCMag and International Accounting Bulletin.
  • The U.S. National Weather Service. Futurism and Vice reported an AI-assisted forecast product for rural Idaho containing town names that do not exist.

A top-tier law firm, an individual practitioner, a Big Four advisory, and a federal agency. Different sectors, different budgets, different review cultures — and the same outcome. That consistency tells you the problem is structural, not a lapse in diligence by any one of them.

Why review catches less than you think

Human review is calibrated to catch human error. A junior associate’s mistakes have a recognisable signature: hedged reasoning, thin analysis, a citation that does not quite support the point. Model output has the opposite signature. It is fluent, confident, formatted correctly, and the fabricated citation looks exactly like a real one — plausible parties, plausible reporter, plausible year. Reviewers who have spent careers reading for weak argument are not equipped to read for invented fact.

Fluency is not a proxy for accuracy, but every review process built before 2023 quietly assumed it was.

The missing control is provenance

In each of these cases, the organisation could not answer a simple question at the moment it mattered: which parts of this document were generated, by what, and who verified them. Not who approved the document — who checked the specific claims a model produced. Without that record, the only way to find a hallucination is for an outsider to try to look something up.

This is the same requirement regulators are converging on for AI systems generally: traceability from output back to the system, the inputs, and the human accountable for the decision. Organisations that build it for their own outbound documents get regulatory readiness as a by-product. Organisations that do not will be reconstructing it under deadline, after an incident, from memory.

The Argorix Evidence Hub records which AI systems touched an output, what policy applied, and who signed off — so provenance is a stored fact rather than a reconstruction exercise when a regulator, a court, or a client asks.

What to put in place

  1. Classify outputs by consequence. A brainstorm draft and a court filing need different controls. Define which document classes may never contain unverified generated content.
  2. Verify citations mechanically. Any factual reference — case, statute, standard, statistic, company — gets resolved against a source of truth, not read for plausibility.
  3. Record the attestation, not just the approval. Who verified which claim, and against what. Approval without attestation is what these organisations already had.
  4. Make disclosure a default, not a decision. Several of these cases became reputational events because the AI involvement was undisclosed, not because it existed.

Sources

Make provenance a stored fact
Start with a 2–4 week assessment.
Start Assessment