Articles

AI Strategy

When AI Writes the Report, Who Checks the Facts?

When Two of the Big Four Got Caught Publishing AI's Hallucinations as Fact

Bharat Kumar5 min read
KPMG report marked retracted after AI hallucinations were published as fact
Table of contents

Everyone is talking about how much AI can generate.

Almost nobody is talking about what happens when nobody checks what it generated before it goes out under a name people already trust.

A few weeks ago, KPMG gave us the clearest, most public answer yet, and the details, once you look closely, are more unsettling than the headline suggests.

The report that never should have existed

In October 2025, KPMG, one of the Big Four, one of the most trusted names in global business advisory, published a report titled "Redefining Excellence in the Age of Agentic AI."

It was positioned as a forward-looking guide for enterprise leaders navigating agentic AI adoption. It named real, recognizable organizations - UBS, NHS Greater Manchester, Swiss Federal Railways, Transport for London, Emirates - and described, in specific operational detail, how each was supposedly deploying advanced AI agents to transform how they worked.

For eight months, the report sat on KPMG's website as thought leadership.

Then, in June 2026, it quietly disappeared.

What the forensic check actually found

The unraveling started with a technology research firm, GPTZero, which ran a forensic check on the report's citations.

What they found was not a handful of sloppy footnotes.

Of the report's 45 citations, only five pointed to real, intact sources.

Twenty-eight were paraphrased titles bolted onto fabricated components, a real-sounding source dressed up with invented details. Twelve more were too vague to verify at all. Roughly half of the factual claims those citations were meant to support turned out to be false or misattributed entirely.

The specific fabrications, once named organizations started responding, were striking:

  • The report claimed UBS operated AI agents "across investment advisory, risk management and compliance monitoring" through "a composable platform co-developed with Microsoft." UBS told the Financial Times this was "factually incorrect."
  • It described Swiss Federal Railways running AI agents that plan, book, and optimize passenger journeys around real-time conditions and carbon impact. The railway said this was "not accurate."
  • It claimed Transport for London used AI agents to predict congestion and coordinate the city's transport network in real time. TfL called the claim "misleading."
  • It described NHS Greater Manchester using AI agents to organize patient records, automate referrals, and predict hospital readmissions. The NHS said the description "doesn't really align" with how the organization actually operates.

Perhaps the most telling fabrication: the report claimed Emirates had deployed a chatbot named "Sara" capable of rebooking passenger flights.

Sara is real. She is a physical robot assistant Emirates introduced in 2023 for airport wayfinding. She is not a chatbot, and she cannot alter bookings.

The report simply invented a capability for a real product that does not exist.

Elsewhere, the report cited a 2019 press release from a Japanese rail operator as evidence of "agentic AI" in production, a term that did not enter common industry use until 2024. The citation predates the concept it was supposedly proving.

Every named organization that responded said the same thing in different words:

This is not what we do, and we did not say it was.

KPMG pulled the report, opened an internal investigation, and issued a statement about "responsible AI guidelines emphasizing human oversight and verification."

The oversight arrived after the report had already been live, under KPMG's name, for eight months.

This wasn't an isolated incident

Weeks before the KPMG story broke, EY had quietly pulled a different report, this one on cyber risk in loyalty programs, after researchers found it contained fabricated data and bogus references, also apparently AI-generated.

GPTZero was the firm that caught that one too.

Two of the world's four largest professional services firms, publishing under some of the most trusted brand names in business advisory, both had AI-generated fabrications reach public, client-facing distribution within weeks of each other.

As GPTZero's CEO put it at the time: error-filled reports from firms of this scale do not just embarrass the firm. They "poison the well of information" for everyone downstream who assumed the name on the cover meant the content had been checked.

The deeper problem: hallucinations that inherit institutional trust

What happened at KPMG was not a typo that slipped through editing.

It is a specific, dangerous mechanism worth naming precisely.

An AI model generates a plausible-sounding claim.

Nobody independently verifies it against a real source.

The claim gets published under the name of an institution people already trust by default.

Readers - mostly executives, journalists, and other consultants building on the report - do not re-verify a KPMG citation the way they might question a random blog post.

They trust the logo.

That is the exact mechanism that let a fabricated Emirates chatbot and an invented UBS-Microsoft platform sit unchallenged in a global report for eight months.

Not because anyone involved intended to deceive, but because the verification step that would have caught it simply did not exist in the workflow.

The cost here is not bad output.

Bad output gets ignored.

The cost is trusted output that turns out to be wrong, because trust is precisely what makes people stop checking.

The shift nobody's talking about

For the past few years, the AI conversation has been dominated by one word:

More.

More content. More reports. More case studies. More speed to publication.

KPMG's report is, in a sense, a triumph of "more" - a multi-country, multi-industry survey of supposed AI adoption, assembled faster than any team of human researchers could have produced it unassisted.

It is also almost entirely wrong.

"More" was always a stand-in for a deeper assumption:

If AI could produce content faster, someone downstream would still catch whatever needed catching before it went out the door.

KPMG and EY, within weeks of each other, demonstrated that this assumption does not hold on its own.

Volume of output and accuracy of output are not the same axis, and optimizing hard for one says nothing about the other.

Why this keeps happening

The same three patterns from the last two issues, but the stakes just moved from an internal budget overrun to a public, named, client-facing failure.

1. Technology feels tangible. Verification does not.

Choosing which AI tool generates the first draft feels like a real decision.

Deciding, specifically, who is personally responsible for checking every citation and every named claim before it goes external is harder, slower, and far less exciting, so it gets skipped, or assumed to be "someone else's job."

2. Activity gets mistaken for progress.

A multi-country report with named case studies and confident claims looks like serious, comprehensive thought leadership.

Looking comprehensive and being accurate are unrelated qualities.

Quick motion - more pages produced, case studies assembled, a launch date hit - is easy to point to.

Truth is not, and nobody was measuring for it.

3. Ownership quietly disappears.

Someone owned drafting the report.

Someone owned the design and formatting.

Someone owned the relationships with the firm's brand and communications team.

Not one person, apparently, owned checking whether the claim about Emirates' chatbot was true before it reached print.

When responsibility is divided across a dozen functions, verification is exactly the kind of unglamorous, invisible task that has no natural owner, until it fails publicly and everyone wants to know who was supposed to be watching.

The AI Value Equation, extended for a world where AI publishes things

The five questions from earlier issues still apply to every AI investment.

When AI-assisted output goes external, to the clients, regulators, the market, or the public, at that time one more question becomes non-negotiable, not optional:

Who verifies this before it goes public?

Not the AI.

Not "the team," diffusely.

A specific, named person, whose job it is to check every material claim against a real source before publication - and who is personally accountable if that check did not happen.

If the honest answer inside your organization is "nobody, exactly", then you are not producing thought leadership.

You are producing a liability with your name already printed on the cover.

Decision of the Week

Before your next AI-assisted report, memo, board deck, or public claim goes out the door, ask one question in the room, out loud:

If this contains a fabricated claim, who catches it before a client, a journalist, or the organization it describes sees it first, and what happens if the answer is nobody?

Two of the largest, most sophisticated professional services firms in the world just found out what happens when that question does not have an answer.

It happened in public, under their own name, and it happened twice within a matter of weeks.

Thanks for reading Issue #3 of Strategic Insights.

Every week, I will break down a real business event - not to report the news, but to uncover the strategic decisions hiding beneath it.

Because in the AI era, technology is becoming abundant. Sound judgment is not.