Insights

2026-07-27 · Article

Ghost AI Numbers: Auditing the AI Statistics in Your Board Pack

Two of the most-quoted AI statistics in circulation fall apart the moment you open the source.

By Frans Vermaak, CEO and AI & Data Architect

Ghost AI Numbers: Auditing the AI Statistics in Your Board Pack

There is a number in your last board pack that nobody in the room could source. Not the finance figures, which have an audit trail a regulator could follow. The AI number: the one in the strategy slide that said something like "95% of AI projects fail" or "85% of AI initiatives never deliver." It arrived with the authority of a fact and the provenance of a rumour, and it may already have shaped a decision.

This is not an argument that AI is working better than the pessimists claim, or worse. It is an argument about a specific and fixable failure: boards are making capital decisions on statistics they have not traced, cannot source, and in at least one common case are quoting in a form the original author never wrote. Two examples, both of which you have probably seen on a slide.

The number everyone cited and few read

In 2025, MIT's Project NANDA published a report called "The GenAI Divide: State of AI in Business 2025." Its headline finding, that roughly 95% of enterprise generative-AI pilots delivered no measurable business return, went around the world in a week. Fortune ran it. It landed in board packs and consultancy decks across every sector, usually rendered as the flat claim "95% of AI fails."

Now open the report. The 95% rests on a sample of around 300 publicly disclosed deployments, 52 structured interviews, and 153 survey responses from senior leaders. That is a legitimate piece of qualitative research. It is not a census of enterprise AI, and it was never presented as one. The distance between "in our sample of 153 leaders, most could not point to measurable returns yet" and "95% of AI fails" is the distance between a finding and a slogan. Analysts who did read it said so; the analysis firm Futuriom published a piece headlined "Why We Don't Believe MIT NANDA's Weird AI Study." None of that caution survived the trip to the boardroom. The number did.

The MIT figure is not a fabrication. It is something more useful to understand: a real finding, from a small sample, stretched into a universal law by people who quoted the headline and skipped the method.

The number that was never what you think it was

The second example is worse, because the original does not say what the boardroom version claims at all.

You will have heard that "Gartner says 85% of AI projects fail." It is one of the most repeated statistics in enterprise technology. Trace it back and you find a Gartner forecast from 2018, which stated that through 2022, 85% of AI projects would deliver erroneous outcomes due to bias in data, in algorithms, or in the teams managing them. Erroneous outcomes. Not failure. Not abandonment. A specific, technical prediction about bias degrading results, dated and bounded to a window that closed years ago.

Somewhere between 2018 and your board pack, "will deliver erroneous outcomes due to bias" became "fail," the date dropped off, and the qualifier vanished. What remains is a phantom statistic: a number that persists through repetition rather than sourcing, cited by people linking to other people who never linked to the original either. It is quoted as a current fact. It is a distortion of a seven-year-old forecast about a different thing.

If a management accountant presented a financial figure with that provenance, they would not survive the audit committee. The AI figure gets waved through because it is about technology, and technology numbers still travel on a different, laxer standard of proof.

Why this keeps happening

Statistics like these spread because they are useful before they are true. "95% of AI fails" is useful to a vendor selling the 5% solution, to a consultant selling a de-risking engagement, and to an executive who wants cover for caution. "85% fail" is useful to anyone arguing for more governance budget. A number that flatters the argument you were already going to make does not get interrogated. It gets forwarded.

The board pack is where this lands, because a slide compresses a claim to its headline and strips the provenance by design. There is no room on the slide for "153 survey responses" or "2018 forecast, erroneous outcomes, not failure." So the caveats fall away, and what reaches the directors is the number alone, wearing the borrowed authority of MIT or Gartner. The name does the persuading. Almost nobody clicks through to check whether the name actually said the thing.

The artefact: the Stat Provenance Test

Run this. Apply the four checks below to any AI statistic before it earns a place in a board decision, holding it to the provenance standard you already hold a financial number to. If a number fails one, it does not drive the choice, and most fail at least one.

The Stat Provenance Test artefact: four checks - named source, actually read, says what we claim, sample and stakes

Named source. Can you name the specific report, its author and its date, or only "studies show" and "Gartner says"? A number without a named, dated source is an opinion wearing a lab coat.

Actually read. Has anyone on your side opened the source itself, not the article that quoted it? The headline and the report frequently disagree, and only one of them is evidence.

Says what we claim. Does the source actually make the claim attributed to it? The Gartner figure did not. Read the original sentence and check that your slide reproduces it, qualifiers and all.

Sample and stakes. What is the sample size and how was it gathered, and would you bet this specific decision on that method? A finding from 153 survey responses may be worth reading and still be far too thin to close a factory or fund a programme.

A few minutes against a number that is about to move real money. Set that against the cost of a strategy built on a statistic that dissolves the moment a journalist, a competitor or a shareholder pulls the same thread you did not.

How we hold ourselves to this

We publish AI research, so it is only fair to say we run this test on our own output before it reaches a client. Every statistic in a Tech Sight board-facing document carries a traceable source, opened and checked against the claim it supports, or it does not ship. The discipline is boring and it is the entire product: an advisory firm that cannot source its own numbers is selling the exact problem this article describes. The Stat Provenance Test is not something we recommend to boards and exempt ourselves from. It is the gate our own work has to pass first.

For the record

The comfortable reading of this is that a couple of statistics got exaggerated, which happens, and the world moves on. That misses the mechanism. These numbers did not slip through despite scrutiny. They spread precisely because there was none: because a recognised name on a slide reads as a substitute for evidence, and because checking a number is dull work that no one is thanked for and everyone assumes someone else did.

A board's job is to be the someone else. The next time an AI statistic appears in a pack with the power to move a decision, the useful question is not whether it is high or low. It is simpler, and almost nobody asks it: who counted, how, and can we see it. If the answer is a shrug, you do not have a fact. You have a number that was sold to the person who put it on the slide, and passed to you at face value.


Sources: MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," 2025, reporting that approximately 95% of enterprise generative-AI pilots showed no measurable profit-and-loss impact, based on a stated sample of roughly 300 publicly disclosed AI deployments, 52 structured interviews, and 153 senior-leader survey responses (as reported by Fortune, 18 August 2025, and the report's own methodology section; sample figures and the "GenAI Divide" framing are the report's own); critical commentary on the study's methodology and sample, including Futuriom, August 2025; Gartner press release, February 2018, forecasting that "through 2022, 85% of AI projects will deliver erroneous outcomes due to bias in data, algorithms or the teams responsible for managing them" (note: the original claim concerns erroneous outcomes from bias, not project failure or abandonment, and is frequently misquoted as an "85% failure" rate); King V Code on Corporate Governance, IoDSA, Principle 10 (data, information and technology governance; oversight proportionate to risk).

Book the two-hour diagnostic

More from Insights

People Will Take the Bot. They Cannot Find the Door.

2026-08-26

The 45 Jobs That Were Not Redundant

2026-08-26

Nobody Published the Denominator

2026-08-26