Insights

2026-07-28 · Article

Why AI Initiatives Fail to Scale: Eighty-Four Per Cent Never Left The Pilot

IBM asked CEOs how their AI programmes were going. A quarter delivered the expected return. Sixteen per cent made it out of the pilot. The interesting question is not why, it is which quarter.

By Frans Vermaak, CEO and AI & Data Architect

Why AI Initiatives Fail to Scale: Eighty-Four Per Cent Never Left The Pilot

Two numbers from IBM's 2026 study of chief executives, and then the rest of this article is about what boards do with them.

Twenty-five per cent of AI initiatives delivered the return on investment that was expected of them.

Sixteen per cent scaled across the enterprise.

Sit with the second one. Eighty-four per cent of AI work in these organisations is still, in some form, a pilot. Not failed, necessarily. Not cancelled. Just permanently in the condition of being about to be rolled out.

The comfortable misreading, and why it is wrong in both directions

There are two ways to get this wrong, and boards manage both.

The pessimist reads 25 per cent and concludes AI does not work. That is the same error the last two articles were about: taking a real figure and stretching it past what it measured. The study says a quarter of initiatives met the return that was expected of them, which is as much a statement about expectations as about outcomes. An initiative can create value and still miss a forecast written by someone who had read a vendor deck.

The optimist reads 25 per cent and says a quarter is rather good for an emerging technology, which is also true, and also not a governance position. "Some of it works" is not something a board can act on. Which some?

The useful reading is narrower. Three quarters of these programmes did not do what the business case said they would do, and the business cases were approved by people who are still in post. That is not a technology finding. It is a capital allocation finding, and capital allocation is squarely the board's problem.

Why the pilot is where things go to live forever

The gap between 25 and 16 is the more revealing one. More initiatives hit their return than made it to enterprise scale, which means some things that demonstrably worked still did not spread.

Pilots are easy to start and almost impossible to kill. They are cheap enough not to need real scrutiny, novel enough to be interesting, and attached to someone's professional reputation. Nothing about a stalled pilot triggers a review. It simply continues, consuming a modest amount of money and a large amount of the one resource nobody meters, which is the attention of your best people.

An organisation with forty AI pilots and six deployments does not have an AI strategy. It has forty small, slow, individually defensible commitments and no mechanism for ending any of them.

The subsidised pricing in the last article makes this worse, not better. When compute is cheap, nothing forces the question. A pilot that costs very little per month is a pilot nobody has to justify, right up until the pricing changes and forty of them reprice at once.

What a board should actually ask

Not "is our AI working." That question produces a status update, and status updates are where uncomfortable facts go to be aggregated into a colour.

Ask instead which specific initiatives have been running longest without either scaling or stopping, who owns each one, and what would have to be true for them to be shut down. The answers are usually available in an afternoon and are frequently the first time anyone has assembled them in one place.

The organisations that got value did not get it by picking better technology than everyone else. They got it by being willing to stop things, which is the least fashionable capability in enterprise technology and the only one that reliably compounds.

The artefact: the ROI Reality Check

Complete this per AI initiative, then read them together. One row per initiative. The portfolio view is the output.

The ROI Reality Check: four fields per AI initiative

What was promised, in a number, by whom, and when. Retrieve the document that authorised the spend, not the current narrative about the initiative. Record the figure, the author and the date. If the document cannot be produced, record that.

What has been measured against that number since. Record the actual measured result on a basis comparable to the original claim, and the date of the most recent measurement. If no measurement exists after approval, record that.

Time since the last stage change. Record how long the initiative has been in its current state, whether pilot, limited rollout or production, and the date that state was entered.

The stop condition. Record the specific observable event that would cause this to be shut down, and the named individual with authority to do it. If either is absent, record that.

How we hold ourselves to this

Tech Sight runs its own AI tooling and we have killed things. A document-search capability we had built and were fond of was retired when we measured actual retrieval quality against what we had claimed for it internally, and the two did not match. The work was not wasted, but continuing it would have been.

We record the stop condition at the point of approval, before anyone is invested in the outcome, because that is the only moment when it can be written honestly. Afterwards it becomes a negotiation with people who would like their project to survive, and that negotiation has one predictable winner.

For the record

The number that will circulate from the IBM study is 25 per cent, because it is the one that supports whichever argument the person quoting it already held. It will be trimmed to "three quarters of AI projects fail" within a month, dropped of its attribution shortly after, and by summer it will be a fact that everybody knows and nobody sourced. You have watched this happen twice in this series already.

The figure that should worry a board is the other one. Sixteen per cent scaled means eighty-four per cent of the AI work inside these organisations is sitting in a state that has no natural end, consuming budget that renews by default, staffed by people who are not doing something else.

That is not a bet on artificial intelligence. It is a portfolio of unclosed options, and the cost of holding it is not the subscription. It is everything those teams were not building instead.


Sources: IBM Institute for Business Value, "CEOs are Reshaping C-suite Roles for the AI Era," annual CEO study, published May 2026, surveying 2,000 chief executives globally and reporting that 25 per cent of AI initiatives delivered the expected return on investment and 16 per cent scaled enterprise-wide (note: the study measures returns against expectations set at approval, which is a statement about both outcomes and forecasting quality; formulations such as "three quarters of AI projects fail" are not what the study reports); OpenAI audited 2025 financial statements as obtained by Ed Zitron and independently verified by the Financial Times, referenced here only for the cost-to-revenue ratio discussed in article 2 of this series; King V Code on Corporate Governance, IoDSA, Principle 10 (data, information and technology governance; oversight proportionate to risk); Companies Act 71 of 2008, section 76 (directors' standard of conduct, in respect of capital allocation oversight).

Book an AI Audit

More from Insights

The Single Vendor Trap: Testing An AI Exit Before You Need One

2026-07-28

The End of Cheap Inference: The Repricing Risk Nobody Has Modelled

2026-07-28

The OpenAI Loss Discrepancy: Reading an AI Vendor's Accounts Before Someone Reads Them To You

2026-07-28