Back to Blog
Article

Why Corporate AI Projects Fail — And It Has Nothing to Do With the Technology

Boroji
14 min read
Why Corporate AI Projects Fail — And It Has Nothing to Do With the Technology

Photo: DigiFusion

Enterprises have burned $30–40 billion on AI pilots. The models worked. What failed was the size of the bet — and the arithmetic of that mistake is three centuries old and far less forgiving than the technology.

By Boroji Adebayo-Hopewell


Let me be blunt about the record, because the record is brutal.

We have just lived through the most expensive experiment in the history of corporate technology. Enterprises poured an estimated $30–40 billion into generative AI, and in 2025 MIT's Project NANDA delivered the autopsy: roughly 95% of those pilots produced no measurable impact on profit and loss. Not disappointing impact. Not delayed impact. No measurable impact.

I want to handle that number carefully, because an article about companies believing unexamined figures has no business quoting one unexamined. The NANDA study is a working paper, not peer-reviewed, and it rests on around three hundred public initiatives, fifty-two interviews and a hundred and fifty-three surveyed leaders. That is a real piece of work and a modest sample, and it has been repeated so often that most people quoting it have never seen its methodology. Treat 95% as a strong signal about direction rather than a precise measurement — and notice that the study's own buried finding is the more useful one anyway: outcomes were determined by approach, not by technology. Which is precisely the argument of this article, and it survives whether the true figure is 95% or seventy.

The wider numbers tell the same story from different angles. By 2025, McKinsey found that 88% of organisations were using AI in at least one function — adoption is nearly universal — yet only 39% could attribute any effect on earnings to it. Gartner projected that more than 40% of "agentic AI" projects will be cancelled by the end of 2027. The industry even coined a name for the condition: pilot purgatory — the state in which a company runs demonstration after impressive demonstration while its costs, cycle times, and margins stay exactly where they were.

Here's what should trouble you most. The technology works. Today's models summarise, draft, classify, extract, and reason at a level that would have looked like sorcery in 2020. The failure is not in the silicon.

The failure is in the companies bolting it on.

The lesson we already learned — a century ago

When MIT's researchers dug into why the 95% failed, they didn't find a technology gap. They found what they called a learning gap: companies had bolted AI tools onto workflows that were never redesigned, so the tools couldn't hold context, couldn't connect to the processes around them, and couldn't compound.

If that sounds familiar, it should. We have run this exact experiment before.

When factories first replaced their steam engines with electric motors, most simply put the new motor where the old engine had been — driving the same overhead shafts, the same belts, the same layout that steam had dictated for a century. And for nearly thirty years, productivity barely moved. Economists puzzled over why a miraculous new power source showed up everywhere except the output statistics.

The gains only arrived when a new generation of engineers asked a different question — not "where do we put the motor?" but "what should a factory be, now that power can go anywhere?" Small motors at every workstation. Machines arranged in the order of the work itself. The technology took five years to install. The rethinking took thirty. The firms that did the rethinking early didn't gain an edge. They gained an era.

Every generation installs its new power source in the floor plan of the old one — and then blames the power source.

We are living inside that story again, almost line for line. The price of routine thinking — reading, classifying, drafting, reconciling — is collapsing, exactly as the price of mechanical power once collapsed. And most companies are responding exactly as the factory owners of 1900 did: bolting the new motor onto the old shafts, dropping copilots into unredesigned workflows, and then reading in their own results that "nothing happened."

The four illusions that burn the money

Across the failed transformations I've watched, the money almost always dies in one of four ways. Learn to spot them and you've learned most of what separates the 5% who succeed from the 95% who don't.

The Demo Illusion. A demo is the most dishonest room in business — not because anyone lies, but because the data is clean, the use case is bounded, and there are no exceptions, no legacy systems, no angry customer on line two. The tool glides. Then it meets your actual company, and it drowns in the friction the demo carefully removed. The tell is a pilot that succeeds and then can't scale. Scaling didn't break the tool. It removed the clean room you built around it.

The Layer Illusion. This is the belief that intelligence added on top of a messy process can compensate for the mess inside it. The team drowning in email gets an AI that drafts replies faster — so the inbox fills faster. The analyst re-keying data between systems gets help re-keying it more fluently. The friction wasn't removed; it was lubricated — made a little more comfortable, and therefore more permanent. A copilot bolted onto chaos doesn't produce order. It produces chaos with better grammar.

The Substitution Illusion. This one lives in the business-case spreadsheet: the process uses 20 people, agents can do 40% of the tasks, so redeploy eight people and bank the savings. It sounds rigorous and it's quietly wrong, because it automates the tasks while leaving the architecture — the handoffs and queues between tasks — completely intact. And the handoffs are where most of the delay and most of the waste actually live. Speed the tasks up and leave the queues in place, and the work just arrives at the same bottlenecks faster, and piles up higher.

The Savior-Vendor Illusion. This is the hope that someone outside your company can take responsibility for the mess inside it. But no vendor has ever seen your exception folklore. No model ships knowing that your quotes over $50,000 secretly wait four days for a director who is always in meetings. A tool can be bought. But the thing that has to be re-engineered is you — and you are the only one who can do it.

A copilot bolted onto a broken process doesn't produce intelligent order. It produces chaos with better grammar.

The fifth illusion, and the one that actually empties the account

Those four explain why a pilot underperforms. None of them explains why underperformance turns into catastrophe, and that distinction matters more than anything else in this article. Plenty of firms have run a disappointing pilot, shrugged, and gone back to work. The ones that got hurt were not the ones whose technology worked least well. They were the ones who had committed too much to it before they knew whether it worked at all.

This is a sizing failure, and sizing has a mathematics that predates every model in the field. Daniel Bernoulli set it out in 1738, arguing that what a rational actor maximises is not expected winnings but the logarithm of wealth — the rate at which it compounds. John Kelly Jr. formalised the consequence at Bell Labs in 1956: for any given edge, there is exactly one fraction of your resources that maximises long-run growth, and the curve around that optimum is not symmetrical.

The asymmetry is the entire lesson. Take a genuine advantage — a fifty-five per cent chance of success, paying two to one — and simulate it across ten thousand runs of a hundred decisions. Commit half the optimal fraction each time and you keep about seventy-six per cent of the growth you could have had. Commit twice the optimal fraction and your growth rate does not merely halve; it turns negative, and more than half of all runs finish below where they started, despite the advantage being real throughout. Underbetting costs you time. Overbetting costs you the firm. Those two errors look symmetrical on a spreadsheet and are nothing of the sort in life.

Now change the nouns and read it again. Replace capital with the discretionary transformation budget. Replace the win rate with the probability that this particular automation initiative delivers. Replace the payoff ratio with recoverable margin over cost at risk, and replace the trade with the engagement. Every sentence holds.

And here is what firms actually did. They committed a quarter of a discretionary budget and eighteen months of senior attention to a single un-instrumented flow, on the strength of a demonstration. That is not a technology decision. It is a bet at roughly twice the defensible size, on a probability nobody in the room had measured — placed by people who would never sign off a currency hedge on the same evidence. The models performed as advertised. The size was lethal.

I want to be careful about how far this travels, because the failure mode of a good analogy is being pushed one step past where it holds. Kelly's result assumes bets that repeat, that can be re-sized freely, and that are independent of one another. A firm's automation programme is none of those things: it is lumpy, it is partly irreversible, and its initiatives fail together rather than separately, because they usually fail from the same unmapped process. So this is not a formula to run your capital budget through. It is a ceiling and a discipline — a way of establishing the number you will not exceed, before enthusiasm establishes it for you.

Which leads to the only question that matters before any of it: how would you know? The optimal fraction depends on a probability, and a probability you have not measured is not an input, it is a wish. Almost nobody in this field measures it. That is the actual scandal, and it is the reason the next section is about legibility rather than about tools.

The real constraint isn't the model. It's legibility.

Here's the idea I most want you to carry away, because it reframes the entire problem.

An AI agent can only execute a process to the degree that the process is legible — mapped, measured, bounded, and honest about its exceptions. Below that threshold, deploying AI produces chaos at machine speed. Above it, deploying AI produces compounding advantage. And that threshold is crossed by engineering, not by purchasing.

A legible process can answer five questions in writing. What are the real steps — not the version in the binder nobody's opened since the audit, but the actual path, including the workaround someone invented in 2019 that everyone now quietly relies on? What exactly goes in and comes out of each step? Where does judgment genuinely live — which decisions follow rules, and which need human context no rule captures? What are the exceptions, and how often does each occur? And how is success measured — in numbers your P&L can feel, or in demo applause?

Most companies discover their operations can't answer these. Not because they're badly run, but because they run on folklore — tribal memory, exceptions handled by instinct, boundaries negotiated meeting by meeting. And an illegible process is exactly what an AI agent cannot execute. This, and not model quality, is why the pilots died.

The uncomfortable part: you cannot self-assess this alone

There is a trap sitting immediately after that paragraph, and most readers walk into it.

Having read the five questions, you will now form a view of how legible your own operation is. That view will be too generous, and this is measurable rather than rhetorical: in one study, only about 22% of enterprises that assessed themselves as AI-ready met the threshold when an outside party checked. Senior leaders systematically score their firms higher than the people doing the work do — not from vanity, but because the view from the top is the view of the process as designed, and the view from the floor is the process as performed. Both are honest. Only one of them is what an agent will actually meet.

Which suggests a test more useful than any score. Take the five questions above — or any legibility instrument you like — and have three people answer independently: yourself, whoever runs the work day to day, and whoever owns it on paper. Then compare.

The gap between their answers is your real finding. If your operations manager rates your exception handling far lower than you do, you have not discovered a difference of opinion. You have measured the distance between the official account of the work and the actual one — and that distance is exactly the space where an AI pilot goes to die.

The frontier of AI is in San Francisco. The frontier of AI value is in your workflow documentation — and only one of those is within your control.

That last line is the strategic heart of it. You cannot out-engineer the frontier AI labs. You can absolutely out-engineer your competitors' willingness to map their own operations honestly. Model capability stopped being the binding constraint years ago. Your legibility is the frontier now — and it's the one frontier entirely within your reach.

What the 5% actually do differently

The companies that escape pilot purgatory aren't smarter, and their vendors are the same vendors. What separates them is sequence. They refuse the reflex to deploy first and fix later, and they run the order the physics demands:

Map the process as it actually runs. Measure its real cost and cycle time. Re-engineer it — kill the steps that only exist to compensate for other broken steps, and decide deliberately which work stays human and which goes to the machine. Then automate what's left. And govern it forever, because a process left alone starts silting up the day you stop watching.

Notice that the first three steps consume most of the calendar and almost none of the technology budget — which is exactly why the plug-and-play culture skips them, and exactly why skipping them turns the remaining money into a bonfire. McKinsey's own data makes the point without meaning to: the small group of companies actually extracting real earnings from AI were about three times more likely than everyone else to have fundamentally redesigned their workflows before scaling the technology.

They didn't buy better AI. They built better engines, and then applied the power.


The takeaway

  • The great majority of enterprise AI pilots produce no profit impact — and the cause is almost never the technology.
  • It's the electrification mistake, repeated: installing a powerful new engine inside the floor plan of the old one.
  • The money dies in four predictable ways: the Demo Illusion, the Layer Illusion, the Substitution Illusion, and the Savior-Vendor Illusion.
  • The real constraint is legibility — how well-mapped and honest your processes are. It's not something you can buy, and it's the one frontier fully within your control.
  • You will overrate your own legibility. Have three people score it independently; the gap between them is the finding.
  • The winners run a sequence — Map, Measure, Re-engineer, Automate, Govern — and spend most of their effort before the technology ever arrives.

Find your legibility threshold — four minutes, free

Everything above is diagnosis without an instrument, which is where most articles of this kind stop.

The FrictionIQ assessment is the instrument. Twelve questions, one per operational domain, scored against written anchors so that two people reading the same business reach the same number. It returns a readiness band, the domains that are blocking you, and the single question about your own operation you could not answer.

Two things it does that matter here. Any domain scoring zero caps your reported readiness regardless of your total — because averages hide blockers, and a firm with excellent process discipline and no data authority is not eighty per cent ready, it is blocked. And at the end you can send the same twelve questions to two colleagues, so you get the divergence test described above rather than one flattering view.

It is genuinely free, there is no call, and if your score comes back low enough I will tell you plainly not to automate yet. That is the least commercial sentence in this article and the most important one.

Take the assessment


This article draws on the argument of my book, The Thermodynamic Firm: Eliminating Entropy in the Age of AI Agents. If you're deploying AI this year — or explaining to a board why last year's pilots went nowhere — the book is the field manual for crossing the legibility threshold and being one of the few. Pre-order the series →


Boroji Adebayo-Hopewell is a techno-economist and enterprise systems architect who has spent two decades rebuilding the operating engines of companies from ten-person shops to global enterprises.


BA

Written by

Boroji Adebayo-Hopewell

Founder and Lead Architect, Digital Fusion Labs

Founder and Lead Architect of Digital Fusion Labs — writing on System thinking, AI automation, business development, and digital media strategy for operators who need answers.