04 May 2026 · 7 min · By Jordan Foord
Why 95% of AI pilots fail (and the boring reason it isn't the model)
The models are fine. The problem is everything your organisation wraps around them, and that's fixable.
There’s a number doing the rounds that deserves more attention than it gets. MIT’s NANDA initiative, in its “GenAI Divide: State of AI in Business 2025” study, found that 95% of enterprise GenAI pilots deliver no measurable P&L impact. Not “underwhelming impact”. No measurable impact. Despite an estimated US$30–40 billion poured in.
That figure sits alongside another one from McKinsey’s November 2025 State of AI: 88% of organisations now use AI regularly, but only around 6% achieve significant enterprise-wide impact: the kind that shows up as five-plus per cent of EBIT. Nearly two-thirds haven’t begun scaling beyond pilots at all.
So adoption is nearly universal, and value is nearly nonexistent. If you’re sitting on a pilot that demoed brilliantly and then quietly died, you are not unusual. You are the statistical norm.
It’s not the model
The reflexive explanation is that the technology isn’t ready. It’s a comfortable explanation, because it means the fix is to wait. It’s also wrong.
The MIT study traced the failure to what it called a “learning gap”: tools that don’t retain feedback, don’t fit existing workflows, and organisations that don’t redesign processes around them. The model performed. The pilot demoed. Then it was dropped into a workflow that was never rebuilt to use it, and the workflow won. Workflows always win.
BCG puts a structure on this with its 10-20-70 rule: successful AI deployment is 10% algorithms, 20% technology and data, and 70% people and process. Most companies spend in exactly the inverse proportion. They buy the 10%, wire up the 20%, and assume the 70% will sort itself out.
It doesn’t. BCG’s own data shows 70% of companies have trained less than a quarter of their workforce on AI tools, and 32% don’t define or measure AI value KPIs at all. You cannot scale what you don’t measure, and you cannot measure what you never defined.
What pilot purgatory actually looks like
We run a software company (nollie, an AI CRM for hospitality, live in four markets) and the day-to-day operations are run by a fleet of AI agents: finance close, customer onboarding, support triage, marketing production. We say this not to show off but because it means our opinions about AI failure come from our own incident logs, not a survey.
Here’s one of ours. We built an agent to handle customer onboarding: preparing a new customer’s account end-to-end so they could go live without waiting on us. The agent worked. The pilot was, by any demo standard, a success.
Then it failed in production for six weeks. Not because the model got anything wrong, but because the onboarding workflow still assumed a human was doing it. Approvals sat in someone’s head. The “definition of done” lived in a Slack thread. The agent would finish its work and the process would simply stall, because the next step had never been written down anywhere an agent could find it.
The fix wasn’t a better model. It was redesigning the workflow around the agent: explicit completion criteria, a written runbook, a defined human checkpoint before anything reached a customer. After that, it ran. The technology was identical before and after. The process was not.
That, in miniature, is what’s happening across the 95%. The pilot proves the model can do the task. Nobody redesigns the process the task lives inside. The pilot dies of organisational neglect and the model takes the blame.
The four failure patterns we keep seeing
Across our own operations and the businesses we talk to, the failures cluster:
The pilot has no owner. Someone sponsored it; nobody operates it. When it breaks at 9pm (and it will), there’s no runbook and no name attached. Compare that to how you’d treat a new hire: you wouldn’t onboard them and then never speak to them again.
The pilot has no number. If you can’t say what metric the pilot was supposed to move, it didn’t fail: it was never really tried. This is the BCG 32% in the wild: no KPI defined, so “success” defaults to vibes, and vibes don’t survive budget season.
The workflow wasn’t redesigned. The AI is bolted onto the side of a process built for humans, with human-shaped gaps (verbal approvals, tribal knowledge, “ask Sandra”) that an agent falls straight through.
The pilot was a demo wearing a pilot’s badge. It was built to impress a steering committee, not to survive contact with a real Tuesday. Real pilots are boring: narrow scope, one workflow, defined exit criteria.
None of these are technology problems. All of them are operational problems, which is genuinely good news, because operational problems can be fixed by mid-sized companies on mid-sized budgets. You can’t out-spend Google on models. You can absolutely out-execute the 95% on process.
The uncomfortable bit about external help
Here’s a finding from the same MIT study that we’d quote even if it didn’t flatter our line of work, but in honesty, it does: pilots built with specialised external partners succeeded around 67% of the time, versus roughly 33% for internal builds. Twice the success rate.
We’d add the caveat the study doesn’t print: “external partner” is not a magic word. Plenty of external partners will sell you a strategy deck and a workshop and leave you exactly where you started, minus the fee. The partners who move the number are the ones doing the unglamorous 70% (workflow redesign, runbooks, training, measurement), not the ones presenting about the 10%.
And to be clear: if your need is narrow and your team is capable, you may not need anyone. A single well-scoped internal build with a named owner and a real KPI beats an expensive engagement with neither. Don’t buy help to avoid doing the thinking.
What the 5% do differently
Inverting the failure patterns gives you the playbook, and it is almost embarrassingly mundane:
- One workflow, not a platform. The successful pilots we’ve seen and run pick a single process (invoice matching, support triage, onboarding) and rebuild that process end-to-end with the AI inside it.
- A number before a build. Hours saved per week, days off the close, response time, error rate. Decided before anyone writes a prompt.
- A named owner with a runbook. Someone who knows how it works, watches it weekly, and can say what happened when it broke.
- A human checkpoint where it matters. Our own rule is that no agent takes an unsupervised external action: nothing reaches a customer, a bank, or a regulator without a human in the loop. Slower on paper. Faster in practice, because trust is what lets you scale scope.
- Process redesign as the actual project. The model integration is the easy fortnight. The workflow rebuild is the real work. Budget accordingly: closer to BCG’s 70% than to the 10% you’ll be tempted to spend it on.
Do this next week
Pick your most disappointing AI pilot: the one that demoed well and went nowhere. Spend thirty minutes answering four questions in writing:
- What single metric was it supposed to move?
- Who owns it operationally, by name?
- What changed about the surrounding workflow when it was deployed?
- What happens, step by step, when it produces a wrong answer?
If you can’t answer all four, you’ve found your problem, and it wasn’t the model. Fix those four for one pilot (just one) before you start another. The 95% isn’t a verdict on the technology. It’s a verdict on treating AI like software procurement instead of operational change, and verdicts like that can be appealed.