Anchor LotusConsulting
  1. Home
  2. Insights
  3. AI

Nobody Fails an AI Pilot Any More. That's the Problem.

The demo worked, the tool shipped, everyone clapped — and the numbers didn't move. Here's why nearly every AI pilot fails to pay for itself, and the four questions that fix it.

The demo works. The team claps. The tool goes live. And the P&L doesn't move a penny. The gap between a working pilot and a captured return is now the defining problem in AI implementation — and it has almost nothing to do with technology.

Here is the strangest thing about the current state of AI in business: the pilots keep working.

The model performs. The demo lands. Someone in the room says the words every project sponsor wants to hear — that would have taken me all afternoon. The tool goes live, the initiative is marked complete, the slide turns green, and everyone moves on feeling rather good about it.

Then the quarter closes, and nothing has changed.

MIT put a number on this that should stop any business owner cold. Drawing on 52 executive interviews, surveys of 153 leaders and analysis of 300 public deployments, the research found that 95% of enterprise AI pilots produced no measurable P&L impact.

But the number is not the interesting part. The definition is.

MIT did not define failure as a technical failure. They defined it as a value-capture gap: pilots that shipped and produced no enterprise-level financial return. These were not projects that broke, ran over, or got cancelled. These were projects that worked — and delivered nothing.

That is a far more uncomfortable finding than "AI is hard", because you cannot fix it by buying a better model, hiring a better vendor, or waiting for the next release. The failure is happening after the technology succeeds.

Part OneThe mechanism — you didn't remove the work, you moved it

Once you see how this happens, you cannot unsee it.

The overwhelming majority of AI deployments automate a task that sits inside a process nobody has agreed to change.

So the first draft now takes four minutes instead of forty. Genuinely useful. But that draft still routes to the same reviewer, still waits for the same Thursday meeting, still collects the same three approvals — and now somebody has to check the AI's output, which is a job that did not previously exist.

You have not removed the work. You have relocated it, and attached a licence fee.

This is the honest explanation for something I hear constantly: teams reporting that AI "saves us loads of time," while cycle time, headcount and margin sit almost exactly where they were a year ago. Both statements are true simultaneously. Time was genuinely saved. It simply had nowhere to go, so it dissipated — absorbed into slightly less pressure, slightly earlier finishes, slightly more slack in a system that was never re-cut to exploit it.

The research bears the mechanism out precisely. MIT found that generic tools — the ChatGPT-shaped ones — are piloted almost universally: roughly 80% explored, close to 40% deployed. But embedded, workflow-specific tools cross into production only around 5% of the time.

Read those two figures together, because that contrast is the entire story. The tools that are easy to adopt don't change the business. The tools that would change the business are hard to adopt. The organisations stuck in the 95% are not the ones that failed to try AI. They are the ones that only ever tried the easy half.

Part TwoWhat the 5% actually do differently

Here is the genuinely encouraging part, and the reason I don't think this is a story about big companies with big budgets winning.

The differentiator is not technology. It is not scale. It is a decision.

High performers are nearly three times more likely to fundamentally redesign their workflows around AI — roughly 55% of them, against about 20% of everyone else. They decline to bolt AI onto the existing process. Instead they ask a harder question: if the constraint that originally shaped this process no longer exists, what should the process look like now?

171%  ·  5.1 months

Average reported ROI on agentic deployments, and the median payback period on enterprise agent deployments. Around three-quarters of executives report a return within the first year.

What makes this compelling as evidence is that it shows up in industries with essentially nothing in common.

Financial services

Klarna's customer-service agent absorbed roughly the workload of 853 people and about $60m in cost. JPMorgan Chase removed some 360,000 hours of manual work from operations annually. Fraud-detection deployments report returns around 77%, largely by catching patterns humans miss while cutting false positives — two things that usually trade off against each other.

None of these are chatbots bolted onto an unchanged queue. In each case the queue itself was reconceived.

Healthcare

Clinical documentation agents have cut documentation time by around 42%, with organisations reporting roughly $3.20 returned for every $1 invested inside 14 months.

Notice where that value comes from. It is not that clinicians type faster. It is that the shape of what happens after a patient interaction was redesigned — which is why the return is measurable at organisational level rather than disappearing into individual convenience.

Retail

Retailers running agentic inventory management report around 14% increases in online sales, driven by availability and dynamic pricing.

That number cannot be achieved by giving a merchandiser a better dashboard. It requires pricing and replenishment decisions to move to a different cadence altogether — decisions that used to be weekly and human becoming continuous and supervised. The workflow changed; that is why the money appeared.

Creative, media and agencies

More than 90% of US agencies have now moved from experimental pilots to daily practical execution, visibly reshaping creative ideation, pitch preparation and production timelines.

The agencies that gained real margin from this were not the ones producing identical deliverables more quickly. They were the ones who changed what the deliverable is — more concepts explored before a pitch, iteration moved earlier in the process, the expensive human hours redeployed to judgement rather than production.

Beauty and fashion

Dove's decision to use generative AI to expose algorithmic bias — rather than to manufacture yet another unattainable model, as several competitors did — is the same principle showing up in brand strategy rather than operations. The technology changed the strategic position, not merely the output. That is workflow redesign wearing different clothes.

Smaller operators

For SMEs the pattern is remarkably consistent: the wins are in the back office. Invoice processing, claims handling, query routing. Early adopters report 20–30% faster workflow cycles and administrative overhead reduced by up to 40%.

This matters because it is where process alignment is easiest — a genuinely useful principle for anyone without an enterprise budget. Start where the process is simplest to change, not where the technology is most impressive.

Six sectors. Almost no common ground between a clinic, a fashion house and a payments business. One shared pattern: value appeared exactly where somebody was willing to change how the work happens.

Part ThreeThe four questions I ask before anyone builds anything

I have watched this go wrong often enough to be blunt about it — including when the person getting it wrong was me.

I once spent the better part of three months helping a client automate a reporting process. The work was good. Technically it was clean. We took a build that consumed two days of manual effort down to about four minutes.

The value captured was close to zero.

Why? Because that report fed a monthly meeting that still ran monthly, attended by the same people, making the same decisions on the same cadence. We had made the inputs to an unchanged decision arrive faster. Nobody's day changed. No decision was made sooner. Nothing downstream moved at all.

It was a technical success and a commercial irrelevance, and it taught me considerably more than any smooth delivery has. It is why I now run every proposed project through the same four questions before a single line of anything is built.

1. What work disappears entirely?

Not "gets faster" — stops happening.

Name the specific task that no longer exists once this is live. Write it down. If you cannot name one, you are not removing a step, you are adding one, and the business case is already fictional.

This single question kills more bad AI projects than any technical review I have ever sat in.

2. Whose day actually changes — and into what?

Name the person. Name what they do with the reclaimed time instead.

"Higher-value work" is not an answer; it is a way of avoiding the answer. If nobody's calendar looks materially different ninety days after go-live, the saving will not land anywhere you can measure. It will be absorbed as slack — which is not worthless, but it is not a return, and you should not have written it on a business case as one.

3. Where does the saving land, and who owns that number?

Choose one metric before you build: cycle time, error rate, cost per unit, throughput, conversion, revenue. Give it one named owner. Measure the baseline first.

This is the discipline MIT's findings point at most directly — the most common root cause of stalled deployments is the absence of a measurable business objective tied to the initiative from day one. Without a production success metric there is no forcing function for the last mile, and the last mile is where all the difficulty lives. Projects without one don't get cancelled. They get quietly declared finished at 80%, which is exactly the state that produces a working pilot and no return.

4. What breaks at ten times the volume?

Governance, audit trail and error handling belong in the design, not in a remediation project eighteen months later.

Pilots move fast precisely because governance is thin — and that is entirely appropriate for a pilot. The failure is treating the pilot's architecture as the production architecture. As one analysis put it, the problem is not governance itself; it is governance that is layered on rather than built in. In any regulated environment — finance, health, legal, anything with a duty of care — this is the difference between a system you can defend and one you quietly stop using the first time somebody asks how a decision was reached.

If a project cannot answer all four questions, it is not ready to build. It is ready to be scoped properly — which is typically a week of work that saves six months of it.

Part FourThe uncomfortable part

One further finding deserves to be sat with, because it cuts against every instinct a sensible manager has.

The organisations that succeed lean into friction. They treat resistance as the price of learning and design for it deliberately, rather than selecting the deployment that nobody will object to.

The AI project that meets no resistance whatsoever is almost always the project that changes nothing.

That matches everything I have seen in practice. Nothing anyone actually depends on has been touched. Resistance is a reasonably reliable signal that you are near something that matters.

This does not mean picking fights. It means expecting the conversation about how a role changes, and having it deliberately and early, rather than hoping it will not come up. The successful deployments redesign the role so that AI removes work — rather than adding a supervisory layer on top of everyone's existing job while calling it augmentation.

That single distinction — removing work versus adding review — is worth more than any tooling decision you will make this year.

In ClosingThe point

The AI question stopped being "does it work?" some time ago. It works. That argument is over, and the pilots prove it twice a quarter in most organisations.

The question that determines whether any of it shows up in your accounts is whether you are willing to change the process the technology sits inside. That, and not model quality, is where the money is. Gartner expects 40% of enterprise applications to include task-specific agents by the end of this year; the tools will keep arriving regardless. What will separate the businesses that profit from them is entirely a question of operating discipline.

A pilot that ships into an unchanged workflow is not a modest win. It is a cost with a success story attached — and, worse, it consumes the organisational appetite that the next, better project will need.

So before your next AI initiative, don't ask what it can automate.

Ask what it lets you stop doing. Then actually stop doing it.

Related Anchor Lotus service

Clarity Audit

If you want to know whether your next AI project will actually land on the P&L or just make a nice demo, start with a Clarity Audit — let's find where the time and money really go before anyone builds anything.

Find out more

Let’s talk

Would this help in your business?

Tell us what you would like to change. A 60-minute conversation, then a clear next step.

Start an enquiry