← All patterns

Pattern

AI pilots that never reach production

You've run pilots, done training, maybe built a chatbot. Nothing has stuck, and the organisation is no more capable than it was six months ago.

The pilot-to-production gap: the accelerated pilot thread stops at the unredesigned joins of verification, handoff and accountability, short of the production thread beyond. THE PILOT — IMPRESSIVE, FAST PRODUCTION where pilots stop THE JOINS NOBODY REDESIGNED VERIFICATION · HANDOFF · ACCOUNTABILITY
The pilot-to-production gap. The demo accelerates the task; the joins on either side of it stay as they were, receive the constraint, and the pilot stops there. An instance of constraint migration at the smallest scale.

Reuse this diagram: SVG · PNG · · CC BY 4.0

Agilist definition — the pilot-to-production gap

The distance between a pilot proving the model can do the task and the organisation absorbing the task being done differently: pilots die at unchanged verification steps, workflow joins, and unowned accountability, not at the model.

In one line

Demos prove the model. Production proves the organisation.

What you see

There is a folder somewhere with three impressive demos in it. Each one had a moment: the all-hands where it summarised a contract in seconds, the workshop where the room went quiet. Each one then entered an afterlife of security review, workflow questions nobody owned, and a sponsor who moved on. Nothing was cancelled. Nothing shipped. The board has started asking what the AI spend produced, and the honest answer is a folder of demos.

What is actually happening

A pilot answers the question “can the model do the task?” Production answers a different question: “can the organisation absorb the task being done differently?” Most plateaued rollouts confused the first question for the second.

The places pilots die are rarely technical. They die at verification (who checks the output, and is that cheaper than doing the work?), at workflow joins (the AI does step three brilliantly but steps two and four still assume a human), and at accountability (when it is wrong, whose name is on it?). None of these appear in a demo, which is precisely why demos succeed and rollouts stop.

Meanwhile your best engineers are using AI heavily as individuals, which proves the capability exists and makes the organisational plateau look even more like a mystery. It is not a mystery. It is an operating model that was never redesigned to receive the capability.

Why AI changes the constraint

The pilot accelerated one step of a workflow and left the joins on either side of it unchanged. That is constraint migration at the smallest scale: the constraint moves from “can we do this?” to the unredesigned handoff next to it, and the pilot sits in the queue that forms there until everyone stops mentioning it. The more impressive the demo, the further the constraint had to move, and the more surely something unowned was waiting to receive it.

Evidence, and what I’ve observed

Published evidence

The pattern is now institutionally documented: McKinsey’s 2025 State of AI survey finds the organisations capturing value from AI are the ones that fundamentally redesigned workflows rather than piloting tools in place, and the UK Government’s response to the AI Champions’ adoption plans finds scaling depends less on the model than on whether workers can interpret and act on its outputs.

Observed in practice

A real engagement, anonymised to sector and scale. Every client gets that deal.

A global consumer brand can buy every AI tool on the market. This one had. That was never the thing that made a team AI-native.

A large product organisation at real scale. Real budget, the kind of place where new tools land constantly and everyone is fluent in the latest one.

They brought me in to author the operating model for AI-native ways of working, and to launch the team into it.

I didn’t start with the tools. I started with the sequence.

First I connected the skills to real problems the team already had, because a skill with no problem attached is a party trick. Then I formed the team around the goals, so everyone was aiming the same capability at the same outcomes. Only then did the tools go in.

That order is the whole point. Same ingredients everyone else has. The sequence is the whole game.

They walked away with two things. A team launched into AI-native ways of working, aimed at real problems rather than the tool of the week. And a playbook to keep working from once I’d gone. That is the start of the journey, not the end of it, which is exactly what an operating model is for.

If that sounds like you: if you’ve bought the tools and the change hasn’t followed, the tools were never the problem. The order was.

How to diagnose it

Do a post-mortem on the last pilot that didn’t ship, and be literal about it: find the specific meeting, review, or unanswered question where it stopped moving. If the answer is a workflow join, a verification step, or an accountability gap, you are in this pattern. If every pilot dies in the same place, you have found your constraint.

When this isn’t the pattern. If the pilot genuinely couldn’t do the task to a usable standard, that is a capability or scoping problem: pick a better-shaped workflow, not a bigger rollout. And if pilots ship but reviewing their output is drowning your senior people, the constraint has already migrated onwards: the verification bottleneck.

What changes

Stop piloting. Pick one workflow that matters, trace where the last pilot actually died, and fix that join: the verification step, the handoff, the accountability. One workflow genuinely absorbing AI end to end teaches the organisation more than ten demos, and it produces the board answer you are currently missing.

The tell that it’s working is that the question changes. Not “can the model do the task?” but “which workflow do we absorb next?” One process runs with AI in it end to end — verification designed, accountability named — and the second conversion costs half as much, because the organisation has learned the shape. The board stops asking what the spend produced. The answer is running in production.

Cite: Robinson, T. (2026). “The pilot-to-production gap”. Agilist. www.agilist.co.uk/patterns/ai-pilots-that-never-reach-production

Is this where you are?

The AI Maturity Diagnosis measures your level and names your constraint in one to two days, for a fixed £4,500, in a document written for your board. There is a 90-minute Maturity Review at £750 if you want a smaller first step.

If you’re wondering which level you’re at — doing the work faster, changing how the work works, or doing work that wasn’t viable before — that question is usually where the useful conversation starts.