AI Manifesto: Stage 5, Pilot
The Prototype That Tells the Truth
AI pilots have become corporate theatre. A few enthusiastic users, a shiny demo, a slide full of “time saved”, and a claim that the company is now “AI enabled”.
Stage 5 exists to prevent that story from turning into regret. Pilot asks: “Have we proven this works, with a human still owning the outcome?”
The framework defines a pilot in practical terms. Start with one team or workflow and agree on clear success criteria before work begins. Assign a named person accountable for every AI-assisted output. Use narrow tools that solve one problem well. Document what the AI gets wrong as carefully as what it gets right. The guardrail turns “human in the loop” into an operational rule: if the pilot can’t run with meaningful human review, it isn’t ready.
“Works” and “works here” are not the same thing
This is the stage where organisations learn the difference between “works” and “works here”. A model can look impressive in a clean test set and still fail in the real workflow because real work is messy. Inputs are incomplete. Policies have exceptions. Customers use odd language. People are distracted. The system is integrated into tools that have their own quirks. All of that matters.
Document what goes wrong as carefully as what goes right
Software reliability engineering offers a useful comparison. Netflix popularised chaos engineering with tools like Chaos Monkey that deliberately terminate instances in production to reveal weaknesses before customers do. The philosophy isn’t expecting failure. It’s recognising that failure already exists.
That’s exactly what Stage 5 is asking for when it tells you to document errors as carefully as successes. The “failure log” becomes one of the most valuable outputs of the pilot. Not just “the AI hallucinated” but where, why, and with what impact. Not just “people liked it” but whether they verified outputs or deferred to them. Not just “productivity rose” but what happened to quality.
Request the full AI Manifesto
Shared accountability is often no accountability
This stage also reinforces the manifesto’s accountability spine. The “named person accountable” requirement is how you avoid a common organisational failure: everyone assumes someone else is checking. A pilot with shared accountability often has no accountability. A pilot with explicit ownership has learning you can trust.
Narrow tools are a governance decision
Choosing narrow tools isn’t simply a technical preference, it’s a governance decision. Narrow tools are easier to evaluate, easier to constrain, and easier to retire if they drift. Broad platforms invite broad use, which makes it harder to distinguish “pilot behaviour” from uncontrolled adoption.
You need evidence that the technology strengthens human judgment rather than quietly replacing it. A credible pilot also has clear success criteria, which Stage 5 insists on. That sounds obvious, but it’s often missing. Time saved is not enough on its own. You need quality metrics. You need evidence that the human review step is real. You need measures that show the tool is augmenting judgment, not bypassing it.
A strong pilot ends with a set of decisions: what continues, what changes, what stops. It should produce training improvements for Upskill, tighter boundaries for Integrate, and honest assessments for Measure. It should also feed governance, because what you learned about data handling, bias risk, and failure modes becomes part of your operational oversight.
What a pilot is really testing
Most importantly, Stage 5 produces organisational truth. It doesn’t simply prove that technology can do it. It reveals how people behave when the technology is introduced. And in AI adoption, behaviour is almost always the greater risk.
Request the full AI Manifesto