Skip to main content
← Back to Currents

How to Run an AI Pilot That Doesn't Waste the Budget

July 28, 2026

AI StrategyAI ImplementationPilots
Illustration: the first solid section of a bridge is bolted in place while a hand-crank crane lowers the next plank; the rest of the span is only a dotted outline.

Most AI pilots do not fail. They conclude. A demo gets shown, everyone agrees it was interesting, and nothing changes. Six months later the budget line is quietly reallocated and the organization files AI under "we tried it." The pilot did not test what needed testing, so it could not produce a decision.

The fix is not more budget or a better model. It is scoping the pilot around the question that actually determines whether AI works in your organization.

Find the binding constraint first

Every organization adopting AI has one constraint that binds before all the others. For some it is data access: the information the AI needs lives in systems nobody can get it out of cleanly. For others it is workflow ownership: nobody can change the process the AI is supposed to improve. For others it is evaluation: nobody can say what a good output looks like, so nobody can say whether the AI produced one.

A pilot that does not touch your binding constraint tells you nothing, no matter how well it goes. The classic version is the chatbot pilot run on public documentation because getting access to the real knowledge base would have required a security review. The pilot succeeds. The production version requires the security review anyway. You have learned that the vendor can build a chatbot, which you already knew, and postponed the hard question by a quarter.

This is why we treat readiness assessment as the step before the pilot, not a substitute for it. The assessment's job is to name the binding constraint so the pilot can be aimed at it.

Scope it to produce a decision

A pilot is an experiment, and experiments need a hypothesis sharper than "AI could help here." The scoping questions we insist on:

What decision does this pilot inform? Acceptable answers look like: "whether extraction accuracy on our real documents is high enough to remove the manual review step." Unacceptable answers look like: "whether the team finds it useful." If no decision hangs on the outcome, the pilot is theater.

What is the baseline? If you do not measure the current process before the pilot, you cannot claim improvement after it, and someone in the budget meeting will be right to say so. Time per task, error rate, backlog size. Measure something, even roughly, before the first prompt is written.

What are the kill criteria? Decide in advance what result means stop. Pilots without kill criteria do not end; they fade, consuming attention and goodwill as they go. Being able to kill a pilot cleanly is a sign of organizational health, and it is what makes the next pilot fundable.

Real data, real users, or it is a rehearsal. Sanitized data and volunteer users produce sanitized results. The pilot should run inside the actual workflow of the people who would live with the production system, on the data that system would see, security review and all. If that is too expensive for a pilot, that fact is itself the finding: your binding constraint is upstream of the AI.

The shape that works

The pilots that produce decisions share a shape: six to eight weeks, one workflow, a handful of real users, explicit baseline, explicit kill criteria, and someone accountable on the client side who owns the workflow being changed. Small enough to kill without embarrassment, real enough that success means something.

Notice what is absent from that shape: model selection debates, platform commitments, enterprise licensing negotiations. Those decisions get dramatically easier after a pilot has established that the workflow improves and the constraint is manageable. Made before, they are expensive guesses.

What "success" should obligate

One more discipline. Before the pilot starts, agree on what a successful result obligates the organization to do. If the pilot hits its numbers and the answer is still "we'll see," the pilot was not a pilot. It was a way of appearing to act. The point of the whole exercise is that AI that ships beats AI that demos, and shipping is a commitment someone has to make in advance.

Our AI implementation framework describes how we take the pilots that earn it into production. The short version: the pilot is stage one of a build, not a standalone science project, and it is designed that way from the first week.

Have a problem worth solving?

Tell us what you are trying to build or modernize, and we will tell you honestly how we would approach it.