The kickoff is going well until the slide with the number on it. The sponsor has told the board that AI will make engineering ten times more productive. Sometimes the number is one hundred. The people in the room know it did not come from anywhere in particular. They also know that a vendor who contradicts it is no longer a vendor, and an engineering lead who contradicts it is having a career conversation. So the scope gets written to a number nobody believes, and the project is mispriced before a single line of code is generated.
We scope AI acceleration engagements for mid-market executive sponsors, which means we live inside this problem rather than observing it from a distance. The way out is not courage in the kickoff meeting. It is metric selection.
You cannot argue with the number, so change what it measures
An inflated multiplier survives because it is unfalsifiable as stated. Ten times more productive, measured over what? Lines generated? Features shipped? Revenue per engineer? Nobody committed to a denominator, which means nobody can be wrong, which means the number will still be alive and unmet two years from now, presiding over a disappointed program.
The practical move is to accept the ambition and volunteer the instrumentation. You do not say the number is impossible. You say: here is how we will know it is working. Then you propose measures that track outcomes the sponsor actually wants, and that engineering can move without cooking the books.
Measure accepted work, not generated work
Generation metrics are the trap. Lines produced, pull requests opened, tokens consumed: all of them jump the moment AI arrives, all of them flatter the initiative, and none of them are delivery. Code that has been generated but not reviewed, integrated, and shipped is inventory, not output, and inventory with a defect rate at that.
The metrics that survive contact with reality sit on the acceptance side of the pipeline.
Time from request to merged, verified change. The full cycle, not the generation step. This is the only number in the set that the business actually feels, and the one an inflated multiplier is implicitly promising to move.
Review cycles per change. How many round trips before work is accepted. If AI-assisted changes take more cycles than they used to, the assistance is manufacturing rework upstream of the merge, and the speedup is an illusion held together by unmerged branches.
Adoption of what shipped. Whether the delivered thing got used. A growing pile of shipped-but-unused features is the multiplier failing quietly, one sprint at a time, in a way no velocity chart will confess to.
Escaped defect trend. Whether speed is being purchased with quality debt that lands next quarter, when the program's credibility will need it least.
None of these require contradicting anyone. Together they convert the multiplier from a slogan into a funnel, and a funnel can be improved stage by stage, reported quarter by quarter, and defended in front of a board.
One more discipline makes the set work: capture the baseline before the tooling lands. A funnel with no before-picture invites the slogan back in through the side door, because any post-adoption number can be narrated as a win. Even a month or two of pre-adoption data on the same four measures turns the later conversation from persuasion into subtraction, and subtraction is the register executives trust. If the tooling is already deployed, reconstruct what you can from the version control and ticket history rather than skipping the step. An imperfect baseline still beats an argument.
Scope the engagement against the binding stage
Once the metrics measure accepted work, scoping follows from them. You size the engagement against the stage of the funnel that is actually binding. On most teams adopting AI heavily, that stage is no longer generation. The queue forms at review, integration, and verification, where senior attention is scarce and the new volume lands. Scoping an engagement as "add more generation capacity" for a team whose review queue is already full is how a 100x mandate produces a 1.1x result and a warehouse of stalled branches.
This is also the honest answer to "why not just buy more seats." Seats accelerate the stage that is rarely the constraint. The work that moves the funnel is usually verification design, integration discipline, and pipeline sequencing, which is the same reason most AI pilots fail before they start: they test the impressive constraint instead of the binding one.
Keeping the sponsor whole
The sponsor who announced the number is not the adversary in this story. They are the budget, and usually the only person in the building who wants the initiative to succeed as much as the vendor does. The metric conversation works when it is framed as protecting their announcement: a funnel showing real, compounding acceleration each quarter is defensible in a way a stalled moonshot never is. You are not lowering their number. You are giving it a denominator it can survive.
What kills projects is the opposite move, and we have watched it happen. The team privately re-scopes to what it believes is possible while publicly reporting against the slogan. The gap widens quietly for two quarters and then becomes a credibility event for everyone attached to the program, sponsor included. The mispricing was never corrected. It was deferred, with interest.
Say yes to the ambition, propose the instrumentation, report the funnel. If you are scoping AI work against an expectation that arrived pre-inflated, this is the framework we run. AI that ships, not AI that demos.

