The AI consulting market has a supply problem, and it is not a shortage. Everyone with a slide deck is an AI consultancy now. The gap between firms that have shipped production AI systems and firms that have run workshops about them is enormous, invisible on a website, and expensive to discover after signature.
We are a vendor in this market, so read what follows with that in mind. But these are the questions we would ask before hiring anyone, including us, and each one exists because we have watched a buyer skip it and pay for the omission.
1. "Walk me through an AI system you built that is running in production today."
Not a pilot, not a demo, not "an engagement with a Fortune 500 company." A system, in production, doing work, today. Then the follow-ups: how long has it run? What broke? Who fixed it? Real production systems generate war stories about drift, cost surprises, and edge cases. A consultant with no war stories has no production systems, whatever the logo slide says. Case studies should name outcomes specific enough to be falsifiable.
2. "What did you build that didn't work?"
Anyone shipping real AI systems has failures: pilots killed at the gate, approaches abandoned mid-build, models that could not hit the accuracy bar. A firm that claims none has either done too little work to accumulate failures or lacks the candor to describe them. Both disqualify. The quality answer includes what the failure taught them and what they now do differently.
3. "How will we measure whether the output is good, and who defines it?"
This question separates engineering from theater faster than any other. Production AI lives or dies on evaluation: a written definition of acceptable output, measured continuously, owned by someone. If the consultant's answer is a vibe ("our clients are very happy with the quality"), the production system will be graded on vibes too, and vibes do not survive contact with your compliance team. The strong answer describes evaluation as a deliverable: built early, run automatically, and handed to you.
4. "What happens to our data?"
Which models see it, under what terms, in whose cloud. Whether it trains anything. What leaves your boundary and what can be kept inside it, including self-hosted options where the data demands it. A firm that cannot answer precisely has not thought about it, and for organizations in healthcare, finance, insurance, or government, that is a disqualifying gap, not a detail to resolve during onboarding.
5. "Who maintains this after you leave, and what do they need to know?"
AI systems are not fire-and-forget. Models get deprecated, prompts drift, costs move, the vendor API changes under you. The right answer includes a handoff plan, documentation as a deliverable, and an honest account of what ongoing ownership costs. If the answer implies permanent dependence on the consultant, you are not buying a system. You are leasing one, on their terms. Ask for the exit plan while you still have leverage, which is before signing.
6. "Why does our existing stack matter to you?"
AI systems earn their keep by connecting to the systems where your business actually lives, and in most established organizations some of those systems are old. A consultant who waves at your legacy estate ("we'll just call an API") has never had to get structured data out of a twenty-year-old application. Firms with real modernization experience ask about your legacy systems early and specifically, because that is where implementation projects go to die. It is why we treat legacy modernization and AI implementation as one practice, not two.
7. "What would make you tell us not to do this project?"
The most revealing question on the list. A consultancy that cannot describe conditions under which it would advise against its own engagement is a sales organization. The credible answer names real disqualifiers: data that is not accessible, no internal owner for the workflow, a use case where errors are intolerable and evaluation is impossible. You want the firm that has walked away from bad projects, because the alternative is a firm that has billed for them.
What the questions have in common
Every question above pulls the conversation from promises to evidence: running systems, named failures, measurable quality, written handoffs. Firms that ship have that evidence in inventory and will show it readily, and none of the seven questions will surprise them.
Ours are on the AI implementation framework page and in our case studies. Bring the list.

