Skip to main content
← Back to Currents

5 Signs Your AI Project Has a Data Problem, Not a Model One

September 28, 2026

AI ReadinessEnterprise AIAI Strategy
The mascot stands on a step ladder syncing one wall clock to another while a puzzled AI robot looks on — two sources of truth disagreeing, fixed at the source.

Your AI pilot underwhelmed, the postmortem blamed the model, and now you're lining up the next pilot on a better one. Before you do, check whether the real constraint is your data. A data problem means the model is doing its job on inputs that are wrong, conflicting, or off-limits. It shows up as five signs: a demo built on a hand-picked extract, no named data owner, over-broad access, sources that disagree, and incident fixes that are all data work. If more than one fits, a new model won't save the next pilot.

You're in good company. In Nasuni's survey of 1,000 enterprises, published May 18, 2026, 94% said they struggle to manage unstructured data, and 46% said AI initiatives had revealed data quality and governance issues. We see the same pattern in stalled projects. It just rarely makes it into the postmortem.

1. The pilot ran on a hand-picked extract. Someone cleaned a dataset for the demo, and the demo worked. Production data is the real test. If nobody can tell you how far live data drifts from that extract, the gap between demo and rollout isn't a model problem. It's cleanup nobody scoped.

2. Nobody can name who owns the data. When the model gives a wrong answer, someone has to rule on what the right answer was. If ownership is a shrug, disputes about correctness have no referee. Trust in the system dies from unresolved arguments long before the model's error rate matters.

3. The system reads more than any employee should. The fastest path to a working pilot is often a service account with access to everything. Security review will stall that rollout, and they'll be right to. Scope access to what the use case needs. It's slower up front, and it's what turns a pilot into a product.

4. Two sources disagree and both feed the model. Duplicate records, a warehouse that lags the source system, an export nobody retired. The model reflects the mess faithfully and takes the blame for it. If the answer changes depending on which copy got read, no model swap fixes that.

5. Every incident fix is data work. Pull up your remediation list after each bad output. If the entries say cleanup, reclassification, permission changes, and source reconciliation, your project already told you where the constraint is. Believe it.

None of these show up in a model benchmark, which is why the postmortem keeps reaching the wrong verdict. Legacy systems make it harder. On a ColdFusion or other long-lived app, fifteen years of business rules often live in queries and application code, not in the schema, so the data only makes sense with the code beside it.

The fix is ownership, classification, and access scoping, done before the next pilot. We laid out the order for that work in our guide to fixing the data foundation before you add an AI agent; this list is how you tell the foundation is the problem in the first place. If the signs look familiar, an AI readiness assessment names your binding constraint before the next budget cycle pays to rediscover it.

Have a problem worth solving?

Tell us what you are trying to build or modernize, and we will tell you honestly how we would approach it.