The Data Problem AI Can't Solve
Tue, 22nd Sep 2026 (Today)
One of the most persistent assumptions driving AI adoption in the enterprise is that better technology will surface better outcomes. In practice, the limiting factor is seldom the sophistication of the model, but rather the quality of the data architecture, the systems integration, and the institutional knowledge that feeds it.
Organisations are deploying AI into workflows that were built around manual processes, disconnected systems, and institutional knowledge that was never formally documented. The expectation is that AI will clean up what came before it. What it actually does is process everything faster, including those inconsistencies and integration failures pre-existing within the manual workflow.
That distinction matters more now than it did a year ago, because the stakes around data quality are rising on both sides of the equation.
The Other Side of the AI Adoption Story
Most enterprise AI conversations focus on internal efficiency: what can be automated, what can be accelerated, what labor can be redeployed. Fewer are accounting for what is happening externally at the same time.
In my area of expertise - tax compliance - regulatory agencies are making the same investments that enterprises are. They are moving toward continuous monitoring, automated data matching, and AI-assisted anomaly detection. The window between when a problem occurs inside an organization and when an external authority identifies it is closing.
That changes the cost calculation around data quality in a fundamental way. When scrutiny arrived on a delayed cycle, gaps in data infrastructure were a manageable internal problem, and there were longer cycles to find inconsistencies, and correct them before it was a problem. When visibility is driving towards real time, that buffer disappears. What was once an internal inefficiency becomes an external exposure, and the systems and processes responsible for producing the data are suddenly carrying a different category of risk.
A Case Study in High-Stakes Data
Consider excise tax compliance. Managing obligations across motor fuel and tobacco alone means navigating more than 10,000 rates across federal, state, and local jurisdictions, each with different product categories, license requirements, and filing obligations. The data feeding that process comes from multiple systems, multiple entities, and years of accumulated manual workarounds that were never designed to serve as a foundation for AI-driven decision-making.
When AI is deployed into that environment without first addressing the data infrastructure underneath it, the tool isn't noting knowledge gaps because it doesn't realize the gaps exist. And what gets processed at scale across hundreds of filings and dozens of jurisdictions becomes significantly harder to unwind than what a well-governed manual review would have caught early.
One organization we've spoken with operating in this space missed a $2.5 million refund for four years. The filing obligation existed only in a retiring employee's head. No system captured it. No automated process flagged it. By the time it was discovered, they were nearly out of statute and will likely recover less than a third of what was owed. That is not a technology failure. It is a data and knowledge infrastructure failure that no AI tool could have prevented, because the information it needed to act on was never structured or accessible to begin with.
What AI Cannot Substitute For
The knowledge gap is the part of the data quality problem that is hardest to close and most consistently underestimated. Organisational knowledge that has accumulated through years of experience, judgment calls, and direct exposure to how external authorities interpret ambiguous requirements exists in people themselves rather than in a structured form AI can access and index.
When those people leave, the knowledge leaves with them unless the organization has deliberately built it into systems and processes first. AI deployed on top of that gap does not fill it. It operates around it, producing outputs that look complete until the moment they aren't.
This is where the assumption that better AI produces better outcomes breaks down most clearly. And in environments where outputs must hold up under external scrutiny, the cost of that gap surfaces as penalties, audit findings, and positions that have to be unwound at significantly greater cost than prevention would have required.
The Questions That Actually Matter
For organizations evaluating where AI creates genuine value versus where it creates the appearance of progress, the most useful questions are about foundation over capability:
- What is the actual quality of the data feeding the AI tools already in use?
- Is the institutional knowledge the organization depends on embedded in systems and processes, or is it concentrated in individuals who may not always be there?
- Can the outputs AI tools produce be explained and defended when external scrutiny arrives?
- And is governance being built before automation is expanded, or is it being planned for after it becomes necessary?
The organizations that will realize the most durable value from AI will be the ones asking these questions and gaining insight into what their data actually contains, where accountability for AI outputs actually lives, and whether governance in place before the environment makes it a requirement. That work is less visible than a deployment announcement. It is also what determines whether AI compounds an organization's advantage or its exposure.