Skip to content
Lorendix

Data and AI

Why AI fails on ungoverned data

A confident answer drawn from records that do not reconcile is worse than no answer, because it gets acted on. What has to be true before an AI initiative is safe to start.

  • Lorendix engineering team
  • Delivery and advisory
  • 22 May 2026
  • 6 min read

The failure mode of an AI project is rarely that it produces nothing. It is that it produces something plausible, quickly, from data that nobody had reconciled, and the organisation acts on it before anyone checks.

Most organisations approaching AI ask a version of the same question: which use case should we start with? It is the wrong first question, and it is wrong in a specific way. It assumes the constraint is imagination. In almost every case we assess, the constraint is that the underlying records cannot support a confident answer, and adding a model to that does not fix it. It removes the friction that used to make the problem visible.

When a person produces a number from a messy data estate, the process is slow and full of hedging. They know which spreadsheet they distrust. They say so in the meeting. When a model produces the same number, it arrives instantly, in fluent prose, with no hedging at all. The underlying uncertainty has not changed. The signals that used to communicate it have simply been removed.

A model does not make bad data worse. It makes bad data persuasive, which is worse.

3

answers the same question typically returns across systems we assess

0

of those systems that can say which answer is authoritative

2 weeks

a data health check takes to establish both

Figures from assessments we have run, not from published research.

Three things that have to be true first

Before an AI initiative is safe to start, three questions need real answers. Not aspirational answers, and not answers that describe the target state in a strategy deck.

Ownership
For each significant dataset, one named person is accountable for whether it is right. Not a department, and not a committee. Where there is no owner, there is no one whose job it is to notice that a field stopped being populated in March.
Quality
You can state, in numbers, how complete and how consistent each dataset is, and the statement is produced by a check that runs rather than by an exercise somebody did once. Duplicate rates, null rates in fields that matter, and referential breaks between systems.
Access
You know who can read each dataset, that answer is enforced by a system rather than by convention, and it survives the data being copied into whatever the model runs on. Most governance failures in AI projects are not model failures, they are a copy of a sensitive table sitting somewhere new.

What a two week health check actually looks at

A data health check is not a strategy exercise, and it should not produce a maturity score. It produces a list of specific defects with a cost attached to each. In two weeks, on a mid-sized estate, that is achievable, and the output is usually uncomfortable in a useful way.

  • Every system that holds a copy of a customer, a product or a transaction, and which one the business treats as authoritative
  • Duplicate and near duplicate rates in the core entities, measured rather than estimated
  • Fields that are populated by convention rather than by constraint, and therefore populated inconsistently
  • Reconciliation breaks between systems that are supposed to agree, quantified
  • Where personal or commercially sensitive data actually sits, including in exports and reporting copies
  • Which questions the estate can answer today with confidence, and which it cannot

That last item is the one clients find most useful. It reframes the conversation from what we would like AI to do, to what our data currently entitles us to ask. Those are very different lists, and the gap between them is the actual project.

The order that works

The sequence we recommend is unglamorous, and the order is the whole point. Each step is what makes the next one hold.

  1. Fix ownership

    First, because without a named owner nothing else stays fixed. Every improvement decays back to where it started.

  2. Measure quality

    You cannot improve what you have not quantified, and a number makes the argument for funding that an opinion does not.

  3. Constrain access

    The moment data becomes useful it also becomes mobile. Access has to be bounded before it is worth moving, not after.

  4. Choose one narrow use case

    Narrow enough that the answer can be checked against something independent. A use case you cannot verify teaches you nothing.

It is worth saying what this is not. It is not an argument for delaying AI until a data estate is perfect, because no estate is perfect and the ones that wait tend to wait indefinitely. It is an argument for knowing which parts are sound enough to build on. A business with three reliable datasets and accurate documentation of the rest is in a far better position than one with forty datasets and no view of which to trust.

The governance question that gets skipped

There is one more thing that ungoverned estates make impossible, and it tends to surface late. When a model produces an answer that turns out to be wrong, and something was decided on the basis of it, somebody will ask what the model was looking at. If the data lineage is not recorded, that question has no answer, and the organisation is left explaining a decision it cannot reconstruct.

Recording lineage is not difficult when it is designed in. It is close to impossible to reconstruct afterwards. That asymmetry is the whole argument for doing the boring work first, and it is the same asymmetry that applies to audit trails, access models and every other structural decision. You are not buying tidiness. You are buying the ability to answer a question you cannot yet predict.

Before you approve an AI initiative

  • Name the person accountable for each dataset it will read
  • Ask for the duplicate and null rates in the core entities, as measured numbers
  • Establish where the data will physically sit once the model can reach it
  • Choose a first use case whose output can be checked against something independent
  • Agree how lineage will be recorded before the first result is presented to anyone

Share this

Bring us the question you have not been able to answer internally

If this raised something specific about your own estate, describe the situation rather than the solution and we will tell you what we would look at first.