All insights
AI Agents·September 29, 2026·6 min

AI-ready data: four tests before you add an AI step

AI-READY DATATHE AI STEP IS ONE BOX IN SIXSourcesYOUR SYSTEMSPipelineDEDUPE · JOINOne sourceFRESH · OWNEDAI stepREADS CLEAN DATAChecksCONFIDENCE · RULESPersonAPPROVES EXCEPTIONSEVERY DECISION LOGGED BACK

Most AI projects do not fail at the model. They fail at the data the model is asked to read. That is not a slogan; it is where the research keeps landing. In February 2025 Gartner predicted that through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data. The same release reported that 63% of organisations either do not have, or are not sure they have, the right data management practices for AI, from a survey of 248 data management leaders.

In April 2026 Gartner followed up from the other side: organisations with successful AI initiatives invest up to four times more in data and analytics foundations. And on 21 September 2026 it added that 60% of organisations that ignore data governance culture will fail to govern AI successfully by 2027. Three different questions, one answer. The work that decides whether AI helps is the work underneath it.

What "not AI-ready" looks like in a normal business

It rarely looks like a disaster. It looks like three systems that each hold a reasonable version of the truth. The CRM calls the customer Acme Supplies Ltd and was updated this morning. Billing calls them ACME Supplies, from an export someone ran six weeks ago. Support has Acme Suppl., with no customer ID at all. Each system is fine on its own. Nobody owns the question of which one is right.

Put an AI assistant on top of that and ask a simple question: what does Acme owe us? It will answer, fluently and with confidence, from whatever it can reach. It may count three customers where there is one, miss the unpaid invoice that lives only in billing, and never tell you it was guessing. The model read exactly what it was given. The fix is upstream.

Same customer, three spellings. One confident answer.CRMAcme Supplies LtdBILLING EMAIL MISSINGUPDATED TODAYBILLINGACME SuppliesACCOUNT OWNER MISSINGEXPORT SIX WEEKS OLDSUPPORTAcme Suppl.CUSTOMER ID MISSINGNO OWNER"Acme has three separateaccounts and owes nothing."CONFIDENT. WRONG.THE MODEL READ EXACTLY WHAT IT WAS GIVEN · THE FIX IS UPSTREAM
An illustrative example of the pattern we see most: one customer, three systems, three spellings, and an answer that sounds right.

The workflow under an AI step that works

When we build an AI step for a client, the AI step is one box in six. Sources come first: the CRM, billing, support, inboxes, forms. A pipeline pulls from them on a schedule, removes duplicates, checks required fields, and joins records on stable IDs rather than names. The result lands in one source of truth, a warehouse table or a well-kept sheet, with a known age and a named owner. Only then does the AI step read, and it reads the clean table, not the raw systems.

After the AI step come checks: a confidence score, the client's own rules, and a list of cases that must always reach a person. The person approves the exceptions. Every decision is written back to a log, so the next run is better informed and anyone can see why an answer was given. That loop is what separates an AI feature that survives its first month from one that is quietly switched off.

Four tests you can run on your data this week

You do not need a data team to find out whether your data is ready. You need four honest answers.

  • One record per thing. Can two rows describe the same customer, product or supplier? If records are matched by name, the answer is yes, and every AI answer built on them inherits the doubt. Stable keys fix this; clever prompts do not.
  • Known freshness. Do you know when each table last updated? An AI step that reads a six-week-old export will be wrong in a way that looks right. Every table should carry its own age, and a stale one should raise an alert instead of feeding an answer.
  • A named owner. When a field goes blank or a feed stops, who hears about it? If the answer is nobody, the data is decaying on a timer. Ownership is a person with an alert, not a line in a policy.
  • Traceable answers. Can you show the row behind any answer the AI gives? If not, you cannot check it, correct it, or defend it to a customer or an auditor. Traceability is also what makes people trust the tool enough to use it.
Four tests of AI-ready data.One record per thingKEYS, NOT NAMESCan two rows describe the same customer?Known freshnessEVERY TABLE HAS AN AGEDo you know when each table last updated?A named ownerSOMEONE FIXES ITWho gets the alert when a field goes blank?Traceable answersEVERY ANSWER POINTS TO A ROWCan you show the row behind any answer?
If any of the four answers is no, fix that before you add an AI step. It is cheaper now than after the tool has been trusted and been wrong.

What to do this month

Pick the one question you most want an AI assistant to answer, and trace where its answer would come from. List the systems involved and check the four tests against each. Usually one or two gaps stand out: a missing key, a manual export, a table nobody owns. Fix those with a scheduled pipeline and a single table the AI step can read. Then add the model, with a confidence threshold and a person on the exceptions from day one.

Fix the pipeline before the prompt. The model is the easy part; the data it reads decides whether anyone keeps using it.

If your team is being asked to "add AI" and the data underneath is spread across tools that disagree, tell us the question you want answered in 2 lines. We will tell you honestly which of the four tests it fails, and whether it wants a pipeline first or is ready for an AI step now.

AI-ready datadata pipelinesAI agentsdata quality

Got a workflow like this?

Tell us what's eating your team's time, we'll tell you honestly whether automation is worth it.

Book a Consultation

We typically respond within 24 hours