Having Data Is Not the Same as Having Evidence
Before promising an AI outcome, validate that the data can support it. What querying a platform I couldn't trust taught me about data readiness.
I have written before about a platform where two reports, pulling from the same underlying data, returned different numbers. That piece is about why I stopped trusting the dashboards. This one is about what I learned once I started producing the numbers myself.
The Reports That Disagreed With Each Other covers the architecture. The short version is that the platform ran on Dynamo and OpenSearch, the two drifted apart, and I went to Snowflake to get answers that held up.
The question comes before the query
When someone on the operations team asked me for a number, I did not open Snowflake first. I made sure I understood what they were trying to find out and what decision the answer would inform. A request for a count is usually a question about something else, and the query depends on which question it really is.
Only then did I write anything. I knew our data well, but I did not always know how to piece together complex queries to get what I was after, so I used AI to help write them. I gave it table names and fields. It got noticeably better once I described the data model, because it started to understand what was stored where.
I never gave it user PII. That was a fixed line, and it never slowed the work down, because a query needs to know the structure of the data and has no use for the records themselves.
I used ChatGPT for this. I don’t think the choice of model was what mattered. The setup and the questions I brought to it mattered far more.
The mistakes were mine
I checked every result before I shared it. I pulled the output and read it against what I knew about the product, the structure of our data, and what values belonged in each column. Knowing the product well was what let me see when a number did not make sense.
When a number was wrong, the cause was almost always me. I had pointed the AI at the wrong table, or I had left out something the prompt needed. The AI wrote a faithful query for the question I gave it. My understanding of the tables was the weak point, and getting better at this meant learning the data more deeply.
I expected the risk to be AI writing bad SQL. What I found was that the quality of the answer depended on how well I understood the data underneath it.
Where the data itself ran out
The same limit showed up at a larger scale. We wanted to report back to customers with real evidence that the service was working for their students: usage, attendance, how consistently educators gave feedback, and how students rated their experience.
We had the data. A lot of it was sitting in Snowflake. But Dynamo is a document store, and some of what it pushed into Snowflake arrived as raw JSON inside a column instead of readable values in their own fields. It could be queried with effort. It could not be built on at scale without first being reworked into a relational shape.
That is why moving to Postgres came first. The goal was one source of truth that people could trust, in a structure that reporting could be built on. That migration finished shortly after I left. I worked toward that reporting capability, but it did not ship before I left. I can’t speak to what it looks like today.
The rule I took from it
Data readiness is part of the product requirement. If we have not validated the data and tested the outcome, the AI capability is still a hypothesis.
The rule I would teach every new PM is this: before promising an AI outcome, validate that the data can support it.
In practice, that means answering four questions with engineering before committing to delivery.
- Do we have the relevant data, and does it contain the information this task needs?
- Can we trust its meaning and quality? Are records correctly linked, current enough, and representative of the users we will serve?
- Are we permitted to use it this way? Access, consent, and user permissions have to carry through to the AI feature.
- What evidence will prove the feature works? Agree on representative test cases, acceptable errors, and what happens when information is missing.
A hypothetical makes the difference concrete. Imagine a tutoring product that promises to recommend what a student should learn next. It might have thousands of session transcripts and still lack reliable student matching, skill mappings, or evidence of mastery. Summarizing what was discussed in a session and determining what a student has mastered are two different promises. The data may support the first long before it supports the second.
Product defines the outcome and the consequences of being wrong. Engineering establishes how the system obtains, validates, and processes the data. Together they test whether the result deserves the trust the product asks users to place in it.
There is a qualification worth stating. You do not need perfect data to experiment, and clean data does not guarantee correct AI output. The standard is sufficient evidence for the specific use and the consequences of being wrong.