You Do Not Need to Read SQL to Catch a Wrong Number

BumbleB Content Team AI Analytics Decision-Making verification ai-analytics business-judgment decision-making business-analytics trust-in-ai

You Do Not Need to Read SQL to Catch a Wrong Number

Checking an AI’s analysis was supposed to require an analyst. Most wrong analyses are not wrong analytically — they are wrong about the business, which is the one thing the room already knows.

The answer arrives in four seconds, and the person who has to act on it spends longer than that deciding whether to believe it. That pause has a familiar shape by now. What is less familiar is what happens when they try to resolve it: they ask the system to show its work, and the system shows them a query.

This is offered as transparency, and in a narrow sense it is. It is also useless to almost everyone it is shown to. A founder deciding whether to keep a channel open is not going to audit a join condition. Handing them the mechanism and calling it verification asks them to check the one thing they are least equipped to check, and quietly skips the thing they are best equipped to check in the world.

The concession that has been sitting there

In July, in The Answer Got Cheap. Checking It Didn’t., the argument here was that verification, not generation, is now the binding constraint — following Christian Catalini, Xiang Hui and Jane Wu, whose model of the transition turns on “a biologically bottlenecked Cost to Verify” and who conclude that “the binding constraint on growth is no longer intelligence but human verification bandwidth.”

That piece ended on a concession it did not resolve. Cheap verification, it argued, requires judgment in the reader: legibility lowers the cost of checking for a person equipped to check, but it does not manufacture the equipment. Which meant the benefit landed most on whoever needed the least help. A team with a strong analyst gets faster. A team without one gets a confident answer and no way in.

That is the right worry and the wrong conclusion, and the gap between them is where most of the value in this market sits.

The inversion

The assumption underneath it is that checking analytical work requires analytical judgment. It is intuitive, and it collapses on contact with how analyses actually go wrong.

Analytical errors — a mis-specified window function, a fan-out from a bad join — are real, and they are not where the body count is. The failures that reach a decision and change it are almost always failures of premise. A segment defined one way by the system and another way by the company. A date range that spans a period when the business changed underneath it. A population that silently excludes the users who matter most, for a reason that made sense to the machine and would not have survived one sentence of contact with the room.

None of those are mistakes in arithmetic. Every one is a claim about the business, made by something that does not know the business, on behalf of people who do.

Which inverts who is qualified. The person who cannot read the query is frequently the only person in the building who can catch the error, because the error is not in the query. It is in a decision made before the query, in territory they own outright.

Judgment-different, not judgment-poor

The founder who has never had an analyst to ask is not short of judgment. They are short of one specific kind and unusually rich in another.

They cannot tell you whether a cohort was computed correctly. They can tell you in under a second that internal accounts are in that number, that the channel in question has reported badly since the migration, that “active” as the system has defined it is not the definition the company runs on, that the comparison period contains the outage. This is not a lesser form of checking. It is the form that catches the errors that actually get made, held by the person who has the most at stake in catching them.

What that asks of a system

If the reader is the person who knows the business, then shown work has to be written for them, and the requirement gets much more specific than “be transparent.”

It is not a translation problem. Rendering the query in plain English is a presentation choice applied after the fact, and it still describes the mechanism. The requirement is about which decisions get surfaced at all — and about surfacing them whether or not anyone knew to ask.

Every analysis is a stack of judgment calls made before any computation happens. What population counts. What window. What gets excluded and why. What the ambiguous term in the question was taken to mean. What was assumed because the data did not say. A system that made those calls knows it made them. Stating them, in the vocabulary of the business, as claims a person can disagree with, is a different artifact from a query plan and a different artifact from a citation.

This is not provenance, which most systems have already shipped. Citing the rows establishes that the arithmetic ran on real data; it says nothing about whether the arithmetic answered the question asked. The decisions sit upstream of the data any citation points at.

Consider a team told that retention fell in one of their segments. The check that matters is not whether the percentage is arithmetically right. It is whether that segment means what this company means by it, whether the users missing a recorded source were dropped or kept, whether the window covers a period when something else moved. State those three choices unprompted and someone who has never opened a query editor can overturn the conclusion in a minute. Leave them unstated and the only way to check the answer is to produce it again — the bottleneck, restored exactly where it was.

You cannot show work you did not do

There is a reason this is harder to copy than it sounds, and it is not about the quality of the writing.

A system that goes from a question to a query to a number did not make separable decisions. It made one leap. Asked afterwards what it assumed, it can produce an account — fluent, plausible, appropriately hedged — but that account is generated about the output rather than from the process. Nothing was decided; something is being described as though it had been.

That is worse than silence. A confident narration of reasoning that never happened is unfalsifiable by exactly the reader it is aimed at. The person who cannot audit the query cannot audit the story about the query either — and now holds a document that feels like grounds for believing the number.

Which means stated premises are not a feature. They are a tell. A system can only surface the choices it made before the answer if those choices existed as things in their own right — if the question was actually decomposed before it was computed. That is a property of investigating rather than retrieving, and it is not something a presentation layer can add afterwards.

This is also why “everyone will claim this by next year” is not the objection it appears to be. Everyone claiming it is what makes the tell worth having. The claim is cheap and the artifact is not, and the difference between them shows up in a form a non-specialist can actually use: whether the decisions arrive before the answer or get assembled after someone asks.

The pushback, and the limit

The strongest objection is that this is a floor and not a ceiling, and that some errors are only visible to someone with analytical training. That is true, and it does not need arguing away. A wrong statistical treatment will not be caught by business context, ever, and a team operating without that skill is exposed to a class of error they cannot see.

But context was never offered as a substitute for expertise. Where an analyst exists this makes their judgment go further rather than replacing it — less of the week reconstructing what a system decided, more of it on whether the decision was sound. And where no analyst exists, the alternative on offer is not rigorous analysis. It is a hunch. Against that baseline, a room that can contest the premises of an analysis is a substantial improvement.

A sharper version of the objection is that stating the choices will bury the reader. An analysis makes dozens of small decisions, and narrating all of them produces a second document nobody reads — with the added risk that a confidently stated assumption is itself wrong. That is why the requirement is not exhaustive disclosure: what has to surface is the contestable few, the choices where a different and equally reasonable answer would have changed the conclusion. Deciding which those are is work, and a system unwilling to do it has not met the requirement, it has offloaded it.

The honest framing is that the ceiling is still made of judgment. What changes is that a large amount of the judgment required turns out to be the kind the room already has.

The sorting

The tools in this category are converging on the same capability. Producing a defensible-looking answer from a question in plain language is becoming table stakes, and table stakes do not sort a market.

What will sort it is whether the work being shown was ever done. A system that narrates its reasoning after the fact has produced a document. A system whose premises existed before the answer did has produced something a person who knows the business can overturn — and the gap between those two is invisible in a demo and decisive in use.

The answer got cheap and checking did not. The way through is not to make everyone an analyst, and it is not better explanations. It is that a decision you can contest had to be made as a decision first, and nothing bolted on afterwards can fake that.

Key Takeaways

  • Analytical errors are real, but the failures that change decisions are almost always failures of premise — claims about the business made by something that does not know the business.
  • The person who cannot read the query is often the only one who can catch the error, because the error is upstream of the query.
  • A founder who has never had an analyst to ask is not judgment-poor. They are judgment-different, and their kind catches the errors that actually get made.
  • You cannot show work you did not do. A system that leapt from question to number can narrate afterwards, but a narration of reasoning that never happened is unfalsifiable by the person it is aimed at.
  • Stated premises are not a feature, they are a tell — the observable difference between a system that investigated and one that retrieved, legible to a buyer who can inspect neither.

The economic framing is drawn from Christian Catalini, Xiang Hui and Jane Wu, “Some Simple Economics of AGI” (arXiv preprint, February 2026), and continues the argument in “The Answer Got Cheap. Checking It Didn’t.”

🐝 BumbleB Crunch™ — your product insight, one conversation away Demo Try free