The $440,000 sentence nobody read
Generation is free now. Verification isn't.
In October 2025, a consulting firm delivered a report to the Australian Department of Employment and Workplace Relations. It cited studies that didn't exist. It attributed a quote to a sitting federal court judge — who never said it. An outside academic happened to read the report closely enough to catch it. Not a reviewer. Not a QC process. An academic, reading it after the fact. The firm partially refunded the engagement: A$440,000. The revised report now discloses that generative AI was used in drafting.
Every error in that report was sitting in a sentence that looked properly sourced. A number, a base size, a citation. That's what makes this hard to catch, and it's the whole reason I'm starting this series here.
Drafting got maybe ten times faster. Review capacity crept up a little.
That's not a research-industry problem. It's the shape of the entire AI build-out right now, and software hit it first, loudly enough that the numbers are public. In May 2026, OpenAI open-sourced Symphony — a spec that turns a project board into a control plane for fleets of autonomous coding agents. The number everyone repeated was throughput: one write-up put the increase in landed pull requests at roughly 500%. Linear is now pivoting toward agent management as a first-class surface, not just human ticket tracking.
The sharpest response to that number wasn't a bigger number. It was one line: "generation scales effortlessly, validation does not." Every team standing up a fleet of agents like this is months, not years, from hitting that exact wall.
Here's the version of that wall that doesn't require any AI vocabulary at all. Pick any process where a person used to draft something slowly and a colleague read it before it shipped: a report, a topline, a memo, a summary. AI can now produce the first draft in a fraction of the time. The number of people qualified to read that draft critically — to check it against the underlying data, catch the sentence that doesn't follow, notice the citation that's decorative rather than load-bearing — has grown only a little. AI helps a reviewer move faster too, just nowhere near as much as it helps a drafter. It's roughly the same senior person it was last year, reading a little faster, and there is still nothing like enough of them.
You cannot close that gap by reading harder, or by reading a little faster with AI's help. The only thing that scales at the same rate as generation is a check that runs at generation's own speed — which still can't be a person reading every line, because even AI-assisted, that capacity is nowhere close to keeping pace.
Why this lands hardest in research and analysis work, specifically
In most published, public work, there's a chance someone outside catches the error eventually — an academic, a journalist, a competitor with an incentive to look. That's cold comfort, and it's also not available to you. In market research, in an internal analysis, in a deck built for one client or one steering committee, there is no outside reader. The first person to actually test a claim against the data is the client — in the room, in a planning meeting, or worse, after money has already been committed against it. Deloitte got to be embarrassed in public and repair its reputation. Most of us don't get that version of the story. We get a quiet, permanent loss of trust from the one client who caught it, and no idea how many others didn't.
What I've been doing about it
I've spent the last several months building and running the thing this series is actually about: a system that checks AI-generated output against the material it's supposed to be grounded in, before that output reaches anyone who'd act on it — not a second opinion from another model, but an independent check with no model in the loop for the parts of the claim that can be checked mechanically. I built the first version of it on my own production software, running with real autonomy, because I wanted to feel the failure modes myself before I asked anyone else to trust a fix for them.
That's the next post: what I actually built, what specifically broke, and the one rule I'd defend hardest out of all of it. After that, this series turns toward what the same problem looks like from inside market research and analytics specifically — because the shape of the failure is identical, and almost nobody is naming it out loud yet.
If you've caught a version of this — a number that looked right and wasn't, a citation that pointed at nothing, a claim that survived review because nobody had time to trace it back — I'd like to hear about it. Let me know on LinkedIn or on Substack. That's exactly the gap this series is about.