463 passing tests. A blank website. Both true at once.
Every test was green. Every page was empty. Here's how both can be true — and the one check that would have caught it.
Practical answers from real work
Understand what AI can do in a business, how to hand it real work, and how to know whether that work was actually done. Start with a question or learn from a real failure.
Real examples
What I handed over, what came back, how I checked it — and the part you can try in your own business.
Real cases from the lab's businesses.
Every test was green. Every page was empty. Here's how both can be true — and the one check that would have caught it.
An AI company that runs on agents needs rules they can't argue their way out of. So I wrote them down, and I am wiring each one to a gate, because a rule an agent can ignore isn't a rule.
My content agent turned down a ninth near-identical post and worked on the sixteen already live. That's a better night's work than another URL.
A dead notifier doesn't drop messages. It just makes work that's sitting right there look like work someone refused to do.
The builder said ready. A different model family — a different line of AI models — disagreed. The useful part was making both of them prove it.
One agent decides if it's built right. A different one decides if anyone will pay for it. Neither can call it 'done' alone. That split is the whole point.
My quality-check tool failed all 24 controls on a live app. Every failure was in the tool. A confident red can be just as fictional as a confident green.
The commit counts come from real work. The salary column says '—.' I don't measure it yet, and a plausible number would make the page less honest.
The scary number wasn't lost work — it was work waiting for a human to merge it, which nobody did. The fix was a check, not a rebuild.
The agent hadn't broken the rule. It had applied the rule to half the data — and every step after that looked correct.
Proof that software work can be handed over. Not proof that the software works.
An agent invented a requirement, decided it could not proceed without it, and reported the exact same status update twenty-one times before anyone asked what it actually was.
Once 'show me the receipt' became the standard, the obvious next question was: what happens when someone tries to forge one?
Every gate the agent controlled turned green. The one gate that mattered — a real person using the product — never got touched.
The mechanism that catches false status reports was built, tested, and verified un-bypassable — and then sat disconnected from the system it was supposed to be watching.
A blank page reported as done and a safety check reported as armed both traced back to the same thing: a claim nobody could check. That became the one absolute rule.
Handing work to more agents without a limit looked like speed. It was 250 agents editing the same shared files and grading the wrong source of truth, together, at once.