Our AI Missed Its Own Mistake
The line nobody asked for
"Falls right first time, so it still runs in ten years."
That line turned up in a web page one of our AIs wrote for a groundworks firm, Coastline Groundworks. Coastline doesn't exist. We invented it for the test. No history, no ten-year-old jobs, nothing in its brand pack about lasting a decade. The AI made the promise anyway.
Then it marked its own work. On the test built specifically to catch invented facts, it gave that page nine out of ten. A judge from the other AI family, run through Codex, gave the same page one out of ten and quoted the line back as the reason.
Nine and one. Same page, same test, same afternoon.
Why we tested it instead of reposting it
Nate Herkelman put out a video called "I Tested Claude Code vs. Codex on Design. It Wasn't Even Close.", in which he reported Codex winning five of six open-ended design rounds. Nate's a builder I've picked up plenty from. We ran the claim ourselves rather than repeating it. His figures are his, ours had to be ours.
How we set it up, in plain English
Everything was written down before a single page existed: four possible verdicts and the threshold each needed to fire, on disk, on 26 August. That's what stops you reading back the answer you wanted.
Coastline Groundworks Ltd is a fictional worked example built by ConstructionX AI. No real firm, person or job is depicted.
Four drafting seats got the same brief for it: three models, one of them run twice under different persona headers, which we labelled Codex and Sol. We ran it twice, once with a tight brief that said what to build, once with a loose one that left room.
Three judges scored every page from anonymised rendered screenshots, against three tests: correct, defensible, and any good. Two of the judges were outside the drafting family and properly blind. Our own Claude judge had drafted too, and it could recognise its own page; we had written that risk down before the run. Before anyone opened the file of which model wrote what, I made my own pick on each run, by eye. The file times back that up, though they don't prove it.
The first run, and the ten-year promise
The ten-year line came out of the tight brief. The judge from the other AI family scored that page one out of ten on correctness, an automatic fabrication fail. The second outside judge flagged the same sentence, plus "itemised" and "We turn up on it", as things the firm had never claimed. Our Claude judge scored the same page nine and never saw any of it.
Across three judges, the run came out at 9.00, 7.56, 7.22 and 5.67, a spread of 3.33. All three ranked the Sol draft first. Both outside judges put the Claude page last and said scrap it.
One more number. The Claude judge scored its own draft 8.33, where the other two gave that same page 5.00 and 3.67. More than four points apart, and we had written that risk down before the run started.
The bit I liked least: my blind pick was the Claude page, the one both outside judges ranked last and wanted killed. The human and the machines disagreed completely.
The second run, and a rule broken by the model that wrote it
The rule sat in the packet the drafts were written from: "One message in the hero and one primary action. Two competing calls to action is a fail on this test."
The Claude seat wrote that rule earlier that afternoon. Then it built a hero with two buttons, "Talk to Sam" and "See what we take on". Both outside judges spotted it separately and put the page last for it. The Claude judge gave its own page seven out of ten on quality and never mentioned the buttons.
Nothing was invented on the second run. Correctness scores ran 8 to 10 with no automatic fail, so the fabrication didn't repeat. The scores went flat too: Fable 8.67, Sol 8.67, Codex 8.44, Claude 7.89, a spread of 0.78. Three judges, three different first places, no majority for anything. My blind pick was the Fable page, so I picked a Claude-family draft both times.
One result held throughout. Both outside judges put the Claude draft last on both runs: eight drafts, two runs, two judges, same position every time.
They caught different faults, though. Our Claude judge found mobile clipping on the Codex page both outside judges walked past. The Codex judge flagged an "across North Norfolk" claim wider than the towns named in the kit. Argument for a mixed panel, not more of one family.
The reviewer that wouldn't sign
The reviewer that signs off our work is never from the same model family as the thing that did the work. A model asked to check itself tends to pass itself, and the nine is what that looks like.
At the close of the session, that signer read our own write-up and returned NOT SIGNED. Twice.
Round one: our write-up said all ten steps had run, while the merged-draft files, the sign-off file, the save manifest and the save commit did not exist. It offered the ordering of the blind-pick files as proof, when that ordering was only consistent with the claim. And the tree carried out-of-scope changes.
Round two: the merged-draft files had six headings and no merged draft. The "all ten steps ran" line was still in the plan after a promise to withdraw it. The tree still wasn't clean.
Keep the distinction, it's the honest one. It refused to sign off the write-up, not the findings, and the verdict arithmetic was confirmed correct in round one. The final verdict on file is still NOT SIGNED, and it was right both times.
What the verdict was, and what it wasn't
Nothing won, and nothing changed.
The result was KEEP-CLAUDE, by the default clause that fires when nothing else does. No verdict condition met its own terms: adopting Codex failed two of its three tests, and KEEP-CLAUDE's own conditions failed too. Codex didn't clear the bar for a change, and Claude didn't win either. Frontend generation stays where it was because nothing earned a move.
One thing in the data I'm not calling a finding. The Codex and Sol seats were the same model on the same settings, split only by a paragraph of persona text, and on the first run they scored 9.00 and 7.22. A gap of 1.78, the largest single effect in the pilot. The second run didn't reproduce it, 8.67 against 8.44, so it's a question for a future test, nothing more.
What this doesn't prove
Quite a lot, and it's worth naming.
The automated bug checker returned zero on all eight drafts and missed a real defect on each run: 401 pixels of page inside a 390 pixel phone screen on the first, clipped hero copy on the second. Both obvious by eye in the render. Honest, and nearly useless.
The Claude seat judged the first run before drafting the second. Contamination across runs, not pre-registered, and it leans towards leaving things as they are.
Two of the three judges turned out to be the same model, so "both outside judges agreed" is one model agreeing with itself twice.
My blind picks are two data points from one bloke, supported rather than proven. And this is two briefs, one fictional firm, one type of page. Stretch it further and you're making things up, which is where we came in.
What an owner should take from it
The nine out of ten isn't a story about a model being wrong. Models are wrong sometimes. What caught the wrongness was the arrangement around the model, not the model.
Three things worth copying.
Your reviewer comes from somewhere else. Same-family review found nothing here and scored it nine.
A rule you care about needs a gate, not a reminder. The two-button rule was written down by the same AI earlier that afternoon, and it broke it anyway. Something other than the thing being checked has to do the checking.
Keep a human pick that's allowed to disagree. Mine did, twice. If your process can't survive a person saying "I don't care what the score says", it's a rubber stamp.
To be straight about our own setup: nothing moved, because nothing won. The one action booked is a dated line in our rule about which model does which job, on Wednesday 2 September.
Look at the pages yourself
All eight pages are up in a gallery on this site, live and scrollable, so you can judge them for yourself. The method is written up in full on the worked example page.
Coastline Groundworks Ltd is a fictional worked example built by ConstructionX AI. No real firm, person or job is depicted.
The test firm is a groundworks company because that's the trade I know. This isn't construction-only. The same systems work in any business where an owner's expertise needs to reach further than the owner can.
PS: I run paid 45-min Opportunity Maps for UK trades and small businesses. If this got you thinking, drop me a line.