The Coastline Pilot

On 26 August 2026 we gave four AI drafting seats the same brief and asked each to build a landing page for a groundworks firm we invented. One wrote a promise the firm could not have made, then scored its own page nine out of ten on the test meant to catch exactly that. A judge from the other AI family gave the same page one out of ten. No model won, nothing changed, and the reviewer that signs our work refused the write-up twice. Here is the method, and here are the pages.

What we tested

The trigger was Nate Herkelman's video "I Tested Claude Code vs. Codex on Design. It Wasn't Even Close.", in which he reported Codex winning five of six open-ended design rounds. That's a real claim about tools we put client work through, so we tested it rather than repeating it.

The shape: four drafting seats, one firm, two briefs (tight and loose), three judges scoring from anonymised screenshots on three tests, one human blind pick per run. Four seats means three models, one of them run twice under two different persona headers. Two of the judges were outside the drafting family and properly blind; our own Claude judge had drafted too and could recognise its own page, a limit we wrote down before the run.

How we tested it

Anyone can copy this. Seven steps, one day.

  1. Write the verdicts down first. Every possible outcome and its threshold, on disk before a single draft exists.
  2. Use a firm that doesn't exist. Nothing true to lean on, and no client exposed.
  3. Run the brief twice, tight and loose. Tight says what to build. Loose leaves room.
  4. Anonymise the drafts. Renders only, no names, mapping held back until scoring ends.
  5. Judge from the rendered page, not the code, and use judges from more than one model family. Ours were one Claude judge and two runs of an OpenAI model through the Codex CLI under different personas, which turned out to be a limit (see below).
  6. Make a human pick before the mapping is opened, and record it so the order is checkable.
  7. Have the write-up signed off by a reviewer from a different model family than the one that did the work.

What happened

On the tight brief, the Claude draft's copy read "Falls right first time, so it still runs in ten years". Coastline has no history, and nothing in its kit promises a decade. The judge from the other AI family scored that page one out of ten on correctness, an automatic fabrication fail. Our Claude judge scored the same page nine.

On the loose brief, the packet carried a rule the Claude seat had written earlier that afternoon: one message in the hero, one primary action, two competing calls to action is a fail. The Claude draft's hero shipped two buttons. Both outside judges named it separately, and the Claude judge scored its own page seven out of ten on quality without mentioning them.

One pattern ran through the lot: both outside judges put the Claude draft last on both runs. Eight drafts, two judges, same position every time.

Then the sign-off. The reviewer, from the other family, returned NOT SIGNED on our write-up twice: for claiming steps that had left no artefact behind, and for offering file ordering as proof when it was only consistent with the claim. It refused the write-up, not the findings, and it was right both times. The verdict on file is still NOT SIGNED.

What it did not show

No model won. Nothing changed. The result was KEEP-CLAUDE by the default clause, the one that fires when no condition meets its own terms, so frontend generation stayed where it was because nothing earned a move.

The limits matter as much as the findings. The automated bug check returned zero on all eight drafts and missed a real defect on each run, both obvious by eye. The Claude seat judged the first run before drafting the second, contamination we hadn't pre-registered. Two of the three judges were the same model. And this is two briefs, one fictional firm, one type of page.

The pages

Every page is below, live and scrollable. The judges scored from fixed desktop and mobile screenshots of these same pages, with the names hidden.

Eight pages, two briefs, one firm that doesn't exist.

Open the gallery full-page

Coastline Groundworks Ltd is a fictional worked example built by ConstructionX AI. No real firm, person or job is depicted.

The loose brief is where the two best ideas came from: a strip-foundation typical section drawing and a "Sketch map, not to scale" of the firm's patch. Neither was asked for, and judges from both families picked them out.

The idea worth keeping

Everything useful here came from the arrangement around the models, not from any model in it. A page was checked by something that hadn't written it, from a different family. A human pick sat outside the scoring and could disagree with it. A reviewer could say no to our own write-up and make it stick. That's the part worth buying or building, because models change and the arrangement doesn't have to.

Credit to Nate Herkelman of Uppit AI, who posts through the AI Automation Society community. His video is why we ran this at all.

The test firm is a groundworks company because that's the trade I know. This isn't construction-only. The same systems work in any business where an owner's expertise needs to reach further than the owner can.

PS: most owners I sit with find time and money leaking through stuff like this. The Opportunity Map is the 45-min paid diagnostic that surfaces it. Drop me a line.