Frontier AI models, and an experienced DPO where one sat the case, were given the same file.

This page follows it through the bench and shows how every answer was marked.

Case file · 0417
Received · Day 0Stage 1 of 4
From
Marketing director
Re
Win-back campaign, lapsed customers

The marketing director wants to email 40,000 lapsed customers by Friday.

The list comes from the old CRM. Nobody can say when these people last heard from us, or what they were told when they signed up. Sales have already drafted the email.

Attached · customer list extract · draft email
Received
Under test
frontier AI models, marked the same way as the people
Reference
an experienced DPO, where one has sat the case
Information
released in four stages, as the week unfolds
Questions
identical, in the same order, for everyone
Answers
free text, in their own words
Marking
a sealed answer sheet, then twelve kinds of judgement
Judge
a different model family from every model under test
No cookies · no tracking Scroll 01 / 10

The job is judgement.

Two answers to the same request. Both are legally correct. Only one of them is the job.

Answer one“No. You can't email people who haven't opted in. The campaign can't go ahead.”Correct · the gatekeeper
Answer two“Here's how you can, safely. Suppress everyone who unsubscribed. Re-permission the rest with a one-line reason. Start with the newest tenth of the list and watch the complaints.”Correct · and the job
Gatekeeper
says yes or no
Enabler
says here is how, and keeps the people whose data it is safe
Measured here
the second, never just the first
Gatekeeper to enablerScroll02 / 10

Same file. Same questions. Same order.

Frontier AI models and, where one has sat the case, an experienced DPO. Nobody sees anyone else's answer until their own is in.

Model ACase file · 0417
Model BCase file · 0417
Model CCase file · 0417
Model …Case file · 0417
A DPOCase file · 0417Models' answers withheld until this desk has answered
Envelopes
identical, delivered at the same moment
Order
the same questions in the same sequence, for everyone
The DPO's desk
nothing from the models until their own answer is in, so nothing can be copied
One pipeline for everyoneScroll03 / 10

The case arrives in stages, like a real week.

New evidence lands, a decision is due before the next reveal, and nobody sees what is coming. Not the models, not the DPO.

Stage 1 · Day 0

The marketing director wants to email 40,000 lapsed customers by Friday.

Answer before stage 2
Stage 2 · Day 2

The list turns out to have come with a company bought in 2019. Nobody kept the consent records.

Answer before stage 3
Stage 3 · Day 5

Sales have already sent a test batch to two thousand names. Three complaints so far.

Answer before stage 4
Stage 4 · Day 9

The board wants a one-line answer: do we tell the regulator, or not?

Final answer
Stages
four, each with its own questions and its own deadline
Ahead
sealed until the previous answer is in
Why
real cases do not arrive complete; judgement is what you do before you know everything
Answer before the next revealScroll04 / 10

Asked openly.

One question. No options shown, to the models or to the DPO. The question, the answers, the key and the ticks shown here are illustrative.

Stage 1 · question 3

What is your first move?

No options shown
Model A“Pause the send until we know where the list came from.”
Model B“Send it. Legitimate interests covers a win-back email.”
A DPO“Ask marketing what Friday is really about before anyone touches the list.”

The reader

option A · hidden
option B · hidden
option C · hidden
none of these → a person

Sorts each answer into a bucket. Never sees who wrote it. Never judges quality.

Sealed answer sheet · written before the exercise
Question 3 · first move
Options A, B, Chidden from everyone
Sealed
Sealed before the exercise Scroll · 05 / 10

Twelve kinds of judgement, scored one at a time.

The judge never asks whether an answer is good. It asks twelve narrower questions, and has to quote the answer before it can give a high mark. The scores shown here are illustrative.

    Axis
    ·waiting for the answer

    The judge scores one axis at a time and must quote the answer to justify any mark of 7 or above.
    1. 10almost never given
    2. 8–9clearly better · a rare insight
    3. 7a good professional response
    4. 5–6meets the professional floor
    5. 3–4partially there
    6. 0–2below the floor
    No model marks its own homework Scroll · 06 / 10

    We test the examiner before we trust the marks.

    The same answer, reworded three ways, must land in the same band. A quietly flawed one must drop. And no model ever marks its own family's homework. The answers and bands shown here are illustrative.

    Wording Aidentity masked“Pause the send until we know where the list came from, then re-permission it.”→ 7
    Wording Bidentity masked“Hold the campaign, establish the list's provenance, and seek fresh consent.”→ 7
    Wording Cidentity masked“Don't send yet. Find out where the names came from and ask them again.”→ 7
    Quietly flawedidentity masked“Pause the send until we know where the list came from. Legitimate interests will cover the rest.”→ 5
    1. 7a good professional response
    2. 5–6meets the professional floor
    3. 3–4partially there

    No model marks its own homeworkScroll · 07 / 10

    Nothing gets rewritten.

    Every published result is frozen with the exact version of the marking that produced it. Change the marking and older results are labelled, never quietly re-scored. The two cards shown here are illustrative.

    Result · case 0417 · axis 3, foresight7
    Marking
    version 3 · strict bands
    Judge
    family C · identity-blind
    Guide
    text kept with the score
    Published
    16 Aug 2026 · result set 2026-08-A
    Frozen
    Result · case 0417 · axis 3, foresight8
    Marking
    version 2 · generous bands
    Judge
    family B
    Status
    legacy · not comparable with version 3
    Published
    2 Jul 2026 · result set 2026-07-A
    Legacy
    Kept with the score
    the guide text, the axis definitions, the judge's instructions, the model version, the scale
    Compared
    only within one marking version
    Changed
    every change to the marking is recorded against the runs it affects
    Snapshot disciplineScroll08 / 10

    Tables vs the AI.

    In the workshop, a room of DPOs sits the same case as the models, then hunts the AI's answer for mistakes. Points for every genuine catch, with a referee. Nobody is scored individually. The teams and scores shown here are illustrative.

    Projector · running totals

    1. The Lawful Basis0
    2. Table Four0
    3. Recital 470
    4. Late Filers0
    5. Article 30 Club0
    A game, not a testScroll · 09 / 10

    What we don't claim.

    The honest part, and it appears wherever a score appears.

    • That any exact score is the true score. Bands are fixed; exact scores are never asserted. The bench claims direction and stability, not decimal places.
    • That a model beats a person, or the other way round. The bench measures models. Where a DPO has sat the case, their answers are a reference, not a contestant.
    • That the marking is finished. Scores move when the marking moves. Every score carries the version that produced it, and versions are never mixed on one report.
    • That a checklist is judgement. Naming every expected item is not enough. Part of every axis stays with the judge's overall reading of the answer.
    • That the reader cannot misread. Sorting free text into buckets is easier than judging quality, not infallible. That is why "none of these" exists and a person maps what the reader cannot.
    The stillness is the point10 / 10