The method

How do we break a tie?

Everyone has an answer. Ours was researched by four AIs, checked by code, and scored by one that only gives odds.

Federer or Nadal, as of 23 September 2026

Rafael Nadal88.2

Roger Federer78.1

Rafael Nadal, clearly.

Every figure below uses this one tie. It never changes.

Chapter I

The scorecard, set before looking

First we decide what “better” means: six things, each with a fixed share of the score and one written reason for it. The shares are set before any evidence is gathered, so nobody can tilt them towards a favourite.

Like a skating panel publishing its marking sheet before the skaters take the ice.

The athlete scorecardrubric v2

  • Peak dominance30%A greatest-ever argument is first about how good someone was at their best.
  • Longevity20%Staying at the top is the second thing everyone cites.
  • Head-to-head15%Beating the strongest rivals separates the great from the merely dominant.
  • Major titles20%Titles are the record people check first.
  • Versatility10%It rewards range, but range matters less than peak.
  • Influence5%Kept low because it is the hardest criterion to evidence and the easiest to argue.
Try other scorecards on a tie →
The detail

A rubric family is one scorecard reused unchanged for every tie of its kind, so two athlete ties are measured with the same yardstick. Weights are published before the run, each with one written reason, and cannot move after it. The family is picked by the kind of question, never by the contenders.

  • Peak dominance: negligible · notable · strong · era-defining · unmatched in the sport's history
  • Longevity: fewer than 3 seasons · 3 to 7 seasons · 8 to 12 seasons · 13 or more seasons
  • Head-to-head: clearly behind · behind · even · ahead · clearly ahead
  • Major titles: few · competitive · leading for their era · record-setting · untouched record
  • Versatility: narrow · adaptable · strong everywhere
  • Influence: minimal · notable · transformative

Chapter II

Four researchers who never compare notes

Four AI models from four different labs research each side on their own, each with its own web search, and none of them sees another's notes. Then plain code checks every quote against its page, word for word, and only what survives goes to the judge.

Four reporters sent out on the same story, separately; only the quotes that check out go to print.

Each kept fact, found by how many of the four

1 of 4121
2 of 424
3 of 47
all four5

112 and 102 quotes reach the judge, about 2,570 words per side.

The detail

The researchers are Claude, GPT, Gemini and Grok, each pinned to one version and each with its own web search. Each works on one side at a time and sees only that side and the scorecard, never another researcher's notes.

Code fetches every cited page and keeps a quote only if it is there word for word. Quotes that say the same thing are merged, and counts, dates and head-to-head records are computed by code. The researchers never score; the judge in chapter IV scores, from the blinded packet alone.

The research as it ran · 8 sessions, four researchers, two sides

ClaudeGPTGeminiGrokchecked by codeFederer's packetNadal's packet
web searches
85
pages found
285
quotes brought back
274
found word for word
216
after merging
157
in the packet
142

Not found word for word: 44 pages would not load, 11 quotes were not on the page, 3 were too long to check.

Left out of the packet: 14 repeated a quote already in it, 1 gave a name away after blinding.

Chapter III

Take the names off

A famous name carries a reputation, so the judge never sees one. Every name becomes Contender A or B, and anything that gives the game away, like a tournament or a nationality, is swapped out.

Blind auditions, where the orchestra hears the player from behind a screen.

Names on: the line as publishedNames off, as judged

Peak dominance

As in the player’sSpaniard’s win/loss record on  Parisian clay – an astounding 96.5% of wins.

Versatility

Contender B,Nadal, the first playerSpaniard to lift the Major ZNorman Brookes trophy, told the rival: "the rival,Federer: "Roger, I know exactly how you feel,”

Head-to-head

Contender BNadal is also the only player to beat the rivalFederer in the finals of three different Grand Slam tournaments — the Major X,French Open, the Major Z,Australian Open, and Major Y).Wimbledon).

The detail

Before the judge reads a word, code removes names, nationalities, home cities and the names of events. The web's own phrasing gives players away in ways a hand-written sheet never did (“the Spaniard”, “the Mallorcan”, “Parisian clay”), so the list grows with every packet.

A packet holding only a famous name scored almost exactly like the full evidence, because the judge has priors about famous names. So the blinded run is the verdict, and the named twin runs beside it as a check, printed in every tie's receipts.

Chapter IV

The judge answers in odds, twelve times

For each criterion the judge spreads its belief across the worded levels, and the score is the balance point of that spread. It is asked in both orders, with the labels swapped, and with the names in: twelve questions for one verdict.

A weather forecast: not “it will rain”, but “70% chance of rain”.

Peak dominance, one contender: the judge's odds

0%
0%
0%
43%
57%
negligiblenotablestrongera-definingunmatched in the sport's history

The score is the balance point, level 3.57 of 4: 89 out of 100.

12 questions, 3 rounds of 4

The verdict round · blinded; this one decides
A and B swapped · checks the label is not evidence
Names in · measures the pull of a famous name
The detail

The judge is Jev, pinned to one version so every verdict traces back to the model that produced it. When its answer is spread across levels, the tie page marks the criterion “the judge was unsure here”, and it still counts.

Chapter V

Adding it up, and how close is close

Plain arithmetic weighs each criterion and adds them up; the margin is the gap between the two totals. Inside 3 points the gap could be noise, so we still name a winner and say it was close.

A photo finish: you still name the winner, and you say it was a photo finish.

← toward Federertoward Nadal →

Head-to-head+9.3
Major titles+0.9
Longevity+0.6
Influence0.0
Versatility−0.2
Peak dominance−0.4

The pulls add up to the margin: Rafael Nadal by 10.1.

  1. 1Who won the most important criterion?Here that is peak dominance, the biggest share.
  2. 2Whose answers was the judge surer about?The side whose odds were less spread out.
  3. 3Who got there first?The side established earlier, so it is never a coin.
The detail
  • 10 points or more: “Nadal, clearly”
  • 3 to 10: “Nadal, by a margin”
  • Under 3: “Nadal, by a whisker”, inside our measured noise

Fifty identical runs gave a margin with a standard deviation of 0.85 points and never flipped the winner. A result inside the fence breaks on the published ladder. Never a coin.

Chapter VI

Is the judge any good?

We wrote 111 questions whose answers we already know and graded the judge: it picked the right level 80% of the time, and was within one level 97% of the time. Suggested rewordings sit the same exam: of 5 so far, 1 got through.

The one kept change (athlete-v1.1) is not in use: its reworded longevity levels ask the judge to count years at the top, and counting is reserved for code. It is adopted only once the packet carries that count as a computed fact.

A forecaster whose 70% days really are wet about 7 times in 10.

111 questions with known answers, one dot each

right: 89 one level off: 19 further: 3

Flagged unsure on 14% of readings, and right on those 57% of the time, against 84% for the rest: when it says it is unsure, it means it.

The everyday judge

Jev

right 80% · within one 97%

The second opinion

Claude

right 75% · within one 99%

The detail

The labelled set measures the judge. An automated loop may tune level wording and level counts against it; it cannot touch the weights, because the exam asks one criterion at a time and a weight changes nothing it measures. A tuned rubric becomes a new version used for new ties only. It never changes a published verdict: there is no ground truth for a verdict to climb toward. A new rubric or model version applies to new ties only. Published verdicts are never re-run; the benchmark is re-run on every version bump.

The second judge is an Anthropic model, the same family as the Claude session that assembles the evidence packets. It controls for quirks of the verdict judge, not for anything the packet's framing introduced, so it is a same-family judge control, not an independent second opinion.

Printed, and it stays printed

A person reads every verdict before it goes out. Once published it is dated and final; votes are counted and shown, and they never change it.

This week's tieAll ties

What we avoid

  • Models debating each other towards a consensus.
  • “Win the argument” debate, which rewards rhetoric over evidence.
  • Asking a model to count or compare numbers.
  • Unpublished rubrics, and names in front of the judge.
  • Choosing the scorecard after seeing who is in the tie. The family is picked by the kind of question, never by the contenders.

What can still go wrong

  • The judge can be wrong, even when its answer looks tidy.
  • The wording of a level moves scores, so each rubric is tested for it.
  • Runs vary a little, which is what the fence is for.
  • The exam's answers are ours; a second labeller is on the way.