6,400 elite performances, 16 events

The tables promise that 1200 points is 1200 points.

Among the world's best, it isn't.

Combined events, athlete-of-the-year votes, every "who is the GOAT" argument — all of them rest on the World Athletics scoring tables treating events equivalently. Tested directly against 6,400 top performances across 16 events, the median score of an event's top 30 ranges over 68 points, from the 5000 m at +24 above the cross-event norm down to the javelin at −44. Kruskal-Wallis rejects equality at H = 236, p ≈ 10⁻⁴².

68 pts

spread in the elite ceiling between the highest-scoring and lowest-scoring event.

+24

the 5000 m, furthest above the cross-event norm. The 1500 m sits close behind it.

−44

the javelin, furthest below. Hammer and high jump sit down there too.

p ≈ 10⁻⁴²

for the hypothesis that all sixteen events score equally at the top. Not a marginal result.

Start · why it matters

An invisible assumption under every comparison

The scoring tables are the plumbing of athletics. Nobody argues about them, because nobody notices them. But the moment anyone compares a decathlete's javelin to their 1500 m, or ranks a shot putter against a miler for an award, the tables are doing the work, and the whole exercise inherits whatever bias they carry.

The claim is testable. If the tables are level, then the best javelin throwers in the world and the best 5000 m runners in the world should be earning roughly the same number of points, because both groups are at the frontier of their event.

↓ the test

The elite ceiling is not level

Take each event's top 30 performances of 2023–24 and look at the median WA score. If the tables were level these medians would cluster. They do not: they spread across 68 points.

Distance and middle distance sit highest, with the 5000 m and 1500 m at the top. Throws and jumps sit well below, with the javelin, hammer and high jump at the bottom. The Kruskal-Wallis test rejects equality overwhelmingly.

Put plainly: the same 1250 points is a merely very good javelin throw and a fairly ordinary 5000 m.

Box plots of elite WA point scores for each of sixteen events, with distance events sitting visibly higher than throws and jumps and the javelin lowest.
Each event's elite scoring ceiling. Level tables would put these boxes in a row.
Bar chart of each event's deviation from the cross-event norm, running from the 5000 metres at plus 24 points down to the javelin at minus 44.
Deviation from the cross-event norm, event by event.

↓ the part most write-ups skip

This does not prove the tables are miscoded

An event's elite score distribution reflects two things at once: how the table is calibrated, and how deep and competitive that event's current field is. A thinly contested event will show a lower ceiling even under a perfectly fair table, simply because fewer athletes are pushing its frontier.

The javelin and hammer are exactly the events you would expect to be thin. So the honest reading of this result is "equal points do not currently mean equal elite standing," not "the tables are wrong." Both readings share the same chart; only one is supported by it.

Separating the two would need per-mark calibration against a fixed physical frontier rather than against the current field. That is the natural next step and it is not what this repository does. The gap is real; its cause is part table and part talent pool.

↓ so what

What it changes in practice

Cross-event awards

A points-based shortlist quietly favours distance runners over throwers at the same standing within their event.

Combined events

The same points gap applies inside the decathlon and heptathlon, where the tables decide the whole result.

Cross-country comparisons

Any national-strength measure built on WA points inherits this tilt — including my own, which is why the two studies are published together.

Finish · how it was built

Where the numbers come from

Data
World Athletics top lists, scraped with mark, WA points and nationality per event
Sample
6,400 performances across 16 events, 2023–24
Points
The WA score each mark was actually awarded, not recomputed
Test
Kruskal-Wallis on the equality of elite score distributions across events
Built with
Python, pandas, scipy, matplotlib
Companion
National Strength Profiles, which spends this currency and inherits its tilt

Equal points, unequal standing