six courses, three frameworks, one index

2:02 in Berlin against 2:05 in Boston.

How much of that gap is the runner, and how much is the road?

The six World Marathon Majors are treated as interchangeable and are obviously not. This measures the difference three separate ways — an energy cost integration over each course's gradient, a within-runner paired comparison of the same athletes across course pairs, and a weather-normalized elite average — then combines them into a single Course Difficulty Index with Berlin set to 1.000. Boston comes out 101 ± 2 seconds slower than Berlin for an equivalent runner.

+101 s

Boston against Berlin, from the cleanest framework, with a bootstrap confidence interval of ± 2 seconds.

+78 s

New York, the second hardest. The two hard courses are hard for different reasons.

35 s

covers Berlin, Chicago, London and Tokyo. Berlin and Chicago are statistically indistinguishable.

60–120 s

the swing a single hot edition produces. Weather moves a course's year more than the course's own character does.

Start · why one framework is not enough

Every obvious method has a hole in it

Compare raw finish times across courses and you are mostly measuring which race attracted the faster field that year. Adjust for elevation only and you miss weather, which turns out to be larger. Compare elite top-tens and you inherit whichever course had the deeper start list.

So the answer here is not one method done carefully, it is three methods with different weaknesses, checked against each other. Where they agree, the finding is probably about the course. Where they disagree, the disagreement is the finding.

1 · Grade-adjusted prediction

Minetti et al. (2002) energy-cost integration over each course's real gradient profile. Sees terrain, blind to everything else.

2 · Within-runner paired

The same athlete's time on two courses within 18 months, weather-normalized, bootstrapped. The cleanest of the three, because the runner is their own control.

3 · Weather-normalized top-10

Maughan / El Helou penalty curves applied to elite finish times, so a hot edition is not counted as a hard course.

↓ the cleanest read

Same runner, two courses, 101 seconds

The within-runner framework is the one to trust most, because genetics, training and physiology cancel out — it is the same person on both sides of the comparison. It puts Boston 101 seconds and New York 78 seconds behind Berlin for an equivalent runner, both with tight bootstrap intervals.

Boston's difficulty is not the Newton hills alone. The course drops sharply early, which shreds quadriceps before the climbs arrive, and then the hills land between 16 and 21 miles. New York's is the bridges and a course that never lets a rhythm settle.

Within-runner time gaps between course pairs with bootstrap confidence intervals, showing Boston and New York clearly slower than Berlin and the remaining courses clustered together.
Same-athlete deltas across course pairs, with bootstrap intervals.

↓ do the three agree?

Same ranking, different magnitudes

All three frameworks put Boston hardest and New York second, and all three cluster the other four. They disagree on how large the gaps are, which is expected given what each can and cannot see. Combined, they produce the Course Difficulty Index.

Comparison of the three frameworks' difficulty estimates per course, showing consistent ordering but differing magnitudes.
Three frameworks — same order, different sizes.
The combined Course Difficulty Index for all six Majors with Berlin normalised to 1.000 and Boston highest.
The CDI — Berlin = 1.000, everything relative to it.
Elevation and grade-adjusted predicted times per course, derived by integrating the Minetti energy-cost curve over each course's gradient profile.
Framework 1 on its own: what the terrain alone predicts.

↓ the finding nobody expects

Weather beats the course

Year-to-year variation within a single course is dominated by conditions, not by anything about the course. Berlin 2022 at 18°C and London 2018 at 23.5°C each moved that edition's winning time by 60 to 120 seconds — comparable to or larger than the permanent gap between most of these courses.

Which is a warning about the whole enterprise. If you compare two courses using single editions and do not normalize for weather, you will mostly measure the weather. That is precisely the error frameworks 2 and 3 are built to avoid, and it is consistent with the one-minute-per-degree effect measured separately across fourteen Bostons.

Weather-normalized elite times by course, with the raw and adjusted values shown so the size of the weather correction is visible.
Elite times before and after the weather penalty is applied.

↓ robustness

What it takes to change the answer

The sensitivity analysis varies the assumptions that could plausibly be wrong — the weather penalty curve, the pairing window, the elite cutoff — and re-runs the index. Boston stays hardest throughout. The four-course middle cluster reshuffles internally, which is the honest reading of a 35-second spread: those four are not really distinguishable.

Sensitivity analysis showing how the difficulty estimates move under alternative assumptions, with Boston remaining hardest across all scenarios.
Sensitivity — the top of the ranking is stable.
An alternative ranking of the six courses under a different weighting of the three frameworks.
Alternative weighting — the middle four swap places.
Raw finish-time distributions by course before any adjustment, showing overlapping spreads that make the naive comparison misleading.
The raw times, for reference. This is the comparison the rest of the work exists to avoid.

↓ what this cannot tell you

Where the numbers are softest

The within-runner framework is the most trustworthy and also the most data-hungry: it needs athletes who ran two different Majors inside 18 months, which is a much smaller set than the full results. Tight bootstrap intervals describe the precision of that sample, not the size of it.

Course elevation profiles come from official PDFs and canonical Strava routes, which do not always agree to the metre, and the Minetti curve is a laboratory relationship applied to a road race. Both push framework 1 toward being indicative rather than exact — one reason it is not the headline.

The Sydney and Cape Town extension is a bonus scored under framework 1 only, and should be read as much weaker evidence than the six main courses.

Finish · how it was built

Where the numbers come from

Results
Top-100 per Major per year, 2015–2024, from World Athletics, ARRS and MarathonGuide
Elevation
Official course PDFs from each Major, plus canonical Strava race routes
Weather
Open-Meteo ERA5 historical API, cross-checked against race-day coverage
Energy cost
Minetti et al. (2002), J. Appl. Physiol. 93:1039–1046
Heat penalty
Maughan (2010); El Helou et al. (2012), PLOS ONE 7(5)
Reproduce
python src/build_data.py then python src/analysis.py — deterministic, seed 42, under 90 seconds
Also ships
A Jupyter notebook, PDF and Word reports, and a standalone HTML article

Boston is hardest, and weather is bigger than you think