Methods & data

The dataset

This research draws on just over three million finish times from Indian road races run between 2017 and 2026: 5Ks, 10Ks, half marathons and full marathons. They come from many public sources across the country, pulled into one place so the whole picture can be read at once, instead of one result at a time. The live counts on the research home page grow as new races are added.

The findings are about racing in India. The big international marathons are included only for their Indian finishers, and are kept out of these counts.

Where the data comes from

We gather results that are published openly after each race, and keep them in three layers. First, an untouched archive of every record exactly as the source published it. These are our receipts, so any finding can be traced back to its raw data. Second, a cleaned, standardised copy that all the analysis reads from. Third, aggregate tables that summarise the whole population.

Cleaning & normalising

Collecting the results is the easy part. The real work is making them line up across all the different sources, and each one records things its own way. A few of the snags: gender often isn't a field of its own. It's tucked into the category label, as in "35–39 yrs Male" or "U17 Boys", and we have to read it out of there. The same four distances turn up spelled a hundred different ways. Cities, ages and finish times have to be dug out of whatever format a source happened to use, and a fair amount of what arrives is junk: "India" typed into the hometown box, a bib number sitting where a name should be, a 45-minute time saved as if it were 45 hours. We catch those and drop them before they skew a result.

Identifying the same runner

To follow a runner across years and races, we first have to recognise that two results belong to the same person, and Indian race data has no shared runner ID. So we match probabilistically: comparing names (allowing for spelling and transliteration differences), gender, finishing pace, and (where a race records them) city and age. Every match carries a confidence score.

We are deliberately cautious. Wrongly merging two different people is worse than missing a match, so when the evidence is thin we leave results unlinked. A very common name only links when something extra agrees, like the same city. Until you claim your results yourself, a match stays a best guess, marked as unconfirmed.

Comparing fairly: the within-athlete method

This is our default way to isolate a single effect (the weather, the terrain, the passing years) from everything else that makes one runner faster than another.

Comparing different runners on hot vs. cool days is confounded: fitness, ability, and which runners show up all vary from race to race. Instead we compare each runner to themselves, in two flavours. For the headline numbers we pair each race with the same runner's previous race at that distance — usually a different event — and study how the change in their time tracks the change in conditions. For per-race drill-downs we subtract the runner's own average finish time, leaving only how far each edition sat above or below their personal normal:

The weather data

The temperature behind the heat findings is not a thermometer at the finish line; few races record one. Instead we use a modelled weather reconstruction for each race's city and date, averaged over the race morning, roughly 4am to noon. It is humidity-aware: a muggy 24°C morning in Mumbai counts as hotter than a dry 24°C in Delhi, because sweat evaporates less and the body sheds heat less easily.

It is also a shade estimate: it has no direct-sun component. Standing in the sun on race morning is harder than these numbers show, so they read as conservative, not inflated.

How the weather is built

One weather morning is computed per city and date, from a public reanalysis dataset (ERA5), and applied to every race in that city that day. The reconstruction tracks local weather-station records closely across the period.

Limits & honesty

Finisher survivorship

The data only sees finishers. Anyone who dropped out on a brutal day is invisible, so every weather-cost number is a lower bound. The runners hurt worst aren't in the sample dragging the average further down.

Fitness drifts over time

Subtracting a runner's own average assumes their ability is constant across the years, but people get fitter or slower. That drift only biases a result if it lines up with which years were hot; real race weather zig-zags year to year, and taking the median over many runners with different trends cancels most of it. Mitigated, not eliminated.

The season travels with the weather

A within-runner comparison that crosses from a cool-season race to a warm-season one also crosses training cycles and race goals, so part of a raw "hotter = slower" number is the racing calendar, not the weather itself. Restricting comparisons to similar times of year shrinks the heat numbers; the ranges we show should be read as the practical cost of a warmer morning, weather plus what travels with it.

We miss many runners

Because linking is cautious, a large share of results (anyone with a common name and little corroborating detail) are never joined to a runner. Findings that need same-runner data are drawn only from those we could link with confidence, which is a fraction of everyone who ran.

Coverage is uneven

Not every race from every provider is complete; some older result lists were published only in part. We record what is available and treat the rest as gaps, not zeros.

Modelled, not measured

The weather is a reconstruction of the race morning, not an on-site reading, and it is a shade figure with no direct-sun term.

Thin marathon data

Few runners repeat a marathon, and when they do it's often on a different course, so marathon-distance numbers are the least certain and are shown as a range.