HGIYS

Method

Every score here comes from data the Office of Rail and Road publishes openly. There is no survey and no private feed.

The data

ORR Table 3130 reports, for every station and every four-week railway period, how many stops were scheduled, how many were actually measured, what share arrived within three minutes of schedule, and what share were cancelled. Table 1410 supplies the three-letter code, region and footfall.

Scores use the most recent 13 periods (Aug 2025 to Jul 2026), roughly a year. That is long enough to survive one bad month without describing a station as it was years ago. A railway period is four weeks; the rail year starts in April, so “2025/26 P05” is the four weeks around August 2025. Charts here show the month instead.

The four measures

MeasureWeightBuilt from
Reliability35%Share of stops arriving within 3 minutes
Cancellations25%Share of scheduled stops cancelled
Frequency20%Scheduled stops per day
Connectivity20%Operators calling, and interchange volume

How a sub-score works

Each measure becomes a percentile rank across all scored stations. A reliability score of 80 means the station beats 80% of British stations on punctuality. It does not mean 80% of its trains are on time. The raw percentages sit next to the score on every station page.

Percentiles are used instead of raw percentages because the distributions are heavily skewed. A few termini handle thousands of trains a day while hundreds of rural halts see a dozen, and a straight average would be dominated by the extremes.

The headline score

The four sub-scores are blended using the weights above, and that blend is then converted to a percentile in its own right. The second step matters: averaging four percentiles pulls every station toward the middle, so the raw blend spans only about 12 to 93 and a blend of 80 would really sit near the 97th percentile. Percentiling it again means the headline number carries the same meaning as the sub-scores — a score of 40 is a station that rates better than 40% of British stations, all four measures taken together.

Because it is a percentile, the scores are spread evenly by construction: a tenth of stations sit below 10, half sit below 50, and the very best and worst dozen round to 100 and 0. A score is a position in the table, not a mark out of 100.

Thin samples

The ORR does not measure every stop. Coverage sits above 95% at most large stations but can drop below half at quiet ones, where a few hundred measured stops can produce a freakishly good or bad figure by chance.

Each station's punctuality is therefore pulled toward the national average in proportion to how little of it was measured. A station with very few measured stops ends up close to average; one with tens of thousands barely shifts. Anything under 60% coverage carries a warning on its page.

What it does not tell you

Sources

Contains public sector information licensed under the Open Government Licence v3.0. Last updated 4 September 2026.