Method
Every score here comes from data the Office of Rail and Road publishes openly. There is no survey and no private feed.
The data
ORR Table 3130 reports, for every station and every four-week railway period, how many stops were scheduled, how many were actually measured, what share arrived within three minutes of schedule, and what share were cancelled. Table 1410 supplies the three-letter code, region and footfall.
Scores use the most recent 13 periods (Aug 2025 to Jul 2026), roughly a year. That is long enough to survive one bad month without describing a station as it was years ago. A railway period is four weeks; the rail year starts in April, so “2025/26 P05” is the four weeks around August 2025. Charts here show the month instead.
The four measures
| Measure | Weight | Built from |
|---|---|---|
| Reliability | 35% | Share of stops arriving within 3 minutes |
| Cancellations | 25% | Share of scheduled stops cancelled |
| Frequency | 20% | Scheduled stops per day |
| Connectivity | 20% | Operators calling, and interchange volume |
How a sub-score works
Each measure becomes a percentile rank across all scored stations. A reliability score of 80 means the station beats 80% of British stations on punctuality. It does not mean 80% of its trains are on time. The raw percentages sit next to the score on every station page.
Percentiles are used instead of raw percentages because the distributions are heavily skewed. A few termini handle thousands of trains a day while hundreds of rural halts see a dozen, and a straight average would be dominated by the extremes.
The headline score
The four sub-scores are blended using the weights above, and that blend is then converted to a percentile in its own right. The second step matters: averaging four percentiles pulls every station toward the middle, so the raw blend spans only about 12 to 93 and a blend of 80 would really sit near the 97th percentile. Percentiling it again means the headline number carries the same meaning as the sub-scores — a score of 40 is a station that rates better than 40% of British stations, all four measures taken together.
Because it is a percentile, the scores are spread evenly by construction: a tenth of stations sit below 10, half sit below 50, and the very best and worst dozen round to 100 and 0. A score is a position in the table, not a mark out of 100.
Thin samples
The ORR does not measure every stop. Coverage sits above 95% at most large stations but can drop below half at quiet ones, where a few hundred measured stops can produce a freakishly good or bad figure by chance.
Each station's punctuality is therefore pulled toward the national average in proportion to how little of it was measured. A station with very few measured stops ends up close to average; one with tens of thousands barely shifts. Anything under 60% coverage carries a warning on its page.
What it does not tell you
- Nothing about time of day or day of week. ORR data is aggregated over four-week periods, so it cannot say whether Friday evenings are worse than Tuesday mornings. That needs per-train history.
- Three minutes is a blunt cutoff. A train 2 minutes 55 seconds late counts as punctual. One at 3 minutes 5 seconds does not.
- Frequency and connectivity reward size. A quiet, perfectly punctual rural station still scores modestly, because it genuinely offers fewer trains and fewer places to go. Read the sub-scores, not just the headline number.
- Nothing about the station itself. Staffing, step-free access, shelter and parking are not included.
- 151 stations are unscored because too little of their service is measured to say anything honest about it.
Sources
- ORR Table 3130, Open Government Licence
- ORR Table 1410, Open Government Licence
Contains public sector information licensed under the Open Government Licence v3.0. Last updated 4 September 2026.