Key takeaways
- BO3 / MD3: larger Glicko gaps → fewer total games, on average. Near-equal ratings ≈ 22.5 mean games; gaps of 500+ ≈ 16.1.
- Around Δ = 200 in MD3: mean ≈ 21 games, median ≈ 19 — not 28.
- BO5 / MD5 (ATP Grand Slam): matches are longer by design. Medians sit near 28–30 games across most gap bands. The gap→games curve is weaker and noisier (smaller sample).
- This is a historical tendency, not a total-games tip. Use it as match-shape context when reading ratings on a match page.
1. What is the Glicko rating gap?
Tennis Glicko estimates each player’s strength with Glicko-2. The overall rating μ is a skill estimate on a display scale centered near 1500.
A gap of 30 points means the players look nearly equal. A gap of 400+ usually means a clear favorite. That difference does not decide the winner by itself — but it does correlate with how long and how competitive the scoreline tends to be.
In this article we use as-of ratings: each player’s Glicko immediately before the match is played, then we look at the final score’s total games and sets.
2. Mean, median, and % of long matches — explained simply
Before reading the tables, three numbers appear often. They answer different questions:
Mean (average)
Sum of total games across all matches ÷ number of matches. A few very long matches (for example a 5-set marathon) pull the mean upward.
Median
The middle value when matches are sorted by total games. Half finished with fewer games, half with more. Medians are more resistant to extreme outliers than means.
% 3+ sets (BO3) / % 4+ sets (BO5)
Share of matches that went the distance enough to need an extra set. In best-of-3, “3+ sets” means the match reached a deciding set. In best-of-5, “4+ sets” means the match did not end 3–0 — it went to at least four sets. We also show “% 5 sets” for full-distance MD5 matches.
Rule of thumb: if mean ≈ median, the sample is fairly stable. If mean is clearly higher than median, there is a long-match tail.
3. Best-of-3 (MD3 / BO3): gap vs average total games
Most tour, Challenger, and ITF matches are best-of-3. Below: pre-match absolute Glicko overall gap versus realized total games. ATP Grand Slam men’s matches are excluded here (those are BO5 and appear in the next section). WTA Grand Slam matches remain (they are BO3).
Sample: n = 135,690 BO3 matches · overall mean total games ≈ 21.1 · ratings from chronological Glicko-2 walk (warm-up 2021–2024, observation 2025–2026 through 2026-06-10).
| |Δ GLICKO| | MEAN GAMES | MEDIAN | % 3+ SETS | N |
|---|---|---|---|---|
| 0–50 | 22.5 | 20 | 34% | 28,778 |
| 50–100 | 22.3 | 20 | 33% | 25,561 |
| 100–150 | 21.9 | 20 | 30% | 21,228 |
| 150–200 | 21.3 | 19 | 27% | 16,136 |
| 200–250 | 20.8 | 18 | 24% | 12,112 |
| 250–300 | 20.0 | 18 | 20% | 9,055 |
| 300–400 | 19.0 | 17 | 16% | 11,175 |
| 400–500 | 17.6 | 16 | 10% | 6,022 |
| 500+ | 16.1 | 15 | 6% | 5,553 |
How to read this: the relationship is real but gentle. Moving from a near-even matchup to a 400–500 point mismatch only shifts the mean by about 5 games. That is useful as context (“likely shorter / longer”), not as a precise over/under call.
Snapshot around Δ ≈ 200 (MD3)
Window 195–205 points: mean 20.9 games · median 19 · typical interquartile range about 16–24 (n ≈ 2,800).
4. Best-of-5 (MD5 / BO5): ATP Grand Slam men
Men’s Grand Slam singles are best-of-5. Matches have a higher floor for total games because three sets are required to win. That is why people often remember “around 28–30 games” — that level is normal for MD5, not for a typical MD3 with a 200-point gap.
Sample: n = 1,420 ATP Grand Slam men’s matches · overall mean ≈ 30.4 total games · mean sets ≈ 3.09. Bins above ~250 points are thin — treat them cautiously.
| |Δ GLICKO| | MEAN GAMES | MEDIAN | % 4+ SETS | % 5 SETS | N |
|---|---|---|---|---|---|
| 0–50 | 29.1 | 28 | 24% | 8% | 439 |
| 50–100 | 30.5 | 28 | 28% | 12% | 389 |
| 100–150 | 30.9 | 29 | 30% | 15% | 255 |
| 150–200 | 32.1 | 30 | 37% | 15% | 150 |
| 200–250 | 32.1 | 30 | 42% | 9% | 77 |
| 250–300 | 32.4 | 30 | 38% | 11% | 37 |
| 300–400 | 29.8 | 30 | 32% | 2% | 41 |
What is different from BO3? In MD5, medians stay near 28–30 across most bands. The gap does not produce the same clean “bigger mismatch → clearly fewer games” staircase you see in BO3. Part of that is format (more sets required); part is smaller sample size.
WTA Grand Slam matches are best-of-3 (mean ≈ 21.8 games in the same window) — do not mix them into MD5 tables.
5. The “Δ 200 → 28 games” intuition — where it comes from
A common mental shortcut is: “about 200 Glicko points apart → roughly 28 games.” Empirically that shortcut mixes formats:
| FORMAT | AROUND Δ ≈ 200 | WHAT 28 MEANS |
|---|---|---|
| MD3 / BO3 | Mean ~21 · Median ~19 | 28 is high — more like a long 3-set match, not the band average |
| MD5 / BO5 | Mean ~32 · Median ~30 | ~28 is a normal MD5 median even for closer gaps |
Keep one rule: always condition on format before comparing total-games numbers.
6. How to use this on a match page (without overclaiming)
On Tennis Glicko match pages, both players’ Glicko ratings are already visible. The gap is something you can read in seconds:
- Small gap: expect a more competitive shape on average — more room for a deciding set in BO3.
- Large gap: expect a shorter scoreline on average in BO3 — not a guarantee of a bagel set.
- Do not treat the table mean as an over/under tip. Residual noise is large (typical absolute error vs a fitted pre-match model is still ~5 games).
For win-probability edge versus the sharp market, use VOPO — a different question from match length.
7. Methodology notes
- Ratings are Glicko-2 overall (pure) as of match time, after chronological updates and weekly RD decay — not end-of-season ratings.
- Total games are parsed from completed set scores (tie-break games counted in the set totals as recorded).
- Walkovers / defaults and incomplete scores are excluded. Carpet/unknown surfaces are rare and omitted from surface slices elsewhere; tables here are format-based.
- Surface-specific Glicko gaps show a similar but weaker pattern than overall gaps in BO3.
8. FAQ
What is a Glicko rating gap in tennis?
The Glicko gap is the absolute difference between two players’ pre-match Glicko-2 overall ratings (|μ_A − μ_B|). A small gap means similar estimated strength; a large gap means a clearer favorite.
How many games does a BO3 (MD3) tennis match average?
Across ~136k best-of-3 matches in our sample, the overall mean is about 21.1 total games. When the Glicko gap is near zero, the mean rises to ~22.5; when the gap exceeds 500 points, it falls to ~16.1.
How many games does a BO5 (MD5) Grand Slam match average?
In ATP Grand Slam men’s singles (best-of-5, n≈1,420), the overall mean is about 30.4 total games, with a typical median near 28–30 across most rating-gap bands.
Does a 200-point Glicko difference mean 28 games?
Not in BO3. Around a 200-point gap in MD3, the mean is ~21 games and the median ~19. A median near 28 games is typical of MD5 (best-of-5) matches, not MD3.
What is the difference between mean and median total games?
The mean is the arithmetic average and is pulled up by very long matches. The median is the middle value: half of matches had fewer games, half had more. When mean ≈ median, the distribution is stable; when mean is higher, a long-match tail is present.
What does % of 4+ set matches mean?
In best-of-5, “% 4+ sets” is the share of matches that reached at least four sets (not finishing 3–0). It is a length/competitiveness indicator for MD5 only.
Next Topic
How Glicko-2 models skill uncertainty — and why that feeds every VOPO score.