Does Surface Really Matter for Tennis Upsets?
I’ve watched enough tennis to have a gut feeling: grass always seemed like the surface where anything could happen. Low bounces, fast points, less time to set up — surely that’s where lower-ranked players catch the favorites off guard more often than they do on clay or hard courts. It felt obvious. I’d never actually checked it, though, so I decided to test the theory using Jeff Sackmann’s open-source ATP match dataset, which covers tour-level matches from 1973 to 2022.
Cleaning Up 50 Years of Match Data
The first step was less glamorous than the analysis itself. The raw dataset had dozens of columns, but a lot of them — player heights, ages, detailed serve stats — were riddled with gaps, especially for older matches where that data was simply never recorded. I stripped it down to only what I actually needed: tournament info, surface, the two players, the score, and both players’ rankings. Any match missing a rank for either player got dropped, since you can’t call something an “upset” without knowing who was actually favored going in.
The Problem With a Simple Upset Definition
My first definition of an upset was simple: if the winner’s rank number was higher than the loser’s, the underdog won. But that definition bothered me almost immediately. Under that rule, the world No. 2 beating No. 1 counts exactly the same as a No. 150 player beating the world No. 1 — and those are obviously not the same kind of event.
So I introduced an upset score: the gap between the winner’s and loser’s rank, and only called it a real upset once that gap exceeded 10. That filtered out the marginal, barely-there reversals and left only the upsets that actually meant something.
Why Carpet Didn’t Make the Cut
Grouping by surface and year with this stricter definition, I plotted the upset rate over time. Carpet was the first thing I cut. It was phased out of the tour years ago, and by the late 2000s so few matches were being played on it that its upset percentages swung wildly from year to year, making it more noise than signal. Keeping it in would have muddied every chart and every test, so I removed it entirely and focused the analysis on the three surfaces that actually define modern professional tennis: Clay, Hard, and Grass.
Does Era Change the Picture?
With a cleaner picture emerging, I also wondered whether era mattered. Tennis in the 1990s looked very different from tennis post-2002 — playing styles shifted, hard courts came to dominate the calendar, and the game generally sped up. If surface was going to show any effect on upsets, it seemed worth checking whether that effect was consistent across eras or specific to one period. So I split the data at 2002 and ran the analysis separately for each half.
What the Numbers Actually Say
Fifty years of year-by-year upset rates, plotted by surface. The lines cross constantly, and no surface holds a consistent lead over the other two for any sustained stretch — that alone is a hint about where this is headed.
In the pre-2002 era, the numbers told a clear story: a p-value of 0.10 — above the standard 0.05 significance threshold — and a Cramér’s V of 0.0083, essentially zero. In plain English: surface had no statistically meaningful relationship with upsets in that period. The differences you see between Clay, Grass, and Hard are within the range you’d expect from random noise.
Post-2002, the picture shifts, just slightly. The p-value drops to 0.002, making the relationship statistically significant for the first time. But Cramér’s V only moves to 0.0143 — still firmly in negligible territory (as a rule of thumb, anything under 0.1 is considered a very weak association). In plain English: with well over 50,000 matches in the sample, even a tiny, practically meaningless difference between surfaces becomes “statistically significant.” A pattern emerges, but it’s barely there — and it’s not the kind of effect you’d want to bet on.
So, Does Surface Matter?
Here’s where the data actually leaves us. Before 2002, surface had no statistically significant relationship with upsets at all. Post-2002, a relationship does emerge, but a Cramér’s V of 0.0143 tells you it’s barely worth calling meaningful.
So across fifty years of ATP tennis, the most honest summary is this: surface doesn’t decide upsets. The post-2002 result is technically significant but practically negligible, and even that faint modern-era pattern is more a statistical curiosity than a real explanation. It turns out tennis upsets probably have more to do with the players on the day than the surface under their feet.
Methodology Note
Analysis based on Jeff Sackmann’s open-source ATP match dataset (1973–2022). Of 188,161 total match records, 47,822 matches were dropped due to missing player rankings, surface, or score. A further 15,237 Carpet matches were excluded, as Carpet was phased out of the tour and had insufficient data in later years. The final analysis used 125,102 matches across Clay, Hard, and Grass. An “upset” is defined as a win where the rank gap between winner and loser exceeds 10 places.
SpinInsight builds opponent scouting and match analysis for tour professionals and college programs. If your team wants this level of detail on your own opponents, book a 20-minute demo.