Numbers · 3 min read
Regression to the Mean: Why Extreme Results Don't Last
Why the worst performers improve and the best slip back, with no cause at all. Regression to the mean, and how it fools us about praise, cures and luck.
By the Cognosc team ·
The pupils who score worst on a test get extra tutoring, and next time they do better. Did the tutoring work?
Maybe. But they’d probably have done better anyway. This is regression to the mean, and it’s behind a surprising number of things people believe.
Try it
The worst scorers improve by themselves
Each dot is a pupil: first test across, second test up. Nobody got any help between them. Pick out the lowest scorers on the first test.
The bottom 15% averaged 42 on the first test and 48 on the second: up 6 points with no help at all. Part of a very low score is bad luck, and luck doesn’t repeat.
The idea
Most results are part skill and part luck. A very low score usually means someone had both lower skill and a bad day. Next time, their skill is the same but their luck is probably more ordinary, so their score moves back towards the average: it regresses to the mean.
The same happens at the top. The best scorers on one test usually include people who had a lucky day, and many of them slip back a little next time.
No cause is needed. It happens whenever you pick people, or anything else, because they were extreme, and then measure them again.
A pilot’s story
The psychologist Daniel Kahneman described teaching flight instructors in the Israeli Air Force that praise works better than criticism. One instructor objected: when he praised a cadet for a great landing, the next one was usually worse; when he shouted at a cadet for a terrible one, the next was usually better. So shouting works.
Kahneman realised the instructor was seeing regression to the mean. An exceptional landing is likely to be followed by a more ordinary one, whatever the instructor says. So is a terrible one. Praise was being blamed, and criticism credited, for what luck would have done anyway.
Where it fools us
- Treatments. People tend to seek help when a problem is at its worst, like back pain, a cold or a slump in mood. Many would improve anyway, so whatever they tried next looks like it worked. That’s one reason medical trials need a comparison group.
- Sport. The “Sports Illustrated cover jinx”: athletes appear on the cover after an exceptional season, which is often followed by a more ordinary one.
- Policy. Speed cameras placed where there were a lot of crashes last year will see fewer next year, partly because last year was unusually bad.
- Business. The worst-performing branch gets a new manager and improves; the best one slips back. Both changes may be mostly luck.
How to avoid being fooled
- Ask how the group was chosen. If it was chosen for being extreme, expect it to move back towards average.
- Compare with a similar group that got nothing. If the untreated worst scorers improve just as much, the treatment didn’t do much.
- Look at more than two measurements. A real effect should persist; regression is a one-off bounce.
Test yourself
The tutoring question is in the free statistics test. The thinking traps test covers the related trap of assuming that what came first caused what came next.