Regression to the Mean: Why Praise and Criticism Both Look Like They Work
In the early 1970s, Daniel Kahneman gave a talk to a group of Israeli Air Force flight instructors on the value of positive reinforcement. One of the senior instructors interrupted him to disagree, and the disagreement is the part of the story everyone remembers.
The instructor said, more or less: "With respect, you have it exactly backwards. Whenever I praise a pilot for a clean manoeuvre, his next attempt is worse. Whenever I tear into a pilot for a sloppy one, his next attempt is better. Praise doesn't work. Punishment does." The room nodded. The instructors had decades of pattern-recognition behind them. They had watched it happen, repeatedly, with their own eyes.
Kahneman, in Thinking, Fast and Slow, calls this the moment that changed how he thought about feedback in general. Because the instructor was right about what he had observed. He was wrong about what was causing it. And what was causing it has nothing to do with pilots, training, praise, or criticism. It is a property of any noisy measurement of an underlying skill, and it appears in places you would never expect it to.
The mechanism
Suppose every pilot has a true underlying skill level. The skill is reasonably stable from day to day. What changes from flight to flight is everything else: the weather, the fatigue, the equipment, the luck. Each individual flight performance is the pilot's true skill plus a noise term. Over many flights, the noise averages out and you can see the skill. On any given flight, the noise can swamp the signal.
An instructor who praises a pilot for a great flight has, almost by definition, selected a flight where the noise was strongly positive. The pilot had a good day, the weather was kind, the manoeuvre clicked. The next flight, the noise term resets to its usual distribution, which is centred on zero. So the next flight is, on average, closer to the pilot's true skill. Closer to true skill is, in this case, lower than the praised flight. The pilot has not gotten worse. The instructor's intervention has not failed. The score has simply regressed to the mean.
Symmetrically, when an instructor criticises a pilot for a terrible flight, they have selected a flight where the noise was strongly negative. The next flight is again drawn from the usual distribution, and again pulls back toward true skill. From the seat of true skill, that's a higher score. From the seat of the criticised flight, that's an improvement. The pilot has not absorbed the lesson. The instructor's intervention has not worked. The score has just regressed.
The picture is the whole proof. The identity line is what you would see if every pilot's next flight were as good (or as bad) as their current one. The actual regression line is shallower, because the next flight is not just a function of skill, it is also a function of fresh noise, and fresh noise is centred on zero. The high-flight-1 pilots tend to drift down. The low-flight-1 pilots tend to drift up. Both at once. Praise and criticism both appear to work, because both target the extremes, and both extremes regress.
The only way to actually tell whether the praise or criticism is doing anything is to give the same feedback to a randomly selected group, watch their flights, and compare. The instructor who has been doing this for twenty years has never run that experiment, because every commander who has ever had to choose between intervening and not intervening has chosen intervening. There is no control group in real life. The mean does the regression invisibly, and the intervention takes the credit.
Your best discovery hole will disappoint you
If you have spent any time around step-out drilling, you have seen this pattern often enough to feel it in your stomach. The discovery hole comes in at twelve grams over six metres. The team is psyched. The next four step-outs come back at three, four, two, and five grams. The quiet sense is that something is wrong. The structure must be pinching. The mineralisation must be lensy. The team starts revising the model.
Sometimes the team is right and the deposit really is small or pinching. Often, though, this is regression to the mean dressed up in a hi-vis vest. The discovery hole was the highest grade you'd hit so far. Almost by definition, it's at the upper end of the noise distribution overlaid on whatever the underlying mineralisation looks like. The next holes are just samples from the same population, drawn afresh. They are, on average, going to be lower. Not because the deposit is smaller, but because the discovery hole was an extreme draw and the step-outs are average draws.
This is a hidden cost of how we run drill programs. A discovery hole that comes in spectacularly high creates an expectation, and the step-outs are then judged against that expectation rather than against the underlying grade distribution of the deposit. When the step-outs come in at the deposit's actual mean grade, which is what the math demands they do, they look like failures. Resource estimates get downgraded. Confidence in the model wavers. Sometimes drill programs get truncated entirely on the basis of step-outs that, statistically, were doing exactly what step-outs are supposed to do.
The opposite version is just as treacherous. A program that opens with a string of disappointing holes is sometimes shelved in favour of a "better" target, when the discovery program might just have been on the wrong side of regression. The next hole, drawn fresh, would have been closer to the average. We mistake the noise for the signal, and then we mistake our reaction to the noise for the cause of the next sample's behaviour.
The leadership version is the same problem
I want to spend the rest of this post on a version of regression to the mean that I have watched up close, because it is the version with high stakes and almost no visibility. It is the leadership version, and it shows up everywhere humans manage other humans with periodic feedback. I will lean on Navy examples because that is the world I know, but the structure is universal: it is the same in any organisation that uses periodic performance reviews.
Imagine a small department. Twelve people, varying skill, varying experience, varying days. Every week there's a deliverable, and every week the manager grades it. Some weeks one person turns in something genuinely excellent. Some weeks another one turns in something that has to be redone. The manager, like the air force instructor, gives praise to the high performer and a tough conversation to the low performer.
The next week, the praised person's deliverable is fine but unspectacular, and the reprimanded person's deliverable is fine but unspectacular. The manager concludes that the praise was a touch too generous (they coasted) and the criticism was bracing and necessary (they shaped up). The cycle repeats. Over time the manager builds a working model of leadership in which their interventions are the dominant cause of their team's variance. This model is mostly noise. Their team would have looked similar without any of the interventions, because regression doesn't care.
The actually useful version of leadership feedback works on a longer time horizon. It looks at trends, not at this week's deliverable. It distinguishes "this person consistently produces below their cohort over six months" from "this person had a bad Tuesday." It calibrates expectations to a person's true skill level, not to their last extreme observation. And critically, it asks what the world would look like if the manager said nothing at all. Some fraction of perceived improvements would happen anyway. Some fraction of perceived regressions would happen anyway. Subtract those off before claiming credit for either.
This is the source, by the way, of one of the most damaging dynamics I've watched: the manager who genuinely believes their constant interventions are the reason the team performs. The team would likely have performed much the same without the interventions, because the team's true skill is doing most of the work and the interventions are a noisy add-on. But the manager has spent years calibrating their internal model on regression to the mean, and they've concluded that without them the wheels would fall off. They are not lying. They have just been running an unblinded experiment on themselves for their entire career.
The honest version of feedback
None of this is an argument against feedback. Feedback genuinely works. There is a real signal underneath the noise. Pilots really do learn. Sailors really do improve. Step-out programs really do refine the resource model. The point is not that nothing helps; the point is that we routinely overestimate how much our individual interventions help, because we are reacting to noise and crediting the response.
A few habits that defuse this:
Define the baseline before the extreme. If you only ever judge an outcome by comparing it to the previous extreme, you'll always be impressed or disappointed by regression. Anchor on a longer baseline: this person's six-month average, this hole's expected grade given the resource block, this team's typical sprint completion. Compare the new observation to the baseline, not to the previous outlier.
Run the counterfactual in your head. Before claiming credit for an intervention, ask what would have happened if you'd done nothing. If the answer is "this would have improved/declined anyway because it was an extreme observation," that is regression doing the work, not you.
Build in actual control. The cleanest defence is to occasionally do nothing. Pick a subset of pilots, sailors, programs, or holes where you deliberately skip the intervention. Watch what happens. The first time you see an extreme low-performance situation correct itself in the absence of any feedback, the lesson lands hard. (This is rarely politically possible. When it is, do it.)
Distinguish the trend from the event. Six bad weeks in a row from one person is a signal worth acting on. One bad week is regression bait. The boundary is fuzzy and depends on the noise level of the work, but the question to ask is "is this observation a trend or a draw?" If you can't tell, treat it as a draw and wait. The next observation will be closer to the trend, regardless of what you do.
Notice your own asymmetry. Most managers have stronger feelings about extreme bad performance than extreme good performance. The reprimand is delivered, the praise often isn't. The result is that managers experience the regression of bad performances strongly (it feels like the reprimand worked) and experience the regression of good performances weakly (the praised person's next deliverable just gets noted as "fine"). The asymmetric record-keeping calcifies the belief that criticism works and praise doesn't, even when, statistically, both are doing the same nothing.
The deepest version
The piece of this that most surprised me when I first thought about it carefully is that regression to the mean is not really a fact about psychology, or training, or noise. It is a fact about how we select cases for attention. Any time you pick out an extreme observation and watch what happens next, you will see regression. It happens whether or not you intervene, whether or not the intervention is any good, and whether the underlying process is people, drillholes, sales numbers, or athletic performance.
This means the appearance of "feedback working" is a structural feature of how we sample reality, not a fact about feedback. When we praise the great hole and the next hole is worse, the praise didn't cause it. When we reprimand the poor performer and they improve, the reprimand didn't cause it. When we tweak our drilling parameters after a string of bad assays and the next assay is better, the tweak didn't cause it. We need to do controlled work to know whether any of those interventions are any good. Otherwise we are reading noise as causation.
The flight instructor was watching regression to the mean. So was every commander I have served under who was sure that their interventions were the reason morale held. So is every drill geologist who is sure that the model revision after the disappointing step-outs was the reason the next program landed. So am I, when I notice that the post I worked the hardest on has fewer reads than the post I dashed off, and conclude that effort doesn't pay. We are all calibrating on noise. The mean is the mean. Things drift toward it. The feedback we give in between is often storytelling we apply to a process that was always going to do the regression on its own.
This post is part of a series on statistical pitfalls in geoscience, drawn from a talk given at GACMAC 2026.