In June, Michael Hallsworth published an essay in his newsletter, The Judgment Gap, called “Clear and obvious errors.” Michael spent years at the Behavioral Insights Team and was a guest on my podcast, so I read most of what he writes, and I expected a smart piece about football. It is that. It is also the clearest description I have found of what goes wrong when we deploy precise technology into work that runs on judgment, and I have not stopped thinking about it since.
His subject is video review in football, the system that freezes a frame and draws lines across the pitch to rule a goal offside by the width of an armpit. His observation is that the technology did not end the arguments about offside; it shrank them and moved them. Officials used to argue about whether a player was level with the defender. Now they argue about which frame shows the instant the ball left the passer’s boot, and whether a shoulder counts as a scoring part of the body. Technology, he writes, seems to make errors fractal: the same mistakes, reproduced at ever-finer scales. The judgment is still there. It has been relocated somewhere harder to see, and the people who relocated it have declared victory.
Michael’s own conclusion is that the answer is not more precision but what he calls system stewardship: accept that judgment cannot be engineered out, decide in advance which calls belong to people, and build rules that can absorb the residual disagreement rather than pretending it is gone. He names artificial intelligence as the most urgent place to apply that thinking. I would narrow it further. The most urgent place is the clinical AI your health system bought last year.
Here is the fractal pattern, as I now see it, in three tools most large systems already run.
Ambient documentation removed the typing and kept the judgment about what the visit was about, but moved that judgment from the act of composing the note to the act of reviewing a draft. Those are different mental tasks. Composing forces you to decide what mattered because you cannot write the note without deciding. Reviewing asks only whether anything on the screen looks wrong, and a draft that has quietly left out the finding you would have led with does not look wrong. It looks finished. The decision did not disappear; it moved to a moment where it is easy to skip.
Predictive alerts did the same thing to the decision to treat, converting it into a smaller, later, more frequent decision about whether this particular alert deserves attention. And imaging AI turned the radiologist’s job from generating a read into second-guessing one that arrived first.
What happens to human performance on the relocated decision? The radiology evidence is uncomfortable. When German researchers gave 27 radiologists mammograms accompanied by purported AI suggestions, a dozen of which were deliberately wrong, even the most experienced readers, who were right 82% of the time with good suggestions, were right only 45% of the time with bad ones. Less experienced readers fell to 20%. Twenty years of practice were worth roughly 25 points of resistance, and the expert still lost the coin flip.
And the relocated skill does not hold its value. A study published last year followed endoscopists at four Polish centers after AI polyp detection was switched on, looking only at the colonoscopies they still did without it. Their unassisted adenoma detection rate dropped from 28% to 22% within three months. The reps went to the machine, and the skill followed the reps.
This is where Michael’s stewardship idea becomes a concrete agenda for a governance committee, and where I think most health systems are doing the fractal thing without noticing. They are measuring the machine with increasing precision and the relocated human decision not at all. The number everyone can quote is the override rate, which counts how often a clinician pushed back against the tool and is silent on whether the pushback was correct, how much time the clinician had, and how the patient fared. That is attendance, not oversight. Stewardship would mean sampling the overrides and the acceptances, tracing the outcomes, checking unassisted performance periodically, and writing down, before deployment, where a model is permitted to participate and where a human keeps the call by policy rather than by accident.
Football, to its credit, wrote that list before its first video review. For eight years it held at four kinds of decision. This summer it grew, which Michael would recognize as the pattern operating on the rulebook itself.
I have laid out the full argument on my site this week, with the evidence and a proposed replacement for the override-rate dashboard.
Also worth your time
Michael Hallsworth, “Clear and obvious errors,” The Judgment Gap, June 2026. Start here. The football is the hook; the argument about relocated judgment and system stewardship is the payload, and it transfers to any domain where a precise tool meets an imprecise task.
Goddard, Roudsari, and Wyatt, “Automation bias: a systematic review of frequency, effect mediators, and mitigators,” JAMIA 2012. The foundational review of why people over-trust automated advice. One finding worth pinning above every decision-support build desk: a recommendation biases the user more than the same information presented without one. Open access.
Jassim and colleagues, “Performance of artificial intelligence in breast cancer screening programmes: a systematic review,” BMJ Open 2025. Thirty-one studies and more than two million examinations. The sentence to read twice says triage works “when thresholds were conservatively calibrated.” Somebody picks the threshold. That somebody is the frame. Open access.


