Blog

The Paradox of Becoming Better at Surgery

The Paradox of Becoming Better at Surgery cover illustration

Why a surgeon's numbers get worse as the surgeon gets better

1. The Worst Injuries

In June 2026, the Journal of the American College of Surgeons published a comparison of Level I and Level II trauma centers in Pennsylvania, covering twenty-five years and more than 360,000 injured patients. [1] Level I is the top tier: the center with the most subspecialists on hand, and the one other hospitals send their worst injuries to. On the raw numbers, more of its patients died, 8.2 percent against 6.9 at Level II.

Is it better, then, to be taken to a Level II center? No. The Level I centers were sent the worst injuries: more transfers from other hospitals, more patients in shock, more people hurt in several parts of the body at once. Once the comparison accounted for how badly each patient was hurt and the condition they arrived in, the study found no difference in overall mortality. For two of the hardest groups, patients who arrived in shock and those with blunt injuries to several body systems, the Level I centers did better.

Statisticians have a name for the extreme version of this: Simpson's paradox. One side does better within every group of similar patients, yet worse once all the patients are added together, because the two sides were not given the same mix of cases. Basketball shows the problem more simply. Take two New York Knicks. In the 2025–26 regular season, Mitchell Robinson made 72 percent of his shots and Jalen Brunson made 47. [2] Robinson is a center who scores on offensive rebounds and dunks: more than nine in ten of his shots came within three feet of the rim, and half of all his shots were dunks. He is very good at that job. Brunson takes the threes, the contested midrange jumpers, and the last shot of a close game, and he did not dunk once all season. The season percentage pools every shot into one number, and the one number mostly measures where the shots came from.

Brunson is the Level I trauma center: he takes the threes and the contested midrange shot at the end of the game. Robinson is the Level II trauma center, good at his work, with more of the rebounds and dunks and fewer of the last shots. Pooled, the trauma center built for the worst injuries looks worse; adjusted for who came through the door, it does as well or better. I read that as the most accurate short description of a surgical career I know.

The better you are, the harder the cases you are sent, and the worse your raw outcomes look. That half of the paradox is arithmetic. The other half is not. A good surgeon does not rest on his laurels, so he stays worried about his outcomes. He measures himself only against himself and is never satisfied with good. The closer he comes to mastering the craft, the further he is from security. The better he becomes, the less satisfied he is with his own skill.

Both halves are true, and they are linked: the harder cases and the worse raw numbers are part of the reason a good surgeon never feels done.

2. The Referral Drift

Early in a career the numbers are generous. The cases are selected, a senior surgeon still stands at your shoulder, and the anatomy matches the atlas. The results are good because training was built to produce exactly this work.

Then the referrals change. A reconstructive practice does not fill with first operations in healthy tissue. It fills with the case that already failed somewhere else: the ureter that has been stented for a year, the urethra that has been cut and dilated until there is little left to sew, the fistula after radiation, the repair of someone else's repair, the patient who has been told nothing more can be done. These cases do not arrive because the surgeon has grown careless. They arrive because he has become the surgeon other surgeons call.

Skill is not the only thing that brings the hard cases. A tertiary center fills with them whoever is operating in it, and seniority brings referrals whether or not skill came with it. It also seems to me that the system sends its hardest patients toward the hands it trusts most, which is what it should do. Whatever the reason, the hard cases are not shared evenly. In one study of 23 heart surgeons in north-west England, all doing first-time bypass operations between 1999 and 2002, high-risk patients made up about one in four of one surgeon's operations and about one in twenty of another's. [3] The authors warned that this makes raw death rates misleading, and that judging surgeons by raw rates can push them to turn high-risk patients away.

Hands that take sicker patients carry worse raw numbers than their skill deserves. Brunson is the player the Knicks give the ball to when the shot clock is running out and the play has broken down. By the raw numbers he is an ordinary shooter. Even counting his threes at their full value, his shooting efficiency over the season was a little below the league average, 53 percent against 55. [2] Now compare only the late-clock shots. In the last seven seconds of the shot clock, the league's efficiency fell to 47 percent. Brunson's stayed at 53. [4] The Knicks kept giving him those shots, and in June he was the Finals MVP when they won their first championship since 1973. [5] Risk adjustment is the fairer box score. It compares each shot with what an average shooter makes from the same spot at the same moment. In the ordinary box score, those hard misses drag his shooting percentage down. In such a comparison they cost him little, and the hard makes count for a great deal.

The pooled numbers count a leak in a field that has been operated on, irradiated, and infected the same way they count a leak in healthy tissue. Risk adjustment narrows that gap and does not close it, because the chart records little of what makes a reoperative pelvis hard.

So the trajectory runs backward. Competence produces reputation. Reputation produces harder work. Harder work produces numbers that look worse than the numbers from five years earlier, when the surgeon was less useful. A clean record can mean mastery. It can also mean a practice that has stopped accepting the patients who most need a surgeon.

I see this in my own schedule. The operations I was proudest of ten years ago are now the operations I would hand to a junior partner. What remains on my list are the problems left over after the standard answer has been tried. If this year's pooled numbers look worse than the numbers from the year I learned an operation, the likelier reason is that the operation worked well enough for harder problems to replace the easy ones. The pooled table cannot see that substitution. I believe I can, which is not the same thing.

"My cases are harder" is the standard defense of every surgeon with bad numbers, and it is not always true, including when I say it. No one else knows a surgeon's record in as much detail, so he has to check the claim against it, one kind of problem at a time, and accept the answer when it goes against him. He cannot use "my cases are harder" as an alibi.

Knowing the arithmetic helps less than it should. I can explain the substitution to anyone who asks, and a worse number still reads to me as a question about my hands. I think that is the right reaction. A worse number should send a surgeon back to his own work first, even when the answer turns out to be the harder cases.

3. No Laurels

The first half of the paradox is statistical. The second is temperamental, and it is harder to write about, because it sounds like a virtue and does not feel like one.

A surgeon who is merely adequate can treat a good year as proof: a high shooting percentage built on open shots he picked for himself. A good surgeon cannot. Last year's results were earned on last year's cases, and next year's list is harder. A championship does not get the champion the same shots next season; the teams he beat spend the summer on his film. He also knows how much of a good result was the case itself, the team, the tissue, and the patients who never came back to report a failure.

So he stays worried. He reviews the anastomosis that is patent and still asks whether he should have done the anti-refluxing kind. He sees the patient who is doing well at six months and wonders what the imaging would show at two years. He does not mistake the absence of a phone call for a durable result.

Worry in this sense is not anxiety, and the difference can be seen from outside. Worry ends in a changed plan: a different graft, earlier imaging, a longer conversation before consent. Anxiety ends where it started. Laurels are for people whose work is finished, and his is not. The next referral is already on the schedule, and it is harder than the last.

4. The Private Scoreboard

The public numbers compare surgeons with each other. That comparison is contaminated by construction: different case mix, different hospitals, different willingness to operate on the patient everyone else declined, different honesty about what counts as a complication. A good surgeon eventually stops using other surgeons as her reference class. Not because they have nothing to teach her, but because their numbers are not her numbers.

Surgery lacks one thing basketball has. Anyone watching the Knicks, at Madison Square Garden or on television, can see which shots Brunson takes and judge whether he is good. Nobody watches a surgeon's game. The referring doctor, the hospital, and the patient choosing a surgeon cannot see the shots, only the shooting percentage. The one person who saw every shot is the surgeon.

She measures against herself.

The remedy for Simpson's paradox is to stop pooling and compare like with like, and a surgeon's own record is the only one sorted by the problems she actually takes. It is her shot chart, split by where each shot came from and how much clock was left. It is not a clean table. One surgeon does not do enough of any single kind of operation for a rate to mean much. If she does a particular redo operation eight times a year, one bad case moves her rate by more than twelve points, the way a few misses swing a shooting percentage built on a handful of attempts. The cases within each kind of operation also get harder from year to year, so even her own record has a Simpson problem over time. The scoreboard I keep is mostly a record of decisions, not rates. Did I recognize poor tissue quality sooner than I would have three years ago? Did I change the plan in time, or did I continue an operation that was doomed to fail? Did the patient leave with a result I would accept for someone I love? A single case can answer a question about a decision. It cannot answer a question about a rate, which is the point of The Red Beads. I argued in The Wicked Problem of Surgical Failure for keeping this count. That argument left out what the count does to the person who keeps it.

On the scoreboard she keeps for herself, what once counted as good is now the minimum. The only result that satisfies her is one that beats the last honest version of her own work, and because that standard only rises, there is never a point where she can stop. Every improvement resets the baseline. The surgeon who was once pleased by patency now wants a shorter stent dwell, an earlier discharge, a patient who has stopped thinking about the operation by the end of the year. The target recedes because she keeps moving it.

This is how the craft improves. It is also how satisfaction is withheld from the people most responsible for the improvement. I cannot say which causes which. Dissatisfaction may produce the improvement as much as the improvement produces the dissatisfaction, and in the careers I know it is probably both. Either way, the better you are, the less satisfied you are with your skill, and the dissatisfaction is not a symptom of failure. It is the instrument working.

5. Further from Security

People outside the field imagine that seniority buys calm: more cases, more pattern recognition, more right to trust your hands. Some of that is true. The thousandth reconstruction is not the tenth.

The first rung of seniority can work the other way. Chris Whitty, England's chief medical officer, has been writing and speaking this year about what he calls risk holding. [6] Over three decades, training has moved decisions that carry risk later and later, away from residents. Then, on the first day as a consultant (an attending, in American terms), all of the risk is handed over at once. That is profoundly uncomfortable, and probably unsafe; new consultants have told Whitty as much. [7, 35:24] Holding risk is a skill in the same way operating is, and it is taught far less systematically.

This comes from British training, but I think the same applies worldwide. Overnight, the new attending holds all of the risk, with the same skill she had the week before. That is one reason seniority does not buy the calm people expect.

Brunson made the same step in 2022, when he left Dallas for New York. His father has said the point was for him to run a team. [5] What changed was how much of the team's result rode on it. [2]

What seniority also buys is a clearer view of how much can still go wrong after a technically sound operation. The young surgeon fears the error she can name. The older surgeon fears the failures she now knows exist and cannot see coming: the air embolism in the middle of a routine step, the radiation necrosis that surfaces months after a repair that looked healed, the recurrence that arrives after a year of believing the problem was solved. Experience lengthens that list.

She also loses her excuse. When a case fails early in a career, part of the explanation is the learning curve, and everyone in the room grants it. New platforms still send senior surgeons back to the start of a curve, as I wrote in The Endless Residency. But when an operation she has done for years, on a platform she knows, fails late in her career, the explanation everyone reaches for is the surgeon, and she reaches for it first. The same failure costs more the better she is, because nobody, including her, believes it was inexperience.

So the paradox closes on itself. The skill that draws the hard cases concentrates the risk, and the self-comparison that improves the work withholds the satisfaction. The closer a surgeon comes to mastering the craft, the further she is from security in it, and the surgeon who has done the most to earn security is the one least able to feel it.

This has a cost, and a wellness slogan does not cover it: a standard that never settles can exhaust the surgeon who holds it. One tempting response is worse: the surgeon protects her numbers by protecting her schedule, measures herself against colleagues with easier lists, and calls the resulting calm mastery.

On a basketball team, that surgeon would be the player who passes up the last shot to protect his percentage, and this is where the analogy breaks down. A player who does that does it in front of a full arena. A surgeon chooses her shots in private. She can decline the referral or send the redo case elsewhere, and nobody sees the shot she did not take. Surgeons know this happens. In a 2000 survey of cardiac surgeons in the United Kingdom, 94 percent of those who answered agreed that high-risk patients were being turned down for surgery. [3] The payment system does not see it either. Insurers pay mostly by the name of the operation, not by how hard it was in that patient, so the surgeon who keeps the straightforward cases is paid more per minute than the one who takes the reoperations. A billing modifier for unusual difficulty exists, but in my experience the extra payment is usually not commensurate with the extra difficulty, when it is paid at all. Who wants to argue with the referee that the basket I just scored should be worth more points? The insurer cannot tell which of the two is the better surgeon, and it rewards the easier list.

Patients do not need the surgeon who protects her numbers. The badly injured patient sent to the Level I trauma center needs her least of all.

6. Coda: The Split Table

Go back to the trauma centers. The raw death rates and the adjusted ones came from the same patients. No new data was needed to get from one to the other, only the question of which center had been sent which injuries. The raw rate was correct arithmetic. It was also the wrong answer for anyone who stopped there.

Every surgical career runs on a pooled table: the department dashboard, the credentialing file, the referring doctor's memory of the last case that went badly. The split table exists in one place, the surgeon's own record, sorted by the problems she chose to take. It is imperfect, and it is still the record in which her improvement is easiest to see. It will never tell her she is finished.

The same arithmetic runs wherever a workload is pooled. It punishes whoever takes the hardest share and flatters whoever quietly takes the easiest. It is why the Finals MVP's season numbers looked ordinary.

That is the trade. The public numbers may get worse as the work gets better. The private numbers will keep raising the standard until satisfaction is always one case away. Neither is a reason to become less good. Both are reasons to stop asking the numbers for peace.

Peace was never what they were for.

Notes

  1. McLaughlin CJ, Song J, Kern JA, Seamon MJ, Martin ND, Chreiman K, Kim PK, Reilly PM, Kaufman EJ. Trauma Center Level and Mortality in Injured Patients with Shock or Multisystem Trauma. J Am Coll Surg. Published online 22 June 2026. PMID 42319113. doi:10.1097/XCS.0000000000002078.
  2. Basketball-Reference.com. 2025–26 regular season. Mitchell Robinson: field goal percentage .723; 92.2% of attempts within 3 feet of the rim; dunks 51.5% of attempts. Jalen Brunson: field goal percentage .467; effective field goal percentage .533; 11.6% of attempts within 3 feet; no dunks; usage rate by season. 2025–26 NBA season summary: league-average effective field goal percentage .546, computed from the league per-game averages (42.0 field goals and 13.3 three-pointers made on 89.1 attempts). Robinson; Brunson; league. Accessed 27 September 2026.
  3. Bridgewater B, Grayson AD, Jackson M, Brooks N, Grotte GJ, Keenan DJ, Millner R, Fabri BM, Jones M; North West Quality Improvement Programme in Cardiac Interventions. Surgeon specific mortality in adult cardiac surgery: comparison between crude and risk stratified data. BMJ. 2003;327(7405):13-17. PMID 12842949. doi:10.1136/bmj.327.7405.13. The 2000 survey of UK cardiac surgeons is reported in its Discussion, citing Keogh B, Kinsman R, National Adult Cardiac Surgical Database Report 2000–2001, Society of Cardiothoracic Surgeons of Great Britain and Ireland, 2002.
  4. Schuhmann J. 4 takeaways: Jalen Brunson stars, Hawks struggle shooting from 3. NBA.com, 29 April 2026. Regular-season effective field goal percentage in the last seven seconds of the shot clock: league 47.1%, Brunson 53.2%, second in the league with 157 baskets. nba.com. Accessed 27 September 2026.
  5. Powell S. Jalen Brunson wins Bill Russell trophy as 2026 NBA Finals MVP. NBA.com, updated 14 June 2026. nba.com. Accessed 27 September 2026.
  6. Whitty CJ, Dickson J, Abraham AR, Mitchell T, Patel M, Brown VT, MacEwen C. We must be more systematic about risk holding throughout doctors' careers. BMJ. 2026;392:s509. PMID 41844254. doi:10.1136/bmj.s509.
  7. Whitty C. Risk, uncertainty, and medical decision-making. Inaugural BMJ lecture, with the Academy of Medical Sciences. The BMJ podcast, 18 September 2026 (44 min). Podcast; video. Timestamps from the video's English auto-generated captions. Accessed 27 September 2026.