top of page

The Red Beads

Why surgical quality has to measure the bin, not just the bead



1. The Conference


Every surgeon knows the room. The lights are down for the imaging. A resident presents the case: a man in his sixties, an operation that went as planned, a recovery that did not. Days later he was dead.

The questions start, the same ones every time. Was the case indicated? The right operation? Did the imaging come soon enough, the complication get caught in time? Each steps a rung up the timeline toward the same charge: somewhere the surgeon chose, and had he chosen otherwise, the patient would have lived.

The conference has learned to talk about systems, root cause, blameless review. It means it. Underneath the vocabulary the room still has an old habit: it works backward from a dead patient to a decision, and does not like to adjourn without one.

Often there is a decision, and finding it is the point. The trouble is the cases where the room finds one anyway: a bad outcome is not always the residue of a bad choice, and a review that needs an answer will find one.

A bad outcome is evidence. It is not a verdict. The conference fails when it forgets the difference.


2. The Red Beads


W. Edwards Deming spent his career on why factories produce defects, first in wartime American industry, then in postwar Japan. He learned that managers constantly mistake luck for skill. A worker had a good month and was praised; another had a bad month and was put on a plan. Next quarter they traded places, and no one asked why: everyone believed the numbers measured the people who produced them.

He had argued this for years to executives who nodded and went home and ranked their people anyway, so he staged a demonstration. [1]

He pulled volunteers from the audience, the Willing Workers, and gave them a job: produce white beads from a bin in which one bead in five was red. The tool was a paddle, a flat scoop with fifty shallow holes. Dipped into the bin, it came up with a bead seated in every hole, fifty at a time, a blind sample. An inspector counted the reds in each paddle. Deming played the boss. He set a standard well below what the bin could give: a fair draw of fifty returns about ten red, and the standard sat under that. He posted the counts, praised the low scorers, warned the high ones, and fired the worst. The rate never moved. The fraction of red was fixed, and nothing a worker did with the paddle could touch it.

That was the proof. The rankings had been rankings of luck, every reward and firing handed out for a result the worker could not control.

Every surgical complication arrives as a bead: the leak, the stroke, the death, concrete and dated and tied to a face and a surgeon. The bin is what we cannot see, the operation and the population and the case mix and the baseline risk, the whole distribution the outcome was drawn from. We are built to see the bead. Quality systems exist because the bin has to be constructed.


3. When One Bead Is Enough


Surgery did not get its quality tools from Deming. The morbidity and mortality conference predates his movement, and the checklist owes more to aviation and the World Health Organization than to any factory. [2] But we talk about complications the way engineers talk about defects. That is why the beads are worth borrowing, and worth getting right.

Start with what a single outcome can tell you, because it is not nothing. Some events are self-identifying failures. A few are sentinel or serious reportable events: wrong-site surgery, retained foreign object. Others are plainer: a critical step omitted, an injury made and not recognized, a deterioration answered too late. These need no denominator. One is one too many, and the conference exists to find them and prevent the next.

A known-risk complication is different. A leak at a defensible rate despite careful surgery is not automatically a failure to find. It is a draw, though rarely a clean one: many events inside defensible rates still carry a correctable contributor, and separating those from irreducible risk is part of the room's work. The error the conference is prone to, the one Deming spent his life on, runs the other way: treating every draw as an omitted step because the outcome was bad and the room needs a cause.


4. The Bead and the Bin


Quality engineers separate two kinds of variation. One is the ordinary scatter a stable process throws off. The other is a specific, findable disturbance: a broken machine, a bad lot, a step done wrong. The red beads are the first. The wrong-site operation is the second, the Swiss cheese model, holes in successive defenses lining up. [10] The known-risk complication is the case that picture misses. No hole lined up. Every defense held. The harm came from the bin.

The usual argument stops at whether the surgeon's hand was at fault. That is the wrong place to stop. Even an honest reviewer cannot read the bin from one bead. The distribution is not in the room.

A surgeon will object that he is not a factory bead-drawer, and he is right. He shapes his bin through selection, exposure, technique, and the decision whether to operate. But shaping the distribution is not the same as reading it from a single draw. Skill moves the bin and narrows it, and two surgeons can still draw the same red from very different bins. Because the surgeon builds his own bin, one bad draw is not silent. It moves the estimate, which a bead never does for the willing worker who cannot touch the jar. But it moves it only a little, and noisily. A single outcome informs without convicting.

That is the maddening part. The work arrives one case at a time. The answer to whether you are any good does not.

Learning curve makes this harder. The novice does not draw from the master's bin, nor the surgeon early in a new operation from the bin he will have in five years. Inexperience adds red beads; so do innovation, a new team, a new device, a new referral pattern. One outcome cannot distinguish ordinary risk from early learning, poor selection, or thin supervision. Review can raise a hypothesis and buy an interim safeguard; separating the causes takes counting the bin over time. The count is slow because the feedback is wicked: delayed by years, censored by lost follow-up, biased toward the patients who come back. [9] Learning curve is not an excuse. It is a reason to count.

To make one outcome a verdict on the surgeon you need what it does not carry: the denominator, the case mix, the baseline risk, the learning curve. Comparing surgeons on raw mortality ignores case mix; credible profiling adjusts for risk before it ranks. [3] One bad result is where the analysis starts, not where it ends.


5. My Own Bead


Here is my own bead.

I performed a robotic ileal conduit. Nothing about the case fought me: trivial blood loss, a routine bowel anastomosis, a stapler I have fired hundreds of times. The patient recovered on schedule. Days later the anastomosis leaked. I took him back to the operating room, and he recovered.

I did what surgeons do. What did I do wrong? I went back to the video, looking for what I had missed: tension on the mesentery, a poorly seated staple line, bowel caught where it should not have been. This is the honest instinct, and the narrow one: the belief that a bad outcome must hold something visible if you look hard enough.

I could not find a technical failure. A video can show an error. It cannot prove the absence of one. The staple line gave way where staple lines are not supposed to, in a construction I had built as I always build it.

The video could tell me whether I had missed something visible, and finding nothing did not let me off. It sent me to a question it could not answer: what my rate in this operation is, and whether it had moved. That answer does not live on one recording. It lives in the count, slower and less consoling than an evening re-watching a single case. A red bead is not a verdict. It is a summons to read your own bin, and a bin takes more than one case, more than memory, and more than guilt.


6. The Hunt for a Cause


Going back to the video is right, and I would do it again. When there is a special cause, review is how you find it. The danger is that the machinery runs whether or not there is anything to find; two habits of mind ensure it usually does.

The first is outcome bias: we judge a decision by its result, and rate the same reasoning worse when the outcome was bad. [4] The second is hindsight: once we know what happened, it looks like we should have seen it coming. [5] Poker has a name for the first, resulting, grading the decision by the card that turned. [6] Underneath both is something simpler: the bead is concrete, the bin abstract, and the mind reaches for what it can see. Put a dead patient and a timeline in a darkened room and the reconstruction will find a decision to condemn, because every step now points backward at the death.

M&M is the institution we built to find special causes, and it is good at that when they exist. Its tragedy is that the same machinery can keep running when no special cause is there, until the need for explanation attaches itself to a name, and the name is a surgeon.


7. The Honest Count


Every quality system depends on someone counting red beads honestly, and the counting happens in the room that has just shown it will pin each one on a person.

Punitive counting produces bad counts. Where a red bead can end your week, the incentive is not to draw fewer, which no one can do, but to record fewer. The complication becomes expected. The death becomes the death of a very sick patient. The leak becomes a known risk of the operation. Sometimes that is the truer classification. Sometimes it is self-protection. A punitive room is poorly built to tell the difference.

The deeper failure is quieter. A crude quality system does not eliminate red beads. It moves them, away from surgeons who take hard referrals and toward wherever the counting is thinner. It teaches the surgeon that the safest way to look excellent is to keep the hardest patients out of the bin: the hostile abdomen, the radiated pelvis, the salvage reconstruction, the frail patient. Those bins already hold more red, and most need someone willing to operate. Punish the surgeon who takes the hardest cases and you have not reduced harm. You have moved it where no one counts.

The bead hides an asymmetry. A surgeon can read his own bin. The patient and the referring doctor cannot; they see the one bead, the case that leaked, and nothing of the distribution behind it. That is the sharpest reason the bin has to be built and shown, not merely known: the people who carry the outcome cannot see the odds.

None of this argues against measurement. It argues for the kind that can see a bin. Some outcomes are true failures and demand direct review. So does a bead that breaks the pattern, or a run of red where the bin was white. The discipline is telling signal from noise: the bead that marks a changed bin from the bead the bin was always going to give.

But one case is too small to read on its own. The shift it gives is real and far too noisy to trust alone; the rate becomes legible only in the bin, counted across many cases and adjusted for risk. [8] The cultures that run hazardous work well protect honest reporting and separate the culpable act from ordinary variation, so a surgeon can report a bad result without being punished for the distribution. [7] The count is the point. The trial is not.


8. Coda: The Bead Generalized


Go back to the darkened room. The bead is on the table: a leak, a stroke, a death. It deserves review and honesty. Sometimes it will reveal an error, a system failure, a learning curve that needs structure. And sometimes it will reveal only what surgery has always carried: risk that survives even careful work.

The mistake is believing the bead can tell us which. It cannot. The bead is visible. The bin has to be built.

Surgery is not unusual in this. The investor is judged on one exit, the coach on one season, the parent on one hard year: a single draw standing in for a distribution no one can see. Surgery only plays for higher stakes, and lays its draws on a table in a darkened room.

A bad outcome proves harm. Some prove the failure outright. Whether the rest prove bad surgery depends on what the bead cannot show.


Notes


  1. Red Bead Experiment and Deming's work: The W. Edwards Deming Institute, Red Bead Experiment and About Dr. Deming. The one-in-five red, fifty-per-paddle, roughly ten expected figures are the standard setup, used here to make the arithmetic explicit; the Institute page prints no bin total or numeric quota, so the work standard is described, not quoted.

  2. Surgical Safety Checklist: World Health Organization, Safe Surgery tools and resources. The aviation lineage: Clay-Williams R, Colligan L. Back to basics: checklists in aviation and healthcare. BMJ Qual Saf. 2015. doi:10.1136/bmjqs-2015-003957.

  3. Provider profiling, public comparison, and the statistical risks of raw rankings: Normand SLT, Shahian DM. Statistical and clinical aspects of hospital outcomes profiling. arXiv:0710.4622.

  4. Outcome bias: Baron J, Hershey JC. Outcome bias in decision evaluation. J Pers Soc Psychol. 1988. doi:10.1037/0022-3514.54.4.569.

  5. Hindsight bias: Fischhoff B. Hindsight is not equal to foresight: The effect of outcome knowledge on judgment under uncertainty. J Exp Psychol Hum Percept Perform. 1975. doi:10.1037/0096-1523.1.3.288.

  6. Resulting: Duke A. Thinking in Bets. 2018. Publisher page.

  7. High reliability and just culture (preoccupation with failure, reluctance to simplify, sensitivity to operations, deference to expertise, commitment to resilience; just culture separates the culpable act from honest reporting): AHRQ PSNet, High Reliability and Culture of Safety.

  8. Risk-adjusted, case-mix adjusted, 30-day surgical outcomes and registry-based quality measurement: American College of Surgeons, National Surgical Quality Improvement Program.

  9. Surgery as a wicked learning environment, where feedback is delayed, noisy, and biased: Zhao L, The Wicked Problem of Surgical Failure. The kind/wicked distinction is Robin Hogarth's.

  10. The Swiss cheese model of organizational accidents: Reason J. Human error: models and management. BMJ. 2000. doi:10.1136/bmj.320.7237.768.

Comments

Couldn’t Load Comments
It looks like there was a technical problem. Try reconnecting or refreshing the page.

Lee C. Zhao MD

This is an independent educational website, not affiliated with NYU Langone Health or New York University. Content is for informational purposes only and does not provide medical advice, diagnosis, or treatment. Viewing this site does not establish a physician-patient relationship.

222 E 41st St 11th Floor, New York, NY 10017

  • Facebook
  • Instagram
  • X
  • TikTok

 

© 2026 by Lee Zhao.

bottom of page