Which of the Six Big Losses Should You Fix First?
Not the Biggest One
The Six Big Losses give you a complete inventory of where production goes. What they do not give you is an order — and ranking them by minutes lost, which is what almost everyone does, is reliably wrong.
The inventory is the easy part
The six categories come out of TPM and they are genuinely comprehensive: equipment failure, setup and adjustment, idling and minor stops, reduced speed, process defects, and reduced yield at startup. Mapped onto OEE, the first two land in Availability, the middle two in Performance, and the last two in Quality.
Most plants can produce that breakdown from their event data in an afternoon. Then the improvement meeting starts, someone points at the largest bar, and the budget follows it. That step is where the money gets misdirected.
Why the biggest loss is usually the wrong target
Three reasons, and they compound.
Losses are not independent
A loss inventory adds minutes as though each stoppage cost the same. It does not. Ten minutes of downtime on a machine feeding a full buffer costs nothing at all; the same ten minutes on a starved constraint costs ten minutes of finished product. The report records both identically, so the ranking it implies is a ranking of occurrence, not of cost.
Position matters more than magnitude
Where a loss sits relative to the constraint decides whether it matters. A large loss on a machine with slack either side may be almost free; a small, frequent loss on the constraint is paid in full every time. Two plants with identical loss tables and different topologies have different correct answers.
Averaging erases the difference
This one has been observed in print. Fischel and Lange, modelling a food plant with over twenty unit operations and up to twenty failure modes each, noted that had they taken the conventional averaged, single-failure-mode-per-operation approach, "the gains and losses would be less accurate and essentially identical to each other."
That is the whole problem in one sentence. Coarse loss accounting does not merely blur the ranking — it flattens it, until every candidate looks about as good as every other. At which point the decision falls to whoever argues most confidently, and the analysis has contributed nothing.
The ranking that actually holds
The question you want answered is not "how many minutes did this cost?" but "how much production comes back if I remove it?" Those are different numbers, and only the second one justifies a budget.
A model answers it by subtraction. Take the validated line, remove one failure mode, re-run, and read the difference in output. Repeat for every failure mode. What you get is a ranking by recoverable throughput — which accounts for position, for interaction, and for the buffers that absorb some losses and not others.
Two rankings of the same line
| Ranked by | What it measures | What it misses |
|---|---|---|
| Minutes lost the loss report | how often and how long | whether the line noticed |
| Recovered throughput remove it and re-run | what you get back | nothing — it is the answer |
The two orderings routinely disagree, and the disagreements are the point. On the demo line, dropping individual interrupts separates the curves by roughly ten efficiency points, while sliding along any single curve — adding storage instead — moves you far less. The dominant failure mode is worth more than any amount of buffer, and no loss table says so.
The counterintuitive ones
Exhaustive sweeps surface findings that nobody nominates in a meeting:
- Frequent short stops beat rare long ones. A stoppage measured in seconds looks negligible on a loss report. On a tightly coupled line it never lets buffers refill, so its cost is carried continuously rather than occasionally.
- Faster is not always more. Raising line rate raises failure rate too, so throughput has an optimum rather than rising monotonically. Fischel and Lange found exactly this. It is invisible to any analysis that treats rate and reliability as separate.
- Some losses are already absorbed. Fixing them returns nothing, because a buffer was covering them. Better to know that before funding the project.
What this needs from your data
The ranking is only as good as the failure-mode detail underneath it. One lumped downtime figure per machine gets you a working model and a flat ranking — the exact failure Fischel and Lange describe. Named failure modes, fitted individually, get you a ranking you can spend against.
If you have a Line Event Data system, you likely already have what is required. The comparison between those two levels of detail is what the bottling line demo models exist to show: same line, same topology, two levels of failure-data resolution, and a different answer about where the bottleneck is.
Related: How big is your hidden factory? · Find the real bottleneck · How big should a buffer be?
Rank your own losses
The sandbox lets you drop interrupts one at a time and watch what comes back.
Open the sandbox → How the method works