Case Study · Published Validation

Within 1% of Measured OEE
A Published Model, Rebuilt and Re-run

Most simulation accuracy claims are a vendor’s word. This one is a paper. At the 2020 Winter Simulation Conference, Fischel and Lange reported a per-failure-mode reliability model of a multi-line food plant that agreed with the plant’s actual OEE to within one percentage point over a full simulated year. We rebuilt their model in ReliaSim: it reproduces the published result, and it runs roughly 1200× faster.

Sector
Food & beverage · salad dressing
The published bar
(Actual − Simulated) < 1% OEE, one-year run
Model scale
20+ unit operations, up to 20 failure modes each
Our part
Rebuilt the model in ReliaSim; ~1200× faster
Evidence
Peer-reviewed paper · plus our own timing

The study

Fischel and Lange’s 2020 Winter Simulation Conference paper models a Hidden Valley Ranch salad dressing plant — a real multi-line food operation, not a teaching example. The method runs in the order any plant already has the material for: line event data, out of the historian the plant is already keeping; distributions fitted per failure mode; then a discrete rate and reliability model of the line.

The scale is the part people underestimate. Up to twenty failure modes on each of twenty-plus distinct unit operations — roughly four hundred modes in one model. That is not modeling for its own sake, and the paper is blunt about why it is necessary:

“Had an averaged, single-failure-mode-per-unit-operation approach been followed, the gains and losses would be less accurate and essentially identical to each other.”
— Fischel & Lange, WSC 2020

Collapse a machine’s failure modes into one average and you do not merely lose resolution. You lose the ability to tell an improvement from a loss, because both come back looking the same.

What the validation gate actually was

The bar the authors set themselves was agreement within 1% of overall system OEE between a one-year simulation and the actual data, per product category. It met it.

The validation is per failure mode, not merely in aggregate — and that distinction is the whole reason the figure is worth anything. A single overall number can be right for the wrong reasons: one mode overstated, another understated, the errors quietly cancelling into a total that looks convincing. Comparing observed against simulated availability for every mode separately cannot hide that.

A reliability block per unit operation — not across the line

This is the detail most often got wrong when the paper is summarised, including by people arguing in its favour. The reliability block diagram sits inside each rate-affecting unit operation, where it carries that unit’s several failure modes and lets its failure rate scale with production rate. The line does not come from a block diagram at all — it comes from discrete-rate coupling through the valves, tanks and conveyors between the units.

So the tidy version — treat the line as blocks in series and multiply the availabilities — is not this method, and quoting it as such misstates the source. Multiplying assumes the units fail independently. Buffers are precisely what breaks that assumption, which is why a plant that multiplies out to 46.7% can measure 54.3%: the arithmetic cannot see the decoupling that the physical line actually has.

What we did with it

The paper’s model was built in ExtendSim. We rebuilt the same model in ReliaSim. It reproduces the published result — verified independently by Lange — and runs the same one-year simulated duration roughly 1200× faster in real clock time.

Two claims, and they have different sources. The 1% agreement is the published paper’s, measured by its authors in the tool they used; we did not produce it and do not claim it as ours. The 1200× is our own measurement from reproducing that model, verified by one of the paper’s authors — the paper reports the validation, not the clock. Keeping those apart is the point: the pairing is a validated method whose runtime no longer constrains the experiment.

Why the clock changes the question

A one-year run that costs an overnight is a run you do three of. You pick the three candidates in advance, and the meeting becomes an argument about which three to pick — an argument about the sampling, dressed up as an argument about the plant.

When the same run costs seconds, the three-scenario habit stops being a constraint and starts being a choice. That is why the buffer sweep behind How big should a buffer be? is 7,770 runs rather than three. The discipline does not change — you still build the model, you still need the failure data, you still validate before any answer is worth having. What collapses is the distance between asking a question and having it answered.

On the evidence

This case study is different in kind from the others here. The rest are our own project records; this one is a published study by other people, whose model we reproduced. Fischel wrote it at Clorox Services Company and Lange at his own consultancy — neither is a ChiAha customer, and nothing here should be read as their endorsement of ReliaSim.

The paper is two pages and open access, so you do not have to take our summary of it: read it directly. It is also the featured case study on ExtendSim’s own reliability block diagram page.

One piece of lineage worth stating plainly, because it cuts against our own interest in the comparison: ExtendSim’s discrete rate module descends from the Extend+Industry work that ReliaSim’s founder originated in the 1990s. The two tools in this comparison are not strangers, and the speed difference is thirty years of the same idea, not a verdict on somebody else’s engineering.

ExtendSim® is a registered trademark of Andritz Inc. It is referenced here for identification and historical purposes only; no affiliation or endorsement is implied.