ReliaSim Reliability engineering simulation

Reliability Engineering

Simulation for Reliability Engineering
availability is a machine number — throughput is the one you get asked for

Your deliverables are MTBF, MTTR and availability, per asset, defensibly calculated. The question in the capital meeting is what the line will make next year. Those are different quantities, and the gap between them is not a reporting problem — it is a modelling one.

1 ppof measured OEE, across a full year — the bar a published model met
0.73×what the higher-ranked failure mode returned when removed
1.21×what the lower-ranked one returned, on near-identical recorded loss

What a reliability engineer actually gets asked for

The reliability function produces asset-level numbers: time between failures, time to repair, availability per unit, a criticality ranking, a spares policy. All of it is correct and all of it is per machine.

Then someone asks whether the line makes the volume, or which of two fixes to fund, or whether a bigger surge tank is worth the floor space. Those are questions about a system under interacting failures, and no amount of per-asset availability adds up to an answer, because the machines do not fail independently of each other's consequences. One stops; its neighbours starve or block; the cost lands somewhere else.

This page is about closing that gap without abandoning the reliability work you already did. The distributions you fitted are exactly the input a line model needs.

Where the block diagram stops

A reliability block diagram in series multiplies availabilities. Five machines at 85% give 44%, and a real line with those machines does better — sometimes much better — because there is storage between them. The RBD has nowhere to put a buffer, so it cannot represent the thing that makes the real number differ.

That is not a flaw in the method; it is the method's scope. An RBD answers a question about a system's configuration. A line's output is a question about its dynamics, and dynamics need a clock.

The long version: reliability block diagrams vs simulation and what a RAM analysis can and cannot tell you.

Reliability as architecture, not as a parameter

In a general-purpose discrete-event model, reliability is usually a number you type into a machine: a downtime percentage, or one MTBF and one MTTR. The model then behaves as though the machine has a single way of failing.

Real machines do not. A capper has a plow area, a glue station, a picker tree, a compression stage, a magazine — each with its own failure behaviour, each racing the others. In ReliaSim the unit operation carries its failure modes natively: every mode gets its own time-to-failure and time-to-repair distribution, fitted from your own event data, and the engine resolves which one fires first.

Why averaging the modes destroys the ranking

The published model behind our validation carries twenty-plus unit operations with up to twenty failure modes each. Collapse those to one averaged mode per machine and the ranking you came for disappears — gains and losses come out, in the paper's own words, essentially identical to each other.

That is the whole argument for keeping modes separate. A ranking is only useful if the differences between its entries survive the modelling.

What reliability does to a line

A failure's cost is not its own downtime. It is its own downtime plus what it does to the machines on either side — and that second part is recorded under their names, not under its.

In our bottling line, 120 one-minute micro stops at the Filler landed as idle time on the Labeler, the Case Packer and the Palletizer. Three machines logged a cost caused by a fourth. An availability calculation cannot produce that quantity, because availability is defined per asset and the cost crossed assets.

See starved vs blocked and micro-stops vs breakdowns.

Rank by what comes back, not by what it cost

This is the payoff, and it is measurable. Take two failure modes with near-identical recorded losses and ask the model what each returns when removed:

Same bar on the Pareto, different decision

Failure modeRecorded lossReturned when removedReturn ÷ loss
Filler Micro Stop — frequent, short6.72%+8.1 pp1.21×
Labeler Misalignment — rare, long6.79%+5.0 pp0.73×

The lower-ranked mode recovers 62% more. One returns more than it cost, because removing it also removes what it was cascading; the other returns less, because some of its downtime was absorbed by the line and cost nothing. Direct loss does not rank the fix.

Worth being precise about what this is and is not. It is not a forecast of when a component will fail — that is a different discipline with its own tools. It is a counterfactual: what the line would produce under a change nobody has made yet. See why condition monitoring cannot answer it.

Validation you can go and check

A model that predicts a configuration nobody has built should first reproduce one somebody measured.

A peer-reviewed WSC 2020 study modelled a real multi-line food plant — twenty-plus unit operations, up to twenty failure modes on each — and agreed with the plant's measured OEE to within one percentage point across a full year. That model was rebuilt in ReliaSim and independently re-validated by Tom Lange, one of its co-authors, to the same bar.

It is two pages and openly available. The point of citing it is that it is falsifiable and was not falsified, which is a different kind of claim from a capability list.

Why speed is an argument, not a spec

Ranking every failure mode by what it returns means one full re-run per mode. Sizing a buffer properly means a sweep, not three sizes somebody nominated in a meeting. Both are unaffordable if a run takes minutes.

ReliaSim simulates a year in 0.154 seconds. A buffer study of 7,770 runs — 1,864 simulated years — completes in 31.8 seconds. That changes which questions are worth asking: you stop nominating scenarios and start sweeping them.

Where each tool fits

These are not substitutes, and most teams that do this well use more than one.

ToolThe question it answers best
RBD tools — BlockSim and similarComponent reliability, redundancy configuration, spares policy, system availability
ExtendSimGeneral-purpose simulation, and the environment where discrete rate modelling first shipped
ReliaSimWhat a line produces under interacting failures, and what a change returns

On the middle row: same inventor. Andrew Siprelle created discrete rate simulation in 1995 and originated the technology behind ExtendSim's rate modules. ExtendSim is the general-purpose environment where the method first shipped; ReliaSim is the reliability-first engine built around it. The validation study above was built in one and rebuilt in the other, which is why the comparison can rest on a shared artefact rather than a contested claim.

What you need to start

The stop and start times your historian already records. That is the raw material: per-mode time-to-failure and time-to-repair distributions are fitted from it, and the Interrupt Designer is where they are set or adjusted. If you have already done a RAM study, you have most of what a line model needs — it was just aggregated to the wrong level.

Frequently asked questions

Can a RAM study predict line OEE?

Not on its own. RAM gives per-asset availability; line OEE depends on how those assets interact through buffers and on where the constraint sits. The RAM output is a valid input to the line model, not a substitute for it.

Is MTBF enough to size a buffer?

No. Two machines with the same MTBF and MTTR — one stopping rarely and long, the other often and briefly — need very different buffers, because a buffer absorbs short frequent interruptions far better than rare long ones. The distributions matter, not just their means. See how big should a buffer be.

What failure data do I need before starting?

Time-stamped stop and start events per machine, ideally with a fault name. A Line Event Data System export is ideal. Where a mode has too few events to fit, it can be represented by an estimate and flagged as one.

How is this different from predictive maintenance software?

Condition monitoring forecasts when a specific asset will fail, from sensors on that asset. This forecasts what the line produces under a configuration that does not exist yet. Both are useful; neither answers the other's question, and a perfect failure prediction still will not tell you what fixing it returns.

Do I need simulation expertise to use it?

No. The models are built from a production graph rather than written as code, and the browser sandbox runs validated bottling-line models with nothing installed.

Run a validated model, not a trial

Eight models of one real bottling line, in the browser, with the failure modes exposed. No download, no licence, no sign-up.

Open the Sandbox →