Run to failure maintenance: when letting an asset fail is the right strategy, what has to be in place for it to be a decision rather than neglect, and how it is recorded

Updated

Run to failure maintenance is the strategy of doing no proactive work on an asset and repairing or replacing it when it fails. Said that way it sounds like neglect, and on many sites it is, because it is the strategy every asset gets by default until someone decides otherwise. Decided deliberately, for the right assets, it is the cheapest correct strategy on the register. This page says when it is right, what has to be in place for it to be a decision, and how it is recorded so that nobody later mistakes it for a task that was forgotten.

When it is the right call

When the consequence of failure is economic only, no safety or environmental effect, no line stopped. When the cost of the failure, repair plus downtime, is less than the cost of the preventive work that would stop it. When there is no task that can predict or prevent the failure mode anyway, which is true of many electronic and some sealed components. And when a spare or a replacement is to hand, so that the failure costs hours rather than weeks. Lighting, small fans, redundant pumps and low-value instruments are the usual candidates.

What has to be in place

A written decision on the register that the asset is run to failure and why. A spare on the shelf or a supplier who can deliver in the time the operation can bear. A way for the failure to be noticed and reported, because an asset that fails silently and stays failed is not run to failure, it is abandoned. And a reactive work order when it does fail, with findings, so that a run-to-failure asset that fails often is noticed and its strategy revisited.

What it is not

It is not the absence of a schedule. It is not the strategy for anything whose failure would hurt someone or stop the operation. It is not a way to save the cost of the spare, because without the spare the downtime is the cost. And it is not permanent: an asset that was cheap to let fail becomes expensive when the operation starts to depend on it, and the decision has to be revisited when that happens.

Recording it

On the asset's register entry: strategy run to failure, the reason, the spare's location, and the date decided. In the CMMS that is a field or a note; on the maintenance log worksheet on this site it is the criticality column and a note. The reactive work orders against that asset then accumulate as its history, and a yearly look at that history is the review. The site's worked example is 180 assets with 60 preventive jobs a month, which means most of the register is on some form of inspection or run to failure, and each of those should be able to say which.

Questions people ask about run to failure maintenance

Is run to failure cheaper?

On the right assets, yes, because preventive work costs hours and the failure costs less. On the wrong assets it is the most expensive strategy there is.

Should critical assets ever be run to failure?

Not on their own. A critical asset with a redundant twin can run to failure while the twin is maintained, which is a common and sensible arrangement.

How do I know if run to failure is working?

The reactive work orders against the asset are rare and cheap, and the downtime is what was accepted. If either grows, the strategy is revisited.

Sources

Related answers

Price your reactive work against plannedSee what your PM schedule costs to run