Reliability and maintenance are two halves of one function. Reliability is the property of the asset: how long it runs before it fails. Maintenance is what the team does to keep that property. Maintenance reliability, the phrase on a job title, is the discipline of running the second to improve the first, and reliability in maintenance is measured from the records the team keeps. This page sets out what the words mean together, how maintenance and reliability are measured from work orders, and where reliability is actually won or lost on a site.
What reliability means for an asset
The probability that it does its job for a given time under the conditions it runs in. In practice that is measured as MTBF, the average operating time between failures, from the reactive work orders against the asset and its run hours. An asset is more reliable when its MTBF rises. The definition matters because reliability is about failures, not about how much maintenance is done; an asset can be maintained heavily and still be unreliable if the tasks are the wrong ones.
Measuring it from the maintenance records
MTBF says how often an asset fails; MTTR says how long each failure costs; the reactive share says what part of the team's work is failures. All three come from work orders that record the type, the asset, and the times. A team that keeps those records can measure its own reliability by asset and by month; one that does not is guessing, however skilled its technicians. The maintenance log worksheet on this site is that record in one sheet.
Where reliability is won
In the schedule. The right preventive task at the right interval on the right asset stops the failures the task was written for, and the history says whether it is working: reactive work orders arriving between preventive tasks mean the interval or the task is wrong. In the findings: a technician who records what was found lets the next one fix the cause instead of the symptom. And in the response to findings: a corrective work order raised and done before the failure is reliability bought cheaply.
Where it is lost
In deferral, which moves failures from planned to reactive. In records that say checked, ok. In schedules copied from the manual and never adjusted. In treating reliability as an engineering project rather than a daily habit. This site's worked example, 84 breakdowns a month at 3.4 hours each against 1.1 planned, is the cost of reliability lost, priced, and the schedule that runs is how it is won back.
Questions people ask about reliability and maintenance
Is a reliability engineer the same as a maintenance manager?
No. The reliability engineer analyses failures and designs the strategy; the maintenance manager runs the schedule and the team. On a small site one person does both, and the records serve both.
What is a reliability program?
A decided strategy per asset, a schedule that carries it out, records that measure it, and a habit of adjusting the schedule from the records. Without the last it is a document.
Can reliability be improved without new equipment?
Usually yes, because most unreliability is a wrong task or a missed one. New equipment is the answer only when the failure mode is designed in.