Reliability
MTBF vs MTTR: formulas, a worked factory example and what good looks like
MTBF and MTTR answer two different questions about a machine. MTBF (mean time between failures) tells you how long it runs, on average, before it breaks down. MTTR (mean time to repair) tells you how long it stays down, on average, each time it does. Read them together: MTBF points you to root cause and preventive maintenance, MTTR points you to spares, skills and response.
MTBF and MTTR in plain words
Picture one machine on your shop floor over a month. It runs, it breaks down, your team repairs it, it runs again. Cut that month into running stretches and down stretches, and you have everything you need for both numbers.
MTBF is the average length of the running stretches. It measures reliability: how often the machine lets you down. Higher is better.
MTTR is the average length of the down stretches caused by failures. It measures maintainability and response: how long it hurts each time. Lower is better.
Neither number is enough alone. A machine that fails once a month for 20 hours and a machine that fails 20 times a month for one hour each lose the same 20 hours. The fix for the first is spares and repair instructions. The fix for the second is root cause and preventive maintenance. MTBF and MTTR tell you which machine is which.
The formulas, and the rules that keep them honest
The arithmetic fits on one line: MTBF = total operating time / number of failures, and MTTR = total repair time / number of failures. The errors come from what you count. Write these rules down, agree them with production, and keep them the same every month.
| Question | Rule that works | Why it matters |
|---|---|---|
| What is operating time? | Hours the machine was scheduled to run, minus the hours it was down for failures | Calendar hours inflate MTBF on a one or two shift machine |
| What counts as a failure? | An unplanned stop that needs maintenance to restore the machine | PMs, setups and changeovers are not failures |
| Do short stops count? | Pick a threshold, for example stops over 10 minutes, and log it | Without a threshold, two shifts count different things |
| When does the repair clock start? | When the machine stops, not when the technician arrives | Starting late hides the response time inside MTTR |
| When does it stop? | When production accepts the machine back | Stopping at "repair done" hides testing and handover |
| Does waiting for parts count? | Yes, it is downtime. Also note it separately | Waiting is often the largest part of MTTR |
| Which items use MTBF? | Machines you repair and put back in service | For parts you throw away, use mean time to failure (MTTF) |
From the two numbers you can also work out inherent availability: MTBF / (MTBF + MTTR). It is a ceiling, because it leaves out planned stops, setups and waiting for an operator. Try your own figures in the MTBF and MTTR calculator.
What sits inside MTTR
MTTR is one number, but it hides several clocks. Someone notices the stop and reports it. A technician reaches the machine. They find the cause. They wait for a part, a contractor or a permit. They repair. Production tests and accepts the machine.
In many plants the hands-on repair is the smallest slice. If you only train technicians to repair faster, you attack the wrong slice. Log the start time of each step for your worst repairs for one month, and the biggest bar will tell you where to act: a critical spare on the shelf, a clear call-out rule for the night shift, or a written repair procedure for a job that only one person knows.
A worked example from a typical Indian plant
Worked example (illustrative)
Take an auto components plant near Chennai running two shifts of 8 hours on 25 working days. Each machine is scheduled for 400 hours in the month. The maintenance team looks at three machines that keep coming up in the morning meeting. All numbers below are illustrative.
Operating time is 400 hours minus the failure downtime. MTBF is operating time divided by failures, MTTR is repair time divided by failures, and inherent availability is MTBF / (MTBF + MTTR).
| Machine (illustrative) | Failures | Repair hours | MTBF (h) | MTTR (h) | Inherent availability |
|---|---|---|---|---|---|
| CNC turning centre | 8 | 6 | 394 / 8 = 49.3 | 6 / 8 = 0.75 | 98.5% |
| Hydraulic press | 4 | 10 | 390 / 4 = 97.5 | 10 / 4 = 2.5 | 97.5% |
| Screw air compressor | 1 | 18 | 382 / 1 = 382 | 18 / 1 = 18 | 95.5% |
What to do with it. The CNC fails most often but recovers fast: it is a root-cause and PM problem. If six of the eight stops are chip conveyor jams, a daily conveyor clean in the operator checklist is the first move. The compressor has the best MTBF and the worst MTTR, and it feeds the whole line, so its 18 hours probably cost the most. Fourteen of those hours were spent waiting for a bearing, so the fix is a critical spare and a named supplier, not a faster technician. One compressor failure in a month is a weak guide on its own, so check the last twelve months of its history before you buy the spare. The press sits in the middle: watch its trend before acting.
What good looks like
You will find published "good MTBF" figures online, but there is no honest universal target. A centrifugal pump and a CNC fail in different ways, duty cycles differ, and every plant counts failures and repair time a little differently. Good is a trend you can trust on the same machine, with the same rules: MTBF rising and MTTR falling over three to six months. Use this grid to decide your first action for each machine.
| Pattern | What it usually means | First action |
|---|---|---|
| Low MTBF, low MTTR | Frequent small failures that are quick to fix | Find the repeat cause; add or change a PM task |
| High MTBF, high MTTR | Rare failures that stop you for a long time | Critical spares, a written repair procedure, a trained backup |
| Low MTBF, high MTTR | Your worst actor on both counts | Put it at the top of the list; consider an overhaul or replacement case |
| High MTBF, low MTTR | Healthy for now | Keep the PM going; check the trend each month |
Compare a machine with its own past first, then with similar machines in your plant on the same rules. A useful extra check for MTTR: how long can the line wait before it misses a delivery? If MTTR on a single-point machine is longer than that, you need a spare or a standby, whatever the average says.
Common mistakes
- Counting planned work as failures. A PM or a die change in the failure count drags MTBF down and makes your best planning look like breakdowns.
- Using calendar hours. A one-shift machine measured on 720 hours a month shows an MTBF about three times too high.
- Leaving repair time inside operating time. The machine is not operating while it is being fixed. Take those hours out.
- Starting the repair clock when the technician arrives. It makes MTTR look better and hides your slowest step.
- Averaging across very different machines. One plant-wide MTBF mixes pumps, presses and cranes into a number nobody can act on. Work machine by machine.
- Judging a machine on one month. With two failures, one more or one less changes MTBF by a third or more. Read the trend.
- Chasing one number only. Pushing technicians to close jobs fast can lower MTTR while the same fault returns next week, so MTBF falls. Watch both.
Sources
- Repairable systems, non-repairable populations and lifetime distribution models (opens in a new tab), NIST/SEMATECH e-Handbook of Statistical Methods
- Homogeneous Poisson Process (HPP) (opens in a new tab), NIST/SEMATECH e-Handbook of Statistical Methods
- Mean time between failures (opens in a new tab), Wikipedia
- Mean time to repair (opens in a new tab), Wikipedia
Frequently asked questions
Is a higher MTBF or a lower MTTR more important?
Neither on its own. Look at total downtime first, then at which number drives it. Frequent short failures call for root cause and PM work. Rare long failures call for spares, procedures and faster response.
What is a good MTBF for a machine?
There is no universal figure, because machines, duty cycles and counting rules differ. A good MTBF is one that is rising on the same machine, measured the same way, over three to six months.
What is a good MTTR?
One that is falling, and one shorter than the time your line can wait before it misses a delivery. If a single-point machine takes longer than that to repair, plan a spare or a standby.
Does MTTR include waiting for spare parts?
In the common definition, yes, because the machine is down the whole time. It helps to record the waiting time separately, since it is often the largest part of the repair.
How do MTBF and MTTR relate to availability and OEE?
Inherent availability is MTBF / (MTBF + MTTR). It is a ceiling for the availability factor in OEE, which also loses time to setups, planned stops and waiting for an operator.
Should planned maintenance count in MTBF?
No. MTBF counts failures only. Planned maintenance, setups and changeovers are recorded separately, otherwise good planning looks like poor reliability.
Keep reading
MTBF and MTTR reliability analytics
MTBF, MTTR and downtime per machine from the breakdowns you log.
ReadPreventive maintenance: types and how to start
Turn a low MTBF into planned PM tasks in one week.
ReadMTBF and MTTR calculator
Work out both numbers and inherent availability from your own figures.
Read