Reliability

MTBF vs MTTR: formulas, a worked factory example and what good looks like

MTBF and MTTR answer two different questions about a machine. MTBF (mean time between failures) tells you how long it runs, on average, before it breaks down. MTTR (mean time to repair) tells you how long it stays down, on average, each time it does. Read them together: MTBF points you to root cause and preventive maintenance, MTTR points you to spares, skills and response.

By Urvish Savalia, Co-founder, Codiot TechnologiesPublished Last reviewed 7 min read

MTBF and MTTR on one machine timeline A pencil sketch of one machine over time. A line steps between running and down. Orange brackets mark the long running stretches between failures, and their average is MTBF. Dark brackets mark the short down stretches while the machine is repaired, and their average is MTTR. RUNNING DOWN TIME MTBF = average of the running stretches MTTR = average of the repair stretches
One machine over time. MTBF is the average length of the running stretches between failures. MTTR is the average length of the repair stretches.

MTBF and MTTR in plain words

Picture one machine on your shop floor over a month. It runs, it breaks down, your team repairs it, it runs again. Cut that month into running stretches and down stretches, and you have everything you need for both numbers.

MTBF is the average length of the running stretches. It measures reliability: how often the machine lets you down. Higher is better.

MTTR is the average length of the down stretches caused by failures. It measures maintainability and response: how long it hurts each time. Lower is better.

Neither number is enough alone. A machine that fails once a month for 20 hours and a machine that fails 20 times a month for one hour each lose the same 20 hours. The fix for the first is spares and repair instructions. The fix for the second is root cause and preventive maintenance. MTBF and MTTR tell you which machine is which.

The formulas, and the rules that keep them honest

The arithmetic fits on one line: MTBF = total operating time / number of failures, and MTTR = total repair time / number of failures. The errors come from what you count. Write these rules down, agree them with production, and keep them the same every month.

Counting rules for MTBF and MTTR you can copy
QuestionRule that worksWhy it matters
What is operating time?Hours the machine was scheduled to run, minus the hours it was down for failuresCalendar hours inflate MTBF on a one or two shift machine
What counts as a failure?An unplanned stop that needs maintenance to restore the machinePMs, setups and changeovers are not failures
Do short stops count?Pick a threshold, for example stops over 10 minutes, and log itWithout a threshold, two shifts count different things
When does the repair clock start?When the machine stops, not when the technician arrivesStarting late hides the response time inside MTTR
When does it stop?When production accepts the machine backStopping at "repair done" hides testing and handover
Does waiting for parts count?Yes, it is downtime. Also note it separatelyWaiting is often the largest part of MTTR
Which items use MTBF?Machines you repair and put back in serviceFor parts you throw away, use mean time to failure (MTTF)

From the two numbers you can also work out inherent availability: MTBF / (MTBF + MTTR). It is a ceiling, because it leaves out planned stops, setups and waiting for an operator. Try your own figures in the MTBF and MTTR calculator.

What sits inside MTTR

What sits inside one 18-hour repair (illustrative) Six bars draw in turn for one compressor breakdown: report the breakdown, half an hour; technician reaches the machine, half an hour; diagnose, one and a half hours; wait for the spare bearing, fourteen hours; repair, one hour; test and hand back, half an hour. The wait for parts is by far the longest bar. Report0.5 H Reach the machine0.5 H Diagnose1.5 H Wait for parts14 H Repair1 H Test, hand back0.5 H
One 18-hour compressor repair, split into its parts. Illustrative numbers.

MTTR is one number, but it hides several clocks. Someone notices the stop and reports it. A technician reaches the machine. They find the cause. They wait for a part, a contractor or a permit. They repair. Production tests and accepts the machine.

In many plants the hands-on repair is the smallest slice. If you only train technicians to repair faster, you attack the wrong slice. Log the start time of each step for your worst repairs for one month, and the biggest bar will tell you where to act: a critical spare on the shelf, a clear call-out rule for the night shift, or a written repair procedure for a job that only one person knows.

A worked example from a typical Indian plant

Worked example (illustrative)

Take an auto components plant near Chennai running two shifts of 8 hours on 25 working days. Each machine is scheduled for 400 hours in the month. The maintenance team looks at three machines that keep coming up in the morning meeting. All numbers below are illustrative.

Operating time is 400 hours minus the failure downtime. MTBF is operating time divided by failures, MTTR is repair time divided by failures, and inherent availability is MTBF / (MTBF + MTTR).

Machine (illustrative)FailuresRepair hoursMTBF (h)MTTR (h)Inherent availability
CNC turning centre86394 / 8 = 49.36 / 8 = 0.7598.5%
Hydraulic press410390 / 4 = 97.510 / 4 = 2.597.5%
Screw air compressor118382 / 1 = 38218 / 1 = 1895.5%

What to do with it. The CNC fails most often but recovers fast: it is a root-cause and PM problem. If six of the eight stops are chip conveyor jams, a daily conveyor clean in the operator checklist is the first move. The compressor has the best MTBF and the worst MTTR, and it feeds the whole line, so its 18 hours probably cost the most. Fourteen of those hours were spent waiting for a bearing, so the fix is a critical spare and a named supplier, not a faster technician. One compressor failure in a month is a weak guide on its own, so check the last twelve months of its history before you buy the spare. The press sits in the middle: watch its trend before acting.

What good looks like

You will find published "good MTBF" figures online, but there is no honest universal target. A centrifugal pump and a CNC fail in different ways, duty cycles differ, and every plant counts failures and repair time a little differently. Good is a trend you can trust on the same machine, with the same rules: MTBF rising and MTTR falling over three to six months. Use this grid to decide your first action for each machine.

Read MTBF and MTTR together: pattern, meaning and first action
PatternWhat it usually meansFirst action
Low MTBF, low MTTRFrequent small failures that are quick to fixFind the repeat cause; add or change a PM task
High MTBF, high MTTRRare failures that stop you for a long timeCritical spares, a written repair procedure, a trained backup
Low MTBF, high MTTRYour worst actor on both countsPut it at the top of the list; consider an overhaul or replacement case
High MTBF, low MTTRHealthy for nowKeep the PM going; check the trend each month

Compare a machine with its own past first, then with similar machines in your plant on the same rules. A useful extra check for MTTR: how long can the line wait before it misses a delivery? If MTTR on a single-point machine is longer than that, you need a spare or a standby, whatever the average says.

Common mistakes

  • Counting planned work as failures. A PM or a die change in the failure count drags MTBF down and makes your best planning look like breakdowns.
  • Using calendar hours. A one-shift machine measured on 720 hours a month shows an MTBF about three times too high.
  • Leaving repair time inside operating time. The machine is not operating while it is being fixed. Take those hours out.
  • Starting the repair clock when the technician arrives. It makes MTTR look better and hides your slowest step.
  • Averaging across very different machines. One plant-wide MTBF mixes pumps, presses and cranes into a number nobody can act on. Work machine by machine.
  • Judging a machine on one month. With two failures, one more or one less changes MTBF by a third or more. Read the trend.
  • Chasing one number only. Pushing technicians to close jobs fast can lower MTTR while the same fault returns next week, so MTBF falls. Watch both.

Frequently asked questions

Is a higher MTBF or a lower MTTR more important?

Neither on its own. Look at total downtime first, then at which number drives it. Frequent short failures call for root cause and PM work. Rare long failures call for spares, procedures and faster response.

What is a good MTBF for a machine?

There is no universal figure, because machines, duty cycles and counting rules differ. A good MTBF is one that is rising on the same machine, measured the same way, over three to six months.

What is a good MTTR?

One that is falling, and one shorter than the time your line can wait before it misses a delivery. If a single-point machine takes longer than that to repair, plan a spare or a standby.

Does MTTR include waiting for spare parts?

In the common definition, yes, because the machine is down the whole time. It helps to record the waiting time separately, since it is often the largest part of the repair.

How do MTBF and MTTR relate to availability and OEE?

Inherent availability is MTBF / (MTBF + MTTR). It is a ceiling for the availability factor in OEE, which also loses time to setups, planned stops and waiting for an operator.

Should planned maintenance count in MTBF?

No. MTBF counts failures only. Planned maintenance, setups and changeovers are recorded separately, otherwise good planning looks like poor reliability.

Fix it before it fails.

Start a free trial today, or book a demo tailored to your plant, assets and standards.

Start with one line. Extend when it works.