Safety

Incident investigation on the shop floor: from report to root cause

A good incident investigation finds out what happened, why it happened, and what will stop it happening again, without hunting for someone to blame. You make the area safe, gather the facts while they are fresh, work back through the immediate, underlying and root causes, and close each cause with an action that someone verifies. Incident management software helps by keeping the report, the evidence, the causes and the actions on one record, so nothing is lost between the shop floor and the safety meeting.

By Dhruvesh, Lead Architect, MaintenanceIQPublished Last reviewed 8 min read

An incident investigation, from report to root cause A pencil sketch of five boxes joined by arrows: make safe and report, gather the facts, find the causes, act, and verify. Under the causes box, three stacked layers are labelled immediate, underlying and root. A curved arrow runs from verify back to the start, labelled lessons feed your risk assessments. Make safeand report Gatherthe facts Find thecauses Act Verify IMMEDIATE UNDERLYING ROOT Lessons feed your risk assessments
From report to root cause: make safe, gather the facts, find the causes, act, then check the action worked. What you learn feeds back into your risk assessments.

What an investigation is for

An incident investigation has one job: stop the same thing, or something like it, from happening again. It is not a hunt for the person to blame. If your people expect a punishment at the end, they stop reporting near misses, and near misses are your cheapest warnings.

It helps to agree the words first. The UK Health and Safety Executive guide HSG245 calls an event that hurt someone an accident, and an event that could have hurt someone but did not a near miss. Both are worth investigating. A hydraulic hose that bursts beside an empty operator station is the same failure as one that bursts beside a full one.

HSG245 also gives you three levels of cause. Keep them in mind for the rest of this post:

  • Immediate cause: the most obvious reason. The guard was missing, the hose burst, the floor was wet.
  • Underlying cause: the unsafe act or condition behind it. Nobody checked the guard at start-up, the hose was rubbing on the frame.
  • Root cause: the management or system failing that let the underlying cause exist. There was no reassembly check after repairs, training needs were never assessed.

Fix only the immediate cause and you fix one machine for one day. Fix the root cause and you prevent a whole family of incidents.

From report to root cause in six steps

HSG245 describes four steps: gather the information, analyse it, identify risk controls, and make an action plan. On a shop floor it helps to add one step before and one after.

  1. 01

    Make safe and report

    Give first aid, stop or isolate the machine, and tell the supervisor. Report anything notifiable to the authorities on time under the rules that apply to your factory. Do not wait for the investigation to finish.

  2. 02

    Decide how deep to go

    Judge the worst outcome that could reasonably have happened, not only what did happen, and how likely it is to happen again. A near miss that could have been a fatality deserves a team investigation.

  3. 03

    Gather the facts while they are fresh

    Photograph the scene before it is cleaned. Keep the broken part. Talk to witnesses separately, on the same shift if you can. Pull the machine history, the permit, the last work order and the training records.

  4. 04

    Build the sequence

    Write down what happened, in order, with times. Gaps in the timeline are where the questions are.

  5. 05

    Find the causes

    For each event in the sequence, ask why until you reach a decision, a missing check or a system gap. You will usually find more than one chain of causes.

  6. 06

    Act, then verify

    Give every cause an action, an owner and a date. Someone other than the owner checks that the action was done and that it worked.

Asking why, from immediate cause to root cause (illustrative) Five answers appear one after another, each joined to the next by a line. Oil sprayed on the walkway. The hose burst. It was rubbing on the press frame. A clamp was not refitted after a seal change. The job card had no reassembly check. The last answer is marked as the root cause. Oil sprayed on the walkway WHY The hose burst WHY It was rubbing on the press frame WHY A clamp was not refitted after a seal change WHY The job card had no reassembly checkROOT CAUSE
Asking why moves you from the immediate cause to the root cause. Illustrative chain from the worked example below.

Five is not a magic number. Stop when the answer is something your plant controls through a procedure, a check, a design or a decision, and when fixing it would also stop other incidents. If your last answer is "the operator was careless", keep going: ask why the system let a careless moment turn into an injury.

A worked example from a typical Indian plant

Worked example (illustrative)

Take a sheet metal plant near Pune with a 250-tonne hydraulic press. At 11:40 on a Tuesday a hydraulic hose bursts and sprays oil across the walkway beside the operator station. The operator has stepped away to fetch blanks, so nobody is hurt. It is a near miss. All details below are illustrative.

The supervisor isolates the press, ropes off the walkway and reports it at 11:55. The EHS officer rates the worst likely outcome as a serious burn or eye injury, and the chance of a repeat as possible, so the plant runs a medium level investigation with the maintenance head and an employee representative.

The team photographs the hose before it is removed and keeps it. The burst is at a point where the hose was rubbing on the press frame. The machine history shows a seal change on the same press three weeks earlier, and the job card has no step for refitting the hose clamps.

Cause levelWhat the team foundAction, owner and date
ImmediateHose burst where it was worn throughReplace the hose and inspect every hose on the press. Maintenance supervisor, same day
UnderlyingA clamp was not refitted after the seal change, so the hose rubbed on the frameRefit the clamps and fit a spray shield on hoses near the operator station. Maintenance head, 7 days
UnderlyingNo spray shield between the hoses and the walkwayAdd hose shields to the HIRA for all presses. EHS officer, 14 days
RootRepair jobs can be closed with no reassembly checkAdd "guards and clamps refitted" to every hydraulic job card, signed by a second person. Plant engineering head, 14 days
CheckIs the new step being used?Review 10 closed hydraulic jobs after 60 days. EHS officer, day 74

What changed. The quick fix was a new hose. The lasting fix was a reassembly check on every hydraulic job, which protects every press in the plant, not just this one. The 60-day review is the step most plants skip, and it is the one that tells you the root cause action actually worked.

Incident investigation checklist you can copy

Print this, or turn it into a form. The times are a sensible default for a medium level investigation. Shorten them for a serious event.

Incident investigation checklist, from report to close
WhenWhat to doWhat to record
First hourGive first aid, make the area safe, isolate the machine, tell the supervisorTime, place, people involved, injury if any, who was told
Same shiftPhotograph the scene, keep broken parts, talk to witnesses one by onePhotos, witness names and statements, the condition of guards and PPE
Within 24 hoursDecide the investigation level and the teamWorst likely outcome, likelihood of a repeat, level chosen and why
Within 3 daysBuild the timeline and pull the recordsSequence of events, machine history, permits, last work order, training records
Within 7 daysFind immediate, underlying and root causesEach cause chain, with the evidence for each step
Within 7 daysAgree actionsOne or more actions per cause, each with an owner and a due date
30 to 90 daysVerify the actions workedWho checked, when, what they found, and whether the case can close
At closeShare the lessonToolbox talk held, risk assessment updated, similar machines checked

What to look for in incident management software

A paper register and a shared spreadsheet can work for a small site. They break down when reports sit in someone's drawer, photos live on a phone, and actions are tracked in a separate sheet that nobody opens. Good incident management software keeps one record from the first report to the verified close. When you compare options, look for these:

  • Easy reporting. Anyone on the floor can report a hazard, near miss or incident in a minute, from a phone or a shared terminal.
  • Investigation built in. Witnesses, photos and documents, an investigation checklist and cause analysis on the same record as the report.
  • Real corrective actions. Each action has an owner, a due date and a separate person who verifies it, plus an effectiveness check before it closes.
  • Links to your assets. If the incident involved a machine, you should see its work orders and repair history without opening another system.
  • Rates from real hours. LTIFR and TRIR worked out from the injuries and man-hours you record, not typed in by hand.
  • An audit trail. Every change recorded, so you can show an auditor who did what and when.

What you can skip: long dashboards nobody reads, and dozens of mandatory fields that make people stop reporting. A short form that gets used beats a perfect form that does not.

Common mistakes

  • Stopping at "human error". It is a description, not a cause. Ask why the system allowed the error, or why the error was not caught.
  • Investigating only injuries. Near misses tell you the same story at no cost. Judge the depth by what could have happened.
  • Cleaning up before you look. Once the oil is mopped and the part is in the scrap bin, the evidence is gone.
  • One cause, one action. Most incidents have several causes. One action against one cause leaves the rest in place.
  • Actions with no owner or date. "Improve training" is not an action. "Train all four press operators on hose checks by 30 November, owner: shift in charge" is.
  • Closing without checking. An action marked done is not an action that worked. Check it after 30 to 90 days.
  • Keeping the lesson in one department. If maintenance found the root cause, production and EHS need to hear it, and the same machine type in other bays needs checking.

Frequently asked questions

What is the difference between an incident and an accident?

In HSG245 terms, an accident is an event that causes injury or ill health. An incident is an event that could have caused harm but did not, such as a near miss. Both are worth investigating.

Who should investigate an incident?

The supervisor of the area for minor events. For more serious events, a team with the line manager, a health and safety adviser and an employee representative, with senior management overseeing the most serious ones.

How soon should an investigation start?

As soon as the area is safe. Photos, broken parts and witness memories are best in the first hours. Report anything notifiable to the authorities on time, without waiting for the investigation to finish.

Is five whys enough for root cause analysis?

It is a good start for most shop floor incidents. Follow every branch, not just one, and stop when you reach a system or management cause your plant can fix. For serious events, use a team and a more structured method.

Should near misses be investigated?

Yes. A near miss has the same causes as the injury it could have been. Set the depth of the investigation by the worst outcome that could reasonably have happened.

What should incident management software do?

Let anyone report quickly, keep the evidence, causes and actions on one record, track each corrective action to a verified close, and work out LTIFR and TRIR from the hours actually worked.

Fix it before it fails.

Start a free trial today, or book a demo tailored to your plant, assets and standards.

Start with one line. Extend when it works.