Safety
Incident investigation on the shop floor: from report to root cause
A good incident investigation finds out what happened, why it happened, and what will stop it happening again, without hunting for someone to blame. You make the area safe, gather the facts while they are fresh, work back through the immediate, underlying and root causes, and close each cause with an action that someone verifies. Incident management software helps by keeping the report, the evidence, the causes and the actions on one record, so nothing is lost between the shop floor and the safety meeting.
What an investigation is for
An incident investigation has one job: stop the same thing, or something like it, from happening again. It is not a hunt for the person to blame. If your people expect a punishment at the end, they stop reporting near misses, and near misses are your cheapest warnings.
It helps to agree the words first. The UK Health and Safety Executive guide HSG245 calls an event that hurt someone an accident, and an event that could have hurt someone but did not a near miss. Both are worth investigating. A hydraulic hose that bursts beside an empty operator station is the same failure as one that bursts beside a full one.
HSG245 also gives you three levels of cause. Keep them in mind for the rest of this post:
- Immediate cause: the most obvious reason. The guard was missing, the hose burst, the floor was wet.
- Underlying cause: the unsafe act or condition behind it. Nobody checked the guard at start-up, the hose was rubbing on the frame.
- Root cause: the management or system failing that let the underlying cause exist. There was no reassembly check after repairs, training needs were never assessed.
Fix only the immediate cause and you fix one machine for one day. Fix the root cause and you prevent a whole family of incidents.
From report to root cause in six steps
HSG245 describes four steps: gather the information, analyse it, identify risk controls, and make an action plan. On a shop floor it helps to add one step before and one after.
- 01
Make safe and report
Give first aid, stop or isolate the machine, and tell the supervisor. Report anything notifiable to the authorities on time under the rules that apply to your factory. Do not wait for the investigation to finish.
- 02
Decide how deep to go
Judge the worst outcome that could reasonably have happened, not only what did happen, and how likely it is to happen again. A near miss that could have been a fatality deserves a team investigation.
- 03
Gather the facts while they are fresh
Photograph the scene before it is cleaned. Keep the broken part. Talk to witnesses separately, on the same shift if you can. Pull the machine history, the permit, the last work order and the training records.
- 04
Build the sequence
Write down what happened, in order, with times. Gaps in the timeline are where the questions are.
- 05
Find the causes
For each event in the sequence, ask why until you reach a decision, a missing check or a system gap. You will usually find more than one chain of causes.
- 06
Act, then verify
Give every cause an action, an owner and a date. Someone other than the owner checks that the action was done and that it worked.
Five is not a magic number. Stop when the answer is something your plant controls through a procedure, a check, a design or a decision, and when fixing it would also stop other incidents. If your last answer is "the operator was careless", keep going: ask why the system let a careless moment turn into an injury.
A worked example from a typical Indian plant
Worked example (illustrative)
Take a sheet metal plant near Pune with a 250-tonne hydraulic press. At 11:40 on a Tuesday a hydraulic hose bursts and sprays oil across the walkway beside the operator station. The operator has stepped away to fetch blanks, so nobody is hurt. It is a near miss. All details below are illustrative.
The supervisor isolates the press, ropes off the walkway and reports it at 11:55. The EHS officer rates the worst likely outcome as a serious burn or eye injury, and the chance of a repeat as possible, so the plant runs a medium level investigation with the maintenance head and an employee representative.
The team photographs the hose before it is removed and keeps it. The burst is at a point where the hose was rubbing on the press frame. The machine history shows a seal change on the same press three weeks earlier, and the job card has no step for refitting the hose clamps.
| Cause level | What the team found | Action, owner and date |
|---|---|---|
| Immediate | Hose burst where it was worn through | Replace the hose and inspect every hose on the press. Maintenance supervisor, same day |
| Underlying | A clamp was not refitted after the seal change, so the hose rubbed on the frame | Refit the clamps and fit a spray shield on hoses near the operator station. Maintenance head, 7 days |
| Underlying | No spray shield between the hoses and the walkway | Add hose shields to the HIRA for all presses. EHS officer, 14 days |
| Root | Repair jobs can be closed with no reassembly check | Add "guards and clamps refitted" to every hydraulic job card, signed by a second person. Plant engineering head, 14 days |
| Check | Is the new step being used? | Review 10 closed hydraulic jobs after 60 days. EHS officer, day 74 |
What changed. The quick fix was a new hose. The lasting fix was a reassembly check on every hydraulic job, which protects every press in the plant, not just this one. The 60-day review is the step most plants skip, and it is the one that tells you the root cause action actually worked.
Incident investigation checklist you can copy
Print this, or turn it into a form. The times are a sensible default for a medium level investigation. Shorten them for a serious event.
| When | What to do | What to record |
|---|---|---|
| First hour | Give first aid, make the area safe, isolate the machine, tell the supervisor | Time, place, people involved, injury if any, who was told |
| Same shift | Photograph the scene, keep broken parts, talk to witnesses one by one | Photos, witness names and statements, the condition of guards and PPE |
| Within 24 hours | Decide the investigation level and the team | Worst likely outcome, likelihood of a repeat, level chosen and why |
| Within 3 days | Build the timeline and pull the records | Sequence of events, machine history, permits, last work order, training records |
| Within 7 days | Find immediate, underlying and root causes | Each cause chain, with the evidence for each step |
| Within 7 days | Agree actions | One or more actions per cause, each with an owner and a due date |
| 30 to 90 days | Verify the actions worked | Who checked, when, what they found, and whether the case can close |
| At close | Share the lesson | Toolbox talk held, risk assessment updated, similar machines checked |
What to look for in incident management software
A paper register and a shared spreadsheet can work for a small site. They break down when reports sit in someone's drawer, photos live on a phone, and actions are tracked in a separate sheet that nobody opens. Good incident management software keeps one record from the first report to the verified close. When you compare options, look for these:
- Easy reporting. Anyone on the floor can report a hazard, near miss or incident in a minute, from a phone or a shared terminal.
- Investigation built in. Witnesses, photos and documents, an investigation checklist and cause analysis on the same record as the report.
- Real corrective actions. Each action has an owner, a due date and a separate person who verifies it, plus an effectiveness check before it closes.
- Links to your assets. If the incident involved a machine, you should see its work orders and repair history without opening another system.
- Rates from real hours. LTIFR and TRIR worked out from the injuries and man-hours you record, not typed in by hand.
- An audit trail. Every change recorded, so you can show an auditor who did what and when.
What you can skip: long dashboards nobody reads, and dozens of mandatory fields that make people stop reporting. A short form that gets used beats a perfect form that does not.
Common mistakes
- Stopping at "human error". It is a description, not a cause. Ask why the system allowed the error, or why the error was not caught.
- Investigating only injuries. Near misses tell you the same story at no cost. Judge the depth by what could have happened.
- Cleaning up before you look. Once the oil is mopped and the part is in the scrap bin, the evidence is gone.
- One cause, one action. Most incidents have several causes. One action against one cause leaves the rest in place.
- Actions with no owner or date. "Improve training" is not an action. "Train all four press operators on hose checks by 30 November, owner: shift in charge" is.
- Closing without checking. An action marked done is not an action that worked. Check it after 30 to 90 days.
- Keeping the lesson in one department. If maintenance found the root cause, production and EHS need to hear it, and the same machine type in other bays needs checking.
Sources
- Investigating accidents and incidents (HSG245) (opens in a new tab), Health and Safety Executive (HSE), UK
- HSG245 full text (PDF) (opens in a new tab), Health and Safety Executive (HSE), UK
- Five whys (opens in a new tab), Wikipedia
- Root cause analysis (opens in a new tab), Wikipedia
- ISO 45001 (opens in a new tab), Wikipedia
Frequently asked questions
What is the difference between an incident and an accident?
In HSG245 terms, an accident is an event that causes injury or ill health. An incident is an event that could have caused harm but did not, such as a near miss. Both are worth investigating.
Who should investigate an incident?
The supervisor of the area for minor events. For more serious events, a team with the line manager, a health and safety adviser and an employee representative, with senior management overseeing the most serious ones.
How soon should an investigation start?
As soon as the area is safe. Photos, broken parts and witness memories are best in the first hours. Report anything notifiable to the authorities on time, without waiting for the investigation to finish.
Is five whys enough for root cause analysis?
It is a good start for most shop floor incidents. Follow every branch, not just one, and stop when you reach a system or management cause your plant can fix. For serious events, use a team and a more structured method.
Should near misses be investigated?
Yes. A near miss has the same causes as the injury it could have been. Set the depth of the investigation by the worst outcome that could reasonably have happened.
What should incident management software do?
Let anyone report quickly, keep the evidence, causes and actions on one record, track each corrective action to a verified close, and work out LTIFR and TRIR from the hours actually worked.