Back to Resources

Incident Investigation Best Practices

How to investigate incidents so they don't happen twice — securing evidence early, finding root causes instead of convenient ones, and turning findings into actions that close.

The purpose of an incident investigation is not to complete a form or find someone to blame — it is to stop the next incident. This guide covers the practices that separate investigations that change things from investigations that file things: securing evidence early, asking why until you reach the management system, and turning findings into actions that actually close.

Why investigate — and why "who did it?" is the wrong question

Every incident is information: the moment where your risk assessment, training, supervision or equipment met reality and lost. Investigations that start from "who failed?" get compliance answers — people protect themselves, detail disappears, and the report lands on the person closest to the injury rather than the system that set them up. Investigations that start from "what made this possible?" get learning. The distinction is not soft; it determines whether your near-miss reporting survives. Workers report what they see treated as learning and hide what they see treated as evidence for discipline.

In the UK the legal frame is twofold. RIDDOR sets what must be reported and when — see our RIDDOR guide for the deadlines — while the HSE's guidance Investigating accidents and incidents (HSG245) describes the investigation itself: gather information, analyse it, identify risk control measures, and make an action plan. An investigation record is also your evidence of due diligence; if an incident later becomes an enforcement matter or a civil claim, the quality of your investigation is examined alongside the incident itself.

The first hours: secure people, then evidence

Immediate priorities in order: make the area safe, get people the care they need, and prevent a recurrence in the next hour — isolate the equipment, stop the task, brief the shift. Then preserve the evidence, because it decays fast:

  • Photograph everything before it moves — wide shots for context, close-ups for detail, and the things that seem irrelevant now.
  • Capture first accounts the same day. Short, factual, in the witness's own words.
  • Quarantine physical evidence — the tool, the fitting, the PPE — and secure CCTV before it is overwritten.
  • Record the boring context: weather, lighting, staffing levels, shift hours, what else was happening. Root causes hide in the boring context.

Finding root causes, not convenient ones

Most investigations stop too early. "Operative slipped on wet floor" is an immediate cause; the investigation's job is to keep asking why the condition existed. Why was the floor wet? — the machine leaks. Why does it leak? — the seal failed. Why wasn't that caught? — maintenance is reactive, and inspections don't cover it. Why? — the maintenance budget was cut and nobody assessed the safety impact. Five whys later you are no longer looking at a floor; you are looking at a management decision, which is where prevention actually lives.

Structured methods keep you honest: the "5 whys" for straightforward events, a fishbone diagram when several factors interact (people, equipment, process, environment, management), and a timeline reconstruction for complex incidents. Whatever the method, test the conclusion the same way: if we fix this, does the incident become impossible — or just less likely at this exact spot?

Corrective actions that actually close

An investigation is worth exactly what its actions achieve. Three disciplines matter:

  • Prefer strong controls over weak ones. "Retrain and remind" is the weakest fix on the hierarchy — it asks people to be more careful in an unchanged situation. Elimination, engineering controls and process changes prevent; briefings decorate.
  • Give every action an owner and a date — and track them to closure. Actions without owners are wishes. This is where a digital system earns its keep: overdue actions surface themselves instead of dying in a spreadsheet.
  • Verify, then share. Check the fix works in practice, then brief the learning to everyone who runs the same risk — the other shift, the other site, the subcontractor. An unlearned lesson will reteach itself.

Finally, feed the findings back into the documents that should have prevented the incident: update the risk assessment, amend the method statement, adjust the audit checklist. That loop — incident to investigation to updated controls — is a safety management system actually managing.

Five common mistakes

  1. Investigating only reportable incidents — the near-misses you ignore are rehearsals for the injury you will investigate.
  2. Stopping at the first plausible cause, usually the person nearest the event.
  3. Letting the scene go — starting the paperwork Monday for Friday's incident.
  4. Writing actions no one owns, or closing them on promise rather than verification.
  5. Keeping the learning local — fixing one site while five others run the same risk.

Frequently asked questions

Which incidents should we investigate?

All of them, proportionately. A RIDDOR-reportable injury warrants a full investigation with root cause analysis; a near-miss might need ten minutes and two questions. The mistake is investigating only what you must report — near-misses are your cheapest lessons, because the learning arrives without the injury.

Who should lead an investigation?

Someone with the competence to analyse the work and the independence to follow the evidence — usually a manager or safety professional not directly responsible for the area involved, supported by a supervisor who understands the task. For serious incidents, senior leadership should review and sign off the findings.

How soon should an investigation start?

Evidence decays fastest in the first hours: the scene changes, memories reshape themselves, and CCTV gets overwritten. Secure the scene and capture photographs and first accounts the same day, even if the full analysis follows later.

What's the difference between immediate and root causes?

The immediate cause is what happened at the moment of the event — the wet floor, the missing guard. Root causes are the management-system reasons those conditions existed — the cleaning regime nobody reviewed, the maintenance backlog, the production pressure. Fix only the immediate cause and the incident will repeat somewhere else in a different costume.