What is Root Cause Analysis?

Root Cause Analysis (RCA) is a structured method for identifying and removing the underlying cause of a failure, incident or deviation, so that the same problem does not happen again. Where a technician replaces a broken pump and gets the line running, RCA asks why the pump failed and why the organisation did not see it coming. In industry, RCA is applied to equipment failures, quality deviations, process safety incidents and, increasingly, to cyber incidents affecting the control system.


🎯 Why is RCA more than fixing a fault?

Fixing a fault treats the symptom; RCA treats the cause. The difference shows up in the numbers: without RCA, the same failures return as unplanned downtime, which directly erodes the availability factor of OEE.

  • Preventing recurrence — a corrective action aimed at the root cause eliminates a whole class of future failures, not just this one
  • Organisational learning — findings are recorded and shared rather than remembered only by the engineer on shift
  • Legal obligation — Seveso establishments must investigate major accidents and near misses, and entities under the Dutch Cybersecurity Act must report the likely cause of a significant incident
  • Prioritisation — RCA shows which measure removes the most risk, instead of trying to fix everything at once

A rule of thumb: RCA pays off for repeat failures, for high-impact events (safety, environment, major production loss) and for every near miss in which a protective layer failed.


🧠 What levels of cause does RCA distinguish?

A good analysis does not stop at the first cause it finds. A common model uses three layers:

Level Question Example for a pump failure
Physical cause What failed technically? Bearing damage from insufficient lubrication
Human cause Which action, or missing action, contributed? A lubrication round was skipped
Latent or organisational cause Which system or decision allowed that error? The lubrication task was missing from the CMMS and its interval was never reviewed after the load increased

Only measures at the latent level prevent recurrence on other assets. “Human error” is therefore never the end point of an RCA, but the starting point for asking why the system allowed the error to happen and go unnoticed.


🔧 Which RCA methods are there?

There is no single RCA method. The right choice depends on the complexity and impact of the event.

Method Origin How it works Best suited for
5 Whys Toyota; described by Taiichi Ohno (1978) Ask “why?” repeatedly until a systemic cause is reached Simple shop-floor problems
Ishikawa (fishbone) Kaoru Ishikawa; popularised in the 1960s Group candidate causes into categories, typically the 6Ms Team brainstorming, quality problems
Fault Tree Analysis (FTA) Bell Labs, 1962; standard IEC 61025 Top-down tree with AND/OR gates, optionally quantitative Complex systems, safety functions
Events & Causal Factors Charting MORT programme of the US Department of Energy (DOE) Timeline of events plus contributing conditions Major incidents with many parties involved
Apollo / RealityCharting Dean Gano, after Three Mile Island Cause-and-effect chains with evidence for every cause Incidents with several interacting causes
TapRooT Mark Paradies and Linda Unger (System Improvements), 1991 Structured root cause tree with a strong human factors focus Process safety, occupational safety, oil and gas
Kepner-Tregoe Problem Analysis Kepner and Tregoe, The Rational Manager (1965) Is / is-not comparison to rule causes out Technical faults with an unclear pattern
Pareto analysis Juran (1941), after Vilfredo Pareto Find the “vital few”: which roughly 20% of causes drive roughly 80% of downtime Prioritising before a deeper RCA

The 6Ms of the Ishikawa diagram are Man, Machine, Method, Material, Measurement and Mother Nature (environment). Fault tree analysis is described in IEC 61025, edition 2.0 from 2006, with a third edition under development. Apollo’s creator Dean Gano stresses that a single root cause rarely exists: several causes usually have to be present at the same time for an event to occur.


🛠️ How do you carry out an RCA step by step?

  1. Secure and collect evidence — capture historian trends, alarm logs, PLC diagnostic buffers, photographs and failed parts before anyone restarts or cleans up
  2. Describe the problem factually — what, where, when and how big, without assigning blame
  3. Assemble a team — operator, maintenance, process engineer and, where digital causes are possible, OT security
  4. Build the timeline and cause chain — with 5 Whys, a fishbone or a fault tree, depending on complexity
  5. Test every cause against evidence — a cause without evidence is still a hypothesis
  6. Define corrective actions per level — physical, human and organisational
  7. Embed and verify effectiveness — record actions, assign an owner and check after a few months whether the problem has really gone, for instance following the PDCA cycle

Worked example: a PLC stops after an unplanned firmware update

A packaging line stops during the night shift and the PLC is in STOP mode. A 5 Whys analysis produces:

Why? Answer
Why did the line stop? The PLC went to STOP with a fault in a function block
Why that fault? After a firmware update, a library function was no longer compatible with the program
Why was the firmware updated? A vendor service engineer applied the update during a remote maintenance session
Why did nobody know? The update bypassed change management and was never tested
Why was that possible? Vendors had standing access without an approval step, and patch management only covered IT

The physical cause is incompatible firmware, the human cause an unplanned action, and the latent cause a missing change process for OT vendors. Only the last corrective action prevents the same stop on the other lines.


🔄 How does RCA relate to maintenance, process safety and OT security?

  • Maintenance — RCA is reactive and FMEA is proactive: FMEA (IEC 60812, 2018 edition) anticipates failure modes, while RCA analyses what actually went wrong and feeds new failure modes back into the FMEA. Reliability Centred Maintenance according to SAE JA1011 (1999) and predictive maintenance use this feedback loop to adjust maintenance strategies. A misleading measurement can also be the cause, so check the calibration of the instruments involved.
  • Process safety — Annex III of the Seveso III Directive requires operators to report and investigate major accidents and near misses, especially where protective measures failed, and to follow up on lessons learnt. In the Netherlands this was implemented through the Brzo 2015 and has been part of the Environment and Planning Act since 2024; see Seveso and Cybersecurity. The CCPS handbook Guidelines for Investigating Process Safety Incidents (third edition, 2019) is the reference work. A failed SIS function or a missed HAZOP scenario is a typical latent cause.
  • OT security — in incident response, the post-incident review or lessons-learned session is the RCA. Under incident reporting under the Dutch Cybersecurity Act, the final report is due within one month of the incident notification and must state the type of threat or the likely root cause. Forensic analysis supplies the evidence for it.

⚠️ What are the common pitfalls?

  • Blame culture — if RCA ends with naming a culprit, people start withholding information and future analyses become worthless
  • Stopping at “human error” — the real question is why the error was possible and why it was not caught
  • Settling on one cause too early — the first plausible explanation is often never tested
  • No evidence — logs are overwritten and parts thrown away before the analysis starts
  • Actions without follow-up — measures stay in a report without an owner or an effectiveness check
  • Over-analysis — a full TapRooT investigation of a sensor that vibrated loose costs more than the failure itself

RCA is embedded in improvement approaches such as Lean and Six Sigma, where it forms the Analyse phase of DMAIC. Software supports the work: RealityCharting, TapRooT software and RCA modules in a CMMS or quality management system.


❓ Frequently asked questions

What is the difference between RCA and FMEA?

Root Cause Analysis is reactive: RCA investigates why something went wrong after a failure or incident. FMEA is proactive and analyses in advance which failure modes can occur and how severe they would be. The two reinforce each other, because RCA findings reveal failure modes the FMEA had missed.

Is 5 Whys enough for a Root Cause Analysis?

The 5 Whys technique is enough for simple problems with one clear chain of cause and effect. For complex incidents with several interacting causes, 5 Whys tends to miss side branches, and a fault tree, Apollo or TapRooT is more suitable. The number five is a guideline: you stop only when you reach a systemic cause you can influence.

How long does a Root Cause Analysis take?

A simple RCA using 5 Whys often takes between an hour and a day. A full investigation after a process safety incident or cyberattack can take weeks, because evidence must be gathered, people interviewed and hypotheses tested. Under the Dutch Cybersecurity Act, the final report on a significant incident is due within one month of the incident notification and must state the type of threat or the likely root cause.

When should you perform a Root Cause Analysis?

A Root Cause Analysis is worthwhile for repeat failures, incidents with safety or environmental impact, major production losses and near misses in which a protective layer failed. Many companies set fixed triggers, such as an RCA for every stoppage longer than an agreed duration.

Why is human error not a root cause?

Human errors are predictable and happen within a system of procedures, training, design and workload. A Root Cause Analysis that ends at human error only leads to “retrain the operator”, while a colleague can make the same mistake tomorrow. The RCA should therefore ask which organisational or design factor made the error possible.

How do you use RCA after an OT cyber incident?

After an OT cyber incident, the Root Cause Analysis is the post-incident review: how did the attacker or unauthorised change get in, why was it not detected sooner, and which control failed. Forensic evidence from logs, network traffic and PLC projects forms the basis. The outcome goes into the final report to the CSIRT and the regulator and into improvements to segmentation, access control and change management.


📌 In summary

Root Cause Analysis looks beyond the symptom to the physical, human and organisational causes of a failure or incident, so that it does not happen again. Choose the method to match the complexity, from 5 Whys to fault trees and TapRooT, back every cause with evidence and treat human error as a starting point rather than a conclusion. For Seveso establishments and entities under the Dutch Cybersecurity Act, investigating the cause of incidents is also a legal obligation.