Framework library · Risk management and assessment
Failure mode and effects analysis (FMEA)
FMEA works through a process or a product one step at a time and asks how each step can fail, what the failure would do, what causes it and what would catch it before it reaches the customer. Each failure gets three scores, for severity, occurrence and detection. The value is in the discipline of going step by step: failures nobody thought to put on a risk register turn up in the middle of the process.
Use it when
- A process where failure is costly or dangerous is being designed or changed, such as power failover in a data centre or a network migration.
- A failure has happened once and you want to find its neighbours before they happen too.
- You need to decide where to spend on monitoring, testing or redundancy, and want the choice to rest on more than the loudest engineer.
- A customer or regulator asks for evidence that failure modes have been worked through systematically.
Avoid it when
- The risk is strategic or commercial, such as a competitor's move. FMEA is for processes and products whose steps can be mapped. Use a risk register or a pre-mortem.
- A failure has already happened and you need its cause. Use five whys or a fishbone diagram first, then FMEA to check the rest of the process.
- The process has not been mapped. Map it first, or the analysis will skip steps.
- You need to show one serious hazard with many causes and consequences. A bow tie shows threats, barriers and consequences more clearly.
How to run it
Map the process as steps
Six to ten steps from trigger to outcome: utility power fails, the transfer switch operates, generators start, the UPS bridges the gap, cooling restarts. Each row of the analysis belongs to one step.
List the ways each step can fail
A failure mode is what goes wrong at that step (the switch does not transfer), not why it happens (a firmware fault) or what follows (the hall loses power). Keep the three apart.
Score severity from the effect
1 to 10, from no noticeable effect to harm or total loss of service. Severity belongs to the effect, so existing controls do not lower it.
Score occurrence from the cause
How often the cause happens, from failure history where you have it. 1 is remote, 10 is almost certain.
Score detection from the current controls
How likely the controls are to catch the cause or the failure before it reaches the customer. 1 means almost certain to catch it; 10 means no way to catch it.
Act on severity first
The priority column marks a row high if severity is 9 or 10 and occurrence or detection is above 3, or if the risk priority number (RPN) is 200 or more. Give every high row an action, an owner and a date.
Re-score after the actions
Severity usually stays the same; occurrence and detection should fall. If they do not, the action did not work.
Work through it
Answer the questions below, or load the worked example to see a finished one. The drawing updates as you type. Export the result as a PowerPoint deck, a Word document, an Excel workbook, a PDF or plain text.
What you type stays in this browser, so you can close the page and come back to it. It is not sent to Blue Prysm or anyone else, and the exports are made here, on your device. Privacy policy.
Mistakes to avoid
- Treating the RPN as a measure of risk. A severity 10 failure with a low RPN can matter more than a nuisance with a high one. The AIAG and VDA handbook moved away from RPN for this reason.
- Setting an RPN threshold, such as 100, and ignoring everything below it. Teams then nudge scores to land just under the line.
- Mixing up mode, cause and effect, so the same failure appears three times with different scores.
- Scoring detection from what the procedure says rather than what has been tested. An alarm nobody has seen work is not a control.
- Doing it once at design and never again. Update the analysis when the process, the equipment or the failure history changes.
Where it comes from
Developed by the US military as procedure MIL-P-1629 (9 November 1949) and revised as MIL-STD-1629A (1980). It was used by NASA in the Apollo programme and taken up by the automotive industry from the 1970s. The current automotive reference is the AIAG & VDA FMEA Handbook (June 2019), which replaces ranking by risk priority number with action priority tables. Source.
Use it with
Work through it with us
The frameworks here are free to use as they stand. If you would rather work through the question behind this one with us, these are the ways an engagement starts.