FREE WEBINAR & DOWNLOAD:  How does your SCTA performance measure up?  Get your SCTA Health Check here.

Incident Investigation for Pharma Manufacturing: What’s on Your DMAIC Tool Belt?

Why do so many pharma deviation investigations end with “retrain and update the SOP”? It usually isn’t because anyone investigated badly. It’s because of what’s on the tool belt.

This blog was co-authored by Dominic Furniss (Human Reliability Associates) and Julie Avery (Chatham Consulting)

In short

  • DMAIC’s five steps are fixed. The tools used inside them are not, and there is no definitive set.
  • That toolset draws on Six Sigma, developed at Motorola in the 1980s, and lean manufacturing, which grew out of the Toyota Production System. Its roots are in efficiency and quality improvement, not in human factors or risk management.
  • As a result, risk is not formally stated as a defined outcome of DMAIC, the human-shaped actions available are often limited to training and procedure updates, and the investigation relies on explicit data rather than on how the work is really done.
  • Adding a small number of tools in the Analyse step, principally Systems Critical Task Analysis, changes what an investigation can find. DMAIC then remains the vehicle for delivering the improvements.
  • Preventive action is where the systemic work gets committed, without overloading the deviation record.

The pattern everyone recognises

A deviation is investigated on time. A root cause is identified. CAPAs are signed off, usually retraining plus a change to the procedure. A few months later something very similar happens on a different line or a different shift.

The investigation was not done badly. In most cases it followed a recognised improvement methodology, such as DMAIC, and it followed it properly. Competent people did competent work. The problem is what they had to work with.

This is not an argument that DMAIC is broken. It is an argument about what is hanging on the DMAIC tool belt. And to be clear about the metaphor, we mean the belt around your waist with the tools on it, not the colour of anyone’s Lean Six Sigma certification.

DMAIC is a process with an undefined toolset

Look at any DMAIC one-pager. Five steps, Define or Diagnose, Measure, Analyse, Improve and Control, with a list of possible activities and possible tools under each. The version we worked from, built up over years of practice, carries around sixty: data driven tools, value stream maps, SIPOC, run charts, control charts, hypothesis testing, measurement system analysis, design of experiments, Kaizen, SMED, 5 Whys, cause and effect diagrams, FMEA and many more.

Sixty is a lot to hold in your head, let alone have hands-on experience of, and that version is not definitive. Another organisation, another trainer, another sector will carry more, fewer or simply different ones.

That is the point. There is no canonical DMAIC toolset. The five steps are a process with an undefined toolset, and what you hang on that process governs what you look for, what you find and what you fix.

It explains something people notice but rarely name. Two investigators can run a faultless DMAIC on the same event and produce completely different levels of insight. The method is systematic in its steps but not in how you carry them out. Nothing in it prevents a superficial answer.

Where the toolset came from

The tools are not a random collection. They come from a tradition: Six Sigma, developed at Motorola in the 1980s, and lean manufacturing, which grew out of the Toyota Production System, brought together under the banners of operational excellence and quality improvement. Kaizen, SMED, value stream mapping and design of experiments all share that parentage, and it is an excellent one for its purpose. If you have a process with a defined output and a measurable gap against a standard, this is the right kit.

But that tradition was not built around human-centred design, human factors or risk management, so those tools were never added. This matters more than it sounds, because a process and a system are not the same thing. A process has a defined output measured against a fixed standard. A system is a dynamic network with connections, feedback loops and interdependencies. An efficiency tradition gathers process measures, which means you can hit the target and pay for it somewhere else in the system that potentially no one is measuring. Efficiency pressure quietly drives drift, and nothing on the belt is watching out for it.

Once you see where DMAIC comes from, the gaps stop looking like oversights. They are exactly the gaps you would predict.

Three predictable consequences

Risk is never formally stated. FMEA does appear in most toolsets, usually under Define and again under Improve, so it is wrong to say risk is absent. But it is one selectable tool among sixty, and nothing requires a risk assessment as an input or an output. You can complete a full DMAIC cycle without ever saying which risks you identified and which you have decided to accept, including risks that have not yet materialised. ICH Q9 expects a quality system to make those decisions explicitly. DMAIC is an improvement tool, not a risk management tool, and using it as one leaves a gap the deviation record will never show.

The only human-shaped tools are the weakest two. Look for people in a DMAIC toolset and you find them in three places: as something to observe while gathering data, as a source of variation under measurement system analysis, and then in Control, as someone to be trained, alongside modifying the SOP.

That explains the pattern in the opening far better than any theory about investigator behaviour. If training and procedure updates are the only human-shaped actions on the belt, those are the CAPAs you will reach for. The two weakest controls in the hierarchy of control are built into the method as its default closing move.

The work as it is really done stays invisible. DMAIC biases towards explicit knowledge: batch records, deviation logs, KPIs, specification limits, the SOP. It puts less emphasis on the knowledge that lives in people’s heads and hands, and on the value of qualitative evidence, which often captures the very activities that enable a dynamic system to succeed.

A site may have six recorded line clearance deviations and hundreds of unrecorded near misses. If the six are all you look at, you may be investigating in the wrong place entirely. And because the data records only failures, a quality system can treat the absence of a detected failure as good performance, without ever knowing what it took to achieve that “good”.

This also sets the framing trap. Traditional root cause analysis asks what should have happened and what did happen. Between them sits the question that explains most events: what normally happens. Miss it, and the problem statement becomes “operators are not following the procedure”, at which point the investigation is over before it starts. What you look for is what you find, and what you find is what you fix.

Tools worth adding to go deeper into human performance

None of this requires sidelining DMAIC. Where we are relying on human performance for important steps, it requires hanging something else on the belt, in a specific place, which is the Analyse step.

A model of human performance. Systems Critical Task Analysis (SCTA), and its retrospective form, Task Analysis Based Incident Evaluation (TABIE), start from the task as it is actually carried out and separate two things investigations usually merge. The failure mode is what went wrong: a step omitted, done wrongly, done too late, done on the wrong item. The Performance Influencing Factors are why it became likely: time pressure, interruptions, similar looking materials, lack of equipment or materials, poor maintenance, procedures that are unclear or not to hand, handover complexity, workload, fatigue. These factors interact. A confusing label is survivable alone, but combine it with a changeover under time pressure and an interruption at the point of selection and it becomes a trap.

Where DMAIC offers thirty ways to analyse a process, this offers one systematic way to stand in the operator’s shoes.

Event mapping that reaches further back. The Accident Sequence and Precursor (ASAP) model maps five phases: the organisational and policy conditions that set the scene, the latent failures lying dormant, the initiating sequence, the consequence, and how the situation was detected and managed afterwards. Today’s deviation is often the product of last year’s decision about resourcing, design or maintenance.

Lines of enquiry instead of a single chain of whys. Rather than drilling down until you reach a person, you develop several connected lines to explore, informed by the task and the PIF model. How many procedures does this task span? Was the procedure usable at that point? What else was the operator handling? How is this normally done? These are pursued through individual and group interviews, walk-throughs and observation, which is qualitative evidence gathered deliberately rather than anecdotally.

Controls ranked against the hierarchy of control. The hierarchy of control comes from occupational health and safety, where it is built into law and standards. The ICH guidelines don’t require it, but some pharma companies already use it, and it brings a useful discipline to choosing CAPAs. Once you can say which factor made which failure mode likely at which step, the options open up beyond training. Can the risk be removed, engineered out through verification at the point of use, or designed out through segregated storage and visually distinct materials, before you fall back on procedures and supervision? You can also better assess the risk and the options to reduce it.

Then DMAIC comes back into its own. If the analysis shows the equipment is generating failures and the procedure cannot keep up, you have two well defined improvement projects, one on the equipment and one on the procedure. That is exactly the work the method is built to deliver. The human factors tools tell you what to improve. DMAIC gets it implemented, measures it and holds it in place.

The CAPA bridge: from correction to prevention

There is a fair objection to all of this. A deviation record exists for a narrow, specific purpose: establishing what happened and whether this batch is safe to release. Load it with detailed systemic analysis and you produce a document that is difficult to use, while delaying a decision that needs making now.

The answer is not to widen the deviation. It is to use the part of the system that already exists for this. Keep the investigation focused on the event. Then commit the systemic work through the preventive half of the CAPA. Corrective action stops it happening again on that line. Preventive action, done well, is not just checking the other lines: it asks what made this possible, and commits to finding out, for example by running an SCTA on the task or on the maintenance regime behind it.

That keeps the investigation clean, and it avoids mixing “what actually happened” with “what could have happened” in the same interview, which muddles recollection and drifts towards blame. And because it is a CAPA commitment, the work is funded, governed under GMP, tracked and closed out rather than remaining a good intention.

A tool kit pictured as a metaphor for the tools requited to effectively complete incident investigations for pharma manufacturing.

Audit your own tool belt

Because there is no standard toolset, the useful exercise is not to judge DMAIC in the abstract. It is to lay out what your investigators actually reach for and ask where the gaps are. Take your last ten deviations where people were involved, and ask:

  1. How many concluded “human error” as the root or direct cause, and how many CAPAs relied on training or a procedure update?
  2. Did any explain why the action made sense to the person at the time?
  3. Did any establish what normally happens, as opposed to what should have happened?
  4. Did the tools used identify the conditions that made the failure likely, and account for how those conditions combine?
  5. Were controls ranked against the hierarchy of control before a CAPA was chosen?
  6. Did any state which risks were identified and which are being accepted?
  7. Did any surface the latent conditions below the waterline and lead to systemic change rather than a local fix? (See our iceberg blog, How deep does your investigation go?)

Severity grading is not a good trigger for going deeper. Plenty of majors are procedurally routine, while a cluster of repeat minors is often where the real story sits. The better questions are whether human performance was central, whether something like this has happened before, and what this means for risks in the future. Can our system hear the signals from near misses, and see the adaptations people make to get a process completed and a batch over the line?

In summary

DMAIC is a strong process with an undefined toolset, and the toolset most organisations inherited was built for efficiency rather than for human performance or risk. That is why investigations involving people so often end where they do, and it is not a reflection on the people running them.

The fix is not a new methodology. It is a focused addition of a few meaningful tools on the belt, used routinely in the Analyse step, and a preventive action commitment that lets the systemic work happen without overloading the deviation record.

If you want somewhere to start this week, take the last ten deviations involving people and run the seven questions above. The answers usually make the case on their own.

We would be glad to compare notes with anyone wrestling with this.


Want to go further?

If you’d like an introduction to this way of thinking, our blog Beyond Human Error explains why “human error” is where a good investigation starts rather than where it ends, and what to look for instead.

We’re also running a free one-hour webinar of the same name, presented by Julie and Dominic, which runs twice to suit different time zones. It covers how to move investigations past retraining and procedure updates, with practical examples from pharmaceutical manufacturing. Sign up to the newsletter to hear when registration opens.

And if you’d like to talk about how your own investigation toolset measures up, we’d be glad to hear from you.

How this blog was developed

This piece grew out of a series of conversations between Dominic and Julie, drawing on Julie’s years of experience running DMAIC and improvement projects in pharmaceutical manufacturing across different organisations, and Dominic’s work in human factors and task analysis. We used Claude, an AI assistant, to help summarise those discussions, organise the arguments and test them against each other, and to draft sections for us to work on. The ideas, examples and judgements are ours, and the piece went through several rounds of review and editing by both authors before publication.

Professionals like you subscribe to our Human factors thinking.

Want specialist HRA insights direct to your inbox? Get the handbook, email series, and monthly insights from our industry-leading thinkers.