The CAPA That Was Really a Design Decision: Tracing Problems Back to Where They Start
Many of the CAPAs that keep coming back aren’t quality problems. They’re design decisions that arrived late and got a quality badge. Here’s how to spot them, why the org chart keeps hiding them, and how to fix them where they start.
The CAPA that keeps coming back
Every quality team has one. A CAPA that was opened, investigated, closed on schedule, and then quietly reopened a few months later under a slightly different description. Same part. Same failure. New number.
When I trace those CAPAs back, the details change but the shape rarely does. A dimension drifts out of tolerance on the line. Operations quarantines the lot. Quality opens a nonconformance, then a CAPA. The investigation finds that operators weren’t following the work instruction closely enough. The corrective action is retraining plus an added inspection step. The CAPA closes. Everyone moves on.
And the root cause stays exactly where it was: upstream, in a drawing. The tolerance was tighter than the process could reliably hold, and nobody challenged it in design review because, at the time, it looked like a reasonable engineering choice.
That is the pattern this article is about. It cuts across all three areas of my work — design and manufacturing engineering, quality systems, and program management — because the problem doesn’t belong to any one of them. That is exactly why it survives.
“Retrain and add inspection” is rarely a fix. It’s a receipt for a design decision you’re still paying for.
How a design decision becomes a quality problem
Design decisions don’t fail where they’re made. They fail later, somewhere else, under someone else’s name. A typical chain looks like this:
A tolerance, material, or feature goes unchallenged in design review.
The production process can’t hold it consistently, so it drifts.
The drift becomes a nonconformance: scrap, rework, holds.
The nonconformance becomes a CAPA, usually closed with retraining and inspection.
The repeat CAPA becomes visible to an auditor or investigator.
Figure 1 — The same root cause changes owners at every hop, and gets more expensive and more visible each time.
At each step the problem gets a new owner, a new budget, and a higher cost. A tolerance that would have taken one conversation to loosen during design now costs scrap, labor, inspection time, and engineering hours on every lot. If it reaches a field issue or an inspection finding, the cost stops being measured in dollars alone.
The frustrating part is that the fix is often cheapest exactly where nobody is looking anymore: at the original design decision.
Why the org chart hides root causes
None of this happens because people are careless. It happens because each function is doing its job well against the metrics it was given.
Design teams are measured on milestones and launch dates. Once a design is transferred, their attention moves to the next program. Operations is measured on yield, output, and cost. Quality is often measured on how fast CAPAs close and how few are overdue. Each of those measures is reasonable on its own. Together they create a system where the fastest path to “closed” is a local fix, and the slowest path is the one that crosses a department boundary.
A CAPA that concludes “change the drawing” needs design engineering time, a design change under change control, possibly re-verification, and an update to the risk file. A CAPA that concludes “retrain and add inspection” can be closed by the people already in the room. Under deadline pressure, the second option wins almost every time.
Same problem. Three departments. Three budgets. No single owner. The org chart is where root causes go to hide.
Why this matters more in 2026
On February 2, 2026, FDA’s Quality Management System Regulation (QMSR) took effect. It amends 21 CFR Part 820 to incorporate ISO 13485:2016 by reference, and FDA retired the Quality System Inspection Technique (QSIT) in favor of a new inspection approach aligned with the QMSR.
ISO 13485 expects a risk-based approach to controlling QMS processes, and its corrective action clause asks organizations to determine causes, take action proportionate to the effects of the nonconformity, and review whether that action was effective. ISO 14971 adds its own expectation: information from production and post-production should feed back into risk management.
My practical takeaway: be ready to show how a recurring CAPA connects to your design history and your risk file, not just to a training record. A repeat CAPA closed with retraining is hard to defend as effective. One that traces back to a design change and an updated risk analysis tells a much stronger story, because it shows the quality system working as a system.
Three questions before you close any CAPA
The most useful habit I know for breaking this pattern is a short gate at CAPA closure. It takes a few minutes and it changes where fixes land.
Figure 2 — A three-question gate that sends root causes back to the phase that created them.
Where in the lifecycle could this have been prevented?
Not where it was found. Where it could have been stopped. Tag every CAPA with its phase of origin: design, design transfer, or production. Most organizations already record where a problem was detected. Very few record where it was born, and that is the field that tells you where to invest.
Does the fix change that phase, or just add a check downstream?
If the problem originated in design and the fix lives entirely in production, be honest about what you’ve done. You’ve built a detection step around a cause that still exists. Sometimes that is the right short-term containment. It is rarely the right permanent corrective action.
Did the risk file learn anything?
If a real failure mode showed up in production and the risk analysis didn’t change, the system didn’t either. Maybe the occurrence estimate was optimistic. Maybe a hazardous situation wasn’t considered. Either way, the risk file is supposed to be a living record, and a CAPA is one of the best sources of evidence it will ever get.
Climb the ladder of fixes
Not all corrective actions are equal. A useful way to judge a proposed fix is to ask how much it depends on people remembering to do the right thing.
Figure 3 — The higher the rung, the more durable the fix. Most recurring CAPAs were closed on the bottom two.
At the bottom, retraining asks people to do better. One step up, inspection catches the defect after it has already been made. Procedure changes clarify the method. Process changes, such as fixtures, error-proofing, and validated parameters, make the wrong outcome harder to produce. At the top, a design change removes the failure mode altogether.
This ordering will feel familiar to anyone who works with ISO 14971, which lists inherent safety by design first among risk control options, ahead of protective measures and information for safety. The same logic applies to corrective actions. Climb as high as the root cause allows. When you can’t climb all the way yet, say so in the record, contain the issue, and schedule the higher fix rather than declaring victory at the bottom.
Inspection catches a problem. It doesn’t remove it.
A practical routine
None of this requires a new system. It requires a few small changes to how CAPAs are opened, reviewed, and measured:
Add a “phase of origin” field to every CAPA and nonconformance, separate from where the issue was detected.
Put a design engineer on the CAPA review board whenever the suspected origin is design or design transfer.
Make “risk file reviewed and updated, or rationale documented” a closure criterion, not an afterthought.
Review the distribution of origins quarterly. If a large share trace back to design, that is a design review and DFM problem, not a training problem.
Feed design-origin CAPAs into your DFM checklists and design review lessons learned so the next program doesn’t repeat them.
Measure recurrence alongside closure time. A fast close that comes back is slower than a careful one that doesn’t.
Signals you’re treating symptoms
A few patterns show up again and again when CAPAs are being closed at the wrong level:
Retraining appears as the corrective action on more than one CAPA for the same part or process.
Inspection steps keep accumulating on a line, and nobody can remember which CAPA added which one.
Design engineering is never invited to CAPA reviews, even when a drawing is clearly involved.
Effectiveness checks confirm that training was completed rather than that the problem stopped happening.
The risk file hasn’t changed in years, while the nonconformance log keeps growing.
What good looks like
In a healthy quality system, the CAPA log gets shorter over time instead of just staying current. Repeat issues are rare, and when they happen they trigger a deeper look rather than another round of retraining. Design engineers see production data and treat it as input to the next design. The risk file reflects what the product actually does in the field. And when an investigator asks how a CAPA was resolved, the answer runs from the symptom all the way back to the decision that caused it.
That is what “Medical Device Excellence” means in practice for quality: not closing CAPAs faster, but closing them at the right level, so they stay closed.
The cheapest place to fix a problem is the place it started. Trace it back, and fix it there.
Work with MEDEVEX - If your CAPA log keeps repeating itself, or your corrective actions keep landing on retraining and inspection, I can help. I work with medical device teams to trace problems back to where they start and fix them there, across design and manufacturing engineering, quality systems and compliance, and program execution.