Causality & Seriousness Assessment · Section 6.5
~6 min read · The Drug Safety Coach — Global PV Career Course
Key points
Choosing between the two most widely used causality scales
| WHO-UMC | Naranjo | |
|---|---|---|
| Output | Six qualitative categories reached via branching criteria | A numeric score (weighted sum of 10 questions) mapped to 4 categories |
| Where it’s favoured | Company/regulatory PV assessment; explicitly recommended in several national programmes (e.g. India’s PvPI) | Clinical and academic settings valued for reproducibility and simplicity |
| Strength | Captures nuance — six categories allow finer distinctions than a 4-band score | Mechanically simple to apply consistently; less room for interpretation of "which branch" |
| Weakness | Requires more judgment at each branching decision, which can reduce inter-rater agreement | A single number can obscure which specific evidence is actually driving the result |
| Handling of rechallenge | Central to reaching "Certain" specifically | Heavily weighted (+2/−1) but blended into an overall sum |
Full text
Having covered both scales individually, the natural question is which one to actually use — and the honest answer is that both remain in active, legitimate use, solving the same underlying problem with different tradeoffs. WHO-UMC’s branching, criteria-based approach captures more nuance: six categories allow finer distinctions than Naranjo’s four score bands, and the structure explicitly separates "not enough information to judge" (Unassessable/Unclassifiable) from "judged and found unlikely" (Unlikely) — a distinction Naranjo’s numeric approach doesn’t make as cleanly, since a low score can result from either genuinely negative evidence or simply missing information.
Naranjo’s advantage is mechanical consistency. Because every question has a fixed point value, two assessors working from the same source information are more likely to compute the same total than two assessors reaching a WHO-UMC category through a branching decision tree, where each branch point involves its own judgment call about how the evidence should be read. That reproducibility is exactly why Naranjo remains popular in academic and clinical research settings, where consistency across many raters and many studies matters enormously.
The research literature bears this tradeoff out in an unexpected way: studies directly comparing inter-rater agreement have found that neither scale eliminates disagreement between trained assessors — agreement levels vary considerably by study and by therapeutic area, and at least one comparison found WHO-UMC’s reproducibility across judges to be lower than might be assumed, precisely because of the judgment embedded in its branching structure. This isn’t an argument against either scale; it’s a reminder that structured causality assessment reduces inconsistency relative to unstructured clinical opinion, without fully eliminating the underlying judgment calls both scales still require.
In practice, many organisations settle this pragmatically rather than philosophically: WHO-UMC as the company’s primary regulatory-facing causality method, since it’s the scale most global regulators and national programmes expect, with Naranjo used as a secondary reference or in specific clinical contexts where its numeric reproducibility is valued. The genuinely useful skill for a PV professional isn’t having a preferred scale — it’s being able to apply either one’s criteria consistently to a new case and articulate exactly which piece of evidence drove the category or score reached, which is precisely what a medical reviewer or an inspector will ask about.
Quick check
Test yourself before moving on — no pressure, just click an answer.
1. What does WHO-UMC’s category structure distinguish that Naranjo’s numeric score doesn’t as cleanly?