Real-World Evidence & Pharmacoepidemiology · Section 15.4
~7 min read · The Drug Safety Coach — Global PV Career Course
Key points
Full text
Choosing a pharmacoepidemiology study design is exactly as consequential as choosing the right disproportionality metric was in Module 7 — the wrong design doesn’t just produce a less efficient study, it can produce a result that looks methodologically sound while systematically misleading whoever relies on it. This lesson and the next work through six designs; this one covers the two most foundational.
A cohort study defines an exposed cohort — patients actually taking the drug in question — and an unexposed (or alternatively-exposed) comparison cohort, then follows both groups forward in time, comparing the incidence rate of the adverse event between them. This design is genuinely well-suited to absolute risk estimation and relative risk calculation with real denominator data, and it supports long follow-up for delayed-onset events. Its dominant validity threat is confounding by indication — a drug prescribed specifically to sicker patients will show an apparent elevated risk for almost any outcome, purely because the patients receiving it were already at higher baseline risk, not because the drug itself caused the excess events. Propensity score methods, covered in depth in Lesson 15.6, are the standard tool for controlling this. A representative worked example: a PASS study for hepatotoxicity might compare the incidence of liver injury in patients taking Drug X against age-and-sex-matched patients on standard-of-care therapy, using a claims database, calculating a hazard ratio and an absolute incidence difference.
Case-control studies invert the logic entirely, and they exist specifically to solve cohort design’s biggest practical weakness: detecting rare outcomes. Rather than following a large cohort forward and waiting for enough events to accumulate — impractical when the outcome is genuinely rare — a case-control study identifies cases (patients who already experienced the adverse event) and controls (patients who didn’t), then looks backward to compare drug exposure rates between the two groups. This is a dramatically more efficient use of data for uncommon events, at the cost of recall bias (when exposure data is collected retrospectively from patients themselves) and the real methodological challenge of selecting appropriate controls. A representative example: investigating a possible signal for Stevens-Johnson Syndrome by identifying SJS cases in a dermatology database, matching controls, and comparing the rate of suspect drug exposure between the two groups to calculate an odds ratio.
Notice that neither design is simply "better" than the other in the abstract — a cohort study can’t efficiently study a genuinely rare outcome, and a case-control study structurally can’t produce a true incidence rate the way a cohort study can, because it starts from cases and controls rather than a defined population followed over time. Matching the design to the actual question — common outcome needing incidence rates versus rare outcome needing efficient detection — is the first, most consequential decision in any pharmacoepidemiology study.
Important
Choosing the wrong study design for a given question isn’t just an inefficiency — it can produce systematically misleading results capable of influencing a safety decision in the wrong direction. A rare-outcome study using cohort design (too few events, no matter how large the cohort) or a study needing incidence rates using case-control design (which structurally cannot produce them) fails for a methodological reason no amount of careful execution can fix.
Quick check
Test yourself before moving on — no pressure, just click an answer.
1. Why is a case-control design generally preferred over a cohort design when investigating a genuinely rare adverse event?
2. What is "confounding by indication," and why is it the dominant threat to cohort study validity?