Real-World Evidence & Pharmacoepidemiology · Section 15.3
~7 min read · The Drug Safety Coach — Global PV Career Course
Key points
Seven RWD source types — strengths, limitations, and primary PV use
| RWD Source | Key Strength | Key Limitation | Primary PV Use |
|---|---|---|---|
| Electronic Health Records (EHRs) | Rich clinical detail, lab-confirmed outcomes, long longitudinal follow-up, captures comorbidities | Variable data quality across institutions; unstructured notes require NLP; smaller population than claims data | PASS study primary data source; signal validation in defined disease populations |
| Insurance & Administrative Claims | Enormous population size (US Sentinel: 100M+ lives); drug exposure well-documented via dispensing | No clinical detail beyond billing codes; no lab values; residual confounding from unmeasured variables | Background incidence estimation; drug utilisation studies; comparative safety analyses |
| Disease-Specific Patient Registries | High data quality; condition-specific clinical depth; can capture outcomes not in routine records | Selection bias (enrolled patients may differ from general population); expensive to maintain | PASS study design; special population safety monitoring (pregnancy, paediatric, rare disease) |
| Wearables & Remote Monitoring | Near-continuous 24/7 physiological monitoring between clinic visits; patient-generated evidence | Clinical interpretation standards still evolving; interoperability with clinical systems limited | Cardiac safety monitoring in oncology; CGM data for hypoglycaemia signals |
| Patient-Reported Outcomes (ePRO) | Direct patient perspective; captures symptoms before formal healthcare contact | Response rate variability; digital literacy required; instrument standardisation needed | Patient-reported safety signals; effectiveness monitoring in RWE studies |
| Pharmacy Dispensing Databases | Actual drug dispensed (not just prescribed); adherence proxies; drug combination data | Dispensing doesn’t confirm consumption; indication not recorded in most systems | Drug utilisation studies; adherence-safety relationship analysis |
| Social Media & Patient Forums | Pre-spontaneous-report signal detection; patient language and perspective; global, multilingual reach | Unvalidated; limited causality inference; sampling bias toward younger, digitally engaged patients | Hypothesis generation; early signal identification |
Full text
Different RWD sources answer genuinely different PV questions, and treating them as interchangeable — reaching for whichever database happens to be available rather than the one that can actually answer the question at hand — is one of the most common and costly methodological errors in pharmacoepidemiology. This lesson works through the seven main RWD source types this module’s comparison table lays out, each with real, specific strengths and real, specific limitations.
Electronic Health Records offer genuinely rich clinical detail — lab-confirmed outcomes, long longitudinal follow-up, visibility into comorbidities and concomitant medications — but that richness comes with variable data quality across different institutions, and unstructured clinical notes require exactly the NLP techniques Module 13 covered to extract usable information at all. Insurance and administrative claims data trades clinical depth for genuinely enormous scale — the US Sentinel system alone covers over 100 million lives — with well-documented drug exposure through dispensing records, but essentially no clinical detail beyond billing codes and no laboratory values whatsoever, meaning any outcome requiring lab confirmation simply can’t be definitively established from claims data alone, regardless of how large the population is.
Disease-specific patient registries offer high data quality and genuine clinical depth for a specific condition, including outcomes routine records might miss entirely, but they carry real selection bias — patients who enrol in a registry may differ systematically from the broader population living with that condition — and they’re expensive to build and maintain. Wearables and remote monitoring devices provide something genuinely new: near-continuous physiological data between clinic visits, capturing events a scheduled appointment would simply miss, though clinical interpretation standards for this kind of continuous data are still actively evolving.
Patient-reported outcomes capture the patient’s own perspective directly, often before any formal healthcare contact happens at all, at the cost of variable response rates and the need for standardised instruments across different studies. Pharmacy dispensing databases confirm what was actually dispensed — a genuinely useful adherence proxy — but dispensing doesn’t confirm the patient actually took the medication, and most dispensing systems don’t record the clinical indication at all. And social media and patient forum data sits closest, in reliability terms, to the spontaneous reporting Module 7 built — genuinely valuable for early hypothesis generation and picking up signals before they reach formal reporting channels, but unvalidated, limited in what it can say about causality, and skewed toward a younger, more digitally engaged population than the medicine’s full user base. The discipline this lesson is really teaching is methodological: define the actual question first, and only then ask which of these seven sources — possibly none of them alone — can genuinely answer it.
Important
Define the study question first, then assess which RWD source can actually answer it — not the reverse. A common, costly methodological error is starting with "we have access to this claims database" and then designing a study around what that source happens to contain, rather than starting with the actual clinical question and only then determining whether any available source — claims, EHR, registry, or otherwise — can genuinely answer it.
Quick check
Test yourself before moving on — no pressure, just click an answer.
1. A researcher needs to confirm cases of drug-induced acute kidney injury, which requires a documented serum creatinine rise. Why would a large insurance claims database alone be an inadequate RWD source for this specific study?