How we build this site
Every number on Vaccine Data Navigator and OpenPV is traceable to a public dataset and a documented transformation. This page is the single place that describes those sources, the statistics we compute, how pages are reviewed, and what the data cannot tell you.
1. Data sources
| Dataset | Used for | How we access it | Refresh |
|---|---|---|---|
| VAERS (CDC/FDA) | Vaccine pages, lot dashboard, flagged pairs | Public CSV downloads; lot-level aggregation in the lot explorer | Extract date shown on each page's data line |
| FAERS via openFDA | OpenPV drug, reaction, class and comparison pages | count= aggregation over patient.drug.openfda.generic_name.exact × patient.reaction.reactionmeddrapt.exact; serious-only, sex and reporter-type strata; 20.7 M reports | Build of 14 September 2026; quarterly target |
| FDA drug labels (openFDA) | Drug class, brands, indication, boxed warning, label screen | Exact generic-name query, prescription label preferred; name validated against the returned product; multi-ingredient (homeopathic) matches rejected | Fetched 8 October 2026 |
| ClinicalTrials.gov API v2 | Trial cross-check on drug pages and assessment cards | Largest completed trials with posted results naming the drug in a study arm; per-arm numAffected/numAtRisk | Fetched 8 October 2026 (top 800 drugs) |
| PubChem | "What it is" sentence (fallback: FDA label Description) | PUG-REST compound description (NCIt, DrugBank, ChEBI) | As available |
| National immunisation schedules | Schedules by country (31 nations) | Hand-compiled from ministry of health / NITAG publications; source line on each page | Data as of date in the page header |
| VICP / CICP tables | Compensation pages | Federal Register versions, versioned | Per version |
| Canada Vigilance, JADER, EudraVigilance | International surveillance sections | Public extracts where available; "Not searched" is stated when a section was not run | Per section |
2. Statistics we compute
All disproportionality measures start from the 2×2 table for one drug and one reaction: a reports with both, b with the drug and another reaction, c with the reaction and another drug, d neither.
| Measure | Definition | Signal convention | Why we show it |
|---|---|---|---|
| PRR | (a/(a+b)) / (c/(c+d)), with 95% CI on log scale | PRR ≥ 2, χ² ≥ 4 (Yates), a ≥ 3 (Evans 2001) | Standard, transparent, comparable with the literature |
| ROR | ad / bc, with 95% CI | lower CI > 1 | Odds-based companion to PRR |
| EBGM / EB05 | MGPS empirical-Bayes geometric mean of the observed/expected ratio; two-component gamma prior fitted by maximum marginal likelihood on all 1.3 M pairs | EB05 ≥ 2 | Shrinks small-count inflation; marks "PRR-only" signals |
| IC / IC025 | BCPNN information component log₂((a+0.5)/(E+0.5)) with closed-form lower bound | IC025 > 0 | The WHO-UMC method; third independent check |
| Serious-only PRR | PRR recomputed on reports flagged serious | Shown, not thresholded | Removes many lab-value and coding artifacts |
| Sex / reporter skew | Share of female reports vs the drug's baseline; share from consumers vs clinicians (n ≥ 20) | Flag at ±20 points / ≥ 70% consumer | Exposes population and stimulated-reporting effects |
Across the full dataset, 25% of pairs that pass the PRR threshold do not reach EB05 ≥ 2. We publish both so the reader can see where the methods disagree.
3. The label screen (what the badges mean)
On every OpenPV drug page each strong signal gets one badge. The badge is a mechanical text lookup, not a judgement: does the reaction term (or a clinical synonym, with UK/US spellings normalised) appear in the FDA label's boxed warning, warnings or adverse-reactions text (labeled), in the indications text (likely indication confounding), is it an administrative or coding term (artifact), a laboratory or disease-marker term (lab value), a death outcome field, or none of these (not in label — unassessed)? The unassessed tier is a queue for proper assessment, not a finding. String matching misses paraphrases; a strong PRR can still be the disease itself, co-medication, or publicity.
4. Causality assessments
The assessment page holds hand-written cards for 30 drug–reaction pairs. Each applies the nine Bradford Hill viewpoints with a status per viewpoint (met / partial / not established / n/a), cites the regulatory or trial basis, and shows ClinicalTrials.gov arm-level counts where a trial exists. Grades are: well established; signal later overturned by trial; contested / confounded by indication; monitoring or coding artifact; expected pharmacological effect. Statistical strength and causal grade are independent axes — the largest PRR on that page is graded a pure artifact. The vaccine evidence dossiers use a separate, PubMed-gated pipeline described on About.
5. Schedules by country
Each of the 31 national schedules is compiled from the country's official immunisation programme document and expressed in a common vocabulary (All, Risk, Shared, None per cell) so that countries can be compared. The United States column carries a note where the August 2026 recommendation text changed. Age labels are the country's own; where a country gives a range we show the range. Every country page lists its source and the data-as-of date.
6. Review and corrections
Pages carry a reviewer line (Matthew Thomas J Halma, Founder, Open Source Medicine Foundation) and a review date. Statements of causality on vaccine pages follow the published DARE-SAFE framework (DOI 10.3390/pharma4020007). Anyone can report an error in a source line; accepted corrections are logged in the changelog with the date the dataset or wording moved. Field definitions are in the data dictionary.
7. What the numbers cannot tell you
Spontaneous reports are not rates. VAERS and FAERS have no denominator (how many people took the product) and no comparator arm. A report does not establish that a vaccine or drug caused the event. Reporting is shaped by time on market, media attention, litigation, indication severity and who files the report. Disproportionality says a reaction is reported more often than expected with a product — it is a prompt to look, not a conclusion.
Trial cross-checks have real denominators but different populations, doses and follow-up than the reporting population, and only a minority of trials post results. Label text lags evidence. Nothing here is medical advice; decisions belong with the patient and clinician.
8. Reuse and citation
Content is published by the Open Source Medicine Foundation for reuse with attribution. Each page has a "Cite this page" block with BibTeX. Bulk data: the handoff lists on the OpenPV assessment page (handoff_top1000_signals.csv) and per-drug tables (CSV download on each drug page). Build scripts and the causality runbook are documented at /openpv/CAUSALITY_ASSESSMENT_RUNBOOK.md.