How we build this site

Every number on Vaccine Data Navigator and OpenPV is traceable to a public dataset and a documented transformation. This page is the single place that describes those sources, the statistics we compute, how pages are reviewed, and what the data cannot tell you.

1. Data sources

DatasetUsed forHow we access itRefresh
VAERS (CDC/FDA)Vaccine pages, lot dashboard, flagged pairsPublic CSV downloads; lot-level aggregation in the lot explorerExtract date shown on each page's data line
FAERS via openFDAOpenPV drug, reaction, class and comparison pagescount= aggregation over patient.drug.openfda.generic_name.exact × patient.reaction.reactionmeddrapt.exact; serious-only, sex and reporter-type strata; 20.7 M reportsBuild of 14 September 2026; quarterly target
FDA drug labels (openFDA)Drug class, brands, indication, boxed warning, label screenExact generic-name query, prescription label preferred; name validated against the returned product; multi-ingredient (homeopathic) matches rejectedFetched 8 October 2026
ClinicalTrials.gov API v2Trial cross-check on drug pages and assessment cardsLargest completed trials with posted results naming the drug in a study arm; per-arm numAffected/numAtRiskFetched 8 October 2026 (top 800 drugs)
PubChem"What it is" sentence (fallback: FDA label Description)PUG-REST compound description (NCIt, DrugBank, ChEBI)As available
National immunisation schedulesSchedules by country (31 nations)Hand-compiled from ministry of health / NITAG publications; source line on each pageData as of date in the page header
VICP / CICP tablesCompensation pagesFederal Register versions, versionedPer version
Canada Vigilance, JADER, EudraVigilanceInternational surveillance sectionsPublic extracts where available; "Not searched" is stated when a section was not runPer section

2. Statistics we compute

All disproportionality measures start from the 2×2 table for one drug and one reaction: a reports with both, b with the drug and another reaction, c with the reaction and another drug, d neither.

MeasureDefinitionSignal conventionWhy we show it
PRR(a/(a+b)) / (c/(c+d)), with 95% CI on log scalePRR ≥ 2, χ² ≥ 4 (Yates), a ≥ 3 (Evans 2001)Standard, transparent, comparable with the literature
RORad / bc, with 95% CIlower CI > 1Odds-based companion to PRR
EBGM / EB05MGPS empirical-Bayes geometric mean of the observed/expected ratio; two-component gamma prior fitted by maximum marginal likelihood on all 1.3 M pairsEB05 ≥ 2Shrinks small-count inflation; marks "PRR-only" signals
IC / IC025BCPNN information component log₂((a+0.5)/(E+0.5)) with closed-form lower boundIC025 > 0The WHO-UMC method; third independent check
Serious-only PRRPRR recomputed on reports flagged seriousShown, not thresholdedRemoves many lab-value and coding artifacts
Sex / reporter skewShare of female reports vs the drug's baseline; share from consumers vs clinicians (n ≥ 20)Flag at ±20 points / ≥ 70% consumerExposes population and stimulated-reporting effects

Across the full dataset, 25% of pairs that pass the PRR threshold do not reach EB05 ≥ 2. We publish both so the reader can see where the methods disagree.

3. The label screen (what the badges mean)

On every OpenPV drug page each strong signal gets one badge. The badge is a mechanical text lookup, not a judgement: does the reaction term (or a clinical synonym, with UK/US spellings normalised) appear in the FDA label's boxed warning, warnings or adverse-reactions text (labeled), in the indications text (likely indication confounding), is it an administrative or coding term (artifact), a laboratory or disease-marker term (lab value), a death outcome field, or none of these (not in label — unassessed)? The unassessed tier is a queue for proper assessment, not a finding. String matching misses paraphrases; a strong PRR can still be the disease itself, co-medication, or publicity.

4. Causality assessments

The assessment page holds hand-written cards for 30 drug–reaction pairs. Each applies the nine Bradford Hill viewpoints with a status per viewpoint (met / partial / not established / n/a), cites the regulatory or trial basis, and shows ClinicalTrials.gov arm-level counts where a trial exists. Grades are: well established; signal later overturned by trial; contested / confounded by indication; monitoring or coding artifact; expected pharmacological effect. Statistical strength and causal grade are independent axes — the largest PRR on that page is graded a pure artifact. The vaccine evidence dossiers use a separate, PubMed-gated pipeline described on About.

5. Schedules by country

Each of the 31 national schedules is compiled from the country's official immunisation programme document and expressed in a common vocabulary (All, Risk, Shared, None per cell) so that countries can be compared. The United States column carries a note where the August 2026 recommendation text changed. Age labels are the country's own; where a country gives a range we show the range. Every country page lists its source and the data-as-of date.

6. Review and corrections

Pages carry a reviewer line (Matthew Thomas J Halma, Founder, Open Source Medicine Foundation) and a review date. Statements of causality on vaccine pages follow the published DARE-SAFE framework (DOI 10.3390/pharma4020007). Anyone can report an error in a source line; accepted corrections are logged in the changelog with the date the dataset or wording moved. Field definitions are in the data dictionary.

7. What the numbers cannot tell you

Spontaneous reports are not rates. VAERS and FAERS have no denominator (how many people took the product) and no comparator arm. A report does not establish that a vaccine or drug caused the event. Reporting is shaped by time on market, media attention, litigation, indication severity and who files the report. Disproportionality says a reaction is reported more often than expected with a product — it is a prompt to look, not a conclusion.

Trial cross-checks have real denominators but different populations, doses and follow-up than the reporting population, and only a minority of trials post results. Label text lags evidence. Nothing here is medical advice; decisions belong with the patient and clinician.

8. Reuse and citation

Content is published by the Open Source Medicine Foundation for reuse with attribution. Each page has a "Cite this page" block with BibTeX. Bulk data: the handoff lists on the OpenPV assessment page (handoff_top1000_signals.csv) and per-drug tables (CSV download on each drug page). Build scripts and the causality runbook are documented at /openpv/CAUSALITY_ASSESSMENT_RUNBOOK.md.