# Runbook: adding a new drug–adverse-event causality assessment to OpenPV

Target file: `openpv/causality-assessment.html` (deployed at
`https://vaccinedatanavigator.org/openpv/causality-assessment.html`).

This document is self-contained. You do not need prior context on this
project to follow it — everything you need is either in this file or in the
paths it points to. Follow the steps in order. **Do not skip the validation
step at the end.**

---

## 0. What you're building, in one paragraph

Each entry on this page is one drug↔reaction pair, graded for causality
using four independent kinds of evidence: (1) this site's own FAERS
disproportionality statistics (PRR/ROR/report count — already computed,
you just read them), (2) a qualitative Bradford Hill 9-viewpoint table
(you write this, from outside literature), (3) real ClinicalTrials.gov
trial data if a relevant trial has posted results (you fetch this live),
and (4) a plain-English paragraph tying it together with real citations.
The whole point of the page is to be **honest about uncertainty** — a
"not found" or "contested" result, clearly labeled, is a correct and
valuable output, not a failure.

---

## 1. Pick a candidate pair

Candidates come from this site's own precomputed FAERS signal data, not
from your own knowledge of famous drug scandals (though if you already
know a good one, that's a fine starting point too — just verify it against
real data before writing anything).

Data location: `faers_signal` project's `site/data/drugs/<slug>.json`
(path may differ in your checkout — find it by searching for a file like
`clozapine.json` under a `data/drugs/` directory). Each file has:

```json
{
  "drug": "CLOZAPINE",
  "reactions": [
    {"reaction": "NEUTROPENIA", "PRR": 27.12, "ROR": 31.78, "a": 18898, "signal_flag": true},
    ...
  ]
}
```

`a` = report count (co-reports of this drug + this reaction). `signal_flag`
is already computed (PRR≥2, chi²≥4, a≥3).

**Selection filter** (do not skip the exclusion list — raw top-PRR is
mostly noise):

- `signal_flag == true`
- `a >= 20` (lower if you have a specific reason; below this, PRR is
  statistically unstable)
- `PRR >= 2.0` and both PRR and ROR must be finite (`math.isfinite`) —
  infinite values mean the reaction never co-occurs with any other drug in
  the dataset, which is a degenerate case, not a real "infinite" signal
- Exclude administrative/lab/procedural MedDRA terms: anything matching
  `product`, `use issue`, `off label`, `expired`, `incorrect dose/route`,
  `drug ineffective`, `device`, `normal`, `negative`, `test(s)`,
  `screening`, `therapy$`, `count (increased|decreased)`,
  `enzyme`, `prophylaxis`, `monitoring`. (Full pattern list is in
  `openpv/build_handoff_200.py`, `EXCLUDE_PATTERNS` — reuse it verbatim,
  don't re-derive it.)
- **Even after that filter, most remaining pairs are still not worth
  writing up.** The single biggest remaining noise pattern: a drug's own
  indication reported back as its "reaction" (e.g. a bladder-cancer drug
  showing "bladder cancer" as its top signal). Check: does the reaction
  term match what the drug is approved to treat? If yes, it's probably
  this artifact, not a real signal — see §5 for how to still use pairs
  like this (as a deliberate counter-example), or just skip it.

**What makes a pair worth a full write-up** (vs. skipping it): an
independent regulatory or trial basis exists — an FDA boxed warning, a
REMS program, a market withdrawal, or a named randomized trial — something
you can verify outside this site's own PRR number. If you can't find that
independent basis after a real literature search (§2), don't publish a
"well established" card for it. Either grade it more cautiously (see the
grade vocabulary in §4) or skip it.

---

## 2. Research the pair

Use your web search tool. Search for, in roughly this order:

1. `"<drug>" "<reaction>" FDA boxed warning` — regulatory basis
2. `"<drug>" "<reaction>" mechanism PubMed` — biological plausibility
3. `"<drug>" "<reaction>" cohort study` or `randomized controlled trial` —
   epidemiological strength/consistency
4. If the drug has a REMS program, search for it by name — REMS programs
   are themselves strong causality evidence (a REMS existing means
   regulators already concluded the link is real enough to mandate
   monitoring).

**Capture while you search**: author names, journal, year, and a PMID or
DOI wherever you can get one. A claim like "published cohort studies have
found X" is much weaker, and should be avoided, if you can instead write
"Jick et al. (2000), a 7,195-patient Saskatchewan cohort, found X (PMID
11030769)." Specificity is the whole point — vague citations are barely
better than no citation.

**Do not fabricate a citation.** If your search doesn't turn up a specific
study, write the claim in general terms and don't attach a fake author/
year/PMID to it. A page full of real-but-vague claims is honest; a page
with invented-sounding specific citations is not, and is worse.

---

## 3. Fetch ClinicalTrials.gov trial evidence (optional but high-value)

This is a live API, no key required, confirmed working from this
environment. Script: `openpv/fetch_trial_ae.py` — read it, it's short and
documented inline. The core pattern:

```python
# Search for completed trials with posted results
GET https://clinicaltrials.gov/api/v2/studies
    ?query.intr=<drug name>
    &aggFilters=results:with
    &filter.overallStatus=COMPLETED
    &pageSize=15
    &fields=NCTId,BriefTitle,EnrollmentCount,OverallStatus,StudyType

# For each candidate NCT ID, pull full results:
GET https://clinicaltrials.gov/api/v2/studies/<NCT_ID>
    ?fields=NCTId,BriefTitle,ResultsSection

# The adverse events live at:
#   resultsSection.adverseEventsModule.eventGroups   (arms: id, title, N at risk)
#   resultsSection.adverseEventsModule.seriousEvents  (list of {term, organSystem, stats: [{groupId, numAffected, numAtRisk}]})
#   resultsSection.adverseEventsModule.otherEvents    (same shape)
```

Search `otherEvents` + `seriousEvents` for your reaction term using a
**loose regex** (e.g. `tardive|dyskinesia|extrapyramidal`, not an exact
string match — MedDRA terms vary in wording).

**Known failure mode, check for it every time**: a loose regex can match
the wrong trial entirely. (Example from this page: an isotretinoin search
for "depress" matched "Depressed level of consciousness" in an unrelated
pediatric oncology trial — a sedation term, not psychiatric depression,
and not even an isotretinoin trial.) Before using a match, read the
trial's `BriefTitle` and confirm it's actually about your drug, in a
population where your reaction term means what you think it means.

**Three honest outcomes, all valid, write up whichever one you get:**

- **Found and relevant** → build a `trial-box` (green, see §6). State the
  real numbers (affected/at-risk per arm), link the study.
- **Searched, nothing relevant found** → build a `trial-none` box (gray,
  see §6). Say so plainly, and say *why* if you can reason about it (e.g.
  "this reaction typically takes months to develop, longer than most
  registered trials run" is a legitimate, informative explanation — not a
  dodge).
- **Not applicable in principle** → if the signal isn't something a trial
  AE table could ever show (e.g. a manufacturing contaminant unrelated to
  the drug's own pharmacology, or a lab-monitoring-protocol artifact with
  no real biological hypothesis), say that instead of searching fruitlessly.

---

## 4. Grade the pair

Use exactly one of these five grades (reuse the existing CSS classes —
don't invent new ones without a reason):

| Grade (badge text) | CSS class | Use when |
|---|---|---|
| Causality: well established | `grade-established` | Independent regulatory/trial basis + plausible mechanism + no major contradicting evidence |
| Signal detected, later overturned by randomized trial | `grade-overturned` | A real trial (not just more spontaneous reports) specifically tested this and returned a null/contradicting result |
| Contested / confounded by indication | `grade-overturned` (reused) | Literature genuinely disagrees, or a strong alternative (non-causal) explanation fits the same data equally well |
| Monitoring-protocol / coding artifact, not a drug effect | `grade-artifact` | The "signal" reflects a testing requirement, the drug's own indication, or similar — not a biological effect at all |
| Expected pharmacological effect | *(use `grade-expected` if writing a new one, or just describe it in a table row — see the extended-scan table pattern for low-stakes expected effects)* | A known, labeled, mechanistically direct effect that isn't really in question (e.g. chemo → hair loss) — doesn't need a full card |

Don't grade something "well established" because the PRR is high. PRR
measures reporting disproportionality, not causal evidence — the highest
PRR entry on this whole page (thalidomide↔hCG) is graded as a pure
artifact. Statistical strength and causal grade are independent axes.

---

## 5. Build the Bradford Hill table

Nine rows, always in this order, always these exact labels. This is Sir
Austin Bradford Hill's 1965 framework. **Temporality is the only one Hill
considered strictly necessary** — mark it so.

| Viewpoint | What it's asking |
|---|---|
| Temporality *(necessary)* | Did exposure precede the event, in a plausible window? |
| Strength | How large/precise is the association? |
| Consistency | Reproduced across independent settings/studies? |
| Specificity | Is it tied to this specific drug/dose/outcome, or a diffuse class effect? |
| Biological gradient | Does risk track with dose or duration? |
| Plausibility | Is there a named, credible biological mechanism? |
| Coherence | Does it fit the rest of what's known about the drug and the disease? |
| Experiment | Does removing/adding exposure (trial, dechallenge, policy change) change the risk? |
| Analogy | Is there a precedent — a similar drug/mechanism causing a similar effect? |

Status per row: `met` (green), `partial` (amber), `not established`
(gray), or `n/a` (light gray, only when the viewpoint genuinely doesn't
apply — e.g. "biological gradient" for a pure testing-protocol artifact).
One short clause of justification per row, citing what you found in §2–3.

**Honesty rule, non-negotiable**: this table is *your own qualitative
read* of the literature, not an automated or citation-gated output. Say so
explicitly in a `<p class="hill-note">` under every table. Do not imply
more rigor than you actually did. This site's vaccine evidence dossiers
(`evidence-dossiers/` directory) run an actual 68-node, PubMed-citation-
gated pipeline where every node is backed by a logged, timestamped query —
that is categorically more rigorous than what you are doing here, and the
page must not blur that distinction. If you ever get access to that real
pipeline and run it for a drug pair, that's a different (better) output —
label it differently.

---

## 6. HTML template

Copy this structure exactly (reuses CSS already defined in the page's
`<style>` block — don't redefine classes). Replace everything in `{braces}`.

```html
  <section class="sig-card" style="border-left:4px solid #166534">
    <span class="grade grade-established">{grade badge text}</span>
    <h3 class="mt-2">{Drug} &harr; {Reaction}</h3>
    <div class="sig-stats">
      <span class="sig-stat">n=<b>{a}</b> &middot; PRR <b>{PRR}</b> &middot; ROR <b>{ROR}</b></span>
    </div>
    <p>{2-5 sentence prose: mechanism, regulatory history, real citations with author/year/PMID where you have them}</p>
    <p class="src">{Regulatory basis + citation line, with PMID/DOI links}</p>
    <table class="hill-table"><tr><th>Bradford Hill viewpoint</th><th>Status</th><th>Note</th></tr>
      <tr><td class="vp">Temporality <span class="req">(necessary)</span></td><td><span class="hv hv-met">met</span></td><td>{note}</td></tr>
      <tr><td class="vp">Strength</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Consistency</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Specificity</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Biological gradient</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Plausibility</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Coherence</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Experiment</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
      <tr><td class="vp">Analogy</td><td><span class="hv hv-{status}">{status label}</span></td><td>{note}</td></tr>
    </table>
    <p class="hill-note">This page's own qualitative read, not an automated/citation-gated assessment.</p>

    <!-- IF trial evidence found: -->
    <div class="trial-box">
      <h4>&#x2713; Trial evidence found &mdash; {trial name} (NCT{id}), N={total enrolled}{, X arms if >2}</h4>
      <table class="trial-table"><tr><th>{event type}</th><th>{Arm 1 name} (n={N})</th><!-- ...more arm columns --></tr>
        <tr><td>{term}</td><td>{numAffected}</td><!-- ... --></tr>
      </table>
      <p style="font-size:11px;color:#166534;margin-top:6px">{1-2 sentence interpretation}. <a href="https://clinicaltrials.gov/study/NCT{id}?tab=results" target="_blank" rel="noopener">View on ClinicalTrials.gov &rarr;</a></p>
    </div>
    <!-- IF searched, nothing found: -->
    <div class="trial-none">&#x2717; No ClinicalTrials.gov results-posting trial in the top {N} search hits for {drug} captured {reaction} as a tracked adverse event. {1 sentence on why, if you can reason about it}.</div>
    <!-- IF not applicable: -->
    <div class="trial-none">Not searched in ClinicalTrials.gov. {1 sentence on why a trial AE table could never show this}.</div>
  </section>
```

Grade badge CSS classes available: `grade-established` (green),
`grade-overturned` (amber), `grade-artifact` (dark orange). Hill status
classes: `hv-met` (green), `hv-partial` (amber), `hv-not` (gray), `hv-na`
(light gray).

---

## 7. Validate before you consider it done

Run this after every edit, not just at the end — it's cheap and catches
mistakes immediately:

```python
import re, json
from html.parser import HTMLParser
s = open('openpv/causality-assessment.html', encoding='utf-8').read()
bad = 0
for m in re.finditer(r'<script type="application/ld\+json">', s):
    start = m.end(); end = s.find('</script>', start)
    try: json.loads(s[start:end])
    except Exception as e: bad += 1; print('JSON-LD ERR', e)
c = HTMLParser()
try: c.feed(s)
except Exception as e: print('HTML ERR', e)
print('jsonld bad:', bad, '| file length:', len(s))
```

Both outputs must read "jsonld bad: 0" and show no HTML ERR line. If the
file doesn't parse, you broke a tag somewhere in your edit — find it
before moving on, don't publish broken HTML.

If you added a new pair, also update:
- The page-head summary count (`<div class="meta">...</div>` near the top)
- The `<meta name="description">`, `og:description`, `twitter:description`
  tags if the count or headline topic materially changed

---

## 8. Deploy

This project deploys via FTPS using `deploy_ftps.py` (`upload(env, files,
dry_run=False)`). Deploy only the files you changed — never a full-site
sweep (the local mirror is known to be incomplete relative to the live
site; a full-site deploy can silently delete live pages that aren't in
your local checkout). If the FTP host is unreachable, that's a known
intermittent issue with this host, not a sign your work is wrong — retry
later rather than reworking the content.

---

## 9. Worked example to copy from

Open `openpv/causality-assessment.html` and look at the **metoclopramide
↔ tardive dyskinesia** card (search for `Metoclopramide &harr; Tardive`).
It has all four components (stats, prose with citation, full Hill table,
and a `trial-none` box with a reasoned explanation) and is a good template
length — not the longest card on the page, not the shortest.

---

## 10. The automated per-drug "label screen" is NOT an assessment

Every page under `openpv/drug/<slug>/` (built by `build_drug_pages_v2.py`)
runs each strong FAERS signal through a mechanical screen and prints a badge:

| Badge | Meaning | How it is decided |
|---|---|---|
| Labeled — boxed warning | regulator-recognised, serious | term or synonym appears in the boxed-warning text |
| Labeled adverse reaction | known, expected | exact MedDRA term appears in label warnings / adverse reactions |
| Labeled (related wording) | known, expected | synonym table (`SYN`) or all content words of the term appear in the label |
| Likely indication confounding | reports describe the disease treated | term (or a 6-letter disease stem) appears in the label's Indications text |
| Reporting / coding artifact | not a reaction | matches the administrative regex (`ADMIN`: product issue, off-label use, overdose, "test normal"...) |
| Lab value / disease marker | not screened further | matches the investigation/lab regex (`LABTERM`) and is not in the label |
| Outcome field, not a reaction | death terms | fixed list (`OUTCOME`) |
| Not in label — unassessed | **candidate for a real assessment** | none of the above |
| No label text to screen | screen could not run | no validated FDA label, or label has < 1,500 chars of AE text |

The inputs are `data/enrich/<slug>.json` (openFDA label text, PubChem
description, FAERS seriousness/sex/death counts — produced by
`enrich_drugs.py`), `trial_ae_data/bulk/<slug>.json` (ClinicalTrials.gov
AE modules — `fetch_trials_bulk.py`) and `data/pair_cache.pkl` (2×2 cells,
95% CIs, chi-square — `build_pair_cache.py`).

Rules for an agent working from these pages:

1. "Not in label — unassessed" is where to look for candidate pairs for
   §1 of this runbook. It is a string-matching miss, not a finding. The
   first thing to do with such a pair is read the actual label section —
   the screen misses paraphrases.
2. Never copy a screen badge into a Bradford Hill card as the grade. The
   card grades (§4) are judgments; the badges are lookups.
3. The trial cross-check column shows "x/N on <drug>" only for trials where
   the drug names a study arm or the trial title. A match is useful for
   *Experiment* and *Strength* rows, but check the trial's population and
   dose — registry studies (e.g. the Humira all-patient investigations)
   are single-arm with no comparator.
4. Re-running the Honduras `faers_signal/batch/build_site.py` would
   overwrite these v2 pages with the old 480 KB single-table pages. The
   deployed copy in `ParentsVaccineGUide/openpv/` is authoritative; rebuild
   drug pages with `build_drug_pages_v2.py`, not with `build_site.py`.

### 10a. Extra columns added later on 8 October 2026

* **EBGM / EB05 / IC025** (`compute_bayes_signals.py`, `data/bayes_cache.pkl`): MGPS and BCPNN
  estimates for the same pair. A row marked *PRR-only* (EB05 < 2 and IC025 ≤ 0) is a small-count
  or masking artifact until proven otherwise — do not open a card on a PRR-only pair without
  explaining why the shrinkage estimators disagree.
* **Serious-only & strata** (`fetch_serious_strata.py`, `data/strata/`): PRR recomputed on
  serious reports; sex skew vs the drug's own baseline; consumer-vs-clinician share. A pair that is
  ≥ 70% consumer-reported with a news spike is the classic stimulated-reporting pattern (see the
  Zantac card). A pair whose serious-only PRR collapses is usually a lab/monitoring artifact.
* **Table tools**: every drug page has "Download CSV" and filters (hide artifacts, unlabeled only,
  hide PRR-only). Use "unlabeled only" + "hide PRR-only" to get the real candidate queue for §1.
* **Class and comparison hubs** (`/openpv/class/`, `/openpv/compare/`): a term flagged across most
  of a class supports *Consistency* and *Plausibility* (shared mechanism) — or indicates a shared
  patient population; check which before citing it in a Hill row.
