US EMS Protocol Census
Methodology

How the EMS Census is built

1,739 documents223 named agencies27,740 dose entriesas of 2026-09-20

Everything on the census comes from protocol documents agencies themselves published. This page says where each document came from, what was read out of it, what was not, and which numbers are therefore safe to quote. As of 2026-09-20.

1,739documents
223named agencies
27,740dose entries
26,551 of 27,740 (96%)machine-parsed
2,182comparable groups
2026-09-20as of

How a document reaches us

Each document carries an origin, one of six values, derived at build time and never guessed: app (a medic uploaded their own agency's protocol), seed (we collected a published document directly), watch (a scheduled re-check of a page an agency publishes to), wayback (an Internet Archive capture), device (recovered from an app corpus), and unknown. A document minted before origin was recorded reads unknown rather than claiming a provenance it cannot prove.

Nothing here is a copy of an agency's PDF. The census links to the agency's own source where one is known, and hosts no protocol documents.

Document identity is a content hash

A document is identified by the hash of its contents, not by its filename, its URL, or the agency's name for it. Two agencies posting byte-identical files are one document; the same protocol re-posted at a new URL is still the same document; a revision is a new one. That is what makes a version history possible at all, and it is why a re-upload of a file we already hold adds nothing.

Classification and what stays unpublished

A model reads each document to find the agency, the state, and the jurisdiction, and records a confidence of high, medium, low, or user (a value a person supplied). A document whose identity is not settled goes to pending review instead of being listed — the reasons are a name that collided with an existing agency, a merge chain that could not be resolved, a document with no readable agency, and a version whose date could not be placed. Of 1,739 documents, 241 are in that state and do not appear anywhere on the census.

An agency is named only where its protocols are a public record — statewide, regional, county, city, or fire-district. Everything else contributes to counts and distributions as an unnamed source and never gets a page. 1,214 documents are listed with a name; 88 contribute to aggregates only.

What is read out of a document, and what is not

Extraction pulls drug, indication, population, dose, route, repeat interval, and standing-order status. It does not read a protocol's narrative, its flowcharts as flowcharts, or anything a human reader would infer from layout.

Two corpus shapes feed the census, and their limits are different. One carries page numbers; the other carries none, so page not captured is the majority case and is not a defect. The second shape also carries no standing-order flag, so standing is absent rather than false on those entries — the census never renders an absent flag as "not a standing order", because absence of a flag is not evidence of the negative. Pediatric age bands were lost upstream on the second shape entirely: 997 of 27,740 entries (4%) are pediatric entries with no age band, and they are excluded from every distribution and every outlier check on this site. They are still counted, and they still appear on their agency's own page as written.

Parsed, partial, and raw

A dose string is either parsed to a number with a unit and a route, parsed partially (a number and a unit but no route), or kept raw — a string like "per medical control" that carries no number at all. Today: 26,551 of 27,740 (96%) parsed, 477 of 27,740 (2%) partial, 712 of 27,740 (3%) raw.

A raw entry never enters a distribution. It is real data and it is shown as written on its agency's page, and it is counted in the entry totals — but it has no number, so putting it in a median would mean inventing one. Every distribution on this site says how many entries under it carry a machine-readable number.

Units, per-kilogram doses, and ranges

Mass units are canonicalized to milligrams: micrograms and grams convert, so 300 mcg and 0.3 mg are the same value in the same group. Nothing else converts. Millilitres need a concentration the documents do not carry, and units, milliequivalents and joules are not doses of a mass at all — each is its own group and is never folded into another.

A weight-based dose and a flat dose are separate groups for the same reason: 0.01 mg/kg and 1 mg are not the same quantity and never share a median.

A range contributes its low end only. "0.3–0.5 mg" enters a distribution as 0.3. The high end is kept and shown on the agency page, but it never enters a distribution or an outlier check — counting both ends would let one entry vote twice, and picking the high end would overstate every range in the census.

Sources and named agencies are two different counts

A source is one protocol document. A named agency is a source whose agency is a public record and is identified on the census. Every distribution is built one value per source — a document that lists a drug five times gets one vote, its own median, not five — so a verbose document cannot decide a median for everyone.

Both counts appear on every published number, in the form "n=<sources> protocols from <agencies> named agencies / <states> states". Sources are always the larger number, and quoting one for the other is the mistake the two-part citation exists to prevent.

A group publishes a distribution only at 5 or more sources. Below that the census shows the count and nothing else: five documents is thin, and four is not a distribution.

Outlier review

Every night, each group's entries are checked against a reference built from all of that group's parsed entries — one value per source, and the median of those values. An entry more than three times that median, or less than a third of it, is flagged, in groups of at least 5 sources. Flagging is one pass with no feedback: clearing or removing an entry changes nothing about the reference or about any other entry's flag, so the same documents produce the same flags every night.

A flagged entry is suppressed everywhere — every page, every distribution, every count of published rows — until a person reviews it and either clears it (it returns) or rejects it (it stays out). It is counted before it is removed, so the size of the review queue is visible: 1,335 of 27,740 entries (5%) are under review right now, 0 have been reviewed and rejected, and 26,405 are published.

The two distributions are named on purpose. Flags are judged against the reference built before any suppression; the numbers this site publishes are computed after it. A flag is a claim about one entry against its peers, not a claim about the published median.

Accuracy: not yet measured

We do not publish a dose-level accuracy number, because we have not measured one. What exists today is a hand-labelled comparison of drug names, indication text, contraindications and adverse effects — no dose value in it is ever compared against a document. Publishing an extraction-success rate or a model-agreement figure in place of accuracy would be quoting a measurement of a different thing, and it would read as the number this section does not have.

What is measured, and lives on this page because it comes from the build itself: the share of entries that parse to a number (96%), the share under outlier review (5%), and the share of pediatric entries excluded for having no age band (4%).

Every page here carries the same warning, and it is the honest one: this is a training reference compiled from published protocols, not a clinical order. Verify against your own agency's document and your medical director.

Corrections and removal

If a listing is wrong, send the current document's public URL from the agency page's correction form and the listing is rebuilt from it. If an agency wants its listing removed, it comes down the same day — no argument about whether the document is a public record.

Freshness

A document more than 24 months old is flagged as possibly outdated on its agency's page, and a newer version awaiting review is disclosed with its date. The build itself refuses to publish when the number of named agencies drops sharply against the last published build — a collapse in coverage is a broken build, and shipping it would quietly replace the census with a smaller one.

The as-of date on every page is the date the data changed, not the date the page was generated. A page that did not change is not rewritten.

Citation

United States EMS Protocol Census, as of 2026-09-20. https://protoquiz.com/census/ · Data license