Head-to-head methodology: how the 98.2% USMLE score was measured.

Pymander Health's clinical AI scored 98.2% on the USMLE across two independent, fully logged runs: 479 correct responses of 488, at $9.30 per run. This page publishes exactly how that number was produced: the corpus, the three configurations compared head to head, the run protocol, and the results. The companion working paper covers the pre-registered reasoning-layer experiment (95.3% baseline to 98.9% sealed held-out); this page covers the public benchmark comparison behind the headline number.

Loading the paper… Open the PDF.

Corpus design

The evaluation corpus is 244 text-only questions drawn from the official published USMLE sample examinations, spanning Step 1, Step 2 CK, and Step 3. Items that depend on an image were excluded, because the system under test is text-native. The corpus is fixed: questions are drawn from the published sample set, not authored or edited by Pymander, and the official answer key is the grading source.

The USMLE is the three-step examination required for medical licensure in the United States: Step 1 on the basic science of medicine, Step 2 CK on clinical knowledge and patient management, Step 3 on practising unsupervised. It is the standard benchmark for medical AI because its questions demand reasoning rather than recall, and because an official answer key exists - which makes a claimed score checkable instead of merely asserted.

How cases were run

Each configuration answered the full 244-question set once per run. The full set was administered twice as independent runs, yielding 488 scored responses per configuration. Every run was fully logged. A run is independent: no state, answers, or grading feedback carries between runs.

Grading is agreement with the official USMLE answer key. There is no partial credit and no adjudication step that can move the headline: an answer matches the key or it does not.

The three configurations, head to head

ConfigurationCostScore
Two-tier cascade, cost-routed$1.7091.4%
Single model with retrieval (corpus + live literature)$23.3697.1%
Single model, no retrieval (shipped)$9.30 per run98.2%

Table 1. Accuracy and cost by configuration, measured on the same 244-item set across two independent runs.

The head-to-head result is why the shipped system is a single model without retrieval: the cheapest configuration lost 6.8 points to save $7.60, and the retrieval configuration paid 2.5x the cost to score lower. The configuration that ships is the one that scored highest on the fixed key.

What this does not mean

It does not mean the system can practise medicine. The USMLE is a closed problem: the relevant facts are contained in the vignette, the differential is bounded, and one of the listed options is correct. Performance on a closed-form examination does not, without further evidence, extend to open-ended clinical care. Safety behavior in live conversation is evaluated separately, on the deterministic safety architecture's red-flag suite.

Data availability

Run transcripts are retained in full and kept out of the public repository by policy. For methodology questions or access to the underlying run data: hello@pymander.app.

Free AI doctor, 24/7 by textStart a free consult