Research

Last updated September 16, 2026.

The published evidence behind Pymander's clinical claims: benchmark protocols, sealed-run results, and the safety architecture, open for anyone to check and rerun.

USMLE Benchmark Working Paper

Pre-registered evaluation of the clinical reasoning layer: 95.3% baseline (n=150) to 98.9% on a sealed held-out set (93/94).

Read →
Head-to-Head Methodology

How the 98.2% USMLE score was measured: 244 official sample questions, three system configurations, two independent runs.

Read →
Safety Architecture Working Paper

Deterministic safety architecture for LLM-based medical consultation: the model proposes, the pipeline disposes. Safety score 1.0 on a 60-case red-flag gate sample.

Read →
USMLE Evaluation

The public eval page: 98.2% across two independent, fully logged runs, and what the number does and does not claim.

Read →

Pymander is not a replacement for a physician and does not provide medical advice, diagnosis, or treatment.

Free AI doctor, 24/7 by textStart a free AI doctor consult