Original 2020 paper; USMLE four-option evaluation
MedQA · USMLE four-option ↗
A medical answer-selection score with a specific denominator.
MedQA & MedMCQA · independent analysis
Independent analysis of MedQA and MedMCQA: verified dataset splits, historical retrieval experiments, answer-selection metrics and an interactive accuracy interval calculator.
Training, development and test partitions from the original papers. Each bar shows the composition of its own question bank.
12,723 questions · Original 2020 paper; USMLE four-option evaluation
Published split counts; these are questions, not patients. They sum to 12,723.
193,155 questions · CHIL 2022 paper; exact split counts from §3/Table 2
The original README swaps validation/test table headings; use exact paper statistics and record the actual downloaded files.
Bar widths describe dataset composition, not model performance. The four-option variant, split and retrieval condition remain part of every score.
Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.
Paper-reported results / selected rows
September 2020 paper, four-option USMLE test; original retrieval/reader protocols.
Paper-reported historical results. No uncertainty intervals are provided in this table; do not present these as current frontier performance.
Source: Table 8 [1]
Paper-reported results / selected rows
CHIL 2022, PubMedBERT reader, original fine-tuning and test evaluation.
Paper-reported historical measurements. Table proportions are multiplied by 100 to display percentage accuracy; no intervals are supplied.
Source: Table 4 [3]
An original analytical tool
Explore a Wilson 95% interval for a hypothetical answer-selection run. The initial denominator is the 1,273-question MedQA USMLE test set; 1,000 successes is an invented illustration, not a model result.
Illustrative counts / edit to explore
Starting success counts are hypothetical. This does not reproduce a model run or account for repeated patients, clustered questions, grader error, or multiple comparisons.
Wilson 95% interval
76.22% – 80.72%
This interval describes sampling uncertainty under the stated binomial assumptions. It is not a test of clinical usefulness or a paired model comparison.
This binomial illustration assumes independent question outcomes. It does not account for contamination, question dependence or paired model comparisons, and it does not estimate clinical benefit. [1][5]
Original 2020 paper; USMLE four-option evaluation
A medical answer-selection score with a specific denominator.
CHIL 2022 paper; exact split counts from §3/Table 2
The split and the retrieval source belong beside the score.
MedQA and MedMCQA evaluate medical examination answer selection. Their scores depend on which questions, options and retrieval resources a system receives. This publication makes those conditions visible: exact denominators, source-version discrepancies, historical experiments and the limits of an accuracy claim. Start with the benchmark dossiers, inspect the illustrative interval calculator, then use the comparison worksheet to document a run. Arcophos provides the analysis and tools; the original research teams created the benchmarks.
Match split, options, retrieval and output parsing before interpreting medical examination accuracy.
Resolve the original source disagreement using exact statistics, file identity and a transparent evaluation receipt.
Use the MedQA denominator and MedMCQA context experiment to separate score resolution, sampling uncertainty and pipeline changes.
MedQA includes English and Chinese medical examination partitions; this site focuses on its English USMLE four-option test. MedMCQA assembles Indian postgraduate medical examination material with examination-based splits. Their source populations and protocols differ.
The paper’s exact statistics give 4,183 development questions and 6,150 test questions. Other source wording conflicts. Record the downloaded file and identifiers, and do not relabel a validation run as test.
A Wilson 95% interval for user-entered binary outcomes. Its starting successes are hypothetical; it does not report model performance or establish clinical readiness.
No. Arcophos publishes independent analysis and tools. The original authors created MedQA and MedMCQA, and their papers and repositories remain the primary sources.
Working tool / saved on this device
Document the actual run before comparing accuracy. Checked items are self-reported protocol decisions, not evidence that a benchmark is uncontaminated or clinically valid.
Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.
Download the evidence ↗