MedQA & MedMCQA · independent analysis

Know what the exam score compares.

Independent analysis of MedQA and MedMCQA: verified dataset splits, historical retrieval experiments, answer-selection metrics and an interactive accuracy interval calculator.

Independent analysis by Arcophos · updated

Two exams. Different denominators.

Dataset dossiers ↗

Training, development and test partitions from the original papers. Each bar shows the composition of its own question bank.

MedQA

12,723 questions · Original 2020 paper; USMLE four-option evaluation

  • 10,178Training [1]
  • 1,272Development [1]
  • 1,273Test [1]

Published split counts; these are questions, not patients. They sum to 12,723.

MedMCQA

193,155 questions · CHIL 2022 paper; exact split counts from §3/Table 2

  • 182,822Training [3]
  • 4,183Validation [3]
  • 6,150Test [3]

The original README swaps validation/test table headings; use exact paper statistics and record the actual downloaded files.

Bar widths describe dataset composition, not model performance. The four-option variant, split and retrieval condition remain part of every score.

Published results

Source record ↗

Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.

Paper-reported results / selected rows

Original paper: selected USMLE baselines

September 2020 paper, four-option USMLE test; original retrieval/reader protocols.

Answer-selection accuracy · %
050100
Reported
IR-CustomRetrieval baseline
36.1
BERT-Base-EnOriginal paper reader pipeline
34.3
BioBERT-BaseOriginal paper reader pipeline
34.1
BioBERT-LargeOriginal paper reader pipeline
36.7

Paper-reported historical results. No uncertainty intervals are provided in this table; do not present these as current frontier performance.

Source: Table 8 [1]

Paper-reported results / selected rows

One reader, three context conditions

CHIL 2022, PubMedBERT reader, original fine-tuning and test evaluation.

Test accuracy · %
050100
Reported
PubMedBERT · no contextPaper value 0.41
41
PubMedBERT · WikipediaPaper value 0.42
42
PubMedBERT · PubMedPaper value 0.47
47

Paper-reported historical measurements. Table proportions are multiplied by 100 to display percentage accuracy; no intervals are supplied.

Source: Table 4 [3]

An original analytical tool

Inspect the accuracy denominator

Inspect the arithmetic

Explore a Wilson 95% interval for a hypothetical answer-selection run. The initial denominator is the 1,273-question MedQA USMLE test set; 1,000 successes is an invented illustration, not a model result.

Illustrative counts / edit to explore

Starting success counts are hypothetical. This does not reproduce a model run or account for repeated patients, clustered questions, grader error, or multiple comparisons.

Observed fraction in this example78.55%

Wilson 95% interval
76.22% – 80.72%

0%50%100%

This interval describes sampling uncertainty under the stated binomial assumptions. It is not a test of clinical usefulness or a paired model comparison.

This binomial illustration assumes independent question outcomes. It does not account for contamination, question dependence or paired model comparisons, and it does not estimate clinical benefit. [1][5]

The benchmark in detail

All dossiers →
Dossier01

Original 2020 paper; USMLE four-option evaluation

MedQA · USMLE four-option ↗

A medical answer-selection score with a specific denominator.

UnitQuestionMeasureAnswer-selection accuracy
Dossier02

CHIL 2022 paper; exact split counts from §3/Table 2

MedMCQA ↗

The split and the retrieval source belong beside the score.

UnitQuestionMeasureAnswer-selection accuracy

What we examine

MedQA and MedMCQA evaluate medical examination answer selection. Their scores depend on which questions, options and retrieval resources a system receives. This publication makes those conditions visible: exact denominators, source-version discrepancies, historical experiments and the limits of an accuracy claim. Start with the benchmark dossiers, inspect the illustrative interval calculator, then use the comparison worksheet to document a run. Arcophos provides the analysis and tools; the original research teams created the benchmarks.

Variant
Identify the language, answer-option format and split before reading a score.
System
Separate the base model from retrieved context, prompts and answer parsing.
Claim
Distinguish answer accuracy from evidence about clinical action or benefit.

Analysis & interpretation

All analyses →

Questions, answered

How do MedQA and MedMCQA differ?

MedQA includes English and Chinese medical examination partitions; this site focuses on its English USMLE four-option test. MedMCQA assembles Indian postgraduate medical examination material with examination-based splits. Their source populations and protocols differ.

Is the MedMCQA validation set the test set?

The paper’s exact statistics give 4,183 development questions and 6,150 test questions. Other source wording conflicts. Record the downloaded file and identifiers, and do not relabel a validation run as test.

What does the accuracy calculator show?

A Wilson 95% interval for user-entered binary outcomes. Its starting successes are hypothetical; it does not report model performance or establish clinical readiness.

Are these official benchmark websites?

No. Arcophos publishes independent analysis and tools. The original authors created MedQA and MedMCQA, and their papers and repositories remain the primary sources.

Prepare a comparison worksheet

Working tool / saved on this device

Prepare a MedQA or MedMCQA comparison receipt

Interactive worksheet

Document the actual run before comparing accuracy. Checked items are self-reported protocol decisions, not evidence that a benchmark is uncontaminated or clinically valid.

Identify the task

Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.

Download the evidence ↗