Join the network

Statistics · 8 min read

LLM medical benchmark results tracker

A running read of how large language models score on the leading clinical benchmarks — MedHELM, HealthBench, and the new HealthBench Professional — set beside the peer-reviewed critique of what those scores do, and do not, tell you about bedside performance. As of August 2026.

The short version

  • Three benchmarks anchor the field as of August 2026: MedHELM (121 clinician-defined clinical tasks), HealthBench (5,000 conversations against 48,562 physician-written rubric criteria), and HealthBench Professional (525 physician-authored clinician-chat tasks, released April 2026).
  • The leaderboard has turned over: on MedHELM's 14 May 2026 refresh, Gemini 3.1 Pro (Preview) leads at a 0.652 mean win rate, ahead of Gemini 3.5 Flash (0.642); the 2025 paper's leaders have fallen away — DeepSeek R1 now sits mid-pack at 0.485, and o3-mini and Claude 3.5 Sonnet no longer appear on the board.
  • On HealthBench Professional, the deployed ChatGPT for Clinicians system (GPT-5.4) scored 59.0 — above base GPT-5.4 (48.1) and above specialist-matched physician responses written with unbounded time and web access (43.7). A vendor-reported result on a vendor-built benchmark, with an open grading harness.
  • The catch: a systematic review of 519 studies found only 5% used real patient-care data, 44.5% tested licensing-exam-style questions, and 95.4% scored accuracy while calibration and uncertainty (1.2%) went nearly unmeasured.
  • A high benchmark score licenses a hypothesis worth testing in your setting — it does not license clinical use. Prospective, local validation is still the bar.
On this page

Every few weeks a new model tops a medical benchmark, and the headline writes itself: the machine "passes the boards." This page tracks what those benchmarks actually score, who currently leads, and — the part the leaderboards leave out — how far a high number sits from a clinical warrant. Every figure is dated and tied to a numbered source below. As of August 2026.

Which benchmarks anchor the field?

Three efforts define the serious end of clinical LLM evaluation right now.

MedHELM, built on the HELM framework at Stanford's Center for Research on Foundation Models, scores models across 121 clinical tasks organised into 5 categories and 22 subcategories — a taxonomy assembled with 29 practising clinicians — using a suite of 35 benchmarks, 17 drawn from existing datasets and 18 newly built 1. Its distinguishing move is coverage: instead of a single exam, it spans diagnostic decision support, note summarisation, patient communication, and administrative work, and it includes tasks grounded in real electronic health records rather than curated question banks.

HealthBench, released by OpenAI, takes the opposite tack — depth over breadth of task type. It grades models on 5,000 multi-turn conversations against 48,562 unique rubric criteria written by 262 physicians who practise across 60 countries and 26 specialties 3. Each conversation is scored against its own rubric, so a model earns credit for the specific things a physician would want said, and loses it for the specific things a physician would flag — including under uncertainty and in emergencies.

HealthBench Professional, released by OpenAI in April 2026 alongside its ChatGPT for Clinicians product, narrows that method to the population that matters most here: clinicians. It contains 525 physician-authored tasks across three use cases — care consult, writing and documentation, and medical research — selected from a candidate pool of 15,079 examples, with difficult examples enriched roughly 3.5-fold and about a third drawn from deliberate adversarial red-teaming; every example and rubric was adjudicated by three or more physicians across three phases 67. It also ships with an unusually strong human baseline: specialist-matched physicians wrote responses for every task with unbounded time and web access 6. Our full read is in HealthBench Professional, explained.

Who currently leads?

Two current readings, one per benchmark family — and both are snapshots.

On MedHELM's live leaderboard, refreshed 14 May 2026 (v5.0.0), Gemini 3.1 Pro (Preview) leads with a 0.652 mean win rate, ahead of Gemini 3.5 Flash (0.642), Muse Spark (0.621), GPT-5.4 mini (0.552), and GPT-5.4 (0.538) 5. The board itself has turned over: o3-mini and Claude 3.5 Sonnet — leaders of the 2025 paper's evaluation — no longer appear on it, and DeepSeek R1, that evaluation's top scorer, now sits mid-pack at 0.485 5. That original nine-model evaluation — DeepSeek R1 at a 66% win-rate, o3-mini at 64%, Claude 3.5 Sonnet comparable at roughly 40% lower estimated compute cost 2 — is now trend history, and its fate is the standing lesson: the specific names turned over within a year, while its structural finding, that reasoning-tuned models do better on multi-step clinical tasks than raw scale predicts, has held.

On HealthBench Professional, the deployed ChatGPT for Clinicians system (built on GPT-5.4) scored 59.0 — above base GPT-5.4 at 48.1, above every other evaluated model, and above the specialist-matched physician responses at 43.7 67. The gap was widest where drafting is the job: 64.1 versus 32.1 for physicians on writing and documentation, against 51.0 versus 42.7 on care consults and 67.0 versus 56.3 on medical research 7. Two reading notes travel with those numbers. They are length-adjusted rubric scores on a 0–100 scale, not percent accuracy 7. And they are vendor-reported — OpenAI scoring its own product on its own benchmark — mitigated, but not erased, by the open grading harness that lets an independent team re-run the comparison.

BenchmarkWhat it scoresScaleCurrent signal (as of August 2026)
MedHELM121 clinical tasks, clinician taxonomy35 benchmarks; live leaderboardGemini 3.1 Pro (Preview) leads the 14 May 2026 refresh at 0.652 mean win rate; Gemini 3.5 Flash 0.642 5
HealthBench5,000 graded conversations48,562 rubric criteria, 262 physiciansLaunch paper documented model progress from 16 (GPT-3.5 Turbo era) to 60 (o3); its Hard subset topped out at 32 at launch 3
HealthBench Professional525 physician-authored clinician-chat tasks3 use cases; rubrics adjudicated by ≥3 physiciansChatGPT for Clinicians (GPT-5.4) 59.0 vs base GPT-5.4 48.1 and physician baseline 43.7 67

What do benchmark scores license clinically — and what don't they?

Here is the part the leaderboards omit. A systematic review of 519 studies published between January 2022 and February 2024 examined how health-care LLMs were actually being evaluated — and the shape of that evaluation is narrow 4.

  • Only 5% used real patient-care data. The overwhelming majority tested models on curated questions rather than live records.
  • 44.5% assessed medical-knowledge and licensing-examination questions, the single most common task, followed by diagnosis at 19.5%. Fully 84.2% of studies were question-answering; summarisation (8.9%) and open dialogue (3.3%) were rare.
  • 95.4% used accuracy as the primary dimension. Fairness and bias were examined in 15.8% of studies, deployment considerations in 4.6%, and calibration and uncertainty in just 1.2%.

Read together, those numbers explain the gap between a benchmark score and a clinical warrant. A model that answers licensing-exam items correctly has demonstrated recall and reasoning on a tidy, single-answer format. That is a real capability — and a weak proxy for drafting a safe note from a rambling encounter, reconciling a contradictory record, or knowing when to say "I am not sure, escalate." The tasks that dominate the benchmarks are the tasks least like the ones that carry clinical risk.

This is why MedHELM's inclusion of real-EHR tasks matters, why HealthBench's rubric grading of uncertainty and emergency behaviour matters, and why HealthBench Professional's adversarial enrichment and physician baseline matter: all three push scoring toward the messier competencies. But even they are laboratory measures. A high score says a model is worth putting in front of a prospective, local evaluation. It does not stand in for that evaluation.

How to read these numbers

Five cautions travel with every score on this page. Leaderboard rank is a snapshot that turns over with each model release — treat any single ranking as perishable; the DeepSeek R1 era lasted about a year. A benchmark measures the tasks its authors chose; coverage gaps are invisible in the headline number. Exam-style performance and bedside performance are different measures, and the review above shows how rarely the second is tested. Almost all scoring optimises accuracy while leaving calibration, bias, and deployment behaviour largely unmeasured — precisely the properties that govern safety. And when a benchmark's author is also the vendor of its winning system — as with HealthBench Professional — treat the headline as a vendor claim until independently replicated, which the open grading harness makes possible.

The practical rule: use benchmarks to shortlist and to falsify — a model that fails MedHELM's coding tasks is unlikely to surprise you in production — but require prospective, in-setting validation before a score touches a clinical decision. Our sibling clinical AI trial results tracker follows the studies that do that testing, and what benchmark scores don't tell you goes deeper on the reading frame.

Sources and method

Figures are drawn from the primary publications listed below: the MedHELM paper 12 and its live leaderboard 5, the HealthBench paper 3, the HealthBench Professional paper 67, and the JAMA systematic review of evaluation practice 4. Vendor-published evaluation numbers are labeled as such wherever they appear. We revisit this page on a ninety-day cycle and whenever a benchmark ships a new version, a frontier model posts new scores, or a peer-reviewed methodology critique lands — it was refreshed on 16 August 2026 to pick up the May 2026 MedHELM leaderboard and HealthBench Professional. As of August 2026.

Questions and answers

  • What is the best LLM for medical tasks according to benchmarks?

    On MedHELM's leaderboard, as of its 14 May 2026 refresh, Gemini 3.1 Pro (Preview) leads with a 0.652 mean win rate, ahead of Gemini 3.5 Flash (0.642), Muse Spark (0.621), GPT-5.4 mini (0.552), and GPT-5.4 (0.538). On clinician-chat tasks, the GPT-5.4-based ChatGPT for Clinicians system posted 59.0 on HealthBench Professional, the strongest reported score. A leaderboard position reflects performance on defined tasks, and the ranking shifts with every new model release, so read it as a snapshot rather than a standing verdict.

  • Do medical LLM benchmarks predict real clinical performance?

    Only partly. A systematic review of 519 studies found that just 5% used real patient-care data and that 44.5% tested medical-knowledge or licensing-exam questions. Exam-style scores measure recall and reasoning on curated questions, which is a weak proxy for how a model behaves on messy records, incomplete histories, and live clinical workflows.

  • What is the difference between MedHELM and HealthBench?

    MedHELM, from Stanford's Center for Research on Foundation Models, scores models across 121 clinical tasks in a clinician-built taxonomy, including tasks drawn from real electronic health records. HealthBench, from OpenAI, grades 5,000 open-ended conversations against 48,562 rubric criteria written by 262 physicians. MedHELM emphasises task coverage; HealthBench emphasises graded, realistic dialogue. HealthBench Professional, added in April 2026, narrows HealthBench's method to 525 physician-authored clinician-chat tasks with harder, adversarially enriched examples.

Sources

  1. Bedi S, Liu Y, Orr-Ewing L, et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nature Medicine. 2026. doi.org/10.1038/s41591-025-04151-2
  2. Bedi S, et al. MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks. arXiv:2505.23802. 2025. arxiv.org/abs/2505.23802
  3. Arora RK, Wei J, Soskin Hicks R, et al. HealthBench: Evaluating Large Language Models Towards Improved Human Health. arXiv:2505.08775. 2025. arxiv.org/abs/2505.08775
  4. Bedi S, Jain SS, Bedi P, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025;333(4):319-328. doi.org/10.1001/jama.2024.21700
  5. Stanford Center for Research on Foundation Models. MedHELM leaderboard, v5.0.0 refresh of 14 May 2026 (accessed 16 August 2026). medhelm.org/
  6. Soskin Hicks R, Trofimov M, Lim D, Arora RK, et al. HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats. arXiv:2604.27470. 2026. arxiv.org/abs/2604.27470
  7. OpenAI. HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats (paper PDF). 2026. cdn.openai.com/dd128428-0184-4e25-b155-3a7686c7d744/HealthBench-Professional.pdf

Every claim on this page is tied to a numbered primary source above. Read how we source and review at our editorial policy.