Join the network

Guides

How to evaluate, deploy, and govern AI in real clinical work

Practical, evidence-first guides for the questions that decide whether a tool belongs in your workflow — written for clinicians and health leaders, organized into six collections.

How to read the numbers and studies behind clinical AI claims — no ML background assumed.

HealthBench Professional, explained: how OpenAI now measures clinician-facing AI

OpenAI's clinician benchmark grew a professional edition in 2026 — physician- authored conversations, multi-stage physician adjudication, and a headline result in which the deployed model outscored specialist-matched physicians. What the benchmark actually measures, how it was built, and what a score like that should and should never change in your decisions. As of August 2026.

LLM evals vs clinical evaluation: two different instruments

Why a benchmark score and a clinical evaluation measure different things, in different units, on different populations — and three published cases where the same models looked strong on one instrument and weak on the other. As of July 2026.

AUROC explained for clinicians

What the area under the ROC curve really measures, worked through two real published models — an AUROC of 0.991 that still forces a trade-off, and one of 0.63 that hid a 67% miss rate — and the four questions it can never answer for you. As of July 2026.

Calibration curves for clinicians

How to read a calibration plot the way it fails — each shape of curve paired with a real published model that failed that way, and the bedside decision each failure would distort. Discrimination tells you the ranking is right; calibration tells you the number is. As of July 2026.

Dataset shift and model decay

A clinical AI model is accurate at one moment, on one population. Both move. This is a working guide to dataset shift and model decay — the three ways a model's performance comes apart, the signal that warns you first, and what the evidence says actually slows it down. As of July 2026.

How to read an AI validation study

A working checklist a clinician can hold against any paper claiming an AI model works — the nine questions that decide whether a result will survive contact with real patients, each tied to the reporting standard that governs it. As of July 2026.

How to read an FDA clearance summary

A section-by-section walk of one real 510(k) summary for a cleared AI head-CT triage device — the clearance letter, the indications, the predicate, the performance table — mapped to what clearance does and does not establish. As of July 2026.

Internal vs external validation

The gap between how an AI model scores on a held-out slice of its own data and how it scores at a second hospital is where most inflated results go to die. This is a working guide to internal versus external validation — what each measures, how far performance typically falls between them, and the published figures that show it. As of July 2026.

Prospective vs retrospective evaluation

A model that looks brilliant on curated historical data has cleared the lowest bar, not the last one. This is a working guide to prospective versus retrospective evaluation — what each can prove, the documented cases where retrospective promise did not survive prospective testing, and the one where it did. As of July 2026.

Sample size and confidence intervals in AI studies

A single accuracy number with no interval around it is a guess wearing a lab coat. This is a working guide to sample size and confidence intervals in clinical AI studies — why so many are underpowered, how wide the real uncertainty is, and the published figures that prove it. As of July 2026.

Subgroup performance and bias audits

A model can post excellent overall accuracy and still fail the specific patients who most need it to work. This is a working guide to subgroup performance and bias audits — the documented failures, why they happen, and the stratified metrics that surface them before deployment. As of July 2026.

Ten red flags of an overfit model claim

Ten warning signs that a model's reported performance will not survive contact with new patients — each one anchored to a documented, published failure, and paired with the specific check that would have caught it. As of July 2026.

What benchmark scores don't tell you

A benchmark score is a real measurement of a narrow thing. This guide reads the peer-reviewed critique of medical AI benchmarks — saturation, contamination, and construct validity — and the one question to ask before a leaderboard number reaches a clinical decision. As of July 2026.

The evidence, economics, and risks of AI scribes and ambient clinical documentation.

Ambient AI reaches nursing documentation: what the inpatient turn means

In 2026 the ambient documentation market pivoted from physician notes to nursing flowsheets — Epic's tool reached bedside nurses in new health systems, and vendors shipped inpatient nursing suites. Nurses spend about a third of a twelve-hour shift in flowsheets; the technology aimed at that number works differently from a scribe, and fails differently too. As of August 2026.

AI scribe vendor landscape 2026

A dated, neutral map of the ambient AI scribe field — what each vendor documents, what it has disclosed raising, and which independent evaluations actually exist — with every cell tied to a source and no leader named. As of July 2026.

AI scribes: what the evidence actually shows

Ambient scribes reached millions of uses before the first randomized trial reported a result. This guide sets the deployment scale and the vendor efficiency claims against the peer-reviewed record — how few evaluations meet real-world-evidence criteria, and what the studies that do exist actually found. As of July 2026.

Ambient AI ROI calculator: a transparent model

A fully worked return model for ambient AI scribes — every input, formula, and example laid out and drawn from published figures — expressed in clinician-time and visit-capacity terms. A model to test locally, never a promise. As of July 2026.

Documentation time and burnout: what the evidence shows

Documentation load is the most consistently measured driver of clinician burnout, and it is the problem ambient scribes are aimed at. This guide separates two kinds of evidence — objective time (observation and EHR logs) and self-reported burnout — and reads the published studies on each. As of July 2026.

Hallucination and omission rates in AI scribe notes: what the studies measured

Published "hallucination rates" for ambient scribe notes range from about 1.5% to 31% — a spread that looks like disagreement and is really a difference of units. This guide keys each number to what it counted, pairs every hallucination figure with its omission counterpart, and gives you one table to compare them honestly. As of July 2026.

The Kaiser Permanente 2.5-million-encounter ambient scribe deployment, analyzed

Two NEJM Catalyst papers document the largest ambient AI scribe rollout in the published record — 7,260 physicians and more than 2.5 million encounters in a single year. This guide reads them strictly on their own terms, separating what the deployment actually measured from what it did not, so the headline number is read as adoption evidence rather than efficacy proof. As of July 2026.

Coding, billing, and upcoding risks of ambient AI notes

Ambient scribes write fuller notes — and a fuller note can support a higher bill. This guide connects the long regulatory record on documentation-driven upcoding to the specific new failure modes of AI-generated notes, and pairs each risk with a control that deployed programs actually run. As of July 2026.

Do ambient AI scribes need FDA regulation?

Two questions hide inside this one. Is an ambient scribe an FDA device today? — a statutory question with a reasonably clear answer. And should documentation tools face oversight given measured note-quality gaps? — an open policy question. This guide keeps them apart and maps each scribe capability to the statute it would or would not cross. As of July 2026.

An implementation checklist for medical groups deploying ambient AI scribes

A deployment checklist where every item traces to a published study or regulation rather than to vendor advice — and, where the evidence measured it, the number a group should expect. Governance, contracting, training, review, monitoring, and wellbeing, in the order they bite. As of July 2026.

Patient consent for ambient recording, by jurisdiction

Ambient AI scribes listen to the visit — which raises three separate legal questions that ranking pages tend to blur into one: recording-consent law, HIPAA authorization, and data-protection lawful basis. A primary-source map across US states, the EU, and the UK. As of July 2026.

Agents in clinical and administrative work — safety, oversight, and what is actually deployed.

Agent safety frameworks for clinical settings

There is no single safety standard for clinical AI agents yet. In its place is a stack of published frameworks — general AI risk, LLM and agent security, healthcare assurance, and regulation — that a hospital has to compose itself. This guide maps which standard governs which layer. As of July 2026.

Agents in nursing workflows: documentation, handover, and the virtual-nursing evidence

Two very different things get marketed as the "AI nurse": documentation assistants that reformat what a nurse says, and virtual-nursing programs that move tasks to a remote colleague. The measured time savings, the completeness gains, and why one of these has much stronger evidence than the other. As of July 2026.

Designing human oversight for clinical agents

How to build supervision around an AI agent that acts on the clinical record — translating EU AI Act Article 14 and FDA guidance into concrete oversight tiers, automation-bias defences, and after-go-live monitoring. As of July 2026.

Documented agentic deployments in healthcare: a tracker

A running tally of the healthcare AI-agent systems that carry a published paper or the operator's own disclosure — what each one does, where the human sits, and what it actually reported. The field is loud; the documented set is small. As of July 2026.

EHR-integrated agents: what it takes, and what the evidence shows

Putting an AI agent inside the electronic health record means three hard jobs — read the record, reason over it, and act on it — each governed by a standard and measured by a benchmark. Here is what the published evidence says about how well today's agents do each one. As of July 2026.

How to evaluate an agent before deployment

A pre-deployment checklist for clinical AI agents — ten dimensions to test before a system touches a patient, each tied to a published benchmark or a named safety framework, with the questions to put to any vendor. As of July 2026.

MCP and interoperability for hospital agents

A hospital agent needs two standards to do its job: one for how the model reaches its tools, and one for the shape of the clinical data it reaches. This guide maps the MCP-plus-FHIR stack from the specifications themselves, and the security duties that live at the seam between them. As of July 2026.

Prior-authorization agents: what the CMS rule and the evidence actually say

Prior authorization is being automated from both ends at once — providers building agents to submit and appeal, payers running algorithms to adjudicate. What the CMS-0057-F final rule now requires, what the published evidence shows an agent can and cannot do, and the one guardrail regulators drew around automated denials. As of July 2026.

Revenue-cycle and coding agents: accuracy, denials, and the compliance stakes

Coding and denial-management agents are being pointed at the most repetitive, highest- volume work in a hospital's back office. What the peer-reviewed accuracy benchmarks actually show, why a fabricated billing code is a compliance event rather than a cosmetic error, and where the audit trail has to stay human. As of July 2026.

What is agentic AI in healthcare

The hub for agentic AI in a clinical setting — the working definition from the peer-reviewed literature, the anatomy of an agent, what benchmarks show these systems can and cannot yet do, and why the definition itself carries governance weight. As of July 2026.

FDA, EU AI Act, MHRA, WHO, HIPAA, state law — and the governance work the rules leave to hospitals.

The CMS WISeR Model, explained: AI prior authorization reaches traditional Medicare

Since January 2026, six technology companies have been running AI-assisted prior authorization on select services in six US states — the first Innovation Center model built around the technology, and the first to survive a Senate repeal vote. How the model actually works, from the Federal Register notice itself. As of August 2026.

EU AI Act on 2 August 2026: what now applies to healthcare

The Digital Omnibus is law, and the AI Act's general application date arrives with the healthcare high-risk obligations deferred to 2027 and 2028 — while the Article 50 transparency duties land on schedule. What each date now covers, tied to the amended Regulation. As of August 2026.

The first FDA-cleared patient-facing LLM: what UpDoc's 510(k) actually says

In December 2025 the FDA cleared UpDoc, a prescription insulin-management device whose patient interface is a large language model — and the clearance went through as a drug dose calculator, with a predetermined change control plan attached. What the record shows, what it leaves open, and what it signals for every LLM heading toward the device pathway. As of August 2026.

US state laws on AI mental-health chatbots: a tracker

Two legislative waves in two years: 2025 brought the first therapy restrictions in Illinois, Nevada and Utah and the first companion-chatbot safety statutes, and 2026 has added five more states restricting AI-delivered therapy — with 98 chatbot bills pending across 34 states. Every row tied to a statute or a primary tracker. As of August 2026.

Building an algorithmovigilance program

A build sequence for watching clinical algorithms after they go live — the operating model, the four signals to instrument, the people to name, the cadence to run, and the escalation path — each step tied to a primary source. As of July 2026.

EU AI Act for healthcare: a living timeline

A dated timeline of the EU AI Act as it lands on healthcare — the exact application dates after the Digital Omnibus entered into force: standalone health AI under Annex III now falls due 2 December 2027, and AI that is a medical device under Annex I on 2 August 2028. Each row tied to the Regulation itself. As of August 2026.

The FDA AI-enabled device list: a statistics tracker

A dated read of the FDA's AI-Enabled Medical Device List — how many devices carry an authorization, which specialties dominate, which pathways they take, and how thinly they report performance and demographics — each figure tied to the FDA or a peer-reviewed census. As of July 2026.

The FDA's total-product-lifecycle draft guidance for AI devices, explained

What the FDA's January 2025 draft guidance on AI-enabled device software functions actually asks of manufacturers across the total product lifecycle — its scope, its thirteen sections, its transparency-and-bias and representativeness demands, and its still-draft status. As of July 2026.

HIPAA and LLMs: what is permitted

A decision map for putting patient data near a large language model under US health-privacy law — read straight from the 45 CFR Part 164 text: when data stops being protected, when a vendor becomes a business associate, what a permitted use is, and why training a model on records is still an open question. As of July 2026.

The hospital AI governance committee playbook

A build sequence for standing up a hospital AI governance committee — charter, membership, intake, tiered review, a local-validation gate, monitoring, and board reporting — with every step cross-walked to a published governance framework. As of July 2026.

Liability when clinical AI errs

When an AI-assisted clinical decision harms a patient, who answers for it — the clinician, the developer, or the health system? A fair-minded map of the three doors a claim can walk through, grounded in the peer-reviewed legal scholarship, and honest that the law is still unsettled. As of July 2026.

Predetermined Change Control Plans in practice

How manufacturers actually build a PCCP and what health systems should check — the three required sections turned into what you write, the five guiding principles as design constraints, and what the first authorized plans reveal. As of July 2026.

Transparency and labeling requirements for clinical AI

What US rules actually force a clinical-AI tool to disclose — the FDA device track and the ONC/ASTP HTI-1 certified-EHR track — mapped onto the single artifact both converge on, the model card, and read against how little devices disclose today. Each requirement tied to primary rule text. As of July 2026.

The UK MHRA AI Airlock, explained

What the MHRA's regulatory sandbox for AI as a medical device actually is, how it works, and what its two published cohorts found — mapping each real regulatory gap to the case study that surfaced it, drawn from the pilot and Phase 2 reports themselves. As of July 2026.

US state laws on AI scribes: a tracker

A dated, two-layer status table of the state laws that reach ambient AI scribes — the recording-consent rules that decide whether a scribe may capture the visit at all, and the newer generative-AI disclosure and governance statutes — each row tied to the legislature's own text. As of July 2026.

WHO guidance on large language models in health

What the World Health Organization's 2024 guidance on large multi-modal models actually says — the five ways it expects these systems to be used in health, the risks it names, who it tells to do what, and the one thing to keep straight: it is advisory, and binding rules sit elsewhere. As of July 2026.

What is cleared and evidenced in each specialty — and what is still promise.

AI in cardiology: what is cleared and evidenced in 2026

Cardiology has the field's rare thing — a randomized trial where an AI alert changed diagnoses in routine care — alongside the largest consumer screening study ever run. A dated read of what is authorized and what the strongest trials measured, each figure tied to a primary source. As of July 2026.

AI in Emergency Care

A guide to artificial intelligence in the emergency department — sepsis and deterioration alerts, triage, and imaging detection for stroke and fracture — reading each result for the one thing that decides its value: whether it changed the workflow and the human response, rather than the model's accuracy alone. Each figure tied to its primary source. As of July 2026.

AI in hospital operations: a 2026 evidence guide

Beds, queues, theatres, and staff rosters are where AI in hospitals has the cleanest data and the clearest payoff — and the widest gap between prediction accuracy and proven operational benefit. What the evidence supports for patient-flow, scheduling, and command-center AI, and why an accurate forecast is only half the job. As of July 2026.

AI in nursing: a 2026 evidence guide

A strained global workforce, a documentation load measured in the hundreds of entries per shift, and the one nurse-facing AI with a randomized mortality result behind it. What the evidence supports for AI in nursing — deterioration detection, documentation, and knowledge tools — and where a nurse still has to own the output. As of July 2026.

AI in oncology: what the evidence actually shows

A guide to where artificial intelligence has earned its place across the cancer-care pathway — screening, detection, pathology — with the randomized evidence separated from the retrospective reader studies that fill vendor decks, each figure tied to its primary source. As of July 2026.

AI in pathology: what is cleared and evidenced in 2026

Pathology has the field's most striking research results and one of its smallest cleared footprints. A dated read of what the FDA has authorized in digital pathology and what the strongest studies actually measured, each figure tied to a primary source. As of July 2026.

AI in pediatrics: a 2026 guide

Pediatrics is the specialty where clinical AI is furthest behind — for structural reasons — yet it holds some of the field's most striking proofs of concept. This guide pairs the landmark evidence with the data and device gaps that explain why so little is cleared for children. As of July 2026.

AI in pharmacy: a 2026 evidence guide

What the published record actually supports for AI in pharmacy — where dispensing automation has measured safety gains, why interruptive drug-interaction alerts are overridden roughly nine times in ten, and what a language model can and cannot be trusted to do with a drug question. As of July 2026.

AI in primary care: exam scores versus patient outcomes

A guide to artificial intelligence at the point of first contact — decision support, conversational diagnosis, autonomous screening, and ambient documentation — reading the field's benchmark hype against the handful of large trials that measured what happened to patients. Each figure tied to its primary source. As of July 2026.

AI in psychiatry and mental-health chatbots: a 2026 guide

Mental-health chatbots are the most consumer-facing and most contested use of AI in healthcare. This guide separates three things the market blurs — structured tools with trial evidence, general-purpose chatbots with documented safety failures, and the fast-moving 2025 regulation — with primary sources throughout. As of July 2026.

AI in radiology: what is cleared and evidenced in 2026

Radiology is where clinical AI is most deployed and most authorized — but the cleared reality and the trial evidence are two different maps. A dated read of the FDA device record and the strongest randomized trial in imaging, each figure tied to a primary source. As of July 2026.

AI in surgery: a 2026 guide

What artificial intelligence actually does in the operating room today — from polyp detection and anatomy guidance to surgical-phase recognition and the first supervised-autonomy demonstrations — with every capability tied to its primary evidence and its regulatory status. As of July 2026.

Side-by-side, sourced comparisons of the major healthcare AI tool categories.

The best AI in healthcare communities, societies, and networks in 2026

Twelve places where the people doing AI in healthcare actually gather — CHAI, Health AI Partnership, AMIA, SIIM, AIME, DiMe, AAIH, the FHIR chat, Medblocks, Health Tech Nerds, Out-Of-Pocket, and AIMOCS — compared on who they serve, what membership really looks like, and every published fee, each figure from the organization's own pages. As of 12 August 2026.

How to get into AI in healthcare: a clinician's guide

The realistic map for a physician, nurse leader, or researcher entering the field in 2026: five destinations, what each actually requires, where formal credentials matter and where they are myth, a self-directed 90-day reading path, and when a community shortens all of it. Every number sourced and dated. As of 12 August 2026.

The best AI in healthcare courses in 2026, compared

Johns Hopkins, Harvard, Stanford, MIT Sloan, and AIMOCS, weighed on what each programme publishes about itself: how current the syllabus is, how deep the generative and agentic coverage goes, what you build, what continues after the certificate, and who carries accredited credit. Every cell sourced and dated. As of 1 August 2026.

AI scribe head-to-head comparison

A dated capability matrix for four ambient AI scribes — what the vendors document and what independent trials actually measured — presented as attributes, with no winner declared. As of July 2026.

AI tools for medical education compared

A neutral, sourced read of the AI systems learners and educators actually reach for — what each scores on medical exams and benchmarks, how it grounds its answers, and why an exam-recall number is a weak proxy for competence. No ranking. As of July 2026.

Clinical reference AI compared

A dated capability matrix for the AI tools clinicians use to answer questions at the point of care — what each vendor documents, and what one blinded benchmark measured — with no winner declared. As of July 2026.

Imaging AI marketplaces compared

A neutral, attribute-by-attribute read of the platforms that put radiology AI algorithms in front of a reading room — where the software runs, whose models it carries, and which regulator cleared them — anchored to the FDA's own device list. No ranking. As of July 2026.

Medical coding AI compared

A dated capability matrix for AI medical-coding tools — what each vendor documents about autonomy and human review, set beside the peer-reviewed evidence on code-assignment accuracy. No winner declared. As of July 2026.

Open vs closed models for hospital deployment

A neutral trade-off matrix for the build-or-buy question underneath clinical AI — data residency, validation burden, transparency, governance, and measured performance — set out attribute by attribute and sourced, with no winner declared. As of July 2026.

Patient communication drafting tools compared

A neutral read of what controlled and deployment studies actually document about AI tools that draft replies to patients — empathy, readability, time, clinician burden, and the human review every one of them keeps. No ranking. As of July 2026.

Patient triage chatbots compared

A dated, safety-first capability matrix for symptom-checker and triage chatbots — what peer-reviewed studies measured about their diagnostic and triage accuracy, presented without a winner. As of July 2026.