An on-premise clinical artificial intelligence (AI) agent achieved 90.04% accuracy on a seven-disease diagnostic task and used repeated-run behavioral consistency to identify a lower-risk subset for simulated autonomous handling, according to a study published in
Nature Medicine.
In the dual-agent framework, a Physician Agent interacted with a Patient Agent, requested findings through clinical tools, and returned a diagnosis with a reasoning trace. Investigators evaluated open-weight large language models on MIRA-v2 (551 cases across seven conditions) and Clinical Decision Making (2,400 cases across four abdominal conditions), both derived from MIMIC-IV, and on the PubMed-derived VivaBench benchmark (990 cases across 10 specialties).
On MIRA-v2, Qwen-3.5 had the highest on-premise accuracy at 90.0%, compared with 90.7% for the cloud-based GPT-5.2 baseline. Qwen-3.5 reached 83.8% on the four-condition benchmark. In the primary reliability analysis using GLM-4.5-Air, diagnostic consistency across repeated runs best distinguished correct from incorrect decisions, with an area under the receiver operating characteristic curve of 0.860 versus 0.747 for an internal probability score.
At a diagnostic-consistency threshold of 0.90, the system retained 272 of 551 cases (49.4%) for autonomous handling, with 98.9% accuracy and three residual errors; 279 cases were deferred for clinician review, including 49 errors. Under an information-scarcity stress test, overall accuracy fell from 90.6% to 70.2%, while consistency remained discriminative (AUC, 0.875). Aggregating five runs increased doctor-agent token use by approximately fivefold.
The authors cautioned that the evaluations were retrospective simulations. Both primary benchmarks came from MIMIC-IV's single-institution data ecology, the workflow was text-only, thresholds varied with implementation choices, and VivaBench did not replace independent or prospective validation. They called for prospective studies and dedicated bias audits to assess clinician reliance, review burden, safety, and fairness.
Source: Zhang L, Wölflein G, Ferber D, et al. On-premise medical AI agents for reliable clinical decision-making.
Nat Med. Published online September 15, 2026.
doi:10.1038/s41591-026-04609-x