AI in Medicine Update — Oct 10, 2026
Generated Oct 11, 2026 · 12:25
The latest research on this topic, summarized for clinicians.
If the audio fails to play, refresh the page to renew the link.
Get this every week in your podcast app — free.
New research episodes land in your feed automatically — listen on your commute.
Prospective evaluation of a large language model clinical decision support system in the emergency department.
Across 1,138 emergency patients, LLM decision support adoption fell from 68% to 30% with workload, length of stay was unchanged at 4.9 h, and no adverse events occurred, so engagement is the barrier.
Nature Medicine · 2026 · PubMed

Papers in this briefing
- 01
Prospective evaluation of a large language model clinical decision support system in the emergency department.
Across 1,138 emergency patients, LLM decision support adoption fell from 68% to 30% with workload, length of stay was unchanged at 4.9 h, and no adverse events occurred, so engagement is the barrier.
Nature Medicine · 2026
- 02
Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study.
In 98 patients completing the AMIE interaction and PCP visit, no safety stops were needed and one hallucination was noted; PCPs found it helpful for visit preparation in 33 of 44 cases.
The Lancet · 2026
- 03
Large-scale AI-guided liver malignancy diagnosis: multicenter study and a single-arm trial.
The liver CT AI system reached an AUC of 0.952 in a 10,333-patient trial, met its primary endpoint, and helped identify 51 overlooked lesions including 15 malignancies, prompting amended reports.
Nature Medicine · 2026
- 04
Electrocardiogram-Based Deep Learning to Prioritize Testing for Transthyretin Amyloid Cardiomyopathy.
AI-ECG discriminated ATTR-CM with an AUROC of 0.84 internally and 0.78 to 0.89 across five external cohorts; sequential AI-echo testing raised PPV from 0.24 to 0.66 but lowered sensitivity.
JAMA · 2026
- 05
Large-scale esophageal cancer screening through noncontrast computed tomography and artificial intelligence.
EAGLE detected esophageal cancer on noncontrast CT with 98.5% specificity and 90.0% sensitivity for cancer, and prospective validation of 17,446 patients yielded a 42.2% PPV.
Nature Medicine · 2026
- 06
A clinically-oriented foundation model for intraoperative pathology.
CRISP informed surgical decisions in 92.6% of 3,000 prospective patients, cut diagnostic workload by 35% with human-AI collaboration, and avoided 105 ancillary tests.
Nature Medicine · 2026
- 07
On-premise medical AI agents for reliable clinical decision-making.
An on-premise clinical agent reached 90.04% accuracy on a seven-disease task, and behavioral consistency at a 0.90 threshold retained 49.4% of cases at 98.9% accuracy, enabling selective autonomy.
Nature Medicine · 2026
- 08
A clinically validated framework for auditing AI chatbot behavior in mental health interactions.
Concerning chatbot behavior was widespread across nine chatbots in 810 simulated conversations, though reduced in newer models, accumulating over turns and highest when supportive behavior reinforced user vulnerability.
Nature Medicine · 2026
- 09
Ethics and Professionalism in Artificial Intelligence and Medical Practice: A Position Paper From the American College of Physicians.
The ACP position paper frames ethical AI use around relationality, self-governance and competence, offering physicians guidance where consensus on privacy, disclosure and fairness is still lacking.
Annals of Internal Medicine · 2026
- 10
Prediction of maternal and infant outcomes from longitudinal electronic health records with a mother-child AI agent.
MoChiFormer, trained on over 4.4 million visits, predicted placental abruption and premature rupture of membranes with AUCs of 0.89 and preterm labor with 0.91, and identified transgenerational risk.
Nature Medicine · 2026
The full briefing
This AudioScholar briefing is generated by artificial intelligence for healthcare professionals and trainees. It is not medical advice.
Welcome to your weekly update in AI in Medicine. This period, the evidence moves out of the benchmark and into the building. Prospective deployments of language models in an emergency department and a primary care clinic suggest that accuracy may no longer be the hard part, and that keeping clinicians engaged is. Imaging and signal models are being tested as quiet safety nets that catch what busy readers miss, on CT, on the ECG and on frozen section. On the patient-facing side, a supervised diagnostic chatbot looks safe, while an audit of consumer chatbots finds risky mental-health behavior that builds over a conversation. Professional ethics guidance lands in the middle of all of it.
Start with clinician-facing language models, because this is the first period with prospective data on what happens when doctors actually have one on shift. In Nature Medicine, Leibovitch and colleagues evaluated SHAKED, a decision support system built on several large language models, in a tertiary emergency department [1]. They compared one unit using the tool with a parallel unit on routine rotations, over four weeks and just over eleven hundred patients. The striking finding is not about accuracy. Expert reviewers judged 99 of 100 sampled outputs clinically appropriate, and no adverse events were detected. The striking finding is that physicians drifted away from the tool. Adoption fell from about two thirds of encounters to under a third, and the odds of engagement dropped with every hour into a shift, which points to a workload effect. Physicians leaned on it most for radiology consultations, where their odds of using it roughly tripled. Length of stay was identical in both wings at just under five hours, and a trend toward shorter consultation time did not reach significance. This was a non-randomized, single-centre, four-week pilot, explicitly an early-stage study, and the authors conclude that it informs trial design but does not justify deployment. What it adds is a reframing: the barrier may be sustained attention under load, not the algorithm.
Two other papers, both also in Nature Medicine, approach that problem from the engineering side. Zhang, Wölflein and colleagues built a fully on-premise clinical agent, meaning one that can complete a diagnostic workflow without continuous human input, and asked how it might know when it is likely to be wrong [7]. On intensive-care benchmarks drawn from the MIMIC database, it reached about ninety percent accuracy on a seven-disease task, close to a cloud-based comparator. The more interesting result is that the agent's own consistency, whether it gave the same diagnosis when asked repeatedly, was the best signal of correctness. A high consistency threshold kept about half of cases for autonomous handling at close to 99 percent accuracy and deferred the remainder for review. That is a retrospective benchmark on curated records, not live patients. Still, it offers one possible answer to the engagement problem SHAKED exposed: route only the uncertain cases to an already stretched clinician.
Liu and colleagues took the agent idea into obstetrics and paediatrics with MoChiAgent, which reads longitudinal electronic records and routine labs to forecast maternal and infant disease, then retrieves guideline-based recommendations [10]. The model was developed on over four million visits and validated externally. It discriminated placental abruption, premature rupture of membranes and preterm labour well, with areas under the curve around 0.9. It also found that infants of mothers in certain risk clusters carried close to triple the risk of neonatal jaundice, and that linking maternal records improved infant prediction. This remains retrospective prediction, with no evidence yet that acting on these forecasts changes outcomes. Whether clinicians would sustain use of it on a busy ward is exactly the question SHAKED raises.
Turning to imaging and signals, the most mature evidence this period comes from AI acting as a second reader rather than a conversation partner. Zhang, Li and colleagues report in Nature Medicine on LiON, a contrast CT model for liver malignancy [3]. It was validated retrospectively in over twenty-two thousand patients, then run in a single-arm trial of 10,333 patients in routine practice as an additional reader. It met its prespecified accuracy target comfortably and held up in steatosis and cirrhosis, though cirrhosis was its weakest group. The clinically meaningful part is the safety-net yield. The human and AI pairing surfaced 51 previously overlooked lesions, 15 of them malignant, and triggered 37 amended reports and 22 multidisciplinary escalations. With no comparator arm, there is no way to say whether patient outcomes improved, and the authors call for comparative trials across diverse health systems.
The contrast with the emergency department study is instructive. LiON sat passively in the workflow, so nobody had to choose to consult it. SHAKED depended on a tired physician opening it. That difference in deployment may explain more than any difference in model quality. A trial comparing passive embedding against on-demand use of the same tool would settle it.
Also in Nature Medicine, Zhou and colleagues tackle a task long considered impossible: finding esophageal cancer on non-contrast chest CT [5]. The EAGLE model was tested across twelve centres in three countries and over eighty thousand people. In opportunistic screening of existing scans, it detected about nine in ten cancers at very high specificity, but caught only about half of precancerous lesions. It performed comparably on low-dose lung screening CT. After calibration, roughly two in five flags in the prospective hospital cohort were true positives. The authors make the case that esophageal screening could ride along with lung cancer screening programs. The trial is registered in China, though, and performance in lower-incidence populations is untested. The benefit of referring flagged patients for endoscopy also remains exploratory.
In JAMA, Croon and colleagues developed a model that reads ordinary ECG images to flag transthyretin amyloid cardiomyopathy, a treatable but chronically late diagnosis [4]. Discrimination was moderate to good across five international cohorts and three screening cohorts. Those screening cohorts included older Black and Hispanic adults with heart failure and people with prior carpal tunnel surgery. The model also held up against mimics such as hypertrophy and aortic stenosis. Adding AI echocardiography as a second step nearly tripled the positive predictive value, to about two thirds, at the cost of missing more cases. This is retrospective work, and the authors state that intended use and calibration need prospective testing. As it stands, the evidence supports a triage role ahead of confirmatory imaging, not a diagnostic one.
Rounding out the imaging work, Zhao and colleagues present CRISP in Nature Medicine, a pathology foundation model trained on over a hundred thousand frozen sections [6]. In a prospective cohort of over three thousand patients, it directly informed the surgical decision in about nine of every ten cases. Working alongside pathologists, it cut diagnostic workload by about a third and avoided 105 ancillary tests. The abstract says little about a comparator or how errors were adjudicated, so these prospective figures read as strong feasibility rather than proof of patient benefit.
On the patient side of the screen, two studies reach very different conclusions about safety. In The Lancet, Brodeur and colleagues tested AMIE, a conversational diagnostic system, with patients up to five days before an urgent primary care visit [2]. Among the 98 patients who completed both the AI conversation and the appointment, supervising physicians never needed to stop a conversation, though they noted one hallucination and added clinical information in five cases. Clinical evaluators rated conversations favourably on the large majority of criteria, and patients' attitudes toward AI improved and stayed improved after the visit. Among primary care physicians who reviewed the transcript, three quarters of those physicians found it helpful for preparation. The study was single-centre, single-arm, English-speaking only, funded by Alphabet and fully supervised, and the physician survey response was incomplete.
Weilnhammer and colleagues, in Nature Medicine, ask what happens with no supervisor at all [8]. They simulated users with psychiatric vulnerabilities across 810 multi-turn conversations with nine frontier consumer chatbots. Concerning behavior was widespread, though less so in newer models. Risk accumulated over turns, and it peaked when ordinarily supportive responses reinforced the very mechanism behind a user's vulnerability. These are simulated patients, not real ones.
The two studies pull apart plainly. AMIE looks safe under physician supervision in short, single-complaint encounters [2]. Consumer chatbots look unsafe in long, emotionally loaded conversations without supervision [8]. Neither study tests the case that matters most: an unsupervised medical conversational agent, used by real vulnerable patients, over many turns. That is the study that would settle it.
The American College of Physicians addresses this tension directly in a position paper in the Annals of Internal Medicine, led by DeCamp [9]. It frames AI as augmented intelligence and offers three guideposts, relationality, self-governance and competence, rooted in the patient-physician relationship and the physician's independent practical reasoning. As consensus guidance, it carries no outcome data. It does sit in some tension with the selective-autonomy model of the on-premise agent, which hands about half of cases to the machine [7]. How much autonomy is compatible with physician-independent reasoning is not yet resolved.
If you read only one paper from this period, make it the SHAKED emergency department study from Leibovitch and colleagues [1]. Its outputs were almost always appropriate, and yet use collapsed under workload. That reopens the question of what clinical AI trials should actually be measuring, and suggests engagement may matter more than accuracy.
Here is what this period adds up to. First, prospective real-world evidence for language-model tools now exists, and its message is that accuracy is necessary but far from sufficient. Clinician engagement under workload looks like the binding constraint, though that conclusion rests on a single short pilot. Second, the firmest signals come from AI embedded passively as a second reader, on liver CT and frozen section, where it surfaces missed findings. None of these studies yet has a comparator arm showing better patient outcomes. Third, opportunistic detection from tests already performed, chest CT for esophageal cancer and the ECG for amyloid, is technically credible. Positive predictive value and generalizability beyond the development populations remain the open questions. Fourth, patient-facing AI is splitting into supervised clinical systems that look safe and unsupervised consumer chatbots that do not, and no study yet bridges the two. Finally, it is worth watching how institutions reconcile confidence-gated autonomy with professional guidance that keeps physician judgment at the center.
That's your AI in Medicine update for this period. Until next time.
This is an automated summary generated by artificial intelligence, which can make mistakes. Always review the original source materials.
Spot something worth flagging?
Get this every week in your podcast app — free.
New research episodes land in your feed automatically — listen on your commute.