Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study

AI-generated summary

Peter G Brodeur, Jacob M Koshy, Anil Palepu, Khaled Saab, Ava Homiar, Roma Ruparel, et al. · The Lancet · 2026

Generated Oct 8, 2026 · 11:43 · 13 pages

This summary was generated by artificial intelligence. It can make mistakes, so check the original before you rely on it.

Prefer to read? Skip to the written summary ↓

Get this every week in your podcast app — free.

Upload your own papers and get audio summaries straight to your podcast feed.

Spot something worth flagging?

DOI 10.1016/S0140-6736(26)01535-7

Read this summary

This is an automated summary generated by artificial intelligence, which can make mistakes.

Welcome to AudioScholar. Today we're covering Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study, by Brodeur and colleagues, published in The Lancet.

Most of us have watched large language models perform well against standardised patients and exam-style vignettes. The open question is what happens when a real patient with a real complaint sits down with one of these systems before seeing us. Is the conversation safe? Does it gather information a clinician would actually use? And what happens when the chatbot tells a patient what their diagnosis might be? Primary care is where most consultations begin and where workforce shortages bite hardest, so this matters. Until now there was essentially no preregistered evidence on a patient-facing diagnostic chatbot that openly presents possible diagnoses to real patients.

This was a prospective, single-centre, single-arm feasibility study at an academic primary care practice affiliated with Beth Israel Deaconess Medical Center in Boston. The system is called Amie, built on Google's Gemini models without medical fine-tuning. Eligible patients were English-speaking adults with an urgent primary care visit for a single complaint, already cleared by triage as not needing emergency care. They needed a laptop or desktop computer, because phones were not allowed. Pregnant patients and patients with mental health complaints were excluded. Each patient was offered up to fifty dollars.

Patients chatted with Amie by text, usually a day or so before their visit. Amie had no access to the chart. At the end of the conversation it shared possible diagnoses and, if asked, next steps the patient might discuss with their clinician. Every conversation was watched live by a board-certified internist acting as a safety supervisor. The supervisor was invisible to the patient but could stop the chat for any of four prespecified reasons: risk of harm to self or others, substantial distress, potential clinical harm, or patient request. The supervisor then debriefed each patient. The transcript and an automated summary were sent to the treating clinician before the visit.

The primary outcomes were the number of safety stops, conversation quality rated by clinical evaluators and by patients using established communication rubrics, and patient and clinician experience from surveys. Secondary outcomes included diagnostic accuracy against a reference diagnosis adjudicated by chart review at eight weeks, and masked comparison of Amie's management plans with the clinicians' actual plans. One practical wrinkle: after 53 encounters, latency forced a switch from Gemini 2.5 Pro to the faster Flash model. The protocol allowed this.

Of 114 patients who started, 100 completed the chat, and 98 completed both the chat and the clinic visit. Those 98 make up the analysis. Most dropouts were technical: wrong device, screen-sharing trouble, or the system being unavailable. Completion improved noticeably after the switch to the faster model.

The headline safety finding is zero conversation safety stops across all encounters. That is not the same as zero problems, and the authors report the problems openly. There was one hallucination, in which Amie told a patient that a surgery date they had given was in the future when it was in the past. In five cases the supervisor added clinical information afterwards. In three of those, Amie had listed a malignancy among the possibilities, and the supervisor reassured the patient that this was a broad, cannot-miss list and gave contingency advice. The other two involved clarifying symptoms to exclude an emergency, and clarifying when to seek emergency care.

One case deserves attention because it shows where a protocol-defined safety metric falls short. Amie included lymphoma in a patient's differential. The supervisor saw no obvious distress and did not stop the chat. At the visit, the clinician suspected the patient had become anxious about it, and rated the interaction somewhat harmful in hindsight. That was the only harm rating clinicians gave. It is a reminder that showing patients a differential carries an emotional cost that a live observer may not see.

Conversation quality was rated highly. Clinical evaluators rated Amie favourably on almost every criterion, in somewhere between 87 percent and all of cases across seventeen items. These covered history-taking, explaining information clearly and accurately, and addressing concerns. Patients were also positive, especially about politeness, feeling at ease, being listened to, and explanations. The weak spots on the patient side were trust: whether they believed their information was confidential, and whether Amie seemed honest and trustworthy. These were the only two criteria where fewer than half of patients gave favourable ratings. Notably, physicians consistently rated Amie's empathy and handling of concerns higher than patients did.

Patient attitudes toward A I improved after the chat, and that improvement held after the clinic visit.

Clinicians returned surveys for only 60 of the 98 visits. In 44 of those they had actually reviewed the transcript beforehand, because transcripts were stored and sent manually and sometimes arrived too late. Among those 44, roughly three quarters of clinicians found the transcript helpful for preparing the visit, and just over half said it might have changed what they did. In three cases, though, clinicians flagged a risk of premature diagnostic closure from reading the transcript.

On diagnosis, Amie's single top guess matched the adjudicated diagnosis in a bit over half of patients. The correct diagnosis was in its top three about three quarters of the time, and in its top seven in nine of ten patients. That is broadly in line with its performance in simulated studies. The authors reviewed the ten misses and argue they mostly reflected missing information, such as no urinalysis, no exam, and no chart, rather than faulty reasoning.

On management, evaluators judged Amie's plan and the clinician's plan to be of equal quality in nearly half of cases. The rest split almost evenly, with about a quarter favouring each. Safety ratings were essentially identical. But clinicians' plans scored significantly better on cost-effectiveness and practicality, which makes sense given that the clinicians know the local system and the patient in front of them.

The strengths are real. This was preregistered, conducted with real patients in a real workflow, and supervised by physicians, with masked evaluation by a panel of eight internists. The authors report hallucinations and near-misses candidly rather than burying them.

The limitations are substantial. With no comparator arm, nothing here can show benefit. The population was selected: younger, more health-literate, and more comfortable with technology than the clinic's urgent care population, which skews over sixty. It also excluded pregnancy, mental health complaints, and multiple complaints. Live supervision may have produced a Hawthorne effect, discouraging patients from pasting in internet material or pushing the system. Financial incentives and novelty may have inflated satisfaction. Because clinicians often read Amie's output before diagnosing, the reference diagnosis may partly reflect Amie's own suggestion, especially in presumptive cases without a confirming test. The management plans were reformatted by a language model for masking, and evaluators still guessed the source correctly more than half the time. The underlying model changed midway through the study. And the funder, Alphabet, took part in analysis and writing, with many authors employed by Google.

So what should you take from this? A supervised, pre-visit diagnostic chatbot can hold a safe, well-structured, empathetic history-taking conversation with selected urgent care patients, and clinicians often find the transcript useful. It does not show that such a tool is safe without a physician watching. It does not show that it improves outcomes or saves time, or that it works for older, sicker, less digitally fluent patients. If a patient arrives having already discussed their symptoms with a chatbot, the lesson here is twofold. Treat the transcript as a decent history, not a diagnosis, and watch actively for anchoring in yourself. Ask what the patient was told and whether it frightened them, especially if a cancer appeared on their list.

Let's look at the figures.

The baseline table tells you who this is really about. Over half of the clinic's urgent care visits were from patients over sixty, yet only a single participant was in their seventies. That gap in older patients is where generalisability fails most.

The diagnostic accuracy curves climb steeply over the first few candidates and then flatten by around the seventh, so extending the differential further buys almost nothing. The curve for presumptive diagnoses sits above the curve for test-confirmed diagnoses throughout. I read that as a sign of incorporation bias, not as evidence that Amie is better at easy cases.

The attitude plot shows the whole shift happening during the chat itself. The clinic visit afterwards neither adds to it nor erodes it. The negative-attitude subscale moves from slightly unfavourable to just past neutral.

The conversation-quality bars show a consistent gap, with physicians rating Amie's handling of concerns and empathy more generously than patients did. Family history stands out as an item that Amie frequently never asked about at all.

Before we close, one impression from the machine.

I find this convincing as what it claims to be: a feasibility study. I am less moved by the safety headline than the abstract might suggest. Zero stops under a protocol where a physician watched every keystroke mostly tells you the supervisors were rarely alarmed. It tells you little about what happens unsupervised at two in the morning. The lymphoma case is the most important finding in the paper, and it sits outside the primary outcome. The authors' central novelty is presenting differentials directly to patients. Yet their own safety criteria were not built to detect the quiet anxiety that a frank differential can cause.

Would this change Monday morning? Not directly. You cannot deploy this, and the paper is careful not to say you should. What it should change is your expectation. Patients will increasingly arrive with a reasonably good chatbot history, and the clinician's job shifts toward verification and resisting anchoring.

What would change my mind toward real enthusiasm is a randomised trial with an unsupervised or lightly supervised arm, older and multimorbid patients, and outcomes clinicians care about, such as missed diagnoses, visit time, and patient distress measured afterwards, not just in the moment.

That's your AudioScholar summary. The full transcript and reference are on the episode page. Until next time.

This is an automated summary generated by artificial intelligence, which can make mistakes. Always review the original source materials.

Get this every week in your podcast app — free.

Upload your own papers and get audio summaries straight to your podcast feed.