AudioScholar

AI in Medicine Watch — Sep 24, 2026

Generated Sep 25, 2026 · 7:20

The week's practice-changing research, summarized for clinicians.

If the audio fails to play, refresh the page to renew the link.

Prefer to read? Skip to the papers and the full briefing ↓

Get this every week in your podcast app — free.

New research episodes land in your feed automatically — listen on your commute.

Papers in this briefing

  1. 01

    Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC.

    Nature Medicine

    PMID 42733093

  2. 02

    Deep Learning-derived Bone Mineral Density and Longitudinal White Matter Microstructure and Cognitive Decline: Multi-Ethnic Study of Atherosclerosis.

    Radiology

    PMID 42770818

  3. 03

    Opportunistic Chest CT Paraspinal Muscle Biomarkers to Predict Longitudinal Thoracic Vertebral Bone Mineral Density Loss: Results from MESA.

    Radiology

    PMID 42742379

  4. 04

    Machine learning-driven spleen imaging and genomics uncover a splenic connection to coronary artery disease.

    Science Translational Medicine

    PMID 42715348

  5. 05

    Multimodal protective and susceptibility clusters in paediatric atopic dermatitis: A machine learning-based, data-driven observational study.

    PLOS Medicine

    PMID 42709851

The full briefing

This AudioScholar briefing is generated by artificial intelligence for healthcare professionals and trainees. It is not medical advice.

Welcome to your weekly update in AI in Medicine. Three threads run through this period's papers. First, multimodal fusion — the idea that stacking imaging, pathology and genomics onto clinical data will outperform clinical data alone — gets its most serious test yet, and the answer is more equivocal than the field expected. Second, opportunistic imaging: two cohort studies show routine chest CT carrying signals about bone, muscle and even brain aging that nobody ordered the scan to find. And third, running underneath both, the widening gap between how these models perform in their own test set and how they behave when they meet someone else's patients and someone else's scanner.

Start with the question of whether more data modalities actually help. In Nature Medicine, Prelaj and colleagues report I3LUNG, the largest international real-world artificial intelligence study in non-small cell lung cancer, enrolling nearly 2,400 patients to predict who benefits from immunotherapy. They built early- and intermediate-fusion models combining clinical and blood data, CT, digital pathology and genomics. The headline finding is двух-edged. Models using clinical and blood variables alone reached discrimination approaching 0.77 in the held-out test set, and significantly outperformed everything currently used at the bedside — PD-L1, performance status, neutrophil-to-lymphocyte ratio, lactate dehydrogenase and the Lung Immune Prognostic Index. In a prospective usability study, both lung cancer experts and non-experts made better predictions when assisted by the explainable version of that tool. But the multimodal models, despite looking stronger in development, did not deliver a reliable incremental benefit in either the test set or external validation, where performance fell considerably. This is real-world retrospective data with a modest prospective usability component, and the authors are explicit that prospective validation in more than two thousand patients is still running. What it supports for now is that a simple, explainable, routinely-available-data model already beats the biomarkers oncologists are using — not that fusion earns its complexity.

Set that against PLOS Medicine, where Zhakparov and colleagues took the opposite lesson from integration. In 217 AmaXhosa children in South Africa, rural and urban, they applied explainable machine learning across environmental, cytokine, antibody and transcriptomic data in paediatric atopic dermatitis. Integration identified three multimodal clusters — one protective, dominated by rural environmental exposures correlating with plasma cytokines and autophagy-related gene expression, and two susceptibility clusters, one transcriptomic, one driven by correlated immunoglobulin E and the cytokines MCP-4 and TARC. So where Prelaj found fusion added little, Zhakparov found structure visible only in combination. The disagreement is real but partly about purpose: one study is optimising prediction, the other generating mechanism. And the atopic dermatitis work is explicitly exploratory, in a single population, with no independent validation cohort. What would settle it is prospective testing with a pre-specified comparison of the multimodal model against the single-modality one — exactly what I3LUNG's ongoing trial is designed to do.

Turning to what routine scans already contain. Two Radiology papers, both secondary analyses of the Multi-Ethnic Study of Atherosclerosis, make the case for opportunistic biomarkers. Hathaway and colleagues used deep-learning three-dimensional segmentation of paraspinal muscle from T1 to T10 on noncontrast chest CT in just over 1,300 participants, with segmentation agreeing closely with manual reference. Low muscle attenuation — fatty infiltration — was the strongest predictor of vertebral bone density years later, outperforming FRAX without bone density, and it improved discrimination for incident vertebral fracture beyond bone density alone. The evidence supports muscle quality on a scan ordered for other reasons carrying independent fracture information.

Momtazmanesh and colleagues pushed the same idea further, extracting deep-learning-derived thoracic vertebral bone density in 715 participants and linking it to longitudinal brain MRI and cognition. Lower baseline bone density tracked with faster white matter hyperintensity accumulation in the corpus callosum, faster microstructural decline in the anterior limb of the internal capsule, and faster global cognitive decline, with the effect amplified in participants with diabetes. Be careful here: the whole-brain white matter signal did not survive correction for multiple comparisons. These are modest, observational associations in an older cohort, not a causal claim and not a screening test.

Which brings us to the translation problem stated most bluntly. In Science Translational Medicine, Kamineni and colleagues used deep learning to extract 107 splenic radiomic features from abdominal MRI in over 42,000 UK Biobank participants, found ten associated with coronary artery disease, and linked them genetically to 219 loci including 9p21 — the strongest and most mechanistically obscure coronary locus known. Internally consistent, genuinely novel biology. But external validation in a clinical biobank of under three thousand patients faltered, and the authors attribute that to imaging protocol variability and patient heterogeneity. That echoes the external drop in I3LUNG, and the authors of both are unusually candid about it.

If you read only one paper from this period, make it the I3LUNG study in Nature Medicine. It is the first large study to show physicians measurably improving with an explainable tool, and simultaneously to reopen the question of whether multimodal fusion is worth its cost.

Here is what this period adds up to. Explainable clinical-and-blood models now have evidence of outperforming established immunotherapy biomarkers and of improving physician judgement — though from a single programme, with prospective confirmation pending. The incremental value of multimodal fusion remains unproven for prediction, even as it looks useful for mechanism. Opportunistic biomarkers from routine chest CT are accumulating observational support across bone, muscle and brain, but none of it is causal and none is in guidelines. And the most consistent finding across all five papers is that external validation is where these models lose ground — protocol variability and population difference, not algorithm design, are the rate-limiting step.

That's your AI in Medicine update for this period. Until next time.

This is an automated summary generated by artificial intelligence, which can make mistakes. Always review the original source materials.

Spot something worth flagging?

Get this every week in your podcast app — free.

New research episodes land in your feed automatically — listen on your commute.