Published
April 21, 2026

Researchers find that radiologists prefer custom AI-generated and human-created impressions over generic AI-generated impressions.
Oncologic imaging reports are often long and detailed. Think 40 CTs per day at about 500 words each, you’re at 100,000 words over a five-day workweek. That’s longer than reading “Life of Pi” or “The Hobbit” – yes, seriously! Imagine that every week, for years.
Not only is this time consuming, it may also contribute to the 44-65% burnout rate seen in radiology, which is leading to the early retirement of many radiologists and the slowing entry of new radiologists into the field.
Recent AI advancements enable automated generation of AI impressions. While initial evaluations have been promising, the published quality measures have been limited, with no publications to date on complex studies and usefulness of abdominal imagers.
“Radiology reports are cognitively demanding, and the impression requires an additional layer of synthesis and judgement,” said Andrew Del Gaizo, MD, CMIO at Rad AI. “That combination contributes meaningfully to workload and burnout, and is exactly where AI has the potential to help.”
In 2026, a group of researchers, including Dr. Del Gaizo, set out to evaluate the quality, safety and clinical utility of AI-generated radiology impressions compared to human-authored impressions across multiple clinical stakeholder groups.
They took a retrospective sample of 200 de-identified CT oncologic reports dictated by four abdominal radiologists — 50 per radiologist. Impression types:
The original radiologist evaluation was blinded to the impression source, and they were rated on completeness, correctness and conciseness. All reviewers were assessed for clarity, clinical utility and likelihood of it leading to patient harm. Quantitative metrics and reliability included word count, item count, BLEU score, generation time and inter-rater reliability (ICC).

“We wanted to move beyond anecdotal impressions and rigorously evaluate how AI-generated impressions perform in real clinical contexts, across the different stakeholders who rely on them,” said Dr. Del Gaizo.
The study resulted in a variety of findings, including:

So, what did the radiologists and oncologists think? Radiologists preferred human and custom AI over generic AI (p < 0.001). Oncologists were more evenly split, with a slight preference for custom AI (p=0.21-0.63).

Custom and generic model impressions demonstrated near parity with human-authored impressions across most quality metrics. The exception: Generic models impressions were significantly less concise. Patient harm ratings were uniformly low.

When it comes to studies in the field, clinical impact is paramount. It’s what matters most to radiologists and, ultimately, impacts patients. Here’s what researchers concluded.
See the full study for more details.
Learn about Rad AI Impressions.
Citation
Phadke, S., Suresh, N., Allen, Z., Balagopal, A., Chan, S., Shah, A., Winter, M., Lam, C., Rose, T., Araujo, C., Ahmed, A., Imanirad, I., Berland, L., & Del Gaizo, Al. (2026). Comparison of AI generated radiology impressions: a multi-stakeholder evaluation. npj Digital Medicine. https://www.nature.com/articles/s41746-026-02586-6
Expand Your Expertise
The Readout explores the technologies, challenges and realities shaping where care goes next.
Radiology-First AI
Actually Works