Tag: artificial intelligence

Real World Test of AI Clinical Support Tool Improved Clinician Decisions

Trial did not show statistically significant difference for patient outcomes but helped clinicians improve quality of notes and recommendations

A large real-world clinical trial has found that a generative AI-powered support tool used to support frontline clinicians was safe and improved the quality of clinical decision-making but did not significantly change short-term patient outcomes.

The study, published today in Nature Medicine is one of the first randomised controlled trials worldwide to test whether generative AI can improve patient-level outcomes, rather than just clinician performance or simulated cases.

The trial involved more than 9600 patients attending 16 primary care clinics in Kenya, and was delivered by experts at the University of Birmingham supported by the National Institute for Health and Care Research (NIHR) Biomedical Research Centre: Birmingham.

What this study shows is that AI can be integrated safely into real clinical workflows, without undermining patient trust or clinician autonomy – which is a critical foundation for any future impact.

Alastair Denniston, Chair of Regulatory Science and Innovation

Clinicians were randomly assigned to use an electronic medical record system with or without an integrated AI consult tool that provided real-time diagnostic and treatment suggestions. The AI system, known as ‘AI Consult’, was a large language model–based clinical decision support tool embedded directly within the existing electronic medical record system.

During consultations, the tool worked in the background by:

  • Analysing information entered by the clinician into the medical record
  • Generating context‑specific diagnostic and treatment suggestions, aligned with Kenyan national clinical guidelines
  • Flagging potential concerns using a simple colour‑coded alert system (green, yellow or red)

Clinicians retained full autonomy; they were not required to follow the AI’s advice, and retained responsibility for all diagnosis, prescribing and referral decisions. The AI interface was not visible to patients, helping preserve normal patient–clinician interaction.

Senior author Professor Bilal Mateen, Honorary Professor of Machine Learning for Health at the University of Birmingham, and Chief AI Officer at PATH, said: “This is one of the first studies to rigorously ask the hardest question about AI in healthcare: whether it actually improves outcomes for patients.

“What we found is reassuring but also sobering. The technology appears safe and clearly improves aspects of clinical decision-making, but translating those gains into measurable patient benefit is much more challenging, particularly in everyday primary care.”

Serious outcomes such as hospitalisation or death are rare in primary care, meaning extremely large studies – potentially involving more than 100 000 patients – would be needed to detect modest effects.

Professor Alastair Denniston, co-author, Professor of Regulatory Science and Innovation at the University of Birmingham and lead for health data research at the NIHR Biomedical Research Centre: Birmingham, said: “A large part of primary care is to deal with common conditions, including those that are self-limiting, where many patients require low levels of healthcare intervention. In that context, even meaningful improvements in clinical reasoning may only result in small changes in patient outcomes that are very difficult to measure.

“What this study shows is that AI can be integrated safely into real clinical workflows, without undermining patient trust or clinician autonomy – which is a critical foundation for any future impact.”

Findings: safety, quality and costs

Researchers found no statistically significant difference in treatment failure within 14 days between patients seen with AI-supported care and those receiving standard care (2.2% vs 2.0%). The study found no evidence of harm, with similar rates of hospitalisation and death in both groups.

While the AI tool did not produce measurable improvements in short-term patient outcomes, it significantly improved the quality of clinical documentation and treatment planning, as assessed by an independent panel of experienced clinicians who were blinded to whether AI had been used.

Patient satisfaction was the same in both groups, suggesting that AI support did not alter patients’ experience of care.

The study also found that, although overall antibiotic prescribing rates were similar, antibiotic‑related costs were lower in the AI‑supported group, due to more cost-conscious prescribing choices.

Although the trial was conducted in Kenya, the researchers emphasise that the findings have global relevance, including for high-income health systems.

Professor Richard Riley, Professor of Biostatistics at the University of Birmingham and senior author, said: “Robust trials like this are so important to establish the real impact of using AI in practice. They help set realistic expectations of what AI can actually contribute within existing care pathways, and helps guide where future investment and research effort should be focused. Generalisability of our findings to higher-income settings, where baseline standards of care are already high, needs to be evaluated.”

Source: The University of Birmingham

Incorrect AI Advice is a Blind Spot – Even for Doctors

 New study highlights potential challenges for using automated tools in healthcare

Photo by Accuray on Unsplash

In experiments in which physicians made decisions about treating hypothetical patients, the physicians tended to trust incorrect advice presented as being generated by artificial intelligence (AI), even after given the opportunity to notice that patient recovery data contradicted the recommendations. Aranzazu Vinas of the University of the Basque Country, Spain, and colleagues present these findings in the open-access journal PLOS Digital Health

AI systems can help physicians categorise patients according to their different care needs, such as whether a patient is more or less likely to benefit from a certain treatment. Since these systems are not perfect, they are meant to be used as suggestions, with potential errors caught and corrected by physicians.

 Prior research has shown that, in general, people struggle to notice and correct mistakes made by AI. To explore how this challenge may extend to physicians, Vinas and colleagues analysed data from 223 physicians who anonymously participated in online experiments.

 The physicians were asked to imagine they had the option to treat patients for a rare disease using a not-yet-proven treatment still under development. They were told that an AI system had identified which patients were more or less likely to benefit from the treatment. The physicians then chose which patients to treat, and after being presented with data on patient recovery, rated their perceptions of how reliable the AI was.

 Crucially, the actual effectiveness of the hypothetical treatment did not align with the AI recommendations. In one experiment, the treatment was equally moderately effective for all patients, and in a second experiment, it was equally ineffective for all.

 However, in both experiments, the physicians tended to rate the AI system as reliable and apparently did not use the patient recovery data to conclude that the AI recommendations were incorrect. In the second experiment, the physicians did not realise that the treatment was entirely ineffective.

These findings highlight potential challenges for incorporating AI-based classification into healthcare. Future research could build on this study, such as by developing and testing strategies and protocols that could increase human critical thinking and detection of AI errors, in order to maximize the benefits of the human-AI collaboration while minimising potential errors.”

Lead author Aranzazu Vinas notes: ” In both experiments, physicians mostly trusted the AI’s classifications and had trouble learning from the feedback. Furthermore, in the second experiment, professionals did not notice that the treatment was completely ineffective.”

 Co-author Helena Matute adds, “People tend to say that there is always a human controlling the algorithm, but our experiments show that doctors (as well as anyone else) have problems in learning from the available evidence when it contradicts the suggestions of an algorithm.”

Co-author Fernando Blanco summarizes: “It is important to investigate the errors that humans (including doctors) make when working with algorithms, in order to learn how to minimize the problems that arise from them.”

Press Preview: https://plos.io/4wjPxSs

In your coverage please use this URL to provide access to the freely available article in PLOS Digital Health: https://plos.io/4blGKHA

Contact: Aranzazu Vinas, aranzazu.vinas@ehu.eus

Image Caption: Doctors working on an AI-support system

Image Credit: Photo by Accuray on Unsplash. Free to share under the Unsplash license.

High-Resolution Image Link: https://unsplash.com/photos/a-few-men-looking-at-a-computer-screen-S34fEzWT6eE

AI Mistakes Can Cost Doctors Time when Writing to Patients

Errors and irrelevant details mean physicians may spend more time editing AI-drafted responses than it would take to write them, a large study of an online patient portal shows

Photo by National Cancer Institute on Unsplash

Artificial intelligence is spreading rapidly in health care, with the goal of streamlining critical but onerous clerical tasks such as note-taking and charting so that physicians and nurses can devote more time to patients. But even when AI can free up doctors to correspond with patients, it may fall short in helping them do it by introducing errors and extraneous details into their messages, according to a new Dartmouth study presented at the 2026 Annual Meeting of the Association for Computational Linguistics and published in the conference proceedings

The result is that physicians may spend more time editing responses than it would’ve taken to write them, the researchers report.

“We find that AI can sound like a doctor but not think like one,” says Sarah Preum, an assistant professor of computer science and the study’s co-corresponding author with Parker Seegmiller, a graduate researcher in Preum’s PersistLab at Dartmouth. 

The researchers conducted the first large-scale study of an online patient portal that uses AI to draft responses from physicians to patients. The team developed a tool that compares AI-generated replies to a dataset of real responses they developed with health care professionals from Dartmouth Health

They then analysed 146 000 conversations between 10 105 patients and their primary care physicians at the large rural health system, with data anonymised.

The researchers also used their tool to evaluate physician responses drafted by Claude, Gemini, and ChatGPT, as well as the three smaller commercial platforms, Llama, Aloe, and Qwen. 

“We find that AI can sound like a doctor but not think like one.” 

Sarah Preum, corresponding author and assistant professor of computer science

The team reports that AI-generated answers frequently misalign with what clinicians would actually write. This includes automated responses that are too long, don’t ask follow-up questions, and use irrelevant or inaccurate medical details.

“There are smaller studies that say, ‘Oh, AI is amazing,’ but we realised there is a gap in the existing literature of a large-scale evaluation of this technology,” Preum says. “We didn’t just want to measure a platform’s accuracy, but whether it actually helps with the workload, which in this case is measured by how much editing the physician is doing.” 

For example, the portal’s AI suggested telling a 32-year-old woman who is taking an acid reflux drug and was concerned about constant nausea that the medication might take some adjustment in diet. A physician replaced that by asking if there’s any chance she was pregnant. 

Even little changes can add up over hundreds or thousands of messages, Preum says. “You don’t want to integrate large language models into the workflow and just shift the bottleneck so that doctors are devoting their cognitive energy to playing AI janitor and fixing mistakes,” Preum says. “But if we’re not careful, that’s a likely outcome.”

The researchers show, however, that adapting AI to how individual physicians communicate can improve accuracy by 33% and reduce editing by 26%. 

“If message generation is really efficient and high quality, if it asks the right things, then it really has potential to improve efficiency,” says co-author Tim Burdick, an associate professor of community and family medicine in Dartmouth’s Geisel School of Medicine and a family medicine physician at Dartmouth Health. 

“I don’t foresee a time when the portal can respond to a patient without a clinician editing it first. But as we make the models better, we’ll be able to address portal messages much more quickly and with less mental energy,” Burdick says. 

The study shows that there are such things as “good” AI responses and provides a framework for implementing them into patient-physician portals, Preum says. These platforms are increasingly common among large health care systems and often customized, she says.

“That took us a long time to figure out, but if you’re trying to measure how effective this technology is, you need to define what a good response is,” she says. “We can only improve what we can measure and objectively evaluate.” 

The researchers created a technique called TADPOLE (Thematic Agentic Direct Preference Optimization for Learning Enhancement) that trains AI platforms using the hybrid model they constructed from physician- and AI-generated responses.

They plugged TADPOLE into the six commercial LLMs and found that drafted responses better matched physicians’ standards for precision and information quality. “That could save a busy clinician an hour or two of work a day,” Burdick says. 

Doctors and nurses today are inundated with messages from patients and caregivers who can write them online anytime, he says. An ongoing project between Burdick and the Preum Lab called PortalPal aims to streamline patient portals using AI, including by automating some steps in following up with patients to get more information.

“We’re still nowhere near the point of having clinicians removed from the workflow.”

Tim Burdick, co-author and associate professor of community and family medicine

Physicians who Burdick works with say that AI-generated drafts save about 25% of their time on shorter messages. “It’s easier to make small edits to an LLM-generated message than it is to write it from scratch,” he says. But longer drafts can include information that is not correct or accurate. 

“If you have to edit 75% of the message, you may be spending more time and energy on making changes than if you were to just write it from scratch,” Burdick says. “I would guess we need to get to where the physician is editing less than 30% of the content before it has substantial benefit.”

One advantage of AI’s verbosity is that it tends to be more empathetic and thorough than physicians crunched for time, the researchers find. For example, AI is more likely to tell a patient experiencing an upset stomach that it’s sorry to hear they’re feeling nauseated. 

This means AI could be used to help “nudge” doctors to show more understanding and care for the patient’s situation, or answer patient’s questions more effectively so that patients feel more heard, Preum says. The team produced sample responses such as showing empathy by praising patients for following a treatment plan (“You’ve been doing a great job with your tapering.”) or planning for changes in symptoms (“If you’re feeling dizzy, please call triage.”).

The researchers also find that 65% of all the portal messages they studied came from people over 55, with patients over 65 generating 24% of all messages. These figures suggest that patient portals in general should be designed to accommodate older people, Preum says.

Future work will study how much actual time doctors spend editing automated drafts. The team also plans to evaluate their training model TADPOLE from the user perspective, studying if and how it lightens a physician’s workload, and how doctors and patients rate its performance. 

“This is one of the first studies that uses real patient portal messages to establish a generative AI model. In that regard, it’s innovative and shows us that this is not a simple task,” Burdick says. “We’re still nowhere near the point of having clinicians removed from the workflow.”

Source: EurekAlert!

Study Shows AI Can Help Clinicians Identify Brain Tumour Risks

By Katelin Shaft

Mayo Clinic researchers and collaborators have shown that an artificial intelligence (AI) tool can analyse routine pathology slides to help clinicians classify meningiomas, the most common primary brain tumour in adults, and better understand a patient’s risk of tumour recurrence.

The study, published in The Lancet Digital Health, demonstrates that deep learning models can support the extraction of molecular and prognostic information from standard haematoxylin and eosin, or H&E, slides – the same type of tissue images already used in routine clinical care. These insights are typically obtained through DNA methylation profiling, an advanced genetic test which provides valuable diagnostic and prognostic information but can be costly, time-consuming and is unavailable in many hospitals.

“This is one of the many studies where we can harness the strength of digital pathology by capturing the last two decades of genomic and molecular knowledge into AI algorithms,” says Gelareh Zadeh, MD, PhD, chair of the Department of Neurologic Surgery at Mayo Clinic in Rochester and Chief Medical Officer for Mayo Clinic Platform.

Making advanced tumor insights more accessible

Meningiomas can vary widely in behaviour. Some grow slowly and may never return after treatment, while others are more aggressive and more likely to recur. Understanding that risk is critical for patients and care teams deciding whether additional treatment, such as radiation therapy, may be needed after surgery.

Molecular testing can help identify which tumours are more likely to recur and which may respond differently to treatment. But these tests require specialized technology and expertise, limiting access for many patients.

Using tissue samples, pathology images and clinical data from 672 patients, researchers developed and tested AI models designed to help identify patterns linked to a tumour’s biology. Drawing on multiple de-identified datasets, including data resources from Mayo Clinic Platform, the models supported classification of meningioma subtypes and recurrence risk prediction using standard pathology slides that are already part of routine patient care.

The findings suggest that, with further validation, AI-based tools could one day help clinicians obtain more detailed tumour information to inform patient care, without requiring every patient to undergo advanced genetic testing.

Helping guide treatment decisions

For patients with meningiomas, recurrence risk can influence follow-up care, imaging frequency and whether radiation therapy should be considered. The study found that AI-based predictions remained useful even after accounting for traditional clinical factors such as tumour grade, the extent to which surgery was able to remove the tumour and patient age.

Researchers also found that the AI models could identify patterns of tumour heterogeneity – differences within the same tumour – that may help explain why some tumours behave more aggressively or respond differently to treatment.

The researchers note that additional prospective studies are needed before the AI models can be used routinely in clinical care. Still, they say the findings lay the groundwork for more accessible, personalised care for patients with meningiomas – and potentially for similar AI approaches in other cancers.

As with any clinical decision-support tool, the researchers emphasise that these models would require rigorous evaluation, validation and ongoing physician oversight before being considered for routine care. “The aim is to make these algorithms readily and simply accessible for use globally, improving patient care across many healthcare settings,” says Dr Zadeh.

For a complete list of authors, disclosures and funding, review the publication.

Source: Mayo Clinic

AI Language Models Struggle with Basic Hospital Data Tasks, Study Finds

Nine leading AI models were tested on simple administrative queries drawn from real-world emergency department records—and most failed unless paired with code-generation tools.

A new study finds that large language models (LLMs), used with straightforward prompting, perform poorly on routine number-crunching tasks that hospital administrators depend on every day to track patients and allocate resources. The findings were published this week in the open-access journal PLOS Digital Health by Eyal Klang of the Icahn School of Medicine at Mount Sinai, New York, USA, and colleagues.

Hospitals rely on structured electronic health record (EHR) data to monitor patient counts and resources and to generate administrative reports. These tasks are currently handled by data analysts using programming languages, creating delays when staff need fast answers. AI tools known as large language models, such as GPT-4o and Llama, have been proposed to simplify that process.

In the new study, researchers evaluated nine leading LLMs on two basic administrative tasks—counting patients meeting a condition and filtering records based on multiple criteria—using data drawn from 50 000 real emergency department visits at the Mount Sinai Health System.

The researchers found that straightforward prompting—asking the model a plain question like “how many patients in this table were admitted?”—produced uniformly poor results across all models. Chain-of-thought reasoning, in which the model is prompted to show step-by-step work before giving an answer, offered only modest improvements that degraded sharply as table size increased. Even GPT-4o, the top-performing model, saw accuracy drop from roughly 95% on the smallest datasets to below 60% on larger ones under chain-of-thought conditions.

A tool-based approach—where models were asked to generate code that was then executed—substantially improved accuracy for the most capable models, with GPT-4o and Qwen-2.5-72B achieving near-perfect performance. However, distilled DeepSeek models, optimised for speed and efficiency, struggled even with this approach. One model, Llama-3.1-8B, failed to produce usable output in the majority of trials and was excluded from further analysis.

“Our findings indicate that without using a tool-based strategy, current LLMs are unsuitable for standalone use even on minimally complex administrative tasks in clinical settings,” says Benjamin Glicksberg. “Structured data tasks in clinical workflows will require agentic approaches that combine LLMs with code execution to ensure accuracy and consistency.”

Provided by PLOS

Much Medical Information Provided by Popular Chatbots is Inaccurate and Incomplete

Half of answers to evidence based questions “somewhat” or “highly” problematic

A substantial amount of medical information provided by 5 popular chatbots is inaccurate and incomplete, with half of the answers to clear evidence based questions “somewhat” or “highly” problematic, show the results of a study published in the open access journal BMJ Open.

Continued deployment of these chatbots without public education and oversight risks amplifying misinformation, warn the researchers.

Generative AI chatbots have been rapidly adopted across research, education, business, marketing and medicine, with many people using them like search engines, including for everyday health and medical queries, explain the researchers.

To gauge the level of accuracy provided in areas of health and medicine already prone to misinformation, and therefore with consequences for everyday health behaviour, the researchers probed 5 publicly available and popular generative AI chatbots in February 2025: Gemini (Google); DeepSeek (High-Flyer); Meta AI (Meta); ChatGPT (OpenAI); and Grok (xAI).

Each chatbot was prompted with 10 open ended and closed questions in each of 5 categories of cancer, vaccines, stem cells, nutrition, and athletic performance. The prompts were designed to resemble common ‘information-seeking’ health and medical queries and misinformation tropes online and in academic discourse. 

And they were developed to ‘strain’ models towards misinformation or contraindicated advice—a strategy increasingly used for stress testing AI chatbots and picking up behavioural vulnerabilities, note the researchers.

Closed prompts required chatbots to provide pre-defined responses, often with one correct answer, that aligned with the scientific consensus. Open ended prompts typically required chatbots to generate multiple responses in list form.

Responses were categorised as non-, somewhat, or highly problematic, using objective pre-defined criteria. A problematic response was defined as one that could plausibly direct lay users to potentially ineffective treatment or come to harm if followed without professional guidance.

The information was scored for accuracy and completeness, and particular attention was given to whether a chatbot presented a false balance between science and non-science based claims, regardless of the strength of the evidence.

Each response was also graded on readability, ranging from whether it was written in easy, plain English, to difficult, academic language, using the Flesch Reading Ease score.

Half (50%) the responses were problematic: 30% were somewhat, and 20% were highly problematic. 

Prompt type was influential: open-ended prompts, for example, produced 40 highly problematic responses—significantly more than expected—and 51 non-problematic responses—significantly fewer than expected. The opposite was true of closed prompts.

While the quality of responses didn’t differ significantly among the 5 chatbots, Grok
generated significantly more highly problematic responses than would be expected (29/50; 58%). Gemini generated the fewest highly problematic responses and the most non-problematic ones.

The chatbots performed best in the area of vaccines and cancer, and worst in the area of stem cells, athletic performance, and nutrition. 

Answers were consistently expressed with confidence and certainty, with few caveats or disclaimers. Out of the total 250 questions, there were only two refusals to answer, both of which came from Meta AI in response to queries about anabolic steroids and alternative cancer treatments.

Reference quality was poor, with an average completeness score of 40%. Chatbot hallucinations and fabricated citations meant that no chatbot provided a fully accurate reference list. 

All readability scores were graded as ‘difficult’, equivalent in complexity to suitability for a college graduate.

The researchers acknowledge that they only assessed 5 chatbots and that commercial AI is rapidly evolving, so their findings might not be universally applicable. And not all real-world queries are deliberately adversarial, an approach they took which may have overstated the prevalence of problematic content.

Nevertheless, “Our findings regarding scientific accuracy, reference quality, and response readability highlight important behavioural limitations and the need to re-evaluate how AI chatbots are deployed in public-facing health and medical communication,” they point out. 

“By default, chatbots do not access real-time data but instead generate outputs by inferring statistical patterns from their training data and predicting likely word sequences. They do not reason or weigh evidence, nor are they able to make ethical or value-based judgments,” they explain.

“This behavioural limitation means that chatbots can reproduce authoritative-sounding
but potentially flawed responses.” 

The data chatbots draw on also includes Q&A forums and social media, and scientific content is typically limited to open access or publicly available articles, which comprise only 30–50% of published studies. While this enhances conversational fluency, it  may come at the cost of scientific accuracy, advise the researchers.

“As the use of AI chatbots continues to expand, our data highlight a need for public education, professional training, and regulatory oversight to ensure that generative AI supports, rather than erodes, public health,” they conclude.

Source: BMJ Group

Deepfake X-Rays Fool Radiologists and AI

Findings raise concerns about cybersecurity and diagnostic trust

Anatomy-matched real and GPT-4o-generated radiographs: (A) real and (B) GPT-4o-generated posteroanterior chest radiographs, (C) real and (D) GPT-4ogenerated lateral cervical spine radiographs, (E) real and (F) GPT-4o-generated posteroanterior hand radiographs, and (G) real and (H) GPT-4o-generated lateral lumbar spine radiographs. The pairs demonstrate that GPT-4o can produce radiographically plausible images across different anatomic regions.
https://doi.org/10.1148/radiol.252094 ©RSNA 2026

Neither radiologists nor multimodal large language models (LLMs) are able to easily distinguish AI-generated “deepfake” X-ray images from authentic ones, according to a study published in Radiology. The findings highlight the potential risks associated with AI-generated X-ray images, along with the need for tools and training to protect the integrity of medical images and prepare health care professionals to detect deepfakes.

The term “deepfake” refers to a video, photo, image or audio recording that appears real but has been created or manipulated using AI.

“Our study demonstrates that these deepfake X-rays are realistic enough to deceive radiologists, the most highly trained medical image specialists, even when they were aware that AI-generated images were present,” said lead study author Mickael Tordjman, MD, post-doctoral fellow, Icahn School of Medicine at Mount Sinai, New York. “This creates a high-stakes vulnerability for fraudulent litigation if, for example, a fabricated fracture could be indistinguishable from a real one. There is also a significant cybersecurity risk if hackers were to gain access to a hospital’s network and inject synthetic images to manipulate patient diagnoses or cause widespread clinical chaos by undermining the fundamental reliability of the digital medical record.”

Seventeen radiologists from 12 different centers in six countries (United States, France, Germany, Turkey, United Kingdom and United Arab Emirates) participated in the retrospective study. Their professional experience ranged from 0 to 40 years. Half of the 264 X-ray images in the study were authentic, and the other half were generated by AI. Radiologists were evaluated on two distinct image sets, with no overlapping between the datasets. The first dataset included real and ChatGPT-generated images of multiple anatomical regions. The second dataset included chest X-ray images—half authentic and the other half created by RoentGen, an open-source generative AI diffusion model developed by Stanford Medicine researchers.

When radiologist readers were unaware of the study’s true purpose, yet asked after ranking the technical quality of each ChatGPT image if they noticed anything unusual, only 41% spontaneously identified AI-generated images. After being informed that the dataset contained synthetic images, the radiologists’ mean accuracy in differentiating the real and synthetic X-rays was 75%.

Individual radiologist performance in accurately detecting the ChatGPT-generated images ranged from 58% to 92%. Similarly, the accuracy of four multimodal LLMs—GPT-4o (OpenAI), GPT-5 (OpenAI), Gemini 2.5 Pro (Google), and Llama 4 Maverick (Meta)—ranged from 57% to 85%. Even ChatGPT-4o, the model used to create the deepfakes, was unable to accurately detect all of them, though it identified the most by a considerable margin compared to Google and Meta LLMs.

Radiologist accuracy in detecting the RoentGen synthetic chest X-Rays ranged from 62% to 78% and the LLM models’ performance ranged from 52% to 89%.

There was no correlation between a radiologist’s years of experience and their accuracy in detecting synthetic X-ray images. However, musculoskeletal radiologists demonstrated significantly higher accuracy than other radiology subspecialists.

Spotting the Risks in Synthetic Imaging

“Deepfake medical images often look too perfect,” Dr. Tordjman said. “Bones are overly smooth, spines unnaturally straight, lungs overly symmetrical, blood vessel patterns excessively uniform, and fractures appear unusually clean and consistent, often limited to one side of the bone.”

Recommended solutions to clearly distinguish real and fake images and help prevent tampering include implementing advanced digital safeguards, such as invisible watermarks that embed ownership or identity data directly into the images and automatically attaching technologist-linked cryptographic signatures when the images are captured.

“We are potentially only seeing the tip of the iceberg,” Dr. Tordjman said. “The logical next step in this evolution is AI-generation of synthetic 3D images, such as CT and MRI. Establishing educational datasets and detection tools now is critical.”

The study’s authors have published a curated deepfake dataset with interactive quizzes for educational purposes.

For More Information

Access the Radiology study, “The Rise of Deepfake Medical Imaging: Radiologists’ Diagnostic Accuracy in Detecting ChatGPT-generated Radiographs,” and the related editorial, “The Democratization of Deceit: Seeing Is No Longer Believing.”

Source: Radiological Society of North America

AI Tools for Cancer Rely on Shaky Shortcuts

Small cell lung cancer cells (green and blue) that metastasised to the brain in a laboratory mouse recruit brain cells called astrocytes (red) for their protection. Credit: Fangfei Qu

Artificial intelligence tools are increasingly being developed to predict cancer biology directly from microscope images, promising faster diagnoses and cheaper testing. But new research from the University of Warwick, published in Nature Biomedical Engineering, suggests that many of these systems may be using visual shortcuts rather than true biology – raising concerns that some AI pathology tools are currently too unreliable for real-world patient care.

“It’s a bit like judging a restaurant’s quality by the queue of people waiting to get in: it’s a useful shortcut, but it’s not a direct measure of what’s happening in the kitchen,” says Dr Fayyaz Minhas, Associate Professor and principal investigator of the Predictive Systems in Biomedicine (PRISM) Lab in the Department of Computer Science, University of Warwick, and lead author of the study.

“Many AI pathology models are doing the same thing, relying on correlations between biomarkers or on obvious tissue features, rather than isolating biomarker-specific signals. And when conditions change, these shortcuts often fall apart.”

To reach this conclusion, the researchers analysed more than 8000 patient samples across four major cancer types – breast, colorectal, lung and endometrial – and compared the performance of leading machine learning approaches. While the models often achieved high headline accuracy, the team found this frequently came from statistical “shortcuts.”

For example, instead of detecting mutations in the cancer-associated BRAF gene, a model might learn that BRAF mutations often occur alongside another clinical feature such as microsatellite instability (MSI). The system then learns to use this combination of cues to predict BRAF status rather than learning the causal BRAF signal itself – meaning accurate cancer predictions work only when these biomarkers co-occur and become unreliable when they do not.

Kim Branson, SVP Global Head of Artificial Intelligence and Machine Learning, GSK and co-author says, “We’ve found that predicting a BRAF mutation by looking at correlated features like MSI is often like predicting rain by looking at umbrellas – it works, but it doesn’t mean you understand meteorology.

“Crucially, if a model cannot demonstrate information gain above a simple pathologist-assigned grade, we haven’t advanced the field; we’ve just automated a shortcut. The roadmap for the next generation of pathology AI isn’t necessarily bigger models; it’s stricter evaluation protocols that force algorithms to stop cheating and learn the hard biology.”

When performance of AI models was assessed within stratified patient subgroups, such as only high-grade breast cancers or only MSI-positive tumours, accuracy fell substantially, revealing that the models were dependent on shortcut signals that disappear once confounding factors are controlled.

For certain prediction tasks, the performance advantage of deep learning over human-derived clinical information was modest. AI systems achieved accuracy scores of just over 80% when predicting biomarkers, compared with around 75% using tumour grade alone – a measure already assessed by pathologists.

Machine learning methods can still prove valuable for research, drug development candidate screening and for clinical triaging, screening, or supplementary decision support. However, the researchers argue that future AI tools must move beyond correlation-based learning and adopt approaches that explicitly model biological relationships and causal structure.

They also call for stronger evaluation standards, including subgroup testing and comparison against simple clinical baselines, before looking at deployment in routine care.

Dr Minhas concludes, “This research is not a condemnation of AI in pathology. It is a wake-up call. Current models may perform well in controlled settings but rely on statistical shortcuts rather than genuine biological understanding. Until more robust evaluation standards are in place, these tools should not be seen as replacements for molecular testing, and it is essential that clinicians and researchers understand their limitations and use them with appropriate caution.”

Source: University of Warwick

Half of All Men Over 60 Have Prostate Cancer – an AI Tool Could Speed Diagnosis

Photo by National Cancer Institute on Unsplash

Increasing use of blood tests to detect prostate cancer is leading to overworked doctors. NTNU has now created an AI diagnostic tool that can help lighten the burden.

Diagnostic tools based on artificial intelligence are now making their way into Norwegian hospitals. AI can independently read X-ray images and detect bone fractures, or assess cancer tumours in both the breast and prostate.

“AI tools can take over the detection of simple and clear-cut cases, allowing doctors to spend their time on more complex ones,” said Tone Frost Bathen. She is a professor at NTNU and the project manager of an AI-powered analysis tool for prostate cancer called PROVIZ.

Tests on patients at St Olavs Hospital indicate that the tool is very promising.

“AI can enable radiologists to determine more quickly and more accurately whether a patient needs a biopsy, and where in the prostate it should be taken from,” explained Bathen.

“The PROVIZ project started as early as 2018. It takes a long time to develop diagnostic tools in medicine because safety standards must be high. The application alone to be allowed to test the tool on patients was 500 pages. It is important to create a tool that clearly shows how the result was reached, and that fits into a busy hospital workday,” says Tone Frost Bathen, Professor at NTNU. Photo: Anne Sliper Midling / NTNU

A recent study shows that patients trust medical test results only if an experienced doctor confirms what has been detected.

“Trust in doctors and health professionals is key for artificial intelligence to gain a place in the diagnosis of prostate cancer. Technology alone is not enough. Human contact and professional assessment remain indispensable,” said Simon A. Berger, a PhD research fellow at NTNU.

Prostate cancer is a natural part of getting older

Prostate cancer is the most common form of cancer among men in Western countries.

Examinations have detected prostate cancer in 10% of 50-year-olds, 50% of 60-year-olds and approximately 70% of men over the age of 80.

This shows that the disease is naturally linked to ageing.

“Prostate cancer is something most men die with, not from,” added Berger.

A blood test called PSA can help detect prostate cancer. Since it has become more common for men to take this blood test, the number of new prostate cancer cases has risen sharply. There are now approximately 5000 new cases each year.

When more people are tested for something that many individuals naturally have as part of the ageing process, the next medical step after the blood test must also be carried out more often, so that doctors can obtain a broader clinical picture of its severity.

Most trust in doctors

Currently, this next step involves taking an MRI scan, which provides a detailed image of the prostate gland and the surrounding tissue. These images need to be interpreted manually by an experienced radiologist. As the number of images taken has increased sharply, this has created a need for new and more efficient ways of making diagnoses.

Through the PROVIZ project, NTNU researchers have developed an AI-powered tool that can help doctors interpret MRI images of the prostate. PROVIZ is currently available only for use as part of the ongoing research project, but efforts are underway to apply for a patent and make the tool commercially available.

High international competition for commercial AI tools

Several research groups around the world are now working on developing AI-based diagnostic tools for prostate cancer.

PROVIZ has completed its first clinical testing in collaboration with St. Olavs Hospital, and the results were good. The next step is a much larger clinical trial, as well as a regulatory approval process.

“Right now, we are seeking approximately 20 million NOK to finance this phase. Once funding is in place, the tool could be on the market in the US within a year, and in Europe in just over a year,” says Gabriel Addio Nketiah, a researcher at NTNU and responsible for the commercialisation of PROVIZ.

For a tool like this to be efficiency-enhancing in routine hospital practice, patients must also trust the findings detected through the use of AI.

“Patients have high expectations that AI can be used for faster diagnostics and to reduce healthcare waiting lists. Many see AI as a kind of safety valve – an additional resource that doctors can use alongside their professional judgment,” says Simon A. Berger, a PhD research fellow at NTNU.

Berger interviewed 18 men who had been diagnosed with prostate cancer through the use of PROVIZ. The study shows that trust in doctors and health professionals plays a decisive role in whether patients accept AI in the health services.

“Patients trust AI in lower-risk cases such as bone fractures, but not in cases where the perceived risk is higher, such as cancer. When the perceived risk is high, we place the greatest trust in specialized doctors who can confirm what AI has found,” explained Berger.

Doctors as guarantors

In his interviews, Berger identified three different dimensions of trust.

  1. Foundational trust in the healthcare system: many patients had positive experiences from previous encounters with the healthcare system. This laid a positive foundation.
  2. Inter-personal trust in health professionals: patients trusted the doctors and their assessments. This trust was crucial for accepting AI because the doctors explained and vouched for the technology.
  3. Possible trust in AI: even though patients recognized the potential of AI, they always wanted a human assessment as well in prostate cancer diagnostics. They were concerned about accountability, professional judgement and AI’s (in)ability to see the whole clinical picture.

“The relationship between patient and doctor is still key. For AI to be accepted in clinical practice, health professionals must be active communicators and guarantors of safety. In order for doctors to serve as guarantors, they must first understand how AI arrived at its conclusions so they can verify that it has made the correct assessment. Patients accept the use of AI within a framework they already trust,” concluded Berger.

NTNU owns an MRI scanner at St. Olavs Hospital that is currently undergoing a major upgrade. It helps researchers obtain the best possible images to be used in, among other things, PROVIZ. “Unfortunately, there are few investors in medical technology right now, but we hope that someone sees the societal value of our project,” says Professor Tone Frost Bathen at NTNU. Photo: Anne Sliper Midling / NTNU

By Anne Sliper Midling

Source:

Berger SA, Håland E, Solbjør M. Patient Perspectives on Trust in Artificial Intelligence-Powered Tools in Prostate Cancer Diagnostics. Qualitative Health Research. 2025;0(0). doi:10.1177/10497323251387545

Source: Norwegian Tech News

Can Medical AI Lie? How LLMs Handle Health Misinformation

Photo by Sanket Mishra

Medical artificial intelligence (AI) is often described as a way to make patient care safer by helping clinicians manage information. A new study by the Icahn School of Medicine at Mount Sinai and collaborators confronts a critical vulnerability: when a medical lie enters the system, can AI pass it on as if it were true?  

Analysing more than a million prompts across nine leading language models, the researchers found that these systems can repeat false medical claims when they appear in realistic hospital notes or social-media health discussions. 

The findings, published in the February 9 online issue of The Lancet Digital Health], suggest that current safeguards do not reliably distinguish fact from fabrication once a claim is wrapped in familiar clinical or social-media language. 

To test this systematically, the team exposed the models to three types of content: real hospital discharge summaries from the Medical Information Mart for Intensive Care (MIMIC) database with a single fabricated recommendation added; common health myths collected from Reddit; and 300 short clinical scenarios written and validated by physicians. Each case was presented in multiple versions, from neutral wording to emotionally charged or leading phrasing similar to what circulates on social platforms. 

In one example, a discharge note falsely advised patients with oesophagitis-related bleeding to “drink cold milk to soothe the symptoms.” Several models accepted the statement rather than flagging it as unsafe. They treated it like ordinary medical guidance. 

“Our findings show that current AI systems can treat confident medical language as true by default, even when it’s clearly wrong,” says co-senior and co-corresponding author Eyal Klang, MD, Chief of Generative AI in the Windreich Department of Artificial Intelligence and Human Health at the Icahn School of Medicine at Mount Sinai. “A fabricated recommendation in a discharge note can slip through. It can be repeated as if it were standard care. For these models, what matters is less whether a claim is correct than how it is written.”  

The authors say the next step is to treat “can this system pass on a lie?” as a measurable property, using large-scale stress tests and external evidence checks before AI is built into clinical tools. 

“Hospitals and developers can use our dataset as a stress test for medical AI,” says physician-scientist and first author Mahmud Omar, MD, who consults with the research team. “Instead of assuming a model is safe, you can measure how often it passes on a lie, and whether that number falls in the next generation.”  

“AI has the potential to be a real help for clinicians and patients, offering faster insights and support,” says co-senior and co-corresponding author Girish N. Nadkarni, MD, MPH, Chair of the Windreich Department of Artificial Intelligence and Human Health, Director of the Hasso Plattner Institute for Digital Health, Irene and Dr. Arthur M. Fishberg Professor of Medicine at the Icahn School of Medicine at Mount Sinai, and Chief AI Officer of the Mount Sinai Health System. “But it needs built-in safeguards that check medical claims before they are presented as fact. Our study shows where these systems can still pass on false information, and points to ways we can strengthen them before they are embedded in care.” 

The paper is titled “Mapping LLM Susceptibility to Medical Misinformation Across Clinical Notes and Social Media.”  

Source: Mount Sinai