CHARGE
/
Insights
/
Speed Kills: UPMC & ECRI Report That AI Adoption, Drift Outpacing Governance Structures
Governance
September 3, 2026

Speed Kills: UPMC & ECRI Report That AI Adoption, Drift Outpacing Governance Structures

31% of healthcare organizations confirmed a faulty AI output over the past year.

By the onset of the SARS-COV-2 pandemic in March 2020, the University of Pennsylvania Healthcare System had employed an ML algorithm designed to predict 180-day mortality rates among outpatients with cancer for over a year without any major hiccups. Since January 2019, UPHS clinicians had initiated serious illness conversations with outpatients assessed a ≥10% risk score by ML. The model, however, couldn't withstand the pandemic: it was built on top of pre-pandemic EHR records. When COVID disrupted care patterns, the model’s true-positive rate for identifying high-risk patients dropped seven percentage points from 80.9% to 73.9%. 

The UPHS ML is a clear example of model drift, where, due to changes in input distribution, AI model performance degrades over time. COVID was, of course, anomalous – models built before COVID couldn’t reasonably anticipate the radical shift in care patterns initiated by the pandemic. Drift, however, characterizes any dissonance between training and real world-data, and the risk isn’t only extraordinary events like pandemics: today, health leaders fear a crisis of algorithmic biases, diagnostic inaccuracy, and model hallucination in their deployed AI models.

Hospitals are unprepared for drift

It’s also the AI risk surface for which hospitals are least prepared, especially as deployment accelerates. According to a recent UPMC-Klaws report, more than 90% of health systems have deployed third-party AI solutions – of those health systems, 92% report that they assess third-party AI tools for diagnostic accuracy before deployment. Many models can perform better in limited sandboxes and controlled settings. Far fewer hospitals, however, possess the capability to test their AI systems numerically after deployment: less than half (44%) of UPMC respondents reported a dedicated data environment for testing AI solutions, and respondents indicated no strong consensus on how to quantify diagnostic AI success. 

Rather, as Peter Pronovost, Chief Quality and Clinical Transformation Officer at University Hospitals puts it: 

Most health systems are monitoring the safety and performance of these tools the same way they governed a new MRI scanner in 2010: a subcommittee, a checklist, a quarterly meeting, an approval or a rejection. The process could take six months or longer. -- Peter Pronovost, MD, PhD
The risk impacts patients

That administrative lethargy is increasingly archaic and increasingly dangerous. ECRI, which helps health systems improve patient safety, reports that over 31% of hospitals encountered a faulty AI output over the past year. Of those health systems, over 9% had confirmed AI errors affected patients directly by impacting care decisions. Most worryingly, 35% of health systems were unsure whether errors had even occurred. Where error modes are visible, they are at least identifiable; where they are hidden, patient risk is opaque and impossible to mitigate. That’s a disastrous pairing for healthcare. 

AI drift, too, occurs most often in stealth: healthcare organizations can rarely detect decay in AI solutions, especially when they sit at the edge of healthcare organizations or as error modes are obscured. Detection oversight, however, is not an inevitable, infrastructural feature of AI; rather, hospitals must tackle a new AI governance challenge. 

The challenge, by the numbers

That challenge is post-deployment monitoring. CHAIRS (the Council for Healthcare AI Responsibility and Safety), was convened precisely to evaluate the depth of the problem. The results of its first study are stark: 

  • no organization conducts manual review of more than about 5% of interactions;
  • roughly 70% of organizations audit patient visits monthly or less; 
  • 96% of organizations rely on clinical notes for assessing patient interactions, creating blind spots around AI omissions

The gap between the 92% of organizations which rigorously test their AI models, and the 5% of interaction makes detecting and taming AI drift the foremost challenge for health leaders. To tackle the scope of the crisis, hospital leaders must turn towards the new governance frontier: post-deployment.

References

[1] UPHS: https://academic.oup.com/jamia/article/30/2/348/6835770

[2] AI Drift: https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2834882

[3] UPMC Report: https://www.upmc.com/media/news/080626-ai-use-in-healthcare-systems-upmc-research

[4] ECRI: https://hitconsultant.net/2026/08/26/ecri-expands-problem-reporting-network-track-healthcare-ai-errors-patient-safety/

THE AUTHOR
CHARGE
Schedule a meeting with a co-founder.
SOURCE

From the CHARGE Newsletter.