AI promises to transform clinical research and practice, but obtuse performance metrics limits implementation. To explore the challenges of AI adoption, CHARGE spoke with Joshua Lampert, MD, Director of Cardiovascular AI at Mount Sinai’s Windreich Department of AI and Human Health, in a wide-ranging interview which touched on AI sycophancy, autonomous implementation, and human-AI partnership.
Dr. Lampert's groundbreaking published work through Windreich forges a path for implementation of independent AI models in research settings, and his robust LinkedIn sketches a forward-looking vision for AI in healthcare which preserves what “we do best: being human.” This interview picks up on his vision – CHARGE thanks Dr. Lampert for his participation.
Q: What first sparked your interest in artificial intelligence, and how did you come to work at the intersection of AI and medicine?
My interest began about seven years ago when I started looking at subclinical biomarkers hidden within ECG waveform data. Traditional ECG analysis relies heavily on manual heuristics, which often forces clinicians to use surrogate markers instead of directly predicting the outcomes that matter most. That gap sparked my curiosity: I wanted to see if deep learning could decode those complex visual patterns and translate raw physiological data into actionable patient care,
Q: How do you think AI will change cardiology practice over the next decade, and what changes will be most noticeable for clinicians and patients?
AI is already helping us better detect disease, risk stratify patients, and has even demonstrated a mortality benefit by prompting action on certain alerts. Change will be informed by patient and physician trust, flexibility across a variety of clinical workflows and practice settings, and mechanisms that support the financial viability of solutions that improve patient care.
The most noticeable changes will be in how clinicians and patients enter and interact with clinical workflows because AI tools streamline care in ways previously impossible.
For example, it is common to arrive at a visit with a primary care physician or specialist only to establish a diagnostic plan. You don’t receive diagnoses or treatment plans, you embark on a diagnostic pathway. With AI tools, it’s possible to coordinate appropriate care so that patients and physicians can discuss actionable next steps – that enhances effective clinical decision making.
Q: In medical research, where do you see AI having the biggest impact: speeding up current methods, or changing how studies are designed and evaluated altogether?
There is certainly an opportunity to achieve all of those goals. The issue really comes down to using the right tools under the often misunderstood umbrella term of "AI" for the specific problem being solved.
Different models have distinct mathematical underpinnings that make them advantageous for certain tasks but not for others. For example, LLMs alone do not compute reliable underlying probabilities. In patient care, they may struggle with edge cases and clinical uncertainty. I’m waiting to see how this may change with the use of dedicated tools or other mechanisms designed to achieve the desired outcome.
Q: In your recent research on AI agents for clinical trials, you describe systems that can perform parts of the research process independently. How should we balance automation with human oversight to ensure reliability and safety?
Any system needs to include guardrails.
The study you mentioned highlights some of the benefits of employing AI agents in research: we leveraged an autonomous agent to calibrate EHR estimates to randomized control trials results using a Bayesian hierarchical model. RCT results do not necessarily translate to local real-world observations since enrollment in an RCT is highly specific, follow-up is augmented and prespecified, and patients receive structured monitoring. Utilizing an autonomous agent enabled repeated, standardized trial replication within a local context, converting EHR-RCT discrepancies into data from which we can learn institution-level transport properties.
As with any balance involving oversight, however, the key is how we quantify and report uncertainty. This is what provides safety by helping determine when a tool should not be applied or whether the scenario it is being tasked with addressing was not represented in its training data or exists at an extreme. Currently, our ability to adequately express and quantify uncertainty remains insufficient for reliable safety and limits autonomous deployments, though developments in this area are emerging.
Q: Your work highlights model sycophancy, where AI systems tend to agree with or reinforce user input rather than provide independent judgment, as a critical limitation in clinical contexts. Why is this problematic in medicine, and how can systems be designed to maintain both safety and clinical utility?
Any scenario in which a tool may feed into, reinforce, or amplify a user’s bias without a quantifiable measure of uncertainty can be dangerous. Part of the solution is to acquire objective data that can be processed and coordinated by deep learning systems to help inform risk assessment, with an escalation pathway to human experts. In any safe system, safety comes from redundancy and escalation, as we have learned from the airline industry.
The challenge in medicine is determining how much additional uncertainty is introduced and how humans interpret and act on the resulting data. Systems must be designed with synergy in mind rather than focusing on a single use case with associated general performance metrics. A good system is not simply a tool with a good-appearing AUROC (which itself can be a misleading performance metric), but rather the composite of all aspects of the system, down to how outputs are represented to end users, which is model agnostic.
Q: Following your work on AI agents in clinical research and decision support systems such as DecidEHR, which provide individualized recommendations that may differ from standard guideline-based approaches. How should responsibility be shared between clinicians and AI systems when such disagreements arise?
It depends on what is meant by responsibility.
There are both ethical and medico-legal implications. Ethically, we as physicians take an oath: first, do no harm. I do not think that will ever change. Our primary responsibility is to ensure that we keep our patients safe. As we become more comfortable with sophisticated tools deployed at scale, this role will become even more important. Likewise, there are many potential sources of error in how models generate or display incorrect outputs, and clinicians should not be solely responsible for that process.
To that end, pragmatic systems need to function like a refrigerator. You do not necessarily need to understand all of the technical details of how the system works (which will never be achievable at scale for all end users), but if you open the door and it is not cold inside, you know there is something wrong. This is where representing uncertainty and providing accessible sanity checks become important for all such systems, particularly those in which errors can scale.
Q: Looking ahead, what will the role of AI in medicine look like in the next 5 to 10 years, and what developments should we be preparing for today?
Over the next 5–10 years, I hope that these novel tools will allow us to reimagine the nature of medical practice. We are in a transitional phase, merely getting our feet wet. Currently, we leverage novel computational approaches that still recapitulate historical assumptions about how medical care should be delivered.
Over the next decade, the technology will mature, as will (hopefully) our perspective on how to best reinvent the delivery of excellent healthcare. We should all remain humble, keep an open mind, and work through the mathematical and operational challenges that have persisted in medicine for centuries. We should prepare ourselves to interact with healthcare systems differently. In some instances, this may involve not interacting with a human in the traditional sense. We should be prepared to question the nature of good judgment. At the health system level, we should recognize that, in the short term, there will be increased costs associated with encouraging the development and adoption of new technologies. We need to make these investments now to build the infrastructure and enable the adoption needed to leverage more synergistic systems in the coming years, allowing us to deliver higher-quality, more efficient, and cost-effective care.
We will undoubtedly be able to detect and treat disease more effectively and efficiently. As we offload the cognitive and operational burdens of tasks that humans do not perform optimally to integrated and streamlined systems, we will be able to focus our time on what we do best: being human.



