With all the talk of AI outperforming and replacing doctors, a new JACC study reminds us – AI diagnostic algorithms don’t perform great when applied to populations different from the ones they’re trained on.
- The algorithm in question is Pathway Labs’ EchoNext, an AI-ECG model trained on multicenter hospital data for detecting structural heart disease (SHD).
- EchoNext launched in June to mixed reception. The New York Times wrote a positive piece about the problem it solves, while some cardiologists called it unnecessary.
What better way to address the critics than with real world data? EchoNext was applied to patients from PREVUE-VALVE, a community-based study of adults aged 65-85 who underwent in-home ECG and echo.
But the details of those participants’ disease matters and they looked fundamentally different from EchoNext’s training data.
- SHD prevalence in the real-world community participants was just 8% versus 43% in the hospital training data.
- The community-based patients’ SHD also presented as less severe.
- Moderate tricuspid regurgitation was more common in these participants than the usual systolic heart failure from the hospital patient group.
With those differences in mind, it’s no wonder EchoNext showed a drop in diagnostic performance among community-based patients.
- Model discrimination fell to an AUC of 71% compared to 83% in the hospital.
- Adjusting for differences in prevalence and case mix narrowed but didn’t close that gap.
- But, AUC did rise to 79% in participants with an abnormal ECG.
- It also reached 76% among patients with worse health status.
In the context of EchoNext’s potential applications, these results are still decent (though an AUC > 80% is preferred for clinical reliability).
There’s a lesson to be learned here, and it goes beyond this one model.
- AI algorithms can only be as good as the data they’re trained on, especially since they lack the intuition of human doctors.
The Takeaway
Diagnostic AI is an eventuality in medicine, but that doesn’t mean that it’ll be perfectly reliable for every patient or population. Rather, physicians encountering and eventually using tools like EchoNext will still need to rely on their gut and experience, no matter AI’s utility.

