Abstract
Heart failure (HF) is a complex, multifactorial, and difficult-to-treat syndrome. Over the past years, a concerning increase in its global prevalence, mortality, costs, and burden on the healthcare system has been observed. The recent development of machine learning (ML), especially unsupervised and deep learning (DL) algorithms, offers a potential way to facilitate diagnosis, enable more precise treatment, and reduce both mortality and costs of HF patients. Especially unsupervised ML and DL present new opportunities for increased efficiency in clinical practice in cardiology. Unsupervised ML, e.g., enables novel phenogrouping of HF patients into high-risk groups and disease outcome prediction. Deep learning algorithms can enhance echocardiographic analysis by improving image quality and ECG interpretation, and by providing assistance and guidance to inexperienced cardiologists. However, substantial challenges related to generalizability, external validation, overfitting, model explainability, and ethical considerations currently severely limit the implementation of AI-based tools in real-world clinical practice. This review critically evaluates current AI models in HF, focusing on their roles in diagnosis, risk stratification, and treatment personalization, as well as the major challenges that restrict their application in clinical practice.
Key words: heart failure, machine learning, deep learning, artificial intelligence, precision medicine
Introduction
Heart failure (HF) is a complex, multifactorial syndrome characterized by impaired cardiac pump function and/or structural or functional abnormalities of the heart, resulting in inadequate metabolic supply to the body.1
In 2021, the European Society of Cardiology (ESC) defined HF as “a clinical syndrome consisting of cardinal symptoms (e.g., breathlessness, ankle swelling, and fatigue) that may be accompanied by signs (e.g., elevated jugular venous pressure, pulmonary crackles, and peripheral edema). It is due to a structural and/or functional abnormality of the heart that results in elevated intracardiac pressures and/or inadequate cardiac output at rest and/or during exercise”.2 Traditionally, HF is classified into 3 phenotypes: HF with preserved ejection fraction (HFpEF), defined as left ventricular ejection fraction (LVEF) >50%; HF with mildly reduced ejection fraction (HFmrEF), defined as LVEF 41–49%; and HF with reduced ejection fraction (HFrEF), defined as ejection fraction (LVEF) <40%.2
The prevalence of HF has been steadily increasing over recent years, reaching more than 64 million cases worldwide.1 Together with an estimated annual cost of up to €25,000 per patient and a 1-year mortality rate of up to 23.1% among patients with acute HF, this represents a substantial burden on the global healthcare system.1
In Poland, the number of patients with HF increased byapprox. 34% over the past 10 years, accompanied by a concerning decline in 1-year survival from 86% in 2014 to 76% in 2021.3 A recent study conducted in Poland highlighted the importance of HF nurses, particularly the impact of their education on the diagnostic and therapeutic management of patients with HF. In 2021, a Polish online educational platform for nurses was introduced, including a certification course for HF nurses. According to the survey findings, this educational initiative led to a significant improvement in the quality of care and management of patients with HF.4 Major challenges in HF include its complexity, multifactorial etiology, and substantial variability in pathomechanisms (Figure 1), which not only make the diagnostic process highly time- and cost-intensive but may also result in suboptimal or ineffective treatment.1
Artificial intelligence (AI) is a rapidly evolving technology with considerable potential in medicine. It is expected to facilitate and accelerate diagnostic processes, individualize treatment, reduce clinicians’ workload, and minimize healthcare expenditure, particularly in complex diseases such as HF.5
Objectives
This narrative review provides an overview of the most important AI models in the field of HF, their role in precision medicine, and the potential limitations of their clinical implementation.
Materials and methods
An electronic database search was conducted using PubMed. The database was searched for AI models developed and tested to support the diagnosis of HF, improve treatment precision in patients with HF, and assist clinical decision-making in HF management. Particular emphasis was placed on randomized controlled trials (RCTs) that included internal and/or external validation of the evaluated AI algorithms, as this is an important factor in assessing their clinical utility and potential implementation.
The search strategy included all possible combinations of the following terms: heart failure AND (artificial intelligence OR machine learning OR supervised learning OR unsupervised learning OR deep learning). As the field of AI extends beyond medicine, the following exclusion criteria were applied: 1) AI models not directly related to HF; 2) AI models without internal and/or external validation; 3) substudies describing the same algorithm; and 4) lack of freely available full text. The inclusion criteria were as follows: 1) AI models directly related to HF; 2) AI models with internal and/or external validation; 3) RCTs published within the last 5 years; and 4) freely available full text. One author (J.R.) screened the titles and abstracts of the identified articles according to the inclusion and exclusion criteria (Figure 2). The selected articles were subsequently assessed for potential selection bias and grouped into 2 categories: unsupervised learning AI algorithms and supervised learning AI algorithms. The most important studies in each category are presented in Table 1.6, 7, 8, 9 and Table 2.10, 11, 12
Techniques and methods of AI
AI is a branch of computer science focused on developing algorithms capable of performing tasks at or above human level.5 In general, AI operates by mathematically transforming input data, so-called “features,” into formats that can be processed by machine learning (ML) algorithms, typically using vector representations. Features are quantifiable data that may, in their simplest form, consist of clinical variables stored in a patient’s health record (Table 3).13
Machine learning has been used in clinical applications since the 1970s and is based on algorithms that identify linear and nonlinear relationships between input data and labeled output data.13 A further distinction can be made between supervised ML, unsupervised ML, and deep learning (DL), as shown in Figure 3.
The most common type of machine learning is supervised ML, in which models are trained on labeled data. Input data are paired with corresponding output labels, enabling the model to learn the relationship between them and predict the output label for previously unseen data.14 A common clinical application of supervised ML is the detection of electrocardiographic (ECG) abnormalities. A model is trained using a large dataset of ECG recordings labeled with diagnostic annotations corresponding to specific pathologies, such as atrial fibrillation or tachycardia. At the end of the training process, the model can predict the presence of previously learned pathological patterns in new ECGs. It is important to note that the model can only recognize patterns on which it has been trained; clinical validation by human experts remains essential.15
In contrast, unsupervised ML is primarily used for large and heterogeneous datasets. It analyzes large volumes of unlabeled data to identify recurring patterns within the input dataset.16 In cardiology, unsupervised ML is used, e.g., for phenotyping patients with HF based on clinical variables, helping to identify high-risk individuals and support treatment individualization.17
Deep learning is a class of machine learning models based on artificial neural networks with multiple hidden layers, designed to enable automated learning of hierarchical data representations. It is a subset of ML and can be implemented using both supervised and unsupervised learning approaches.14 Data flow through multiple neural network layers, where each neuron performs non-linear transformations, enabling learning with minimal manual feature engineering. Compared with traditional ML, DL is better suited for processing large and complex datasets, such as images and videos, and for performing highly complex predictive tasks with high accuracy. In the context of HF, DL algorithms are used, e.g., to automate echocardiographic analysis, identify predictive patterns in imaging data associated with future cardiovascular events, and support diagnosis and outcome prediction in patients with HF.18
AI and heart failure
The recent emergence of unsupervised ML and DL has opened new avenues for the diagnosis and management of HF.19 AI-driven methodologies offer the potential to overcome current limitations, such as multifactorial pathophysiology, imprecise classification, and suboptimal therapeutic approaches. Table 1 and Table 26, 7, 8, 9, 10, 11, 12 provide an overview of the most recent advances in ML-based algorithms for HF.
Phenogrouping of patients with heart failure
To date, patients with HF have been classified primarily according to LVEF. However, this classification does not adequately reflect the complexity, multifactorial etiology, and burden of comorbidities in individual patients with HF. As a result, current management often focuses predominantly on symptom control rather than addressing the underlying causes of HF.
Several studies have proposed unsupervised ML models to identify novel phenogroups among patients with HF. These models incorporate clinical characteristics, comorbidities, and underlying pathophysiological mechanisms rather than relying solely on LVEF, offering the potential for more individualized outcome prediction and treatment strategies.
Gaevert et al., e.g., used AI-based cluster analysis to identify 6 phenogroups of patients with HF and stratify outcomes into risk groups. Each phenogroup was characterized by a predominant comorbidity: coronary heart disease (CAD), valvular heart disease, atrial fibrillation (AF), sleep apnea, chronic obstructive pulmonary disease (COPD), or few comorbidities. All groups were independent of LVEF. A 12-month follow-up analysis of the primary endpoints suggested that this comorbidity-based classification could predict HF outcomes more accurately than the traditional LVEF-based classification.6
A similar study was performed by Urban et al., who developed an unsupervised ML-based model identifying 6 phenogroups of patients with acute HF based on 63 different parameters. The mortality rate of each phenogroup was evaluated over the following year, and the researchers observed statistically significant differences in 1-year mortality (p = 0.002), suggesting potential value for future risk stratification in patients with HF.17
In addition to phenotyping patients who have already developed HF, researchers have sought to identify individuals at increased risk of developing HF, thereby enabling more effective preventive strategies in high-risk populations. For example, a study from the USA analyzed prediabetic and diabetic patients using an AI-based random forest model to predict individual risk of developing HF (area under the curve (AUC) = 0.978 in the training set and AUC = 0.865 in the test set). The study identified age, poverty-to-income ratio, prior myocardial infarction, CAD, chest pain, and glucose-lowering medication use as independent predictors of HF (p < 0.05).11
Although these models show promising potential for improving outcome prediction in patients with HF and supporting treatment individualization, most phenogrouping models remain hypothesis-generating and have not been implemented in clinical practice. This is due to several critical limitations, including limited generalizability, overfitting, algorithmic bias, and a lack of prospective validation studies. For example, hierarchical clustering algorithms, which are commonly used in phenogrouping models, may prioritize categorical variables over continuous ones, leading to inappropriate weighting of certain variables.
AI clinical support system
Choi et al. developed an AI-based clinical decision support system (AI-CDSS) for HF diagnosis and evaluated its diagnostic accuracy. The concordance rate between the AI-CDSS and HF specialists was 98%, whereas the concordance rate between non-HF specialists and HF specialists was 76%. These findings suggest that AI-based models may be useful in identifying HF, particularly in settings with limited access to specialists and constrained healthcare resources. However, it is important to note that the AI-CDSS is a supportive tool and has only seen regional use due to limited generalizability.20
Transthoracic echocardiography
Echocardiography is a time-consuming and costly tool for the diagnosis of HF, with a substantial risk of variability due to differences in operator expertise and subjective interpretation of echocardiographic findings. AI-powered models, particularly ML-based approaches, may reduce inter-observer variability and provide a cost-effective, automated screening tool.16 In 2021, a deep convolutional neural network (DCNN) model was developed in China to generate more stable, noise-reduced real-time echocardiographic images during conventional echocardiographic examination. Transthoracic echocardiography (TTE) supported by an ML-based program demonstrated higher diagnostic accuracy than conventional echocardiography alone, potentially reducing the risk of cardiovascular events in patients with HF while lowering mortality, diagnostic costs, and treatment-related expenditures.18
AGILE-Echo has developed an innovative program designed to assist less experienced physicians in performing TTE assessments by integrating AI-based guidance during the examination. AI-guided TTE demonstrated superior accuracy compared with non-AI-assisted examinations, highlighting its potential to enhance the quality and consistency of cardiac imaging. This advancement holds significant promise for improving the timely and accurate diagnosis of HF and valvular heart disease, particularly in rural, remote, or resource-limited settings.
Despite the positive results of these studies, neither program has yet been implemented in routine clinical practice due to limited generalizability, logistical constraints, lack of long-term follow-up data, and insufficient time-to-event data. Moreover, neither model has undergone external validation while all carry a risk of overfitting, and therefore cannot currently be considered generalizable.21
ECG and heart tones
Kagiyama et al. recently published a noteworthy study describing the development of an AI-assisted program that analyzes ECG data to identify early left ventricular diastolic dysfunction (LVDD) by quantifying myocardial relaxation. The model accurately predicted LVEF across HF categories, with AUCs of 0.84, 0.80, and 0.81. This approach has the potential to serve as an ECG-free, cost-effective screening tool for the early detection of LVDD. However, as the model was based on a limited set of clinical parameters and lacked follow-up data analysis, further refinement is required before clinical implementation.22
Another model, based on a convolutional neural network (CNN), was developed to detect early LVDD by analyzing recorded heart sounds. The model achieved high performance metrics (accuracy: 0.987, sensitivity: 0.986, specificity: 0.988). Nevertheless, several limitations – including single-center data collection, sensitivity to recording conditions, lack of external validation, its adjunctive nature, and the limited interpretability of DL models – currently hinder its clinical implementation.23
Invasive treatment methods
In 2018, an unsupervised ML-based model was developed to support patient selection for cardiac resynchronization therapy (CRT). The model demonstrated superior predictive performance compared with traditional selection criteria when evaluated against clinical outcomes.8 Another model using unsupervised ML was developed to identify positive and negative risk factors for HF development in elderly patients undergoing coronary rotational atherectomy (CRA).24 Although both models demonstrated promising predictive performance, they were developed retrospectively using small datasets and have not undergone external validation, limiting their potential clinical implementation.
Pharmacological treatment
Machine learning-based algorithms offer considerable potential for identifying high-risk patients and supporting treatment individualization. A recent study by Bayes-Genis et al. used artificial neural network (ANN) algorithms to identify the most probable mechanism of action of empagliflozin, a sodium-glucose co-transporter 2 inhibitor (SGLT2i), and subsequently linked these findings to the gene expression profiles of patients with HFpEF. The identified mechanism of action of empagliflozin may contribute to a better understanding of the clinical benefits of SGLT2i in HFpEF and help guide more individualized treatment strategies. However, a major limitation of this study is its small sample size, which may result in an inaccurate representation of the relative contribution of different mechanisms of action of SGLT2i. Further research with larger sample sizes is needed.25
The HOMAGE trial developed a model to identify high-risk patients with HF based on echocardiographic features and evaluated their response to spironolactone, finding that a phenogroup characterized by a high proportion of women and elevated blood pressure appeared to benefit from the anti-remodeling effects of spironolactone. However, as this was a retrospective analysis, the findings should be considered hypothesis-generating.26
AI in clinical practice
Clinical implementation and regulatory approval
Despite their considerable potential, the implementation of AI-based models in real-world clinical practice remains in its early stages. Supervised and unsupervised ML models have experimentally demonstrated the potential to outperform conventional diagnostic and risk stratification methods, improve clinical workflow and efficiency, and provide clinician support, while generally receiving positive acceptance among physicians.27, 28 Moreover, AI-based models may reduce costs and lessen the burden on the healthcare system (Figure 4).29 However, the number of AI-based models approved for and implemented in clinical practice remains low compared with the number of experimentally developed models. Currently, regulatory approval of AI-based tools in HF is largely limited to screening, automated imaging quantification, and clinical decision support functions. Examples of approved tools in these areas include EchoGo® Heart Failure (cleared in 2022), an AI-based tool designed to assist in the diagnosis of HFpEF using echocardiographic features.30
Additional examples include Eko Low EF AI (cleared in 2024), an AI tool integrated into a stethoscope that analyzes ECG signals and heart sounds during auscultation,31 and Anumana ECG-AI LEF (cleared in 2023), medical software that analyzes 12-lead ECGs to assess LVEF and detect EF < 40%.19, 32
Regulatory approval of AI tools in HF is subject to strict requirements: human oversight is always required, and AI tools cannot replace standard HF diagnostic processes; a narrowly defined intended use must be specified, and only locked algorithms may be used, meaning that the algorithm cannot continue learning during clinical deployment. These models are approved as adjunctive tools for clinicians rather than autonomous diagnostic or therapeutic systems. It is important to note that regulatory approval of AI models does not necessarily translate into improved patient outcomes.
Limitations of AI models in heart failure
Limited generalizability is a major barrier to the clinical implementation of AI-based models. These models are often developed and trained in single-center settings and evaluated only in internal test environments rather than in real-world clinical practice. As a result, models may demonstrate high predictive accuracy in controlled test settings but substantially lower performance when applied to broader patient populations. This is particularly problematic in countries such as the USA, where certain ethnic groups, including Hispanic and Black populations, remain underrepresented in clinical trials, thereby increasing the risk of perpetuating bias in ML-based algorithms.
Moreover, there is concern that entire countries or regions may be excluded from AI research due to financial disparities.33 External validation is essential to demonstrate a model’s generalizability; however, in real-world settings, it is often limited by the difficulty of obtaining sufficiently large and heterogeneous datasets.30
Another major challenge is the risk of overfitting in ML-based models. These models require large and diverse datasets to ensure robust pattern recognition. Particularly in complex, multifactorial medical problems, insufficient training data may lead algorithms to capture noise, spurious associations, and random fluctuations rather than true underlying relationships. This results in impaired model performance and violates the principle of parsimony.34
Limited explainability, often referred to as the “black box” nature of algorithms such as deep neural networks (DNNs), represents another major barrier to the implementation of AI models. This term refers to opaque and difficult-to-interpret internal decision-making processes, which may reduce trust and acceptance among physicians, patients, and legal stakeholders.35
Outliers may also adversely affect AI-driven models, as extreme values can distort pattern recognition and reduce predictive performance. Increasing dataset size and diversity may mitigate the impact of outliers, thereby improving overall model robustness.34
Future directions
Despite the initial clinical implementation of AI-based models for HF screening, image analysis, and clinical decision support, several major challenges remain. Current AI-based models cannot replace complex therapeutic decision-making and therefore serve primarily as adjunctive tools. Multimorbidity is not yet adequately captured by existing AI models, limiting the scope of their clinical implementation. Furthermore, ethical considerations, including patient preferences and data security, remain unresolved. Addressing these issues should be a priority in future model development and validation.36
Conclusions
AI shows considerable potential for enhancing HF care by improving diagnostic accuracy, risk stratification, treatment individualization, and workflow efficiency, while reducing costs and the burden on the healthcare system. However, currently implemented AI tools remain adjunctive and cannot replace clinical judgment, comprehensive decision-making, or individualized care in complex, multimorbid patients. Clinicians should regard AI as a supportive tool that can supplement, but not replace, specialist clinical assessment and decision-making, while remaining mindful of its limitations regarding generalizability and ethical considerations.
Use of AI and AI-assisted technologies
Not applicable.







