Artificial intelligence (AI) in cardiology: interpreting evidence across atrial fibrillation, myocardial infarction and PCI
AI is moving from a broad cardiology concept towards specific clinical tasks: finding an atrial fibrillation signal hidden within a sinus-rhythm electrocardiogram, interpreting patterns in routine blood tests, and quantifying coronary anatomy during stent implantation. Three recent studies show why that progress is notable—and why specialist interpretation, validation and outcome evidence remain essential.
Why this matters
The most persuasive question in cardiovascular AI is not whether an algorithm can generate a score. It is whether a score, classification or measurement adds reliable information at a moment when it can be used appropriately. The source studies address exactly those moments: before non-cardiac surgery, during assessment for acute myocardial infarction, and at the time of percutaneous coronary intervention.
Their appeal is practical. The atrial fibrillation model uses a single routine 12-lead electrocardiogram. The acute myocardial infarction model uses routinely measured haematological information and derived ratios. The AI-based quantitative coronary angiography system uses coronary angiograms already acquired during intervention. These are potentially lower-friction applications than tools requiring wholly new data collection. Yet convenience should not be mistaken for established clinical benefit.
Implementation will be shaped by local realities. Electrocardiogram protocols, laboratory methods, angiographic acquisition and sensor availability are not uniform. Acute stroke pathways are not the focus of these studies, but they illustrate how any time-sensitive clinical pathway depends on clearly defined imaging, decision and escalation timelines. Dataset transfer can be affected by patient mix, scanner vendor and imaging field strength where relevant, as well as by differences in routine care. A model can perform well retrospectively, or an automated measurement can work during a trial, while still requiring careful testing within existing local preoperative, emergency and catheterisation-laboratory systems. Adoption depends on workflow fit, oversight and evidence that use of the output improves decisions or outcomes.
Evidence at a glance
| Study | Setting/population | Clinical question | AI method | Endpoint | Main result | Key limitation | Practical relevance |
|---|---|---|---|---|---|---|---|
| Lee et al. AI-electrocardiogram study | 13,687 eligible non-cardiac surgical patients in Taiwan | Can a sinus-rhythm electrocardiogram detect hidden atrial fibrillation risk? | Deep-learning AI-electrocardiogram | 30-day postoperative new-onset atrial fibrillation and outcomes | High-risk classification was associated with higher postoperative event rates; AUC 0.8107 for new-onset atrial fibrillation | Retrospective design and incomplete postoperative ascertainment | Could support research into targeted preoperative assessment |
| Altıntop haematological classifier | 981 individuals in an open-access dataset | Can haematological variables classify acute myocardial infarction and subtypes? | ReliefF plus Random Forest and AdaBoost | Five-fold cross-validated classification | Random Forest reached 84.92% binary accuracy and 80.75% balanced multiclass accuracy | Single-source data; healthy-control comparator; no external validation | A potentially reproducible model-development approach |
| FLASH randomised trial | 400 patients at 13 South Korean centres | Can AI-based quantitative coronary angiography guide stenting noninferiorly to optical coherence tomography? | Fully automated quantitative coronary angiography | Post-intervention minimal stent area | 6.3 versus 6.2 mm²; noninferiority P<0.001 | Surrogate endpoint and exclusion of complex anatomy | Supports further procedural evaluation in selected lesions |
Study-by-study clinical interpretation
Lee et al.: a preoperative electrocardiographic signal for hidden atrial fibrillation
Atrial fibrillation can be clinically silent, intermittent and difficult to identify from standard episodic testing. This creates a relevant perioperative challenge, particularly when patients present in sinus rhythm before non-cardiac surgery. Lee et al. examined whether deep learning applied to a single 10-second, 12-lead sinus-rhythm electrocardiogram could identify patients with pre-existing atrial fibrillation or a high likelihood of postoperative new-onset atrial fibrillation within 30 days.
The retrospective cohort came from two hospitals in Taipei. Of 107,903 screened individuals, 13,687 were included after exclusions. There were 98 patients with untreated pre-existing atrial fibrillation, 54 with postoperative new-onset atrial fibrillation and 13,526 controls. The reported 30-day incidence of postoperative new-onset atrial fibrillation was 0.4%, with higher incidence in higher-risk surgery and in some surgical specialties.
The model performed well in its reported development and validation analyses, with AUC values ranging from 0.87 to 0.91 across hidden atrial fibrillation-related tasks. In the perioperative cohort without known pre-existing atrial fibrillation, it classified 1.3% of patients as high risk and 11.7% as medium risk. Ten high-risk patients developed postoperative new-onset atrial fibrillation within one month. The high-risk group also had a reported 17.33-fold higher 30-day all-cause mortality hazard than the low-risk group; the confidence interval was wide, as expected for a small high-risk subgroup.
Compared with C₂HEST and the Taiwan AF Score, the AI-electrocardiogram had higher reported discrimination for 30-day postoperative new-onset atrial fibrillation and mortality. This is an important signal because conventional scores rely on clinical information and were designed mainly for longer-term atrial fibrillation prediction. The study suggests that an electrocardiogram may carry information about atrial substrate or broader cardiovascular vulnerability that is not captured by those scores alone.
However, the interpretation should remain cautious. The model’s high-risk output is not synonymous with a diagnosis of atrial fibrillation. Atrial fibrillation ascertainment relied on 12-lead electrocardiograms, not continuous monitoring, and postoperative electrocardiograms were performed when clinically indicated rather than routinely. Undetected episodes may therefore have remained in the control group. The observed association with adverse outcomes is clinically meaningful, but does not prove that the electrocardiographic features themselves cause those outcomes or that intervening on a score would alter them.
Clinical interpretation: this study offers a compelling research case for AI-supported preoperative risk stratification from existing electrocardiograms. It does not yet establish a standard monitoring protocol, anticoagulation approach or outcome benefit from AI-guided care.
Altıntop: turning routine blood variables into an acute myocardial infarction classifier
The Altıntop study takes a different route. Rather than analysing raw signals, it combines familiar blood variables with engineered haematological indices and interpretable machine-learning methods. The dataset included 477 people with acute myocardial infarction and 504 healthy controls, collected between January 2020 and May 2023. Acute myocardial infarction cases included both ST-segment elevation and non-ST-segment elevation myocardial infarction.
The investigators assessed 31 features. Alongside routine blood measures, they examined seven derived indices: mean platelet volume-to-platelet count ratio, platelet-to-white blood cell ratio, red cell distribution width-to-lymphocyte ratio, red cell distribution width-to-platelet ratio, Systemic Immune-Inflammation Index, Systemic Inflammation Response Index and derived neutrophil-to-lymphocyte ratio. These variables were selected to reflect inflammatory, platelet and red-cell patterns that may be associated with myocardial infarction.
ReliefF was used to rank variables, with Random Forest and AdaBoost forming the principal models. In binary classification, Random Forest achieved 84.92% ± 2.02% mean accuracy, while AdaBoost achieved the highest reported AUC at 91.45% ± 1.40%. The distinction between binary and multiclass results is important. Accuracy fell to approximately 73% in the unbalanced three-class problem, then increased to around 81% after use of the Synthetic Minority Oversampling Technique (Synthetic Minority Over-sampling Technique (SMOTE)).
That improvement demonstrates a familiar machine-learning issue: class distribution changes performance estimates. SMOTE generated synthetic observations to balance less frequent classes; it did not add independent patient examples. Likewise, low fold-to-fold standard deviations support consistency within the cross-validation scheme but cannot demonstrate generalisability outside the source dataset.
The novel variables were statistically different between groups in many analyses, but the ReliefF results also showed that conventional features—such as basophils, monocytes and white blood cells—often had greater weight. The model’s apparent utility is therefore multivariable and contextual. It should not be read as evidence that any one derived index is a definitive biomarker.
Clinical interpretation: the framework makes an accessible case for further research because it relies on routinely available variables and reports an interpretable feature-selection approach. It does not test performance in an unselected emergency population, replace standard diagnostic evaluation or establish an effect on time to treatment or patient outcomes.
FLASH: automated angiography meets a procedural imaging comparator
The FLASH trial brings AI into the catheterisation laboratory. Unlike the first two studies, it was prospective, multicentre and randomised. The investigators enrolled 400 participants at 13 South Korean hospitals and compared AI-based quantitative coronary angiography-assisted percutaneous coronary intervention with optical coherence tomography-guided intervention.
AI-based quantitative coronary angiography automatically provided measurements of lesion length, lumen diameters and stenosis from angiographic images. In the trial protocol, these measurements guided stent sizing and high-pressure postdilatation. The comparator arm used optical coherence tomography measurements for procedural planning. Both strategies were protocolised, which is important: the trial evaluates not just software but an AI-based quantitative coronary angiography-guided procedural approach.
The primary endpoint was minimal stent area measured by final optical coherence tomography. Of 395 participants with analysable final images, mean minimal stent area was 6.3 ± 2.2 mm² in the automated angiography group and 6.2 ± 2.2 mm² in the optical coherence tomography group. The result met the prespecified noninferiority criterion, while superiority was not demonstrated. Procedural duration and contrast volume did not differ significantly.
The trial also highlights a limitation of interpreting a single favourable endpoint. Stent malapposition was more common in the automated angiography group, at 13.6% compared with 5.6%. Final optical coherence tomography in the automated group identified findings that prompted additional treatment in 33 patients after primary-endpoint measurement. Six-month clinical events were rare and not significantly different, but low event rates and limited follow-up prevent firm conclusions about long-term comparative safety or efficacy.
The population was deliberately selective. Left main lesions, chronic total occlusions, graft lesions and bifurcation lesions requiring two stents were excluded. This reflects known difficulties in defining suitable reference segments for quantitative angiography in complex anatomy.
Clinical interpretation: FLASH is an important randomised demonstration that automated angiographic measurement can guide selected percutaneous coronary intervention to a noninferior minimal stent area compared with optical coherence tomography guidance. It does not establish a replacement role for intracoronary imaging in complex disease or prove long-term clinical equivalence.
Why the studies should or should not be compared directly
The studies share an AI label but not a common endpoint or evidence standard. Lee et al. assessed retrospective prediction and outcome associations in surgical patients. Altıntop assessed internally cross-validated diagnostic classification in a dataset of cases and controls. FLASH tested a procedural strategy in a randomised trial and used an imaging-derived surrogate endpoint.
Their apparent performance measures are therefore not interchangeable. The area under the curve of a preoperative atrial fibrillation model, the accuracy of an acute myocardial infarction classifier and noninferiority of post-procedure minimal stent area each describe different properties. Their inputs, cohorts, reference standards, thresholds and clinical settings differ. A balanced reading treats them as complementary examples of AI applied to different cardiology problems, rather than as competing demonstrations of the same benefit.
What this means for clinical practice
For specialists, the practical opportunity is targeted support rather than wholesale substitution. An AI-electrocardiogram could be evaluated as an additional preoperative signal where routine electrocardiograms already exist. A haematological classifier could be studied within laboratory and emergency workflows, provided it is tested against the clinical populations in which it would be used. Automated quantitative angiography could provide objective, real-time measurements during selected percutaneous coronary intervention, particularly where its trial population and procedural protocol are relevant.
The practical caution is equally clear. Risk stratification must be connected to a safe, evidence-supported response; diagnostic classification must be tested against meaningful alternatives; and procedural guidance must be judged beyond a single surrogate. The source does not establish that any of these technologies should replace clinical assessment, conventional testing or specialist procedural judgement.
What remains uncertain
Each application needs local validation. The atrial fibrillation study was based in Taiwan, the acute myocardial infarction dataset was single-centre, and FLASH was conducted in South Korea. Before implementation, organisations would need to consider the governance of electrocardiographic, laboratory and angiographic data; the documentation of AI outputs; and the ability to audit model performance and clinical responses. Accountability for a decision remains particularly important where an output may influence surveillance, triage or procedural optimisation. These operational questions are not answered by the supplied studies.
Generalisability also remains unresolved for patients under-represented by age, ethnicity, comorbidity or local practice patterns. The source specifically identifies potential bias related to age, sex, race and ethnicity. Subgroup performance, calibration and appropriate safeguards need prospective assessment. Future studies should determine not only whether models retain their technical performance, but whether they improve real-world decisions and patient-centred outcomes across diverse services.
Conclusion
Across three cardiology settings, AI appears capable of extracting clinically relevant patterns from established data sources. The AI-electrocardiogram study identifies a perioperative risk signal associated with postoperative atrial fibrillation and adverse outcomes; the haematological study reports internally cross-validated acute myocardial infarction classification; and FLASH demonstrates noninferior minimal stent area with automated angiographic guidance in selected percutaneous coronary intervention. The evidence is encouraging because it is specific, not because it is universal. Retrospective associations, cross-validation metrics and a surrogate-endpoint noninferiority result each have value, but they answer different questions. The next evidence requirements are clear: external validation, prospective workflow evaluation, assessment in diverse populations and demonstration of meaningful clinical impact.
Educational content only. It does not replace clinical assessment, current guidelines, or patient-specific professional advice.