ORCID ID
0009-0003-1067-1318
Date of Award
8-2026
Degree Type
Thesis-Restricted
Degree Name
M.S.
Degree Program
Computer Science
Department
Computer Science
Major Professor
Dr. Md. Tamjidul Hoque
Second Advisor
Dr. Abdullah Al Redwan Newaz
Third Advisor
Dr. Ben Samuel
Abstract
Abstract
Cardiovascular disease (CVD) remains one of the leading causes of mortality worldwide, emphasizing the need for accurate and clinically useful risk prediction models. This thesis aims to develop and evaluate machine learning predictive frameworks for predicting multiple CVD outcomes using data from the Atherosclerosis Risk in Communities (ARIC) study. Four clinically relevant outcomes were considered: fatal coronary heart disease (Fatal CHD), myocardial infarction (MI), stroke, and overall CVD within the next 10 years.
Several machine learning algorithms were implemented and evaluated, including various classical models - random forests (RF), extremely randomized trees (Extra Trees), support vector machines (SVM), logistic regression (LR), extreme gradient boosting (XGBoost) and deep neural network architectures, including a feed-forward ANN, a 1D convolutional network (CNN), a ResNet-34 residual network, a WaveNet network and Stacking ensemble models. Each model was assigned a dedicated preprocessing pipeline to match its structural requirements. These pipelines ensured compatibility with the clinical feature set, and engineered interaction features, allowing consistent and appropriate data preparation across all model types.
All models were evaluated using accuracy, balanced accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, with emphasis on PR-AUC because the cardiovascular outcomes were rare and imbalanced. The results show that no single model consistently dominated across all four outcomes. ANN-SELU achieved the strongest individual performance for fatal coronary heart disease and stroke, tree-based ensemble models were highly competitive for myocardial infarction, and logistic regression, SVM, ANN-SELU, and WaveNet-1D performed similarly for the any-CVD outcome. Stacking ensembles provided modest improvements for some outcomes, especially myocardial infarction, stroke, and any CVD, but did not produce a substantial improvement over the strongest single models.
Additionally, interpretability was analyzed through SHAP-based feature importance, which identified clinically meaningful predictors such as hypertension, sex, plaque measures, diabetes, age-related interactions, smoking, and blood pressure features. Overall, this study shows that machine learning can support multi-outcome cardiovascular risk prediction in ARIC, but improvements are moderate and depend strongly on the outcome, feature representation, and evaluation strategy.
Recommended Citation
Nguyen, Anh, "A Comprehensive Machine Learning Framework for Multi-Outcome Cardiovascular Risk Prediction Using the ARIC Cohort" (2026). LSU New Orleans Theses and Dissertations. 3403.
https://scholarworks.uno.edu/td/3403
Rights
The University of New Orleans and its agents retain the non-exclusive license to archive and make accessible this dissertation or thesis in whole or in part in all forms of media, now or hereafter known. The author retains all other ownership rights to the copyright of the thesis or dissertation.