A COMPARATIVE ANALYSIS OF MACHINE LEARNING ALGORITHMS FOR PREDICTING DIABETES MELLITUS
Abstract
Diabetes mellitus continues to pose a significant global health challenge, with increasing
prevalence across both high and low income countries. The inability to detect and treat diabetes
effectively at its early stages has fueled the need for reliable predictive models. This research
addresses the problem of determining the most accurate machine learning algorithm for predicting
diabetes. The aim is to perform a comparative analysis of four machine learning algorithms,
Random Forest (RF), Extreme Gradient Boost Classifier (XGBC), Support Vector Machine
(SVM), and Naïve Bayes Classifier (NBC) to assess their predictive accuracy. The study uses
diabetes datasets from the UCI Machine Learning Repository, applying feature selection,
normalization, and statistical evaluation metrics like accuracy, precision, recall, and F1-Score to
compare these models. The Results showed that the Random Forest algorithm outperforms the
others, achieving 99% accuracy, followed by XGBoost at 93%, while SVM and Naive Bayes
performed less effectively with accuracies of 90% and 83%, respectively. This project provides
valuable insights into the effectiveness of different machine learning techniques for diabetes
prediction, guiding healthcare professionals in model selection for early diagnosis and
management.
Full-Text Access Notice
In accordance with the NERD Policy on promoting peer-reviewed publication, public access to the full text of a project, thesis or dissertation is restricted for three years, allowing the author and supervisors sufficient time to pursue peer-reviewed publication.
During this period, researchers with legitimate academic or research purposes may request authorisation directly from the author to enable NERD to release the indexed full texts of the work using the form below.