Predicting Type-2 Diabetes Risk among Adults in Ethiopia: A Machine Learning Approach

Authors

  • Yonas Sileshi Bultum Author
  • Eyob Negussie Author

Keywords:

Machine Learning, Type 2 Diabetes, Classification, Imbalanced Data, Chronic Disease, Diabetes Risk

Abstract

This study aims to develop a machine learning model to predict the risk of Type 2 diabetes (T2D) among Ethiopian adults, using demographic, behavioral, physical, and biochemical data. The research responds to the increasing diabetes burden in Ethiopia, utilizing data from the EPHI NCD STEPS survey, which includes 9,800 instances. Seven machine learning 
models—Decision Tree, Random Forest, Logistic Regression, SVM, KNN, ANN, and XGBoost were employed to assess predictive accuracy, incorporating appropriate preprocessing, class balancing, and evaluation techniques. The Random Forest model consistently outperformed others, achieving an accuracy of 87.79% without resampling. When using SMOTE and 10-fold cross-validation, its accuracy peaked at 93.70%, with precision and recall also at  93.69%, and an F1-score of 93.69%. Furthermore, it yielded an impressive ROC-AUC score of 98.92%, establishing it as the most effective classifier. The findings demonstrate the potential of machine learning to enhance early diabetes risk detection in Ethiopia, highlighting key lifestyle predictors such as alcohol use, fruit and vegetable intake, waist-hip ratio, and BMI. The study advocates for the integration of machine learning-driven screening tools into national health programs, proposing a scalable, cost-effective strategy for the early diagnosis and prevention of Type 2 diabetes. 

Published

2026-05-20