| Communications on Applied Electronics |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 8 - Number 2 |
| Year of Publication: 2026 |
| Authors: Oluwaponmile H. Adefuwa, Olatunde D. Akinrolabu, Ayodeji O. Ayeni |
10.5120/cae1e52adc932f2
|
Oluwaponmile H. Adefuwa, Olatunde D. Akinrolabu, Ayodeji O. Ayeni . An Inclusive Diabetes Detection System using Machine Learning and Explainable Artificial Intelligence. Communications on Applied Electronics. 8, 2 ( Sep 2026), 9-21. DOI=10.5120/cae1e52adc932f2
Diabetes mellitus is a chronic metabolic disorder affecting over 537 million adults globally, with prevalence expected to reach 783 million by 2045. Even with advances in diagnostic technology, equitable and early detection remains a challenge, particularly across demographically diverse populations. This study presents the design and development of an inclusive diabetes detection system that combines machine-learning classification, Local Interpretable Model-agnostic Explanations (LIME)-based explainability, and interactive Streamlit web deployment. A publicly available synthetic dataset of 100,000 patient records was preprocessed through encoding, Min-Max normalisation, and SMOTE-based class balancing. Four machine learning models (XGBoost, Random Forest, Naive Bayes, and Long Short-Term Memory (LSTM)) were trained and evaluated using accuracy, precision, recall, F1-score, and AUC-ROC. Hyperparameter tuning was conducted using Grid Search with 5-fold cross-validation. XGBoost achieved the best performance with 91.99% accuracy, precision of 1.0000, recall of 0.8665, F1-score of 0.9285, and AUC-ROC of 0.9438, clearly exceeding the research target of 70% accuracy. Demographic subgroup analysis across gender, ethnicity, age group, and income level confirmed equitable performance, with F1-score variation below 0.004 across most demographic dimensions. LIME generated clinically meaningful individual-level explanations, consistently identifying HbA1c and fasting glucose as the dominant predictive features. A broader evaluation was done, the trained XGBoost model was further tested on the Pima Indians Diabetes Dataset, a widely used benchmark with only 8 features, achieving 72.53% accuracy and AUC-ROC of 0.7969 using only 6 of 38 model features, demonstrating generalisation capability beyond the original training distribution. The complete system was deployed as a browser-based Streamlit application enabling real-time prediction and transparent explanation for clinical and non-clinical users.