Logo
American Journal of Medical Research & Reviews

American Journal of Medical Research & Reviews

Research Article

Machine-Learning and Neural-Network Prediction of Diabetes Status Using Glycemic, Lipid, Renal, and Anthropometric Biomarkers: Analysis of a Public Kaggle Dataset

Authors: Ahed J Alkhatib*


Abstract

Background: Diabetes mellitus is an important global metabolic disorder, and early identification remains important for reducing long-term cardiovascular, renal, and metabolic complications.


Objective: This study evaluated clinical, biochemical, and machine-learning markers for differentiating normal, prediabetic and diabetic status focusing on prediabetes diagnosis and models performance on exclusion of HbA1c.


Methods: The cleaned data set was analyzed for 1,000 participants of which 103 normal, 53 prediabetic, 844 diabetics. We used various statistics and tests to make comparisons across diabetes classes like Spearman Correlation and more. For analysis in machine-learning, repeats of clinical profiles were removed to obtain 826 records. A set of classical machine-learning models and neural networks were evaluated on binary and multiclass classification using accuracy, balanced accuracy, sensitivity/recall, F1-score, and ROC-AUC. To avoid
any diagnostic circularity, we prioritized models that did not utilize HbA1c.


Results: After cleaning, the dataset included 103 normal, 53 prediabetic, and 844 diabetic records. After clinical deduplication, 826 unique records remained for machine-learning analysis. HbA1c and BMI showed the strongest differences across diabetes classes, with large Kruskal-Wallis effect sizes. HbA1c correlated moderately with BMI and age. In multivariable logistic regression excluding HbA1c, BMI, cholesterol, and age independently predicted diabetic status. In binary machine-learning analysis without HbA1c, random forest achieved the best performance, with 95.6% accuracy, 94.7% balanced accuracy, 96.1% sensitivity, 93.4% specificity, 97.4% F1- score, and 98.2% ROC-AUC. Neural networks performed well but did not outperform random forest; the best neural network achieved 90.4% balanced accuracy and 95.9% ROC-AUC. Multiclass classification was more difficult, mainly due to the small prediabetic group.


Conclusion: Information about BMI, age, cholesterol, triglycerides and TG/HDL ratio can help detect diabetes and prediabetes. Overall, the random forest models performed the best, but the identification of prediabetes was still difficult due to class imbalance and overlapping metabolic profiles.

Citation: Ahed J Alkhatib, Machine-Learning and Neural-Network Prediction of Diabetes Status Using Glycemic, Lipid, Renal, and Anthropometric Biomarkers: Analysis of a Public Kaggle Dataset. American J Med Res Rev. 2026; 1(1), 01-18.