Research Article
Determining Loan Eligibility in the Banking Sector Using a Hybrid Model of Support Vector Machine and Extreme Gradient Boosting
James Githaiga Muhoro*,
Harun Mwangi Gitonga,
Josephine Njeri Ngure
Issue:
Volume 12, Issue 3, June 2026
Pages:
37-51
Received:
20 May 2026
Accepted:
30 May 2026
Published:
3 July 2026
DOI:
10.11648/j.ijdsa.20261203.11
Downloads:
Views:
Abstract: The loan prediction models are ever-changing due to changes in technology, whereby financial institutions are adopting various technologies to automate the loan process. The surge of loan applicants with diverse attributes has accelerated the need for machine learning models which can incorporate different applicant attributes and improve accuracy in determining loan eligibility. While the use of individual machine learning models was more robust and accurate, these models have some limitations that may hinder the achievement of optimal results when establishing loan eligibility. Thus, the need to combine two or more individual models and leverage their strengths to improve accuracy and robustness. This study aimed to develop an eXtreme Gradient Boosting (XGBoost)-Support Vector Machine (SVM) hybrid model to determine loan eligibility in the banking sector. This work utilized secondary financial dataset from Google Kaggle. The dataset was preprocessed, transformed and used to train and test the models. Evaluation metrics, namely accuracy, precision, F1-score, recall and Receiver Operating Characteristics-Area Under Curve (ROC-AUC), were used to evaluate the performance and reliability of the XGBoost, SVM and the XGBoost-SVM hybrid model. The XGBoost-SVM hybrid model posted strong performance metrics results with accuracy of 0.78, precision of 0.27, a balanced recall of 0.49, F1-Score of 0.34 and AUC-ROC curve of 0.74. It was therefore evident that the hybrid model was able to leverage the standalone model's strength for better performance in determining loan eligibility.
Abstract: The loan prediction models are ever-changing due to changes in technology, whereby financial institutions are adopting various technologies to automate the loan process. The surge of loan applicants with diverse attributes has accelerated the need for machine learning models which can incorporate different applicant attributes and improve accuracy ...
Show More
Research Article
Random Forest Classification for Type 2 Diabetes Risk Prediction in Kenya: Evidence from the 2022 Kenya Demographic and Health Survey
Diana Boro*
Issue:
Volume 12, Issue 3, June 2026
Pages:
52-63
Received:
29 June 2026
Accepted:
9 July 2026
Published:
24 July 2026
DOI:
10.11648/j.ijdsa.20261203.12
Downloads:
Views:
Abstract: Type 2 diabetes mellitus (T2DM) is an escalating public health burden in sub-Saharan Africa, where over half of affected individuals remain undiagnosed at clinical presentation. In Kenya, early risk identification is constrained by limited population-level screening infrastructure and the underutilization of data-driven predictive tools. This study applied a Random Forest (RF) machine learning classifier to the nationally representative 2022 Kenya Demographic and Health Survey (KDHS 2022) dataset to develop and evaluate a predictive model for T2DM risk among Kenyan adults aged 15–54 years. The analytical sample comprised 31,354 respondents drawn from all 47 counties of Kenya, of whom 272 (0.87%) were classified as diabetic, reflecting a severe class imbalance ratio of 114.3:1. Data preprocessing encompassed median imputation for missing values, min-max normalisation of continuous predictors, one-hot encoding of categorical variables, and Recursive Feature Elimination with Cross-Validation (RFECV) for feature selection. The Synthetic Minority Over-sampling Technique (SMOTE) was applied exclusively to the training partition to address class imbalance without contaminating test-set evaluation. The RF model was optimised via exhaustive grid search with five-fold stratified cross-validation, yielding a cross-validation Area Under the Receiver Operating Characteristic Curve (AUC-ROC) of 0.9895 (standard deviation (SD = 0.0004). At a calibrated classification threshold of 0.20, the model achieved a test-set sensitivity of 62.96%, specificity of 69.81%, balanced accuracy of 66.40%, and Matthews Correlation Coefficient (MCC) of 0.0659, correctly identifying 34 of 54 diabetic cases in the held-out test set. SHAP (SHapley Additive exPlanations) analysis identified age, wealth index, hypertension, employment status, and Body Mass Index (BMI) as the dominant predictors of T2DM risk. These findings establish a reproducible, nationally representative RF-based screening framework with direct implications for targeted public health intervention and early detection policy in Kenya.
Abstract: Type 2 diabetes mellitus (T2DM) is an escalating public health burden in sub-Saharan Africa, where over half of affected individuals remain undiagnosed at clinical presentation. In Kenya, early risk identification is constrained by limited population-level screening infrastructure and the underutilization of data-driven predictive tools. This study...
Show More