Loading

Customer Churn Prediction Using Machine Learning Models: A Comparative Study Using R SoftwareCROSSMARK Color horizontal
Uppu Venkata Subbarao1, Tedlapu Narayana Rao2, Vantaku Bala3, K.B Rajeswara Rao4

1Dr. Uppu Venkata Subbarao, Department of Basic Sciences, N.S Raju Institute of Technology, Autonomous, Visakhapatnam (Andhra Pradesh), India.

2Dr. Tedlapu Narayana Rao, Department of Management Studies, N.S Raju Institute of Technology, Autonomous, Visakhapatnam (Andhra Pradesh), India.

3Dr. Vantaku Bala, Department of Management Studies, N.S Raju Institute of Technology, Autonomous, Visakhapatnam (Andhra Pradesh), India.

4K.B Rajeswara Rao, Department of Computer Science and Engineering Lendi Institute of Technology, Autonomous Vizianagaram (Andhra Pradesh), India.

Manuscript received on 28 July 2026 | First Revised Manuscript received on 01 August 2026 | Second Revised Manuscript received on 10 August 2026 | Manuscript Accepted on 15 August 2026 | Manuscript published on 30 August 2026 | PP: 5-13 | Volume-12 Issue-12, August 2026 | Retrieval Number: 100.1/ijmh.L189312120826 | DOI: 10.35940/ijmh.L1893.12120826

Open Access | Editorial and Publishing Policies | Cite | Zenodo | OJS | Indexing and Abstracting
© The Authors. Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP). This is an open access article under the CC-BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)

Abstract: Predicting customer churn is essential for improving retention and supporting long-term business growth. In this study, we compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R. Our approach included data preprocessing, exploratory analysis, model development, performance evaluation, and further analysis. We developed and evaluated four classification algorithms: Logistic Regression, Decision Tree, Random Forest, and Extreme Gradient Boosting, using an 80:20 train-test split. We assessed each model’s accuracy, precision, recall, and F1-score. Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset. Feature importance analysis indicated that contract type, customer tenure, monthly charges, total charges, and internet service were the key factors influencing churn. These results suggest that explainable machine learning offers both strong predictive performance and greater transparency. The R-based framework we present provides a practical, reproducible approach to support customer retention strategies and help managers make evidence-based decisions in customer relationship management.

Keywords: Customer Churn; Explainable Machine Learning; Management Science; R programming; Logistic Regression; Random Forest; XGBoost; Decision Tree; Customer Analytics; Decision Support.
Scope of the Article: Business Strategy & Policy