EXPLAINABLE MACHINE LEARNING FOR PHISHING URL DETECTION USING SHAP INTERPRETATION
Main Article Content
Abstract
Phishing attacks delivered through malicious URLs represent an increasingly prevalent cyber threat capable of causing significant harm to users. Although machine-learning–based phishing detection has been widely explored, most existing models still operate as black boxes, making their classification decisions difficult to interpret. This study proposes an Explainable Machine Learning framework for phishing URL detection by integrating five algorithms—XGBoost, Random Forest, Gradient Boosting, Decision Tree, and K-Nearest Neighbors—augmented with SHAP (SHapley Additive Explanations) for interpretability. The dataset includes structural URL features such as character length, special symbol counts, number of subdomains, and string entropy. Model performance was evaluated using accuracy, precision, recall, F1-score, and confusion matrix to enable comparative assessment among algorithms. The results show that XGBoost achieves the best performance, obtaining 97.8% accuracy, an F1-score of 0.976, and stable predictions across all classes. Random Forest ranks second with 96.4% accuracy, followed by Gradient Boosting at 95.7%. Meanwhile, Decision Tree and KNN exhibit lower performance due to their higher sensitivity to data variation. SHAP analysis reveals that the most influential features in phishing prediction include URL length, special character frequency, entropy levels, and the number of subdomains. These findings demonstrate that integrating XAI not only enhances model transparency but also ensures that phishing detection systems remain accurate, interpretable, and accountable.
Article Details
References
[2] Fauzan, A. V. Vitianingsih, D. Cahyono, A. L. Maukar, and Y. A. B. Suprio, “Penerapan Algoritma Klasifikasi pada Machine Learning untuk Deteksi Phishing: Application of Classification Algorithms in Machine Learning for Phishing Detection”, MALCOM, vol. 5, no. 2, pp. 531-540, Mar. 2025.
[3] A. F. Mahmud and S. Wirawan, “Deteksi Phishing Website menggunakan ML,” Jurnal Sistemasi, vol. 13, no. 4, 2024. doi: 10.32520/stmsi.v13i4.3456.
[4] G. A. Rahmat, R. Sarno, K. R. Sungkono, A. T. Haryono, A. F. Septiyanto, and Sholiq, “Detecting Phishing Website using Machine Learning Methods,” in Proceedings of the ICoCSETI 2025 - International Conference on Computer Sciences, Engineering, and Technology Innovation, F. W. Wibowo, Ed. Institute of Electrical and Electronics Engineers (IEEE), 2025, pp. 351–356. doi: 10.1109/ICoCSETI63724.2025.11019127.
[5] A. Khairunnisya, L. Lindawati, and S. Zefi, “Phishing URL Detection System Using Random Forest and Gradient Boosting for Cybercrime Prevention,” CSRID (Computer Science Research and Its Development Journal), vol. 17, no. 3, pp. 296–310, 2025. doi: 10.22303/csrid-.17.3.2025.296-310.
[6] M. Irsan, F. Febriana, H. H. Nuha and H. R. Putra Sailellah, "Phishing Detection on URL Data Using K-Nearest Neighbors Method," 2024 International Conference on Innovation and Intelligence for Informatics, Computing, and Technologies (3ICT), Sakhir, Bahrain, 2024, pp. 792-797, doi: 10.1109/3ict64318.2024.10824630. keywords: {Uniform resource locators;Technological innovation;Accuracy;Phishing;Computational modeling;Nearest neighbor methods;Feature extraction;Data models;Decision trees;Overfitting;Cybersecurity;Phishing Detection;K-Nearest Neighbor (KNN);Decision Trees (DT);URLs},
[7] Preeti and P. Sharma, “Enhancing phishing URL detection through comprehensive feature selection: a comparative analysis across diverse datasets,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 36, pp. 1182–1188, 2024. doi: 10.11591/ijeecs.v36.i2.pp1182-1188.
[8] Tianqi Chen and Carlos Guestrin. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '16). Association for Computing Machinery, New York, NY, USA, 785–794. https://doi.org/10.1145/2939672.2939785
[9] O. K. Sahingoz, E. Buber, O. Demir, and B. Diri, “Machine learning based phishing detection from URLs,” Expert Systems with Applications, vol. 117, pp. 345–357, 2019, doi: 10.1016/j.eswa.2018.09.029.
[10] Rantamaa, H.-R., Kangas, J., Kumar, S. K., Mehtonen, H., Järnstedt, J., & Raisamo, R. (2023). Comparison of a VR Stylus with a Controller, Hand Tracking, and a Mouse for Object Manipulation and Medical Marking Tasks in Virtual Reality. Applied Sciences, 13(4), 2251. https://doi.org/10.3390/app13042251
[11] P. R. E. Gabriel, “Deteksi URL phishing menggunakan hybrid deep learning CNN dan XGBoost dengan teknik balancing SMOTE-ENN,” in Prosiding Seminar Nasional Informatika Bela Negara (SANTIKA), vol. 5, no. 2, pp. 70–79, Dec. 2025.
[12] S. Romadi, “Perbandingan kinerja algoritma gradient boosting dan multilayer perceptron dalam klasifikasi website phishing,” Diss., Universitas Mercu Buana Jakarta, 2025.
[13] H. S. Lallie, L. A. Shepherd, J. R. C. Nurse, A. Erola, G. Epiphaniou, C. Maple, and X. Bellekens, “Cyber security in the age of COVID-19: A timeline and analysis of cyber-crime and cyber-attacks during the pandemic,” Computers & Security, vol. 105, Art. no. 102248, 2021, doi: 10.1016/j.cose.2021.102248.
[14] “KLASIFIKASI EMAIL PHISHING MENGGUNAKAN ALGORITMA K-NEAREST NEIGHBOR”, Restikom, vol. 5, no. 2, pp. 148–157, Aug. 2023, doi: 10.52005/restikom.v5i2.152.
[15] Mafaza, r. V. Klasifikasi malicious url pada file menggunakan metode k-nearest neighbor berdasarkan lexical feature extraction.
[16] S. Sunaryono, “Penelitian komparasi algoritma klasifikasi dalam menentukan website palsu,” Teknikom: Teknologi Informasi, Ilmu Komputer dan Manajemen, vol. 1, no. 1, pp. 1–11, 2017.