A Comparative Study of SMOTE Variants and Particle Swarm Optimization for Feature Selection and Hyperparameter Tuning on Decision Tree in Breast Cancer Classification
DOI:
https://doi.org/10.70656/ijcse.v3i01.661Keywords:
Breast Cancer, Decision Tree, Particle Swarm Op-timization, SMOTE-Tomek, Feature Selection, Hyperparameter TuningAbstract
Breast cancer is one of the most prevalent types of cancer in Indonesia, making accurate early detection highly crucial to reduce mortality rates. However, classification effectiveness is often hindered by high-dimensional data and class imbalance issues, which can introduce bias into predictive models. This study proposes the integration of the Decision Tree algorithm with hybrid sampling techniques and Particle Swarm Optimization (PSO). PSO is employed to perform a dual role, namely simultaneous feature selection and hyperparameter tuning. Experimental results show that using 30 particles in PSO successfully reduced the data dimensionality significantly from 30 features to only 6 essential features: texture1, symmetry1, radius2, area3, smoothness3, and symmetry3. The combination of SMOTE-Tomek and PSO-based feature selection emerged as the best-performing scenario, achieving an accuracy of 99.68% while producing only one False Negative prediction. In addition to superior precision, the model demonstrated remarkable computational efficiency with an average latency of 0.0036 ms and a throughput of 274,897 samples per second. Explainable AI (XAI) analysis using SHAP confirmed that area3 was the most dominant feature, which is clinically consistent with indicators of cancer cell proliferation. This study proves that the synergy between data balancing techniques and metaheuristic optimization can produce an accurate, transparent, and efficient diagnostic model suitable for real-time medical implementation.
Downloads
References
[1] A. Latifah, A. Alrizal, and Y. Fendriani, “Klasifikasi Penyakit Kanker Payudara pada Citra Mammogram Menggunakan Algoritma Convolu-tional Neural Network (CNN) dan Random Forest,” Journal of Online Physics, vol. 10, no. 3, pp. 95–103, Jul. 2025.
[2] H. Oktavianto and R. P. Handri, “Analisis Klasifikasi Kanker Payudara Menggunakan Algoritma Naive Bayes,” INFORMAL Informatics Jour-nal, vol. 4, no. 3, p. 117, 2020.
[3] I. N. Atthalla, A. Jovandy, and H. Habibie, “Klasifikasi Penyakit Kanker Payudara Menggunakan Metode K-Nearest Neighbor,” in Proceedings of Annual Research Seminar, vol. 4, no. 1, 2018, pp. 978–979.
[4] D. Kashyap, D. Pal, R. Sharma, V. K. Garg, N. Goel, D. Koundal,
A. Zaguia, S. Koundal, and A. Belay, “Global Increase in Breast Cancer Incidence: Risk Factors and Preventive Measures,” BioMed Research International, vol. 2022, pp. 1–16, 2022.
[5] World Health Organization, “Breast Cancer,” https://www.who.int/news-room/fact-sheets/detail/breast-cancer, 2024, [Online; accessed 10-Apr-2026].
[6] A. Rizka, M. K. Akbar, and N. A. Putri, “Carcinoma Mammae Sinistra T4bN2M1 Metastasis Pleura,” AVERROUS: Jurnal Kedokteran dan Kesehatan Malikussaleh, vol. 8, no. 1, p. 23, 2022.
[7] C. Chazar and B. Erawan, “Machine Learning Diagnosis Kanker Payu-dara Menggunakan Algoritma Support Vector Machine,” Informasi: Jurnal Informatika dan Sistem Informasi, vol. 12, no. 1, pp. 67–80, 2020.
[8] J. J. Pangaribuan and V. Angkasa, “Komparasi Tingkat Akurasi Random Forest dan KNN untuk Mendiagnosis Penyakit Kanker Payudara,” Journal Information System Development (ISD), vol. 7, no. 1, Jan. 2022.
[9] J. O. Afolayan, M. O. Adebiyi, M. O. Arowolo, C. Chakraborty, and
A. A. Adebiyi, “Breast Cancer Detection Using Particle Swarm Opti-mization and Decision Tree Machine Learning Technique,” in Intelligent Healthcare. Springer, Jun. 2022, pp. 61–83.
[10] C. Alna and C. R. Gunawan, “Pendekatan Machine Learning untuk Meningkatkan Akurasi Klasifikasi pada Dataset Tidak Seimbang,” JID (Jurnal Info Digit), vol. 2, no. 3, Sep. 2024.
[11] M. A. Latief, L. R. Nabila, W. Miftakhurrahman, S. Ma’rufatullah, and H. Tantyoko, “Handling Imbalance Data Using Hybrid Sampling SMOTE-ENN in Lung Cancer Classification,” International Journal of Engineering and Computer Science Applications (IJECSA), vol. 3, no. 1,
pp. 11–18, Mar. 2024.
[12] M. Rahman, “Perbandingan SMOTE-Variants untuk Mengatasi Ketidak-seimbangan Data pada Prediksi Cacat Software,” Banjarbaru, Nov. 2023.
[13] T. O. Omotehinwa, D. O. Oyewola, and E. G. Dada, “A Light Gradient-Boosting Machine Algorithm with Tree-Structured Parzen Estimator for Breast Cancer Diagnosis,” Healthcare Analytics, vol. 4, p. 100218, Dec. 2023.
[14] S. Widodo, H. Brawijaya, and S. Samudi, “Stratified K-Fold Cross Validation Optimization on Machine Learning for Prediction,” Sinkron, vol. 6, no. 4, pp. 2407–2414, Oct. 2022.
[15] Wijiyanto, A. I. Pradana, Sopingi, and V. Atina, “Teknik K-Fold Cross Validation untuk Mengevaluasi Kinerja Mahasiswa,” Jurnal Algoritma, vol. 21, no. 1, 2023. [Online]. Available: https://jurnal.itg.ac.id/index.php/algoritma
[16] M. A. S. Aritonang, M. J. Simanulang, T. P. Batubara, I. Zega, and
M. H. Afrizal, “Natural Language Processing (NLP) and Support Vector Machine (SVM) Optimization in Detecting Phishing Website URLs,” Jurnal Teknik Informatika (JUTIF), vol. 7, no. 1, pp. 552–570, Feb. 2026.
[17] R. Ghorbani and R. Ghousi, “Comparing Different Resampling Methods in Predicting Students’ Performance Using Machine Learning Tech-niques,” IEEE Access, vol. 8, pp. 67 899–67 911, Apr. 2020.
[18] H. Ding, L. Chen, D. Liang, Z. Fu, and X. Cui, “Imbalanced Data Classification: A KNN and Generative Adversarial Networks-Based Hybrid Approach for Intrusion Detection,” Future Generation Computer Systems, vol. 131, p. 7, Feb. 2022.
[19] F. Li, W. Ma, H. Li, and J. Li, “Improving Intrusion Detection System Using Ensemble Methods and Over-Sampling Technique,” in Proceed-ings of the 4th International Academic Exchange Conference on Science and Technology Innovation (IAECST), Dec. 2022.
[20] D. S. Assyifa and A. Luthfiarta, “SMOTE-Tomek Re-sampling Based on Random Forest Method to Overcome Unbalanced Data for Multi-class Classification,” Inform: Jurnal Ilmiah Bidang Teknologi Informasi dan Komunikasi, vol. 9, no. 2, p. 151, Jul. 2024.
[21] A. S. Biyantoro and B. Prasetiyo, “Application of Decision Tree for Health Status Classification, Compared to KNN and Naive Bayes,” IJIRSE: Indonesian Journal of Informatic Research and Software Engi-neering, vol. 4, no. 1, pp. 47–55, Mar. 2024.
[22] B. P. Lohani, A. Dagur, and D. K. Shukla, “An Efficient Approach for Diabetes Classification Using Feature Selection and Hyperparameter Tuning,” Recent Advances in Electrical & Electronic Engineering, vol. 18, no. 7, pp. 793–807, Aug. 2025.
[23] I. P. A. W. W. Putra, I. P. G. H. Suputra, and I. B. G. Sarasvananda, “Optimasi Hyperparameter CART Menggunakan Particle Swarm Op-timization (PSO) untuk Klasifikasi Penyakit Stroke,” JNATIA, vol. 4, no. 1, pp. 161–170, Nov. 2025.
[24] A. Widiyanto, M. Prameswari, and M. A. Latief, “Gambling Comments Detection on YouTube: A Comparative Study of Tree-Based Boosting, LSTM and GRU Models,” JUTI: Jurnal Ilmiah Teknologi Informasi, vol. 23, no. 2, Jul. 2025.
[25] F. H. Putra, A. R. Albar, A. S. N. Akbar, A. Kurniawan, L. Okta,
F. Fitriah, and M. Muntahanah, “Analysis of the Effectiveness of Voice Command Recognition Using Machine Learning Algorithms to Support Digital Learning in the Bengkulu Region,” MESTAKA: Jurnal Pengabdian Kepada Masyarakat, vol. 4, no. 4, pp. 452–459, Aug. 2025.
[26] W. Wolberg, O. Mangasarian, N. Street, and W. Street, “Breast Cancer Wisconsin (Diagnostic),” UCI Machine Learning Repository, 1993.
[27] W. N. Street, W. H. Wolberg, and O. L. Mangasarian, “Nuclear Feature Extraction for Breast Tumor Diagnosis,” in Proceedings of SPIE: Electronic Imaging, Jul. 1993.
[28] P. Michalakis, D. Vasilaki, A. J. Abdallah, C. Asikis, A. Niakou,
A. Stratos, A. Tsouknidas, E. Johnstone, and K. Michalakis, “An Insight into Cancer Cells and Disease Progression Through the Lens of Mathematical Modeling,” Current Issues in Molecular Biology, vol. 47, no. 7, p. 477, Jun. 2025.



.png)
.png)



.png)
.png
)


