Comparative Evaluation of MobileNetV2 and Vision Transformer for Pneumonia Classification on Chest X-Ray Images: Fine-Tuning, Focal Loss, and Decision-Threshold Analysis

Authors

  • Hendra Marcos Universitas Amikom Purwokerto https://orcid.org/0000-0002-2284-2503
  • Ali Novian Universitas Amikom Purwokerto
  • Ma`dan Ma`dan Shomsomi Universitas Amikom Purwokerto
  • Widhaksa Triawan Universitas Amikom Purwokerto

DOI:

https://doi.org/10.70656/ijcse.v3i01.857

Keywords:

Chest X-Ray, Pneumonia, MobileNetV2, Vision Transformer, Fine-Tuning, Focal Loss, Decision Threshold

Abstract

Pneumonia remains a major respiratory infection that imposes a substantial health burden and requires rapid and accurate detection. This study evaluates five deep-learning configurations for binary classification of chest X-ray (CXR) images into Normal and Pneumonia classes: MobileNetV2 with a frozen backbone, fine-tuned MobileNetV2, fine-tuned MobileNetV2 with Focal Loss, Vision Transformer (ViT) without full fine-tuning, and fine-tuned ViT. The public dataset used in the experiments contains 5,856 images, comprising 5,216 training images, 16 validation images, and 624 test images. All metrics in this revised version were recalculated consistently from the experimental confusion matrices. ViT without full fine-tuning achieved the best overall performance at a threshold of 0.5, with 92.63% accuracy, 98.72% sensitivity, 82.48% specificity, a 94.36% F1-score, 90.60% balanced accuracy, and a Matthews correlation coefficient (MCC) of 0.845. MobileNetV2 with Focal Loss achieved 92.15% accuracy, 98.97% sensitivity, 80.77% specificity, and a 94.03% F1-score, providing slightly higher sensitivity with only four false-negative cases. In contrast, fine-tuned ViT achieved 100% sensitivity but only 57.69% specificity at the 0.5 threshold, indicating a shift in the predicted probability distribution and imbalanced generalization. Post hoc threshold analysis showed that changing the decision threshold can improve the error trade-off; however, it was not used to select the primary model because the threshold sweep was evaluated on the test set. These findings demonstrate that fine-tuning strategy and loss function affect model error characteristics differently, while screening-oriented evaluation should consider sensitivity and specificity together with accuracy

Downloads

Download data is not yet available.

References

[1] GBD 2021 Lower Respiratory Infections and Antimicrobial Resistance Collaborators, "Global, regional, and national incidence and mortality burden of non-COVID-19 lower respiratory infections and aetiologies, 1990-2021: a systematic analysis from the Global Burden of Disease Study 2021," Lancet Infect. Dis., vol. 24, no. 9, pp. 974-1002, 2024, doi: 10.1016/S1473-3099(24)00176-2.

[2] D. S. Kermany, M. Goldbaum, W. Cai, et al., "Identifying medical diagnoses and treatable diseases by image-based deep learning," Cell, vol. 172, no. 5, pp. 1122-1131.e9, 2018, doi: 10.1016/j.cell.2018.02.010.

[3] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "MobileNetV2: Inverted residuals and linear bottlenecks," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 4510-4520.

[4] A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., "An image is worth 16x16 words: Transformers for image recognition at scale," in Proc. Int. Conf. Learn. Represent. (ICLR), 2021.

[5] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, "Focal loss for dense object detection," in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017, pp. 2980-2988.

[6] D. Avola, A. Bacciu, L. Cinque, A. Fagioli, M. R. Marini, and R. Taiello, "Study on transfer learning capabilities for pneumonia classification in chest-x-rays images," Comput. Methods Programs Biomed., vol. 221, Art. no. 106833, 2022, doi: 10.1016/j.cmpb.2022.106833.

[7] G. I. Okolo, S. Katsigiannis, and N. Ramzan, "IEViT: An enhanced vision transformer architecture for chest X-ray image classification," Comput. Methods Programs Biomed., vol. 226, Art. no. 107141, 2022, doi: 10.1016/j.cmpb.2022.107141.

[8] S. Singh, M. Kumar, A. Kumar, B. K. Verma, K. Abhishek, and S. Selvarajan, "Efficient pneumonia detection using Vision Transformers on chest X-rays," Sci. Rep., vol. 14, Art. no. 2487, 2024, doi: 10.1038/s41598-024-52703-2.

[9] A. S. Tejani, M. E. Klontzas, A. A. Gatti, et al., "Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 update," Radiol. Artif. Intell., vol. 6, no. 4, 2024, doi: 10.1148/ryai.240300.

[10] P. Szepesi and L. Szilagyi, "Detection of pneumonia using convolutional neural networks and deep learning," Biocybern. Biomed. Eng., vol. 42, no. 3, pp. 1012-1022, 2022, doi: 10.1016/j.bbe.2022.08.001.

[11] J. Huang, "DSSViT: Multi-scale adaptive fusion Vision Transformer with dense feature reuse for robust pneumonia detection in chest radiography," Int. J. Imaging Syst. Technol., vol. 35, no. 3, Art. no. e70127, 2025, doi: 10.1002/ima.70127.

[12] A. Dash, S. Panigrahi, D. S. K. Nayak, A. P. Dash, and T. Swarnkar, "Vi-GeN 1.0: GAN-Augmented Vision Transformer pipeline for diversified pneumonia classification," J. Transform. Technol. Sustain. Dev., vol. 9, Art. no. 14, 2025, doi: 10.1007/s41314-025-00081-6.

[13] F. Garcea, A. Serra, F. Lamberti, and L. Morra, "Data augmentation for medical imaging: A systematic literature review," Comput. Biol. Med., vol. 152, Art. no. 106391, 2023.

[14] H. Aljuaid, H. H. A. Adlan, B. Alkebsi, B. S. Alfurhood, A. Liotta, and L. Cavallaro, "An experimental comparison of deep learning models for pneumonia classification from chest X-ray images," Biomed. Signal Process. Control, vol. 112, Art. no. 108742, 2026, doi: 10.1016/j.bspc.2025.108742.

Downloads

Published

2026-09-03

How to Cite

Marcos, H., Novian, A., Ma`dan Shomsomi, M., & Triawan, W. (2026). Comparative Evaluation of MobileNetV2 and Vision Transformer for Pneumonia Classification on Chest X-Ray Images: Fine-Tuning, Focal Loss, and Decision-Threshold Analysis. Indonesian Journal of Computer Science and Engineering, 3(01), 32–37. https://doi.org/10.70656/ijcse.v3i01.857