Patient-Disjoint Evaluation for Trustworthy and Explainable Breast Cancer Histopathology Classification


Authors

  • Ronal Watrianthos Politeknik Negeri Padang, Padang, Indonesia
  • Arif Rizki Marsa Politeknik Negeri Padang, Padang, Indonesia
  • Dian Eka Putra Politeknik Negeri Padang, Padang, Indonesia
  • Rozi Meri Politeknik Negeri Padang, Padang, Indonesia
  • Novi Efendi Politeknik Negeri Padang, Padang, Indonesia
  • Sofia Yosse Politeknik Negeri Padang, Padang, Indonesia

DOI:

https://doi.org/10.64366/ijids.v3i2.613

Keywords:

BreaKHis Dataset; Data Leakage; EfficientNet; Explainable Artificial Intelligence; Grad-CAM

Abstract

Breast cancer diagnosis from histopathological images increasingly relies on deep learning, with recent studies reporting classification accuracies above 95% on the benchmark BreaKHis dataset. However, much of this literature uses image-level splits that let the same patient's images appear in both training and test sets a source of inflation that is not merely statistical but potentially clinically misleading if taken as evidence of real-world reliability and rarely pairs high accuracy with a rigorous account of model interpretability. This study addresses that gap with an explainable deep learning pipeline for binary (benign/malignant) BreaKHis classification under a patient-disjoint split. An EfficientNetB0 backbone, pretrained on ImageNet, was fine-tuned in two phases with class-weighted loss and Reinhard-based stain normalization. Evaluated under this leakage-safe protocol, the baseline model achieved a test-set balanced accuracy of 0.765 markedly lower than the above-95% figures reported elsewhere, yet a more trustworthy generalization estimate together with an F1-score of 0.777 for the malignant class, a ROC-AUC of 0.849, and a PR-AUC of 0.940, with the best checkpoint selected before overfitting deepened during fine-tuning. A regularization ablation against a more heavily regularized variant showed a narrower generalization gap but lower test performance, confirming the unregularized configuration as the more suitable final model. Grad-CAM was applied to the model's final convolutional layer to visualize the regions driving individual predictions, supporting qualitative inspection of correct and misclassified cases. These findings argue for combining leakage-aware evaluation with explainability as a joint, rather than separate, requirement for trustworthy histopathological classification models.

Downloads

Download data is not yet available.

References

R. Watrianthos, R. Rayendra, E. Asri, Y. Yuhefizar, and H. Humaira, “Threshold-Optimized and Calibrated Logistic Regression for Breast Cancer Classification,” Salud, Ciencia y Tecnología, vol. 5, p. 2241, Oct. 2025, doi: 10.56294/saludcyt20252241.

Agus Perdana Windarto, Anjar Wanto, S Solikhun, and Ronal Watrianthos, “A Comprehensive Bibliometric Analysis of Deep Learning Techniques for Breast Cancer Segmentation: Trends and Topic Exploration (2019-2023),” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 5, pp. 1155–1164, Oct. 2023, doi: 10.29207/resti.v7i5.5274.

R. L. Siegel, A. N. Giaquinto, and A. Jemal, “Cancer statistics, 2024. Cancer Facts & Figures 2024.,” CA Cancer J. Clin., 2024.

B. Jiang, L. Bao, S. He, X. Chen, Z. Jin, and Y. Ye, “Deep learning applications in breast cancer histopathological imaging: diagnosis, treatment, and prognosis,” Breast Cancer Research, vol. 26, no. 1, p. 137, Sep. 2024, doi: 10.1186/s13058-024-01895-6.

E. M. Senan, F. W. Alsaade, M. I. A. Al-Mashhadani, T. H. H. Aldhyani, and M. H. Al-Adhaileh, “Classification of histopathological images for early detection of breast cancer using deep learning,” Journal of Applied Science and Engineering (Taiwan), vol. 24, no. 3, 2021, doi: 10.6180/jase.202106_24(3).0007.

L. Aldakhil, H. Alhasson, and S. Alharbi, “Attention-Based Deep Learning Approach for Breast Cancer Histopathological Image Multi-Classification,” Diagnostics, vol. 14, no. 13, p. 1402, Jul. 2024, doi: 10.3390/diagnostics14131402.

T. A. Toma et al., “Breast Cancer Detection Based on Simplified Deep Learning Technique With Histopathological Image Using BreaKHis Database,” Radio Sci., vol. 58, no. 11, Nov. 2023, doi: 10.1029/2023RS007761.

J. G. Elmore et al., “Diagnostic concordance among pathologists interpreting breast biopsy specimens,” JAMA - Journal of the American Medical Association, vol. 313, no. 11, 2015, doi: 10.1001/jama.2015.1405.

F. A. Spanhol, L. S. Oliveira, C. Petitjean, and L. Heutte, “A Dataset for Breast Cancer Histopathological Image Classification,” IEEE Trans. Biomed. Eng., vol. 63, no. 7, pp. 1455–1462, Jul. 2016, doi: 10.1109/TBME.2015.2496264.

T. T. Brunyé et al., “Machine learning classification of diagnostic accuracy in pathologists interpreting breast biopsies,” Journal of the American Medical Informatics Association, vol. 31, no. 3, 2024, doi: 10.1093/jamia/ocad232.

R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,” Int. J. Comput. Vis., vol. 128, no. 2, pp. 336–359, Feb. 2020, doi: 10.1007/s11263-019-01228-7.

A. Alqutayfi et al., “Explainable Disease Classification: Exploring Grad-CAM Analysis of CNNs and ViTs,” Journal of Advances in Information Technology, vol. 16, no. 2, 2025, doi: 10.12720/jait.16.2.264-273.

M. Azmoodeh-Kalati, H. Shabani, M. S. Maghareh, Z. Barzegar, and R. Lashgari, “Leveraging an ensemble of EfficientNetV1 and EfficientNetV2 models for classification and interpretation of breast cancer histopathology images,” Sci. Rep., vol. 15, no. 1, 2025, doi: 10.1038/s41598-025-06853-6.

N. Bussola, A. Marcolini, V. Maggio, G. Jurman, and C. Furlanello, “AI Slipping on Tiles: Data Leakage in Digital Pathology,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2021. doi: 10.1007/978-3-030-68763-2_13.

S. Liu, G. M. S. Himel, and J. Wang, “Breast Cancer Classification With Enhanced Interpretability: DALAResNet50 and DT Grad-CAM,” IEEE Access, vol. 12, pp. 196647–196659, 2024, doi: 10.1109/ACCESS.2024.3520608.

H. Guo, J. Zhang, Y. Li, X. Pan, and C. Sun, “Advanced pathological subtype classification of thyroid cancer using efficientNetB0,” Diagnostic Pathology , vol. 20, no. 1, 2025, doi: 10.1186/s13000-025-01621-6.

H. T. Vo, N. N. Thien, K. C. Mui, and P. P. Tien, “Enhancing Confidence in Brain Tumor Classification Models With Grad-CAM and Grad-CAM++,” Indonesian Journal of Electrical Engineering and Informatics, vol. 12, no. 4, 2024, doi: 10.52549/ijeei.v12i4.5977.

M. Harahap et al., “Skin cancer classification using EfficientNet architecture,” Bulletin of Electrical Engineering and Informatics, vol. 13, no. 4, 2024, doi: 10.11591/eei.v13i4.7159.

M. I. Rosadi, L. Hakim, and M. Faishol A., “Classification of Coffee Leaf Diseases using the Convolutional Neural Network (CNN) EfficientNet Model,” Conference Series, vol. 4, no. 1, 2023, doi: 10.34306/conferenceseries.v4i1.627.

I. N. Alam, I. H. Kartowisastro, and P. Wicaksono, “Transfer Learning Technique with EfficientNet for Facial Expression Recognition System,” Revue d’Intelligence Artificielle, vol. 36, no. 4, 2022, doi: 10.18280/ria.360405.

H. Ali, N. Shifa, R. Benlamri, A. A. Farooque, and R. Yaqub, “A fine tuned EfficientNet-B0 convolutional neural network for accurate and efficient classification of apple leaf diseases,” Sci. Rep., vol. 15, no. 1, 2025, doi: 10.1038/s41598-025-04479-2.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Patient-Disjoint Evaluation for Trustworthy and Explainable Breast Cancer Histopathology Classification

Dimensions Badge

ARTICLE HISTORY

Published: 2026-06-30

Abstract View: 1 times
PDF Download: 2 times

How to Cite

Watrianthos, R., Arif Rizki Marsa, Dian Eka Putra, Rozi Meri, Novi Efendi, & Sofia Yosse. (2026). Patient-Disjoint Evaluation for Trustworthy and Explainable Breast Cancer Histopathology Classification. International Journal of Informatics and Data Science, 3(2), 81-88. https://doi.org/10.64366/ijids.v3i2.613

Issue

Section

Articles

Most read articles by the same author(s)