Comparative Study of TF-IDF-SVM and IndoBERT for Imbalanced Indonesian Hoax Detection


Authors

  • Yono Cahyono Universitas Pamulang, South Tangerang, Indonesia
  • Nardiono Universitas Pamulang, South Tangerang, Indonesia
  • Hendri Ardiansyah Universitas Pamulang, South Tangerang, Indonesia
  • Cintia Septiani Universitas Pamulang, South Tangerang, Indonesia

DOI:

https://doi.org/10.64366/ijids.v3i2.575

Keywords:

Indonesian Hoax Detection; TF-IDF; Support Vector Machine; IndoBERT; Fake News Classification; Class Imbalance

Abstract

Indonesian hoax detection remains challenging because classification performance is influenced by text representation, class imbalance, and experimental design. This study comparatively evaluates a conventional Term Frequency–Inverse Document Frequency with Support Vector Machine (TF-IDF–SVM) approach and the transformer-based IndoBERT model under the same experimental protocol on a naturally imbalanced Indonesian hoax dataset. News titles and narratives were combined as textual input, duplicate records were removed, and the resulting dataset was stratified into training, validation, and held-out test sets. Model configurations were selected exclusively using validation Macro F1-score to prevent test-set information from influencing model selection. The final SVM employed unigram–bigram TF-IDF features with balanced LinearSVC, whereas the selected IndoBERT configuration used unweighted cross-entropy loss with a maximum sequence length of 128 tokens. On the test set, TF-IDF–SVM achieved an accuracy of 0.8299, Macro Precision of 0.7123, Macro Recall of 0.7065, and Macro F1-score of 0.7093, while IndoBERT achieved 0.8047, 0.6665, 0.6573, and 0.6616, respectively. The stronger numerical performance of TF-IDF–SVM suggests that discriminative lexical patterns remained highly informative in this dataset, while the contextual advantages of IndoBERT may not have been fully realized under the substantial class imbalance and the investigated training configuration. However, the exact McNemar test produced a p-value of 0.1371, indicating that the difference in paired classification correctness was not statistically significant at the 0.05 level. These findings demonstrate that TF-IDF–SVM remains a competitive and computationally simpler baseline for Indonesian hoax detection under imbalanced data conditions, while the benefit of more complex contextual models remains dependent on dataset characteristics and experimental settings.

Downloads

Download data is not yet available.

References

L. A. Pekandi, R. G. Widjaja, A. Ananta, J. Harefa, and K. Jingga, “Evaluating IndoBERT for Indonesian Hoax News Detection: A Comparative Study with Ensemble and CNN-LSTM Models,” in Procedia Computer Science, Elsevier B.V., 2025, pp. 1625–1633. doi: 10.1016/j.procs.2025.09.105.

K. S. Sreekar Datta, G. Narasimha Naidu, S. Abhishek, and T. Anjali, “Enhancing Veracity: Empirical Evaluation of Fake News Detection Techniques,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 97–107. doi: 10.1016/j.procs.2024.03.199.

B. Prastyapradipta, K. Naoko, R. A. B. Himawan, M. A. Ibrahim, and G. A. Tarigan, “Analyzing Indonesian political hoax detection system on social media using deep learning and natural language processing,” in Procedia Computer Science, Elsevier B.V., 2025, pp. 1742–1751. doi: 10.1016/j.procs.2025.09.117.

P. W. Cahyo, U. S. Aesyi, W. A. Setianto, and T. Sulaiman, “A Novel Named Entity Recognition approach of Indonesian fake news using part of speech and BERT model on presidential election,” Dec. 01, 2025, Elsevier B.V. doi: 10.1016/j.jjimei.2025.100354.

Q. H. Zhou and T. Cai, “Adaptive gate residual connection and multi-scale RCNN for fake news detection,” Machine Learning with Applications, vol. 19, Mar. 2025, doi: 10.1016/j.mlwa.2024.100612.

B. P. Nayoga, R. Adipradana, R. Suryadi, and D. Suhartono, “Hoax Analyzer for Indonesian News Using Deep Learning Models,” in Procedia Computer Science, Elsevier B.V., 2021, pp. 704–712. doi: 10.1016/j.procs.2021.01.059.

T. Wahyuningsih, D. Manongga, I. Sembiring, and S. Wijono, “Comparison of Effectiveness of Logistic Regression, Naive Bayes, and Random Forest Algorithms in Predicting Student Arguments,” in Procedia Computer Science, Elsevier B.V., 2024, pp. 349–356. doi: 10.1016/j.procs.2024.03.014.

S. T. Hamidou and A. Mehdi, “Enhancing IDS performance through a comparative analysis of Random Forest, XGBoost, and Deep Neural Networks,” Machine Learning with Applications, vol. 22, p. 100738, Dec. 2025, doi: 10.1016/j.mlwa.2025.100738.

R. Pradina Kusumawardani, M. R. Lukman, M. Ibrahim, A. Ilham, C. Waladsae Adiena, and R. Prayoga, “ScienceDirect The Effect of the Back Translation Data Augmentation Technique for Indonesian-Language Hoax News Detection Model,” Procedia Comput. Sci., vol. 284, pp. 1479–1486, 2026, [Online]. Available: www.sciencedirect.com

Amandeep and S. Suresh, “Transforming Fake News Detection: Leveraging DistilBERT Models for Enhanced Accuracy,” in Procedia Computer Science, Elsevier B.V., 2025, pp. 283–290. doi: 10.1016/j.procs.2025.03.203.

O. Embarak, “Deep Learning for Fake News Detection: Analysing Facebook’s Misinformation Networks,” in Procedia Computer Science, Elsevier B.V., 2025, pp. 199–208. doi: 10.1016/j.procs.2025.07.173.

J. Alghamdi, Y. Lin, and S. Luo, “Unveiling the hidden patterns: A novel semantic deep learning approach to fake news detection on social media,” Eng. Appl. Artif. Intell., vol. 137, Nov. 2024, doi: 10.1016/j.engappai.2024.109240.

A. Malik, D. K. Behera, J. Hota, and A. R. Swain, “Ensemble graph neural networks for fake news detection using user engagement and text features,” Results in Engineering, vol. 24, Dec. 2024, doi: 10.1016/j.rineng.2024.103081.

F. J. Lara-Abelenda, D. Chushig-Muzo, C. B. Acosta, A. M. Wägner, C. Granja, and C. Soguero-Ruiz, “Evaluating Time Series Classification Models for Nocturnal Hypoglycemia: From Predictive Performance to Environmental Impact,” IEEE Access, vol. 13, no. September, pp. 150756–150771, 2025, doi: 10.1109/ACCESS.2025.3600917.

A. Chahal et al., “Predictive analytics technique based on hybrid sampling to manage unbalanced data in smart cities,” Heliyon, vol. 10, no. 24, Dec. 2024, doi: 10.1016/j.heliyon.2024.e39275.

M. A. Umar, Z. Chen, K. Shuaib, and Y. Liu, “Effects of feature selection and normalization on network intrusion detection,” Data Science and Management, vol. 8, no. 1, pp. 23–39, Mar. 2025, doi: 10.1016/j.dsm.2024.08.001.

W. Zhou and D. Yang, “Analysis and comparison of automatic image focusing algorithms in digital image processing,” J. Radiat. Res. Appl. Sci., vol. 16, no. 4, p. 100672, Dec. 2023, doi: 10.1016/j.jrras.2023.100672.

R. Ranjan et al., “A novel local optimal oriented pattern for image splicing detection leveraging deep learning and SVM,” Franklin Open, vol. 16, Sep. 2026, doi: 10.1016/j.fraope.2026.100634.

S. Nouas, L. Oukid, and F. Boumahdi, “Enhancing imbalanced text classification: an overlap-based refinement approach,” Data Science and Management, vol. 8, no. 4, pp. 474–484, Dec. 2025, doi: 10.1016/j.dsm.2025.03.001.

B. Widmer, G. Moffa, H. E. Viehweger, and M. Sangeux, “Use of natural language processing to predict the diagnoses of patients with gait impairments,” Inform. Med. Unlocked, vol. 59, Jan. 2025, doi: 10.1016/j.imu.2025.101712.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Comparative Study of TF-IDF-SVM and IndoBERT for Imbalanced Indonesian Hoax Detection

Dimensions Badge

ARTICLE HISTORY

Published: 2026-06-30

Abstract View: 1 times
PDF Download: 0 times

How to Cite

Yono Cahyono, Nardiono, Hendri Ardiansyah, & Cintia Septiani. (2026). Comparative Study of TF-IDF-SVM and IndoBERT for Imbalanced Indonesian Hoax Detection. International Journal of Informatics and Data Science, 3(2), 115-126. https://doi.org/10.64366/ijids.v3i2.575

Issue

Section

Articles