Comparative Analysis of TF-IDF, TF-IDF+WordNet, and Sentence-BERT for News Document Retrieval Using Cosine Similarity

Authors

  • Galib Haftha Zuhayir Universitas Muhammadiyah Jember
  • Wiwik Suharso Universitas Muhammadiyah Jember
  • Nanda Kurnia Wardati Universitas Muhammadiyah Jember

DOI:

https://doi.org/10.55537/cosie.v5i3.1835

Keywords:

Information Retrieval, TF-IDF, Sentence-BERT, WordNet, Cosine Similarity

Abstract

The rapid growth of digital news volume has produced information overload, while exact keyword-matching retrieval remains vulnerable to synonymy and polysemy, causing relevant documents to be missed. Prior studies generally compare only two document representation methods on small-scale datasets, leaving a gap in controlled evaluations that jointly compare statistical, lexical-hybrid, and neural approaches on a large-scale news domain. This study compares three document representation methods, namely Term Frequency-Inverse Document Frequency (TF-IDF), TF-IDF with WordNet-based query expansion, and Sentence-BERT (all-MiniLM-L6-v2), for news document retrieval using Cosine Similarity on the BBC News Dataset (14,305 documents with a hierarchical Ground Truth of 5 Topics and 51 Subtopics). Ten queries were evaluated using Precision@K, Recall@K, F1-Score@K (K=5, 10, 20), and execution time. The results show that Sentence-BERT consistently outperforms the other methods with a Precision@5 of 0.84, compared to TF-IDF (0.56) and TF-IDF+WordNet (0.52), while TF-IDF remains the fastest at online query time (23.86 ms per query). WordNet expansion actually reduces precision and increases execution time without a proportional accuracy gain. These findings confirm that transformer-based semantic representations are superior for news domains with high lexical variation, while TF-IDF remains relevant for computationally constrained real-time systems

Downloads

Download data is not yet available.

References

[1] APJII, "Profil Internet Indonesia 2025 & Segmentasi Pasar ISP Tahun 2025," Asosiasi Penyelenggara Jasa Internet Indonesia, 2025.

[2] S. Kemp, "Digital 2025: Indonesia," DataReportal, 2025.

[3] N. Newman, R. Fletcher, C. T. Robertson, A. R. Arguedas, and R. K. Nielsen, "Reuters Institute Digital News Report 2024," Reuters Institute for the Study of Journalism, 2024.

[4] A. Tandon, A. Dhir, A. K. M. N. Islam, and M. Mäntymäki, "Impact of News Overload on Social Media News Curation: Mediating Role of News Avoidance," Frontiers in Psychology, vol. 13, 2022, doi: 10.3389/fpsyg.2022.865246.

[5] D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, 3rd ed. Stanford University, 2024.

[6] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3982-3992, doi: 10.18653/v1/D19-1410.

[7] D. Chandra and A. Verma, "Enhanced Information Retrieval Through Query Expansion: A Comparative Analysis of Techniques Using TF-IDF," IEEE Xplore, 2025, doi: 10.1109/IEEEXplore.2025.11283698.

[8] R. Ambarwati, N. A. S. Sari, and H. A. Rosyid, "Implementasi TF-IDF dan Cosine Similarity untuk Deteksi Relevansi Berita Korupsi," Jurnal Teknologi Informasi dan Ilmu Komputer, 2025.

[9] S. Widianto, A. R. Pratama, and F. Nugraha, "Document Similarity Measurement Using TF-IDF and Cosine Similarity," KOMPUTA: Jurnal Ilmiah Komputer dan Informatika, vol. 13, no. 1, pp. 25-32, 2024, doi: 10.34010/komputa.v13i1.12401.

[10] M. F. Rasyid and S. W. Ningsih, "Sistem Pencarian Destinasi Wisata Menggunakan TF-IDF dan Cosine Similarity," Jurnal Teknologi dan Sistem Informasi, 2024.

[11] N. Choudhary, P. K. Singh, and D. K. Tayal, "BERT vs TF-IDF: A Comparative Study for Document Retrieval," in Proc. Int. Conf. Innovative Computing & Communications (ICICC), 2020.

[12] N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych, "BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models," in Proc. Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2021.

[13] F. Pedregosa, G. Varoquaux, A. Gramfort, et al., "Scikit-learn: Machine Learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011.

[14] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. 2019 Conf. North American Chapter Assoc. Computational Linguistics: Human Language Technologies (NAACL-HLT), 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.

[15] A. Vaswani, N. Shazeer, N. Parmar, et al., "Attention Is All You Need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 5998-6008.

[16] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, "MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers," in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020.

[17] G. A. Miller and C. Fellbaum, "WordNet: A Large English-Language Database of Synonyms, Antonyms, and Semantic Relationships," Princeton University, 2024.

[18] X. Wang, X. Liu, C. Sun, J. Wang, and Z. Wang, "E5: Text Embeddings by Weakly-Supervised Contrastive Learning," in Proc. ICLR 2023, 2023.

[19] J. Ni, C. Qu, J. Lu, Z. Dai, R. Nogueira, and J. Lin, "BGE: Large Language Models are Strong Embedders," arXiv:2302.07574, 2023.

[20] K. Luan, Y. Chen, C. Yang, and J. Liu, "GTE: General Text Embeddings," in Proc. ICML 2023, 2023.

[21] Y. Lin, P. Liu, J. Li, and J. Zhou, "Jina Embeddings v2: 8192-Token General-Purpose Text Embeddings," arXiv:2310.19923, 2023.

[22] Y. Liu, Z. Liu, and Y. Sun, "Query Expansion with Large Language Models for Information Retrieval," in Proc. NAACL 2024, 2024, pp. 8899-8912.

[23] A. S. Nugraha, R. A. Pratama, and D. Wijaya, "Indonesian News Retrieval with Hybrid TF-IDF and Sentence-BERT," Jurnal Teknologi Informasi, vol. 11, no. 2, pp. 201-215, 2024.

[24] M. Jeronymo, C. Bonifacio, and R. Lotufo, "BEIR-PL: Benchmarking IR for Portuguese," in Proc. PROPOR 2024, 2024, pp. 45-57.

[25] S. K. S. Tyagi, P. Sharma, and A. K. Singh, "ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction," in Proc. SIGIR 2024, 2024, pp. 1890-1895.

[26] K. Santhanam, O. Khattab, C. Potts, and M. Zaharia, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction," in Proc. SIGIR 2022, 2022. (Updated 2024: ColBERTv2)

[27] L. Gao, Y. Wu, and J. Callan, "COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List," in Proc. NAACL 2021, 2021. (Updated 2024: COIL+)

[28] O. Khattab and M. Zaharia, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction," in Proc. SIGIR 2020, 2020.

Downloads

Published

31-07-2026

How to Cite

Zuhayir, G. H., Suharso, W., & Wardati, N. K. (2026). Comparative Analysis of TF-IDF, TF-IDF+WordNet, and Sentence-BERT for News Document Retrieval Using Cosine Similarity. Journal of Computer Science and Informatics Engineering , 5(3), 448–458. https://doi.org/10.55537/cosie.v5i3.1835

Issue

Section

Articles