Comparative Analysis of TF-IDF, TF-IDF+WordNet, and Sentence-BERT for News Document Retrieval Using Cosine Similarity
DOI:
https://doi.org/10.55537/cosie.v5i3.1835Keywords:
Information Retrieval, TF-IDF, Sentence-BERT, WordNet, Cosine SimilarityAbstract
The rapid growth of digital news volume has produced information overload, while exact keyword-matching retrieval remains vulnerable to synonymy and polysemy, causing relevant documents to be missed. Prior studies generally compare only two document representation methods on small-scale datasets, leaving a gap in controlled evaluations that jointly compare statistical, lexical-hybrid, and neural approaches on a large-scale news domain. This study compares three document representation methods, namely Term Frequency-Inverse Document Frequency (TF-IDF), TF-IDF with WordNet-based query expansion, and Sentence-BERT (all-MiniLM-L6-v2), for news document retrieval using Cosine Similarity on the BBC News Dataset (14,305 documents with a hierarchical Ground Truth of 5 Topics and 51 Subtopics). Ten queries were evaluated using Precision@K, Recall@K, F1-Score@K (K=5, 10, 20), and execution time. The results show that Sentence-BERT consistently outperforms the other methods with a Precision@5 of 0.84, compared to TF-IDF (0.56) and TF-IDF+WordNet (0.52), while TF-IDF remains the fastest at online query time (23.86 ms per query). WordNet expansion actually reduces precision and increases execution time without a proportional accuracy gain. These findings confirm that transformer-based semantic representations are superior for news domains with high lexical variation, while TF-IDF remains relevant for computationally constrained real-time systems
Downloads
References
[1] APJII, "Profil Internet Indonesia 2025 & Segmentasi Pasar ISP Tahun 2025," Asosiasi Penyelenggara Jasa Internet Indonesia, 2025.
[2] S. Kemp, "Digital 2025: Indonesia," DataReportal, 2025.
[3] N. Newman, R. Fletcher, C. T. Robertson, A. R. Arguedas, and R. K. Nielsen, "Reuters Institute Digital News Report 2024," Reuters Institute for the Study of Journalism, 2024.
[4] A. Tandon, A. Dhir, A. K. M. N. Islam, and M. Mäntymäki, "Impact of News Overload on Social Media News Curation: Mediating Role of News Avoidance," Frontiers in Psychology, vol. 13, 2022, doi: 10.3389/fpsyg.2022.865246.
[5] D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, 3rd ed. Stanford University, 2024.
[6] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3982-3992, doi: 10.18653/v1/D19-1410.
[7] D. Chandra and A. Verma, "Enhanced Information Retrieval Through Query Expansion: A Comparative Analysis of Techniques Using TF-IDF," IEEE Xplore, 2025, doi: 10.1109/IEEEXplore.2025.11283698.
[8] R. Ambarwati, N. A. S. Sari, and H. A. Rosyid, "Implementasi TF-IDF dan Cosine Similarity untuk Deteksi Relevansi Berita Korupsi," Jurnal Teknologi Informasi dan Ilmu Komputer, 2025.
[9] S. Widianto, A. R. Pratama, and F. Nugraha, "Document Similarity Measurement Using TF-IDF and Cosine Similarity," KOMPUTA: Jurnal Ilmiah Komputer dan Informatika, vol. 13, no. 1, pp. 25-32, 2024, doi: 10.34010/komputa.v13i1.12401.
[10] M. F. Rasyid and S. W. Ningsih, "Sistem Pencarian Destinasi Wisata Menggunakan TF-IDF dan Cosine Similarity," Jurnal Teknologi dan Sistem Informasi, 2024.
[11] N. Choudhary, P. K. Singh, and D. K. Tayal, "BERT vs TF-IDF: A Comparative Study for Document Retrieval," in Proc. Int. Conf. Innovative Computing & Communications (ICICC), 2020.
[12] N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych, "BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models," in Proc. Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2021.
[13] F. Pedregosa, G. Varoquaux, A. Gramfort, et al., "Scikit-learn: Machine Learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011.
[14] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. 2019 Conf. North American Chapter Assoc. Computational Linguistics: Human Language Technologies (NAACL-HLT), 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.
[15] A. Vaswani, N. Shazeer, N. Parmar, et al., "Attention Is All You Need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 5998-6008.
[16] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, "MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers," in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020.
[17] G. A. Miller and C. Fellbaum, "WordNet: A Large English-Language Database of Synonyms, Antonyms, and Semantic Relationships," Princeton University, 2024.
[18] X. Wang, X. Liu, C. Sun, J. Wang, and Z. Wang, "E5: Text Embeddings by Weakly-Supervised Contrastive Learning," in Proc. ICLR 2023, 2023.
[19] J. Ni, C. Qu, J. Lu, Z. Dai, R. Nogueira, and J. Lin, "BGE: Large Language Models are Strong Embedders," arXiv:2302.07574, 2023.
[20] K. Luan, Y. Chen, C. Yang, and J. Liu, "GTE: General Text Embeddings," in Proc. ICML 2023, 2023.
[21] Y. Lin, P. Liu, J. Li, and J. Zhou, "Jina Embeddings v2: 8192-Token General-Purpose Text Embeddings," arXiv:2310.19923, 2023.
[22] Y. Liu, Z. Liu, and Y. Sun, "Query Expansion with Large Language Models for Information Retrieval," in Proc. NAACL 2024, 2024, pp. 8899-8912.
[23] A. S. Nugraha, R. A. Pratama, and D. Wijaya, "Indonesian News Retrieval with Hybrid TF-IDF and Sentence-BERT," Jurnal Teknologi Informasi, vol. 11, no. 2, pp. 201-215, 2024.
[24] M. Jeronymo, C. Bonifacio, and R. Lotufo, "BEIR-PL: Benchmarking IR for Portuguese," in Proc. PROPOR 2024, 2024, pp. 45-57.
[25] S. K. S. Tyagi, P. Sharma, and A. K. Singh, "ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction," in Proc. SIGIR 2024, 2024, pp. 1890-1895.
[26] K. Santhanam, O. Khattab, C. Potts, and M. Zaharia, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction," in Proc. SIGIR 2022, 2022. (Updated 2024: ColBERTv2)
[27] L. Gao, Y. Wu, and J. Callan, "COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List," in Proc. NAACL 2021, 2021. (Updated 2024: COIL+)
[28] O. Khattab and M. Zaharia, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction," in Proc. SIGIR 2020, 2020.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Galib Haftha Zuhayir, Wiwik Suharso, Nanda Kurnia Wardati

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


