Obesity Level Classification Using the Gaussian Naive Bayes and 10-Fold Cross-Validation

Authors

  • Rivaldi Prima Nanda Universitas Islam Negeri Sumatera Utara
  • Muhammad Siddik Hasibuan Universitas Islam Negeri Sumatera Utara
  • Mhd. Ikhsan Rifki Universitas Islam Negeri Sumatera Utara

DOI:

https://doi.org/10.55537/cosie.v5i4.1855

Keywords:

Obesity Classification, Gaussian Naïve Bayes, Machine Learning, K-Fold Cross Validation, Data Mining

Abstract

Obesity classification is an important task in health-related data analysis for identifying obesity levels based on individual characteristics. This study aims to implement the Gaussian Naïve Bayes algorithm for obesity level classification and evaluate differences in performance estimates obtained using a 70:30 train-test split and 10-Fold Cross Validation. The dataset consisted of 1,000 observational records containing five predictor variables: age, height, weight, body mass index (BMI), and physical activity level. Obesity levels were categorized into four classes: Underweight, Normal, Overweight, and Obese. The model was implemented using Python and evaluated through two validation scenarios. Experimental results showed that the Gaussian Naïve Bayes model achieved an accuracy of 95.33% using the train-test split approach. Evaluation using 10-Fold Cross Validation produced an average accuracy of 96.30%, precision of 96.66%, recall of 96.45%, and F1-score of 96.48%. These findings indicate that different validation strategies can produce different performance estimates. On the dataset used in this study, Gaussian Naïve Bayes demonstrated high classification performance, while 10-Fold Cross Validation provided performance estimates across multiple data partitions, offering a broader assessment of model performance. This study provides empirical evidence regarding the use of different validation strategies for obesity level classification

Downloads

Download data is not yet available.

References

[1] W. Kurdanti et al., “Jurnal Gizi Klinik Indonesia Faktor-Faktor yang Mempengaruhi Kejadian Obesitas pada Remaja,” Jurnal Gizi Klinik Indonesia, vol. 11, no. 04, pp. 179–190, 2025.

[2] Juliatin and Hasniar, “Faktor-Faktor Penyebab Terjadinya Obesitas pada Remaja,” Student Research Journal (SRJ), vol. 2, no. 6, pp. 137–149, 2024, doi: 10.55606/srj-yappi.v2i6.1633.

[3] Nurzakiah, E. L. Achadi, and R. A. D. Sartika, “Faktor Risiko Obesitas pada Orang Dewasa Urban dan Rural,” Kesmas, vol. 5, no. 1, pp. 29–34, 2023, doi: 10.21109/kesmas.v5i1.159.

[4] S. Aziz, Y. Pramana, and Sukarni, “Hubungan Aktivitas Fisik Dengan Kejadian Obesitas Pada Remaja,” Mahesa: Malahayati Health Student Journal, vol. 3, pp. 1115–1124, 2023.

[5] J. M. Gorriz, R. Martin-Clemente, F. Segovia, J. Ramirez, A. Ortiz, and J. Suckling, “Is K-fold cross validation the best model selection method for machine learning?,” Information Fusion, vol. 135, p. 104404, Nov. 2026, doi: 10.1016/j.inffus.2026.104404.

[6] M. N. Fahmi, “Implementasi Machine Learning menggunakan Python Library: Scikit-Learn,” Sains Data, vol. 2, pp. 87–96, 2023.

[7] R. S. Nurhalizah, R. Ardianto, and Purwono, “Analisis Supervised dan Unsupervised Learning pada Machine Learning: Systematic Literature Review,” JIKI, vol. 4, no. 1, pp. 61–72, 2024.

[8] M. S. Hasibuan, Y. R. Nasution, I. Komputer, U. Islam, and N. Sumatera, “OPTIMASI MODEL SEMI-SUPERVISED LEARNING DENGAN SVM,” vol. 9, no. 2, pp. 231–239, 2024.

[9] M. S. Hasibuan and Suhardi, “Higher Education and Health Breakthroughs: Residual Network for Pneumonia Detection from Chest X-ray Images,” Int. J. Adv. Sci. Eng. Inf. Technol., vol. 15, no. 5, 2025, doi: 10.18517/ijaseit.15.5.20483.

[10] J. Oriana, W. Sinaga, M. Fathurahman, S. Wahyuningsih, and M. N. Hayati, “Evaluating Different K Values in K-Fold Cross Validation for Binary Logistic Regression to Classify Poverty,” Jurnal Gaussian, vol. 8, no. 2, pp. 189–198, 2025.

[11] M. A. Faradeya and E. R. Subhiyakto, “Klasifikasi Penyakit Gagal Jantung Menggunakan Algoritma Naive Bayes,” Jurnal Algoritma, pp. 115–127, 2025, doi: 10.33364/algoritma/v.22-1.2178.

[12] I. Akil and I. Chaidir, “Classification of Heart Disease Diagnoses Using Gaussian Naive Bayes,” Komputasi, vol. 21, pp. 31–36, 2024.

[13] Nurainun, E. Haerani, F. Syafria, and L. Oktavia, “Penerapan Algoritma Naive Bayes Classifier dalam Klasifikasi Status Gizi Balita dengan Pengujian K-Fold Cross Validation,” JOSYC, vol. 4, no. 3, 2023, doi: 10.47065/josyc.v4i3.3414.

[14] H. Hafid, “Penerapan K-Fold Cross Validation untuk Menganalisis Kinerja Algoritma K-Nearest Neighbor pada Data Kasus Covid-19 di Indonesia,” Journal of Mathematics, Computations, and Statistics, vol. 6, no. 2, pp. 161–168, 2023.

[15] M. Andani, J. Triloka, S. Y. Irianto, and H. W. Nugroho, “Performance Comparison of K-Nearest Neighbor, Naive Bayes, and Random Forest Algorithms in Obesity Prediction,” Sinkron, vol. 9, no. 1, pp. 502–510, 2025.

[16] S. Khoerunnisa, D. F. Shiddiq, and D. Nurhayati, “Penerapan Algoritma Naive Bayes dengan Teknik TF-IDF dan Cross Validation untuk Analisis Sentimen Terhadap Starlink,” MALCOM, vol. 5, pp. 566–577, 2025.

[17] M. Jamali et al., “Applications of traditional machine learning and deep learning algorithms in obesity prediction or classification: a systematic review of comparative performance,” NPJ Digit. Med., 2026, doi: 10.1038/s41746-026-02986-8.

[18] H. T. Santoso, F. A. Felmidi, A. Nur, A. Ristyawan, and E. Daniati, “Analisis Kinerja Algoritma Data Mining pada Klasifikasi Tingkat Obesitas dengan K-Fold Cross Validation dan AUC,” Inotek, vol. 8, pp. 113–122, 2024.

[19] T. H. Pinem and Z. P. Putra, “Evaluasi Kinerja Algoritma Klasifikasi Deep Learning dalam Prediksi Diabetes,” vol. 17, no. 1, pp. 17–28, 2025, doi: 10.22441/fifo.2025.v17i1.003.

[20] M. M. A. Zaid and A. A. Mohammed, “Diabetes Prediction Using Machine Learning: Methods, Challenges, and Insights from a Systematic Literature Review,” JKMC, vol. 12, no. 2, pp. 21–27, 2025.

[21] Wijiyanto, A. I. Pradana, Sopingi, and V. Atina, “Teknik K-Fold Cross Validation untuk Mengevaluasi Kinerja Mahasiswa,” pp. 239–248, 2024, doi: 10.33364/algoritma/v.21-1.1618.

[22] J. O. W. Sinaga, M. Fathurahman, S. Wahyuningsih, and M. N. Hayati, “Evaluating Different K Values in K-Fold Cross Validation for Binary Logistic Regression to Classify Poverty,” Jurnal Varian, vol. 8, no. 2, pp. 189–198, 2025, doi: 10.30812/varian.v8i2.4403.

[23] R. S. Pall, S. Yadav, S. Bhalerao, S. Sahu, R. Ahluwalia, and B. Awadhiya, “Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis,” in 2026 International Conference on Intelligent Processing, Hardware, Electronics, and Radio Systems (CIPHER), IEEE, Feb. 2026, pp. 1–6. doi: 10.1109/cipher70417.2026.11523789.

Downloads

Published

13-09-2026

How to Cite

Nanda, R. P., Hasibuan, M. S., & Rifki, M. I. (2026). Obesity Level Classification Using the Gaussian Naive Bayes and 10-Fold Cross-Validation. Journal of Computer Science and Informatics Engineering , 5(4), 540–550. https://doi.org/10.55537/cosie.v5i4.1855

Issue

Section

Articles