Comparison of Naive Bayes, SVM, and Logistic Regression for Sentiment Analysis of the Makan Bergizi Gratis Program
DOI:
https://doi.org/10.65780/bima.v1i6.34Keywords:
Sentiment Analysis, MBG Program, Naive Bayes, SVM, Logistic RegressionAbstract
The Makan Bergizi Gratis Program (MBG) became one of the widely discussed public issues on platform X and generated diverse responses from users. These responses included supportive, critical, and neutral opinions, making sentiment analysis relevant for understanding public opinion toward the program. This study compares three text classification algorithms, namely Naive Bayes, LinearSVC based Support Vector Machine, and Logistic Regression, to analyze the sentiment of tweets concerning the Makan Bergizi Gratis Program. Data were collected from platform X on 19 June 2026 using tweet crawling and were subsequently filtered and manually annotated. The final dataset consisted of 1,061 Indonesian language tweets classified into negative, neutral, and positive sentiment. Two annotators independently assigned sentiment labels, followed by discussion to resolve disagreements. The inter annotator agreement reached 94.62%, with a Cohen's Kappa coefficient of 0.9097. Text preprocessing consisted of cleaning, normalization, tokenization, stopword removal, and stemming. TF-IDF features were generated within a pipeline to ensure that feature fitting was performed only on the training portion of each evaluation fold. The data were divided into training and testing sets using an 80:20 stratified split. Hyperparameter optimization was performed for each classifier, followed by evaluation using accuracy, macro precision, macro recall, macro F1 score, confusion matrix, and stratified 5 fold cross validation. The test results show that SVM achieved the highest accuracy of 72.77%, precision of 75.11%, recall of 65.17%, and macro F1 score of 68.28%. Logistic Regression obtained a macro F1 score of 67.05%, while Naive Bayes obtained 63.69%. SVM also achieved the highest average macro F1 score in cross validation at 60.25%, compared with 59.94% for Logistic Regression and 56.53% for Naive Bayes. These results indicate that SVM provided the strongest overall performance for the dataset examined in this study.
Downloads
References
[1] M. Rodríguez-Ibánez, A. Casánez-Ventura, F. Castejón-Mateos, and P. M. Cuenca-Jiménez, “A review on sentiment analysis from social media platforms,” Expert Syst. Appl., vol. 223, no. March, 2023, doi: 10.1016/j.eswa.2023.119862.
[2] B. Setiawan, “A Review of Sentiment Analysis Applications in Indonesia Between 2023-2024,” J. Inf. Eng. Educ. Technol., vol. 8, no. 2, pp. 71–83, 2025, doi: 10.26740/jieet.v8n2.p71-83.
[3] A. Gasparetto, M. Marcuzzo, A. Zangari, and A. Albarelli, “Survey on Text Classification Algorithms: From Text to Predictions,” Inf., vol. 13, no. 2, pp. 1–39, 2022, doi: 10.3390/info13020083.
[4] S. N. Khan, S. U. Khan, H. Aznaoui, C. B. Şahin, and Ö. B. Dinler, “Generalization of linear and non-linear support vector machine in multiple fields: a review,” Comput. Sci. Inf. Technol., vol. 4, no. 3, pp. 226–239, 2023, doi: 10.11591/csit.v4i3.p226-239.
[5] M. Dawud, “Klasifikasi sentimen terhadap program makan bergizi gratis menggunakan metode Logistic Regression,” pp. 400–412, 2025, doi: https://doi.org/10.35889/progresif.v22i2.3610.
[6] M.Arif Firmansyah, Asep Saeppani, and Irfan Fadil, “Analisis Sentimen Publik Program Makan Bergizi Gratis Menggunakan Support Vector Machine,” JPNM J. Pustaka Nusant. Multidisiplin, vol. 3, no. 4, pp. 14–20, 2025, doi: 10.59945/jpnm.v3i4.805.
[7] S. Jatmiko and C. Dometian, “Analisis Sentimen Ulasan Google Play Store: Studi Komparatif Algoritma SVM, Naïve Bayes, dan Logistic Regression,” J. FASILKOM, vol. 15, no. 3, pp. 610–620, Jan. 2026, doi: 10.37859/jf.v15i3.10016.
[8] S. Pradeepa et al., “HGATT_LR: transforming review text classification with hypergraphs attention layer and logistic regression,” Sci. Rep., vol. 14, no. 1, pp. 1–14, 2024, doi: 10.1038/s41598-024-70565-6.
[9] R. Obiedat et al., “Sentiment Analysis of Customers’ Reviews Using a Hybrid Evolutionary SVM-Based Approach in an Imbalanced Data Distribution,” IEEE Access, vol. 10, pp. 22260–22273, 2022, doi: 10.1109/ACCESS.2022.3149482.
[10] H. Hairani, T. Widiyaningtyas, and D. D. Prasetya, “Addressing Class Imbalance of Health Data: a Systematic Literature Review on Modified Synthetic Minority Oversampling Technique (SMOTE) Strategies,” Int. J. Informatics Vis., vol. 8, no. 3, pp. 1310–1318, 2024, doi: 10.62527/joiv.8.3.2283.
[11] I. Amal and Jayanta, “Perbandingan Pelabelan Otomatis Dan Manual Untuk Analisis Sentimen Terhadap Kenaikan Harga BBM Pertamina Pada Twitter Menggunakan Algoritma Support Vector Machine,” Semin. Nas. Mhs. Ilmu Komput. dan Apl., vol. 4, no. 2, pp. 473–487, 2023, doi: 10.1109/SENAMIKA57008.2023.10338748.
[12] H. D. Abubakar and M. Umar, “Sentiment Classification: Review of Text Vectorization Methods: Bag of Words, Tf-Idf, Word2vec and Doc2vec,” SLU J. Sci. Technol., vol. 4, no. 1&2, pp. 27–33, 2022, doi: 10.56471/slujst.v4i.266.
[13] I. Hendrawan Rifky, E. Utami, and A. Hartanto Dwi, “Analisis Perbandingan Metode Tf-Idf dan Word2vec pada Klasifikasi Teks Sentimen Masyarakat Terhadap Produk Lokal di Indonesia,” Smart Comp Jurnalnya Orang Pint. Komput., vol. 11, no. 3, pp. 497–503, 2022, doi: 10.30591/smartcomp.v11i3.3902.
[14] S. Syafrizal, M. Afdal, and R. Novita, “Analisis Sentimen Ulasan Aplikasi PLN Mobile Menggunakan Algoritma Naïve Bayes Classifier dan K-Nearest Neighbor,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 4, no. 1, pp. 10–19, Dec. 2023, doi: 10.57152/malcom.v4i1.983.
[15] K. Walji, A. Erraissi, A. Zakrani, and M. Banane, “From Review to Practice : A Comparative Study and Decision-Support Framework for Sentiment Classification Models,” vol. 16, no. 9, pp. 699–709, 2025, doi: 10.14569/IJACSA.2025.0160967.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 BIMA (Bulletin of Intelligent Machines and Algorithms)

This work is licensed under a Creative Commons Attribution 4.0 International License.