Comparative Analysis of Self-Training XGBoost and Label Propagation for Blue Tourism Sentimen Analysis in Lampung Province

  • Selvia Ayu Syaputri Universitas Teknokrat Indonesia
Keywords: Sentiment Analysis, Semi-Supervised Learning, Self-Training, XGBoost, Label Propagation, TF-IDF, Blue Tourism.

Abstract

Blue tourism in Lampung Province is one of the tourism sectors with significant potential to support regional economic development. Tourists' perceptions of tourism destinations can be analyzed through reviews published on various digital platforms, such as Twitter/X, Instagram, and Google Maps Reviews. However, the limited availability of labeled data presents a major challenge in developing sentiment analysis models, making it necessary to employ approaches that can utilize both labeled and unlabeled data simultaneously. This research aims to compare the performance of the XGBoost-based Self-Training method and Label Propagation within a Semi-Supervised Learning approach for sentiment classification of blue tourism reviews in Lampung Province. The dataset consists of 3,947 reviews, including 500 labeled data and 3,447 unlabeled data. The research process includes text preprocessing, feature extraction using TF-IDF, partitioning labeled data into training and testing subsets with an 80:20 ratio, model training, and performance evaluation based on accuracy, precision, recall, and F1-score metrics. The trial outcomes show that XGBoost-based Self-Training reached 83% accuracy, 83.14% precision, 83% recall, and 82.98% F1-score. Conversely, Label Propagation yielded 55% accuracy, 71.31% precision, 55% recall, alongside a 50.11% F1-score. These findings demonstrate that the XGBoost-based Self-Training method outperforms Label Propagation, making it a more effective approach for sentiment classification of blue tourism reviews within a Semi-Supervised Learning framework.

References

DAFTAR PUSTAKA
[1] H. Fadilla, “Pengembangan Sektor Pariwisata untuk meningkatkan Pendapata Daerah Di Indonesia,” J. Business, Econ. Financ., vol. 2, 2024, doi: 10.31080/BENEFIT.
[2] S. Rahmadila and D. Alita, “Perbandingan Metode Naive Bayes Classifier dan Support Vector Machine Pada Analisis Sentimen Wisata Biru Berdasarkan Ulasan Twitter, Instagram, dan Google Maps Review,” Technol. Sci., vol. 7, no. 3, 2025, doi: 10.47065/bits.v7i3.8791.
[3] X. Guo and J. A. Pesonen, “The Role of Online Travel Reviews in Evolving Tourists’ Perceived Destination Image,” Scand. J. Hosp. Tour., vol. 22, no. 4–5, pp. 372–392, 2022, doi: 10.1080/15022250.2022.2112414.
[4] B. Liu, J. Zhao, K. Liu, and L. Xu, c. Press, 2015. doi: 10.1162/COLI.
[5] J. E. van Engelen and H. H. Hoos, “A Survey on Semi-supervised Learning,” Mach. Learn., vol. 109, no. 2, pp. 373–440, 2020, doi: 10.1007/s10994-019-05855-6.
[6] I. G. S. M. U. F. S. V. H. S. H. A. F. Dyasa, Machine Learning. Media Nusa Creative (MNC Publishing), 2026.
[7] M. W. Dunham, A. Malcolm, and J. K. Welford, “Improved Well-log Classification Using Semisupervised Label Propagation and Self-training, with Comparisons to Popular Supervised Algorithms,” Geophysics, vol. 85, pp. OI1–OI15, 2020, doi: 10.1190/GEO2019-0238.1.
[8] H. Liu, S. K. Rallabandi, Y. Wu, P. P. Dakle, and P. Raghavan, “Self-training Strategies for Sentiment Analysis: An Empirical Study,” in Findings of the Association for Computational Linguistics, 2024.
[9] T. Yang, L. Hu, C. Shi, H. Ji, X. Li, and L. Nie, “HGAT: Heterogeneous Graph Attention Networks for Semi-supervised Short Text Classification,” ACM Trans. Inf. Syst., vol. 39, no. 3, 2021, doi: 10.1145/3450352.
[10] C. Agustina, P. Purwanto, and F. Farikhin, “Enhancing Sentiment Analysis Accuracy in Borobudur Temple Visitor Reviews through Semi-Supervised Learning and SMOTE Upsampling,” J. Adv. Inf. Technol., vol. 15, no. 4, pp. 492–499, 2024, doi: 10.12720/jait.15.4.492-499.
[11] R. Kusumaningrum, A. A. Herlambang, W. Afifah, A. Wibowo, Sutikno, and P. S. Sasongko, “Semi-Supervised Learning vs. Few-Shot Learning: Which is Better for Sentiment Analysis on Hotel Reviews Towards a Small Labeled Training Data?,” Int. J. Adv. Comput. Sci. Appl., vol. 16, no. 11, pp. 671–677, 2025, doi: 10.14569/IJACSA.2025.0161168.
[12] D. Guidotti, L. Pandolfo, and L. Pulina, “Discovering sentiment insights: streamlining tourism review analysis with Large Language Models,” Inf. Technol. Tour., vol. 27, pp. 227–261, 2025, doi: 10.1007/s40558-024-00309-9.
[13] C. Xu, M. Wang, and S. Zhu, “Tourism sentiment quadruple extraction via new neural ordinary differential equation networks with multitask learning and implicit sentiment enhancement,” Expert Syst. Appl., vol. 270, p. 126417, 2025, doi: 10.1016/j.eswa.2025.126417.
[14] M. Hajiabadi, H. Vahdat-Nejad, and H. Hajiabadi, “Enhancing tourism sentiment analysis with deep learning: a comprehensive study on social media data,” Curr. Issues Tour., 2025, doi: 10.1080/13683500.2025.2583465.
[15] C. R. Utami and I. Santiko, “Comparison of Support Vector Machine and XGBoost Algorithms in Sentiment Analysis of Visitor Reviews of Baturraden Tourism Forest,” J. Multimed. Trend Technol., vol. 4, no. 2, 2025.
Published
2026-08-29
How to Cite
Syaputri, S. A. (2026). Comparative Analysis of Self-Training XGBoost and Label Propagation for Blue Tourism Sentimen Analysis in Lampung Province. MUSTEK ANIM HA, 15(02), 70 - 76. https://doi.org/10.35724/mustek.v15i02.7836