Article Detail

Loading...

Sentiment Analysis on the IMDb Dataset: A Comparative Study of Classical, Deep Learning, and Transformer-Based Models

Osman Toplu

Keywords

sentiment analysis IMDb dataset machine learning deep learning transformers BERT RoBERTa

Doi : 10.71350/jere.2026186

Abstract

Text data analysis, particularly sentiment analysis, has become a critical research area due to the rapid growth of user-generated content on digital platforms. Advances in natural language processing (NLP) have enabled the automatic analysis of opinions expressed in movie reviews, social media posts, and customer feedback, making sentiment classification an essential task in both academic research and industrial applications. This study aims to conduct a comparative analysis of classical machine learning, deep learning, and transformer-based models for sentiment analysis using the IMDb 50K dataset.

References

  1. Alaparthi, S., & Mishra, M. (2020). BERT: A sentiment analysis odyssey. International Journal of Engineering and Advanced Technology, 9(4), 263–268.
  2. Atayolu, Y., Kutlu, Y. (2023). Effective Use of Content as a Feature in IMDB Dataset Analysis. Journal of Artificial Intelligence with Applications, 4(1), 1-8, doi : 10.5281/zenodo.14587323
  3. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) (pp. 4171–4186). Association for Computational Linguistics.
  4. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780.
  5. Jurafsky, D., & Martin, J. H. (2023). Speech and language processing (3rd ed., draft). Stanford University.
  6. Kim, Y. (2014). Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1746–1751). Association for Computational Linguistics.
  7. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  8. Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., & Potts, C. (2011). Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (pp. 142–150). Association for Computational Linguistics.
  9. Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
  10. Oncu, M., Kutlu, Y. (2024). Performance Comparison of Known AI Translation Tools for Turkish Language. Tethys Environmental Science, 1(3), 117-126, doi : 10.5281/zenodo.13269233
  11. Pang, B., & Lee, L. (2008). Opinion mining and sentiment analysis. Foundations and Trends in Information Retrieval, 2(1–2), 1–135.
  12. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
  13. Urban, D. (2024). Integrating Textual and Visual Data For Social Media Analysis: Insights Into Sentiment, Behavior, and Trends. Journal of Artificial Intelligence with Applications, 5(1), 18-21, doi : 10.5281/zenodo.14644001

23 20

Article Summery

ISSN : 3108-608X

Volume 2 Issue 1

Submission Date: 2026-05-27

Accepted Date : 2026-06-29

Available Online : 2026-06-30

Publication Date :2026-06-30



How to Cite

Cite as :

Toplu, O. (2026). Sentiment Analysis on the IMDb Dataset: A Comparative Study of Classical, Deep Learning, and Transformer-Based Models. Journal of Environmental Research and Engineering, 2(1), 1-6, doi : 10.71350/jere.2026186