THESIS

Integration of IndoBERT and Machine Learning Features to Improve the Performance of Indonesian

Recognizing Textual Entailment

 

ABSTRACT

This research aims to develop a Recognizing Textual Entailment (RTE) model in the Indonesian language, named Hybrid-IndoBERT-RTE. The model is designed to address challenges in recognizing textual entailment, which is a critical task in Natural Language Processing (NLP). The architecture of Hybrid-IndoBERT-RTE is built by  modifying IndoBERT-large-p1, a language model that has proven effective in various NLP tasks in the Indonesian language. In this modification, the output vectors generated by IndoBERT-large-p1 are combined with machine learning features from a feature rich classifiers, enabling the model to capture richer and deeper information The classification head of this model consists of 1 input layer, 3 hidden layers, 1 dropout layer, and 1 output layer, which are designed to enhance the model’s predictive performance. To test the model’s performance, this research uses the Wiki Revisions Edits Textual Entailment (WRETE) dataset, which consists of 450 data samples, with 300 data samples used for training, 50 for validation, and 100 for testing. Experimental results show that Hybrid-IndoBERT-RTE achieved an F1-score of 85%, indicating that the model has a strong capability in recognizing textual entailment in Indonesian In addition to good performance, the Hybrid-IndoBERT-RTE odel also demonstrates efficiency in computational resource usage. During the training process this model utilized Video Random Access Memory Graphics Processing Unit (VRAM GPU) resources 42 times more efficiently on average compared to IndoBERT-large- pl used in previous IndoNLU research. Moreover, the training time of this model is 44.44 times faster, allowing for quicker experimentation and more iterations. This efficiency is crucial in the context of RTE model development, where saving computational resources and training time can accelerate innovation and fiurther applications.

 

PUBLICATION
  • Incorporationof IndoBERT and Machine Learning Features to Improve the Performance of Indonesian Textual EntailmentRecognition (link)