Hybrid Transformer-RNN Model for Classification of Indonesian Regional Language
ABSTRACT:
Transformer-based language models have been widely adopted in natural language processing tasks, yet their ability to capture sequential dependencies remains limited-particularly in Indonesia’s regional languages, which are low-resource and morphologically hybrid architecture that integrates rich. This study proposes a NusaBERT with BiLSTM and BiGRU layers on top of Transformer representations, while exploring various pooling strategies (CLS token, last hidden state, mean, and Experiments were conducted on three benchmark datasets-NusaParagraph max) Emotion, Rhetoric. Topic) NusaTranslation (Emotion, Sentiment), and Nusax Sentiment)-for multi-class and cross-lingual classification tasks. Results show that the hybrid models consistently outperform baselines, with the best variant NusaBERTLarge BiGRU, mean pooling, batch size 8) achieving macro FI scores of 77.09% (Emotion), 53.38% (Rhetoric), and 88.81% (Topic) on NusaParagraph, 71.03% (Emotion) and 88.71% (Sentiment on NusaTranslation; and 83.26% on NusaX (average across languages). Additional evaluation on the phenomenon of catastrophic forgetting shows that the hybrid model maintains more stable performance when sequential fine-tuning is applied across languages. These findings demonstrate that combining contextual representations from Transformers with the sequential modeling capabilities of RNNs can improve both performance and robustness in multilingual NLP scenarios, while supporting the development of more inclusive and adaptive language technologies for regional language preservation in Indonesia.
PUBLICATION
Hybrid Transformer-RNN Model for Classification of Indonesian Regional Language (link)