From Bag-of-Words to Transformers: a Systematic Review of Multilingual Text Classification for Low-Resource Languages

Contenido principal del artículo

Masooda Masarat
Prof.Rumaan Bashir

Resumen

Text classification, a solution to the booming availability of digital textual data that has greatly amplified the demand of automated methods in order to structure and extract meaning of vast amounts of unstructured text is one of the fundamental processes of Natural Language Processing (NLP), which involves the automated placement of written texts into designated semantic categories. The paper provides a systematic literature review on multilingual text classification, including the traditional machine learning, deep learning, and transformer-based methods. It explores the methods of feature representation such as Bag-of-Words, TF-IDF as well as the more advanced embedding methods of feature representation that utilizes semantic and contextual information. The paper describes how statistical models like Naive Bayes, Support Vector Machines and Logistic Regression were developed to neural networks like CNNs, RNNs, LSTMs, and transformer-based networks including BERT. The text classification of international and Indian languages is compared and an overview has been given with data sets, techniques, and major findings. Special consideration is made to languages with low resources, especially Kashmiri, where other issues, such as the scarcity of annotated data, morphological complexity and the absence of standardized resources, still prevail. The review reveals the gaps in research that are considered critical and highlights the necessity of the development of datasets and hybrid methods. On the whole, it provides a brief summary of the current approaches and presents the further perspectives of multilingual and low-resource text classification.

Detalles del artículo

Sección
Articles