Improving short text classification through global augmentation methods

Marivate, Vukosi; Sefara, Tshephisho

Improving short text classification through global augmentation methods

Files

Marivate_Improving_2020.pdf (532.54 KB)

Date

2020-08

Authors

Marivate, Vukosi

Sefara, Tshephisho

Publisher

Springer

Abstract

We study the effect of different approaches to text augmentation. To do this we use three datasets that include social media and formal text in the form of news articles. Our goal is to provide insights for practitioners and researchers on making choices for augmentation for classification use cases. We observe that Word2Vec-based augmentation is a viable option when one does not have access to a formal synonym model (like WordNet-based augmentation). The use of mixup further improves performance of all text based augmentations and reduces the effects of overfitting on a tested deep learning model. Round-trip translation with a translation service proves to be harder to use due to cost and as such is less accessible for both normal and low resource use-cases.

Keywords

Natural language processing (NLP), Data augmentation, Text classification, Deep neural network (DNN)

Citation

Marivate V., Sefara T. (2020) Improving Short Text Classification Through Global Augmentation Methods. In: Holzinger A., Kieseberg P., Tjoa A., Weippl E. (eds) Machine Learning and Knowledge Extraction. CD-MAKE 2020. Lecture Notes in Computer Science, vol 12279. Springer, Cham. https://doi.org/10.1007/978-3-030-57321-8_21.

URI

http://hdl.handle.net/2263/76628

Collections

Research Articles (Computer Science)
Research Articles (University of Pretoria)

Full item page

Improving short text classification through global augmentation methods

Files

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Description

Keywords

Sustainable Development Goals

Citation

URI

Collections