Identifying financial risk through natural language processing of company annual reports

Identifying financial risk through natural language processing of company annual reports

Files

Theron_Identifying_2020.pdf (8.31 MB)

Date

2020

Publisher

University of Pretoria

Abstract

A pipeline was developed to source annual reports of South African banks and convert them into a novel corpus. Plain text was extracted from unstructured reports whilst maintaining lineage to its coordinates in the original Portable Document Format (PDF). Initial experiments with Natural Language Processing (NLP) and machine learning classification aim at exposing financial risk inherent in the text as opposed to analysing the numerical financial values. Failed financial or governance events related to banks in the public domain were used to label annual reports as high risk. The balance of the reports were annotated as low risk to formulate a binary classification problem for machine learning. Bag of words and word embedding techniques were applied and supplemented with linguistic features like tone, uncertainty and causality based on available wordlists. Classifiers were built using traditional logistic regression and Support Vector Machine (SVM), as well as modern Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN) deep learning models. The corpus and initial findings provide a baseline for further research. Applications include an early warning system for regulators as well as question answering based on the content.

Description

Mini Dissertation (MIT (Big Data Science))--University of Pretoria, 2020.

Keywords

UCTD, Financial risk, Company annual reports, Natural language processing (NLP), Machine learning, Classification, Closed domain question answering

Citation

*

URI

http://hdl.handle.net/2263/83181

Collections

Theses and Dissertations (University of Pretoria)

Full item page

Identifying financial risk through natural language processing of company annual reports

Files

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Description

Keywords

Sustainable Development Goals

Citation

URI

Collections