Supervised and Unsupervised Categorization of an Imbalanced Italian Crime News Dataset

Contributo in Atti di convegno

Data di Pubblicazione:

2022

Citazione:

Supervised and Unsupervised Categorization of an Imbalanced Italian Crime News Dataset / Rollo, F.; Bonisoli, G.; Po, L.. - 442:(2022), pp. 117-139. ( 16th Conference on Information Systems Management, ISM 2021 and Information Systems and Technologies conference track, FedCSIS-IST 2021 Held as Part of 16th Conference on Computer Science and Information Systems, FedCSIS 2021 Virtual, Online 2021) [10.1007/978-3-030-98997-2_6].

Abstract:

The automatic categorization of crime news is useful to create statistics on the type of crimes occurring in a certain area. This assignment can be treated as a text categorization problem. Several studies have shown that the use of word embeddings improves outcomes in many Natural Language Processing (NLP), including text categorization. The scope of this paper is to explore the use of word embeddings for Italian crime news text categorization. The approach followed is to compare different document pre-processing, Word2Vec models and methods to obtain word embeddings, including the extraction of bigrams and keyphrases. Then, supervised and unsupervised Machine Learning categorization algorithms have been applied and compared. In addition, the imbalance issue of the input dataset has been addressed by using Synthetic Minority Oversampling Technique (SMOTE) to oversample the elements in the minority classes. Experiments conducted on an Italian dataset of 17,500 crime news articles collected from 2011 till 2021 show very promising results. The supervised categorization has proven to be better than the unsupervised categorization, overcoming 80% both in precision and recall, reaching an accuracy of 0.86. Furthermore, lemmatization, bigrams and keyphrase extraction are not so decisive. In the end, the availability of our model on GitHub together with the code we used to extract word embeddings allows replicating our approach to other corpus either in Italian or other languages.

Tipologia CRIS:

Relazione in Atti di Convegno

Keywords:

Crime category; Keyphrase extraction; Text categorization; Word embeddings; Word2Vec

Elenco autori: