Skip to Main Content (Press Enter)

Logo UNIMORE
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Professioni
  • Persone
  • Pubblicazioni
  • Strutture
  • Terza Missione
  • Attività
  • Competenze

UNI-FIND
Logo UNIMORE

|

UNI-FIND

unimore.it
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Professioni
  • Persone
  • Pubblicazioni
  • Strutture
  • Terza Missione
  • Attività
  • Competenze
  1. Pubblicazioni

Supervised and Unsupervised Categorization of an Imbalanced Italian Crime News Dataset

Contributo in Atti di convegno
Data di Pubblicazione:
2022
Citazione:
Supervised and Unsupervised Categorization of an Imbalanced Italian Crime News Dataset / Rollo, F.; Bonisoli, G.; Po, L.. - 442:(2022), pp. 117-139. ( 16th Conference on Information Systems Management, ISM 2021 and Information Systems and Technologies conference track, FedCSIS-IST 2021 Held as Part of 16th Conference on Computer Science and Information Systems, FedCSIS 2021 Virtual, Online 2021) [10.1007/978-3-030-98997-2_6].
Abstract:
The automatic categorization of crime news is useful to create statistics on the type of crimes occurring in a certain area. This assignment can be treated as a text categorization problem. Several studies have shown that the use of word embeddings improves outcomes in many Natural Language Processing (NLP), including text categorization. The scope of this paper is to explore the use of word embeddings for Italian crime news text categorization. The approach followed is to compare different document pre-processing, Word2Vec models and methods to obtain word embeddings, including the extraction of bigrams and keyphrases. Then, supervised and unsupervised Machine Learning categorization algorithms have been applied and compared. In addition, the imbalance issue of the input dataset has been addressed by using Synthetic Minority Oversampling Technique (SMOTE) to oversample the elements in the minority classes. Experiments conducted on an Italian dataset of 17,500 crime news articles collected from 2011 till 2021 show very promising results. The supervised categorization has proven to be better than the unsupervised categorization, overcoming 80% both in precision and recall, reaching an accuracy of 0.86. Furthermore, lemmatization, bigrams and keyphrase extraction are not so decisive. In the end, the availability of our model on GitHub together with the code we used to extract word embeddings allows replicating our approach to other corpus either in Italian or other languages.
Tipologia CRIS:
Relazione in Atti di Convegno
Keywords:
Crime category; Keyphrase extraction; Text categorization; Word embeddings; Word2Vec
Elenco autori:
Rollo, F.; Bonisoli, G.; Po, L.
Autori di Ateneo:
BONISOLI GIOVANNI
PO Laura
ROLLO FEDERICA
Link alla scheda completa:
https://iris.unimore.it/handle/11380/1288751
Link al Full Text:
https://iris.unimore.it//retrieve/handle/11380/1288751/447203/Rollo%20et%20al%20-%20Supervised%20and%20Unsupervised%20Categorization%20of%20an%20Imbalanced%20Italian%20Crime%20News%20Dataset.pdf
Titolo del libro:
Lecture Notes in Business Information Processing
Pubblicato in:
LECTURE NOTES IN BUSINESS INFORMATION PROCESSING
Series
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.6.0.0