Skip to Main Content (Press Enter)

Logo UNIMORE
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Professioni
  • Persone
  • Pubblicazioni
  • Strutture
  • Terza Missione
  • Attività
  • Competenze

UNI-FIND
Logo UNIMORE

|

UNI-FIND

unimore.it
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Professioni
  • Persone
  • Pubblicazioni
  • Strutture
  • Terza Missione
  • Attività
  • Competenze
  1. Pubblicazioni

Fully-Attentive Iterative Networks for Region-based Controllable Image and Video Captioning

Articolo
Data di Pubblicazione:
2023
Citazione:
Fully-Attentive Iterative Networks for Region-based Controllable Image and Video Captioning / Cornia, M., Baraldi, L., Ayellet, T., Cucchiara, R.. - In: COMPUTER VISION AND IMAGE UNDERSTANDING. - ISSN 1077-3142. - 237:(2023), pp. 1-10. [10.1016/j.cviu.2023.103857]
Abstract:
Controllable image captioning has recently gained attention as a way to increase the diversity and the applicability to real-world scenarios of image captioning algorithms. In this task, a captioner is conditioned on an external control signal, which needs to be followed during the generation of the caption. We aim to overcome the limitations of current controllable captioning methods by proposing a fully-attentive and iterative network that can generate grounded and controllable captions from a control signal given as a sequence of visual regions from the image. Our architecture is based on a set of novel attention operators, which take into account the hierarchical nature of the control signal, and is endowed with a decoder which explicitly focuses on each part of the control signal. We demonstrate the effectiveness of the proposed approach by conducting experiments on three datasets, where our model surpasses the performances of previous methods and achieves a new state of the art on both image and video controllable captioning.
Tipologia CRIS:
Articolo su rivista
Keywords:
Controllable captioning; Image captioning; Video captioning; Vision-and-language;
Elenco autori:
Cornia, Marcella; Baraldi, Lorenzo; Ayellet, Tal; Cucchiara, Rita
Autori di Ateneo:
BARALDI LORENZO
CORNIA MARCELLA
CUCCHIARA Rita
Link alla scheda completa:
https://iris.unimore.it/handle/11380/1319266
Pubblicato in:
COMPUTER VISION AND IMAGE UNDERSTANDING
Journal
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.7.2.0