PhD Dissertation
I completed my Ph.D. in Artificial Intelligence at the AImageLab , focusing on scalable multimodal foundation models, retrieval-augmented generation, and multimodal evaluation.
I am a Postdoctoral Researcher at AImageLab , University of Modena and Reggio Emilia, working on multimodal LLMs and retrieval-augmented generation.
My research focuses on multimodal architectures, retrieval techniques, and the evaluation of vision-and-language models, with particular interest in hallucination. I also work on training multimodal large language models, including LLaVA and its derivatives.
I completed my Ph.D. at the University of Modena and Reggio Emilia, advised by Prof. Rita Cucchiara , Prof. Lorenzo Baraldi , and Prof. Marcella Cornia . During my Ph.D., I spent six months as a research intern at Amazon in London.
AImageLab, University of Modena and Reggio Emilia
AImageLab, University of Modena and Reggio Emilia
6-month internship at Amazon London.
"Retrieval-augmented Transformer for Image Captioningβ
AImageLab, University of Modena and Reggio Emilia
Premio alla Memoria Davide Rabotti
University of Modena and Reggio Emilia.
I completed my Ph.D. in Artificial Intelligence at the AImageLab , focusing on scalable multimodal foundation models, retrieval-augmented generation, and multimodal evaluation.
Proud to share that I was invited to speak at the Max Planck Institute for Informatics about my research on retrieval-augmented generation!
π Happy to share that I will have a session on Retrieval-Augmentation in Multimodal Large Language models π
π Happy to share that our paper "ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering" has been accepted at CVPR 2026 in Denver! π
Happy to share that we will host the 2026 edition of IRCDL, the Conference on Information and Research Science Connecting to Digital and Library Science. Visit the website
π Happy to share that our paper "Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval" has been accepted at CVPR 2025 in Nashville! π
Our project "VISTA β Versatile Intelligent Systems for Tailored and Adaptive Next-Generation Multimodal AI" was accepted for the EuroHPC Extreme Scale grant , with an allocation of almost 1M GPU hours. Read the news on the UNIMORE website
π We are introducing LLaVA-MORE , a family of models that enhances LLaVA by integrating LLaMA 3.1 as the language model. Check out our Github repo ! π
Participation to National and European Projects
Contributing to the European initiative for developing multimodal AI systems.
Participating in the Italian National Research Project on multimodal reasoning and vision-language alignment.