Selected Publications

CVPR 2025
A. Compagnoni*, M. Morini*, Sara Sarto*, F. Cocchi, D. Caffagni, M. Cornia, L. Baraldi, R. Cucchiara · CVPR 2026 (Highlight)
A novel Reasoning‑Augmented Multimodal RAG approach that combines coarse‑ and fine‑grained retrieval with a critic model that filters irrel‑ evant passages, ensuring high‑quality additional context. The model follows a multi‑stage training strategy leveraging reinforcement learning to enhance reasoning over retrieved content, while supervised fine‑tuning serves only as a cold start.
CVPR 2025
Davide Caffagni*, Sara Sarto*, M. Cornia, L. Baraldi, R. Cucchiara · CVPR 2025
An approach enabling multimodal queries — image + text — to search multimodal document collections through a novel Transformer-based recurrent cell integrating textual and visual features across layers.
ICCV Workshop 2025
Federico Cocchi*, Nicholas Moratelli*, Davide Caffagni*, Sara Sarto*, M. Cornia, L. Baraldi, R. Cucchiara · ICCV Workshop 2025
A new family of MLLMs integrating modern language models with diverse visual backbones.
IJCAI 2025
Sara Sarto, M. Cornia, R. Cucchiara · IJCAI 2025
An overview of image captioning evaluation, discussing metric evolution, limitations, challenges from longer MLLM captions, and metric adaptability.
CVPR Workshop 2024
D. Caffagni*, F. Cocchi*, N. Moratelli*, Sara Sarto*, M. Cornia, L. Baraldi, R. Cucchiara · CVPR Workshop 2024
Integration of external document knowledge into an MLLM through hierarchical retrieval.
ACL 2024
D. Caffagni*, F. Cocchi*, L. Barsellotti*, N. Moratelli*, Sara Sarto*, L. Baraldi*, M. Cornia, L. Baraldi, R. Cucchiara · ACL Findings 2024
A comprehensive review of recent visual-based MLLMs, analyzing architectures, alignment strategies, and training techniques.
CBMI 2022
Sara Sarto*, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara · CBMI 2022
An image captioning approach with a kNN memory, with retrieval from an external corpus to aid the generation process.
Overview of the PAC-S++ positive-augmented contrastive learning method
Sara Sarto*, N. Moratelli*, M. Cornia, L. Baraldi, R. Cucchiara · International journals of Computer Vision (IJCV) 2025
PAC-S++ strengthens CLIP-based caption evaluation with generated positive visual and textual samples, improving alignment with human judgments and serving as a reward for captioning models.
WF-Net weight-filtering layers for CNN and Transformer architectures
S. Poppi, Sara Sarto, M. Cornia, L. Baraldi, R. Cucchiara · IEEE Intelligent Systems 2024
A single-round framework that uses weight-filtering layers to selectively unlearn any image class in CNNs or Vision Transformers while exposing class-component relationships.
Retrieval-augmented self-attention and cross-attention captioning architectures
Sara Sarto, M. Cornia, L. Baraldi, A. Nicolosi, R. Cucchiara · ACM Transactions on Multimedia Computing, Communications and Applications (TOMM) 2024
Two retrieval-augmented image captioning variants combine a visual kNN retriever with either self-attention prefixes or kNN cross-attention over captions stored in an external memory.
Privacy-preserving action recognition and image description pipeline
R. Cucchiara, L. Baraldi, M. Cornia, Sara Sarto · IEEE Computer 2024
A survey and experimental perspective showing how anonymization and knowledge distillation can support action recognition and image description without exposing personal identity.