The Art of Saying ‘Maybe’: A Conformal Lens for Uncertainty Benchmarking in VLMs
- 1Bangladesh University of Engineering and Technology, Dhaka, Bangladesh
- 2Qatar Computing Research Institute, HBKU, Doha, Qatar
Abstract
Vision-language models have made substantial progress in complex visual understanding, but their ability to quantify uncertainty has received less attention. This study evaluates 18 open- and closed-source models across six multimodal datasets and three scoring functions, including instruction-guided likelihood proxies for API-only models without token-level probability access.
The results show that larger models generally quantify uncertainty more effectively and that predictions made with greater certainty tend to be more accurate. Mathematical and reasoning-intensive tasks remain particularly challenging across models. The study establishes a broad foundation for evaluating uncertainty and reliability in multimodal systems.
Citation
Azad, A., Hossain, M. S., Shanto, M. S. H., Rahman, M. S., & Parvez, M. R. (2026). The Art of Saying ‘Maybe’: A Conformal Lens for Uncertainty Benchmarking in VLMs. Findings of the Association for Computational Linguistics: EACL 2026, 5185–5201.
