Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

Abstract

Street-level geolocalization is a challenging task requiring the fusion of visual landscape clues with spatial reference data. In this paper, we propose a novel multimodal framework that integrates state-of-the-art Vision-Language Models (VLMs) with Retrieval-Augmented Generation (RAG) to localize street-level imagery. By combining hierarchical spatial indexing and visual semantics, our approach significantly improves localization precision and provides interpretable reasoning traces for geographical predictions.

Publication
arXiv preprint arXiv:2509.01341
Yunus Serhat Bıçakçı
Yunus Serhat Bıçakçı
Assistant Professor

Assistant Professor specializing in GeoAI, Multimodal Vision-Language Models, and Spatial Data Science.

Related