This study presents a comprehensive zero-shot benchmark evaluating seven state-of-the-art open-source multimodal Vision-Language Models (VLMs) with Turkish language support—including Aya Vision 32B, Gemma 3 27B, InternVL3 38B, and Qwen2.5-VL—on cultural culinary image classification using the TurkishFoods-15 and TurkishFoods-25 benchmark datasets. The results provide empirical evidence on multimodal reasoning capabilities across regional gastronomic imagery.