Visual language models (VLMs) are increasingly driving transformation across multiple domains, including culinary research and food science. This study investigates the performance of open-weight large VLMs with Turkish language support in extracting recipes from culturally specific cooking videos. Experiments conducted on short-form culinary videos evaluate automated recipe generation, ingredient extraction, and cooking duration estimation, discussing future integration of retrieval-augmented generation (RAG).