Foundation models have transformed natural language processing and computer vision, yet integrating geospatial inductive biases remains an open frontier. This chapter outlines a comprehensive conceptual and architectural framework towards GeoAI Foundation Models. We synthesize key methodologies spanning multimodal self-supervised pre-training, spatial-temporal graph encoders, and vision-language architectures tailored for geographic and Earth observation domains.