Multimodal Foundation Models
This course aims to provide an in-depth understanding of fundamental multimodal models and generative artificial intelligence technologies. Students will explore the theoretical foundations, recent advancements, and practical applications of deep learning models in multimodal contexts. By the end of the course, participants will have gained a solid understanding of modern technologies such as Transformers, CLIP models, multimodal generation, Visual Question Answering (VQA), and advanced techniques such as pruning and quantization.