Generative AI

Module aims

This module covers the deep learning methods and systems architectures driving modern generative AI. Students first bridge classical spatial and temporal representations (CNNs and RNNs) to the Attention Mechanism and Transformer Architectures that power today’s foundation models. The curriculum then develops the mathematical, probabilistic, and adversarial frameworks underpinning generative modelling—tracing the evolution from VAEs and GANs to modern Diffusion Models, Flow Matching, and Rectified Flow.
Building on these foundations, students explore applications in Multimodal Media (Image and 3D Generation) before studying Embodied AI, focusing on how Vision-Language-Action (VLA) models ground multimodal representations into physical control loops. Finally, the module addresses the engineering of Scalable and Compute-Constrained Systems, covering scaling laws, parameter-efficient fine-tuning (PEFT), and high-throughput inference optimization. Assessment combines a practical generative modelling project with a written examination.

Learning outcomes

On successful completion of this module, students will be able to:

1. Explain the probabilistic foundations of modern generative models, including latent-variable, adversarial, and score-based formulations, and derive the core objectives that underpin them.
2. Compare and contrast the principal families of generative model (variational autoencoders, generative adversarial networks, diffusion, and flow matching), evaluating their assumptions, trade-offs, and failure modes.
3. Describe the attention mechanism and transformer architecture, and analyse how they differ from convolutional and recurrent approaches as inductive biases for sequence and vision tasks.
4. Explain the pretrain-then-adapt paradigm, including self-supervised objectives, scaling laws, and parameter-efficient fine-tuning, and reason about when each is appropriate.
5. Apply generative and transformer-based methods to multimodal problems spanning image, 3D, and vision-language-action generation.
6. Evaluate the robustness, calibration, and computational cost of large generative models, and select techniques for efficient training and inference under hardware constraints.
7. Design, implement, and empirically evaluate a deep generative model for an image-generation task, and interpret its performance against quantitative metrics.    

Module syllabus

1. CNNs & RNNs
2. Attention Mechanism & Transformer Architectures
3. VAEs and GANs
4. Diffusion models
5. Flow Matching and Rectified Flow
6. Multimodal Media: Image & 3D Generation
7. Embodied AI: Vision-Language-Action (VLA) Models
8. Scalability & Compute-Constrained Systems

Pre-requisites


Teaching methods

The module is delivered over a single term through a combination of pre-recorded video lectures and in-class sessions. Each week, students first work through a pre-recorded video lecture that introduces the core concepts and mathematical foundations of the week's topic. This is followed by a timetabled in-class session that consolidates the material through worked examples, discussion, and tutorial exercises, and provides the opportunity to raise questions with the teaching team. Delivering the foundational content by video allows the contact time to be used actively for problem solving, clarification, and deeper engagement with the more challenging material.
The taught content is reinforced by a single practical coursework task that runs across the whole term. Students develop and refine a deep generative model incrementally as the relevant methods are introduced in lectures, applying each new concept as it is taught rather than in a single concentrated effort at the end. Supervised laboratory support is available through the in-class sessions and an online discussion forum, where graduate teaching assistants provide technical guidance.

Assessments

The module is assessed by a written examination worth 80 percent and practical coursework worth 20 percent.
The examination tests understanding of the theoretical and mathematical foundations of the module, including generative model formulations, attention and transformer architectures, the pretrain-then-adapt paradigm, and the principles of robustness and efficient computation. It assesses the outcomes concerned with explaining, comparing, and analysing the methods covered.
The coursework is practical and runs across the term, in which students design, implement, and evaluate a deep generative model for an image-generation problem. It assesses the applied outcomes, requiring students to translate taught methods into working implementations and interpret their results against quantitative metrics.

Written feedback on the coursework is returned via the standard submission system, identifying strengths and areas for improvement in both the implementation and the interpretation of results. Because the coursework runs across the term, students also receive ongoing formative feedback during the weekly in-class sessions and through the online discussion forum, where the teaching team responds to questions as the work develops. Cohort-wide feedback on the examination and coursework highlights common errors and gives general guidance for improvement.

Module leaders

Dr Harry Coppock
Dr Jiankang Deng
Dr Rolandos Potamias
Professor Bernhard Kainz

Reading list

To be advised - module reading list in Leganto