Training module
M4 — Local LLMs & image generation on GPU
Run an LLM and Stable Diffusion on local GPU: quantization, context, throughput, LoRA, identity pipeline.
Duration : 1 to 2 days (depending on hardware)
- Objectives: run an LLM on a 12–24 GB card (llama.cpp and derivatives), understand quantization, context and throughput (tok/s), and turn it into a local service with an OpenAI-compatible API; generate images with Stable Diffusion while keeping control (LoRA, post-processing) — and, if time allows, an identity pipeline that keeps a face from a photo in the generated image.
- Audience: ML/AI engineers, data scientists, infra teams who have hardware (or know how to size it) and confidentiality constraints.
- Prerequisites: basic Linux, Docker, an NVIDIA GPU (or access to a remote machine; the workshop can run on provided hardware).
- Deliverables / measurable outcomes:
- a measured local LLM server (tok/s throughput and retained context written in the session report — we learn to measure, not to trust other people’s benchmarks);
- a locally generated image + a reproducible script;
- a written sizing: which GPU, which quantization, which model for their case (document taken home per pair).