Advanced
For engineers who build and run LLM systems: model internals, inference cost, adapting models, retrieval quality, agents, reliability, releases and attacks. It assumes the beginner track or similar experience, and ends with a project that takes the beginner assistant to production.
Paid track · coming soonInside the Model
What the beginner diagram of a transformer leaves out.
After this level you can decide whether a reasoning model, a longer context or a multimodal model is worth its cost for your task.
- Attention variants: MHA, GQA, MQA
- Mixture of ExpertsFree
- Positional encoding and long context
- Reasoning models and test-time compute
- Multimodality
Inference Performance
Where the latency and the money go.
After this level you can decide where your latency and GPU money go, and whether batching, quantization or routing to smaller models will cut them.
Adapting Models
When and how to change the weights.
After this level you can decide whether to change the model's weights at all, and if so how: full fine-tuning, LoRA, distillation or preference tuning.
Retrieval at Depth
The pipeline runs; why are the answers mediocre?
After this level you can decide why retrieval returns mediocre results, and whether rewriting queries, reranking or a different embedding model will fix it.
Agent Architecture
What happens when one call becomes forty.
After this level you can decide whether a job needs several agents, a plan or checkpoints, and how to judge the whole run rather than one answer.
Production Reliability
Failures that have nothing to do with the model.
After this level you can decide what your system does when a dependency fails or traffic spikes: retries, rate limits, circuit breakers and fallbacks.
- Idempotency & RetriesFree
- Rate Limiting & BackpressureFree
- Circuit BreakerFree
- Fallback and graceful degradation
Ship and Observe
Releasing what you cannot unit-test.
After this level you can decide how to release a change safely, and how you'll notice when quality or cost drifts after launch.
Adversarial and Hardening
How people try to break an LLM system, and what stops them.
After this level you can decide how someone would attack your system, and which layers of defense stop injected instructions and data leaking through tools.
- Red teaming an LLM system
- Defense in depth for prompt injection
- Data exfiltration and tool permissions
Capstone
The advanced lessons applied to one system, run for real users.
After this level you can decide how to run one system for real users under cost, load and attack at the same time.