The five disciplines of LLM optimization
A 4:45 script, stable chapters, the HTML transcript below and a vector visual. Published video URL: none.
Accessible media
No video is presented as published. These resources are ready to record and retain their sources, chapters and limitations.
A 4:45 script, stable chapters, the HTML transcript below and a vector visual. Published video URL: none.
“LLM optimization” can refer to five different jobs. The map prevents SEO, application architecture, model compression, evaluation and agents from being confused.
The optimized object is a page or website. Work covers crawling, indexing, structure, entities, self-contained passages and observable mentions.
The optimized object is a product or pipeline. Task quality, cost and latency are measured together. Context, retrieval, routing, tools and caching remain connected to evals.
Fine-tuning, quantization, distillation and runtime selection change trade-offs between quality, memory and throughput. A theoretical estimate does not replace a benchmark.
An improvement only exists relative to a criterion. A versioned eval set retains cases, graders, failures and regressions.
An agent orchestrates model, state and actions. MCP connects data, tools and workflows without certifying their security or accuracy.