AI
7.0
Chip Huyen explains how to cut inference costs without new hardware
Chip Huyen's P99 CONF keynote distills inference optimization strategies, arguing that training-to-inference cost ratios of 1:10 to 1:100 make inference the dominant expense over a model's lifetime. The talk covers key latency metrics like time-to-first-token (TTFT) and practical techniques to reduce token burn without hardware upgrades, directly addressing why frontier model economics remain unprofitable for most users.
Read article →