Final replay · 12 stages
One prompt, end to end
Follow one request through model readiness, inference, GPU execution, and streamed output without changing timelines.
System spine
Switch zoom levels while preserving the exact replay stage.
System view follows routing, readiness, scheduling, inference phases, and response streaming.
Core replay
One prompt: system view
PromptHow do GPUs work?
Illustrative responseGPUs run many calculations in parallel.
ClientHow do GPUs work?
→Gateway + schedulernot submitted
→Model workercoldartifact available
→GPUready
Streamed response
No output token yet
waitingStage 1 of 12Model artifact exists
Training has already produced a versioned model artifact.
Models, serving systems, and GPU internals are three zoom levels of one token-producing process.