Stop trusting averages. Learn why token metrics, queue depths, and tail latency are the only reliable ways to monitor LLM inference performance in production.