Amazon SageMaker HyperPod Inference Gateway cuts LLM latency by up to 82%
Amazon just dropped SageMaker HyperPod Inference Gateway, a Kubernetes-native routing system that replaces dumb round-robin load balancing with real-time inference signals—think KV cache utilization and queue depth. It's available now for per-cluster routing across all AWS regions where HyperPod inference is supported, with cross-cluster and cross-region magic coming soon.
source: [aws/whats-new]