News: AI infrastructure

AWS Adds an Inference Gateway to SageMaker HyperPod
The EKS add-on routes model requests using workload signals such as queue length and cache utilization.

Fastly Introduces Controls for AI API Traffic
AI Runtime Control combines virtual keys, spending controls, request visibility and provider failover.

Google's GKE Agent Sandbox reaches general availability
Warm pools and snapshots aim to reduce the cost and delay of running short-lived AI jobs.

NVIDIA releases Dynamo 1.0 for AI inference
The open-source release coordinates GPU memory and requests across inference clusters.