bartek@aws: ~/news
$ whoami
$ AWS Architect · DevOps · Cloud

tag: GPU inference

show all
Thursday, July 23, 2026

AWS Brings G7e GPU Instances to Asia and Europe on SageMaker

Amazon's beefed-up G7e instances are now live in Seoul, London, and Tokyo on SageMaker AI inference. These monsters pack up to 8 NVIDIA RTX PRO 6000 Blackwell GPUs with 96GB each, delivering 2.3x better performance than the previous G6e generation. You can now run massive 70B parameter language models without breaking a sweat or splitting across multiple nodes. Perfect for slashing latency on your generative AI workloads across Asia and Europe.

AWS Brings G6 GPU Instances to GovCloud for Faster AI Inference

AWS just dropped G6 instances in GovCloud (US-East) for SageMaker AI inference, and it's a solid move for government agencies. These beasts pack up to 8 NVIDIA L4 GPUs with 24GB memory each, delivering 2x better deep learning performance than G4dn instances. Perfect for running generative AI workloads—think language models, image generation, and computer vision—while keeping your data locked down with strict compliance requirements. If your inference models fit in 24GB of VRAM, G6 offers killer price-performance for production deployments.