.CLOUD
SYSTEM OPERATIONAL
← Back to case studies Machine Learning / AI · Series B ML platform company · 6 weeks

$180k/mo GPU spend → down 38% in 6 weeks

How we reduced $180k/mo GPU spend by 38% in 6 weeks through MIG partitioning and inference scheduling optimization.


$180,000 Monthly GPU spend before
$111,600 Monthly GPU spend after
38% Reduction
2 weeks Time to first savings

Context: A Series B machine learning platform was running a mixed workload of training and inference across three separate GKE clusters. The infrastructure relied entirely on 24× A100 80GB instances. The monthly GPU bill had reached $180,000, and growing inference load was driving linear GPU provisioning, severely impacting unit economics.

Constraint: The environment presented strict operational boundaries. Training jobs were long-running and could not be interrupted without significant data loss. Inference latency SLAs required p99 response times of under 200ms. The engineering team consisted of two ML engineers with no dedicated infrastructure personnel, meaning any changes had to be entirely non-disruptive and require minimal ongoing maintenance.

What we changed:

  • Profiled GPU utilization using the DCGM exporter, revealing a stark contrast: inference pods utilized only 15-30% of the SM capacity, while training jobs utilized 85-95%.
  • Configured MIG (Multi-Instance GPU) partitioning on inference nodes to align hardware allocation with actual requirements. We established 3g.40gb slices for large models and 1g.10gb slices for smaller, supporting models.
  • Migrated training workloads to preemptible A100 instances, implementing automated checkpointing since the training framework was already inherently fault-tolerant.
  • Implemented Karpenter for dynamic inference node auto-provisioning, configured specifically for GPU-aware bin-packing.
  • Deployed a comprehensive DCGM dashboard to provide the engineering team with ongoing, high-fidelity utilization monitoring.

Measured result: Total GPU spend dropped from $180,000 to $111,600 per month, representing a 38% reduction. Crucially, inference latency remained unchanged, actually tightening slightly from a p99 of 185ms to 182ms. Training throughput saw a 4% improvement due to the transition to dedicated, specialized node pools.