Self-Hosting vs Cloud GPU: When Buying Hardware Pays Off
A100 at $1,366/month on AWS. Buying the card pays off in 10 months. Real math is more complex.
Self-Hosting vs Cloud GPU: When Does Buying Your Own Hardware Pay Off?
Cloud GPUs are expensive if you run 24/7. A single A100 instance costs about $1,366/month on AWS ($1.90/hr × 24 × 30). At that rate, an A100 card ($12,000-$15,000 retail) pays for itself in about 10 months.
But the real math is more complicated.
The Hardware Costs
| Component | Cost |
|---|---|
| NVIDIA A100 80GB (PCIe) | ~$13,000 |
| Server (dual Xeon, 512GB RAM, PSU) | ~$4,000 |
| Networking (10GbE switch, NICs) | ~$800 |
| Rack/cooling/power setup | ~$1,500 |
| Total upfront | ~$19,300 |
This is for one GPU. If you need 2-4 GPUs, the server cost increases but per-GPU cost drops. A 4× A100 server costs about $57,000 upfront.
The Ongoing Costs
| Item | Monthly |
|---|---|
| Electricity (A100 draws ~300W under load × 24h × $0.12/kWh) | $26 |
| Cooling (HVAC overhead, ~30% of power) | $8 |
| Colocation (1U rack space, power, bandwidth) | $100-200 |
| Internet (business fiber, 1Gbps) | $100-300 |
| Hardware maintenance (spare parts, failures) | ~$100 amortized |
| Total monthly | ~$450 |
Break-Even Analysis
1× A100, 24/7 inference:
| Approach | Months 1-12 | Months 12-24 | Months 24-36 |
|---|---|---|---|
| Cloud (AWS, on-demand) | $16,392 | $16,392 | $16,392 |
| Cloud (AWS, 1yr reserved) | $12,000 | $12,000 | $12,000 |
| Self-hosted | $24,700 | $5,400 | $5,400 |
Break-even:
- vs on-demand: month 15
- vs reserved: month 20
If you run the GPU 24/7 for more than 15-20 months, buying beats renting. If you run it for less, cloud wins.
What This Analysis Misses
Cloud Advantages (Not in the Math)
- No upfront capital: $19,300 at once vs monthly billing
- Scalability: Need 4 more GPUs next week? Cloud gives them. Self-hosted: order hardware, wait 4-6 weeks.
- No hardware obsolescence risk: NVIDIA releases a new GPU generation. Your cloud provider upgrades. Your self-hosted A100 is now a depreciating asset.
- Geographic distribution: 5 regions costs the same as 1 region ($0 more per region). Self-hosting in 5 locations costs 5×.
- No hardware failure risk: Cloud handles dead GPUs, PSU failures, RAM errors. Self-hosting: you handle them.
- No staffing overhead: Someone has to manage hardware, even if it's just you. Cloud: none.
Self-Hosting Advantages
- Fixed cost: $19,300 upfront, then $450/month. Predictable budgeting.
- No rate limits: Run at 100% utilization 24/7 with no throttling. Cloud providers don't throttle, but spot instances can be reclaimed.
- Data sovereignty: Hardware is physically in your location. No cloud provider has access.
- No vendor dependency: You're not tied to AWS/GCP/Azure pricing changes. NVIDIA raised prices? Too late, you already own the card.
- Full control: Overclock if you want. Custom cooling. Direct hardware access for performance tuning.
The Practical Threshold
Self-hosting starts making financial sense at:
- 3+ GPUs running 24/7 — at this scale, the cloud bill is $3,500-$5,500/month. Hardware pays off in under 12 months.
- Predictable, stable workloads — fine-tuning jobs, batch inference, always-on endpoints
- Existing data center/office space — if you already pay for space and power, the incremental cost is lower
- In-house hardware expertise — if nobody on the team has managed server hardware, factor in the learning curve
For most ML teams deploying 1-4 models, cloud is still cheaper when you account for the hidden costs of self-hosting (your time, reliability, scaling flexibility).
The Hybrid Approach
Run baseline workloads on owned hardware. Burst to cloud for spikes. This is how most mid-size ML teams operate:
- 2× A100s in colocation: handles steady traffic (80% of requests)
- Cloud GPU instances for spikes: handles peak traffic (20% of requests)
- Roptal routes traffic across both based on utilization and cost
This gives you the cost advantage of owned hardware without the scaling limitations.