REBRANDING NOTICE: EzDeploy is now Roptal. Platform launching on roptal.com. Docs: docs.roptal.com.REBRANDING NOTICE: EzDeploy is now Roptal. Platform launching on roptal.com. Docs: docs.roptal.com.REBRANDING NOTICE: EzDeploy is now Roptal. Platform launching on roptal.com. Docs: docs.roptal.com.REBRANDING NOTICE: EzDeploy is now Roptal. Platform launching on roptal.com. Docs: docs.roptal.com.
Oryvo
← All articles

Self-Hosting vs Cloud GPU: When Buying Hardware Pays Off

A100 at $1,366/month on AWS. Buying the card pays off in 10 months. Real math is more complex.

Self-Hosting vs Cloud GPU: When Does Buying Your Own Hardware Pay Off?

Cloud GPUs are expensive if you run 24/7. A single A100 instance costs about $1,366/month on AWS ($1.90/hr × 24 × 30). At that rate, an A100 card ($12,000-$15,000 retail) pays for itself in about 10 months.

But the real math is more complicated.

The Hardware Costs

ComponentCost
NVIDIA A100 80GB (PCIe)~$13,000
Server (dual Xeon, 512GB RAM, PSU)~$4,000
Networking (10GbE switch, NICs)~$800
Rack/cooling/power setup~$1,500
Total upfront~$19,300

This is for one GPU. If you need 2-4 GPUs, the server cost increases but per-GPU cost drops. A 4× A100 server costs about $57,000 upfront.

The Ongoing Costs

ItemMonthly
Electricity (A100 draws ~300W under load × 24h × $0.12/kWh)$26
Cooling (HVAC overhead, ~30% of power)$8
Colocation (1U rack space, power, bandwidth)$100-200
Internet (business fiber, 1Gbps)$100-300
Hardware maintenance (spare parts, failures)~$100 amortized
Total monthly~$450

Break-Even Analysis

1× A100, 24/7 inference:

ApproachMonths 1-12Months 12-24Months 24-36
Cloud (AWS, on-demand)$16,392$16,392$16,392
Cloud (AWS, 1yr reserved)$12,000$12,000$12,000
Self-hosted$24,700$5,400$5,400

Break-even:

  • vs on-demand: month 15
  • vs reserved: month 20

If you run the GPU 24/7 for more than 15-20 months, buying beats renting. If you run it for less, cloud wins.

What This Analysis Misses

Cloud Advantages (Not in the Math)

  • No upfront capital: $19,300 at once vs monthly billing
  • Scalability: Need 4 more GPUs next week? Cloud gives them. Self-hosted: order hardware, wait 4-6 weeks.
  • No hardware obsolescence risk: NVIDIA releases a new GPU generation. Your cloud provider upgrades. Your self-hosted A100 is now a depreciating asset.
  • Geographic distribution: 5 regions costs the same as 1 region ($0 more per region). Self-hosting in 5 locations costs 5×.
  • No hardware failure risk: Cloud handles dead GPUs, PSU failures, RAM errors. Self-hosting: you handle them.
  • No staffing overhead: Someone has to manage hardware, even if it's just you. Cloud: none.

Self-Hosting Advantages

  • Fixed cost: $19,300 upfront, then $450/month. Predictable budgeting.
  • No rate limits: Run at 100% utilization 24/7 with no throttling. Cloud providers don't throttle, but spot instances can be reclaimed.
  • Data sovereignty: Hardware is physically in your location. No cloud provider has access.
  • No vendor dependency: You're not tied to AWS/GCP/Azure pricing changes. NVIDIA raised prices? Too late, you already own the card.
  • Full control: Overclock if you want. Custom cooling. Direct hardware access for performance tuning.

The Practical Threshold

Self-hosting starts making financial sense at:

  • 3+ GPUs running 24/7 — at this scale, the cloud bill is $3,500-$5,500/month. Hardware pays off in under 12 months.
  • Predictable, stable workloads — fine-tuning jobs, batch inference, always-on endpoints
  • Existing data center/office space — if you already pay for space and power, the incremental cost is lower
  • In-house hardware expertise — if nobody on the team has managed server hardware, factor in the learning curve

For most ML teams deploying 1-4 models, cloud is still cheaper when you account for the hidden costs of self-hosting (your time, reliability, scaling flexibility).

The Hybrid Approach

Run baseline workloads on owned hardware. Burst to cloud for spikes. This is how most mid-size ML teams operate:

  • 2× A100s in colocation: handles steady traffic (80% of requests)
  • Cloud GPU instances for spikes: handles peak traffic (20% of requests)
  • Roptal routes traffic across both based on utilization and cost

This gives you the cost advantage of owned hardware without the scaling limitations.

Self-Hosting vs Cloud GPU: When Buying Hardware Pays Off — Oryvo AI Blog