The decision in 60 seconds · Rent or buy?
H100s aren’t scarce anymore. So, the question now is how to choose from 300+ providers, when buying makes sense, and which compromises matter for your specific workload.
If you've spent any time researching the title question, you've probably seen the same list a dozen times: AWS, GCP, Azure, Lambda Labs, CoreWeave, RunPod, Vast.ai. Each of those providers makes different tradeoffs, and the right choice for a startup fine-tuning Llama 3 will be different compared to the right choice for a public company running compliance-regulated inference at scale.
This guide explains the decisions you'll need to make, in roughly the order most teams hit them: rent or buy, which tier of provider, which specific provider in that tier, and how to avoid the procurement traps that catch most first-time GPU buyers.
The dominant answer in 2026 is rent. Multiple independent analyses converge on the same conclusion. Buying H100s only makes more sense only when GPU utilization sustains above 60–70% for 24+ months.
Real-world utilization across most AI teams averages 30–50% due to development cycles, maintenance windows, and workload variation.

The breakeven point is roughly 15,000 GPU-hours of use against a $30,000 card at $2.00/hr rental. That’s about 21 months of continuous 24/7 operation, or four years at 50% utilization. By the time you amortize an H100, the B200 would be mid-cycle and the Rubin generation will be approaching.
Even then, it only makes sense if most of these are true: sustained utilization above 60% around the clock, stable predictable workload for 2+ years, capital and staff to manage physical infrastructure, comfort with 18–24 month depreciation, and either data sovereignty requirements prevent cloud use or compliance constraints mandate on-premise hardware.
For everyone else, the question collapses to which cloud provider to rent from. That's where it gets interesting.
Despite the 300+ providers now on the market, they cluster into three meaningful tiers. The clusters aren’t based on pricing alone.

The price gap between tiers is real. The cheapest marketplace rate can be as much as 6× lower than the most expensive hyperscaler rate. But you’re not comparing identical offerings. Reliability, networking, support, and ecosystem integration differ too.
Cheap H100 hours on Vast.ai are not the same product as expensive H100 hours on Azure, and the difference can make you end up with broken production deployments or six-figure overruns.
The cloud you already use, sold to you with H100s attached.
Hyperscalers are the easy answer if you need compliance certifications (SOC 2, HIPAA, FedRAMP), enterprise procurement contracts (negotiated rates, invoicing, support SLAs), deep ecosystem integration (Sagemaker, Vertex AI, Azure ML), or multi-region deployment with consistent tooling. For most large companies, the answer is "we're already on AWS, so we're using P5". That's a perfect choice for them even at premium pricing.

Hyperscaler pricing looks straightforward until you encounter egress fees ($0.08–$0.12/GB), inter-AZ networking charges, storage IOPS costs, and startup latency that can exceed 10 minutes for large GPU nodes.
Real total cost is typically 15–25% higher than the GPU-hour rate alone suggests. For workloads that move large datasets in and out of the cloud, bear in mind that egress can dwarf compute.
Lean operations built specifically for AI compute. The sweet spot for most production ML workloads.
Choose a specialist if you're running production ML training or inference but:
CoreWeave — Specialist enterprise. Strong InfiniBand fabric, reserved discounts up to 60%, on-demand around $6.16/hr but reserved rates bring this near $2.50/hr. Best for sustained large-scale training.
Lambda Labs — The researcher favorite. H100 SXM pricing is $3.49/GPU-hr on demand for an 8× H100 cluster. Three-year reservations for large deployments can bring the rate down to $1.89/GPU-hr. Strong ML-focused tooling, reliable uptime, simple UX. Often the first specialist teams try.
GMI Cloud — $2.00/hr on-demand. NVIDIA Reference Cloud Platform Provider with optimized InfiniBand and no hidden egress fees.
Voltage Park — $2.00/hr on-demand. Specialized GPU-only operation, good price-performance for steady workloads.
Hyperbolic — $3.19/hr. Until a few weeks ago, the price was among the lowest specialist rates available. However, Hyperbolic has recently shifted away from its standout-cheap positioning.
Paperspace / DigitalOcean — H100 Droplets starting around $5.95/hr. Strong UX for startups, simpler than CoreWeave or Lambda for solo developers and small teams.
Together.ai — $3.99/GPU-hr on-demand. Strong for inference workloads with their optimized serving stack.
Specialist clouds are missing the things hyperscalers bundle. They don’t offer managed MLOps, multi-region failover, deep ecosystem services, and/or enterprise procurement contracts. For organizations that don't need those things, it makes sense that you pay only for the GPUs.
The cheapest tier and the most variable.
Choose a marketplace if you're running experimentation, hyperparameter sweeps, or interruptible batch jobs, need to minimize cost above all else, can tolerate preemption with frequent checkpointing, or are doing early-stage development before production scale-up.
Vast.ai — SXM listings run $1.73–$3.49/hr. Independent hosts set their own prices. Pricing and availability vary by host, so network performance, hardware configuration, and host reliability matter alongside the headline rate.
RunPod — H100 SXM at $3.49/GPU-hr. Offers both Secure Cloud and Community Cloud capacity, plus serverless GPU endpoints. Polished UX, popular with individual developers and small teams.
TensorDock — H100 SXM5 from $2.25/GPU-hr on demand, with Spot instances from $1.91/GPU-hr. Marketplace model with pricing that varies based on the CPU and RAM attached to the GPU.
SaladCloud — Distributed GPU infrastructure with particularly low-cost consumer GPU capacity. Its interruptible pricing model is better suited to fault-tolerant workloads than sustained H100 training. H100s, in particular, may not be available.
Marketplace economics favor low prices over reliability. Expect variable performance, sometimes-broken networking, occasionally questionable hardware provenance, and limited support. In most cases, it’s fine for experimentation but not for production inference.
If you've worked through the rent-vs-buy analysis and concluded buying is right for your situation, that is, you need predictable 24/7 utilization, sufficient capital, in-house infrastructure, here's the procurement landscape.
Direct OEMs (Dell, HPE, Supermicro, Lenovo) — The most straightforward route for complete HGX systems. For an 8-GPU H100 system, expect roughly $250,000–$320,000, or about $31,000–$40,000 per GPU once the server, NVSwitch fabric, networking, and integration are included.
Authorized distributors (Ingram Micro, TD SYNNEX, Uvation) — Useful for single-card or smaller-quantity purchases, particularly for H100 PCIe. Current market pricing is roughly $25,000–$30,000 per PCIe card, although availability and quotes vary considerably.
Workstation builders (Lambda, Bizon, Exxact, Comino) — Pre-integrated systems that bundle the GPUs with the appropriate chassis, cooling, power, and supporting components. You're paying a premium for that integration, but it can make sense when you don't have the infrastructure expertise or procurement volume to build the system yourself.
NVIDIA DGX systems — NVIDIA's fully integrated 8-GPU systems, with NVLink/NVSwitch, networking, software, and validated architecture. Expect roughly $300,000–$400,000+ depending on configuration and support.
Most teams buying their first system should look for an 8-GPU HGX H100 SXM5 server from a major OEM, with InfiniBand networking, proper liquid cooling, and a 3-year support contract. PCIe cards work for single-GPU inference but lose 30–40% of throughput on multi-GPU training due to interconnect bandwidth limits.
The GPU isn't the whole infrastructure bill. An H100 SXM5 can draw up to 700W, and an 8-GPU H100 node has a roughly 10.2 kW power envelope before you start thinking about the wider facility.
Then there's the infrastructure around it:
An 8-GPU H100 system itself is currently estimated at roughly $250K–$320K, before you account for the broader operational environment around it.
A practical complication: most production teams end up running across more than one provider. Compliance pushes some workloads onto hyperscalers. Data residency rules in Europe push others onto EU-native clouds.

Cost-sensitive batch jobs migrate to whichever marketplace is cheapest that week. The result is fragmented operations — separate consoles, inconsistent CUDA images, isolated cost views, varied governance models.
Multi-cloud control planes like emma's GPU virtual machines have emerged specifically to consolidate this: a single provisioning flow across AWS, GCP, Azure, and EU-native providers, with pre-validated CUDA images so each new VM doesn't begin with a half-day of driver debugging.
Whether you adopt a platform layer or build the orchestration in-house with Terraform, the operational complexity of multi-cloud GPU fleets is now a real line item. Sometimes, it’s larger than the hourly rate savings that motivated multi-cloud in the first place.
Here are some other costs that are important to consider:
Data egress fees. Data egress. Moving 2 TB out of AWS can cost roughly $175 in transfer charges alone. For workloads moving data between clouds, egress quickly adds up Private cross-cloud networking can cut this by 60–80%, though it requires explicit setup.
Startup latency. GPU provisioning can take minutes, while some specialist providers report 30–90 second startup times. For latency-sensitive inference, that difference matters. Driver and CUDA mismatches can add another failure point. Pre-validated CUDA images solve this but aren't standard across providers.
Networking quality variance. Distributed training depends heavily on the interconnect. InfiniBand/RDMA can deliver substantially better multi-node performance than conventional Ethernet, depending on the workload.
Idle utilization waste. A $2.50/hr H100 running at 50% useful utilization effectively costs $5.00 per productive GPU-hour. Most teams have no centralized view of GPU utilization across clouds. Utilization data lives in CloudWatch on AWS, Cloud Monitoring on GCP, Azure Monitor, and nothing at all on most specialist providers.
Governance gaps. Ad-hoc GPU purchases can fall outside RBAC, tagging and cost controls, leaving teams with unattributed or unapproved spend.
Vendor lock-in. Provider-specific images, monitoring and networking can make migration harder. Containers, Kubernetes and Terraform can reduce that friction, but don't eliminate provider dependencies.
H100 cloud rental prices have dropped 64–75% from 2024 peaks, with market rates that reached around $8–10/GPU-hour in 2024 now commonly available for roughly $2–$3.50/GPU-hour on specialist providers. There's likely another 10–20% of erosion coming as Blackwell (B200, B300) capacity ramps and frontier workloads migrate up the stack. Expect H100 spot pricing to settle in the $1.00–$1.50/hr range by late 2026.
This is broadly good for buyers and bad for anyone holding inventory. If you're considering a purchase, the secondary market for used H100s will likely soften further over the next 12–18 months — another argument for renting.
The H100 will remain the cost-performance sweet spot for most production workloads through at least mid-2027. It's no longer the frontier chip, but for the workloads most teams actually run, like fine-tuning, inference, and mid-scale training, it's now the obvious choice.
The cheapest H100 rates are generally found on marketplaces and specialist clouds. Current examples include RunPod at $1.99/hr, Voltage Park at $1.99/hr, and GMI Cloud from $2.00/GPU-hour.
But these aren't necessarily apples-to-apples. PCIe vs. SXM, spot vs. on-demand, networking, reliability, and support all affect the real cost.
For most teams, renting is the safer choice. Buying starts to make more economic sense only when you have predictable 24/7 workloads, $400K+ in capital, in-house infrastructure expertise, and accept 18–24 month depreciation cycles.
A single H100 80GB GPU costs $25,000–$30,000 (PCIe variant) or $27,000–$40,000 (SXM5 variant) as of 2026. Complete 8-GPU HGX systems range from $250,000–$320,000 including chassis, cooling, and networking. But also consider the pricing for power infrastructure, cooling upgrades, and operations to get realistic deployment costs.
There's no single best provider, and it mostly depends on the workload.
For experimentation: RunPod or Vast.ai. For production ML training: Lambda Labs or CoreWeave. For compliance-regulated enterprise workloads: AWS, GCP, or Azure. For inference at scale: Together.ai or specialist clouds with InfiniBand. The right answer depends on whether you're optimizing for cost, reliability, ecosystem, or compliance.
es. AWS offers H100s through P5 instances at around $6.88/GPU-hr on demand in US regions. Google Cloud's A3 High instances work out to roughly $11.06/GPU-hour, while Azure's H100 pricing varies significantly by region and configuration.
All three offer spot/preemptible pricing at 30–50% discounts and reserved capacity at 30–60% discounts for 1–3 year commitments.
Teams running H100 workloads across multiple clouds have to deal with fragmented consoles, CUDA images, cost views, and governance models.
Options to avoid this fragmentation include building custom Terraform modules and tagging conventions internally, or using a multi-cloud control plane like emma's GPU virtual machines, which consolidates provisioning, governance, and cost attribution across AWS, GCP, Azure, and EU-native providers behind a single workflow.
The SXM5 variant uses NVLink 4.0 at 900 GB/s for inter-GPU communication, versus PCIe 5.0 at ~128 GB/s — a 7× advantage in inter-GPU bandwidth at multi-GPU scale. SXM5 also has higher power limits (700W vs 350–400W) and greater power headroom.
Rent SXM5 for any multi-GPU training. PCIe is fine for single-GPU inference, where the cost savings of 20% to over 40% depending on provider outweigh the networking penalty.
Yes, for most workloads. The H100 has become the cost-performance sweet spot as supply caught up and prices dropped 64–75% from 2024 peaks
The B200 offers higher throughput but commands a 2–3× price premium with less mature deployment tooling. For training models under 200B parameters and standard inference workloads, the H100 typically delivers lower cost-per-token than newer Blackwell GPUs.
Yes, through multiple channels. Hyperscalers (AWS, GCP, Azure) operate EU regions with H100 capacity for data residency. EU-native providers (OVHcloud, Scaleway, and several smaller specialists) offer H100s with explicit GDPR and EU data sovereignty positioning. For organizations with strict EU data residency requirements, EU-native providers are often preferable to US hyperscaler EU regions.
Pricing verified across primary provider pricing pages and independent trackers as of September 2026. Cloud GPU rates fluctuate weekly; always confirm current rates before provisioning or signing reserved contracts.