Head to head · verified Sep 5, 2026

H100 vs H200: rent the memory you can name

H200 is Hopper with a bigger tank. Same family, same NVLink generation, more HBM3e. I rent H100 when 80 GB holds the job. I rent H200 when I can point at the bytes that spill. If you cannot name the spill, the extra hour is a brochure.

H10080 GB · $1.73
vs
H200141 GB · $3.00
H100$1.73
VRAM80 / 141 GB
H200$3.00
The resultDefault: H100

The result

Default: H100

Pick: H100

Start on H100. Upgrade to H200 only after a context window, KV cache, or optimizer state does not fit 80 GB. Two H100s vs one H200 is a real comparison. “H200 is newer” is not.

In brief

What to know before you pick

Same fields

The comparison

Catalog facts first. The pick above is editorial. Confirm the live on-demand row before you provision.

FieldH100H200
VRAM80 GB141 GBWins this row
MemoryHBM3TieHBM3eTie
Bandwidth3,350 GB/s4,800 GB/sWins this row
Catalog from-price$1.73Wins this row$3.00
Cheapest indexed onVast.aiKoyeb
On-demand rows in sample1510
In-stock rows106
Architecturehopperhopper

Rent H100 if

The 70B fits. Fine-tunes, serving with a sane context, single-node training that actually uses NVLink. You want stock this week, not a memory souvenir.

Skip it if Context windows that blow the KV cache, multi-node jobs that need a reserved fabric, or anyone shopping a 4090 workload with Hopper money.

Rent H200 if

You can name what fell out of 80 GB: KV cache at long context, Adam states, a 70B that you refuse to quantize. Then the H200 hour can drop GPU count and win the bill.

Skip it if Batches that already fit H100, budgets that only work at Vast.ai floor H100, or anyone who needed NVLink 5 and FP4.

H100

  • The SKU every catalog actually stocks. You can compare an hour, not a waitlist
  • FP8 Transformer Engine is the reason this generation still trains at a sane $/step
  • MIG exists when you want to split an 80 GB card instead of overpaying for a second one

H200

  • 141 GB is the difference between one GPU and a sharded headache for many 70B stacks
  • 4.8 TB/s HBM3e moves KV cache without pretending bandwidth is free
  • Same Hopper software path as H100, fewer “new architecture” surprises

The only test is whether 80 GB holds

H200 is Hopper with a bigger tank. Same SM family, same NVLink generation, more HBM3e and more bandwidth. If the model, optimizer states and KV cache already sit inside 80 GB, you are renting headroom. Headroom is not a workload.

I watch nvidia-smi, not the brochure. If you are shrinking context to survive, or sharding a 70B only because of memory, the H200 hour can drop a GPU and win the bill. If you are compute-bound inside 80 GB, the extra dollars buy a prettier screenshot.

Two H100s vs one H200 is the real fork

People compare one H100 to one H200 because the URLs look like that. The job often is: do I buy a second Hopper, or do I buy memory. One H200 that actually fits can beat two H100s on interconnect tax, checkpointing, and the fact that two GPUs is a distributed job.

Do the multiplication on 720 hours. Include the second GPU, the fabric, and the engineering time. Then open both listings and check in-stock on-demand. A cheaper H200 that is waitlist is not cheaper.

SXM vs PCIe still matters more than the badge

H100 SXM and H100 PCIe are different products. H200 is the same trap with a bigger number. Confirm form factor, interconnect and TDP on the row you will paste a card into. Marketplace “H200” at an H100 price is often an 80 GB mislabel or a host you should not train on.

Filter on-demand. Read SXM vs PCIe. Then decide. The catalog sample on this page is the map, not the reservation.

The trap

How this comparison usually goes wrong

Paying H200 prices for an H100 job. Marketplace “H200” rows that are 80 GB mislabels. Mixing a waitlist H200 with an in-stock H100 and calling H200 cheaper.

What I would do

The move

Open both listings. Filter on-demand, in-stock. Confirm SXM vs PCIe. If the spill is real, rent H200. If it is not, keep the H100 and spend the delta on another GPU or a longer run.

Catalog listings

H100 and H200 in this catalog

On-demand hours first. Spot and waitlists are labeled. From-prices are catalog samples, not a reservation.

ProvidersRegionBilling typeInterconnect/GPU/hr
Vast.aiIn stockUS-EastOn-demandPCIe$1.73Open provider
Packet.aiLimitedUS-WestOn-demandNVLink$2.19Open provider
RunPodIn stockUS-EastOn-demandNVLink$2.39Open provider
LambdaIn stockUS-WestOn-demandNVLink$2.49Open provider
FluidstackIn stockEU-WestOn-demandPCIe$2.69Open provider
KoyebIn stockEU-WestOn-demandNVLink$2.85Open provider
HyperstackIn stockEU-WestOn-demandNVLink$2.99Open provider
CrusoeIn stockUS-WestOn-demandNVLink$2.99Open provider
NebiusIn stockEU-NorthOn-demandNVLink$3.10Open provider
Voltage ParkIn stockUS-WestOn-demandNVLink$3.35Open provider
Together AIIn stockUS-WestOn-demandNVLink$3.49Open provider
CoreWeaveLimitedUS-EastOn-demandInfiniBand NDR$3.89Open provider
AWSLimitedUS-EastOn-demandNVLink$4.10Open provider
Google CloudLimitedUS-WestOn-demandNVLink$4.85Open provider
AzureLimitedUS-EastOn-demandNVLink$4.98Open provider

Full review

The two reviews

Same product, full page. Specs, listings, who it is for. Confirm the live hour there.

FAQ

H100 vs H200 questions

Is H200 just an H100 with more VRAM?+

Mostly. Hopper compute, 141 GB HBM3e, higher bandwidth. Do not expect Blackwell TFLOPS. That is the point of the SKU.

Will H200 always beat H100 on Llama 70B?+

No. If you were already compute-bound inside 80 GB, you bought headroom. If you were shrinking context to survive, H200 can drop a GPU and win.

Which is cheaper per hour?+

H100, in this catalog sample. Cheap is not the test. Bytes that do not fit are the test. Confirm the live on-demand row before you provision.

When do two H100s beat one H200?+

When the job is compute-bound inside 80 GB, when you already have Hopper software, or when H200 is waitlist and H100 is in stock. Memory-bound 70B serving with a fat KV cache is the other way around.

H100 vs H200 for fine-tuning?+

LoRA that fits 80 GB: H100. Full fine-tune whose Adam states spill: H200, or more GPUs. Measure the envelope, not the generation year.

Keep comparing

Other matchups