Most GPU Cloud News about H100 vs H200 vs B200 starts with a FLOPs chart. I start with a memory sentence. If you cannot name what spills out of 80 GB, you do not need H200. If you cannot name what Hopper’s fabric or precision is blocking, you do not need B200. The rest is brochure.
I compare these three because they are the cards people actually argue about in 2026. Open the live rows while you read: H100, H200, B200. Then use Compare with the same GPU count and the same billing flag.
The one-screen version
- H100 80 GB: default Hopper hour. Rent it when the job fits and you can fill the SMs.
- H200 141 GB: same family, more and faster HBM. Rent it when the KV cache, weights, or optimizer states are the bill.
- B200 180–192 GB: Blackwell. Rent it when FP4 or the new fabric is the reason, and the row is in stock.
Catalog sample on GPUBeacon, on-demand, per GPU: H100 from $1.73, H200 from $3.00, B200 from $3.75. Those are not list prices for every region. They are the cheapest published on-demand hours in the index. Confirm the live SKU. A waitlist B200 at $3.75 is not cheaper than an in-stock H100 at $2.10.
What actually changed on the silicon
H100
Hopper, 80 GB on the SXM you see in most GPU clouds, Transformer Engine, NVLink 4 in the right node. It is still the card I default to for 70B work that fits. The SM count is the family. You are not buying a toy. You are buying the hour that exists.
H200
H200 is Hopper with a memory upgrade: 141 GB HBM3e and more bandwidth. It is not Blackwell. It will not magically double your tokens per second on a job that already sat inside 80 GB. It will stop you from shrinking context, dropping batch, or adding a second GPU just to hold the KV cache.
B200
B200 is Blackwell. Catalog rows often say 180 GB; NVIDIA’s SXM story is typically 192 GB HBM3e. Confirm the SKU. FP4 dense around 9 PFLOPS is why the slide exists. Your stack has to speak FP4. Power is part of the product: these nodes sit near a kilowatt. They are not marketplace 4090s with a new badge.
VRAM vs FLOPs, applied to this fork
Fine-tunes and long-context serving die on memory first. If Adam states plus activations do not fit, extra TFLOPS are an expensive fan. I wrote GPU Cloud News this way because vendor blogs still sell generation names. I sell fit.
- Weights + optimizer + activations + KV cache = the VRAM floor.
- If the floor is ≤ 80 GB, H100 is the hour to beat.
- If the floor is 80–141 GB, H200 can drop GPU count. That is how it wins the bill.
- If you need FP4, a newer interconnect, or 180 GB+ without stacking Hopper, look at B200.
A pair of cheaper Hopper cards can still beat one Blackwell card if you pay a software tax, a waitlist, or an 8× node you cannot fill. Count GPUs. Count hours. Count whether NCCL will even start.
When I rent H100
The job fits 80 GB. I can fill the SMs. I need Hopper software that already works. I need stock today. I do not want to debug a new precision format on a deadline. That is most 7B–70B fine-tunes and a lot of serving if you are honest about context length.
I also rent H100 when H200 is 70% more expensive and I cannot point at the extra 61 GB. “Headroom” is not a workload. Read how to rent an H100 on demand if you are still mixing SKUs.
When I rent H200
Long context. Fat KV cache. A 70B+ weight that spills. A job that was two H100s only because of memory, not because of compute. Then one H200 can be cheaper than two H100s, even if the per-GPU hour is higher. Do the multiplication. Include interconnect. Include the fact that two GPUs is a distributed job.
I treat the H200 hour as a memory purchase. If I cannot point at the bytes that did not fit on H100, I do not upgrade. A GPU Cloud News post that says “H200 is better” without a VRAM sentence is an ad.
When I rent B200
The run is large enough that Hopper memory and NVLink 4 are the bottleneck. The stack is CUDA 12.8+ and actually uses FP4. The row is in stock, not a quota form. I do not put B200 on a demo that needs to boot tomorrow. I do not put it on a LoRA that fits a 24 GB card.
B300 exists too, with even more memory, if you are in that part of the catalog. For most readers of this page, B200 is already the upgrade they will not provision this week. Check B300 only after B200 is a real SKU you can start.
How to compare the three hours without lying to yourself
Same GPU count. Same region if you can. Same billing: on-demand vs on-demand. Same interconnect class. Then look at in-stock published hours. Then look at whether the job needs one GPU or eight. Then look at software: if you cannot run FP4, B200’s slide does not apply to you.
- Filter on-demand on GPUBeacon.
- Open H100, H200, B200 with the same GPU count.
- Write down VRAM, interconnect, stock, hour.
- Name the spill. If you cannot, stay on H100.
- Confirm the live rate. Then rent.
Marketplace H100 vs managed B200 is not a comparison. Vast.ai host risk is not CoreWeave cluster risk. Say the business model first. Provider index exists so you stop mixing them.
AMD sits next to this argument
MI300X is 192 GB HBM3. It is the CUDA-free way to fit a 70B on one GPU. If you are not ready for ROCm, it is an expensive science project. If you are, it can beat an H200 on memory without paying Blackwell tax. I do not pretend it is a drop-in H100.
FAQ: H100 vs H200 vs B200
Is H200 just an H100 with more VRAM?
Mostly yes: Hopper compute, 141 GB HBM3e, higher bandwidth. That is the point. Do not expect Blackwell TFLOPS. Do expect fewer memory-bound compromises.
Will B200 always beat H100?
No. Not if you are compute-bound inside 80 GB and H100 is in stock at half the hour. Not if B200 is a waitlist. Not if your serving stack does not use FP4. Silicon you cannot start is not faster.
H100 vs H200 for Llama 70B?
Adapters can live on H100. Full fine-tune with optimizer states often wants more than 80 GB or more than one GPU. Measure the states. If you were swapping or shrinking context to survive, H200 can drop GPU count and win. If you already fit, stay.
Where do I see live prices?
The catalog, not this paragraph. H100, H200, B200, then compare. This GPU Cloud News article is the decision tree. The index is the number.
Pick the card whose memory matches the job. Pay the in-stock on-demand hour. Upgrade generations last. That is the whole market once you ignore the landing pages.
Some outbound links in these articles may be affiliate links. Rankings on GPUBeacon stay independent.


