NVIDIA

Cloud GPU · verified 2026-09-05

NVIDIA L4

Ada Lovelace24 GB

Ada 24 GB at 72 W. The efficiency play: video, inference, and GCP-shaped catalogs. You rent it to pack density, not to win a FLOPs argument with a 4090.

VRAM24 GB
From /GPU/hr$0.39
ArchitectureAda Lovelace
VendorNVIDIA

In brief

What to know before you rent

Catalog listings

L4 Pricing and Availability

Published on-demand rates first. Spot and waitlists are labeled. Confirm the live rate before you provision.

ProvidersRegionBilling typeInterconnect/GPU/hr
Google CloudIn stockUS-WestOn-demandPCIe$0.39Open provider
Vast.aiLimitedUS-EastOn-demandPCIe$0.41Open provider
RunPodIn stockUS-EastOn-demandPCIe$0.43Open provider
AWSIn stockUS-EastOn-demandPCIe$0.81Open provider
AzureIn stockUS-EastOn-demandPCIe$0.98Open provider

Our reading

Who this GPU is for

I rent L4 for always-on inference and media pipelines where 72 W density beats a 450 W 4090. I do not rent L4 to fine-tune a 13B because a GCP tutorial used it. Throughput per watt is not throughput per deadline.

Rent it if Endpoints, batch inference, video, people packing many GPUs per node.

Skip it if Training that should be on 4090/L40S/H100, bandwidth-bound 24 GB jobs, NVLink dreams.

How we compare hours

Datasheet

Technical specifications

NVIDIA L4 · Per GPU · PCIe

Compute

FP32
30.3 TFLOPS
FP16 Tensor
121 TFLOPS242 with sparsity
FP8 Tensor
242 TFLOPS485 with sparsity
CUDA cores
7,680
Precision support
FP8, FP16, BF16, TF32, FP32, INT8

Memory

Capacity
24 GB GDDR6
Bandwidth
300 GB/s
ECC
Yes

Silicon

Architecture
Ada Lovelace
Class
Datacenter PCIe inference

Fabric and host

GPU interconnect
None
Host interface
PCIe 4.0 x16

Power

Board power
72 W

L4 is not L40S. One letter and 24 GB vs 48 GB and 72 W vs 350 W. Catalog mix-ups are common. If the hour looks like a 4090, you are probably not on L4.

Trade-offs

Pros and cons of L4

What this GPU does well, and when the hour is the wrong buy.

Strengths

  • 72 W is how you actually pack a server
  • Ada Tensor + NVENC/NVDEC for media+AI
  • 24 GB matches A10 with a generation bump
  • Cheap to leave on, the idle story matters for endpoints

Limits

  • 300 GB/s memory is modest
  • Not a training card
  • Easy to confuse with L40S (48 GB, 350 W)
  • Hyperscaler quota still gates the “cheap” hour

Density is a feature

People dunk on L4 TFLOPS. They are measuring the wrong unit. If you run 8 inference replicas, 72 W × 8 is a different room than 450 W × 8. I pick L4 when the fleet is the product.

If you run one replica and care about tokens/sec, price a 4090 or L40S. Do not virtue-signal watts on a single GPU job.

How I actually rent it

  1. 01

    Count replicas, not TFLOPS

    Fleet math first.

  2. 02

    Do not train by tutorial

    GCP sample notebooks are not a training plan.

  3. 03

    Compare A10 and L40S

    A10 if you already have quota. L40S if 48 GB. L4 if watts.

FAQ

L4 questions

Who has the cheapest L4?+

As of September 5, 2026, the cheapest verified in-stock on-demand L4 tracked by GPUBeacon is $0.39 per GPU per hour from Google Cloud.

How much does the L4 cost per month?+

As of September 5, 2026, at 720 hours per month, one L4 costs an estimated $310 at the median on-demand price in this sample. Cheapest verified in stock: $281 per month on-demand.

Where can I rent an L4?+

GPUBeacon currently tracks L4 configs from 5 providers. The cheapest verified in-stock on-demand configs come from Google Cloud, RunPod, AWS and Azure. See the full price comparison above for every provider and config.

Are L4 prices going up or down?+

GPUBeacon publishes a catalog sample, not a 90-day price index. As of September 5, 2026, the median on-demand L4 in this sample is $0.43 per GPU per hour. The cheapest verified in-stock hour is $0.39.

Why choose the L4?+

Endpoints, batch inference, video, people packing many GPUs per node.

When is the L4 not a good fit?+

Training that should be on 4090/L40S/H100, bandwidth-bound 24 GB jobs, NVLink dreams.

L4 or L40S?+

L4: 24 GB, 72 W, density. L40S: 48 GB, 350 W, more card. Different jobs.

Is L4 good for LLMs?+

Small ones, yes, especially serving. 70B full precision, no. 7B quantized endpoints, often yes.

Keep comparing

Alternatives to NVIDIA L4

Same job, different envelope. Compare VRAM, bandwidth and the hour before you lock a badge.

NVIDIA A10

24 GB GDDR6 · 2.0x the bandwidth of the L4

From$0.45/ GPU / hr

Compare

NVIDIA L40S

48 GB GDDR6 · 2.9x the bandwidth of the L4

From$0.79/ GPU / hr

Compare

NVIDIA RTX 4090

24 GB GDDR6X · 3.4x the bandwidth of the L4

From$0.29/ GPU / hr

Compare

NVIDIA A100

80 GB HBM2e · 6.8x the bandwidth of the L4

From$1.19/ GPU / hr

Compare

Browse all GPUs

Sources

Where these numbers come from

TFLOPS and TDP are from the vendor datasheet. Hourly prices are from the GPUBeacon catalog sample. Open the originals, then confirm the live SKU with the provider before you rent.