• 4 minuti
  • Pubblicato

NVIDIA RTX PRO 4500 Blackwell Server Edition shakes up edge AI performance

Matteo Sala Giornalista e analista tecnologico QWERTYmag

Scritto da Matteo Sala

NVIDIA RTX PRO 4500 Blackwell Server Edition shakes up edge AI performance QWERTYmag © www.qwertymag.it
NVIDIA RTX PRO 4500 Blackwell Server Edition shakes up edge AI performance © www.qwertymag.it

NVIDIA's RTX PRO 4500 Blackwell Server Edition brings a 2.7x jump in throughput and slashes latency compared to the L4 in the same HPE ProLiant DL145 Gen11 edge server. Edge AI workloads just got a new benchmark.

Edge AI just got a serious upgrade. Drop the NVIDIA RTX PRO 4500 Blackwell Server Edition into the HPE ProLiant DL145 Gen11, and the numbers speak for themselves. This card doesn't just edge out the L4-it leaves it behind. Median throughput jumps 2.7 times. Latency drops to a fraction of what the L4 could handle. For anyone running heavy AI inference at the edge, this isn't a small step. It's a big change in what a single-slot, passively cooled GPU can do outside the data center.

Here's what that looks like. On two compact BF16 models-Qwen3.5-4B and Gemma 4 E4B-the RTX PRO 4500 pushed out 2.7 times more output tokens per second than the L4 at 8 to 32 concurrent requests. At 256 requests, the gap widened to 3.2 times. Latency, always a pain point for real-time AI, dropped hard: the 4500 kept it at 12 to 15ms per token. The L4 lagged at 35 to 40ms. Even when hammered with 256 concurrent requests, the 4500's median time to first token stayed between 264ms and 650ms. The L4, by comparison, shot up to 9 and even 39 seconds. That's not a technicality-it's the difference between a snappy AI service and one that leaves users hanging.

This kind of leap does mean more power draw. The DL145's usage climbed from 196W with the L4 to 301W with the 4500 at 32 concurrent sessions. But efficiency tells a different story: output tokens per system watt nearly doubled, hitting 5.1 on Qwen3.5-4B and 5.6 on Gemma 4 E4B. For edge setups where every watt matters, that's a trade worth making.

RTX PRO 4500 Blackwell Server Edition is not just an accelerated L4, but a more power- and memory-dense server card, purpose-built for edge AI and inference services with increased parallelism.

NVIDIA / DellVendor & OEM Analysis

Memory and precision set the RTX PRO 4500 apart. It packs 32GB of GDDR7 and supports FP4 Tensor Cores, letting it run models the L4 can't touch. On Gemma 4 26B-A4B in NVFP4 with an FP8 KV cache, the 4500 hit 1,613 output tokens per second at 32 concurrent sessions, all while staying inside its 165W power envelope. The L4 can't keep up, either in memory or speed.

Unpacking the hardware leap

NVIDIA calls the RTX PRO 4500 Blackwell Server Edition the new entry point for its Blackwell-based enterprise GPUs. It's aimed at both mainstream data centers and edge or standard server deployments. The card brings 32GB ECC GDDR7, a 165W TDP, single-slot PCIe Gen5, passive cooling, and FP4 Tensor Core support. These specs set it apart from the L4 and from the workstation version, which runs at 200W and has video outputs. This server edition is built for rackmount and edge setups where density and efficiency matter most. It's already showing up in OEM builds from GIGABYTE and Dell AI Factory, with a full rollout expected by 2026.

The HPE ProLiant DL145 Gen11, used for these tests, is made for edge sites without classic server rooms. It runs quietly, can be wall-mounted, and is tough enough for harsh conditions. The test setup paired an AMD EPYC 8534P CPU and 96GB DDR5 memory with the RTX PRO 4500, making a compact but powerful edge AI box. The DL145 Gen11 is built on the AMD EPYC 8004 Siena platform, which fits low-power edge deployments. That makes the L4-to-4500 comparison directly relevant for real-world edge use.

Real-world testing and competitive context

Tests ran on the Metrum AI Bench Platform with the vLLM serving engine, covering a range of concurrency and prompt shapes on Qwen3.5-4B and Gemma 4 E4B. The results were clear. The RTX PRO 4500 beat the L4 in every throughput and latency test. It kept per-user speed high even with 256 concurrent chat sessions, delivering 11 to 22 tokens per second per user. The L4 dropped to 3.7 to 8.7.

The DL145 can fit up to three L4s, but even then, three L4s at 32 sessions each only get close to the throughput of a single 4500. The 4500's real edge is in per-user speed, latency, and its single 32GB memory pool. For those watching the AI hardware race, these results echo the kind of jump seen in workstation upgrades, as covered earlier in desktop AI systems.

Implications for edge AI deployments

The RTX PRO 4500 Blackwell Server Edition isn't just a spec bump. It changes what's possible for edge AI. Sites already running L4s in HPE edge boxes can expect a clear jump in throughput, a big drop in latency, and the ability to run larger, more complex models. The extra 105W at the wall is balanced by nearly double the efficiency per watt. For edge deployments where speed and responsiveness are must-haves, the 4500 is the obvious upgrade-if the chassis can handle the riser and power needs.

With the RTX PRO 4500, NVIDIA has set a new bar for edge AI hardware. Edge servers no longer have to settle for hand-me-downs from the data center. This card brings performance that changes expectations. For organizations serious about AI at the edge, it's the new standard. Full specs are on the NVIDIA RTX PRO 4500 Blackwell Server Edition Product Page.

Articoli correlati