AMD EPYC Venice vs NVIDIA Vera: What the SPEC 2026 Numbers Actually Show

Posted by Reda Fornera on 2026-09-23
Estimated Reading Time 18 Minutes
Words 3k In Total

Data center aisle lined with glowing server racks — generic stock photo representing the 2026 AMD EPYC Venice vs NVIDIA Vera server CPU head-to-head, not an actual depiction of either chip

The AMD EPYC Venice vs NVIDIA Vera matchup is the server-CPU fight of 2026. Two vendor white papers in three months, each with the first official SPEC CPU 2026 numbers for a brand-new server CPU, each aimed squarely at the other. In July 2026, NVIDIA published estimated SPEC CPU 2026 Integer results for its Vera CPU, as covered by ServeTheHome. In September, AMD answered — the first official EPYC “Venice” benchmarks since the July launch — and the target was not Intel. As hwbusters put it: “The target is not Intel. It is NVIDIA’s Vera.”

The headline: AMD claims its 256-core EPYC 9996 delivers 2.24× the platform throughput of an 88-core Vera system in SPECrate 2026 Integer. Skeptics, including eWeek’s own analysis, say the race is much closer once you count cores, compilers, and watts. Both things are true, depending on how you read the chart.

Meet the contenders

One framing point first: these are not the same kind of product.

EPYC 9006 “Venice”: AMD’s Zen 6 general-purpose workhorse

Venice is the sixth generation of AMD’s EPYC server line, built on the Zen 6 architecture and — per hwbusters’ report on the white paper — TSMC’s N2 node, making it “the first high-performance x86 part on 2nm-class silicon.” The stack comes in two flavors: a core-density branch topped by the 256-core/512-thread EPYC 9996 (Zen 6c cores, 2.55 GHz base, 4.1 GHz boost, 1,024 MB of L3, 600 W TDP), and a high-frequency branch topping out at 96 Zen 6 cores running past 5 GHz. Venice launched in July 2026 and, per eWeek, also serves as the CPU in AMD’s Helios rack-scale platform.

Typical use cases: virtualization, cloud, databases, HPC, and the general-purpose compute layer of AI infrastructure. It’s a generalist.

NVIDIA Vera: a purpose-built host CPU for AI factories

Vera is NVIDIA’s first CPU with a fully in-house core design — the Arm-compatible “Olympus” core — and it ships as the host CPU in NVIDIA’s rack-scale systems, alongside Rubin accelerators. That rack-scale interconnect story is part of NVIDIA’s broader ecosystem push, including its investment in MediaTek to advance NVLink Fusion. NVIDIA’s developer blog describes the design target bluntly: strong single-threaded performance for the sequential, latency-bound critical path of agentic AI workloads, plus enough concurrency to absorb bursts of parallel tool calls and sandboxes. Each Vera CPU has 88 Olympus cores with Spatial Multithreading (176 threads), up to 1.5 TB of LPDDR5X on detachable SOCAMM modules, and up to 1.2 TB/s of memory bandwidth. Per NVIDIA, Vera systems arrive from major OEMs in H2 2026, and — per Tom’s Hardware’s reporting (via Pipedot’s republication) — NVIDIA plans a single 88-core Vera SKU. Vera is a specialist, sold as part of an integrated rack story.

That asymmetry — generalist versus specialist — explains most of the benchmark controversy that follows.

AMD EPYC Venice vs NVIDIA Vera: quick comparison table

All figures below come from fetched vendor disclosures and covering outlets; performance figures are estimates.

Close-up of a green circuit board with microchips — generic stock photo used to illustrate the AMD EPYC Venice vs NVIDIA Vera spec comparison below, not an actual spec table or either processor

Spec AMD EPYC 9996 “Venice” NVIDIA Vera CPU
Architecture Zen 6c (6th-Gen EPYC), TSMC N2 NVIDIA Olympus core, Arm v9.2-compatible
Cores / threads 256 cores / 512 threads 88 cores / 176 threads (Spatial Multithreading)
Clocks 2.55 GHz base / 4.1 GHz boost (hwbusters); high-freq branch past 5 GHz Not disclosed (ServeTheHome’s analysis)
Cache 1,024 MB L3 Shared L3 via 2nd-gen Scalable Coherency Fabric (164 MB per CPU, per NVIDIA’s developer blog)
TDP 600 W (SP7 socket) [UNVERIFIED: official Vera CPU TDP from a fetched primary source]
Memory 16-channel DDR5 (MRDIMMs, DDR5-12800 per reader analysis on Tom’s Hardware) Up to 1.5 TB LPDDR5X, SOCAMM modules (NVIDIA developer blog)
Peak memory bandwidth ~1.64 TB/s theoretical; ~18% Stream lead over Vera claimed (reader analysis on Tom’s Hardware / hwbusters) Up to 1.2 TB/s (NVIDIA developer blog); ~1.23 TB/s theoretical
I/O 128 PCIe 6.0 lanes, CXL 3.1 (hwbusters) PCIe 6.4, 88 lanes per CPU (Tom’s Hardware deep dive)
Claimed SPECrate 2026 Integer (2P, estimated) 2,070 (AMD, per eWeek) 925 (NVIDIA white paper, per eWeek and ServeTheHome)
Availability / pricing Launched July 2026; pricing [UNVERIFIED: official list pricing for either platform] H2 2026 from major OEMs (NVIDIA developer blog); pricing [UNVERIFIED: official list pricing]

Platform throughput: the 2.24× claim, dissected

Here is the raw disclosure, as eWeek reported it: an estimated SPECrate 2026 Integer score of 2,070 for a two-socket EPYC 9996 system versus 925 for a two-socket Vera platform. Divide and you get the 2.24× platform-throughput claim.

Sounds like a rout. It isn’t, for one obvious reason: core count. The two-socket EPYC 9996 system has 512 cores and 1,024 threads; the two-socket Vera system has barely more than a third as many cores (176 vs 512). eWeek’s analysis states it plainly: “The 2.24× throughput estimate therefore does not mean an individual Venice core is more than twice as fast.” hwbusters flips it: getting 2.24× out of ~2.9× the threads “is a respectable scaling result for a 256-core part, but it is not the blowout the number suggests at a glance.”

AMD’s counterargument, as Tom’s Hardware’s reporting relays it: the comparison is fair because NVIDIA only offers Vera as a single 88-core SKU — a throughput race Vera could never win, and NVIDIA has argued it isn’t trying to. NVIDIA’s developer blog makes the same point from the design side: “High-core-count systems might appear efficient on paper, but they often force a trade-off,” sacrificing single-thread performance for density. And per ServeTheHome, NVIDIA’s own white paper emphasized single-threaded results; its full-chip figure — 925 versus 898 for AMD’s EPYC 9755, a 3% lead — appeared only at the end, un-headlined.

Per-core performance: Venice’s stronger ground

This is where AMD’s case holds up much better — on its opponent’s turf.

Down-cored to 96 cores within the same 600 W budget, AMD claims roughly 20% more SPECrate 2026 Integer throughput than Vera, with about 1.2× the per-core performance (hwbusters). The underlying scores, as eWeek reported: 1,210 for the 96-core Venice configuration (6.3 per core) versus 925 for Vera (5.3 per core). Pipedot’s analysis re-ran the math and got “about an 18.8% lead. That’s not far off enough to say AMD was maliciously juicing its own numbers, but it’s important to note.”

Abstract macro shot of silicon chip circuitry — generic stock photo representing the SPECrate 2026 Integer benchmark comparison discussed here, not an actual bar chart or benchmark result

AMD’s own people leaned into this. Ravi Kuppuswamy, AMD’s corporate VP of compute and enterprise solutions, was quoted in Pipedot’s coverage: “We are very happy that Nvidia published their Vera performance [numbers]… we thought we’d have at least a 10% advantage. What we’re finding is… we have 20% advantage, and we have not even finished completely tuning.”

Two caveats:

  • The SKU mismatch. Pipedot’s analysis notes the 96-core configuration that scored 1,210 “doesn’t seem to exist” as shipped: AMD’s actual high-frequency 96-core part, the EPYC 9686F, carries a 500 W TDP — and the benchmark ran at 600 W (The Clarity Today flags the same gap). The number describes a config you can’t buy as-is.
  • AMD’s own figures have wobbled across documents. In earlier rack-scale material (per Data Center News), AMD estimated a 27% per-core advantage for a 64-core Venice part and an 11% lead for a 96-core chip; the September white paper settles on ~20% for 96 cores. Read charitably: modeling improved. Read skeptically: the number depends on which AMD document you open.

NVIDIA’s developer blog claims the mirror image: “NVIDIA Vera CPU delivers up to 1.5x the per-core performance of AMD Venice across four agentic workloads” — with a footnote revealing that AMD Venice results were estimated from the 2,070 SPECrate score “with components normalized based on internal Turin measurements.” Both vendors are now estimating the other’s chips.

Memory bandwidth and the rack-density counterargument

Memory is Vera’s strong ground, and AMD’s paper mostly concedes it. That’s no accident: memory bandwidth and memory cost have become the defining constraint of AI infrastructure economics — see our analysis of AI chip memory costs and HBM dominance.

Vera’s LPDDR5X SOCAMM subsystem delivers up to 1.2 TB/s while using, per NVIDIA’s developer blog, “less than half the memory power of traditional DDR configurations,” with a 3.4 TB/s bisection bandwidth through the Scalable Coherency Fabric sustaining over 90% of peak bandwidth under load. AMD’s Stream comparison — per hwbusters and The Clarity Today — gives 96-core Venice an ~18% lead in total bandwidth and ~8% per core, with the Vera side borrowed from Phoronix’s independent testing. A reader analysis on Tom’s Hardware’s article added a sharper angle: the raw theoretical bandwidths are ~1.64 TB/s (AMD’s 16-channel MRDIMM DDR5-12800) versus ~1.23 TB/s (Vera’s SOCAMM2 LPDDR5X-9600) — meaning NVIDIA converts a smaller pipe far more efficiently (~89% vs ~79%), and does it with lower power.

Then there’s the rack-level counterargument, AMD’s strongest card. AMD’s rack-scale modeling, as reported by Data Center News, uses a modeled 100 kW rack and claims EPYC 9965 “Turin” systems deliver 2.37× the rack-level throughput of a Vera baseline (~1.6× Intel’s Xeon 6980P), with Venice projected to extend the Vera comparison to 3.30×. eWeek’s account of the September paper rounds that to up to 3.4×, across six workloads — explicitly a model, not a measured rack. In AMD’s density math, a Turin-based liquid-cooled rack could host more than 27,000 CPU cores, with Venice designed to exceed 36,000.

NVIDIA’s counter is a different definition of density. Its Vera CPU Rack scales to 256 liquid-cooled Vera CPUs per rack running, per NVIDIA’s developer blog, more than 22,500 concurrent sandboxes — which NVIDIA claims delivers over 4× the capacity and 2× the performance-per-watt of x86 server racks. The per-rack cores-versus-memory framing sometimes quoted in this debate (704 cores and ~12 TB versus 512 cores and ~8 TB) [UNVERIFIED: exact per-rack core and memory configuration figures attributed to a specific fetched source]; each side’s “rack” optimizes for something different — AMD’s for total CPU throughput, NVIDIA’s for sandbox/orchestration capacity.

Tangled network cables in a server rack — generic stock photo representing the rack-density discussion, not a diagram of AMD or NVIDIA rack configurations

The compiler controversy

The sharpest technical criticism of AMD’s disclosure concerns compiler versions — and the answer is more nuanced than either “AMD cheated” or “nothing to see here.”

The mismatch: the new per-core benchmarks in AMD’s white paper ran with GCC 16.1, while the Vera scores AMD references were produced with GCC 15.2 (hwbusters; The Clarity Today). GCC 16 is the release that added Zen 6 support, so the comparison “folds architecture and compiler maturity into a single ratio with no clean way to separate them” — though hwbusters itself concedes this is “less sleight of hand than an unavoidable consequence of benchmarking a part whose compiler support only just landed.”

There’s an important wrinkle: for the headline 2.24× platform comparison, eWeek reported that AMD’s methodology “lists GCC 15.2 for both systems” — so the compiler mismatch applies to the newer per-core comparisons, not (per that reading) the flagship throughput ratio. Tom’s Hardware’s take, as quoted in The Clarity Today’s coverage: “it’s not best practice to compare benchmarks using two different compiler versions.” The Clarity Today adds the buyer’s warning: the per-core margin “transfers only if both sides are built with the compiler and flags you plan to ship.”

A legitimate defense exists: as one technical reader noted on Tom’s Hardware’s article, SPEC CPU has always allowed each vendor to supply its own hardware and software stack, and CPU-specific compiler tuning is standard practice. GCC 16.1 is simply the first compiler that knows what Zen 6 is. The honest summary: the mismatch is understandable, but until someone rebuilds Vera’s run on GCC 16.1, the ~20% per-core number bundles compiler and silicon together. The one apples-to-apples comparison in AMD’s paper, meanwhile, is generational: the EPYC 9996 is about 78% faster than AMD’s own 192-core EPYC 9965 — same vendor, same benchmark, same methodology. As hwbusters put it: “That one you can more or less take at face value.”

Verdict: AMD EPYC Venice vs NVIDIA Vera — closer than the headline

Caveats first:

  1. Nothing here is independently validated. Every figure from both white papers is an estimate from pre-production or engineering hardware; neither vendor can file official SPEC submissions for unreleased products (hwbusters; Pipedot’s note on SPEC’s reporting rules).
  2. The configurations don’t match. 256-core/600 W versus 88-core Vera is a platform comparison, not a chip comparison (eWeek’s analysis); the 96-core proxy ran at 600 W against a 500 W shipping SKU (Pipedot; The Clarity Today).
  3. The compilers don’t match in the per-core comparison (GCC 16.1 vs 15.2), and NVIDIA has published no floating-point results for Vera at all (ServeTheHome).
  4. Independent context cuts against both sides. ServeTheHome notes published dual-socket EPYC 9755 systems on spec.org scoring over 1,000 on SPECrate 2026 Integer with AMD’s own compiler — above Vera’s 925 — and Turin Dense systems over 1,200.

With all that said, a caveated read of the fetched data: at equal-ish core counts, Venice’s per-core integer advantage looks real but modest (~20%, by AMD’s own estimate), while Vera’s bet is strong single-threaded performance and memory efficiency rather than raw throughput. The 2.24× number is best understood as a statement about core density, not per-core speed. On power efficiency, NVIDIA’s LPDDR5X approach has a credible claim that AMD’s Stream lead doesn’t address. If your buying decision hinges on any of this, wait for audited, matched-compiler results on shipping hardware.

Which should you buy?

The AMD EPYC Venice vs NVIDIA Vera decision depends less on benchmark charts than on which of these scenarios matches your build.

Choose EPYC Venice if…

  • You’re running general-purpose x86 workloads — virtualization, databases, Java, web serving, caching — where the mature x86 software ecosystem (the same foundation behind Microsoft’s Project Zenith Windows 11 dev PC) is a real asset. AMD’s rack-scale benchmarks (SPEC CPU 2017 Integer Rate, SPECjbb-derived Java, NGINX, Redis, Memcached, TPROC-C on MySQL, per Data Center News) are aimed exactly here.
  • You need maximum core density per socket or per rack, on platforms you can deploy now.
  • You’re building an AMD Helios rack or any GPU fleet where the host CPU is a separate purchase rather than part of an NVIDIA SKU.
  • Per-core integer performance matters more than memory power draw.

Choose Vera if…

  • You’re already committed to the NVIDIA rack ecosystem — Vera Rubin NVL72 or the Vera CPU Rack — and want one coherent platform with NVLink-C2C and BlueField-4 integration.
  • Your workloads are agentic AI orchestration, sandboxes, RL post-training evaluation, and other sequential, latency-bound, lightly-threaded tasks — precisely the profile NVIDIA’s developer blog designs for.
  • Memory power draw and bandwidth-per-core are decisive: LPDDR5X SOCAMM at ~1.2 TB/s and NVIDIA’s claimed 2× performance-per-watt versus x86 racks matter if you’re power-constrained. Extracting more from less bandwidth is the same principle behind running a 400B-parameter LLM locally on an iPhone 17 Pro.
  • You accept Arm software compatibility as a solved problem, and H2 2026 availability from major OEMs fits your timeline.

And a third option nobody’s white paper covers: if neither story fits, Intel’s 256-core Diamond Rapids Xeon is also targeting this battleground, per eWeek. For teams sizing the best server CPU for AI infrastructure in 2026, the honest answer is that all three platforms are still claiming, not proving.

What to watch next

  1. Official SPEC submissions. Watch spec.org’s published results for both platforms — audited, reproducible, matched-disclosure runs. ServeTheHome’s comparison against published EPYC 9755 results shows why.
  2. Independent testing on shipping hardware. hwbusters’ bottom line: the argument “will be settled by audited submissions on shipping hardware, run through matched compilers, by people who do not sell either chip.”
  3. Pricing and real deployments. Neither platform has public pricing yet, and Vera Rubin deployments (including OpenAI-backed data centers, per eWeek) are committed. TCO, not peak throughput, is how server purchases get made.

The 2.24× is a marketing number with an interesting engineering story underneath. The ~20% per-core figure is the one worth arguing about. And the rack-density argument is really a question about what kind of AI infrastructure you’re building — answer that first, and the CPU choice mostly makes itself.


References and further reading


Please let us know if you enjoyed this blog post. Share it with others to spread the knowledge! If you believe any images in this post infringe your copyright, please contact us promptly so we can remove them.



// adding consent banner