Hot Chips 2026: NVIDIA Vera and Why CPU Performance Is More Than Core Count
News – 01/10/26
At the 2026 Hot Chips Symposium, NVIDIA presented the Vera CPU as part of its broader computing platform for agentic AI.
One aspect of the presentation particularly stood out to us: NVIDIA's emphasis on performance per core rather than core count.
It highlights a broader CPU architecture question.
How much useful performance can additional cores actually deliver as processors scale?
More cores do not automatically lead to proportional performance.
Increasing CPU core counts has been one of the industry's primary approaches to increasing compute capacity as improvements in single-core performance have become increasingly difficult.
But adding cores does not mean application performance scales proportionally.
As more cores work on a parallel workload, they need to access data, communicate, coordinate work, and synchronize execution. Those operations introduce overhead.
Memory access can become a bottleneck, especially when a growing number of cores need to access the same data. Synchronization can leave cores waiting for one another. Communication between cores can increasingly consume execution time.
Eventually, the question becomes less about how many cores a processor contains and more about how effectively the architecture can turn those cores into useful performance.
NVIDIA's focus on per-core performance with Vera reflects one approach to this broader challenge. Vera combines 88 custom Olympus cores with NVIDIA Spatial Multithreading, high-bandwidth memory, and a scalable coherency fabric designed to sustain performance across heavily utilized AI systems.
This reflects a broader theme emerging around Hot Chips. Industry commentary has increasingly focused on system architecture and the importance of addressing data movement and memory bottlenecks rather than looking only at raw compute resources. Computerworld's Hot Chips analysis, for example, highlighted analyst observations that system architectures are changing as designers work to reduce data-movement bottlenecks and bring computation closer to memory.
We approach another part of this CPU performance challenge with a complementary acceleration architecture.
Rethinking parallel CPU performance.
We believe that increasing conventional CPU core counts alone does not resolve architectural bottlenecks limiting parallel performance.
Flow PPU is designed as a general-purpose parallel-processing architecture that works alongside the CPU.
The division of work is important.
The CPU continues to handle sequential execution, control, and existing software. Suitable parallel sections can instead be executed using Flow PPU, which is designed specifically for scalable general-purpose parallel computation.
Rather than replacing the CPU or changing its fundamental role, Flow PPU adds closely coupled parallel acceleration capabilities alongside conventional CPU cores for workloads where scalable general-purpose parallel processing can add value.
Early Flow PPU performance results demonstrate the potential of this approach. In our recent RISC-V performance model, an entry-level 16-core Flow PPU paired with a single RISC-V CPU core delivered an average 18.6x speedup compared with a conventional 4-core RISC-V CPU across the tested parallel workloads.
Why this matters for AI.
CPU scaling discussions are particularly relevant as AI infrastructure evolves.
GPUs remain critical to modern AI. They excel at the specialized, massively parallel computation behind training and many inference workloads.
But AI systems do not run on GPUs alone.
Agentic AI introduces significant CPU-side work around model execution, including orchestration, tool calls, code execution, data processing, and other general-purpose workloads. NVIDIA itself describes Vera as the CPU responsible for much of this work, while its GPUs handle accelerated AI computation.
This is why we do not see the future as CPU vs. GPU.
We see an increasingly heterogeneous computing environment in which different architectures perform the work they are best suited to execute.
CPU for general-purpose and sequential computation. GPU for specialized accelerated computation. Flow PPU for scalable general-purpose parallel computation alongside the CPU.
These architectures are complementary rather than competing. Flow relies on the CPU remaining at the center of general-purpose computing and adds parallel acceleration capabilities for workloads where conventional CPU execution can benefit from additional architectural support.
Beyond the core-count.
NVIDIA's Vera presentation at the Hot Chips Symposium raises a crucial point for the broader CPU industry: core count alone tells us relatively little about how effectively a processor will execute a real workload.
As AI and other workloads become increasingly parallel and computationally demanding, processor performance will depend on more than how many cores can fit on a chip.
It will depend on how effectively those computing resources can work together.
NVIDIA is approaching that challenge with a CPU designed around strong per-core NVIDIA is strengthening CPU performance through Vera's focus on high performance per core and an architecture designed for heavily utilized AI systems. Flow complements advances at the CPU-core level by adding closely coupled parallel acceleration alongside the CPU for suitable general-purpose parallel workloads.
The broader direction is CPU + the right architecture for each workload. A strong CPU remains fundamental, while complementary acceleration architectures can extend the range of workloads it can execute efficiently.
That leads to an increasingly important question for the industry:
How do we get more useful performance from the CPU without relying only on additional conventional cores?