Hot Chips 2026: Fujitsu MONAKA and Why CPU Throughput Still Matters for AI and HPC

At Hot Chips 2026, Fujitsu presented FUJITSU-MONAKA, its next-generation Arm-based CPU designed for AI and high-performance computing.

One of the most interesting messages in the presentation was also one of the simplest: as AI and HPC systems evolve, CPUs remain fundamental.

Fujitsu highlighted how agentic AI is increasing CPU usage for orchestration, AI pre- and post-processing such as RAG, and general-purpose workloads such as databases and code execution. In HPC, meanwhile, CPUs continue to play an important role in high-throughput and high-precision computing.

The challenge, then, isn't simply whether CPUs remain relevant. It's how CPU architectures continue delivering more useful throughput as workloads become increasingly parallel and data-intensive.

Designing for throughput

MONAKA approaches this challenge at several levels.

Its Arm-based architecture combines 144 cores per socket with SVE2 execution capabilities and a 3D chiplet architecture. Fujitsu has also designed its core microarchitecture specifically around AI, HPC and general-purpose workloads.

The presentation pays particular attention to throughput. Fujitsu describes techniques for sustaining 256-bit SIMD load throughput and increasing gather throughput for HPC workloads.

But compute resources alone aren't the whole performance equation.

MONAKA's memory architecture includes configurable NUMA regions and memory localization designed to balance latency and throughput depending on software characteristics.

That reflects a broader challenge in modern processor design: adding compute capability doesn't automatically produce proportional application performance if moving, accessing and coordinating data becomes the bottleneck.

Industry commentary around Hot Chips echoed this point. Computerworld's analysis highlighted observations from industry analysts that processor and system architectures are changing as designers address data-movement bottlenecks and move computation closer to memory.

A complementary approach to CPU performance

This is particularly relevant from Flow's perspective.

Traditional multicore CPUs can struggle to scale parallel workloads efficiently as memory access, synchronization and cache-coherence overhead increase. Our Parallel Processing Unit® (Flow PPU) approaches these bottlenecks by adding a closely coupled architecture for scalable general-purpose parallel computation alongside the CPU.

The CPU handles sequential execution, control and existing software, while the PPU executes suitable parallel sections of the workload.

The goal isn't to replace the CPU. It's to expand what CPU-powered systems can efficiently execute.

This makes the approach particularly relevant to server, cloud and HPC environments running scientific computing, data processing, preprocessing, inference-support and other parallel CPU workloads.

Architecture for the workload

An interesting moment in the MONAKA Q&A reinforced this broader theme.

Fujitsu was asked why MONAKA uses a 256-bit SIMD datapath after the 512-bit design used in A64FX. The answer came down to the intended workload: MONAKA targets data centers, and Fujitsu wanted to reduce core size while achieving better cost/performance for general-purpose workloads.

That's an important architectural principle. The objective isn't necessarily to maximize every individual specification. It's to design the architecture around the work it needs to perform.

MONAKA does that through innovations across the CPU core, memory hierarchy, chiplets and software ecosystem.

We approach another part of the same broader CPU-performance challenge with a complementary architecture: adding scalable parallel acceleration alongside conventional CPU cores.

As AI and HPC workloads continue to evolve, CPUs aren't disappearing from the picture. If anything, they're being asked to do more.

The next generation of CPU performance will depend not only on stronger cores, but on architectures that allow CPU-powered systems to scale with the workloads around them.

Explore Flow's architecture and latest performance results.→

Contact usX