Flow PPU Performance Results.

Speed backed by data.

Benchmark results demonstrating how Flow Computing®'s Parallel Processing Unit® (Flow PPU) delivers significant performance gains beyond conventional multicore CPU scaling.

Flow PPU commercial development.

Flow PPU delivers scalable performance beyond the limits of conventional multicore CPUs. From compute-heavy tasks to high-throughput workloads, our benchmark results demonstrate significant improvements in throughput while reducing synchronization overhead and latency.

Flow is progressing toward a commercial implementation of Flow PPU. To validate the commercial architecture, we benchmarked a 16-core Flow PPU paired with a single RISC-V CPU core against a conventional 4-core RISC-V processor across representative parallel workloads.

Benchmark results demonstrate consistent improvements as the commercial implementation matures. The current entry-level Flow PPU already delivers significant performance gains compared with conventional CPU execution while maintaining the simplicity and programmability of a CPU-based architecture.

The ongoing development of the commercial Flow PPU supports the transition toward production-ready implementations for AI, HPC, cloud infrastructure, embedded systems, and edge computing.

Outstanding performance gains for existing code.

Flow PPU supports existing software, and can even accelerate fully sequential code that it can parallelize. Using the TSVC (Test Suite for Vectorizing Compilers) benchmark suite, Flow evaluated how existing sequential code benefits from automatic parallelization and execution on Flow PPU.

TSVC is a widely used benchmark suite consisting of 151 tests that evaluate the parallelization of sequential code. These workloads represent common building blocks used across AI, machine learning, simulation, rendering, and sensor/signal-processing applications.

The top benchmark results demonstrate speedups of up to 55x compared with sequential execution. These results highlight Flow PPU's ability to accelerate existing sequential code bases while reducing the effort required to utilize the benefits of accelerated performance from parallel computing.

Speedup potential of high-end Flow PPU against latest processors.

Flow demonstrated how its 256-core PPU accelerates real workloads across AI, HPC, and hyperscale use cases using proof-of-concept against state-of-the-art datacenter CPUs in cloud and consumer CPUs in laptops.

The results show a near-linear scaling of performance as tasks are distributed across PPU units, eliminating long-standing CPU bottlenecks caused by synchronization, cache coherence, and thread-management overhead.

Breaking through the CPU bottleneck.

Modern workloads have outgrown the limits of sequential and multi-core CPU scaling.

Flow’s architecture introduces a new execution model that brings true parallelism inside the CPU, achieving high throughput without relying on external accelerators.

The result is a scalable, efficient foundation for next-generation computing, from AI and cloud infrastructure to embedded and edge systems.

Request access to the complete results and poster.

Fill out the form below to access our benchmark set, technical notes, and the official performance poster. 

Thank you for being awesome!

We appreciate you contacting Flow. Our team will get in touch with you soon! Have a great day!

Close

Contact usX