Skip to main content

Introducing CS-4: The Fastest AI Accelerator in the Industry Learn more >>

Aug 25 2026

Ultrafast Frontier Inference: Cerebras Deep Dive at Hot Chips 2026

Last week at Supernova 2026, Cerebras introduced CS-4: the fastest AI accelerator in the industry and the first system built on the new Cerebras Nexus rack-scale platform. We shared major advances in token speed, throughput, efficiency, and scalability that make CS-4 the new foundation for frontier AI.

Today at the HOT CHIPS conference in Palo Alto, we shared how CS-4 achieves those gains and previewed the technology roadmap behind CS-5 and CS-6.

CS-4 is the first system built on Nexus, our reusable rack-scale platform. Nexus supports three Wafer-Scale Engines, each housed in a modular compute backpack at the rear of the rack. Its modular power, cooling, and I/O architecture allows us to improve each part of the system independently, giving us a path to double token-generation speed year over year for the next several years while dramatically improving throughput and efficiency.

Each compute backpack contains one WSE and its dedicated power, cooling, and I/O. The backpack reinvents the server as a modular unit designed for faster datacenter deployment: the front power rack is installed in the datacenter, and the compute backpacks can then be dropped in on-site.

More CS-4 system architecture details unveiled

With advances in CS-4’s power delivery, cooling, and rack integration, the WSE delivers ultrafast AI performance reliably at datacenter scale.

Power delivery built around the wafer

Moving high current across a circuit board creates resistance and wastes power as heat. Conventional GPU systems place their power converters about 50 millimeters from the silicon, requiring current to travel through multiple layers of copper before reaching the processor.

CS-4 places AC/DC converters about 0.5 millimeters from the wafer (100x less than GPUs), without a printed circuit board in the final power-delivery path. This allows CS-4 to deliver nearly twice as much power at almost the same voltage, with little additional resistive loss in the delivery path.

Cooling is built into each compute backpack

Each CS-4 compute backpack contains its own water-conditioning system, making installation and maintenance faster in hyperscale datacenters. An energy meter tracks flow and inlet and outlet temperatures, while an actuator adjusts the flow to the cold plates. Dry quick-disconnect valves allow technicians to connect or remove a backpack without draining the system.

Leak and condensation sensors can automatically place the backpack in a safe state and cut power to its power-supply modules. Water and compute remain at the rear of the rack, separated from the high-voltage AC equipment in the front. Supply and return manifolds run along the sides, while protected conduits keep fiber connections out of the service path. This allows a compute backpack to be replaced without disturbing the rack’s shared water or network infrastructure.

Power is centralized at the front of the rack

The front of the CS-4 rack contains the shared power infrastructure for all three compute backpacks. The design supports several power and redundancy configurations, making it easier to integrate into different hyperscale datacenter environments.

Each compute backpack can draw from up to 30 dedicated, air-cooled AC/DC power-supply modules. The modules accept up to 277 volts AC and deliver 54.5 volts DC. They support 5+1, 4+1, 3+1, and 4+2 feed-redundancy configurations, with each module protected by its own 30-amp circuit breaker.

Up to six hard-wired AC feeds enter from the top of the rack. An integrated interconnect distributes power down both sides to the breakers and power supplies, eliminating additional rack-internal power cabling during field installation. The feeds share the load for each backpack, and the AC interconnect is fully phase balanced.

Nexus is designed to support multiple generations of Cerebras systems, beginning with CS-4. Its modular power, cooling, and I/O architecture allows each part of the platform to advance independently, so new technologies can be developed and deployed faster. Nexus was also co-designed to support our next-generation WSE, which will debut in CS-5.

CS-5: the next speed standard

CS-5, targeted for 2027, is designed to generate up to 10,000 output tokens per second per user on leading open-source models, including Gemma 4 31B and gpt-oss-120b. For agentic workloads, which often require many sequential model calls, faster output at each step can sharply reduce total task-completion time.

For the largest frontier models, including multi-trillion-parameter models such as Kimi and GPT-5.6 Sol, CS-5 targets up to 5,000 output tokens per second per user and 3 million tokens per second per megawatt. The same architecture is designed to support models with more than 50 trillion parameters while maintaining interactive speeds.

CS-6: wafer scale goes 3D

Building at wafer scale required rethinking packaging, power delivery, cooling, interconnect, and system design. Over the past decade, Cerebras has solved each of these challenges, turning the world’s first and only wafer-scale processor from a radical idea into a product shipping at scale.

That experience has brought us to the next great frontier in computing.

A wafer-scale processor already fills the largest practical area available in two dimensions. Adding substantially more memory means building upward while preserving the data locality that makes wafer scale fast.

In 2024, we began turning that vision into CS-6. By integrating wafer-scale SRAM and compute with 3D-stacked DRAM through ultra-high-bandwidth connections, CS-6 is designed to dramatically expand memory capacity without sacrificing the locality that makes wafer scale fast.

Wafer-scale SRAM already scales economically to accelerate even the largest models, including GPT-5.6 Sol and beyond. With tightly integrated DRAM, more of each model can reside on each system, reducing the infrastructure required to run it. The result is ultrafast inference in an order-of-magnitude smaller system footprint, bringing ultrafast AI speeds to everyone.

The architecture is the advantage

AI performance depends heavily on how quickly data moves between compute cores. Conventional systems scale by connecting many GPUs and coordinating them through external links and switches. Every transfer requires data to leave one chip, cross the system, and arrive at the next chip in time for the following operation.

Serialization, synchronization, switching, and software coordination all consume power and add latency. At small batch sizes, this communication overhead can take as long as the computation itself.

The scale is visible in today’s rack designs. NVIDIA specifies 260 terabytes per second of rack-level NVLink bandwidth for an entire 72-GPU Rubin rack, whose NVLink spine comprises roughly 5,000 internal cables. A single Cerebras WSE-3T provides 53.5 petabytes per second of aggregate on-wafer fabric bandwidth, more than 200 times the NVL72 rack’s scale-up bandwidth.

In a single NVIDIA Rubin NVL72 rack, there are thousands of cables connecting all the GPUs together. NVIDIA markets this as a good thing, claiming these cables have more scale-up or fabric memory bandwidth (260 terabytes per second) than the entire global internet. But all this cabling has real costs, in terms of performance, power, cost, and reliability.

Because communication between cores happens on the wafer, Cerebras does not need thousands of cables to make many separate processors behave like one. This reduces communication latency, power consumption, hardware cost, and potential points of failure.

When a model spans multiple Cerebras systems, execution is pipelined to keep high-volume tensor and expert communication within each wafer. Only lower-volume data, primarily activations, moves between wafers. This becomes increasingly important as models grow, mixture-of-experts routing expands, context windows lengthen, and batch sizes shrink.

The wafer-scale advantage compounds

The advantage of Cerebras wafer-scale architecture compounds across generations.

CS-4 is the first system built on Nexus.

CS-5 will pair Nexus with our next-generation WSE to set a new standard for token-generation speed.

CS-6 will push the frontier of 3D integration by stacking DRAM at wafer-scale to deliver the next step-change in performance and efficiency.

Conventional architectures scale by adding more processors, switches, and links. Cerebras scales by advancing compute, memory, power, cooling, and I/O together on one co-designed, integrated platform. That is the architecture and the roadmap for frontier AI.

Forward-Looking Statements

This blog contains "forward-looking statements" within the meaning of applicable securities laws. All statements other than statements of historical fact and any assumptions relating to such statements could be deemed to be forward-looking. The words "may," "will," "shall," "should," "expects," "plans," "anticipates," "could," "intends," "target," "projects," "contemplates," "believes," "estimates," "predicts," "potential," "objective," or "continue," or the negative of these words or other similar terms or expressions that concern our expectations, strategy, plans, or intentions are intended to identify forward-looking statements, although not all forward-looking statements contain these identifying words. These forward-looking statements are subject to a number of risks and uncertainties, many of which involve factors or circumstances that are beyond Cerebras’ control. These risks and uncertainties include, but are not limited to: Cerebras’ ability to sustain and manage its growth, access borrowings and other sources of capital on acceptable terms, and deploy available capital to support growth; its history of net losses and ability to achieve and maintain profitability; its limited operating history at its current scale and ability to accurately forecast revenue and appropriately budget and manage expenses; its dependence on a limited number of significant customers, including OpenAI, Group 42 Holding Ltd, Mohamed bin Zayed University of Artificial Intelligence, and AWS, and the potential impact of any reduction in demand from, material adverse development in its relationships with, or failure to meet its obligations to, such customers, including under its Master Relationship Agreement with OpenAI; the timing, execution and expected benefits of its strategic customer, partner and financing arrangements; its historical reliance on sales of hardware systems and the early-stage, rapidly evolving market for its cloud-based offerings and AI infrastructure; its ability to secure sufficient data center capacity and capital to support its cloud-based offerings; its ability to launch new offerings and add new product capabilities; and its ability to compete effectively in the rapidly evolving and competitive market for AI computing solutions.

Cerebras’ actual results could differ materially from those stated or implied in forward-looking statements due to a number of factors. Accordingly, undue reliance should not be placed on such statements. These forward-looking statements are made as of the date they were first issued and are based on information available to Cerebras together with Cerebras’ expectations, estimates, forecasts, projections, beliefs, and assumptions as of such date. These forward-looking statements should not be relied upon as representing Cerebras’ views as of any date subsequent to the date of this blog. Past performance is not necessarily indicative of future results. Cerebras undertakes no intention or obligation to update or revise any forward-looking statements, whether as a result of new information, future events, or otherwise, except as required by law.

Further information on potential risks that could affect actual results is included in Cerebras’ most recent filings with the SEC, including in Cerebras’ most recent Quarterly Report on Form 10-Q, copies of which may be obtained by visiting Cerebras’ Investor Relations website at investors.cerebras.ai or the SEC’s website at www.sec.gov.

Performance comparisons are based on third-party benchmarking or internal testing. Observed inference speed improvements versus GPU-based systems may vary depending on workload, configuration, date and models being tested.

1237 E. Arques Ave
 Sunnyvale, CA 94085

© 2026 Cerebras.
All rights reserved.