Aiming at NVIDIA Corporation! Cerebras (CBRS.US) has released the CS-4 system, the "fastest AI accelerator," claiming an inference speed 30 times that of GPU solutions.

date
18:50 19/08/2026
avatar
GMT Eight
Cerebras has launched the CS-4 system, claiming it to be the fastest AI accelerator in the industry.
AI chip company Cerebras Systems (CBRS.US) officially launched its next-generation rack-scale AI system, Cerebras CS-4, at the SUPERNOVA event held in San Francisco, claiming it to be "the fastest AI accelerator in the industry." This system targets the AI inference track, directly competing with NVIDIA Corporation (NVDA.US) GPU server products. CS-4 is expected to begin shipping this quarter (Q3 2026), with a limited number of customers currently testing samples. Technical Specifications: Three Wafer-Scale Processors Deliver 750 PFLOPs of Computing Power CS-4 is the first product of Cerebras's new generation Nexus rack-scale platform architecture, powered by three entirely new Wafer Scale Engine 3 Turbo (WSE-3T) wafer-scale processors. WSE-3T is the largest AI processor to date, with each chip integrating 40 trillion transistors and 900,000 AI-optimized cores, covering 46,225 square millimeters of silicon area and featuring 44GB of SRAM. The CS-4 system offers 750 PFLOPs of AI computing power, 129.6 PB per second of memory bandwidth, and 7.2 Tbps of I/O throughput. Compared to the previous generation CS-3, CS-4 achieves a twofold speed improvement and a tenfold increase in throughput per watt, significantly enhancing the economic efficiency of data centers. Compared to the previous generation WSE-3, WSE-3T doubles the AI computing power per wafer from 125 PFLOPS to 250 PFLOPS, while memory bandwidth increases from 21.6 PB/s to 43.2 PB/s. The power consumption per chip also jumps from 15 kilowatts to about 33 kilowatts. In terms of architectural design, CS-4 adopts a modular Nexus platform structure, separating computing, power, and I/O into independent components. The power conversion module has been reduced in distance from about 50 millimeters from the processor on traditional GPU boards to about 0.5 millimeters, nearly eliminating board-level power loss. The pluggable "backpack" design integrates power conversion, liquid cooling, and control electronics into one unit, reducing deployment time from days to hours. Performance Breakthrough: Over 4,400 Tokens per User per Second in GPT-OSS-120B Tests Cerebras claims that CS-4 has set a new industry benchmark for inference speed. In direct comparison tests on the GPT-OSS-120B model, CS-4 generates over 4,400 tokens per user per second, reaching speeds up to 30 times those of GPU solutions. The system can support models with over 50 trillion parameters. Sean Lie, Chief Technology Officer and co-founder of Cerebras, stated, Being 30 times faster isnt just about making responses feel smoother; it provides AI systems with over an order of magnitude more space to perform reasoning, validation, or tool calls within the same actual time. Cerebras CEO Andrew Feldman pointed out, Historically, fast inference meant using smaller, weaker models. CS-4 achieves leading speeds on the largest frontline models, fundamentally transforming this paradigm. In terms of energy efficiency, CS-4's throughput per watt is up to ten times that of CS-3, significantly improving the economic viability of data centers. Cerebras CEO Andrew Feldman noted that because fast tokens are more valuable than slow ones, CS-4 can deliver higher-quality tokens and a greater total number of tokens within a given power budget, fundamentally reshaping the paradigm. CS-4 also features a design that reduces the number of components by 50%, making it easier and faster for customers to deploy. In the system architecture, the new Nexus platform modularizes computing, power, and I/Opower conversion has been shortened from about 50 millimeters from the processor on traditional GPU motherboards to about 0.5 millimeters, greatly reducing board-level power consumption and supporting higher operating frequencies. The pluggable "backpack" design integrates power conversion, liquid cooling, and control electronics into a single unit, reducing deployment time from days to hours. Notably, Cerebras's wafer-scale system uses Static Random Access Memory (SRAM) rather than the commonly used DRAM. SRAM is much faster than DRAM but is larger in size and more expensiveCerebras's platter-sized wafers provide the physical space required to utilize SRAM. Additionally, the single-chip design allows for significantly shorter data transmission distances compared to multi-chip systems from NVIDIA Corporation or AMD. Ecosystem Partnerships: Joint Efforts with AMD Helios, Targeting 5x Efficiency Improvement The programmable I/O subsystem of CS-4 supports Ethernet-based RoCE v2 RDMA standards, enabling integration with heterogeneous systems. Cerebras has partnered with AMD (AMD.US) on this architecture, pairing AMD's Helios rack-level solution with the WSE, generating five times the number of tokens per watt per second compared to purely Cerebras configurations. AWS Trainium is also listed as a partner. Dylan Patel, founder and CEO of SemiAnalysis, commented that the enhancements in deployability and networking capabilities of CS-4 allow it to scale to larger models, creating a "large-scale token factory." Market Background: Stock Price Down 35% from Recent Highs, CS-4 as a Key Turning Point Cerebras went public on Nasdaq on May 14, 2026, at a price of $185 per share, opening at $350 on the first day and closing at $311.07, a gain of 68%. However, the stock price has since declined continuously, reporting about $220 on Tuesday, down approximately 35% from its peak. The companys Q2 financial report showed mixed results: core revenue was $209.9 million, an increase of 103% year-over-year; however, earnings per share were a loss of $2.98, compared to a profit of $1.91 for the same period last year. The company's guidance for tripling core revenue by 2027 and the better-than-expected Q3 outlook failed to fully impress investors. Nonetheless, the company's remaining performance obligations stand at $25.4 billion, providing strong visibility for future revenues. Industry Significance: AI Inference Market Landscape Set to Transform The launch of CS-4 comes at a critical juncture of explosive growth in AI inference demand. As large models transition from training to large-scale deployment, inference efficiency is becoming a core competitive dimension of AI infrastructure. If Cerebras's claimed 30-fold speed advantage can be validated in independent benchmark tests, it may change the procurement decisions of hyperscale data center operators. The release of Cerebras CS-4 signifies that the competition in AI chips is accelerating from model training towards the inference market. Whether this system can genuinely disrupt NVIDIA Corporation's dominance in the AI inference field will depend on whether its performance claims can be validated in independent benchmark testsand whether CS-4 can be delivered at a sufficient scale quickly enough to meet the rapidly urgent demand for high-speed token generation from hyperscale data center operators. Cerebras's cloud inference business achieved a year-over-year growth of 287% in Q2, reaching $127.7 million in core cloud revenue. The launch of CS-4 marks a new phase for Cerebras as it transitions from hardware sales to a business model centered on leasing inference capabilities.