Skip to main content

Introducing CS-4: The Fastest AI Accelerator in the Industry Learn more >>

now upgraded with glm 4.7

Cerebras Code: Fast AI Coding

Stop waiting on your model. Cerebras runs GLM 4.7 — the best-in-class model for code generation, at 1,000 tokens+ per second — so you can stay in flow.

State of the Art Frontier Model

GLM-4.7 is one of the world’s top open coding models: #1 for tool calling on the Berkeley Function Calling Leaderboard and on par with Sonnet 4.5 in web-dev performance.

Bring Your Own AI Code Editor

Use Cerebras Code Pro with any AI-friendly editor or agent that accepts your API key. Works out of the box with Cline, OpenCode, Crush, and more. Integrate instantly and code without switching tools.

Free Trial

Limited Free Credits

GLM 4.7 access with limited tokens and requests.
Great for trying out Cerebras inference or building a small demo in your favorite AI Code Editor.

Pro

$50

GLM4.7 access with fast, high-context completions. Send up to 24million tokens per day, enough for 3–4 hours of uninterrupted vibe coding.

Ideal for indie devs, simple agentic workflows, and weekend projects.

Max

$200

GLM4.7 access for heavy coding workflows. Send up to 120 million tokens/day.

Ideal for full-time development, IDE integrations, code refactoring, and multi-agent systems.

Follow us on X and subscribe to our newsletter to be the first to find out when new plans drop.​

Performance comparisons are based on third-party benchmarking or internal testing. Observed inference speed improvements versus GPU-based systems may vary depending on workload, configuration, date and models being tested.

1237 E. Arques Ave
 Sunnyvale, CA 94085

© 2026 Cerebras.
All rights reserved.