Skip to main content

General Compute Selects Cerebras to Bring Ultra-Fast Inference to Agentic Coding >>

Oct 01 2026

How AlphaSense Uses Fast Inference to Make Agentic Research Interactive

Cerebras: Alec McLean and Griffin Marge

AlphaSense: Aimee Hirsch, Chris Ackerson, Daniel Campos, Eldar Tinjic, James Chen, Matt Lapointe, Michelle Turner, Mohit Sharma, Pete Daly, Prashant Budania, and Yumo Luo

The new bar for enterprise research through architectural trade-offs in low-latency GenAI

Enterprise users no longer want a search box that returns a list of documents and leaves the synthesis to them. They want a research partner - a system that can understand a question, break it down, retrieve relevant evidence, and produce a clear, supported answer.

Consider an equity analyst preparing for an earnings call. Instead of manually stitching together transcripts, filings, and news, they ask: “What are the key risks to margin expansion this quarter?” The system pulls relevant passages, weighs conflicting signals, and returns a structured answer with supporting evidence. What used to take hours becomes something they can iterate on in minutes.

That is the ambition behind AlphaSense’s Generative Search, to make deep, evidence-based research practical in everyday workflows. But enterprise research products face a core trade-off. Higher-quality answers require more planning, retrieval, and evaluation, and each step adds latency. Users expect responses to feel interactive, even as the system does more work behind the scenes.

To address this, AlphaSense partnered with Cerebras to power the parts of its platform that rely on many fast, repeated model calls. By reducing the latency of steps like routing, query generation, and evidence evaluation, Cerebras makes it possible to include more reasoning within the same response window.

The result is a shift in architecture. Instead of collapsing everything into a single prompt or limiting the number of steps, Generative Search can compose multiple stages of planning, retrieval, and evaluation into a single responsive experience.

Business challenge: research has to feel interactive

AlphaSense’s users work in environments where research quality and speed both matter. A user may ask for a view on a company’s competitive position, a market trend, a regulatory development, or the implications of a management change. Answering accurately often requires several layered tasks:

  • Identify the relevant companies, people, markets, products, and events
  • Determine which sources are appropriate for the question
  • Retrieve passages from structured and unstructured content
  • Evaluate which snippets are relevant, redundant, stale, or contradictory
  • Synthesize the strongest evidence into a clear answer

Traditional search architectures can return documents quickly, but they leave much of this reasoning to the user. Agentic research architectures can do more of the work, but they introduce a new bottleneck. Each additional reasoning, routing, or evaluation step usually requires another model call.

That creates a product and architecture challenge. Users expect the default experience to feel fast, but the highest-quality answers often require more intermediate reasoning. The system has to decide how much work to do before the user notices the delay.

Why latency becomes an architecture constraint

As the Generative Search team expanded beyond traditional retrieval into planning, routing, evidence evaluation, and adaptive research modes, they encountered a new constraint. Each additional capability typically required another model call, and every model call consumed part of the user's response-time budget.

This is where inference latency becomes more than a performance metric. In agentic applications, it becomes an architectural design constraint. The slower each model call is, the fewer model-mediated decisions an application can afford before the experience no longer feels interactive.

For systems like Generative Search, that often forces difficult architectural trade-offs:

  • Collapse multiple responsibilities into one prompt. Rather than using specialized models for routing, retrieval planning, evidence evaluation, and synthesis, these tasks are combined into a single prompt to minimize the number of model calls.
  • Replace model-driven decisions with heuristics. Intent classification, tool selection, and workflow routing are handled with static rules or regex because another LLM call introduces too much latency.
  • Evaluate less evidence. Although retrieval may return hundreds of candidate passages, only a small subset can be inspected before synthesis, increasing the risk that relevant or conflicting information is overlooked.
  • Serialize work that could otherwise run in parallel. Independent tasks are executed sequentially to reduce the number of concurrent model calls and simplify latency management.
  • Limit workflow adaptability. Every request follows a similar execution path because dynamically escalating from a lightweight answer to a deeper research workflow introduces unpredictable delays.

For AlphaSense, reducing inference latency changed these trade-offs. Instead of collapsing the workflow into fewer, larger prompts, Generative Search could decompose research into smaller, specialized model calls for planning, retrieval, and evidence evaluation while still delivering an interactive experience.

More broadly, this illustrates an important principle for agentic AI systems - Fast inference doesn't simply make an existing workflow faster, it expands the architectural design space. As the cost of individual model calls falls, developers gain the freedom to compose richer workflows, evaluate more evidence, and adapt the execution path to each user's question without exceeding the application's response-time budget.

Solution overview

Rather than treating every inquiry the same, AlphaSense redesigned Generative Search around a planned, adaptive research workflow. Instead of relying on a single prompt to perform retrieval, reasoning, and synthesis, the system decomposes research into a series of specialized decisions that can be adjusted based on the complexity of the user's request.

The first decision is determining how much research the question actually requires. A straightforward question may only need a lightweight retrieval and synthesis path, while a more complex request (such as comparing competing analyst opinions or evaluating a company's strategic position) benefits from additional planning, retrieval, and evidence evaluation before generating an answer.

To support these different workloads, Generative Search provides three research modes:

  • Auto: Optimized for fast, interactive responses by minimizing unnecessary planning while still producing useful, evidence-based answers.
  • Think Longer: Allocates additional time for planning and evidence evaluation when a question benefits from deeper analysis.
  • Deep Research: Executes a more comprehensive workflow with broader retrieval, more intermediate reasoning, and deeper synthesis across a larger set of trusted sources.

A key architectural decision is that these are not separate products or independent pipelines. They share a common research architecture that dynamically adjusts the amount of planning, retrieval, tool use, evidence evaluation, and synthesis performed based on the user's intent and latency expectations.

Cerebras accelerates the parts of this workflow that depend on many fast, repeated model calls. These calls often happen before the user ever sees an answer, classifying intent, selecting research modes, generating retrieval queries, evaluating candidate evidence, and determining what information should ultimately reach the synthesis model. While they may be invisible to the user, they have an outsized impact on both the quality of the final answer and how responsive the overall experience feels.

Illustrative research workflow

This is not intended to describe AlphaSense’s exact implementation. It illustrates the architectural pattern; use fast inference for repeated, latency-sensitive decisions in the orchestration and evidence layers, then reserve deeper reasoning for planning-heavy or synthesis-heavy steps where it creates the most user-visible value.

How latency shapes application architecture

As the AlphaSense Generative Search team evaluated different architectures, it wasn’t simply optimizing for faster model responses. It was deciding how much reasoning the system could afford to perform before users perceived a delay.

Every interactive AI application operates within a finite response-time budget. For Generative Search, that budget must accommodate retrieval, orchestration, evidence evaluation, synthesis, networking, and every intermediate model call.

The faster each individual model call becomes, the more architectural options fit within that same user experience.

For AlphaSense, this meant the architecture no longer had to treat every model call as an expensive resource. Instead, Generative Search could use fast, specialized model calls throughout the workflow, for intent classification, query generation, evidence scoring, and context selection, before reserving deeper reasoning for the final synthesis stage.

The broader lesson extends well beyond enterprise research. The fundamental trade-off isn't latency versus cost; it's architectural flexibility. Lower inference latency expands the set of application designs that are practical in production. Instead of relying on a single, monolithic prompt to perform every task, developers can compose smaller, purpose-built model calls that are easier to optimize, scale, and evolve independently while maintaining an interactive user experience.

Where Cerebras helps

Within the Generative Search architecture, AlphaSense identified two categories of model calls where latency had an outsized impact on the overall user experience. These became natural candidates for Cerebras-backed inference.

Fast routing and tool selection

Routing and tool-selection calls shape the entire workflow. They help determine whether a request should stay in a lightweight path or move into a more involved research mode, which tools should be called, and how downstream work should be allocated. When these decisions happen quickly, the product can be more adaptive without feeling slower.

High-volume snippet evaluation

Research workflows can retrieve far more evidence than they can ultimately show to the user. The value comes from inspecting that evidence quickly enough to decide what matters. Faster inference makes it practical to evaluate more candidates, apply more nuanced relevance checks, and pass stronger context into synthesis.

Cerebras enables AlphaSense to focus inference spend on the latency-sensitive workflows where speed creates the most customer value, making these features possible:

  • Partnering with Cerebras enables AlphaSense to deliver initial results in as little as 10 to 20 seconds. By accelerating critical behind-the-scenes tasks, like planning and refining research, Cerebras keeps complex workflows responsive; with other inference providers, the same processes can take up to twice as long.
  • AlphaSense can now offer an "ultra-fast" search mode alongside standard options. This gives users the flexibility to prioritize speed or depth, a capability made possible by the efficiency and speed of Cerebras's technology.
  • For complex research tasks, the system can repeatedly check and refine its findings to ensure high accuracy. This iterative process, which involves multiple rounds of analysis, relies on the rapid performance provided by Cerebras to maintain a smooth and efficient user experience.

Results and impact

Faster inference affects three dimensions: responsiveness, throughput, and product ambition.

  • Responsiveness: Auto mode can make routing and tool-selection decisions quickly, helping preserve the feeling that the product is reacting immediately to the user’s question. Using Cerebras reduces p90 time to first token by 88% from 19.5 seconds to 2.3 seconds, serving the same model 8.5x faster.
  • Throughput: Snippet evaluation is naturally high-volume. Cerebras enables reviewing 3x the volume of evidence without increasing latency, giving the model a broader evidence base from which to produce a more accurate, well-supported answer.
  • Product ambition: Lower latency gives AlphaSense more room to design around adaptive routing, scalable evidence evaluation, and deeper synthesis across trusted sources. AlphaSense has also observed that faster responses are associated with lower abandonment and higher usage, helping drive product adoption and retention.

Lessons learned

The AlphaSense redesign highlights several lessons for teams building agentic enterprise AI systems.

Design the latency budget before the prompt: The call graph determines what the product can afford to do. Before optimizing prompts, teams should understand the response-time budget, the number of model calls required, which calls can run in parallel, and which calls are critical to answer quality.

Separate control-plane, evidence-plane, and synthesis work: Not every model call has the same purpose. Routing and planning calls should be fast and composable. Evidence evaluation calls should scale across many candidates. Synthesis calls should focus on producing the final user-facing answer from the best available context.

Use fast models where repeated decisions matter: Fast inference is most valuable where the system needs many decisions quickly such as classifying intent, selecting tools, generating retrieval queries, scoring snippets, and filtering context. These are the invisible steps users may never see, but they strongly influence answer quality.

Give product modes a shared architecture: Auto, Think Longer, and Deep Research should not require entirely separate stacks. A shared architecture that can dial depth up or down makes it easier to give users control over the trade-off between speed and completeness.

What’s next

AlphaSense's redesign lays the foundation for increasingly capable research workflows. Because planning, retrieval, evidence evaluation, and synthesis are composed from a common architecture, the platform can continue to add richer reasoning and agent behaviors without fundamentally changing how requests are processed.

As inference performance continues to improve, more of these intermediate decisions can happen within the same interactive response window. Rather than choosing between fast answers and thorough research, future iterations of Generative Search can continue to expand the amount of planning, evidence evaluation, and adaptive execution performed before the user perceives additional latency.

This architectural direction is reflected in AlphaSense's broader AI strategy. The recently announced SuperAnalyst extends these concepts beyond conversational search to produce finished, decision-grade work products grounded in trusted, curated enterprise content. It represents a natural evolution of the same design philosophy, using fast, composable AI workflows to automate increasingly sophisticated research tasks.

Conclusion

The AlphaSense and Cerebras collaboration is not just about making existing prompts run faster. It is about using faster inference to make a more capable Generative Search architecture possible.

By reducing the latency cost of routing, tool selection, and snippet evaluation, Cerebras helps AlphaSense preserve the speed users expect while expanding the depth of research the system can perform. For enterprise AI builders, the broader lesson is clear: as applications become more agentic, inference speed becomes a core design constraint.

Faster inference allows teams to use more model calls in the places where they matter most, evaluate more evidence, and build workflows that feel both intelligent and responsive.

See what this architecture makes possible for your own research. Sign up for a free trial of AlphaSense to put SuperAnalyst to work on your team's questions.

If you want to work on cutting-edge engineering challenges like these, AlphaSense is hiring. Explore open roles at AlphaSense Careers.

Cerebras is also hiring. Explore open roles at Cerebras Careers.

Performance comparisons are based on third-party benchmarking or internal testing. Observed inference speed improvements versus GPU-based systems may vary depending on workload, configuration, date and models being tested.

1237 E. Arques Ave
 Sunnyvale, CA 94085

© 2026 Cerebras.
All rights reserved.