TechPro
Harju maakond, Tallinn, Kesklinna linnaosa, Tartu mnt 25-46, 10117 smarttek.ou@gmail.com

CoreWeave is making major moves to cement its position in high-performance cloud computing, unveiling full support for Nvidia’s next-generation Vera Rubin NVL72 rack-scale architecture alongside a brand-new development environment called Forge.

The announcements, made at the company’s Fully Connected summit in San Francisco, highlight a clear strategic shift: optimizing infrastructure for “agentic AI”—autonomous systems that execute complex code, multi-step workflows, and continuous tool usage.

Early Benchmarks: A Massive Leap for Autonomous Agents

Cognition—the team behind the AI software engineer Devin—has signed on as the flagship client for CoreWeave’s Vera Rubin cluster. After deploying the hardware in early September, Cognition ran the platform’s first customer-led inference tests, reporting dramatic performance leaps compared to Nvidia’s current GB200 NVL72 setup:

  • 4.8x boost in total token throughput on SWE-2 software engineering inference.

  • 3.8x gain in output-token throughput during reinforcement learning tasks.

According to CoreWeave’s EVP of Product and Engineering, Chen Goldberg, these gains allow developers to scale autonomous tasks drastically. The increased efficiency translates directly into more simultaneous user sessions per GPU, accelerated research timelines, and reduced operational costs per session without sacrificing output speed.

Goldberg emphasized that CoreWeave’s underlying architecture allows enterprise clients to roll out new rack-scale hardware within days, eliminating the need to re-engineer core infrastructure with every hardware generation.

Hardware Upgrades: CPU Muscle for Complex Workflows

Beyond raw GPU power, CoreWeave is leaning heavily into Nvidia’s new Vera CPU—a chip tailored specifically to handle the secondary execution demands of modern AI systems.

While GPUs manage heavy model training and inference, CPUs handle critical auxiliary tasks, including:

  • Isolated execution sandboxes and code runs

  • Data processing pipelines

  • Multi-step tool calls and API integrations

  • Reinforcement learning loops

To support these workloads, CoreWeave’s initial Vera setup packs 128 Vera CPUs (11,264 individual cores) into a single rack-scale environment, backed by Nvidia BlueField-4 DPUs and Spectrum-X Ethernet. The company estimates this single-rack configuration can run over 11,000 isolated agent environments concurrently.

Existing clients utilizing CoreWeave’s GB200 and GB300 NVL72 systems (including Dell PowerRack deployments) can adopt the new Vera Rubin setup using their existing management tools.

Unifying the Stack with “Forge”

On the software front, CoreWeave introduced Forge, an all-in-one AI platform aimed at closing the loop between training, evaluation, inference, and real-time observability.

Traditionally, AI engineering teams stitch together fragmented tools from multiple vendors, creating friction when feeding production insights back into model retraining. Forge consolidates these steps into a continuous feedback loop:

  1. Deploy & Observe: Run agents and track live performance metrics.

  2. Curate & Refine: Use production data to improve model prompts and logic.

  3. Evaluate & Iterate: Test improvements in sandboxed environments before redeploying.

Forge is designed to remain vendor-agnostic, integrating with various open-source frameworks, model architectures, and third-party cloud environments. It integrates natively with CoreWeave’s underlying tools—including its managed Kubernetes service (SUNK), Mission Control, and serverless inference layers—allowing teams to upgrade underlying hardware generations without breaking their operational workflow.

CoreWeave has officially rolled out Forge, offering a 30-day free trial for its Pro tier.

Photo by Steve A Johnson on Unsplash

Share:

Leave a Reply

Your email address will not be published. Required fields are marked *