equitiesworkspace_premiumPremium

NVIDIA's Groq Racks Deploy as AI Inference Race Accelerates

By James NakamuraPublished September 8, 20265 min read
Data center server rack aisle with blue ambient lighting

NVIDIA is shipping hardware from its $20 billion Groq acquisition, and the first deployment will arrive at a single cloud provider later this year. The Groq 3 LPX rack delivers 3,400 tokens per second in inference workloads.

That is more than 4x faster than OpenAI's Cerebras-powered solution. This gives NVIDIA a quantified speed advantage in the race to monetize low-latency AI inference.

The financial impact will be negligible in the near term. Only one customer, neocloud Nebius, is publicly committed to deploying the rack. But the move extends NVIDIA's reach across the AI compute stack.

It adds specialized inference capability to NVIDIA's dominant training franchise. This vertical integration…

workspace_premium

Continue reading with Premium

The full equities analysis, with levels, positioning, and what changed, continues below the line.

  • checkFull deep-dive reports while they're current: levels, positioning, and conviction scores
  • checkWatchlist changes as our analysts make them
  • checkExclusive investigative reports
$29/month or $199/year · cancel anytime
Subscribe to Premium

Not ready? Create a free account for extended previews · Already a member? Log in

Secure checkout via Stripe