
NVIDIA Vera Rubin’s $9B Ramp Exposes a Cloud Depreciation Dilemma
NVIDIA disclosed on August 24 that its Vera Rubin NVL72 delivers up to 30× greater throughput per megawatt and 35× lower token cost than the current GB300, measured on a SemiAnalysis AgentX DeepSeek V4-Pro agentic workload at demanding latency targets. Morgan Stanley expects Rubin to contribute nearly $9 billion in NVIDIA's fiscal third quarter ending October 2026—roughly 9% of consensus quarterly sales almost immediately after launch. Racks are running at CoreWeave, Google Cloud, Microsoft Azure, Oracle, and Nebius.
Those numbers require careful reading. The 30× figure appears at ~160 tokens/sec/user on a narrow benchmark still pending third-party review. NVIDIA's broader product claims, using Kimi-K2-Thinking with standard sequence lengths, center on approximately 10× more tokens per megawatt. The honest modeling range: 2–10× for most workloads, 30–35× in select high-interactivity agentic regimes. Large, but conditional.
The Depreciation Mismatch
Hyperscalers and neoclouds carry GPU servers on their books for five to six years. Microsoft reports server and network assets at $215.9 billion gross carrying cost. Alphabet depreciates servers over six years and spent $80.6 billion of capex in just the first half of 2026. CoreWeave uses a six-year schedule for technology equipment that reached $33.8 billion by June 30. Meta extended most server lives to 5.5 years in January 2025.
NVIDIA ships major new architectures on roughly annual cycles.
Amazon has already acted on this tension. Effective January 2025, it shortened a subset of AI server lives from six years to five, citing the accelerated pace of AI development, and recorded approximately $920 million of accelerated depreciation and related charges in the fourth quarter of 2024 with another $600 million expected during 2025. That confirms the mechanism—faster hardware cadence compresses useful life on the balance sheet—without confirming catastrophe. Amazon moved from six years to five, not six to two.
Why Blackwell Won't Simply Disappear
CoreWeave recently disclosed a customer contract for NVIDIA A100 GPUs extending into 2029. The A100 launched in 2020. CoreWeave's CEO said its Ampere and Hopper fleets remain largely sold out. Rental data corroborates this: B200 spot pricing near $5.66/hour still commands a heavy premium over H100 at $2.69/hour and A100 at $1.60/hour. No fire sale is visible.
The explanation is mechanical. Once a Blackwell cluster is purchased, its price is sunk. Operators keep running it as long as revenue covers electricity, cooling, networking, and maintenance. Depreciation hits the income statement, not the cash operating decision. Rubin's superiority will erode return on invested capital before it idles a single GPU. Accounting charges—shorter useful lives, higher annual depreciation—will precede any physical shutdown by years.
The Two-Speed Capital Market
This is where the analysis sharpens into an investable thesis. Rubin does not annihilate Blackwell. It splits the GPU capital market into tiers: Rubin captures scarce premium power and latency-sensitive agentic workloads. Blackwell retains utilization at lower pricing for mainstream inference. Hopper and Ampere migrate further down—batch processing, fine-tuning, academic compute, any job where older silicon satisfies latency requirements cheaply.
Revenue per megawatt governs the split. If Rubin produces 10× the throughput at a given power constraint, token prices can fall 80% and revenue per megawatt still doubles, assuming demand exists. Cheaper tokens generate more tokens—NVIDIA cites OpenRouter data showing agentic interactions consuming around 15× the tokens of conventional chat—so aggregate demand can absorb efficiency gains without idling older hardware.
Financial stress concentrates where financing duration exceeds the GPU's premium economic life. Leveraged neoclouds carrying six-year depreciation schedules and debt maturing beyond contracted customer terms face the greatest exposure. GPU-backed private credit structures underwritten to optimistic residual values are similarly vulnerable; the Bank of England has flagged maturity mismatch and uncertain GPU collateral curves. CoreWeave's $104 billion backlog and 112% year-over-year revenue growth look operationally healthy, but its $35 billion of combined debt and $1.4 billion of quarterly depreciation make it the clearest public test case for whether this financing model survives repeated architecture transitions.
The decisive indicator is a simultaneous three-part signal: Blackwell rental-price compression, weakening recontracting economics, and cloud operators shortening depreciation lives. Until all three converge, Rubin is a margin event inside a growing market—not a liquidation trigger.
not investment advice