Nvidia Is Selling You the Knife

IBM had built a beautiful recurring revenue business renting compute time. Then the minicomputer arrived. The customers who could run their workloads locally stopped renting. The ones who couldn't stayed. The market bifurcated along capability lines, and the capability line kept moving down toward

Nvidia Is Selling You the Knife
Photo by Thomas Foster / Unsplash

IBM had built a beautiful recurring revenue business renting compute time. Then the minicomputer arrived. The customers who could run their workloads locally stopped renting. The ones who couldn't stayed. The market bifurcated along capability lines, and the capability line kept moving down toward cheaper hardware every two years. Nobody at IBM called it a threat until it was a restructuring.

Nvidia is doing this right now, and the evidence is sitting in benchmark threads on Reddit.

An RTX 5080 paired with an older RTX 3090 is hitting 80 tokens per second on Qwen 3.6 27B at Q8 quantization. That's a 27-billion-parameter model running at conversational speed on consumer hardware that fits under a desk. Dual DGX Sparks are hitting 350 tokens per second aggregate on DeepSeek V4 Flash, with 40 tokens per second on a single 1-million-token context window. On the image side, an RTX 5090 is generating 720x720 resolution video frames at 161 frames on SCAIL-2 — which was a data center workload eighteen months ago, not a thing someone does at home on a Tuesday.

That's not a hobbyist number.

The DGX Spark is $3,000. The RTX 5090 is around $2,000. These are not enterprise procurement line items. These are credit card purchases.

A platform company builds something dominant. Partners build on top of it. The platform company watches the partners get rich. Then the platform company starts selling directly. Microsoft did it with Office, then Azure, then Teams. Apple did it with apps, then its own apps, then its own chips. Amazon did it with third-party sellers, then Amazon Basics, then AWS undercutting the very startups it hosted. The tell is always the same: the platform company starts offering infrastructure that makes the layer above it unnecessary. They never announce this. They just quietly start selling the knife.

Nvidia is not building these products for hobbyists, even if hobbyists are the ones stress-testing them on Reddit. The DGX Spark is a deliberate product decision. Nvidia looked at the inference market, looked at what OpenAI and Anthropic and Google are charging per token, and decided to sell the alternative directly to the people paying those bills. That's not a hardware company expanding its SKU lineup. That's a platform company repositioning itself one layer up the stack.

A company doing serious volume on GPT-4o or Claude Sonnet is paying somewhere between $3 and $15 per million tokens depending on the model and tier. At 80 tokens per second on local hardware, a single RTX 5080 setup handles roughly 250 million tokens a month running eight hours a day. At $5 per million tokens, that's $1,250 a month in API costs replaced by a one-time hardware purchase. The hardware pays for itself in under two years on conservative assumptions, and in under a year if you're running anything close to capacity. CFOs doing budget reviews are going to find that math very quickly.

DeepSeek, Qwen, and their successors are getting good enough, fast enough, that the gap between "what you can run locally" and "what you need the frontier API for" is narrowing in the direction that matters most to enterprise buyers: cost-sensitive, high-volume, latency-sensitive workloads. Legal document review. Code completion. Customer service triage. These don't need GPT-4o. They need 80 tokens per second and no per-token billing.

What's forming right now is a sweet spot around 25-30B quantized models that didn't exist six months ago. The hardware supports it. Q8 quantization is mature enough to deliver real throughput without catastrophic quality loss. What's missing is a dominant 25-30B model the community has actually optimized for this hardware tier — Qwen 3 27B is close, but not there yet. When someone closes that gap, the local inference case for mid-market enterprises stops being a hobbyist argument and becomes a procurement conversation. I don't know exactly when that happens. I thought it would have happened already.

They sell H100 clusters to OpenAI. They sell DGX Sparks to the developers who are trying to route around OpenAI. They collect from both sides of a conflict they are quietly accelerating. The cloud providers cannot retaliate without cannibalizing their own revenue. They can't drop API prices to zero. They can't make frontier models small enough to run locally without undermining the case for their own APIs. The only move is to keep pushing model size up and hope the capability gap stays wide enough that enterprises stay on the API. But DeepSeek V4 Flash running at 350 tokens per second on two $3,000 boxes is not a narrow gap.

The partners always think they have more runway than they do. The signal isn't the product announcement. It's when the platform company starts making it easy for customers to buy the alternative. Nvidia didn't just release fast GPUs. They released the DGX Spark with a price point and a form factor designed for the same buyers who are currently paying cloud inference bills. That's a specific decision about who the customer is.

The hobbyists on r/LocalLLaMA are running the benchmarks. The CFOs are going to read the spreadsheet version in about eighteen months. By then, the 25-30B local model will exist, the tooling will be mature, and the conversation will have already shifted. Nvidia will have sold GPUs to both sides and the cloud providers will be figuring out what they're actually selling when the compute arbitrage closes.

But "something else" is not a business model.