#inference
← All postsNVIDIA Dynamo, fleet-wide
A Modelplane cluster can now serve models with NVIDIA Dynamo's components. It's a per-cluster platform choice, and the API an ML team writes stays the same.

Modelplane v0.3: Vultr, the Anthropic Messages API, and testing without a GPU
Modelplane v0.3 adds Vultr VKE as an inference cluster provider, serves the Anthropic Messages API end to end so tools like Claude Code run against your own GPUs, improves multi-node scheduling, and adds local end-to-end testing that needs no cloud and no GPU.

Why Day 0 for Nemotron 3.5 Lightning wasn't a scramble
NVIDIA released Nemotron-3.5-Lightning this morning. It was running on Modelplane by the afternoon, without a line of new Modelplane code, because day-zero model support is built into the design, not a scramble by the team.

Anthropic is subsidizing our AI coding at 13x. How long will it last?
We measured what our team's Claude Code usage would cost at API prices. It runs about 13x our seat price on average, and 52x for our heaviest engineer. Here are the numbers, how we measure them, and the script to measure your own.

Modelplane v0.2: more clouds, and traffic you can direct
Modelplane v0.2 adds Nebius and Azure AKS as inference cluster providers, weighted routing for safe model rollouts, and cluster taints for reserving and draining capacity.

Any Engine, Any Topology, Any Infrastructure: How We Designed Modelplane
How we designed Modelplane's fleet-level inference API to fit any engine, in any topology, on any infrastructure — and what's under the hood now that v0.1 has shipped.

Introducing Modelplane: the control plane for AI inference
Today we're open sourcing Modelplane, a control plane that operates AI inference across a fleet of GPU clusters, on cloud, neocloud, and on-premise, as one inference platform.



