At today's cloud prices, a continuously used H100 pays for itself in under a year. The interesting question is no longer 'cloud or on-prem' — it is at what utilization the switch happens.
The break-even, roughly
A rented H100 costs $2–4 per GPU-hour depending on commitment. An owned HGX node — hardware, power, cooling, space, operations — lands near $1 per GPU-hour over four years at high utilization. If your GPUs run above roughly 40–50% of the time, ownership wins on raw cost; below 20%, renting wins; in between, the answer is a hybrid.
What the spreadsheets miss
- Data gravity — moving tens of terabytes to a cloud region every training run has its own bill and its own latency.
- Compliance — banks and government bodies increasingly cannot ship training data off-premises at all.
- Queue risk — spot capacity evaporates exactly when a frontier model release makes everyone train at once.
- Exit cost — leaving a cloud after two years of accumulated tooling is a project of its own.
The pattern that works
Own the baseline, rent the burst: a private cluster sized for steady training load, with cloud overflow for experiments.
A typical starting point we deliver: two to four HGX H200 nodes with liquid-ready racks, 400G interconnect and storage sized for checkpoints — expandable in place, delivered and burned-in as one project. From there, cloud becomes a tactical tool instead of a monthly surprise.