The question looks simple. It rarely is. The cost of running an AI model is the sum of hardware, power, cooling, connectivity, and the facility supporting them, shaped by whether you are training, running inference at scale, or both.
We get asked about this regularly from organisations planning AI deployments, either trying to understand why their current cloud bills have spiralled, or working out whether owning and colocating the hardware would change the economics. This post sets out where the costs actually sit and why infrastructure decisions have a bigger effect on the final number than most organisations realise when they start.
Training vs Inference: Two Very Different Cost Problems
Before any infrastructure conversation makes sense, it is worth separating these two workload types. Both use GPUs, both require the same facility infrastructure, and the hardware is often identical. The cost profile difference is one of utilisation pattern and duration, not hardware type, which is why the two tend to get conflated in budget conversations when they should not be.
Training a large AI model is a sustained, intensive compute job. It requires thousands of GPUs running at or near full utilisation, continuously, for weeks or months. The power draw is substantial: a single NVIDIA H100 GPU draws approximately 700W at peak load, and a rack of four DGX H100 systems, each containing eight GPUs, pulls 40 to 42kW. Scale that to a modest training cluster, and you are looking at tens of megawatts of sustained draw, before cooling adds another 30 to 50% on top.
Training GPT-4, by OpenAI's own reporting, required approximately 25,000 A100 GPUs and consumed around 50 gigawatt-hours over the training run. Meta's Llama 3 405B model consumed roughly 30 gigawatt-hours over 54 days on 16,384 H100s. MIT Technology Review's analysis of AI energy use confirms that this scale of demand is driving data centre construction at a pace the broader energy infrastructure is struggling to keep up with. These are hyperscaler-scale numbers, not everyday enterprise AI costs. But they illustrate what sustained GPU compute at high density demands from the physical infrastructure supporting it.
Inference, running a trained model to generate responses, is where most organisations encounter AI costs day to day. The IEA's 2025 Energy and AI report estimates that querying a language model currently requires around 2 watt-hours per response, with short video generation requiring roughly 25 times as much.
The issue is not the cost per query. It is the volume. Inference now accounts for approximately 80 to 90% of total AI compute demand globally. As organisations embed AI into search, customer service, document processing, and internal tooling, the cumulative GPU time adds up continuously. An inference workload is not a job you schedule and end. It runs persistently, and it grows.
The Four Cost Components of AI Infrastructure
Breaking down the total cost reveals four components where higher density in one drives cost in the next: more powerful hardware demands more power, more power at density demands more capable cooling, and more capable cooling raises facility overhead on top of the base electricity cost.
GPU hardware . A single NVIDIA H100 SXM GPU costs approximately £20,000 to £32,000 to purchase. A complete 8-GPU DGX H100 server runs £160,000 to £360,000. For a serious inference cluster or training capability, hardware investment reaches tens of millions before the facility is considered. GPU generations cycle quickly: the B200 succeeds the H100 with 2.5x the inference throughput at 1.4x the power draw, which means hardware depreciation is a real cost over any three to five-year TCO horizon.
Power. An H100 draws 700W. A B200 draws 1,000W. At rack densities of 40 to 100kW, electricity costs accumulate every hour the hardware is running. The IEA reports that data centres consumed approximately 415 terawatt-hours globally in 2024, growing at 12% per year, with AI as the primary driver. The efficiency of the facility compounds this directly. A facility running at PUE 1.5 spends 50% more electricity per unit of IT load than one at PUE 1.0. For a rack drawing 50kW, the difference between PUE 1.5 and PUE 1.1 is roughly 20kW of overhead running continuously.
Cooling. Standard air cooling handles rack densities up to 10 to 15kW comfortably. GPU clusters running at 40kW or above require Direct-to-Chip or immersion cooling, which pushes PUE down to 1.1 to 1.15. The cost consequence is not only efficiency. An air-cooled facility physically cannot sustain a 100kW rack. Organisations that plan GPU deployments and then discover their intended facility cannot support the density, either accept degraded performance, spread workloads across more racks, or move. Our post on air cooling vs liquid cooling sets out the technical and cost tradeoffs in detail.
Connectivity and egress. For organisations running inference on cloud GPU instances, data egress charges are often the line item that makes the cost model opaque. Cloud providers charge for data leaving their environment, and for AI workloads that process large documents, serve responses at volume, or query external data sources; those charges accumulate quickly. Owning hardware in a colocation facility eliminates egress charges on traffic between your own systems.
Cloud vs Owned Hardware: Where the Numbers Land
The right model depends on utilisation. The table below summarises how the two approaches compare across the factors that actually determine total cost.
The key takeaway is that for workloads running at 70% utilisation or above continuously, owned hardware in a colocation facility consistently reaches a lower total cost of ownership within 18 to 24 months. For workloads that are seasonal, experimental, or not yet at steady state, the cloud retains the advantage.
One caution worth noting: cloud GPU pricing rarely includes the managed service overhead, storage, egress, and support tier costs that appear on actual bills. Colocation costs include hardware depreciation and procurement. Both totals need to be calculated honestly to compare like-for-like. Our post on colocation vs cloud sets out the full decision framework.
What This Means for Infrastructure Planning
In the power assessments we run, the organisations with the most predictable AI infrastructure costs tend to do three things differently from those who are surprised by their bills. They separate training and inference workloads early and size infrastructure for each independently. They run total cost of ownership calculations that extend at least three years and include egress, support, and hardware depreciation. And they choose facilities based on power density capability, not headline rack price.
For any deployment drawing above 15kW per rack consistently, air cooling is a constraint. For any inference workload running at sustained high utilisation, cloud GPU costs will generally exceed owned hardware within two years.
How Carbon-Z Supports AI Deployments
Two patterns come up repeatedly in the power assessments we run. The first is organisations that sized their initial infrastructure for the model they were deploying in year one, without accounting for the power and cooling requirements of the workloads that followed it. The second is organisations that started on cloud GPU instances and found their egress and storage costs were doubling the headline compute price, often without anyone having noticed until the quarterly bill arrived.
Our Immersion Cooling environments run GPU clusters and AI training workloads at 25kW to 100kW per pod, with Direct-to-Chip options for precision thermal control, and are live and accepting deployments now. Standard enterprise and networking infrastructure runs alongside in air-cooled racks from 4kW to 10kW, within the same certified facility, so a deployment that starts with standard racks can scale into liquid-cooled GPU infrastructure without a facility change. All environments are ISO 27001, ISO 14001, and ISO 45001 certified, with N+1 redundancy, carrier-neutral connectivity, 24/7 manned access, and on-site smart hands. We charge for power consumed, not rack footprint.
If you are running GPU workloads or planning an AI deployment that needs liquid-cooled infrastructure, explore our Immersion Cooling environments and see what a deployment would look like at your required density.
If you are planning an AI deployment and want to model what the infrastructure costs would look like against your current setup, book a free power assessment , and we will work through the numbers with you.
For anything more immediate, request a call back , and we will respond within 24 hours.


