All news

What Does It Actually Cost to Run an AI Model?

Running an AI model costs more than most organisations expect. We break down GPU hardware, power, cooling, and egress to show where the money actually goes.

What Does It Actually Cost to Run an AI Model?

The question looks simple. It rarely is. The cost of running an AI model is the sum of hardware, power, cooling, connectivity, and the facility supporting them, shaped by whether you are training, running inference at scale, or both.

We get asked about this regularly from organisations planning AI deployments, either trying to understand why their current cloud bills have spiralled, or working out whether owning and colocating the hardware would change the economics. This post sets out where the costs actually sit and why infrastructure decisions have a bigger effect on the final number than most organisations realise when they start.

Training vs Inference: Two Very Different Cost Problems

Before any infrastructure conversation makes sense, it is worth separating these two workload types. Both use GPUs, both require the same facility infrastructure, and the hardware is often identical. The cost profile difference is one of utilisation pattern and duration, not hardware type, which is why the two tend to get conflated in budget conversations when they should not be.

Training a large AI model is a sustained, intensive compute job. It requires thousands of GPUs running at or near full utilisation, continuously, for weeks or months. The power draw is substantial: a single NVIDIA H100 GPU draws approximately 700W at peak load, and a rack of four DGX H100 systems, each containing eight GPUs, pulls 40 to 42kW. Scale that to a modest training cluster, and you are looking at tens of megawatts of sustained draw, before cooling adds another 30 to 50% on top.

Training GPT-4, by OpenAI's own reporting, required approximately 25,000 A100 GPUs and consumed around 50 gigawatt-hours over the training run. Meta's Llama 3 405B model consumed roughly 30 gigawatt-hours over 54 days on 16,384 H100s. MIT Technology Review's analysis of AI energy use confirms that this scale of demand is driving data centre construction at a pace the broader energy infrastructure is struggling to keep up with. These are hyperscaler-scale numbers, not everyday enterprise AI costs. But they illustrate what sustained GPU compute at high density demands from the physical infrastructure supporting it.

Inference, running a trained model to generate responses, is where most organisations encounter AI costs day to day. The IEA's 2025 Energy and AI report estimates that querying a language model currently requires around 2 watt-hours per response, with short video generation requiring roughly 25 times as much.

The issue is not the cost per query. It is the volume. Inference now accounts for approximately 80 to 90% of total AI compute demand globally. As organisations embed AI into search, customer service, document processing, and internal tooling, the cumulative GPU time adds up continuously. An inference workload is not a job you schedule and end. It runs persistently, and it grows.

The Four Cost Components of AI Infrastructure

Breaking down the total cost reveals four components where higher density in one drives cost in the next: more powerful hardware demands more power, more power at density demands more capable cooling, and more capable cooling raises facility overhead on top of the base electricity cost.

GPU hardware . A single NVIDIA H100 SXM GPU costs approximately £20,000 to £32,000 to purchase. A complete 8-GPU DGX H100 server runs £160,000 to £360,000. For a serious inference cluster or training capability, hardware investment reaches tens of millions before the facility is considered. GPU generations cycle quickly: the B200 succeeds the H100 with 2.5x the inference throughput at 1.4x the power draw, which means hardware depreciation is a real cost over any three to five-year TCO horizon.

Power. An H100 draws 700W. A B200 draws 1,000W. At rack densities of 40 to 100kW, electricity costs accumulate every hour the hardware is running. The IEA reports that data centres consumed approximately 415 terawatt-hours globally in 2024, growing at 12% per year, with AI as the primary driver. The efficiency of the facility compounds this directly. A facility running at PUE 1.5 spends 50% more electricity per unit of IT load than one at PUE 1.0. For a rack drawing 50kW, the difference between PUE 1.5 and PUE 1.1 is roughly 20kW of overhead running continuously.

Cooling. Standard air cooling handles rack densities up to 10 to 15kW comfortably. GPU clusters running at 40kW or above require Direct-to-Chip or immersion cooling, which pushes PUE down to 1.1 to 1.15. The cost consequence is not only efficiency. An air-cooled facility physically cannot sustain a 100kW rack. Organisations that plan GPU deployments and then discover their intended facility cannot support the density, either accept degraded performance, spread workloads across more racks, or move. Our post on air cooling vs liquid cooling sets out the technical and cost tradeoffs in detail.

Connectivity and egress. For organisations running inference on cloud GPU instances, data egress charges are often the line item that makes the cost model opaque. Cloud providers charge for data leaving their environment, and for AI workloads that process large documents, serve responses at volume, or query external data sources; those charges accumulate quickly. Owning hardware in a colocation facility eliminates egress charges on traffic between your own systems.

Cloud vs Owned Hardware: Where the Numbers Land

The right model depends on utilisation. The table below summarises how the two approaches compare across the factors that actually determine total cost.

The key takeaway is that for workloads running at 70% utilisation or above continuously, owned hardware in a colocation facility consistently reaches a lower total cost of ownership within 18 to 24 months. For workloads that are seasonal, experimental, or not yet at steady state, the cloud retains the advantage.

One caution worth noting: cloud GPU pricing rarely includes the managed service overhead, storage, egress, and support tier costs that appear on actual bills. Colocation costs include hardware depreciation and procurement. Both totals need to be calculated honestly to compare like-for-like. Our post on colocation vs cloud sets out the full decision framework.

What This Means for Infrastructure Planning

In the power assessments we run, the organisations with the most predictable AI infrastructure costs tend to do three things differently from those who are surprised by their bills. They separate training and inference workloads early and size infrastructure for each independently. They run total cost of ownership calculations that extend at least three years and include egress, support, and hardware depreciation. And they choose facilities based on power density capability, not headline rack price.

For any deployment drawing above 15kW per rack consistently, air cooling is a constraint. For any inference workload running at sustained high utilisation, cloud GPU costs will generally exceed owned hardware within two years.

How Carbon-Z Supports AI Deployments

Two patterns come up repeatedly in the power assessments we run. The first is organisations that sized their initial infrastructure for the model they were deploying in year one, without accounting for the power and cooling requirements of the workloads that followed it. The second is organisations that started on cloud GPU instances and found their egress and storage costs were doubling the headline compute price, often without anyone having noticed until the quarterly bill arrived.

Our Immersion Cooling environments run GPU clusters and AI training workloads at 25kW to 100kW per pod, with Direct-to-Chip options for precision thermal control, and are live and accepting deployments now. Standard enterprise and networking infrastructure runs alongside in air-cooled racks from 4kW to 10kW, within the same certified facility, so a deployment that starts with standard racks can scale into liquid-cooled GPU infrastructure without a facility change. All environments are ISO 27001, ISO 14001, and ISO 45001 certified, with N+1 redundancy, carrier-neutral connectivity, 24/7 manned access, and on-site smart hands. We charge for power consumed, not rack footprint.

If you are running GPU workloads or planning an AI deployment that needs liquid-cooled infrastructure, explore our Immersion Cooling environments and see what a deployment would look like at your required density.

If you are planning an AI deployment and want to model what the infrastructure costs would look like against your current setup, book a free power assessment , and we will work through the numbers with you.

For anything more immediate, request a call back , and we will respond within 24 hours.

Related articles

Data Centre Migration Checklist - How to Plan a Low-Risk MoveUse this data centre migration checklist to plan dependencies, power, connectivity, rollback and validation for a lower-risk move.How Much Does Immersion Cooling Cost? Capex, Opex, and TCO ExplainedSee what drives immersion cooling cost in the UK, from CapEx and OpEx to TCO, and compare the real cost of supporting high-density compute.What Is a Coolant Distribution Unit? A Data Centre CDU GuideWhat is a coolant distribution unit? Learn how CDUs manage coolant flow, heat transfer, and pressure in liquid-cooled data centres.AI Colocation in the UK: How to Choose Infrastructure That Will Not Hold Your GPUs BackCarbon-Z delivers AI colocation in the UK with liquid cooling up to 120kW per rack. Built for GPU clusters, AI training, and sustained high-density workloads.From 8kW To 120kW: When Your GPU Cluster Outgrows Standard ColocationGPU clusters scaling from 8kW to 120kW often outgrow standard colocation. Learn the warning signs and what infrastructure changes are needed for dense compute.Liquid Cooling For Data Centres: A Buyer's GuideLiquid Cooling for Data Centres explained. Compare cooling options, buyer checks and key questions before planning high-density infrastructure.What 120kW Per Rack Actually Looks Like: Power, Cooling, And Cabling SpecificationsSee what 120kW per rack means for power, cooling, cabling and monitoring before planning high-density data centre infrastructure.Are Colocation Data Centres the Same as Servers?Colocation data centres and servers fill different roles in IT infrastructure. Learn how each works, when colocation is the right choice and what to look for.Carrier-Neutral Data Centre Benefits: Why Network Choice MattersExplore carrier-neutral data centre benefits, from provider choice and route diversity to stronger hybrid connectivity.What Is Immersion Cooling? A Practical Guide for High-Density InfrastructureWhat is immersion cooling? Learn how it works, when it makes sense, and how it supports high-density infrastructure.How to Improve Network Resilience in Data CentresHow to improve network resilience in data centres through diverse connectivity, tested failover, configuration control and wider observability.Where to Colocate in the UK: A Guide to the Top Data Centre HubsChoosing a UK colocation hub now turns on power and cooling, not postcode. We map the four hub types and how to match each to your workload.Why AI and HPC Workloads Need Immersion CoolingAI and HPC racks now draw 40 to 140 kilowatts. We explain why air cooling has hit its ceiling and where immersion genuinely earns its place.What Is Colocation? The Complete UK Guide 2026Colocation lets you house your servers in a managed UK data centre. Our guide covers costs, cooling, security, cloud comparisons, and how to choose a providerAir Cooling vs Liquid Cooling: Which Does Your Infrastructure Actually Need?Air cooling vs liquid cooling: which does your infrastructure need? We break down rack density, PUE, and total cost to help you make the right call.Colocation vs Cloud - Where Your Workloads Actually BelongColocation vs cloud isn't a philosophy debate. We break down the real cost, compliance, and performance factors that determine where your workloads belong.The thirst for AIAI is revolutionary in its capabilities. It is becoming integrated to all the applications that we use…Combating obsolete Data CentresDive into how immersion cooling slashes energy use and unlocks high rack densities for AI, GPU and HPC workloads.Open DayExciting News! Join us for a Journey into the Future of Hosting and Cooling at Swindon Data Centre Open Day!AtomsCarbon-Z Atoms are modular, build-on-demand data centre units with up to 1MW capacity and flexible cooling options built for rapid deployment and scalability.

Ready to upgrade your infrastructure?

Stop overpaying for legacy efficiency. Get a quote for colocation, immersion or a custom build in under 24 hours.