An 8kW rack and a 120kW rack may both sit inside a data centre, but they are not the same class of infrastructure.
At lower densities, standard colocation can work well. You get secure rack space, resilient power, cooling, connectivity and operational support without running your own facility. For many enterprise systems, that model is still the right answer.
GPU clusters make the rack behave differently.
As the hardware gets denser, the rack stops acting like a normal cabinet and starts acting like a concentrated power and heat zone. The decision is no longer just "where can we host this?" It becomes "can this environment support the cluster under sustained load without forcing compromises on power, cooling, cabling or maintenance access?"
That is usually the point where standard colocation starts to stop fitting the workload.
The Density Climb: From Standard Racks To GPU Clusters
Most buyers do not jump from 8kW to 120kW overnight. The problem usually appears in stages.
A first GPU deployment may fit into a standard colocation footprint. Then another node is added. Then a second rack. Then newer GPUs arrive with higher power draw, more demanding interconnects and tighter cooling requirements. Before long, the original rack plan starts to look like it was written with a quill.
NVIDIA's official guide for the DGX GB200 rack-scale system states that rack power consumption is approximately 120kW. That gives buyers a practical reference point: these densities are no longer theoretical; they are part of current accelerated compute planning.
What Standard Colocation Still Does Well
Standard colocation is not the problem. It remains a strong fit for many workloads.
If you are running traditional enterprise systems, network equipment, storage, modest virtualisation clusters, or lower-density compute, standard colocation can provide a sensible balance of control, security and operational convenience.
Our guide on what a colocation data centre is explains the basic model: you own the hardware, while the facility provides the building, power, cooling, security, connectivity and wider infrastructure.
That model works well while the rack density matches the facility design. Problems start when a GPU cluster asks a standard rack hall to behave like a high-performance compute environment.
Signs Your GPU Cluster Has Outgrown Standard Colocation
The signs usually show up in day-to-day operations first.
You may have outgrown standard colocation if:
The facility cannot allocate enough power per rack
Hardware is being spread across more racks just to manage heat
GPU performance may be constrained by thermal conditions
Cabling is becoming harder to route, label or service
Airflow requirements dictate the rack layout
The provider cannot clearly support direct liquid cooling or immersion cooling where the workload needs it
The next hardware refresh will exceed current power limits
Engineers cannot access the power, network and cooling routes safely
The cost model rewards low density but penalises consolidation
The clearest sign is when density problems force workarounds. If you are using more racks, more cross-connects and more operational effort just to avoid the density problem, the environment is no longer a good fit.
What Usually Breaks First
GPU clusters put pressure on several systems at once. The first constraint is usually power, cooling or cabling.
GPU infrastructure needs sustained electrical capacity, not just a headline allocation. Buyers need to understand peak draw, sustained draw, redundancy, feed design, UPS support and rack-level metering.
At higher densities, assumptions become expensive. A data centre power assessment gives you a practical starting point before committing to high-density hardware or migration planning.
Standard air cooling can become difficult to scale as rack density rises. The issue is not only supplying cold air. It is removing enough heat from a concentrated footprint while keeping the equipment stable and serviceable.
Uptime Institute's analysis of density choices for AI training explains how newer AI training infrastructure is pushing rack densities into ranges where power distribution, floor loading and cooling design become major planning constraints.
This is where direct-to-chip or immersion cooling may become relevant. Our article on why AI and HPC workloads need immersion cooling explains when immersion starts to make practical sense.
GPU clusters also demand careful interconnect planning.
High-speed networking, storage links, management cabling, power feeds and cooling hoses all need clear physical routing. Poor cabling design makes maintenance slower, increases fault-finding time and can restrict airflow or access.
At low density, untidy routing is annoying. At high density, it becomes an operational risk.
The Hidden Cost Of Staying In The Wrong Environment
When a GPU cluster outgrows standard colocation, the cost is not always obvious at first.
You may still be able to run the hardware, but you start paying in other ways:
More racks than the workload should need
More internal cabling and cross-connect complexity
Slower deployment of new nodes
Reduced maintenance access
More thermal monitoring workarounds
Less room for future hardware generations
Greater risk during upgrades or hardware swaps
A commercial model that charges for footprint rather than actual density needs
This is why rack density should be treated as a business decision, not just a technical preference. The cheapest rack price may not be the cheapest operating model if the workload is forced to spread sideways.
What To Check Before Moving A GPU Cluster
Before moving from standard colocation into a high-density environment, buyers should pressure-test the specification.
Ask:
What is the confirmed power draw per rack?
Is the workload steady, bursty, training-heavy or inference-heavy?
What GPU hardware will be used now and at refresh?
Does the cluster need air, direct-to-chip or immersion cooling?
Can storage and management systems run at a lower density?
How will fibre, power and coolant routes be separated?
What monitoring is available at the rack and cooling level?
What happens if a feed, cooling loop or network path fails?
Can engineers access the rack safely during maintenance?
Does the commercial model charge for space, power or both?
A mixed-density design may be the most sensible route for many clusters. GPU nodes may need high-density liquid-cooled zones, while storage, CPU compute, and management infrastructure may still sit comfortably in lower-density racks.
HPC Colocation For GPU Clusters That Need More
For GPU clusters, the service question is not just "how many racks do we need?" It is "what density can the environment support safely, and how should different parts of the cluster be housed?"
Carbon-Z's HPC colocation service supports rack densities up to 120kW, with air, direct-to-chip and immersion cooling options available within the same facility. This can help buyers plan GPU nodes, CPU compute, storage and network infrastructure around the cooling and power profile each part actually needs.
The aim is to house each part of the cluster according to its power, cooling and service requirements, rather than forcing the whole environment into one compromise. At GPU scale, compromise usually shows up later as heat, cost or awkward service access.
It is better to validate the cluster plan before the next hardware decision is locked in.
When Standard Colocation Stops Fitting
A GPU cluster outgrows standard colocation when the rack density, cooling requirement and operational complexity no longer match the environment.
At 8kW, standard colocation may be the right tool. At 60kW, 80kW or 120kW per rack, the conversation changes. Power, cooling, cabling, monitoring and maintenance all need to be designed as one system.
The right move is not always a full rebuild. It may be a staged migration into high-density colocation, a mixed-density cluster layout, or a cooling strategy that separates GPU nodes from lower-density infrastructure.
If your GPU cluster is starting to strain the limits of standard colocation, talk to Carbon-Z about reviewing the power, cooling and density requirements before the next hardware decision is locked in.


