Data center thermal management: A practical guide for AI workloads

Author: Kyle Geib, Application Engineering Manager at Motivair by Schneider Electric

Executive Summary

This blog answers common questions about AI data center thermal management, including:

  • What is data center thermal management? How cooling systems manage heat and humidity.
  • Why is AI changing data center cooling? Higher GPU power and rack densities generate more heat.
  • When is liquid cooling needed? High-density AI workloads can far exceed the capabilities of air cooling.
  • How are high-density AI racks cooled? Rear door heat exchangers (RDHx) and direct-to-chip cooling more efficiently manage increasing heat loads.
  • How can cooling keep pace with more powerful AI chips? New cooling technologies can handle the heat generated by evolving chips.

With the rapid growth of artificial intelligence (AI) applications, data centers are undergoing significant changes, including in data center thermal management. As AI workloads become increasingly demanding, cooling systems designed for previous generations of IT equipment can struggle to keep pace with the heat generated by today’s high-density AI servers and GPUs.

We deal with this issue every day in our work on modern data centers, and we often field questions about data center thermal management, GPU thermal management, and AI data center cooling. I thought it’d be useful to answer some of the most common questions and explain how cooling strategies are evolving to support AI infrastructure.

data center thermal management

What is data center thermal management?

Data center thermal management encompasses the systems and strategies used to control temperature and humidity throughout a data center. That includes cooling the “white space,” where servers and other IT equipment operate, as well as general room air conditioning for the rest of the facility.

When it comes to cooling computing loads, data center heat management generally means ensuring servers don’t get too hot, which would reduce their performance or, in extreme cases, even cause them to shut down. It’s no different from what you feel if you leave your laptop on your lap too long – it gets hot because its airflow is restricted.

For traditional IT loads of under 20 kW per rack, air-cooled systems provide sufficient cooling capacity. This involves computer room air-conditioning/air-handling (CRAC/CRAH) systems that deliver cool air across servers and reject heat loads via a containment-type system. Such systems can take various forms, including in-row cooling, which ensures that cool air is enclosed within computing racks and warm air is rejected so it doesn’t mix with the cool air. Such systems are far more effective than trying to cool the entire white space.

How do I manage heat in AI data centers?

As AI servers become more powerful, generating more heat, and rack densities increase, air cooling alone is no longer sufficient. Higher GPU power densities are creating new GPU thermal management challenges, requiring more robust cooling systems that use liquid cooling to remove heat more efficiently.

Water, for example, has excellent heat transfer properties, meaning it’s good at taking heat from one place (servers) and moving it elsewhere. In fact, water is 23.5 times more efficient than air in transferring heat and can absorb about 3,500 times as much heat as a comparable volume of air.

In data center liquid cooling systems, water is typically mixed with propylene glycol (PG) in a 25% concentration known as PG25. At that concentration, PG inhibits corrosion and (along with other additives) the growth of bacteria and algae. In AI and other data centers with powerful servers, PG25 may be used in various ways, including in cold plates that directly cool servers, cooling distribution units (CDUs), and rear-door heat exchanger (RDHx) systems.

What thermal approach do I use for current AI architecture?

The latest AI chips are driving rack densities well beyond what traditional air cooling can efficiently support, reaching into the 200 kW range or more. For example, NVIDIA’s newest Vera Rubin chip can reach 227kW per rack. The challenge isn’t simply that cooling needs to move closer to the heat source. As AI chips consume more power, they generate more heat, necessitating a more efficient way to dissipate it. Because liquid can carry 23.5 times more heat than air, direct-to-chip liquid cooling using cold plates is increasingly essential for managing the higher thermal loads generated by AI workloads.

Cold plates are metal plates upon which computer chips sit. A chilled PG25 solution circulates inside the plates, cooling them and the chips that sit on them. As the solution heats up, it is transferred from the plate to a CDU. The CDU works with the facility’s cooling system, such as a chiller or a cool water tower, to ensure the PG25 solution maintains the proper temperature and humidity properties as it flows to and from the computing rack.

Such systems can transfer 80% or more of the total rack heat load to the cold plates. The remaining 20% can be captured by a rear door heat exchanger (RDHx), such as the Motivair by Schneider Electric ChilledDoor®. Fans within the servers pull cool air through the front of the rack and push warm air out the back, where it passes over metal fins in the rear door. A cool

PG25 solution circulating inside those fins absorbs heat, cooling the air back down to the same temperature as the data hall, creating a room neutral solution where all the heat from the system is rejected to water.

AI data center thermal management must evolve

Effective AI data center thermal management requires more than simply adding cooling capacity. Cooling architecture must evolve alongside the chips, servers, and racks they support. The goal is to stay in lockstep with chipmakers so that, as AI infrastructure becomes more powerful, data center operators have a thermal management strategy that supports the next generation of AI.

To ensure cooling systems work effectively with new generations of AI chips, Schneider Electric engineers collaborate closely with chip manufacturers like NVIDIA to understand the heat their chips generate and what that’ll mean for the heat-rejection ratios and flow rates required for proper cooling. Based on that collaboration, we co-develop validated reference designs that detail the exact power, cooling, and other requirements for successful AI data center operation and optimized performance. Learn more about how liquid cooling reference designs optimize and accelerate AI data center deployments

Add a comment

All fields are required.