GPU cluster in cooling data center

1344×768 · AVIF · CC BY 4.0

GPU cluster in cooling data center in editorial style

GPU cluster in a cooled data center: critical infrastructure for training AI models.

About this subject

A GPU cluster is a configuration of multiple graphics processing units working in parallel, primarily used for high-performance computing such as training deep neural networks. In data centers, these clusters are mounted in specialized racks and require robust cooling systems, as modern GPUs can dissipate hundreds of watts each. For instance, the NVIDIA H100, a common GPU in AI clusters, has a TDP of up to 700W. Without adequate cooling, the generated heat can reduce component lifespan and cause instability.

Data center cooling can be air-based (CRAC) or liquid-based. Liquid cooling, such as immersion or cold plates, is more efficient for dense clusters. Companies like CoreWeave and Lambda Labs operate data centers with thousands of GPUs for rent, serving AI startups. The largest known cluster, NVIDIA's Selene, uses over 4,000 A100 GPUs. Power consumption is enormous: a single rack can draw tens of kilowatts, requiring careful electrical planning.

Data center location also matters. Cold regions, like Northern Europe, reduce cooling costs. Google and Microsoft have invested in underwater data centers and arid regions with renewable energy. In Brazil, data centers in São Paulo and Rio de Janeiro use air cooling, but liquid cooling adoption grows due to high GPU density. The trend is for GPU clusters to become even denser with the advance of generative AI, demanding innovations in cooling and energy efficiency.

Frequently Asked Questions

How many GPUs does a typical cluster have?

A cluster can have from dozens to thousands of GPUs. For example, NVIDIA's Selene cluster has over 4,000 A100 GPUs. Smaller clusters used by startups may have 8 to 64 GPUs.

What is the difference between air and liquid cooling?

Air cooling uses fans and air conditioning, being simpler and cheaper, but less efficient for high density. Liquid cooling uses water or dielectric fluid, transferring heat more effectively, allowing denser clusters and energy savings.

How much does it cost to operate a GPU cluster?

Costs vary widely. Renting an H100 GPU in the cloud costs about US$ 2-3 per hour. A cluster of 1,000 GPUs can consume over 700 kW, resulting in electricity bills of hundreds of thousands of dollars per month.

Download

Download AVIF

110 KB · 1344×768

Direct URL

https://pub-c7d6a6ea828543ac903a74a341ccb2e1.r2.dev/imagens/gpu-cluster-in-cooling-data-center-editorial-portrait-overcast-soft-daylight.avif

How to credit

Include a visible link back to UtilizAí. Copy one of the snippets below:

HTML
<a href="https://xn--utiliza-eza.com/en/midia/imagens/gpu-cluster-in-cooling-data-center-editorial-portrait-overcast-soft-daylight">GPU cluster in cooling data center</a> by <a href="https://xn--utiliza-eza.com">UtilizAí</a>, licensed under <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>.
Markdown
[GPU cluster in cooling data center](https://xn--utiliza-eza.com/en/midia/imagens/gpu-cluster-in-cooling-data-center-editorial-portrait-overcast-soft-daylight) by [UtilizAí](https://xn--utiliza-eza.com), CC BY 4.0

License: CC-BY-4.0

Tags

Related images