Redundant Air-Cooled Cooling Loop for Datacenter Reliability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Datacenter cooling systems face challenges in providing redundant and efficient cooling solutions, particularly in addressing sudden high heat requirements and failures in primary cooling loops, which can lead to inadequate heat management and downtime.

Innovation Solution

An intelligent and redundant air-cooled cooling loop system is introduced, featuring a primary cooling loop supported by a secondary loop with a liquid-to-air heat exchanger and a fluid source, enabling alternate coolant circulation to maintain cooling even in primary loop failures, using a dual-cooling cold plate and flow controllers to manage coolant flow.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If a primary cooling loop is used for datacenter cooling, then cooling efficiency is improved, but system reliability deteriorates due to lack of redundancy

Engineering Contradiction:
Improvecooling efficiencyVSAvoidsystem reliability
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The cooling system is divided into two independent loops: a primary cooling loop and a secondary cooling loop. Each loop has its own coolant circulation path, heat exchangers, and control mechanisms. This segmentation allows the system to maintain cooling functionality even when one loop fails, thus improving reliability while maintaining cooling efficiency through the primary loop.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The secondary cooling loop is prepared in advance as a standby system with all necessary cooling components pre-installed and configured. When the primary loop fails, the secondary loop can immediately activate to provide cooling coverage, cushioning against the harmful effect of complete cooling system failure and ensuring continuous operation.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

2Reliability

If a redundant cooling loop is added, then system reliability is improved, but device complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The secondary cooling loop is designed with universal components that can serve both as a standby system and as an active cooling system. The same heat exchangers, coolant channels, and control mechanisms are used in both loops, reducing the need for entirely separate redundant components and thereby limiting the increase in device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The primary and secondary cooling loops share common infrastructure elements such as the coolant reservoir, pump housing, and control system architecture. By merging these common elements, the system achieves redundancy while minimizing the additional complexity that would result from completely separate systems.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If coolant flow is redirected to secondary loop, then redundancy is activated, but cooling efficiency may deteriorate

Engineering Contradiction:
Improveredundancy activationVSAvoidcooling efficiency
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system employs dynamic flow control mechanisms that can adjust coolant distribution between the primary and secondary loops in real-time. When the primary loop is operational, the majority of coolant flow is directed through it for optimal efficiency. When the primary loop fails, the flow control dynamically redirects coolant to the secondary loop, maintaining adequate cooling performance while activating redundancy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The control system continuously monitors the operational status of the primary cooling loop and provides feedback to the flow control mechanisms. Based on this feedback, the system automatically adjusts coolant flow distribution to maintain optimal cooling efficiency during normal operation and ensures adequate cooling performance when redundancy is activated.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution ensures continuous cooling operations by providing a redundant path for heat removal, reducing downtime and maintaining efficient temperature control during primary cooling loop failures, thus enhancing datacenter reliability and efficiency.

Implementation Method 1

a liquid-to-air heat exchanger to cool alternate coolant

Methodology Applied
Scientific EffectHeat exchanger: Heat Exchanger

Implementation Method 2

a cooling tower or other external heat exchanger that receives heated coolant from the datacenter and that disperses the heat by forced air or other means to the environment

Methodology Applied
Scientific EffectForced convection: Forced Convection

Data Source

PatentUS11822398B2Intelligent and redundant air-cooled cooling loop for datacenter cooling systems
Publication Date: 2023.11.21 NVIDIA CORP
  • US11822398B2 patent drawing
  • US11822398B2 patent drawing
  • US11822398B2 patent drawing

AI summary

Systems and methods for cooling a datacenter are disclosed. In at least one embodiment, an alternate cooling loop with its own fluid source and a liquid-to-air heat exchanger is used to provide cooling for the at least one computing component alternatively from a secondary cooling loop that is associated with a primary cooling loop and a chilling facility.