Rack-Level Thermal Caching for Liquid-Cooled Data Centers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing liquid cooling systems for high-performance electronic devices are inefficient, costly, and prone to failure, leading to potential damage or data loss due to overheating in the event of a primary cooling system failure, as they rely on pumps that require maintenance and can disrupt coolant flow.
Innovation Solution
A supplemental liquid cooling system within a rack creates an independent closed loop using a standby pump and increased coolant volume, with a monitoring device that detects flow reduction or temperature rise to activate a bypass valve, establishing a secondary cooling loop to maintain heat removal and initiate a controlled shutdown.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Temperature
If a single-phase single loop pumped data center liquid cooling system is used, then thermal energy output of high-powered processors is managed effectively, but the system becomes costly, complex, and prone to failure due to pump requirements and maintenance needs
Solution Approach 1:
The cooling system is divided into two independent segments: a primary data center liquid cooling system for normal operation and a secondary rack-based cooling system for failure scenarios. Each segment has its own pump, heat exchangers, and coolant loops, allowing the secondary system to function independently when the primary system fails without requiring integration or complex coordination between the two segments
Solution Approach 2:
The secondary cooling system is pre-configured within the rack with standby pump, coolant reservoir, and heat exchangers positioned and ready before failure occurs. The monitoring device continuously monitors primary system parameters and can immediately activate the secondary system upon detecting pump failure or coolant flow disruption, eliminating the need for rapid manual intervention or system reconfiguration
Solution Approach 3:
The system incorporates a coolant reservoir and standby pump that maintain a cushion of cooling capacity within the rack. This pre-positioned cooling infrastructure acts as a safety buffer that activates automatically when the primary system fails, preventing catastrophic overheating and providing sufficient cooling headroom for controlled shutdown procedures
2Reliability
If pumps are used in liquid cooling systems, then coolant flow is maintained, but the pumps require maintenance and can break down or leak, disrupting cooling and potentially damaging electronic components
Solution Approach 1:
A monitoring device acts as an intermediary between the primary cooling system and the secondary cooling system. It continuously monitors pump operation, coolant flow, and temperature parameters, detecting failures and automatically triggering the secondary system's standby pump. This intermediary mechanism eliminates the need for manual monitoring and rapid manual response to pump failures
Solution Approach 2:
The secondary cooling system is a simplified copy of the primary system, replicated within the rack with its own pump, heat exchangers, and coolant loop. This duplicate infrastructure provides a backup cooling capability that can take over immediately when the primary pump fails, without requiring complex redundancy mechanisms or shared components that would increase maintenance burden
3Temperature
If air cooling methods are used, then cooling is provided, but the environmental control system becomes economically infeasible to maintain consistent temperature as heat generation increases
Solution Approach 1:
The cooling infrastructure is extracted from the external environmental control system and relocated inside the rack itself. The secondary cooling system uses the rack structure as part of the heat dissipation pathway, with heat exchangers positioned to transfer heat directly from the coolant to the rack housing and surrounding air, eliminating the need for expensive centralized environmental control
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution effectively reduces temperature rise and prevents data loss by maintaining cooling capabilities during primary system failures, allowing for organized shutdowns and minimizing damage to components.
Implementation Method 1
Heat is rejected from the electronic components into the working fluid passing through the cold plate
Implementation Method 2
the emerging working fluid is then run through an air-cooled heat exchanger where the heat is rejected from the working fluid to an air-stream
Data Source
AI summary
The failure of a data center liquid cooling system can result in a rapid temperature rise in electronic components that may either damage a component, result in the loss of data housed within the component, or both. A supplemental liquid cooling system is placed within each rack serviced by a data center cooling system to mitigate such a failure. Coolant flow is monitored to determine whether the data center liquid cooling system has, for a particular rack, failed. Upon determination of failure, the supplemental liquid cooling system is initiated to reduce the thermal rise of the electronic components within the rack allowing them to conduct an organized and complete shutdown.


