Liquid Cooling Pump Control for Heterogeneous Server Racks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional air-cooling solutions struggle to manage thermal challenges in high power density GPU racks, leading to inefficient performance and increased energy consumption in data centers, while liquid cooling systems offer better performance but require optimization to minimize power consumption.
Innovation Solution
A liquid cooling system for data center racks that includes a coolant distribution unit and a rack management unit, which determines an optimal pump speed based on operating parameters to minimize power consumption by balancing the power usage of CPU and GPU servers and cooling equipment, using a controller to adjust the pump speed dynamically.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Temperature
If liquid cooling system is used to cool high power density GPU racks, then cooling performance is improved, but power consumption of the cooling system increases
Solution Approach 1:
The patent implements dynamic pump speed adjustment based on real-time thermal conditions and workload. The control system continuously monitors GPU temperatures and cooling water flow rates, then dynamically adjusts the pump speed to match the actual cooling demand, avoiding excessive power consumption while maintaining adequate cooling performance.
Solution Approach 2:
The system changes operating parameters (pump speed, cooling water flow rate) based on varying thermal conditions and GPU workload. By adjusting these parameters dynamically rather than operating at fixed settings, the system optimizes the balance between cooling effectiveness and power consumption.
2Temperature
If pump speed is increased to improve cooling efficiency, then temperature control is improved, but power consumption of the pump increases
Solution Approach 1:
The control system incorporates feedback mechanisms that monitor GPU temperatures and cooling water flow rates, then use this information to adjust pump speed. This closed-loop control ensures the pump operates at the minimum necessary speed to maintain adequate cooling, preventing excessive power consumption while avoiding temperature excursions.
Solution Approach 2:
The system uses the thermal feedback from the GPU racks themselves to determine the appropriate cooling provision. The cooling system essentially serves itself by using temperature and flow rate data to automatically adjust its own operation, ensuring optimal balance between cooling performance and energy consumption.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach optimizes the liquid cooling system by minimizing total power consumption in data center racks by adjusting pump speed to match the thermal and power needs of both CPU and GPU servers, enhancing cooling efficiency and reducing energy usage.
Implementation Method 1
cooling liquid flow rate to the processors
Implementation Method 2
removing the heat generated by the chips
Implementation Method 3
a liquid pump to pump the cooling liquid
Data Source
AI summary
An electronic rack includes an array of server blades arranged in a stack. Each server blade contains one or more servers and each server includes one or more processors to provide data processing services. The electronic rack includes a coolant distribution unit (CDU) and a rack management unit (RMU). The CDU supplies cooling liquid to the processors and receives the cooling liquid carrying heat from the processors. The CDU includes a liquid pump to pump the cooling liquid. The RMU is configured to manage the operations of the components within the electronic rack such as CDU, etc. The RMU includes control logic to determine an optimal pump speed to minimize the total power consumption of the pump, acceleration servers and the host servers, based on the one or more parameters and the association between temperature and power consumption of the acceleration servers and the host servers. The RMU then controls the liquid pump based on the optimal pump speed.


