Liquid Cooling Module Protocol for N+1 Server Redundancy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current liquid cooling systems for computer servers face challenges in balancing redundancy and space efficiency, with 1+1 redundancy systems either oversizing for normal operation or being fragile in case of module failure, and 2+1 redundancy systems risking instability due to master/slave architectures and potential inconsistencies in non-hierarchical communication.
Innovation Solution
Implementing a collaborative communication protocol among interchangeable liquid cooling modules without a master/slave hierarchy, ensuring consistent data exchange and stability checks to maintain system robustness and efficiency, allowing seamless module replacement without shutting down the cooling system or servers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If 1+1 redundancy configuration is used, then robustness in case of module failure is improved, but space efficiency deteriorates due to significant oversizing
Solution Approach 1:
The cooling system is divided into multiple independent cooling modules (at least three modules) that can operate independently. Each module handles a portion of the cooling load, allowing the system to maintain robustness through redundancy while reducing the size of individual modules compared to a single oversized module, thus improving space efficiency.
2Area of stationary object
If module redundancy is not maintained, then space efficiency is improved by sizing precisely for normal operation, but robustness deteriorates in case of module failure
Solution Approach 1:
The system pre-configures at least three cooling modules with defined roles (master and slave) before failure occurs. The slave modules are prepared to take over cooling duties immediately upon master module failure, ensuring robustness is maintained without requiring oversized modules for normal operation, thus optimizing space efficiency.
3Ease of operation
If master/slave architecture is used, then ease of control is improved, but reliability deteriorates due to vulnerability to master module failure
Solution Approach 1:
The system pre-establishes a master module among the at least three cooling modules with clear control responsibilities. This master module coordinates the operation of slave modules, providing ease of control through centralized management while maintaining reliability through the preparedness of slave modules to take over if the master fails.
Solution Approach 2:
The system implements monitoring and communication mechanisms where cooling modules exchange operational status information. This feedback enables automatic detection of master module failure and triggers the slave module takeover protocol, maintaining reliability while preserving the simplicity of master/slave control architecture.
4Reliability
If collaborative protocol without master/slave is used, then reliability is improved by avoiding single point of failure, but device complexity increases due to non-hierarchical communication
Solution Approach 1:
The control architecture is segmented into distinct master and slave module roles, creating a hierarchical structure that simplifies communication protocols. This segmentation reduces the complexity of inter-module communication while maintaining reliability through the preparedness of slave modules to assume master responsibilities upon failure detection.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
The invention relates to a method for communication between a plurality of liquid cooling modules (4, 5, 6) of a system for cooling at least one computer server (3), characterised in that the cooling modules (4, 5, 6) communicate with each other in such a way as to operate under N+1 redundancy where N is higher than or equal to 2, so as to be able to carry out a standard replacement of any one of said cooling modules (4, 5, 6) without stopping the cooling and without stopping the operation of the at least one server (3), said communication being ensured by a collaborative protocol without master/slave technology, before switching (35) from an active mode (16, 17) in which it cools, to a backup mode (20) in which it no longer cools, the redundant cooling module (6) previously checking (33) that a set of data is coherent between all of said cooling modules (4, 5, 6) and that said coherence is maintained over a predetermined period.