Liquid Cooling Module Protocol for N+1 Server Redundancy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current liquid cooling systems for computer servers face challenges in balancing redundancy and space efficiency, with 1+1 redundancy systems either overengineering for normal mode or being sensitive to single module failures, leading to potential shutdowns and inefficiencies.
Innovation Solution
Implementing a collaborative protocol among multiple interchangeable liquid cooling modules operating in N+1 redundancy, where all modules communicate equally without a master/slave hierarchy, allowing seamless replacement and maintaining system operation without shutdowns, and verifying data consistency to prevent instability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If 1+1 redundancy is implemented with two large cooling modules, then robustness against failure is improved, but space efficiency deteriorates due to one module remaining inactive
Solution Approach 1:
The cooling system is divided into multiple independent cooling modules, each capable of operating autonomously. Instead of using two large modules where one sits idle, the system segments the cooling capacity across multiple smaller modules that can be dynamically activated based on actual cooling needs, thereby reducing the space occupied by inactive hardware while maintaining reliability.
Solution Approach 2:
The system transitions from a static redundancy configuration (where one module is permanently inactive) to a dynamic configuration where modules can be activated or deactivated based on real-time cooling demands and operational status. This dynamic approach allows the system to optimize space utilization by only activating the necessary number of modules at any given time.
2Area of stationary object
If N+1 redundancy with multiple interchangeable modules is implemented, then space efficiency is improved, but system complexity increases due to collaborative protocol requirements
Solution Approach 1:
All cooling modules are designed with identical functionality and interfaces, making them universally interchangeable. Each module can assume any role (active or standby) based on operational needs. This universality simplifies the collaborative protocol by eliminating the need for complex role management and hierarchical control, as any module can seamlessly replace any other module without reconfiguration.
Solution Approach 2:
The cooling modules autonomously manage their own state transitions and coordinate with each other through a simplified peer-to-peer protocol. Each module independently monitors its own operational status and automatically activates or deactivates based on system needs, reducing the need for complex centralized control logic and minimizing overall system complexity.
3Reliability
If master/slave architecture is used with 2+1 redundancy, then robustness is improved, but adaptability deteriorates due to inventory requirements for different module types
Solution Approach 1:
The system employs universally interchangeable cooling modules that all conform to the same interface and functional specifications. This eliminates the need for different types of master and slave modules, allowing any module to serve any position in the system. Consequently, the inventory requirement is reduced to a single module type, significantly improving adaptability and simplifying maintenance operations.
Data Source
AI summary
Disclosed is a method of communication between a plurality of liquid cooling modules of a cooling system for one or more one computer servers, in which: the cooling modules communicate with each other in a manner that operates in N+1 redundancy where N is greater than or equal to 2, so as to enable a standard replacement of any one of these cooling modules without stopping the cooling and without stopping the operation of the server or servers, this communication being ensured by a collaborative protocol without master/slave, before switching from an active mode where it is cooling to a backup mode where it is no longer cooling, the redundant cooling module verifying beforehand that a data set is consistent across all these cooling modules and that this consistency is maintained for a predetermined duration.


