Shared Chassis Cooling Failover for Server Rack Reliability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing datacenter cooling systems lack redundancy and flexibility to efficiently manage failures and varying usage levels across server chassis, potentially leading to inadequate cooling resources and reduced system reliability.
Innovation Solution
A controller system that detects failures in one server chassis and dynamically allocates cooling resources from a redundant system to maintain optimal cooling, including adjusting pump speed and power distribution through an auxiliary input manifold, ensuring continuous operation and increased redundancy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If cooling systems are designed with redundancy for each server chassis, then system reliability is improved, but device complexity and cost increase
Solution Approach 1:
The patent merges cooling resources across multiple server chassis by enabling a first chassis to provide cooling resources to a second chassis. This consolidation allows the system to maintain reliability through resource sharing rather than requiring complete redundancy at each chassis level, thereby reducing overall system complexity while preserving failover capabilities.
Solution Approach 2:
The cooling system is designed with universal resources that can serve multiple chassis. The first chassis cooling system can function as both its own primary cooler and as a backup for the second chassis. This multi-functionality allows a single cooling infrastructure to provide both dedicated and redundant cooling services, improving reliability without proportionally increasing complexity.
2Temperature
If each server chassis has dedicated cooling resources, then cooling performance is ensured, but adaptability to failures and varying usage levels is reduced
Solution Approach 1:
The patent implements dynamic cooling resource allocation where the system can adaptively redirect cooling capacity based on real-time needs. When the second chassis experiences high usage levels or failures, the first chassis dynamically increases its cooling resource provision. This dynamic adjustment maintains optimal cooling performance while providing flexibility to handle varying operational conditions and failures.
Solution Approach 2:
The system incorporates feedback mechanisms that monitor usage levels and cooling needs across chassis. Based on this feedback, the controller adjusts cooling resource distribution in real-time, enabling the system to adapt to failures and varying usage patterns while maintaining temperature performance.
3Reliability
If cooling resources are shared across chassis, then redundancy and flexibility are improved, but control complexity increases
Solution Approach 1:
The patent introduces a controller as an intermediary that manages cooling resource allocation between chassis. This centralized control mechanism simplifies the complexity of coordinating shared cooling resources by providing a single point of decision-making. The controller handles the complexity of failover logic and resource distribution, allowing the physical cooling infrastructure to remain relatively simple while achieving high reliability through intelligent control.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances redundancy and flexibility in datacenter cooling systems by providing failover capabilities and adaptive resource allocation, supporting additional computing functionalities like turbo mode and maintaining system reliability even under failure conditions.
Implementation Method 1
a pump of a third chassis delivers cooling resources to a first chassis and a second chassis
Implementation Method 2
cooling resources to a first chassis and a second chassis. In some examples, each of the server chassis can include a plurality of server blades
Implementation Method 3
The CDUs within a datacenter can operate with redundancy to ensure that if one or more CDUs malfunction that additional CDUs can maintain cooling of the racks
Data Source
AI summary
Example implementations relate to a chassis cooling resource. In some examples, a chassis cooling resource includes a controller, comprising instructions to detect a failure corresponding to a first cooling system of a first chassis coupled to a server rack, and alter settings of a second cooling system of a second chassis coupled the server rack to provide additional cooling resources to the first cooling system in response to the detected failure.


