Server Thermal Management During Cooling Failure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face power consumption increases and potential server shutdowns due to cooling failures, leading to overheating and reduced operational uptime.
Innovation Solution
Implementing a system within servers that monitors temperature and rate of change, allowing for gradual reduction in power consumption by throttling processor speed, memory bandwidth, and fan speed to minimize heat production, thereby extending operational time during cooling failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Temperature
If fan speed is increased to cool servers during cooling failure, then cooling effectiveness is improved, but power consumption increases and more heat is produced
Solution Approach 1:
The system dynamically changes operational parameters (processor speed, memory bandwidth, fan speed) in response to cooling failure conditions. By throttling processor speed and memory bandwidth, the system reduces heat generation at the source, thereby reducing the demand for active cooling and overall power consumption while maintaining acceptable temperature levels.
2Productivity
If processor speed is maintained at full capacity, then computational productivity is improved, but heat production increases leading to server shutdown
Solution Approach 1:
The system applies partial throttling to processor speed and memory bandwidth rather than complete shutdown. This partial reduction in operational capacity generates sufficient heat reduction to prevent shutdown while maintaining a reduced level of productivity, allowing the server to continue operating at diminished capacity during cooling failures.
3Duration of action of moving object
If power consumption is reduced during cooling failure, then heat production is decreased extending operational time, but productivity is reduced
Solution Approach 1:
The system dynamically adjusts processor speed, memory bandwidth, and fan speed based on real-time temperature conditions and cooling failure status. This dynamic adaptation allows the server to optimize the balance between productivity and heat generation continuously, extending operational uptime while maintaining the highest possible productivity level under constrained thermal conditions.
Data Source
AI summary
A method includes detecting that a rate of temperature change in a server is above a threshold rate, changing the server to a lowest system performance state when the rate of the temperature change in the server is above the threshold rate, and reducing a fan speed to a minimum fan speed level when the rate of temperature change is above the threshold rate.


