Datacenter Thermal Management via Risk-Based Hardware Shutdown
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Datacenters face challenges in managing temperature when cooling systems malfunction, leading to excessive heat that can damage computing equipment and disrupt operations, as existing methods lack efficient prioritization strategies for shutting down hardware units to maintain optimal temperatures.
Innovation Solution
A computer-implemented method that uses thermal imaging data, access frequency, and mission critical application information to calculate a weighted score for each hardware unit, ordering them for sequential shutdown to reduce datacenter temperature effectively when cooling systems fail.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If cooling systems are used to maintain datacenter temperature, then equipment operates in healthy state, but system complexity increases and energy consumption rises
Solution Approach 1:
The patent pre-calculates risk scores for each hardware unit based on thermal imaging data, access frequency, and mission critical application information before cooling failure occurs. This preliminary risk assessment enables immediate prioritized shutdown decisions when cooling systems fail, resolving the contradiction by preparing mitigation strategies in advance rather than adding complex real-time control systems
Solution Approach 2:
The system uses existing infrastructure components (thermal imaging cameras, network resource monitors, application deployment data) that already serve other purposes to also provide cooling risk assessment functionality. This self-service approach eliminates the need for dedicated complex cooling management systems while maintaining equipment reliability
2Temperature
If hardware units are shut down to reduce temperature, then temperature control is achieved, but loss of service and productivity occur
Solution Approach 1:
The patent applies differentiated shutdown priorities to different hardware units based on their specific risk scores. Instead of uniform shutdown protocols, each hardware unit receives localized treatment according to its thermal characteristics, access frequency, and mission critical status, achieving temperature control while minimizing overall productivity loss
Solution Approach 2:
The system dynamically changes the operational status parameter of hardware units from active to shutdown based on calculated risk scores and current temperature conditions. This parameter-based control approach optimizes the balance between temperature management and service continuity by shutting down only the least critical units first
3Ease of operation
If thermal imaging and risk assessment systems are implemented, then prioritized shutdown strategy is achieved, but system complexity and measurement requirements increase
Solution Approach 1:
The patent creates a multi-functional risk assessment system that uses thermal imaging data, network resource information, and application deployment data to simultaneously evaluate multiple risk factors. This universal approach consolidates what could be separate complex systems into a single integrated framework, achieving ease of operation through unified risk-based control
Solution Approach 2:
The system continuously monitors thermal imaging data, access frequency, and application status to calculate updated risk scores for each hardware unit. This feedback mechanism enables dynamic adjustment of shutdown priorities based on real-time conditions, simplifying operational decision-making through automated prioritization while managing complexity through systematic data collection
Data Source
AI summary
Systems, methods and/or computer program products for managing the temperature of datacenter during a period of malfunction or inoperability of the cooling system responsible for maintaining temperatures within the datacenter. Temperatures of the datacenter's computing systems are monitored by thermal imaging systems and/or sensors. Computing systems are also monitored for how frequently the systems are accessed during a defined period of time and the number of mission critical deployments by each computing system. Collected parameters, including temperature, frequency of access and number of running mission critical applications are imputed into a scoring algorithm which uses the collected parameters and weightings to generate a ranking of computing systems to shutdown sequentially in response to rising temperatures. As temperatures rise above a target temperature, computing systems with the highest score (and the lowest amount of risk) are shutdown, lowering the temperature of the datacenter until the cooling system becomes operation again.


