Datacenter Thermal Management via Risk-Based Hardware Shutdown

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Datacenters face challenges in managing temperature when cooling systems malfunction, leading to excessive heat that can damage computing equipment and disrupt operations, as existing methods lack efficient prioritization strategies for shutting down hardware units to maintain optimal temperatures.

Innovation Solution

A computer-implemented method that uses thermal imaging data, access frequency, and mission critical application information to calculate a weighted score for each hardware unit, ordering them for sequential shutdown to reduce datacenter temperature effectively when cooling systems fail.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If cooling systems are used to maintain datacenter temperature, then equipment operates in healthy state, but system complexity increases and energy consumption rises

Engineering Contradiction:
Improveequipment operation reliabilityVSAvoidcooling system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent pre-calculates risk scores for each hardware unit based on thermal imaging data, access frequency, and mission critical application information before cooling failure occurs. This preliminary risk assessment enables immediate prioritized shutdown decisions when cooling systems fail, resolving the contradiction by preparing mitigation strategies in advance rather than adding complex real-time control systems

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses existing infrastructure components (thermal imaging cameras, network resource monitors, application deployment data) that already serve other purposes to also provide cooling risk assessment functionality. This self-service approach eliminates the need for dedicated complex cooling management systems while maintaining equipment reliability

Inventive Principle:
Principle #25Self-service

2Temperature

If hardware units are shut down to reduce temperature, then temperature control is achieved, but loss of service and productivity occur

Engineering Contradiction:
Improvedatacenter temperatureVSAvoidhardware unit availability
Core Design Contradiction:
TemperatureVSProductivity

Solution Approach 1:

The patent applies differentiated shutdown priorities to different hardware units based on their specific risk scores. Instead of uniform shutdown protocols, each hardware unit receives localized treatment according to its thermal characteristics, access frequency, and mission critical status, achieving temperature control while minimizing overall productivity loss

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes the operational status parameter of hardware units from active to shutdown based on calculated risk scores and current temperature conditions. This parameter-based control approach optimizes the balance between temperature management and service continuity by shutting down only the least critical units first

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If thermal imaging and risk assessment systems are implemented, then prioritized shutdown strategy is achieved, but system complexity and measurement requirements increase

Engineering Contradiction:
Improveprioritized shutdown operationVSAvoidtemperature management system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent creates a multi-functional risk assessment system that uses thermal imaging data, network resource information, and application deployment data to simultaneously evaluate multiple risk factors. This universal approach consolidates what could be separate complex systems into a single integrated framework, achieving ease of operation through unified risk-based control

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system continuously monitors thermal imaging data, access frequency, and application status to calculate updated risk scores for each hardware unit. This feedback mechanism enables dynamic adjustment of shutdown priorities based on real-time conditions, simplifying operational decision-making through automated prioritization while managing complexity through systematic data collection

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12144155B2Datacenter temperature control and management
Publication Date: 2024.11.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12144155B2 patent drawing
  • US12144155B2 patent drawing
  • US12144155B2 patent drawing

AI summary

Systems, methods and/or computer program products for managing the temperature of datacenter during a period of malfunction or inoperability of the cooling system responsible for maintaining temperatures within the datacenter. Temperatures of the datacenter's computing systems are monitored by thermal imaging systems and/or sensors. Computing systems are also monitored for how frequently the systems are accessed during a defined period of time and the number of mission critical deployments by each computing system. Collected parameters, including temperature, frequency of access and number of running mission critical applications are imputed into a scoring algorithm which uses the collected parameters and weightings to generate a ranking of computing systems to shutdown sequentially in response to rising temperatures. As temperatures rise above a target temperature, computing systems with the highest score (and the lowest amount of risk) are shutdown, lowering the temperature of the datacenter until the cooling system becomes operation again.