Auto Tuning Data Center Cooling Based on Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data centers face inefficiencies in cooling high-performance computing devices due to manual temperature control methods, which are reactive and inefficient, especially when dealing with varying workloads across different plots of servers, leading to increased costs and potential hardware damage from overheating.

Innovation Solution

Implementing an automated environmental control system that uses monitored temperatures and upcoming workload data to proactively adjust ambient conditions, such as temperature and humidity, across plots of servers, allowing for dynamic cooling management without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual temperature control methods are used, then simplicity of system is maintained, but cooling efficiency deteriorates and energy consumption increases

Engineering Contradiction:
Improvecooling efficiencyVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by stationary object

Solution Approach 1:

The system proactively adjusts cooling before heat buildup occurs by monitoring workload metrics (CPU utilization, memory usage, I/O operations) and predicting future heat generation. This allows the environmental control system to prepare cooling capacity in advance, responding to computational demands before they manifest as thermal issues, thereby improving cooling efficiency while avoiding energy waste from reactive over-cooling.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors both workload metrics and temperature data, creating a closed-loop feedback mechanism. The environmental control system receives real-time information about computational demands and actual thermal conditions, then dynamically adjusts cooling parameters. This feedback-driven approach ensures cooling efficiency matches actual needs, preventing both overheating and unnecessary energy consumption from excessive cooling.

Inventive Principle:
Principle #23Feedback

2Speed

If manual temperature control methods are used, then system complexity is reduced, but response time to workload changes deteriorates

Engineering Contradiction:
Improveresponse timeVSAvoidsystem complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system proactively adjusts cooling before heat buildup occurs by monitoring workload metrics (CPU utilization, memory usage, I/O operations) and predicting future heat generation. This allows the environmental control system to prepare cooling capacity in advance, responding to computational demands before they manifest as thermal issues, thereby improving cooling efficiency while avoiding energy waste from reactive over-cooling.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors both workload metrics and temperature data, creating a closed-loop feedback mechanism. The environmental control system receives real-time information about computational demands and actual thermal conditions, then dynamically adjusts cooling parameters. This feedback-driven approach ensures cooling efficiency matches actual needs, preventing both overheating and unnecessary energy consumption from excessive cooling.

Inventive Principle:
Principle #23Feedback

3Reliability

If proactive cooling adjustment is implemented, then hardware protection is improved, but system complexity increases

Engineering Contradiction:
Improvehardware protectionVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables the data center environment to self-regulate by automatically translating workload metrics into environmental control decisions without human intervention. The workload monitoring system and environmental control system work together autonomously, with the former detecting computational demands and the latter executing appropriate cooling adjustments. This self-service capability improves hardware protection while managing complexity through automation rather than manual processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system introduces a workload monitoring system as an intermediary between computational tasks and environmental control. This intermediary layer translates abstract workload metrics (CPU utilization, memory usage) into concrete thermal management decisions, bridging the gap between computational demands and physical cooling requirements. This intermediary approach protects hardware by ensuring appropriate cooling while managing system complexity through a dedicated translation layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230141508A1Methods and apparatus for auto tuning cooling of compute devices based on workloads
Publication Date: 2023.05.11 INTEL CORP
  • US20230141508A1 patent drawing
  • US20230141508A1 patent drawing
  • US20230141508A1 patent drawing

AI summary

Methods, apparatus, systems, and articles of manufacture are disclosed to auto tune data center cooling based on workloads. Disclosed herein is an apparatus including processor circuitry to determine environmental conditions setpoint(s) based on at least one of a future workload status or a current workload status of devices within a data center.