Phase-Aware Compute Scheduling With Thermal Energy Budgeting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face significant energy inefficiencies due to the rapid growth of energy dissipation in servers and cooling systems, with traditional solutions being reactive and failing to optimize energy use across computing and cooling resources effectively.
Innovation Solution
A holistic approach treating energy as a first-class resource, using fast thermal models and global schedulers to allocate energy budgets dynamically, combined with proactive and reactive control mechanisms to optimize cooling and computing operations, ensuring energy-efficient use of resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If servers are overprovisioned with high capacity, then peak performance is improved, but energy efficiency deteriorates due to operation at low utilization
Solution Approach 1:
The patent implements dynamic performance settings that allow servers to adapt their operational state based on workload conditions. The system transitions between different performance modes (high-performance, balanced, energy-efficient) depending on utilization levels, enabling servers to maintain peak performance when needed while operating efficiently at lower loads
Solution Approach 2:
The system changes operational parameters including CPU frequency, voltage, and performance mode settings based on workload demands. By dynamically adjusting these parameters, the system optimizes the trade-off between performance and energy consumption, preventing servers from operating inefficiently at low utilization while maintaining peak capability when required
2Temperature
If cooling capacity is increased to handle peak loads, then thermal management is improved, but energy consumption deteriorates during low-utilization periods
Solution Approach 1:
The system implements feedback mechanisms that monitor server utilization, workload characteristics, and thermal conditions in real-time. Based on this feedback, the cooling system dynamically adjusts its capacity and operational parameters, reducing cooling energy consumption during low-utilization periods while maintaining adequate thermal management during peak loads
Solution Approach 2:
The cooling system operates dynamically rather than statically, adapting its capacity and operational mode based on actual thermal demands driven by workload conditions. This enables the system to maintain effective thermal management when needed while minimizing energy consumption during low-utilization periods
3Reliability
If reactive energy management is used based on sensor feedback, then response to thermal conditions is improved, but optimization of overall energy efficiency deteriorates due to lack of proactive control
Solution Approach 1:
The system performs preliminary actions by predicting future workload patterns and thermal conditions based on historical data and current trends. It proactively adjusts performance settings and cooling capacity before thermal issues arise or energy inefficiency occurs, rather than merely reacting to sensor feedback after conditions have developed
Solution Approach 2:
The system introduces an intermediary layer of intelligence that processes workload patterns, predicts thermal conditions, and coordinates between computing operations and cooling systems. This intermediary enables holistic energy optimization by considering both computing and cooling energy consumption together rather than separately
Data Source
AI summary
A system, and method for controlling a computing system, comprising: reading a stored energy-performance characteristic of a plurality of different phases of execution of software, an execution of each phase being associated with a consumption of a variable amount of energy in dependence on at least a processing system performance state, the performance state being defined by a selectable performance-energy consumption optimization for at least two processing system components; scheduling a plurality of phases of execution of the software, in dependence on the stored energy-performance characteristics, for each of the respective phases of execution of the software and at least one system-level energy criterion; and executing the phases of execution of the software in accordance with the scheduling.


