Datacenter Load Shedding Using ML-Guided Power Threshold Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional datacenter power management techniques lack insight into workload and customer impacts during load shedding, leading to suboptimal power reduction strategies that can cause unnecessary disruptions and inefficiencies.
Innovation Solution
A method involving a computer system that identifies and manages an aggregate power threshold by monitoring temperature control systems and using machine-learning models to determine power consumption adjustments, enabling dynamic and orchestrated power reduction actions such as power capping, workload migration, and host shutdowns based on real-time conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If manual load shedding is performed by shutting down host racks or individual devices one by one, then power consumption is reduced, but the operator has no insight into the workloads or customers impacted and the extent of effects
Solution Approach 1:
The system implements automated feedback loops where the load shedding manager continuously monitors power consumption, temperature control system status, and workload information. When load shedding is required, the system uses this feedback to intelligently select which hosts to shut down based on minimal impact to customers and workloads, rather than random or manual selection.
Solution Approach 2:
The load shedding manager operates autonomously to manage power consumption without requiring manual operator intervention. It automatically detects when power reduction is needed, determines the optimal hosts to shut down using machine learning models and real-time data, and executes the load shedding actions while minimizing disruption to services.
2Loss of energy
If emergency shut off switch is pressed to turn off power to the whole or large portion of the datacenter, then power consumption is rapidly reduced, but widespread power failure and significant disruption occur
Solution Approach 1:
Instead of shutting down the entire datacenter or large portions at once, the system segments the load shedding into individual host-level actions. The load shedding manager selectively shuts down specific hosts one at a time based on real-time conditions, distributing the power reduction across multiple small increments rather than a single large-scale shutdown.
Solution Approach 2:
The system dynamically adjusts the load shedding strategy based on real-time monitoring of power consumption, temperature control capacity, and workload distribution. The machine learning models continuously learn from operational data to optimize which hosts should be shut down at any given moment, making the power reduction process adaptive rather than static.
3Loss of energy
If conventional load shedding methods are used without considering temperature control system capacity, then power consumption is reduced, but the system fails to account for the relationship between power consumption and cooling requirements
Solution Approach 1:
The load shedding manager integrates multiple functions into a single system: it monitors power consumption, tracks temperature control system capacity, predicts future conditions using machine learning models, and executes load shedding actions. This multi-functional approach ensures that power reduction decisions always consider the interrelationship between computing load and cooling requirements.
Solution Approach 2:
The system performs preliminary analysis using machine learning models to predict future power consumption and temperature control system capacity before making load shedding decisions. This allows the system to proactively adjust power consumption in advance of potential overheating or power failures, rather than reacting after problems occur.
Data Source
AI summary
Disclosed techniques relate to orchestrating power consumption reductions across a number of hosts. A current value for an aggregate power threshold of a plurality of hosts may be identified. During a first time period, an aggregate power consumption of the plurality of hosts may be managed using the current value for the aggregate power threshold. A triggering event indicating a modification to the aggregate power threshold is needed may be detected. A new value for the aggregate power threshold may be determined based on the triggering event. During a second time period, the aggregate power consumption of the plurality of hosts may be managed using the new value for the aggregate power threshold.


