ML-Based Power Manager for Data Center Node Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face significant energy consumption challenges due to constant processor core activity, even during low network traffic, leading to wasted power and increased carbon footprint, which complicates meeting green energy goals and service level agreements.
Innovation Solution
Implementing an adaptive power optimizer that uses machine learning to predict traffic criticality across nodes in a data center cluster, allowing for dynamic adjustment of processor core frequency, pausing, or changing power states based on workload and network function metrics to reduce power usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If direct polling is used with dedicated processor cores to ensure high network function performance, then network performance is improved, but power consumption increases due to continuous full-capacity operation regardless of traffic load
Solution Approach 1:
The patent applies dynamics by transitioning from static full-capacity processor operation to dynamic power adjustment. The system continuously monitors network traffic conditions and adjusts processor frequency and power states accordingly, allowing processors to operate at high performance when needed and reduce power consumption when traffic is low, thus resolving the contradiction between performance and energy usage
Solution Approach 2:
The patent changes the operating parameters of processor cores by adjusting frequency and power states based on traffic conditions. When traffic is heavy, processors run at higher frequencies for optimal performance; when traffic is light, processors reduce frequency and enter power-saving states, directly addressing the contradiction between maintaining performance and reducing energy consumption
2Reliability
If dedicated processor cores run at full capacity continuously to handle network functions, then network reliability is maintained, but energy waste increases due to empty polls during low traffic periods
Solution Approach 1:
The system implements feedback by continuously monitoring network traffic conditions and using this information to adjust processor power states. The feedback loop ensures that processors maintain sufficient capacity for network functions while avoiding empty polls during low traffic periods, thus maintaining reliability while reducing energy waste
Solution Approach 2:
The patent applies preliminary action by predicting future traffic patterns and proactively adjusting processor power states before actual traffic changes occur. This allows the system to maintain network reliability by having processors ready when needed while avoiding energy waste by reducing power consumption during predicted low-traffic periods
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Example systems and techniques are disclosed for power management. An example system includes one or more memories and one or more processors. The one or more processors are configured to obtain workload metrics from a plurality of nodes of a cluster. The one or more processors are configured to obtain network function metrics from the plurality of nodes of the cluster. The one or more processors are configured to execute at least one machine learning model to predict a corresponding measure of criticality of traffic of each node. The one or more processors are configured to determine, based on the corresponding measure of criticality of traffic of each node, a corresponding power mode for at least one processing core of each node. The one or more processors are configured to recommend or apply the corresponding power mode to the at least one processing core of each node.