Dynamic Power Limit Management in Disaggregated Data Centers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In data centers with disaggregated architectures, resources like compute devices and storage devices are physically separate, leading to challenges in meeting quality of service targets due to power usage limits, which can hamper performance and thermal management, especially in distributed data storage systems where a single resource's limitations affect the entire cluster.

Innovation Solution

The implementation of a system that dynamically manages power usage limits across resources using a pod manager and power management logic units, allowing for temporary adjustments to power consumption based on workload demands, enabling resources to either increase or decrease power usage to maintain performance and thermal balance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Temperature

If hard power usage limits are imposed on resources to control thermal conditions and wear, then thermal management and resource longevity are improved, but quality of service targets cannot be met when additional power is needed

Engineering Contradiction:
Improvethermal conditionsVSAvoidquality of service target
Core Design Contradiction:
TemperatureVSReliability

Solution Approach 1:

The system dynamically adjusts power usage limits based on real-time conditions. Resources can temporarily exceed their hard power limits when needed to meet QoS targets, and the system coordinates these increases across the disaggregated architecture to maintain thermal balance. This transforms the static hard limit into a dynamic soft limit that adapts to workload demands.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The pod manager acts as an intermediary between individual resources and the overall system thermal management. It receives QoS target information, coordinates power limit increases across multiple resources, and ensures that thermal conditions are maintained system-wide. This mediator enables resources to temporarily exceed local power limits while maintaining global thermal balance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If power usage limits are increased to meet quality of service targets, then performance is improved, but thermal conditions and resource wear are worsened

Engineering Contradiction:
Improvequality of service targetVSAvoidthermal conditions
Core Design Contradiction:
ReliabilityVSTemperature

Solution Approach 1:

The system continuously monitors thermal conditions and power usage across all resources. When a resource increases its power usage to meet QoS targets, the pod manager receives feedback about the resulting thermal conditions and coordinates compensatory actions with other resources to maintain overall thermal balance, preventing excessive heat accumulation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The pod manager coordinates power limit adjustments across the disaggregated architecture, allowing individual resources to increase power usage when needed while ensuring that the overall system maintains thermal balance. It acts as a mediator that balances local performance needs with global thermal constraints.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If power usage limits are fixed to simplify management, then device complexity is reduced, but adaptability to varying workload demands is limited

Engineering Contradiction:
Improvepower management complexityVSAvoidworkload demand adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The pod manager provides a universal power management mechanism that handles multiple functions: coordinating power limit increases across resources, monitoring thermal conditions, and ensuring QoS targets are met. This single multi-functional component manages the complexity of dynamic power adjustment across the entire disaggregated system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11537191B2Technologies for providing advanced management of power usage limits in a disaggregated architecture
Publication Date: 2022.12.27 INTEL CORP
  • US11537191B2 patent drawing
  • US11537191B2 patent drawing
  • US11537191B2 patent drawing

AI summary

Technologies for providing advanced management of power usage limits in a disaggregated architecture include a compute device. The compute device includes circuitry configured to execute operations associated with a workload in a disaggregated system. The circuitry is also configured to determine whether a present power usage of the compute device is within a predefined range of a power usage limit assigned to the compute device. Additionally, the circuitry is configured to send, to a device in the disaggregated system and in response to a determination that the present power usage of the present compute device is not within the predefined range of the power usage limit assigned to the present compute device, offer data indicative of an offer to reduce the power usage limit assigned to the present compute device to enable a second power utilization limit of another compute device in the disaggregated system to be increased.