Dynamic Power Profile Scheduling for Data Center Stranded Capacity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-performance computing (HPC) data centers face inefficiencies due to overprovisioning of power, leading to 'stranded power capacity' where actual power consumption is lower than provisioned, resulting in suboptimal resource utilization and increased costs.
Innovation Solution
Creating power profiles for applications to determine peak power consumption rates and strategically scheduling jobs to limit power usage within these profiles, allowing freed-up power to be utilized elsewhere in the data center, thereby reducing stranded power capacity without impacting performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If power capacity is overprovisioned to meet peak demand, then reliability is improved, but use of energy by stationary object worsens due to stranded power capacity
Solution Approach 1:
The system dynamically adjusts power allocation by creating application-specific power profiles that capture peak power consumption rates. The job scheduler dynamically decides which jobs to execute based on current power capacity availability, transforming the static overprovisioned power system into a dynamic one that adapts to actual workload demands, thereby reducing stranded power capacity while maintaining reliability
Solution Approach 2:
The system implements feedback mechanisms by monitoring actual power consumption of applications and using this information to create accurate power profiles. These profiles feed back into the job scheduling decisions, allowing the system to learn from past consumption patterns and optimize future power allocation, reducing the gap between provisioned and actual power usage
2Productivity
If power is overprovisioned to ensure sufficient capacity, then productivity is improved, but device complexity increases due to power management overhead
Solution Approach 1:
Instead of directly managing complex real-time power constraints, the system creates simplified power profiles that copy and represent the essential power consumption characteristics of applications. These profiles serve as abstracted models that enable straightforward scheduling decisions without requiring complex real-time power analysis, thus maintaining productivity while reducing management complexity
Solution Approach 2:
The system performs preliminary power profile creation and analysis before job scheduling decisions are made. By pre-characterizing application power consumption patterns and pre-calculating compatible job combinations within power constraints, the system avoids complex real-time calculations during scheduling, reducing operational complexity while preserving productivity
3Loss of energy
If power allocation is optimized to reduce stranded capacity, then use of energy by stationary object is improved, but measurement precision requirements increase
Solution Approach 1:
The system measures power consumption at sufficient precision to capture peak rates needed for scheduling decisions, rather than requiring continuous high-precision measurement at all times. By focusing measurement efforts on critical moments when power consumption peaks occur, the system achieves adequate measurement precision for optimization purposes without the overhead of continuous high-precision monitoring
Data Source
AI summary
Systems and methods described herein make previously stranded power capacity (power that is provisioned for a data center according to a computing system's nameplate power consumption but is currently not useable) available to the data center. Systems described herein generate empirical power profiles that specify expected upper bounds for the power consumption levels that applications trigger. Using the upper bounds for application power-consumption levels, a computing system described herein can reliably release part of its provisioned nameplate power for other systems or data center consumers, reducing the amount of stranded power in a data center. The method described herein avoids performance penalties for most jobs by using sensor measurements made at a rapid rate explained herein to ensure that a system power cap based on running application's measured peak power consumption is reliable with reference to the power capacitance inherent in the computing system.


