SLA-Based Accelerator Power Throttling Under Data Center Power Caps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies fail to provide granular control over power distribution in data centers, leading to overprovisioned hardware issues, outages, and non-compliance with SLAs, especially in high-computational tasks like AI workflows.
Innovation Solution
Dynamically control power distribution to individual components like GPUs based on service-level agreements (SLAs), throttling lower-priority tasks to maintain power policy limits and ensure high-priority tasks are uninterrupted.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If computational resources are increased to meet growing AI demand, then service capability is improved, but hardware overprovisioning and power consumption increase
Solution Approach 1:
The system dynamically adjusts power allocation to computational resources based on real-time workload demands and service-level agreements. Power limits are not static but adapt to changing conditions, allowing the system to meet AI computational demands while avoiding permanent overprovisioning of hardware capacity.
Solution Approach 2:
The invention changes the power consumption parameter of computational resources by implementing fine-grained power capping at the accelerator level. This allows selective reduction of power limits for specific hardware components based on their current task priority and SLA requirements, rather than uniformly increasing or decreasing all power consumption.
2Ease of operation
If power is distributed uniformly to all components, then hardware utilization is simplified, but granular control over individual components is lost
Solution Approach 1:
The system segments power control into fine-grained allocations at the accelerator level rather than uniform distribution across entire systems or racks. Each accelerator receives individualized power limits based on its specific workload and SLA requirements, enabling precise control over individual components while maintaining manageable complexity through automated policies.
Solution Approach 2:
The invention introduces an intermediary power management layer that sits between the power supply and computational resources. This intermediary layer implements sophisticated power capping logic and SLA enforcement without requiring direct complex control of each individual component, thereby maintaining ease of operation while achieving granular control.
3Speed
If all tasks are allowed to run at full power, then task completion speed is improved, but SLA compliance and hardware lifespan are compromised
Solution Approach 1:
The system applies different power quality characteristics to different tasks and accelerators based on their SLA requirements. High-priority tasks with strict SLAs receive guaranteed power allocation to ensure completion speed, while lower-priority tasks experience dynamic power throttling. This local differentiation of power quality maintains SLA compliance without unnecessarily limiting overall system throughput.
Solution Approach 2:
The invention implements feedback mechanisms that continuously monitor task progress, power consumption, and SLA compliance. Based on this feedback, the system dynamically adjusts power allocation to maintain both task completion speed and SLA adherence. When tasks approach completion or when SLA thresholds are near violation, power is adjusted accordingly to prevent violations while maximizing throughput.
Data Source
AI summary
Various embodiments described herein dynamically control the distribution of power to individual components of a node in an overprovisioned rack, node, or accelerators of a data center based on service-level agreements (SLAs) defining priorities for workloads for certain user accounts. The SLA is used to determine a throttling order for throttling the accelerators. Controlling the distribution of power in the node includes throttling at an accelerator or coprocessor, based on the throttling order or SLA, until the power consumption is at or below a power policy limit. In this manner, various embodiments discussed herein provide (1) granular control over the execution of tasks in an overprovisioned rack and (2) a user experience consistent with a priority level defined by an SLA, while complying with power policy limit(s) to improve the lifespan and operation of hardware, as well as to reduce the wear and tear experienced by overprovisioned hardware.


