GPU Power Limit Tuning for Thermal-Efficient Workload Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed systems with large numbers of GPUs face challenges in achieving efficient resource management and scalability due to high power consumption and temperature issues, leading to thermal trips and hardware degradation, while operating at lower utilization results in decreased performance.

Innovation Solution

A system that adjusts GPU power consumption and workload distribution in real-time using telemetry data to balance power, performance, and thermal profiles, employing an adaptive power tuner and workload tuner to optimize resource management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If GPU power limits are increased to improve performance, then productivity increases, but use of energy and temperature increase leading to thermal trips and hardware degradation

Engineering Contradiction:
Improveworkload processing performanceVSAvoidGPU power consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts GPU power limits in real-time based on monitored workload performance and thermal conditions. The power limit configuration is not static but changes adaptively during operation, allowing the system to optimize between performance and energy consumption by responding to actual runtime conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements a feedback mechanism that continuously monitors telemetry data including workload performance metrics and thermal conditions. Based on this feedback, the system automatically adjusts power limit configurations to maintain optimal performance while preventing thermal trips and hardware degradation through closed-loop control.

Inventive Principle:
Principle #23Feedback

2Productivity

If GPU power limits are increased to improve performance, then productivity increases, but temperature increases leading to thermal trips and hardware degradation

Engineering Contradiction:
Improveworkload processing performanceVSAvoidGPU operating temperature
Core Design Contradiction:
ProductivityVSTemperature

Solution Approach 1:

The system dynamically adjusts GPU power limits in real-time based on monitored workload performance and thermal conditions. The power limit configuration is not static but changes adaptively during operation, allowing the system to optimize between performance and thermal management by responding to actual runtime conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements a feedback mechanism that continuously monitors telemetry data including workload performance metrics and thermal conditions. Based on this feedback, the system automatically adjusts power limit configurations to maintain optimal performance while preventing thermal trips and hardware degradation through closed-loop control.

Inventive Principle:
Principle #23Feedback

3Use of energy by moving object

If GPU utilization is reduced to decrease power consumption, then use of energy decreases, but performance decreases

Engineering Contradiction:
ImproveGPU power consumptionVSAvoidworkload processing performance
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The system changes the power limit parameter dynamically based on workload characteristics and system conditions. By adjusting this key parameter, the system optimizes the balance between power consumption and performance, allowing higher utilization when performance is critical and lower utilization when energy efficiency is prioritized.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12591287B2Workload based graphics processing unit (GPU) performance adjustment for energy efficiency
Publication Date: 2026.03.31 DELL PROD LP
  • US12591287B2 patent drawing
  • US12591287B2 patent drawing
  • US12591287B2 patent drawing

AI summary

Techniques for improving energy efficiency in a distributed system including a plurality of processing units are described. One example method includes configuring an initial power limit for each processing unit in the distributed system; initiating processing of a workload by plurality of processing units in the distributed system, wherein the workload is associated with a target workload performance level; and during the processing of the workload: identifying a peak workload performance level associated with the processing of the workload based on telemetry data received from the plurality of processing units in the distributed system; determining that the peak workload performance level is less than the target workload performance level; and in response to determining that the peak workload performance level is less than the target workload performance level, configuring an increased power limit for at least a portion of the processing units in the distributed system.