Dynamic Clock Frequency Management for GPU Power Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High Performance Computing (HPC) server graphics processing units (GPUs) face challenges in dynamically optimizing performance across varying workloads due to static tuning assumptions, which do not account for workload behavior changes over time, and require dynamic allocation of power between I/O and compute subsystems based on runtime telemetry.

Innovation Solution

A system that monitors runtime telemetry to determine if a GPU is I/O bounded or compute bounded, allowing for dynamic allocation of power between the I/O subsystem and the compute subsystem, adjusting clock frequencies of I/O and compute chiplets based on workload characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If static tuning is used to configure the system for each type of workload, then performance is improved under specific workload conditions, but the system cannot adapt when workload behavior changes over time

Engineering Contradiction:
ImproveperformanceVSAvoidadaptability to workload changes
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic frequency adjustment for I/O and compute circuitries based on runtime telemetry data. The system continuously monitors workload characteristics and adjusts clock frequencies accordingly, transitioning from static to dynamic configuration. This allows the system to adapt to changing workload conditions in real-time, resolving the contradiction between optimized performance for specific workloads and adaptability to workload changes.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system employs feedback mechanisms by monitoring runtime telemetry to determine whether the GPU is I/O bounded or compute bounded. Based on this feedback, the system dynamically calculates and adjusts power distribution to I/O and compute subsystems, enabling continuous optimization of performance as workload conditions change.

Inventive Principle:
Principle #23Feedback

2Productivity

If power is allocated statically to I/O and compute subsystems, then power management is simplified, but performance optimization under varying workloads is limited

Engineering Contradiction:
Improveperformance optimizationVSAvoidpower allocation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system uses runtime telemetry feedback to dynamically determine whether the GPU is I/O bounded or compute bounded, automatically adjusting power allocation between subsystems. This feedback-driven approach enables performance optimization under varying workloads while keeping the control logic relatively simple, as the system automatically adapts based on monitored conditions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service by autonomously monitoring its own performance characteristics through telemetry and automatically adjusting power distribution without external intervention. The GPU driver calculates the distribution of power to I/O and compute subsystems based on runtime conditions, enabling the system to optimize its own performance dynamically.

Inventive Principle:
Principle #25Self-service

3Speed

If clock frequencies are increased for both I/O and compute chiplets, then processing speed is improved, but power consumption increases beyond power limits

Engineering Contradiction:
Improveprocessing speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by selectively adjusting clock frequencies of specific circuitries (I/O or compute) based on their current performance needs. Instead of uniformly increasing frequencies across all components, the system dynamically allocates higher frequencies only to the subsystem that is currently performance-bounded, while maintaining lower frequencies in other subsystems. This localized frequency adjustment optimizes processing speed while constraining overall power consumption within limits.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes operational parameters (clock frequencies) of different circuitries based on runtime conditions. By calculating the distribution of power and adjusting frequencies according to whether the system is I/O bounded or compute bounded, the system optimizes processing speed while maintaining power consumption within acceptable limits.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240053789A1Clock frequency management of multiple circuitries
Publication Date: 2024.02.15 INTEL CORP
  • US20240053789A1 patent drawing
  • US20240053789A1 patent drawing
  • US20240053789A1 patent drawing

AI summary

A system that includes first circuitries to operate at a first clock frequency, second circuitries to operate at a second clock frequency, and circuitry to adjust the first and second clock frequencies. In some examples, the circuitry is to selectively adjust the first and second clock frequencies provided to the respective first circuitries and the second circuitries according to a target ratio based on temperature and power consumption of the first circuitries and the second circuitries, wherein the target ratio is based on clock frequencies of the first circuitries and the second circuitries, stall time of the first circuitries, and dynamic capacitance of the first circuitries and the second circuitries.