GPU Separate Clocking for Workload-Based Power and Thermal Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional single/universal clocking schemes in GPUs fail to account for varying processing workloads across different components, leading to inefficient thermal and power performance due to uniform clock frequency application.
Innovation Solution
Implementing separate clocking for shader engine modules and non-shader-engine modules within a GPU using programmable dividers, adjusting frequencies based on performance counter data to optimize thermal and power efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If a single universal clocking scheme is used for all GPU components, then device complexity is reduced and ease of operation is improved, but thermal performance and power efficiency deteriorate due to inability to scale frequencies independently
Solution Approach 1:
The GPU is divided into multiple component groups (e.g., shader engine modules and non-shader-engine modules) with separate clocking control. Each group can have its clock frequency independently scaled based on workload demands, allowing power-efficient operation by reducing frequencies of underutilized components while maintaining high frequencies for actively used components.
Solution Approach 2:
The clocking scheme transitions from static uniform frequency to dynamic independent frequency control. Monitor circuits continuously track workload characteristics and adjust clock frequencies in real-time based on actual component utilization, enabling the system to adapt power consumption to actual performance requirements.
2Productivity
If clock frequency is uniformly scaled for all components, then ease of operation is maintained, but productivity deteriorates because high-frequency components cannot operate at maximum speed when needed
Solution Approach 1:
The GPU components are segmented into separate clockable groups with independent frequency control. This allows each group to operate at the optimal frequency for its current workload, maximizing overall processing throughput without requiring complex manual configuration.
Solution Approach 2:
Monitor circuits automatically detect workload characteristics and trigger appropriate clock frequency scaling for each component group without external intervention. The system self-regulates to maintain optimal performance across different workload scenarios.
3Temperature
If separate independent clocking is implemented for different component groups, then power efficiency and thermal performance improve, but device complexity increases due to additional control circuitry
Solution Approach 1:
The GPU is segmented into independently clocked component groups, allowing thermal management through selective frequency scaling. Components generating excessive heat can have their frequencies reduced independently, improving thermal performance without affecting overall system operation.
Solution Approach 2:
Monitor circuits act as intermediaries between workload sources and clock control mechanisms. These monitors detect workload characteristics and automatically trigger appropriate clock scaling decisions, eliminating the need for complex external control systems while achieving optimal thermal management.
4Use of energy by moving object
If clock frequency scaling is applied uniformly, then power consumption is reduced during low workload, but productivity suffers because high-frequency operation cannot be maintained for intensive tasks
Solution Approach 1:
The clocking system dynamically adjusts frequencies of individual component groups based on real-time workload monitoring. During low-intensity tasks, frequencies are reduced to save power; during intensive tasks, frequencies are maintained or increased to maximize performance, with transitions handled automatically by monitor circuits.
Data Source
AI summary
Systems and methods related to controlling clock signals for clocking shader engines modules (SEs) and non-shader-engine modules (nSEs) of a graphics processing unit (GPU) are provided. One or more dividers receive a clock signal CLK and output a clock signal CLKA to the SEs and output a clock signal CLKB to the nSEs. The frequencies of CLKA and CLKB are independently selected based on sets of performance counter data monitored at the SEs and nSEs, respectively. The clock signal frequency for either the SEs or the nSEs is reduced when the corresponding sets of performance counter data indicates a comparatively lower processing workload for the SEs or for the nSEs.


