Graphics Processing Unit Execution Units Variable Clock Rate Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing demand for high-performance graphics processing units with a large number of parallel execution units leads to significant costs and semiconductor area consumption, as existing technologies struggle to balance operational efficiency with circuitry complexity and size limitations.

Innovation Solution

The method involves operating a reduced number of processing engines at an increased clock rate, where P execution units are operated using a first clock rate, and Q execution units, with Q being less than P, are operated at a second clock rate twice that of the first, allowing for efficient execution of instructions across multiple execution units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a large number of parallel execution units are used to increase processing capability, then productivity is improved, but device complexity and semiconductor area consumption increase

Engineering Contradiction:
Improveprocessing capabilityVSAvoidcircuitry complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent changes the operational parameters of execution units by implementing variable clock rates. Different execution units operate at different clock frequencies (first clock rate for some units, second clock rate for others), allowing the system to achieve high productivity without proportionally increasing complexity. This parameter change enables flexible resource allocation and optimization.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The execution units are segmented into different groups that can be operated independently at different clock rates. The system divides the processing workload and allocates different clock frequencies to different segments of execution units, allowing optimized performance without requiring all units to operate at maximum speed, thus reducing overall complexity.

Inventive Principle:
Principle #1Segmentation

2Productivity

If a large number of parallel execution units are used to increase processing capability, then productivity is improved, but semiconductor area consumption increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidsemiconductor area
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

By implementing variable clock rates, the system optimizes the utilization of execution units. Not all execution units need to be fully operational at maximum speed simultaneously, so reducing the clock rate of some units decreases their required circuitry size and semiconductor area while maintaining overall processing capability through the faster-operating units.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses a partial subset of execution units operating at high speed rather than requiring all units to operate at full capacity. This partial action approach allows the system to achieve desired productivity with fewer actively engaged units, reducing the total semiconductor area required compared to having all units fully equipped and operational.

Inventive Principle:
Principle #16Partial or excessive action

3Speed

If execution units operate at higher clock rates to increase speed, then productivity is improved, but device complexity increases

Engineering Contradiction:
Improveclock rateVSAvoidcircuitry complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent implements variable clock rates as a key parameter change, allowing different execution units to operate at different frequencies. This enables the system to achieve high speed where needed without uniformly increasing complexity across all units, as lower-clock units can use simplified circuitry designs.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7484076B1Executing an SIMD instruction requiring P operations on an execution unit that performs Q operations at a time (Q<P)
Publication Date: 2009.01.27 NVIDIA CORP
  • US7484076B1 patent drawing
  • US7484076B1 patent drawing
  • US7484076B1 patent drawing

AI summary

Methods, apparatuses, and systems are presented for performing instructions using multiple execution units in a graphics processing unit involving issuing an instruction for P executions of the instruction wherein each execution uses different data, P being a positive integer, the instruction being issued based on a first clock having a first clock rate, operating Q execution units to achieve the P executions of the instruction, Q being a positive integer less than P and greater than one, each of the execution units being operated based on a second clock having a second clock rate higher than the first clock rate of the first clock, and wherein the second clock rate of the second clock is equal to the first clock rate of the first clock multiplied by the ratio P/Q.