Dual ALU Pipeline SIMD Unit Wavefront Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing graphics processing units (GPUs) face inefficiencies when executing wavefronts across multiple arithmetic logic unit (ALU) pipelines, particularly in scenarios where the number of work items exceeds the number of ALUs, leading to extended execution cycles and idle pipelines.

Innovation Solution

The implementation of dual ALU pipeline processing within a single instruction multiple data (SIMD) unit, where two ALU pipelines can execute instructions independently in a single cycle, and a cache system that stores wavefronts to enable simultaneous execution without increasing VGPR bandwidth.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the number of work items in a wavefront exceeds the number of ALUs in an ALU pipeline, then execution extends over more than one execution cycle, but this increases execution time and reduces productivity

Engineering Contradiction:
Improveexecution throughputVSAvoidexecution cycle time
Core Design Contradiction:
ProductivityVSDuration of action of moving object

Solution Approach 1:

The patent divides the execution of wavefronts into multiple independent ALU pipelines. Each pipeline can process a portion of the wavefront simultaneously, segmenting the workload to enable parallel execution and reduce overall execution time

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a second dimension of parallelism by implementing dual ALU pipelines that can execute independently. This transforms single-cycle sequential execution into multi-pipeline parallel execution, effectively adding a dimensional layer to the processing architecture

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If multiple ALU pipelines are used to execute wavefronts, then throughput is improved, but idle pipelines occur when the number of work items does not充分利用 all pipelines

Engineering Contradiction:
Improveprocessing throughputVSAvoididle pipeline resources
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent implements dynamic work item distribution mechanisms that adaptively allocate wavefronts to available ALU pipelines based on current workload. This dynamic scheduling ensures that pipelines remain actively utilized and minimizes idle time across the processing units

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent designs ALU pipelines with universal functionality to handle diverse work item types. This multi-functionality allows any pipeline to process any wavefront, enabling flexible load balancing and reducing the occurrence of idle pipelines when work item distribution is suboptimal

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If dual ALU pipeline processing is implemented, then execution efficiency is improved, but device complexity increases

Engineering Contradiction:
Improveinstruction execution rateVSAvoidALU pipeline architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges the control logic for multiple ALU pipelines into a unified instruction dispatch mechanism. By combining control functions and sharing common resources such as instruction buffers and register files, the architecture achieves dual-pipeline throughput without proportionally increasing overall system complexity

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12299413B2Dual vector arithmetic logic unit
Publication Date: 2025.05.13 ADVANCED MICRO DEVICES INC
  • US12299413B2 patent drawing
  • US12299413B2 patent drawing
  • US12299413B2 patent drawing

AI summary

A processing system executes wavefronts at multiple arithmetic logic unit (ALU) pipelines of a single instruction multiple data (SIMD) unit in a single execution cycle. The ALU pipelines each include a number of ALUs that execute instructions on wavefront operands that are collected from vector general process register (VGPR) banks at a cache and output results of the instructions executed on the wavefronts at a buffer. By storing wavefronts supplied by the VGPR banks at the cache, a greater number of wavefronts can be made available to the SIMD unit without increasing the VGPR bandwidth, enabling multiple ALU pipelines to execute instructions during a single execution cycle.