Configurable Clustered Systolic Array for GPU Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current graphics processing units (GPUs) face limitations in efficiently processing graphics and machine-learning operations due to fixed function computational units and inadequate parallel processing techniques, leading to suboptimal performance in handling diverse workloads.

Innovation Solution

The implementation of a general-purpose GPU with a dynamically configurable systolic array architecture, where systolic units are consolidated and interconnected via high-bandwidth interconnects to enhance data sharing and power efficiency, allowing for unified, dynamically configurable systolic arrays that improve compute operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If fixed function computational units are used in GPUs, then specific graphics operations can be performed, but the system lacks adaptability to handle diverse workloads efficiently

Engineering Contradiction:
Improveworkload handling capabilityVSAvoidcomputational unit configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a unified systolic array architecture where a single configurable computational unit can perform multiple functions including matrix multiplication, vector operations, and other compute tasks. This eliminates the need for separate fixed-function units for different operations, providing versatility while maintaining a relatively simple underlying structure.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The systolic array units are designed to be dynamically configurable through programming, allowing the same hardware structure to adapt its behavior based on the specific workload requirements. This dynamic reconfigurability enables efficient handling of diverse workloads without requiring complex static hardware design.

Inventive Principle:
Principle #15Dynamics

2Productivity

If traditional parallel processing techniques are implemented, then processing throughput can be increased, but data movement overhead and power consumption increase

Engineering Contradiction:
Improveprocessing throughputVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent consolidates multiple systolic units into a unified array structure where adjacent units share common data pathways and computational resources. This merging reduces redundant data movement between separate processing units and minimizes the overall power consumption while maintaining high parallel processing throughput.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The systolic array introduces intermediate processing stages where data can be partially processed within the array before being moved to external memory or other processing units. This reduces the frequency and volume of data transfers to and from high-bandwidth memory, thereby reducing power consumption associated with data movement.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If systolic units are distributed across multiple clusters, then processing capacity increases, but data sharing efficiency and power efficiency decrease

Engineering Contradiction:
Improveprocessing capacityVSAvoiddata movement energy
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent merges systolic units from multiple clusters into a unified systolic array architecture, allowing processing capacity to scale while maintaining efficient data sharing through shared internal pathways. This unified structure eliminates the inefficiencies of distributed data sharing across separate clusters.

Inventive Principle:
Principle #5Merging (Combining)

4Productivity

If high-bandwidth interconnects are implemented for systolic units, then data sharing efficiency improves, but device complexity increases

Engineering Contradiction:
Improvedata sharing efficiencyVSAvoidinterconnect structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines multiple interconnect pathways into a unified high-bandwidth interconnect structure that serves the entire systolic array. This merged interconnect provides efficient data sharing across all systolic units while avoiding the complexity of implementing separate high-bandwidth interconnects for each unit or cluster.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240220448A1Scalable and configurable clustered systolic array
Publication Date: 2024.07.04 INTEL CORP
  • US20240220448A1 patent drawing
  • US20240220448A1 patent drawing
  • US20240220448A1 patent drawing

AI summary

A scalable and configurable clustered systolic array is described. An example of apparatus includes a cluster including multiple cores; and a cache memory coupled with the cluster, wherein each core includes multiple processing resources, a memory coupled with the plurality of processing resources, a systolic array coupled with the memory, and one or more interconnects with one or more other cores of the plurality of cores; and wherein the systolic arrays of the cores are configurable by the apparatus to form a logically combined systolic array for processing of an operation by a cooperative group of threads running on one or more of the plurality of cores in the cluster.