Graphics Core Command Partitioning for Multi-Chiplet Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributing instructions across multiple chiplets in processing systems increases the likelihood of clock and voltage domain crossings, leading to increased complexity, reduced efficiency, and scalability limitations due to bottlenecks and the need for additional processing to mitigate these crossings.

Innovation Solution

A processing system with multiple graphics cores on distinct dies, each equipped with a command processor and front-end circuitry, divides command packets into partitions and assigns them to individual cores for independent execution, using synchronization circuitry to maintain core synchronization with minimal communication, reducing the need for inter-core connections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple chiplets are used to distribute processing tasks, then processing capability is improved, but clock domain crossing and voltage domain crossing likelihood increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidclock domain crossing and voltage domain crossing
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the processing system into multiple chiplets, each with its own graphics core and command processor. This segmentation allows independent processing on each chiplet while maintaining separate clock and voltage domains, thereby improving processing capability without proportionally increasing domain crossing complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary command distribution mechanism that efficiently distributes commands to multiple chiplets without requiring complex synchronization infrastructure. This intermediary approach reduces the need for extensive clock and voltage domain crossing mitigation circuitry

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If a central command processor distributes instructions to multiple chiplets, then instruction distribution is achieved, but system complexity increases

Engineering Contradiction:
Improveinstruction distributionVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the command processing function by giving each chiplet its own command processor. This eliminates the need for a single complex central command processor and reduces overall system complexity by distributing intelligence throughout the system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements command processors on each chiplet as copies of the same functional unit. This copying approach simplifies the overall system architecture compared to a centralized processor, as each chiplet can independently process commands without requiring complex inter-chiplet communication infrastructure

Inventive Principle:
Principle #26Copying

3Productivity

If multiple chiplets are supported by a central command processor, then processing throughput is improved, but bottlenecks increase

Engineering Contradiction:
Improveprocessing throughputVSAvoidbottlenecks
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the command processing function across multiple independent chiplets, each with its own command processor. This eliminates the bottleneck of a single central command processor that would need to service multiple chiplets sequentially, allowing parallel command processing without creating new bottlenecks

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a dynamic command distribution system where commands can be efficiently routed to any chiplet without traversing a complex centralized hierarchy. This dynamic approach maintains high throughput by adapting command distribution to minimize wait times and avoid bottlenecks

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250111461A1Concurrent processing of command partitions using groups of graphics cores
Publication Date: 2025.04.03 ADVANCED MICRO DEVICES INC
  • US20250111461A1 patent drawing
  • US20250111461A1 patent drawing
  • US20250111461A1 patent drawing

AI summary

A processing system includes two or more graphics cores each disposed on respective dies and configured for concurrent processing of command packets. To this end, the processing system is configured to determine two or more command partitions associated with a command packet and to assign each command partition to a graphics core. Each graphics core then executes the same command packet by only performing instructions of the command packet associated with the command partitions assigned to the graphics core. Further, after executing an instructions of the command packet based on one or more assigned partitions, each graphics core adjusts one or more counters used to synchronize the execution of the command packet across the graphics cores.