Graphics Core Command Partitioning for Multi-Chiplet Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributing instructions across multiple chiplets in processing systems increases the likelihood of clock and voltage domain crossings, leading to increased complexity, reduced efficiency, and scalability limitations due to bottlenecks and the need for additional processing to mitigate these crossings.
Innovation Solution
A processing system with multiple graphics cores on distinct dies, each equipped with a command processor and front-end circuitry, divides command packets into partitions and assigns them to individual cores for independent execution, using synchronization circuitry to maintain core synchronization with minimal communication, reducing the need for inter-core connections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple chiplets are used to distribute processing tasks, then processing capability is improved, but clock domain crossing and voltage domain crossing likelihood increases
Solution Approach 1:
The patent segments the processing system into multiple chiplets, each with its own graphics core and command processor. This segmentation allows independent processing on each chiplet while maintaining separate clock and voltage domains, thereby improving processing capability without proportionally increasing domain crossing complexity
Solution Approach 2:
The patent introduces an intermediary command distribution mechanism that efficiently distributes commands to multiple chiplets without requiring complex synchronization infrastructure. This intermediary approach reduces the need for extensive clock and voltage domain crossing mitigation circuitry
2Ease of operation
If a central command processor distributes instructions to multiple chiplets, then instruction distribution is achieved, but system complexity increases
Solution Approach 1:
The patent segments the command processing function by giving each chiplet its own command processor. This eliminates the need for a single complex central command processor and reduces overall system complexity by distributing intelligence throughout the system
Solution Approach 2:
The patent implements command processors on each chiplet as copies of the same functional unit. This copying approach simplifies the overall system architecture compared to a centralized processor, as each chiplet can independently process commands without requiring complex inter-chiplet communication infrastructure
3Productivity
If multiple chiplets are supported by a central command processor, then processing throughput is improved, but bottlenecks increase
Solution Approach 1:
The patent segments the command processing function across multiple independent chiplets, each with its own command processor. This eliminates the bottleneck of a single central command processor that would need to service multiple chiplets sequentially, allowing parallel command processing without creating new bottlenecks
Solution Approach 2:
The patent implements a dynamic command distribution system where commands can be efficiently routed to any chiplet without traversing a complex centralized hierarchy. This dynamic approach maintains high throughput by adapting command distribution to minimize wait times and avoid bottlenecks
Data Source
AI summary
A processing system includes two or more graphics cores each disposed on respective dies and configured for concurrent processing of command packets. To this end, the processing system is configured to determine two or more command partitions associated with a command packet and to assign each command partition to a graphics core. Each graphics core then executes the same command packet by only performing instructions of the command packet associated with the command partitions assigned to the graphics core. Further, after executing an instructions of the command packet based on one or more assigned partitions, each graphics core adjusts one or more counters used to synchronize the execution of the command packet across the graphics cores.


