Load-Balanced Tessellation Distribution for Parallel GPU Pipelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In graphics processing units (GPUs) with multiple geometry and setup fixed-function pipelines (GSPs), the requirement for large buffers to maintain in-order processing of triangles and patches leads to inefficiencies, particularly when dealing with varying tessellation rates, causing idle GSPs and queueing issues due to the need for significant data storage and synchronization across multiple pipelines.

Innovation Solution

Implementing a load-balanced tessellation distribution architecture that splits high-tessellation-rate patches into sub-patches, allowing multiple GSPs to work in parallel, thereby reducing buffer requirements and improving workload distribution, with patch splitters and subpatch/patch distribution blocks ensuring even processing across GSP back end blocks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple GSPs work in parallel with strict in-order processing, then processing throughput is improved, but buffer size requirements increase significantly

Engineering Contradiction:
Improveprocessing throughputVSAvoidbuffer size
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent segments the processing architecture into multiple independent GSP pipelines, each capable of autonomous operation. By dividing the monolithic processing unit into parallel segments (GSP0-GSP7), the system achieves higher throughput while each segment maintains its own smaller buffer, eliminating the need for one enormous buffer that would be required for a single processing unit handling all patches sequentially.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of parallelism by processing multiple patches simultaneously across different GSPs rather than sequentially in a single pipeline. This dimensional shift from temporal sequencing to spatial parallelism allows the system to achieve high throughput while keeping individual buffers small, as each GSP handles only its assigned subset of patches.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Quantity of substance

If a single GSP processes all patches sequentially, then buffer requirements are reduced, but processing time increases significantly

Engineering Contradiction:
Improvebuffer sizeVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent divides the workload into segments assigned to multiple GSPs, allowing concurrent processing of different patch subsets. This segmentation enables the system to maintain small buffers per GSP while dramatically reducing total processing time through parallel execution, rather than having one GSP process all patches sequentially with a large buffer.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent ensures continuous useful action by keeping all GSPs actively processing patches simultaneously rather than having them idle or waiting. Each GSP continuously receives and processes its assigned patches, eliminating idle time and maximizing throughput while maintaining small buffer sizes through efficient workload distribution.

Inventive Principle:
Principle #20Continuity of useful action

3Manufacturing precision

If one patch has very high tessellation rate, then processing complexity increases, but this causes idle GSPs and queueing delays in parallel architecture

Engineering Contradiction:
Improvetessellation accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent segments high-tessellation-rate patches into smaller sub-patches and distributes them across multiple GSPs. This segmentation prevents any single GSP from being overloaded with an excessively complex patch, thereby eliminating idle GSPs and queueing delays while maintaining the required tessellation accuracy through coordinated processing of the sub-patches.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic workload distribution where the system adaptively assigns patches and sub-patches to GSPs based on current processing capacity and patch complexity. This dynamic allocation ensures that high-tessellation-rate patches are broken down and distributed to balance the workload, preventing any single GSP from becoming a bottleneck while maintaining processing efficiency and accuracy.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3513373B1Load-balanced tessellation distribution for parallel architectures
Publication Date: 2023.02.22 INTEL CORP
  • EP3513373B1 patent drawingFigure 1
  • EP3513373B1 patent drawingFigure 2
  • EP3513373B1 patent drawingFigure 3

AI summary

Briefly, in accordance with one or more embodiments, an architecture to load balance tessellation distribution apparatus comprises a memory to store one or more patches representing an object in an image, and a processor, coupled to the memory, to perform one or more tessellation operations on the one or more patches. The one or more tessellation operations including splitting one or more of the patches into one or more subpatches, and load balancing the one or more patches and the one or more subpatches among two or more geometry and setup fixed-function pipelines (GSPs).