Tile-Based Graphics Processor With Hierarchical Bounding Boxes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing tile-based graphics processors face performance bottlenecks and scalability issues due to the serial generation and writing out of primitive lists, which hinder efficient rendering and scaling to different tiling performance levels.

Innovation Solution

A tile-based graphics processor that generates a multi-level hierarchical bounding box data structure using atomic operations, allowing parallel processing in the first pass, and uses this data in a second pass to determine and process primitives for rendering tiles, thereby eliminating the need for serial primitive list generation and writing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If serial primitive list generation and writing is used, then device complexity is reduced, but productivity deteriorates due to performance bottlenecks

Engineering Contradiction:
Improverendering throughputVSAvoidprocessing circuit complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The primitive list generation process is segmented into multiple independent backend circuits (first backend circuit, second backend circuit, etc.), each capable of generating bounding box information for different subsets of primitives simultaneously. This segmentation enables parallel processing while maintaining manageable complexity in each individual backend circuit.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention transitions from serial (one-dimensional) primitive list generation to parallel (multi-dimensional) processing by introducing multiple backend circuits that operate simultaneously. Each backend circuit processes a different portion of the primitive data in parallel, effectively adding a temporal dimension to the processing throughput.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If parallel processing is implemented, then productivity improves, but device complexity increases due to multiple backend circuits

Engineering Contradiction:
Improvebounding box generation speedVSAvoidnumber of backend circuits
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The parallel processing architecture is achieved through segmentation into multiple identical or modular backend circuits. Each circuit is a simplified unit that generates bounding box information for a specific subset of primitives, reducing the complexity burden on any single circuit while enabling collective parallel throughput.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each backend circuit is designed as a universal processing unit capable of handling primitive data and generating bounding box information. The circuits are functionally identical or similar, allowing for easy scaling and reducing design complexity compared to creating entirely different processing units for each parallel channel.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If serial primitive list writing is used, then ease of operation is maintained, but loss of time increases due to bottleneck delays

Engineering Contradiction:
Improveprimitive list generation timeVSAvoiddata structure management simplicity
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The primitive data set is divided into multiple segments, each assigned to a different backend circuit for parallel bounding box generation. This segmentation reduces the time each circuit needs to process its portion, eliminating the serial bottleneck while the hierarchical data structure manages the distributed results efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Bounding box information is generated and stored in a hierarchical data structure in advance of the actual rendering operation. This preliminary parallel generation of spatial information allows the rendering stage to proceed more efficiently by having spatial relationships pre-computed and organized, reducing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12602844B2Graphics processor
Publication Date: 2026.04.14 ARM LTD
  • US12602844B2 patent drawing
  • US12602844B2 patent drawing
  • US12602844B2 patent drawing

AI summary

A tile-based graphics processor performs first and second processing passes to generate a render output. The first processing pass generates and writes out information representative of a set of bounding boxes, and the second processing pass uses the bounding box information to determine which primitives to process for which rendering tiles.