Interruptible GPU Context Switching via Run Lists

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current graphics processing units (GPUs) face challenges in efficiently handling multiple graphically intensive applications simultaneously due to inefficient scheduling and lack of mechanisms for precise interruption, leading to bottlenecks when resources are shared among applications.

Innovation Solution

A GPU is configured to be interruptible, allowing it to switch between multiple contexts by creating a run list with ring buffers and command stream processors, enabling efficient processing and memory access management to handle multiple applications concurrently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a GPU executes operations in serialized order as received, then the GPU can process operations sequentially without complex scheduling, but the GPU becomes a bottleneck when multiple applications with differing priorities need to access the same resources

Engineering Contradiction:
Improvescheduling mechanism complexityVSAvoidmulti-application processing throughput
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the processing workflow into distinct phases: command submission phase, command execution phase, and result retrieval phase. Multiple applications can submit commands asynchronously to a command queue while the GPU executes them in order, allowing prioritization and pre-fetching without disrupting the execution sequence. This segmentation enables complex scheduling behavior at the submission level while maintaining simple ordered execution at the processing level.

Inventive Principle:
Principle #1Segmentation

2Productivity

If a GPU is configured for precise interruption and context switching, then the GPU can efficiently handle multiple applications concurrently, but the device complexity increases due to additional hardware components like reorder buffers and extra pipeline stages

Engineering Contradiction:
Improvemulti-application processing throughputVSAvoidhardware architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements pre-fetching mechanisms that anticipate future command needs and prepare data in advance. When the GPU is about to execute a command that requires data from memory, the system proactively fetches that data beforehand, hiding memory latency. This preliminary action allows the GPU to maintain high throughput without requiring complex interruption mechanisms, as the pipeline remains continuously fed with ready data.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If a GPU processes calculation-heavy graphics operations to free the CPU, then the CPU can perform other functions, but the GPU becomes a bottleneck when multiple applications attempt to use the GPU simultaneously

Engineering Contradiction:
ImproveCPU-GPU task distributionVSAvoidGPU resource utilization for multiple applications
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent introduces a command queue and descriptor ring buffer as intermediary structures between the CPU and GPU. Applications submit standardized command descriptors to the queue, which the GPU processes in order. This intermediary layer decouples the CPU from direct GPU control, allowing the CPU to efficiently hand off tasks without managing detailed GPU operations. Multiple applications can interact with this standardized interface simultaneously, improving GPU resource utilization while maintaining CPU simplicity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7580040B2Interruptible GPU and method for processing multiple contexts and runlists
Publication Date: 2009.08.25 VIA TECH INC
  • US7580040B2 patent drawing
  • US7580040B2 patent drawing
  • US7580040B2 patent drawing

AI summary

A graphics processing unit (“GPU”) is configured to interrupt processing of a first context and to initiate processing of a second context upon command so that multiple programs can be executed by the GPU. The CPU creates and the GPU stores a run list containing a plurality of contexts for execution, where each context has a ring buffer of commands and pointers for processing. The GPU initiates processing of a first context in the run list and retrieves memory access commands and pointers referencing data associated with the first context. The GPU's pipeline processes data associated with first context until empty or interrupted. If emptied, the GPU switches to a next context in the run list for processing data associated with that next context. When the last context in the run list is completed, the GPU may switch to another run list containing a new list of contexts for processing.