Multi-Context Thread Distribution in GPUs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As integrated circuit fabrication advances, increased functionality on a single chip leads to higher heat generation and power consumption, limiting device efficiency, longevity, and usage models, particularly for battery-powered devices, and existing parallel graphics processing techniques face inefficiencies in thread distribution and power management.

Innovation Solution

Implementing efficient multi-context thread distribution techniques in graphics processing units (GPUs) that allocate work efficiently across processing clusters, using intelligent scheduling algorithms and streaming microprocessors to optimize thread execution and reduce power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If additional functionality is integrated onto a single IC chip, then processing capability is improved, but heat generation and power consumption increase

Engineering Contradiction:
Improveprocessing capabilityVSAvoidheat generation
Core Design Contradiction:
ProductivityVSTemperature

Solution Approach 1:

The patent segments the processing workload by distributing threads across multiple contexts and processing clusters rather than concentrating all processing on a single chip. This segmentation reduces thermal density on any one chip while maintaining overall processing capability through coordinated multi-chip operation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a multi-dimensional processing architecture with multiple contexts, clusters, and chips operating in parallel. By moving from single-chip monolithic processing to distributed multi-chip processing with multiple execution contexts, the system achieves higher processing capability without proportional increases in heat generation at any single location.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If additional functionality is integrated onto a single IC chip, then processing capability is improved, but power consumption increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent segments processing work across multiple contexts and clusters, enabling selective activation of processing resources based on actual workload demands. This segmentation allows the system to achieve high processing capability when needed while reducing power consumption during lower-demand periods through more efficient resource utilization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic workload distribution and context switching that adapts processing allocation to actual needs. The system dynamically adjusts which contexts and clusters are active based on workload characteristics, optimizing the balance between processing capability and power consumption in real-time.

Inventive Principle:
Principle #15Dynamics

3Productivity

If threads are distributed across multiple contexts in GPU, then processing efficiency is improved, but thread distribution complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidthread distribution complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where the hardware automatically manages thread distribution across contexts and clusters based on workload characteristics. The system autonomously determines optimal thread placement and context switching without requiring complex external control, thereby reducing operational complexity while maintaining high processing efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent utilizes parameter-based classification of threads (such as thread type, priority, and workload characteristics) to automatically guide distribution decisions. By changing and utilizing these parameters for intelligent scheduling, the system achieves efficient thread distribution while keeping the control logic manageable through parameter-driven decision making.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If multiple contexts are executed in GPU, then processing efficiency is improved, but power management difficulty increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidpower management difficulty
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments power management by associating power control with specific contexts and clusters rather than managing power uniformly across the entire GPU. This allows independent power management of individual contexts, enabling the system to maintain high processing efficiency through parallel context execution while simplifying power management through granular, context-specific control.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic power management that adapts to the actual execution state of multiple contexts. The system dynamically adjusts power supply to active contexts based on their workload and execution status, maintaining processing efficiency while managing power consumption more effectively through real-time adaptation rather than static power allocation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10908905B2Efficient multi-context thread distribution
Publication Date: 2021.02.02 INTEL CORP
  • US10908905B2 patent drawing
  • US10908905B2 patent drawing
  • US10908905B2 patent drawing

AI summary

Methods and apparatus relating to techniques for avoiding cache lookup for cold cache. In an example, an apparatus comprises logic, at least partially comprising hardware logic, to determine a first number of threads to be scheduled for each context of a plurality of contexts in a multi-context processing system, allocate a second number of streaming multiprocessors (SMs) to the respective plurality of contexts, and dispatch threads from the plurality of contexts only to the streaming multiprocessor(s) allocated to the respective plurality of contexts. Other embodiments are also disclosed and claimed.