Parallel Processor Predicate Register File for Reduced Per-Thread State

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing parallel processor designs face challenges in minimizing per-thread state and instruction encoding costs for predicated execution, particularly in SIMT and SIMD architectures, leading to increased chip area and power consumption.

Innovation Solution

A method is introduced that involves receiving instructions for execution by a thread group, computing predicate results, and storing them in dedicated predicate registers, reducing the need for per-thread state and instruction bits, and allowing optional negation of predicates to save additional bits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If 4-bit condition code (CC) registers are used for each thread or lane instance with 7-bit instruction guards, then predicated execution capability is achieved, but per-thread state cost and instruction encoding space increase

Engineering Contradiction:
Improvepredicated execution capabilityVSAvoidper-thread state cost
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the condition code functionality from traditional per-thread CC registers and implements it through a shared CC register file with predicate bits. Instead of having dedicated 4-bit CC registers for each thread (16 bits per thread), the system uses a shared register file where each thread accesses condition codes through predicate bits, significantly reducing per-thread state while maintaining predicated execution capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The shared condition code register file serves multiple threads simultaneously, making the CC storage mechanism universal. The same physical register file is shared across all thread instances, and the same predicate bits are used for both condition code storage and predication control, eliminating the need for separate per-thread CC registers

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If 7-bit instruction guards and 3-bit CC register encoding are used, then conditional execution control is achieved, but instruction encoding space is consumed

Engineering Contradiction:
Improveconditional execution controlVSAvoidinstruction encoding space
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges the condition code selection and predication functionality into a unified predicate bit mechanism. Instead of separate 7-bit guards for CC selection and 3-bit encoding for destination CC registers, the system uses compact predicate bits that combine both functions, reducing the instruction encoding overhead while maintaining full conditional execution control

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If per-thread CC registers are implemented for hundreds of parallel threads, then individual thread predication is enabled, but chip area and power consumption increase

Engineering Contradiction:
Improveindividual thread predicationVSAvoidchip area
Core Design Contradiction:
Adaptability or versatilityVSArea of stationary object

Solution Approach 1:

The patent extracts the per-thread CC register allocation and replaces it with a shared CC register file accessed through thread identifiers and predicate bits. This eliminates the need for hundreds of dedicated CC registers, reducing chip area while preserving individual thread predication capability through the shared resource model

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system discards the traditional model of permanent per-thread CC registers and recovers the functionality through a shared register file with predicate bits. Threads temporarily access condition codes through the shared file based on their current execution context, eliminating the need for persistent per-thread storage

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS10360039B2Predicted instruction execution in parallel processors with reduced per-thread state information including choosing a minimum or maximum of two operands based on a predicate value
Publication Date: 2019.07.23 NVIDIA CORP
  • US10360039B2 patent drawing
  • US10360039B2 patent drawing
  • US10360039B2 patent drawing

AI summary

A mechanism for predicated execution of instructions within a parallel processor executing multiple threads or data lanes is disclosed. Each thread or data lane executing within the parallel processor is associated with a predicate register that stores a set of 1-bit predicates. Each of these predicates can be set using different types of predicate-setting instructions, where each predicate setting instruction specifies one or more source operands, at least one operation to be performed on the source operands, and one or more destination predicates for storing the result of the operation. An instruction can be guarded by a predicate that may influence whether the instruction is executed for a particular thread or data lane or how the instruction is executed for a particular thread or data lane.