Parallel Processor Predicate Register File for Reduced Per-Thread State
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parallel processor designs face challenges in minimizing per-thread state and instruction encoding costs for predicated execution, particularly in SIMT and SIMD architectures, leading to increased chip area and power consumption.
Innovation Solution
A method is introduced that involves receiving instructions for execution by a thread group, computing predicate results, and storing them in dedicated predicate registers, reducing the need for per-thread state and instruction bits, and allowing optional negation of predicates to save additional bits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If 4-bit condition code (CC) registers are used for each thread or lane instance with 7-bit instruction guards, then predicated execution capability is achieved, but per-thread state cost and instruction encoding space increase
Solution Approach 1:
The patent extracts the condition code functionality from traditional per-thread CC registers and implements it through a shared CC register file with predicate bits. Instead of having dedicated 4-bit CC registers for each thread (16 bits per thread), the system uses a shared register file where each thread accesses condition codes through predicate bits, significantly reducing per-thread state while maintaining predicated execution capability
Solution Approach 2:
The shared condition code register file serves multiple threads simultaneously, making the CC storage mechanism universal. The same physical register file is shared across all thread instances, and the same predicate bits are used for both condition code storage and predication control, eliminating the need for separate per-thread CC registers
2Ease of operation
If 7-bit instruction guards and 3-bit CC register encoding are used, then conditional execution control is achieved, but instruction encoding space is consumed
Solution Approach 1:
The patent merges the condition code selection and predication functionality into a unified predicate bit mechanism. Instead of separate 7-bit guards for CC selection and 3-bit encoding for destination CC registers, the system uses compact predicate bits that combine both functions, reducing the instruction encoding overhead while maintaining full conditional execution control
3Adaptability or versatility
If per-thread CC registers are implemented for hundreds of parallel threads, then individual thread predication is enabled, but chip area and power consumption increase
Solution Approach 1:
The patent extracts the per-thread CC register allocation and replaces it with a shared CC register file accessed through thread identifiers and predicate bits. This eliminates the need for hundreds of dedicated CC registers, reducing chip area while preserving individual thread predication capability through the shared resource model
Solution Approach 2:
The system discards the traditional model of permanent per-thread CC registers and recovers the functionality through a shared register file with predicate bits. Threads temporarily access condition codes through the shared file based on their current execution context, eliminating the need for persistent per-thread storage
Data Source
AI summary
A mechanism for predicated execution of instructions within a parallel processor executing multiple threads or data lanes is disclosed. Each thread or data lane executing within the parallel processor is associated with a predicate register that stores a set of 1-bit predicates. Each of these predicates can be set using different types of predicate-setting instructions, where each predicate setting instruction specifies one or more source operands, at least one operation to be performed on the source operands, and one or more destination predicates for storing the result of the operation. An instruction can be guarded by a predicate that may influence whether the instruction is executed for a particular thread or data lane or how the instruction is executed for a particular thread or data lane.


