Superscalar Microprocessor Execution Queues for Dependency Handling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer processor architectures face limitations in increasing instructions per cycle (IPC) due to instruction stalling caused by memory access issues and dependencies, leading to power consumption challenges in portable devices.
Innovation Solution
A superscalar microprocessor architecture with in-order execution queues and out-of-order execution across queues, utilizing multiple load queues to handle dependent instructions and employing duplicate free lists to manage instruction dependencies, allowing for concurrent execution of instructions and reducing power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If instructions are executed in-order with single execution queues, then simplicity of control is maintained, but instructions per cycle (IPC) is limited due to stalling
Solution Approach 1:
The execution queue is segmented into multiple load queues (first load queue, second load queue, etc.) that can independently store and manage instructions. Each queue operates semi-independently, allowing instructions to be issued to different queues based on their specific requirements, thereby increasing the number of instructions that can be processed in parallel without requiring a completely complex reorganization of the execution pipeline.
Solution Approach 2:
The patent introduces a new dimension to instruction execution by organizing queues in a multi-dimensional structure where instructions can be routed to different queues based on their dependencies and execution requirements. This dimensional organization allows for more flexible instruction scheduling and reduces stalling by providing multiple execution paths simultaneously.
2Productivity
If out-of-order execution is implemented to increase IPC, then more instructions can be executed concurrently, but stalled instructions block independent instructions as queues fill up
Solution Approach 1:
The patent implements feedback mechanisms through duplicate free lists that track the status of instructions across multiple queues. When an instruction is executed or its results become available, this information is fed back to update the duplicate free lists, which then trigger the issuance of dependent instructions that were previously stalled. This feedback system ensures correct execution order while enabling out-of-order processing.
Solution Approach 2:
The duplicate free lists are pre-configured with information about instruction dependencies and potential duplicate instructions. This preliminary organization allows the processor to quickly determine which instructions can be safely executed out-of-order without compromising correctness, as the dependency relationships are already established and tracked before execution begins.
3Productivity
If multiple execution queues are used to handle dependencies, then more instructions can be processed in parallel, but power consumption increases
Solution Approach 1:
The patent uses multiple load queues, but not all queues need to be fully populated or actively executing instructions at any given time. The system performs partial action by activating only the necessary queues based on current instruction flow and dependency requirements, thereby reducing overall power consumption while still achieving parallel processing benefits when needed.
Solution Approach 2:
The execution queue structure is dynamic rather than static - queues can be activated or deactivated based on current workload and dependency requirements. This dynamic configuration allows the processor to optimize power consumption by keeping only the necessary queues active at any given moment, rather than maintaining all queues in a constant ready state.
4Adaptability or versatility
If duplicate instructions are stored in multiple queues to handle dependencies, then instruction dependencies are managed, but redundant storage increases area requirements
Solution Approach 1:
The patent implements a mechanism where duplicate instruction copies in different queues are tracked using duplicate free lists. When an instruction is executed or its results are no longer needed, the corresponding duplicate entries are identified and discarded from the queues. This recovering of storage space ensures that redundant instructions do not permanently occupy queue entries, thereby reducing the overall area requirements while maintaining the ability to handle complex dependencies.
Data Source
AI summary
A processor includes an instruction unit which provides instructions for execution by the processor, a decode/issue unit which decodes instructions received from the instruction unit and issues the instructions, and a plurality of execution queues coupled to the decode/issue unit. Each issued instruction from the decode/issue unit is stored into an entry of at least one queue of the plurality of execution queues, wherein each entry of the plurality of execution queues is configured to store an issued instruction and a duplicate indicator corresponding to the issued instruction which indicates whether or not a duplicate instruction of the issued instruction is also stored in an entry of another queue of the plurality of execution queues.


