Superscalar Microprocessor Execution Queues for Dependency Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computer processor architectures face limitations in increasing instructions per cycle (IPC) due to instruction stalling caused by memory access issues and dependencies, leading to power consumption challenges in portable devices.

Innovation Solution

A superscalar microprocessor architecture with in-order execution queues and out-of-order execution across queues, utilizing multiple load queues to handle dependent instructions and employing duplicate free lists to manage instruction dependencies, allowing for concurrent execution of instructions and reducing power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If instructions are executed in-order with single execution queues, then simplicity of control is maintained, but instructions per cycle (IPC) is limited due to stalling

Engineering Contradiction:
Improveinstructions per cycleVSAvoidexecution queue structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The execution queue is segmented into multiple load queues (first load queue, second load queue, etc.) that can independently store and manage instructions. Each queue operates semi-independently, allowing instructions to be issued to different queues based on their specific requirements, thereby increasing the number of instructions that can be processed in parallel without requiring a completely complex reorganization of the execution pipeline.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension to instruction execution by organizing queues in a multi-dimensional structure where instructions can be routed to different queues based on their dependencies and execution requirements. This dimensional organization allows for more flexible instruction scheduling and reduces stalling by providing multiple execution paths simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If out-of-order execution is implemented to increase IPC, then more instructions can be executed concurrently, but stalled instructions block independent instructions as queues fill up

Engineering Contradiction:
Improveinstructions per cycleVSAvoidinstruction execution correctness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms through duplicate free lists that track the status of instructions across multiple queues. When an instruction is executed or its results become available, this information is fed back to update the duplicate free lists, which then trigger the issuance of dependent instructions that were previously stalled. This feedback system ensures correct execution order while enabling out-of-order processing.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The duplicate free lists are pre-configured with information about instruction dependencies and potential duplicate instructions. This preliminary organization allows the processor to quickly determine which instructions can be safely executed out-of-order without compromising correctness, as the dependency relationships are already established and tracked before execution begins.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If multiple execution queues are used to handle dependencies, then more instructions can be processed in parallel, but power consumption increases

Engineering Contradiction:
Improveinstructions per cycleVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent uses multiple load queues, but not all queues need to be fully populated or actively executing instructions at any given time. The system performs partial action by activating only the necessary queues based on current instruction flow and dependency requirements, thereby reducing overall power consumption while still achieving parallel processing benefits when needed.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The execution queue structure is dynamic rather than static - queues can be activated or deactivated based on current workload and dependency requirements. This dynamic configuration allows the processor to optimize power consumption by keeping only the necessary queues active at any given moment, rather than maintaining all queues in a constant ready state.

Inventive Principle:
Principle #15Dynamics

4Adaptability or versatility

If duplicate instructions are stored in multiple queues to handle dependencies, then instruction dependencies are managed, but redundant storage increases area requirements

Engineering Contradiction:
Improvedependency handling capabilityVSAvoidqueue storage area
Core Design Contradiction:
Adaptability or versatilityVSArea of stationary object

Solution Approach 1:

The patent implements a mechanism where duplicate instruction copies in different queues are tracked using duplicate free lists. When an instruction is executed or its results are no longer needed, the corresponding duplicate entries are identified and discarded from the queues. This recovering of storage space ensures that redundant instructions do not permanently occupy queue entries, thereby reducing the overall area requirements while maintaining the ability to handle complex dependencies.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS8904150B2Microprocessor systems and methods for handling instructions with multiple dependencies
Publication Date: 2014.12.02 ASCALE TECHNOLOGIES LLC
  • US8904150B2 patent drawing
  • US8904150B2 patent drawing
  • US8904150B2 patent drawing

AI summary

A processor includes an instruction unit which provides instructions for execution by the processor, a decode/issue unit which decodes instructions received from the instruction unit and issues the instructions, and a plurality of execution queues coupled to the decode/issue unit. Each issued instruction from the decode/issue unit is stored into an entry of at least one queue of the plurality of execution queues, wherein each entry of the plurality of execution queues is configured to store an issued instruction and a duplicate indicator corresponding to the issued instruction which indicates whether or not a duplicate instruction of the issued instruction is also stored in an entry of another queue of the plurality of execution queues.