Multi-level Dispatch Circuit for Superscalar Processor Timing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-performance processors face challenges in maintaining high clock cycle rates due to the pressure of processing large numbers of instructions, particularly in wide-issue superscalar processors, where efficiently distributing instructions to multiple parallel pipelines is complex and timing pressures are significant.

Innovation Solution

A multi-level dispatch circuit with multiple dispatch buffers, each coupled to multiple reservation stations, simplifies the selection and distribution of operations, approximating even distribution to reservation stations with available entries, thereby relieving timing pressures and enhancing frequency operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a wide-issue superscalar processor processes large numbers of instructions to increase execution rate, then instruction throughput is improved, but timing pressures and complexity of distributing instructions to multiple parallel pipelines increase

Engineering Contradiction:
Improveinstruction throughputVSAvoidcomplexity of distributing instructions
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The dispatch circuit is segmented into multiple levels: first dispatch buffers receive instructions from the instruction queue, then second dispatch buffers receive instructions from the first dispatch buffers, and finally reservation stations receive instructions from the second dispatch buffers. This multi-level segmentation distributes the complexity of instruction allocation across hierarchical stages, reducing the timing pressure on any single dispatch stage while maintaining high instruction throughput to multiple parallel pipelines

Inventive Principle:
Principle #1Segmentation

2Productivity

If multiple reservation stations are used to supply operations to parallel execution pipelines, then parallel processing capability is improved, but the pressure of locating and distributing large numbers of instructions quickly increases

Engineering Contradiction:
Improveparallel processing capabilityVSAvoidtime to locate and distribute instructions
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The first dispatch buffers perform preliminary action by receiving and holding instructions from the instruction queue before they are needed by the reservation stations. This advance preparation allows instructions to be staged and organized in advance, reducing the time pressure on the final dispatch stage to reservation stations and enabling faster instruction location and distribution when needed

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If direct transmission of operations to reservation stations is attempted, then distribution precision is improved, but the complexity of selection mechanisms increases

Engineering Contradiction:
Improveprecision of operation distributionVSAvoidcomplexity of selection mechanisms
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The second dispatch buffers serve as intermediary structures between the first dispatch buffers and the reservation stations. These intermediaries simplify the selection mechanism by providing a staged approach: the first dispatch buffers allocate to second dispatch buffers using one selection mechanism, and the second dispatch buffers allocate to reservation stations using another selection mechanism. This breaks down the complex direct selection problem into simpler staged selections, maintaining distribution precision while reducing overall mechanism complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9336003B2Multi-level dispatch for a superscalar processor
Publication Date: 2016.05.10 APPLE INC
  • US9336003B2 patent drawing
  • US9336003B2 patent drawing
  • US9336003B2 patent drawing

AI summary

In an embodiment, a processor includes a multi-level dispatch circuit configured to supply operations for execution by multiple parallel execution pipelines. The multi-level dispatch circuit may include multiple dispatch buffers, each of which is coupled to multiple reservation stations. Each reservation station may be coupled to a respective execution pipeline and may be configured to schedule instruction operations (ops) for execution in the respective execution pipeline. The sets of reservation stations coupled to each dispatch buffer may be non-overlapping. Thus, if a given op is to be executed in a given execution pipeline, the op may be sent to the dispatch buffer which is coupled to the reservation station that provides ops to the given execution pipeline.