Out-of-Order Clustered Decoding Load Balancing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern processors face challenges in efficiently executing complex instructions, such as floating-point operations and load/store operations, which require more execution time and resources, leading to reduced overall throughput.

Innovation Solution

The implementation of out-of-order clustered decoding with load balancing, where instructions are decoded in parallel by multiple clusters and then reordered for execution, allowing for increased pipeline throughput and improved performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If instructions are decoded in parallel by multiple clusters, then pipeline throughput is improved, but load imbalance occurs causing some clusters to be underutilized

Engineering Contradiction:
Improvepipeline throughputVSAvoidload balancing
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system dynamically adjusts the distribution of instructions to decode clusters based on current load conditions. The out-of-order execution mechanism allows the processor to adaptively balance the workload across multiple decode clusters, ensuring that instructions are evenly distributed and all clusters remain actively utilized, thereby maintaining high pipeline throughput without load imbalance

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Instructions are speculatively decoded ahead of time by multiple decode clusters before they are actually needed for execution. This preliminary decoding action allows the system to prepare multiple instruction streams in parallel, improving throughput while the out-of-order mechanism ensures that decoded instructions are properly managed and executed when ready, preventing cluster underutilization

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If complex instructions such as floating-point operations are executed, then processing capability is improved, but execution time increases reducing overall throughput

Engineering Contradiction:
Improveprocessing capabilityVSAvoidoverall throughput
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

Complex instructions are broken down into multiple simpler micro-operations that can be executed in parallel by different execution units. This segmentation allows floating-point operations and other complex instructions to be processed through multiple stages simultaneously, maintaining processing capability while improving overall throughput by reducing the time any single instruction blocks the pipeline

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An out-of-order execution buffer acts as an intermediary between instruction decoding and execution. This buffer allows instructions to be held and reordered, enabling complex instructions to be executed as soon as their dependencies are resolved rather than waiting for sequential order, thereby maintaining versatility while improving throughput by keeping execution units continuously busy

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10331454B2System and method for load balancing in out-of-order clustered decoding
Publication Date: 2019.06.25 INTEL CORP
  • US10331454B2 patent drawing
  • US10331454B2 patent drawing
  • US10331454B2 patent drawing

AI summary

A processor includes a back end to execute decoded instructions and a front end. The front end includes two decode clusters and circuitry to receive data elements representing undecoded instructions, in program order, and to direct subsets of the data elements to the decode clusters. An IP generator directs one subset of data elements to the first cluster, detects a condition indicating that a load balancing action should be taken, and directs a subset of data elements immediately following the first subset in program order to the first or second decode cluster dependent on the action taken. The action may include annotating a BTB entry, inserting a fake branch in the BTB, forcing a cluster switch, or suppressing a cluster switch. The detected condition may be a predicated taken branch or an annotation thereof, or a heuristic based on a queue state, a count of uops, or a latency value.