Microprocessor Branch Execution Logic for Cache Miss Handling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In out-of-order execution microprocessors, branch instructions pose a significant performance penalty due to misprediction and the need for correcting the instruction stream, which can lead to inefficient re-fetching and re-processing of instructions.
Innovation Solution
A pipelined out-of-order execution in-order retire microprocessor with a branch predictor, fetch unit, and execution unit that detects mispredicted branch instructions and selectively executes or refrains from executing them based on the presence of unretired load instructions, allowing for early correction and avoiding detrimental re-processing by flushing and replaying instructions without re-fetching from the cache.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If the microprocessor executes a mispredicted branch instruction to enable early correction, then the branch penalty is reduced, but instructions fetched at the predicted target address must be flushed and replayed without re-fetching from cache
Solution Approach 1:
The microprocessor executes the mispredicted branch instruction in advance (before the older load instruction completes) to enable early correction of the branch target. This preliminary execution allows the fetch unit to be redirected to the correct target address sooner, reducing the overall branch penalty despite the need to flush previously fetched instructions.
Solution Approach 2:
The patent extracts the branch execution logic from the normal sequential flow and allows it to execute independently out of order. The mispredicted branch is separated from the dependency chain of the older load instruction, enabling it to be resolved and corrected earlier without waiting for the load to complete, thus reducing the time loss associated with branch misprediction.
2Productivity
If the microprocessor allows out-of-order execution of branch instructions, then performance is improved by keeping execution units supplied with instructions, but the architectural state must be updated in program order requiring additional pipeline stages
Solution Approach 1:
The microprocessor pipeline is segmented into distinct functional units: a front end (fetch unit, branch predictor) that handles instruction fetching and prediction, and a back end (execution units, reorder buffer, retire unit) that handles out-of-order execution and in-order retirement. This segmentation allows the front end to speculate and fetch out of order while the back end maintains program order for retirement, enabling high instruction throughput without compromising architectural correctness.
Solution Approach 2:
The reorder buffer acts as an intermediary structure between the out-of-order execution units and the in-order retire mechanism. It buffers instructions that have executed out of order and manages their retirement in program order, allowing the execution units to operate at full capacity while maintaining correct architectural state updates.
3Reliability
If the microprocessor refrains from executing the mispredicted branch when an older load instruction is pending, then correctness is maintained, but the branch penalty increases due to later correction
Solution Approach 1:
The microprocessor dynamically determines whether to execute a mispredicted branch based on the status of older instructions in the instruction queue. The execution unit monitors the completion status of older load instructions and makes a real-time decision: if the older load is complete, the mispredicted branch executes immediately for early correction; if the older load is still pending, execution is deferred to maintain correctness. This dynamic adaptation optimizes the balance between speed and correctness.
Data Source
AI summary
A pipelined out-of-order execution in-order retire microprocessor includes a branch predictor that predicts a target address of a branch instruction, a fetch unit that fetches instructions at the predicted target address, and an execution unit that: resolves a target address of the branch instruction and detects that the predicted and resolved target addresses are different; determines whether there is an unretired instruction that must be corrected and that is older in program order than the branch instruction, in response to detecting that the predicted and resolved target addresses are different; execute the branch instruction by flushing instructions fetched at the predicted target address and causing the fetch unit to fetch from the resolved target address, if there is not an unretired instruction that must be corrected and that is older in program order than the branch instruction; and otherwise, refrain from executing the branch instruction.


