Conditional instruction prediction

By introducing bias prediction circuits and instruction prediction circuits into the processor, using bias tables and multiple tables to predict conditional instructions, the problem of insufficient accuracy of conditional instruction prediction by existing processors is solved, and higher processor efficiency and lower error prediction processing delay is achieved.

CN120216028APending Publication Date: 2025-06-27APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510294884.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-02-01
Filing Date
2023-01-31
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The lack of accuracy in conditional instruction prediction of existing processors leads to discarding speculative work and degrading processor performance, especially when the width and depth of the execution pipeline increases.

Method used

Using a processor design including a bias prediction circuit and an instruction prediction circuit, the bias prediction and instruction prediction of conditional instructions are provided through the use of bias tables and multiple tables, thereby improving the processing efficiency of conditional instructions.

Benefits of technology

By improving the prediction accuracy of conditional instructions, reducing speculative work discards caused by misprediction, improving the overall efficiency of the processor, and reducing the processor's error prediction processing delay in the cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216028A_ABST
    Figure CN120216028A_ABST
Patent Text Reader

Abstract

The invention relates to conditional instruction prediction. A processor may include bias prediction circuitry and instruction prediction circuitry to provide respective predictions for conditional instructions. The bias prediction circuit may provide a bias prediction that the condition of the conditional instruction is bias true or bias false. The instruction prediction circuitry may provide an instruction prediction of whether the condition of the conditional instruction is true or false. In response to a bias prediction that the condition of the conditional instruction is bias true or bias false, the processor may speculatively process the conditional instruction using the bias prediction from the bias prediction circuitry. Otherwise, the processor may speculatively process the conditional instruction using the instruction prediction from the instruction prediction circuitry.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application date of January 31, 2023, the application number of 202380019404.2, and the title of "Conditional Instruction Prediction". Technical Field

[0002] The embodiments described herein relate to processors, and more particularly, to processors including circuitry for predicting the outcome of conditional instructions and / or processing conditional instructions based on the prediction. Background Art

[0003] Computing systems typically include one or more processors that serve as a central processing unit (CPU). The CPU executes control software (e.g., an operating system) that controls the operation of various peripheral devices. The CPU may also execute application programs that provide user functionality in the system. Sometimes, a processor may implement an instruction pipeline including multiple stages, where instructions are divided into a series of steps that are executed at corresponding stages of the pipeline. Thus, the instruction pipeline may execute multiple instructions in parallel. To improve efficiency, a processor may also implement a conditional instruction prediction circuit (also referred to as a "conditional instruction predictor") that can predict the condition of a conditional instruction. Based on the prediction, the processor may speculatively fetch instructions from a target address for execution. However, if a conditional instruction is mispredicted, the speculative work must be discarded, and the processor may have to refetch the instructions from the correct target address for execution. Therefore, the accuracy of the prediction of conditional instructions can play a key role in the performance of a processor, and thus techniques with improved prediction accuracy are needed. Additionally, as the width and depth of the execution pipeline increase, a processor may process multiple conditional instructions and / or mispredictions in a cycle. Therefore, techniques for improving the efficiency of conditional instruction processing in a processor are also needed. Brief Description of the Drawings

[0004] The following detailed description refers to the accompanying drawings, which are briefly described now.

[0005] Figure 1 is a block diagram of one embodiment that is part of a processor including a bias prediction circuit and an instruction prediction circuit.

[0006] Figure 2A shows one embodiment of a bias table of the bias prediction circuit.

[0007] Figure 2B shows one embodiment of a basic table of the instruction prediction circuit.

[0008] Figure 3A is a block diagram of one embodiment of the operation of the bias prediction circuit.

[0009] Figure 3BIt is a block diagram of an embodiment of the operation of an instruction prediction circuit.

[0010] Figure 4 It is a flowchart illustrating an embodiment of the operation of a processor including a bias prediction circuit and an instruction prediction circuit.

[0011] Figure 5 It is a block diagram of an embodiment of a part of a processor including an instruction distribution circuit and a plurality of execution pipelines.

[0012] Figure 6 It is a flowchart illustrating an embodiment of the operation of a processor including an instruction distribution circuit and a plurality of execution pipelines.

[0013] Figure 7 It is a block diagram of another embodiment of a part of a processor including an instruction distribution circuit and a plurality of execution pipelines.

[0014] Figure 8 It is a flowchart illustrating another embodiment of the operation of a processor including an instruction distribution circuit and a plurality of execution pipelines.

[0015] Figure 9 It is Figures 1 to 8 A block diagram of an embodiment of a processor including a bias prediction circuit, an instruction prediction circuit, and / or an instruction distribution circuit as shown in

[0016] Figure 10 It may include Figure 9 A block diagram of an embodiment of a system - on - chip (SOC) that may include one or more processors as shown in

[0017] Figure 11 It is a block diagram of an embodiment of a system used in various contexts.

[0018] Figure 12 It is a block diagram of a computer - accessible storage medium.

[0019] Although the embodiments described in this disclosure may be subject to various modified forms and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that the drawings and the specific implementation thereof are not intended to limit the embodiments to the particular forms disclosed, but rather, the present invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the appended claims. The headings used herein are for organizational purposes only and are not intended to limit the scope of the specification. Detailed Description

[0020] Now turning to Figure 1, a block diagram of an embodiment of a portion of a processor 30 including a bias prediction circuit 156 and an instruction prediction circuit 160 is shown. In the illustrated embodiment, the bias prediction circuit 156 and the instruction prediction circuit 160 may be implemented as part of the fetch and decode circuit 100 of the processor 30. Alternatively, in some other embodiments, the bias prediction circuit 156 and / or the instruction prediction circuit 160 may be implemented as components separate from the fetch and decode circuit 100.

[0021] As Figure 1 indicated, in the illustrated embodiment, the fetch and decode circuit 100 may be implemented as a pipeline having several stages. For example, to process an instruction, the fetch and decode circuit 100 may first use the prefetch circuit 150 to load the instruction from the memory or cache 12 into the instruction cache (Icache) 102 (hereinafter referred to as the "prefetch" stage). Then, the instruction may be fetched by the fetch circuit 152 from the Icache 102 to the decoder 154 for decoding (hereinafter referred to as the "fetch" stage). The decoder 154 may decode the instruction, convert it into operations and / or micro-operations (hereinafter referred to as the "decode" stage), and send the operations and / or micro-operations to the execution pipeline 164 for execution. Note that sometimes the instruction may already be present in the Icache 102. For example, when the Icache 102 stores the instruction previously loaded from the memory or cache 12 in the Icache 102. In this case, the prefetch stage may be avoided, and the instruction may be directly fetched from the Icache 102 for execution. In the illustrated embodiment, the execution pipeline 164 may be implemented using Figure 9 the execution units 112 (e.g., integer, floating-point, and / or vector execution units) and the associated reservation stations 110 described in Figure 1 . Additionally, for illustrative purposes,

[0022] the execution of code including conditional instructions may depend on the conditions of the conditional instructions. When the condition of a conditional instruction is true, the first instruction from the first target address may be loaded, fetched, and executed. Conversely, when the condition of the conditional instruction is false, the second instruction from the second target address may be loaded, fetched, and executed. For illustrative purposes, the following is an example code including conditional instructions:

[0023] ==========================================If(a>b) / / Conditional instruction

[0024] {

[0025] x = 1; / / Instruction to be executed when the condition is true

[0026] }

[0027] else

[0028] {

[0029] x = 2; / / Instruction to be executed when the condition is false

[0030] }

[0031] ==========================================

[0032] In this example, the conditional instruction only involves comparing the values of two variables, "a" and "b". If the condition of the conditional instruction is true (i.e., the value of "a" is greater than the value of "b"), then the first instruction from the first target address can be executed to assign the value of variable "x" to 1. Conversely, if the condition of the conditional instruction is false (i.e., the value of "a" is less than or equal to the value of "b"), then the second instruction from the second target address can be executed to assign the value of variable "x" to 2.

[0033] In the illustrated embodiment, the fetch and decode circuit 100 may presumably process the conditional instruction. For example, the fetch and decode circuit 100 may predict the condition of the conditional instruction before the (actual) execution of the conditional instruction, and based on the prediction, speculatively determine the target address, from which subsequent instructions can be obtained for execution. As described above, the target address may reside in the memory or cache 12 or the Icache 102. In addition, as illustrated in the above example, the subsequent instructions may or may not immediately follow the conditional instruction. To improve efficiency, the fetch and decode circuit 100 may also use the bias prediction circuit 156 with a bias table 158 to provide a bias prediction as to whether the conditional instruction is bias true or bias false. When the conditional instruction is predicted to be bias true or bias false, the fetch and decode circuit 100 may use the bias prediction from the bias prediction circuit 156 to process the conditional instruction. Conversely, when the conditional instruction is predicted not to be bias true or bias false, the fetch and decode circuit 100 may use the instruction prediction circuit 160 with one or more tables 162 to provide another prediction (such as an instruction prediction, regardless of whether the condition of the conditional instruction is true or false), and use the instruction prediction to speculatively process the conditional instruction.

[0034] In the illustrated example, the bias prediction circuit 156 and the instruction prediction circuit 160 may perform corresponding predictions at different stages of the processing of conditional instructions in the fetch and decode circuit 100. For example, in the illustrated embodiment, when a conditional instruction is loaded from the memory or cache 12 into the Icache 102, the bias prediction circuit may provide a bias prediction for the conditional instruction at the prefetch stage. By comparison, in the illustrated embodiment, the instruction prediction circuit 160 may provide an instruction prediction at a relatively "later" stage (such as the fetch stage when the conditional instruction is fetched from the Icache 102 to the decoder 154). Note that the above is provided for illustrative purposes only. In some embodiments, the bias prediction circuit 156 and the instruction prediction circuit 160 may provide their respective predictions approximately simultaneously, e.g., both at the same stage (such as the prefetch stage, the fetch stage, etc.).

[0035] Sometimes, when a conditional instruction is predicted to be bias-true or bias-false, the fetch and decode circuit 100 may cause the conditional instruction to "bypass" the instruction prediction circuit 160 only. In other words, the instruction prediction circuit 160 may not necessarily provide a second prediction such as an instruction prediction. Alternatively, sometimes the fetch and decode circuit 100 may still use the instruction prediction circuit 160 to provide an instruction prediction. However, when a conditional instruction is predicted to be bias-true or bias-false, the fetch and decode circuit 100 may ignore the instruction prediction from the instruction prediction circuit 160 and instead use the bias prediction from the bias prediction circuit 156 to speculatively process the conditional instruction, as described above.

[0036] In the illustrated embodiment, the bias prediction from the bias prediction circuit 156 and the instruction prediction from the instruction prediction circuit 160 may indicate different natures of the conditional instruction. In addition, as described in FIGS. 2 to 3, they may be generated in different ways. In the illustrated embodiment, the bias prediction from the bias prediction circuit 156 may indicate whether the conditional instruction is predicted to be biased (e.g., bias-true or bias-false). A conditional instruction being biased refers to a scenario where the condition of the conditional instruction is always true or false. For example, if the condition is always true, the condition of the conditional instruction is considered bias-true. Conversely, if it is always false, the condition is considered bias-false. Referring back to the bias prediction, when the condition of the conditional instruction is predicted to be bias-true (or bias-false), it means that the condition of the conditional instruction is predicted to always be true (or always be false), and thus it is assumed that the conditional instruction always operates in one way (or the other way). By comparison, the instruction prediction from the instruction prediction circuit 160 may indicate that the condition of the conditional instruction is predicted to be true or false. However, different from the bias prediction, the instruction prediction may not necessarily indicate whether the conditional instruction is bias-true or bias-false, or in other words, always true or always false.

[0037] Note that both bias prediction and instruction prediction are just predictions. Therefore, either of them may be incorrect. In the illustrated embodiment, the quality of the prediction may be determined, for example, after a conditional instruction is executed by the execution pipeline 164. Considering the above example code, once the values of the operands (e.g., variables "a" and "b") are obtained and the operator (e.g., comparator ">") is applied to the operands, the processor 30 may be able to determine whether the condition of the conditional instruction is actually true or false, and thus evaluate whether the bias prediction and / or the instruction prediction is correct. In the illustrated embodiment, the bias prediction circuit 156 and / or the instruction prediction circuit 160 may be updated based on the evaluation of the conditional instruction. For example, when the bias prediction and / or the instruction prediction is an incorrect prediction, the bias table 158 of the bias prediction circuit 156 and / or the table 162 of the instruction prediction circuit 160 may be updated.

[0038] When an incorrect prediction occurs, the processor 30 may have to discard the speculative work and obtain another instruction from the correct target address for execution. For example, the execution pipeline 164 may discard the instructions speculatively fetched within the execution pipeline, and the fetch and decode circuit 100 may have to redirect the prefetch circuit 150 and / or the fetch circuit 152 to obtain an instruction from the correct target address for execution (also referred to as refetching). Sometimes, this causes additional latency in the operation of the processor 30. However, in practice, most conditional instructions may be bias instructions. Therefore, even with the above losses due to incorrect predictions, using an additional bias prediction circuit can still increase the overall efficiency of the processor 30. In particular, if the processor 30 allows a predictive bias conditional instruction to "bypass" the instruction prediction circuit 160, this can greatly reduce the total workload and improve the efficiency of the processor 30.

[0039] In the illustrated embodiment, the bias prediction circuit 156 may use the bias table 158 to provide a bias prediction for a conditional instruction. Figure 2AAn example bias table 158 is shown. In the figure, the bias table 158 can be organized into one or more entries, where each entry can be identified by a corresponding index and includes a corresponding value. In the illustrated embodiment, the index of the bias table 158 can be associated with the address of a conditional instruction. For example, an index can be created by hashing the address of the conditional instruction using a hash function. In the context of hashing, the address of the conditional instruction can be regarded as the "key", and the value in the entry can be regarded as the "value", and the two can be associated with each other via the index (and the hash function). Thus, for a given conditional instruction, the bias prediction circuit 156 can identify the value (e.g., the "value") in the corresponding entry of the bias table 158 based on the address of the conditional instruction (e.g., the "key"), and then provide a bias prediction for the conditional instruction based on the identified value in the bias table 158. For example, when the bias prediction circuit 156 receives a conditional instruction, the bias prediction circuit 156 can obtain the address of the conditional instruction from, for example, a program counter (PC). The bias prediction circuit 156 can determine an index based on the address of the conditional instruction (e.g., using a hash function). The bias prediction circuit 156 can then use the index to search the bias table 158 to find an entry that matches the index, identify the value in the entry, and use the value to determine the bias prediction for the conditional instruction. Note that sometimes the index of the bias table 158 may suffer from hash collisions, for example, the phenomenon that different addresses of different conditional instructions may be hashed to the same index. In other words, different keys can correspond to the same value in the bias table 158. Sometimes, hash collisions may result in incorrect predictions of conditional instructions.

[0040] In Figure 2A the illustrated embodiment, the value in the bias table 158 can be a 2-bit value indicating different predictions regarding the bias of a conditional instruction. For example, the value "00" can indicate that the bias prediction circuit 156 has not previously encountered a conditional instruction corresponding to the entry with this value. The value "01" can indicate that the condition of the conditional instruction is bias false. The value "10" can indicate that the condition of the conditional instruction is bias true. And the value "11" can indicate that the condition of the conditional instruction is not biased (e.g., neither bias true nor bias false), although the bias prediction circuit 156 has previously encountered a conditional instruction corresponding to the entry with this value. Note that Figure 2A the bias table 158 in

[0041] In an illustrative embodiment, the instruction prediction circuit 160 may also use one or more tables 162 to predict the instruction prediction of conditional instructions. However, unlike the bias prediction circuit 156, at least some of the tables in the table 162 may be highly associated with the previous prediction history (e.g., by the instruction prediction circuit 160) and / or the evaluation history of the conditional instruction. In addition, sometimes the history may relate to the history of a particular conditional instruction, but may also relate to the history of other conditional instructions in the same code. For example, sometimes the instruction prediction circuit 162 may be a tagged geometric length predictor (also known as a TAGE predictor), which includes a basic predictor T0 and a set of (partial) tagged predictors T i (1 ≤ i ≤ M). The basic predictor T0 may use the basic table 162(0) to provide a basic prediction. In an illustrative embodiment, the index of the basic table 162(0) may be generated by hashing the address of the conditional instruction. By comparison, the tagged predictor T i (1 ≤ i ≤ M) may each have a table 162(i) (1 ≤ i ≤ M), the index of which may be created by hashing (a) the address of the conditional instruction and (b) the previous prediction and / or evaluation history of the conditional instruction. This history may be considered a geometric sequence. For example, the address of the conditional instruction may be concatenated with the history, and then the two may be hashed together to generate an index. The tables 162(i) of different tagged predictors T i (1 ≤ i ≤ M) may be associated with different history lengths. For example, the higher the order of the tagged predictor (e.g., the larger i), the longer the history may be used to generate the index of the table 162(i) of the tagged predictor T i (1 ≤ i ≤ M). Thus, the tagged predictors T i (1 ≤ i ≤ M) may use their respective tables 162(i) (1 ≤ i ≤ M) to provide corresponding predictions for the conditional instruction. Sometimes, the hash functions of the basic table 158 of the bias prediction circuit 156 and the basic table 162(0) of the instruction prediction circuit 160 may be different. In addition, sometimes the hash functions of the different tables 162(i) for different predictors T i (0 ≤ i ≤ M) may also be different. Additionally, the above hash functions may be implemented based on any suitable hash function, including the exclusive OR (or XOR) operation.

[0042] In an illustrative embodiment, for a given conditional instruction, to provide instruction prediction, the instruction prediction circuit 160 may determine the indices of the corresponding (M + 1) predictors (0 ≤ i ≤ M) based on the address and history of the conditional instruction (for the tag predictor only), identify the matching predictor with the longest history (e.g., with the highest order), and use the prediction from the matching predictor as the (final) instruction prediction for the conditional instruction. From the above description, it can be seen that the instruction prediction circuit 160 may be more complex than the bias prediction circuit 156 and thus consume more time to make a prediction. Therefore, using the additional bias prediction circuit 156 to allow predictive biasing of conditional instructions to "bypass" the instruction prediction circuit 160 may reduce the total workload and improve the efficiency of the processor 30.

[0043] Figure 2B An example basic table 162(0) of the instruction prediction circuit 160 is shown. For illustrative purposes, the basic table 162(0) is also provided as an example to illustrate the tag predictor T i (1 ≤ i ≤ M) of the table 162(i). In the illustrative embodiment, the table 162(i) of the tag predictor T i (1 ≤ i ≤ M) may be similar to the basic table 162(0). For example, the values at each entry are also provided, but additional information such as a history-dependent geometric sequence is also included. Additionally, the basic table 162(0) may also illustrate the difference between the bias prediction circuit 156 and the instruction prediction circuit 160. As Figure 2B indicated, in the illustrative embodiment, the values in the basic table 162(0) may be 2-bit values. For example, the value "00" may indicate that the condition of the conditional instruction is strongly false. The value "01" may indicate that the condition of the conditional instruction is weakly false. The value "10" may indicate that the condition of the conditional instruction is weakly true. The value "11" may indicate that the condition of the conditional instruction is strongly true. Therefore, the values in the table 162(0) of the instruction prediction circuit 160 may not necessarily indicate the bias of the conditional instruction, but only indicate whether it is true or false in a certain relativity. For example, compared with the value "01", the value "00" may indicate that the conditional instruction is predictively more likely to be false. Similarly, compared with the value "10", the value "11" may indicate that the conditional instruction is predictively more likely to be true. Note that Figure 2B the basic table 162(0) is provided only as an example for illustrative purposes. In some embodiments, the values in the basic table 162(0) and / or the table 162(i) of the tag predictor T i (1 ≤ i ≤ M) may have fewer or more bits.

[0044] Now turning to Figure 3A and Figure 3B , state machines of the bias prediction circuit 156 and the instruction prediction circuit 160 are shown to illustrate the operation of the corresponding prediction circuits. As Figure 3AAs indicated, circles 302, 304, 306, and 308 may correspond to Figure 2A four possible predictions in bias table 158 of bias prediction circuit 156 in Figure 3B . Similarly, in Figure 2B , circles 312, 314, 316, and 318 may correspond to

[0045] Referring back to Figure 3A , in the illustrated embodiment, the value "00" may be designed as the initial state or default value of a conditional instruction. For example, at startup, the value of the conditional instruction in bias table 158 may be set to the default value "00". When the conditional instruction is first loaded from memory or cache 12 into Icache 102, assuming there is no hash conflict for the conditional instruction yet, bias prediction circuit 156 may encounter the conditional instruction corresponding to the entry of the conditional instruction in bias table 158 for the first time. Thus, the value of the conditional instruction in bias table 158 may be "00" (e.g., corresponding to circle 302). Since the value "00" does not indicate that the condition of the conditional instruction is bias-true or bias-false, acquisition and decoding circuit 100 may also use instruction prediction circuit 160 to provide a second prediction for the conditional instruction, such as instruction prediction. As described above, instruction prediction circuit 160 may use table 162 to provide instruction prediction. Similarly, in the illustrated embodiment, processor 30 may designate one of the four possible states as the initial state or default value of the conditional instruction. For illustrative purposes, assume that the initial state or default value of the conditional instruction is "10" (e.g., corresponding to circle 316), thereby indicating that the condition is predicted to be weakly true. Based on the instruction prediction from instruction prediction circuit 160, acquisition and decoding circuit 100 may determine the target address, and based on the target address, subsequent instructions may be speculatively obtained for execution. Consider the above example code including the conditional instruction "if(a>b)". Since the conditional instruction is predicted to be "weakly true", acquisition and decoding circuit 100 may speculatively obtain the subsequent instruction "x = 1" for execution.

[0046] After executing an instruction (e.g., in execution pipeline 164), the condition of a conditional instruction may be actually determined, and the bias prediction from the bias prediction circuit 156 and the instruction prediction from the instruction prediction circuit 160 may be evaluated based on the execution result of the conditional instruction. In the illustrated embodiment, the bias table 158 of the bias prediction circuit 156 and / or the table 162 of the instruction prediction circuit 160 may be updated based on the evaluation. For example, when the result of the evaluation is that the condition of the conditional instruction is actually true, this means that the previous bias prediction from the bias prediction circuit 156 (which is the initial state or default value "00") is an incorrect prediction. Thus, in the bias table 158, the value of the conditional instruction may be changed from "00" (e.g., the initial state) to "10" (e.g., biased true). In Figure 3A this is illustrated by the change from circle 302 (e.g., corresponding to "00") to circle 306 (e.g., corresponding to "10"). By comparison, the evaluation of the conditional instruction may confirm that the previous instruction prediction from the instruction prediction circuit 160 is not an incorrect prediction. Thus, in the table 162, the value of the conditional instruction may be changed from "10" (e.g., weakly true) to "11" (e.g., strongly true), indicating that the instruction prediction circuit 160 receives a reward. In Figure 3B this is illustrated by the change from circle 316 (e.g., corresponding to "10") to circle 318 (e.g., corresponding to "11").

[0047] Conversely, when the result of the evaluation of the conditional instruction is that the condition of the conditional instruction is actually false, this means that the previous bias prediction from the bias prediction circuit 156 is an incorrect prediction. Thus, in the bias table 158, the value of the conditional instruction may be changed from "00" (e.g., the initial state) to "01" (e.g., biased false). In Figure 3A this is illustrated by the change from circle 302 (e.g., corresponding to "00") to circle 304 (e.g., corresponding to "01"). Additionally, the evaluation of the conditional instruction may indicate that the previous instruction prediction from the instruction prediction circuit 160 is also an incorrect prediction. Thus, in the table 162, the value of the conditional instruction may be changed from "10" (e.g., weakly true) to "01" (e.g., weakly true), indicating that the instruction prediction circuit 160 receives a penalty. In Figure 3B this is illustrated by the change from circle 316 (e.g., corresponding to "10") to circle 314 (e.g., corresponding to "01").

[0048] As Figure 3AAs indicated, once updated to the value "10" (e.g., biased true) or "01" (e.g., biased false), the value of the conditional instruction in the bias table 158 can be maintained as "10" or "01" until an incorrect prediction occurs. In other words, once the value of the conditional instruction in the bias table 158 is updated from its initial state, the bias prediction circuit 156 can inhibit changing it to another value until an incorrect prediction occurs. From an operational perspective, this means that the bias prediction circuit 158 can predict the condition of the conditional instruction unconditionally in the same way until the evaluation of the conditional instruction indicates that the bias prediction is an incorrect prediction. When this incorrect prediction occurs, the value of the conditional instruction in the bias table 158 can be updated from "10" or "01" to "11" (e.g., unbiased). In Figure 3A this is illustrated by the change from circle 304 (e.g., corresponding to "01") or 306 (e.g., corresponding to "10") to circle 308 (e.g., corresponding to "11"). Additionally, once updated to the value "11", the value of the conditional instruction in the bias table 158 can be maintained as "11" (e.g., unbiased) until the prediction circuit 156 resets the value to the initial state or default value "00".

[0049] As Figure 3B indicated, the value of the conditional instruction in table 162 can change from one value to another upon update, depending on whether the instruction prediction circuit 160 receives a reward or a penalty. For example, when the evaluation of the conditional instruction confirms that the instruction prediction from the instruction prediction circuit 160 is not an incorrect prediction, the instruction prediction circuit 160 can receive a reward to change the value of the conditional instruction in table 162 from a relatively weak prediction to a relatively strong prediction (e.g., from weakly true to strongly true, or from weakly false to strongly false), or remain at a relatively strong prediction (e.g., strongly true or strongly false). Conversely, when the evaluation of the conditional instruction indicates that the instruction prediction from the instruction prediction circuit 160 is an incorrect prediction, the instruction prediction circuit 160 can receive a penalty to change the value of the conditional instruction in table 162 from a relatively strong prediction to a relatively weak prediction (e.g., from strongly true to weakly true, or from strongly false to weakly false), or even from a relatively weak prediction to the opposite relatively weak prediction (e.g., from weakly true or weakly false, and vice versa). Note that in the illustrated embodiment, the value of the conditional instruction in table 162 may not change directly from one relatively strong prediction to the opposite relatively strong prediction (e.g., from strongly true to strongly false, and vice versa). Thus, table 162 can be considered to have a certain level of hysteresis.

[0050] In addition, as described above, when the condition of a conditional instruction is predicted to be biased true or biased false, e.g., when the value of the conditional instruction in the bias table 158 is "10" or "01", the fetch and decode circuit 100 may use the bias prediction from the bias predictor 156 to speculatively process the conditional instruction. Conversely, when the condition of a conditional instruction is predicted to not be biased true or biased false, e.g., when the value of the conditional instruction in the bias table 158 is "00" or "11", the fetch and decode circuit 100 may use the instruction prediction from the instruction prediction circuit 160 to speculatively process the conditional instruction.

[0051] In the illustrated embodiment, the bias table 158 and / or the table 162 may be implemented using one or more registers. Additionally, the fetch and decode circuit 100 may encode the bias prediction from the bias prediction circuit 156 and / or the instruction prediction from the instruction prediction circuit 160 in the instruction line containing the conditional instruction. For example, the fetch and decode circuit 100 may append the value of the conditional instruction (e.g., a 2-bit value) from the bias table 158 and / or the table 162 to the machine code of the instruction line that includes the conditional instruction in front, behind, or in the middle. Alternatively, the fetch and decode circuit 100 may re-encode the machine code of the instruction line that includes the conditional instruction to embed the prediction for the conditional instruction. For example, the fetch and decode circuit 100 may change the value of one or more bits of the machine code. Thus, when an instruction with the appended value is received at the Icache 102 and / or the decoder 154, the Icache 102 and / or the decoder 154 may identify the prediction of the conditional instruction and speculatively process the conditional instruction based on the prediction as described above.

[0052] In the illustrated embodiment, when a conditional instruction is predicted to be biased true or biased false, sometimes the fetch and decode circuit 100 may cause the conditional instruction to "bypass" the instruction prediction circuit 160. In the illustrated embodiment, to implement the "bypass", the fetch and decode circuit 100 may re-encode the conditional instruction as an unconditional instruction. Thus, the instruction prediction circuit 160 may treat the conditional instruction as an unconditional instruction and may not necessarily provide an instruction prediction for the conditional instruction being logged.

[0053] As described above, the bias prediction circuit 156 and / or the instruction prediction circuit 160 may incorrectly predict conditional instructions. Thus, the bias prediction circuit 156 and / or the instruction prediction circuit 160 may saturate. For example, when the code is executed by the processor 30 for a relatively long time, the bias prediction circuit 156 may experience enough incorrect predictions for one or more conditional instructions of the code. Thus, the values of the conditional instructions in the bias table 158 may change to the value "11". As described above, once the values change to "11", they may remain "11" until reset. Thus, to address saturation, in the illustrated embodiment, the bias prediction circuit 156 and / or the instruction prediction circuit 160 may respectively detect the occurrence of saturation and responsively reset the bias table 158 and / or the table 162. For example, the bias prediction circuit 156 may monitor the number of values "11" in the bias table 158. When it reaches a specified threshold (e.g., a specified percentage), the bias prediction circuit 156 may determine that the bias table 158 has saturated. Thus, the bias prediction circuit 156 may reset those values "11" to the initial state "00". Sometimes, the bias prediction circuit 156 may also reset other values in the bias table 158 (e.g., the entire bias table 158) to the initial state "00".

[0054] Now turning to Figure 4 , a flowchart illustrating an embodiment of the operation of the processor 30 including the bias prediction circuit 156 and the instruction prediction circuit 160 is shown. In the illustrated embodiment, a conditional instruction may be received at the fetch and decode circuit 100, as indicated in block 402. As described above, the conditional instruction may be loaded from the memory or cache 12 into the Icache 102, or fetched from the Icache 102 to the decoder 154. The fetch and decode circuit 100 may use the bias prediction circuit 156 to provide a bias prediction as to whether the condition of the conditional instruction is bias true or bias false, as indicated in block 404. In the illustrated embodiment, the bias prediction circuit 156 may use the bias table 158 for the bias prediction, as indicated in block 404. As described above, when the condition of the conditional instruction is predicted to be bias true or bias false, it means that the bias prediction circuit 156 predicts that the condition of the conditional instruction is always true or always false.

[0055] When the bias prediction from the bias prediction circuit 156 does not predict that the condition of the conditional instruction is bias true or bias false, the fetch and decode circuit 100 may use the instruction prediction circuit 160 to provide an instruction prediction as to whether the condition of the conditional instruction is true or false, as indicated in block 406. As described above, in the illustrated embodiment, the instruction prediction circuit 160 may be a TAGE predictor having a total of (M + 1) predictors, such as a base predictor T0 having a base table 162(0) and one or more additional (partial) tag predictors T having corresponding tables 162(i) (1 ≤ i ≤ M) i。Marker predictor T i (1 ≤ i ≤ M), Table 162(i) can be associated with a historical geometric sequence of a corresponding historical length.

[0056] As described above, in the illustrated embodiment, when the bias prediction circuit 156 predicts that the condition of the conditional instruction is bias-true or bias-false, the fetch and decode circuit 156 can cause the conditional instruction to "bypass" the instruction prediction circuit 160. Thus, the operation in block 406 can be avoided. For example, the fetch and decode circuit 100 can re-encode the conditional instruction as an unconditional instruction. Additionally, as described above, in the illustrated embodiment, the bias prediction from the bias prediction circuit 156 and the instruction prediction from the instruction prediction circuit 160 can be provided at different stages of processing the conditional instruction in the fetch and decode circuit. For example, when the conditional instruction is loaded from the memory or cache 12 into the Icache 102, the bias prediction circuit 156 can provide the bias prediction at the prefetch stage, and when the conditional instruction is fetched from the Icache 102 to the decoder 154, the instruction prediction circuit 160 can perform the instruction prediction at the fetch stage.

[0057] In the illustrated embodiment, the fetch and decode circuit 100 can use either the bias prediction from the bias prediction circuit 156 or the instruction prediction from the instruction prediction circuit 160 to speculatively determine the target address of the conditional instruction, as indicated in block 408. For example, the fetch and decode circuit 100 can speculatively determine the target address of the conditional instruction from which subsequent instructions can be obtained for execution based on the bias prediction from the bias prediction circuit 156 or the instruction prediction from the instruction prediction circuit 160.

[0058] In the illustrated embodiment, the fetch and decode circuit 100 can send the conditional instruction to the execution pipeline 164 for execution, as indicated in block 410. Additionally, the fetch and decode circuit 100 can receive an evaluation of the conditional instruction based on the execution result of the conditional instruction, as indicated in block 412. As described above, the execution of the conditional instruction can determine whether the condition of the conditional instruction is actually true or false, and thus determine whether the previous bias prediction from the bias prediction circuit 156 and / or the previous instruction prediction from the instruction prediction circuit 160 was an incorrect prediction.

[0059] In an illustrative embodiment, the bias prediction circuit 156 and / or the instruction prediction circuit 160 update their respective bias tables 158 and 162 based on the evaluation of a conditional instruction, as indicated in blocks 414 and 416. As described above in FIGS. 2-3, when the evaluation indicates that the condition of the conditional instruction is actually true or false, the update of the bias table 158 can change the value of the conditional instruction from an initial state such as "00" to "01" (e.g., indicating a bias false) or "10" (e.g., indicating a bias true), respectively, or change the value from "01" or "10" to "11" (e.g., indicating unbiased) when the evaluation indicates that a previous bias true or bias false prediction was actually an incorrect prediction. By comparison, when the evaluation confirms that the instruction prediction from the instruction prediction circuit 160 is not an incorrect prediction, the instruction prediction circuit 160 can update the value of the conditional instruction in the table 162(i) of the base and tag predictors Ti (0 ≤ i ≤ M) from a relatively weak prediction to a relatively strong prediction (e.g., from weak true "10" to strong true "11", or from weak false "01" to strong false "00"), or maintain the value at a relatively strong prediction (e.g., strong true "11" or strong false "00"); or when the evaluation indicates that the instruction prediction from the instruction prediction circuit 160 is an incorrect prediction, from a relatively strong prediction to a relatively weak prediction (e.g., from strong true "11" to weak true "10", or from strong false "00" to weak false "01"), or from a relatively weak prediction to the opposite relatively weak prediction (e.g., from weak true "10" or weak false "01", and vice versa).

[0060] Turning now to Figure 5 , a block diagram of an embodiment of a portion of a processor 30 including an instruction distribution circuit 520 and execution pipelines 504 and 506 is shown. In Figure 5 , the instruction distribution circuit 520 can receive a conditional instruction associated with a prediction from the fetch and decode circuit 100 and distribute the conditional instruction to one of a plurality of execution pipelines (such as 504 and 506) based on the predicted confidence of the conditional instruction. When it is determined that the conditional instruction has a relatively high confidence, the instruction distribution circuit can distribute the conditional instruction to the first execution pipeline 504. Conversely, when it is determined that the conditional instruction has a relatively low confidence, the instruction distribution circuit can distribute the conditional instruction to the second execution pipeline 506. One difference between the execution pipeline 504 and the execution pipeline 506 can be that the execution pipeline 506 (but not the execution pipeline 504) can have the ability to redirect the fetch and decode 100 to obtain instructions for execution from the correct target address when a conditional instruction is incorrectly predicted (also referred to as re-fetch).

[0061] Thus, when the execution pipeline 504 detects a misprediction of a conditional instruction, the execution pipeline 504 may have to use the execution pipeline 506 to re-fetch the instruction from the correct target address for execution. For example, the execution pipeline 504 may create a bubble in the execution pipeline 506 and then insert the conditional instruction into the bubble for the execution pipeline 506 to execute. Once the execution pipeline 506 executes the conditional instruction and also determines that the conditional instruction was mispredicted, the execution pipeline 506 may redirect the fetch and decode 100 to re-fetch the instruction from the correct target address for execution. For example, the execution pipeline 506 (instead of the execution pipeline 504) may have a communication path to the fetch and decode circuitry 100, and the execution pipeline 506 may instruct the fetch and decode circuitry 100 via the communication path to perform the re-fetch. Given that the conditional instruction has been executed in the execution pipeline 504, the second execution of the conditional instruction in the execution pipeline 506 may also be considered a re-execution or replay of the conditional instruction. Additionally, in the illustrated embodiment, the execution pipeline 506 may also use the bubble to execute one or more non-conditional instructions as well as the mispredicted conditional instruction. For example, the execution pipeline 506 may execute one or more non-conditional instructions created by the bubble in the same cycle as the mispredicted conditional instruction.

[0062] In the illustrated embodiment, when it detects a mispredicted conditional instruction, the execution pipeline 504 may not necessarily write the result back to any register or memory until the instruction from the correct target address is successfully executed by the execution pipeline 504. This ensures that only the correct result is written to the register or memory. However, this may also delay the exit of the execution pipeline 504 and thus cause an additional delay to the execution pipeline 404. In comparison, when the mispredicted conditional instruction is initially assigned to the execution pipeline 506, the execution pipeline 506 may detect the misprediction and directly cause the fetch and decode circuitry 100 to obtain the execution from the correct target address for execution, thus causing a minimal delay to the execution. Therefore, due to the different latencies in the processing of the mispredicted conditional instruction, the execution pipeline 504 may be considered a "slow" execution pipeline, while the execution pipeline 506 may be considered a "fast" execution pipeline. Sometimes, the execution pipeline 504 and the execution pipeline 506 may include the same stages or the same number of stages. In other words, for a conditional instruction without misprediction, the execution pipeline 504 and the execution pipeline 506 may not necessarily have different latencies, and there are only different latencies for the mispredicted conditional instruction because the execution pipeline 504 lacks the ability to directly instruct the fetch and decode circuitry 100 to perform the re-fetch. Alternatively, sometimes the execution pipeline 504 may have more stages or a greater number of stages than the execution pipeline 506. Therefore, regardless of whether the conditional instruction is mispredicted, the execution pipeline 504 may always have a greater latency than the execution pipeline 506.

[0063] In an illustrative embodiment, the prediction of a conditional instruction used by the instruction distribution circuit 502 to distribute a conditional instruction may be (a) a bias prediction from the bias prediction circuit 156 or (b) an instruction prediction from the instruction prediction circuit 160. For example, as described above, when the condition for the bias prediction circuit 156 to provide a conditional instruction is a bias prediction of bias true or bias false, the fetch and decode circuit 100 may use the bias prediction to speculatively process the conditional instruction. In this case, the instruction distribution circuit 502 may use the bias prediction from the bias prediction circuit 156 to determine the distribution of the conditional instruction. Conversely, when the bias prediction circuit 156 predicts that the condition of the conditional instruction is not bias true or bias false, the fetch and decode circuit 100 may use the instruction prediction from the instruction prediction circuit 160 to speculatively process the conditional instruction. In this case, the instruction distribution circuit 502 may use the instruction prediction from the instruction prediction circuit 160 to determine the distribution of the conditional instruction. In other words, the prediction of the conditional instruction disclosed herein may be a prediction of the conditional instruction, and the fetch and decode circuit 100 speculatively processes the conditional instruction based on the prediction.

[0064] In an illustrative embodiment, the confidence of the prediction may be determined with respect to one or more criteria. For example, when the prediction of the conditional instruction is a bias prediction from the bias prediction circuit 156 (e.g., when the conditional instruction is predicted to be bias true or bias false), the instruction distribution circuit 502 may determine that the prediction has a high confidence. Additionally, when the prediction is an instruction prediction from the instruction prediction circuit 160 (e.g., when the conditional instruction is not predicted to be bias true or bias false), if the instruction prediction is provided by a tag predictor T i with a saturation counter or a tag predictor T i with a higher-order table (e.g., when the instruction prediction circuit 160 is a TAGE predictor), then the instruction distribution circuit 502 may determine that the prediction has a high confidence. Otherwise, when the prediction of the conditional instruction fails to meet one or more of the above criteria, the instruction distribution circuit 502 may determine that the prediction has a low confidence.

[0065] When the confidence level is high, the instruction distribution circuit 502 may distribute conditional instructions to the "slow" execution pipeline 504. Conversely, when the confidence level is low, the instruction distribution circuit 502 may distribute conditional instructions to the execution pipeline 506 (e.g., a "fast" execution pipeline). From an operational perspective, this means that when a conditional instruction is predicted with high confidence, the instruction distribution circuit 502 may assume that the conditional instruction is less likely to be mispredicted, and thus the execution of the conditional instruction in the execution pipeline 504 (e.g., the "slow" execution pipeline) may have a lower probability of causing a refetch. In comparison, when a conditional instruction is predicted with low confidence, the instruction distribution circuit 502 may assume that the prediction is more likely to be incorrect. Therefore, the instruction distribution circuit 502 may distribute the conditional instruction to the execution pipeline 506 (e.g., the "fast" execution pipeline) to reduce the potential latency of a refetch.

[0066] Sometimes, the instruction distribution circuit 502 may perform load balancing between the execution pipeline 504 and the execution pipeline 506. For example, the instruction distribution circuit 502 may distribute conditional instructions to the execution pipeline 504 and the execution pipeline 506 based on the occupancy of the execution pipelines rather than the prediction of the conditional instructions. For example, when the execution pipeline 504 is overloaded and the execution pipeline 506 is underutilized, the instruction distribution circuit 502 may distribute conditional instructions associated with a high-confidence prediction to the execution pipeline 506 for execution.

[0067] Now turning Figure 6 , a flowchart illustrating an embodiment of the operation of a processor 30 including an instruction distribution circuit 502 and different execution pipelines 504 and 506 is shown. In the illustrated embodiment, a conditional instruction associated with a prediction may be received at the instruction distribution circuit 502, as indicated in block 602. As described above, the prediction of the conditional instruction may be (a) a bias prediction from the bias prediction circuit 156 or (b) an instruction prediction from the instruction prediction circuit 160.

[0068] In the illustrated embodiment, the instruction distribution circuit 502 may evaluate the prediction of the conditional instruction with respect to one or more criteria to determine the confidence level of the prediction, as indicated in block 604. For example, the instruction distribution circuit 502 may determine whether the prediction is a bias prediction (e.g., bias true or bias false) provided by the bias prediction circuit 156 or an instruction prediction (e.g., true or false) provided by a tag predictor T with a saturating counter i or a tag predictor T with a higher-order table of the instruction prediction circuit 160 i (e.g., when the instruction prediction circuit 160 is a TAGE predictor). If so, the instruction distribution circuit 502 may determine that the conditional instruction has a high confidence level. Otherwise, the instruction distribution circuit 502 may determine that the conditional instruction has a low confidence level.

[0069] The instruction distribution circuit 502 may assign a conditional instruction to one of a plurality of execution pipelines according to the predicted confidence of the conditional instruction regarding one or more criteria. For example, when the confidence is high, the instruction distribution circuit 502 may assign the conditional instruction to the execution pipeline 504 (e.g., the "slow" execution pipeline) for execution, as indicated in block 606. Otherwise, when the confidence is low, the instruction distribution circuit 502 may assign the conditional instruction to the execution pipeline 506 (e.g., the "fast" execution pipeline) for execution, as indicated in block 610.

[0070] When the conditional instruction is assigned to the execution pipeline 504, the execution of the conditional instruction may determine that the prediction of the conditional instruction is a misprediction, as indicated in block 608. In response, the execution pipeline 504 may cause the mispredicted conditional instruction to be re-executed or replayed in the execution pipeline 506, as indicated in block 610. As described above, in the illustrated embodiment, the execution pipeline 504 may create a bubble in the execution pipeline 506 and insert the mispredicted conditional instruction into the bubble for execution by the execution pipeline 506. As described above, the execution of the conditional instruction in the execution pipeline 506 may determine that the conditional instruction is mispredicted, as indicated in block 612. The execution pipeline 506 may direct the fetch and decode circuit 100 to obtain instructions from the correct target address of the conditional instruction for execution, as indicated in block 614.

[0071] Now turning to Figure 7 , a block diagram of an embodiment of a portion of a processor 30 including an instruction distribution circuit 520 and execution pipelines 504, 506, 708, and 710 is shown. In the illustrated embodiment, the execution pipeline 708 may be similar to the execution pipeline 504 (e.g., the "slow" execution pipeline), such that the execution pipeline 708 lacks the ability to directly re-fetch instructions by the fetch and decode circuit 100 for mispredicted conditional instructions. By comparison, the execution pipeline 710 may be similar to the execution pipeline 506 (e.g., the "fast" execution pipeline), such that the execution pipeline 710 may also be capable of directly re-fetching instructions by the fetch and decode circuit 100 for mispredicted conditional instructions. For example, similar to the execution pipeline 506, the execution pipeline 710 may also have a communication path to the fetch and decode circuit to direct re-fetching for mispredicted conditional instructions. Thus, in Figure 7 , the processor 30 includes two "slow" execution pipelines (e.g., the execution pipeline 504 and the execution pipeline 708) and two "fast" execution pipelines (e.g., the execution pipeline 507 and the execution pipeline 710). Note that Figure 7 is provided only as an example for illustrative purposes. Sometimes, the processor 30 may include fewer or more "slow" execution pipelines, and / or fewer or more "fast" execution pipelines.

[0072] AsFigure 7 As indicated, instruction distribution circuit 502 may distribute conditional instructions to execution pipelines 504, 506, 708, and 710 according to the predicted confidence of the conditional instructions. In the illustrated embodiment, the predicted confidence of a conditional instruction may be determined with respect to one or more criteria, as described above in Figures 5 to 6 As described. Thus, instruction distribution circuit 502 may distribute conditional instructions associated with a prediction of high confidence to execution pipeline 504 and execution pipeline 708 (e.g., the "slow" execution pipelines) for execution, and distribute conditional instructions associated with a prediction of low confidence to execution pipeline 506 and execution pipeline 710 (e.g., the "fast" execution pipelines) for execution.

[0073] In the illustrated embodiment, execution pipelines 504, 506, 708, and 710 may operate in parallel and thus process one or more conditional instructions approximately simultaneously. However, in the illustrated embodiment, only one of the "fast" execution pipelines (such as execution pipeline 506) may be used to re-execute or replay mispredicted conditional instructions provided from the "slow" execution pipelines (such as execution pipeline 504 and execution pipeline 708) (so as to cause re-fetching). Thus, when both execution pipeline 504 and execution pipeline 708 respectively detect mispredicted conditional instructions, processor 30 may use first misprediction selection circuit 712 to select one of the mispredicted conditional instructions from execution pipeline 504 and execution pipeline 708 for re-execution or replay in execution pipeline 506, as Figure 7 As indicated.

[0074] In the illustrated embodiment, the selection may be made respectively according to the ages of two mispredicted conditional instructions of execution pipeline 504 and execution pipeline 708. For example, first misprediction selection circuit 712 may compare the age of the first mispredicted conditional instruction in execution pipeline 504 with the age of the second mispredicted conditional instruction in execution pipeline 708, and cause the older of the two conditional instructions to be executed in execution pipeline 506. The age of a conditional instruction may be obtained in one of various ways. For example, fetch and decode circuit 100 may assign a number (such as Gnum) to a conditional instruction when the conditional instruction is decoded by decoder 154. For each instruction, Gnum may be a unique, monotonically increasing (or decreasing) number. Thus, a newer instruction may be assigned a smaller Gnum (or a larger Gnum), while an older instruction may be assigned a larger Gnum (or a smaller Gnum). Thus, first misprediction selection circuit 712 may compare the Gnums of two conditional instructions to select the older conditional instruction. Additionally, sometimes the age of a conditional instruction may also be determined based on the order of the conditional instruction in the reorder buffer (ROB) 108 of processor 30.

[0075] Once the first misprediction selection circuit 712 makes a selection, the corresponding execution pipeline (e.g., execution pipeline 504) may create a bubble in execution pipeline 506 and insert the selected conditional instruction into the bubble for execution by execution pipeline 506. Once execution pipeline 506 executes the conditional instruction and detects that it was mispredicted, execution pipeline 506 may direct fetch and decode circuit 100 to obtain an instruction from the correct target address of the mispredicted conditional instruction for execution, as described above. Note that the selection by the first misprediction selection circuit 712 may not necessarily mean that the unselected mispredicted conditional instruction will not be re-executed or replayed by execution pipeline 506. Instead, this only means that when both the "slow" execution pipeline 504 and execution pipeline 708 detect a misprediction approximately simultaneously, to resolve the conflict, one of the conditional instructions may be selected to cause a re-fetch first. Thereafter, the other unselected conditional instruction may be re-executed or replayed by execution pipeline 506 to direct another re-fetch.

[0076] However, in the illustrated embodiment, given that a plurality of execution pipelines including two "fast" execution pipelines 506 and 710 may process instructions in parallel, it is possible that execution pipeline 710 (e.g., the second "fast" execution pipeline) may also detect a mispredicted conditional instruction approximately simultaneously with execution pipeline 506 (e.g., the first "fast" execution pipeline) detecting a mispredicted conditional instruction. This may also create a conflict. As Figure 7 indicated, in such a case, the processor 30 may use a second prediction selection circuit 714 to select one of the two mispredicted conditional instructions from the two "fast" execution pipelines 506 and 710 for re-fetch. For example, the second misprediction circuit 714 may compare the age of the first mispredicted conditional instruction in execution pipeline 506 with the age of the second mispredicted conditional instruction in execution pipeline 710 and select the older of the two conditional instructions to direct fetch and decode circuit 100 to obtain an instruction from the target address for execution.

[0077] Now turning to Figure 8 , a flowchart is shown illustrating the operation of one embodiment of a processor 30 including an instruction distribution circuit 502 and different execution pipelines 504, 506, 708, and 710. In the illustrated embodiment, conditional instructions associated with a prediction may be received at instruction distribution circuit 502, as indicated in block 802. As described above, the prediction of the conditional instruction may be (a) a bias prediction from bias prediction circuit 156 or (b) an instruction prediction from instruction prediction circuit 160.

[0078] In an illustrative embodiment, the instruction distribution circuit 502 may evaluate the prediction of a conditional instruction with respect to one or more criteria to determine the confidence in the prediction, as indicated in block 704. The instruction distribution circuit 502 may assign the conditional instruction to one of a plurality of execution pipelines based on the confidence in the prediction of the conditional instruction with respect to one or more criteria. For example, when the confidence is high, the instruction distribution circuit 502 may assign the conditional instruction to one of execution pipeline 504 and execution pipeline 708 (e.g., the "slow" execution pipeline) for execution, as indicated in block 806. Otherwise, when the confidence is low, the instruction distribution circuit 502 may assign the conditional instruction to one of execution pipeline 506 and execution pipeline 710 (e.g., the "fast" execution pipeline) for execution, as indicated in block 812.

[0079] When the conditional instruction is assigned to one of execution pipeline 504 and execution pipeline 708, the execution of the conditional instruction may determine that the prediction of the conditional instruction is an incorrect prediction, as indicated in block 808. However, the other of execution pipeline 504 and execution pipeline 708 may also detect the incorrectly predicted conditional instruction at approximately the same time. Thus, to resolve the conflict, the processor 30 may use the first incorrect prediction selection circuit 712 to select one of the two incorrectly predicted conditional instructions to be re-executed or replayed by execution pipeline 506, as indicated in block 810. In an illustrative embodiment, the selection may be made based on the ages of the two conditional instructions. For example, the first incorrect prediction selection circuit 712 may compare the age of the first incorrectly predicted conditional instruction in execution pipeline 504 with the age of the second incorrectly predicted conditional instruction in execution pipeline 708, and cause the older of the two conditional instructions to be executed in execution pipeline 506, as indicated in block 812.

[0080] In an illustrative embodiment, the execution of the conditional instruction in execution pipeline 506 may determine that the conditional instruction was incorrectly predicted, as indicated in block 814. Additionally, another execution pipeline 710 (e.g., the second "fast" execution pipeline) may also detect the incorrectly predicted conditional instruction at approximately the same time as execution pipeline 506 detects the incorrectly predicted conditional instruction. Thus, the processor 30 may use the second incorrect prediction circuit 714 to select one of the two incorrectly predicted conditional instructions from execution pipeline 506 and execution pipeline 710, as indicated in block 816. Thus, the execution pipeline 506 or execution pipeline 710 of the selected incorrectly predicted conditional instruction may instruct the fetch and decode circuit 100 to obtain instructions from the correct target address of the selected incorrectly predicted conditional instruction for execution, as indicated in block 818.

[0081] Figure 9 is a block diagram of one embodiment of a processor 30 that includes Figures 1 to 8The bias prediction circuit 156, instruction prediction circuit 160, and / or instruction allocation circuit 502 described therein. Note that Figure 9 are provided only by way of example for illustrative purposes. Thus, sometimes the processor 30 may not include all but only some of the illustrated components. For example, sometimes the processor 30 may include the bias prediction circuit 156 and the instruction prediction circuit 160, but not the instruction allocation circuit 502.

[0082] In the illustrated embodiment, the processor 30 includes an fetch and decode unit 100 (including an instruction cache or ICache 102), a map-dispatch-rename (MDR) unit 106 (including a reorder buffer (ROB) 108), one or more reservation stations 110, one or more execution units 112, a register file 114, a data cache (DCache) 104, a load / store unit (LSU) 118, a reservation station (RS) 116 for the load / store unit, and a core interface unit (CIF) 122. The fetch and decode unit 100 is coupled to the MDR unit 106, which is coupled to the reservation stations 110, the reservation station 116, and the LSU 118. The reservation station 110 is coupled to the execution unit 28. The register file 114 is coupled to the execution unit 112 and the LSU 118. The LSU 118 is also coupled to the DCache 104, which is coupled to the CIF 122 and the register file 114. The LSU 118 includes a store queue 120 (STQ 120) and a load queue (LDQ 124).

[0083] The fetch and decode unit 100 may be configured to fetch instructions for execution by the processor 30 and decode the instructions into ops for execution. More specifically, the fetch and decode unit 100 may be configured to cache instructions previously fetched from memory (via the CIF 122) in the ICache 102 and may be configured to fetch the speculative path of instructions for the processor 30. As described above, in the illustrated embodiment, the fetch and decode unit 100 may include a bias prediction circuit 156 and an instruction prediction circuit 160 to provide corresponding predictions for conditional instructions. The fetch and decode unit 100 may implement various prediction structures to predict the fetch path. For example, a next fetch predictor may be used to predict the fetch address based on previously executed instructions. Various types of branch predictors may be used to verify the next fetch prediction or, if the next fetch predictor is not used, may be used to predict the next fetch address. The fetch and decode unit 100 may be configured to decode instructions into instruction operations. In some embodiments, a given instruction may be decoded into one or more instruction operations, depending on the complexity of the instruction. In some embodiments, particularly complex instructions may be microcoded. In such embodiments, the microcoding routine for the instruction may be encoded in the instruction operation. In other embodiments, each instruction in the instruction set architecture implemented by the processor 30 may be decoded into a single instruction operation, and thus the instruction operation may be substantially synonymous with the instruction (although its form may be modified by the decoder). The term "instruction operation" may be more briefly referred to herein as "operation" or "op".

[0084] The MDR unit 106 may be configured to map the op to speculative resources (e.g., physical registers) to allow out-of-order and / or speculative execution and may dispatch the op to the reservation stations 110 and 116. As Figure 9 indicated, in the illustrated embodiment, the MDR unit 106 may include an instruction allocation circuit 502. The op may be mapped from the architectural registers used in the corresponding instruction to physical registers in the register file 114. That is, the register file 114 may implement a set of physical registers, the number of which may be greater than the architectural registers specified by the instruction set architecture implemented by the processor 30. The MDR unit 106 may manage the mapping of architectural registers to physical registers. In one embodiment, there may be separate physical registers for different operand types (e.g., integer, media, floating point, etc.). In other embodiments, the physical registers may be shared among the operand types. The MDR unit 106 may also be responsible for tracking speculative execution and retiring ops or flushing mis-speculated ops. The reorder buffer 108 may be used to track the program order of the ops and manage retirement / flushing. That is, the reorder buffer 108 may be configured to track a plurality of instruction operations corresponding to instructions fetched by the processor and not retired by the processor.

[0085] When the source operands for an op are ready, the op can be scheduled for execution. In the illustrated embodiment, distributed scheduling is used for each execution unit in execution unit 28 and LSU 118, e.g., in reservation stations 116 and reservation stations 110. Other embodiments may implement a centralized scheduler if desired.

[0086] LSU 118 can be configured to execute load / store memory ops. Generally speaking, a memory operation (memory op) can be an instruction operation that specifies an access to memory (although the memory access may be done in a cache such as DCache 104). A load memory operation can specify a data transfer from a memory location to a register, while a store memory operation can specify a data transfer from a register to a memory location. A load memory operation can be referred to as a load memory op, a load op, or a load; and a store memory operation can be referred to as a store memory op, a store op, or a store. In one embodiment, a store op can be executed as a store address op and a store data op. A store address op can be defined as generating the address of the store, probing the cache to determine an initial hit / miss, and using the address and cache information to update the store queue. Thus, a store address op can take an address operand as a source operand. A store data op can be defined as delivering the store data to the store queue. Thus, a store data op may not take an address operand as a source operand, but may take a store data operand as a source operand. In many cases, the address operand of the store may be available before the store data operand, and thus the address can be determined and made available earlier than the store data. In some embodiments, e.g., if the store data operand is provided before one or more of the operands in the store address operand, it is possible for the store data op to execute before the corresponding store address op. Although in some embodiments a store op can be executed as a store address op and a store data op, other embodiments may not implement a store address / store data split. As an example, the remainder of the present disclosure will generally use the store address op (and store data op), but specific implementations that do not use the store address / store data optimization are also contemplated. The address generated via the execution of the store address op can be referred to as the address corresponding to the store op.

[0087] A load / store op may be received in reservation station 116, which may be configured to monitor the source operands of the operation to determine when they are available, and then issue these operations to the load or store pipeline respectively. When an operation is received in reservation station 116, some of the source operands may be available, and these operations may be indicated in the data received by reservation station 116 from MDR unit 106 for the corresponding operation. Other operands may become available via the execution of operations by other execution units 112 or even via the execution of earlier load ops. The operands may be collected by reservation station 116, or may be read from register file 114 when issued from reservation station 116, as Figure 6 shown.

[0088] In one embodiment, reservation station 116 may be configured to issue load / store ops out of order (different from their original order in the code sequence executed by processor 30 (referred to as "program order")) when these operands become available. To ensure that there is space in LDQ 124 or STQ 120 for the older ops bypassed by the newer ops in reservation station 116, MDR unit 106 may include circuitry for preallocating LDQ 124 or STQ 120 entries to the ops being transferred to load / store unit 118. If there are no available LDQ entries for a load being processed in MDR unit 106, MDR unit 106 may stop dispatching the load op and subsequent ops in program order until one or more LDQ entries become available. Similarly, if there are no STQ entries available for a store, MDR unit 106 may stop op dispatching until one or more STQ entries become available. In other embodiments, reservation station 116 may issue operations in program order, and the LRQ 46 / STQ 120 assignment may occur when issued from reservation station 116.

[0089] LDQ 124 may track the load from the initial execution of LSU 118 to the exit. LDQ 124 may be responsible for ensuring that memory ordering rules are not violated (between out-of-order executed loads, and between loads and stores). If a memory ordering violation is detected, LDQ 124 may signal a redirect for the corresponding load. The redirect may cause processor 30 to refresh the load op and subsequent ops in program order, and re-fetch the corresponding instruction. The speculative state of the load op and subsequent ops may be discarded, and these ops may be re-fetched and re-processed by fetch and decode unit 100 for execution again.

[0090] When a load / store address op is issued by reservation station 116, LSU 118 may be configured to generate the address accessed by the load / store and may be configured to translate the address from the effective address or virtual address created by the address operand of the load / store address op to the physical address actually used for memory addressing. LSU 118 may be configured to generate an access to DCache 104. For a load operation that hits DCache 104, data may be speculatively forwarded from DCache 104 to the destination operand of the load operation (e.g., a register in register file 114), unless the address hits a previous operation in STQ 120 (i.e., an older store in program order) or the load is replayed. Data may also be forwarded to a dependent op that has been speculatively scheduled and is now in execution unit 112. In such cases, execution unit 112 may bypass the forwarded data rather than the data output from register file 114. If store data is available for forwarding upon a hit in STQ, data output by STQ 120 may be forwarded rather than cache data. Cache misses and STQ hits where data cannot be forwarded may be reasons for replay, and in those cases load data cannot be forwarded. The cache hit / miss status from DCache 104 may be recorded in STQ 120 or LDQ 124 for later processing.

[0091] LSU 118 may implement multiple load pipelines. For example, in one embodiment, three load pipelines ("pipelines") may be implemented, but more or fewer pipelines may be implemented in other embodiments. Each pipeline may independently and in parallel with other loads execute different loads. That is, RS 116 may issue any number of loads in the same clock cycle, up to the number of load pipelines. LSU 118 may also implement one or more store pipelines, and specifically may implement multiple store pipelines. However, the number of store pipelines need not be equal to the number of load pipelines. In one embodiment, for example, two store pipelines may be used. Reservation station 116 may independently and in parallel with the store pipelines issue store address ops and store data ops. The store pipelines may be coupled to STQ 120, which may be configured to hold store operations that have been executed but not yet committed.

[0092] The CIF 122 may represent that the processor 30 is responsible for communicating with the rest of the system including the processor 30. For example, the CIF 122 may be configured to request data for which the DCache 104 and the ICache 102 miss. When the data is returned, the CIF 122 may signal the corresponding cache to fill the cache. For DCache filling, the CIF 122 may also notify the LSU 118. The LDQ 124 may attempt to schedule the replayed load waiting on the cache fill such that the replayed load may forward the filled data (referred to as a fill forwarding operation) when the filled data is provided to the DCache 104. If the replayed load is not successfully replayed during the fill, the replayed load may then be scheduled and replayed through the DCache 104 as a cache hit. The CIF 122 may also write back the modified cache line evicted by the DCache 104, merge the stored data that cannot be cached, etc. In another example, the CIF 122 may transmit interrupt-related signals to the processor 30, such as interrupt requests and / or acknowledgement / non-acknowledgement signals from / to the peripherals of the system including the processor 30.

[0093] In various embodiments, the execution unit 112 may include any type of execution unit. For example, the execution unit 112 may include integer, floating-point, and / or vector execution units. The integer execution unit may be configured to execute integer ops. Generally speaking, an integer op is an op that performs a defined operation (e.g., arithmetic, logical, shift / rotate, etc.) on integer operands. An integer may be a numerical value, where each value corresponds to a mathematical integer. The integer execution unit may include branch processing hardware for handling branch ops, or there may be a separate branch execution unit. As described above, the execution unit 112 and the associated reservation station 110 may implement one or more execution pipelines 164, 504, 506, 708, and / or 710 as Figures 1 to 8 described therein.

[0094] The floating-point execution unit may be configured to execute floating-point ops. Generally speaking, a floating-point op may be an op that has been defined to operate on floating-point operands. A floating-point operand is an operand that is represented as a mantissa (or significand) multiplied by a base to an exponent power. The exponent, the sign of the operand, and the mantissa / significand may be explicitly represented in the operand, and the base (e.g., in one embodiment, the base 2) may be implicit.

[0095] The vector execution unit may be configured to execute vector ops. Vector ops may be used, for example, to process media data (e.g., image data such as pixels, audio data, etc.). Media processing may be characterized by performing the same processing on a large amount of data, where each data is a relatively small value (e.g., 8 bits or 16 bits, as compared to 32 bits to 64 bits for integers). Thus, vector ops include single instruction multiple data (SIMD) or vector operations on operands representing multiple media data.

[0096] Thus, each execution unit 112 may include hardware configured to perform operations that are defined as the ops to be processed by a particular execution unit. Execution units are generally independent of each other in the sense that each execution unit may be configured to operate on ops issued for that execution unit and independent of other execution units. From another perspective, each execution unit may be an independent pipeline for executing ops. Different execution units may have different execution latencies (e.g., different pipeline lengths). In addition, different execution units may have different latencies for pipeline stages where bypassing occurs, and the clock cycles for speculative scheduling of dependent ops based on a load op may vary based on the type of the op and the execution unit 28 that will execute the op.

[0097] Note that any number and type of execution units 112 may be included in various embodiments, including embodiments having one execution unit and embodiments having multiple execution units.

[0098] A cache line may be an allocation / deallocation unit in a cache. That is, data within a cache line may be allocated / deallocated in the cache as a unit. The size of a cache line may vary (e.g., 32 bytes, 64 bytes, 128 bytes, or cache lines that are larger or smaller). Different caches may have different cache line sizes. ICache 102 and DCache 104 may each be a cache having any desired capacity, cache line size, and configuration. In various embodiments, there may be more additional levels of cache between DCache 104 / ICache 102 and the main memory.

[0099] At various points, load / store operations are referred to as being newer or older than other load / store operations. If a first operation is after a second operation in program order, the first operation may be newer than the second operation. Similarly, if a first operation precedes a second operation in program order, the first operation may be older than the second operation.

[0100] Now turning to Figure 10, which is a block diagram of one embodiment of a system 10 that may include one or more processors 30. In the illustrated embodiment, the system 10 may be implemented as a system-on-chip (SOC) 10 coupled to a memory 12. As the name implies, the components of the SOC 10 may be integrated onto a single semiconductor substrate that is an integrated circuit “chip”. In some embodiments, these components may be implemented on two or more discrete chips in the system. However, the SOC 10 will be used as an example herein. In the illustrated embodiment, the components of the SOC 10 include multiple processor clusters 14A - 14n, an interrupt controller 20, one or more peripheral components 18 (more briefly, “peripherals”), a memory controller 22, and a communication fabric 27. The components 14A - 14n, 18, 20, and 22 may all be coupled to the communication fabric 27. The memory controller 22 may be coupled to the memory 12 during use. In some embodiments, there may be more than one memory controller coupled to a corresponding memory. The memory address space may be mapped across the memory controllers in any desired manner. In the illustrated embodiment, the processor clusters 14A - 14n may include respective multiple processors (P) 30, and the respective processors (P) 30 may also include respective bias prediction circuits 156, respective instruction prediction circuits 160, and / or respective instruction allocation circuits 502 as described in Figures 1 to 9 The processor 30 may form the central processing unit (CPU) of the SOC 10. In one embodiment, one or more of the processor clusters 14A - 14n may not be used as the CPU.

[0101] As described above, the processor clusters 14A - 14n may include one or more processors 30 that may be used as the CPU of the SOC 10. The CPU of the system includes a processor that executes the main control software of the system, such as an operating system. Generally, the software executed by the CPU during use may control other components of the system to achieve the desired functions of the system. The processor may also execute other software such as application programs. The application programs may provide user functions and may rely on the operating system for lower-level device control, scheduling, memory management, etc. Thus, the processor may also be referred to as an application processor.

[0102] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor. A processor may encompass a processor core implemented on an integrated circuit having other components at a system-on-chip (SOC 10) or other level of integration. A processor may also include a discrete microprocessor, a processor core, and / or a microprocessor integrated into a multi-chip module implementation, a processor implemented as multiple integrated circuits, etc.

[0103] The memory controller 22 generally may include circuitry for receiving memory operations from other components of the SOC 10 and for accessing the memory 12 to complete the memory operations. The memory controller 22 may be configured to access any type of memory 12. For example, the memory 12 may be static random access memory (SRAM), dynamic RAM (DRAM) such as synchronous DRAM (SDRAM) including double data rate (DDR, DDR2, DDR3, DDR4, etc.) DRAM. Low power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controller 22 may include a memory operation queue for sorting (and potentially reordering) these operations and presenting these operations to the memory 12. The memory controller 22 may also include a data buffer for storing write data waiting to be written to the memory and read data waiting to be returned to the memory operation source. In some embodiments, the memory controller 22 may include a memory cache for storing recently accessed memory data. For example, in an SOC implementation, the memory cache may reduce power consumption in the SOC by avoiding re-accessing data from the memory 12 in cases where it is expected to be accessed again soon. In some cases, the memory cache may also be referred to as a system cache, which is different from private caches such as L2 caches or caches in a processor, which only serve certain components. Additionally, in some embodiments, the system cache need not be located within the memory controller 22.

[0104] The peripherals 18 may be any collection of additional hardware functions included in the SOC 10. For example, the peripherals 18 may include video peripherals such as: an image signal processor configured to process image capture data from a camera or other image sensor; a GPU; a video encoder / decoder; a scaler; a rotator; a mixer; a display controller, etc. The peripherals may include audio peripherals such as: a microphone; a speaker; an interface to the microphone and speaker; an audio processor; a digital signal processor; a mixer, etc. The peripherals may include interface controllers for various interfaces external to the SOC 10 (including interfaces such as universal serial bus (USB), peripheral component interconnect (PCI) (including PCI Express (PCIe)), serial and parallel ports, etc.). The peripherals may include networking peripherals such as a media access controller (MAC). Any set of hardware may be included.

[0105] The communication fabric 27 can be any communication interconnect and protocol for communicating among the components of the SOC 10. The communication fabric 27 can be bus-based, including shared bus configurations, crossbar switches, and hierarchical buses with bridges. The communication fabric 27 can also be packet-based and can be hierarchical with bridges, crossbar switches, point-to-point, or other interconnects.

[0106] Note that the number of components of the SOC 10 (and Figure 4 the number of sub-components of those components shown in Figure 4 , such as the processors 30 in each processor cluster 14A - 14n) can vary depending on the implementation. Additionally, the number of processors 30 in one processor cluster 14A - 14n can be different from the number of processors 30 in another processor cluster 14A - 14n. The number of each component / sub-component can be more or less than

[0107] Computer system

[0108] Turning now to Figure 11 , a block diagram of one implementation of the system 700 is shown. In the illustrated implementation, the system 700 includes at least one instance of a system-on-chip (SOC) 10 coupled to one or more peripheral devices 704 and an external memory 702, as Figure 10 described. A power management unit (PMU) 708 is provided that supplies a power supply voltage to the SOC 10 and one or more power supply voltages to the memory 702 and / or the peripheral device 154. In some implementations, more than one instance of the SOC 10 (e.g., SOC 10A - 10q) can be included (and more than one memory 702 can also be included). In one implementation, the memory 702 can include Figures 1 to 10 the memory 12 illustrated in

[0109] Depending on the type of the system 700, the peripheral device 704 can include any desired circuitry. For example, in one implementation, the system 704 can be a mobile device (e.g., a personal digital assistant (PDA), a smart phone, etc.), and the peripheral device 704 can include devices for various types of wireless communication, such as Wi-Fi, Bluetooth, cellular, global positioning system, etc. The peripheral device 704 can also include additional storage, which includes RAM storage, solid-state storage, or disk storage. The peripheral device 704 can include user interface devices, such as a display screen, which includes a touch display screen or a multi-touch display screen, a keyboard or other input devices, a microphone, a speaker, etc. In other implementations, the system 700 can be any type of computing system (e.g., a desktop personal computer, a laptop computer, a workstation, a network set-top box, etc.).

[0110] The external memory 702 may include any type of memory. For example, the external memory 702 may be SRAM, dynamic RAM (DRAM) (such as synchronous DRAM (SDRAM)), double data rate (DDR, DDR2, DDR3, etc.) SDRAM, RAMBUS DRAM, low-power versions of DDR DRAM (e.g., LPDDR, mDDR, etc.), and the like. The external memory 702 may include one or more memory modules to which memory devices may be mounted, such as single in-line memory modules (SIMMs), dual in-line memory modules (DIMMs), and the like. Alternatively, the external memory 702 may include one or more memory devices mounted on the SOC 10 implemented in a die-on-die or package-on-package manner.

[0111] As illustrated, the system 700 is shown to have applications in a wide range of fields. For example, the system 700 may be used as part of the chip, circuitry, components, etc. of a desktop computer 710, a laptop computer 720, a tablet computer 730, a cellular or mobile phone 740, or a television 750 (or a set-top box coupled to the television). Also illustrated are a smartwatch and a health monitoring device 760. In some embodiments, the smartwatch may include various general computing-related functions. For example, the smartwatch may provide access to email, cellular service, the user's calendar, and the like. In various embodiments, the health monitoring device may be a dedicated medical device or otherwise include dedicated health-related functions. For example, the health monitoring device may monitor the user's vital signs, track the user's proximity to other users for epidemiological social distancing purposes, contact tracing, provide communication to emergency services in the event of a health crisis, and the like. In various embodiments, the aforementioned smartwatch may or may not include some or any health monitoring-related functions. Other wearable devices are also envisioned, such as devices worn around the neck, devices implantable in the human body, glasses designed to provide augmented and / or virtual reality experiences, and the like.

[0112] The system 700 may also be used as part of a cloud-based service 770. For example, the previously mentioned devices and / or other devices may access computing resources in the cloud (i.e., remotely located hardware and / or software resources). Further, the system 700 may be used in one or more devices in the home other than those previously mentioned. For example, household appliances may monitor and detect notable situations. For example, various devices in the home (e.g., a refrigerator, a cooling system, etc.) may monitor the status of the device and provide an alert to the homeowner (or, for example, a repair agency) in the event of a specific event being detected. Alternatively, a thermostat may monitor the temperature in the home and may automate the adjustment of the heating / cooling system based on the homeowner's historical responses to various situations. Figure 11Also illustrated in the present disclosure is the application of the system 700 to various modes of transportation. For example, the system 700 can be used in the control and / or entertainment systems of airplanes, trains, buses, rental cars, private cars, watercraft from personal boats to cruise ships, scooters (for rental or private use), etc. In various cases, the system 700 can be used to provide automated guidance (e.g., self-driving vehicles), general system control, etc. Any number of other such embodiments are possible and contemplated. Note that Figure 11 The devices and applications illustrated are merely illustrative and are not intended to be limiting. Other devices are possible and contemplated.

[0113] Computer-readable storage medium

[0114] Now turning to Figure 12 , a block diagram of one embodiment of a computer-readable storage medium 800 is shown. Generally speaking, a computer-accessible storage medium can include any storage medium that can be accessed by a computer during use to provide instructions and / or data to the computer. For example, a computer-accessible storage medium can include storage media such as magnetic or optical media, e.g., disks (fixed or removable), tapes, CD-ROMs, DVD-ROMs, CD-Rs, CD-RWs, DVD-Rs, DVD-RWs, or Blu-ray. The storage medium can also include volatile or non-volatile memory media such as RAM (e.g., synchronous dynamic RAM (SDRAM), Rambus DRAM (RDRAM), static RAM (SRAM), etc.), ROM, or flash memory. The storage medium can be physically included within the computer to which the storage medium provides the instructions / data. Alternatively, the storage medium can be connected to the computer. For example, the storage medium can be connected to the computer via a network or wireless link such as a network-attached storage device. The storage medium can be connected via a peripheral interface such as a universal serial bus (USB). Generally speaking, the computer-accessible storage medium 800 can store data in a non-transitory manner, where non-transitory in this context can mean that instructions / data are not transmitted via a signal. For example, a non-transitory storage device can be volatile (and may lose stored instructions / data in response to a power outage) or non-volatile.

[0115] Figure 12 The computer-accessible storage medium 800 in Figure 10The database 804 of the SOC 10 described in [reference]. Generally speaking, the database 804 can be a database that can be read by a program and directly or indirectly used to manufacture the hardware including the SOC 10. For example, the database can be a behavioral-level description or a register transfer level (RTL) description of the hardware functions in a high-level design language (HDL) such as Verilog or VHDL. This description can be read by a synthesis tool, which can synthesize this description to generate a netlist including a list of gates from a synthesis library. The netlist includes a set of gates, which also represents the functions of the hardware including the SOC 10. Then the netlist can be placed and routed to generate a data set for describing the geometry to be applied to the mask. Then the mask can be used in various semiconductor manufacturing steps to produce one or more semiconductor circuits corresponding to the SOC 10. Alternatively, as needed, the database 804 on the computer-accessible storage medium 800 can be a netlist (with or without a synthesis library) or a data set.

[0116] Although the computer-accessible storage medium 800 stores a representation of the SOC 10, other embodiments may carry a representation of any part of the SOC 10 as needed, including Figure 4 any subset of the components shown. The database 804 can represent any of the above parts.

[0117] ***

[0118] This disclosure includes references to "embodiments" or groups of "embodiments" (e.g., "some embodiments" or "various embodiments"). Embodiments are different specific implementations or instances of the disclosed concepts. References to "an embodiment", "one embodiment", "a particular embodiment", etc. do not necessarily refer to the same embodiment. A large number of possible embodiments are envisioned, including those specifically disclosed, as well as modifications or alternatives that fall within the spirit or scope of this disclosure.

[0119] The present disclosure may discuss potential advantages that may result from the disclosed embodiments. Not all implementations of these embodiments will necessarily exhibit any or all of the potential advantages. Whether a particular implementation achieves an advantage depends on many factors, some of which are outside the scope of the present disclosure. In fact, there are many reasons why a particular implementation that falls within the scope of the claims may not exhibit some or all of the disclosed advantages. For example, a particular implementation may include other circuitry outside the scope of the present disclosure that, in combination with one of the disclosed embodiments, negates or diminishes one or more of the disclosed advantages. Additionally, sub-optimal design implementation of a particular implementation (e.g., a particular implementation technique or tool) may also negate or diminish the disclosed advantages. Even assuming an implementation of the technology, the achievement of an advantage may still depend on other factors, such as the circumstances of the environment in which the implementation is deployed. For example, the input provided to a particular implementation may prevent one or more of the problems addressed in the present disclosure from occurring on a particular occasion, and as a result, the benefits of its solution may not be realized. Given the existence of possible factors outside the scope of the present disclosure, it is hereby expressly stated that any potential advantages described herein should not be construed as claim limitations that must be met in order to establish infringement. Instead, the identification of such potential advantages is intended to illustrate the types of improvements available to designers who benefit from the present disclosure. Permanently describing such advantages (e.g., stating that a particular advantage "may occur") is not intended to convey doubt as to whether such advantages can actually be achieved, but rather to recognize the technical reality that the achievement of such advantages generally depends on additional factors.

[0120] Unless otherwise specified, the embodiments are non-limiting. That is, the disclosed embodiments are not intended to limit the scope of the claims drafted based on the present disclosure, even in cases where only a single example is described with respect to a particular feature. The embodiments disclosed in the present invention are intended to be illustrative rather than restrictive, without any contrary statement in the present disclosure. Accordingly, this application is intended to permit claims that cover the disclosed embodiments, as well as such alternatives, modifications, and equivalents, which will be apparent to those skilled in the art who are aware of the beneficial effects of the present disclosure.

[0121] For example, the features in this application may be combined in any suitable manner. Accordingly, new claims may be made during the prosecution of this application (or an application claiming priority therefrom) for any such combination of features. Specifically, with reference to the appended claims, the features of dependent claims may, where appropriate, be combined with the features of other dependent claims, including claims that depend from other independent claims. Similarly, the features from corresponding independent claims may be combined where appropriate.

[0122] Accordingly, while the appended dependent claims may be drafted such that each dependent claim depends from a single other claim, additional dependencies are contemplated. Any combination of dependent features consistent with the present disclosure is contemplated, and such combinations may be claimed in this application or in another application. In short, the combinations are not limited to those specifically recited in the appended claims.

[0123] In appropriate cases, it is also contemplated that claims drafted in one format or statutory type (e.g., apparatus) are intended to support corresponding claims in another format or statutory type (e.g., method).

[0124] ***

[0125] Because the present disclosure is a legal document, various terms and phrases may be subject to regulatory and judicial interpretation. Notice is hereby given that the following paragraphs, as well as the definitions provided throughout the present disclosure, will be used to determine how claims drafted based on the present disclosure are to be interpreted.

[0126] References to items in the singular form (i.e., a noun or noun phrase preceded by "a", "an", or "the") are intended to mean "one or more" unless the context clearly dictates otherwise. Thus, without accompanying context, a reference to an "item" in a claim does not exclude additional instances of that item. A "plurality" of items means a collection of two or more items.

[0127] The word "may" is used herein in an enabling sense (i.e., having the potential to, being able to), rather than in a mandatory sense (i.e., must).

[0128] The terms "comprising" and "including" and their forms are open-ended and mean "including but not limited to".

[0129] When the term "or" is used in this disclosure with respect to a list of options, it will generally be understood to be used in an inclusive sense unless the context provides otherwise. Thus, the statement "x or y" is equivalent to "x or y, or both", and thus encompasses 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, phrases such as "either x or y, but not both" make it clear that "or" is used in an exclusive sense.

[0130] The expressions "w, x, y, or z, or any combination thereof" or "... at least one of w, x, y, and z" are intended to cover all possibilities of a single element up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element in the set (e.g., w but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. The phrase "... at least one of w, x, y, and z" thus refers to at least one element in the set [w, x, y, z], thereby covering all possible combinations in this list of elements. This phrase should not be construed as requiring the existence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0131] In the present disclosure, various "labels" may precede a noun or noun phrase. Unless the context provides otherwise, different labels used for features (e.g., "first circuit", "second circuit", "specific circuit", "given circuit", etc.) refer to different instances of the feature. Additionally, unless otherwise stated, the labels "first", "second", and "third" do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) when applied to features.

[0132] The phrase "based on" or is used to describe one or more factors that influence a determination. This term does not exclude the possibility that additional factors may influence the determination. That is, the determination may be based solely on the specified factors or on the specified factors and other unspecified factors. Consider the phrase "determine A based on B". This phrase specifies that B is a factor used to determine A or that B influences the determination of A. This phrase does not exclude the possibility that the determination of A may also be based on some other factor such as C. This phrase is also intended to cover embodiments in which A is determined based solely on B. As used herein, the phrase "based on" is synonymous with the phrase "at least partially based on".

[0133] The phrases "responsive to" and "responsive" describe one or more factors that trigger an effect. This phrase does not exclude the possibility that additional factors may influence or otherwise trigger the effect, either in conjunction with or independent of the specified factors. That is, the effect may be responsive solely to these factors, or it may be responsive to the specified factors and other unspecified factors. Consider the phrase "perform A responsive to B". This phrase specifies that B is a factor that triggers the performance of A or triggers a particular result of A. This phrase does not exclude the possibility that the performance of A may also be responsive to some other factor, such as C. This phrase also does not exclude the possibility that the performance of A may be performed in response to B and C in conjunction. This phrase is also intended to cover embodiments in which A is performed responsive solely to B. As used herein, the phrase "responsive" is synonymous with the phrase "at least partially responsive to". Similarly, the phrase "responsive to" is synonymous with the phrase "at least partially responsive to".

[0134] ***

[0135] Within this disclosure, different entities (which may be variously referred to as "units", "circuits", other components, etc.) may be described or claimed as "configured to" perform one or more tasks or operations. This expression - [entity] [configured to [perform one or more tasks]] - is used herein to refer to a structure (i.e., a physical thing). More specifically, this expression is used to indicate that this structure is arranged to perform one or more tasks during operation. A structure may be considered "configured to" perform a certain task even if the structure is not currently being operated. Thus, an entity described or stated as "configured to" perform a certain task refers to a physical thing for implementing that task, such as a device, a circuit, a system having a processor unit, and a memory storing executable program instructions, etc. This phrase is not used herein to refer to intangible things.

[0136] In some cases, various units / circuits / components may be described herein as performing a set of tasks or operations. It should be understood that these entities are "configured to" perform those tasks / operations even if not specifically stated.

[0137] The term "configured to" is not intended to mean "configurable to". For example, an unprogrammed FPGA is not considered to be "configured to" perform a specific function. However, the unprogrammed FPGA can be "configurable to" perform that function. After being appropriately programmed, the FPGA can then be considered "configured to" perform a specific function.

[0138] For the purposes of U.S. patent applications based on this disclosure, stating in a claim that a structure is "configured to" perform one or more tasks is expressly intended not to invoke 35 U.S.C. § 112(f) for that claim element. If an applicant wishes to invoke section 112(f) during the prosecution of a U.S. patent application based on this disclosure, it will use the "means for [performing a function]" structure to phrase the claim element.

[0139] Different "circuits" may be described in this disclosure. These circuits or "circuit systems" constitute hardware that includes various types of circuit elements, such as combinational logic, clock storage devices (e.g., flip - flops, registers, latches, etc.), finite state machines, memories (e.g., random access memories, embedded dynamic random access memories), programmable logic arrays, etc. The circuit system can be custom - designed or taken from a standard library. In various specific embodiments, the circuit system may include digital components, analog components, or a combination of both as appropriate. Certain types of circuits may generally be referred to as "units" (e.g., decoding units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units also refer to circuits or circuit systems.

[0140] Accordingly, the disclosed circuits / units / components and other elements illustrated in the drawings and described herein include hardware elements, such as those described in the preceding paragraphs. In many cases, the internal arrangement of the hardware elements in a particular circuit can be specified by describing the function of that circuit. For example, a particular "decoding unit" may be described as performing the function of "processing the opcode of an instruction and routing the instruction to one or more of a plurality of functional units", which means that the decoding unit "is configured to" perform that function. For those skilled in the art of computers, this functional specification is sufficient to imply a set of possible structures for the circuit.

[0141] In various embodiments, as discussed in the preceding paragraphs, circuits, cells, and other elements defined by the functions or operations they are configured to implement form, with respect to their arrangement relative to one another and the manner in which such circuits / cells / components interact, a microarchitecture definition of the hardware that is ultimately fabricated in an integrated circuit or programmed into an FPGA to form a physical implementation of the microarchitecture definition. Thus, the microarchitecture definition is considered by those skilled in the art to be a structure from which many physical implementations can be derived, all of which physical implementations fall within the broader structure described by the microarchitecture definition. That is, one skilled in the art having a microarchitecture definition provided according to the present disclosure can, without undue experimentation and using the applications of an ordinary skilled person, implement the structure by encoding a description of the circuits / cells / components in a hardware description language (HDL) such as Verilog or VHDL. HDL descriptions are often expressed in a manner that can appear functional. However, to those skilled in the art, the HDL description is a means for translating the structure of a circuit, cell, or component into the next level of implementation details. Such HDL descriptions can take the form of behavioral code (which is typically non-synthesizable), register transfer language (RTL) code (which is typically synthesizable compared to behavioral code), or structural code (e.g., a netlist specifying logic gates and their connectivity). The HDL description can be synthesized sequentially for a cell library designed for a given integrated circuit manufacturing technology and can be modified for timing, power, and other reasons to obtain a final design database that is transmitted to a foundry to generate masks and ultimately produce an integrated circuit. Some hardware circuits or portions thereof can also be custom designed in a schematic editor and captured into an integrated circuit system design along with the synthesized circuitry. The integrated circuit can include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), as well as interconnectors between the transistors and circuit elements. Some embodiments can implement multiple integrated circuits coupled together to implement a hardware circuit, and / or discrete components can be used in some embodiments. Alternatively, the HDL design can be synthesized into a programmable logic array such as a field programmable gate array (FPGA) and implemented in the FPGA. This decoupling between the design of a set of circuits and the subsequent lower-level implementation of those circuits typically results in a situation where the circuit or logic designer never specifies a particular set of structures for the lower-level implementation beyond a description of what the circuit is configured to do, since the process is carried out at different stages of the circuit implementation process.

[0142] The fact that many different low-level combinations of circuit elements can be used to implement the same specifications of a circuit results in a large number of equivalent structures of the circuit. As noted, these low-level circuit implementations can vary depending on manufacturing technology, the foundry selected for manufacturing the integrated circuit, the cell library provided for a particular project, etc. In many cases, the choice of generating these different implementations through different design tools or methods can be arbitrary.

[0143] In addition, for a given implementation, a single implementation of a particular functional specification of a circuit typically includes a large number of devices (e.g., millions of transistors). Thus, the sheer volume of this information makes it impractical to provide a complete narrative of the low-level structures for implementing a single implementation, let alone the large number of equivalent possible implementations. For this reason, the present disclosure describes the structure of a circuit using functional shorthand commonly used in the industry.

[0144] The following clauses describe example implementations consistent with the accompanying drawings and the above description.

[0145] 1. A processor, the processor comprising:

[0146] An instruction distribution circuit configured to:

[0147] Receive a first conditional instruction associated with a prediction; and

[0148] Distribute the first conditional instruction to one of a plurality of execution pipelines for execution according to the prediction of the first conditional instruction, wherein the plurality of execution pipelines includes a first execution pipeline and a second execution pipeline, and the first execution pipeline and the second execution pipeline provide different time delays for conditional instructions handling mispredictions; and

[0149] The first execution pipeline is configured to:

[0150] In response to receiving the first conditional instruction:

[0151] Execute the first conditional instruction;

[0152] Determine that the prediction of the first conditional instruction is a misprediction; and

[0153] In response to determining that the prediction of the first conditional instruction is a misprediction, cause the first conditional instruction to be executed in the second execution pipeline. 2. The processor according to clause 1, wherein the second execution pipeline is configured to:

[0154] Execute the first conditional instruction;

[0155] Determine that the prediction of the first conditional instruction is a misprediction; and

[0156] In response to determining that the prediction of the first conditional instruction is an incorrect prediction, causing an instruction to be obtained from the target address of the first conditional instruction for execution.

[0157] 3. The processor according to clause 1, wherein the prediction of the first conditional instruction is associated with a confidence level regarding one or more criteria, and wherein the one or more criteria include at least one of (a) the prediction is provided by a biased prediction circuit using a bias table or (b) the prediction is provided by an instruction prediction circuit that at least partially uses one or more tables based on the prediction history of conditional instructions.

[0158] 4. The processor according to clause 3, wherein the instruction allocation circuit is configured to allocate the first conditional instruction to the first execution pipeline for execution when the prediction of the first conditional instruction is associated with a high confidence level, or to allocate the first conditional instruction to the second execution pipeline for execution when the prediction of the first conditional instruction is associated with a low confidence level.

[0159] 5. The processor according to clause 3, the processor further comprising:

[0160] A third execution pipeline configured to:

[0161] Execute a second conditional instruction associated with a high-confidence prediction; and

[0162] Determine that the prediction of the second conditional instruction is an incorrect prediction; and

[0163] A first incorrect prediction selection circuit configured to:

[0165] Before causing the first conditional instruction to be executed in the second execution pipeline, compare the age of the first conditional instruction from the first execution pipeline with the age of the second conditional instruction from the third execution pipeline; and

[0166] In response to determining that the first conditional instruction is older than the second conditional instruction, cause the first conditional instruction to be executed in the second execution pipeline.

[0167] 6. The processor according to clause 3, the processor further comprising:

[0168] A fourth execution pipeline configured to:

[0169] Execute a third conditional instruction associated with a low-confidence prediction; and

[0170] Determine that the prediction of the third conditional instruction is an incorrect prediction; and

[0171] A second misprediction selection circuit, the second misprediction selection circuit being configured to:

[0173] Compare the age of the first conditional instruction from the second execution pipeline with the age of the third conditional instruction from the fourth execution pipeline; and

[0174] Cause an instruction to be fetched for execution from the target address of the older one of the first conditional instruction and the third conditional instruction.

[0175] 7. The processor according to clause 1, wherein the second execution pipeline provides less latency for processing the mispredicted conditional instruction than the first execution pipeline, and wherein, to provide the less latency, the second execution pipeline is configured to directly instruct another circuit of the processor to fetch an instruction for execution from the target address of the first conditional instruction in response to determining that the prediction of the first conditional instruction is a misprediction

[0176] from the target address of the first conditional instruction.

[0177] 8. The processor according to clause 1, wherein the second execution pipeline provides less latency for processing the mispredicted conditional instruction than the first execution pipeline, and wherein the second execution pipeline includes fewer pipeline stages than the first execution pipeline.

[0178] 9. The processor according to clause 1, wherein, to cause the first conditional instruction to be executed in the second execution pipeline, the first execution pipeline is configured to:

[0179] Create a bubble in the second execution pipeline; and

[0180] Cause the first conditional instruction to be inserted into the bubble so that the first conditional instruction is executed by the second execution pipeline.

[0181] 10. The processor according to clause 9, wherein the second execution pipeline is further configured to:

[0182] Execute an unconditional instruction in the same cycle as the first conditional instruction.

[0183] 11. A method, the method comprising:

[0184] Receiving, at an instruction distribution circuit of a processor, a first conditional instruction associated with a prediction;

[0185] Using the instruction distribution circuit, distribute the first conditional instruction to one of a plurality of execution pipelines for execution according to the prediction of the first conditional instruction, wherein the plurality of execution pipelines includes a first execution pipeline and a second execution pipeline, and the first execution pipeline and the second execution pipeline provide different time delays for processing conditionally predicted instructions with incorrect predictions;

[0186] Execute the first conditional instruction using the first execution pipeline;

[0187] Using the first execution pipeline, determine that the prediction of the first conditional instruction is an incorrect prediction; and

[0188] In response to determining that the prediction of the first conditional instruction is an incorrect prediction, cause the first conditional instruction to be executed in the second execution pipeline.

[0189] 12. The method according to clause 11, the method further comprising:

[0190] Using the second execution pipeline, cause an instruction to be obtained from the target address of the first conditional instruction for execution.

[0191] 13. The method according to clause 13, wherein the prediction of the first conditional instruction is associated with a confidence level with respect to one or more criteria, and wherein the one or more criteria include at least one of (a) the prediction is provided by a bias prediction circuit using a bias table or (b) the prediction is provided by an instruction prediction circuit that at least partially uses one or more tables based on the prediction history of the conditional instruction.

[0192] 14. The method according to clause 13, the method further comprising:

[0193] When the prediction of the first conditional instruction satisfies the one or more criteria, distribute the first conditional instruction to the first execution pipeline for execution, or when the prediction of the first conditional instruction fails to satisfy the one or more criteria, distribute the first conditional instruction to the second execution pipeline for execution.

[0194] 15. The method according to clause 13, the method further comprising:

[0195] Execute a second conditional instruction associated with a prediction that satisfies the criteria using a third execution pipeline;

[0196] Determine that the prediction of the second conditional instruction is an incorrect prediction,

[0197] wherein causing the first conditional instruction to be executed in the second execution pipeline further includes:

[0198] Compare the age of the first conditional instruction with the age of the second conditional instruction using a first misprediction selection circuit; and

[0199] In response to determining that the first conditional instruction is older than the second conditional instruction, cause the first conditional instruction to be executed in the second execution pipeline.

[0200] 16. The method according to clause 13, the method further comprising:

[0201] Execute a third conditional instruction associated with a prediction that fails to meet the criterion using a fourth execution pipeline;

[0202] Determine that the prediction of the third conditional instruction is a misprediction;

[0203] Compare the age of the first conditional instruction with the age of the third conditional instruction using a second misprediction selection circuit; and

[0204] Cause an instruction to be obtained for execution from the target address of the older one of the first conditional instruction and the third conditional instruction.

[0205] 17. The method according to clause 13, the method further comprising:

[0206] Receive a fourth conditional instruction associated with a prediction that meets the criterion at the instruction allocation circuit; and

[0207] Allocate the fourth conditional instruction to the second execution pipeline for execution based on the occupancy of the second execution pipeline.

[0208] 18. The method according to clause 11, wherein the second execution pipeline provides less latency than the first execution pipeline for processing the mispredicted conditional instruction, and wherein the second execution pipeline has a communication path to directly obtain an instruction for execution from a target address in response to determining that the prediction of the first conditional instruction is a misprediction.

[0209] 19. A system, the system comprising:

[0210] One or more processors, each of the one or more processors comprising:

[0211] An instruction allocation circuit, the instruction allocation circuit being configured to:

[0212] Receive a first conditional instruction associated with a prediction; and select one of a plurality of execution pipelines for executing the first conditional instruction according to the prediction of the first conditional instruction, wherein the plurality of execution pipelines includes a first execution pipeline and a second execution pipeline, and the first execution pipeline and the second execution pipeline provide different time delays for processing conditional instructions with incorrect predictions; and

[0213] The first execution pipeline is configured to:

[0214] In response to receiving the first conditional instruction:

[0215] Execute the first conditional instruction;

[0216] Determine that the prediction of the first conditional instruction is an incorrect prediction;

[0217] And

[0218] In response to determining that the prediction of the first conditional instruction is an incorrect prediction, cause the first conditional instruction to be executed in the second execution pipeline.

[0219] 20. The system according to clause 19, wherein the second execution pipeline is configured to:

[0220] Execute the first conditional instruction;

[0221] Determine that the prediction of the first conditional instruction is an incorrect prediction; and

[0222] In response to determining that the prediction of the first conditional instruction is an incorrect prediction, cause an instruction to be obtained from the target address of the first conditional instruction for execution.

[0223] Once the above disclosure is fully understood, many variations and modifications will become apparent to those skilled in the art. It is intended that the following claims be interpreted to cover all such variations and modifications.

Claims

1. A processor, comprising: A bias prediction circuit configured to predict a respective condition of each conditional instruction based on one or more previous executions of the respective conditional instructions, wherein the predicted conditions each include a bias condition or a non-bias condition, and wherein the bias condition indicates that the respective condition is always true or always false; and An acquisition and decoding circuit configured to encode information identifying the conditional instruction as biased in a cache in response to the bias prediction circuit predicting a bias condition for the conditional instruction.

2. The processor according to claim 1, wherein, in order to encode information identifying the conditional instruction as biased, the acquisition and decoding circuit is configured to re-encode the conditional instruction as an unconditional instruction in the cache.

3. The processor according to claim 1, wherein, in order to encode information identifying the conditional instruction as biased, the acquisition and decoding circuit is configured to: Encode the conditional instruction in an entry of the cache; and Append information identifying that a bias condition has been predicted to the entry of the cache.

4. The processor according to claim 1, wherein, in order to predict the condition of a given conditional instruction among the respective conditional instructions, the bias prediction circuit is configured to: Identify a value for the conditional instruction in a bias table based on the address of the given conditional instruction; and Provide the predicted condition based on the identified value for the conditional instruction in the bias table.

5. The processor according to claim 4, wherein the value for the given conditional instruction in the bias table is one of a plurality of values including the following values: A first value indicating that the bias prediction circuit has not encountered a conditional instruction corresponding to the value in the bias table; A second value indicating that the condition of the conditional instruction is biased to false; A third value indicating that the condition of the conditional instruction is biased to true; And A fourth value indicating that the condition of the conditional instruction is not biased.

6. The processor according to claim 5, wherein when the execution of a given conditional branch indicates that the condition of the given conditional instruction is true or false respectively, the value for the given conditional instruction in the bias table is set from the first value to the second value or the third value, or when the execution of the conditional instruction indicates that the value for the conditional instruction in the bias table does not match the result of the execution of the given conditional instruction, the value for the given conditional instruction in the bias table is set from the second value or the third value to the fourth value.

7. The processor according to claim 1, wherein the bias prediction circuit is configured to predict the respective conditions of the respective conditional instructions at a first level of the processor corresponding to loading the respective conditional instructions into the cache of the processor, and wherein the processor further includes an instruction prediction circuit configured to provide respective instruction predictions at a second level of the processor corresponding to fetching from the cache the respective conditional instructions for which the bias prediction circuit has predicted non-bias conditions respectively.

8. A method, comprising: using a bias prediction circuit of a processor to predict the respective conditions of the respective conditional instructions based on one or more previous executions of the respective conditional instructions, wherein the predicted conditions each include a bias condition or a non-bias condition, and wherein the bias condition indicates that the respective condition is always true or always false; and using the fetch and decode circuitry of the processor to encode information identifying the conditional instruction as biased in the cache of the processor in response to the bias prediction circuit predicting a bias condition for the conditional instruction.

9. The method according to claim 8, wherein encoding the information identifying the conditional instruction as biased includes re-encoding the conditional instruction as an unconditional instruction in the cache.

10. The method according to claim 8, wherein encoding the information identifying the conditional instruction as biased includes: encoding the conditional instruction in an entry of the cache; and appending information identifying that a bias condition has been predicted to the entry of the cache.

11. The method according to claim 8, wherein predicting the condition of a given conditional instruction among the respective conditional instructions includes: identifying a value for the conditional instruction in a bias table based on the address of the given conditional instruction; and providing the predicted condition based on the identified value for the conditional instruction in the bias table.

12. The method according to claim 11, wherein the value for the given conditional instruction in the bias table is one of a plurality of values including: a first value indicating that the bias prediction circuit has not encountered a conditional instruction corresponding to the entry in the bias table for the value; a second value indicating that the condition of the conditional instruction is biased false; a third value indicating that the condition of the conditional instruction is biased true; and a fourth value indicating that the condition of the conditional instruction is not biased.

13. The method according to claim 12, wherein when the execution of a given conditional branch indicates that the condition of the given conditional instruction is true or false respectively, the value for the given conditional instruction in the bias table is set from the first value to the second value or the third value, or when the execution of the conditional instruction indicates that the value for the conditional instruction in the bias table does not match the result of the execution of the given conditional instruction, the value for the given conditional instruction in the bias table is set from the second value or the third value to the fourth value.

14. The method according to claim 8, wherein predicting the respective conditions of the respective conditional instructions is performed at a first level of the processor corresponding to loading the respective conditional instructions into a cache of the processor, and wherein the method further includes providing respective instruction predictions at a second level of the processor corresponding to fetching from the cache the respective conditional instructions for which a non-biased condition has been predicted by the bias prediction circuit using an instruction prediction circuit of the processor.

15. A system, comprising: one or more processors, each of the one or more processors including: a bias prediction circuit configured to predict the respective conditions of the respective conditional instructions based on one or more previous executions of the respective conditional instructions, wherein the predicted conditions each include a biased condition or a non-biased condition, and wherein the biased condition indicates that the respective condition is always true or always false; and a fetch and decode circuit configured to encode information identifying the conditional instruction as biased in a cache in response to the bias prediction circuit predicting a biased condition for the conditional instruction.

16. The system according to claim 15, wherein, in order to encode information identifying the conditional instruction as biased, the fetch and decode circuit is configured to re-encode the conditional instruction as a non-conditional instruction in the cache.

17. The system according to claim 15, wherein, in order to encode information identifying the conditional instruction as biased, the fetch and decode circuit is configured to: encode the conditional instruction in an entry of the cache; and append information identifying that a biased condition has been predicted to the entry of the cache.

18. The system according to claim 15, wherein, in order to predict the condition of a given conditional instruction among the respective conditional instructions, the bias prediction circuit is configured to: identify a value for the conditional instruction in a bias table based on the address of the given conditional instruction; and provide the predicted condition based on the identified value for the conditional instruction in the bias table.

19. The system according to claim 18, wherein the value for the given conditional instruction in the bias table is one of a plurality of values including the following values: a first value indicating that the bias prediction circuit has not encountered a conditional instruction corresponding to the value in the bias table; a second value indicating that the condition of the conditional instruction is biased to false; a third value indicating that the condition of the conditional instruction is biased to true; and a fourth value indicating that the condition of the conditional instruction is not biased.

20. The system according to claim 19, wherein when the execution of a given conditional branch indicates that the conditions of the given conditional instruction are true or false respectively, the value in the bias table for the given conditional instruction is set from the first value to the second value or the third value, or when the execution of the conditional instruction indicates that the value in the bias table for the conditional instruction does not match the result of the execution of the given conditional instruction, the value in the bias table for the given conditional instruction is set from the second value or the third value to the fourth value.

21. A processor, comprising: A hardware instruction distribution circuit; And A plurality of execution pipelines, including a first execution pipeline and a second execution pipeline that provide different latencies for processing mispredicted conditional instructions, Wherein the hardware instruction distribution circuit is configured to: Receive a first conditional instruction associated with a prediction; and Allocate the first conditional instruction to one of the first execution pipeline and the second execution pipeline for execution according to the prediction of the first conditional instruction; and Wherein the first execution pipeline is configured to: In response to receiving the first conditional instruction: Execute the first conditional instruction; Determine that the prediction of the first conditional instruction is a misprediction; and In response to determining that the prediction of the first conditional instruction is a misprediction, cause the first conditional instruction to be executed in the second execution pipeline.

22. The processor according to claim 21, wherein the second execution pipeline is configured to: Execute the first conditional instruction; Determine that the prediction of the first conditional instruction is a misprediction; and In response to determining that the prediction of the first conditional instruction is a misprediction, cause an instruction to be obtained from the target address of the first conditional instruction for execution.

23. The processor according to claim 21, wherein the prediction of the first conditional instruction is associated with a confidence level regarding one or more criteria, and wherein the one or more criteria include at least one of the following: (a) the prediction is provided by a bias prediction circuit using a bias table or (b) the prediction is provided by an instruction prediction circuit that at least partially bases on the prediction history of the conditional instruction using one or more tables.

24. The processor according to claim 23, wherein the instruction distribution circuit is configured to allocate the first conditional instruction to the first execution pipeline for execution when the prediction of the first conditional instruction is associated with a high confidence level, or allocate the first conditional instruction to the second execution pipeline for execution when the prediction of the first conditional instruction is associated with a low confidence level.

25. The processor according to claim 23, further comprising: A third execution pipeline, the third execution pipeline being configured to: Execute a second conditional instruction associated with a high-confidence prediction; And Determine that the prediction of the second conditional instruction is a misprediction; And A first misprediction selection circuit, the first misprediction selection circuit being configured to: Before causing the first conditional instruction to be executed in the second execution pipeline, compare the age of the first conditional instruction from the first execution pipeline with the age of the second conditional instruction from the third execution pipeline; And In response to determining that the first conditional instruction is older than the second conditional instruction, cause the first conditional instruction to be executed in the second execution pipeline.

26. The processor according to claim 23, further comprising: A fourth execution pipeline, the fourth execution pipeline being configured to: Execute a third conditional instruction associated with a prediction of low confidence; And Determine that the prediction of the third conditional instruction is an incorrect prediction; And A second incorrect prediction selection circuit, the second incorrect prediction selection circuit being configured to: Compare the age of the first conditional instruction from the second execution pipeline with the age of the third conditional instruction from the fourth execution pipeline; And Cause an instruction to be obtained for execution from the target address of the older one of the first conditional instruction and the third conditional instruction.

27. The processor according to claim 21, wherein for a conditional instruction handling the incorrect prediction, the second execution pipeline provides less latency than the first execution pipeline, and wherein, to provide less latency, the second execution pipeline is configured to directly instruct another circuit of the processor to obtain an instruction for execution from the target address of the first conditional instruction in response to determining that the prediction of the first conditional instruction is an incorrect prediction.

28. The processor according to claim 21, wherein for a conditional instruction handling the incorrect prediction, the second execution pipeline provides less latency than the first execution pipeline, and wherein the second execution pipeline includes fewer pipeline stages than the first execution pipeline.

29. The processor according to claim 21, wherein, to cause the first conditional instruction to be executed in the second execution pipeline, the first execution pipeline is configured to: Create a bubble in the second execution pipeline; and Cause the first conditional instruction to be inserted into the bubble so that the first conditional instruction is executed by the second execution pipeline.

30. The processor according to claim 29, wherein the second execution pipeline is further configured to: Execute an unconditional instruction in the same cycle as the first conditional instruction.

31. A method, comprising: Receiving, at an instruction distribution circuit of a processor, a first conditional instruction associated with a prediction; Using the instruction distribution circuit to allocate the first conditional instruction to one of a plurality of execution pipelines of the processor for execution according to the prediction of the first conditional instruction, wherein the plurality of execution pipelines includes a first execution pipeline and a second execution pipeline, and the first execution pipeline and the second execution pipeline provide different latencies for handling conditional instructions with incorrect predictions; Executing the first conditional instruction using the first execution pipeline; Using the first execution pipeline to determine that the prediction of the first conditional instruction is an incorrect prediction; And In response to determining that the prediction of the first conditional instruction is an incorrect prediction, causing the first conditional instruction to be executed in the second execution pipeline.

32. A system, comprising: One or more processors, wherein each of the one or more processors includes: A hardware instruction distribution circuit; and A plurality of execution pipelines, including a first execution pipeline and a second execution pipeline that provide different latencies for handling conditional instructions with incorrect predictions, Wherein the hardware instruction distribution circuit is configured to: Receive a first conditional instruction associated with a prediction; and Select one of a first execution pipeline and a second execution pipeline for executing a first conditional instruction according to a prediction of a first conditional instruction; and wherein the first execution pipeline is configured to: in response to receiving the first conditional instruction: execute the first conditional instruction; determine that the prediction of the first conditional instruction is an incorrect prediction; and in response to determining that the prediction of the first conditional instruction is an incorrect prediction, cause the first conditional instruction to be executed in the second execution pipeline.

33. A processor, comprising: a hardware instruction allocation circuit; and a plurality of execution pipelines, including a first execution pipeline and a second execution pipeline that provide different latencies for handling conditional instructions with incorrect predictions, wherein the hardware instruction allocation circuit is configured to: receive a first conditional instruction associated with a prediction confidence level among a plurality of prediction confidence levels including a low confidence level and a high confidence level; in response to determining that the prediction confidence level of the first conditional instruction is the low confidence level, allocate the first conditional instruction to the first execution pipeline for execution; and in response to determining that the prediction confidence level of the first conditional instruction is the high confidence level, allocate the first conditional instruction to the second execution pipeline for execution.