Predict-block skipping and early injection
By employing predict block skipping and early injection in branch prediction logic, the energy consumption and efficiency of out-of-order processors are improved through the identification and immediate processing of sequential instruction blocks.
Patent Information
- Application Number
- PCT/US2023/086349
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
Existing branch prediction mechanisms in out-of-order processors consume significant energy and can be improved for increased efficiency and reduced power consumption.
Implementing branch prediction logic that includes predict block skipping for sequential instruction blocks and early injection of subsequent blocks into the pipeline, utilizing branch target buffers to identify and skip branch prediction for sequential blocks and immediately process them.
Reduces power consumption and enhances instruction pipeline throughput by identifying and processing sequential blocks without unnecessary branch prediction, thereby improving overall processor efficiency.
Smart Images

Figure US2023086349_03072025_PF_FP_ABST
Abstract
Description
[0001] Predict-Block Skipping and Early Injection
[0002] BACKGROUND
[0003] This specification relates to computational devices that employ a branch predictor as part of an out-of-order (OoO) processor. A branch predictor anticipates the outcome of a branch instruction, allowing the processor to speculatively fetch and execute subsequent instructions before the outcome of the current branch is known.
[0004] In modem out-of-order (OoO) processors, a branch predictor includes one or more branch target buffers (BTB) that act as caches for target address of previously encountered branch instructions. When a branch instruction is encountered and executed, the target of the branch instruction is saved in a BTB which is queried with every new input branch instruction. When a previously encountered branch instruction is encountered, the target address is directly accessed in the BTB, which can help predict the outcome of the branch and speculatively execute subsequent instructions.
[0005] By predicting branch outcomes, processors can speculatively fetch and execute instructions without waiting for the branch to resolve. This approach maximizes pipeline utilization and increases instruction throughput. Despite the benefits realized by executing a branch predictor in a modem OoO processor, the approach comes with its own energy consumption and an opportunity to further improve branch prediction efficiency. An increase in prediction performance and a decrease in energy consumption would further contribute to the benefit to OoO processors that employ a branch predictor.
[0006] SUMMARY
[0007] This specification describes systems and methods for implementing branch prediction that includes predict block skipping to decrease power consumption and early injection to increase the instruction pipeline efficiency.
[0008] The fetch stage of a processor is responsible for fetching blocks of instructions to be decoded and executed. A block of instructions can include one or more branch instructions which can re-route the processor to execute a block of instmctions that is stored at a different address in memory. The processor might not know immediately if it should take a branch, as this can depend on the outcome of other instructions. However, the processor can use branch prediction logic to process a block of instructions, referred to as a “predict block”, before the fetch stage to guess if the branch will be taken and speculatively fetch and execute a block of instmctions before it's certain about this decision. Branch prediction logic often includes the use of one or more branch target buffers (BTB). A BTB is a specialized cache that stores the addresses of previously encountered branch instructions, the target addresses of the associated target instruction blocks, and a prediction of whether or not the branches are to be taken. When a branch instruction is encountered, a processor can query' one or more BTBs to predict the outcome of the branch decision and the location of the target instructions. This specification introduces an additional field to the one or more BTBs that indicates if the target instruction block related to a branch instruction is “sequential”, where a “sequential” instruction block is defined as a set of instructions that does not include a branch instruction. The processor can perform predict block skipping and early injection when a block of instructions is known to be sequential. Predict block skipping is the skipping of branch prediction logic for sequential predict blocks. Early injection is the immediate injection of a sequential predict block into the instruction processing pipeline and the early inj ection of the next instruction block into the fetch stage of the processor.
[0009] A processor that includes predict block skipping can consume less power compared to a processor that initiates branch prediction logic for every block of instructions. If the processor is aware of sequential predict blocks when they are retrieved by the fetch stage, it can skip the BTB queries and immediately move the block of instructions to the decode and execution steps of the processing pipeline.
[0010] A processor that includes early injection of instruction blocks can exhibit a greater instruction pipeline throughput compared to a processor that does not include early injection, where instruction pipeline throughput is defined as the number of instructions processed per unit time. By detecting a sequential block of instructions and processing it immediately along with the current block of instructions (the block of instructions that contains the branch instruction), the processor can immediately fetch the next block of instructions to be processed.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS
[0012] FIG. 1 is an overview of an example system implementation.
[0013] FIG. 2 is an example process where instructions are processed by the example system implementation.
[0014] FIG. 3 is an example process where instructions are processed by the example system implementation. FIG. 4 is an example process where instructions are processed by the example system implementation.
[0015] FIG. 5 is a detailed view of an example system implementation.
[0016] DETAILED DESCRIPTION
[0017] FIG. 1 is an overview of logical stages of an example processor system that can perform sequential block detection, early injection, and predict block skipping. The system is an example of a system that can implement out-of-order execution and can generally execute instructions through five main stages: a branch prediction stage 100, a fetch stage 102, a decode stage 104, a dispatch stage 106, an execution stage 108, and a commit stage 112. Each of these stages can be implemented using digital logic circuitry of any appropriate device, e.g., an integrated circuit device. Each stage can be performed by one or more functional modules of the device. In some implementations, each stage is implemented by a distinct functional module. The various modules as mentioned above may be implemented using various logic circuitry7components, to include AND, OR, NOT, NAND, or XOR gates. Other implementations may choose to use other circuitry7components or data processing apparatus.
[0018] The branch prediction stage 100 can predict the outcome of branch instructions in order to support speculative execution. For example, the branch prediction stage 100 can make a prediction of whether a branch will be taken before the conditional information that determines the decision is available. The value that determines if the branch is taken may still be executing before the branch prediction stage makes a prediction. The ability to make accurate predictions of whether the branch is taken before the conditional information is available helps ensure the instruction pipeline, e.g., a sequence of instructions ready to be executed, is maximally utilized.
[0019] The fetch stage 102 is responsible for retrieving each subsequent block of instructions for execution. The instructions can be fetched from an instruction cache if they have been previously executed, from the output of the branch prediction stage 100, or from the main memory7if the instruction address is not found in the instruction cache. As illustrated, the branch prediction stage 100 can perform predict block skipping 116 and early injection 118. The fetch stage 102 can perform sequential block detection 120. When the branch prediction stage 100 detects a branch instruction in the set of incoming instructions referred to as a predict block, the processor can perform branch prediction logic to predict whether the branch will be taken.
[0020] If the branch prediction stage 100 predicts a branch to be taken, the anticipated target address of the branch instruction can be determined, and the processor can fetch the set of instructions that corresponds to the target address of the predicted branch to be taken. The corresponding instructions stored at the target address can be speculatively executed before the current set of instructions is finished executing. If the system later determines that the prediction was incorrect, a branch misprediction occurs which can lead to performance penalties in processor execution because the processor may have speculatively executed instructions along the wrong path based on the incorrect prediction. In the case of a branch misprediction, the processor can roll back the results of the speculatively executed instructions and redirect execution to the correct path.
[0021] The branch prediction stage 100, which can implement branch prediction logic, can query a branch target buffer (BTB), a specialized cache in a processor that stores the target addresses of previously executed branch instructions to speed up the process of branch prediction, to determine if a target for the taken branch is known. If the target of the taken branch is known, the fetch stage 102 can fetch the instructions corresponding to the target address and the processor can speculatively execute them.
[0022] In general, BTBs can be indexed by the instruction block address, which is the location in memory where the processor will direct the program to continue execution, and contain one or more fields. A BTB entry typically contains a prediction of whether the branch is taken. In addition, a BTB record typically contains the address of the target block of instructions. In this specification, the BTBs can also include a field that indicates if the target block of instructions is sequential, where a sequential block of instructions is a block of instructions that does not include a branch instruction. In other words, the BTB entries can include the target instruction block address, w hether the branch is taken or not taken, and if the target block of instructions is sequential. An example subset of BTB entries is illustrated in reference to Table 1. The significance of the "‘Sequential Target” field in the BTB entries will be described in detail in reference to the Figures below and in reference to predict block skipping 116, early injection 118, and sequential block detection 120.
[0023] TABLE 1. EXAMPLE BTB ENTRIES.
[0024] When the processor queries one or more BTBs and receives a record at the branch instruction address, the successful query is referred to as a “hit”. When the processor does not receive a record at the branch instruction address from at least one BTB, the unsuccessful query is referred to as a “miss”. A miss indicates the branch instruction address has not been previously encountered or the instruction address does not point to a branch instruction. In addition, the processor can receive a hit from the one or more BTBs with a “not taken” determination. In this case, the processor can continue executing sequential instructions.
[0025] An advantage of maintaining this information in the BTB is that it provides the ability for a processor to perform predict block skipping 116. Predict block skipping 116 allows the processor to skip the branch prediction stage 100 when the predict block is known to be sequential. For example, if the branch prediction stage 100 predicts a branch will be taken and the target block of instructions is sequential, e.g.. the first entry of the example BTB entries illustrated in Table 1, the processor can implement instruction fetch logic that results in predict block skipping 116. Predict block skipping 116 leads to power savings because specific modules that are related to branch prediction 100 are not activated. Predict block skipping 116 is described in detail with reference to FIG. 2.
[0026] Another advantage of maintaining this information in the BTB is that it provides the ability for a processor to perform early injection 118 of instruction blocks. Early injection 118 allows the processor to immediately inject the predict block into the fetch stage 102 along with an additional block of instructions. Since the predict block will not enter the branch prediction stage 100 at the beginning of the processing like a normal block of instructions, the processor can pre-emptively fill the beginning of the pipeline with the next set of instructions. Early injection 118 leads to performance improvement through increased pipeline efficiency because the predict block can be immediately injected into the pipeline and the next block of instructions in the instruction queue can be injected into the beginning of the pipeline immediately. The instruction fetch logic that includes early injection 118 is described in detail in relation to FIG. 3.
[0027] The example fetch stage 102 can contain a sequential block detection module 120, with which the system can determine if the current set of instructions contains a branch instruction. If the current set of instructions is determined to be sequential, this information can be saved in a database, e.g., a BTB, to be queried the next time the set of instructions is encountered. When the current set of instructions is the target of a future branch instruction, the database holds the information that the set of instructions is sequential and therefore eligible for early injection and predict block skipping.
[0028] The decode stage 104 can analyze the incoming instructions and convert them into a format that can be understood and executed by the execution units 108. The dispatch module 106 can act as a buffer that holds decoded instructions before they are sent to execution stage 108. The dispatch module 106 may perform several functions including register renaming, dependency checking, and issuing instructions to available execution units.
[0029] The execution units 108 can perform the operations as defined by the instructions.
[0030] The execution units 108 can be designed to perform specialized operations such as arithmetic, logic operations, or memory access. As a superscalar processor, the system can execute instructions out-of-order to maximize the number of instructions processed per unit time. The superscalar processor can have multiple execution units 108 that can carry out the operations in parallel.
[0031] The reorder buffer 110 is responsible for keeping track of the correct order of instruction execution. The reorder buffer 110 can communicate with the dispatch module 106 to understand the expected order of instruction execution. The reorder buffer 110 can hold the results of executed instructions until they are ready to be committed by the commit module 112.
[0032] The commit stage 112 is responsible for finalizing the state of the processor according to the results of executed instructions. The primary function of the commit stage is to manage and clean up the various data structures that track the status of instructions currently in flight. This cleanup process can involve updating the processor's state to reflect the completion of these instructions, ensuring that the changes they have made are permanent and in line with the original program order. The commit stage thus plays a crucial role in maintaining the consistency and integrity of the processor's state, marking the point at which the execution of instructions is officially completed, and their effects are fully integrated into the system's ongoing operation.
[0033] FIG. 2 is a flow diagram of an example process that can implement predict block skipping. For convenience, the process will be described as being performed by any appropriate processor configured to operate in accordance with this specification.
[0034] The processor receives instruction addresses (202). As described above, the branch prediction stage is responsible for retrieving each subsequent block of instruction addresses before sending the instruction addresses to the fetch stage and the remainder of the instruction processing pipeline.
[0035] The processor can query one or more branch target buffers (204) for each instruction address in the block of instruction addresses. Unless specifically instructed to skip the branch prediction process, which includes executing queries to the one or more BTBs, the processor can query the one or more BTBs 204 to determine if one or more instruction addresses of the incoming set of instruction addresses have been previously encountered and determined to be a branch instruction. As illustrated in FIG. 2, in the case of a hit, e.g., an instruction address within the set of incoming instruction addresses is a previously encountered branch, the processor can check if the branch is predicted to be taken 206. If the processor determines the branch is to be taken, the processor injects the predict block to the standard fetch stage 207. The processor can check if the target block is sequential 208. If the processor predicts the target block to be sequential, the processor can inject the target block into the fetch stage and skip branch prediction 210. Since it is known that the target block is sequential, e.g., it does not contain a branch instruction, branch prediction is an unnecessary’ use of processing power which can be skipped.
[0036] The processor can generate a signal that is registered during the branch prediction stage 100 that indicates an incoming set of instructions is sequential. In other words, when a target block is determined to be sequential and injected into the instruction processing pipeline, a signal can be established to direct the processor to skip the branch prediction stage for the next block of instructions.
[0037] One or more BTBs can return a miss, e.g., they can fail to return an entry that corresponds to a queried instruction address. A BTB miss may occur when a branch instruction has not been encountered yet or if the instruction does not correspond to a branch instruction. Alternatively, the BTBs can return an entry for a particular instruction address with a “not taken” value. In both cases, where the BTBs return a miss and the BTBs return a not taken value, the processor can process sequential instructions 212.
[0038] Alternatively, one or more BTBs can return a hit, contain a target instruction block address, predict the branch to be taken, but return a “not sequential” or “unknown if sequential” value. In this case, the processor recorded the target instruction block to be nonsequential or undetermined. The processor can inject the target instruction block to the standard fetch stage 214 without skipping branch prediction for the incoming predict block. Since the target block of instructions may contain a branch instruction, branch prediction is initiated. FIG. 3 is a flow diagram of an example process that can implement early injection of instruction blocks. For convenience, the process will be described as being performed by any appropriate processor configured to operate in accordance with this specification.
[0039] The processor receives instruction addresses (302). As described above, the branch prediction stage is responsible for retrieving each subsequent block of instruction addresses before sending the instruction addresses to the fetch stage and the remainder of the instruction processing pipeline.
[0040] The processor can query one or more branch target buffers (304) for each instruction address in the block of instruction addresses. Unless specifically instructed to skip the branch prediction process, which includes executing queries to the one or more BTBs, the processor can query the one or more BTBs 304 to determine if one or more instruction addresses of the incoming set of instruction addresses have been previously encountered and determined to be a branch instruction. As illustrated in FIG. 3, in the case of a hit, e.g., an instruction address within the set of incoming instruction addresses is a previously encountered branch, the processor can check if the branch is predicted to be taken 306. If the branch is predicted to be taken, the processor can inject the predict block into the standard fetch stage 307. If the branch is not predicted to be taken, the processor can process the sequential instructions 314. If the branch is predicted to be taken, the processor can check if the target block is sequential 308. In other words, the processor can check if the block of instructions that is the target of the taken branch instruction contains a branch instruction. If the target instruction block is determined to be sequential, the processor can inject an additional block of instructions to the fetch stage 310. If the target instruction block is determined to be not sequential, the process does nothing 316 and continues to process the instructions in the instruction queue sequentially.
[0041] For example, consider a block of instructionsC'A” that contains a single branch instruction. Consider the scenario in which the processor queries a BTB, receives a hit, determines the branch is to be taken, and that the target block “B”, which corresponds to the target of the branch instruction contained in A, is sequential. The processor can identify this pattern and continue through the instruction processing pipeline with a new set of instructions that includes the original block of instructions A truncated at the branch instruction, and the target block of instructions B. In other words, the block of instructions that is fetched can consist of instructions A (truncated at the branch instruction) and instructions B (sequential target block). The processor is responsible for providing each subsequent block of instruction addresses to the branch prediction stage 100 or skipping the branch prediction stage 100 and providing the block of instruction addresses directly to the fetch stage 102. If the processor injects the predict block early, the processor can immediately fetch the next block of instructions from the instruction queue. Since the processor can inject the block of instructions early and can inject the subsequent block of instructions in the following cycle, the instruction pipeline efficiency is improved.
[0042] In addition to one or more BTBs returning a hit, one or more BTBs can return a miss, e.g. they can fail to return an entry' that corresponds to an instruction address. This may occur when a branch instruction has not been encountered yet or if the instruction does not correspond to a branch instruction. Alternatively, the BTBs can return an entry for a particular instruction address with a ‘‘not taken” value. In both cases, where one or more BTBs return a miss and one or more BTBs return a not taken value, the processor should process the sequential instructions 314 in the current set of instructions.
[0043] Alternatively, one or more BTBs can return a hit, contain a target block address, predict the branch to be taken, but return a “not sequential” value. In this case, the processor previously recorded the target block to be non-sequential, e g., to contain a branch instruction within the set of instructions. The processor can inject the target block of instructions through the standard instruction processing pipeline without skipping the branch prediction. Since a branch instruction may be present in the block of instructions, branch prediction is initiated.
[0044] FIG. 4 is a flow diagram of an example process that can implement sequential block detection. For convenience, the process will be described as being performed by any appropriate processor configured to operate in accordance with this specification.
[0045] The processor receives instruction addresses 402. As described above, the branch prediction stage is responsible for retrieving each subsequent block of instruction addresses before sending the instruction addresses to the fetch stage and the remainder of the instruction processing pipeline. The block of instructions may include one or more branch instructions. The processes described in relation to FIG. 2 and FIG. 3 assume one or more BTBs contain information in relation to whether a block of instructions is sequential, where a sequential block of instructions does not contain a branch instruction.
[0046] By analyzing the block of instructions after the block of instructions has been fetched and decoded, the processor can detect if a set of instructions is sequential 404. The processor can update one or more BTBs 406 with the determination of whether a set of instructions is sequential for the corresponding entries. FIG. 5 is a flow diagram of a detailed example process that can include branch prediction, early injection detection, and sequential instruction detection. For convenience, the process will be described as being performed by any appropriate processor configured to operate in accordance with this specification.
[0047] The processor receives instruction addresses 502. As described above, the branch prediction stage is responsible for retrieving each subsequent block of instruction addresses before sending the instruction addresses to the fetch stage and the remainder of the instruction processing pipeline. If the processor does not consider the block of instruction addresses to be sequential as described in relation to FIG. 2 and FIG. 3, the processor can query one or more BTBs to determine if the block of instruction addresses contains a previously encountered branch instruction.
[0048] To accommodate for variations in system design and memory requirements, several BTBs of various levels can store branch prediction information.
[0049] The level of the BTB can indicate the size or speed of the BTB. If the processor queries the first level BTB 504a which does not contain an entry for the instruction address, the processor can query a second level BTB 504b where the second level BTB can be larger and in some cases, slower than the first level BTB. Similarly, if the processor queries the second level BTB 504b and it does not contain an entry' for the instruction address, the processor can query a third level BTB 504c where the third level BTB can be larger and in some cases, slower than both the first and second level BTBs. To accommodate specific design requirements and memory requirements, the processor can query additional BTBs depending on the particular architecture of the branch prediction module. If the processor does not find entries associated with the instruction address in the plurality of BTBs, the processor can assume that the one or more BTBs do not have information about the instruction associated with the instruction address.
[0050] If the processor detects a sequential target instruction block and injects the block of instructions early into the instruction processing pipeline, the processor detects early injection 506a-b eligibility in response to a BTB hit, as previously described in relation to FIG. 3. The timing of the early injection detection depends on which BTB returns a hit with a sequential instruction block as its target. If the block of instructions is eligible for early injection, it can be directly injected into the current pipeline cycle along with an additional instruction block, as described in relation to FIG. 3.
[0051] The processor can detect a sequential block 508. In other words, a processor can be configured to detect if a sequence of instructions in an instruction block contains a branch instruction. If a block of instructions is determined to be sequential by the sequential instruction detector, the processor can update one or more BTBs 510 to include the sequential block information, as described in relation to FIG. 4. The processor can use the sequential block information in the future for power saving in relation to the predict block skipping configuration as described in relation to FIG. 2 or for increased instruction pipeline efficiency in relation to the early injection configuration as described in relation to FIG. 3.
[0052] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[0053] The term “data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0054] In addition to the embodiments described above, the following embodiments are also innovative:
[0055] Embodiment 1 is a computing device comprising: a branch target buffer configured to store target addresses of previously executed branch instructions; logic configured to fetch blocks of instructions, wherein the logic is configured to perform operations comprising: fetching a first block of first instructions according to an instruction block size; determining that the first block is a sequential instruction block; in response, designating the first block in the branch target buffer as being a sequential instruction block; receiving a request to execute a second block of instructions; determining, from querying the branch target buffer, that the second block of instructions is a sequential instruction block; and in response, fetching a group of second instructions having a size larger than the instruction block size.
[0056] Embodiment 2 is the processor of embodiment 1, wherein the group of second instructions starts at an address associated with the branch instruction in the branch target buffer.
[0057] Embodiment 3 is the processor of embodiment 2, wherein fetching the group of second instructions comprises fetching two or more instruction blocks from the address associated with the branch instruction in the branch target buffer.
[0058] Embodiment 4 is the processor of any one of embodiments 1-3, wherein the logic is configured to query' the branch target buffer for ordinary' instruction blocks and to bypass querying the branch target buffer for sequential instruction blocks.
[0059] Embodiment 5 is the processor of embodiment 4, wherein the operations further comprise executing the block of second instructions yvithout querying the branch target buffer.
[0060] Embodiment 6 is the processor of any one of embodiments 1-5, wherein a sequential instruction block is an instruction block that does not include any branch instructions.
[0061] Embodiment 7 is the processor of any one of embodiments 1-6, wherein the operations further comprise: determining that a block of instructions whose address is associated with a branch instruction in an entry in the branch target buffer was a sequential instruction block; and in response, modifying the entry in the branch target buffer to indicate that the block of instructions is a sequential instruction block.
[0062] Embodiment 8 is the processor of any one of embodiments 1-7, wherein the branch target buffer has a plurality’ of levels, and wherein the logic is configured to query the plurality of levels in a sequence over a plurality of cycles.
[0063] Embodiment 9 is the processor of any one of embodiments 1-8, wherein the logic is configured to query' a first levels of the plurality of levels on a first cycle and to fetch the group of second instructions on a subsequent second cycle.
[0064] Embodiment 10 is the processor of embodiment 9, yvherein whenever the logic determines on a first cycle that a first level of the plurality of levels does not have an entry for the branch instruction, the instruction fetch logic is configured to query a second level of the plurality of levels on a second cycle occurring after the first cycle and to fetch the group of second instructions on a third cycle occurring after the second cycle. Embodiment 11 is a method performed by a processor comprising a branch target buffer configured to store target addresses of previously executed branch instructions and logic configured to fetch blocks of instructions, the method comprising: fetching, by the processor, a first block of first instructions according to an instruction block size; determining that the first block is a sequential instruction block; in response, designating the first block in the branch target buffer as being a sequential instruction block; receiving a request to execute a second block of instructions; determining, from querying the branch target buffer, that the second block of instructions is a sequential instruction block; and in response, fetching a group of second instructions having a size larger than the instruction block size.
[0065] Embodiment 12 is the method of embodiment 11, wherein the group of second instructions starts at an address associated with the branch instruction in the branch target buffer.
[0066] Embodiment 13 is the method of embodiment 12, wherein fetching the group of second instructions comprises fetching two or more instruction blocks from the address associated with the branch instruction in the branch target buffer.
[0067] Embodiment 14 is the method of any one of embodiments 11-13, wherein the logic is configured to query the branch target buffer for ordinary instruction blocks and to bypass query ing the branch target buffer for sequential instruction blocks.
[0068] Embodiment 15 is the method of embodiment 14, wherein the operations further comprise executing the block of second instructions without querying the branch target buffer.
[0069] Embodiment 16 is the method of any one of embodiments 11-15, wherein a sequential instruction block is an instruction block that does not include any branch instructions.
[0070] Embodiment 17 is the method of any one of embodiments 11-16, wherein the operations further comprise: determining that a block of instructions whose address is associated with a branch instruction in an entry in the branch target buffer was a sequential instruction block; and in response, modifying the entry in the branch target buffer to indicate that the block of instructions is a sequential instruction block. Embodiment 18 is the method of any one of embodiments 11-17, wherein the branch target buffer has a plurality of levels, and wherein the logic is configured to query the plurality of levels in a sequence over a plurality of cycles.
[0071] Embodiment 19 is the method of any one of embodiments 11-18, wherein the logic is configured to query' a first levels of the plurality of levels on a first cycle and to fetch the group of second instructions on a subsequent second cycle.
[0072] Embodiment 20 is the method of embodiment 19. wherein whenever the logic determines on a first cycle that a first level of the plurality of levels does not have an entry for the branch instruction, the instruction fetch logic is configured to query a second level of the plurality of levels on a second cycle occurring after the first cycle and to fetch the group of second instructions on a third cycle occurring after the second cycle.
[0073] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0074] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0075] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.
[0076] What is claimed is:
Claims
CLAIMS1. A processor comprising: a branch target buffer configured to store target addresses of previously executed branch instructions; logic configured to fetch blocks of instructions, wherein the logic is configured to perform operations comprising: fetching a first block of first instructions according to an instruction block size; determining that the first block is a sequential instruction block; in response, designating the first block in the branch target buffer as being a sequential instruction block; receiving a request to execute a second block of instructions; determining, from querying the branch target buffer, that the second block of instructions is a sequential instruction block; and in response, fetching a group of second instructions having a size larger than the instruction block size.
2. The processor of claim 1, wherein the group of second instructions starts at an address associated with the branch instruction in the branch target buffer.
3. The processor of claim 2, wherein fetching the group of second instructions comprises fetching two or more instruction blocks from the address associated with the branch instruction in the branch target buffer.
4. The processor of any one of claims 1-3, wherein the logic is configured to query the branch target buffer for ordinary' instruction blocks and to bypass query ing the branch target buffer for sequential instruction blocks.
5. The processor of claim 4, wherein the operations further comprise executing the block of second instructions without query ing the branch target buffer.
6. The processor of any one of claims 1-5, wherein a sequential instruction block is an instruction block that does not include any branch instructions.
7. The processor of any one of claims 1-6, wherein the operations further comprise: determining that a block of instructions whose address is associated with a branch instruction in an entry in the branch target buffer was a sequential instruction block; and in response, modifying the entry in the branch target buffer to indicate that the block of instructions is a sequential instruction block.
8. The processor of any one of claims 1-7, wherein the branch target buffer has a plurality of levels, and wherein the logic is configured to query the plurality of levels in a sequence over a plurality of cycles.
9. The processor of any one of claims 1-8, wherein the logic is configured to query7a first levels of the plurality of levels on a first cycle and to fetch the group of second instructions on a subsequent second cycle.
10. The processor of claim 9, wherein whenever the logic determines on a first cycle that a first level of the plurality of levels does not have an entry for the branch instruction, the instruction fetch logic is configured to query7a second level of the plurality of levels on a second cycle occurring after the first cycle and to fetch the group of second instructions on a third cycle occurring after the second cycle.
11. A method performed by a processor comprising a branch target buffer configured to store target addresses of previously executed branch instructions and logic configured to fetch blocks of instructions, the method comprising: fetching, by the processor, a first block of first instructions according to an instruction block size; determining that the first block is a sequential instruction block; in response, designating the first block in the branch target buffer as being a sequential instruction block; receiving a request to execute a second block of instructions; determining, from query ing the branch target buffer, that the second block of instructions is a sequential instruction block; andin response, fetching a group of second instructions having a size larger than the instruction block size.
12. The method of claim 11, wherein the group of second instructions starts at an address associated with the branch instruction in the branch target buffer.
13. The method of claim 12. wherein fetching the group of second instructions comprises fetching two or more instruction blocks from the address associated with the branch instruction in the branch target buffer.
14. The method of any one of claims 11-13, wherein the logic is configured to query the branch target buffer for ordinary instruction blocks and to bypass querying the branch target buffer for sequential instruction blocks.
15. The method of claim 14, wherein the operations further comprise executing the block of second instructions without querying the branch target buffer.
16. The method of any one of claims 11-15, wherein a sequential instruction block is an instruction block that does not include any branch instructions.
17. The method of any one of claims 1 1 -1 , wherein the operations further comprise: determining that a block of instructions whose address is associated with a branch instruction in an entry in the branch target buffer was a sequential instruction block; and in response, modifying the entry in the branch target buffer to indicate that the block of instructions is a sequential instruction block.
18. The method of any one of claims 11-17, wherein the branch target buffer has a plurality of levels, and wherein the logic is configured to query the plurality’ of levels in a sequence over a plurality of cycles.
19. The method of any one of claims 11-18, wherein the logic is configured to query a first levels of the plurality of levels on a first cycle and to fetch the group of second instructions on a subsequent second cycle.
20. The method of claim 19, wherein whenever the logic determines on a first cycle that a first level of the plurality of levels does not have an entry for the branch instruction, the instruction fetch logic is configured to query a second level of the plurality of levels on a second cycle occurring after the first cycle and to fetch the group of second instructions on a third cycle occurring after the second cycle.
Citation Information
Patent Citations
Controlling access to a branch prediction unit for a sequence of fetch groups
JP7397858B2
Branch prediction suppression for blocks of instructions predicted to not include a branch instruction
US10289417B2
Cited By
Instruction prediction method and device, electronic equipment, medium and product
CN121523740A