Control for accessing the branch prediction unit with a sequence for extracting a group

By introducing sequential extraction logical function blocks into the processor, using the records associated with CTI to determine the number of extraction groups and preventing the access of branch prediction function blocks, the dynamic energy waste problem caused by frequent access to branch prediction function blocks in the prior art is solved, and lower power consumption and higher user satisfaction are achieved.

CN112673346BActive Publication Date: 2025-05-30ADVANCED MICRO DEVICES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980058993.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-10
Filing Date
2019-06-18
Publication Date
2025-05-30
Estimated Expiration
2039-06-18

AI Technical Summary

Technical Problem

When existing processors prepare for instruction execution, they frequently access branch prediction function blocks, resulting in waste of dynamic energy, especially multiple accesses to non-CTI instructions.

Method used

By introducing sequential extraction logical function blocks in the processor, the records associated with the CTI are used to determine the number of extract groups that do not include the CTI sequentially extracted after the CTI, and access to the branch prediction function blocks is prevented when this number is reached.

Benefits of technology

It effectively avoids unnecessary access to branch prediction function blocks, reduces the power consumption of the processor, and thus reduces the cost and delay of the electronic device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112673346B_ABST
    Figure CN112673346B_ABST
Patent Text Reader

Abstract

Describe an electronic device that disposes of control transfer instructions (CTIs) when executing instructions in program code. The electronic device has a processor, and the processor includes a branch prediction function block and a sequential fetch logic function block. The sequential fetch logic function block determines, based on a record associated with the CTI, that a specified number of instruction fetch groups previously determined not to include CTIs will be fetched to be sequentially executed after the CIT. When each of the specified number of fetch groups is fetched and ready for execution, the sequential fetch logic blocks a corresponding access to the branch prediction function block to obtain branch prediction information about the instructions in the fetch group.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Government Rights

[0002] This invention was made with government support under PathForward, a program planned by Lawrence Livermore National Laboratory and awarded by the U.S. Department of Energy (DOE) (prime contract No. DE-AC52-07NA27344, subcontract No. B620717). The government has certain rights in this invention. Background Art

[0003] Related Art

[0004] Many processors in electronic devices (e.g., microprocessors) include functional blocks that perform operations to improve the efficiency of executing instructions in program code. For example, some processors include a prediction functional block that is used to predict the path or flow of instruction execution (i.e., the sequence of addresses in memory from which instructions will be fetched for execution) based on a record of one or more previous instances of the execution of an instruction. A common prediction functional block is a branch prediction functional block that predicts the resolution of control transfer instructions (CTIs) such as jumps and returns in program code. The branch prediction functional block monitors and records the behavior of a CTI when the CTI is executed, such as the "taken" or "not taken" resolution of the CTI, the target instruction of the taken CTI, etc. After a CTI is encountered again during the execution of program code (e.g., a CTI is detected in the fetched instructions, etc.), the previously recorded behavior of the CTI is used to predict the current execution resolution of the CTI. Based on the predicted resolution, the processor speculatively fetches instructions and prepares the instructions to be executed along the predicted path after the CTI, while preparing and executing the CTI itself. Compared with a processor that waits to determine the resolution of a CTI before processing or speculatively follows a fixed choice of path from the CTI, such a processor can speculatively follow the path from the CTI that is more likely to be the path that will be followed when the CTI is executed, resulting in lower latency and / or fewer recovery operations.

[0005] In some processors, for all instructions at the beginning of the process of preparing the fetched instructions for execution, the branch prediction functional block will be accessed automatically to ensure that the predicted resolution of any CTI can be used to guide the path of program code execution as quickly as possible. However, because CTI instructions typically form only a small portion of program code, multiple accesses to the branch prediction functional block are for instructions that are not CTIs, and thus waste dynamic energy. Assuming that each access to the branch prediction functional block has an associated cost with respect to the power consumed, etc., unnecessary accesses to the branch prediction functional block need to be avoided. Brief Description of the Drawings

[0006] Figure 1A block diagram of an illustrated electronic device is presented in accordance with some embodiments.

[0007] Figure 2 A block diagram of an illustrated processor is presented in accordance with some embodiments.

[0008] Figure 3 A block diagram of an illustrated branch prediction unit is presented in accordance with some embodiments.

[0009] Figure 4 A block diagram of an illustrated branch target buffer is presented in accordance with some embodiments.

[0010] Figure 5 A block diagram of an illustrated sequential fetch logic is presented in accordance with some embodiments.

[0011] Figure 6 A block diagram of an illustrated sequential fetch table is presented in accordance with some embodiments.

[0012] Figure 7 A flowchart of a process for using a record associated with a control transfer instruction to block access to a branch prediction unit is presented in accordance with some embodiments.

[0013] Figure 8 A flowchart of an operation performed to maintain a record when using a branch target buffer to store a record associated with a control transfer instruction is presented in accordance with some embodiments.

[0014] Figure 9 A timeline diagram of an operation for updating a branch target buffer with a count of fetch groups that do not include a control transfer instruction sequentially fetched after a corresponding control transfer instruction is presented in accordance with some embodiments.

[0015] Figure 10 A flowchart of an operation performed to maintain a record when using a sequential fetch table to store a record associated with a control transfer instruction is presented in accordance with some embodiments.

[0016] Figure 11 A timeline diagram of an operation for updating a sequential fetch table with a count of fetch groups that do not include a CTI sequentially fetched after a corresponding CTI is presented in accordance with some embodiments.

[0017] Figure 12 A flowchart of a process for determining a specified number of fetch groups that do not include a CTI to be sequentially fetched after a CTI and using the specified number to block access to a branch prediction unit is presented in accordance with some embodiments.

[0018] Throughout the figures and the description, like reference numerals refer to like elements. Detailed Description

[0019] The following description is presented to enable any person skilled in the art to make and use the described embodiments, and is provided in the context of a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications. Thus, the described embodiments are not limited to the embodiments shown, but are to be accorded the widest scope consistent with the principles and features disclosed herein.

[0020] Term

[0021] In the following description, various terms are used to describe embodiments. The following is a simplified and general description of some of these terms. Note that these terms may have important additional aspects that are not enumerated herein for the sake of clarity and brevity, and thus the description is not intended to limit these terms.

[0022] Functional block: A functional block refers to a group, collection, and / or set of one or more interrelated circuit elements (e.g., integrated circuit elements, discrete circuit elements, etc.). Circuit elements are "interrelated" because the circuit elements share at least one property. For example, interrelated circuit elements may be included in a particular integrated circuit chip or a portion thereof, fabricated on a particular integrated circuit chip or a portion thereof, or otherwise coupled to a particular integrated circuit chip or a portion thereof, may be involved in the execution of a given function (computing or processing function, memory function, etc.), may be controlled by a common control element, etc. A functional block may include any number of circuit elements, from a single circuit element (e.g., a single integrated circuit logic gate) to millions or billions of circuit elements (e.g., an integrated circuit memory).

[0023] Control Transfer Instruction: A control transfer instruction (CTI) is an instruction in program code that, when executed, causes / enables a jump, displacement, or discontinuity in another sequential flow of instruction execution. CTIs include "unconditional" CTIs such as jumps, calls, returns, etc., which automatically cause the instruction execution to jump from the instruction (CTI) at the first memory address to the instruction at the second memory address or the "target instruction". CTIs also include "conditional" CTIs such as conditional jump instructions, which include conditions such as greater than, equal to, non-zero, etc., associated with or depending on the condition. When the corresponding condition (e.g., true, false, etc.) is satisfied, the conditional CTI causes a jump in the instruction execution from the CTI to the instruction at the second memory address, but when the condition is not satisfied, the instruction execution continues sequentially after the CTI. For example, a conditional branch instruction can be implemented using a conditional check instruction and a conditional CTI (or a single combined instruction), where when the condition is satisfied, the branch is "taken" and the instruction execution jumps to the target instruction, and when the condition is not satisfied, the branch instruction is "not taken" or "falls through" and the instruction execution continues sequentially. CTIs include "indirect" unconditional CTIs and conditional CTIs for which the address of the target instruction is specified dynamically at runtime. For example, the address of the target instruction of a CTI can be calculated and stored in a processor register or other location by a prior instruction, and then the address is used to determine the address to which the instruction execution is to jump after the CTI is executed.

[0024] Overview

[0025] The described embodiments include a processor in an electronic device. The processor includes functional blocks such as, for example, central processing unit (CPU) cores, one or more cache memories, communication interfaces, etc., which perform computing operations, memory operations, communication with other functional blocks and devices, etc. The processor also includes sequential fetch logic functional blocks that perform operations to avoid (when possible) accessing a branch prediction functional block in the processor to obtain branch prediction information. In the described embodiments, a fetch group of instruction blocks of a specified size (e.g., 32 bytes, 64 bytes, etc.) is fetched as a group from a cache memory or main memory and prepared for execution in the instruction execution pipeline of the processor. When a given fetch group is fetched and ready for execution, the sequential fetch logic checks a record associated with the given fetch group (if such a record exists) to determine the number of previously determined fetch groups that do not include control transfer instructions ("CTIs") and that will be sequentially fetched after the given fetch group. For example, in some embodiments, the sequential fetch logic checks a record associated with a CTI in a given fetch group to determine the number of fetch groups. As another example, in some embodiments, the sequential fetch logic checks a record associated with the target instruction of a CTI in a given fetch group to determine the number of fetch groups. Since there are no CTIs in the number of fetch groups, the fetch groups will be sequentially fetched - and branch prediction information will not be needed. Since each of the number of fetch groups is sequentially fetched and ready for execution, the sequential fetch logic prevents access to the branch prediction functional block to obtain branch prediction information. For example, the sequential fetch logic may prevent checks in a branch target buffer (BTB), a branch direction predictor, etc. in the branch prediction functional block for obtaining branch prediction information.

[0026] In some embodiments, the BTB is used to store some or all of the above records indicating the number of fetch groups without CTIs that will be sequentially fetched after respective CTIs. In these embodiments, in addition to the solution information of the CTI, the BTB also includes information about the number of fetch groups without CTIs (if any) that will be sequentially fetched after the CTI. During operation, after encountering a specified type of CTI in a given fetch group (e.g., detecting a specified type of CTI within a given fetch group), the sequential fetch logic functional block maintains a count of the number of fetch groups to be fetched sequentially after the given fetch group before the subsequent CTI retires (i.e., completes execution and is ready for delivery to the architectural state of the processor). The sequential fetch logic then stores the count of the number of fetch groups in an entry associated with the CTI in the BTB. The count in the entry in the BTB is then used as described herein to avoid accessing the branch prediction functional block for the number of fetch groups to obtain branch prediction information.

[0027] In some embodiments, a Sequential Fetch Table (SFT) is used to store some or all of the above records indicating the number of fetch groups without a CTI that will be sequentially fetched after respective CTIs. In these embodiments, an entry in the SFT is associated with the target instruction of the CTI, such that each record in the SFT includes an indication of the number of fetch groups without a CTI that will be sequentially fetched after a particular path / solution of the corresponding CTI. During operation, after encountering the retirement of the target instruction of a CTI in a given fetch group (e.g., whether the instruction targeted by the CTI is on the taken or not-taken path from the CTI), the sequential fetch logic function block maintains a count of the number of fetch groups to be fetched before the subsequent CTI retires for sequential execution. The sequential fetch logic then stores the count of the number of fetch groups in the entry in the SFT associated with the target instruction. The count in the entry in the SFT is then used as described herein to avoid accessing the branch prediction function block for the number of fetch groups to obtain branch prediction information.

[0028] In some embodiments, when using the above records to prevent access to the branch prediction function block to obtain branch prediction information, the sequential fetch logic obtains the number of fetch groups from the corresponding record (from the BTB, SFT, or elsewhere) and sets a counter equal to the number of fetch groups. As each subsequent fetch group is fetched and ready for execution, the sequential fetch logic decrements the counter and prevents access to the branch prediction function block as described above. After the counter reaches zero, since one or more subsequent fetch groups are fetched and ready for execution, the sequential fetch logic permits the corresponding access to the branch prediction function block to obtain branch prediction information about the instructions in the one or more subsequent fetch groups. In other words, the sequential fetch logic uses the counter to prevent access to the branch prediction function block for a specified number of fetch groups and then begins to perform access to the branch prediction function block for subsequent fetch groups. In some of these embodiments, the sequential fetch logic also pauses the inspection of the record itself as long as the counter is greater than zero – and the portion of the function block storing the record (e.g., BTB, SFT, etc.) can be powered down, placed in a reduced power mode (e.g., supplied with a reduced supply voltage, etc.), or otherwise set to reduce power consumption (e.g., clock gating, etc.).

[0029] By obtaining branch prediction information of an extraction group by using a record indicating the number of extraction groups excluding a CTI that will be sequentially extracted after a given CTI to block access to a branch prediction functional block, the described embodiments can avoid unnecessary access to the branch prediction functional block. This can help reduce power consumption in a processor and more generally in an electronic device. The reduced power consumption can result in a lower usage cost of the electronic device, more efficient use of battery power, etc., which can lead to higher user satisfaction with the electronic device.

[0030] Electronic device

[0031] Figure 1 According to some embodiments, a block diagram illustrating an electronic device 100 is presented. As Figure 1 can be seen, the electronic device 100 includes a processor 102 and a memory 104. Generally, the processor 102 and the memory 104 are implemented in hardware (i.e., using various circuit elements and devices). For example, the processor 102 and the memory 104 can all be fabricated on one or more semiconductor chips (including fabricated on one or more separate semiconductor chips), can be made of semiconductor chips with combined discrete circuit elements, can be individually fabricated by discrete circuit elements, etc. As described herein, the processor 102 and the memory 104 perform operations for blocking access to a branch prediction functional block for certain extraction groups.

[0032] The processor 102 is a functional block in the electronic device 100 that performs computational operations and other operations (e.g., control operations, configuration operations, etc.). For example, the processor 102 can be or include one or more microprocessors, central processing unit (CPU) cores, and / or another processing mechanism.

[0033] The memory 104 is a functional block in the electronic device 100 that performs operations on the memory of the electronic device 100 (e.g., “main” memory). The memory 104 includes: a volatile memory circuit, such as a fourth-generation double data rate synchronous DRAM (DDR4 SDRAM) and / or other types of memory circuits, which is used to store data and instructions to be used by functional blocks in the electronic device 100; and a control circuit, which is used to handle access to the data and instructions stored in the memory circuit and perform other control or configuration operations.

[0034] For illustrative purposes, the electronic device 100 is simplified. However, in some embodiments, the electronic device 100 includes additional or different functional blocks, subsystems, elements, and / or communication paths. For example, the electronic device 100 can include a display subsystem, a power subsystem, an input-output (I / O) subsystem, etc. The electronic device 100 generally includes sufficient functional blocks, etc. to perform the operations described herein.

[0035] The electronic device 100 can be or can be included in any device that performs computing operations. For example, the electronic device 100 can be or can be included in the following: a desktop computer, a laptop computer, a wearable computing device, a tablet computer, a piece of virtual or augmented reality equipment, a smart phone, an artificial intelligence (AI) or machine learning device, a server, a network device, a toy, a piece of audio-visual equipment, a household appliance, a vehicle, etc. and / or a combination thereof.

[0036] Processor

[0037] As described above, the electronic device 100 includes a processor 102, which can be a microprocessor, a CPU core, and / or another processing mechanism. Figure 2 According to some embodiments, a block diagram illustrating the processor 102 is presented. Although certain functional blocks are shown in Figure 2 , in some embodiments, different arrangements, connectivities, numbers, and / or types of functional blocks may be present in the processor 102. Generally, the processor 102 includes sufficient functional blocks to perform the operations described herein.

[0038] As Figure 2 can be seen, the functional blocks in the processor 102 can be regarded as part of a front-end subsystem 200, an execution subsystem 202, or a memory subsystem 204. The front-end subsystem 200 includes functional blocks that perform operations for fetching instructions from a cache memory or a main memory in or communicating with the memory subsystem 204, and prepares the instructions for dispatching to the execution unit functional blocks in the execution subsystem 202.

[0039] The front-end subsystem 200 includes an instruction decoder 206, which is a functional block that performs operations related to decoding and preparing the fetched instructions for execution. The instruction decoder 206 fetches or otherwise receives instructions in an N-byte fetch group (e.g., four instructions in a 32-byte fetch group, etc.) from the L1 instruction cache 216, the L2 cache 218, an L3 cache (not shown), or a main memory (not shown). The instruction decoder 206 then possibly decodes the instructions in the fetch group into individual micro-operations in parallel. The instruction decoder 206 next dispatches the micro-operations to an instruction dispatcher 208 for forwarding to the appropriate execution units in the execution subsystem 202 for execution.

[0040] The front-end subsystem 200 also includes a next PC 210, which is a functional block that performs operations for determining the program counter or the address in memory from which the next fetch group will be fetched. The next PC 210 calculates the next sequential value of the program counter based on the initial or current value of the program counter. For example, given a 32-byte fetch group, the next PC 210 may calculate the next address = current address + 32 bytes. When CTI is adopted without changing the program flow, the front-end subsystem 200 uses the sequential value of the program counter calculated by the next PC 210 to fetch the fetch group from the corresponding sequential address in memory.

[0041] The front-end subsystem 200 also includes a branch prediction unit 212, which is a functional block that performs operations for predicting the solution for CTI in the fetch group and modifying the program counter and thus the address in memory from which the subsequent fetch group will be fetched. In other words, the branch prediction unit 212 uses one or more records of CTI behavior to predict the "taken" or "not taken" solution for CTI and provides the predicted target address for taking CTI. When the branch prediction unit 212 predicts that CTI will be taken, the target address returned by the branch prediction unit 212 may be used to replace the next or subsequent program counter provided by the next PC 210.

[0042] Figure 3 According to some embodiments, a block diagram illustrating the branch prediction unit 212 is presented. Although the branch prediction unit 212 is shown with various functional blocks in Figure 3 For the purposes of this specification, the branch prediction unit 212 is simplified; in some embodiments, different functional blocks are present in the branch prediction unit 212. For example, in some embodiments, multi-level branch prediction, branch pattern predictors, multi-level branch target buffers, and / or direction predictors and / or other branch prediction mechanisms or techniques are used, and the corresponding functional blocks are included in the branch prediction unit 212. Generally, the branch prediction unit 212 includes sufficient functional blocks to perform the operations described herein.

[0043] As in Figure 3It can be seen that the functional blocks in the branch prediction unit 212 include a controller 300, a direction predictor 302, and a branch target buffer (BTB) 304. The controller 300 includes circuit elements for performing operations of the branch prediction unit 212, such as updating the direction predictor 302 and the branch target buffer 304 and looking up in the direction predictor 302 and the branch target buffer 304, communicating with other functional blocks, etc. The direction predictor 302 includes a record having a number of entries, such as a lookup table, a list, etc., each entry for storing an address associated with a CTI and an indication of a taken or not taken solution for the CTI. For example, for a CTI at address A, the direction predictor 302 may include an entry associating address A or a value based on the address with a corresponding prediction (e.g., a saturating counter, etc.) of a taken or not taken solution for the CTI. The branch target buffer 304 includes a record having a number of entries, such as a lookup table, a list, etc., each entry for storing an address associated with a CTI and an indication of the target address of the CTI. For example, for a CTI at address A, the branch target buffer 304 may include an entry associating address A or a value based on the address with the corresponding absolute or relative address of the target instruction of the CTI. When executing an instruction, the controller 300 may store and / or update the corresponding entry in the direction predictor 302 and / or the branch target buffer 304 based on the actual result of the CTI instruction, thereby storing the value used in the above prediction of the CTI instruction solution.

[0044] In some embodiments, in addition to branch target information, the branch target buffer 304 is also used to store a record of a count of the number of extraction groups that are extracted prior to a subsequent CTI retirement for sequential execution after the CTI. In these embodiments, each entry includes the location where a count (e.g., a number, a string, etc.) associated with the corresponding CTI is stored, if such a count is available. For example, the count may be stored in eight bits reserved for this purpose in the entry. Figure 4 A block diagram illustrating the branch target buffer 304 is presented in accordance with some embodiments. Although the branch target buffer 304 is shown as storing certain information, for the purposes of this specification, the branch target buffer 304 is simplified; in some embodiments, different arrangements of information are stored in the entries of the branch target buffer 304. Generally, the branch target buffer 304 stores sufficient information to perform the operations described herein.

[0045] As in Figure 4It can be seen that the branch target buffer 304 includes a number of entries 408, each entry including an address (ADDR) 400, a branch target 402, a count 404, and metadata 406. The address 400 is used to store the CTI or an address for or otherwise associated with the CTI or a value based on the address, and the entry stores information regarding the CTI. The branch target 402 is used to store the address of the target instruction for the CTI or a value based on the address, and the entry stores information regarding the CTI - and thus the CTI can be used to predict the address of the target instruction when the CTI is executed again. The count 404 is used to store a count of the number of fetch groups to be fetched and executed sequentially after the CTI before the subsequent CTI retires, and the entry stores information for the CTI. The metadata 406 is used to store information regarding the entry, the count, and / or the CTI or information associated with the entry, the count, and / or the CTI, such as a valid bit, an enable bit, etc.

[0046] Return Figure 2 , the front - end subsystem 200 further includes sequential fetch logic 214, which is a functional block that performs operations to avoid accessing the branch prediction functional block to obtain branch prediction information when possible. The sequential fetch logic 214 uses records associated with the fetch groups to determine the number of previously determined fetch groups that do not include a CTI and that will be fetched sequentially after a given fetch group (or an instruction in the fetch group). Since each of the number of fetch groups is fetched sequentially, the sequential fetch logic 214 prevents access to the branch prediction unit 212 to obtain branch prediction information.

[0047] Figure 5 According to some embodiments, a block diagram illustrating the sequential fetch logic 214 is presented. Although the sequential fetch logic 214 is shown with various functional blocks in Figure 5 , for the purposes of this specification, the sequential fetch logic 214 is simplified; in some embodiments, different functional blocks are present in the sequential fetch logic 214. For example, although the sequential fetch logic 214 is shown as including a sequential fetch table 502, in some embodiments, the sequential fetch logic 214 does not include or does not use a sequential fetch table. Instead, the sequential fetch logic 214 uses the count information in the entries stored in the branch target buffer 304 to perform the corresponding operations. Generally, the sequential fetch logic 214 includes sufficient functional blocks to perform the operations described herein.

[0048] As in Figure 5It can be seen that the functional blocks in the sequential fetch logic 214 include a controller 500 and a sequential fetch table 502. The controller 500 includes circuit elements for performing the operations of the sequential fetch logic 214, such as updating the sequential fetch table 502 (or the branch target buffer 304) and looking up in the sequential fetch table 502 (or the branch target buffer 304), communicating with other functional blocks, etc. The sequential fetch table 502 includes records having a number of entries, such as a lookup table, a list, etc., and each entry is a record for storing the address associated with the target instruction of the CTI and the count of the number of fetch groups to be fetched for sequential execution before the subsequent CTI retires. Figure 6 According to some embodiments, a block diagram illustrating the sequential fetch table 502 is presented. Although the sequential fetch table 502 is shown as storing certain information, for the purposes of this specification, the sequential fetch table 502 is simplified; in some embodiments, different arrangements of information are stored in the entries of the sequential fetch table 502. Generally, the sequential fetch table 502 stores sufficient information to perform the operations described herein.

[0049] As shown in Figure 6 It can be seen that the sequential fetch table 502 includes a number of entries 606, and each entry includes a CTI target address (ADDR) 600, a count 602, and metadata 604. The address 600 is used to store the address of the CTI or an address for or otherwise associated with the CTI or a value based on the address, and the entry stores information about the CTI. Generally, the "target" instruction of the CTI is an instruction and thus an address in the memory, and when the CTI is executed, the program flow jumps to that address. For an unconditional CTI with a statically defined target, such as a static target jump instruction, there is only one target instruction - and thus the program flow always jumps from the CTI instruction to the same address in the memory. However, for a conditional CTI, there are at least two target instructions (on the taken path and the not-taken path in the program code), and there can be any number of target instructions. For example, an indirect CTI (whose target instruction is specified at runtime) can have any number of target instructions on the taken path, although the not-taken path is sequential. Since the sequential fetch table 502 stores records associated with the target instructions of the CTI, each possible target instruction for a given CTI (and thus the path starting from the given CTI) can have an associated separate entry in the sequential fetch table 502. The count 602 is used to store the count of the number of fetch groups to be fetched for sequential execution after the corresponding CTI before the subsequent CTI retires. The metadata 604 is used to store information about the entry and / or the count or information associated with the entry and / or the count, such as a valid bit, an enable bit, etc.

[0050] In an embodiment where the sequential fetch table 502 is used to store the above-mentioned records for determining the number of fetch groups, after receiving the program counter (i.e., the address where a given fetch group is to be fetched), the controller 500 performs a lookup in the sequential fetch table 502 to determine whether there is an entry in the sequential fetch table 502 with an associated address. In other words, the lookup determines whether an address (i.e., the target instruction of the CTI) within the address range of the instructions present in a given fetch group will be found in the sequential fetch table 502. If so, the controller 500 obtains the corresponding count from the counter 602 and then uses the count as the number of fetch groups that block access to the branch prediction unit 212 to obtain branch prediction information. Otherwise, when no matching address is found in the sequential fetch table 502, the controller 500 does not block access to the branch prediction unit 212, that is, allows the acquisition of branch prediction information to proceed normally.

[0051] In an embodiment where the branch target buffer 304 is used to store the above-mentioned records for determining the number of fetch groups, as part of performing a lookup for branch prediction information regarding the CTI in a fetch group, the branch prediction unit 212 obtains the address of the predicted target instruction of the CTI from the branch target buffer 304 if such a target address exists. The branch prediction unit 212 also obtains the count 404 from the branch target buffer 304 and passes the count back to the controller 500 in the sequential fetch logic 214. The controller 500 uses the count as the number of fetch groups that block access to the branch prediction unit 212 to obtain branch prediction information. Otherwise, when no matching address and / or count is found in the branch target buffer 304, the controller 500 does not block access to the branch prediction unit 212, that is, allows the acquisition of branch prediction information to proceed normally.

[0052] In some embodiments, the sequential fetch table 502 includes only a limited number of entries (e.g., 32 entries, 64 entries, etc.) and may thus reach full capacity during the operation of the processor 102. When the sequential fetch table 502 is full, it will be necessary to orderly overwrite the existing information in the entries to store new information in the sequential fetch table 502. In some embodiments, the controller 500 manages the entries in the sequential fetch table 502 using one or more replacement policies, guidelines, etc. In these embodiments, when selecting an entry to overwrite, the entry is selected according to the replacement policy, guidelines, etc. For example, the controller 500 may use the least recently used replacement policy to manage the information in the entries of the sequential fetch table 502.

[0053] Note that although the sequential fetch logic 214 is shown as being associated with Figure 2a single functional block separate from other functional blocks in [the context], but in some embodiments, some or all of the sequential extraction logic 214 may be included in Figure 2 the other functional blocks shown. In these embodiments, the operations attributed to the sequential extraction logic 214 may be performed by circuit elements in the other functional blocks. Generally, the sequential extraction logic 214 includes various circuit elements for performing the described operations, and there is no limitation on the specific location of the circuit elements in the Figure 2 processor 102 shown.

[0054] Return Figure 2 , the execution subsystem 202 includes an integer execution unit 222 and a floating-point execution unit 224 (collectively referred to as "execution units"), which are functional blocks that perform operations for executing integer instructions and floating-point instructions, respectively. The execution units include elements such as renaming hardware, an execution scheduler, an arithmetic logic unit (ALU), floating-point multiplication and addition components (in the floating-point execution unit 224), a register file, and the like.

[0055] The execution subsystem also includes a retirement queue 226, which is a functional block in which the results of the executed instructions are held after the corresponding instructions have completed execution but before the results of the executed instructions are delivered to the architectural state of the processor 102 (e.g., written to a cache or memory and made available for use in other operations). In some embodiments, certain instructions may not be executed in program order, and the retirement queue 226 is used to ensure that the results of out-of-order instructions will be retired appropriately relative to other out-of-order instructions.

[0056] In some embodiments, the retirement queue 226 performs at least some of the operations for maintaining a count of the number of extraction groups that are extracted before subsequent CTI retirement for sequential execution after CTI. For example, in some embodiments, in addition to the CTI to be executed in the execution subsystem 202, the front-end subsystem 200 (e.g., instruction decoding 206) also includes an indication of whether an instruction is (or is not) a CTI. For example, the front-end subsystem 200 may set a specified flag bit in the metadata bits associated with the instructions in the execution subsystem 202. After encountering an instruction indicated as a CTI, the retirement queue 226 may begin to maintain a count of the extraction groups (or, more generally, individual instructions) that are retired before the subsequent instruction indicated as a CTI is retired. The count may then be communicated to the sequential extraction logic 214 for storage for future use as described herein. In some embodiments, the count is not reported by the retirement queue 226 (or not used to update any records) unless the count exceeds a corresponding threshold.

[0057] The memory subsystem 204 includes a cache hierarchy, where the cache is a functional block that includes a volatile memory circuit for storing a limited number of copies of instructions and / or data that are close to the functional block using the instructions and / or data, and a control circuit for handling operations such as accessing data. The hierarchy includes two levels, with the level-1 (L1) instruction cache 216 and the L1 data cache 220 belonging to the first level, and the L2 cache 218 belonging to the second level. The memory subsystem 204 is communicatively coupled to the memory 104 and may be coupled to an external L3 cache (not shown). The memory 104 may be coupled to a non-volatile mass storage device (e.g., a disk drive or a solid state drive) (not shown) that serves as a long-term memory for instructions and / or data.

[0058] Record using a sequential extraction group associated with a control transfer instruction to avoid accessing a branch predictor

[0059] In the described embodiments, a processor (e.g., processor 102) in an electronic device uses a record associated with the CTI to determine the number of extraction groups that do not include the CTI and that will be sequentially extracted after the CTI, and for that number of extraction groups, blocks access to the branch prediction unit to obtain corresponding branch prediction information. Figure 7 According to some embodiments, a flowchart is presented that illustrates a process for using a record associated with the CTI to block access to the branch prediction unit. Note that Figure 7 the operations shown are presented as general examples of operations performed by some embodiments. Operations performed by other embodiments include different operations and / or operations in a different order. For Figure 7 the example in, a processor in an electronic device having an internal arrangement similar to that of processor 102 is described as performing various operations. However, in some embodiments, a processor having a different internal arrangement performs the described operations.

[0060] Figure 7The operations shown begin when the processor maintains a record associated with each of one or more CTIs, each record indicating a specified number of extraction groups excluding the CTIs that are to be sequentially extracted after the corresponding CTI (step 700). During this operation, a controller in the sequential extraction logic (e.g., controller 500 in sequential extraction logic 214) receives an indication of the CTI or an identifier of the CTI from the retirement queue (e.g., retirement queue 226) and the number of extraction groups that are to be sequentially extracted after the CTI. For example, assuming an implementation where an extraction group includes four instructions, if the retirement queue encounters a CTI and then counts 129 instructions before the next CTI retires, in addition to the identification of the CTI, the retirement queue may communicate the value 129 or another value, such as 32 (32 is 129 / 4 rounded to the nearest integer to represent the number of extraction groups), to the sequential extraction logic. (As used herein, "encountering" a CTI includes detecting the CTI in a retired instruction based on a processor flag associated with the CTI, detecting a behavior or pattern associated with the CTI in the instruction stream, and / or subsequent instructions or result values, etc.) The controller then updates the record associated with the CTI to indicate the specified number of extraction groups.

[0061] Figure 8 and Figure 10 presents a flowchart illustrating the operations performed to maintain records associated with one or more CTIs as described in step 700 with respect to Figure 7 . Figure 8 According to some embodiments, a flowchart is presented illustrating the operations performed to maintain the records when a branch target buffer is used to store records associated with CTIs. Figure 10 According to some embodiments, a flowchart is presented illustrating the operations performed to maintain the records when a sequential extraction table is used to store records associated with CTIs. Note that Figure 8 and Figure 10 the operations shown are presented as general examples of operations performed by some embodiments. Operations performed by other embodiments include different operations, operations performed in a different order, and / or operations performed by different functional blocks.

[0062] Figure 8 The operations shown begin when the retirement queue begins maintaining a count of the number of extraction groups (or, more generally, individual instructions retired) that are sequentially retired before the retirement of the next CTI after the retirement of a specified type of CTI (step 800). For this operation, the retirement queue monitors the retirement of CTI instructions of the specified type and then maintains a count of the subsequent instructions retired before the next CTI instruction. For example, the retirement queue may increment a counter for each retired instruction after a CTI instruction of the specified type until the next CTI retires.

[0063] As described above, the branch target buffer is used to store Figure 8 the records in the illustrated embodiments. In some embodiments, the branch target buffer stores branch target records indexed by an identifier of the CTI (e.g., the address of the CTI instruction in memory) or the extraction group in which the CTI is included. For this purpose, in these embodiments, the count information in the entries in the branch target buffer will be associated only with the CTI or the extraction group, and not with the (possibly many) target instructions of the CTI. In other words, the count information described above is a single data piece that represents only a single decision path through the corresponding target instruction starting from the CTI. These records in the branch target buffer thus cannot be reliably used to determine the count of the sequential extraction groups for all possible solutions regarding static target instruction CTIs and / or CTIs with dynamically specified target instructions. Thus, in some embodiments, a "specified" type of CTI is an unconditional static target instruction CTI or an untaken path from a conditional CTI.

[0064] The retirement queue then conveys the count to the sequential extraction logic, which stores a record indicating the count of the number of extraction groups in an entry in the branch target buffer associated with the CTI or the extraction group (step 802). For this operation, the sequential extraction logic determines the entry in the branch target buffer into which the record will be stored, e.g., by determining the address of the CTI or a value based on the address and determining the corresponding entry in the branch target buffer. The sequential extraction logic then stores the record indicating the count of the number of extraction groups in the determined entry. For example, when the entry in the branch target buffer includes space for an eight-bit value, the sequential extraction logic stores an eight-bit value representing the count in the entry. In cases where the entry is not large enough to store the count (e.g., for an eight-bit count value, the count is greater than 255), a default or zero value can be stored in the entry, thereby indicating that access to the branch predictor will be performed (not blocked) for a particular CTI. Alternatively, a maximum value can be stored in the entry, thereby enabling avoidance of at least some of the accesses to the branch predictor.

[0065] Figure 9 A timeline diagram of an operation for updating the branch target buffer with the count of extraction groups that do not include the CTI sequentially extracted after the corresponding CTI is presented according to some embodiments. As Figure 9 shown, time advances from left to right, and during that time, many extraction groups (FGs), each including a separate set of instructions from the program code, are extracted, prepared for execution (e.g., decoded, dispatched, etc.), executed, and then retired. Each extraction group includes many individual instructions (e.g., four, six, etc.), and there are three CTI instructions among the instructions in the extraction group.

[0066] In Figure 9 , the first CTI instruction CTI1, which is the second instruction in fetch group 902, is a static unconditional branch instruction that causes the instruction stream to jump from fetch group 902 at a first address to fetch group 904 at a second non-sequential address, as illustrated by the arrow from the second instruction in fetch group 902 to the initial instruction in fetch group 904. After detecting the retirement of CTI1 (CTI1 is one of the designated types of CTI), the retirement queue maintains a count of the number of fetch groups that will retire sequentially before the subsequent CTI retires (i.e., for which the count constitutes instruction retirement) – the subsequent CTI is an un-taken conditional CTI2 with a static target address and is the third instruction in fetch group 910. Since three fetch groups (fetch groups 904 to 908) retire before CTI2 retires, the retirement queue conveys the identification of CTI1 (e.g., its address, processor instruction tag, or internal identifier, etc.) and a count value of three or another value (e.g., the number of retired instructions, 14) to the sequential fetch logic. After receiving the identifier of CTI1 and the value, the sequential fetch logic updates the corresponding entry in the branch target buffer, as shown in the entry for CTI1 in the example of the branch target buffer in Figure 9 . Note that, as shown in the branch target buffer in Figure 9 , CTI1 and CTI2 can be the addresses of CTI1 and CTI2, or other values that represent or identify CTI1 and CTI2, such as the addresses associated with the corresponding fetch groups.

[0067] The second CTI instruction CTI2, which is the end of the count for CTI1, also causes the retirement queue to start a second / corresponding count. As described above, CTI2 is un-taken, so the fetch groups will be fetched sequentially from the next address (as by the next PC function block), and the retirement queue maintains the corresponding count until the subsequent CTI CTI3 in fetch group 916 (which directs the program flow to fetch group 918) retires. Since two complete fetch groups (fetch groups 912 to 914) retire before CTI3 retires, the retirement queue conveys the identification of CTI2 and a count value of two or another value (e.g., the number of retired instructions, 10) to the sequential fetch logic. After receiving the identifier of CTI2 and the value, the sequential fetch logic updates the corresponding entry in the branch target buffer, shown as CTI2 in the example of the branch target buffer in Figure 9 . Note that CTI2 is an un-taken path case of a conditional CTI and is thus one of the designated types of CTI. In some embodiments, similar tracking and recording are not performed for the taken path.

[0068] Proceed to Figure 10 , Figure 10The operation shown begins when the retirement queue starts keeping count (step 1000) of the number of fetch groups (or, more generally, individual retired instructions) that retire sequentially before the next CTI retirement after the retirement of the target instruction of the CTI. For this operation, the retirement queue monitors the retirement of the CTI and the target instruction of the CTI (i.e., the instruction that is sequentially next after the CTI), and keeps count of the subsequent instructions that retire before the next CTI instruction. For example, the retirement queue can increment a counter for each retired instruction after a specified type of target instruction until the next CTI retirement.

[0069] As described above, a sequential fetch table is used to store Figure 10 the records in the illustrated embodiments. In some embodiments, the sequential fetch table stores records associated with the target instruction of the CTI. In other words, in these embodiments, the sequential fetch table can include records associated with each possible target instruction of the CTI, and thus includes each path taken from the CTI. For this reason, in these embodiments, various types of CTIs can be handled (i.e., can have a sequential fetch count of records). Remember, this is different from embodiments where the count information is stored in a branch target buffer, and thus only a specified type of CTI can be handled. The branch target buffer stores count information associated with the CTI or the fetch group, but not the corresponding target instruction.

[0070] The retirement queue then conveys the count to the sequential fetch logic, which stores a record indicating the count of the number of fetch groups in the entry associated with the target instruction in the sequential fetch table (step 1002). For this operation, the sequential fetch logic determines the entry in the sequential fetch table into which the record will be stored, for example by determining the address of the target instruction or a value based on the address and determining the corresponding entry in the sequential fetch table. The sequential fetch logic then stores the record indicating the count of the number of fetch groups in the determined entry. For example, when the entry in the sequential fetch table includes space for an eight-bit value, the sequential fetch logic stores the eight-bit value representing the count in the entry. In cases where the entry is not large enough to store the count (e.g., for an eight-bit count value, the count is greater than 255), a default or zero value can be stored in the entry, thereby indicating that for a particular target instruction, an access to the branch predictor will be made (cannot be avoided). Alternatively, the maximum value can be stored in the entry, thereby enabling avoidance of at least some of the accesses to the branch predictor.

[0071] Figure 11 A timeline diagram of an operation to update a sequential fetch table with the count of fetch groups that do not include a CTI and that are fetched sequentially after a corresponding CTI is presented according to some embodiments. As Figure 11As shown, time advances from left to right, and during said time, many fetch groups (FGs), each including a separate set of instructions from program code, are fetched, prepared for execution (e.g., decoded, dispatched, etc.), executed, and then retired. Each fetch group includes many individual instructions (e.g., four, six, etc.), and there are three CTI instructions among the instructions in the fetch group.

[0072] In Figure 11 , the first CTI instruction CTI1, which is the second instruction in fetch group 1102, is a static unconditional branch instruction that causes the instruction stream to jump from fetch group 1102 at a first address to fetch group 1104 at a second non-sequential address (ADDR1). This process is illustrated by an arrow from the second instruction in fetch group 1102 to the initial instruction (at ADDR1) of fetch group 1104. After detecting the retirement of the target instruction (the initial instruction in fetch group 1104) at ADDR1, the retirement queue maintains a count of the number of fetch groups that will retire sequentially before the subsequent CTI retires – the subsequent CTI is an un-taken conditional CTI2 with a static target address and is the first instruction in fetch group 1112. Since three complete fetch groups (fetch groups 1104 to 1108) retire before CTI2 retires, the retirement queue conveys to the sequential fetch logic the identification of the target instruction at ADDR1 (e.g., its address, processor instruction tag, or internal identifier, etc.) and a count value of three or another value (e.g., the number of retired instructions, 13). After receiving the identifier of the target instruction at ADDR1 and said value, the sequential fetch logic updates the corresponding entry in the sequential fetch table, as shown in the entry regarding ADDR1 in the example of the sequential fetch table in Figure 11 . Note that as shown in the sequential fetch table in Figure 11 , ADDR1 and ADDR2 can be the addresses ADDR1 and ADDR2, or other values that represent or identify the corresponding target instructions.

[0073] Retirement of the target instruction for CTI 2 at ADDR2 causes the retirement queue to begin a second / corresponding count. As described above, CTI2 is not adopted, so the target instruction is at the next sequential address ADDR2, and a fetch group is fetched from ADDR2. The retirement queue maintains the corresponding count until a subsequent CTI (CTI3 in fetch group 1116) (which directs the program flow to fetch group 1118) retires. Since two complete fetch groups (fetch groups 1112 to 1114) retire before CTI3 retires, the retirement queue conveys to the sequential fetch logic the identification of the target instruction at ADDR2 (e.g., its address, processor instruction label, or internal identifier, etc.) and a count value of two or another value (e.g., the number of retired instructions, 13). After receiving the identifier of the target instruction at ADDR2 and the value, the sequential fetch logic updates the corresponding entry in the sequential fetch table, as Figure 11 shown in the entry for ADDR2 in the example of the sequential fetch table in

[0074] In some embodiments, the retirement queue employs at least one threshold to report the count to the branch target buffer and / or the sequential fetch logic. In these embodiments, when fewer than the threshold number of fetch groups are counted, the retirement queue does not report the CTI and / or the count to the sequential fetch logic. In this way, the retirement queue can avoid jittering the record (i.e., the branch target buffer or the sequential fetch table) where the count is stored by causing entries to be overwritten quickly. In some embodiments, the threshold is dynamic and can be set / resetted conditionally at runtime. In some embodiments, when fewer than the threshold number of fetch groups appear between a corresponding CTI or target instruction and a subsequent CTI, an entry in the record (e.g., in the branch target buffer) can mark the count of the entry as invalid or set it to a default value (e.g., 0).

[0075] Although the example in Figures 8 to 11 describes the retirement queue as maintaining a count of the number of fetch groups (or, more generally, individual retired instructions) that retire sequentially before a subsequent CTI retires, in some embodiments, different functional blocks maintain the count. Generally, in the described embodiments, any functional block that can identify individual CTI instructions and / or a particular type of CTI instruction and count the number of instructions between the identified CTI instructions can maintain the count. For example, in some embodiments, the decode unit and / or the branch prediction unit perform some or all of the operations for maintaining the count, although for mispredicted branches, various rollback operations and recovery operations may be performed.

[0076] Return Figure 7, when the program code is subsequently executed, the sequential fetch logic determines, based on the records associated with a given CTI, that a specified number of fetch groups without CTI will be fetched after the CTI for sequential execution (step 702). As used herein, a record "associated with a CTI" can be a record associated with the CTI itself or the corresponding fetch group for an implementation that uses a branch target buffer to store records, or a record associated with the target instruction of the CTI for an implementation that uses a sequential fetch table to store records. Generally, after encountering a CTI or target instruction (e.g., detecting a CTI or target instruction in a fetch group based on an address or other information associated with the instruction), the sequential fetch logic obtains, from the appropriate record, a count of the number of fetch groups without CTI that will be fetched sequentially. As described above, the count can be a numerical value, a string, or other value indicating a specific number of fetch groups.

[0077] When each of the specified number of fetch groups has been fetched and is ready for execution, the sequential fetch logic then prevents a corresponding access to the branch prediction unit (e.g., branch prediction unit 212) to obtain branch prediction information about the instructions in that fetch group (step 704). Generally, during this operation, the sequential fetch table inhibits, blocks, or otherwise prevents access to the branch prediction unit, thereby avoiding unnecessarily consuming power, etc. In some embodiments, in addition to preventing access to the branch prediction unit, the sequential fetch logic also avoids performing the check against the count (as shown in step 702) until the operation in step 704 is complete, and may place functional blocks such as part or all of the sequential fetch table or the branch prediction unit in a low power mode.

[0078] Figure 12 According to some embodiments, a flowchart is presented that illustrates a process for determining a specified number of fetch groups without CTI that will be fetched sequentially after a CTI and using the specified number to prevent access to a branch prediction unit. Note that Figure 12 the operations shown are presented as general examples of operations performed by some embodiments. Operations performed by other embodiments include different operations, operations performed in a different order, and / or operations performed by different functional blocks. Figure 12 The operations of Figure 7 will be generally described with respect to steps 702 to 704 of Figure 12 and thus

[0079] For Figure 12 the operations, it is assumed that a sequential fetch record associated with a CTI instruction (or, more generally, its target instruction), i.e., a record of the number of fetch groups without the CTI instruction that will be fetched sequentially after the CTI, is stored in the sequential fetch table, as described above with respect toFigure 7 and Figures 10 to 11 as described. However, note that for embodiments where sequential fetch records are stored in a branch target buffer, a similar operation will be performed as described elsewhere herein. It is also assumed that: prior to step 1200, the program counter (supplied by the next PC function block) indicates a fetch group that will fetch the target instruction including the CTI. It is further assumed that: the fetch group includes four instructions. However, these values and conditions are used to present an example in Figure 12 and are not the same in all embodiments.

[0080] Figure 12 The operation in

[0081] begins when the sequential fetch logic (e.g., sequential fetch logic 214) obtains a specified number of fetch groups that will be sequentially fetched after the fetch group (step 1200) from the sequential fetch table based on the addresses or identifiers of the instructions in the fetch group. For this operation, the sequential fetch logic uses the program counter to check if there is a match or "hit" in the sequential fetch table for any of the four instructions in the fetch group. For example, the sequential fetch logic may calculate a hash value or another index value based on the program counter (i.e., the fetch address), which will be compared with some or all of the entries in the sequential fetch table. As another example, the sequential fetch logic may calculate all the addresses of the instructions in the fetch group based on the program counter and may compare each calculated address with each entry in the sequential fetch table. As described above, the instructions in the fetch group are the target instructions of the CTI, so a match and hit occur in the sequential fetch table. The sequential fetch logic thus reads the specified number of fetch groups from the matching entry in the sequential fetch table.

[0081] The sequential fetch logic then sets a counter to be equal to the specified number of the fetch group (step 1202). For example, the sequential fetch logic may store the specified number of the fetch group or its representation in a dedicated counter register or other memory location.

[0082] When each fetch group is fetched and ready to execute, the sequential fetch logic blocks corresponding accesses to the branch prediction unit (step 1204). For example, the sequential fetch logic may assert one or more control signals to block circuit elements in the branch prediction unit from performing access operations, may block sending addresses or related values to the branch prediction unit, may pause the clock, power down circuit elements, and / or perform other operations to block corresponding accesses to the branch prediction unit. For this operation, in some embodiments, by using some or all of the above techniques to block individual functional blocks in the branch prediction unit from performing related operations, each of two or more separate and possibly parallel accesses to the branch prediction unit is completely blocked, such as branch prediction solutions, branch address acquisition, etc.

[0083] Although in Figure 12Not shown, but in some embodiments, for extraction groups with a non - zero counter, access to the sequential extraction table is also blocked. This is done because it is known that the extraction group does not include CTI and thus does not include the target instruction of CTI, making such lookups unnecessary. In some embodiments, when the counter is non - zero, the sequential extraction table is placed in a reduced - power mode. For example, the sequential extraction table can pause the control clock (e.g., via clock gating), reduce the electrical power, de - assert the enable signal, etc.

[0084] When each extraction group is fetched and ready for execution, the sequential extraction logic also decrements the counter (step 1206). For example, the sequential extraction logic can decrement the value of the counter in a dedicated counter register or other memory location by one, can convert the counter to the next lower value or its representation, etc.

[0085] Before the counter reaches zero (step 1208), the sequential extraction logic continues to block access to the branch prediction unit (step 1204) and decrements the counter as extraction groups are fetched and ready for execution (step 1206). After the counter reaches zero, i.e., after the last of the specified number of extraction groups has been fetched and is being readied for execution, the sequential extraction logic permits corresponding access to the branch prediction unit to obtain branch prediction information due to one or more subsequent extraction groups being fetched and ready for execution (step 1210). In other words, when the counter equals zero, the sequential extraction logic permits the execution of normal branch prediction operations, such as branch target and branch direction prediction. In this way, the sequential extraction logic blocks branch prediction access (and possibly sequential extraction table access) when the counter is non - zero to avoid unnecessary access to the branch prediction unit (and possibly the sequential extraction table).

[0086] Multi-threaded processor

[0087] In some embodiments, the processor 102 in the electronic device 100 is a multi-threaded processor and thus supports two or more separate threads of instruction execution. Generally, a multi-threaded processor includes functional blocks and / or hardware structures dedicated to each separate thread, but may also include functional blocks and / or hardware structures shared among the threads and / or performing corresponding operations for more than one thread. For example, functional blocks such as a branch prediction unit and an in-order fetch unit may perform branch prediction operations and branch prediction unit access prevention respectively for all processes. As another example, the in-order fetch table (in embodiments where an in-order fetch table is used) may be implemented as a single in-order fetch table for all threads (or a certain combination of multiple threads), or may be implemented on a per-thread basis such that each thread has a corresponding separate in-order fetch table. In these embodiments, the records in each in-order fetch table will pertain to the associated thread and may differ from the records kept in the in-order fetch table regarding other records. As yet another example, the in-order fetch logic may prevent access to the branch prediction unit on a per-thread basis and may thus maintain a separate and independent counter for each thread, which counter is used as described herein to prevent access to the branch prediction unit regarding the corresponding thread.

[0088] As described above, in some embodiments, while preventing access to the branch prediction unit for extraction groups that do not include a CTI and are sequentially fetched after the CTI, the in-order fetch logic may also prevent access to the in-order fetch table and place the in-order fetch table in a reduced power mode. In these embodiments, when multiple threads depend on the in-order fetch table (when one in-order fetch table is used to store the records of two or more threads), the in-order fetch table may remain in the full power mode / active to serve other threads (and thus will not be switched to the reduced power mode). The same is true for the branch prediction unit; when only a single thread is using the branch prediction unit, the branch prediction unit may be placed in the reduced power mode when access is being blocked. However, when the branch prediction unit is being used by two or more threads, the branch prediction unit may remain in the full power mode / active to serve other threads. However, as described herein, no specific access is made to a particular thread.

[0089] In some embodiments, an electronic device (e.g., electronic device 100 and / or a portion thereof) performs some or all of the operations described herein using code and / or data stored on a non-transitory computer-readable storage medium. More specifically, when performing the described operations, the electronic device reads the code and / or data from the computer-readable storage medium and executes the code and / or uses the data. The computer-readable storage medium can be any device, medium, or combination thereof that stores code and / or data for use by the electronic device. By way of example, the computer-readable storage medium can include, but is not limited to, volatile or non-volatile memory, including flash memory, random access memory (eDRAM, RAM, SRAM, DRAM, DDR, DDR2 / DDR3 / DDR4 SDRAM, etc.), read-only memory (ROM), and / or magnetic or optical storage media (e.g., disk drives, tapes, CDs, DVDs).

[0090] In some embodiments, one or more hardware modules perform some or all of the operations described herein. By way of example, the hardware modules can include, but are not limited to, one or more processors / cores / central processing units (CPUs), application specific integrated circuit (ASIC) chips, field programmable gate arrays (FPGAs), computing units, embedded processors, graphics processing units (GPUs) / graphics cores, pipelines, accelerated processing units (APUs), functional blocks, system management units, power controllers, and / or other programmable logic devices. When activated, the hardware modules perform some or all of the operations. In some embodiments, the hardware modules include one or more general-purpose circuits that are configured by executing instructions (program code, firmware, etc.) to perform the operations.

[0091] In some embodiments, data structures representing some or all of the structures and mechanisms described herein (e.g., processor 102, memory 104, and / or portions thereof) are stored on a non-transitory computer-readable storage medium that includes a database or other data structure that can be read by an electronic device and used, directly or indirectly, to fabricate hardware that includes the structures and mechanisms. For example, the data structure can be a behavioral or register transfer level (RTL) description of the functionality of hardware in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool that can synthesize the description to produce a netlist that includes a listing of gates / circuit elements from a synthesis library that represents the functionality of hardware that includes the above-described structures and mechanisms. The netlist can then be placed and routed to produce a dataset that describes the geometry to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce one or more semiconductor circuits (e.g., integrated circuits) corresponding to the above-described structures and mechanisms. Alternatively, the database on the computer-accessible storage medium can be a netlist (with or without a synthesis library) or a dataset (as needed), or a graphics data system (GDS) II data.

[0092] In this specification, variables or unspecified values (i.e., general descriptions of values in the absence of a specific instance of a value) are represented by letters such as N. As used herein, although similar letters may be used at different locations in this specification, the variables and unspecified values in each case are not necessarily the same, i.e., there may be different variable amounts and values intended for some or all of the general variables and unspecified values. In other words, in this specification, N and any other letter used to represent variables and unspecified values are not necessarily related to each other.

[0093] As used herein, the expressions “et cetera” or “etc.” are intended to present one and / or cases, i.e., the equivalent of “at least one” of the elements associated with the etc. in a list. For example, in the statement “the system performs a first operation, a second operation, etc.,” the system performs at least one of the first operation, the second operation, and other operations. Additionally, the elements associated with etc. in a list are only instances in a set of instances – and at least some of the instances may not appear in some embodiments.

[0094] The foregoing description of the embodiments has been presented for purposes of illustration and description. The foregoing description is not intended to be exhaustive or to limit the embodiments to the disclosed form. Accordingly, many modifications and variations will be apparent to practitioners skilled in the art. Additionally, the above disclosure is not intended to limit the embodiments. The scope of the embodiments is defined by the appended claims.

Claims

1. An electronic device for handling control transfer instructions when executing instructions in program code, the electronic device comprising: a processor, the processor comprising: a branch prediction circuit; and sequential fetch logic circuitry, wherein the sequential fetch logic circuitry: obtains a record associated with a control transfer instruction, the record including a value representing the number of instruction fetch groups previously determined not to include control transfer instructions that will be fetched after the control transfer instruction; determines, based on the value from the record, a given number of fetch groups not including control transfer instructions to be fetched for sequential execution after the control transfer instruction; and when each of the given number of fetch groups is fetched and ready for execution, blocks corresponding access to the branch prediction circuit to obtain branch prediction information about the instructions in the fetch group.

2. The electronic device according to claim 1, wherein the processor further comprises: a branch target buffer associated with the branch prediction circuit; wherein the sequential fetch logic circuitry: maintains a count of the number of fetch groups that are sequentially retired before a subsequent control transfer instruction retires after detecting a control transfer instruction of a specified type; and stores a record indicating the count of the number of fetch groups in an entry associated with the control transfer instruction in the branch target buffer.

3. The electronic device according to claim 2, wherein the control transfer instruction of the specified type includes an unconditional control transfer instruction having a static target instruction, or a conditional control transfer instruction in which an untaken path is directed to a static target instruction.

4. The electronic device according to claim 2, wherein the sequential fetch logic circuitry: stores the record indicating the count of the number of fetch groups in the entry associated with the control transfer instruction in the branch target buffer only when the count of the number of fetch groups is greater than a threshold.

5. The electronic device according to claim 1, wherein the processor further comprises: a sequential fetch table; wherein the sequential fetch logic circuitry: maintains a count of the number of fetch groups that are sequentially retired before a subsequent control transfer instruction retires after a target instruction for the control transfer instruction retires; and stores a record indicating the count of the number of fetch groups in an entry associated with the target instruction of the control transfer instruction in the sequential fetch table.

6. The electronic device according to claim 5, wherein the sequential fetch logic circuitry: stores the record indicating the count of the number of fetch groups in the entry associated with the target instruction of the control transfer instruction in the sequential fetch table only when the count of the number of fetch groups is greater than a threshold.

7. The electronic device according to claim 1, wherein, when blocking corresponding access to the branch prediction circuit to obtain branch prediction information when each of the given number of fetch groups is fetched and ready for execution, the sequential fetch logic circuitry is configured to: set a counter equal to the given number of fetch groups; Decrement the counter when each fetch group is fetched and ready for execution and the corresponding access to the branch prediction circuit is blocked; And After the counter reaches zero, due to one or more subsequent fetch groups being fetched and ready for execution, access to the corresponding access of the branch prediction circuit is permitted to obtain branch prediction information about the instructions in the one or more subsequent fetch groups.

8. The electronic device according to claim 1, wherein the branch prediction circuit Comprises: A branch target buffer; And A branch direction predictor circuit; Wherein when the access to the branch prediction circuit is blocked, the sequential fetch logic circuit blocks access to at least the branch target buffer and the branch direction predictor circuit.

9. The electronic device according to claim 1, wherein when each of the given number of fetch groups is fetched and ready for execution, the sequential fetch logic circuit: Blocks access to one or more circuits storing records associated with control transfer instructions; and Places the one or more circuits in a reduced power mode.

10. The electronic device according to claim 1, wherein each fetch group includes a reserved number of instructions fetched in the same transfer operation for execution in the processor.

11. The electronic device according to claim 1, wherein blocking the corresponding access to the branch prediction circuit includes blocking at least one operation that consumes power, thereby saving power.

12. A method for handling control transfer instructions in an electronic device when executing instructions in program code, the electronic device having a processor including a branch prediction circuit and a sequential fetch logic circuit, the method Comprises: Obtaining, by the sequential fetch logic circuit, a record associated with a control transfer instruction, the record including a value representing the number of instruction fetch groups previously determined not to include control transfer instructions that will be fetched after the control transfer instruction; Determining, by the sequential fetch logic circuit, based on the value from the record, a given number of fetch groups that do not include control transfer instructions to be fetched for sequential execution after the control transfer instruction; And When each of the given number of fetch groups is fetched and ready for execution, blocking, by the sequential fetch logic circuit, the corresponding access to the branch prediction circuit to obtain branch prediction information about the instructions in the fetch group.

13. The method according to claim 12, wherein the processor further includes a branch target buffer associated with the branch prediction circuit, and wherein the method further Comprises: After detecting a specified type of control transfer instruction, maintaining, by the processor, a count of the number of fetch groups that are sequentially retired before a subsequent control transfer instruction retires; And Storing, by the sequential fetch logic circuit, a record indicating the count of the number of fetch groups in an entry associated with the control transfer instruction in the branch target buffer.

14. The method according to claim 13, wherein the control transfer instruction of the specified type includes an unconditional control transfer instruction having a static target instruction, or a conditional control transfer instruction whose non-taken path is directed to a static target instruction.

15. The method according to claim 13, the method further comprises: storing the record indicating the count of the number of the extraction groups in the entry associated with the control transfer instruction in the branch target buffer only when the count of the number of the extraction groups is greater than a threshold.

16. The method according to claim 12, wherein the processor further includes a sequential extraction table, and wherein the method further comprises: after the target instruction of the control transfer instruction retires, maintaining, by the processor, a count of the number of extraction groups that are sequentially retired before the subsequent control transfer instruction retires; and storing, by the sequential extraction logic circuit, a record indicating the count of the number of the extraction groups in the entry in the sequential extraction table associated with the target instruction of the control transfer instruction.

17. The method according to claim 16, the method further comprises: storing the record indicating the count of the number of the extraction groups in the entry in the sequential extraction table associated with the target instruction of the control transfer instruction only when the count of the number of the extraction groups is greater than a threshold.

18. The method according to claim 12, wherein preventing a corresponding access to the branch prediction circuit to obtain branch prediction information when each of the given number of extraction groups is extracted and ready for execution comprises: setting, by the sequential extraction logic circuit, a counter to be equal to the given number of extraction groups; decrementing, by the sequential extraction logic circuit, the counter when each extraction group is extracted and ready for execution and the corresponding access to the branch prediction circuit is prevented; and after the counter reaches zero, permitting, by the sequential extraction logic circuit, a corresponding access to the branch prediction circuit to obtain branch prediction information about instructions in the one or more subsequent extraction groups since the one or more subsequent extraction groups are extracted and ready for execution.

19. The method according to claim 12, wherein the branch prediction circuit includes a branch target buffer and a branch direction predictor circuit, and wherein preventing the access to the branch prediction circuit includes at least preventing access to the branch target buffer and the branch direction predictor circuit.

20. The method according to claim 12, wherein each extraction group includes a reserved number of instructions extracted in the same transfer operation for preparation for execution in the processor.

21. The method according to claim 12, the method further comprises: when each of the given number of extraction groups is extracted and ready for execution: preventing, by the sequential extraction logic circuit, access to one or more circuits storing records associated with the control transfer instruction; and Place the one or more circuits in a reduced power mode via the sequential extraction logic circuit.

22. The method of claim 12, wherein preventing corresponding access to the branch prediction circuit includes preventing at least one operation that consumes power, thereby saving power.

Citation Information

Patent Citations

  • glow wire candle

    DE620717C

  • Selectively blocking branch instruction prediction

    CN103513963A