Instruction Cache Prefetch Throttling

By tracking misses in the branch target buffer and instruction cache, and adjusting prefetch activity through a throttling toggle, the technique addresses branch prediction delays and improves cache efficiency, reducing power consumption and optimizing instruction fetch throughput.

JP7676393B2Active Publication Date: 2025-05-14ADVANCED MICRO DEVICES INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022532021
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-10
Filing Date
2020-11-19
Publication Date
2025-05-14
Estimated Expiration
2040-11-19

AI Technical Summary

Technical Problem

Branch prediction in microprocessors often leads to delays due to incorrect predictions and the resulting need for sequential prefetching of instructions, which can result in wasted power consumption and reduced cache efficiency.

Method used

A technique for controlling instruction cache prefetching by tracking branch target buffer misses and instruction cache misses, modifying a throttling toggle based on this tracking, and adjusting prefetch activity accordingly to prevent unnecessary prefetching.

Benefits of technology

This approach reduces power consumption and improves cache efficiency by minimizing the prefetching of unused instructions, thereby optimizing instruction fetch throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676393000001
    Figure 0007676393000001
  • Figure 0007676393000002
    Figure 0007676393000002
  • Figure 0007676393000003
    Figure 0007676393000003
Patent Text Reader

Abstract

Techniques for controlling the prefetching of instructions into an instruction cache are provided. Embodiments describe techniques for tracking branch target buffer misses and / or instruction cache misses, modifying a throttling toggle based on the tracking, and adjusting prefetch activity based on the throttling toggle.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Patent Application No. 16 / 709,831, filed December 10, 2019, the contents of which are incorporated herein by reference. [Background technology]

[0002] In a microprocessor, instructions are fetched sequentially for execution until a branch occurs. Branches cause a change in the address at which instructions are fetched and can contribute to delays in instruction fetch throughput. For example, a branch may need to be evaluated to determine not only whether it is taken, but what the branch's destination is. However, a branch cannot be evaluated until it enters the instruction execution pipeline. Branch latency is related to the difference between the time a branch is fetched and the time the branch is evaluated to determine the outcome of the branch and, as a result, what instructions need to be fetched next.

[0003] Branch prediction helps reduce this delay by predicting the presence and outcome of a branch instruction based on the instruction address. A branch target buffer stores information that associates a program counter address with a branch target. The presence of an entry in the branch target buffer implies that a branch is present in the program counter associated with that entry. An instruction fetch unit may prefetch instructions into an instruction cache based on the contents of the branch target buffer. Corrections to the branch target buffer occur frequently.

[0004] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which: [Brief description of the drawings]

[0005] [Figure 1]FIG. 1 is a block diagram of an example device capable of implementing one or more disclosed embodiments. [Diagram 2] FIG. 2 is a block diagram of an instruction execution pipeline within the processor of FIG. 1. [Diagram 3] 1 is a block diagram of an exemplary branch target buffer, according to one example. [Figure 4] FIG. 13 illustrates an example of an operation for throttling instruction cache prefetching. [Diagram 5] 1 is a flow diagram of a method for prefetching instructions into an instruction cache, according to an example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0006] Techniques for controlling the prefetching of instructions into an instruction cache are provided, the techniques including tracking either or both of branch target buffer misses and instruction cache misses, modifying a throttling toggle based on the tracking, and adjusting prefetch activity based on the throttling toggle.

[0007] 1 is a block diagram of an example device 100 in which aspects of the disclosure may be implemented. Device 100 may include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, or a tablet computer. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Device 100 may also include optional input drivers 112 and output drivers 114. It will be understood that device 100 may include additional components not shown in FIG. 1.

[0008] The processor 102 may include a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU on the same die, or one or more processor cores, each of which may be a CPU or a GPU. The memory 104 may be located on the same die as the processor 102 or may be separate from the processor 102. The memory 104 may include volatile or non-volatile memory (e.g., random access memory (RAM), dynamic RAM, cache).

[0009] Storage devices 106 include fixed or removable storage (e.g., hard disk drives, solid state drives, optical disks, flash drives). Input devices 108 include a keyboard, keypad, touch screen, touch pad, detector, microphone, accelerometer, gyroscope, biometric scanner, or network connection (e.g., wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals). Output devices 110 include a display, speaker, printer, haptic feedback device, one or more lights, antenna, or network connection (e.g., wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).

[0010] The input driver 112 communicates with the processor 102 and the input device 108, allowing the processor 102 to receive input from the input device 108. The output driver 114 communicates with the processor 102 and the output device 110, allowing the processor 102 to send output to the output device 110. It should be noted that the input driver 112 and the output driver 114 are optional components, and the device 100 would operate similarly if the input driver 112 and the output driver 114 were not present.

[0011] Figure 2 is a block diagram of an instruction execution pipeline 200 within processor 102 of Figure 1. While one particular configuration of instruction execution pipeline 200 is shown, it should be understood that any instruction execution pipeline 200 that uses a branch target buffer to prefetch instructions into an instruction cache is within the scope of this disclosure. Instruction execution pipeline 200 fetches instructions from memory, executes the instructions, and outputs data to memory as well as changing the state of elements associated with instruction execution pipeline 200, such as registers in register file 218.

[0012] The instruction execution pipeline 200 includes an instruction fetch unit 204 that fetches instructions from a system memory (such as memory 104) using an instruction cache 202, a decoder 208 that decodes the fetched instructions, functional units 216 that perform calculations to process the instructions, a load store unit 214 that loads and stores data from and to the system memory via a data cache 220, and a register file 218 that includes registers that store working data for instructions. A reorder buffer 210 tracks instructions that are currently in-flight and ensures in-order retirement of instructions while allowing out-of-order execution while in-flight. The term "in-flight instruction" refers to an instruction that has been received by the reorder buffer 210 but that has not yet had its results committed to the processor's architectural state (e.g., written to a register file). A reservation station 212 maintains the in-flight instructions and tracks the instruction operands. When all operands are ready for execution of a particular instruction, reservation stations 212 send the instruction to functional units 216 or load / store units 214 for execution. A completed instruction is retired when it is marked for retirement in reorder buffer 210 and is at the head of reorder buffer queue 210. Retirement refers to the act of committing the result of an instruction to the architectural state of the processor. For example, writing an add result to a register via an add instruction, writing a loaded value to a register via a load instruction, or causing the instruction flow to jump to a new location via a branch instruction are all examples of retiring an instruction.

[0013] The various elements of instruction execution pipeline 200 communicate via a common data bus 222. For example, functional units 216 and load / store unit 214 write results onto common data bus 222 which may be read by reservation stations 212 for execution of dependent instructions or by reorder buffer 210 as final processing results for in-flight instructions that have completed execution. Load / store unit 214 also reads data from common data bus 222. For example, for store instructions, load / store unit 214 reads the results of the completed instruction from common data bus 222 and writes the results to memory via data cache 220.

[0014] In general, the instruction fetch unit 204 fetches instructions in memory sequentially. The sequential control flow may be interrupted by branch instructions, thus causing the instruction execution pipeline 200 to fetch instructions from non-sequential addresses. Branch instructions may be conditional (which causes the branch to occur only if a certain condition is met) or unconditional, and may specify targets directly or indirectly. Direct targets are specified by constants in the instruction byte itself, while indirect targets are specified by values ​​in registers or memory. Direct and indirect branches may be conditional or unconditional.

[0015] Sequential fetching of instructions is relatively straightforward for instruction execution pipeline 200. Instruction fetch unit 204 sequentially fetches large chunks of contiguously stored instructions for execution. However, branch instructions may interrupt such fetching for a number of reasons. More specifically, depending on the type of branch instruction, execution of the branch instruction may involve instruction decoder 208 determining that the instruction is in fact a branch instruction, functional unit 216 calculating the target of the branch instruction, and / or functional unit 216 evaluating the condition of the branch instruction. Because there is a delay between when a branch instruction is fetched and issued for execution by instruction fetch unit 204 and when the branch instruction is actually executed by instruction execution pipeline 200, instruction fetch unit 204 includes branch target buffer 206. In essence, branch target buffer 206 caches predicted block addresses of previously encountered branch or jump instructions along with the target of the branch or jump instruction. When fetching instructions to be executed, instruction fetch unit 204 provides one or more addresses corresponding to one or more instructions to be executed next to branch target buffer 206. If an entry corresponding to the provided address exists in BTB 206, this means that BTB 206 predicts the presence or absence of a branch instruction. In this example, instruction fetch unit 204 retrieves the target of the taken branch from BTB 206 and starts fetching from the target if the branch is an unconditional branch or a conditional branch that is predicted to be taken.

[0016] In some embodiments, instruction fetch unit 204 performs a prefetch operation during which instruction fetch unit 204 prefetches instructions into instruction cache 202 in anticipation of the instructions eventually being provided to the remainder of instruction execution pipeline 200. In some embodiments, performing such a prefetch operation involves examining upcoming instructions in instruction cache 202 for a hit in branch target buffer 206. If a hit occurs, instruction fetch unit 204 prefetches an instruction from the target of the hit (i.e., the target address specified by an entry in branch target buffer 206) into instruction cache 202; if a hit does not occur, instruction fetch unit 204 prefetches instructions sequentially into instruction cache 202. Prefetching instructions sequentially means prefetching instructions at an address following the address of a previously prefetched instruction. In one example, the instruction fetch unit 204 prefetches a first cache line into the instruction cache 202, does not detect a BTB hit in that cache line, and prefetches a second cache line into the instruction cache 202 that is contiguous to the first cache line.

[0017] 3 is a block diagram of an exemplary branch target buffer 206, according to one example. The branch target buffer 206 includes multiple BTB entries 302. Each BTB entry 302 includes a predicted block address 304, a target 306, and a branch type 308. The predicted block address 304 indicates the starting address of a block of instructions that contains a previously encountered branch. In some embodiments, a new predicted block begins when a cache line boundary is crossed. The target 306 is the target address of the branch. The branch type 308 encodes the type of branch. Examples of branch types include a conditional branch that branches to a target address based on the outcome of a conditional evaluation, or an unconditional branch (or "jump") instruction that always results in a jump to the instruction's target. In various embodiments, the BTB 206 is direct-mapped, set-associative, or fully associative.

[0018] In operation, when instruction fetch unit 204 fetches an instruction into instruction cache 202, it compares the address of the fetched instruction to the predicted block address 304 in BTB entry 302. A match is called a hit and will result in instruction fetch unit 204 fetching the instruction at target 306 of the hit BTB entry 302 if the branch type 308 is unconditional, or if the branch type 308 is conditional and predicted to be taken by conditional branch predictor 209. If there is no match in BTB 206, this is called a miss, or for a conditional branch that is predicted not taken, it is called a hit, and instruction fetch unit 204 fetches instructions sequentially into instruction cache 202. Sequential fetching refers to fetching instructions that are contiguous in memory to the instruction that caused the BTB miss. For example, if the BTB 206 examines an entire cache line of instructions for a hit and no hit occurs, a sequential fetch involves fetching a cache line at an aligned memory address of the cache line subsequent to the examined cache line.

[0019] The predictions made by the branch target buffer 206 may be inaccurate. In one example, an entry in the branch target buffer 206 indicates that an instruction at an address will result in control flowing to the indicated target. However, when the instruction is actually evaluated by the functional units 216, the instruction instead flows control to a different address. In one example, a BTB entry for an instruction's address is initially placed in the BTB 206 due to a conditional branch being taken, but the branch instruction does not take the branch when the conditional branch instruction is next executed. In this example, the BTB 206 updates the conditional predictor to improve the accuracy of the take / not take selection. In another example, a miss occurs in the BTB 206 because no entry exists in the BTB 206 corresponding to the address of the branch instruction, but the branch instruction actually results in a non-sequential control flow when evaluated by the functional units 216. In this example, the BTB 206 generates a new BTB entry 302 for a newly encountered branch instruction that indicates the target of the branch instruction.

[0020] As described above, when a miss occurs in the BTB 206, the instruction fetch unit 204 sequentially prefetches instructions into the instruction cache 202. In certain circumstances, this type of activity is beneficial because a miss in the BTB 206 generally means that there are no branches for the currently examined instruction set, and therefore the instructions are fetched sequentially. Even if some entries 302 in the BTB 206 are incorrect, or some BTB entries 302 for the branches of the instruction being examined are not present in the BTB 206, it is often expected that at least some of the predictions made by the BTB 206 will be correct, and thus the instructions prefetched into the instruction cache 202 will be used for a later fetch before being evicted from the instruction cache. However, if control flows to an entirely new section that has not been recently encountered, most or all of the branches in that new section will not have corresponding entries in the BTB 206. Thus, the instruction fetch unit 204 will cause a sequential prefetch in the instruction cache 202. One problem with this mode of operation is that the prefetching that occurs is likely to prefetch into cache 202 a large number of instructions that will be unused in subsequent execution. This activity may be undesirable for at least the following reasons: cache traffic consumes power, and cache traffic to prefetch unused instructions represents wasted power consumption. In some embodiments, instruction cache 202 shares bandwidth with other caches, and therefore using that bandwidth for instruction cache 202 reduces the bandwidth available for other cache traffic. Putting instructions into instruction cache 202 that will not be used later will evict other instructions that may be used, resulting in delays to refetch the evicted instructions and additional cache traffic in the future.

[0021] 4 illustrates an example of operations for throttling instruction cache prefetches. According to these operations, the instruction fetch unit 204 throttles prefetches to the instruction cache 202 based on a throttling toggle 404 and, if used, a throttling degree 406. More specifically, if the throttling toggle 404 indicates that the instruction fetch unit 204 throttles the instruction cache prefetches, the instruction fetch unit 204 throttles the instruction cache prefetches, and if the throttling toggle 404 indicates that the instruction fetch unit 204 does not throttle the instruction cache prefetches, the instruction fetch unit 204 does not throttle the instruction cache prefetches. The throttling degree 406, if used, indicates the degree to which throttling occurs. The throttling toggle 404 and the throttling degree 406 represent data stored in memory locations such as registers, memory, etc. Throttling toggle 404 and throttling degree 406 are set according to the techniques described elsewhere herein, such as in the following paragraphs.

[0022] Further example details of the instruction cache prefetch throttling technique are now provided.

[0023] When throttling toggle 404 indicates that a prefetch into instruction cache 202 is to occur, such a prefetch occurs as follows: Branch target buffer 206 receives instruction addresses from instruction fetch unit 204 and determines the prefetch destination. In some embodiments, the instruction address provided by instruction fetch unit 204 follows the branch target buffer 206 prediction until corrected by functional unit 216. More specifically, instruction fetch unit 204 identifies the next predicted address to be fetched based on whether there is a hit or miss in branch target buffer 206, and then prefetches from that target, repeating this operation. If the prefetched instructions include a branch that was not predicted by branch target buffer 206, or if branch target buffer 206 predicted a branch that does not exist or was not taken, functional unit 216 detects such an error and instruction fetch unit 204 begins prefetching from the correct address specified by functional unit 216. It should be noted that functional unit 216 is the unit that actually "executes" branch instructions, e.g., by performing target address calculations and performing conditional evaluations, and thus corrects prediction errors made by instruction fetch unit 204 and branch target buffer 206. Thus, prefetching into instruction cache 202 occurs based on instruction addresses provided by the execution path predicted by branch target buffer 206.

[0024] As described above, when the throttling toggle 404 indicates that throttling is to occur, the instruction fetch unit 204 throttles instruction cache prefetches. In one example, throttling of instruction cache prefetches occurs by limiting the number of outstanding prefetches waiting for a response from a lower level cache (e.g., the number of level 1 cache fills waiting for a response from a level 2 cache). In another example, limiting the number of outstanding prefetches excludes prefetches that are requested prior to a branch prediction correction from a functional unit. In another example, throttling of instruction cache prefetches occurs by limiting the rate at which prefetches are requested. In another example, throttling of instruction cache prefetches occurs by prefetching instructions into fewer caches in the instruction cache hierarchy (e.g., prefetching into level 3 cache but not into level 2 cache or level 1 cache, prefetching into level 3 cache and level 2 cache but not into level 1 cache, etc.). In examples where throttling degree 406 is used, the throttling degree indicates the number of outstanding prefetches, the rate at which prefetches are being requested, and the number of caches in the instruction cache hierarchy from which instructions are prefetched.

[0025] Next, techniques for setting the throttling toggle 404 and, if used, the throttling degree 406 are described. The miss tracker 402 detects and records misses in the branch target buffer 206 and misses in the instruction cache 202, indicating lines that are likely not previously encountered. Based on the misses encountered in the branch target buffer 206 and the misses encountered in the instruction cache 202, the miss tracker 402 varies the throttling toggle 404 and, in embodiments that vary the degree of throttling, the throttling degree 406.

[0026] Next, some examples of how the miss tracker 402 may modify the throttling toggle 404 are shown. In a first example, the miss tracker 402 tracks the number of consecutive lookups that miss in the branch target buffer 206 and the instruction cache (IC) 202. If the number of such consecutive misses exceeds a threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that the prefetch is throttled. If the number of consecutive misses does not exceed the threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that the prefetch is not throttled. Consecutive misses in the BTB 206 are misses that occur after another miss with no hit between the two misses, where the term "after" refers to the subsequent memory address. In one example, a miss occurs to the first memory address of the BTB 206, and then another miss occurs to the address immediately following it. In this example, the two misses are considered consecutive misses. In another example, a miss occurs on the first memory address in the BTB 206 and instruction cache (IC) 202, then a hit occurs on the address immediately following that, and then a miss occurs on the address immediately following that. In this example, the two misses are not considered consecutive. Note that in some examples, the addresses are cache line addresses and the BTB 206 includes entries for individual cache lines. A BTB 206 hit means that the address of the cache line is a hit and that the BTB 206 predicts at least one instruction in the cache line to be a branch. A BTB 206 miss means that the address of the cache line is a miss and that the BTB 206 predicts that no instruction in the cache line is a branch.

[0027] In a second example, the miss tracker 402 tracks the number of consecutive misses in the branch target buffer 206 and the instruction cache 202, and also tracks the total number of outstanding misses that have missed in both the branch target buffer 206 and the instruction cache 202. A outstanding miss in this case is a request sent by the instruction fetch unit to a lower level cache after missing in the instruction cache 202 and the branch target buffer 206. The outstanding miss counter is decremented when the lower level cache returns data for any of these misses. This data is then installed in the instruction cache 202.

[0028] In a second example, if the number of consecutive misses exceeds a first threshold and the number of outstanding misses exceeds a second threshold, the miss tracker 402 changes the throttling toggle 404 to indicate that instruction prefetch throttling for the instruction cache 202 will occur. If the number of consecutive misses is less than the first threshold or the number of outstanding misses is less than the second threshold, the miss tracker 402 changes the throttling toggle 404 to indicate that instruction prefetch throttling for the instruction cache 202 will not occur. In an embodiment in which a throttling degree 406 is used, the miss tracker 402 increases the throttling degree 406 if the number of consecutive misses or outstanding misses increases and decreases the throttling degree 406 if the number of consecutive misses or outstanding misses decreases.

[0029] In a third example, the miss tracker 402 tracks the number of outstanding misses. In response to the number of outstanding misses exceeding a first threshold, the miss tracker 402 begins tracking the number of consecutive misses. In response to the number of consecutive misses exceeding a second threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that throttling is enabled. In response to the number of consecutive misses being less than a third threshold or the number of outstanding misses being less than a fourth threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that throttling is disabled. In an embodiment in which a throttling degree 406 is used, the miss tracker 402 increases the throttling degree 406 if the number of consecutive misses or outstanding misses increases and decreases the throttling degree 406 if the number of consecutive misses or outstanding misses decreases.

[0030] In some embodiments, in an extension of the above technology, including any of the first, second, or third examples above, the miss tracker 402 tracks instruction cache misses and does not track BTB misses. In embodiments in which the miss tracker 402 tracks instruction cache misses, an outstanding miss is defined as an instruction cache miss that resulted in a request for bytes to a lower level of the cache where the request has not yet been satisfied. In such an embodiment, the miss tracker 402 tracks the total number of outstanding instruction cache misses and the number of consecutive instruction cache misses. If the number of consecutive misses exceeds a first threshold and the number of instruction cache misses exceeds a second threshold, the miss tracker 402 sets a throttling toggle to indicate that throttling will occur. If the number of instruction cache misses is less than the first threshold or if the number of consecutive instruction cache misses is less than the second threshold, the miss tracker 402 sets a throttling toggle 404 to indicate that throttling will not occur.

[0031] In some embodiments, BTB lookups are performed on a cache line basis. More specifically, addresses provided by the instruction fetch unit 204 to the branch target buffer 206 are aligned to the size of a cache line. Furthermore, entries in the BTB 206 store one or more branch targets for an entire cache line. A miss occurs when a cache line address is provided to such a BTB 206 and there is no entry corresponding to that cache line address. Such a miss is considered outstanding until the instruction execution pipeline 200 either verifies that there are no branches within the cache line or identifies one or more branches within the cache line and their targets. A consecutive miss occurs when there are misses to two or more sequential cache line addresses. Some BTBs 206 contain entries that store up to a certain number of branches per cache line, so there can be multiple BTB entries per cache line if the cache line contains more than the maximum number of branches in a BTB entry.

[0032] It should be noted that although branch target buffer 206 is shown in FIG. 2 as being included in instruction fetch unit 204, it should be understood that interactions between instruction fetch unit 204 and branch target buffer 206, such as applying an instruction address to branch target buffer 206, are performed by a suitable entity on instruction fetch unit 204, such as fixed function circuitry or programmable circuitry.

[0033] In some examples, when a branch misprediction occurs, the count of outstanding prefetches is reset and / or the count of outstanding instruction cache misses is reset.

[0034] 5 is a flow diagram of a method 500 for prefetching instructions into an instruction cache, according to one example. Although method 500 is described with respect to the systems of FIGS. 1-4, one skilled in the art will recognize that any system configured to perform the steps of method 500 in any order that is technically feasible is within the scope of the present disclosure.

[0035] The method 500 begins at step 502, where the miss tracker 402 tracks misses in the BTB 206 and / or the instruction cache 202. There are many ways in which the miss tracker 402 can track such misses. In one example, the miss tracker 402 tracks the number of consecutive misses in the BTB 206 and the instruction cache 202. As described elsewhere herein, consecutive misses in the BTB 206 and the IC 202 are two or more misses that occur without an intervening hit in the BTB 206 or the IC 202. An intervening hit is a hit that occurs to an address that is between two consecutive misses. In a second example, the miss tracker 402 tracks the number of outstanding misses and consecutive misses in the BTB 206 and the IC 202. In a third example, the miss tracker 402 tracks the number of outstanding misses and tracks the number of consecutive misses if the number of outstanding misses exceeds a threshold. In some embodiments, as a variation of either of the above examples, the miss tracker 402 tracks only cache misses in the instruction cache 202.

[0036] In step 504, the miss tracker 402 modifies the throttling toggle 404 based on the tracking. In an example where the miss tracker 402 tracks the number of consecutive BTB misses, if the number of consecutive misses exceeds a first threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that prefetching is throttled. If the number of consecutive misses does not exceed a second threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that prefetching is not throttled. In some examples, the first threshold is the same as the second threshold, and in other examples, the first threshold is different from the second threshold.

[0037] In embodiments where the miss tracker 402 tracks both the number of consecutive misses in the branch target buffer 206 and the number of outstanding misses in the branch target buffer 206, if the number of consecutive misses exceeds a first threshold and the number of outstanding misses exceeds a second threshold, the miss tracker 402 changes the throttling toggle 404 to indicate that instruction prefetch throttling for the instruction cache 202 will occur. If the number of consecutive misses is less than a third threshold or the number of outstanding misses is less than a fourth threshold, the miss tracker 402 changes the throttling toggle 404 to indicate that instruction prefetch throttling for the instruction cache 202 will not occur. In some embodiments, the third threshold is the same as the first threshold. In some embodiments, the third threshold is different from the first threshold. In some embodiments, the second threshold is the same as the fourth threshold. In some embodiments, the second threshold is different from the fourth threshold. In embodiments in which the throttling degree 406 is used, the miss tracker 402 increases the throttling degree 406 if the number of consecutive misses or outstanding misses increases and decreases the throttling degree 406 if the number of consecutive misses or outstanding misses decreases.

[0038] In another example, the miss tracker 402 tracks the number of outstanding misses, and in response to the number of outstanding misses exceeding a first threshold, the miss tracker 402 begins tracking the number of consecutive misses. In this example, in response to the number of consecutive misses exceeding a second threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that throttling is enabled. In response to the number of consecutive misses being less than a third threshold or the number of outstanding misses being less than a fourth threshold, the miss tracker 402 sets the throttling toggle 404 to indicate that throttling is disabled. In an embodiment in which a throttling degree 406 is used, the miss tracker 402 increases the throttling degree 406 if the number of consecutive misses or outstanding misses increases and decreases the throttling degree 406 if the number of consecutive misses or outstanding misses decreases.

[0039] In some embodiments, as a variation of the above technique, including any of the first, second or third examples above, the miss tracker 402 tracks only instruction cache misses.

[0040] At step 506, the instruction fetch unit 204 adjusts its instruction prefetch activity according to the throttling toggle 404. In embodiments where the throttling degree 406 is used, the instruction fetch unit 204 adjusts its instruction prefetch activity according to the throttling degree 406. Adjusting the prefetch activity based on the throttling toggle includes switching off instruction prefetching to the instruction cache 202 if the throttling toggle 404 is on, and switching on instruction prefetching to the instruction cache 202 if the throttling toggle 404 is off. In embodiments where the throttling degree 406 is used, the instruction fetch unit 204 adjusts the degree to which prefetching occurs based on the throttling degree 406. Generally, a higher throttling degree 406 means less prefetching occurs, and a lower throttling degree 406 means more prefetching occurs. In some instances, more prefetching is associated with prefetching more instructions than less prefetching, and in other instances, more prefetching is associated with prefetching more instructions into the cache than less prefetching.

[0041] Presented herein is a method for controlling prefetching of instructions into an instruction cache. The method includes tracking either or both of branch target buffer misses and instruction cache misses. The method includes modifying a throttling toggle based on the tracking. The method also includes adjusting prefetch activity based on the throttling toggle.

[0042] In some embodiments, the method includes modifying a throttling degree based on the tracking, and adjustments of prefetch activity are also performed based on the throttling degree. In some embodiments, the tracking includes detecting that a number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold, and in response, setting a throttling toggle to indicate that instruction cache prefetching is throttled.

[0043] In some embodiments, the tracking includes detecting that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold, and in response, setting a throttling toggle to indicate that instruction cache prefetching is not throttled. In some embodiments, the tracking includes detecting that a number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold and that a number of such outstanding branch target buffer misses and instruction cache misses exceeds a second threshold, and in response, setting a throttling toggle to indicate that instruction cache prefetching is throttled. In some embodiments, the tracking includes detecting either that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold, or that a number of outstanding branch target buffer misses and instruction cache misses does not exceed a second threshold, and in response, setting a throttling toggle to indicate that instruction cache prefetching is not throttled. In some embodiments, the tracking includes, in response to detecting a number of outstanding branch target buffer misses and instruction cache misses exceeding a first threshold, tracking a number of such consecutive branch target buffer misses and instruction cache misses, and in response to detecting a number of consecutive branch target buffer misses exceeding a second threshold, setting a throttling toggle to indicate that instruction cache prefetching is throttled. In some embodiments, the tracking includes tracking outstanding instruction cache misses and consecutive instruction cache misses and not tracking branch target buffer misses.

[0044] It should be understood that many variations are possible based on the disclosure of this specification. Although features and elements have been described above in particular combinations, each feature or element can be used alone without the other features and elements or in various combinations with the other features and elements.

[0045] The various functional units illustrated in the figures and / or described herein (processor 102, input drivers 112, input devices 108, output drivers 114, output devices 110, instruction cache 202, instruction fetch unit 204, branch target buffer 206, decoder 208, reorder buffer 210, reservation stations 212, data cache 220, load / store unit 214, functional units 216, register file 218, common data bus 222, and miss tracker 402, as appropriate) may be implemented as hardware circuits, software running on a programmable processor, or a combination of hardware and software. The methods provided may be performed by a general purpose computer, processor, or processor core. Suitable processors include, by way of example, general purpose processors, special purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and / or a state machine. Such processors may be produced by configuring a manufacturing process with the results of processed hardware description language (HDL) instructions and other intermediate data including a netlist (such instructions may be stored on a computer readable medium). The results of such processing may be a maskwork used in a semiconductor manufacturing process to produce a processor embodying aspects of an embodiment.

[0046] The methods or flow charts provided herein may be implemented in a computer program, software, or firmware embodied in a non-transitory computer-readable storage medium for execution by a general purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, and optical media such as magneto-optical media, e.g., CD-ROM disks and digital versatile disks (DVDs).

Claims

1. 1. A method for controlling prefetching of instructions into an instruction cache, comprising: tracking either or both of branch target buffer misses and instruction cache misses; modifying a throttling toggle and a throttling degree based on the tracking, the throttling degree indicating either a number of instructions to be prefetched or a number of caches from which instruction prefetching occurs; and adjusting a prefetch activity of the instruction cache based on the throttling toggle and the throttling degree. method.

2. The tracking comprises: Detecting a number of consecutive branch target buffer misses and instruction cache misses exceeding a first threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is throttled.

2. The method of claim 1.

3. The tracking comprises: detecting that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is not throttled.

2. The method of claim 1.

4. The tracking comprises: Detecting a number of consecutive branch target buffer misses and instruction cache misses exceeding a first threshold and a number of outstanding branch target buffer misses and instruction cache misses exceeding a second threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is throttled.

2. The method of claim 1.

5. The tracking comprises: Detecting either that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold or that a number of outstanding branch target buffer misses and instruction cache misses does not exceed a second threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is not throttled.

2. The method of claim 1.

6. The tracking comprises: responsive to detecting that the number of outstanding branch target buffer misses and instruction cache misses exceeds a first threshold, tracking a number of consecutive branch target buffer misses and instruction cache misses; in response to detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a second threshold, setting the throttling toggle to indicate that instruction cache prefetching is throttled.

2. The method of claim 1.

7. said tracking including not tracking a number of misses in said instruction cache.

2. The method of claim 1.

8. and tracking includes not tracking a number of misses in a branch target buffer.

2. The method of claim 1.

9. 1. An instruction fetch system for controlling prefetching of instructions into an instruction cache, comprising: The instruction fetch system comprises: A branch target buffer; Equipped with Miss Tracker, The mist tracker is tracking either or both of branch target buffer misses and instruction cache misses for the branch target buffer; modifying a throttling toggle and a throttling degree based on the tracking, the throttling degree indicating either a number of instructions to be prefetched or a number of caches from which instruction prefetching occurs; adjusting a prefetch activity of the instruction cache based on the throttling toggle and the throttling degree; 4. The method of claim 3, Instruction fetch system.

10. The tracking comprises: Detecting a number of consecutive branch target buffer misses and instruction cache misses exceeding a first threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is throttled.

10. The instruction fetch system of claim 9.

11. The tracking comprises: detecting that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is not throttled.

10. The instruction fetch system of claim 9.

12. The tracking comprises: Detecting a number of consecutive branch target buffer misses and instruction cache misses exceeding a first threshold and a number of outstanding branch target buffer misses and instruction cache misses exceeding a second threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is throttled.

10. The instruction fetch system of claim 9.

13. The tracking comprises: Detecting either that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold or that a number of outstanding branch target buffer misses and instruction cache misses does not exceed a second threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is not throttled.

10. The instruction fetch system of claim 9.

14. The tracking comprises: responsive to detecting that the number of outstanding branch target buffer misses and instruction cache misses exceeds a first threshold, tracking a number of consecutive branch target buffer misses and instruction cache misses; in response to detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a second threshold, setting the throttling toggle to indicate that instruction cache prefetching is throttled.

10. The instruction fetch system of claim 9.

15. said tracking including not tracking a number of misses in said instruction cache.

10. The instruction fetch system of claim 9.

16. and tracking includes not tracking a number of misses in a branch target buffer.

10. The instruction fetch system of claim 9.

17. An instruction cache; A branch target buffer; Equipped with Miss Tracker, The mist tracker is tracking either or both of branch target buffer misses and instruction cache misses for the branch target buffer; modifying a throttling toggle and a throttling degree based on the tracking, the throttling degree indicating either a number of instructions to be prefetched or a number of caches from which instruction prefetching occurs; adjusting a prefetch activity of the instruction cache based on the throttling toggle and the throttling degree; 4. The method of claim 3, Processor.

18. The tracking comprises: Detecting a number of consecutive branch target buffer misses and instruction cache misses exceeding a first threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is throttled.

20. The processor of claim 17.

19. The tracking comprises: detecting that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is not throttled.

20. The processor of claim 17.

20. The tracking comprises: Detecting a number of consecutive branch target buffer misses and instruction cache misses exceeding a first threshold and a number of outstanding branch target buffer misses and instruction cache misses exceeding a second threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is throttled.

20. The processor of claim 17.

21. The tracking comprises: Detecting either that a number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold or that a number of outstanding branch target buffer misses and instruction cache misses does not exceed a second threshold; and correspondingly setting the throttling toggle to indicate that instruction cache prefetching is not throttled.

20. The processor of claim 17.

22. The tracking comprises: responsive to detecting that the number of outstanding branch target buffer misses and instruction cache misses exceeds a first threshold, tracking a number of consecutive branch target buffer misses and instruction cache misses; in response to detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a second threshold, setting the throttling toggle to indicate that instruction cache prefetching is throttled.

20. The processor of claim 17.

23. said tracking including not tracking a number of misses in said instruction cache.

20. The processor of claim 17.

24. and tracking includes not tracking a number of misses in a branch target buffer.

20. The processor of claim 17.

Citation Information

Patent Citations

  • Processor, method for operating processor, and information processing system

    JP2009140502A

  • Method and apparatus for saving power by efficiently disabling ways for a set-associative cache

    US20080082753A1

  • Prefetch optimization in shared resource multi-core systems

    US20140136795A1

  • Branch target buffer for a data processing apparatus

    WO2018142140A1