Instruction cache prefetch throttling
Patent Information
- Application Number
- CN202080083173.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-10
- Filing Date
- 2020-11-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2040-11-19
Smart Images

Figure CN114761922B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. nonprovisional patent application No. 16 / 709,831, filed December 10, 2019, the contents of which are hereby incorporated by reference. Background Technology
[0003] In a microprocessor, instructions are fetched for sequential execution until a branch is encountered. A branch causes a change in the address from which the fetched instructions are located and can be associated with latency in instruction fetch throughput. For example, a branch might need to be evaluated to determine whether to adopt it and what its destination is. However, a branch cannot be evaluated until it enters the instruction execution pipeline. Branch latency is the difference between the time it takes for a branch to be fetched and the time it takes to evaluate the branch to determine its outcome and thus determine which instructions need to be fetched next.
[0004] Branch prediction helps mitigate this latency by predicting the existence and outcome of branch instructions based on instruction addresses. The branch target buffer stores information that associates the program counter address with the branch target. The presence of an entry in the branch target buffer implicitly indicates the existence of a branch at the program counter associated with that entry. The instruction fetch unit can prefetch instructions into the instruction cache based on the contents of the branch target buffer. The branch target buffer is frequently modified. Attached Figure Description
[0005] A more detailed understanding can be obtained from the following description, given by way of example in conjunction with the accompanying drawings:
[0006] Figure 1 It is a block diagram of an exemplary apparatus in which one or more of the disclosed embodiments may be implemented;
[0007] Figure 2 It is located in Figure 1 A block diagram of the instruction execution pipeline within the processor;
[0008] Figure 3 It is a block diagram of the example branch target buffer based on an example;
[0009] Figure 4 This illustrates an example operation for rate limiting of instruction cache prefetching;
[0010] Figure 5 This is a flowchart illustrating an example of a method for prefetching instructions into an instruction cache. Detailed Implementation
[0011] A technique is provided for controlling instruction prefetching into an instruction cache. The technique includes tracking either or both of branch target buffer misses and instruction cache misses, modifying a rate-limiting trigger based on the tracking, and adjusting prefetching activity based on the rate-limiting trigger.
[0012] Figure 1 This is a block diagram of an example device 100 that implements several aspects of the present disclosure. Device 100 includes, for example, a computer, gaming device, handheld device, set-top box, television, mobile phone, or tablet computer. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Device 100 may also optionally include an input driver 112 and an output driver 114. It should be understood that device 100 may include... Figure 1 Additional components not shown.
[0013] Processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and GPU located on the same die, or one or more processor cores, wherein each processor core is a CPU or a GPU. Memory 104 may be located on the same die as processor 102, or may be located separately from processor 102. Memory 104 includes volatile or non-volatile memory, such as random access memory (RAM), dynamic RAM, or cache.
[0014] Storage device 106 includes fixed or removable storage devices, such as hard disk drives, solid-state drives, optical disks, or flash drives. Input device 108 includes a keyboard, keypad, touchscreen, touchpad, detector, microphone, accelerometer, gyroscope, biometric scanner, or network connection (e.g., a wireless LAN card for transmitting and / or receiving wireless IEEE 802 signals). Output device 110 includes a display, speaker, printer, haptic feedback device, one or more lights, antenna, or network connection (e.g., a wireless LAN card for transmitting and / or receiving wireless IEEE 802 signals).
[0015] Input driver 112 communicates with processor 102 and input device 108, and allows processor 102 to receive input from input device 108. Output driver 114 communicates with processor 102 and output device 110, and allows processor 102 to send output to output device 110. Note that input driver 112 and output driver 114 are optional components, and device 100 will operate in the same manner without input driver 112 and output driver 114.
[0016] Figure 2 It is located in Figure 1A block diagram of the instruction execution pipeline 200 within processor 102 is shown. While a specific configuration for the instruction execution pipeline 200 is illustrated, it should be understood that any instruction execution pipeline 200 that uses a branch target buffer to prefetch instructions into an instruction cache is within the scope of this disclosure. The instruction execution pipeline 200 retrieves instructions from memory and executes those instructions, outputs data to memory, and modifies the state of elements associated with the instruction execution pipeline 200 (e.g., registers within register file 218).
[0017] The instruction execution pipeline 200 includes an instruction fetch unit 204 that fetches instructions from system memory (e.g., memory 104) using an instruction cache 202, a decoder 208 that decodes the fetched instructions, a function unit 216 that performs computations to process the instructions, a load / store unit 214 that loads data from system memory or stores data into system memory via a data cache 220, and a register file 218 that includes registers storing working data for the instructions. A reordering buffer 210 tracks currently in-process instructions and ensures their in-order termination, although out-of-order execution is allowed while in process. The term "in-process instruction" refers to an instruction that has been received by the reordering buffer 210 but whose results have not yet been committed to the processor's architectural state (e.g., the results have been written to the register file, etc.). A reservation station 212 maintains in-process instructions and tracks instruction operands. When all operands are ready to execute a specific instruction, the reservation station 212 sends the instruction to function unit 216 or load / store unit 214 for execution. Completed instructions are marked as retired in reorder buffer 210 and retired while at the front of reorder buffer queue 210. Retirement refers to the act of submitting the result of an instruction to the processor's architectural state. Examples of instruction retirement include writing additional results to a register via an add instruction, writing loaded values to a register via a load instruction, or causing the instruction flow to jump to a new location via a branch instruction.
[0018] Various components of the instruction execution pipeline 200 communicate via a common data bus 222. For example, functional unit 216 and load / store unit 214 write results to the common data bus 222, which can be read by the holding station 212 to execute related instructions and by the reordering buffer 210 as the final processing result of an ongoing instruction that has been completed. Load / store unit 214 also reads data from the common data bus 222. For example, load / store unit 214 reads the result of a completed instruction from the common data bus 222 and writes the result to memory via data cache 220 to store the instruction.
[0019] Typically, instruction fetch unit 204 fetches instructions from memory sequentially. Sequential control flow may be interrupted by branch instructions, causing instruction pipeline 200 to fetch instructions from non-sequential addresses. Branch instructions can be conditional, causing a branch only if a specific condition is met, or unconditional, and can specify a target directly or indirectly. Direct targets are specified by constants within the instruction byte itself, while indirect targets are specified by values in registers or memory. Direct and indirect branches can be conditional or unconditional.
[0020] For instruction execution pipeline 200, sequential instruction fetching is relatively straightforward. Instruction fetch unit 204 sequentially fetches large blocks of contiguous stored instructions for execution. However, branch instructions may interrupt such fetching for several reasons. More specifically, depending on the type of branch instruction, for the execution of a branch instruction, any or all of the following may occur: instruction decoder 208 determines that the instruction is actually a branch instruction, functional unit 216 calculates the target of the branch instruction, and functional unit 216 evaluates the conditions of the branch instruction. Because there is a delay between fetching and issuing branch instructions for execution by instruction fetch unit 204 and when the instruction execution pipeline 200 actually executes the branch instruction, instruction fetch unit 204 includes branch target buffer 206. In short, branch target buffer 206 caches the predicted block address of previously encountered branch or jump instructions and the target of those branch or jump instructions. When fetching instructions to be executed, instruction fetch unit 204 provides one or more addresses corresponding to the next one or more instructions to be executed to branch target buffer 206. If there is an entry in BTB 206 corresponding to the provided address, it means that BTB 206 predicts the existence of a branch instruction. In this case, instruction fetching unit 204 obtains the target of the branch to be adopted from BTB 206, and if the branch is an unconditional branch or a conditional branch to be adopted, it begins fetching from the target.
[0021] In some implementations, instruction fetch unit 204 performs a prefetch operation, whereby instruction fetch unit 204 prefetches instructions into instruction cache 202, expecting those instructions to eventually be served to the remainder of instruction execution pipeline 200. In some implementations, performing such a prefetch operation includes testing whether an upcoming instruction in instruction cache 202 is hit in branch target buffer 206. If a hit occurs, instruction fetch unit 204 prefetches the instruction from the hit target (i.e., the target address specified by the entry in branch target buffer 206) into instruction cache 202; and if no hit occurs, instruction fetch unit 204 sequentially prefetches instructions into instruction cache 202. Sequential prefetching of instructions means prefetching instructions whose addresses follow those of previously prefetched instructions. In one example, instruction fetch unit 204 prefetches a first cache line into instruction cache 202, detects no BTB hit in the cache line, and prefetches a second cache line that is contiguous with the first cache line into instruction cache 202.
[0022] Figure 3 This is a block diagram of an example branch target buffer 206 based on one example. The branch target buffer 206 includes multiple BTB entries 302. Each BTB entry 302 includes a prediction block address 304, a target 306, and a branch type 308. The prediction block address 304 indicates the starting address of the instruction block containing the previously encountered branch. In some implementations, a new prediction block begins if cache line boundaries cross. The target 306 is the target address of the branch. The branch type 308 encodes the branch type. Examples of branch types include conditional branches that branch to the target address based on the result of conditional evaluation, or unconditional branch (or "jump") instructions that always result in a jump to the target of the instruction. In various implementations, BTB 206 is directly mapped, set-associative, or fully associative.
[0023] In operation, when instruction fetch unit 204 fetches instructions into instruction cache 202, instruction fetch unit 204 compares the address of the fetched instruction with the prediction block address 304 of BTB entry 302. A match is called a hit and causes instruction fetch unit 204 to fetch the instruction at target 306 of BTB entry 302, where a hit occurs if branch type 308 is unconditional or if branch type 308 is conditional and expected to be adopted by conditional branch predictor 209. If there is no match in BTB 206 (called a miss) or a hit for a conditional branch that is not expected to be adopted, instruction fetch unit 204 fetches instructions sequentially into instruction cache 202. Sequential fetching means that the instructions in memory that caused the BTB miss are fetched in order. For example, if the entire cache line of the BTB 206 test instruction is either hit or not hit, then sequential fetching includes fetching the cache line at the memory address aligned with the cache line following the one being tested.
[0024] The predictions made by the branch target buffer 206 are sometimes incorrect. In one example, an entry in the branch target buffer 206 indicates that an instruction at a certain address causes control flow to the indicated target. However, when the instruction is actually evaluated by the functional unit 216, the instruction instead causes control flow to a different address. In one example, a BTB entry for the address of an instruction is initially placed in BTB 206 because a conditional branch is adopted, but the next time that conditional branch instruction is executed, the branch instruction is not adopted. In this case, BTB 206 updates the conditional predictor to improve adoption / non-adoption accuracy. In another example, a miss occurs in BTB 206 because there is no entry in BTB 206 corresponding to the address of the branch instruction, and the branch instruction actually causes out-of-sequence control flow when evaluated by the functional unit 216. In this case, BTB 206 generates a new BTB entry 302 for the newly encountered branch instruction, which indicates the target of the branch instruction.
[0025] As described above, when a miss occurs in BTB 206, instruction fetch unit 204 sequentially prefetches instructions into instruction cache 202. In some cases, this activity is advantageous because a miss in BTB 206 generally means that no branch exists for the instruction set currently being checked, and therefore the instructions will be fetched sequentially. Even when some entries 302 in BTB 206 are incorrect, or when some BTB entries 302 for the instruction branch being checked do not exist in BTB 206, it is generally expected that at least some predictions made by BTB 206 will be correct, and the instructions prefetched into instruction cache 202 will be used for subsequent fetches before being evicted from the instruction cache. However, when control flow reaches a completely new segment that has not been recently encountered, most or all branches in that new segment will not have corresponding entries in BTB 206. Therefore, instruction fetch unit 204 will cause sequential prefetching to occur in instruction cache 202. One problem with this operating mode is that the prefetching that occurs is likely to prefetch a large number of instructions that will not be used in subsequent executions into cache 202. This activity is sometimes undesirable for at least the following reasons: cache operations consume power, and prefetching cache operations that do not need the instructions represent wasted power consumption; in some implementations, instruction cache 202 shares bandwidth with other caches, so using that bandwidth for instruction cache 202 reduces the bandwidth available for other cache operations; and putting instructions that will not be used later into instruction cache 202 and evicting other instructions that may be used later causes delays in prefetching the evicted instructions and leads to additional cache operations in the future.
[0026] Figure 4 Example operations for limiting instruction cache prefetching are shown. According to these operations, instruction fetch unit 204 limits the prefetching rate to instruction cache 202 based on a rate limiting trigger 404 and a rate limiting level 406 (if used). More specifically, instruction fetch unit 204 rate-limits instruction cache prefetching if the rate limiting trigger 404 indicates that instruction fetch unit 204 will rate-limit instruction cache prefetching, and instruction fetch unit 204 does not rate-limit instruction cache prefetching if the rate limiting trigger 404 indicates that instruction fetch unit 204 will not rate-limit instruction cache prefetching. If used, the rate limiting level 406 indicates the degree to which rate limiting occurs. The rate limiting trigger 404 and the rate limiting level 406 represent data stored in memory locations (such as registers, memory, etc.). The rate limiting trigger 404 and the rate limiting level 406 are set according to techniques described elsewhere in this document (such as in the paragraphs below).
[0027] Now let's discuss additional example details regarding instruction cache prefetching rate limiting techniques.
[0028] If the rate-limiting trigger 404 indicates that a prefetch will occur in the instruction cache 202, such a prefetch is performed as follows: The branch target buffer 206 receives the instruction address from the instruction fetch unit 204 to determine where to perform the prefetch. In some implementations, the instruction address provided by the instruction fetch unit 204 follows the prediction of the branch target buffer 206 until corrected by the functional unit 216. More specifically, the instruction fetch unit 204 identifies the predicted next address to be fetched based on whether there is a hit or miss in the branch target buffer 206, and then prefetches from that target, repeating this process. If the prefetched instruction includes a branch that was not predicted by the branch target buffer 206, or if the branch target buffer 206 predicts that a branch does not exist or will not be used, the functional unit 216 detects such an error, and the instruction fetch unit 204 begins prefetching from the correct address specified by the functional unit 216. Please note that functional unit 216 is the unit that actually "executes" branch instructions, for example, by performing target address calculation and execution condition evaluation, and thus corrects prediction errors made by instruction fetch unit 204 and branch target buffer 206. Therefore, prefetching to instruction cache 202 is based on the instruction address provided by the execution path predicted by branch target buffer 206.
[0029] As described above, if the rate limiting trigger 404 indicates that rate limiting will occur, the instruction fetch unit 204 rate limits the instruction cache prefetch. In one example, rate limiting of the instruction cache prefetch is performed by limiting the number of incomplete prefetches waiting for a response from a lower-level cache (e.g., the number of L1 cache fills waiting for a response from a L2 cache). In another example, limiting the number of incomplete prefetches excludes any prefetch requested before a branch prediction correction from a functional unit. In yet another example, rate limiting of the instruction cache prefetch is performed by limiting the rate at which prefetch requests are made. In yet another example, rate limiting of the instruction cache prefetch is performed by prefetching instructions into a small number of caches in the instruction cache hierarchy (e.g., prefetching to a L3 cache but not to a L2 or L1 cache, prefetching to both a L3 and L2 cache but not to a L1 cache, etc.). In the example using rate limiting level 406, the rate limiting level indicates the number of incomplete prefetches, the rate of requested prefetches, and the number of cached instructions in the instruction cache hierarchy that will be prefetched.
[0030] The techniques used to set the rate limiting trigger 404 and the rate limiting level 406 (if used) are now discussed. The miss tracker 402 detects and records misses in the branch target buffer 206 and the instruction cache 202, indicating lines that were likely not previously encountered. Based on the misses encountered in the branch target buffer 206 and the instruction cache 202, the miss tracker 402 modifies the rate limiting trigger 404, and, in changing the implementation of the rate limiting level, modifies the rate limiting level 406.
[0031] Several examples are now provided of how the miss tracker 402 modifies the rate limiting trigger 404. In a first example, the miss tracker 402 tracks the number of consecutive lookups that miss in the branch target buffer 206 and the instruction cache (“IC”) 202. If the number of such consecutive misses is higher than a threshold, the miss tracker 402 sets the rate limiting trigger 404 to indicate that prefetching will be rate-limited. If the number of consecutive misses is not higher than the threshold, the miss tracker 402 sets the rate limiting trigger 404 to indicate that prefetching will not be rate-limited. A consecutive miss in BTB 206 is a miss that occurs after another miss and there is no miss between the two misses, where the term “after” refers to sequential memory addresses. In one example, a miss occurs for a first memory address in BTB 206, and subsequently, another miss occurs for the immediately following address. In this example, the two misses are considered consecutive misses. In another example, a miss occurs for a first memory address in BTB 206 and instruction cache (“IC”) 202, followed by a hit for the next address, and then another miss for the next address. In this example, the two misses are not considered consecutive. Note that in some examples, the address is the address of a cache line, and BTB 206 includes entries for individual cache lines. A hit in BTB 206 is a hit for an address of a cache line and means that BTB 206 predicts that at least one instruction in the cache line is a branch. A miss in BTB 206 is a miss for an address of a cache line and means that BTB 206 predicts that no instruction in the cache line is a branch.
[0032] In the second example, the miss tracker 402 tracks the number of consecutive misses in the branch target buffer 206 and the instruction cache 202, and also tracks the total number of such incomplete misses in the branch target buffer 206 and the instruction cache 202. In this case, an incomplete miss is a request sent by the instruction fetch unit to a lower-level cache after a miss in the instruction cache 202 and the branch target buffer 206. When the lower-level cache returns data for one of these misses, the incomplete miss counter is decremented. The data is then loaded into the instruction cache 202.
[0033] In the second example, if the number of consecutive misses exceeds a first threshold and the number of incomplete misses exceeds a second threshold, the miss tracker 402 modifies the rate limiting trigger 404 to indicate that instruction prefetch rate limiting for the instruction cache 202 will occur. If the number of consecutive misses is below the first threshold or the number of incomplete misses is below the second threshold, the miss tracker 402 modifies the rate limiting trigger 404 to indicate that instruction prefetch rate limiting for the instruction cache 202 will not occur. In an implementation using rate limiting level 406, if the number of consecutive or incomplete misses increases, the miss tracker 402 increases the rate limiting level 406, and if the number of consecutive or incomplete misses decreases, the rate limiting level 406 decreases.
[0034] In the third example, the miss tracker 402 tracks the number of incomplete misses. In response to the number of incomplete misses exceeding a first threshold, the miss tracker 402 begins tracking the number of consecutive misses. In response to the number of consecutive misses exceeding a second threshold, the miss tracker 402 sets the rate limiting trigger 404 to indicate that rate limiting is enabled. In response to the number of consecutive misses falling below a third threshold or the number of incomplete misses falling below a fourth threshold, the miss tracker 402 sets the rate limiting trigger 404 to indicate that rate limiting is disabled. In an implementation using the rate limiting level 406, if the number of consecutive or incomplete misses increases, the miss tracker 402 increases the rate limiting level 406, and if the number of consecutive or incomplete misses decreases, the rate limiting level 406 decreases.
[0035] In some implementations, as an extension of the above-described techniques (including any of the first, second, or third examples above), the miss tracker 402 tracks instruction cache misses but not BTB misses. In the implementation where the miss tracker 402 tracks instruction cache misses, an incomplete miss is defined as an instruction cache miss that causes a request for a byte to a lower-level cache, but the request has not yet been completed. In such implementations, the miss tracker 402 tracks the total number of incomplete instruction cache misses and the number of consecutive instruction cache misses. If the number of consecutive misses is higher than a first threshold and the number of instruction cache misses is higher than a second threshold, the miss tracker 402 sets a rate limiting trigger to indicate that rate limiting will occur. If the number of instruction cache misses is lower than the first threshold or the number of consecutive instruction cache misses is lower than the second threshold, the miss tracker 402 sets a rate limiting trigger 404 to indicate that rate limiting will not occur.
[0036] In some implementations, BTB lookups are performed per cache line. More specifically, the address provided by instruction fetch unit 204 to branch target buffer 206 is aligned to the size of the cache line. Furthermore, entries in BTB 206 store one or more branch targets for all cache lines. A miss occurs if a cache line address is provided to such BTB 206 and no entry corresponds to that cache line address. Such misses are considered incomplete until instruction pipeline 200 confirms that there are no branches in the cache line or identifies one or more branches in the cache line and the targets of those branches. In the event of misses for two or more sequential cache line addresses, consecutive misses occur. Some BTBs 206 include entries for a certain number of branches stored per cache line; therefore, if a cache line includes more than the maximum number of branches in its BTB entries, there may be multiple BTB entries per cache line.
[0037] Please note that the branch target buffer 206 is included as shown in the figure. Figure 2 In the instruction fetch unit 204, however, it should be understood that the interaction between the instruction fetch unit 204 and the branch target buffer 206 (such as applying instruction addresses to the branch target buffer 206) is performed by the appropriate entity on the instruction fetch unit 204 (such as fixed-function or programmable circuitry).
[0038] In some examples, when a branch misprediction occurs, the count of the number of incomplete prefetches is reset, and / or the count of the number of incomplete instruction cache misses is reset.
[0039] Figure 5This is a flowchart based on an example method 500 for prefetching instructions into an instruction cache. Although... Figures 1 to 4 The system described herein is method 500, but those skilled in the art will recognize that any system configured to perform the steps of method 500 in any technically feasible order is within the scope of this disclosure.
[0040] Method 500 begins at step 502, where a miss tracker 402 tracks BBT 206 misses and / or instruction cache 202 misses. The miss tracker 402 can track such misses in several ways. In one example, the miss tracker 402 tracks the number of consecutive misses in BTB 206 and instruction cache 202. As described elsewhere in this document, consecutive misses in BTB 206 and IC 202 are two or more misses that occur without an intervening hit in BTB 206 or IC 202. An intervening hit is a hit that occurs at an address between two consecutive misses. In a second example, the miss tracker 402 tracks the number of incomplete misses and consecutive misses in BTB 206 and IC 202. In a third example, the miss tracker 402 tracks the number of incomplete misses and, if the number of incomplete misses exceeds a threshold, tracks the number of consecutive misses. In some implementations, as a variation of any of the above examples, the miss tracker 402 only tracks cache misses in the instruction cache 202.
[0041] In step 504, the miss tracker 402 modifies the rate limiting trigger 404 based on the tracking. In an example where the miss tracker 402 tracks the number of consecutive BTB misses, if the number of consecutive misses is higher than a first threshold, the miss tracker 402 sets the rate limiting trigger 404 to indicate that prefetching will be rate-limited. If the number of consecutive misses is not higher than a second threshold, the miss tracker 402 sets the rate limiting trigger 404 to indicate that prefetching will not be rate-limited. In some examples, the first threshold and the second threshold are the same, and in other examples, the first threshold and the second threshold are different.
[0042] In the example where the miss tracker 402 tracks the number of consecutive misses in branch target buffer 206, if the number of consecutive misses is higher than a first threshold and the number of incomplete misses is higher than a second threshold, the miss tracker 402 modifies the rate limiting trigger 404 to indicate that instruction prefetch rate limiting for instruction cache 202 will occur. If the number of consecutive misses is lower than a third threshold or the number of incomplete misses is lower than a fourth threshold, the miss tracker 402 modifies the rate limiting trigger 404 to indicate that instruction prefetch rate limiting for instruction cache 202 will not occur. In some implementations, the third threshold is the same as the first threshold. In some implementations, the third threshold is different from the first threshold. In some implementations, the second threshold is the same as the fourth threshold. In some implementations, the second threshold is different from the fourth threshold. In the implementation of rate limiting level 406, if the number of consecutive or incomplete misses increases, the miss tracker 402 increases the rate limiting level 406, and if the number of consecutive or incomplete misses decreases, the rate limiting level 406 decreases.
[0043] In another example, miss tracker 402 tracks the number of incomplete misses, and in response to the number of incomplete misses exceeding a first threshold, miss tracker 402 begins tracking the number of consecutive misses. In this example, in response to the number of consecutive misses exceeding a second threshold, miss tracker 402 sets rate limiting trigger 404 to indicate that rate limiting is enabled. In response to the number of consecutive misses falling below a third threshold or in response to the number of incomplete misses falling below a fourth threshold, miss tracker 402 sets rate limiting trigger 404 to indicate that rate limiting is disabled. In an implementation using rate limiting level 406, if the number of consecutive or incomplete misses increases, miss tracker 402 increases rate limiting level 406, and if the number of consecutive or incomplete misses decreases, rate limiting level 406 decreases.
[0044] In some implementations, as a variation of the above techniques (including any of the first, second, or third examples above), the miss tracker 402 only tracks instruction cache misses.
[0045] In step 506, the instruction fetch unit 204 adjusts the instruction prefetch activity according to the rate limiting trigger 404. In an implementation using the rate limiting level 406, the instruction fetch unit 204 adjusts the instruction prefetch activity according to the rate limiting level 406. Adjusting the prefetch activity based on the rate limiting trigger includes enabling instruction prefetching into the instruction cache 202 if the rate limiting trigger 404 is enabled, and disabling instruction prefetching into the instruction cache 202 if the rate limiting trigger 404 is disabled. In an implementation using the rate limiting level 406, the instruction fetch unit 204 adjusts the degree of prefetching based on the rate limiting level 406. Generally, a higher rate limiting level 406 means fewer prefetches, and a lower rate limiting level 406 means more prefetches. In some examples, more prefetching is associated with prefetching more instructions than fewer prefetches, and in other examples, more prefetching is associated with prefetching instructions into caches that are prefetched more frequently than fewer prefetches.
[0046] This paper proposes a method for controlling instruction prefetching into an instruction cache. The method includes tracking either or both of branch target buffer misses and instruction cache misses. The method includes modifying a rate-limiting trigger based on the tracking. The method also includes adjusting prefetching activity based on the rate-limiting trigger.
[0047] In some implementations, the method includes modifying the rate limiting level based on the tracing, wherein adjusting the prefetch activity is also performed based on the rate limiting level. In some implementations, the tracing includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold, and in response, setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
[0048] In some implementations, the tracing includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold, and in response, setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited. In some implementations, the tracing includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds the first threshold and the number of such incomplete branch target buffer misses and instruction cache misses exceeds a second threshold, and in response, setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited. In some implementations, the tracing includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed the first threshold or the number of incomplete branch target buffer misses and instruction cache misses does not exceed the second threshold, and in response, setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited. In some implementations, the tracing includes tracking the number of consecutive branch target buffer misses and instruction cache misses in response to detecting that the number of incomplete branch target buffer misses and instruction cache misses exceeds a first threshold, and setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited in response to detecting that the number of consecutive branch target buffer misses exceeds a second threshold. In some implementations, the tracing includes tracking incomplete and consecutive instruction cache misses but not tracking branch target buffer misses.
[0049] It should be understood that many variations are possible based on the disclosure herein. Although features and elements are described above in specific combinations, each feature or element can be used alone without other features and elements, or in various combinations with or without other features and elements.
[0050] The various functional units shown in the figures and / or described herein (including, where appropriate, processor 102, input driver 112, input device 108, output driver 114, output device 110, instruction cache 202, instruction fetch unit 204, branch target buffer 206, decoder 208, reordering buffer 210, reservation station 212, data cache 220, load / store unit 214, functional unit 216, register file 218, common data bus 222, and miss tracker 402) can be implemented as hardware circuitry, software executing on a programmable processor, or a combination of hardware and software. The provided methods can be implemented in a general-purpose computer, processor, or processor core. Suitable processors include, for example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate array (FPGA) circuitry, any other type of integrated circuit (IC), and / or state machine. Such a processor can be manufactured by configuring a manufacturing process using the results of processed Hardware Description Language (HDL) instructions and other intermediate data, including netlists (such instructions can be stored on a computer-readable medium). The result of such processing can be a mask, which is then used in the semiconductor manufacturing process to manufacture the processor implementing various aspects of the implementation scheme.
[0051] The methods or flowcharts provided herein can be implemented using computer programs, software, or firmware incorporated in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media such as CD-ROMs and digital versatile optical discs (DVDs).
Claims
1. A method for controlling instruction prefetching into an instruction cache, the method comprising: Track either or both of the number of branch target buffer misses and the number of instruction cache misses; The rate limiting trigger and rate limiting level are modified based on the tracing, wherein the rate limiting level indicates one of the number of instructions prefetched and the number of instructions prefetched into the cache therein. as well as The prefetching activity for the instruction cache is adjusted based on the rate limiting trigger and the rate limiting level.
2. The method of claim 1, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
3. The method of claim 1, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited.
4. The method of claim 1, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold and the number of incomplete branch target buffer misses and instruction cache misses exceeds a second threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
5. The method of claim 1, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold or the number of such incomplete branch target buffer misses and instruction cache misses does not exceed a second threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited.
6. The method of claim 1, wherein: The tracking includes tracking the number of consecutive branch target buffer misses and instruction cache misses in response to detecting that the number of incomplete branch target buffer misses and instruction cache misses exceeds a first threshold; and In response to detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a second threshold, the rate limiting trigger is set to indicate that instruction cache prefetching will be rate-limited.
7. The method of claim 1, wherein the tracking comprises: The number of misses in the instruction cache is not tracked.
8. The method of claim 1, wherein the tracking comprises: The stated number of misses in the branch target buffer is not tracked.
9. An instruction fetching system for controlling instruction prefetching into an instruction cache, the instruction fetching system comprising: Branch target buffer; as well as Missed tracker, the missed tracker is configured to: Track either or both of the number of branch target buffer misses and the number of instruction cache misses in the branch target buffer. The rate limiting trigger and rate limiting level are modified based on the tracing, wherein the rate limiting level indicates one of the number of instructions prefetched and the number of instructions prefetched into the cache therein. as well as An instruction fetching unit is configured to adjust prefetching activity for the instruction cache based on the rate limiting trigger and the rate limiting level.
10. The instruction extraction system as described in claim 9, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
11. The instruction extraction system as described in claim 9, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited.
12. The instruction extraction system as described in claim 9, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold and the number of incomplete branch target buffer misses and instruction cache misses exceeds a second threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
13. The instruction extraction system as described in claim 9, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold or the number of incomplete branch target buffer misses and instruction cache misses does not exceed a second threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited.
14. The instruction extraction system as described in claim 9, wherein: The tracking includes tracking the number of consecutive branch target buffer misses and instruction cache misses in response to detecting that the number of incomplete branch target buffer misses and instruction cache misses exceeds a first threshold; and In response to detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a second threshold, the rate limiting trigger is set to indicate that instruction cache prefetching will be rate-limited.
15. The instruction extraction system as described in claim 9, wherein: The tracking includes not tracking the number of misses in the instruction cache.
16. The instruction extraction system as described in claim 9, wherein: The tracking includes not tracking the number of misses in the branch target buffer.
17. A processor configured to prefetch control instructions, the processor comprising: Instruction cache; Branch target buffer; as well as Tracker miss, the tracker being configured to perform the following operations: Tracking either or both of the number of branch target buffer misses and the number of instruction cache misses in the branch target buffer; and The rate limiting trigger and rate limiting level are modified based on the tracing, wherein the rate limiting level indicates one of the number of instructions prefetched and the number of instructions prefetched into the cache therein. as well as An instruction fetching unit is configured to adjust prefetching activity for the instruction cache based on the rate limiting trigger and the rate limiting level.
18. The processor of claim 17, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
19. The processor of claim 17, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited.
20. The processor of claim 17, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a first threshold and the number of incomplete branch target buffer misses and instruction cache misses exceeds a second threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will be rate-limited.
21. The processor of claim 17, wherein: The tracking includes detecting that the number of consecutive branch target buffer misses and instruction cache misses does not exceed a first threshold or the number of incomplete branch target buffer misses and instruction cache misses does not exceed a second threshold, and in response, the modification includes setting the rate limiting trigger to indicate that instruction cache prefetching will not be rate-limited.
22. The processor of claim 17, wherein: The tracking includes tracking the number of consecutive branch target buffer misses and instruction cache misses in response to detecting that the number of incomplete branch target buffer misses and instruction cache misses exceeds a first threshold; and In response to detecting that the number of consecutive branch target buffer misses and instruction cache misses exceeds a second threshold, the rate limiting trigger is set to indicate that instruction cache prefetching will be rate-limited.
23. The processor of claim 17, wherein: The tracking includes not tracking the number of misses in the instruction cache.
24. The processor of claim 17, wherein: The tracking includes not tracking the number of misses in the branch target buffer.
Citation Information
Patent Citations
Branch target buffer for a data processing apparatus
CN110520836A
Prefetch optimization in shared resource multi-core systems
US20140136795A1