Method for prefetching instructions, information processing apparatus, device, and storage medium
Through the prefetch instruction method, the instructions are prefetched into the first-level cache memory in advance, solving the pipeline delay problem caused by the processor due to cache misses and improving the processor's execution efficiency.
Patent Information
- Application Number
- CN202210570597.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-05-24
AI Technical Summary
In the prior art, the problem of increasing pipeline delay caused by cache miss when the processor retrieves instructions.
Through the prefetch instruction method, the first instruction is received and decoded, determined as a prefetch instruction, and obtained the prefetch address information, prefetch the instruction into the first level cache memory in advance, reducing the probability of cache miss.
It effectively reduces cache misses, improves the processor's operating efficiency and overall system processing performance.
Smart Images

Figure CN114924797B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a prefetch instruction method and an information processing apparatus. Background Art
[0002] In a related central processing unit (CPU) architecture, program instructions and data can be stored in a dynamic random access memory (DRAM). Summary of the Invention
[0003] Embodiments of the present disclosure provide a method and an information processing apparatus for prefetching instructions in a computer to solve the problem of increased pipeline latency caused by cache misses when a processor fetches instructions in the prior art.
[0004] At least one aspect of the present disclosure provides a method for prefetching instructions, including: receiving a first instruction; decoding the first instruction to determine that the first instruction is a prefetch instruction and obtaining prefetch address information in the first instruction; and performing a prefetch operation on a prefetch address based on the prefetch address information.
[0005] In one embodiment, the method further includes: in response to the first instruction being a prefetch instruction, marking the first instruction as having been executed and completed in a retirement unit.
[0006] In one embodiment, the first instruction is included in a first instruction group, and the method further includes: determining a position where the first instruction is inserted in the first instruction group based on a relationship between a size of the first instruction group and a capacity of a first-level cache memory.
[0007] In one embodiment, the first instruction is included in a first instruction group, and the method further includes: inserting the first instruction at any position in the first instruction group.
[0008] In one embodiment, performing a prefetch operation on the prefetch address based on the prefetch address information further includes: prefetching the prefetch address into a first-level cache memory before the first instruction group is executed completely.
[0009] In one embodiment, performing a prefetch operation on the prefetch address based on the prefetch address information further includes: prefetching the prefetch address into a first-level cache memory after the first instruction group is executed completely.
[0010] In one embodiment, the prefetch address information is an absolute address of the prefetch instruction or a relative address indicating the absolute address of the prefetch instruction.
[0011] In one embodiment, based on the prefetch address information, the prefetch operation on the prefetch address further includes: obtaining the virtual address of the prefetched instruction based on the prefetch address information, and sending the virtual address to the instruction translation lookaside buffer unit, where: in response to the prefetch address information being the absolute address of the prefetched instruction, sending the absolute address as the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit; or in response to the prefetch address information being a relative address indicating the absolute address of the prefetched instruction, adding the virtual address of the prefetch instruction to the relative address to obtain the virtual address of the prefetched instruction, and sending the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit.
[0012] In one embodiment, based on the prefetch address information, the prefetch operation on the prefetch address further includes: after sending the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit, converting the virtual address of the prefetched instruction into a physical address, and sending the physical address to the first-level cache memory, where: in response to the virtual address of the prefetched instruction existing in the instruction translation lookaside buffer unit, obtaining the physical address corresponding to the virtual address of the prefetched instruction; in response to the virtual address of the prefetched instruction not existing in the instruction translation lookaside buffer unit, sending an address translation request to the translation lookaside buffer miss status tracking register, obtaining the physical address corresponding to the virtual address of the prefetched instruction from the page table unit based on the address translation request, and returning the obtained physical address to the instruction translation lookaside buffer unit; and in response to the virtual address of the prefetched instruction not existing in the page table unit, ending the prefetch operation.
[0013] In one embodiment, based on the prefetch address information, the prefetch operation on the prefetch address further includes: determining whether the instruction data corresponding to the physical address is in the first-level cache memory, where: in response to the instruction data corresponding to the physical address being in the first-level cache memory, ending the prefetch operation; in response to the instruction data corresponding to the physical address not being in the first-level cache memory, obtaining the instruction data corresponding to the physical address from the lower-level cache memory or memory via the instruction address miss status tracking register, and returning the prefetched instruction data to the first-level cache memory, and ending the prefetch operation.
[0014] At least one aspect of the present disclosure also provides an information processing apparatus, including: a cache memory unit configured to receive a first instruction; a decoding unit configured to decode the first instruction; a distribution unit configured to determine that the first instruction is a prefetch instruction and send the prefetch instruction to a prefetch processing unit; and a prefetch processing unit configured to obtain prefetch address information in the first instruction and perform a prefetch operation on a prefetch address based on the prefetch address information.
[0015] In one embodiment, the apparatus further includes: a retirement unit configured to mark the first instruction as having been executed and completed in the retirement unit in response to the first instruction being a prefetch instruction.
[0016] In one embodiment, the cache memory unit further includes a first-level cache memory, wherein the first instruction is included in a first instruction group, and a position where the first instruction is inserted in the first instruction group is determined based on a relationship between a size of the first instruction group and a capacity of the first-level cache memory.
[0017] In one embodiment, the first instruction is included in a first instruction group, and the first instruction is inserted at any position in the first instruction group.
[0018] In one embodiment, the apparatus further includes an execution unit, and the prefetch processing unit is further configured to prefetch the prefetch address into the first-level cache memory before the execution unit finishes executing the first instruction group.
[0019] In one embodiment, the apparatus further includes an execution unit, and the prefetch processing unit is further configured to prefetch the prefetch address into the first-level cache memory after the execution unit finishes executing the first instruction group.
[0020] In one embodiment, the prefetch address information is an absolute address of the prefetched instruction or a relative address indicating the absolute address of the prefetched instruction.
[0021] In one embodiment, the apparatus further includes an address translation unit, the address translation unit includes an instruction TLB unit, and the prefetch processing unit is further configured to obtain the virtual address of the prefetched instruction based on the prefetch address information and send the virtual address to the instruction TLB unit, where: in response to the prefetch address information being the absolute address of the prefetched instruction, sending the absolute address as the virtual address of the prefetched instruction to the instruction TLB unit; or in response to the prefetch address information being a relative address indicating the absolute address of the prefetched instruction, adding the virtual address of the prefetch instruction to the relative address to obtain the virtual address of the prefetched instruction, and sending the virtual address of the prefetched instruction to the instruction TLB unit.
[0022] In one embodiment, the address translation unit further includes a TLB address miss status tracking register, and the instruction TLB unit is configured to convert the virtual address of the prefetched instruction into a physical address and send the physical address to the first-level cache memory, where: in response to the virtual address of the prefetched instruction existing in the instruction TLB unit, obtaining the physical address corresponding to the virtual address of the prefetched instruction; in response to the virtual address of the prefetched instruction not existing in the instruction TLB unit, sending an address translation request to the TLB address miss status tracking register, obtaining the physical address corresponding to the virtual address of the prefetched instruction from the page table unit based on the address translation request, and returning the obtained physical address to the instruction TLB unit; and in response to the virtual address of the prefetched instruction not existing in the page table unit, ending the prefetch operation.
[0023] In one embodiment, the cache memory unit further includes an instruction address miss status tracking register, and the cache memory unit is further configured to determine whether the instruction data corresponding to the physical address is in the first-level cache memory, where: in response to the instruction data corresponding to the physical address being in the first-level cache memory, ending the prefetch operation; in response to the instruction data corresponding to the physical address not being in the first-level cache memory, obtaining the instruction data corresponding to the physical address from the lower-level cache memory or the memory via the instruction address miss status tracking register, and returning the prefetched instruction data to the first-level cache memory, and ending the prefetch operation.
[0024] At least one aspect of the present disclosure also provides a device including a processor and a non-transitory memory having instructions thereon, where when the instructions are executed by the processor, the processor is caused to implement any one of the above methods.
[0025] At least one aspect of the present disclosure also provides a computer-readable storage medium storing computer-readable instructions, the computer-readable instructions including program code for performing any of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0027] Figure 1 is a schematic diagram showing a processor architecture;
[0028] Figure 2 is showing Figure 1 a flowchart of the processor architecture reading and executing instruction data;
[0029] Figure 3 is a schematic diagram of a processor architecture provided according to at least one embodiment of the present disclosure;
[0030] Figure 4 is a flowchart of a prefetch operation for instructions provided according to at least one embodiment of the present disclosure;
[0031] Figure 5 is a flowchart of a further operation example of a prefetch method provided according to at least one embodiment of the present disclosure;
[0032] Figure 6 is a schematic diagram of an information processing device provided according to at least one embodiment of the present disclosure;
[0033] Figure 7 is a schematic diagram of a device provided according to at least one embodiment of the present disclosure;
[0034] Figure 8 is a schematic diagram of a computer-readable storage medium provided according to at least one embodiment of the present disclosure DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0036] Unless otherwise defined, technical or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second", and similar terms used in the specification and claims of this disclosure do not denote any order, quantity, or importance, but are merely used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a limitation of quantity, but rather indicate the presence of at least one. Similarly, words such as "comprising" or "including" mean that the elements or items appearing before the word encompass the elements or items listed after the word and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" do not necessarily refer to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0037] Explanations of terms that may be involved in at least one embodiment of this disclosure are as follows.
[0038] The introduction to instruction prefetch is as follows. In a CPU architecture, program instructions are stored in memory (e.g., DRAM). The operating frequency of the CPU core is much higher than that of the memory. Therefore, it takes hundreds of CPU core clock cycles to fetch instructions from the memory, which often causes the CPU core to idle because it cannot continue to execute relevant instructions, resulting in performance loss. In view of this, high-performance processors all use an architecture that includes multiple levels of cache memories to store recently accessed data, and prefetch the instruction codes required by the program into the cache memory in advance before the program fetches instructions, thereby improving the execution efficiency of the processor.
[0039] The introduction to address translation is as follows. The operating system often supports multiple processes running simultaneously. To simplify multi-process management and enhance security, applications use a complete virtual address space. For example, an application with 32-bit addressing has up to 2 32 = 4GB of virtual address space available for use. When a program is running, these virtual addresses are mapped to multiple memory pages, and each page has its own physical storage address. When an application accesses instructions and data, it must first translate the virtual addresses of the instructions and data into physical addresses, and check whether the application's access to the page is legal, and then obtain the corresponding data from the cache memory or memory and transfer it to the CPU core. The process of converting from a virtual address to a physical address is called address translation.
[0040] The introduction of the Table Lookaside Buffer (TLB) is as follows. The mapping relationship from virtual address to physical address is stored in the page table in memory. Accessing these page tables in memory may also take hundreds of clock cycles. To reduce these memory accesses, the CPU core uses multi-level caches internally to store the most recently used mappings. These caches that improve the virtual address to physical address translation speed are called the Table Lookaside Buffer.
[0041] The introduction of the Missing Status Handling Register is as follows. When a read / write request, prefetch request, or mapping relationship request is not in a certain cache and needs to be read from the next-level storage unit, the request and its corresponding attributes are saved in the Missing Status Handling Register until the data of the request is returned by the next-level cache, so as not to hinder the subsequent pipeline processing in the cache.
[0042] Figure 1 is a schematic diagram showing a processor architecture. As Figure 1 shown, the processor architecture 100 includes a Branch Predictor (BP) 101, an Instruction Table Lookaside Buffer (ITLB) 102, a Translation Work Cache (TWC) 1021, a first-level Instruction Cache (IC) 103, an Instruction Miss Status Handling Register (IMSHR) 1031, a second-level cache (L2C) 1032, a third-level cache (L3C) 1033, a Memory (MEM) 1034, a Decode Unit (DE) 104, a Dispatch Unit (DI) 105, an Execution Unit (EX) 106, a Load / Store Unit (LS) 107, a Retirement Unit (RT) 108, a Data Cache (DC) 109, and a Data Miss Status Handling Register (DC MSHR) 1091.
[0043] Figure 2 is a flowchart showing Figure 1 how the processor architecture reads and executes instruction data. The following will refer to Figure 1 and Figure 2 to illustrate an example of the process of the processor reading and executing instruction data. In one embodiment, the process may include the following steps S01 - step S06.
[0044] Step S01: The Branch Predictor 101 sends the virtual address of the target instruction group to the Instruction Table Lookaside Buffer 102.
[0045] In one embodiment, the target instruction group may include multiple instructions, and the Branch Predictor 101 may sequentially input the predicted virtual address of each of the multiple instructions into the Instruction Table Lookaside Buffer 102.
[0046] Step S02: The instruction translation lookaside buffer unit 102 translates the virtual address of the received instruction into a corresponding physical address.
[0047] In one embodiment, if the physical address corresponding to the virtual address is not found in the instruction translation lookaside buffer unit 102 (i.e., ITLB miss), the instruction translation lookaside buffer unit 102 sends the virtual address to the page table unit 1021. Then, the page table unit 1021 translates it to obtain the corresponding physical address and returns the physical address to the instruction translation lookaside buffer unit 102.
[0048] Step S03: The instruction translation lookaside buffer unit 102 sends the translated physical address to the first-level instruction cache memory 103, and the first-level instruction cache memory 103 determines whether the corresponding instruction data exists in the first-level instruction cache memory 103 based on the physical address.
[0049] In one embodiment, if the instruction data already exists in the first-level instruction cache memory 103 (i.e., instruction cache hit (IC Hit)), the instruction data corresponding to the physical address is obtained from the first-level instruction cache memory 103.
[0050] In another embodiment, if the instruction data does not exist in the first-level instruction cache memory 103 (i.e., instruction cache miss (IC Miss)), the first-level instruction cache memory 103 can apply for a storage entry from the instruction address miss status tracking register 1031 and allocate the storage entry to the above cache miss request. The instruction address miss status tracking register 1031 requests the corresponding instruction data from the next-level cache memory, such as the second-level cache memory 1032 (L2 Cache), based on the storage entry. When the second-level cache memory 1032 obtains the requested instruction data, the second-level cache memory 1032 returns the instruction data to the first-level instruction cache memory 103 via the instruction address miss status tracking register 1031.
[0051] In another embodiment, if the requested instruction data is not stored in the second-level cache memory 1032, the second-level cache memory 1032 can obtain the requested instruction data from the next-level memory. For example, the memory located at the next level of the second-level cache memory can be the third-level cache memory 1033 (L3 Cache), the fourth-level cache memory (e.g., the last-level cache memory, Last Level Cache, not shown), or the memory 1034 (e.g., DRAM), etc. In one embodiment, the obtained instruction data can be returned to the first-level instruction cache memory 103 via the above memory or multi-level cache memory.
[0052] Step S04: The first-level instruction cache memory 103 sends the instruction data obtained as above to the decoding unit 104 for decoding operations, thereby obtaining the corresponding instructions.
[0053] Step S05: The decoding unit 104 sends the decoded instructions to the distribution unit 105, and the distribution unit 105 sends the instructions to the backend execution unit 106 or the memory access unit 107 for execution and storage operations based on the different types of the instructions.
[0054] In one embodiment, when the memory access unit 107 cannot obtain the data required for execution from the data memory 109 during the execution of an instruction (i.e., data miss), it can apply to the data address miss status tracking register 1091 for a storage entry and allocate this storage entry to the above data miss request. The data address miss status tracking register 1091 requests the corresponding data from the next-level cache memory, for example, the second-level cache memory 1032 (L2 Cache) based on this storage entry. When the second-level cache memory 1032 obtains the requested data, the second-level cache memory 1032 returns this data to the data memory 109 via the data address miss status tracking register 1091.
[0055] Similarly, if the requested data is not stored in the second-level cache memory 1032, the second-level cache memory 1032 can obtain the requested instruction data from the multi-level cache memory or the memory 1034 located below the second-level cache memory. In one embodiment, the obtained data can be returned to the data memory 109 via the above memory or multi-level cache memory for execution or for the memory access unit to execute.
[0056] Step S06: When an instruction is executed, it will be sent to the retirement unit 108, and the retirement unit is used to retire the executed instruction, that is, it indicates that the micro-instruction has been actually executed and completed.
[0057] In one embodiment, the target instruction stream that has undergone the operations of steps S01 - S06 as above may include multiple instruction groups and their storage addresses (fetch addresses) in the following format: the first instruction group Function0 that implements the first operation (such as a function, a loop, etc.), the second instruction group Function1 that implements the second operation, the third instruction group Function2 that implements the third operation...
[0058]
[0059]
[0060] In one embodiment, the first instruction group Function0 that implements the first operation is located in the memory line starting at the address 0x80000, and the first instruction group Function0 that implements the first operation may include one or more instructions, and each instruction may indicate the instruction content of different operations, such as arithmetic operations, fetch instruction operations, etc.
[0061] In one embodiment, the branch predictor 101 may send the virtual address corresponding to the first instruction group that implements the first operation, starting from the address 0x80000, to the instruction translation lookaside buffer unit 102 to start the processing steps of the above-mentioned step S01 until the first instruction group that implements the first operation is retired in step S06 after execution.
[0062] For example, during the execution of the first instruction group Function0 that implements the first operation, a second operation is called, so it will jump to the second instruction group Function1 that implements the second operation, that is, the virtual address corresponding to the second instruction group Function1 that implements the second operation, starting from the address 0x18000000, is sent to the instruction translation lookaside buffer unit 102 to start the processing steps of the above-mentioned step S01 until the second instruction group Function1 is retired in step S06 after execution.
[0063] However, during the execution of the second instruction group Function1 that implements the second operation, a third operation is called, so it will jump to the third instruction group Function2 that implements the third operation, that is, the virtual address corresponding to the third instruction group Function2, starting from the address 0x700000000, is sent to the instruction translation lookaside buffer unit 102 to start the processing steps of the above-mentioned step S01 until the third instruction group Function2 is retired in step S06 after execution.
[0064] Therefore, in the above process, the first operation calls the second operation, and the second operation calls the third operation, performing nested calls. When the lower-level operation is completed, it will return to the higher-level operation. Repeating the above steps can execute multiple instructions.
[0065] In the instruction fetching example described above, for instance, before the processor decodes the instruction "call Function1" in the first instruction group Function0, it doesn't know what the instruction indicates. Thus, when the instruction "call Function1" is executed and the jump is made to the second instruction group Function1, it's possible that the instruction data corresponding to the second instruction group Function1 is not in the first-level instruction cache memory 103, resulting in a cache miss (IC miss). This forces the processor to retrieve data from the second-level cache memory 1032, the third-level cache memory 1033, or even the memory 1034. However, in a multi-level cache memory architecture, the first-level cache memory has the fastest access speed but the smallest capacity, the last-level (e.g., the third-level) cache memory has the largest capacity but the slowest access speed, and the second-level cache memory has an access speed and capacity between those of the first-level cache memory and the last-level cache memory. Therefore, in the case of a first-level cache memory miss, it's necessary to wait for the data to be retrieved from the slower-accessing lower-level memories (e.g., 1032, 1033) or even from the memory 1034 and then returned to the first-level cache memory 103, which may take a dozen or even hundreds of clock cycles, blocking the processor's pipeline and thus affecting the processor's execution efficiency.
[0066] The inventors of the present disclosure noticed that to address the performance loss of the processor caused by the above problems, the processor can prefetch the instruction data required by the program into the first-level cache memory in advance before fetching the program instructions. This can effectively reduce the occurrence of cache misses, thereby reducing the number of clock cycles the CPU core waits for data and enhancing the overall performance of the processor.
[0067] At least one embodiment of the present disclosure provides a method and a device for prefetching instructions. The prefetching method at least includes: receiving a first instruction; decoding the first instruction to determine that the first instruction is a prefetch instruction and obtaining the prefetch address information in the first instruction; and performing a prefetch operation on the prefetch address based on the prefetch address information.
[0068] For example, the method and device for prefetching instructions provided by at least one embodiment of the present disclosure can effectively prefetch the instructions to be executed through the cooperation of software and hardware, reducing the probability of cache misses in the cache memory and TLB cache misses in the TLB unit, thereby improving the efficiency of the processor operation and enhancing the processing performance of the overall system.
[0069] The prefetch instruction method provided according to at least one embodiment of the present disclosure will be described below by way of several examples and embodiments in a non-limiting manner. As described below, different features in these specific examples and embodiments can be combined with each other without conflict, so as to obtain new examples and embodiments, and these new examples and embodiments also fall within the scope of protection of the present disclosure.
[0070] An example of the instruction stream provided according to the present disclosure may include multiple instruction groups such as a fourth instruction group Function3 that sequentially executes to implement a fourth operation (such as a function, a loop, etc.), a fifth instruction group Function4 that implements a fifth operation, and a sixth instruction group Function5 that implements a sixth operation. Among them, the fourth instruction group Function3 that implements the fourth operation is located at the position starting from address 0x80000, the fifth instruction group Function4 that implements the fifth operation is located at the position starting from address 0x18000000, and the sixth instruction group Function5 that implements the sixth operation is located at the position starting from address 0x700000000, as follows:
[0071]
[0072] In one embodiment, the fourth instruction group Function3 that implements the fourth operation may include a prefetch instruction, and the format of the prefetch instruction is: iprefetch mem8. For example, the fourth instruction group Function3 may include a prefetch instruction iprefetch 0x18000000 that indicates the address of the fifth instruction group Function4. In one embodiment, the prefetch instruction may include the type information of the instruction and the prefetch address information of the fifth instruction group. In this example, the instruction name iprefetch in the prefetch instruction iprefetch mem8 indicates that the type of this instruction is a prefetch instruction, and the parameter mem8 indicates the prefetch address information of the instruction to be prefetched.
[0073] In one embodiment, the address information of the prefetched instruction may include the absolute address of the prefetched instruction or a relative address indicating the absolute address of the prefetched instruction. That is, the parameter mem8 may represent the absolute address (prefetch address) or an offset value indicating the absolute address (prefetch address). For example, in the above example, the prefetched instruction iprefetch 0x18000000 inserted in the fourth instruction group Function3 implementing the fourth operation indicates the absolute address of the fifth instruction group Function4 to be executed after the fourth instruction group, that is, the starting position of the virtual address: 0x18000000. Thus, when the processor starts to execute the fourth instruction group Function3 implementing the fourth operation, it can know the position of the virtual address of the fifth instruction group to be executed after the execution of the fourth instruction group Function3. Therefore, the instruction data of the fifth instruction group implementing the fifth operation can be prefetched into the first-level cache memory in advance to reduce the possibility of cache misses.
[0074] In one embodiment, the parameter mem8 may also represent the relative address of the instruction to be prefetched, such as an offset value relative to the virtual address of the prefetched instruction. In addition, in another embodiment, the parameter mem8 may also represent an offset value relative to the virtual address of other instructions. It should be understood that the present disclosure does not limit the specific format of the parameter mem8, as long as it can indicate the address position of the instruction to be prefetched.
[0075] In one embodiment, the position of the prefetched instruction iprefetch in the instruction group Function3 implementing the fourth operation may be determined based on the relationship between the size of the instruction group Function3 implementing the fourth operation and the capacity of the first-level cache memory.
[0076] For example, in the instruction example shown above, if the size of the fourth instruction group implementing the fourth operation is smaller than the capacity of the first-level cache memory, the prefetched instruction iprefetch may be inserted at a position immediately after the starting virtual address of the fourth instruction group Function3 implementing the fourth operation, so that when starting to execute the fourth instruction group implementing the fourth operation, the prefetched operation for the fifth instruction group implementing the fifth operation is started, leaving sufficient clock cycles for the prefetched operation of the fifth instruction group implementing the fifth operation.
[0077] However, the insertion position of the prefetch instruction iprefetch in the above example is only exemplary, and the prefetch instruction iprefetch can also be inserted into other positions in the fourth instruction group. For example, in one embodiment, if the size of the fourth instruction group is larger than the capacity of the first-level cache memory, the insertion position of the prefetch instruction iprefetch can be determined according to the maximum number of clock cycles required to prefetch the fifth instruction group. For example, if it takes 200 clock cycles to fetch the instruction data of the fifth instruction group from the memory and return it to the first-level cache memory, the prefetch instruction can be inserted into the position 200 instructions before the last instruction in Function4 of the fourth instruction group, so that before the fourth instruction group is executed, the fifth instruction group can be prefetched into the first-level cache memory, and at the same time, it can be ensured that the processor does not prefetch the fifth instruction group for implementing the fifth operation into the first-level cache memory too early, and when the remaining part of the fourth instruction group is executed subsequently, the instruction data corresponding to the fifth instruction group is removed from the first-level cache memory due to being replaced.
[0078] It should be understood that the above insertion position is only exemplary, and the prefetch instruction can be inserted into any position in the fourth instruction group via the compiler. Because even if the instruction data of the fifth instruction group is sent to the first-level cache memory after the fourth instruction group is executed, since the operation of prefetching the fifth instruction group has entered the pipeline before the operation of fetching the fifth instruction group, when the fifth instruction group is actually executed, the latency of obtaining the instruction data of the fifth instruction group from the cache memory or the memory will also be reduced, so that the performance can also be improved to a certain extent.
[0079] Figure 3 It is a schematic diagram of a processor architecture 300 provided according to at least one embodiment of the present disclosure. The processor architecture 300 provided according to an embodiment of the present disclosure is applicable to processing the above instruction structure, so as to be able to implement an instruction prefetch operation to reduce the performance loss caused by cache misses.
[0080] As Figure 3 shown, Figure 3The processor architecture 300 may include a branch predictor (BP) 301, an instruction translation lookaside buffer unit (ITLB) 302, a translation walk cache (TWC) 3021, an ITLB miss status handling register (ITLB MSHR) 3022, a first-level instruction cache memory (IC) 303, an instruction miss status handling register (IMSHR) 3031, a second-level cache memory (L2C) 3032, a third-level cache memory (L3C) 3033, a memory (MEM) 3034, a decoding unit (DE) 304, a distribution unit (DI) 305, an execution unit (EX) 306, a load / store unit (LS) 307, a retirement unit (RT) 308, a data cache (DC) 309, a data cache miss status handling register (DC MSHR) 3091, and a prefetch processing unit (PREF) 310.
[0081] Figure 4 is a flowchart of a prefetch operation for instructions provided according to at least one embodiment of the present disclosure. The following will refer to Figure 3 and Figure 4 to illustrate the prefetch operation provided according to at least one embodiment of the present disclosure.
[0082] Referring to Figure 4 , the prefetch operation may include the following steps S11 - step S16.
[0083] Step S11: The branch predictor 301 sends the virtual address of the target instruction group to the instruction translation lookaside buffer unit 302.
[0084] For example, in one embodiment, the branch predictor 301 may sequentially send the fourth instruction group for implementing the fourth operation, the fifth instruction group for implementing the fifth operation, and the sixth instruction group for implementing the sixth operation to the instruction translation lookaside buffer unit 302. In one embodiment, each of the fourth instruction group, the fifth instruction group, and the sixth instruction group may include one or more instructions. For example, it may include multiple instructions such as prefetch instructions, arithmetic instructions, call instructions, or other instructions. In the following example, the fourth instruction group Function3 for implementing the fourth operation will be used as the target instruction group to illustrate the prefetch operation process.
[0085] In one embodiment, the branch predictor 301 may sequentially send the virtual addresses of multiple instructions starting from 0x80000 of the fourth instruction group Function3 to the instruction translation lookaside buffer unit 302.
[0086] Step S12: The instruction translation lookaside buffer unit 302 may translate the virtual addresses of the above multiple instructions into corresponding physical addresses.
[0087] In one embodiment, the fourth instruction group Function3 may include a prefetch instruction iprefetch0x18000000, a call instruction call Function4, and a plurality of other instructions such as arithmetic instructions therebetween. The instruction translation lookaside buffer unit 302 may sequentially translate the virtual addresses of the plurality of instructions in the fourth instruction group Function3 into corresponding physical addresses.
[0088] For example, in one embodiment, if the requested virtual address is in the instruction translation lookaside buffer unit 302, i.e., the instruction translation lookaside buffer cache hit (ITLB hit), the instruction translation lookaside buffer unit 302 may translate the virtual addresses of the plurality of instructions in the fourth instruction group into physical addresses through the translation lookaside buffer unit and send the physical addresses to the first-level cache memory 303.
[0089] For example, in another embodiment, if the requested virtual address is not in the instruction translation lookaside buffer unit 302, i.e., the instruction translation lookaside buffer cache miss (ITLB miss), the instruction translation lookaside buffer unit 302 applies for a storage entry from the translation lookaside buffer address miss status tracking register 3022 and assigns the storage entry to the above address translation request. The translation lookaside buffer address miss status tracking register 3022 may send the address translation request to the page table unit 3021 for address translation. Since the address translation request is cached in the translation lookaside buffer address miss status tracking register 3022, it does not prevent the instruction translation lookaside buffer unit 302 from processing subsequent instructions, thereby avoiding the latency caused by the blocking of the processing pipeline.
[0090] Step S13: The instruction translation lookaside buffer unit 302 sends the translated physical addresses to the first-level instruction cache memory 303, and the first-level instruction cache memory 303 determines whether the instruction data corresponding to the physical addresses exists in the first-level instruction cache memory 303. For example, if the requested instruction data exists in the first-level cache memory 303, the instruction data corresponding to the physical address is directly obtained from the first-level cache memory 303. In another embodiment, if the requested instruction data does not exist in the first-level cache memory 303, the instruction data corresponding to the physical address is obtained from a lower-level memory (e.g., the second-level cache memory 3032, the third-level cache memory 3033, or the memory 3034) and returned to the first-level instruction cache memory 303.
[0091] Step S14: The first-level instruction cache memory sends the instruction data corresponding to the physical address obtained as above to the decoding unit 304 for decoding operation to obtain the corresponding instruction.
[0092] Step S15: The distribution unit 305 determines the type of the instruction and decides whether to send the instruction to the backend execution unit 306 or the prefetch processing unit 310 according to different instruction types.
[0093] In one embodiment, the types of instructions may include instructions of a first type and instructions of a second type. Among them, the instructions of the first type may be prefetch instructions, and the instructions of the second type may be other instructions, such as call instructions or arithmetic instructions, etc.
[0094] In one embodiment, in response to determining that the type of the instruction is an instruction of the first type, for example, the distribution unit 305 determines that the instruction in the fourth instruction group is a prefetch instruction, the process proceeds to step S16, that is, the distribution unit 305 sends the prefetch instruction to the prefetch processing unit 310 for prefetch operation. The prefetch operation in step S16 will be further described below in conjunction with Figure 5 The prefetch operation in step S16 is further described.
[0095] In one embodiment, in response to determining that the type of the instruction is an instruction of the second type, for example, the distribution unit 305 determines that the instruction in the fourth instruction group is an arithmetic instruction or other instruction, the process proceeds to step S17, that is, the distribution unit 305 sends the instruction to the execution unit 306 or the memory access unit 307 for execution and storage. In step S17, when the type of the instruction is a memory access instruction, if the requested data is not stored in the data memory 309, the data memory 309 will request data from the second-level cache memory 3032, the third-level cache memory 3033, or the memory 3034 via the data address miss status tracking register 3091 and return it to the data memory 309.
[0096] Step S18: When the distribution unit 305 determines that the type of the instruction is an instruction of the first type, that is, a prefetch instruction, the prefetch instruction is directly marked as executed and completed in the retirement unit 308. So that the prefetch instruction will not enter the execution unit to cause an execution error. In addition, when other instructions such as arithmetic instructions are executed, they will also be sent to the retirement unit 308 to mark that the instruction has been actually executed and completed.
[0097] Figure 5 It is a flowchart of a further operation example of the prefetch method provided according to at least one embodiment of the present disclosure. Figure 5 Shows Figure 4 A further operation example of the instruction prefetch operation in step S16 in Figure 5 As shown, the prefetch operation may include the following steps S21 - S25:
[0098] Step S21: The prefetch processing unit 310 obtains the prefetch instruction prefetch address information based on the prefetch instruction. For example, the virtual address, and sends the obtained virtual address to the instruction translation lookaside buffer unit 302.
[0099] In one embodiment, the prefetch processing unit 310 may first obtain the parameter mem8 in the prefetch instruction iprefetch. For example, in one embodiment, when the parameter mem8 indicates the absolute address of the instruction to be prefetched, the prefetch processing unit 310 may directly send the absolute address in the parameter mem8 to the instruction translation lookaside buffer unit 302 as the virtual address.
[0100] In another embodiment, when the parameter mem8 indicates the relative address of the instruction to be prefetched. For example, the parameter mem8 may be an offset value relative to the virtual address of the prefetch instruction. The prefetch processing unit 310 may perform an addition operation on the virtual address of the prefetch instruction (e.g., the virtual address of the iprefetch instruction) and the value of the parameter mem8 to calculate a new virtual address, and send the calculated virtual address to the instruction translation lookaside buffer unit 302.
[0101] In another embodiment, when the parameter mem8 may also be an offset value relative to the virtual address of other instructions, the virtual address of this instruction may be added to the value of the parameter mem8 at this time to calculate a new virtual address. It should be understood that those skilled in the art may perform corresponding other transformations or adjustments on the offset value of the relative address, not limited to the above examples.
[0102] Step S22: The instruction translation lookaside buffer unit 302 generates an access request for the instruction translation lookaside buffer based on the newly generated virtual address to obtain the corresponding physical address, and sends the physical address to the first-level cache memory 303.
[0103] For example, in one embodiment, if the requested virtual address is in the instruction translation lookaside buffer unit 302, that is, the translation lookaside buffer cache hits, the instruction translation lookaside buffer unit 302 may translate the newly generated virtual address into a physical address through the mapping relationship stored in the translation lookaside buffer unit, and send the physical address to the first-level cache memory 303.
[0104] For example, in one embodiment, if the requested virtual address is not in the instruction translation lookaside buffer unit 302, i.e., the translation lookaside buffer cache miss occurs, the instruction translation lookaside buffer unit 302 applies for a storage entry from the translation lookaside buffer address miss status tracking register 3022 and assigns the storage entry to the above address translation request. The translation lookaside buffer address miss status tracking register 3022 may send the address translation request to the page table unit 3021 for address translation. Since the address translation request is cached in the translation lookaside buffer address miss status tracking register 3022, it does not prevent the instruction translation lookaside buffer unit 302 from processing subsequent instructions, thereby avoiding the latency caused by the pipeline blockage.
[0105] In one embodiment, if the page table unit 3021 can translate the virtual address to obtain the physical address, it returns the address translation result to the translation lookaside buffer address miss status tracking register 3022, and the translation lookaside buffer address miss status tracking register 3022 sends the physical address to the first-level cache memory 303 when the first-level cache memory 303 is not operating.
[0106] In another embodiment, if the page table unit 3021 cannot obtain the physical address based on the virtual address, for example, a translation error (fault) occurs, the error is reported to the translation lookaside buffer address miss status tracking register 3022, and at the same time, the prefetch request is abandoned.
[0107] Step S23: The first-level cache memory 303 checks whether the instruction data corresponding to the physical address is in the first-level cache memory 303 based on the physical address.
[0108] In one embodiment, if the instruction data corresponding to the physical address is already in the first-level cache memory 303 (i.e., cache hit), it means that the code required by the application program in the future already exists in the first-level cache memory 303, and the request processing of the instruction prefetch is completed, and the process proceeds to step S25; if the instruction data corresponding to the physical address is not in the first-level cache memory 303 (i.e., cache miss), the process proceeds to step S24.
[0109] Step S24: If the instruction data does not exist in the first-level instruction cache memory 303, the first-level instruction cache memory 303 applies for a storage entry from the instruction address miss status tracking register 3031 and allocates the storage entry to the above cache miss request. The instruction address miss status tracking register 3031 requests the corresponding instruction data from the next-level cache memory based on the storage entry, for example, the second-level cache memory 3032. The second-level cache memory 3032 obtains the requested instruction data and returns the instruction data to the first-level instruction cache memory 303 via the instruction address miss status tracking register 3031. The request processing for instruction prefetch is completed, and the process proceeds to step S25.
[0110] In one embodiment, if the requested instruction data is not stored in the second-level cache memory 3032, the second-level cache memory 3032 may obtain the requested instruction data from a memory at the next level below the second-level cache memory. For example, the memory at the next level below the second-level cache memory may be a third-level cache memory 3033 or a memory 3034, etc. In one embodiment, the obtained instruction data may be returned to the first-level instruction cache memory 303 via the above memory or multi-level cache memories. The request processing for instruction prefetch is completed, and the process proceeds to step S25.
[0111] Step S25: The prefetch request processing is completed.
[0112] The prefetch operation performed as above may, for example, when executing the fourth instruction group implementing the fourth operation, prefetch the address of the fifth instruction group to be fetched after the fourth instruction group into the cache memory of the processor in advance, so as to reduce the probability of cache memory miss when executing the fifth instruction group, and further improve the execution efficiency of the processor.
[0113] Figure 6 It is a schematic diagram of an information processing device provided according to at least one embodiment of the present disclosure.
[0114] As Figure 6 shown, the information processing device 600 may at least include an address translation unit 601, a cache memory unit 602, a decoding unit 603, a distribution unit 604, a prefetch processing unit 605, an execution unit 606, and a retirement unit 607. Among them, the cache memory unit may further include a first-level cache memory, an instruction address miss status tracking register, and a second-level cache memory. In addition, the address translation unit 601 may further include an instruction translation lookaside buffer unit and a translation lookaside buffer address miss status tracking register. For example, the processing device 600 may be a single-core central processing unit or a certain processing core (CPU core) of a multi-core central processing unit, and the embodiments of the present disclosure are not limited thereto.
[0115] In one embodiment, an address translation unit 601 is configured to translate a virtual address of a received first instruction into a physical address; a cache memory unit 602 is configured to receive the physical address of the first instruction; a decoding unit 603 is configured to decode the first instruction; a distribution unit 604 is configured to determine whether the decoded first instruction is a prefetch instruction, and when it is determined that the first instruction is a prefetch instruction, send the prefetch instruction to a prefetch processing unit 605; and the prefetch processing unit 605 is configured to obtain prefetch address information in the first instruction and perform a prefetch operation on the prefetch address based on the prefetch address information; an execution unit 606 is configured to execute a corresponding first instruction or second instruction; and a retirement unit 607 is configured to retire an instruction that has been executed, that is, indicate that the instruction has been actually executed and completed.
[0116] Figure 7 is a schematic diagram of a device provided according to at least one embodiment of the present disclosure. As Figure 7 shown, the device 700 includes a processor 702 and a non-transitory memory 703. Among them, instructions 701 are stored on the non-transitory memory 703. In one embodiment, when the processor 702 executes the instructions 701, one or more steps in the method for prefetch instructions described above can be implemented.
[0117] Figure 8 is a schematic diagram of a computer-readable storage medium provided according to at least one embodiment of the present disclosure. As Figure 8 shown, the computer-readable storage medium 800 non-transitorily stores computer-readable instructions 801. For example, when the computer-readable instructions 801 are executed by a computer, one or more steps in the method for prefetch instructions described above can be executed.
[0118] For example, the computer-readable storage medium 800 can be applied to the above-mentioned device 700. For example, the computer-readable storage medium 800 can be Figure 7 the non-transitory memory 703 in the device 700 shown.
[0119] Each operation of the method described above can be performed by any suitable means capable of performing the corresponding function. The means can include various hardware and / or software components and / or modules, including but not limited to hardware circuits, application-specific integrated circuits (ASICs), or processors.
[0120] The various illustrative logical blocks, modules, and circuits described herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0121] The steps of a method or algorithm described in connection with the present disclosure may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in any form of tangible storage medium. Some examples of storage media that may be used include random access memory (RAM), read only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, etc. The storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The software module may be a single instruction or many instructions and may be distributed over several different code segments, different programs, and across multiple storage media.
[0122] Accordingly, a computer program product may perform the operations presented herein. For example, such a computer program product may be a computer readable tangible medium having tangible storage (and / or encoding) thereon of instructions executable by one or more processors to perform the operations described herein. The computer program product may include packaging material.
[0123] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the invention. Thus, the invention is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for prefetching instructions, comprising: Receiving a first instruction; Decoding the first instruction to determine that the first instruction is a prefetch instruction and obtaining prefetch address information in the first instruction; Performing a prefetch operation on the prefetch address based on the prefetch address information; Wherein, the first instruction is included in a first instruction group, and the method further includes: the position where the first instruction is inserted in the first instruction group is determined based on the relationship between the size of the first instruction group and the capacity of the first-level cache memory; Wherein, the position where the first instruction is inserted in the first instruction group is determined based on the relationship between the size of the first instruction group and the capacity of the first-level cache memory, including: In the case where the size of the first instruction group is smaller than the capacity of the first-level cache memory, inserting the first instruction at the starting position of the first instruction group; and In the case where the size of the first instruction group is larger than the capacity of the first-level cache memory, determining the position where the first instruction is inserted in the first instruction group based on the time required for the prefetch operation.
2. The method according to claim 1, further comprising: In response to the first instruction being a prefetch instruction, marking the first instruction as having been executed and completed in the retirement unit.
3. The method according to claim 1, wherein, Performing a prefetch operation on the prefetch address based on the prefetch address information further includes: prefetching the prefetch address into the first-level cache memory before the first instruction group is executed.
4. The method according to claim 1, wherein Performing a prefetch operation on the prefetch address based on the prefetch address information further includes: prefetching the prefetch address into the first-level cache memory after the first instruction group is executed.
5. The method according to claim 1, wherein, The prefetch address information is the absolute address of the prefetched instruction or a relative address indicating the absolute address of the prefetched instruction.
6. The method according to claim 5, wherein, Performing a prefetch operation on the prefetch address based on the prefetch address information further includes: obtaining the virtual address of the prefetched instruction based on the prefetch address information and sending the virtual address to the instruction translation lookaside buffer unit, wherein: In response to the prefetch address information being the absolute address of the prefetched instruction, sending the absolute address as the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit; or In response to the prefetch address information being a relative address indicating the absolute address of the prefetched instruction, adding the virtual address of the prefetch instruction to the relative address to obtain the virtual address of the prefetched instruction and sending the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit.
7. The method according to claim 6, wherein based on the prefetch address information, the prefetch operation on the prefetch address further includes: After sending the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit, converting the virtual address of the prefetched instruction into a physical address and sending the physical address to the first-level cache memory, wherein: In response to the virtual address of the prefetched instruction existing in the instruction translation lookaside buffer unit, obtaining the physical address corresponding to the virtual address of the prefetched instruction; In response to the virtual address of the pre-fetched instruction not existing in the instruction translation lookaside buffer unit, send an address translation request to the translation lookaside buffer address miss status tracking register, obtain the physical address corresponding to the virtual address of the pre-fetched instruction from the page table unit based on the address translation request, and return the obtained physical address to the instruction translation lookaside buffer unit; and In response to the virtual address of the pre-fetched instruction not existing in the page table unit, end the pre-fetch operation.
8. The method according to claim 7, wherein based on the prefetch address information, the prefetch operation on the prefetch address further includes: Determine whether the instruction data corresponding to the physical address is in the level-1 cache memory, where:[[]] In response to the instruction data corresponding to the physical address being in the level-1 cache memory, end the pre-fetch operation; In response to the instruction data corresponding to the physical address not being in the level-1 cache memory, obtain the instruction data corresponding to the physical address from the lower-level cache memory or memory via the instruction address miss status tracking register, and return the pre-fetched instruction data to the level-1 cache memory, and end the pre-fetch operation.
9. An information processing apparatus, comprising: A cache memory unit configured to receive a first instruction; A decoding unit configured to decode the first instruction; A distribution unit configured to determine that the first instruction is a pre-fetch instruction and send the pre-fetch instruction to a pre-fetch processing unit; And A pre-fetch processing unit configured to obtain pre-fetch address information in the first instruction and perform a pre-fetch operation on a pre-fetch address based on the pre-fetch address information; wherein the cache memory unit further includes a level-1 cache memory, and the first instruction is included in a first instruction group, and the position where the first instruction is inserted in the first instruction group is determined based on the relationship between the size of the first instruction group and the capacity of the level-1 cache memory; wherein the position where the first instruction is inserted in the first instruction group is determined based on the relationship between the size of the first instruction group and the capacity of the level-1 cache memory, including: In the case where the size of the first instruction group is smaller than the capacity of the level-1 cache memory, insert the first instruction at the start position of the first instruction group; and In the case where the size of the first instruction group is larger than the capacity of the level-1 cache memory, determine the position where the first instruction is inserted in the first instruction group based on the time required for the pre-fetch operation.
10. The apparatus according to claim 9, further comprising: A retirement unit configured to mark the first instruction as having been executed and completed in the retirement unit in response to the first instruction being a pre-fetch instruction.
11. The apparatus according to claim 9, further comprising an execution unit, and the pre-fetch processing unit is further configured to pre-fetch the pre-fetch address to the level-1 cache memory before the execution unit finishes executing the first instruction group.
12. The apparatus according to claim 9, further comprising an execution unit, and the pre-fetch processing unit is further configured to pre-fetch the pre-fetch address to the level-1 cache memory after the execution unit finishes executing the first instruction group.
13. The device according to claim 12, wherein, The prefetch address information is the absolute address of the prefetched instruction or a relative address indicating the absolute address of the prefetched instruction.
14. The apparatus according to claim 13, further comprising an address translation unit, the address translation unit including an instruction translation lookaside buffer unit, and the prefetch processing unit is further configured to obtain the virtual address of the prefetched instruction based on the prefetch address information and send the virtual address to the instruction translation lookaside buffer unit, wherein: in response to the prefetch address information being the absolute address of the prefetched instruction, sending the absolute address as the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit; or in response to the prefetch address information being a relative address indicating the absolute address of the prefetched instruction, adding the virtual address of the prefetch instruction to the relative address to obtain the virtual address of the prefetched instruction, and sending the virtual address of the prefetched instruction to the instruction translation lookaside buffer unit.
15. The apparatus according to claim 14, the address translation unit further includes a translation lookaside buffer address miss status tracking register, and the instruction translation lookaside buffer unit is configured to convert the virtual address of the prefetched instruction into a physical address and send the physical address to a first-level cache memory, wherein: in response to the virtual address of the prefetched instruction existing in the instruction translation lookaside buffer unit, obtaining the physical address corresponding to the virtual address of the prefetched instruction; in response to the virtual address of the prefetched instruction not existing in the instruction translation lookaside buffer unit, sending an address translation request to the translation lookaside buffer address miss status tracking register, obtaining the physical address corresponding to the virtual address of the prefetched instruction from a page table unit based on the address translation request, and returning the obtained physical address to the instruction translation lookaside buffer unit; and in response to the virtual address of the prefetched instruction not existing in the page table unit, ending the prefetch operation.
16. The apparatus according to claim 15, wherein the cache memory unit further includes an instruction address miss status tracking register, and the cache memory unit is further configured to determine whether the instruction data corresponding to the physical address is in the first-level cache memory, wherein: in response to the instruction data corresponding to the physical address being in the first-level cache memory, ending the prefetch operation; in response to the instruction data corresponding to the physical address not being in the first-level cache memory, obtaining the instruction data corresponding to the physical address from a lower-level cache memory or memory via the instruction address miss status tracking register, and returning the prefetched instruction data to the first-level cache memory, ending the prefetch operation.
17. An apparatus, comprising: a processor; and a non-transitory memory storing executable instructions wherein, when the executable instructions are executed by the processor, to perform the method according to any one of claims 1-8.
18. A computer-readable storage medium storing computer-readable instructions, the computer-readable instructions including program code for performing the method according to any one of claims 1-8.
Citation Information
Patent Citations
A code pre-fetch instruction
CN113568663A