Prefetching Method and Device, Prefetching Training Method and Device, Storage Medium

By introducing pointer value reading instruction cache and read architecture register tables into the CPU core, the problem of pointer data prefetching in the existing technology is solved, and more efficient pointer data reading is achieved, and system performance is improved.

CN115934170BActive Publication Date: 2025-07-11HYGON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211726101.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2025-07-11
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

Existing data prefetchers are difficult to effectively prefetch pointer data, resulting in long delays when the CPU core is reading pointer data, affecting system performance.

Method used

By introducing pointer value read instruction cache (PLC) and reading architecture register tables into the CPU core, recording the information of pointer value read instruction, and calculating the pointer data address based on this information, realizing prefetching of pointer data.

Benefits of technology

It reduces the delay of CPU core waiting for reading pointer data, improves system performance, and improves the coverage and timeliness of pointer data prefetching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115934170B_ABST
    Figure CN115934170B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a pointer data pre-fetching method and device, a pre-fetching training method and device, and a storage medium. The pre-fetching method includes: querying a pointer value read instruction cache to hit a first data read request, wherein the pointer value read instruction cache is used to cache at least one backup pointer value read request item, and each backup pointer value read request item includes pointer data address calculation information; executing a first data read request to obtain first read data; using the first pointer data address calculation information corresponding to the first data read request in the pointer value read instruction cache and the first read data to calculate a first pointer data pre-fetch address; using the first pointer data pre-fetch address to issue a first pointer data pre-fetch request. The pre-fetching method can realize the pre-fetching of pointer data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a method and apparatus for prefetching pointer data, a method and apparatus for prefetching training of pointer data, and a storage medium. Background Art

[0002] A processor core (such as a CPU core) of a single-core processor or a multi-core processor improves instruction-level parallelism through pipelining technology. Figure 1 FIG. shows a schematic diagram of a pipeline of a processor core. The dashed line with an arrow in the figure represents a redirected instruction stream. As Figure 1 shown, the inside of the processor core includes multiple pipeline stages. For example, after the program counter from various sources is fed into the pipeline and the next program counter (PC) is selected through a multiplexer (Mux), the instruction corresponding to the program counter has to go through branch prediction, instruction fetch, instruction decode, instruction dispatch and rename, instruction execution, instruction retirement, etc. Waiting queues are set between each pipeline stage as needed, and these queues are usually first-in-first-out (FIFO) queues. For example, after the branch prediction unit, a branch prediction (BP) FIFO queue is set to store the branch prediction result; after the instruction fetch unit, an instruction cache (IC) FIFO is set to cache the fetched instructions; after the instruction decode unit, a decode (DE) FIFO is set to cache the decoded instructions; after the instruction dispatch and rename unit, a retirement (RT) FIFO is set to cache the instructions waiting for confirmation of retirement after execution. At the same time, the pipeline of the processor core also includes an instruction queue to cache the instructions waiting for the instruction execution unit to execute after instruction dispatch and rename. To support a high operating frequency, each pipeline stage may further include multiple pipeline levels (clock cycles). Although each pipeline level performs limited operations, in this way, each clock can be made the shortest, and the performance of the CPU core is improved by increasing the operating frequency of the CPU. Each pipeline level can also further improve the performance of the processor core by accommodating more instructions (i.e., superscalar technology). Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a prefetch method for pointer data, the prefetch method comprising: querying a first data read request in a pointer value read instruction cache (PLC), wherein the pointer value read instruction cache (PLC) is used to cache at least one backup pointer value read request item, each of the backup pointer value read request items including pointer data address calculation information; executing the first data read request to obtain first read data; using the first pointer data address calculation information corresponding to the first data read request in the pointer value read instruction cache (PLC) and the first read data to calculate a first pointer data prefetch address; and issuing a first pointer data prefetch request using the first pointer data prefetch address.

[0004] At least one embodiment of the present disclosure provides a pre-fetch training method for pointer data, comprising: receiving a first data read instruction, wherein the first data read instruction includes a first source register; querying and hitting the first source register in a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register item, and each candidate architecture register item includes information of a past pointer value read instruction using the corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instruction; according to the first data read instruction, obtaining first pointer data address calculation information based on a pointer value corresponding to the first source register and a destination pointer data address; in a pointer value read instruction cache (PLC), updating a first lookup pointer value read request item corresponding to the past pointer value read instruction of the first source register, wherein the pointer value read instruction cache (PLC) is used to cache at least one lookup pointer value read request item, and each lookup pointer value read request item includes pointer data address calculation information; writing the first pointer data address calculation information into the first lookup pointer value read request item to generate a pointer data pre-fetch request.

[0005] At least one embodiment of the present disclosure provides a prefetch training method for pointer data, the prefetch training method comprising: receiving a first instruction and obtaining a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register item, and each candidate architecture register item includes information of a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instruction, and is used to update a pointer value read instruction cache (PLC) for pointer data prefetch operations; in response to the first instruction being a read instruction, creating or updating a first in the destination register corresponding to the read instruction in the read architecture register table; an alternative architecture register entry, recording information of the read instruction in the first alternative architecture register entry, or, in response to the first instruction being a calculation instruction, obtaining first pointer data address calculation information (Opinfo) of the calculation instruction based on a pointer value corresponding to a source register of the calculation instruction, recording the first pointer data address calculation information (Opinfo) in a second alternative architecture register entry in the read architecture register table corresponding to the source register of the calculation instruction, creating or updating a third alternative architecture register entry corresponding to a destination register of the calculation instruction in the read architecture register table, and copying the content recorded in the second alternative architecture register entry to the third alternative architecture register entry.

[0006] At least one embodiment of the present disclosure provides a pre-fetching device for pointer data, the pre-fetching device comprising:

[0007] A query module is configured to query a pointer value read instruction cache (PLC) to hit a first data read request, wherein the pointer value read instruction cache (PLC) is used to cache at least one reserve pointer value read request item, and each reserve pointer value read request item includes pointer data address calculation information;

[0008] an execution module, configured to execute the first data read request to obtain first read data;

[0009] An address calculation module is configured to use the pointer value to read first pointer data address calculation information corresponding to the first data read request and the first read data in a command cache (PLC) to calculate a first pointer data prefetch address;

[0010] The request issuing module is configured to issue a first pointer data prefetch request using the first pointer data prefetch address.

[0011] At least one embodiment of the present disclosure further provides a pre-fetch training device for pointer data, the pre-fetch training device comprising:

[0012] A receiving module, configured to receive a first data read instruction, wherein the first data read instruction includes a first source register;

[0013] a query module configured to query and hit the first source register in a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register entry, and each candidate architecture register entry includes information of a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instruction;

[0014] an acquisition module, configured to acquire, according to the first data read instruction, first pointer data address calculation information based on the pointer value corresponding to the first source register and the destination pointer data address;

[0015] an updating module, configured to update, in a pointer value read instruction cache (PLC), a first reserve pointer value read request item corresponding to a past pointer value read instruction of the first source register, wherein the pointer value read instruction cache (PLC) is used to cache at least one reserve pointer value read request item, and each reserve pointer value read request item includes pointer data address calculation information;

[0016] A writing module is configured to write the first pointer data address calculation information into the first reference pointer value read request item to generate a pointer data prefetch request.

[0017] At least one embodiment of the present disclosure further provides a pre-fetch training device for pointer data, the pre-fetch training device comprising:

[0018] a receiving module configured to receive a first instruction and obtain a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register item, and each candidate architecture register item includes information of a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instruction, and is used to update a pointer value read instruction cache (PLC) for a pointer data prefetch operation;

[0019] a creation / update module configured to, in response to the first instruction being a read instruction, create or update a first candidate architecture register entry in the read architecture register table corresponding to the destination register of the read instruction, and record information of the read instruction in the first candidate architecture register entry, or

[0020] In response to the first instruction being a calculation instruction, obtain first pointer data address calculation information (Opinfo) of the calculation instruction based on a pointer value corresponding to a source register of the calculation instruction, record the first pointer data address calculation information (Opinfo) in a second alternative architecture register entry in the read architecture register table corresponding to the source register of the calculation instruction, create or update a third alternative architecture register entry corresponding to the destination register of the calculation instruction in the read architecture register table, and copy the content recorded in the second alternative architecture register entry to the third alternative architecture register entry.

[0021] At least one embodiment of the present disclosure further provides a processing device for a computer program, including a processing unit and a memory, and one or more computer program modules are stored on the memory, wherein the one or more computer program modules are configured to implement the prefetch method of any of the above embodiments or the prefetch training method of any of the above embodiments when executed by the processing unit.

[0022] At least one embodiment of the present disclosure further provides a non-transitory readable storage medium, wherein computer instructions are stored on the non-transitory readable storage medium, and the computer instructions implement the prefetch method of any of the above embodiments or the prefetch training method of any of the above embodiments when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0024] Figure 1 Shows a schematic diagram of a pipeline of a processor core.

[0025] Figure 2 Is a schematic diagram of an address translation process using a page table in a computer system.

[0026] Figure 3 Shows an exemplary schematic diagram of a pointer array access pattern.

[0027] Figure 4 Shows a schematic diagram of a pointer data prefetch method based on a pointer value reading instruction cache provided by at least one embodiment of the present disclosure.

[0028] Figure 5 Shows a schematic diagram of a pointer value reading instruction cache provided by at least one embodiment of the present disclosure.

[0029] Figure 6 Shows a flowchart of a prefetch method of pointer data according to at least one embodiment of the present disclosure.

[0030] Figure 7 The flowchart of a prefetching method for pointer data according to an example is shown.

[0031] Figure 8 The flowchart of a prefetching training method for pointer data according to at least one embodiment of the present disclosure is shown.

[0032] Figure 9 The flowchart of a prefetching training method for pointer data according to an example is shown.

[0033] Figure 10 The flowchart of a prefetching training method for pointer data according to some embodiments of the present disclosure is shown.

[0034] Figure 11 The flowchart of a prefetching training method for pointer data according to an example is shown.

[0035] Figure 12 The schematic diagram of an electronic device provided for at least one embodiment of the present disclosure. Detailed implementation manners

[0036] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0037] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second" and similar terms used in the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "include" or "comprise" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationships may also change accordingly.

[0038] In existing CPU core architectures, programs and data are both stored in memory (such as DRAM), so there are a large number of memory read instructions (Load instructions) in the program. Since the operating frequency of the CPU core is much higher than that of the memory, it takes hundreds of CPU core clocks to obtain data from the memory, which often causes the CPU core to idle because it cannot continue to execute relevant instructions, resulting in performance loss. High-performance CPU cores usually include multiple levels of cache to shorten the memory access latency and accelerate the running speed of the CPU core. However, when reading data that has never been accessed or data that has been evicted due to cache size limitations, the CPU core still needs to wait for dozens or even hundreds of clock cycles, which will cause performance loss.

[0039] High-performance CPU cores not only include a multi-level cache architecture to store recently accessed data, but also use a prefetcher to discover the patterns of data and instruction access by the CPU core, and prefetch the data and instructions to be accessed into the cache in advance. If the prefetched content is an instruction, it is called instruction prefetch, and the corresponding prefetcher is an instruction prefetcher; if the prefetched content is data, it is called data prefetch, and the corresponding prefetcher is a data prefetcher. The latter can be further divided into L1D data prefetcher (prefetching to the first-level data (L1D) cache), L2 data prefetcher, LLC data prefetcher (prefetching to the last-level cache (Last Level Cache)), etc. according to the target cache location.

[0040] An important function of a computer operating system is memory management. In a multi-process operating system, each process has its own virtual address space and can use any virtual address within the system-specified range. The address used by the CPU to execute an application program is the virtual address. When the operating system allocates memory to a process, it needs to map the used virtual address to a physical address, and the physical address is the real physical memory access address. This has several advantages. First, it simplifies program compilation, and the compiler can compile the program based on a continuous and sufficient virtual address space. Second, the virtual addresses of different processes are assigned to different physical addresses, enabling the system to run multiple processes simultaneously, thereby improving the operating efficiency of the entire computer system. Finally, since an application program can use but cannot change the address translation, one process cannot access the memory content of another process, thus increasing the security of the system.

[0041] Figure 2 It is a schematic diagram of the address translation process using a page table in a computer system, which shows the address translation process using a four-level page table. As Figure 2As shown, a virtual address in the system is divided into several segments. For example, they are respectively represented as EXT, OFFSET_lvl4, OFFSET_lvl3, OFFSET_lvl2, OFFSET_lvl1, and OFFSET_pg. In this example, the high-order virtual address segment EXT is not used. The virtual address segments OFFSET_lvl4, OFFSET_lvl3, OFFSET_lvl2, and OFFSET_lvl1 respectively represent the offset values of the four-level page table. That is, the virtual address segment OFFSET_lvl4 represents the offset value of the fourth-level page table, the virtual address segment OFFSET_lvl3 represents the offset value of the third-level page table, the virtual address segment OFFSET_lvl2 represents the offset value of the second-level page table, and the virtual address segment OFFSET_lvl1 represents the offset value of the first-level page table.

[0042] The initial address of the highest-level page table (i.e., the fourth-level page table) is stored in the architecture register REG_pt, and its content is set by the operating system and cannot be changed by the application program. In the second-level page table, third-level page table, and fourth-level page table, the starting address of the next-level page table is stored in each page table entry of each level of the page table. The first-level page table entry (Page Table Entry, PTE) stores the high-order bits of the physical address of the corresponding memory page, and it can be combined with the virtual address offset (OFFSET_pg) of a virtual address to obtain the physical address corresponding to the virtual address. Thus, by obtaining the starting address of the next-level page table level by level in this way, the first-level page table entry (PTE) can be finally obtained, and then the corresponding physical address can be further obtained, realizing the translation from the virtual address to the physical address. It should be noted that although Figure 2 a 4-level page table is shown, the embodiments of the present disclosure are not limited thereto, and any number of multi-level page tables can be adopted.

[0043] Address translation is a very time-consuming process. Figure 2 In the example, in the worst case, it is necessary to access the memory four times to obtain the corresponding physical address. To save address translation time and improve the performance of the computer system, many CPU cores include a Translation Lookaside Buffer (TLB) to store the previously used first-level page table entries PTE. The address translation that hits the TLB can immediately obtain the corresponding physical address. Similar to the cache architecture used for CPU cores, the TLB can also have various architectures, such as Fully Associative, Set Associative, Directly Indexed, etc. The TLB architecture can also be a multi-level structure. The size of the lowest-level TLB is the smallest and the speed is the fastest. When the lowest-level TLB misses, the next-level TLB is searched.

[0044] A pointer is a special type of data whose content is the virtual address of another block of data. Software can use pointers to implement complex data structures such as linked lists, graphs, etc. In this disclosure, "pointed data" refers to the data content stored in the data block pointed to by the pointer; "pointer value" refers to the content of the pointer itself, which is used to calculate the virtual address of the "pointed data", that is, the "pointed data address", through which the "pointed data" can be found.

[0045] In a computer system, the memory access virtual addresses of many memory access instructions are dynamically generated using one or more (architecture) registers. There are three common forms of address generation as follows:

[0046] (1) One register serves as the base, providing the base address; an immediate number provides the offset (a value with a limited size), and the virtual address is the sum of the two (i.e., base + offset);

[0047] (2) One register serves as the base, providing the base address; another register serves as the index, providing the index value, and supports limited scaling (such as x1, x2, x4, x8, etc., usually defaulting to 1); an immediate number provides the offset (a value with a limited size), and the virtual address is the sum of the three (i.e., base + index * scale + offset);

[0048] (3) One register serves as the index, providing the index value, and supports limited scaling (such as x1, x2, x4, x8, etc., usually defaulting to 1); an immediate number provides the offset (a value with a limited size), and the virtual address is the sum of the two (i.e., index * scale + offset).

[0049] As described above, complex software will extensively use pointers to implement complex data structures such as binary trees, hash tables, linked lists, graphs, etc. In this regard, the inventors of this disclosure have noticed that the reading of pointer values themselves usually follows certain rules in many applications, such as pointer arrays. A pointer array refers to an array that stores pointer values or offset values, where the offset value is a type that, after addition / subtraction / shift, is added to the same base value to obtain the pointer value. The access to pointer arrays has an obvious stride pattern and can be well prefetched by the Stride prefetcher. At the same time, there are also many other access patterns for pointer value reading that can be learned by various different prefetcher and can more effectively issue data prefetch requests.

[0050] Figure 3 An exemplary schematic diagram showing a pointer array access pattern is presented. Figure 3 The upper-middle part is the pseudo-code, and the lower part is the schematic diagram of the pointer array and pointer data. As Figure 3 shown, the continuous solid squares are the pointer array in memory, and the discrete dashed boxes are the memory addresses corresponding to the pointer array, where the corresponding pointer data are stored respectively. For example, in the pseudo-code in the upper part of the figure, the Addq pointer is decremented by 8 each time in the loop, and the ldq instruction is used to access the pointer value existing in the pointer array, and its access has an obvious pattern. After obtaining the pointer value, the ldl instruction uses the pointer value as the address to access the pointer data. For example, the Stride prefetcher in the CPU core can well capture the memory access pattern of the ldq instruction and issue a data prefetch request in advance.

[0051] However, it is difficult for existing data prefetcher to prefetch pointer data using pointers, which is caused by several reasons. First, there is often no particularly obvious pattern for prefetching between the addresses of pointer data access. For example, see Figure 3 shown, the dashed boxes are discrete from each other, which means that these pointer data may be scattered at different positions in memory. Second, when an instruction obtains a pointer value, the pointer value is often immediately used, so that the CPU core has no time to prefetch the corresponding pointer data, nor can it train the data prefetcher well. Therefore, due to the lack of an effective prefetch method, the process of obtaining pointer data from memory often causes long latency, and this long latency is an important factor restricting the performance of the CPU system.

[0052] One or more embodiments of the present disclosure provide a pointer data prefetch method and a pointer data prefetcher (prefetch device). The pointer data prefetch method can, after obtaining the pointer value, immediately generate one or more prefetch requests for pointer data according to the prefetch information of the pointer data established by, for example, the internal mode recognition of the CPU core. For example, the pointer data prefetch method can be applicable to all cases where the reading pattern of pointer values can be learned by other prefetcher and prefetch requests are sent, including but not limited to pointer arrays. The pointer data prefetch method can greatly reduce the latency of the CPU core waiting to read pointer data and improve the system performance.

[0053] The pointer data prefetching method and pointer data prefetching device according to one or more embodiments of the present disclosure can cooperate with other prefetchers in a CPU core. For a pointer value prefetching request issued by another prefetcher, the pointer data prefetching device can access and obtain pointer data prefetching information. When the pointer value prefetching returns a prefetching result, the prefetching result can be used to calculate the pointer data address of the pointer data to be prefetched, and the pointer data prefetching request can be issued using this pointer data address. In this case, the pointer data prefetching method and pointer data prefetching device according to the embodiments of the present disclosure can well solve the timeliness problem of pointer data prefetching and improve the coverage rate of pointer data prefetching through the pointer value prefetching triggered by other prefetchers.

[0054] The pointer prefetcher (prefetching device) provided by at least one embodiment of the present disclosure can be placed outside the level-1 data (L1D) cache (i.e., directly connected to the L1D cache). However, in other embodiments, the pointer prefetcher (prefetching device) can be placed in the L2 cache, LLC cache, or even in the memory. Compared with being placed in the L1D, the pointer value or pointer value prefetching data can be obtained earlier, so that the pointer data prefetching request can be issued earlier.

[0055] One or more embodiments of the present disclosure also provide a training method and method device for the pointer data prefetching method and pointer data prefetching device of the above embodiments. The training method and method device can track the mutual relationship between the pointer value reading instruction and the pointer data reading instruction, the address calculation operations required between the two, and information such as the data size and sign extension of the pointer value reading, so as to more accurately estimate the pointer value to obtain the pointer data address for pointer data prefetching.

[0056] In one or more embodiments of the present disclosure, a pointer value reading instruction cache (PointerLoad Cache, PLC) is provided for the CPU core to store information of past pointer value reading instructions, so as to implement pointer data prefetching based on the information of past pointer value reading instructions.

[0057] Figure 4 The figure shows a schematic diagram of a pointer data prefetching method based on a pointer value reading instruction cache provided by at least one embodiment of the present disclosure. As Figure 4As shown, other prefetchers (such as the Stride prefetcher) generate a pointer value prefetch request, use the pointer value prefetch request to access the PLC, and if a hit occurs, return pointer data prefetch information, and use the pointer value prefetch request to perform pointer value prefetch. When the prefetched pointer value is returned to, for example, the first-level cache (L1 cache), the pointer value prefetch backfill is performed; the pointer data prefetcher uses the prefetched pointer value and the pointer data prefetch information to generate a pointer data prefetch request, perform pointer data prefetch, and after the system obtains the prefetched pointer data, fills the pointer data into the first-level cache for future reading.

[0058] Figure 5 A schematic diagram of a pointer value read instruction cache provided by at least one embodiment of the present disclosure is shown. The pointer value read instruction cache can be implemented using hardware to record at least part of the execution information of one or more pointer value read instructions for subsequent query and use, and the recorded content can be updated (e.g., updated in real time) according to the subsequent use of the pointer value read by the one or more pointer value read instructions for read address calculation.

[0059] like Figure 4 As shown, the pointer value read instruction cache (PLC) may include multiple reference pointer value read request items, and each reference pointer value read request item includes the following multiple domains (fields):

[0060] Tag: a portion of the instruction address of the corresponding pointer value read instruction (e.g., the high-order portion of the virtual address), used to query the corresponding pointer value read instruction in the pointer value read instruction cache;

[0061] Data size (DS): the data size of the pointer value used to calculate the pointer data address in the data read by the corresponding pointer value read instruction;

[0062] One or more offsets (offset 0-n): During the calculation process, the offset between the virtual address of the pointer data and the pointer value can have different offset values ​​for different subsequent instructions using the pointer value;

[0063] One or more scales (scale 0-n): during the calculation process, the scale between the virtual address of the pointer data and the pointer value. Different subsequent instructions using the pointer value may have different scales, corresponding to the different offset values ​​described above.

[0064] One or more confidence levels (conf 0-n): confidence levels for one or more combinations of the one or more offsets and one or more scaling amounts. When the confidence level is greater than a threshold, it indicates that the combination of offsets and scaling amounts is valid and can be used to calculate a pointer data address for prefetching.

[0065] For example, a pointer value read instruction cache may adopt an architecture similar to a cache (e.g., a first-level cache), such as a fully associative, set associative, or directly indexed architecture, and may use tags and the like for retrieval and other operations. For example, during use, the replacement strategy adopted when using the most recently accessed data to fill a certain reference pointer value read request item may also include Least Recently Used (LRU), Least-Frequently Used (LFU), etc. The embodiments of the present disclosure do not limit this.

[0066] The value of the tag in the pointer value read instruction cache can be the same or similar to the tag used when querying the first-level cache (L1 Cache), for example, the high-order part of the instruction fetch address of the first instruction is used as the tag for query. In addition, if a record line in the pointer value read instruction cache includes multiple query pointer value read request items, the low-order part of the instruction fetch address can also be used as the offset value in the line to locate the corresponding query pointer value read request item.

[0067] According to the above example, since the "fixed value" in the corresponding record item will be used for predictive execution only when the confidence (Conf) is greater than the threshold set by the system, when the system determines whether to prefetch pointer data, it needs to detect whether the prefetch conditions are met, and then process according to the detection results.

[0068] In the above fields, if the system has a fixed value for the data size of the pointer value, this item can be left unset (i.e. omitted); if the system does not have a fixed value for the data size of the pointer value, the software can select a different data size, such as a byte, a word, or a double word.

[0069] In one or more embodiments of the present disclosure, in order to transfer various address calculation information between a pointer value read instruction and a pointer data read instruction, a read architecture register table is also provided for the CPU core. The read architecture register table includes one or more alternative architecture register entries, each alternative architecture register entry being an architecture register, and the value recorded therein is directly or indirectly derived from the pointer value read by a certain pointer value read instruction. It should be noted that in the present disclosure, the read architecture register table is an exemplary solution for transferring various information between the pointer value read instruction and the pointer data read instruction, and other solutions for transferring various address calculation information between the pointer value read instruction and the pointer data read instruction can also be used.

[0070] When the CPU core receives the data returned by the pointer value read instruction, different calculation operations are often required to calculate and generate the pointer data address (i.e., the virtual address of the pointer data). In one or more embodiments of the present disclosure, in-core pointer mode recognition is provided. By reading the architecture register table, the calculation information between the pointer value and the pointer data address is saved and summarized as prefetch information and written into the pointer value read instruction cache (PLC).

[0071] Different instruction sets usually define different architecture registers. The address calculation between the pointer value read instruction and the pointer data read instruction usually requires multiple architecture registers as a bridge. In the embodiments of the present disclosure, the main function of the read architecture register table is to record the operations of each relevant instruction between the pointer value and the pointer data address calculation and the corresponding architecture registers involved, and propagate this information along with the instruction stream. In one or more embodiments of the present disclosure, the calculation operations between the pointer value and the pointer data address include but are not limited to the above three methods, namely [base+offset], [base+index*scale+offset], and [index*scale+offset], and these methods can be superimposed and combined multiple times.

[0072] When the pointer data read instruction queries the read architecture register table, it can be known that its source register comes from a pointer value read instruction, and the calculation operation between the pointer value and the pointer data address can be obtained. Through the logical operation of these operations, the calculation operation between the pointer value and the pointer data virtual address can be summarized as prefetch information and saved in each domain of the PLC, and the corresponding confidence level (i.e., one of conf0~n) is incremented by 1.

[0073] For a process running in a computer system, since the base address is basically fixed and maintained in the base register, in the above embodiments, the PLC does not include a field for storing the value of the base address required for calculating the address; in addition, in address calculation, the pointer value itself is used as the index quantity, so in the above embodiments, there is also no need to include a field for storing the value of the index quantity. Embodiments of the present disclosure are not limited to the above examples, and more fields may be included as needed to record the base address and / or the index quantity.

[0074] In the read architecture register table, the content of each alternative architecture register entry may include the following multiple fields:

[0075] · Valid: Used to indicate whether this alternative architecture register entry is valid. For example,

[0076] "1" indicates valid and represents that the source of this architecture register directly or indirectly comes from a past pointer value read instruction, while "0" indicates invalid;

[0077] · Program Counter (PC): The fetch address of the past pointer value read instruction;

[0078] · Data Size (DS): The data size of the pointer value read by the past pointer value read instruction;

[0079] · Operation Information (Opinfo): Records one or more calculation information from the pointer value read by the past pointer value read instruction to the pointer data address.

[0080] Similarly, in the above fields, if the system has a fixed value for the data size of the pointer value, this item may not be set; if the system does not have a fixed value for the data size of the pointer value, the software can select different data sizes, such as one byte, one word, or double word, etc.

[0081] In at least one embodiment of the present disclosure, the read architecture register table may be a separately formed data table. For example, it can be implemented by hardware, or an existing architecture register table in the CPU core can be reused. For example, the retire register table can be selected, and the above new fields are added to each architecture register entry in this register table and the above information is recorded.

[0082] At least one embodiment of the present disclosure provides a prefetch method for pointer data, Figure 6The flowchart of the pre-fetching method is shown. As shown in the figure, the pre-fetching method includes the following steps 101 to 104:

[0083] Step 101: query and hit the first data read request in the pointer value read instruction cache (PLC).

[0084] As described above, the pointer value read instruction cache (PLC) is used to cache at least one pending pointer value read request item, each pending pointer value read request item including pointer data address calculation information. When the pointer value read instruction cache (PLC) query hits the first data read request, it indicates that the first data read request is a pointer value read request (e.g., a pointer value read instruction).

[0085] Step 102: Execute a first data read request to obtain first read data.

[0086] The first data read request is a pointer value read request, and the first read data is the pointer value.

[0087] Step 103: Use the pointer value to read the first pointer data address calculation information corresponding to the first data read request and the first read data in the instruction cache (PLC), and calculate the first pointer data prefetch address.

[0088] A first pointer data prefetch address is calculated using the acquired first read data (pointer value) and information calculated based on the first pointer data address (including an offset, a scaling amount, and a confidence level). For example, in a case where confidence level is included, a combination of an offset and a scaling amount with a confidence level greater than a threshold is selected to calculate the pointer data prefetch address. For example, if there is more than one combination of an offset and a scaling amount with a confidence level greater than a threshold, multiple pointer data prefetch addresses can be calculated.

[0089] Step 104: Use the first pointer data prefetch address to issue a first pointer data prefetch request.

[0090] As described above, in the case where multiple pointer data prefetch addresses are calculated, the multiple pointer data prefetch addresses can be used to generate and issue multiple pointer data prefetch requests, thereby generating multiple data prefetches, thereby improving coverage.

[0091] In at least one example, the first data read request comes from the program to be executed or is generated by a prefetcher. For example, the first data read request is a pointer value read instruction to be executed in the instruction stream of the program to be executed, or a pointer value prefetch request generated by other prefetchers (such as a Stride prefetcher) for, for example, a pointer array.

[0092] In at least one example, in response to the prefetcher indicating that the first read data obtained according to the data prefetch request includes at least one pointer value other than the target pointer value, the first pointer data prefetch address is calculated using the destination pointer value, and at least one other pointer data prefetch address is calculated using the at least one other pointer value, and at least one other pointer data prefetch request is issued using the at least one other pointer data prefetch address.

[0093] For example, the data returned for a pointer value request or prefetch request often is in units of cache lines, such as 64 bytes, which is usually larger than the size of a single pointer value itself. If a pointer prefetch request is generated by a prefetch request of another prefetcher, it generally indicates that there is a certain regularity in the reading of pointer values. For example, if the prefetch request generated by the Stride prefetcher hits the PLC, it means that there is a fixed stride for the reading of pointer values. Therefore, there may be multiple pointer values in a cache line, and the number n = cache line size / stride. Thus, when the data requested to be prefetched is returned, multiple pointer values can be extracted from the cache line according to the stride information. These multiple pointer values include both the destination pointer value targeted by the prefetch request itself and other pointer values. Based on these pointer values, multiple pointer data addresses can be calculated. Therefore, multiple pointer data prefetch requests can be generated through one pointer value prefetch.

[0094] Other prefetcher than the Stride prefetcher can also capture the regularity of pointer value reading. If the regularity implies that there are multiple pointer values in a cache line, multiple pointer prefetch requests can also be issued.

[0095] In at least one example, the above-mentioned method for prefetching pointer data further includes: after a first data read request hits in the pointer value read instruction cache (PLC) and before executing the first data read request to obtain the first read data, setting a pointer data prefetch flag in the first data read request, where the pointer data prefetch flag is used to trigger the first pointer data prefetch request.

[0096] For example, the pointer data prefetch flag can be set by selecting an empty bit in the first data read request, or by attaching a flag bit to the first data read request.

[0097] In at least one example, during the process of executing the first data read request to obtain the first read data, for the first read data (pointer value) to be read, if it directly hits in the first-level cache (for example, the first-level data cache, L1D), the first read data can be directly obtained from the first-level cache, or if it does not hit in the first-level cache (i.e., a miss), the first read data is obtained from a subsequent cache level (such as a second-level cache, a third-level cache, etc.) after the first-level cache and even from the memory.

[0098] For example, for the first-level cache, if the query is missed, the information of the target data to be read is written into the corresponding miss-status handling registers (MSHR), and then the target data is requested from the next-level cache (the second-level cache for the first-level cache). After the second-level cache returns the target data, the target data is filled into the first-level cache and, for example, the item corresponding to the target data in the MSHR is cleared.

[0099] In at least one example, querying the pointer value read instruction cache (PLC) to hit the first data read request includes: querying the pointer value read instruction cache (PLC) based on at least a portion of the instruction fetch address of the first data read request. More specifically, in at least one example, each of the query pointer value read request items further includes a tag for querying, and the tag is at least a portion of the instruction fetch address of the pointer value read instruction corresponding to each of the query pointer value read request items.

[0100] In at least one example, each backup pointer value read request item also includes a pointer value data size (DS) to be read. As described above, the pointer value data size (DS) to be read is the size of the pointer value data read by the pointer value read instruction corresponding to each backup pointer value read request item.

[0101] Figure 7 FIG. 1 shows a flow chart of a method for prefetching pointer data according to an example. Figure 7 As shown, in this example,

[0102] In step 701, a read request is received;

[0103] In step 702, a query is made in the PLC to determine whether the read request is matched. If matched, it indicates that the read request is a pointer value read request and the process proceeds to step 703. If matched, the process ends.

[0104] In step 703, the pointer data pre-fetch information (including pointer data address calculation information) recorded in the hit item is read from the PLC;

[0105] In step 704, a pointer data prefetch flag is set in the read request, and the CPU core determines whether to perform pointer data prefetch according to the pointer data prefetch flag;

[0106] In step 705, a read request is executed. If the requested pointer value is not in the L1D Cache, the process proceeds to step 706, otherwise proceeds to step 711;

[0107] In step 706, the pointer value read information is written into the MSHR of the L1D cache, and then a read request is issued to the lower level cache (secondary cache);

[0108] In step 707, the L1D cache receives the read pointer value returned from the lower level cache and fills it into the corresponding MSHR entry;

[0109] At step 708, read the pointer value in the corresponding MSHR entry in the L1D cache;

[0110] At step 709, the pointer data prefetcher calculates the pointer data prefetch address;

[0111] At step 710, the pointer data prefetcher uses the calculated pointer data prefetch address to issue a pointer data prefetch request;

[0112] In step 711, that is, when the pointer value to be read by the read request hits in the L1D cache, the pointer value is read from the L1D cache;

[0113] In step 712 , the pointer value obtained by reading from the L1D cache is sent to the pointer data prefetcher, and then the process enters step 710 .

[0114] Corresponding to the above-mentioned prefetching method, at least one embodiment of the present disclosure provides a prefetching device for pointer data, the prefetching device comprising:

[0115] A query module is configured to query a pointer value read instruction cache (PLC) to hit a first data read request, wherein the pointer value read instruction cache (PLC) is used to cache at least one reserve pointer value read request item, and each reserve pointer value read request item includes pointer data address calculation information;

[0116] An execution module, configured to execute a first data read request to obtain first read data;

[0117] An address calculation module is configured to use a pointer value to read first pointer data address calculation information corresponding to a first data read request and first read data in a command cache (PLC) to calculate a first pointer data prefetch address;

[0118] The request issuing module is configured to issue a first pointer data prefetch request using a first pointer data prefetch address.

[0119] For example, in the prefetching device of at least one embodiment, the pointer data prefetcher described in the above example includes the above address calculation module and request issuing module.

[0120] For example, in the pre-fetch device of at least one embodiment, for the execution module, executing the first data read request to obtain the first read data includes: directly obtaining the first read data from the first-level cache, or obtaining the first read data from the subsequent-level cache or memory after the first-level cache.

[0121] For example, in the prefetch device of at least one embodiment, the prefetch device also includes a setting module, which is configured to set a pointer data prefetch flag in the first data read request after a pointer value read instruction cache query hits the first data read request and before executing the first data read request to obtain the first read data, wherein the pointer data prefetch flag is used to trigger the first pointer data prefetch request.

[0122] For example, in the prefetch device of at least one embodiment, for the query module, querying the pointer value read instruction cache to hit the first data read request includes: querying the pointer value read instruction cache based on at least part of the instruction fetch address of the first data read request.

[0123] For example, in the pre-fetch device of at least one embodiment, each backup pointer value read request item further includes a tag for querying, and the tag is at least a part of the instruction fetch address of the pointer value read instruction corresponding to each backup pointer value read request item.

[0124] For example, in the pre-fetch device of at least one embodiment, each backup pointer value read request item also includes the size of the pointer value data to be read, and the size of the pointer value data to be read is the size of the pointer value data read by the pointer value read instruction corresponding to each backup pointer value read request item.

[0125] For example, in the prefetch device of at least one embodiment, the pointer data address calculation information included in each reserve pointer value read request item includes one or more offset values, one or more scaling amounts, and one or more confidence levels, and the one or more confidence levels respectively correspond to one or more combinations obtained by combining one or more offset values ​​and one or more scaling amounts; the one or more offset values ​​and one or more scaling amounts are used for instruction fetch address calculation, and the one or more confidence levels are used to respectively determine the credibility of one or more combinations for instruction fetch address calculation.

[0126] For example, in the prefetch device of at least one embodiment, for the address calculation module, the first pointer data address calculation information corresponding to the first data read request and the first read data in the instruction cache are read using the pointer value to calculate the first pointer data prefetch address, including: selecting a combination of an offset and a scaling amount with a confidence level greater than a threshold to calculate the first pointer data prefetch address.

[0127] For example, in the prefetching device of at least one embodiment, the first data read request comes from a program to be executed or a data prefetching request generated by a prefetcher.

[0128] For example, in the prefetch device of at least one embodiment, in response to the prefetcher indicating that the first read data obtained according to the data prefetch request includes at least one other pointer value besides the target pointer value, the first pointer data prefetch address is calculated using the destination pointer value, and at least one other pointer data prefetch address is calculated using at least one other pointer value, and at least one other pointer data prefetch request is issued using the at least one other pointer data prefetch address.

[0129] Some embodiments of the present disclosure also provide a pre-fetch training method for pointer data, the pre-fetch training method is used to maintain a pointer value read instruction cache, Figure 8 FIG. 4 shows a flow chart of the pre-fetching method. Figure 8 As shown, the method includes the following steps 201 to 205:

[0130] Step 201: Receive a first data read instruction.

[0131] For example, a first data read instruction includes a first source register.

[0132] Step 202: Query and hit a first source register in the read architecture register table.

[0133] As described above, the read architecture register table is used to record at least one candidate architecture register entry, and each candidate architecture register entry includes information of past pointer value read instructions using the corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instructions.

[0134] Step 203: According to the first data read instruction, first pointer data address calculation information based on the pointer value corresponding to the first source register and the destination pointer data address is obtained.

[0135] Step 204: In the pointer value read instruction cache (PLC), update a first look-up pointer value read request item corresponding to the past pointer value read instruction of the first source register.

[0136] As described above, the pointer value read instruction cache (PLC) is used to cache at least one backup pointer value read request item, and each backup pointer value read request item includes pointer data address calculation information.

[0137] Step 205: Write the first pointer data address calculation information into the first reference pointer value read request item.

[0138] Here, the first pointer data address calculation information is used to generate a pointer data pre-fetch request.

[0139] In at least one example, the pre-fetch training method further includes: writing first pointer data address calculation information corresponding to the first data read instruction in an architecture register entry corresponding to the first source register in the architecture register table.

[0140] In at least one example, each candidate architecture register entry also includes a valid identifier for identifying whether the current candidate architecture register entry is valid, and only when the valid identifier of the candidate architecture register entry corresponding to the first source register is a valid value, is it confirmed in the read architecture register table that the query hits the first source register.

[0141] In at least one example, each reserve pointer value read request item also includes a pointer value data size (DS) to be read, where the pointer value data size (DS) to be read is the size of the pointer value data read by a past pointer value read instruction; the information of the past pointer value read instruction using the corresponding architectural register includes at least a portion of the value address (PC) of the past pointer value read instruction.

[0142] In at least one example, writing the first pointer data address calculation information in the first reference pointer value read request item includes: in response to the first reference pointer value read request item having recorded the same content as the first pointer data address calculation information, increasing the confidence corresponding to the first pointer data address calculation information. For example, if a certain offset and scaling combination does not exist in a reference pointer value read request item, then the combination is established, and if the same offset and scaling combination already exists, then the confidence corresponding to the combination is increased (e.g., +1).

[0143] In at least one example, the pre-fetch training method further includes updating a read architecture register table.

[0144] In at least one example, the above-mentioned prefetch training method also includes, when actually executing the first pointer data read instruction using the first read data, if a cache query miss (for example, L1D cache) occurs, then the lookup pointer value read request item corresponding to the first data read request in the PLC is reset, for example, the lookup pointer value read request item is cleared, or the confidence corresponding to the combination of the relevant offset and scaling amount in the lookup pointer value read request item is set to zero.

[0145] Figure 9 FIG. 1 shows a flowchart of a method for pre-fetching training pointer data according to an example. Figure 9 As shown, in this example,

[0146] In step 901, an instruction is received;

[0147] In step 902, it is determined whether the instruction is a read instruction. If so, the process proceeds to step 903; otherwise, the process ends.

[0148] In step 903, according to the source register Rs of the read instruction, an item corresponding to the source register Rs is queried in the read architecture register table.

[0149] In step 904, it is determined whether the query is successful and the queried item is valid. If so, it indicates that the read instruction is a pointer data read instruction that uses a pointer value to obtain the pointer data address, and the process proceeds to step 905; otherwise, the process ends.

[0150] In step 905, relevant information in the item corresponding to the source register Rs is obtained from the read architecture register table. The relevant information includes, for example, information in domains such as PC, DS, and Opinfo, which respectively represent the fetch address of the pointer value read instruction that generates the pointer value, the data size of the read pointer value, and the address calculation information.

[0151] In step 906, according to the read instruction itself, the calculation information between the pointer value and the pointer data address is processed to obtain the calculation information. Different pointer data read instructions may generate addresses in different ways (the three methods described above).

[0152] In step 907, using the value in the PC domain of the relevant information obtained above, a query is made in the PLC through labels. After the query is successful, the calculation information is filled into the item of the PLC.

[0153] Moreover, the obtained calculation information can also be filled into the item corresponding to the source register Rs in the read architecture register table for subsequent training use.

[0154] Corresponding to the above training method, at least one embodiment of the present disclosure further provides a prefetch training device for pointer data. The prefetch training device includes:

[0155] A receiving module configured to receive a first data read instruction, where the first data read instruction includes a first source register.

[0156] A creating / updating module configured to query and hit the first source register in the read architecture register table, where the read architecture register table is used to record at least one alternative architecture register item, and each alternative architecture register item includes information of past pointer value read instructions using the corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instructions.

[0157] An acquisition module configured to acquire, according to a first data read instruction, first pointer data address calculation information based on a pointer value corresponding to the first source register and a destination pointer data address;

[0158] an updating module configured to update, in a pointer value read instruction cache (PLC), a first reserve pointer value read request item corresponding to a past pointer value read instruction of a first source register, wherein the pointer value read instruction cache (PLC) is used to cache at least one reserve pointer value read request item, and each reserve pointer value read request item includes pointer data address calculation information;

[0159] The writing module is configured to write first pointer data address calculation information into the first reference pointer value read request item, so as to generate a pointer data pre-fetch request.

[0160] For example, in at least one embodiment of the prefetch device, the prefetch device also includes a second write module, which is configured to write first pointer data address calculation information corresponding to the first data read instruction in the architecture register item corresponding to the first source register in the architecture register table.

[0161] For example, in the pre-fetch device of at least one embodiment, each candidate architecture register item also includes a valid identifier for identifying whether the current candidate architecture register item is valid. Only when the valid identifier of the candidate architecture register item corresponding to the first source register is a valid value, is it confirmed in the read architecture register table that the query hits the first source register.

[0162] For example, in the prefetch device of at least one embodiment, each reserve pointer value read request item also includes the size of the pointer value data to be read, and the size of the pointer value data to be read is the size of the pointer value data read by the past pointer value read instruction; the information of the past pointer value read instruction using the corresponding architectural register includes at least part of the value address of the past pointer value read instruction.

[0163] For example, in the pre-fetch device of at least one embodiment, for the write module, writing the first pointer data address calculation information in the first lookup pointer value read request item includes: in response to the first lookup pointer value read request item having recorded the same content as the first pointer data address calculation information, increasing the confidence corresponding to the first pointer data address calculation information.

[0164] Some embodiments of the present disclosure also provide a pre-fetch training method for pointer data. Figure 10 FIG. 1 shows a flow chart of the pre-fetch method, which is used to maintain the read architecture register table. Figure 10 As shown, the method includes the following steps 301 to 305:

[0165] Step 301: Receive a first instruction and obtain a read architecture register table.

[0166] As described above, the read architecture register table is used to record at least one alternative architecture register entry, and each alternative architecture register entry includes information on reading an instruction using a past pointer value of a corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value for reading the instruction, and is used to update a pointer value read instruction cache (PLC) for a pointer data prefetch operation.

[0167] Processing in the case where the first instruction is a read instruction. This processing includes: creating or updating a first alternative architecture register entry in the read architecture register table corresponding to the destination register of the read instruction, and recording the information of the read instruction in the first alternative architecture register entry.

[0168] Processing in the case where the first instruction is a calculation instruction. This processing includes: obtaining first pointer data address calculation information (Opinfo) of the calculation instruction based on a pointer value corresponding to the source register of the calculation instruction, recording the first pointer data address calculation information (Opinfo) in a second alternative architecture register entry in the read architecture register table corresponding to the source register of the calculation instruction, creating or updating a third alternative architecture register entry in the read architecture register table corresponding to the destination register of the calculation instruction, and copying the content recorded in the second alternative architecture register entry to the third alternative architecture register entry.

[0169] The above steps 302 and 303 are mutually exclusive. If it is a read instruction, a new entry can be created or an existing entry can be updated in the read architecture register table according to its destination register. If it is a calculation instruction, and it uses the read result of a previous read instruction, and the read result is one of the source operands, then the destination register of the calculation instruction can inherit the attributes of the source operand, so the corresponding information of the entry in the read architecture register table corresponding to the source register can be copied to the entry corresponding to the destination register, thereby realizing the propagation of information.

[0170] In at least one example, each alternative architecture register entry further includes a valid flag for identifying whether the current alternative architecture register entry is valid, and the first pointer data address calculation information (Opinfo) is recorded in the second alternative architecture register entry in the read architecture register table corresponding to the source register of the calculation instruction only when the valid flag of the second alternative architecture register entry is a valid value.

[0171] In at least one example, updating the first alternative architecture register entry in the destination register corresponding to the read instruction includes: clearing the original first alternative architecture register entry in the read architecture register table.

[0172] In at least one example, the information of the past pointer value read instruction includes at least a part of the fetch address (PC) of the past pointer value read instruction; the pointer data address calculation information (Opinfo) of the past pointer value read instruction includes the historical information of using the pointer value read by the past pointer value read instruction for read address calculation.

[0173] Figure 11 The flowchart of the prefetch training method of pointer data according to an example is shown. As Figure 11 shown, in this example,

[0174] In step 1101, an instruction is received;

[0175] In step 1102, it is determined whether the instruction is a load instruction. If so, the process proceeds to step 1103; otherwise, it proceeds to step 1105.

[0176] In step 1103, the original content in the entry corresponding to the destination register of the read instruction in the read architecture register table is cleared.

[0177] In step 1104, the fetch address (PC) and the read data size of the current read instruction are written into the entry corresponding to the destination register in the read architecture register table, and then the process proceeds to step 1111.

[0178] In step 1105, it is continuously determined whether the instruction is a calculation instruction. If so, the process proceeds to step 1106; otherwise, the process ends.

[0179] In step 1106, in the read architecture register table, the entry corresponding to the source register of the calculation instruction is queried.

[0180] In step 1107, if the query hits and there is an entry corresponding to the source register, the process proceeds to step 1108; otherwise, the process ends.

[0181] In step 1108, the address calculation information for calculating the address using the source register in the calculation instruction is extracted.

[0182] In step 1109, the address calculation information is written into the entry corresponding to the source register in the read architecture register table.

[0183] In step 1110, the PC and the DS field of the entry corresponding to the source register are copied to the entry corresponding to the destination register.

[0184] In step 1111, the entry corresponding to the destination architecture register in the read architecture register table is set to valid.

[0185] Corresponding to the above training method, at least one embodiment of the present disclosure further provides a prefetch training device for pointer data, and the prefetch training device includes:

[0186] A receiving module, configured to receive a first instruction and obtain a read architecture register table, where the read architecture register table is used to record at least one alternative architecture register entry, and each alternative architecture register entry includes information on a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information (Opinfo) based on the past pointer value read instruction, and is used to update a pointer value read instruction cache (PLC) for pointer data prefetch operations;

[0187] A creating / updating module, configured to, in response to the first instruction being a read instruction, create or update a first alternative architecture register entry corresponding to the destination register of the read instruction in the read architecture register table, and record the information of the read instruction in the first alternative architecture register entry, or

[0188] in response to the first instruction being a calculation instruction, obtain first pointer data address calculation information (Opinfo) of the calculation instruction based on the pointer value corresponding to the source register of the calculation instruction, record the first pointer data address calculation information (Opinfo) in a second alternative architecture register entry corresponding to the source register of the calculation instruction in the read architecture register table, create or update a third alternative architecture register entry corresponding to the destination register of the calculation instruction in the read architecture register table, and copy the content recorded in the second alternative architecture register entry to the third alternative architecture register entry.

[0189] For example, in the prefetch device of at least one embodiment, each alternative architecture register entry further includes a valid flag for identifying whether the current alternative architecture register entry is valid, and the first pointer data address calculation information is recorded in the second alternative architecture register entry corresponding to the source register of the calculation instruction in the read architecture register table only when the valid flag of the second alternative architecture register entry is a valid value.

[0190] For example, in the prefetch device of at least one embodiment, for the updating module, updating the first alternative architecture register entry corresponding to the destination register of the read instruction includes: clearing the original first alternative architecture register entry in the read architecture register table.

[0191] For example, in the prefetch device of at least one embodiment, the information of the past pointer value read instruction includes at least part of the fetch address of the past pointer value read instruction; the pointer data address calculation information of the past pointer value read instruction includes historical information on performing read address calculation using the pointer value read by the past pointer value read instruction.

[0192] At least one embodiment of the present disclosure further provides a processing device for a computer program, including a processing unit and a memory, on which one or more computer program modules are stored, wherein the one or more computer program modules are configured to implement the prefetching method of any of the above embodiments or the prefetching training method of any of the above embodiments when executed by the processing unit.

[0193] At least one embodiment of the present disclosure further provides a non-transitory readable storage medium, on which computer instructions are stored, wherein the computer instructions implement the prefetching method of any of the above embodiments or the prefetching training method of any of the above embodiments when executed by a processor.

[0194] Some embodiments of the present disclosure further provide an electronic device, which includes the processing device of any of the above embodiments or can execute the processing method of any of the above embodiments.

[0195] Figure 12 It is a schematic diagram of an electronic device provided for at least one embodiment of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), etc., and fixed terminals such as desktop computers.

[0196] Figure 12 The illustrated electronic device 1000 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure. For example, as Figure 12 shown, in some examples, the electronic device 1000 includes the image processing device of the embodiments of the present disclosure, and the image processing device can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003, such as the processing method of the computer program of the embodiments of the present disclosure. In the RAM 1003, various programs and data required for the operation of the computer system are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other through the bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0197] For example, the following components can be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; a communication device 1009 which can also include, for example, a network interface card such as a LAN card, a modem, etc. The communication device 1009 can allow the electronic device 1000 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The driver 1010 is also connected to the I / O interface 1005 as needed. A removable storage medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 1010 as needed so that a computer program read from it can be installed into the storage device 1008 as needed.

[0198] Although Figure 12 the electronic device 1000 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices can be implemented or included.

[0199] For example, the electronic device 1000 can further include a peripheral interface (not shown in the figure), etc. The peripheral interface can be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 1009 can communicate with a network and other devices through wireless communication. The network can be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication can use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.

[0200] For the present disclosure, the following points also need to be noted:

[0201] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0202] (2) Where there is no conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.

[0203] The above are only exemplary embodiments of the present disclosure, rather than being used to limit the protection scope of the present disclosure. The protection scope of the present disclosure is determined by the appended claims.

Claims

1. A method for prefetching pointer data, comprising: A first data read request is hit in a pointer value read instruction cache query, wherein the pointer value read instruction cache is used to cache at least one backup pointer value read request item, and each of the backup pointer value read request items includes pointer data address calculation information; executing the first data read request to obtain first read data; Using the pointer value to read first pointer data address calculation information corresponding to the first data read request and the first read data in the instruction cache, calculate a first pointer data prefetch address; A first pointer data prefetch request is issued using the first pointer data prefetch address.

2. The prefetching method according to claim 1, wherein, Executing the first data read request to obtain the first read data includes: directly obtain the first read data from the first level cache, or, The first read data is obtained from a subsequent level cache or a memory after the first level cache.

3. The pre-fetching method according to claim 1, further comprising: After the pointer value read instruction cache query hits the first data read request and before executing the first data read request to obtain the first read data, a pointer data prefetch flag is set in the first data read request, wherein the pointer data prefetch flag is used to trigger the first pointer data prefetch request.

4. The prefetching method according to claim 1, wherein, The query of the pointer value reading instruction cache hitting the first data read request includes: The instruction cache is queried based on at least a portion of the instruction fetch address of the first data read request.

5. The prefetching method according to claim 4, wherein, Each of the read pointer value request items further includes a tag for querying. The tag is at least a portion of the instruction fetch address of the pointer value read instruction corresponding to each of the reserve pointer value read request items.

6. The prefetching method according to claim 1, wherein, Each of the reference pointer value read request items further includes a size of the pointer value data to be read, and the size of the pointer value data to be read is the size of the pointer value data read by the pointer value read instruction corresponding to each of the reference pointer value read request items.

7. The prefetching method according to claim 1, wherein, The pointer data address calculation information included in each of the reference pointer value read request items includes one or more offset values, one or more scaling amounts, and one or more confidence levels. The one or more confidence levels correspond to one or more combinations obtained by combining the one or more offset values ​​and the one or more scaling amounts, respectively; The one or more offset values ​​and the one or more scaling amounts are used for instruction fetch address calculation, and the one or more confidence levels are used to respectively determine the confidence levels of the one or more combinations for the instruction fetch address calculation.

8. The prefetching method according to claim 7, wherein, Using the pointer value to read first pointer data address calculation information corresponding to the first data read request and the first read data in the instruction cache to calculate a first pointer data prefetch address includes: A combination of an offset and a scaling amount with a confidence level greater than a threshold is selected to calculate the first pointer data prefetch address.

9. The prefetching method according to claim 1, wherein, The first data read request comes from a program to be executed or a data prefetch request generated by a prefetcher.

10. The prefetching method according to claim 9, wherein, In response to the prefetcher indicating that the first read data obtained according to the data prefetch request includes at least one other pointer value in addition to the target pointer value, the first pointer data prefetch address is calculated using the destination pointer value, and at least one other pointer data prefetch address is calculated using the at least one other pointer value, and at least one other pointer data prefetch request is issued using the at least one other pointer data prefetch address.

11. A pre-fetch training method for pointer data, comprising: receiving a first data read instruction, wherein the first data read instruction includes a first source register; querying and hitting the first source register in a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register entry, and each candidate architecture register entry includes information of a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information based on the past pointer value read instruction; According to the first data read instruction, first pointer data address calculation information based on the pointer value corresponding to the first source register and the destination pointer data address is obtained; In a pointer value read instruction cache, updating a first lookup pointer value read request item corresponding to a past pointer value read instruction of the first source register, wherein the pointer value read instruction cache is used to cache at least one lookup pointer value read request item, each of the lookup pointer value read request items including pointer data address calculation information; The first pointer data address calculation information is written into the first look-up pointer value read request item to generate a pointer data prefetch request.

12. The pre-fetch training method according to claim 11, further comprising: In the architecture register entry corresponding to the first source register in the architecture register table, first pointer data address calculation information corresponding to the first data read instruction is written.

13. The prefetch training method according to claim 11, wherein, Each of the candidate architecture register entries also includes a valid flag for identifying whether the current candidate architecture register entry is valid. Only when the valid identifier of the candidate architecture register item corresponding to the first source register is a valid value, it is confirmed in the read architecture register table that the query hits the first source register.

14. The prefetch training method according to claim 11, wherein, Each of the reference pointer value read request items further includes the size of the pointer value data to be read, and the size of the pointer value data to be read is the size of the pointer value data read by the previous pointer value read instruction; The information of the past pointer value read instruction using the corresponding architectural register includes at least a portion of the value fetch address of the past pointer value read instruction.

15. The prefetch training method according to claim 11, wherein, Writing the first pointer data address calculation information into the first reference pointer value read request item includes: In response to the first reference pointer value read request item having recorded the same content as the first pointer data address calculation information, the confidence level corresponding to the first pointer data address calculation information is increased.

16. A pre-fetch training method for pointer data, comprising: receiving a first instruction and obtaining a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register entry, and each candidate architecture register entry includes information of a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information based on the past pointer value read instruction, and is used to update a pointer value read instruction cache for a pointer data prefetch operation; In response to the first instruction being a read instruction, creating or updating a first candidate architectural register entry in the destination register corresponding to the read instruction in the read architectural register table, and recording information of the read instruction in the first candidate architectural register entry, or, In response to the first instruction being a calculation instruction, first pointer data address calculation information of the calculation instruction based on a pointer value corresponding to a source register of the calculation instruction is obtained, the first pointer data address calculation information is recorded in a second alternative architecture register entry in the read architecture register table corresponding to the source register of the calculation instruction, a third alternative architecture register entry corresponding to a destination register of the calculation instruction is created or updated in the read architecture register table, and the content recorded in the second alternative architecture register entry is copied to the third alternative architecture register entry.

17. The prefetch training method according to claim 16, wherein, Each of the candidate architecture register entries also includes a valid flag for identifying whether the current candidate architecture register entry is valid. Only when the valid flag of the second candidate architecture register entry is a valid value, the first pointer data address calculation information is recorded in the second candidate architecture register entry in the source register corresponding to the calculation instruction in the read architecture register table.

18. The prefetch training method according to claim 16, wherein, The updating of the first candidate architecture register entry in the destination register corresponding to the read instruction comprises: Clear the original first candidate architecture register entry in the read architecture register table.

19. The prefetch training method according to claim 16, wherein, The information of the past pointer value read instruction includes at least a part of the value address of the past pointer value read instruction; The pointer data address calculation information of the past pointer value read instruction includes history information of performing read address calculation using the pointer value read by the past pointer value read instruction.

20. A pre-fetching device for pointer data, comprising: A query module is configured to query a pointer value read instruction cache to hit a first data read request, wherein the pointer value read instruction cache is used to cache at least one reserve pointer value read request item, and each reserve pointer value read request item includes pointer data address calculation information; an execution module, configured to execute the first data read request to obtain first read data; An address calculation module is configured to use the pointer value to read first pointer data address calculation information corresponding to the first data read request and the first read data in the instruction cache to calculate a first pointer data prefetch address; The request issuing module is configured to issue a first pointer data prefetch request using the first pointer data prefetch address.

21. The prefetching device according to claim 20, wherein, The execution module is further configured as: directly obtain the first read data from the first level cache, or, The first read data is obtained from a subsequent level cache or a memory after the first level cache.

22. The prefetching device according to claim 20, wherein, The pointer data address calculation information included in each of the reference pointer value read request items includes one or more offset values, one or more scaling amounts, and one or more confidence levels. The one or more confidence levels correspond to one or more combinations obtained by combining the one or more offset values ​​and the one or more scaling amounts, respectively; The one or more offset values ​​and the one or more scaling amounts are used for instruction fetch address calculation, and the one or more confidence levels are used to respectively determine the confidence levels of the one or more combinations for the instruction fetch address calculation.

23. The prefetching device according to claim 20, wherein The first data read request comes from a program to be executed or a data prefetch request generated by a prefetcher.

24. A pre-fetch training device for pointer data, comprising: A receiving module, configured to receive a first data read instruction, wherein the first data read instruction includes a first source register; a query module configured to query and hit the first source register in a read architecture register table, wherein the read architecture register table is used to record at least one candidate architecture register entry, and each candidate architecture register entry includes information of a past pointer value read instruction using a corresponding architecture register and pointer data address calculation information based on the past pointer value read instruction; an acquisition module, configured to acquire, according to the first data read instruction, first pointer data address calculation information based on the pointer value corresponding to the first source register and the destination pointer data address; an updating module, configured to update, in a pointer value read instruction cache, a first reserve pointer value read request item corresponding to a past pointer value read instruction of the first source register, wherein the pointer value read instruction cache is used to cache at least one reserve pointer value read request item, and each reserve pointer value read request item includes pointer data address calculation information; A writing module is configured to write the first pointer data address calculation information into the first reference pointer value read request item to generate a pointer data prefetch request.

25. The prefetch training device according to claim 24, wherein, Each of the reference pointer value read request items further includes the size of the pointer value data to be read, and the size of the pointer value data to be read is the size of the pointer value data read by the previous pointer value read instruction; The information of the past pointer value read instruction using the corresponding architectural register includes at least a portion of the value fetch address of the past pointer value read instruction.

26. The prefetch training device according to claim 24, wherein, The writing module is further configured as follows: In response to the first reference pointer value read request item having recorded the same content as the first pointer data address calculation information, the confidence level corresponding to the first pointer data address calculation information is increased.

27. A pre-fetch training device for pointer data, comprising: A receiving module, configured to receive a first instruction and obtain a read architecture register table, wherein the read architecture register table is used to record at least one alternative architecture register entry, and each of the alternative architecture register entries includes information on reading an instruction using a past pointer value of a corresponding architecture register and pointer data address calculation information based on the past pointer value reading instruction, and is used to update a pointer value reading instruction cache for a pointer data prefetch operation; A creating / updating module, configured to, in response to the first instruction being a read instruction, create or update a first alternative architecture register entry in the destination register corresponding to the read instruction in the read architecture register table, and record the information of the read instruction in the first alternative architecture register entry, or in response to the first instruction being a calculation instruction, obtain first pointer data address calculation information of the calculation instruction based on a pointer value corresponding to the source register of the calculation instruction, record the first pointer data address calculation information in a second alternative architecture register entry in the source register corresponding to the calculation instruction in the read architecture register table, create or update a third alternative architecture register entry in the destination register corresponding to the calculation instruction in the read architecture register table, and copy the content recorded in the second alternative architecture register entry to the third alternative architecture register entry.

28. The prefetch training device according to claim 27, wherein, The creating / updating module is further configured to: Empty the original first alternative architecture register entry in the read architecture register table to update the first alternative architecture register entry in the destination register corresponding to the read instruction.

29. The prefetch training device according to claim 27, wherein, The information of the past pointer value reading instruction includes at least part of the fetch address of the past pointer value reading instruction; The pointer data address calculation information of the past pointer value reading instruction includes historical information on performing a read address calculation using the pointer value read by the past pointer value reading instruction.

30. A processing device for a computer program, comprising: A processing unit; A memory, on which one or more computer program modules are stored, wherein the one or more computer program modules are configured to implement the prefetching method according to any one of claims 1-10 or the prefetching training method according to any one of claims 11-19 when executed by the processing unit.

31. A non-transitory readable storage medium, wherein, Computer instructions are stored on the non-transitory readable storage medium, wherein the computer instructions, when executed by a processor, implement the prefetching method according to any one of claims 1-10 or the prefetching training method according to any one of claims 11-19.

Citation Information

Patent Citations

  • Apparatus and method of prefetching data

    CN101467135A

  • Data processing method and data processing device

    CN115080464A