A Data Prefetching Compilation Optimization Method Based on Multi-Level Cache

By inserting hierarchical prefetch instructions into the compiler, the MSHR resource allocation and prefetch distance of multi-level Cache is optimized, and the problem of insufficient utilization of multi-level cache resources in the existing technology is solved, and the processor performance and data prefetching effect are improved.

CN115794111BActive Publication Date: 2025-07-04WUXI ADVANCED TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211446513.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-07-04
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

The existing compiler prefetch technology fails to make full use of the MSHR resources of multi-level cache, resulting in the first-level cache resource limitation in the multi-memory stream competition resource scenario, the memory latency increases, and the data prefetch effect is reduced.

Method used

The data prefetching and compilation optimization method based on multi-level Cache is adopted. By inserting hierarchical prefetch instructions, MSHR resources are allocated in order of first-level Cache, second-level Cache and third-level Cache, and the prefetch distances at each level are calculated. The access points with large element steps are allocated first, and the prefetch instructions of the corresponding level are transmitted.

Benefits of technology

It improves memory bandwidth utilization, alleviates the competitive pressure of first-level prefetch MSHR resources in multi-memory flow scenarios, improves processor performance, and significantly improves the effect of data prefetching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115794111B_ABST
    Figure CN115794111B_ABST
Patent Text Reader

Abstract

The present invention discloses a data prefetching compilation optimization method, device and storage medium based on a multi-level cache, including calculating the estimated execution time of a loop body; screening out memory access points in the loop body that are eligible to issue hierarchical prefetch instructions; sequentially allocating the MSHR resources in the multi-level cache to each of the screened memory access points in the order of the first-level cache, the second-level cache, and the third-level cache until the MSHR resources are allocated; issuing first-level prefetch instructions to the memory access points allocated with first-level cache prefetch resources according to the calculated first-level prefetch distance; if there is a second-level prefetch instruction template of the target machine in the compiler, issuing second-level prefetch instructions to the memory access points allocated with second-level cache prefetch resources according to the calculated second-level prefetch distance; if there is a third-level prefetch instruction template of the target machine in the compiler, issuing third-level prefetch instructions to the memory access points allocated with third-level cache prefetch resources according to the calculated third-level prefetch distance. The present invention makes full use of the MSHR resources in the multi-level cache and improves the performance of the processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of compiler optimization, and particularly relates to a data prefetching compilation optimization method, device and storage medium based on a multi-level cache. Background Art

[0002] The currently mainstream compilers mostly adopt the data prefetching compilation algorithm designed by Mowry.T.C. et al., that is, inserting prefetch instructions for loop-level array accesses based on locality analysis. The research on compiler prefetching technology is mainly divided into two categories. The first category is based on code locality analysis, mainly including three prefetching strategies: greedy prefetching, historical pointer prefetching, and data linearization prefetching proposed by Chi-Keung Luk et al.; the second category is implemented based on profiling technology, mainly including Chi-Keung Luk implementing a data prefetching framework based on profiling technology on the Alpha architecture, Youfeng Wu proposing a stride prefetching strategy based on profiling technology, and Fengbin Qi et al. implementing data prefetching of a feedback-guided chained data structure based on the feedback-guided compilation optimization technology of the ORC compiler; all the above research directly prefetches data from the main memory to the first-level cache.

[0003] Modern computers adopt a memory hierarchy method to organize the memory system. Referring to Figure 1 As shown is a typical memory hierarchy, which includes levels such as registers, first-level cache (Level 1 Cache, abbreviated as L1 Cache), second-level cache (Level 2 Cache, abbreviated as L2 Cache), third-level cache (Level 3 Cache, abbreviated as L3 Cache), and main memory. The distance from the register at the top to the processor core (referred to as CPU) becomes farther and farther down in this hierarchical structure, the device access speed becomes slower and slower, the capacity becomes larger and larger, and the cost per byte also becomes cheaper and cheaper; its caching principle is that the faster and smaller storage device at layer k (Level 0 to 4) serves as the cache of the larger and slower storage device at layer k + 1, that is, each layer in the hierarchical structure caches data objects from the lower layer; the closer the cache is to the CPU, the fewer the number of miss status holding registers (MSHRs) it holds. However, the commonly used compiler prefetching mechanism currently only directly prefetches data from the main memory to the L1 Cache. Since the L1 Cache holds fewer MSHRs, it does not make full use of the MSHRs held by the multi-level cache, which limits the actual memory bandwidth of the CPU. When there are multiple memory access streams in the program loop body, if data is directly prefetched from the main memory to the L1 Cache, it will multiply the number of requests for MSHR resources in the same period of time, making the already scarce MSHR resources become a bottleneck, resulting in the suspension of memory access instructions on the pipeline and reducing the effect of data prefetching. Summary of the Invention

[0004] The object of the present invention is to overcome the deficiencies in the prior art, and provide a data prefetching compilation optimization method, device and storage medium based on a multi-level cache, so as to solve the technical problems in the prior art such as the first-level cache resource limitation, hiding the memory access latency, and reducing the data prefetching effect easily in the case of multi-memory access stream competing for resource fields.

[0005] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:

[0006] In a first aspect, the present invention provides a data prefetching compilation optimization method based on a multi-level cache, and the method includes:

[0007] Obtain the arrangement order of each loop body in the loop queue of the function body;

[0008] According to the arrangement order, perform the operation of inserting hierarchical prefetch instructions on each of the loop bodies in reverse order;

[0009] Among them, the operation of inserting hierarchical prefetch instructions on the loop body includes:

[0010] Calculate the estimated execution time of the loop body;

[0011] Filter out the memory access points in the loop body that are eligible to issue hierarchical prefetch instructions;

[0012] Allocate the MSHR resources in the multi-level cache to each of the filtered memory access points in the order of the first-level cache, the second-level cache, and the third-level cache until the MSHR resources are allocated;

[0013] Issue first-level prefetch instructions to the memory access points allocated with first-level cache prefetch resources according to the calculated first-level prefetch distance;

[0014] If there is a second-level prefetch instruction template of the target machine in the compiler, issue second-level prefetch instructions to the memory access points allocated with second-level cache prefetch resources according to the calculated second-level prefetch distance;

[0015] If there is a third-level prefetch instruction template of the target machine in the compiler, issue third-level prefetch instructions to the memory access points allocated with third-level cache prefetch resources according to the calculated third-level prefetch distance.

[0016] Combined with the first aspect, preferably, when allocating cache prefetch resources to the memory access points, the memory access points with a larger element stride are given priority to obtain MSHR resources.

[0017] Combined with the first aspect, preferably, the calculation formula for the estimated execution time of the loop body is:

[0018]

[0019] In the formula: T loop represents the estimated execution time of the loop body loop, m represents the total number of intermediate representation statements of the loop body loop, and i = 1, 2... m; S i represents the i-th intermediate representation statement in the loop body loop, and W i represents the weight corresponding to the type to which the intermediate representation statement i belongs.

[0020] Combined with the first aspect, preferably, the step of screening out the memory access points in the loop body that are eligible to issue hierarchical prefetch instructions includes:

[0021] Obtaining the number of memory access points and the number of instructions of each basic block in the loop body respectively;

[0022] Substituting the number of memory access points and the number of instructions of each of the basic blocks into formula (2) for calculation:

[0023]

[0024] In the formula: bb_refs represents the number of memory access points in the basic block, bb_insnsn represents the number of instructions included in the basic block, RATIOx represents the minimum parallelism of calculation and memory access that can be accepted during the x-level Cache prefetch, x represents the level of the Cache, and the value of x is 1, 2, or 3; MSHR_NUMx represents the MSHR resource number of the x-level Cache, and α is a tuning coefficient;

[0025] Granting all memory access points in the basic block whose calculation results meet the two conditions in formula (2) the qualification to issue hierarchical prefetch instructions.

[0026] Combined with the first aspect, preferably, it is characterized in that the calculation formula for the first-level prefetch distance of the memory access point is:

[0027]

[0028] In the formula: PD L1 represents the first-level prefetch distance of the calculated memory access point, Lat MEM represents the main memory access latency, Lat l1 represents the first-level Cache access latency, Lat l2 represents the second-level Cache access latency, Lat l3It represents the access latency of the L3 cache, step represents the element step size of the memory access point; l3pf represents the L3 prefetch flag of the memory access point. When the calculated memory access point is allocated L3 cache prefetch resources, the value of l3pf is 1, otherwise the value of l3pf is 0; l2pf represents the L2 prefetch flag of the memory access point. When the calculated memory access point is allocated L2 cache prefetch resources, the value of l2pf is 1, otherwise the value of l2pf is 0.

[0029] Combined with the first aspect, preferably, the calculation formula for the L2 prefetch distance of the memory access point is:

[0030]

[0031] In the formula: PD L2 represents the L2 prefetch distance of the calculated memory access point.

[0032] Combined with the first aspect, preferably, the calculation formula for the L3 prefetch distance of the memory access point is:

[0033]

[0034] In the formula: PD L3 represents the L3 prefetch distance of the calculated memory access point.

[0035] In the second aspect, the present invention provides a data prefetch compilation optimization device based on a multi-level cache. The device includes:

[0036] An acquisition module, configured to acquire the arrangement order of each loop body in the function body loop queue;

[0037] An insert hierarchical prefetch instruction module, configured to perform an operation of inserting hierarchical prefetch instructions on each of the loop bodies in reverse order according to the arrangement order;

[0038] The insert hierarchical prefetch instruction module includes:

[0039] A calculation unit, configured to calculate the estimated execution time of the loop body;

[0040] A screening unit, configured to screen out the memory access points in the loop body that are eligible to issue hierarchical prefetch instructions;

[0041] A resource allocation unit, configured to allocate the MSHR resources in the multi-level cache to each of the screened memory access points in the order of the first-level cache, the second-level cache, and the third-level cache until the MSHR resources are allocated;

[0042] A first-level prefetch instruction emission unit, configured to emit first-level prefetch instructions to the memory access points allocated with first-level cache prefetch resources according to the calculated first-level prefetch distance;

[0043] A second-level prefetch instruction issuing unit, which is used to issue second-level prefetch instructions to the memory access points allocated with second-level Cache prefetch resources according to the calculated second-level prefetch distance if there is a second-level prefetch instruction template of the target machine in the compiler;

[0044] A third-level prefetch instruction issuing unit, which is used to issue third-level prefetch instructions to the memory access points allocated with third-level Cache prefetch resources according to the calculated third-level prefetch distance if there is a third-level prefetch instruction template of the target machine in the compiler.

[0045] In a third aspect, the present invention provides a data prefetch compilation optimization device based on multi-level Cache, including a processor and a storage medium;

[0046] The storage medium is used to store instructions;

[0047] The processor is used to operate according to the instructions to execute the steps of the data prefetch compilation optimization method based on multi-level Cache according to any one of the first aspects.

[0048] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data prefetch compilation optimization method based on multi-level Cache according to any one of the first aspects are realized.

[0049] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0050] 1. The present invention inserts hierarchical prefetch instructions for collaborative work for multi-level caches, improving the utilization rate of memory bandwidth:

[0051] With the continuous development of multi-core processors, their memory hierarchy has become increasingly complex. Only by coordinating the work among multi-level caches can the timeliness of data prefetch and the maximum performance of Cache resources be ensured, further improving the data prefetch optimization effect; for caches farther from the CPU, the MSHR resources increase while the cost decreases; the present invention makes full use of the MSHR resources in multi-level Cache, and by inserting hierarchical prefetch instructions, alleviates the problem of first-level prefetch MSHR resource competition pressure that may occur in the scenario of multiple memory access streams, improving the processor performance.

[0052] 2. The present invention adds a dynamically adjustable offset to the original prefetch distance, sets a more appropriate prefetch distance, and fully exploits the potential of data prefetch:

[0053] For collaborative prefetching for multi-level caches, it is necessary to calculate the prefetch distances of each level of cache for each prefetch object to ensure effective cooperation between levels of cache and timely prefetch data into each level of cache. Considering the memory access latency between multi-level cache storage structures, the present invention adds a secondary prefetch fixed offset and a tertiary prefetch fixed offset to the original prefetch distance and introduces a dynamically adjustable offset to obtain a more accurate prefetch distance, thereby further improving performance on the basis of primary prefetching. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a schematic diagram of a typical cache-based memory hierarchy;

[0055] Figure 2 is the overall flowchart of a data prefetch compilation optimization method based on multi-level caches provided by an embodiment of the present invention;

[0056] Figure 3 is the flowchart of the operation of inserting hierarchical prefetch instructions into a loop body provided by an embodiment of the present invention;

[0057] Figure 4 is the structural principle block diagram of a data prefetch compilation optimization device based on multi-level caches provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The technical solution of the present invention will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0059] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0060] Embodiment 1:

[0061] Referring to Figure 2 as shown, this embodiment introduces a data prefetch compilation optimization method based on multi-level caches. First, it is necessary to check whether the target machine has a cache that supports prefetching. If prefetching is not supported, this optimization process ends; if prefetching is supported, the optimization process provided in this embodiment is executed, which specifically includes the following steps:

[0062] Step S1: Obtain the arrangement order of each loop body in the function body loop queue;

[0063] Step S2: According to the arrangement order, perform the operation of inserting hierarchical prefetch instructions on each of the loop bodies in reverse order;

[0064] As an embodiment of the present invention, referring to Figure 3 as shown, the operation of inserting hierarchical prefetch instructions into the loop body includes:

[0065] Step 2.1: Calculate the estimated execution time of the loop body;

[0066] Specifically, the calculation formula for the estimated execution time of the loop body is:

[0067]

[0068] In the formula: T loop represents the estimated execution time of loop body loop, m represents the total number of intermediate representation statements of loop body loop, i = 1, 2... m; S i represents the i-th intermediate representation statement in loop body loop, W i represents the weight corresponding to the type of intermediate representation statement i.

[0069] Step 2.2: Screen out the memory access points in the loop body that are eligible to emit hierarchical prefetch instructions;

[0070] Furthermore, step 2.2 evaluates whether the MSHR resources competition of each level of cache is crowded in units of basic blocks in the loop body, specifically including the following steps:

[0071] Step 2.2.1: Respectively obtain the number of memory access points and the number of instructions of each basic block in the loop body;

[0072] Step 2.2.2: Substitute the number of memory access points and the number of instructions of each of the basic blocks into formula (2) for calculation:

[0073]

[0074] In the formula: bb_refs represents the number of memory access points in the basic block, bb_insnsn represents the number of instructions included in the basic block, RATIOx represents the minimum parallelism of calculation and memory access that can be accepted during the prefetch of the x-th level of Cache, x represents the level of Cache, and the value of x is 1, 2 or 3; MSHR_NUMx represents the number of MSHR resources of the x-th level of Cache, and α is a tuning coefficient;

[0075] Step 2.2.3: Entitle all memory access points in the basic blocks whose calculation results meet the two conditions in formula (2) to have the qualification to issue hierarchical prefetch instructions.

[0076] Step 2.3: Allocate the MSHR resources in the multi-level Cache to each screened memory access point in the order of the first-level Cache, the second-level Cache, and the third-level Cache until the MSHR resources are completely allocated;

[0077] It should be further noted that Step 2.3 specifically includes:

[0078] Step 2.3.1: Allocate the MSHR resources in the first-level Cache to the memory access points, and give priority to the memory access points with a larger element stride to obtain the resources. Set the first-level prefetch flag l1pff of the memory access points that obtain the MSHR resources in the first-level Cache from 0 to 1. After the MSHR resources in the first-level Cache are completely allocated, execute Step 2.3.2;

[0079] Step 2.3.2: Allocate the MSHR resources in the second-level Cache to the memory access points, and give priority to the memory access points with a larger element stride to obtain the resources. Set the second-level prefetch flag l2pf of the memory access points that obtain the MSHR resources in the second-level Cache from 0 to 1. After the MSHR resources in the second-level Cache are completely allocated, execute Step 2.3.3;

[0080] Step 2.3.3: Allocate the MSHR resources in the third-level Cache to the memory access points, and give priority to the memory access points with a larger element stride to obtain the resources. Set the third-level prefetch flag l3pf of the memory access points that obtain the MSHR resources in the third-level Cache from 0 to 1.

[0081] Step 2.4: Issue first-level prefetch instructions to the memory access points that have been allocated first-level Cache prefetch resources according to the calculated first-level prefetch distance;

[0082] Furthermore, in combination with the estimated execution time of the loop body calculated in Step 2.1, the calculation formula for the first-level prefetch distance of the memory access point is:

[0083]

[0084] In the formula: PD L1 represents the first-level prefetch distance of the calculated memory access point, Lat MEM represents the main memory access latency, Lat l1 represents the first-level Cache access latency, Lat l2 represents the second-level Cache access latency, Lat l3It represents the access latency of the three - level cache, step represents the element step size of the memory access point; l3pf represents the three - level pre - fetch flag of the memory access point. When the calculated memory access point is allocated three - level cache pre - fetch resources, l3pf takes the value of 1, otherwise l3pf takes the value of 0; l2pf represents the two - level pre - fetch flag of the memory access point. When the calculated memory access point is allocated two - level cache pre - fetch resources, l2pf takes the value of 1, otherwise l2pf takes the value of 0.

[0085] Step 2.5: If there is a two - level pre - fetch instruction template of the target machine in the compiler, then issue two - level pre - fetch instructions for the memory access points allocated with two - level cache pre - fetch resources according to the calculated two - level pre - fetch distance.

[0086] Furthermore, in combination with the estimated execution time of the loop body calculated in Step 2.1, the calculation formula for the two - level pre - fetch distance of the memory access point is:

[0087]

[0088] In the formula: PD L2 represents the two - level pre - fetch distance of the calculated memory access point.

[0089] Step 2.6: If there is a three - level pre - fetch instruction template of the target machine in the compiler, then issue three - level pre - fetch instructions for the memory access points allocated with three - level cache pre - fetch resources according to the calculated three - level pre - fetch distance.

[0090] Furthermore, in combination with the estimated execution time of the loop body calculated in Step 2.1, the calculation formula for the three - level pre - fetch distance of the memory access point is:

[0091]

[0092] In the formula: PD L3 represents the three - level pre - fetch distance of the calculated memory access point.

[0093] It should be further noted that before executing Step 2.5, first determine whether there is a two - level pre - fetch instruction template of the target machine in the compiler. If there is a two - level pre - fetch instruction template, then execute Step 2.5; if there is no two - level pre - fetch instruction template, then determine whether there is a three - level pre - fetch instruction template of the target machine in the compiler. If there is a three - level pre - fetch instruction template, then execute Step 2.6; if there is no three - level pre - fetch instruction template, then end the operation of inserting hierarchical pre - fetch instructions for this loop body and continue the operation of inserting hierarchical pre - fetch instructions for the next loop body until all loop bodies are traversed, and finally end this optimization process.

[0094] In summary, the data prefetch compilation optimization method based on multi-level caches provided by the embodiments of the present invention makes full use of the MSHR resources at all levels of the memory. By inserting hierarchical prefetch instructions and adding offsets to the prefetch distances at each level, the pressure on the MSHR resources in the first-level cache is alleviated, and phenomena such as the suspension of memory access instructions in the instruction pipeline are effectively prevented, enabling the data prefetch to obtain positive benefits and further improving the performance on the basis of the first-level prefetch. At the same time, the distribution and quantity of the hierarchical prefetch instructions are restricted, thus avoiding exceeding the processing capabilities of the MSHR resources at each level, and the problems of first-level cache resource limitation and hidden memory access latency can be avoided in the scenario of multi-memory access stream competing for resources, further improving the processor performance. The embodiments of the present invention also conducted experiments on the memory bandwidth test program STREAM and the standard performance evaluation program SPEC2006. The results show that the hierarchical prefetch provided by the present invention can significantly improve the performance evaluation of the subject, effectively solve the deficiencies of the first-level prefetch, and fully tap the potential of data prefetch.

[0095] Embodiment 2:

[0096] As Figure 4 shown, the embodiments of the present invention provide a data prefetch compilation optimization device based on multi-level caches, which can be used to implement the method described in Embodiment 1, and specifically includes:

[0097] An acquisition module, configured to acquire the arrangement order of each loop body in the function body loop queue;

[0098] An insert hierarchical prefetch instruction module, configured to perform an operation of inserting hierarchical prefetch instructions on each of the loop bodies in reverse order according to the arrangement order;

[0099] The insert hierarchical prefetch instruction module includes:

[0100] A calculation unit, configured to calculate the estimated execution time of the loop body;

[0101] A screening unit, configured to screen out the memory access points in the loop body that are eligible to emit hierarchical prefetch instructions;

[0102] A resource allocation unit, configured to allocate the MSHR resources in the multi-level cache to each of the screened memory access points in the order of the first-level cache, the second-level cache, and the third-level cache until the MSHR resource allocation is completed;

[0103] A first-level prefetch instruction emission unit, configured to emit first-level prefetch instructions to the memory access points allocated with the first-level cache prefetch resources according to the calculated first-level prefetch distance;

[0104] A second-level prefetch instruction unit for, if there is a second-level prefetch instruction template of the target machine in the compiler, issuing a second-level prefetch instruction to the memory access points allocated with second-level cache prefetch resources according to the calculated second-level prefetch distance;

[0105] A third-level prefetch instruction unit for, if there is a third-level prefetch instruction template of the target machine in the compiler, issuing a third-level prefetch instruction to the memory access points allocated with third-level cache prefetch resources according to the calculated third-level prefetch distance.

[0106] The data prefetch compilation optimization device based on multi-level caches provided by the embodiments of the present invention and the data prefetch compilation optimization method based on multi-level caches provided by the first embodiment are based on the same technical concept, can produce beneficial effects as described in the first embodiment, and the content not described in detail in this embodiment can be referred to the first embodiment.

[0107] Embodiment 3:

[0108] The embodiments of the present invention provide a data prefetch compilation optimization device based on multi-level caches, including a processor and a storage medium;

[0109] The storage medium is used to store instructions;

[0110] The processor is used to operate according to the instructions to execute the steps of any one of the methods according to the first embodiment.

[0111] Embodiment 4:

[0112] The embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the methods as implemented in the first embodiment are realized.

[0113] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0114] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or multiple blocks.

[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or multiple blocks.

[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 means for implementing the functions specified in one or more blocks or multiple blocks.

[0117] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A data prefetching compilation optimization method based on multi-level Cache, characterized in that The method includes: Obtaining the arrangement order of each loop body in the function body loop queue; According to the arrangement order, successively performing operations of inserting hierarchical prefetch instructions on each of the loop bodies in reverse order; Among them, the operation of inserting hierarchical prefetch instructions for a loop body includes: Calculating the estimated execution time of the loop body; Filtering out the memory access points in the loop body that are eligible to issue hierarchical prefetch instructions; Successively allocating the MSHR resources in the multi-level cache to each of the filtered memory access points in the order of the first-level cache, the second-level cache, and the third-level cache until the MSHR resources are completely allocated; Issuing first-level prefetch instructions to the memory access points allocated with first-level cache prefetch resources according to the calculated first-level prefetch distance; If there is a second-level prefetch instruction template of the target machine in the compiler, issuing second-level prefetch instructions to the memory access points allocated with second-level cache prefetch resources according to the calculated second-level prefetch distance; If there is a third-level prefetch instruction template of the target machine in the compiler, issuing third-level prefetch instructions to the memory access points allocated with third-level cache prefetch resources according to the calculated third-level prefetch distance.

2. The data prefetching compilation optimization method based on multi-level Cache according to claim 1, wherein When allocating cache prefetch resources to the memory access points, the memory access points with a larger element stride are given priority to obtain MSHR resources.

3. The data prefetching compilation optimization method based on a multi-level cache according to claim 1, wherein The calculation formula for the estimated execution time of the loop body is: Where: T loop represents the estimated execution time of the loop body loop, m represents the total number of intermediate representation statements of the loop body loop, and i = 1, 2... m; S i represents the i-th intermediate representation statement in the loop body loop, and W i represents the weight corresponding to the type to which the intermediate representation statement i belongs.

4. The data prefetching compilation optimization method based on a multi-level cache according to claim 1, wherein The step of filtering out the memory access points in the loop body that are eligible to issue hierarchical prefetch instructions includes: Respectively obtaining the number of memory access points and the number of instructions of each basic block in the loop body; Substituting the number of memory access points and the number of instructions of each basic block into formula (2) for calculation respectively: In the formula: bb_refs represents the number of memory access points in the basic block, bb_insnsn represents the number of instructions included in the basic block, RATIOx represents the minimum parallelism of calculation and memory access that can be accepted during the x-level cache prefetch, x represents the level of the cache, and the value of x is 1, 2, or 3; MSHR_NUMx represents the number of MSHR resources of the x-level cache, and α is a tuning coefficient; Giving all the memory access points in the basic block whose calculation results meet the two conditions in formula (2) the qualification to issue hierarchical prefetch instructions.

5. The data prefetching compilation optimization method based on multi-level Cache according to any one of claims 1 to 4, characterized in that The calculation formula for the first-level prefetch distance of the memory access point is: Where: PD L1 represents the first-level prefetch distance of the calculated memory access point, Lat MEM represents the main memory access latency, Lat l1 represents the first-level cache access latency, Lat l2 represents the second-level cache access latency, Lat l3 represents the third-level cache access latency, step represents the element step size of the memory access point; l3pf represents the third-level prefetch flag of the memory access point. When the calculated memory access point is allocated third-level cache prefetch resources, l3pf takes the value of 1, otherwise l3pf takes the value of 0; l2pf represents the second-level prefetch flag of the memory access point. When the calculated memory access point is allocated second-level cache prefetch resources, l2pf takes the value of 1, otherwise l2pf takes the value of 0.

6. The data prefetching compilation optimization method based on a multi-level cache according to claim 5, characterized in that The calculation formula for the second-level prefetch distance of the memory access point is: where: PD L2 represents the second-level prefetch distance of the calculated memory access point.

7. The data prefetching compilation optimization method based on multi-level Cache according to claim 6, wherein The calculation formula for the third-level prefetch distance of the memory access point is: Where: PD L3 represents the three-level prefetch distance of the calculated memory access point.

8. A data prefetching compilation optimization device based on a multi-level cache, characterized in that, The device includes: An obtaining module, configured to obtain the arrangement order of each loop body in the function body loop queue; An inserting hierarchical prefetch instruction module, configured to successively perform operations of inserting hierarchical prefetch instructions on each of the loop bodies in reverse order according to the arrangement order; The inserting hierarchical prefetch instruction module includes: A calculating unit, configured to calculate the estimated execution time of the loop body; A filtering unit, configured to filter out the memory access points in the loop body that are eligible to issue hierarchical prefetch instructions; A resource allocating unit, configured to successively allocate the MSHR resources in the multi-level cache to each of the filtered memory access points in the order of the first-level cache, the second-level cache, and the third-level cache until the MSHR resources are completely allocated; A first-level prefetch instruction issuing unit, configured to issue first-level prefetch instructions to memory access points allocated with first-level Cache prefetch resources according to the calculated first-level prefetch distance; A second-level prefetch instruction issuing unit, configured to issue second-level prefetch instructions to memory access points allocated with second-level Cache prefetch resources according to the calculated second-level prefetch distance if there is a second-level prefetch instruction template of the target machine in the compiler; A third-level prefetch instruction issuing unit, configured to issue third-level prefetch instructions to memory access points allocated with third-level Cache prefetch resources according to the calculated third-level prefetch distance if there is a third-level prefetch instruction template of the target machine in the compiler.

9. A data prefetching compilation optimization device based on a multi-level cache, characterized in that Comprising a processor and a storage medium; The storage medium is used for storing instructions; The processor is configured to operate according to the instructions to execute the steps of the multi-level Cache-based data prefetch compilation optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the multi-level Cache-based data prefetch compilation optimization method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device of pre-fetching data of compiler

    CN102981883A

  • Cache prefetching method and system

    CN110059025A