Data prefetching method and device

By using the control flow graph and the number of events recorded by the performance monitoring unit in the cache between the processor and memory, the target path is determined and the data is written to the cache, which solves the problem of low data access efficiency caused by static branch prediction and improves the data access efficiency of the processor.

CN120832172APending Publication Date: 2025-10-24HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410504625.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

In the cache between the processor and memory, static branch prediction leads to inaccurate branch paths, resulting in data stored in the cache that is not frequently accessed by the processor, thus reducing the overall efficiency of processor data access.

Method used

By acquiring the basic block events in the control flow graph and using the event counts recorded by the performance monitoring unit, the target path is determined. Before executing the basic block in the control flow graph, the target data is written to the cache. The sum of the event counts of the basic blocks in the target path is greater than or equal to a first threshold, and the target data access frequency is greater than or equal to a second threshold.

Benefits of technology

It improves the efficiency of data access when the processor executes according to the control flow graph, ensures that the data written to the cache is frequently accessed by the processor, reduces unnecessary data access, and improves overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832172A_ABST
    Figure CN120832172A_ABST
Patent Text Reader

Abstract

The invention discloses a data prefetching method and device, and relates to the technical field of computers. The data prefetching method comprises the steps that a control flow diagram and events, recorded by a PMU, of basic blocks in the control flow diagram are obtained, target paths of multiple paths in the control flow diagram are determined according to the occurrence frequency of the events of the basic blocks in the control flow diagram, and then target data corresponding to a first basic block are written into a buffer before the first basic block is executed. Wherein the sum of the occurrence frequency of all events of the basic block in the target path is greater than or equal to a first threshold value, and the target data is data with the access frequency or the access frequency greater than or equal to a second threshold value in the data needing to be accessed when the target path is executed. The sum of the occurrence times corresponding to the target path is greater than or equal to the first threshold value, so that the probability that the target path is frequently executed is relatively high, the target data determined according to the target path is written into the cache in advance, and the data access efficiency when the processor executes according to the control flow diagram can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a data prefetching method and device. BACKGROUND

[0002] In a processor architecture, a cache is arranged between a processor and a memory, which can temporarily store data between the processor and the memory. The capacity of the cache is limited. In order to balance the capacity of the cache and the efficiency of the processor accessing data, the data frequently accessed by the processor is usually stored in the cache. At present, a static branch prediction can be performed on source code by using a compiler. Before the source code is executed, the execution probability of each branch of the source code is obtained according to the static branch prediction, the branch path with the maximum execution probability is determined, and then the data to be stored in the cache is the data frequently accessed when the branch path is executed.

[0003] However, in the static branch prediction process, when the local branch probabilities are equal or zero, it is impossible to determine which branch is executed when the source code is executed, which leads to inaccurate determination of the branch path, and then the data determined according to the branch path is not necessarily the data frequently called by the processor. Storing the data determined in the cache will lead to low overall efficiency of the processor accessing data. SUMMARY

[0004] The present application provides a data prefetching method and device to solve the problem that the data written in the cache is not necessarily the data frequently called by the processor, which leads to low overall efficiency of the processor accessing data.

[0005] The present application adopts the following technical solutions.

[0006] In a first aspect, the present application provides a data prefetching method. The data prefetching method can be applied to a computer system or a computing device for implementing the data prefetching method in the computer system. The computing device is, for example, a server or a terminal. The data prefetching method comprises: obtaining a control flow graph and events of basic blocks in the control flow graph recorded by a performance monitoring unit (PMU), determining a target path of a plurality of paths in the control flow graph according to the number of occurrences of the events of the basic blocks in the control flow graph. Then, before a first basic block in the control flow graph is executed, target data corresponding to the first basic block is written into a cache. Wherein, the sum of the number of occurrences of all events of the basic blocks in the target path is greater than or equal to a first threshold value, and the target data is data with a frequency of access or a frequency of access greater than or equal to a second threshold value in the data accessed when the target path is executed.

[0007] In the present application, the PMU records the number of occurrences of various events when the processor executes instructions in a basic block, the more times the PMU records the number of occurrences of events of a basic block, the more frequently the instructions in the basic block are executed. Further, according to the number of occurrences of events of the basic block recorded by the PMU, a target path including a plurality of basic blocks whose sum of the number of occurrences of events is greater than or equal to a first threshold value is determined from a plurality of paths included in the control flow graph, the probability of the target path being frequently executed is higher, and the access frequency or access frequency of the target data determined according to the target path is also higher. Therefore, the data (target data) accessed by the first basic block in the control flow graph for execution is written into the cache, which can improve the overall efficiency of data access when the processor executes according to the control flow graph.

[0008] In a possible example, the first basic block includes instructions for accessing the target data.

[0009] In a possible example, the hint instruction can be deployed in the first basic block, and when the hint instruction is executed, the target data is written into the cache and is stored persistently.

[0010] For example, the hint instruction can be deployed in the adjacent line of the instruction for accessing the target data included in the first basic block.

[0011] In a possible case, the basic blocks in the control flow graph include basic blocks within nested loops.

[0012] In the present application, since the hot function or code segment is usually a nested loop structure, the processor determines the control flow graph including the nested loop (structure) from all control flow graphs, which avoids the subsequent target data identification and data prefetching of all control flow graphs, and improves the efficiency of writing the data accessed by the hot function or code segment into the cache.

[0013] In a possible case, the sum of the number of occurrences of all events of the basic blocks included in each path of the plurality of paths is a candidate value, and one of the plurality of candidate values of the plurality of paths is the first threshold value.

[0014] In a possible example, the first threshold value is the maximum value in the plurality of candidate values. That is, the target path is the path with the maximum sum of the number of occurrences of all events in the plurality of paths.

[0015] In the present application, a path with the maximum sum of occurrence times of all events is determined as a target path from a plurality of paths, and the probability of the target path being frequently executed is the highest among the plurality of paths. Therefore, the access frequency or access frequency of the target data determined according to the target path is also the highest. Thus, the target data required to be accessed by the first basic block in the control flow graph is written into the cache, so that the overall efficiency of data access when the processor executes according to the control flow graph can be improved.

[0016] In another possible example, the first threshold is any one of the candidate values arranged in the first N candidate values. N is a positive integer greater than or equal to 1.

[0017] In a possible implementation, the target path of the plurality of paths in the control flow graph is determined according to the occurrence times of the events of the basic blocks in the control flow graph, comprising: searching a plurality of paths in the control flow graph starting from the initial block, and then determining the sum of the occurrence times of the events of the basic blocks included in each path in the plurality of paths. Thus, the target path with the sum of the occurrence times of the events greater than or equal to the first threshold is determined from the plurality of paths according to the sum of the occurrence times of the events of the basic blocks included in each path in the plurality of paths. Wherein, the exit basic block of each path in the plurality of paths does not belong to the basic blocks in the nested loop in the control flow graph.

[0018] In the present application, the control flow graph is searched starting from the initial block in the control flow graph to obtain a plurality of paths, and the exit basic block of each path in the plurality of paths does not belong to the basic blocks in the nested loop in the control flow graph, which can ensure that the plurality of paths obtained by searching are longer paths (including paths with a larger number of basic blocks), reduce the additional performance loss caused by judging the paths including a smaller number of basic blocks, and improve the efficiency of determining the target path.

[0019] In a possible example, depth-first search can be used to search the control flow graph.

[0020] In a possible case, each path in the plurality of paths includes all the basic blocks in the nested loop in the control flow graph.

[0021] In the present application, since each path includes all the basic blocks in the nested loop, when searching the paths in the control flow graph, the exit basic block of each path does not exit from the last basic block in the nested loop to the exit basic block in the control flow graph, that is, all the basic blocks in the nested loop are completely included in each path. Thus, it is ensured that each path is a longer path (including a path with a larger number of basic blocks), and the shorter path is avoided to be included in the judgment of the sum of the occurrence times of the events, so as to cause additional performance loss and improve the efficiency of determining the target path.

[0022] In a possible implementation, before executing a first basic block in the control flow graph, the target data corresponding to the first basic block is written into the cache, including: parsing the source code corresponding to the target path to obtain the target data, and then determining the first basic block with the highest priority from the plurality of basic blocks included in the control flow graph, so as to write the target data into the cache before executing the first instruction included in the first basic block. The reuse score of the target data is the reuse score of the M data with the highest priority in the reuse scores of the plurality of data indicated by the source code, and M is an integer greater than or equal to 1. The reuse score of the target data is used to indicate the reuse degree of the target data, and each basic block in the plurality of basic blocks includes the first instruction for accessing the target data.

[0023] In the present application, since the first basic block is the basic block with the highest priority in the plurality of basic blocks, and the priority indicates the priority of the time when the target data is written into the cache, the target data is written into the cache before the first instruction included in the first basic block is executed, which can improve the overall efficiency of data access.

[0024] In a possible case, the reuse score of the data can be determined according to one or more of the following: data reuse mode, memory access regularity, read-write type, and space occupation limit.

[0025] In a possible case, the priority of the basic block is determined according to one or more of the following indicators: the loop depth indicated by the first instruction, the access regularity of the first instruction, the number of occurrences of the event of the basic block in which the first instruction is located, the line number of the first instruction in the source code, whether the first instruction is used to indicate reading or writing of the target data, and whether the first instruction is vectorized memory access or scalarized memory access.

[0026] In a possible example, the priority of the basic block is determined according to one or more of the following indicators: the loop depth indicated by the first instruction, the access regularity of the first instruction, the number of occurrences of the event of the basic block in which the first instruction is located, the line number of the first instruction in the source code, whether the first instruction is used to indicate reading or writing of the target data, and whether the first instruction is vectorized memory access or scalarized memory access.

[0027] Taking the first instruction in the basic block a as instruction a and the first instruction in the basic block b as instruction b as an example. First, the loop depth of instruction a is compared with the loop depth of instruction b. When the loop depths of instruction a and instruction b are different, the priority of the basic block corresponding to the instruction with the higher loop depth is higher. If the priority cannot be determined according to the loop depth, the priority is determined according to the access regularity, the number of occurrences, the line number, the read / write target data, and the vectorized memory access / scalarized memory access in sequence until the priority of instruction a and instruction b is determined.

[0028] More details about the present example can be found in the following detailed description, which will not be repeated here.

[0029] In a possible implementation, before the first basic block of the control flow graph is executed, the data prefetching method further comprises: deploying a second instruction in the first basic block, the second instruction being used to indicate that the target data is written into the cache, and deploying a calculation basic block including a third instruction used to calculate the occupied space of the target data before the initial basic block of the control flow graph. The next basic block of the calculation basic block is the initial basic block of the control flow graph and the initial basic block of the control flow graph including the second instruction. The occupied space of the target data is used to indicate that the initial basic block of the control flow graph is executed, or the initial basic block of the control flow graph including the second instruction is executed.

[0030] In the present application, in the case that the size of the target data is unknown, the size of the target data is calculated in advance, and it is determined whether the target data is written into the cache in advance, so as to balance the efficiency of data access and the capacity of the cache, and to improve the overall efficiency of data access while meeting the capacity of the cache.

[0031] In a second aspect, the present application provides a data prefetching device. The data prefetching device is applied to a computer system or a computing device supporting the computer system to implement the data prefetching method. The data prefetching device includes various modules for executing the data prefetching method in the first aspect or any optional implementation of the first aspect. For example, the data prefetching device includes an obtaining module, a determining module and a writing module. Wherein,

[0032] The obtaining module is used to obtain the control flow graph and the events of the basic blocks in the control flow graph recorded by the performance monitoring unit (PMU).

[0033] The determining module is used to determine a target path of the plurality of paths in the control flow graph according to the number of occurrences of the events of the basic blocks in the control flow graph, the sum of the number of occurrences of all the events of the basic blocks in the target path being greater than or equal to a first threshold value.

[0034] The writing module is used to write target data corresponding to a first basic block into a cache before the first basic block in the control flow graph is executed, the target data being data required to be accessed by the first basic block, and the target data being data required to be accessed when the target path is executed and having a frequency of access or a frequency of access greater than or equal to a second threshold value.

[0035] In a possible implementation, the sum of the number of occurrences of all the events of the basic blocks included in each path of the plurality of paths is a candidate value, and one of the plurality of candidate values of the plurality of paths is the first threshold value.

[0036] In a possible implementation, the first threshold value is a maximum value in the plurality of candidate values.

[0037] In a possible implementation, the determining module is specifically configured to search for a plurality of paths in the control flow graph starting from the initial basic block; an exit basic block of each path in the plurality of paths does not belong to a basic block in the control flow graph that is in a nested loop; determine a sum of occurrence times of the event of the basic blocks included in each path in the plurality of paths; and determine a target path in which the sum of occurrence times of the event is greater than or equal to the first threshold value from the plurality of paths according to the sum of occurrence times of the event of the basic blocks included in each path in the plurality of paths.

[0038] In a possible implementation, the writing module is specifically configured to parse source code corresponding to the target path to obtain target data, and then determine a first basic block with the highest priority from the plurality of basic blocks included in the control flow graph, so as to write the target data into the cache before executing a first instruction included in the first basic block. A reuse score of the target data is a reuse score of M pieces of data in the source code, where M is an integer greater than or equal to 1. The reuse score of the target data is used to indicate a reuse degree of the target data, and each basic block in the plurality of basic blocks includes a first instruction for accessing the target data.

[0039] In a possible implementation, the priority of the basic block is determined according to one or more of the following indicators: a loop depth indicated by the first instruction, an access rule of the first instruction, an occurrence time of an event of a basic block in which the first instruction is located, a line number of the first instruction in the source code, whether the first instruction is used to indicate reading or writing of the target data, and whether the first instruction is vectorized memory access or scalar memory access.

[0040] In a possible implementation, the priority of the basic block is determined according to one or more of the following indicators: a loop depth indicated by the first instruction, an access rule of the first instruction, an occurrence time of an event of a basic block in which the first instruction is located, a line number of the first instruction in the source code, whether the first instruction is used to indicate reading or writing of the target data, and whether the first instruction is vectorized memory access or scalar memory access.

[0041] In a possible implementation, the basic block in the control flow graph includes a basic block in a nested loop.

[0042] In a possible implementation, the apparatus further includes a dynamic prediction module; the dynamic prediction module is configured to deploy, in the first basic block, a second instruction for indicating that the target data is written into the cache; and deploy, before the initial basic block of the control flow graph, a calculation basic block including a third instruction for calculating the occupation space of the target data. Wherein, the next basic block of the calculation basic block is the initial basic block of the control flow graph and the initial basic block of the control flow graph including the second instruction; the occupation space of the target data is used to indicate that the initial basic block of the control flow graph is executed, or the initial basic block of the control flow graph including the second instruction is executed.

[0043] In a possible implementation, each path in the plurality of paths includes all the basic blocks in the control flow graph that are within a nested loop.

[0044] In a third aspect, the present application provides a chip, including: a processor and a power supply circuit; the power supply circuit is configured to supply power for the processor, and the processor is configured to execute the method in the first aspect or any possible implementation manner of the first aspect.

[0045] In a fourth aspect, the present application provides a computing device, including a memory and a processor, the memory is configured to store computer instructions; the processor is configured to execute the computer instructions to implement the method in the first aspect or any possible implementation manner of the first aspect.

[0046] In a fifth aspect, the present application provides a computer readable storage medium, the storage medium stores computer programs or instructions; when the computer programs or instructions are executed by a processing device, the method in the first aspect or any possible implementation manner of the first aspect is implemented.

[0047] In a sixth aspect, the present application provides a computer program product, the computer program product includes computer programs or instructions; when the computer programs or instructions are executed by a processing device, the method in the first aspect or any possible implementation manner of the first aspect is implemented.

[0048] The beneficial effects of the above second aspect to sixth aspect can refer to the first aspect or any possible implementation manner of the first aspect, which will not be repeated here. On the basis of the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 It is a schematic diagram of a cache structure;

[0050] Figure 2 It is a schematic diagram of an HBM allocation system architecture;

[0051] Figure 3 It is a schematic diagram of a computer system provided by the present application;

[0052] Figure 4 A flowchart of a compiling method provided by the present application is shown in the following figure;

[0053] Figure 5 A flowchart of a data processing method provided by the present application is shown in the following figure;

[0054] Figure 6 A flowchart of a target path determining method provided by the present application is shown in the following figure;

[0055] Figure 7 A flowchart of a method for writing target data into cache provided by the present application is shown in the following figure;

[0056] Figure 8 A flowchart of a priority determining method provided by the present application is shown in the following figure;

[0057] Figure 9 A complete flowchart of a data prefetching method provided by the present application is shown in the following figure;

[0058] Figure 10 A conversion diagram of a directed acyclic graph provided by the present application is shown in the following figure;

[0059] Figure 11 A control flow diagram of dynamic multi-branch insertion provided by the present application is shown in the following figure;

[0060] Figure 12 A structure diagram of a data prefetching device provided by the present application is shown in the following figure; Figure 1 ;

[0061] Figure 13 A structure diagram of a data prefetching device provided by the present application is shown in the following figure; Figure 2 ;

[0062] Figure 14 A structure diagram of a computing device provided by the present application is shown in the following figure. DETAILED DESCRIPTION

[0063] For the convenience of understanding, first, the technical terms involved in the present application are introduced.

[0064] Source code, source code is a text file written by programmers. Programmers write code as source code in human-readable languages such as Java, C++, C#, etc. Source code is written according to the conventions and rules of a certain language. Source code is human-readable, but computing devices (such as processors) cannot recognize it.

[0065] Executable code, executable code is executable by computing devices, which contains binary instructions that computing devices can recognize. Among them, the compiler can convert the source code into executable code.

[0066] PMU, PMU is a kind of hardware device, usually set in the processor. PMU is mainly used to track, count some underlying hardware events of the processor, such as processor related events (such as the number of times of execution of each instruction, the number of processor captured exception conditions, the number of clock cycles of the processor, etc.), cache related events (such as the number of times of access to each level of cache in the cache, the number of times of miss in the cache, etc.). These events can represent the behavior of the processor in the process of executing executable code.

[0067] The source code needs to meet certain writing specifications in the writing process, such as object-oriented programming (OOP), procedure-oriented programming (POP), etc. The program languages of object-oriented programming include C++, Java, C#, etc. The program languages of procedure-oriented programming include fortran, C, etc.

[0068] The source code written in a certain writing specification can include function / method, field / variable and other code elements.

[0069] Function / method refers to a kind of subprogram in a class. A method is usually composed of a series of statements and can complete a function. In the writing specification of object-oriented programming, it is called method, and in the writing specification of procedure-oriented programming, it is called function. In the embodiments of the present application, it is collectively referred to as "function".

[0070] Loop is a kind of program that may be included in the function. Loop refers to a program that needs to be executed multiple times when the condition is met, and can jump out of the loop when the condition is not met. Alternatively, loop refers to a program that needs to be executed multiple times when the condition is not met, and can jump out of the loop when the condition is met.

[0071] Field / variable stores some data, such as integer and character, string, hash table, pointer, etc. Field / variable can be used as the data mentioned in the embodiments of the present application, such as target data.

[0072] Memory access refers to the operation of data access in the source code. Among them, memory access can include vectorized memory access and scalarized memory access. Vectorized memory access refers to the operation of multiple memory addresses at a time, that is, the processor can obtain multiple data at a time, and then perform parallel calculation. Scalarized memory access refers to accessing memory addresses one by one, that is, the processor obtains single data in sequence and calculates one by one.

[0073] The memory access node includes a base address and an index. The memory access node can be represented as [A, index], where A represents the base address, which is the address used to access data, and index indicates the index, which represents the position of the data to be accessed at the base address.

[0074] The memory access instruction is an instruction including a memory access node. There are many memory access instructions. Here, one of them is listed. For example, in the assignment instruction, the data at the position indicated by the index I at the base address A is assigned to the variable P. In the execution of the assignment instruction, the data at the position indicated by the index I at the base address A needs to be read to determine the value assigned to the variable P. The memory access nodes included in memory access instructions with different semantics are different, and the number of memory access nodes included is also different.

[0075] The control flow graph (CFG) can also be referred to as a control flow diagram, and is an abstract representation of a process or program, which is an abstract data structure used in a compiler. It is maintained internally by the compiler and represents all paths that will be traversed during the execution of a program. It represents the possible flow direction of all basic blocks in a process in the form of a graph, and can also reflect the real-time execution process of a process.

[0076] The basic block (BB) can also be referred to as a base block. For any function, the compiler splits the function into multiple independent instruction sets according to predetermined standards. Any independent instruction set is a BB. The first BB executed in a function is the entry BB or initial BB of the function, and the last BB executed is the exit BB or exit BB of the function. The first BB executed in a loop is the header basic block (header BB) of the loop. Each BB represents a continuous code segment in the program without branches or jumps.

[0077] A function has a position where it can jump to multiple execution paths. This position is usually the position of a statement such as an if statement, a while statement, an else statement, etc. that has a judgment condition. When executing at this position, only one of the multiple execution paths can be executed each time. Any execution path can be referred to as a branch. The probability of executing any branch is referred to as the branch probability.

[0078] The branch probability is related to the statement of the judging condition at the position. In the present application, the compiler can analyze the specific semantics of the statement of the judging condition at the position to estimate the branch probability of each branch. For example, the value range of x is an integer from 1 to 10, and if the statement of the judging condition at the position is if x>8, it can be indicated that this position will be divided into two branches, one branch is x>8 branch, and one branch is x<8 branch, and based on the value range of x, the branch probability of x>8 branch is 20%, and the branch probability of x<8 branch is 80%.

[0079] High performance computing (HPC) refers to using supercomputers, parallel computing clusters or other high-performance computing systems to perform large-scale and complex computing tasks.

[0080] Cache is used to realize high-speed data buffering between a processor and a memory. The cache in the processor can include a first-level cache (L1 cache or L1), a second-level cache (L2 cache or L2), a third-level cache (L3 cache or L3), and a last-level cache (LLC). Generally, the read-write speed of the L1 cache, the L2 cache, the L3 cache and the LLC decreases in turn, and the cache capacity gradually increases. The LLC is the last-level cache between the processor core and the main memory. The LLC is the farthest cache from the processor core and has the largest capacity and the largest access delay among the above-mentioned caches. For example, the LLC can include a high bandwidth memory (HBM).

[0081] As an on-chip resource, the HBM has a small capacity due to the limitation of hardware manufacturing process. Compared with the case of using a static random-access memory (SRAM) as the last-level cache, the HBM supports large-capacity data storage and can effectively expand the cache capacity to 200 GB (Gigabyte). The HBM also has a higher bandwidth, which can effectively improve the data read-write rate of the cache and is more suitable for concurrent access scenarios. The HBM is usually applied to the fields of high-performance computing, graphics processing unit (GPU), artificial intelligence (AI), such as AI servers, to meet the growing demand for memory bandwidth and energy efficiency.

[0082] Footprint, also known as memory footprint, represents the maximum cache capacity used by a program during running.

[0083] Kernel function represents a hot function or code snippet (usually a code snippet with multiple nested loops).

[0084] In HPC scenarios, the scale of HPC operations is typically large. While horizontal scaling is used to distribute HPC tasks to different node processes as much as possible in parallel, the workload within a single node remains relatively large to minimize cross-node communication and synchronization overhead. Therefore, maximizing the computing power of a single node is a key consideration in the architectural design of HPC hardware platforms.

[0085] The common data access patterns in HPC services are usually as follows: (1) multiple iterative operations on data over a large range.

[0086] (2) Multiple threads or processes concurrently access hot data on a single node.

[0087] For data iteration, hardware cache technology is often used to ensure that hot data resides as close to the chip as possible. This effectively reduces the latency increase caused by bursty memory accesses without user intervention. For concurrent data access, HBM is often used to mitigate the memory bandwidth bottleneck caused by high concurrency, thereby reducing the latency degradation caused by memory access. The HBM cache, which combines these two technologies, was subsequently developed to effectively reduce the latency degradation caused by high concurrent access to hot data without user intervention.

[0088] In traditional central processing unit (CPU) architectures, memory access from dynamic random-access memory (DRAM) to the processor core typically passes through multiple levels of cache (such as L1, L2, and L3). In CPU architectures designed for high-performance computing or high-memory access, HBM is often added to accelerate memory access. HBM can typically be configured in three usage modes: flat mode, which treats the entire HBM as DDR; cache mode; and hybrid mode, which splits the HBM into DDR and cache as needed. In cache mode, the HBM is configured as LLC.

[0089] like Figure 1 As shown, Figure 1A schematic diagram of a cache structure. L1, L2 and unified cache (UC) are arranged from near to far according to the distance from the processor core, and the DRAM is connected with the UC. When the HBM is in the cache mode, the UC is connected with the HBM, serving as the LLC of the processor core. The register is closest to the processor core.

[0090] At present, in order to fully utilize the storage resources of the HBM, the hot data in the source code, i.e. the frequently accessed data, is often identified. Then, the frequently accessed data is stored in the HBM, and the processor core obtains the frequently accessed data from the closer HBM, so as to improve the data access efficiency.

[0091] As follows, a scheme of identifying the hot data in the source code and storing the hot data in the HBM is provided. That is, the source code is statically analyzed by the compiler, the target scene (i.e. the memory access intensive within the function range and the Footprint greater than the capacity of the HBM) is identified, the data reuse situation in the target scene is modeled and analyzed, the cold and hot scores or the order of the data are determined, and then the hot data with higher cold and hot scores (frequently accessed) or in front of the order (arranged in descending order) is generated HBM prefetch instruction (such as hint instruction) to inform the HBM to improve the residence level of the hot data in the HBM, so as to improve the access efficiency of the processor core to the data.

[0092] As shown in Figure 2 , a schematic diagram of the architecture of an HBM allocation system is provided. Figure 2

[0093] ①: The compiler can determine the memory access intensive kernel in the source code by static analysis.

[0094] ②: The probability branch path identification, i.e. the combination and ordering of the kernel, the execution probability of each kernel determined by static analysis, and the kernel with high probability of execution within the function are connected in series. The data (variables) with higher reuse degree in the kernel after connection are determined, such as the original variable group is induced by tracing the memory access node in the kernel, and the reuse degree of each variable group is calculated.

[0095] ③: The hint instruction is generated to indicate that the data with higher reuse degree is resident in the HBM. The hint instruction is inserted at the first memory access node (i.e. the first appearance position) of the data with higher reuse degree, and then when the processor core executes the corresponding hint instruction, the HBM is prompted to reside the data with higher reuse degree. The hint instruction can be vectorized or non-vectorized.

[0096] Among them, the compiler can also obtain the events recorded by the PMU in the CPU, and update the above prediction probability according to the event feedback of the PMU record.​Figure 2 The profile in the PMU records events of the PMU.

[0097] However, in the static branch prediction process, the determined execution path is inaccurate in the case of local branch probability being equal or zero, and the data determined according to the execution path is not necessarily the data frequently called by the processor, and storing the determined data to the cache will result in low overall efficiency of the processor in data access.

[0098] Therefore, the present application provides a data prefetching method. The data prefetching method comprises: obtaining a control flow graph, and events of basic blocks in the control flow graph recorded by a PMU, determining a target path of a plurality of paths in the control flow graph according to the number of occurrences of the events of the basic blocks in the control flow graph, and writing target data corresponding to a first basic block in the control flow graph into a buffer before the first basic block is executed. The sum of the number of occurrences of all events of the basic blocks in the target path is greater than or equal to a first threshold, and the target data is data with a frequency of access or a frequency of access greater than or equal to a second threshold in data accessed when the target path is executed.

[0099] In the present application, the PMU records the number of occurrences of various events when the processor executes instructions in the basic blocks, and the more the number of occurrences of the events of a basic block recorded by the PMU, the more frequently the instructions in the basic block are executed. Then, according to the number of occurrences of the events of the basic blocks recorded by the PMU, a target path including a plurality of basic blocks with a sum of the number of occurrences of the events greater than or equal to a first threshold is determined from a plurality of paths included in the control flow graph, the probability of the target path being frequently executed is higher, and the frequency of access or the frequency of access of the target data determined according to the target path is also higher. Therefore, the data (target data) accessed when the first basic block in the control flow graph is executed is written into the cache, which can improve the overall efficiency of the processor in data access when the control flow graph is executed.

[0100] The above method can be applied to Figure 3 The computer system (or computing device) shown, Figure 3 A computer system provided by the present application is shown. The computer system comprises a processor 100 and a cache 200. Optionally, it can also comprise a memory 300.

[0101] The memory 300 can be understood as the memory of the processor 100, and the processor 100 can comprise a compiler 110 and a PMU 120.

[0102] It is worth noting that the present application does not limit the specific form of the system, and the system can be a computer chip, such as a system on chip (SoC).

[0103] Figure 3 The processor 100 shown in FIG. 1 includes a compiler 110 with compiling capability. The compiler 110 can be a software program running on the processor or a hardware module on the processor. Here, only an example is described in which the compiler 110 is part of the processor 100.

[0104] In some possible embodiments, the compiler 110 can also be a software program or a hardware module independent of the processor 100, such as a software program or a hardware module deployed on another processor. In this scenario, the compiler 110 can send executable code to the processor 100 after compiling the executable code, so that the processor 100 executes the executable code.

[0105] The following content is only described with reference to Figure 3 The computer system architecture shown in FIG. 1 is used as an example to describe the functions of the compiler, the processor, and the PMU in the scenario in which the compiler is independent of the processor. Figure 3 The functions of the compiler 110, the processor 100, and the PMU 120 shown in the computer system architecture shown in FIG. 1 are consistent.

[0106] The present application does not limit the type of the processor 100, and any processor 100 capable of executing executable code is applicable to the embodiments of the present application. The processor 100 can be a CPU, a GPU, or the like. The processor 100 can also be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), which can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0107] In a possible embodiment, the cache 200 in the computer system comprises a multi-level cache. Among them, the last level cache (LLC) 210 in the multi-level cache can be an HBM. The "multi-level cache" divides the cache 200 into multiple levels, the cache level closer to the core of the processor 100 is smaller, the read-write speed is faster, and the capacity is relatively small. That is, the LLC 210 in the multi-level cache is the cache farthest from the core of the processor 100 and has the largest capacity in the multi-level cache. The LLC 210 can be used to cache data of the interaction between the processor 100 and the memory. From the perspective of the processor 100, the LLC 210 can be understood as the cache of the processor 100. From the perspective of the memory, the LLC can also be used as the cache of the memory.

[0108] Compared with the case where the SRAM is used as the LLC 210, the HBM supports large-capacity data storage, which can effectively expand the capacity of the cache 200. The HBM also has higher bandwidth, which can effectively improve the data read-write speed of the cache 200 and is more suitable for concurrent access scenarios.

[0109] In this application, the compiler 110 also has an analysis capability, which can determine the target data that needs to be frequently accessed or frequently accessed by analyzing the source code, generate a hint instruction for the target data in the source code, and insert the hint instruction into the corresponding position in the source code. The hint instruction is used to instruct the target data to be persistently stored in the LLC. Then, the compiler 110 compiles the source code inserted with the hint instruction to generate executable code that needs to be executed by the processor 100.

[0110] It is worth noting that the present embodiment takes the hint instruction indicating that the target data needs to be persistently stored in the LLC 210 as an example for illustration. In other embodiments, the hint instruction can also indicate that the target data is persistently stored in other levels of cache included in the cache 200, such as L3.

[0111] The processor 100 can obtain the executable code compiled by the compiler 110 and execute the executable code. In the process of executing the executable code, the processor 100 can migrate the data to be processed from the memory 300 to the cache 200, and can also store the processed data in the cache 200. When the processor 100 executes the code corresponding to the hint instruction in the executable code, the processor 100 can initiate a residence instruction to the LLC 210, and the residence instruction is used to instruct the LLC 210 to persistently store the target data.

[0112] In the process of the processor 100 executing the executable code, the PMU 120 can monitor the process, record events related to the processor 100 (especially events in the process of the processor 100 executing the executable code), and events related to the cache 200.

[0113] The compiler 110 can call the events recorded by the PMU 120, update the determined target data, add hint instructions in the source code for the updated target data, and compile to generate new executable code. The processor 100 can acquire and execute the new executable code.

[0114] The memory 300 is generally used to store computer program codes and the like required by the processor 100 to execute. In the embodiments of the present application, the memory 300 can store source codes required to be compiled by the compiler 110, and the compiler 110 can call the source codes from the memory 300 to analyze and compile the source codes. The memory 300 can also store executable codes generated after the source codes are compiled. After the compiler 110 compiles to obtain the executable codes, the compiler 110 can store the executable codes in the memory 300, and the processor 100 can call and execute the executable codes.

[0115] The memory 300 is generally a DRAM as the memory 300. In addition to the DRAM, the memory 300 can also be other random access memories such as SRAM and the like. In addition, the memory 300 can also be a read only memory (ROM). For the read only memory, for example, it can be a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), and the like. The memory 300 can also be a dual in-line memory module (DIMM), that is, a module composed of dynamic random access memories (DRAM), and can also be a solid state disk (SSD) and the like. The memory 300 can also be a combination of the above-mentioned memories, and the embodiments of the present application do not limit the number and type of the memory 300.

[0116] In the above description, the target data is obtained by the compiler 110 analyzing the loops included in the function in the source code, and can be all or part of the data required to be called by the function in the source code during running. In actual application, the programmer can mark some frequently called data in the process of writing the source code. The compiler 110 can recognize the mark and take the data with the mark as the target data.

[0117] For the compiler 110 shown in the above description, a method for the compiler 110 to compile the source code is provided as follows. Figure 3

[0118] As shown in the above description, Figure 4 Figure 4 A flowchart of a compiling method provided in the present application is shown in the above description. Figure 4 The compiling method shown in the above description includes the following steps ①-⑦.

[0119] ①: The compiler 110 receives the profile sent by the core in the processor 100, such as the events collected by the PMU.

[0120] ②: The compiler 110 performs lexical / syntactic analysis on the source code according to the compiling configuration file to obtain a syntax tree or other intermediate representation form, such as an abstract syntax tree or intermediate code (IR).

[0121] ③: The compiler 110 can convert the above-mentioned intermediate representation form to obtain a control flow graph (CFG) or IR in the form of static single assignment (SSA).

[0122] ④: The compiler 110 performs static single assignment optimization on the CFG or IR in the form of SSA to obtain the optimized CFG or IR in the form of SSA.

[0123] ⑤: The compiler 110 performs inter-procedural optimization on the optimized CFG or IR in the form of SSA to obtain the further optimized CFG or IR in the form of SSA. The aforementioned SSA optimization can include dead code elimination or constant propagation, etc.

[0124] ⑥: The compiler 110 performs register transfer language (RTL) optimization according to the optimized CFG or IR in the form of SSA to obtain the executable code corresponding to the source code.

[0125] ​​⑦: The processor core in the processor 100 executes the executable code. The executable code includes code corresponding to the hint instruction. When the processor core executes the code corresponding to the hint instruction, a resident instruction is initiated to the LLC 210. The resident instruction is used to instruct the LLC 210 to store the target data, that is, to persistently store the target data in the LLC 210.

[0126] The implementation of the embodiment of the present application is described in detail below with reference to the accompanying drawings.

[0127] As shown in the flowchart of the data processing method provided by the present application. Figure 5 Figure 5 The method shown can be executed by the processor 100 in the computer system 1000, specifically, by the compiler 110, and the content shown can be applied in the step ④ shown. Figure 5 Figure 3 The method shown can be executed by the processor 100 in the computer system 1000, specifically, by the compiler 110, and the content shown can be applied in the step ④ shown. Figure 5 Figure 4 The method shown can be executed by the processor 100 in the computer system 1000, specifically, by the compiler 110, and the content shown can be applied in the step ④ shown. Figure 5 The method shown can be executed by the processor 100 in the computer system 1000, specifically, by the compiler 110, and the content shown can be applied in the step ④ shown.

[0128] S510, the processor 100 acquires the control flow graph.

[0129] In a possible implementation, the processor 100 acquires the intermediate data generated by the compiler 110 in the compilation phase, such as the control flow graph.

[0130] In a possible case, the intermediate data generated by the compiler 110 in the compilation phase can be stored in the storage 300 or the cache 200. Therefore, the processor 100 can acquire the control flow graph from the storage 300 or the cache 200.

[0131] The control flow graph shows the control flow of the program in the form of a graph, so that the control logic of the program is more intuitive and visible, facilitating understanding and analysis.

[0132] As shown in the flowchart of the data processing method provided by the present application. Figure 5 The control flow graph acquired by the processor 100 includes: BB-6, BB-3, BB-4, BB-7, and BB-5. Among them, BB-3, BB-4, and BB-7 are in loop 1, BB-3 can exit to BB-5, and BB-4 can exit to BB-5.

[0133] In a possible case, the basic blocks in the control flow graph include basic blocks in nested loops.

[0134] ​​​In the present application, since the hotspot function or code segment is usually a nested loop structure, the processor 100 determines the control flow graph including the nested loop (structure) from all the control flow graphs, avoiding the subsequent target data identification and data prefetching of all the control flow graphs, and improving the efficiency of writing the data accessed by the hotspot function or code segment into the cache.

[0135] Further, the control flow graph obtained by the present application is a control flow graph including a nested loop, and therefore the basic block in the control flow graph obtained by the present application includes a basic block in the nested loop.

[0136] In a possible implementation, the processor 100 obtains the control flow graph, including: the processor 100 identifies the nested loop (structure) in the source code, and then determines the basic block included by the nested loop, and determines the control flow graph including the basic block included by the nested loop from all the control flow graphs.

[0137] For the content of the processor 100 identifying the nested loop in the source code, refer to the description under step ① in the following Figure 9 , which will not be repeated here.

[0138] S520, the processor 100 obtains the events of the basic block in the control flow graph recorded by the PMU.

[0139] The processor 100 records various events related to the performance of the processor, such as instruction execution events, cache access events, branch prediction events, memory access events, pipeline related events, cache events, instruction cycle count, timer events, when executing the executable code corresponding to the program or code segment.

[0140] In a possible case, since the basic block is a continuous code segment without branches or jumps in the program, when the PMU records, the type and occurrence number (also called collection number) of the events triggered by the basic block when executing the executable code corresponding to the basic block are recorded.

[0141] For example, the type and sending number of the events triggered by the processor 100 when executing the executable code corresponding to the basic block a are: instruction execution events-3 times, cache access events-5 times, memory access events-2 times, etc.

[0142] It is worth noting that the types of events recorded by the PMU for basic blocks in the control flow graph by the processor 100 may be one or more of the aforementioned events, and this application does not limit this. For example, in one possible embodiment, the events recorded by the PMU for basic blocks in the control flow graph are cache access events and memory access events. In another possible embodiment, the events recorded by the PMU for basic blocks in the control flow graph are only cache access events.

[0143] In a possible implementation, the processor 100 obtains the events of the basic blocks in the control flow graph recorded by the PMU, including: the processor 100 obtains the events of the basic blocks in the control flow graph recorded by the PMU from the memory 300 .

[0144] For example, the memory 300 is a memory connected to the processor 100 .

[0145] For example, since the PMU usually sends recorded events to a memory connected to the processor 100 for storage, the processor 100 can obtain the events of the basic blocks stored in the memory through a bus or other interfaces.

[0146] It is worth noting that, to distinguish events between basic blocks in the control flow graph recorded by the PMU, processor 100 can determine basic block boundaries by detecting changes in instruction addresses during program execution. When program execution reaches the entry point of a basic block, the address of the instruction can be recorded and checked for address changes in subsequent event records. If an address change is detected, it indicates that execution has entered a different basic block. Therefore, the events recorded before the address change can be used as the type and occurrence count of the event for the corresponding basic block.

[0147] like Figure 5 As shown, the processor 100 obtains that the number of occurrences of the event of BB-6 is 20 times, the number of occurrences of the event of BB-3 is 34 times, the number of occurrences of the event of BB-4 is 37 times, the number of occurrences of the event of BB-7 is 23 times, and the number of occurrences of the event of BB-5 is 9 times.

[0148] S530: The processor 100 determines target paths of multiple paths in the control flow graph according to the number of occurrences of events of the basic blocks in the control flow graph.

[0149] The sum of the occurrence times of all events of the basic blocks in the target path is greater than or equal to a first threshold.

[0150] In a possible implementation, the processor 100 searches a plurality of paths in the control flow graph, and then determines the sum of the occurrence times of all events of all basic blocks included in each path in the plurality of paths, so as to determine a target path according to the sum of the occurrence times of all events of all basic blocks included in each path in the plurality of paths, wherein the sum of the occurrence times of all events of all basic blocks included in each path in the plurality of paths is greater than or equal to the first threshold value.

[0151] In a possible example, the processor 100 starts searching from an initial basic block of the control flow graph.

[0152] For detailed description of the above possible implementation, refer to the following Figure 6 The content is not described here.

[0153] In a possible case, each path in the plurality of paths includes all basic blocks in the nested loop in the control flow graph.

[0154] Since each path includes all basic blocks in the nested loop, when the processor 100 searches the paths in the control flow graph, the processor 100 does not exit from the last basic block in the nested loop to the exit basic block in the control flow graph, that is, all basic blocks in the nested loop are completely included in each path. Therefore, each path is a longer path (including a path with a larger number of basic blocks), and the shorter path is avoided to be included in the sum of the occurrence times of all events, and the efficiency of determining the target path is improved.

[0155] As Figure 5 As shown in the figure, the processor 100 can obtain a path BB-6→BB-3→BB-4→BB-7→BB-5 by searching the control flow graph. The processor 100 avoids the paths BB-6→BB-3→BB-5 and BB-6→BB-3→BB-4→BB-5 when searching, and then improves the efficiency of determining the target path according to the sum of the occurrence times of all basic blocks included in the path.

[0156] It is worth noting that each path includes all basic blocks in the nested loop only once, that is, there is no loop. For example, BB-6→BB-3→BB-4→BB-7→BB-5, there is no path BB-6→BB-3→BB-4→BB-7→BB-3→BB-5.

[0157] In a possible case, the sum of the occurrence times of all events of the basic blocks included in each path in the plurality of paths is a candidate value, and one of the plurality of candidate values of the plurality of paths is the first threshold value.

[0158] For example, the processor 100 selects a path corresponding to a larger one of the candidate value a and the candidate value b as the target path.

[0159] If the candidate value a is larger than the candidate value b and the candidate value b is selected as the first threshold, the processor 100 can select a path from the path a and the path b as the target path, or select the path a with a larger sum of occurrence times as the target path.

[0160] If the candidate value a is larger than the candidate value b and the candidate value a is selected as the first threshold, the processor 100 can select the path a with a larger sum of occurrence times as the target path.

[0161] In a possible example, the first threshold is a maximum value in the plurality of candidate values. The processor 100 determines the target path by selecting the maximum value in the plurality of candidate values as the first threshold, and determines the target data by using the target path. The target data is the most frequently accessed data. Thus, writing the target data into the cache can improve the overall efficiency of data access when the processor 100 executes according to the control flow graph.

[0162] For example, the processor 100 selects a path corresponding to a larger one of the candidate value a and the candidate value b as the target path.

[0163] For the content of determining the target path of the plurality of paths in the control flow graph according to the occurrence times of the events of the basic blocks in the control flow graph by the processor 100, a possible example is provided as follows. Figure 6 Figure 6 FIG. 5 is a flowchart of a target path determination method provided in the present application. The S530 in FIG. 5 can include the following steps S610-S630. Figure 5 S610, the processor 100 searches a plurality of paths in the control flow graph starting from an initial basic block.

[0164] Each of the plurality of paths has an exit basic block that does not belong to a basic block in the control flow graph in a nested loop.

[0165] For example, the processor 100 uses a depth first search (DFS) to search from the initial basic block in the control flow graph to obtain the plurality of paths.

[0166]

[0167] ​​In the above search process, a series of search conditions (also referred to as trace-back conditions) are provided. For example, in the DFS process performed by the processor 100, the search does not proceed from the exit edge of the nested loop when the searched basic block does not return to the header BB of the nested loop. When the searched basic block returns to the header BB, the processor 100 traverses the exit edge of the nested loop to perform trace-back.

[0168] Generally, the exit basic block is not in the nested loop but outside the nested loop. This is because the exit basic block usually includes instructions for exiting the loop.

[0169] For example, a loop structure generally includes an entry block, one or more loop body blocks, and a loop exit condition block. Correspondingly, an exit basic block is provided outside the loop structure. The exit basic block is the first basic block after exiting the loop or the last basic block of the control flow graph.

[0170] It is worth noting that if the above loop is a nested loop, an exit basic block is provided outside each loop structure. In this case, the exit basic block of each path does not belong to the basic block in the nested loop in the control flow graph, which means that the last basic block of each path does not belong to the basic block in the nested loop. Generally, the last basic block is usually used to indicate the return of a function or the end of a program.

[0171] For a detailed description of the above search process, refer to the content shown in step ③ of the following Figure 9 , which will not be described here.

[0172] S620, the processor 100 determines the sum of the occurrence times of the events of the basic blocks included in each path of the plurality of paths.

[0173] For example, the processor 100 determines the sum of the occurrence times of all the basic blocks according to all the basic blocks included in each path of the plurality of paths.

[0174] For example, the path a is BB-6→BB-3→BB-4→BB-7→BB-5, and the processor 100 adds the occurrence times of the events corresponding to BB-6, BB-3, BB-4, BB-7, and BB-5 respectively to obtain the sum of the occurrence times.

[0175] S630, the processor 100 determines a target path from the plurality of paths according to the sum of the occurrence times of the events of the basic blocks included in each path of the plurality of paths, wherein the sum of the occurrence times of the events of the target path is greater than or equal to a first threshold value.

[0176] For example, the processor 100 determines the target path by comparing the sum of the occurrence times of the events of the basic blocks included in each path in the plurality of paths, and determining the target path in which the sum of the occurrence times of the events is greater than or equal to the first threshold value.

[0177] For example, as shown in Figure 6 The control flow graph searched by the processor 100 includes BB-6, BB-3, BB-4, BB-7, and BB-5. Among them, BB-3, BB-4, and BB-7 are in loop1, BB-3 can exit to BB-5, and BB-4 can exit to BB-5.

[0178] The processor 100 searches the above control flow graph to obtain path a: BB-6→BB-3→BB-4→BB-7→BB-5, path b: BB-6→BB-3→BB-5, and path c: BB-6→BB-3→BB-4→BB-5. Then, the sum of the occurrence times of the events included in path a is determined to be 123 times, the sum of the occurrence times of the events included in path b is determined to be 63 times, and the sum of the occurrence times of the events included in path c is determined to be 100 times. Thus, path a with the largest sum of occurrence times is determined as the target path.

[0179] Please continue to refer to the method shown in Figure 5 , Figure 5 The method further includes the following step S540.

[0180] S540, the processor 100 writes the target data corresponding to the first basic block into the cache before executing the first basic block in the control flow graph.

[0181] The target data is data required to be accessed by the execution of the first basic block, and the target data is data with a frequency of access or a frequency of access greater than or equal to a second threshold value among data required to be accessed when the target path is executed.

[0182] In one possible case, the processor 100 writes the target data corresponding to the first basic block into the cache before the processor 100 executes all instructions included in the first basic block according to the execution of the control flow graph.

[0183] For example, the processor 100 can insert a hint instruction in the first basic block, write the target data into the cache when the processor executes the hint instruction, and perform persistent storage, such as storing for 1 minute or 10 minutes. The specific storage time is not limited in the present application. The hint instruction is used to instruct to write the target data into the cache and perform persistent storage.

[0184] In one possible example, the processor 100 can insert a hint instruction in the adjacent line of the first instruction that accesses the target data included in the first basic block.

[0185] For example, the processor 100 can insert a hint instruction in a line before the first instruction.

[0186] For the content of writing the target data corresponding to the first basic block into the cache before the processor 100 executes the first basic block in the target path, a possible embodiment is provided as follows. As shown in Figure 7 Figure 7 A flowchart of a method for writing target data into a cache provided in the present application. The S540 in the above Figure 5 may include the following steps S710-S730.

[0187] S710, the processor 100 parses the source code corresponding to the target path to obtain the target data.

[0188] Among them, the multiplexing score of the target data is the M front in the multiplexing score of the plurality of data indicated by the source code, M is an integer greater than or equal to 1, and the multiplexing score of the target data is used to indicate the multiplexing degree of the target data.

[0189] In a possible implementation, the multiplexing score of the data can be determined according to one or more of the following: data multiplexing mode (such as parallel multiplexing or serial multiplexing), memory access regularity (such as regular or random), read-write type (such as read or write), footprint limit.

[0190] For the content of determining the multiplexing score of the data according to the data multiplexing mode, the memory access regularity, the read-write type, and the footprint limit, the content of the following formulas 2-11 is referred to, which is not described here.

[0191] In a possible implementation, the processor 100 parses the source code corresponding to the target path to obtain the target data, including: the processor 100 determines the multiplexing score of the plurality of data indicated by the source code corresponding to the target path, sorts the multiplexing score of the plurality of data, and takes the M front data as the target data.

[0192] For example, the source code corresponding to the target path shows that data a, data b and data c need to be accessed, the processor 100 determines the multiplexing scores of data a, data b and data c to be 100, 150 and 120 respectively, and then the processor 100 sorts the multiplexing scores of data a, data b and data c to obtain the sorting result: data b-150, data c-120, data a-100, so that the processor 100 takes the data (data b and data c) in the front two as the target data.

[0193] For example, in the above example, if M is 1, the data (data b) in the front one is taken as the target data.

[0194] ​In another possible implementation, the processor 100 parses the source code corresponding to the target path to obtain the target data, including: the processor 100 determines the multiplexing scores of a plurality of data indicated by the source code corresponding to the target path, and determines M data with higher multiplexing scores in the plurality of data as the target data.

[0195] The process in which the processor 100 obtains the target data in the above implementation is only an optional example provided by the embodiments of the present application, and should not be understood as a limitation to the present application. In another possible implementation, the processor 100 determines, according to the multiplexing scores of a plurality of data, data corresponding to a preset ratio (such as 1% or 3%, etc.) of the multiplexing scores in the plurality of data as the target data. The aforementioned preset ratio can be set according to the needs of the user, and the present application does not limit this.

[0196] S720, the processor 100 determines a first basic block with the highest priority from a plurality of basic blocks included in the control flow graph.

[0197] Each of the plurality of basic blocks includes a first instruction for accessing the target data.

[0198] In a possible implementation, the processor 100 determines a first basic block with the highest priority from a plurality of basic blocks included in the control flow graph, including: the processor 100 determines a basic block including a first instruction for accessing the target data from the control flow graph, and further determines the priorities between the basic blocks, and determines a first basic block with the highest priority according to the priorities between the basic blocks.

[0199] It is worth noting that the aforementioned first instruction for accessing the target data can be a first instruction for accessing part of the target data, or a first instruction for accessing all of the target data. For example, the target data is a 3x4 matrix data, and the first instruction can be a first instruction for accessing a 2x2 matrix data in the 3x4 matrix, or a first instruction for accessing all data in the 3x4 matrix.

[0200] In a possible case, the priorities between the basic blocks can be determined according to one or more of the following indicators: a loop depth indicated by the first instruction, an access rule of the first instruction, a number of occurrences of an event of a basic block in which the first instruction is located, a line number of the first instruction in the source code, whether the first instruction is used to indicate reading the target data or writing the target data, whether the first instruction is a vectorized access or a scalarized access.

[0201] The loop depth refers to the number of nested loops in a program or a piece of code. For example, if one loop is nested in another loop in a program, the loop depth of the program is 2. The greater the loop depth, the more levels of loop nesting exist in the program, which usually leads to an increase in the execution time of the program. The loop depth indicated by the first instruction indicates the loop depth in which the first instruction is located. For example, if the first instruction is located in the outer loop of the two-level loop in the above example, the loop depth indicated by the first instruction is 1. If the first instruction is located in the inner loop of the two-level loop in the above example, the loop depth indicated by the first instruction is 2.

[0202] Two possible examples are provided below for determining the priority of each basic block.

[0203] Example 1: The processor 100 can determine the priority of the plurality of basic blocks according to only the loop depth indicated by the first instruction. For example, the loop depth indicated by the first instruction in the basic block a is 1, and the loop depth indicated by the first instruction in the basic block b is 2. Since the greater the set loop depth, the higher the priority of the basic block, the processor 100 determines that the basic block b is the first basic block with the highest priority.

[0204] Example 2: The processor 100 can determine the priority of the plurality of basic blocks according to the loop depth indicated by the first instruction and the access pattern of the first instruction. For example, the loop depth indicated by the first instruction in the basic block a is 2, and the access pattern is regular access, the loop depth indicated by the first instruction in the basic block b is 2, and the access pattern is random access. The processor 100 preferentially compares the loop depths, and the basic block including the first instruction with the greater loop depth is determined as the first basic block with the highest priority. Since the loop depths of the first instructions included in the basic block a and the basic block b are consistent, the access pattern of the first instruction in the basic block a and the access pattern of the first instruction in the basic block b are further determined. The access pattern of the first instruction is regular access, and the priority of the corresponding basic block is higher. Therefore, the processor 100 determines that the basic block a is the first basic block with the highest priority.

[0205] It is worth noting that more examples can be included in other embodiments of the present application to determine the priority of the plurality of basic blocks by using the loop depth indicated by the first instruction, the access pattern of the first instruction, the number of occurrences of the event of the basic block in which the first instruction is located, the line number of the first instruction in the source code, the first instruction for indicating the read target data or the write target data, the first instruction for indicating the single or multiple indicators in the vectorized access or the scalarized access. In addition, only two basic blocks are taken as examples in the above examples, and more basic blocks, such as 5 or 10, can be included in the control flow graph in other embodiments of the present application, which is not limited in the present application.

[0206] In another possible case, the priority between the basic blocks can be determined according to one or more of the following indexes in sequence: the loop depth indicated by the first instruction (which can also be referred to as the loop depth at which the first instruction is located), the access regularity of the first instruction, the number of occurrences of the event of the basic block at which the first instruction is located, the line number of the first instruction in the source code, whether the first instruction is used to indicate read target data or write target data, whether the first instruction is vectorized access or scalarized access.

[0207] For examples of this case, refer to the content shown in the following Figure 8 , which will not be described here in detail.

[0208] S730, the processor 100 writes the target data into the cache before executing the first instruction in the first basic block that accesses the target data.

[0209] In a possible implementation, the processor 100 writes the target data into the cache before executing the first instruction in the first basic block that accesses the target data, including: the processor 100 inserts a hint instruction before the first instruction in the first basic block. The processor 100 compiles the basic block after the hint instruction to obtain corresponding executable code. When the processor 100 executes the executable code corresponding to the hint instruction, the target data is written into the cache in advance, and the target data is stored in the cache persistently.

[0210] In a possible case, the processor 100 can insert the hint instruction before the line before the first instruction.

[0211] In another possible case, the processor 100 can insert the hint instruction after the first line of instruction in the multiple lines of instructions included in the first basic block and before the first instruction.

[0212] In another possible embodiment, the processor 100 can insert the hint instruction after the first instruction in the multiple lines of instructions included in the first basic block and before the last line of instruction in the multiple lines of instructions included in the first basic block.

[0213] As Figure 8 shown, Figure 8 the flowchart of the priority determination method provided in the present application. Figure 8 The content shown in the following Figure 7 illustrates the description of determining the priority between the basic blocks in sequence according to the indexes in S720 in the above Figure 8 , taking the case that the first instruction that accesses the target data is located in two different basic blocks (basic block a and basic block b) as an example. The first instruction in the basic block a can be referred to as instruction a, and the first instruction in the basic block b can be referred to as instruction b. Figure 8The content shown includes the following steps ① to ⑥.

[0214] Step ①: The processor 100 compares the loop depth of instruction a with the loop depth of instruction b.

[0215] When the loop depths of instruction a and instruction b are different, the basic block corresponding to the instruction at the higher loop depth has a higher priority. For the same target data, without lightweight instrumentation on the actual upper and lower bounds of the loop, the instruction at the higher loop depth has a higher probability of obtaining a higher footprint.

[0216] For example, if the loop depth of instruction a is greater than the loop depth of instruction b, the priority of basic block a is higher, and the processor 100 takes basic block a as the first basic block.

[0217] If the loop depth of instruction a is less than the loop depth of instruction b, the priority of basic block b is higher, and the processor 100 takes basic block b as the first basic block.

[0218] If the loop depth of instruction a is equal to the loop depth of instruction b, the processor 100 performs the content of step ②.

[0219] Step ②: The processor 100 compares the access regularity of instruction a and instruction b.

[0220] The access regularity includes regular access (also referred to as regular memory access or direct access / memory access) and random access (also referred to as random memory access), the probability of continuous access / memory access is higher for regular access, and pre-fetching can be performed using a vectorized instruction, that is, a vectorized instruction is used to instruct to write multiple data into a cache, and the data access efficiency is relatively high.

[0221] If the access regularities of instruction a and instruction b are inconsistent, the processor 100 takes the basic block corresponding to the instruction with regular access among instruction a and instruction b as the first basic block.

[0222] For example, when the access regularity of instruction a is regular access and the access regularity of instruction b is random access, the processor 100 takes basic block a as the first basic block.

[0223] When the access regularity of instruction b is regular access and the access regularity of instruction a is random access, the processor 100 takes basic block b as the first basic block.

[0224] If the access regularities of instruction a and instruction b are consistent, the processor performs the content of step ③.

[0225] Step ③: The processor 100 compares the occurrence frequency a of the event corresponding to basic block a with the occurrence frequency b of the event corresponding to basic block b.

[0226] If the difference between the occurrence number a and the occurrence number b is greater than the threshold value A, the basic block corresponding to the larger one of the occurrence number a and the occurrence number b is taken as the first basic block.

[0227] For example, if the difference between the occurrence number a and the occurrence number b is greater than the threshold value A, and the occurrence number a is larger, the basic block a is taken as the first basic block.

[0228] If the difference between the occurrence number a and the occurrence number b is greater than the threshold value A, and the occurrence number b is larger, the basic block b is taken as the first basic block.

[0229] If the difference between the occurrence number a and the occurrence number b is less than the threshold value A, the content of step 4 is executed.

[0230] If the difference between the occurrence number a and the occurrence number b is equal to the threshold value A, the processor 100 can take the basic block corresponding to the larger one of the occurrence number a and the occurrence number b as the first basic block according to the user's needs, or execute the content of step 4.

[0231] In one possible case, when the occurrence number a and the occurrence number b satisfy the following formula 1, the basic block corresponding to the larger one of the occurrence number a and the occurrence number b is the first basic block.

[0232]

[0233] Wherein, bb_count(a) represents the occurrence number a, bb_count(b) represents the occurrence number b, and gap represents the threshold value A, which is set by the user.

[0234] Step 4: The processor 100 compares the line numbers of the instructions a and b in the source code.

[0235] In one possible case, the processor 100 takes the basic block corresponding to the instruction with the smaller line number as the first basic block.

[0236] For example, if the line number of the instruction a is smaller than the line number of the instruction b, the processor 100 takes the basic block a as the first basic block.

[0237] If the line number of the instruction a is greater than the line number of the instruction b, the processor 100 takes the basic block b as the first basic block.

[0238] If the line number of the instruction a is the same as the line number of the instruction b, the processor 100 executes the content of step 5.

[0239] In another possible case, if the difference between the line number of instruction a and the line number of instruction b is greater than threshold B, the processor 100 takes the basic block corresponding to the smaller one of occurrence number a and occurrence number b as the first basic block.

[0240] If the difference between the line number of instruction a and the line number of instruction b is less than threshold B, the processor 100 performs the content of step 5.

[0241] If the difference between the line number of instruction a and the line number of instruction b is equal to threshold B, the processor 100 takes the basic block corresponding to the smaller one of occurrence number a and occurrence number b as the first basic block, or performs the content of step 5.

[0242] Step 5: The processor 100 compares the read-write types of instruction a and instruction b.

[0243] Since the overhead of write operation is large, the target data accessed by instruction a and instruction b of write operation (write type) is written into the cache.

[0244] If the read-write types of instruction a and instruction b are different, the processor 100 takes the basic block corresponding to the instruction of write type among instruction a and instruction b as the first basic block.

[0245] For example, the read-write types of instruction a and instruction b are different, and the read-write type of instruction a is write type, the processor 100 takes basic block a as the first basic block.

[0246] If the read-write types of instruction a and instruction b are the same, the processor 100 performs the content of step 7.

[0247] Step 7: The processor 100 compares whether instruction a and instruction b are vectorized memory access or scalarized memory access.

[0248] If the memory access types of instruction a and instruction b are different, the processor 100 takes the basic block corresponding to the instruction of vectorized memory access among instruction a and instruction b as the first basic block.

[0249] If the memory access types of instruction a and instruction b are the same, the processor 100 takes basic block a or basic block b as the first basic block.

[0250] In a possible embodiment, if the size of the target data occupied space cannot be determined, the target data cannot be directly stored in the cache, so as to avoid allocating too much space in the cache to store the target data, which causes other data to be frequently replaced in the cache, affects the service life of the cache, and also causes the overall data access efficiency of the processor core to be low.

[0251] Based on this, in the embodiment, the processor 100 determines the size of the target data occupied space according to the size of the basic block corresponding to the instruction of the target data. Figure 5S540, the data prefetching method further includes:

[0252] The processor 100 copies the control flow graph obtained in S510, and then deploys a calculation basic block including a third instruction for calculating the target data occupancy space before the initial basic block of the two identical control flow graphs (control flow graph a and control flow graph b). The next basic block of the calculation basic block is the initial basic block of the control flow graph a and the basic block in the control flow graph b. The first basic block in any one of the control flow graph a and the control flow graph b is deployed with a second instruction for indicating writing the target data to the swap.

[0253] In a possible case, the processor 100 compares the occupancy space with a threshold A according to the occupancy space of the target data determined by the calculation basic block. Then, the processor 100 determines to execute according to the control flow graph a or the control flow graph b according to the size of the occupancy space and the threshold A.

[0254] For example, the first basic block in the control flow graph b is deployed with the second instruction. If the processor 100 determines that the occupancy space of the target data is greater than the threshold A, the processor 100 executes according to the control flow graph a. If the processor 100 determines that the occupancy space of the target data is less than the threshold A, the processor 100 executes according to the control flow graph b.

[0255] If the processor 100 determines that the occupancy space of the target data is equal to the threshold A, the processor 100 executes according to the control flow graph a or the control flow graph b according to the user demand.

[0256] For the content of the data prefetching method, the following provides a possible embodiment. As shown in Figure 9 , the content shown in Figure 9 is a complete flowchart of a data prefetching method provided in the present application. Figure 9 The content shown in may include the following steps ① to ⑥.

[0257] Step 1: The processor 100 identifies and analyzes the nested loops in the source code.

[0258] For the content of the processor 100 identifying the nested loops in the source code, the following provides two possible examples.

[0259] Example 1, the nested loops are usually represented by indentation, and each nested level will increase one level of indentation. The nested loops usually have different indentation levels, and then the processor 100 identifies the nested loops by determining the indentation of the source code.

[0260] Example 2, the processor 100 can determine the nested loops by searching for for, while or similar loop structures in the source code.

[0261] In a possible implementation, the processor 100 analyzes the nested loops in the source code, including: after the processor 100 identifies the memory access node of the innermost loop in the nested loops, the processor 100 determines the data dimension (for example, an array or a matrix, etc.) contained by the traced memory access and the boundary iteration number of the loop according to the information of the current memory access trace.

[0262] In a possible case, the processor 100 determines the data dimension contained by the memory access, including: the processor 100 traces the base address and the index of the direct memory access respectively, and obtains the independent variable (iv) of the current loop from the index.

[0263] For example, in the process of the loop, each loop accesses an element in the data, and the position of the current loop iteration can be determined by the index, and then the value of the IV corresponding to the position is obtained.

[0264] For example, the loop traverses the elements in the data one by one according to certain conditions (such as the iteration number of the for loop or the condition judgment of the while loop). In each loop process, the position of the current iteration can be determined by the index. The index usually starts from 0 and increases by one for each loop iteration, until the length of the data is reached or the loop end condition is met. The processor 100 can access the value of the independent variable at the corresponding position in the data through the index.

[0265] Further, the processor 100 analyzes the boundary iteration number of each loop dimension according to the data dimension.

[0266] For example, for one-dimensional data (such as an array), the iteration number of the loop is usually equal to the length of the data (the number of elements).

[0267] For two-dimensional data (such as a two-dimensional array or a matrix), the iteration number of the loop is usually equal to the product of the number of rows and the number of columns of the data structure.

[0268] For multi-dimensional data, the iteration number of the loop depends on the size of each dimension of the data, that is, the iteration number of the loop is equal to the product of the size of each dimension. For example, for a three-dimensional array, the iteration number of the loop is equal to the product of the size of the first dimension, the size of the second dimension, and the size of the third dimension.

[0269] The boundary iteration number analysis includes the analysis of the boundary iteration number of the non-vectorized loop and the analysis of the boundary iteration number of the vectorized loop.

[0270] For the analysis of the boundary iteration number of the non-vectorized loop, refer to the content of the above examples, which will not be repeated here.

[0271] For the analysis of the boundary iteration number of the vectorized loop, since in the scenario of the vectorized loop, the processor 100 can process multiple elements, multiple rows or multiple columns, multiple dimensional elements in the data at a time, therefore, the boundary iteration number of the vectorized loop will be much smaller than the boundary iteration number of the non-vectorized loop, and the specific value is determined by the parallel computing capability of the processor.

[0272] Further, the processor 100 determines which loops are the outermost loops of the multi-dimensional data according to the boundary iteration number, these loops are usually the loops that directly process the multi-dimensional data, and checks the memory access nodes used in the boundary expression, to ensure that they are all dominated nodes of the outermost loops of the multi-dimensional data. For each memory access node, it is ensured that they are accessible when the boundary is calculated. The boundary of the memory access node can represent the dimension size of the multi-dimensional data, which can be a specific value or an induction variable.

[0273] If the boundary iteration number calculation involves dynamic multi-branch (for example, different boundary calculation methods are selected according to the condition), it is ensured that the memory access nodes corresponding to the selected boundary calculation method are correctly accessible.

[0274] Wherein, the dominated node refers to the node that controls the number of loop executions in the loop nesting.

[0275] In a possible case, for the scenario where the boundary iteration number of the innermost loop is unknown, the processor 100 can identify the scenario and determine the boundary by expanding the outer loop, so as to obtain more memory access nodes and boundary iteration numbers.

[0276] In a possible embodiment, the processor 100 can determine the size of the data in the loop according to the boundary iteration number determined above. Further, the processor 100 can determine whether to write the data into the cache (such as HBM) in advance, and persist the cache according to the size, and determine how much storage space in the HBM to allocate to the data.

[0277] Step 2: The processor 100 obtains the control flow graph including the nested loop, and the occurrence number of events of the basic blocks in the control flow graph recorded by the PMU.

[0278] For the detailed content of this step, please refer to the description of S510 and S520 in the above Figure 5 , which will not be repeated here.

[0279] For example, the control flow graph acquired by the processor 100 is BB-A, BB-B, BB-C, BB-D, BB-E, BB-F, and BB-G. Among them, BB-C and BB-D are in loop a, and BB-B, BB-C, BB-D, BB-E, and BB-F are in loop b. The number of occurrences of the event of BB-A is 10, the number of occurrences of the event of BB-B is 5, the number of occurrences of the event of BB-C is 5, the number of occurrences of the event of BB-D is 6, the number of occurrences of the event of BB-E is 5, the number of occurrences of the event of BB-F is 2, and the number of occurrences of the event of BB-G is 4.

[0280] Step 3: The processor 100 determines the target path of the plurality of paths in the control flow graph according to the number of occurrences of the event of each basic block in the control flow graph.

[0281] Since the more the number of occurrences of the event obtained by a basic block during sampling, the longer the residence time of the program in the basic block (code segment), the greater the bottleneck and the potential improvement. The processor 100 can concatenate the basic blocks in the function for which a large number of occurrences of the event are collected and the basic blocks on the path that must be passed through, obtain the path (target path) with the largest sum of the number of occurrences of the event of the upper basic blocks through DFS plus dynamic programming, and filter and combine the Kernels passed through the target path, thereby converting the combination of the function path and the Kernel into a longest path problem.

[0282] For example, the processor 100 uses DFS plus dynamic programming to solve the longest path problem, and the time complexity is O(|V|+|E|). At the same time, since the number of exit edges of a single basic block is basically not more than 2 (in a few cases, due to switch statements and the like, the number of exit edges is between 2 and n-1), the time complexity of this method is basically O(|V|) according to the standard of amortized analysis. When the longest path problem is solved, all the basic blocks traced by the path can be traced back.

[0283] The following shows the search process of DFS:

[0284] Step 1: The processor 100 starts searching from the initial basic block in the control flow graph.

[0285] Step 2: When the basic block is not the header BB of a loop appearing in the recorded list, is not marked as visited, and does not appear in the recorded list, the BB is put into the recorded list and its child node BB is traversed through DFS. After all the child node BBs are traversed, the BB is marked as visited and added to the recorded list.

[0286] Step 3, when the BB is a header BB of a cycle in the recorded, the processor 100 traverses all exit nodes BB of the cycle by DFS.

[0287] Step 4, when the basic block BB is other cases, stop searching.

[0288] Step 5, repeat steps 2-4 until the exit basic block in the control flow graph is traversed.

[0289] Since the traversal record of the basic block is in the form of post order, the reverse order of the recorded of the basic block is the topological order.

[0290] After the processor 100 obtains the topological order of the basic block in the program running time (i.e., multiple paths), the processor 100 determines the sum of the occurrence times of the events of all basic blocks included in each path in the multiple paths.

[0291] For the content of determining the sum of the occurrence times of the events of all basic blocks of each path in the multiple paths, refer to the description of S530 in the above Figure 5 and Figure 6 , which will not be repeated here.

[0292] As shown in Figure 9 , the determined target path is BB-A→BB-B→BB-C→BB-D→BB-E→BB-F→BB-G.

[0293] In one possible embodiment, in the scenario of a directed acyclic graph (DAG), the longest path problem has a fixed linear solution, and the time complexity is O(|V|+|E|). Since there is a cycle in the CFG, the processor 100 converts the CFG into a DAG. The converted DAG has a one-to-one correspondence between the following mapping relationship and the basic blocks of the original CFG:

[0294] G=CFG graph, V={v_1,v_2,…,v_n}, E={e_1,e_2,…,e_m}, weight(e_j)=bb_count(v_i).

[0295] Where v_i is a basic block in the CFG, the index is i, e_j is an exit edge in the CFG, and bb_count represents the occurrence time of the event.

[0296] As shown in Figure 10 , Figure 10 is a conversion diagram of a directed acyclic graph provided by the present application. Figure 10a is the original control flow graph, and b is the control flow graph in the form of a directed acyclic graph converted from the control flow graph. The following shows the process of converting the original control flow graph into the control flow graph in the form of a directed acyclic graph by the processor 100, including the following steps A to C.

[0297] Step A, in the DFS process, the searched basic block does not return to the header BB in the loop (such as BB-3 in FIG. 1A), and the search is not performed from the exit edge outward (i.e., the path from BB-3 to BB-5 and from BB-4 to BB-5 is not searched). Figure 10

[0298] Step B, when the searched basic block returns to the header BB in the loop, the search of the exit edge of the loop is started (such as the path from BB-3 to BB-5 and from BB-4 to BB-5 shown in FIG. 1A). Figure 10

[0299] Step C, the processor 100 connects the start of the exit edge with the back edge of the loop (i.e., the path from BB-7 to BB-3 shown in FIG. 1A), thereby converting the original control flow graph into the control flow graph in the form of a directed acyclic graph. Figure 10

[0300] Further, the processor 100 can determine the longest path (target path) in the control flow graph in the form of a directed acyclic graph by using topological sorting.

[0301] Step IV: the processor 100 determines the target data according to the data required to be accessed by the target path.

[0302] The target data is X in FIG. 1A, for example. Figure 9

[0303] The processor 100 calculates the reuse score of the data (memory access group) required to be accessed by the target path. The reuse score of the data can be determined by the data reuse mode, memory access regularity, read-write type, space occupation limit, etc. For example, the processor 100 can calculate the reuse score of a single memory access (element in the data) in the memory access group according to the following formula 2:

[0304] Level(mem ref) = parallel(x) * reuse(x) * regular(x) Formula 2

[0305] Wherein, parallel(x) is used to represent the parallel reuse score, reuse(x) represents the serial reuse score, and regular(x) represents the memory access regularity score.

[0306] ​​​​For parallel(x), the processor 100 can detect the loop of OpenMP parallel and identify the shared data, and the specific score is as follows formula 3:

[0307]

[0308] For reuse(x), the processor 100 determines whether the def of the trace memory is the same object, and calculates the serial reuse score. Specifically, first, according to the mem ref of the data, it is divided into 8 groups of mem ref types in three dimensions (read / write, regular memory access, and parallel), and then according to whether the same memory access group appears in the basic block or loop of the mem ref, the basic reuse times base_reuse(x) is calculated, as follows formula 4:

[0309]

[0310] The processor 100 converts the score of the memory access according to the granularity of the loop structure to which the memory access belongs. For example, when the same type of memory access is located in the same BB, since the BB granularity is small, the reuse score of the subsequent same type of memory access located in the "same BB" is not high, and the overall score is not more than 0.1.

[0311] According to the dimension identified by the data itself and the loop depth where the data is located, the reuse weight of the data is calculated, and the serial reuse score of the data is calculated, as follows formula 5 to formula 7:

[0312] used_dim(x)=min(loop_depth,var_dim) formula 5

[0313] power(x)=log2(max(0,loop_depth-used_dim(x))+2) formula 6

[0314]

[0315] Wherein, loop_depth represents the loop depth, var_dim represents the dimension of the data, used_dim(x) represents the footprint dimension used by the data in the memory access, and power(x) represents the reuse of the footprint of the data in the loop.

[0316] Without the extra information of lightweight instrumentation, on one hand, the processor 100 estimates the footprint dimension, used_dim, of the data used by the data access according to the dimension of the data and the loop depth where the data access is located. The larger the footprint dimension used by the data access, the larger the footprint of the data used by the data access, and the higher the score of the data access. On the other hand, loop_depth-used_dim reflects the reuse of the footprint of the data within the loop, and the more reuse, the higher the score of the data access. The relative weight of the two factors in the score of the data access is adjusted according to the growth rate, O(n^2) for the former and log(n) for the latter. Finally, the serial reuse score of a single data access is obtained by multiplying the score of the two factors and the basic score.

[0317] For regular(x), in the case of equivalent HBM and DDR delay, there is the following rule, rule access: 1 cacheline in the cache can be used for s data access instructions, random access: 1 cacheline is generally used for only 1 data access instruction, and more bandwidth is occupied. Based on this, the data access rule score of a single data access can be determined according to the following formula 8:

[0318]

[0319] Further, the processor 100 calculates the reuse score of the data access group according to the accumulated reuse score of a single data access according to the following formula 9:

[0320] Level(refgroup) = write_cost(x) * footprint_cost(x) *∑Level(mem ref) Formula 9

[0321] Wherein, write_cost(x) represents the score of the read-write mode of the data, and footprint_cost(x) represents the score of the footprint restriction.

[0322] For write_cost(x), the processor 100 can determine it in the following formula 10:

[0323]

[0324] For footprint_cost(x), the processor 100 preferentially selects the data size in The data prefetching effect of data in a specific size range is better than that of data outside the range. (When the data is smaller than the L2 cache, the data may complete the memory access within the core, the prefetch instruction cannot be out of the core, and the impact on the program footprint is small; when the data is too large, the allocation of the prefetch instruction affects the bandwidth of the DRR and the HBM, and a large part of the capacity in the HBM is occupied by the data, resulting in an increase in the number of passive replacement times of the remaining data that needs to be saved for a long time.) The processor 100 can determine the footprint_cost(x) in the following formula 11:

[0325]

[0326] Step 5: The processor 100 determines the first basic block with the highest priority from the plurality of basic blocks included in the control flow graph.

[0327] Each of the plurality of basic blocks includes a first instruction for accessing target data.

[0328] For details of step 5, refer to the above Figure 8 description, which is not repeated here.

[0329] As Figure 9 shown, the determined first basic block is BB-C.

[0330] Step 6: The processor 100 deploys a second instruction for instructing to write the target data to the cache in the first basic block.

[0331] For example, the processor 100 deploys the second instruction in the first basic block adjacent to the first instruction. The second instruction is a hint instruction.

[0332] When the processor 100 executes the second instruction, the target data can be written in advance to the cache (such as HBM) and stored persistently in the HBM, thereby improving the efficiency of the core in the processor 100 accessing the target data.

[0333] In a possible embodiment, in the static analysis process, if the memory access object is an incoming parameter or input data, that is, the data accessed by the memory access is an incoming parameter or input data, and the boundary iteration number is unknown in the source code and the compilation process, the boundary iteration number size will be determined by the runtime data. In order to accurately know the memory size, reasonably analyze and allocate data, and avoid allocating too much data in the HBM cache to cause frequent mutual replacement, a dynamic multi-branch design is added.

[0334] On the basis of static analysis, the runtime judgment of the loop footprint is added, and the outermost loop to which the attribution variable in the memory access node can point is analyzed and used as the judgment boundary of the footprint. The basic block for calculating the size of the footprint of the data is added outside the control flow graph, and the corresponding hint is inserted into the branch to realize the runtime decision of the static optimization.

[0335] As shown in Figure 11 , Figure 11 The control flow graph provided in the present application is dynamically multi-branch inserted. Figure 11 The BB-82 in the present application is a calculation basic block deployed outside the control flow graph for calculating the footprint of the data. The calculation basic block includes instructions for calculating the footprint of the data (such as the target data mentioned above).

[0336] In a possible example, the calculation basic block includes instructions that can obtain the boundary iteration number of the runtime of the memory access node by reading the induction variable of the loop, and then calculate the size of the footprint according to the boundary iteration number.

[0337] In a possible case, the processor 100 generates the above-mentioned instructions, and the instructions are intermediate representation instructions.

[0338] Figure 11 In the present application, the calculation basic block is deployed before the control flow graph a and the control flow graph b. Specifically, the next basic block of the calculation basic block is the initial basic block (such as BB-83) of the control flow graph a and the initial basic block (such as BB-84) of the control flow graph b.

[0339] In the present application, the control flow graph a includes BB-83, BB-13, BB-16, BB-19, BB-37, BB-44, BB-36, BB-43, BB-35, BB-73, and BB-28. The control flow graph b includes BB-84, BB-74, BB-75, BB-76, BB-77, BB-78, BB-79, BB-80, BB-81, BB-73, and BB-28.

[0340] Among them, BB-73 and BB-28 are basic blocks shared by control flow graphs a and b. BB-19 and BB-37 are in loop 13, BB-16, BB-19, BB-37, BB-44, and BB-36 are in loop 12, BB-13, BB-16, BB-19, BB-37, BB-44, BB-36, BB-43, and BB-35 are in loop 11, BB-76 and BB-77 are in loop 19, BB-75, BB-76, BB-77, BB-78, and BB-79 are in loop 18, and BB-74, BB-75, BB-76, BB-77, BB-78, BB-79, BB-80, and BB-81 are in loop 20.

[0341] In a possible example, the control flow graph b is obtained by the processor 100 copying the control flow graph a, and the first basic block in the control flow graph b also deploys the second instruction. For the content of deploying the second instruction in the first basic block in the control flow graph b, please refer to the above Figure 9 The description of step ⑥ in the above will not be repeated here.

[0342] The processor 100 determines the target data's occupied space A. If occupied space A is less than threshold A, the processor 100 executes according to control flow graph b, pre-writing the target data to the cache and persisting it in storage. If occupied space A is greater than threshold A, the processor 100 executes according to control flow graph a without pre-writing the target data to the cache. If occupied space A is equal to threshold A, the processor 100 executes according to either control flow graph a or control flow graph b.

[0343] Combined with the above Figures 3 to 11 , describes in detail the data prefetching method provided by this application, and will be combined with Figure 12 , Figure 12 A schematic diagram of the structure of a data pre-fetching device provided in this application Figure 1 , describes the data pre-fetching device provided by the present application. The data pre-fetching device 1200 can be used to implement the functions of the processor 100 in the above method embodiment, and thus can also achieve the beneficial effects possessed by the above method embodiment.

[0344] like Figure 12 As shown, the data pre-fetching device 1200 includes an acquisition module 1210, a determination module 1220 and a writing module 1230. The data pre-fetching device 1200 is used to implement the above Figures 3 to 11 The function of the processor 100 in the corresponding method embodiment. In a possible example, the specific process of the data pre-fetching device 1200 for implementing the above-mentioned data pre-fetching method includes the following process:

[0345] The acquisition module 1210 is configured to acquire a control flow graph and events of basic blocks in the control flow graph recorded by a PMU.

[0346] The determination module 1220 is configured to determine a target path of a plurality of paths in the control flow graph according to the number of occurrences of the events of the basic blocks in the control flow graph, the sum of the number of occurrences of all events of the basic blocks in the target path being greater than or equal to a first threshold value.

[0347] The writing module 1230 is configured to write target data corresponding to a first basic block in the control flow graph into a cache before the first basic block is executed, the target data being data required to be accessed by the first basic block for execution, the target data being data required to be accessed when the target path is executed and having a frequency of access or a frequency of being accessed greater than or equal to a second threshold value.

[0348] To further implement the functions in the method embodiments shown in the above Figures 3 to 11 embodiments, the present application also provides a data prefetching apparatus, as shown in Figure 13 , Figure 13 The data prefetching apparatus 1200 provided by the present application has the structure shown in Figure 2 The data prefetching apparatus 1200 further includes a dynamic prediction module 1240.

[0349] The dynamic prediction module 1240 is configured to deploy a second instruction for indicating that the target data is written into the cache in the first basic block. Before an initial basic block included in the control flow graph, a calculation basic block including a third instruction for calculating an occupied space of the target data is deployed; the next basic block of the calculation basic block is the initial basic block of the control flow graph and the initial basic block of the control flow graph including the second instruction. The occupied space of the target data is used to indicate that the initial basic block of the control flow graph is executed, or the initial basic block of the control flow graph including the second instruction is executed.

[0350] As an example of the module as a hardware functional unit, the acquisition module 1210 can include at least one computing device, such as a server, etc. Alternatively, the acquisition module 1210 can also be a device implemented by ASIC or PLD, etc. The above-mentioned PLD can be CPLD, FPGA, GAL or any combination thereof.

[0351] It should be noted that in other embodiments, the obtaining module 1210 can be configured to perform any of the steps of the data prefetching method, the determining module 1220 can be configured to perform any of the steps of the data prefetching method, and the writing module 1230 can be configured to perform any of the steps of the data prefetching method. The steps implemented by the obtaining module 1210, the determining module 1220, and the writing module 1230 can be specified as needed, and the entire function of the data prefetching apparatus can be implemented by the obtaining module 1210, the determining module 1220, and the writing module 1230 implementing different steps of the data prefetching method.

[0352] It should be noted that the processor 100 of the foregoing embodiments can correspond to the data prefetching apparatus 1200, and can correspond to a computer system (including the processor 100) that executes the method according to the embodiments of the present application Figures 3 to 11 corresponding subject, and the operations and / or functions of each module in the data prefetching apparatus 1200 are respectively implemented in order to implement Figures 3 to 11 the corresponding processes of each method corresponding to the embodiments in the foregoing method, and for brevity, will not be described again here.

[0353] In addition, Figure 12 , Figure 13 The data prefetching apparatus shown in the foregoing embodiments can also be implemented by a communication device, where the communication device can be a computer system (including the processor 100) in the foregoing embodiments, or when the communication device is a chip or chip system applied to a computing device, the data prefetching apparatus can also be implemented by the chip or chip system.

[0354] The embodiments of the present application also provide a chip system, which includes a control circuit and an interface circuit, the interface circuit is configured to obtain a control flow graph, and the control circuit is configured to implement the functions of the computing device in the foregoing method according to the control flow graph.

[0355] In a possible design, the chip system further includes a memory configured to store program instructions and / or data. The chip system can be composed of a chip, or can include a chip and other discrete devices.

[0356] The present application also provides a computing device. As Figure 14 shown, Figure 14 a structural schematic diagram of a computing device provided by the present application, the computing device 1400 includes a bus 1402, a processor 1404, a memory 1406, and a communication interface 1408. The processor 1404, the memory 1406, and the communication interface 1408 communicate through the bus 1402. The computing device 1400 can be a server, a terminal device, and the computing device 1400 can be the computer system shown in the foregoing Figure 3 It should be noted that the number of processors and memories in the computing device 1400 is not limited in the present application.

[0357] The bus 1402 can be, but is not limited to, a PCIe bus, a universal serial bus (USB), or an inter-integrated circuit (I2C), an EISA bus, a UB, a CXL, a CCIX, etc. The bus 1402 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 14 Only one line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The bus 1402 can include a path for transmitting information between various components of the computing device 1400 (for example, the memory 1406, the processor 1404, the communication interface 1408).

[0358] The processor 1404 can include any one or more of a CPU, a GPU, a microprocessor (MP), or a digital signal processor (DSP), etc.

[0359] The memory 1406 can include a volatile memory (for example, a random access memory (RAM)), and can also include a non-volatile memory (for example, a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD), or a solid state drive (SSD)).

[0360] The memory 1406 stores executable program code, and the processor 1404 executes the executable program code to respectively implement the functions of the foregoing acquisition module, determination module, and writing module, thereby implementing the foregoing data prefetching method. That is, the memory 1406 stores instructions for executing the data prefetching method.

[0361] The communication interface 1408 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1400 and other devices or communication networks.

[0362] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device 1400. The memory 1406 in one or more computing devices 1400 in the computing device cluster can store the same instructions for executing the data prefetching method.

[0363] The computing device 1400 can be a server, for example, a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device 1400 can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0364] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc.

[0365] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or a program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the data prefetching method.

[0366] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium can be any available medium that the computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium contains instructions that instruct the computing device to perform the data prefetching method.

[0367] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are performed. The computer can be a general purpose computer, a special purpose computer, a computer network, a network device, a user equipment or other programmable apparatus. The computer programs or instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer programs or instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through wired or wireless manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center and the like integrated with one or more available media. The available media can be a magnetic medium, for example, a floppy disk, a hard disk, a magnetic tape; or an optical medium, for example, a digital video disc (digital video disc, DVD); or a semiconductor medium, for example, a solid state disk (solid state drive, SSD).

[0368] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data prefetching method, characterized by, The method comprises: acquiring a control flow graph and events of basic blocks in the control flow graph recorded by a performance monitoring unit (PMU); determining a target path of a plurality of paths in the control flow graph according to the number of occurrences of the events of the basic blocks in the control flow graph, wherein the sum of the number of occurrences of all events of the basic blocks in the target path is greater than or equal to a first threshold value; writing target data corresponding to a first basic block in the control flow graph into a cache before executing the first basic block, wherein the target data is data required to be accessed by the execution of the first basic block, and the target data is data required to be accessed when the target path is executed, and the frequency or the frequency of accessing the data is greater than or equal to a second threshold value.

2. The method of claim 1, wherein, The sum of the number of occurrences of all events of the basic blocks included in each path of the plurality of paths is a candidate value, and one of the plurality of candidate values is the first threshold value.

3. The method of claim 2, wherein, The first threshold value is the maximum value of the plurality of candidate values.

4. The method according to any one of claims 1 to 3, characterized in that, The determining of the target path of the plurality of paths in the control flow graph according to the number of occurrences of the events of the basic blocks in the control flow graph comprises: searching for a plurality of paths in the control flow graph starting from an initial basic block, wherein the exit basic block of each path of the plurality of paths does not belong to a basic block in a nested loop in the control flow graph; determining the sum of the number of occurrences of the events of the basic blocks included in each path of the plurality of paths; determining the target path of the plurality of paths in which the sum of the number of occurrences of the events is greater than or equal to the first threshold value according to the sum of the number of occurrences of the events of the basic blocks included in each path of the plurality of paths.

5. The method according to any one of claims 1 to 4, characterized in that, The writing of the target data corresponding to the first basic block into the cache before the execution of the first basic block in the control flow graph comprises: parsing source code corresponding to the target path to obtain target data, wherein the reuse score of the target data is the Mth in the order of the reuse scores of a plurality of data indicated by the source code, M is an integer greater than or equal to 1, and the reuse score of the target data is used to indicate the reuse degree of the target data; determining the first basic block with the highest priority from a plurality of basic blocks included in the control flow graph, wherein each basic block of the plurality of basic blocks includes a first instruction for accessing the target data; writing the target data into the cache before executing the first instruction included in the first basic block.

6. The method of claim 5, wherein, The priority of the basic block is determined according to one or more indicators. The loop depth indicated by the first instruction, the access rule of the first instruction, the number of occurrences of the events of the basic block in which the first instruction is located, the line number of the first instruction in the source code, the first instruction for indicating reading the target data or writing the target data, and the first instruction for indicating vectorized memory access or scalar memory access.

7. The method of claim 6, wherein, The priority of the basic block is determined according to one or more indicators in turn. The first instruction indicates a loop depth, an access rule of the first instruction, a number of occurrences of an event of a basic block in which the first instruction is located, a line number of the first instruction in the source code, whether the first instruction is used to indicate reading the target data or writing the target data, whether the first instruction is vectorized access or scalarized access.

8. The method according to any one of claims 1 to 7, characterized in that, The basic blocks in the control flow graph include basic blocks in nested loops.

9. The method according to any one of claims 1 to 8, characterized in that, Before executing a first basic block in the control flow graph, the method further comprises: deploying a second instruction for indicating writing the target data into the cache in the first basic block; deploying a calculation basic block including a third instruction for calculating an occupation space of the target data before an initial basic block of the control flow graph; a next basic block of the calculation basic block is the initial basic block of the control flow graph and the initial basic block of the control flow graph including the second instruction; the occupation space of the target data is used to indicate executing the initial basic block of the control flow graph or executing the initial basic block of the control flow graph including the second instruction.

10. The method according to any one of claims 1 to 9, characterized in that, Each of the plurality of paths includes all basic blocks in nested loops in the control flow graph.

11. A data prefetching apparatus, comprising: The apparatus comprises: an obtaining module configured to obtain a control flow graph and events of basic blocks in the control flow graph recorded by a performance monitoring unit (PMU); a determining module configured to determine a target path of a plurality of paths in the control flow graph according to a number of occurrences of events of basic blocks in the control flow graph, wherein a sum of the numbers of occurrences of all events of basic blocks in the target path is greater than or equal to a first threshold value; a writing module configured to write target data corresponding to a first basic block in the control flow graph into a cache before executing the first basic block, wherein the target data is data required to be accessed by the first basic block, and the target data is data with a frequency of access or a frequency of access greater than or equal to a second threshold value in data required to be accessed when the target path is executed.

12. The apparatus of claim 11, wherein, A sum of the numbers of occurrences of all events of basic blocks included in each of the plurality of paths is a candidate value, and one of a plurality of candidate values of the plurality of paths is the first threshold value.

13. The apparatus of claim 12, wherein, The first threshold value is a maximum value in the plurality of candidate values.

14. The apparatus of any one of claims 11 to 13, wherein, The determining module is specifically configured to search for a plurality of paths starting from an initial basic block in the control flow graph, wherein an exit basic block of each of the plurality of paths does not belong to basic blocks in nested loops in the control flow graph, determine a sum of the numbers of occurrences of events of basic blocks included in each of the plurality of paths, and determine a target path in which the sum of the numbers of occurrences of the events is greater than or equal to the first threshold value from the plurality of paths according to the sum of the numbers of occurrences of the events of the basic blocks included in each of the plurality of paths.

15. The apparatus of any one of claims 11 to 14, wherein, The writing module is specifically configured to parse source code corresponding to the target path to obtain target data; a multiplexing score of the target data is M of a plurality of multiplexing scores of data indicated by the source code, M being an integer greater than or equal to 1, and the multiplexing score of the target data being used to indicate a multiplexing degree of the target data; the first basic block with the highest priority is determined from a plurality of basic blocks included in the control flow graph, each basic block in the plurality of basic blocks including a first instruction for accessing the target data; and the target data is written into a cache before the first instruction included in the first basic block is executed.

16. The apparatus of claim 15, wherein, The priority of the basic block is determined according to one or more of the following indexes: The first instruction indicates a loop depth, an access rule of the first instruction, a number of occurrences of an event of a basic block in which the first instruction is located, a line number of the first instruction in the source code, the first instruction being used to indicate reading or writing the target data, and the first instruction being vectorized access or scalar access.

17. The apparatus of claim 16, wherein, The priority of the basic block is determined according to one or more of the following indexes in sequence: The first instruction indicates a loop depth, an access rule of the first instruction, a number of occurrences of an event of a basic block in which the first instruction is located, a line number of the first instruction in the source code, the first instruction being used to indicate reading or writing the target data, and the first instruction being vectorized access or scalar access.

18. The apparatus of any one of claims 11-17, wherein, The basic blocks in the control flow graph include basic blocks in nested loops.

19. The apparatus of any of claims 11 to 18, wherein, The apparatus further includes a dynamic prediction module. The dynamic prediction module is configured to deploy, in the first basic block, a second instruction for indicating writing the target data into the cache; and deploy, before an initial basic block included in the control flow graph, a calculation basic block including a third instruction for calculating an occupation space of the target data; a next basic block of the calculation basic block is the initial basic block of the control flow graph and an initial basic block of the control flow graph including the second instruction; and the occupation space of the target data is used to indicate executing the initial basic block of the control flow graph or executing the initial basic block of the control flow graph including the second instruction.

20. The apparatus of any one of claims 11-19, wherein, Each of the plurality of paths includes all basic blocks in nested loops in the control flow graph.

21. A chip, characterized by Comprise: A processor and a power supply circuit; The power supply circuit is configured to supply power to the processor; The processor is configured to execute the method of any one of claims 1 to 10.

22. A computing device, comprising: Comprise a memory and a processor, the memory is configured to store computer instructions; the processor executes the computer instructions to implement the method of any one of claims 1 to 10.

23. A computer-readable storage medium, characterized in that, The storage medium stores a computer program or instructions, when the computer program or instructions are executed by a processing device, the method of any one of claims 1 to 10 is implemented.

24. A computer program product comprising computer programs or instructions, characterized in that, When the computer program or instructions are executed by a processing device, the method of any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Access control method of memory, memory and electronic equipment

    CN121597614A