Data prefetching method and computer device

CN122816698APending Publication Date: 2026-09-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511376627.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2025-09-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]在编译器的典型数据预取方法中,编译器可能不当地插入无益的预取指令,并将大量时间花费在仅占极小比例的关键代码优化上,同时增加代码体积

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816698A_ABST
    Figure CN122816698A_ABST
Patent Text Reader

Abstract

The application provides a data prefetching method and a computer device. The method is applied to a compiler and includes the following steps: obtaining a computer program, wherein a plurality of subprograms in the computer program form a loop structure and are used for processing a plurality of to-be-processed data; determining a loop calling number N of a first subprogram for processing the plurality of to-be-processed data in the loop structure according to a calling relationship between the plurality of subprograms and a loop body in the subprogram; determining a first gradual increasing rate of a running time length of the loop structure according to the N; and inserting a prefetching instruction into at least one subprogram in a case where the first gradual increasing rate is greater than or equal to a preset gradual increasing rate. The application can determine whether a prefetching instruction needs to be inserted into the loop structure, thereby improving the running efficiency of the computer program.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Russian Federal Patent Application No. RU2025106951, filed on March 24, 2025, entitled "method and device for data prefetching", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and more specifically, to a data prefetching method and a computer device. Background Technology

[0003] Data prefetching improves system performance by reducing memory latency. It can be implemented by hardware units such as the processor or by software modules such as the compiler. At the hardware level, the processor's built-in prefetching mechanism automatically predicts and loads upcoming data into the cache, thus reducing waiting time. At the software level, the compiler analyzes the program's memory access patterns and inserts prefetch instructions at appropriate locations to load data into the cache in advance, reducing memory latency and improving execution efficiency.

[0004] In typical compiler data prefetching methods, the compiler may inappropriately insert useless prefetch instructions and spend a significant amount of time on critical code optimizations that account for only a small percentage of the total time, while increasing code size. Furthermore, the compiler cannot capture optimization opportunities in loop bodies and memory access patterns that span multiple functions. Therefore, prefetch instructions can only be inserted within the current function scope; when data access patterns and desired prefetch locations are distributed across different functions, the compiler cannot effectively perform such optimizations. Summary of the Invention

[0005] This application provides a data prefetching method and a computer device to insert prefetch instructions at appropriate locations in computer program code, thereby improving the running efficiency of the computer program.

[0006] In a first aspect, a data prefetching method is provided, the method being applied to a compiler, the method comprising: acquiring a computer program, wherein multiple subroutines in the computer program form a loop structure and are used to process multiple pieces of data to be processed; determining, based on the calling relationships between the multiple subroutines and the loop bodies in the subroutines, the number of loop calls N of a first subroutine used to process the multiple pieces of data to be processed in the loop structure; determining, based on N, a first asymptotic growth rate of the runtime of the loop structure; and, if the first asymptotic growth rate is greater than or equal to a preset asymptotic growth rate, inserting a prefetch instruction in at least one subroutine.

[0007] In this embodiment of the application, the computer program code that forms a loop structure and is used to process multiple data to be processed is distributed in multiple subroutines. The number of loop calls of the first subroutine in the loop structure can be used to calculate the incremental growth rate of the loop structure's runtime, so as to determine whether it is necessary to insert prefetch instructions in the loop structure, thereby improving the running efficiency of the computer program.

[0008] In some implementations of the first aspect, the first subroutine uses random access to the memory addresses of multiple data to be processed.

[0009] In hardware-level data prefetching methods, the processor's built-in prefetching mechanism can automatically predict and load data to be accessed in sequential and cyclic access. In this implementation, the data prefetching method of this application can focus only on the data to be processed and the prefetch instructions related to random access, avoiding the inappropriate insertion of useless prefetch instructions, thereby improving the running efficiency of the computer program.

[0010] In some implementations of the first aspect, determining the number of loop calls N of the first subroutine in the loop structure includes: determining at least one second subroutine among a plurality of subroutines, wherein the at least one second subroutine is used to perform N loop calls on the first subroutine based on a first dataset, and one of the at least one second subroutines is used to perform M loop calls on the first subroutine in a loop body with a stroke count of M, the first dataset being used to determine the target address of the data to be processed; determining the first subroutine, wherein the first subroutine is used to determine the target address of the data to be processed based on the first dataset and the inductive variable of the loop body, and to reference the target address in memory; determining a third subroutine called by the first subroutine, wherein in the i-th loop call of the first subroutine by the at least one second subroutine, the third subroutine is used to instruct the at least one second subroutine to perform the (i+1)-th loop call of the first subroutine, i≤N, M≤N, and i, M, and N are positive integers; determining the relationship between N and the stroke count M of the loop body of the at least one second subroutine based on the calling relationship of the first subroutine, the at least one second subroutine, and the third subroutine, wherein the relationship between N and the at least one M is used to determine a first asymptotic growth rate.

[0011] In this implementation, the first, second, and third subroutines, which form a loop structure and are used to process multiple unprocessed data, correspond to specific functions and have complex calling relationships. This application can use the sorted calling relationships to determine the asymptotic growth rate of the number of loop calls of the first subroutine in the loop structure, so as to calculate the asymptotic growth rate of the runtime of the loop structure, thereby improving the efficiency of determining whether prefetch instructions need to be inserted into the loop structure.

[0012] In some implementations of the first aspect, determining the relationship between N and at least one M comprises: determining a directed acyclic graph based on a plurality of subprograms, wherein the plurality of subprograms are respectively used as a plurality of graph nodes of the directed acyclic graph, and call relationships between the plurality of subprograms are used as at least one directed edge of the directed acyclic graph, wherein a second subprogram is an ancestor node of a first subprogram, the first subprogram is an ancestor node of a third subprogram, and a trip count M of a loop body of the second subprogram is used for determining a first weight of a directed edge from the second subprogram to the first subprogram; and determining the relationship between N and the at least one M based on the first weight of the at least one directed edge.

[0013] In such implementations, the number of times each subprogram is circularly called by its parent node can be determined according to the first weight of the directed edge between subprograms, so that N can be represented by a plurality of trip counts M, thereby improving the efficiency of determining whether a prefetch instruction needs to be inserted into a loop structure.

[0014] In some implementations of the first aspect, determining the relationship between N and the at least one M based on the first weight of the at least one directed edge comprises: determining the relationship between N and the at least one M according to a rank of the first subprogram, wherein the rank of a subprogram is determined according to a second weight of at least one incoming edge of the subprogram, and the second weight of the incoming edge of the subprogram is determined according to the first weight of the incoming edge and the rank of a parent node of the subprogram in the directed acyclic graph.

[0015] In some implementations of the first aspect, the first weight of a directed edge, the second weight of a directed edge and the rank of a subprogram are all non-negative numbers, and the rank of the first subprogram is a sum of second weights of at least one incoming edge of the first subprogram, wherein: when the rank of the parent node of the first subprogram in the directed acyclic graph is equal to 0, the second weight of the incoming edge is the first weight of the incoming edge; when the first weight of the incoming edge is equal to 0, the second weight of the incoming edge is the rank of the parent node; and when both the first weight of the incoming edge and the rank of the parent node are greater than 0, the second weight of the incoming edge is a product of the first weight of the incoming edge and the rank of the parent node.

[0016] In such implementations, the rank of a subprogram represents the number of circular calls of the subprogram in the loop structure, and the definition of the rank of the first subprogram is the same as that of N, so the rank of the first subprogram can be used as N, thereby improving the efficiency of determining whether a prefetch instruction needs to be inserted into the loop structure.

[0017] In some implementations of the first aspect, inserting a prefetch instruction in at least one subprogram comprises: inserting a prefetch instruction in the second subprogram or the first subprogram, wherein in a j-th cyclic call of the first subprogram by the at least one second subprogram, a prefetch address of the prefetch instruction is a target address of data to be processed in a (j+t)-th cyclic call, where j<j+t≤N, and j and t are positive integers.

[0018] In a second aspect, a computer device is provided, the computer device including at least one processor and at least one memory; the at least one memory stores at least one computer program, the at least one computer program including instructions that, when executed by the at least one processor, cause the method as described in any implementation of the first aspect to be performed.

[0019] Thirdly, a chip is provided, the chip including a processor and a communication interface, the communication interface being used to receive a signal and transmit the signal to the processor, the processor being used to process the signal such that the method as described in any implementation of the first aspect is executed.

[0020] Fourthly, a computer-readable storage medium is provided, wherein computer instructions are stored therein, which, when executed on a computer device, cause the method described in any implementation of the first aspect to be performed.

[0021] Fifthly, a computer program product comprising instructions is provided, which, when executed by a computer device, cause the computer device to perform the method as described in any implementation of the first aspect. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a data prefetching method provided in an embodiment of this application.

[0023] Figure 2 This is a call diagram of a computer program provided in an embodiment of this application.

[0024] Figure 3 This is a directed acyclic graph representing the calling relationship of computer programs, provided in an embodiment of this application.

[0025] Figure 4 This is a directed acyclic graph representing the calling relationship of computer programs, provided in an embodiment of this application.

[0026] Figure 5 This is a computer program code fragment provided in an embodiment of this application.

[0027] Figure 6 This is a schematic block diagram of a chip system provided in an embodiment of this application.

[0028] Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of this application.

[0029] Figure 8 This is a schematic block diagram of a computer program product provided in an embodiment of this application. Detailed Implementation

[0030] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0031] A directed acyclic graph (DAG) contains a plurality of nodes and directed edges, and there are no closed paths, meaning there is no path that starts from a node, goes through several edges, and eventually returns to that node. Furthermore, because a DAG has no cycles, it can be topologically sorted, meaning a linear order can be found such that for any pair of nodes u and v, if there exists a path from u to v, then node u will appear before node v in the sorted order. Node v is called a descendant node of node u, and node u is called an ancestor node of node v.

[0032] Furthermore, a node without an ancestor node is called a source node, and a node without descendants is called a sink node. If there is a directed edge between node u and node v, with node u directly pointing to node v, then node u is considered the parent node of node v, and node v is the child node of node u. This edge is called the outgoing edge of node u and the incoming edge of node v; node u is the starting node of the directed edge, and node v is the ending node. If we consider a directed edge as an arrow, with its starting position as the tail and its pointing position as the head, the starting node (source node) is also called the tail node, and the ending node (target node) is also called the head node.

[0033] In one possible implementation, a combination of an adjacency matrix and a node feature matrix can be used to represent a DAG, where the node feature matrix includes at least the identifiers of the nodes.

[0034] A subroutine of a complete computer program is also called a callable unit. Subroutines can be functions or methods. In this application, unless otherwise specified, subroutines, callable units, functions, and methods have similar meanings and can be used interchangeably.

[0035] Memory latency refers to the time difference between when the processor issues a data request and when the data actually arrives at the processor. Especially in scenarios requiring frequent access to main memory, it is a key factor affecting program execution efficiency. High memory latency can cause the processor to wait for data, thereby reducing overall system performance.

[0036] Data prefetching improves system performance by reducing memory latency. It can be implemented by hardware units such as processors or software modules such as compilers. At the hardware level, the processor's built-in prefetching mechanism automatically predicts and loads upcoming data into the cache, thus shortening waiting time. At the software level, the compiler analyzes the program's memory access patterns and inserts prefetch instructions at appropriate locations to load data into the cache in advance, reducing memory latency and improving execution efficiency. In this application's embodiments, the central processing unit (CPU) is used as an example to illustrate the processor's data prefetching method, but this should not be construed as a limitation of the technical solution of this application.

[0037] Memory access patterns refer to the patterns or rules by which a program accesses computer memory during runtime. Sequential access refers to a program accessing contiguous memory addresses sequentially. Random access refers to a program accessing memory locations without a clear pattern or order; this is common in applications that require frequent lookups and updates of different data structures (such as hash tables and binary trees). Due to its lack of predictability, memory latency issues with random access are usually difficult to optimize. Circular access refers to a program repeatedly accessing a fixed set of memory locations. This may occur within a loop's local variable usage or in operations on a fixed-size buffer. If the dataset accessed in a loop is small and can be completely cached, memory latency can be significantly reduced.

[0038] The prefetch address refers to the memory address predicted to be accessed in the future after analyzing program behavior. The prefetch distance refers to the difference or offset between the currently accessed data address and the prefetch address during data prefetching. It determines when and how data is preloaded into the cache to reduce memory latency. An ideal prefetch distance ensures that data to be used is loaded into the cache exactly when needed, maximizing cache hit rate and minimizing latency. If the prefetch distance is too small, data may be loaded into the cache too early, wasting cache resources or causing data to be overwritten; conversely, if the prefetch distance is too large, data may not arrive in the cache in time, causing the processor to wait for data loading and increasing memory access latency.

[0039] Memory references (MR) refer to the access operations on memory addresses during program execution, including loading and storing data. These references act as a bridge between program logic and physical memory, and have a crucial impact on program performance. Through memory references, programs can store and retrieve information about variables, array elements, and other data structures. Typically, the address targeted by memory reference operations is called the target address or access address. In data prefetching mechanisms, the predicted address is called the prefetch address when predicting the next memory address to be accessed.

[0040] A prefetch instruction is a special memory reference that is generated when the compiler or processor predicts that a program will access a specific memory location in the future, in order to begin reading the data at the prefetch address.

[0041] A typical data prefetching method used by a compiler mainly involves three steps: First, acquiring memory references. Specifically, the compiler identifies all memory references in the program, especially those in loops that may become performance bottlenecks. Next, analyzing memory references involves evaluating the memory access patterns of these operations, determining which memory addresses and their access intervals are suitable for prefetching, and calculating the optimal prefetch distance to maximize cache hit rate while minimizing prefetch overhead. The final step is inserting prefetch instructions. Based on the previous analysis results, prefetch instructions are automatically inserted at appropriate locations to load the required data into the cache in advance, reducing the time the processor waits for data and thus improving program execution efficiency.

[0042] However, when acquiring memory references, the compiler is limited by its inability to fully understand the entire program and its critical parts. This makes it difficult to accurately identify memory references crucial for optimization, potentially leading to the inappropriate insertion of useless prefetch instructions and wasting significant time on optimizing only a small percentage of critical code, while increasing code size. Furthermore, the process of analyzing memory references focuses only on the loop tree and memory references of the current function, failing to capture optimization opportunities in loop bodies and memory access patterns distributed across multiple functions. Therefore, during the prefetch instruction insertion phase, prefetch instructions can only be inserted within the current function's scope. When data access patterns and desired prefetch locations are distributed across different functions, the compiler cannot effectively perform such optimizations.

[0043] In view of this, embodiments of this application provide a data prefetching method, which is applied to a compiler. The method includes: acquiring a computer program, wherein multiple subroutines in the computer program form a loop structure and are used to process multiple pieces of data to be processed; determining, based on the calling relationship between the multiple subroutines and the loop body in the subroutines, the number of loop calls N of a first subroutine used to process the multiple pieces of data to be processed in the loop structure; determining, based on N, a first asymptotic growth rate of the runtime of the loop structure; and, if the first asymptotic growth rate is greater than or equal to a preset asymptotic growth rate, inserting a prefetch instruction in at least one subroutine.

[0044] In this embodiment of the application, the computer program code that forms a loop structure and is used to process multiple data to be processed is distributed in multiple subroutines. The number of loop calls of the first subroutine in the loop structure can be used to calculate the incremental growth rate of the loop structure's runtime, so as to determine whether it is necessary to insert prefetch instructions in the loop structure, thereby improving the running efficiency of the computer program.

[0045] The following is combined Figure 1 The flowchart shown illustrates the data prefetching method and explains the specific process by which the compiler executes the data prefetching method described in this application.

[0046] S110, Obtain a computer program, wherein multiple subroutines in the computer program form a loop structure and are used to process multiple data to be processed.

[0047] A call graph is a graph that represents the calling relationships between subroutines in a program. In a call graph, each node represents a subroutine, and each edge represents a calling relationship between two subroutines. For example, in the main procedure (such as the `main` function in C), if subroutine #A calls subroutine #B, then the call graph includes at least node #A corresponding to subroutine #A, node #B corresponding to subroutine #B, and directed edges from node #A to node #B. Call graphs help compilers or analysis tools understand the overall structure of a program, identify the calling relationships between subroutines, and thus perform more effective optimizations.

[0048] It should be understood that in conventional computer programs, subroutines typically do not include recursive call relationships; that is, the call graph is usually a directed acyclic graph. The method described in the embodiments of this application is mainly applied to computer programs that do not include recursive calls.

[0049] The following is combined Figure 2 The call diagram shown illustrates the process of applying the data prefetching method described in this application to a specific computer program.

[0050] In one possible implementation, such as Figure 2As shown in the topmost code block, the computer program first defines a special array ( Figure 2 The Array structure in the database), this array includes the base address ( Figure 2 (buffer in the array), number of elements in the array ( Figure 2 The size in the text) and the cursor indicating the current position ( Figure 2 (cursor in the array). After processing the i-th data in the array, the cursor can be moved to the (i+1)-th data. That is, when processing the data in the array in the loop body, the loop body does not need to set a specific inductive variable, and the cursor of the array can play a similar role to the inductive variable.

[0051] In one possible implementation, Figure 2 The call diagram shown illustrates the main program ( Figure 2 The main function in the code (for subroutines #1 to #7) Figure 2 The call relationship between funcA1 to funcA7 in the code. Figure 2 The code shown is only a portion of the total code for the corresponding subroutines, and this application embodiment does not limit this. It can be seen that these subroutines use the Obj structure to pass shared data between them, and some subroutines include cyclic calls to other subroutines.

[0052] S120, based on the calling relationship between multiple subroutines and the loop body in the subroutines, determine the number of loop calls N of the first subroutine used to process multiple data to be processed in the loop structure.

[0053] A control flow graph (CFG) is used to represent the possible paths during the execution of a subroutine. In an CFG, nodes represent basic blocks, which are sequential sequences of code that are executed without internal jumps, while directed edges represent control transfer paths from one basic block to another, including conditional branches, loops, and unconditional jumps.

[0054] Specifically, control flow graphs can be used to analyze in detail the number of times loops or branch paths are executed, quantifying the dynamic behavior within subroutines. In a control flow graph, the number of times subroutine #B is called by subroutine #A in a loop is called the trip count of subroutine #B within the loop body of subroutine #A; an induction variable (IV) is a variable that increments or decrements according to a fixed pattern within the loop body. These two types of variables are often used together; they help identify regular memory access sequences, enabling the compiler to perform effective data prefetching optimizations.

[0055] First, let's illustrate fixed-number calls and cyclic calls between subroutines with examples. A fixed-number call from subroutine #A to subroutine #B refers to the explicit and limited number of calls to subroutine #B during the execution of subroutine #A. This number of calls does not change due to external conditions or variations in input data. For example, in one scenario, subroutine #A needs to first use a verification function (i.e., subroutine #B) to verify whether data stream #1 conforms to a certain specification; then, it generates data stream #2 based on data stream #1 and calls the verification function again to check whether data stream #2 conforms to the specification. Whether subroutine #A verifies only data stream #1, only data stream #2, or both simultaneously, these operations constitute a predefined logical structure in which the number of calls to subroutine #B is always fixed.

[0056] In contrast, the cyclic calls of subroutine #A to subroutine #B are typically dynamic, with the number of calls depending on the program's input or other runtime conditions. For example, subroutine #A includes a loop body that makes M loop calls. In the i-th loop call, subroutine #A first uses a verification function (i.e., subroutine #B) to verify whether the i-th data stream #1 conforms to a certain specification. Then, it generates the i-th data stream #2 based on the i-th data stream #1 and calls the verification function again to check whether the i-th data stream #2 conforms to the specification. In each loop call, subroutine #A calls subroutine #B a fixed number of times; therefore, the stroke count of subroutine #B within the loop body of subroutine #A is M, not 2M.

[0057] Furthermore, the operation of updating the inductive variable inside the loop body is called an update expression, such as "i = i + 1" or "i++" in C language.

[0058] The base address typically refers to the starting memory address of an array, structure, or other data structure. It is a reference point relative to the addresses of other elements within that data structure. For example, in an array, the memory address of the first element is the base address of the array.

[0059] S121, determine at least one second subroutine among a plurality of subroutines, wherein at least one second subroutine is used to perform N loop calls on the first subroutine according to the first dataset, and one of the at least one second subroutines is used to perform M loop calls on the first subroutine in a loop body with a stroke count of M, and the first dataset is used to determine the target address of the data to be processed.

[0060] In one possible implementation, the first dataset is Figure 2In each subroutine, the variable `obj` has the attribute `array1`, and the data to be processed is the attribute `array2` of the variable `obj`. Specifically, subroutines #1 to #5... Figure 2 None of funcA1 to funcA5 involve operations on any element of array2. Therefore, it can be assumed that subroutines #1 to #5 do not contain code related to program performance; their actual function is to count the number of iterations based on the loop body. Figure 2 M1 and M2) and inductive variables ( Figure 2 The cursor attribute of the variable array1 in the variable is used to determine subroutine #6. Figure 2 The timing and number of calls to `funcA6`. In other words, subroutines #1 to #5 are all second subroutines.

[0061] S122, determine the first subroutine, wherein the first subroutine is used to determine the target address of the data to be processed based on the first dataset and the inductive variable of the loop body, and to reference the target address in memory.

[0062] In one possible implementation, the first subroutine uses random access to the memory addresses of multiple data items to be processed. Specifically, as shown below... Figure 2 As shown, subroutine #6 ( Figure 2 The first dataset in funcA6) Figure 2 The cursor of variable `array1` is used to indicate the i-th loop call of subroutine #6 by the second subroutine. First, subroutine #6 determines the cursor value of the first dataset to be `i`, which is the current value of the inductive variable. Then, it obtains `loc1` and `loc2` based on `funcD1` and `funcD2` respectively. Next, it determines the target addresses of the two data items to be processed in `array2` based on the base address of variable `array2` and `loc1` and `loc2`. Finally, it processes `param1` and `param2` obtained from the two target addresses respectively and writes the processed `param1` and `param2` back to their corresponding target addresses. In other words, subroutine #6 is the first subroutine.

[0063] Obviously, due to the specific process of determining the target address of the data to be processed ( Figure 2 The values ​​of funcD1 and funcD2 in the dataset are arbitrarily set. In most cases, the first dataset is accessed sequentially. Figure 2 The target address obtained from the elements in the variable array1 is usually equivalent to the base address of the random access target address. Figure 2 The variable array2 contains elements in the buffer attribute. Therefore, subroutine #6 includes code related to program performance; the data prefetching method in this application is mainly used to prefetch the data to be processed in array2.

[0064] In hardware-level data prefetching methods, the processor's built-in prefetching mechanism can automatically predict and load data to be accessed in sequential and cyclic access. In this implementation, the data prefetching method of this application can focus only on the data to be processed and the prefetch instructions related to random access, avoiding the inappropriate insertion of useless prefetch instructions, thereby improving the running efficiency of the computer program.

[0065] S123, determine the third subroutine called by the first subroutine. In the i-th cyclic call of the first subroutine by at least one second subroutine, the third subroutine is used to instruct at least one second subroutine to make the (i+1)-th cyclic call of the first subroutine.

[0066] In one possible implementation, the third subroutine includes an update expression for updating the inductive variable of the loop body. Specifically, as follows: Figure 2 As shown, subroutine #7 ( Figure 2 The statement "obj->array1->cursor++" in funcA7 is used to make the first dataset ( Figure 2 The cursor of the variable array1 is incremented to point to the next element, indicating the (i+1)th call of the first subroutine to the second subroutine. In other words, subroutine #7 is the third subroutine.

[0067] It should be understood that, according to the foregoing embodiments, i≤N, M≤N, and i, M, and N are positive integers.

[0068] In this implementation, the first, second, and third subroutines, which form a loop structure and are used to process multiple unprocessed data, correspond to specific functions and have complex calling relationships. This application can use the sorted calling relationships to determine the asymptotic growth rate of the number of loop calls of the first subroutine in the loop structure, so as to calculate the asymptotic growth rate of the runtime of the loop structure, thereby improving the efficiency of determining whether prefetch instructions need to be inserted into the loop structure.

[0069] S124, based on the calling relationship of the first subroutine, at least one second subroutine, and the third subroutine, determine the relationship between N and the stroke count M of the loop body of at least one second subroutine, wherein the asymptotic growth rate of N with respect to at least one M is used to determine the first asymptotic growth rate.

[0070] Data flow analysis deals with how variables are passed between different functions and how they are used throughout the program's lifecycle. This analysis focuses on the definition, use, transmission, and possible range of variables to understand how data flows between different parts of the program.

[0071] Inter-procedural analysis (IPA) is used to analyze the interactions and effects between different functions or methods in a program. Unlike local analysis, which is limited to the interior of a single function, IPA can cross function call boundaries, considering factors such as parameter passing, return value usage, and the impact of global variables to provide more accurate data flow and control flow information. Through this cross-procedural analysis, the compiler can identify more optimization opportunities, thereby improving the overall performance of the program. Specifically, IPA uses call graphs to identify the propagation paths of memory reference instructions throughout the program and uses data flow analysis to identify the behavior of these memory reference instructions inside functions.

[0072] S1241, a directed acyclic graph is determined based on multiple subroutines, each subroutine is used as a graph node of the directed acyclic graph, and the calling relationship between the multiple subroutines is used as at least one directed edge of the directed acyclic graph.

[0073] The method of combining call graphs and control flow graphs to obtain information such as call relationships between subroutines, stroke counts of subroutine loop bodies, and the number of loop calls of subroutines within the entire loop structure can be called inter-process analysis. Specifically, it can combine... Figure 2 The call graph shown and the control flow graph of each subroutine of the computer program yield the following results: Figure 3 The diagram shows a directed acyclic graph.

[0074] The following is combined Figure 3 The directed acyclic graph shown illustrates the process of applying the data prefetching method described in this application to a specific computer program.

[0075] In some embodiments, the second subroutine is an ancestor node of the first subroutine, the first subroutine is an ancestor node of the third subroutine, and the travel count M of the loop body of the second subroutine is used to determine the first weight of the directed edge from the second subroutine to the first subroutine.

[0076] In one possible implementation, Figure 2 funcA1 to funcA7 are respectively used as Figure 3 Subroutines #1 to #7 are included. Figure 3 In the directed acyclic graph shown, subroutines #1 to #5 are all ancestor nodes of subroutines #6. According to the aforementioned embodiment, subroutines #1 to #5 are all second subroutines, therefore the second subroutines are ancestor nodes of the first subroutines. Subroutines #6 (the first subroutines) call subroutines #7 (the third subroutines) after referencing the target address in memory; therefore, the first subroutines are ancestor nodes of the third subroutines.

[0077] In one possible implementation, according to such Figure 2The call graph shown indicates that subroutine #1 includes M1 loop calls to subroutine #2, subroutine #2 includes a fixed number of calls to subroutine #3, and M2 loop calls to subroutine #4. The remaining call edges in the program represent subroutine call relationships with a fixed number of calls. M1 and M2 can be determined based on the number of elements in the first dataset; that is, M1 and M2 are determined based on the input of the computer program.

[0078] It should be understood that Figure 2 The loop body in the call diagram shown explicitly specifies the stroke counts M1 and M2. In this embodiment, the actual stroke count of a loop body that does not explicitly specify the stroke count can also be obtained according to the IPA, such as the stroke count of a while loop body.

[0079] In one possible implementation, Figure 3 The first weight of the directed edge is shown, denoted by W1. Subroutine #1 includes M1 loop calls to subroutine #2 (i.e., the stroke count M of the loop body of subroutine #1 is M1), so the first weight of the directed edge from subroutine #1 to subroutine #2 can be M1; similarly, subroutine #2 includes M2 loop calls to subroutine #4, so the first weight of the directed edge from subroutine #2 to subroutine #4 can be M2.

[0080] Furthermore, since subroutine #4 includes a fixed number of calls to subroutine #6, the weight of the directed edge from subroutine #4 to subroutine #6 is 0.

[0081] In this implementation, the number of times each subroutine is called by the parent node in a loop can be determined based on the first weight of the directed edges between subroutines, and N can be represented by multiple process counts M, thereby improving the efficiency of determining whether prefetch instructions need to be inserted into the loop structure.

[0082] In some embodiments, the rank of a plurality of subroutines is determined based on a first weight of at least one directed edge, wherein the rank of a subroutine is determined based on a second weight of at least one incoming edge of the subroutine, the second weight of the incoming edge of the subroutine being determined based on the first weight of the incoming edge and the rank of the parent node of the subroutine in the directed acyclic graph.

[0083] Specifically, Figure 4 It shows the relationship with Figure 3 The same directed acyclic graph is shown, with the second weight of the directed edges illustrated, denoted by W2. For ease of explanation, Figure 4 The subroutines in the diagram illustrate the first weight, the second weight, and the value of the subroutines for the directed edges, with the rank denoted by R. According to the aforementioned embodiment, Figure 4The directed acyclic graph shown has only one source node, which does not include any incoming edges, while the other graph nodes include one or more incoming edges.

[0084] In one possible approach, when the rank of the parent node is equal to 0, the second weight of the incoming edge is the same as the first weight of the incoming edge. Specifically, Figure 4 In the given equation, subroutine #1 is the parent node of subroutine #2, and subroutine #1 is the source node. Since it does not include incoming edges, the rank of subroutine #1 is 0. The second weight of the directed edge from subroutine #1 to subroutine #2 (i.e., the incoming edge of subroutine #2) is its own first weight, which is M1. Similarly, the second weight of the directed edge from subroutine #1 to subroutine #5 is its own first weight, which is 0.

[0085] In one possible approach, if the first weight of the incoming edge is equal to 0, the second weight of the incoming edge is the rank of the parent node. Specifically, Figure 4 Subroutine #2 is the parent node of subroutine #3. The first weight of the directed edge from subroutine #2 to subroutine #3 is 0. Therefore, its second weight is the rank of subroutine #2, which is M1.

[0086] In one possible approach, if both the first weight of the incoming edge and the rank of the parent node are greater than 0, the second weight of the incoming edge is the product of the first weight of the incoming edge and the rank of the parent node. Specifically, Figure 4 Subroutine #2 is the parent node of subroutine #4. The first weight of the directed edge from subroutine #2 to subroutine #4 is greater than 0, and the rank of subroutine #2 is greater than 0. Therefore, the second weight of this edge is the product of its own first weight and the rank of subroutine #2, which is M1*M2.

[0087] In one possible approach, the rank of the subroutine is the sum of the second weights of at least one incoming edge of the subroutine. According to the foregoing embodiment, the first weight is the stroke count of the subroutine's loop body or 0, and therefore must be a non-negative number. Figure 4 The rank of subroutine #4 can be the sum of the second weights of all incoming edges to subroutine #4. Specifically, subroutine #4 has three incoming edges: the second weight of the directed edge from subroutine #2 to subroutine #4 is M1*M2, the second weight of the directed edge from subroutine #3 to subroutine #4 is M1, and the second weight of the directed edge from subroutine #5 to subroutine #4 is 0. Therefore, the rank of subroutine #4 is the sum of these three values, i.e., M1*M2+M1.

[0088] It should be understood that, according to the foregoing embodiments, the first weight of the directed edge, the second weight of the directed edge, and the rank of the subroutine are all non-negative numbers.

[0089] S1242, determining the relationship between N and at least one M according to the first weight of at least one directed edge.

[0090] In a possible implementation, the relationship between N and at least one M is determined according to the rank of the first subroutine.

[0091] Specifically, Figure 4 the rank of a subroutine therein indicates the number of times the subroutine is circularly called in a computer program. Therefore, the number of times subroutine #6 is circularly called is M1*M2+M1, that is, N=M1*M2+M1.

[0092] Further, subroutine #2 and subroutine #3 are called M1 times in the loop, then the trip count of the loop bodies of subroutine #2 and subroutine #3 is M1, that is, M1 is used as M, which satisfies M1<N; the trip count of the loop body of subroutine #4 is M1*M2+M1, that is, M1*M2+M1 is used as M, which satisfies M1=N.

[0093] S130, determining a first asymptotic growth rate of the running duration of the loop structure according to N.

[0094] In some embodiments, the first asymptotic growth rate is determined according to the asymptotic growth rate of N with respect to at least one M. Specifically, the first subroutine is the main time-consuming operation in the entire loop structure, and N is the number of circular calls of the first subroutine in the loop structure. Therefore, the growth rate of the number of circular calls of N with respect to M can be used as the asymptotic growth rate of the running duration of the loop structure.

[0095] In a possible implementation, the asymptotic growth rate of N may be the asymptotic growth rate of the time complexity of N, or the asymptotic growth rate of the number of circular calls of N. In this embodiment, the asymptotic growth rate of N is O(M1*M2+M1). In the case of a detailed compiler configuration file, the preset asymptotic growth rate may be set as O(M1*M2), which is substantially the same asymptotic growth rate as that of N, therefore a prefetch instruction needs to be inserted into the computer program. If in another embodiment, the asymptotic growth rate of N is O(M1+M2), then the asymptotic growth rate of N is smaller than the preset asymptotic growth rate, and no prefetch instruction needs to be inserted.

[0096] In the case that the compiler configuration file adopts default options, the compiler can identify that both M1 and M2 are trip counts M of loop bodies, but cannot perform more detailed analysis. Therefore, in this embodiment, the asymptotic growth rate of N is O(M2+M). Further, the preset asymptotic growth rate may be set as O(M2), which is substantially the same asymptotic growth rate as that of N, therefore a prefetch instruction needs to be inserted into the computer program.

[0097] In this implementation manner, the rank of a subroutine indicates the number of loop calls of the subroutine in the loop structure, and the definition of the rank of the first subroutine is the same as the definition of N. Therefore, the rank of the first subroutine can be used as N, thereby improving the efficiency of determining whether a prefetch instruction needs to be inserted into the loop structure.

[0098] S140, inserting a prefetch instruction into at least one subroutine forming the loop structure when the first asymptotic growth rate is greater than or equal to a preset asymptotic growth rate.

[0099] In some embodiments, a prefetch instruction is inserted into the second subroutine or the first subroutine, and in the j-th cyclic call of at least one second subroutine to the first subroutine, the prefetch address of the prefetch instruction is the target address of the data to be processed in the (j+t)-th cyclic call, where j < j+t ≤ N, and j and t are positive integers.

[0100] It should be understood that the method for determining that the asymptotic growth rate of N is greater than or equal to the preset asymptotic growth rate has been described in detail in the foregoing embodiments, and will not be repeated here. The steps of this embodiment are all performed when the asymptotic growth rate of N is greater than or equal to the preset asymptotic growth rate.

[0101] In some embodiments, according to Figure 4 the call relationship between subroutines shown, the prefetch instruction can be inserted into subroutine #4 or subroutine #6. Since these subroutines may include a large amount of code with long execution time, it is possible that the prefetch distance set when inserting a prefetch instruction at the start position of a subroutine is different from the prefetch distance set when inserting a prefetch instruction at the end position of the subroutine. That is, inserting a prefetch instruction into subroutine #4 or subroutine #6 substantially only causes a change in the required prefetch distance. Therefore, when considering the position to insert the prefetch instruction, the readability of the code should be prioritized, for example, inserting the prefetch instruction into the definition domain of the variable related to the prefetch address.

[0102] In a possible implementation manner, inserting a prefetch instruction into subroutine #6 (i.e., the first subroutine), as Figure 5 shown, funcA8 is subroutine #6 after inserting the prefetch instruction, PREFETCH_DISTANCE is the prefetch distance, which is a constant in the code, and the specific value of the constant can be determined by performing performance analysis on a specific computer program obtained after compilation. For example, if PREFETCH_DISTANCE = 2, then in the j-th cyclic call of the second subroutine to the first subroutine, the prefetch address of the prefetch instruction is the target address of the (j+2)-th cyclic call.

[0103] In a possible implementation manner, the specific method for determining the prefetch address according to the value of j+t is similar to the method for determining the target address according to the first data set and the value of j. For example, as Figure 5 As shown, during the normal processing of data at the target address, the current inductive variable has a value of j. Based on funcD1 and funcD2, loc1 and loc2 are obtained respectively. Then, the target address is determined based on the base address of variable array2 and loc1 and loc2. Similarly, when determining the prefetch address, the prefetch inductive variable has a value of j+2. pLoc1 and pLoc2 can be obtained based on funcD1 and funcD2 respectively. Then, the prefetch address is determined based on the base address of variable array2 and pLoc1 and pLoc2.

[0104] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the embodiments of this application include corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0105] This application also provides a chip system 800, such as... Figure 6 As shown, the chip system 800 includes at least one processor and at least one interface circuit. As an example, when the chip system 800 includes one processor and one interface circuit, the processor can be... Figure 6 The processor 810 shown in the solid box (or the processor 810 shown in the dashed box) can have an interface circuit that is... Figure 6 The interface circuit 820 is shown in the solid box (or the interface circuit 820 is shown in the dashed box).

[0106] When the chip system 800 includes two processors and two interface circuits, then the two processors include Figure 6 The processor 810 shown in the solid box and the processor 810 shown in the dashed box, these two interface circuits include Figure 6 Interface circuit 820 is shown in both solid and dashed boxes. This is not a limitation. Processor 810 and interface circuit 820 can be interconnected via lines. For example, interface circuit 820 can be used to receive signals (e.g., instructions stored in memory). As another example, interface circuit 820 can be used to send signals to other devices (e.g., processor 810).

[0107] For example, interface circuit 820 can read instructions stored in memory and send those instructions to processor 810. When the instruction is executed by processor 810, it can cause a device for accessing memory or a device for accessing memory to perform the steps in the above embodiments. Of course, the chip system 800 may also include other discrete devices, and this application embodiment does not specifically limit this.

[0108] This application also provides a computing device 900. For example... Figure 7 As shown, the computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate with each other via the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 900.

[0109] The 902 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus 904 is represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 904 may include a path for transmitting information between various components of the computing device 900 (e.g., memory 906, processor 904, communication interface 908).

[0110] Processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0111] The memory 906 may include volatile memory, such as random access memory (RAM). The processor 904 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0112] The memory 906 stores executable program code, which the processor 904 executes to implement the aforementioned data prefetching methods. That is, the memory 906 stores instructions for executing the data prefetching methods.

[0113] The communication interface 908 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 900 and other devices or communication networks.

[0114] Another embodiment of this application provides a computer-readable storage medium storing instructions that, when executed on a means for accessing memory, perform the various steps of the method flow shown in the above-described method embodiments. In some embodiments, the disclosed method can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of art.

[0115] Figure 8 A conceptual partial view of a computer program product provided in an embodiment of this application is shown schematically. The computer program product includes a computer program for executing computer processes on a computing device.

[0116] In one embodiment, a computer program product is provided using a signal bearer medium 1000. The signal bearer medium 1000 may include one or more program instructions that, when run by one or more processors, can execute the data prefetching method provided in this application embodiment.

[0117] In some examples, the signal carrying medium 1000 may include a computer-readable medium 1001, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), a digital magnetic tape, a memory, a read-only memory (ROM), or a random access memory (RAM), etc.

[0118] In some implementations, the signal carrying medium 1000 may include a computer recordable medium 1002, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, and so on.

[0119] In some implementations, the signal-bearing medium 1000 may include a communication medium 1003, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.). The signal-bearing medium 1000 may be transmitted by a wireless communication medium 1003 (e.g., a wireless communication medium conforming to the IEEE 1502.11 standard or other transmission protocols). One or more program instructions may be, for example, computer-executable instructions or logical implementation instructions.

[0120] In some examples, the computing device may be configured to provide various operations, functions, or actions in response to one or more program instructions in a computer-readable medium 1001, a computer-recordable medium 1002, and / or a communication medium 1003.

[0121] It should be understood that the arrangements described herein are for illustrative purposes only. Therefore, those skilled in the art will understand that other arrangements and other elements (e.g., machines, interfaces, functions, sequences, and functional groups, etc.) can be used instead, and some elements may be omitted depending on the desired outcome. Furthermore, many of the described elements are functional entities that can be implemented as discrete or distributed components, or in any suitable combination and location with other components.

[0122] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, it can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When executed on a computer and when the computer execution instructions are executed, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0123] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access, or it can contain one or more data storage devices such as servers or data centers that can be integrated with that medium. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0124] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data prefetching method, characterized in that, The method is applied to a compiler, and the method includes: A computer program is obtained, wherein multiple subroutines in the computer program form a loop structure and are used to process multiple pieces of data to be processed; Based on the calling relationships between the multiple subroutines and the loop bodies in the subroutines, determine the number of loop calls N of the first subroutine used to process the multiple data to be processed in the loop structure; The first asymptotic growth rate of the runtime of the loop structure is determined based on N; If the first progressive growth rate is greater than or equal to a preset progressive growth rate, a prefetch instruction is inserted in at least one of the subroutines.

2. The method according to claim 1, characterized in that, The first subroutine uses random access to the memory address of the target address of the multiple data to be processed.

3. The method according to claim 1 or 2, characterized in that, Determining the number of times N is called in the loop structure for the first subroutine includes: At least one second subroutine is determined among the plurality of subroutines, wherein the at least one second subroutine is used to call the first subroutine N times in a loop according to the first dataset, and one of the at least one second subroutines is used to call the first subroutine M times in a loop body with a stroke count of M, wherein the first dataset is used to determine the target address of the data to be processed. The first subroutine is determined, wherein the first subroutine is used to determine the target address of the data to be processed based on the first dataset and the inductive variable of the loop body, and to reference the target address in memory; Determine the third subroutine called by the first subroutine. In the i-th cyclic call of the first subroutine by the at least one second subroutine, the third subroutine is used to instruct the at least one second subroutine to make the (i+1)-th cyclic call of the first subroutine, where i≤N, M≤N, and i, M, and N are positive integers. Based on the calling relationship between the first subroutine, the at least one second subroutine, and the third subroutine, the relationship between N and the run count M of the loop body of the at least one second subroutine is determined, wherein the relationship between N and the at least one M is used to determine the first progressive growth rate.

4. The method according to claim 3, characterized in that, Determining the relationship between N and the at least one M includes: A directed acyclic graph is determined based on the plurality of subroutines, wherein the plurality of subroutines are respectively used as a plurality of graph nodes of the directed acyclic graph, and the calling relationship between the plurality of subroutines is used as at least one directed edge of the directed acyclic graph, wherein the second subroutine is an ancestor node of the first subroutine, the first subroutine is an ancestor node of the third subroutine, and the travel count M of the loop body of the second subroutine is used to determine the first weight of the directed edge from the second subroutine to the first subroutine; The relationship between N and the at least one M is determined based on the first weight of the at least one directed edge.

5. The method according to claim 4, characterized in that, Determining the relationship between N and at least M based on the first weight of the at least one directed edge includes: Determining a relationship between N and the at least one M according to the rank of the first subprogram, wherein the rank of the subprogram is determined according to the second weight of at least one incoming edge of the subprogram, and the second weight of the incoming edge of the subprogram is determined according to the first weight of the incoming edge and the rank of a parent node of the subprogram in the directed acyclic graph.

6. The method according to claim 5, characterized in that, The first weight of the directed edge, the second weight of the directed edge and the rank of the subprogram are all non-negative numbers, and the rank of the first subprogram is a sum of the second weights of at least one incoming edge of the first subprogram, wherein: In a case where the rank of the parent node of the first subprogram in the directed acyclic graph is equal to 0, the second weight of the incoming edge is the first weight of the incoming edge; In a case where the first weight of the incoming edge is equal to 0, the second weight of the incoming edge is the rank of the parent node; In a case where both the first weight of the incoming edge and the rank of the parent node are greater than 0, the second weight of the incoming edge is a product of the first weight of the incoming edge and the rank of the parent node.

7. The method according to any one of claims 3 to 6, characterized in that, The inserting a prefetch instruction into at least one of the subprograms comprises: inserting a prefetch instruction into the second subprogram or the first subprogram, wherein in the j-th cyclic call to the first subprogram by the at least one second subprogram, a prefetch address of the prefetch instruction is a target address of the to-be-processed data in the (j+t)-th cyclic call, and j<j+t≤N, wherein j and t are positive integers.

8. A computer device, characterized in that, The computer device comprises at least one processor and at least one memory; the at least one memory stores at least one computer program, the at least one computer program comprises instructions, and when the instructions are executed by the at least one processor, the method according to any one of claims 1 to 7 is performed.

9. A chip, characterized in that, The chip comprises a processor and a communication interface, wherein the communication interface is configured to receive a signal and transmit the signal to the processor, and the processor is configured to process the signal, such that the method according to any one of claims 1 to 7 is performed.

10. A computer-readable storage medium, characterized in that, Computer instructions are stored in the computer-readable storage medium, and when the computer instructions run on a computer device, the method according to any one of claims 1 to 7 is performed.

11. A computer program product containing instructions, characterized in that, When the instructions are run by a computer device, the instructions cause the computer device to perform the method according to any one of claims 1 to 7.