Parallel Optimization Method for Heterogeneous Platforms for Hierarchical Stepping Transformation Loops with Dependencies
By extracting the outer layer dependent data and adjusting the loop structure into a one-layer loop, the problem of difficult to parallelize the layered step transformation loops in the CPU-GPU heterogeneous platform is solved, and the execution efficiency of the program is improved.
Patent Information
- Application Number
- CN202510429446.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the CPU-GPU heterogeneous platform, hierarchical step transformation loops with dependencies are difficult to efficiently parallelize, resulting in frequent start and stop of kernel functions and idle thread counts, resulting in wasted time.
By extracting the outer layer dependent data and caching it into an array, adjusting the loop structure into a one-layer loop, reducing discontinuous memory access and improving spatial locality.
It reduces the start and stop frequency of kernel functions, optimizes thread allocation, avoids thread idleness, and significantly improves program execution efficiency.
Smart Images

Figure CN119938281B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a heterogeneous platform parallel optimization method for a layered step transformation loop containing dependencies, and belongs to the field of electronic information technology. Background Art
[0002] As computing demands continue to increase, traditional single processor architectures are gradually unable to meet the requirements of efficiently processing large-scale computing tasks. In recent years, heterogeneous computing platforms (CPU combined with GPU) have gradually become an important way to solve this problem. The central processing unit (CPU) is the computing and control core of the computer system. The graphics processing unit (GPU) was originally designed to accelerate 3D image rendering and graphics drawing such as video, but because of its powerful parallel capabilities, it has been widely used in complex computing problems. Compared with traditional CPUs, GPUs have large-scale parallel processing capabilities and can perform multiple computing tasks simultaneously on a large scale.
[0003] CUDA (Compute Unified Device Architecture) is a parallel computing platform and programming model developed by NVIDIA that allows developers to use languages such as C, C++, and Fortran to perform high-performance computing on GPUs. The CUDA programming model is divided into the host side and the device side. The device side code running on the GPU is called a kernel. Before the kernel function runs, it first submits the data stored on the host side to the device side's video memory, then calls the GPU's multithreading to calculate the kernel function, and finally transfers the calculated data back to the host side's memory. When the kernel is launched to the GPU for runtime, at the logical level, such as Figure 1 As shown in the CUDA programming model, NVIDIA GPUs use thread blocks and grids to organize and manage threads so that the kernel gets the correct index when running. The total number of threads in the kernel is equal to the size of the grid multiplied by the size of the thread block. However, at the hardware level, the warp scheduler schedules GPU threads in a thread block to execute instructions on a specified device in units of warps. Warps are the basic unit for GPU scheduling and operation, and consist of 32 threads in the same thread block. Even if there are less than 32 threads remaining in the block, they will be scheduled as a warp.
[0004] In CPU-GPU heterogeneous parallel computing, the common approach is to partition tasks by granularity and assign the computationally intensive parts suitable for data parallelism to the GPU for execution, while the CPU undertakes input / output processing and serial computing tasks. In complex numerical computing applications, many tasks involve hierarchical stepwise transformation loops, which usually exhibit characteristics such as strong data dependence, time-variance, and uncertainty. These characteristics make it difficult for traditional parallelization strategies (such as static task decomposition and static scheduling) to handle efficiently.
[0005] In the optimization of CPU-GPU heterogeneous platforms, the core objective is to transplant large-scale computationally intensive tasks that can be parallelized in the program to the GPU and use CUDA for parallel acceleration. Many complex applications contain hierarchical stepwise transformation loops, where the number of iterations and step size in each layer are usually dynamically determined by input parameters or external input values. This makes the number of iterations of the loop structure irregular, possibly resulting in a large variation in the number of iterations between different layers. In addition, there are data associations and dependencies between the outer and inner layer computations of this loop structure.
[0006] For the parallel optimization of such loop structures, ordinary optimization schemes cannot select the entire hierarchical loop for parallelization. They can only assign threads to some of the loops that control the inner layer computations, and the number of threads contained in each thread block in each dimension is a fixed number. This ordinary GPU optimization scheme will cause the situation where the outer loop nests kernel functions in the program, resulting in frequent calls to the kernel functions, thus wasting time. At the same time, the number of threads assigned to each thread block in each dimension cannot be changed during each loop. Facing the dynamic number of iterations, there will be a situation where a large number of threads are idle. Summary of the Invention
[0007] Aiming at the deficiencies of the prior art, the present invention provides a parallel optimization method for heterogeneous platforms for hierarchical stepwise transformation loops with dependencies;
[0008] The present invention proposes a novel computing optimization method. By analyzing the data interaction and dependency relationships between different levels, the outer-layer numerical calculations are extracted before the loop execution, and the calculation results are cached in an array. This optimization can ensure that more loop levels can be efficiently mapped to the GPU, improving the parallel computing efficiency. For multi-level strided iterative loops, a strategy of simplifying the nested loops into a single-layer iteration is adopted. Specifically, first, the total number of iterations and its logical relationship of the hierarchical loop are determined. Based on this relationship, the boundaries of the one-dimensional loop are delimited, and the execution order of the original loop array is adjusted to make it consecutive access; second, to avoid multiple accesses between different positions of the same array in deep-level calculations, a formula is proposed for execution judgment, thus significantly reducing unnecessary access operations; finally, for the problem of non-consecutive array access in the outer-layer calculations, an improved calculation and call strategy is proposed to reduce non-consecutive memory access and improve spatial locality. Through the above optimization method, the number of iterations of the data becomes more determined, so that thread blocks and grids can be more efficiently allocated during the GPU transplantation process. The present invention reduces the startup and shutdown frequency of the kernel function, optimizes the thread allocation, avoids thread idleness, and significantly improves the execution efficiency of the program.
[0009] When optimizing a program on a CPU-GPU heterogeneous system, when there are hierarchical stepped transformation loops with dependencies, it is difficult to select a suitable dimension for parallelism and reasonably allocate the number of threads for each dimension during the parallelization according to the ordinary GPU optimization method during the transplantation process, resulting in idle threads and thus wasting a large amount of time. For programs with the above problems, selecting the solution of the present invention for optimization can effectively solve the problem and improve the execution efficiency of the program. At the same time, the present invention also designs an automated interface SLHO (Solving Loop Hierarchical Optimization) for convenient direct call by programmers.
[0010] The technical solution of the present invention is as follows:
[0011] A heterogeneous platform parallel optimization method for hierarchical stepped transformation loops with dependencies, including:
[0012] 1) The CPU preprocesses the outer-layer data dependencies;
[0013] In the hierarchical stepped transformation loop, the outer-layer dependent data is extracted from the entire hierarchical stepped transformation loop and independently executed for calculation on the CPU side;
[0014] 2) The hierarchical loop is adjusted to a single-layer loop;
[0015] 3) Remap the dependent arrays after extraction.
[0016] Preferably according to the present invention, the data calculated by the outer layer is extracted from the entire hierarchical step - transformation loop and independently executed on the CPU side; including:
[0017] The outer - layer dependent data is stored in an array after extraction, and when the kernel function is executed, the outer - layer dependent data is transferred from the CPU to the GPU.
[0018] Further preferably, when extracting the outer - layer dependent data, by adjusting the calculation method, the data storage method is adjusted to a continuous access mode; specifically including:
[0019] Let p represent the array subscript of the stride iteration, ip represent the stride step size of the original calculation loop, the number of loop executions be x, i represent the value of the continuously accessed array subscript, and there is p=(x - 1)*ip + 1;
[0020] After adjusting the loop step size to 1, the array is calculated and stored continuously. The data originally stored at position p is stored at position x, and at this time, the value of x is equal to the value of i;
[0021] After extracting the outer - layer dependent data, the array still executes according to the original stride - loop logic during the call process, and the dependency relationship between data is restored through the mapping formula (1):
[0022] (1).
[0023] Preferably according to the present invention, a loop with N layers of indefinite iteration times and stride step sizes is changed to a single - layer loop; N≥3; including:
[0024] First, by analyzing the loop, clarify the iteration times of each layer in the hierarchical loop, and determine the boundary n of the single - layer loop through the formula (2):
[0025] (2);
[0026] Let the loop variable of the outermost loop be i1, the starting value be 1, the stride step size be ip1, and the boundary value be n1; the loop variable of the second - layer loop be i2, the starting value be i1, the stride step size be ip2, and the boundary value be n2; the loop variable of the innermost loop be i3, the starting value be i2, the stride step size be ip3, and the boundary value be n3;
[0027] After determining the single - layer loop boundary n, set the single - layer loop index as k1. When k1 < n is satisfied, execute the array calculation. After each calculation is executed, k1 is incremented by 1. At this time, the entire loop control condition is changed to single - layer control;
[0028] Perform data access relationship mapping according to the formula (3), adjust the data access form, and change it to sequential access:
[0029] (3);
[0030] Among them, let j be the execution order of the hierarchical loop, and j1 be the execution order of the first-level loop. At this time, the data executed for the j-th time in the hierarchical loop is adjusted to be executed for the j1-th time after the first-level loop;
[0031] Secondly, for the case where different positions in the same array in the deep calculation are correlated with each other, formula (4) is used to judge the execution situation of a certain position in the array:
[0032] (4);
[0033] Among them, k represents the number of executions, and k2 represents the difference between two associated data; before performing the calculation, it is judged first. When k = 0, it means that the data at the current position has not been executed, so the calculation is performed, thereby ensuring that the data at a position is only accessed once;
[0034] Finally, through formula (5), ensure the correct call of the outer dependent array; let i be the subscript of the dependent array:
[0035] (5).
[0036] According to the preference of the present invention, an heterogeneous platform parallel optimization method for a hierarchical step transformation loop with dependencies is realized through an automated interface SLHO (Solving Loop Hierarchical Optimization); including:
[0037] First, find out the dependent array, the step size of the step, and the loop boundary, set them as interface parameters, extract the outer dependent data part and the required loop part, and adjust the execution order according to the step size of the step;
[0038] Secondly, by setting the deep calculation array, the dependent array, the step size of the step, and the execution boundary as interface parameters, adjust the hierarchical step loop that has processed the outer dependent data to a first-level loop.
[0039] A computer device includes a memory and a processor. When the processor executes the computer program, the steps of the heterogeneous platform parallel optimization method for a hierarchical step transformation loop with dependencies are realized.
[0040] A computer-readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the steps of the heterogeneous platform parallel optimization method for a hierarchical step transformation loop with dependencies are realized.
[0041] The beneficial effects of the present invention are:
[0042] 1. The present invention takes into account the existence of a hierarchical step-by-step transformation loop that depends on data. Due to the existence of outer-layer dependent data, it may not be possible to put the entire hierarchical loop into the GPU for parallel optimization. To solve this problem, the outer-layer dependent data is extracted from the entire loop body and pre-computed on the CPU. This avoids multiple starts and stops of the kernel function, resulting in wasted time.
[0043] 2. The present invention takes into account the situation where the step size and data length of each layer of the hierarchical step-by-step transformation loop are uncertain after splitting the dependent data. During the GPU parallel process, it is not easy to allocate the number of threads, which may result in a large number of idle threads. To solve this problem, a method and steps for changing the hierarchical loop into a single-layer loop are proposed. At the same time, the data access method is modified, reducing the frequent transmission of uncertain data and improving the execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIG. 1 is a schematic diagram of the CUDA programming model;
[0045] Figure 2 FIG. is a schematic diagram of the mapping relationship for continuous access of the stride iteration dependent array;
[0046] Figure 3 FIG. is a schematic diagram of the data access relationship mapping;
[0047] Figure 4 FIG. is a schematic diagram of changing a loop with three layers of indefinite iteration times and stride lengths into a single-layer loop. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but is not limited thereto.
[0049] Embodiment 1
[0050] A heterogeneous platform parallel optimization method for a hierarchical step-by-step transformation loop with dependencies includes:
[0051] 1) CPU preprocesses the outer-layer data dependencies;
[0052] During the task parallelization process, task decomposition is a key step. And during the task decomposition process, the most important goal is to select as many dimensions as possible for parallelization. In the hierarchical step-by-step transformation loop, the data interaction between different layers and the front-back dependency relationship between the outer-layer data make it impossible to directly map to the GPU for parallel optimization. To solve this problem, the present invention selects to extract the outer-layer dependent data from the entire hierarchical step-by-step transformation loop and perform the calculation independently on the CPU side; thus, more loop dimensions can be mapped to the GPU for parallel execution, reducing the start-up and stop times of the kernel function and the frequent overhead of data transmission.
[0053] 2) Adjust the hierarchical loop to a single-layer loop;
[0054] After solving the outer loop dependence by the above method, although the entire hierarchical loop can be mapped to the GPU to reduce the start and stop of the kernel function. However, due to the uncertainty and data correlation of the hierarchical step transformation loop, during the GPU parallel process, there are not only a large number of transmissions of uncertain data, but also it becomes difficult to allocate the number of threads for each dimension.
[0055] To solve this problem, the present invention changes the hierarchical step transformation loop after solving the dependent data into a single loop, and at the same time proposes a formula to avoid multiple accesses to related data;
[0056] 3) Perform remapping of the post-extraction dependent array.
[0057] Embodiment 2
[0058] The heterogeneous platform parallel optimization method for the hierarchical step transformation loop with dependence according to Embodiment 1 is characterized in that:
[0059] Extract the data calculated in the outer layer from the entire hierarchical step transformation loop and independently execute the calculation on the CPU side; including:
[0060] The outer layer dependent data is stored in an array after extraction, and when the kernel function is executed, the outer layer dependent data is transferred from the CPU to the GPU.
[0061] When extracting the outer layer dependent data, if a certain original for loop is executed in sequence, the calculation can be directly performed according to this logic; however, when it comes to a loop with strided iteration, directly extracting and calculating will result in non-continuous memory access, thereby increasing the storage requirement, and when the number of strides is large, it may cause the shared memory capacity on the GPU side to exceed the limit.
[0062] To solve this problem of non-continuous memory access and storage space, the present invention optimizes the original algorithm. Specifically, when extracting the outer layer dependent data, by adjusting the calculation method, the data storage method is adjusted to a continuous access mode; as Figure 2 shown, specifically including:
[0063] Let p represent the array subscript of strided iteration, ip represent the stride length of the original calculation loop, the number of loop executions is x, i represent the value of the continuously accessed array subscript, and there is p = (x - 1) * ip + 1;
[0064] In the original calculation logic, the difference between the subscripts of two adjacent data in the array is ip, and the value of ip is uncertain. After adjusting the loop stride to 1, the array is calculated and stored in a continuous manner, and the data originally stored at the p position is stored at the x position, and at this time, the values of x and i are equal;
[0065] After extracting the outer-layer dependent data, the array still executes according to the original stride loop logic during the call. To ensure the correctness of the results and effectively map the logical relationships, a method is proposed to restore the dependency relationships between data through mapping formula (1): ensuring that each calculation step is executed in the original logical order.
[0066] (1).
[0067] Change the loop with N layers of indefinite iteration times and stride lengths to a single-layer loop; N≥3; including:
[0068] Taking the N-layer loop as an example, an algorithm optimization method is proposed to change the loop with N layers of indefinite iteration times and stride lengths to a single-layer loop. At the same time, a formula is proposed to avoid multiple accesses to related data, and finally, a remapping of the extracted data is performed. This method aims to reduce the uncertainty during the program execution and improve the execution efficiency of the program. Figure 4 It is a schematic diagram for changing the loop with three layers of indefinite iteration times and stride lengths to a single-layer loop.
[0069] First, by analyzing the loop, clarify the iteration times of each layer in the hierarchical loop, and determine the boundary n of the single-layer loop through formula (2):
[0070] (2);
[0071] Let the loop variable of the outermost loop be i1, the starting value be 1, the stride length be ip1, and the boundary value be n1; the loop variable of the second layer loop be i2, the starting value be i1, the stride length be ip2, and the boundary value be n2; the loop variable of the innermost loop be i3, the starting value be i2, the stride length be ip3, and the boundary value be n3;
[0072] After determining the boundary n of the single-layer loop, set the single-layer loop index as k1. When k1 < n is satisfied, perform array calculations. After each calculation, k1 is incremented by 1. At this time, the entire loop control condition is changed to single-layer control;
[0073] Because the array calculations with the original hierarchical control are accessed in a non-continuous manner, therefore, according to formula (3), perform data access relationship mapping, adjust the data access form, and change it to sequential access:
[0074] (3);
[0075] Among them, let j be the execution order of the hierarchical loop, and j1 be the execution order of the single-layer loop. At this time, the data executed for the jth time in the hierarchical loop is adjusted to be executed for the j1th time after changing to the single-layer loop; the data access relationship mapping is as Figure 3 shown.
[0076] Secondly, for the case where different positions in the same array are correlated in deep computing, specifically referring to the situation where data at different positions in the same array in deep computing needs to interact, such as swapping orders, performing numerical calculations with data at two positions being correlated with each other, etc. After converting to a single-layer loop, it is easy to have the problem of accessing the same position in the array multiple times, resulting in calculation errors. Therefore, formula (4) is proposed to determine the execution situation of a certain position in the array through formula (4):
[0077] (4);
[0078] Among them, k represents the number of executions, and k2 represents the difference between two correlated data. Before performing the calculation, a judgment is made first. When k = 0, it means that the data at the current position has not been executed, so the calculation is performed, thereby ensuring that the data at one position is only accessed once;
[0079] Finally, since the extracted data is indexed according to the subscripts of the original hierarchical loop calculation, after converting to a single-layer loop, the calling situation needs to be readjusted. Therefore, through formula (5), the correct calling of the outer-layer dependent array is ensured. Let i be the subscript of the dependent array:
[0080] (5).
[0081] The advantage of this method is that by changing the hierarchical step-by-step transformation loop to a single-layer loop, the original loop with variable step sizes and data lengths for each layer controlled by N layers, and with a very large step size for the cross-step, can be changed to a single-layer loop with a fixed small step size. Through this transformation, it is convenient to allocate the number of threads during parallelization, avoiding the problem of a large number of idle threads, and at the same time avoiding the problem of a large amount of data needing to be frequently transmitted, improving the memory access efficiency of the data and greatly saving the running time.
[0082] Implement a heterogeneous platform parallel optimization method for hierarchical step-by-step transformation loops with dependencies through the automated interface SLHO (Solving Loop Hierarchical Optimization); including:
[0083] First, find out the dependent array, cross-step size, and loop boundary and set them as interface parameters, extract the outer-layer dependent data part and its required loop part, and adjust the execution order according to the cross-step size;
[0084] Secondly, by setting the deep computing array, dependent array, cross-step size, and execution boundary as interface parameters, adjust the hierarchical step-by-step loop after processing the outer-layer dependent data to a single-layer loop.
[0085] The implementation of the automated interface SLHO reduces the programming difficulty, and programmers can directly call the automated interface SLHO to solve problems.
[0086] The description of the automated interface SLHO is shown in Table 1 as follows:
[0087] Table 1 SLHO Interface Description
[0088]
[0089] An example of the main core calling the SLHO interface is as follows:
[0090] / / Call the SLHO interface on the host side (CPU)
[0091] program main
[0092] use SLHO
[0093] … / / Other operations are omitted
[0094] / / Initialize the interface
[0095] call SLHO_Init()
[0096] / / Call the interface
[0097] call SLHO_Dependent_Data(wr,n1,ip1)
[0098] call SLHO_Transformation_Loop(value,wr,ip1,i)
[0099] / / Close the interface
[0100] call SLHO_Finalize()
[0101] … / / Other operations are omitted
[0102] end program main
[0103] Through the implementation of the above method, the present invention tests the inventive method in a CPU-GPU heterogeneous system.
[0104] Taking the algorithm for the evolution of the dynamic characteristics of polymer nanomaterials as an example, this test conducts time tests when the number of iterations is 10 times, 100 times, and 1000 times respectively with the same input data.
[0105] Compare the time and speedup ratios of the CPU serial program, the general GPU optimization program, and the inventive optimization program. The experimental results prove that, compared with the CPU serial program and the general GPU optimization program, the method of the present invention has an obvious acceleration effect. The specific results and comparisons are shown in the following table. The running times and speedup ratios of the test program before and after optimization when the number of iterations is 10 are shown in Table 2;
[0106] Table 2 Comparison of speedup ratios when the number of iterations is 10 (unit: s);
[0107]
[0108] The running times and speedup ratios of the test program before and after optimization when the number of iterations is 100 are shown in Table 3;
[0109] Table 3 Comparison of speedup ratios when the number of iterations is 100 (unit: s);
[0110]
[0111] The running times and speedup ratios of the test program before and after optimization when the number of iterations is 1000 are shown in Table 4;
[0112] Table 4 Comparison of speedup ratios when the number of iterations is 1000 (unit: s);
[0113] 。
[0114] Example 3
[0115] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the heterogeneous platform parallel optimization method for the hierarchical step transformation loop with dependencies described in Example 1 or 2 are implemented.
[0116] Example 4
[0117] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the heterogeneous platform parallel optimization method for the hierarchical step transformation loop with dependencies described in Example 1 or 2 are implemented.
Claims
1. A method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies, characterized in that: Including: 1) The CPU preprocesses the outer-layer data dependencies; In the hierarchical step transformation loop, the outer-layer dependent data is extracted from the entire hierarchical step transformation loop and independently executed on the CPU side for calculation; 2) Adjust the hierarchical loop to a single-layer loop; 3) Remap the post-extraction dependency array; Extract the data of the outer-layer calculation from the entire hierarchical step transformation loop and independently execute the calculation on the CPU side; Including: The outer-layer dependent data is stored in an array after extraction, and when the kernel function is executed, the outer-layer dependent data is transferred from the CPU to the GPU; Change the loop with N layers of indefinite iteration times and cross-step lengths to a single-layer loop; N≥3; including: First, by analyzing the loop, clarify the iteration times of each layer in the hierarchical loop, and determine the boundary n of the single-layer loop through formula (2): Let the loop variable of the outermost loop be i1, the starting value be 1, the cross-step length be ip1, and the boundary value be n1; the loop variable of the second layer loop be i2, the starting value be i1, the cross-step length be ip2, and the boundary value be n2; the loop variable of the innermost loop be i3, the starting value be i2, the cross-step length be ip3, and the boundary value be n3; After determining the boundary n of the single-layer loop, set the single-layer loop index to k1. When k1 < n is satisfied, execute the array calculation. After each calculation is executed, k1 is incremented, and at this time, the entire loop control condition is changed to single-layer control; Perform data access relationship mapping according to formula (3), adjust the data access form to sequential access: Among them, let j be the execution order of the hierarchical loop, and j1 be the execution order of the single-layer loop. At this time, the data executed for the jth time in the hierarchical loop is adjusted to be executed for the j1th time after the single-layer loop; Secondly, for the situation where different positions in the same array in the deep calculation are associated with each other, judge the execution situation of a certain position in the array through formula (4): Among them, k represents the number of executions, and k2 represents the difference between two associated data; judge before executing the calculation. When k = 0, it means that the data at the current position has not been executed, so execute the calculation, thereby ensuring that the data at a position is only accessed once; Finally, through formula (5), ensure the correct call of the outer-layer dependency array; let i be the subscript of the dependency array:
2. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to claim 1, characterized in that: When extracting the outer-layer dependent data, by adjusting the calculation method, adjust the data storage method to a continuous access mode; specifically including: Let p represent the subscript of the array for stride iteration, ip represent the cross-step length of the original calculation loop, the number of loop executions be x, i be the value of the subscript for continuous access to the array, and there is p = (x - 1)*ip + 1; After adjusting the loop step length to 1, the array is calculated and stored continuously. The data originally stored at the p position is stored at the x position, and at this time, x is equal to the i value.
3. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to claim 1, characterized in that: The array still executes according to the original stride loop logic during the call; that is: restore the dependency relationship between data through the mapping formula (1):
4. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to any one of claims 1 to 3, characterized in that: Implement the heterogeneous platform parallel optimization method for the hierarchical step transformation loop with dependencies through the automated interface SLHO; including: Find out the dependent array, stride length, and loop boundary and set them as interface parameters, extract the outer dependent data part and its required loop part, and adjust the execution order according to the stride length; By setting the deep calculation array, dependency array, stride length, and execution boundary as interface parameters, the hierarchical step loop that processes the outer layer dependency data is adjusted to a single layer loop.
Citation Information
Patent Citations
Deep learning compiler optimization method special for CNN accelerator
CN114995822A
Compilation optimization method, device and equipment based on hybrid precision tensor operation instruction
CN117270870A