Heterogeneous platform parallel optimization method for layering stepping transformation circulation with dependence

By extracting and preprocessing outer-layer dependent data on the CPU side and adjusting the layered step transformation cycle into one-layer loop, the problem of inefficient loop optimization in heterogeneous platforms is solved, and more efficient parallel computing is achieved.

CN119938281AActive Publication Date: 2025-05-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510429446.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the CPU-GPU heterogeneous platform, hierarchical step transformation loops with dependencies are difficult to optimize efficiently in parallel, resulting in frequent calls to kernel functions, idle thread count and inefficient execution.

Method used

By extracting the outer dependent data preprocessing on the CPU side and cacheing it into an array, the loop structure is adjusted into a one-layer loop, reducing unnecessary memory access and improving spatial locality.

Benefits of technology

It reduces the start and stop frequency of kernel functions, optimizes thread allocation, avoids thread idleness, and significantly improves program execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938281A_ABST
    Figure CN119938281A_ABST
Patent Text Reader

Abstract

The invention relates to a heterogeneous platform parallel optimization method for layering stepping transformation circulation with dependence, and belongs to the technical field of electronic information. Comprising the following steps: 1) CPU preprocessing outer layer data dependence; in the layered step-by-step transformation circulation, outer layer dependency data is extracted from the whole layered step-by-step transformation circulation, and calculation is independently executed at a CPU end; (2) layered circulation is adjusted into one-layer circulation; and 3) remapping the extracted dependent array. According to the method, outer layer dependency data is extracted from the whole loop body and is calculated in advance in a CPU (Central Processing Unit). Therefore, time waste caused by multiple times of start and stop of the kernel function is avoided. The invention provides a method and steps for changing layered stepping circulation into one-layer circulation, meanwhile, a data access mode is modified, frequent transmission of uncertain data is reduced, and the execution efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a heterogeneous platform parallel optimization method for a layered step transformation loop containing dependencies, and belongs to the field of electronic information technology. Background Art

[0002] As computing demands continue to increase, traditional single processor architectures are gradually unable to meet the requirements of efficiently processing large-scale computing tasks. In recent years, heterogeneous computing platforms (CPU combined with GPU) have gradually become an important way to solve this problem. The central processing unit (CPU) is the computing and control core of the computer system. The graphics processing unit (GPU) was originally designed to accelerate 3D image rendering and graphics drawing such as video, but because of its powerful parallel capabilities, it has been widely used in complex computing problems. Compared with traditional CPUs, GPUs have large-scale parallel processing capabilities and can perform multiple computing tasks simultaneously on a large scale.

[0003] CUDA (Compute Unified Device Architecture) is a parallel computing platform and programming model developed by NVIDIA that allows developers to use languages ​​such as C, C++, and Fortran to perform high-performance computing on GPUs. The CUDA programming model is divided into the host side and the device side. The device side code running on the GPU is called a kernel. Before the kernel function runs, it first submits the data stored on the host side to the device side's video memory, then calls the GPU's multithreading to calculate the kernel function, and finally transfers the calculated data back to the host side's memory. When the kernel is launched to the GPU for runtime, at the logical level, such as Figure 1 As shown in the CUDA programming model, NVIDIA GPUs use thread blocks and grids to organize and manage threads so that the kernel gets the correct index when running. The total number of threads in the kernel is equal to the size of the grid multiplied by the size of the thread block. However, at the hardware level, the warp scheduler schedules GPU threads in a thread block to execute instructions on a specified device in units of warps. Warps are the basic unit for GPU scheduling and operation, and consist of 32 threads in the same thread block. Even if there are less than 32 threads remaining in the block, they will be scheduled as a warp.

[0004] In CPU-GPU heterogeneous parallel computing, the common practice is to divide the tasks into granularities and assign the computing parts suitable for data parallelism to the GPU, while the CPU is responsible for input and output processing and serial computing tasks. In complex numerical computing applications, many tasks involve hierarchical step transformation loops, and this computing structure usually has characteristics such as strong data dependence, time-varying, and uncertainty. These characteristics make it difficult for traditional parallelization strategies (such as static task decomposition, static scheduling, etc.) to cope with them efficiently.

[0005] In CPU-GPU heterogeneous platform optimization, the core goal is to port large-scale computing tasks that can be parallelized in the program to the GPU and use CUDA for parallel acceleration. Many complex applications contain layered step transformation loops, and the number of iterations and step size of each layer are usually dynamically determined by incoming parameters or external input values. This makes the number of iterations of the loop structure irregular, which may cause the number of iterations of different layers to vary significantly. In addition, there is data association and dependency between the outer and inner calculations of this loop structure.

[0006] For parallel optimization of this loop structure, ordinary optimization schemes cannot select the entire hierarchical loop for parallelization. They can only allocate threads to the loop that controls the inner calculation, and the number of threads contained in each dimension of the thread block is fixed. This ordinary GPU optimization scheme will cause the outer loop to nest kernel functions in the program, resulting in the kernel function being called frequently, thus wasting time. At the same time, the number of threads allocated to each dimension of the thread block cannot be changed in each loop. In the face of a dynamic number of iterations, a large number of threads will be idle. Summary of the invention

[0007] In view of the deficiencies of the prior art, the present invention provides a heterogeneous platform parallel optimization method for hierarchical step transformation loops containing dependencies; The present invention proposes a novel calculation optimization method. By analyzing the data interaction and dependency between different levels, the outer layer numerical calculation is extracted before the loop execution, and the calculation result is cached in an array. This optimization can ensure that more loop levels can be efficiently mapped to the GPU, thereby improving the efficiency of parallel computing. For multi-layer stride iterative loops, a strategy of simplifying nested loops into single-layer iterations is adopted. Specifically, firstly, the total number of iterations of the hierarchical loop and its logical relationship are determined, and the boundary of the one-dimensional loop is delineated based on this relationship, and the execution order of the original loop array is adjusted to change it into continuous access; secondly, in order to avoid multiple accesses between different positions of the same array in deep-level calculations, a formula is proposed for execution judgment, thereby significantly reducing unnecessary access operations; finally, for the problem of non-continuous array access in the outer layer calculation, an improved calculation and calling strategy is proposed to reduce non-continuous memory access and improve spatial locality. Through the above optimization method, the number of data iterations becomes more certain, so that thread blocks and grids can be allocated more efficiently during the GPU transplantation process. The present invention reduces the start and stop frequency of kernel functions, optimizes thread allocation, avoids thread idleness, and significantly improves the execution efficiency of the program.

[0008] When optimizing a program on a CPU-GPU heterogeneous system, when there is a hierarchical step transformation loop with dependencies, it is difficult to select a suitable dimension for parallelization during the porting process according to the common GPU optimization method, and it is impossible to reasonably allocate the number of threads in each dimension, resulting in idle threads, which wastes a lot of time. For programs with the above problems, the solution of the present invention is selected for optimization, which can effectively solve the problem and improve the execution efficiency of the program. At the same time, the present invention also designs an automatic interface SLHO (Solving Loop Hierarchical Optimization) to facilitate programmers to call directly.

[0009] The technical solution of the present invention is: A heterogeneous platform parallel optimization method for hierarchical step transformation loops with dependencies, including: 1) CPU preprocesses outer layer data dependencies; In the hierarchical step transformation loop, the outer layer dependent data is extracted from the entire hierarchical step transformation loop and the calculation is performed independently on the CPU side; 2) The layered loop is adjusted to a single layer loop; 3) Remapping of the dependent array after extraction.

[0010] Preferably, according to the present invention, the data of the outer layer calculation is extracted from the entire hierarchical step transformation cycle, and the calculation is performed independently on the CPU side; including: The outer-layer dependent data is stored in an array after extraction, and when the kernel function is executed, the outer-layer dependent data is transferred from the CPU to the GPU.

[0011] Further preferably, when extracting the outer-layer dependent data, by adjusting the calculation method, the data storage method is adjusted to a sequential access mode; specifically including: Let p represent the array subscript of the strided iteration, ip represent the stride of the original calculation loop, the number of loop executions be x, i represent the value of the sequentially accessed array subscript, and there is p = (x - 1) * ip + 1; After adjusting the loop stride to 1, the array is calculated and stored sequentially. The data originally stored at position p is stored at position x, and at this time, the values of x and i are equal; After extracting the outer-layer dependent data, the array still executes according to the original strided loop logic during the call, and the dependency relationship between the data is restored through the mapping formula (1): (1).

[0012] According to the preference of the present invention, the loop with N layers of indefinite iteration times and stride is changed to a single-layer loop; N ≥ 3; including: First, by analyzing the loop, clarify the iteration times of each layer in the hierarchical loop, and determine the boundary n of the single-layer loop through the formula (2): (2); Let the loop variable of the outermost loop be i1, the starting value be 1, the stride be ip1, and the boundary value be n1; the loop variable of the second layer loop be i2, the starting value be i1, the stride be ip2, and the boundary value be n2; the loop variable of the innermost loop be i3, the starting value be i2, the stride be ip3, and the boundary value be n3; After determining the boundary n of the single-layer loop, set the single-layer loop index as k1. When k1 < n is satisfied, execute the array calculation. After each calculation is executed, k1 is incremented by 1. At this time, the entire loop control condition is changed to single-layer control; Perform data access relationship mapping according to the formula (3), adjust the data access form, and change it to sequential access: (3); Among them, let j be the execution order of the hierarchical loop, and j1 be the execution order of the single-layer loop. At this time, the data executed for the jth time in the hierarchical loop is adjusted to be executed for the j1th time after being changed to a single-layer loop; Secondly, for the case where different positions in the same array in the deep calculation are associated with each other, judge the execution situation of a certain position in the array through the formula (4): (4); Where k represents the number of executions, and k2 represents the difference between two associated data. Before executing the calculation, a judgment is made. When k=0, it means that the data at the current position has not been executed, so the calculation is performed to ensure that the data at a position is only accessed once. Finally, use formula (5) to ensure that the outer dependency array is called correctly; let i be the dependency array index: (5).

[0013] Preferably, according to the present invention, a heterogeneous platform parallel optimization method for hierarchical step transformation loops containing dependencies is implemented through an automation interface SLHO (Solving Loop Hierarchical Optimization); comprising: First, find out the dependent array, stride length, and loop boundary, set them as interface parameters, extract the outer layer dependent data part and its required loop part, and adjust the execution order according to the stride length; Secondly, by setting the deep calculation array, dependency array, stride length, and execution boundary as interface parameters, the hierarchical step loop that processes the outer layer dependency data is adjusted to a single layer loop.

[0014] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a heterogeneous platform parallel optimization method for a hierarchical step transformation loop containing dependencies are implemented.

[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a heterogeneous platform parallel optimization method for a hierarchical step transformation loop containing dependencies.

[0016] The beneficial effects of the present invention are: 1. The present invention takes into account the existence of a data-dependent hierarchical step transformation loop. Due to the existence of outer layer dependent data, it may not be possible to put the entire hierarchical loop into the GPU for parallel optimization. To solve this problem, the outer layer dependent data is extracted from the entire loop body and calculated in advance in the CPU. This avoids multiple starts and stops of the kernel function, which causes a waste of time.

[0017] 2. The present invention takes into account the situation that the stride length and data length of each layer of the layered step transformation loop after splitting the dependent data are uncertain, and it is not easy to allocate the number of threads in the GPU parallel process, which may cause a large number of threads to be idle. To solve this problem, a method and steps for changing the layered step loop to a layer loop are proposed, and at the same time, the data access method is modified to reduce the frequent transmission of uncertain data and improve the execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a schematic diagram of the CUDA programming model; Figure 2 A diagram showing the mapping relationship for continuous access to stride iteration dependent arrays; Figure 3 It is a diagram of data access relationship mapping; Figure 4 Schematic diagram of changing a three-layer loop with indefinite number of iterations and stride length into a one-layer loop. DETAILED DESCRIPTION

[0019] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.

[0020] Example 1 A heterogeneous platform parallel optimization method for hierarchical step transformation loops with dependencies, including: 1) CPU preprocesses outer layer data dependencies; In the process of task parallelization, task decomposition is a key step, and in the process of task decomposition, the most important goal is to select as many dimensions as possible for parallelization. In the hierarchical step transformation loop, the interaction of data at different levels and the front-to-back dependencies between outer data make it impossible to directly map them to the GPU for parallel optimization. To solve this problem, the present invention chooses to extract the outer layer dependent data from the entire hierarchical step transformation loop and perform calculations independently on the CPU side; thereby, more loop dimensions can be mapped to the GPU for parallel execution, reducing the number of kernel function starts and stops and the frequent overhead of data transmission.

[0021] 2) The layered loop is adjusted to a single layer loop; After solving the outer loop dependency through the above method, the entire layered loop can be mapped to the GPU to reduce the start and stop of the kernel function. However, due to the uncertainty and data correlation of the layered step transformation loop, in the GPU parallel process, not only a large amount of uncertain data is transmitted, but it also becomes difficult to assign the number of threads to each dimension.

[0022] To solve this problem, the present invention changes the hierarchical step transformation loop after solving the dependent data into a layer loop, and proposes a formula to avoid multiple accesses to the associated data; 3) Perform remapping of the dependency array after extraction.

[0023] Example 2 The difference between the heterogeneous platform parallel optimization method for the hierarchical step transformation loop containing dependencies described in Example 1 is that: Extract the data of the outer layer calculation from the entire hierarchical step transformation loop and perform the calculation independently on the CPU side; including: The outer dependency data is stored in an array after extraction, and is passed from the CPU to the GPU when the kernel function is executed.

[0024] When extracting outer-layer dependent data, if the original for loop is executed in sequence, the calculation can be performed directly according to this logic; however, when it involves a stride iteration loop, direct extraction calculation will cause non-continuous memory access, thereby increasing storage requirements, and when the number of strides is large, it may cause the shared memory capacity on the GPU to exceed the limit.

[0025] To solve this problem of discontinuous memory access and storage space, the present invention optimizes the original algorithm. Specifically, when extracting outer layer dependent data, the data storage mode is adjusted to a continuous access mode by adjusting the calculation method; Figure 2 As shown, specifically including: Let p represent the array index of the stride iteration, ip represent the stride length of the original calculation loop, the number of loop executions is x, i is the value of the consecutive access array index, and there exists p=(x-1)*ip+1; In the original calculation logic, the subscript difference between two adjacent data in the array is ip, and the value of ip is uncertain. After the loop step is adjusted to 1, the array is calculated and stored in a continuous manner, and the data originally stored in position p is stored in position x. At this time, the values ​​of x and i are equal; After extracting the outer layer dependency data, the array is still executed according to the original stride loop logic during the call process. In order to ensure the correctness of the result and effectively map the logical relationship, it is proposed to restore the dependency relationship between the data through the mapping formula (1): to ensure that each calculation step is executed according to the original logical order.

[0026] (1).

[0027] Change N layers of loops with indefinite number of iterations and stride lengths into one layer of loops; N ≥ 3; including: Taking N-layer loop as an example, an algorithm optimization method is proposed to change the N-layer loop with indefinite number of iterations and stride length into a single-layer loop. At the same time, a formula is proposed to avoid multiple accesses to associated data, and finally the extracted data is remapped. This method aims to reduce the uncertainty in the program running process and improve the execution efficiency of the program. Figure 4 Schematic diagram of changing a three-layer loop with indefinite number of iterations and stride length into a one-layer loop.

[0028] First, by analyzing the loop, we can sort out the number of iterations of each layer in the hierarchical loop and determine the boundary n of a layer of loop using formula (2): (2); Let the loop variable of the outermost loop be i1, with the starting value of 1, the step size of ip1 for each step, and the boundary value of n1; the loop variable of the second layer loop be i2, with the starting value of i1, the step size of ip2 for each step, and the boundary value of n2; the loop variable of the innermost loop be i3, with the starting value of i2, the step size of ip3 for each step, and the boundary value of n3; After determining the loop boundary n of one layer, set the loop index of one layer as k1. When k1 < n is satisfied, perform array calculations. After each calculation, k1 is incremented by 1. At this time, the entire loop control condition is changed to one-layer control; Because the array calculations with original hierarchical control were accessed in a non - continuous manner, so, according to formula (3), perform data access relationship mapping, adjust the data access form to sequential access: (3); Among them, let j be the execution order of the hierarchical loop, and j1 be the execution order of one - layer loop. At this time, the data executed for the j - th time in the hierarchical loop is adjusted to be executed for the j1 - th time after being changed to one - layer loop; the data access relationship mapping is as Figure 3 shown.

[0029] Secondly, for the situation where different positions in the same array in deep - layer calculations are correlated with each other, specifically referring to: for the situation where data at different positions in the same array in deep - layer calculations need to interact, such as swapping orders, performing numerical calculations with data at two positions being correlated with each other, etc. After being changed to one - layer loop, it is easy to have the problem of accessing the same position of the array multiple times, resulting in calculation errors. Therefore, formula (4) is proposed to judge the execution situation of a certain position in the array through formula (4): (4); Among them, k represents the number of executions, and k2 represents the difference between two related data; before performing the calculation, make a judgment first. When k = 0, it means that the data at the current position has not been executed, so the calculation is performed, thereby ensuring that the data at one position is only accessed once; Finally, since the extracted data is indexed according to the subscripts calculated by the original hierarchical loop, after being changed to one - layer loop, the calling situation needs to be readjusted. Therefore, through formula (5), ensure the correct calling of the outer - layer dependent array; let i be the subscript of the dependent array: (5).

[0030] The advantage of this method is that by changing the layered step transformation loop to a single-layer loop, the original N-layer control loop with variable stride and data length and large stride length can be changed into a single-layer loop with fixed length and small stride. Through this transformation, it is easier to allocate the number of threads in the parallel process, avoiding the problem of a large number of threads being idle, and also avoiding the problem of a large amount of data needing to be frequently transmitted, thus improving the data access efficiency and greatly saving the running time.

[0031] The automation interface SLHO (Solving Loop Hierarchical Optimization) is used to implement a parallel optimization method for heterogeneous platforms with hierarchical step transformation loops containing dependencies; including: First, find out the dependent array, stride length, and loop boundary and set them as interface parameters, extract the outer dependent data part and its required loop part, and adjust the execution order according to the stride length; Secondly, by setting the deep calculation array, dependency array, stride length, and execution boundary as interface parameters, the hierarchical step loop that processes the outer layer dependency data is adjusted to a single layer loop.

[0032] The implementation of the automation interface SLHO reduces the programming difficulty, and programmers can directly call the automation interface SLHO to solve problems.

[0033] The description of the automation interface SLHO is shown in Table 1: Table 1 SLHO interface description

[0034] The main core calls the SLHO interface example as follows: / / Call the SLHO interface on the host side (CPU) Program Main use SLHO ... / / Omit other operations / / Initialize the interface call SLHO_Init() / / Call interface call SLHO_Dependent_Data(wr,n1,ip1) call SLHO_Transformation_Loop(value,wr,ip1,i) / / Close the interface call SLHO_Finalize() ... / / Omit other operations end program main By implementing the above method, the present invention has tested the inventive method in a CPU-GPU heterogeneous system.

[0035] This test takes the dynamic characteristic evolution algorithm of polymer nanomaterials as an example. When the same data is input, time tests are performed with 10, 100, and 1000 iterations respectively.

[0036] The time and acceleration ratio of the CPU serial program, the GPU ordinary optimization program and the invention optimization program are compared. The experimental results show that compared with the CPU serial program and the ordinary GPU optimization program, the method described in the invention has a significant acceleration effect. The specific results and comparisons are shown in the following table. The running time and acceleration ratio of the test program before and after optimization when the number of iterations is 10 are shown in Table 2; Table 2 Speedup comparison when the number of iterations is 10 (unit: s);

[0037] The running time and speedup of the test program before and after optimization when the number of iterations is 100 are shown in Table 3; Table 3 Speedup comparison when the number of iterations is 100 (unit: s);

[0038] The running time and speedup ratio of the test program before and after optimization when the number of iterations is 1000 are shown in Table 4; Table 4 Speedup comparison when the number of iterations is 1000 (unit: s); .

[0039] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the heterogeneous platform parallel optimization method for a hierarchical step transformation loop containing dependencies described in embodiment 1 or 2 are implemented.

[0040] Example 4 A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the heterogeneous platform parallel optimization method for a hierarchical step transformation loop containing dependencies described in embodiment 1 or 2.

Claims

1. A method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies, characterized in that: Including: 1) The CPU preprocesses the outer-layer data dependencies; In the hierarchical step transformation loop, the outer-layer dependent data is extracted from the entire hierarchical step transformation loop and independently executed for calculation on the CPU side; 2) The hierarchical loop is adjusted to a single-layer loop; 3) Remapping of the extracted post-dependency array is performed.

2. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to claim 1, characterized in that: The data for outer-layer calculations is extracted from the entire hierarchical step transformation loop and independently executed for calculation on the CPU side; Including: The outer-layer dependent data is stored in an array after extraction, and when the kernel function is executed, the outer-layer dependent data is transferred from the CPU to the GPU.

3. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to claim 1, characterized in that: When extracting the outer-layer dependent data, by adjusting the calculation method, the data storage method is adjusted to a continuous access mode; specifically including: Let p represent the array subscript for strided iteration, ip represent the stride of the original calculation loop, the number of loop executions be x, i represent the value of the continuously accessed array subscript, and there exists p = (x - 1) * ip + 1; After adjusting the loop stride to 1, the array is calculated and stored in a continuous manner. The data originally stored at position p is stored at position x, and at this time, x is equal to the value of i.

4. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to claim 1, characterized in that: The array still executes according to the original strided loop logic during the call; that is: the dependency relationship between data is restored through mapping formula (1): (1)。 5. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to claim 1, characterized in that: Changing the loop with N layers of indefinite iteration times and stride to a single-layer loop; N ≥ 3; including: First, by analyzing the loop, clarify the number of iterations of each layer in the hierarchical loop, and determine the boundary n of the single-layer loop through formula (2): (2); Let the loop variable of the outermost loop be i1, the starting value be 1, the stride be ip1, and the boundary value be n1; the loop variable of the second layer loop be i2, the starting value be i1, the stride be ip2, and the boundary value be n2; the loop variable of the innermost loop be i3, the starting value be i2, the stride be ip3, and the boundary value be n3; After determining the single-layer loop boundary n, set the single-layer loop index as k1. When k1 < n is satisfied, perform array calculations. After each calculation is executed, k1 is incremented, and at this time, the entire loop control condition is changed to single-layer control; Perform data access relationship mapping according to formula (3), adjust the data access form to sequential access: (3); Among them, let j be the execution order of the hierarchical loop, and j1 be the execution order of the single-layer loop. At this time, the data executed for the j-th time in the hierarchical loop is adjusted to be executed for the j1-th time after being changed to a single-layer loop; Secondly, for the situation where different positions in the same array in deep calculations are associated with each other, judge the execution situation of a certain position in the array through formula (4): (4); Among them, k represents the number of executions, and k2 represents the difference between two associated data; judge before executing the calculation. When k = 0, it means that the data at the current position has not been executed, so execute the calculation, thereby ensuring that the data at one position is only accessed once; Finally, through formula (5), ensure the correct call of the outer-layer dependent array; let i be the dependent array subscript: (5)。 6. The method for parallel optimization of heterogeneous platforms for hierarchical step transformation loops with dependencies according to any one of claims 1 to 5, characterized in that: Implementing a heterogeneous platform parallel optimization method for a hierarchical step transformation loop with dependencies through the automated interface SLHO; including: Find out the dependent array, stride length, and loop boundary and set them as interface parameters, extract the outer dependent data part and its required loop part, and adjust the execution order according to the stride length; By setting the deep calculation array, dependency array, stride length, and execution boundary as interface parameters, the hierarchical step loop that processes the outer layer dependency data is adjusted to a single layer loop.

Citation Information

Patent Citations

  • Deep learning compiler optimization method special for CNN accelerator

    CN114995822A

  • Compilation optimization method, device and equipment based on hybrid precision tensor operation instruction

    CN117270870A

  • Slave core local storage limitation optimization method based on new generation SW many-core processor

    CN118245118A

  • System, methods and apparatus for program optimization for multi-threaded processor architectures

    US20100218196A1

  • Automatic low level operator loop generation, parallelization and vectorization for tensor computations

    US20240028802A1