Data stream architecture-based Sencil calculation compiling optimization method and electronic equipment
By adding guidance statements to C language to describe the data dependence between the timelines of Stencil calculations, and iteratively allocate the timeline to adjacent parallel computing nodes to generate data flow instructions, the problems of insufficient parallelism of timelines and large synchronization overhead in the existing technology are solved, and the performance of Stencil calculations is improved.
Patent Information
- Application Number
- CN202510286202.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology has not fully explored the coarse-grained timeline parallelism in Stencil calculations. The synchronization overhead between multiple cores is large, and the data blocking technology has redundancy in data memory access and computing instructions. Each cycle iteration of the timeline needs to be synchronized once, which leads to huge time overhead.
By adding guidance statements to describe the data dependencies between the timelines, the compiler parses these relationships and iteratively allocates the timeline to adjacent parallel computing nodes, generates data flow instructions, realizes asynchronous communication and parallel execution between the timelines, and reduces synchronization overhead.
It realizes parallel execution between the timelines, reduces memory access, improves program performance, reduces redundant computing and memory access overhead, and improves computing efficiency.
Smart Images

Figure CN120295630A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a Stencil computing compilation optimization method and an electronic device based on a data flow architecture. Background Art
[0002] Stencil computing (modular computing) plays a very important role in the fields of scientific computing, engineering simulation, image processing, astronomy and meteorology, fluid mechanics, heat conduction and diffusion, etc. The essence of Stencil computing is to perform local operations on each element in a grid, and these operations depend on the values of the element and its neighboring nodes. Stencil computing includes not only one-dimensional and two-dimensional, but also can be extended to three-dimensional and even higher-dimensional computing; Stencil computing is very suitable for parallelism, and currently many research works focus on the parallel optimization of Stencil computing.
[0003] The prior art uses the method of loop tiling to accelerate the parallel computing of Stencil. In order to reduce the synchronization between tiles, the methods of data redundant loading and data redundant computing are adopted, reducing the synchronization overhead and improving the execution performance of Stencil computing on the GPU architecture; although this method improves data locality and parallelism, this method will introduce redundant memory access and computing instructions. For certain scales of 3D Stencil computing, the proportion of such redundant memory access and computing instructions will be larger, and there is a possibility of resource overrun for hardware with limited resources.
[0004] The prior art also adopts a multi-level parallel execution strategy, a data tiling strategy, buffer optimization for data transmission and computing, and manual optimization of the Stencil kernel to improve the performance of Stencil computing, achieving performance improvement on multi-core DSPs (Digital Signal Processors); however, the multi-level parallel optimization, data tiling optimization, data transmission, and computing buffer optimization adopted are all strategies for parallel optimization of the inner loop of Stencil computing, not optimized for time-axis parallelism, and data tiling optimization will introduce redundant memory access.
[0005] The prior art also provides a parallel code automatic generation system for heterogeneous architectures based on a polyhedral compilation model, which can perform fine-grained parallel optimization on Stencil operations and improve the performance of Stencil computing. However, this method does not tile the Stencil time axis, and this parallel execution method results in the time axis being executed only on the main core, and each loop iteration of the time axis will start and synchronize a slave core thread group once, and this process means huge time overhead.
[0006] In summary, it is obvious that there are inconveniences and defects in the existing technology in actual use, so it is necessary to improve it. Summary of the Invention
[0007] In view of the above defects, the purpose of the present invention is to provide a Stencil computing compilation optimization method and an electronic device based on a data flow architecture, which are used to solve the problems of insufficient exploration of coarse-grained timeline parallelism, large synchronization overhead between multiple cores, and redundancy of data access and computing instructions in the data chunking technology.
[0008] In order to solve the above technical problems, on the one hand, the present invention provides a Stencil computing compilation optimization method based on a data flow architecture, including the following steps:
[0009] Write a Stencil program by adding directive statements in C language; wherein, the directive statements are used to indicate the type of Stencil computing and the data dependency relationship between timelines.
[0010] Parse the directive statements through a compiler, analyze the data dependency relationship between the timelines, and determine the startup interval of each timeline relative to the adjacent timeline and the data required to flow.
[0011] According to the relationship between the timeline iteration times and the number of parallel data flow computing nodes, allocate the timeline operations to adjacent parallel computing nodes, so that consecutive timeline iterations are allocated to adjacent data flow computing nodes.
[0012] Generate data flow instructions for the operations within each timeline, and control the asynchronous communication and parallel execution between nodes through a data flow graph; wherein, the data flow instructions include information on the sending node and the receiving node for data transmission between adjacent nodes.
[0013] Further, the step of allocating the timeline operations to adjacent parallel computing nodes according to the timeline iteration times, so that consecutive timeline iterations are allocated to adjacent data flow computing nodes includes:
[0014] If the timeline iteration times are less than or equal to the number of parallel data flow computing nodes, map each timeline iteration to adjacent data flow computing nodes in sequence.
[0015] If the timeline iteration times are greater than the number of parallel data flow computing nodes, allocate the timelines to adjacent data flow computing nodes in ascending order in sequence, and then allocate the other timelines that have not been allocated to adjacent data flow computing nodes in ascending order in sequence.
[0016] Further, the step of controlling the asynchronous communication and parallel execution between nodes through a data flow graph includes:
[0017] Data interaction is achieved between adjacent data stream calculation nodes through the data stream instructions, and the execution of the receiving node depends on the data transmitted by the sending node;
[0018] The hardware loop is used to control the multiple executions of the data flow graph to achieve the parallelism between the time axes.
[0019] Further, after the step of generating data flow instructions for the operations within each time axis and controlling the asynchronous communication and parallel execution between nodes through the data flow graph, the method further includes:
[0020] Execute the optimized Stencil program in the data flow architecture, where only the calculation nodes of the first time axis access the main memory, and the data of the remaining nodes is transmitted through adjacent nodes.
[0021] Further, the data flow architecture includes:
[0022] A main control core, configured to transmit data and the data flow graph to the data flow processing unit array through direct memory access;
[0023] The data flow processing unit array is composed of multiple data stream calculation nodes that can be dynamically combined, and each data stream calculation node is assigned an independent data flow graph;
[0024] Further, the directive statements include #pragma stencil_wait and #pragma stencil_send for describing the data dependence relationship between time axes, and directives for declaring the Stencil calculation type.
[0025] In a second aspect, an embodiment of the present invention provides an electronic device, including a storage medium, a processor, and a program stored on the memory and executable on the processor, characterized in that when the processor executes the program, the above-mentioned Stencil calculation compilation optimization method based on the data flow architecture is implemented.
[0026] The method of adding guidance in the present invention through C language describes the data dependency relationship between time axes and the types of Stencil calculations in the program, thereby guiding the compiler to identify Stencil calculations and analyze the data dependency relationship between time axes; this method can describe the data flow between time axes of Stencil calculations, enabling users to add a small amount of guidance on the basis of the C language code of Stencil calculations, and then guiding the compiler to perform parallel optimization between time axes of Stencil calculations; by analyzing the data dependency relationship between time axes, the operations between time axes are allocated to consecutive parallel data stream calculation nodes, and data stream instructions are generated for the operations within each time axis, enabling parallel execution between time axes; this execution method allows multiple time axes to be executed in parallel at the same time, reducing memory access, and within each time axis, the execution time of data stream instructions is masked by calculation instructions, improving the program performance in multiple aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic flowchart of a Stencil calculation compilation optimization method based on a data stream architecture provided in Embodiment 1 of the present invention;
[0028] Figure 2 It is a schematic diagram of the data stream architecture of a Stencil calculation compilation optimization method based on a data stream architecture provided in Embodiment 1 of the present invention;
[0029] Figure 3 It is a schematic diagram of the data dependency relationship between time axes of a Stencil calculation compilation optimization method based on a data stream architecture provided in Embodiment 1 of the present invention;
[0030] Figure 4 It is the data flow diagram of a Stencil calculation compilation optimization method based on a data stream architecture provided in Embodiment 1 of the present invention;
[0031] Figure 5 It is a schematic diagram of the hardware structure of an electronic device provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] It should be noted that the references to "one embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures or characteristics, but not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures or characteristics with an embodiment, it has been shown that combining such features, structures or characteristics with other embodiments is within the knowledge of those skilled in the art, whether or not explicitly described.
[0034] In addition, in the specification and subsequent claims, certain terms are used to refer to specific components or parts. Those of ordinary skill in the art should understand that manufacturers may use different nouns or terms to refer to the same component or part. This specification and subsequent claims do not use the difference in name as a way to distinguish components or parts, but use the difference in function of components or parts as the criterion for distinction. The terms "including" and "comprising" mentioned throughout the specification and subsequent claims are open-ended terms, and should be interpreted as "including but not limited to".
[0035] When programming and compiling Stencil calculations based on a general data flow architecture, it is found that although there are cyclic dependencies between the outermost time axes of Stencil calculations, by distributing the operations of the time axes to multiple different data flow parallel computing nodes, parallelism of the time axes can be achieved. In addition, the data transfer speed between the data flow computing nodes of the data flow architecture is fast, which is very suitable for parallel execution between time axes. Most of the current research focuses on loop optimization for the inner loops within the time axis and does not fully explore the parallelism between time axes; and in terms of programming, it is difficult for Stencil calculations to use existing parallel programming models to describe the parallelism and synchronization between time axes.
[0036] Before describing the embodiments of the present application in detail, first briefly describe the technical concept of the present application: The present application aims to solve the problems of insufficient parallelism between time axes, redundant memory access and large synchronization overhead in the prior art; its core concept is: through programming guidance and compilation optimization strategies, realize pipelined parallel execution between time axes in the data flow architecture. Specifically, it includes: First, in the programming stage, guidance statements are embedded in the C language to describe the data dependencies between time axes and the type of Stencil calculation, providing a basis for parallel optimization for the compiler. Second, in the compilation optimization stage, parse the guidance statements, analyze the dependencies between time axes, and generate data flow information; distribute the time axis iterations to adjacent parallel computing nodes to achieve coarse-grained time axis parallelism; generate data flow instructions for each time axis, construct a data flow graph for asynchronous communication, and execute through hardware loop control pipelining.
[0037] The implementation process of the Stencil computing compilation optimization method based on the data flow architecture of the present application will be described below according to specific embodiments.
[0038] Figure 1 A Stencil computing compilation optimization method based on the data flow architecture provided by an embodiment of the present invention is shown, including the following steps:
[0039] S101: Write a Stencil program by adding directive statements in C language; among them, the directive statements are used to indicate the type of Stencil calculation and the data dependence relationship between time axes; that is, programmatically, by adding directives in C language, the data dependence relationship between time axes and the types of Stencil calculations are described in the program, so as to guide the compiler to identify Stencil calculations and analyze the data dependence relationship between time axes in subsequent steps; this programming method can describe the data flow between time axes of Stencil calculations, so that users only need to add a small amount of directives on the basis of the C language code of Stencil calculations to guide the compiler to perform parallel optimization between time axes of Stencil calculations.
[0040] Specifically, the directive statements include #pragma stencil_wait and #pragma stencil_send for describing the data dependence relationship between time axes, and directives for declaring the type of Stencil calculation; that is, the data flow between time axes is described by adding #pragma stencil_wait() and #pragma stencil_send() directive statements, and the type of Stencil calculation is described by adding directives such as #pragma stencil_1D3P_*.
[0041] S102: Analyze the data dependence relationship between time axes by the compiler parsing the directive statements, and determine the start interval of each time axis relative to the adjacent time axis and the data required to flow. In order to generate the data flow information required for parallel execution of time axes, in this embodiment, the compiler is used to analyze the directive statements in step S101, parse the dependence relationship between time axis iterations, and obtain the data flow information between time axes, that is, the start interval of each time axis relative to the adjacent time axis and the data that needs to flow between time axes.
[0042] S103: According to the relationship between the number of time axis iterations and the number of parallel data flow calculation nodes, allocate time axis operations to adjacent parallel calculation nodes, so that consecutive time axis iterations are allocated to adjacent data flow calculation nodes, thereby achieving coarse-grained parallelism between time axes and reducing synchronization overhead.
[0043] Specifically, step S103 includes:
[0044] If the number of iterations of the time axis is less than or equal to the number of parallel data stream computing nodes, each iteration of the time axis is mapped to adjacent data stream computing nodes in sequence; in this case, the iterations of the time axis are mapped one by one to each data stream computing node, and adjacent iterations of the time axis are assigned to adjacent data stream computing nodes.
[0045] If the number of iterations of the time axis is greater than the number of parallel data stream computing nodes, the time axes are sequentially assigned to adjacent data stream computing nodes in ascending order, and then the other unassigned time axes are sequentially assigned to adjacent data stream computing nodes in ascending order; that is, in this case, the remaining time axis iterations are cyclically assigned to adjacent data stream computing nodes.
[0046] S104: Generate data stream instructions for the operations within each time axis, and control the asynchronous communication and parallel execution between nodes through the data stream graph; among them, the data stream instructions include information on the sending node and the receiving node for data transmission between adjacent nodes. That is, in order to improve the parallel effect, in this embodiment, a parallel optimization method of pipelining between time axes is used to assign time axis iterations to multiple data stream parallel computing nodes for parallel execution, without introducing redundant calculations, reducing memory access, and improving performance.
[0047] When generating the data stream graph, the data interaction between two adjacent data stream graph nodes is represented by nodes send and wait for the sending node and the receiving node respectively, and the positions of send and wait in the data stream instruction are specifically determined by the node assignment result of the above step S103; and the execution of the wait node depends on the data sent by the send node. When the wait node receives the data sent by the send node, the wait node starts to execute.
[0048] Specifically, the control of asynchronous communication and parallel execution between nodes through the data stream graph includes: realizing data interaction between adjacent data stream computing nodes through the data stream instruction, and the execution of the receiving node depends on the data transmitted by the sending node; controlling the multiple executions of the data stream graph through hardware loops to achieve parallelism between time axes.
[0049] In terms of compilation optimization in this embodiment, by analyzing the data dependence relationship between time axes, each time axis operation is compiled and assigned to a separate parallel computing node, and continuous time axis operations are ensured to be assigned to adjacent parallel computing nodes; within the data stream parallel computing node, for the operations within each time axis, data stream instructions are generated to send and receive data from adjacent time axis computing nodes, thereby realizing parallel execution between time axes.
[0050] Further, after step S104, it further includes: executing the optimized Stencil program in the data flow architecture, where only the first time-axis calculation node accesses the main memory, and the data of the remaining nodes is transmitted through adjacent nodes; thus, among all the operation nodes, only the first operation node needs to access the memory for data, and the calculation data of other nodes comes from adjacent nodes without accessing the memory, reducing the memory access overhead.
[0051] Preferably, within each time axis, the data transmission time can be masked by means of computing and data flow coincidence, further improving the performance. Different from the synchronous barrier method, this parallel execution method between time axes only requires the computing nodes to communicate with adjacent computing nodes, and this flexibility reduces the time spent by the computing nodes waiting for other computing nodes, improving the parallelism.
[0052] See Figure 2 , the data flow architecture of this embodiment includes a main control core, a main memory, a direct memory access, an array of data flow processing units, an on-chip memory, etc.; where: the main control core is used to transmit data and the data flow graph to the array of data flow processing units through direct memory access; the array of data flow processing units is composed of multiple data flow calculation nodes that can be dynamically combined, and each data flow calculation node is assigned an independent data flow graph; the on-chip memory is used to cache the transmission data between the data flow calculation nodes.
[0053] Among them, the scale of the array of data flow processing units can be combined and adjusted as needed, and the method provided in this embodiment is not restricted by the number of data flow processing unit arrays. During the calculation process, the main control core passes the data and the data flow graph to the array of data flow processing units through direct memory access, and each data flow operation node will be assigned a corresponding data flow graph; the data flow graph describes the data dependency relationship of the flow graph nodes. When the data of the data flow graph nodes is ready, the nodes can start to execute. The synchronization method of this data flow graph architecture is flexible and suitable for parallel computing.
[0054] This embodiment will specifically take 1D3P Stencil as an example for illustration, and its specific implementation process is as follows:
[0055] 1) Write the 1D3P Stencil code in a C plus directive manner. The definitions of the directive statements are as follows:
[0056] #pragma stencil_wait() means waiting for the data of adjacent data flow calculation nodes to arrive;
[0057] #pragma stencil_send() means sending the data to adjacent data flow calculation nodes;
[0058] #pragma stencil_1D3P_begin and #pragma stencil_1D3P_end indicate that this section of code is for 1D3P Stencil calculation; meanwhile, the number of iterations of the time axis for Stencil calculation is 4, the number of array elements is 128, and the number of PEs is 4; the specific 1D3P Stencil code is as follows:
[0059] “#pragma stencil_lD3P_begin for(i = 0; i < 4; i++){
[0060] for(j = 1; j < 31; j++){
[0061] if(i > 0) #pragma stencil_wait(i - 1);
[0062] Y[j] = X[j - 1] + X[j] + X[j + 1];
[0063] if(i < 3) #pragma stencil_send(i + 1);
[0064] }
[0065] Assign X = Y;
[0066] }
[0067] #pragma stencil_1D3P_end”
[0068] 2) The compiler analyzes the data dependency relationship between time axes based on the guiding information. The data dependency relationship between time axes in the above example is specifically as Figure 3 shown. Among them, the points on each time axis that are calculated on the same red diagonal line have no dependency relationship and can be executed in parallel. For example, when calculating the data of (i:0, j:3), the operation of the data of (i:1, j:1) can be executed simultaneously on another parallel computing node. When calculating the data of (i:0, j:5), the operation of the data of (i:1, j:3) and the calculation of (i:2, j:1) can be executed simultaneously on another parallel computing node, thus achieving parallel execution between time axes.
[0069] In addition, by analyzing the dependency relationship between time axes, only data transfer through data flow instructions between time axes is required to meet the calculation requirements of each parallel computing node. The three orange data on the i = 0 time axis in the figure can be transferred to the operations on the i = 1 time axis executed on adjacent nodes through data flow instructions, reducing the data access on the i = 1 time axis.
[0070] 3) Analyze the relationship between the number of iterations of the time axis and the number of PE arrays. In this example, the number of iterations of the time axis is 4, and the number of PE arrays is 4. Since the two numbers are the same, there is a one-to-one correspondence between the time axis loop i and the PE array, that is, i = PE_number. Therefore, the time axis operation with i = 0 is assigned to the PE0 parallel computing node, the time axis operation with i = 1 is assigned to the PE1 parallel computing node, the time axis operation with i = 2 is assigned to the PE2 parallel computing node, and the time axis operation with i = 3 is assigned to the PE3 parallel computing node.
[0071] 4) Data flow graph generation. The data flow graph of this Stencil calculation is as Figure 4 shown; the compiler compiles the 4 outer time axis calculations into a data flow graph that is executed in parallel on 4 PEs. The generated data flow graph consists of two data flow graphs, graph1 and graph2. graph1 is executed first, once. graph2 is executed later, and the hardware loop controls graph2 to loop and execute multiple times.
[0072] The following will take the execution phases of the PE0 parallel computing node and the PE1 parallel computing node as examples to introduce the parallel execution process when the time axis i = 0 and i = 1:
[0073] The first step: Execute the graph1 data flow graph. The IMM instructions of PE0, PE1, PE2, and PE3 are executed; the LD and Send of PE0 are executed, and the value of the R0 register is sequentially passed to PE1, PE2, and PE3;
[0074] The second step: Start to execute the graph2 data flow graph. The PE0 ADD j,j,1 instruction is executed;
[0075] The third step: The PE0 Send(X[j - 1]+X[j]+X[j + 1]) / 3,R(j%3) instruction is executed, the PE1 Wait instruction is executed, and the PE1 R1 register is updated;
[0076] The fourth step: The PE0 ADD j,j,1 instruction is looped and executed; the PE1 ADD Count,Count,1 instruction is executed;
[0077] The fifth step: The PE0 Send(X[j - 1]+X[j]+X[j + 1]) / 3,R(j%3) instruction is looped and executed, the PE1 Wait is executed, and the PE1 R2 register is updated;
[0078] Step 6: After the three participating R0, R1, and R2 registers in Timeline 1 on the PE1 parallel computing node are all transferred through the data stream instructions on Timeline 0 of the PE0 parallel computing node, numerical calculations can begin. Subsequently, each time PE0 calculates a new value, it will be passed to the register of PE1 through the data stream instruction, enabling the continuous execution of the calculation on PE1, which in turn prompts the subsequent execution of PE2 and PE3.
[0079] In summary, the present invention forms a closed loop through guidance programming - dependency analysis - timeline allocation - data flow graph generation - optimized execution; adopts a Stencil programming method based on adding data flow control semantics to the C language to express the data dependencies between timelines in Stencil calculations; at the same time, by analyzing the dependencies between timelines, the calculations of timelines are allocated to multiple parallel computing nodes, achieving coarse-grained parallel execution between timelines, reducing memory access and synchronization overheads, and improving the performance of Stencil calculations; for the optimized Stencil calculation, only the first timeline calculation node needs to perform data memory access, and the data of other calculation nodes comes from adjacent nodes, reducing the memory access overhead. Moreover, the entire calculation process does not require all data flow parallel computing nodes to synchronize. A certain calculation node only needs to communicate with adjacent nodes, greatly improving the program performance; at the same time, by executing calculations and data transmission simultaneously, the data transmission time between timeline parallel computing nodes is hidden, further improving the performance.
[0080] According to an embodiment of the present invention, the present invention further provides an electronic device 300 as shown in Figure 5 The electronic device 300 may optionally include a storage medium 100 for storing a program and a processor 200 for executing the program. When the program is executed by the processor 200, it implements the above-mentioned Stencil calculation compilation optimization method based on the data flow architecture, triggers the electronic device 300 to execute the methods and / or technical solutions based on the foregoing multiple embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0081] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose processor, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the above steps or functions. Similarly, the software program (including related data structures) of the present invention can be stored in a readable medium, such as a RAM memory.
[0082] In an alternative embodiment, the program includes program code components suitable for performing all steps of the method according to the present invention when the program runs on a processor. Optionally, the program is embodied on a readable medium.
[0083] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method or apparatus comprising a series of elements not only includes those elements but also other elements not explicitly listed, or elements inherent to such process, method or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method or apparatus comprising such element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0084] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A Stencil computation compilation optimization method based on a data flow architecture, characterized in that, Including the following steps: Writing a Stencil program by adding guidance statements in C language; wherein, the guidance statements are used to indicate the type of Stencil calculation and the data dependency relationship between time axes; Parsing the guidance statements through a compiler, analyzing the data dependency relationship between the time axes, and determining the start interval of each time axis relative to the adjacent time axis and the data required for flow; According to the relationship between the number of time axis iterations and the number of parallel data flow calculation nodes, distributing the time axis operations to adjacent parallel calculation nodes, so that consecutive time axis iterations are distributed to adjacent data flow calculation nodes; Generating data flow instructions for the operations within each time axis, and controlling the asynchronous communication and parallel execution between nodes through a data flow graph; wherein, the data flow instructions include information on the sending node and the receiving node for data transmission between adjacent nodes.
2. The Stencil calculation compilation optimization method based on a data flow architecture according to claim 1, characterized in that The distributing the time axis operations to adjacent parallel calculation nodes according to the number of time axis iterations, so that consecutive time axis iterations are distributed to adjacent data flow calculation nodes includes: If the number of time axis iterations is less than or equal to the number of parallel data flow calculation nodes, mapping each time axis iteration to adjacent data flow calculation nodes in sequence; If the number of time axis iterations is greater than the number of parallel data flow calculation nodes, distributing the time axes to adjacent data flow calculation nodes in ascending order in sequence, and then distributing the other time axes that have not been distributed to adjacent data flow calculation nodes in ascending order in sequence.
3. The Stencil computation compilation optimization method based on a data flow architecture according to claim 1, wherein The controlling the asynchronous communication and parallel execution between nodes through a data flow graph includes: Implementing data interaction between adjacent data flow calculation nodes through the data flow instructions, and the execution of the receiving node depends on the data transmitted by the sending node; Controlling the multiple executions of the data flow graph through a hardware loop to achieve the parallelism between the time axes.
4. The Stencil computation compilation optimization method based on a data flow architecture according to claim 1, wherein, After the step of generating data flow instructions for the operations within each time axis and controlling the asynchronous communication and parallel execution between nodes through a data flow graph, it further includes: Executing the optimized Stencil program in the data flow architecture, wherein only the first time axis calculation node accesses the on-chip storage unit, and the data of the remaining nodes is transmitted through adjacent nodes.
5. The Stencil computation compilation optimization method based on a data flow architecture according to claim 4, wherein The data flow architecture includes: A main control core for transmitting data and the data flow graph to the data flow processing unit array through direct memory access; The data flow processing unit array, which is composed of multiple data flow calculation nodes that can be dynamically combined, and each data flow calculation node is assigned an independent data flow graph; An on-chip memory for caching the transmission data between the data flow calculation nodes.
6. The Stencil computation compilation optimization method based on a data flow architecture according to claim 1, characterized in that The guidance statements include #pragma stencil_wait and #pragma stencil_send for describing the data dependency relationship between time axes, and guidance for declaring the type of Stencil calculation.
7. An electronic device, comprising a storage medium, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the Stencil calculation compilation optimization method based on the data flow architecture described in claims 1 to 6.