Space-time adaptive grid division and scheduling optimization method for high-speed combustion simulation application of heterogeneous multi-domain processor
Through the space-time adaptive mesh division and scheduling optimization method for heterogeneous multi-domain processors, the problems of load imbalance and communication overhead in combustion simulation are solved, and efficient computing performance and simulation accuracy are achieved, which is suitable for the scientific computing needs of heterogeneous multi-domain processors.
Patent Information
- Application Number
- CN202510179583.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-23
AI Technical Summary
In large-scale high-precision combustion simulation, the load imbalance of the area calculation with severe reactions leads to inefficient parallel efficiency. Traditional methods have problems with insufficient accuracy and large communication overhead when triggering AMR at the time step. Especially in heterogeneous multi-domain processors, how to achieve efficient space-time adaptive meshing and load balancing scheduling has become a difficult point.
A spatiotemporal adaptive meshing and scheduling optimization method for heterogeneous multi-domain processors is proposed. Through spatiotemporal adaptive meshing algorithm, grid redistribution scheduling algorithm and multi-level adaptive load balancing scheduling algorithm, dynamic load balancing and meshing optimization is realized, combining the coordinated calculation of CPU and accelerated domain clusters to improve computing efficiency.
It significantly improves the computing performance, achieves a significant reduction in the computational volume and communication volume while ensuring simulation accuracy, improves simulation efficiency, and can effectively utilize the performance of heterogeneous multi-domain processors.
Smart Images

Figure CN120029783A_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the fields of fluid mechanics and high performance computing, and in particular to a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors. [Background technology]
[0002] Numerical simulation has become a core means of designing complex combustion systems in situations where experimental costs are high and experimental conditions are limited, such as in aerospace engines and explosive blasting. High-speed combustion has very high thermal efficiency and extremely fast heat release rates, and can propagate supersonically at speeds exceeding 1000 m / s. In order to accurately capture the high-speed combustion process, it is necessary to use: (1) complex chemical models to calculate the key features of the chain; (2) a sufficiently large computational domain to capture shock wave propagation and reaction zone development; (3) an extremely short time step to meet the CFL condition; and (4) a sufficiently fine grid to avoid the influence of numerical diffusion of shock waves and reaction zones. These factors result in extremely high computing power requirements for high-speed combustion simulations, which require the use of a large-scale computing environment for efficient numerical solutions.
[0003] However, in large-scale high-precision simulations, areas with intense reactions involve more calculations of reaction equations and material changes, resulting in an unbalanced computational load and severely restricting parallel efficiency. Dynamic Load Balancing (DLB) and Adaptive Mesh Refinement (AMR) have gradually become key technologies for these challenges. Dynamic Load Balancing alleviates the problem of load imbalance by optimizing resource allocation, while Adaptive Mesh Refinement uses meshes of different resolutions for simulations of key and non-key areas of the combustion reaction, thereby significantly reducing the amount of computation while maintaining computational accuracy.
[0004] Although adaptive meshing technology can significantly reduce computing requirements, the communication overhead of meshing increases rapidly under the conditions of a large number of grids and large-scale computing nodes, making the time consumption of an AMR process equivalent to the time consumption of a time-step grid solution. In addition, the traditional method of triggering AMR with a fixed time step has limitations. When the combustion is gentle, it may cause over-computation; when the combustion is intense, it may miss key physical processes and reduce simulation accuracy.
[0005] In addition, with the rapid development of E-class supercomputer systems, heterogeneous computing architecture has gradually become the mainstream trend of high-performance computing. How to efficiently implement adaptive meshing and load balancing scheduling optimization in the time and space dimensions based on the numerical simulation software OpenFOAM on the heterogeneous multi-domain processors used in my country's new generation of E-class supercomputers to fully utilize the performance of heterogeneous multi-domain processors is a technical problem that needs to be solved urgently. [Summary of the invention]
[0006] In view of the deficiencies in the prior art, the present application provides a spatiotemporal adaptive grid partitioning and scheduling optimization method and system for high-speed combustion simulation applications for heterogeneous multi-domain processors.
[0007] The present application provides a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors, comprising the following steps:
[0008] S1, execute the space-time adaptive meshing algorithm, based on the grid load change rate of the current time step, perform space adaptive meshing according to the refinement / coarsening criteria;
[0009] S2, according to the load imbalance rate of the node, based on the grid load and grid position relationship, execute the grid redistribution scheduling algorithm to achieve static load balance before calculation;
[0010] S3, based on the grid load and computing resource power differences, execute a multi-level adaptive load balancing scheduling algorithm to achieve dynamic load balancing during calculation;
[0011] S4. The CPU and the acceleration domain cluster jointly perform grid task calculations. The CPU summarizes the solution results and updates the reaction rate, and returns to step S1.
[0012] Preferably, a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors comprises the following steps:
[0013] S1, executing the spatiotemporal adaptive meshing algorithm, judging whether the current time step triggers spatial adaptive meshing based on the grid load change rate; and,
[0014] If not, execute step S3;
[0015] If so, perform spatial adaptive meshing according to the refinement / coarsening criteria;
[0016] S2, judging whether grid redistribution is required according to the load imbalance rate of the node; and,
[0017] If not, proceed to step S3;
[0018] If so, based on the grid load and grid position relationship, a grid redistribution scheduling algorithm is executed to achieve static load balance before grid task calculation;
[0019] S3, based on the grid load and computing resource power differences, execute a multi-level adaptive load balancing scheduling algorithm to achieve dynamic load balancing during the calculation process;
[0020] S4. The CPU and the acceleration domain cluster jointly perform grid task calculations. The CPU summarizes the solution results and updates the reaction rate, and returns to step S1.
[0021] Preferably, the heterogeneous multi-domain processor includes the following features:
[0022] (1) Heterogeneous multi-domain architecture based on regional autonomy, including general area and acceleration area. The general area consists of 16 CPU cores, each of which has its own L1 cache (32KB) and L2 cache (512KB); the acceleration area consists of 96 control cores (Ctrl) and 1536 acceleration cores (Acc), which are evenly divided into 4 acceleration domain clusters. Each acceleration domain cluster includes 24 control cores (Ctrl), 384 acceleration cores (Acc), as well as on-chip shared memory (GSM), high-bandwidth shared memory (HBSM) and off-chip memory space (DDR). Each acceleration domain cluster executes independently. The CPU in the general area can access all HBSM and DDR spaces of different acceleration domain clusters, while the control core and acceleration core can only access the GSM, HBSM and DDR within their own acceleration domain cluster;
[0023] (2) High-energy-efficiency vector processing architecture. One control core and 16 acceleration cores form an acceleration array. The 16 acceleration cores in the acceleration array execute in parallel in a vector manner, using explicit on-chip storage space management, supporting double, single, and half-precision floating-point calculations. Each acceleration core contains three floating-point multiplication and addition units (MACs);
[0024] (3) High-bandwidth and high-reliability parallel memory. The parallel memory supports 16×2-way (half-word, word, double-word) parallel memory access (512B / Cycle), which can provide up to 512 bytes of vector data to 16 accelerator cores at the same time;
[0025] (4) Hierarchical heterogeneous interconnection network, including a two-dimensional mesh network in the general domain, a full cross network in the acceleration domain, and a top-level interconnection channel between domains. The two-dimensional mesh network connects 16 CPUs that support cache coherence; the full cross network connects the acceleration arrays in the acceleration domain cluster and has an ultra-high bandwidth of 2.4Tbps; the top-level interconnection channel connects the two-dimensional mesh network and the full cross network through 8 high-bandwidth data paths, serving as a data sharing bridge between the general area and the acceleration area.
[0026] Data sharing between different acceleration domain clusters is achieved through the CPU side. The general area is capable of overall task control, operating system startup and general processing, while the acceleration area is designed for computationally intensive tasks.
[0027] Preferably, the grid constitutes the entire computational domain for combustion simulation; the refinement / coarsening criterion refers to a threshold set according to the magnitude or rate of change of physical quantities (such as gradients, errors, flow field characteristics, etc.). When the refinement criterion is met, the grid is refined, and when the coarsening criterion is met, the grid is coarsened; the node is an independent computer containing a heterogeneous multi-domain processor; the grid load represents the time consumed by the grid task for solving the grid in the previous time step; the grid task is an independent task after packing the variables required for solving the ordinary differential equation of the grid calculation; the ordinary differential equation refers to the equation represented by the rate of change of the thermal chemical composition over time; the reaction rate refers to the speed at which a chemical reaction proceeds, expressed as the change in the amount of substance concentration per unit time; the substance concentration refers to the amount of a certain chemical substance per unit volume, including the concentrations of fuel, oxidizer, intermediate products, and final products.
[0028] Preferably, the spatio-temporal adaptive grid partitioning algorithm belongs to the spatio-temporal adaptive grid partitioning module, which includes a time adaptive grid partitioning sub-module and a space adaptive grid partitioning sub-module. The space adaptive grid partitioning sub-module adopts an MPI+OpenMP hybrid parallel structure to improve the parallel efficiency. The spatio-temporal adaptive grid partitioning algorithm includes the following steps:
[0029] (1) Time adaptive grid partitioning calculates the change in the grid load of each grid compared to that after the previous space adaptive grid partitioning before the start of each time step calculation.
[0030] (2) Collect the change rates of all grid loads and calculate the mean and standard deviation.
[0031] (3) Calculate the load change threshold β based on the mean and standard deviation. If the proportion of grids higher than the load change threshold exceeds the threshold α, perform space adaptive grid partitioning; otherwise, directly execute the multi-level adaptive load balancing scheduling algorithm.
[0032] (4) The refinement and coarsening triggering mechanism of space adaptive grid partitioning is based on the refinement and coarsening criteria defined in the simulation input file. Starting from the grid with the highest refinement level, traverse and determine whether the grid at the current level meets the coarsening condition; then, traverse the grids at lower refinement levels step by step and mark the grids that need to be refined or coarsened; among them, the refinement level refers to the number of times the initial grid is refined into smaller grids.
[0033] (5) All grids marked for refinement / coarsening are refined / coarsened according to the quadtree / octree partitioning rules, and the data is mapped from the original grid to the new grid; among them, the quadtree / octree partitioning represents that in two-dimensional and three-dimensional cases, the grid is refined into 4 / 8 sub-grids of equal size, or 4 / 8 sub-grids are merged into one parent grid.
[0034] (6) Construct a buffer layer, and traverse the grids again from the highest refinement level. The grids at the highest refinement level are marked as the retention set, and the refinement level remains unchanged; determine the search range based on the preset number of buffer layers, find all adjacent grids in the retention set, and include them in the neighbor grid set; use the updated neighbor grid set to repeatedly search for neighbor grids; for these neighbor grids, reduce the current refinement level of the retention set by 1 and use it as the target refinement level; change the neighbor grid set into a new retention set, and repeat the above process; wherein the buffer layer is a transition grid set to ensure a smooth transition between grids of different refinement levels.
[0035] Preferably, the grid redistribution scheduling algorithm belongs to a grid load balancing scheduling module, comprising the following steps:
[0036] (1) Calculate the average grid load of each node and the maximum load imbalance rate. If the maximum load imbalance rate exceeds the static load imbalance threshold, grid redistribution is required. Otherwise, the multi-level adaptive load balancing scheduling algorithm is directly executed.
[0037] (2) Execute the Scotch grid allocation method, represent all grids and their positional relationships as a graph structure, assign a weight to each point in the graph to represent the grid load, and assign a weight to each edge to represent the communication overhead between grids. The grid allocation method makes the node load close while minimizing the size of the boundary between nodes. First, generate the initial partition based on the geometric position, and then use the Kernighan-Lin algorithm to iteratively optimize the initial partition. By exchanging nodes and adjusting boundaries, the communication overhead is gradually reduced and the load balancing is improved.
[0038] Preferably, the multi-level adaptive load balancing scheduling algorithm belongs to a grid load balancing scheduling module, comprising the following steps:
[0039] (1) Node-level load balancing scheduling. The thermochemical variables and other variables required for solving ordinary differential equations in the grid are packaged into a grid task set. The grid loads of all nodes are sorted and the average node load is calculated. Nodes with a load greater than the average send some grid tasks to nodes with a load less than the average until all nodes reach the global average load, and the grid task set is updated;
[0040] (2) CPU and acceleration domain cluster-level load balancing scheduling. According to the time t that the CPU waited for the acceleration domain cluster in the previous time step, wait and the waiting time threshold t thresold Calculate the load distribution ratio S p ; Among them, S p =S p-1 +a×(t wait -t threshod ), a is a constant coefficient; according to Sp , select grid tasks with appropriate load from the grid task set and assign them to the CPU and acceleration domain cluster; after the solution is completed, record the length of time the CPU waits for the acceleration domain cluster for use in the CPU and acceleration domain cluster-level load balancing scheduling in the next time step;
[0041] (3) Thread-level load balancing scheduling. The grid tasks assigned to the acceleration domain clusters in the load balancing scheduling between the CPU and the acceleration domain clusters are evenly divided according to the load and sent to each acceleration domain cluster; the threads of each acceleration domain cluster are placed in a thread pool to jointly solve the grid tasks; if there are idle threads in the thread pool and there are grid tasks to be solved, the idle threads will take a grid task to solve, and return to the thread pool after the solution is completed, until all grid tasks are solved; the acceleration domain cluster transmits the solution results back to the CPU.
[0042] Preferably, the grid task calculation belongs to a grid task calculation module, and includes the following steps:
[0043] (1) Before solving the first time step, the CPU transfers the thermodynamic constants and other constants to the on-chip shared memory of the acceleration domain cluster for storage, facilitating subsequent data reuse;
[0044] (2) Split the grid tasks to be sent by the CPU to the acceleration domain cluster into multiple independent subtask sets, and send them to the acceleration domain cluster in sequence in a stream manner. After starting the acceleration domain cluster, the CPU can parallelly calculate the grid tasks through the CPU threads, achieving the overlap of calculation and communication;
[0045] (3) The acceleration domain cluster threads parallelize the grid tasks, construct vectors during the material concentration calculation process, and the acceleration cores parallelize the material concentration changes;
[0046] (4) After the calculation is completed, the CPU summarizes the solution results and transmits the grid task back to the original node, updating the grid reaction rate.
[0047] The present application also provides a spatiotemporal adaptive grid partitioning and scheduling optimization system for high-speed combustion simulation applications for heterogeneous multi-domain processors, including:
[0048] A time-space adaptive meshing module, comprising a time-adaptive meshing submodule for determining whether meshing needs to be performed at the current time step and a space-adaptive meshing submodule for performing meshing, and for performing space-adaptive meshing according to a refinement / coarsening criterion based on a rate of change of mesh load;
[0049] A grid load balancing scheduling module, including a grid redistribution submodule for performing static load balancing before calculation and a multi-level adaptive load balancing submodule for performing dynamic load balancing during chemical calculation, for executing a grid redistribution scheduling algorithm based on the load imbalance rate of the nodes, the grid load and the grid position relationship to achieve static load balancing before calculation; and executing a multi-level adaptive load balancing scheduling algorithm based on the grid load and the computing power difference of the computing resources to achieve dynamic load balancing during calculation;
[0050] The grid task calculation module is used for the CPU and the acceleration domain cluster to jointly perform grid task calculations, obtain solution results and update reaction rates.
[0051] Preferably, the present application provides a computing device, comprising a processor and a memory; the memory is used to store computer-executable instructions; the processor is used to execute the computer-executable instructions so that the computing device executes any of the methods described above.
[0052] Preferably, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes computer program instructions, and when the computer program instructions are executed by a computing device, the computing device executes any one of the methods described above.
[0053] Experiments show that the computational performance of 100 nodes (a total of 11,200 integrated computing units) using a spatiotemporal adaptive meshing and scheduling optimization method for high-speed combustion simulation applications on heterogeneous multi-domain processors is 12.30 times higher than that of the original adaptive meshing with a fixed time step and the original high-speed combustion solver for 100 nodes.
[0054] From the above description, it can be concluded that the present invention realizes the spatiotemporal adaptive grid division and scheduling optimization and heterogeneous multi-domain collaborative computing of large-scale heterogeneous multi-domain processors while ensuring the accuracy of high-speed combustion simulation, thereby greatly improving the overall computing performance.
[0055] This method uses the time spent solving the grid task in the previous time step as the grid load in the current time step. By sensing the change in grid load, the changing grid computational amount has a basis for measurement, and adaptive grid division and load balancing scheduling in both time and space are realized based on the grid load. Time-adaptive grid division reduces unnecessary AMR overhead, and the combined use of static and dynamic load balancing reduces the communication overhead of scheduling and improves resource utilization. Ultimately, the method significantly optimizes the computational amount and communication volume, improves simulation efficiency, and meets the needs of large-scale scientific computing for high-speed combustion simulation while ensuring accuracy.
Brief Description of the Drawings
[0056] Figure 1A flow chart of a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors according to the present invention;
[0057] Figure 2 A multi-level adaptive load balancing scheduling step diagram of a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors according to the present invention;
[0058] Figure 3 It is an acceleration domain framework diagram of a heterogeneous multi-domain processor of the present invention, which is a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation application of a heterogeneous multi-domain processor;
[0059] Figure 4 This is a framework diagram of a spatiotemporal adaptive grid partitioning and scheduling optimization system for high-speed combustion simulation applications for heterogeneous multi-domain processors according to the present invention. [Specific implementation method]
[0060] The present invention discloses a method and system for spatiotemporal adaptive grid division and scheduling optimization for high-speed combustion simulation applications for heterogeneous multi-domain processors. High-speed detonation combustion is a process in which, when the energy released by the combustion reaction is strong enough, a shock wave is generated that propagates at supersonic speed in the unburned gas, compresses and heats the unburned gas, and causes it to react rapidly. This application can improve the performance of high-speed combustion simulation while ensuring accuracy, and can be widely used in the research of fields such as aircraft engine propulsion technology and explosive blasting, as well as optimizing the combustion process and improving fuel utilization efficiency in the energy field.
[0061] This application can significantly improve computing performance while ensuring accuracy. The specific implementation of this application will be described in detail below in conjunction with the accompanying drawings. It is worth noting that what is described in this implementation is only a common use of this application. The technology used in this application should be protected until practitioners in related fields can develop better solutions.
[0062] In one embodiment, Figure 1 This is a flow chart of a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors of the present invention. It can be seen that
[0063] Step 101, time adaptive meshing determines whether the current time step triggers space adaptive meshing based on the mesh load change rate; and
[0064] If not, execute step 105;
[0065] If yes, execute step 102;
[0066] Step 102, performing spatial adaptive meshing according to a refinement / coarsening criterion;
[0067] Step 103, judging whether grid redistribution is required according to the load imbalance rate of the node; and,
[0068] If not, execute step 105;
[0069] If yes, execute step 104;
[0070] Step 104, executing a grid redistribution scheduling algorithm according to the grid load and the grid position relationship to achieve static load balance before calculation;
[0071] Step 105, executing a multi-level adaptive load balancing scheduling algorithm according to the grid load and computing resource computing power differences to achieve dynamic load balancing during calculation;
[0072] Step 106: The CPU and the acceleration domain cluster jointly perform grid task calculations, the CPU summarizes the solution results and updates the reaction rate, and then returns to step S1.
[0073] In step 101, the time adaptive meshing submodule determines whether the current time step needs to perform space adaptive meshing, including the following steps:
[0074] (1) Before solving each time step, compare the change of each grid load compared with the last spatial adaptive grid division R i :
[0075]
[0076] in, is the load of the ith grid after N time steps of the last spatial adaptive grid division, is the load of the ith grid during the last spatial adaptive grid division.
[0077] (2) Collect all grid load change rates R i , and calculate the mean μ R and standard deviation σ R ;
[0078] (3) Calculate the load change threshold β based on the mean and standard deviation:
[0079] β=μ R +b·σ R
[0080] Where b is an adjustable parameter used to control the sensitivity of the threshold to capture grids with load variations above the average level. i>β, if the grid ratio above the load change threshold exceeds the threshold α, the spatial adaptive grid division submodule is executed, otherwise the multi-level adaptive load balancing submodule is directly executed.
[0081] In step 102, the space adaptive meshing submodule performs space adaptive meshing, and uses the MPI+OpenMP hybrid parallel structure to improve parallel efficiency. For a large number of loop iteration tasks, multi-threaded synchronous execution is achieved by introducing OpenMP compilation instructions outside the loop. For the data competition problem caused by multi-threading, unnecessary shared variables are converted into thread private variables to avoid simultaneous modification by multiple threads; for necessary shared variables, OpenMP atomic operation instructions are used to avoid competition conditions. The space adaptive meshing submodule includes the following steps:
[0082] (1) The refinement and coarsening trigger mechanism of spatial adaptive meshing is based on the refinement and coarsening criteria defined in the simulation input file. Starting from the grid with the highest refinement level, the grid at the current level is traversed and judged whether it meets the coarsening condition; then, the grids with lower refinement levels are traversed step by step, marking the grids that need to be refined or coarsened; wherein the refinement level refers to the number of times the initial grid is refined into smaller grids.
[0083] (2) All grids marked as refined / coarsened are refined / coarsened according to the quadtree / octree partitioning rule, and the data is mapped from the original grid to the new grid; wherein the quadtree / octree partitioning means, in two-dimensional and three-dimensional cases, refining the grid into 4 / 8 sub-grids of equal size, or merging 4 / 8 sub-grids into one parent grid.
[0084] (3) Construct a buffer layer, traverse the grids again from the highest refinement level, and mark the grids at the highest refinement level as the retention set, keeping the refinement level unchanged; determine the search range based on the preset number of buffer layers, find all adjacent grids in the retention set, and include them in the neighbor grid set; use the updated neighbor grid set to repeatedly search for neighbor grids; for these neighbor grids, reduce the current refinement level of the retention set by 1 and use it as the target refinement level; change the neighbor grid set into a new retention set, and repeat the above process; wherein the buffer layer is a transition grid set to ensure a smooth transition between grids of different refinement levels.
[0085] In step 103 and step 104, the grid redistribution scheduling algorithm belongs to the grid load balancing scheduling module, and includes the following steps:
[0086] (1) Calculate the average grid load of each node and the maximum load imbalance rate D max :
[0087]
[0088] Among them, N node is the number of nodes. If the maximum load imbalance rate of the node is D max Exceeding the static load imbalance threshold D threshold , then grid redistribution is required, otherwise the multi-level adaptive load balancing scheduling algorithm is directly executed;
[0089] (2) Use the parallel Scotch method to allocate grids. All grids and their positional relationships are represented as a graph structure. A weight is assigned to each point in the graph to represent the grid load. A weight is assigned to each edge to represent the communication overhead between grids. The grid allocation method makes the node load close while minimizing the size of the boundary between nodes. First, the initial partition is generated based on the geometric position. Then, the Kernighan-Lin algorithm is used to iteratively optimize the initial partition. By exchanging nodes and adjusting boundaries, the communication overhead is gradually reduced and the load balancing is improved. Finally, the optimized graph partition result is output. Each subgraph corresponds to a grid to which a node is assigned.
[0090] In step 105, the multi-level adaptive load balancing scheduling algorithm belongs to the grid load balancing scheduling module, and includes the following steps:
[0091] (1) Node-level load balancing scheduling. The thermochemical variables and other variables required for solving ordinary differential equations in the grid are packaged into a grid task set. The content of the grid task is shown in Table 1. Here, N sp is the number of substances involved in the chemical reaction. Sort the mesh load of all nodes and calculate the average node load The load is greater than The node is the sender, and the load is less than The node is the receiver, and the sender sends part of the grid task to the receiver until all nodes reach the global average load. Update the grid task collection;
[0092] Table 1
[0093]
[0094]
[0095] (2) CPU and acceleration domain cluster-level load balancing scheduling. According to the time t that the CPU waited for the acceleration domain cluster in the previous time step, wait and the waiting time threshold t threshold Calculate the load distribution ratio S p ; Among them, S p =S p-1 +a×(t wait -t threshold ), a is a constant coefficient; according to Sp , select from the grid task set the task that satisfies l cpu The grid tasks with the highest load are assigned to the CPU, and the remaining grid tasks are assigned to the acceleration domain cluster; l is the load of the current node, N core 、N ctrl and N cluster are the number of CPU cores, control cores, and acceleration domain clusters in the current node. After the solution is completed, the length of time the CPU waits for the acceleration domain cluster is recorded for use in the CPU and acceleration domain cluster-level load balancing scheduling in the next time step.
[0096] (3) Thread-level load balancing scheduling. The grid tasks assigned to the acceleration domain clusters in the load balancing scheduling between the CPU and the acceleration domain clusters are evenly divided according to the load and sent to each acceleration domain cluster; the threads of each acceleration domain cluster are placed in a thread pool to jointly solve the grid tasks; if there are idle threads in the thread pool and there are grid tasks to be solved, the idle threads will take a grid task to solve, and return to the thread pool after the solution is completed, until all grid tasks are solved; the acceleration domain cluster transmits the solution results back to the CPU.
[0097] In step 106, the grid task calculation belongs to the grid task calculation module, and includes the following steps:
[0098] (1) Before solving the first time step, the CPU transfers the thermodynamic constants and other constants to the on-chip shared memory of the acceleration domain cluster for storage, facilitating subsequent data reuse;
[0099] (2) Split the grid tasks to be sent by the CPU to the acceleration domain cluster into multiple independent subtask sets, and send them to the acceleration domain cluster in sequence in a stream manner. After starting the acceleration domain cluster, the CPU can parallelly calculate the grid tasks through the CPU threads, achieving the overlap of calculation and communication;
[0100] (3) The acceleration domain cluster threads parallelize the grid tasks, construct vectors during the material concentration calculation process, and the acceleration cores parallelize the material concentration changes;
[0101] (4) After the calculation is completed, the CPU summarizes the solution results and transmits the grid task back to the original node, updating the grid reaction rate.
[0102] In one embodiment, Figure 2 This is a multi-level adaptive load balancing scheduling step diagram of a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors of the present invention. It can be seen that
[0103] Step 201, perform node-level load balancing scheduling, first determine the sender, receiver and grid task amount according to the load of each node and the average load of all nodes, unify the grid tasks sent from the sender to the receiver through MPI, and the load sent is at least greater than 1% of the sender node load, that is, ignore very small inter-node load balancing communication;
[0104] Step 202: Since there is a 16-core CPU and 4 independent acceleration domain clusters in the node, 4 CPU cores control 1 acceleration domain cluster respectively, and after performing load balancing scheduling between the CPU and the acceleration domain clusters, the process proceeds to step 203; the remaining 12 CPU cores directly proceed to step 203;
[0105] Step 203: The 12 CPU cores that directly enter step 203 each create a thread to perform grid task calculations; the remaining CPU cores also serve as threads to perform grid task calculations after completing step 202; the four acceleration domain clusters each create a thread pool consisting of 24 acceleration arrays to perform thread-level load balancing scheduling and grid task calculations;
[0106] Among them, the steps that require data communication are Figure 2 Indicated by black arrows.
[0107] In one embodiment, Figure 3 The figure shows the overall framework of a heterogeneous multi-domain processor of a spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications of a heterogeneous multi-domain processor. The characteristics of the heterogeneous multi-domain processor are as follows:
[0108] (1) Heterogeneous multi-domain architecture based on regional autonomy, including general area and acceleration area. The general area consists of 16 CPU cores, each with its own L1 cache (32KB) and L2 cache (512KB); the acceleration area consists of 96 control cores (Ctrl) and 1536 acceleration cores (Acc), which are evenly divided into 4 acceleration domain clusters. Each acceleration domain cluster includes 24 control cores (Ctrl), 384 acceleration cores (Acc), as well as on-chip shared memory (GSM), high-bandwidth shared memory (HBSM) and off-chip memory space (DDR). Each acceleration domain cluster executes independently. The CPU in the general area can access all HBSM and DDR spaces of different acceleration domain clusters, while the control core and acceleration core can only access the GSM, HBSM and DDR in their own acceleration domain cluster.
[0109] (2) High-efficiency vector processing architecture. One control core and 16 acceleration cores form an acceleration array. The 16 acceleration cores in the acceleration array execute in parallel in a vector manner, using explicit on-chip storage space management, supporting double, single, and half-precision floating-point calculations, and each acceleration core contains three floating-point multiplication and addition units (MACs).
[0110] (3) High-bandwidth and high-reliability parallel memory. The parallel memory supports 16×2-way (half-word, word, double-word) parallel memory access (512B / Cycle), which can provide up to 512 bytes of vector data to 16 accelerator cores at the same time.
[0111] (4) Hierarchical heterogeneous interconnection network, including a two-dimensional mesh network in the general domain, a full cross network in the acceleration domain, and a top-level interconnection channel between domains. The two-dimensional mesh network connects 16 CPUs that support cache coherence; the full cross network connects the acceleration arrays in the acceleration domain cluster and has an ultra-high bandwidth of 2.4Tbps; the top-level interconnection channel connects the two-dimensional mesh network and the full cross network through 8 high-bandwidth data paths, serving as a data sharing bridge between the general area and the acceleration area.
[0112] Data sharing between different acceleration domain clusters is achieved through the CPU side. The general area is capable of overall task control, operating system startup and general processing, while the acceleration area is designed for computationally intensive tasks.
[0113] In one embodiment, Figure 4 The figure shows a framework diagram of a spatiotemporal adaptive grid partitioning and scheduling optimization system for high-speed combustion simulation applications for heterogeneous multi-domain processors. It can be seen that
[0114] Between the spatiotemporal adaptive grid division module in step 401 and the grid load balancing scheduling module in step 402, after each grid division is completed, the current grid load of the node needs to be evaluated. If the static load imbalance threshold is not reached, there is no need to perform grid redistribution; otherwise, grid redistribution needs to be performed;
[0115] Between the spatiotemporal adaptive grid division module in step 401 and the grid task calculation module in step 403, it is necessary to determine whether adaptive grid division is required after grid calculation is completed at each time step. The result of grid division will directly affect the efficiency and accuracy of grid calculation;
[0116] Between the grid task calculation module in step 403 and the grid load balancing scheduling module in step 402, the grid task calculation starts after the multi-level adaptive load balancing scheduling; the time taken for each grid task calculation will affect the result of the next multi-level adaptive load balancing scheduling.
[0117] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0118] Finally, it should be noted that although the present invention is mainly based on specific cases in terms of example display and design, and the software involved is in the form of a system. However, the various components involved in this system should still be protected, including but not limited to modules such as resource monitors and task schedulers and their codes, so any modification, improvement, use, etc. of the implementation of the present invention should be covered in the claims and description of the present invention.
Claims
1. A spatiotemporal adaptive grid partitioning and scheduling optimization method for high-speed combustion simulation applications for heterogeneous multi-domain processors, characterized in that: The following steps are involved: S1, execute the space-time adaptive meshing algorithm, based on the grid load change rate of the current time step, perform space adaptive meshing according to the refinement / coarsening criteria; S2. According to the load imbalance rate of the nodes, based on the grid load and the grid position relationship, the grid redistribution scheduling algorithm is executed to achieve static load balance before grid task calculation; S3, based on the grid load and computing resource power differences, execute a multi-level adaptive load balancing scheduling algorithm to achieve dynamic load balancing during the calculation process; S4, the CPU and the acceleration domain cluster jointly perform grid task calculations, the CPU summarizes the solution results and updates the reaction rate, and returns to step S1; The grid load represents the time consumed in solving the grid task of the grid in the previous time step.
2. The method according to claim 1, characterized in that The following steps are involved: S1, execute the spatiotemporal adaptive meshing algorithm, and determine whether the current time step triggers spatial adaptive meshing based on the grid load change rate; as well as, If not, execute step S3; If so, perform spatial adaptive meshing according to the refinement / coarsening criteria; S2, judging whether grid redistribution is needed according to the load imbalance rate of the node; as well as, If not, execute step S3; If so, based on the grid load and grid position relationship, a grid redistribution scheduling algorithm is executed to achieve static load balance before calculation; S3, based on the grid load and computing resource power differences, execute a multi-level adaptive load balancing scheduling algorithm to achieve dynamic load balancing during calculation; S4. The CPU and the acceleration domain cluster jointly perform grid task calculations. The CPU summarizes the solution results and updates the reaction rate, and returns to step S1.
3. The method according to claim 1, characterized in that The heterogeneous multi-domain processor includes a general area and an acceleration area. The general area includes 1 CPU, and the CPU includes 16 CPU cores; the acceleration area includes 4 identical acceleration domain clusters, and the acceleration domain clusters include 24 control cores and 384 acceleration cores. The storage device includes on-chip shared memory, high-bandwidth shared memory and off-chip memory.
4. The method according to claim 1, characterized in that: The time-space adaptive grid division algorithm includes time adaptive grid division and space adaptive grid division, and includes the following steps: S11, time adaptive meshing Before each time step, calculate the rate of change R of each grid load compared to the grid load after the last spatial adaptive meshing. i : in, is the load of the ith grid after N time steps of the last spatial adaptive grid division, is the load of the ith grid during the last spatial adaptive grid division; S12. Collect all grid load change rates and calculate the mean grid load change rate μ R and standard deviation σ R ; S13, according to the mean and standard deviation, a grid load change rate threshold β is obtained, and R is counted. i >β grid ratio, if the grid ratio exceeds the threshold α, then perform spatial adaptive grid division; wherein, β=μ R +b·s R Where b is an adjustable parameter used to control the sensitivity of the threshold to capture grids with load variations above the average level; S14, starting from the grid of the highest refinement level, traverse and determine whether the grid of the current level meets the coarsening condition according to the refinement / coarsening standard; traverse the grids of the lower refinement levels step by step, and mark the grids that need to be refined or coarsened; wherein the refinement level refers to the number of times the initial grid is refined into smaller grids; S15, refining / coarsening all grids marked as refined / coarsened according to the quadtree / octree partitioning rule, and mapping the data from the original grid to the new grid; wherein the quadtree / octree partitioning means, in two-dimensional and three-dimensional cases, refining the grid into 4 / 8 sub-grids of equal size, or merging 4 / 8 sub-grids into one parent grid; S16, construct a buffer layer, traverse the grids again from the highest refinement level, mark the grids at the highest refinement level as the retention set, and keep the refinement level unchanged; perform the operation of searching neighbor grids according to the preset number of buffer layers, find all adjacent grids in the retention set, and include them in the neighbor grid set; use the updated neighbor grid set to repeatedly perform the operation of searching neighbor grids; for these neighbor grids, reduce the current refinement level of the retention set by 1 and use it as the target refinement level; change the neighbor grid set into a new retention set, and repeat the above process; wherein, the buffer layer is a transition grid set to ensure a smooth transition between grids of different refinement levels.
5. The method according to claim 1, characterized in that The grid redistribution scheduling algorithm comprises the following steps: (1) Calculate the average value and maximum load imbalance rate of each node's grid load. If the maximum load imbalance rate exceeds the static load imbalance threshold, grid redistribution is performed. Otherwise, multi-level adaptive load balancing is directly performed. (2) Execute the Scotch grid allocation method to represent all grids and their positional relationships as a graph structure. Assign a weight to each point in the graph to represent the grid load, and assign a weight to each edge to represent the communication overhead between grids. The grid allocation method makes the node load close while minimizing the size of the boundary between nodes. First, generate the initial partition based on the geometric position, and then use the Kernighan-Lin algorithm to iteratively optimize the initial partition. By exchanging nodes and adjusting boundaries, the communication overhead is gradually reduced to improve load balancing.
6. The method according to claim 1, characterized in that The multi-level adaptive load balancing scheduling algorithm comprises the following steps: S31, node-level load balancing scheduling: package the thermochemical variables and other variables required for solving ordinary differential equations in the grid into a grid task set; sort the grid loads of all nodes and calculate the average node load, and nodes with a load greater than the average load send part of the grid tasks to nodes with a load less than the average load until all nodes reach the global average load, and update the grid task set; S32, CPU and acceleration domain cluster-level load balancing scheduling: According to the CPU waiting time t of the acceleration domain cluster in the previous time step wait and the waiting time threshold t threshold Calculate the load distribution ratio S p ; Among them, S p =S p-1 +a×(t wait -t threshold ), a is a constant coefficient; according to S p , select grid tasks with appropriate load from the grid task set and assign them to the CPU and acceleration domain cluster; after the solution is completed, record the length of time the CPU waits for the acceleration domain cluster for use in the CPU and acceleration domain cluster-level load balancing scheduling in the next time step; S33, thread-level load balancing scheduling: the grid tasks assigned to the acceleration domain cluster in the load balancing scheduling between the CPU and the acceleration domain cluster are evenly divided according to the load and sent to each acceleration domain cluster; the threads of each acceleration domain cluster are placed in a thread pool to jointly solve the grid tasks; if there are idle threads in the thread pool and there are grid tasks to be solved, the idle thread takes a grid task to solve, and returns to the thread pool after the solution is completed, until all grid tasks are solved; the acceleration domain cluster transmits the solution results back to the CPU.
7. The method according to claim 1, characterized in that The grid task calculation includes the following steps: S41, before the solution of the first time step begins, the CPU transfers the thermodynamic constants and other constants to the on-chip shared memory of the acceleration domain cluster for storage, so as to facilitate subsequent data reuse; S42, dividing the grid task to be sent by the CPU to the acceleration domain cluster into multiple independent subtask sets, and sending them to the acceleration domain cluster in sequence in a stream manner; after starting the acceleration domain cluster, the CPU calculates the grid task in parallel through the CPU thread to achieve the overlap of calculation and communication; S43, the acceleration domain cluster threads parallelly calculate the grid tasks, construct vectors during the material concentration calculation process, and the acceleration cores parallelly calculate the material concentration changes; S44. After the calculation is completed, the CPU summarizes the solution results and transmits the grid task back to the original node, and updates the grid reaction rate.
8. A spatiotemporal adaptive grid partitioning and scheduling optimization system for high-speed combustion simulation applications for heterogeneous multi-domain processors according to any one of claims 1 to 7, characterized in that: include: A time-space adaptive meshing module, comprising a time-adaptive meshing submodule for determining whether meshing needs to be performed at the current time step and a space-adaptive meshing submodule for performing meshing, and for performing space-adaptive meshing according to a refinement / coarsening criterion based on a rate of change of mesh load; A grid load balancing scheduling module, including a grid redistribution submodule for performing static load balancing before calculation and a multi-level adaptive load balancing submodule for performing dynamic load balancing during chemical calculation, for executing a grid redistribution scheduling algorithm based on the load imbalance rate of the nodes, the grid load and the grid position relationship, to achieve static load balancing before calculation; And based on the grid load and computing resource power differences, a multi-level adaptive load balancing scheduling algorithm is executed to achieve dynamic load balancing during calculation; The grid task calculation module is used for the CPU and the acceleration domain cluster to jointly perform grid task calculations, obtain solution results and update reaction rates.
9. A computing device, characterized in that The method comprises a processor and a memory; the memory is used to store computer-executable instructions; the processor is used to execute the computer-executable instructions so that the computing device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises computer program instructions, and when the computer program instructions are executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 7.
Citation Information
Cited By
AMRM-FEM-SPH-based shaped jet simulation method
CN120524762A
Data center network topology dynamic adjustment method, storage medium and equipment
CN120856575A