Computing node load balancing method and device based on octree
By load balancing the computing nodes based on octree, the problem of load imbalance in the material point simulation process is solved, and more efficient calculation and data transmission is achieved.
Patent Information
- Application Number
- CN202510114440.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-24
AI Technical Summary
During the material point method simulation process, as the simulation process progresses, the calculation tasks carried in the subspace area change, resulting in unbalanced loads of the calculation tasks of each process and the load needs to be adjusted frequently, resulting in large calculation and communication overhead.
The octree-based computing node load balancing method is adopted to recursively divide the calculation area, build an octree structure, determine the spatial fill curve, allocate leaf nodes to the computing node according to the filling order and load value, ensure load balancing, and recalculate and adjust when the load gap is greater than the preset value.
It effectively reduces the computing and communication overhead of load adjustment of computing nodes, optimizes the load distribution of each computing node, and improves the computing efficiency and data transmission efficiency in the simulation process.
Smart Images

Figure CN120012508A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of computer technology, and more particularly, to an octree-based computing node load balancing method and device. Background Art
[0002] The material point method (MPM) discretizes material regions through particles, and each particle carries all material information, such as mass, position, momentum, etc. In hypervelocity collisions, the material point method can accurately simulate the deformation, fragmentation, and sputtering of objects during hypervelocity collisions. Based on the material point method, the coupled material point finite difference method is developed for fluid-solid coupling processes with strong nonlinearity, and the coupled material point finite difference finite element method is developed for small deformation problems. In the initialization stage of the simulation, the above method can achieve load balancing of the computing nodes by reasonably dividing the subspace region and carrying the computing tasks of multiple subspace regions on the computing nodes. However, as the simulation process proceeds, the computing tasks carried in the subspace region change, resulting in an unbalanced load on the computing tasks of each process. When it is necessary to adjust the load of the computing tasks of each process, it may be necessary to re-divide the entire space region, resulting in large computing and communication overheads.
[0003] Application Contents
[0004] This application describes an octree-based computing node load balancing method and device, which can solve the above technical problems.
[0005] According to a first aspect, a method for balancing the load of computing nodes based on an octree is provided, comprising: obtaining a computing area, dividing the computing area into a plurality of sub-areas, wherein the computing area is a space involved in computing during a simulation process of an explosion event;
[0006] Recursively divide the calculation area until the load values of the smallest areas in the sub-areas are all less than the preset load values, and obtain an octree structure corresponding to the calculation area, wherein the leaf nodes in the octree structure correspond to the smallest areas in the sub-areas;
[0007] Determine a space filling curve of the computing region according to the octree structure, and distribute the leaf nodes to different computing nodes in sequence according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each computing node is balanced;
[0008] When the gap value of the computing node load value is greater than the preset gap value, the load value of the leaf node is recalculated, and the leaf node is reallocated to the computing node according to the load value of the leaf node, so that the load value of the computing node is rebalanced.
[0009] Based on the above embodiment, determining the space filling curve of the calculation area according to the octree structure specifically includes:
[0010] Traversing from the root node of the octree structure, assigning a corresponding code to each node, recursively accessing its child nodes until a leaf node is reached, and assigning a code to the leaf node;
[0011] The codes of the nodes in the path from the root node to the leaf node are expressed as a code sequence, and the code sequence is the Z-curve code of the leaf node;
[0012] Traversing the octree structure by a depth-first search method to obtain Z-curve codes of all leaf nodes in the octree structure;
[0013] The Z-curve codes of the leaf nodes are sorted using a lexicographic order to obtain a one-dimensional leaf node queue, where the leaf node queue is the space-filling curve of the calculation region.
[0014] Based on the above further embodiment, the leaf nodes are sequentially allocated to different computing nodes according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each computing node is balanced, specifically including:
[0015] Determine the load value of each of the minimum areas according to the computing tasks carried by the minimum area in the sub-area, where the load value of the minimum area is the load value of the corresponding leaf node;
[0016] According to the filling order of the space filling curve, the leaf nodes corresponding to the current filling area are sequentially allocated to the computing nodes until the load of the computing node reaches a preset load threshold;
[0017] Continue to assign the leaf node corresponding to the next filling area to another computing node until all filling areas are assigned.
[0018] Based on the above embodiment, when the difference value of the calculated node load value is greater than the preset difference value, recalculating the load value of the leaf node specifically includes:
[0019] During the simulation process, according to a preset time interval, it is compared whether the difference between the maximum load value and the minimum load value in the calculation node is greater than the preset difference value;
[0020] If so, determining whether the load value of the leaf node in the octree structure is greater than a preset load value;
[0021] If so, the minimum area corresponding to the leaf node with a load value greater than the preset load value is divided until the load value of the minimum area is less than the preset load value, a new octree structure is obtained, the space filling curve is updated, and the load value of the leaf node in the new octree structure is recalculated;
[0022] Otherwise, the load values of the leaf nodes in the octree structure are recalculated.
[0023] Based on the above embodiment, the method further includes:
[0024] quantifying the complexity of different computational tasks involved in the computational region, the computational tasks comprising one or more of a material point method computational task, a coupled material point finite difference method computational task, and a coupled material point finite difference finite element method computational task;
[0025] The load values of the sub-areas in the calculation area and the load values of the leaf nodes are determined according to the number of material points in the material point method calculation task, the number of fluid grids in the coupled material point finite difference method calculation task and the number of rod units in the coupled material point finite difference finite element method calculation task.
[0026] Based on the above embodiment, the method further includes:
[0027] A temporary buffer is set between the computing nodes, and the temporary buffer is used to store the data to be transmitted of the computing nodes.
[0028] Based on the above embodiment, the method further includes: the computing nodes communicate with each other in a non-blocking manner.
[0029] According to a second aspect, a computing node load balancing device based on an octree is provided, the device comprising:
[0030] The first processing module is used to obtain a calculation area and divide the calculation area into a plurality of sub-areas, wherein the calculation area is a space involved in the calculation of the explosion event simulation process;
[0031] A second processing module is used to recursively divide the calculation area until the load values of the smallest areas in the sub-areas are all less than the preset load values, and obtain an octree structure corresponding to the calculation area, wherein the leaf nodes in the octree structure correspond to the smallest areas in the sub-areas;
[0032] A third processing module is used to determine a space filling curve of the calculation area according to the octree structure, and distribute the leaf nodes to different calculation nodes in turn according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each calculation node is balanced;
[0033] The fourth processing module is used to recalculate the load value of the leaf node when the gap value of the load value of the computing node is greater than the preset gap value, and reallocate the leaf node to the computing node according to the load value of the leaf node, so that the load value of the computing node is rebalanced.
[0034] Based on the above embodiment, the third processing module is specifically used to traverse from the root node of the octree structure, assign a corresponding code to each node, recursively access its child nodes until a leaf node is reached, and assign a code to the leaf node;
[0035] The codes of the nodes in the path from the root node to the leaf node are expressed as a code sequence, and the code sequence is the Z-curve code of the leaf node;
[0036] Traversing the octree structure by a depth-first search method to obtain Z-curve codes of all leaf nodes in the octree structure;
[0037] The Z-curve codes of the leaf nodes are sorted using a lexicographic order to obtain a one-dimensional leaf node queue, where the leaf node queue is the space-filling curve of the calculation region.
[0038] Based on the above embodiment, the third processing module is specifically used to determine the load value of each of the minimum areas according to the computing tasks carried by the minimum area in the sub-area, and the load value of the minimum area is the load value of the corresponding leaf node;
[0039] According to the filling order of the space filling curve, the leaf nodes corresponding to the current filling area are sequentially allocated to the computing nodes until the load of the computing node reaches a preset load threshold;
[0040] Continue to assign the leaf node corresponding to the next filling area to another computing node until all filling areas are assigned.
[0041] Based on the above embodiment, the fourth processing module is specifically used to compare whether the difference between the maximum load value and the minimum load value in the calculation node is greater than the preset difference value according to the preset time interval during the simulation process;
[0042] If so, determining whether the load value of the leaf node in the octree structure is greater than a preset load value;
[0043] If so, the minimum area corresponding to the leaf node with a load value greater than the preset load value is divided until the load value of the minimum area is less than the preset load value, a new octree structure is obtained, the space filling curve is updated, and the load value of the leaf node in the new octree structure is recalculated;
[0044] Otherwise, the load values of the leaf nodes in the octree structure are recalculated.
[0045] Based on the above embodiment, a fifth processing module is further included, which is used to quantify the complexity of different computing tasks involved in the computing area, wherein the computing tasks include one or more of a material point method computing task, a coupled material point finite difference method computing task, and a coupled material point finite difference finite element method computing task;
[0046] The load values of the sub-areas in the calculation area and the load values of the leaf nodes are determined according to the number of material points in the material point method calculation task, the number of fluid grids in the coupled material point finite difference method calculation task and the number of rod units in the coupled material point finite difference finite element method calculation task.
[0047] Based on the above embodiment, the third processing module is further used to set a temporary buffer between the computing nodes, and the temporary buffer is used to store the data to be transmitted of the computing nodes.
[0048] Based on the above embodiment, the third processing module is also used for the computing nodes to communicate in a non-blocking manner.
[0049] According to a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the octree-based computing node load balancing method as described in the above technical solution is implemented.
[0050] According to a fourth aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the octree-based computing node load balancing method as described in the above technical solution is implemented.
[0051] In the above system and method provided in the embodiments of this specification, based on the calculation load evaluation results, the processes with unbalanced loads are divided into octree grids, mainly between the processes of adjacent background grids, to optimize the load distribution of each computing node, and facilitate the data transmission of adjacent sub-areas during the calculation process. By using buffer arrays and non-blocking communication, communication can be completed during the execution of computing tasks without the need for special waiting time, and the amount of communication is effectively reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1It is a schematic diagram of a computing node load balancing method based on an octree provided by the present invention;
[0054] Figure 2 It is a schematic diagram of octree grid division provided by the present invention;
[0055] Figure 3 It is a flow chart of a computing node load balancing method based on octree provided by the present invention;
[0056] Figure 4 It is a schematic diagram of an octree-based computing node load balancing device provided by the present invention. DETAILED DESCRIPTION
[0057] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0058] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0059] In the description of the embodiments of the present application, words such as "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a concrete way.
[0060] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more.
[0061] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprises", "has" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0062] The material point method (MPM) discretizes material regions through particles, and each particle carries all material information, such as mass, position, momentum, etc. The material point method has become an effective method for solving strong nonlinear problems such as hypervelocity collision, explosion, and impact penetration. In hypervelocity collision, the material point method can accurately simulate the deformation, fragmentation, and sputtering of objects during hypervelocity collision. In explosion simulation, the material point method can capture complex physical phenomena such as shock waves and fragment scattering generated by the explosion. In impact penetration, the material point method can simulate the dynamic response of materials under impact loads, such as the process of projectiles penetrating a target plate. In dynamic fracture simulation, the material point method can capture the processes of crack initiation, expansion, and penetration. Based on the material point method, the coupled material point finite difference method has been developed for fluid-solid coupling processes with strong nonlinearity, and the coupled material point finite difference finite element method has been developed for small deformation problems. The above method can achieve load balancing of the computational tasks of each process by reasonably dividing the subspace regions in the initialization stage of the simulation. According to the computational tasks of each subspace region, each process loads the computational tasks of multiple subspace regions. However, as the simulation process proceeds, the material points frequently move at high speed in each subspace region, and the computational tasks of each subspace region change, resulting in an unbalanced load of the computational tasks of each process. Similarly, the fluid grid that affects the computational complexity of the finite difference method for coupled material points and the rod unit that affects the computational complexity of the finite difference finite element method for coupled material points may also move in each subspace region, causing the computational tasks of each subspace region to change, resulting in an unbalanced load of the computational tasks of each process. When it is necessary to adjust the load of the computational tasks of each process, it may be necessary to redivide the entire space region, resulting in large computational and communication overheads.
[0063] In view of this, the present application proposes a load balancing method for computing tasks based on octrees, which can solve the technical problem that the frequent high-speed movement of material points, the change of rod units and the movement of fluid grids during the simulation process will lead to unbalanced loads of computing tasks executed between processes. Figure 1 As shown, the process of load balancing of computing tasks based on octree may include the following four stages / steps.
[0064] Step 1, computational load evaluation, specifically includes: quantifying the complexity of different computational tasks of the material point method, coupled material point finite difference method and coupled material point finite difference finite element method. For the material point method, the number of material points directly affects the computational complexity. The more material points a subspace region carries, the greater the computational load of the subspace region. For the coupled material point finite difference method, the more fluid grids there are in the subspace region, the greater the computational load of the subspace region. For the coupled material point finite difference finite element method (MPM-FDM-FEM), the more rod elements there are in the subspace region, the greater the computational load of the subspace region. The total load of each subspace region is obtained by comprehensively considering the number of material points, the number of fluid grids and the number of rod elements in each subspace region. In this step, a computational load evaluation method is constructed to achieve accurate evaluation of the computational load, laying a good foundation for load balancing division.
[0065] Step 2, octree grid division, specifically includes: constructing an octree structure, dividing the three-dimensional space involved in the event simulation into multiple sub-regions, and recursively subdividing these sub-regions until the smallest division unit, the leaf node, is reached, that is, each leaf node corresponds to a subspace region, and the load of each leaf node includes the number of material points, the number of fluid grids, and the number of rod units in the subspace region. Starting from the root node of the octree structure, traversing, for each node, recursively visit all its child nodes, and when reaching a leaf node, record the load of the leaf node. Assign a unique code to each node in the octree structure, and the path from the root node to the leaf node can be expressed as a coding sequence, which is the Z curve code of each leaf node. Through DFS traversal, collect the Z curve codes of all leaf nodes of the octree structure, and use the dictionary order to sort the Z curve codes of the leaf nodes to form a one-dimensional leaf node queue, namely the space filling curve. The space filling curve maintains the proximity relationship between adjacent sub-regions, which is conducive to the data transmission of adjacent sub-regions during the material point method calculation process, and realizes the load balancing adjustment of adjacent regions.
[0066] Calculate the total load of each leaf node: total load = material point workload + rod unit workload + grid workload. According to the load of the leaf node, the leaf node is assigned to different computing nodes according to the order in the leaf node queue to ensure that the load of each computing node is roughly equal. In this way, the load can be evenly distributed while maintaining the spatial proximity of the leaf nodes and reducing communication overhead. When the material points move frequently and at high speed, the rod units change, and the fluid grid moves, the load of the leaf node changes. The load of each leaf node is recalculated, and according to the load of each leaf node, the leaf node is reallocated to the computing node according to the leaf node queue until the load of each computing node is close to the average load. In this step, based on the constructed computing load evaluation method, the computing load of each process is evaluated. For processes with unbalanced loads, the octree grid is re-divided to give full play to the characteristics of the octree grid management, mainly re-dividing the grid between processes of adjacent background grids.
[0067] Step 3, grid data transmission. In order to avoid frequent transmission of data dependencies between computing nodes, a temporary buffer array is opened to package the data to be transmitted. For example, if there are 1,000 dependent data between process 0 and process 1, the mass and momentum of these 1,000 data are packaged and then transmitted to process 1. Furthermore, in a non-blocking manner, when process 2 transmits data to process 3, process 0 transmits data to process 1 at the same time, and they do not affect each other. At the same time, non-blocking communication can also be used so that communication can be completed during the execution of computing tasks without the need for special waiting time. For example, according to the calculation process of material points, basic material point physical quantities, such as the velocity and mass of material points, are transmitted, and other physical quantities can be calculated through basic physical quantities, effectively reducing the amount of communication.
[0068] Step 4, simulation, specifically includes: in the simulation process, through the threshold setting method, after a specific number of iterations, it is determined whether the load needs to be adjusted according to the difference in the computing load between processes. If the load needs to be adjusted, the octree division is continued. If the minimum division is achieved, the load between the computing nodes is directly adjusted according to the computing load. When it is determined that the octree reaches a critical point, the computing load is evenly distributed to other computing nodes to achieve load balancing.
[0069] Through the above method, when the computing tasks of each subspace area change, resulting in an unbalanced load of the computing tasks of each process, and the load of the computing tasks of each process needs to be adjusted, an evaluation method is constructed according to the load characteristics of different computing methods to provide a basis for subsequent load balancing division. Based on the computing load evaluation results, the octree grid is re-divided for the process with unbalanced load, mainly between the processes of adjacent background grids, to optimize the load distribution of each computing node, and facilitate the data transmission of adjacent sub-areas during the calculation process.
[0070] It should be noted that in a high-performance computing (HPC) cluster or supercomputer, a computing node usually refers to an independent server or computer unit in the cluster. In parallel computing, a computing node can refer to a logical unit or a thread. In network computing or distributed systems, a computing node can refer to a device or service in a network. In load balancing, a computing node is the basic unit of load distribution, responsible for processing local tasks or forwarding tasks to other nodes to achieve distributed computing and load balancing.
[0071] Below Figure 2This is a specific example to illustrate how to use space filling curves and octree partitioning to achieve load balancing of computing nodes. This example specifically describes the entire process from space partitioning to load distribution between processes. First, the space involved in the calculation is divided into four first-level sub-areas, namely I, II, III and IV, and the load of each first-level sub-area is calculated respectively. The load of area I is 3, the load of area II is 4, the load of area III is 4, and the load of area IV is 8. Visit each area in turn to ensure that adjacent areas are also adjacent in the encoding, and you can get the space filling curve of the first division. For the first-level sub-area with a load greater than 4, continue to divide it twice, divide the first-level sub-area I into four second-level sub-areas, and calculate the load of the second-level sub-areas respectively. The results are 1, 0, 2, and 1 respectively. Divide the first-level sub-area IV into four second-level sub-areas, and the load of each second-level sub-area is 4, 0, 0, 0, and the load of the second-level sub-area I in the first-level sub-area IV is 4. It is divided three times into four third-level sub-areas, and the load of the third-level sub-areas is 1, 1, 1, 1. Finally, the subspace area divided three times is obtained. By encoding the leaf nodes, we can get 11,12,13,14,2,311,312,313,314,32,33,34,4. The corresponding loads of each leaf node are 1,0,2,1,3,1,1,1,1,0,0,0,2. Sorting the leaf nodes in lexicographic order, we get the space filling curve: 11->12->13->14->2->311->312->313->314->32->33->34->4. According to the load of each process not exceeding 4, we can get leaf nodes 11,12,13,14 are allocated to process 0, which is represented by red. If we continue to allocate, leaf node 2 is allocated to process 1, which is represented by green. If we continue to allocate, leaf nodes 311,312,313,314,32,33,34,4 are allocated to process 2, which is represented by blue. Processes 0,1,2 roughly achieve load balancing. When the material point moves, the load of the first-level sub-area II decreases from 3 to 1, and the load of the first-level sub-area IV increases from 0 to 2. The load of each process needs to be redistributed. According to the space filling curve, leaf nodes 11, 12, 13, and 14 continue to be assigned to process 0, leaf nodes 2, 311, and 312 are assigned to process 1, and leaf nodes 313, 314, 32, 33, 34, and 4 are assigned to process 2, thus re-achieving load balancing of each process.
[0072] Combine the following Figure 3 The flowchart of the octree-based computing node load balancing method is described in detail, specifically:
[0073] 110. Obtain a calculation area, and divide the calculation area into a plurality of sub-areas, wherein the calculation area is a space involved in the calculation of the explosion event simulation process.
[0074] 120. Recursively divide the calculation area until the load value of the smallest area in the sub-area is less than the preset load value, and obtain an octree structure corresponding to the calculation area, in which the leaf node in the octree structure corresponds to the smallest area in the sub-area.
[0075] Specifically, the complexity of different computational tasks involved in the computational region is quantified, and the computational tasks include one or more of a material point method computational task, a coupled material point finite difference method computational task, and a coupled material point finite difference finite element method computational task.
[0076] The load values of the sub-areas and leaf nodes in the calculation area are determined according to the number of material points in the material point method calculation task, the number of fluid grids in the coupled material point finite difference method calculation task, and the number of rod elements in the coupled material point finite difference finite element method calculation task.
[0077] It should be noted that not every material point needs to have its weight calculated, for example, material points that are eroded or out of the boundary are not counted. Generally speaking, the weight of a background grid without material points is set to 0, but this may lead to a large number of grid transfers, affecting program efficiency. Usually the number of grids is greater than the number of material points, so in order to make the number of material points in each process more evenly distributed and the background grid does not need to be excessively transferred, a certain weight is set for the grid, but the proportion of material points will be relatively larger.
[0078] 130. According to the octree structure, determine the space filling curve of the calculation area, and distribute the leaf nodes to different calculation nodes in turn according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each calculation node is balanced.
[0079] Specifically, the steps include:
[0080] 131. Start traversing from the root node of the octree structure, assign a corresponding code to each node, recursively visit its child nodes until a leaf node is reached, and assign a code to the leaf node;
[0081] 132. The codes of the nodes in the path from the root node to the leaf node are represented as a code sequence, and the code sequence is the Z-curve code of the leaf node;
[0082] 133. Traverse the octree structure through the depth-first search method to obtain the Z-curve encoding of all leaf nodes in the octree structure;
[0083] 134. Use the lexicographical order to sort the Z-curve codes of the leaf nodes to obtain a one-dimensional leaf node queue, which is the space-filling curve of the calculation area.
[0084] 135. Determine the load value of each minimum area according to the computing task carried by the minimum area in the sub-area, and the load value of the minimum area is the load value of the corresponding leaf node;
[0085] 136. Allocate leaf nodes corresponding to the current filling area to computing nodes in sequence according to the filling order of the space filling curve until the load of the computing node reaches a preset load threshold;
[0086] 137. Continue to allocate the leaf node corresponding to the next filling area to another computing node until all filling areas are allocated.
[0087] 140. When the gap value of the computing node load value is greater than the preset gap value, the load value of the leaf node is recalculated, and the leaf node is reallocated to the computing node according to the load value of the leaf node, so that the load value of the computing node is rebalanced.
[0088] Specifically, during the simulation process, according to a preset time interval, it is compared whether the difference between the maximum load value and the minimum load value in the calculation node is greater than a preset difference value;
[0089] If so, determine whether the load value of the leaf node in the octree structure is greater than the preset load value;
[0090] If so, the minimum area corresponding to the leaf node with a load value greater than the preset value is divided until the load value of the minimum area is less than the preset load value, a new octree structure is obtained, the space filling curve is updated, and the load value of the leaf node in the new octree structure is recalculated;
[0091] According to the filling order of the updated space filling curve and the load value of the leaf node, the leaf nodes are sequentially allocated to different computing nodes so that the load value of each computing node is balanced.
[0092] Otherwise, recalculate the load value of the leaf node in the octree structure.
[0093] According to the filling order of the previous space filling curve and the load value of the leaf node after the update, the leaf nodes are sequentially allocated to different computing nodes so that the load value of each computing node is balanced.
[0094] When the area corresponding to the computing node changes, a temporary buffer is set between the computing nodes, and the temporary buffer is used to store the data to be transmitted between the computing nodes.
[0095] Furthermore, computing nodes communicate with each other in a non-blocking manner.
[0096] In the above method provided in the embodiment of this specification, based on the calculation load evaluation result, the process with unbalanced load is divided into octree grids, mainly adjusting between the processes of adjacent background grids, optimizing the load distribution of each computing node, and facilitating the data transmission of adjacent sub-areas during the calculation process. By using buffer arrays and non-blocking communication, communication can be completed during the execution of computing tasks without the need for special waiting time, and the communication volume is effectively reduced.
[0097] Combine the following Figure 4 The schematic diagram of the computing node load balancing device based on the octree is described in detail, specifically:
[0098] The first processing module is used to obtain a calculation area and divide the calculation area into a plurality of sub-areas, wherein the calculation area is a space involved in the calculation of the explosion event simulation process;
[0099] The second processing module is used to recursively divide the calculation area until the load value of the smallest area in the sub-area is less than the preset load value, and obtain an octree structure corresponding to the calculation area, in which the leaf node in the octree structure corresponds to the smallest area in the sub-area;
[0100] A third processing module is used to determine a space filling curve of the calculation area according to the octree structure, and distribute the leaf nodes to different calculation nodes in turn according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each calculation node is balanced;
[0101] The fourth processing module is used to recalculate the load value of the leaf node when the gap value of the load value of the computing node is greater than the preset gap value, and reallocate the leaf node to the computing node according to the load value of the leaf node, so that the load value of the computing node is rebalanced.
[0102] Based on the above embodiment, the third processing module is specifically used to traverse from the root node of the octree structure, assign a corresponding code to each node, recursively access its child nodes until a leaf node is reached, and assign a code to the leaf node;
[0103] The codes of the nodes in the path from the root node to the leaf node are expressed as a code sequence, and the code sequence is the Z-curve code of the leaf node;
[0104] Traversing the octree structure by a depth-first search method to obtain Z-curve codes of all leaf nodes in the octree structure;
[0105] The Z-curve codes of the leaf nodes are sorted using a lexicographic order to obtain a one-dimensional leaf node queue, where the leaf node queue is the space-filling curve of the calculation region.
[0106] Based on the above embodiment, the third processing module is specifically used to determine the load value of each of the minimum areas according to the computing tasks carried by the minimum area in the sub-area, and the load value of the minimum area is the load value of the corresponding leaf node;
[0107] According to the filling order of the space filling curve, the leaf nodes corresponding to the current filling area are sequentially allocated to the computing nodes until the load of the computing node reaches a preset load threshold;
[0108] Continue to assign the leaf node corresponding to the next filling area to another computing node until all filling areas are assigned.
[0109] Based on the above embodiment, the fourth processing module is specifically used to compare whether the difference between the maximum load value and the minimum load value in the calculation node is greater than the preset difference value according to the preset time interval during the simulation process;
[0110] If so, determining whether the load value of the leaf node in the octree structure is greater than a preset load value;
[0111] If so, the minimum area corresponding to the leaf node with a load value greater than the preset load value is divided until the load value of the minimum area is less than the preset load value, a new octree structure is obtained, the space filling curve is updated, and the load value of the leaf node in the new octree structure is recalculated;
[0112] Otherwise, the load values of the leaf nodes in the octree structure are recalculated.
[0113] Based on the above embodiment, a fifth processing module is further included, which is used to quantify the complexity of different computing tasks involved in the computing area, wherein the computing tasks include one or more of a material point method computing task, a coupled material point finite difference method computing task, and a coupled material point finite difference finite element method computing task;
[0114] The load values of the sub-areas in the calculation area and the load values of the leaf nodes are determined according to the number of material points in the material point method calculation task, the number of fluid grids in the coupled material point finite difference method calculation task and the number of rod units in the coupled material point finite difference finite element method calculation task.
[0115] Based on the above embodiment, the third processing module is further used to set a temporary buffer between the computing nodes, and the temporary buffer is used to store the data to be transmitted of the computing nodes.
[0116] Based on the above embodiment, the third processing module is also used for the computing nodes to communicate in a non-blocking manner.
[0117] In the above device provided in the embodiment of the present specification, based on the calculation load evaluation result, the process with unbalanced load is divided into octree grids, mainly adjusting between the processes of adjacent background grids, optimizing the load distribution of each computing node, and facilitating the data transmission of adjacent sub-areas during the calculation process. By using the buffer array and non-blocking communication, the communication can be completed during the execution of the computing task without the need for special waiting time, and the communication volume is effectively reduced.
[0118] According to an embodiment of another aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the computing node load balancing method based on the octree in the above technical solution is implemented.
[0119] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute a method for load balancing computing nodes based on an octree.
[0120] Those skilled in the art should be aware that in one or more of the above examples, the functions described in this application can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0121] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation methods of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the present application should be included in the scope of protection of the present application.
Claims
1. A computing node load balancing method based on octree, characterized in that: The method comprises: Obtain a calculation area, and divide the calculation area into a plurality of sub-areas, wherein the calculation area is a space involved in the calculation of the explosion event simulation process; Recursively divide the calculation area until the load values of the smallest areas in the sub-areas are all less than the preset load values, and obtain an octree structure corresponding to the calculation area, wherein the leaf nodes in the octree structure correspond to the smallest areas in the sub-areas; According to the octree structure, the space filling curve of the calculation area is determined, and the leaf nodes are allocated to different calculation nodes in turn according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each calculation node is balanced; when the gap value of the load value of the calculation node is greater than the preset gap value, the load value of the leaf node is recalculated, and according to the load value of the leaf node, the leaf node is reallocated to the calculation node, so that the load value of the calculation node is rebalanced.
2. The method according to claim 1, characterized in that: Determining the space filling curve of the calculation area according to the octree structure specifically includes: Traversing from the root node of the octree structure, assigning a corresponding code to each node, recursively accessing its child nodes until a leaf node is reached, and assigning a code to the leaf node; The codes of the nodes in the path from the root node to the leaf node are expressed as a code sequence, and the code sequence is the Z-curve code of the leaf node; Traversing the octree structure by a depth-first search method to obtain Z-curve codes of all leaf nodes in the octree structure; The Z-curve codes of the leaf nodes are sorted using a lexicographic order to obtain a one-dimensional leaf node queue, where the leaf node queue is the space-filling curve of the calculation region.
3. The method according to claim 1, characterized in that The allocating leaf nodes to different computing nodes in sequence according to the filling order of the space filling curve and the load value of the leaf node so that the load value of each computing node is balanced specifically includes: Determine the load value of each of the minimum areas according to the computing tasks carried by the minimum area in the sub-area, where the load value of the minimum area is the load value of the corresponding leaf node; According to the filling order of the space filling curve, the leaf nodes corresponding to the current filling area are sequentially allocated to the computing nodes until the load of the computing node reaches a preset load threshold; Continue to assign the leaf node corresponding to the next filling area to another computing node until all filling areas are assigned.
4. The method according to claim 1, characterized in that: When the difference value of the calculated node load value is greater than the preset difference value, recalculating the load value of the leaf node specifically includes: During the simulation process, according to a preset time interval, it is compared whether the difference between the maximum load value and the minimum load value in the calculation node is greater than the preset difference value; If so, determining whether the load value of the leaf node in the octree structure is greater than a preset load value; If so, the minimum area corresponding to the leaf node with a load value greater than the preset load value is divided until the load value of the minimum area is less than the preset load value, a new octree structure is obtained, the space filling curve is updated, and the load value of the leaf node in the new octree structure is recalculated; Otherwise, the load values of the leaf nodes in the octree structure are recalculated.
5. The method according to claim 1, characterized in that: The method further comprises: quantifying the complexity of different computational tasks involved in the computational region, the computational tasks comprising one or more of a material point method computational task, a coupled material point finite difference method computational task, and a coupled material point finite difference finite element method computational task; The load values of the sub-areas and the load values of the leaf nodes in the calculation area are determined according to the number of material points in the material point method calculation task, the number of fluid grids in the coupled material point finite difference method calculation task and the number of rod units in the coupled material point finite difference finite element method calculation task.
6. The method according to claim 1, characterized in that The method further comprises: A temporary buffer is set between the computing nodes, and the temporary buffer is used to store the data to be transmitted of the computing nodes.
7. The method according to claim 1, characterized in that The method further comprises: The computing nodes communicate with each other in a non-blocking manner.
8. A computing node load balancing device based on octree, characterized in that: The device comprises: The first processing module is used to obtain a calculation area and divide the calculation area into a plurality of sub-areas, wherein the calculation area is a space involved in the calculation of the explosion event simulation process; A second processing module is used to recursively divide the calculation area until the load values of the smallest areas in the sub-areas are all less than the preset load values, and obtain an octree structure corresponding to the calculation area, wherein the leaf nodes in the octree structure correspond to the smallest areas in the sub-areas; A third processing module is used to determine a space filling curve of the calculation area according to the octree structure, and sequentially distribute the leaf nodes to different calculation nodes according to the filling order of the space filling curve and the load value of the leaf node, so that the load value of each calculation node is balanced; The fourth processing module is used to recalculate the load value of the leaf node when the gap value of the load value of the computing node is greater than the preset gap value, and reallocate the leaf node to the computing node according to the load value of the leaf node, so that the load value of the computing node is rebalanced.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the octree-based computing node load balancing method as described in any one of claims 1-7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the octree-based computing node load balancing method as described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Distributed parallel SPH simulation method
CN106528989A
Octree grid layer-by-layer load balancing method
CN113918348A
Load-balanced large-volume three-dimensional scene LOD construction method and apparatus, and electronic device
CN115311412A
Substance point simulation parallel method based on three-dimensional adaptive division
CN116911147A
Domain decomposition using a multi-dimensional spacepartitioning tree
US20150331964A1