Multi-level grid re-division method for Elmer file type
Through the two-level parallel mesh re-division method based on Hiber space fill curve, the problems of load imbalance and communication optimization in parallel mesh generation are solved, efficient mesh re-division and parallel solution are realized, and computing efficiency and simulation accuracy are improved.
Patent Information
- Application Number
- CN202510350966.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to achieve load balancing and communication optimization in parallel grid generation, resulting in inefficient parallel solution, and traditional graph-based partitioning technology is time-consuming and poor in effect.
A two-level parallel large-scale mesh re-division method based on Hiber space fill curve is adopted. The grid files in Elmer format are processed separately through surface threads, body threads, point threads and shared point threads. Synchronous processing is achieved using signal mechanisms, and partitioning is performed based on the centroid spatial distribution of mesh units, ignoring topological interactions, and improving parallel performance.
It realizes efficient grid re-division, shortens grid division time, improves parallel solution performance, and significantly improves computing efficiency and simulation accuracy. Especially in the simulation of complex aircraft appearance flow field simulation, the calculation time is shortened by 60%, and the simulation result error is reduced by 30%.
Smart Images

Figure CN120335991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large-scale grid remeshing, and specifically to a two-level parallel large-scale grid remeshing method. Background Art
[0002] Parallel grid generation aims to divide a complex computational domain into multiple sub-domains and generate grid cells synchronously or asynchronously on multiple processors to accelerate the overall grid construction process. The subsequent parallel solution process depends on the previously generated distributed grid, distributes the computational tasks to the sub-blocks corresponding to each processor, and obtains the final result through collaborative operations. In this process, load balancing is a key factor to ensure the overall efficiency of the system. Ideally, parallel solution requires that each sub-block of the input distributed grid has the same grid scale and the minimum number of shared nodes between sub-blocks. In this way, both "load balancing" of each processor can be achieved, avoiding some processors becoming bottlenecks due to excessive tasks, and "minimizing communication" can be achieved, reducing unnecessary data transmission overhead between processors, thereby ensuring high-efficiency solution.
[0003] However, in actual operation, the distributed grid obtained by parallel grid generation generally cannot directly meet the strict load balancing conditions of parallel solution. This is mainly because there are great challenges in accurately estimating the load of grid generation tasks. When attempting to approximate the sub-task load with the finally generated grid cells N in a sub-region, it will be found that the value of grid cells N is not determined by a single factor, but is affected by the interaction of multiple factors such as the size of the sub-region, the cell size function, and the grid generation algorithm used. For example, different regional geometries, boundary condition settings, and different grid generation algorithms will result in very different numbers and distributions of the finally generated grid cells, making it almost impossible to accurately control the load in advance, and thus the generated grid is difficult to be directly used for efficient parallel solution. Moreover, the disadvantage of the traditional graph-based partitioning technology is that: before grid remeshing, it is necessary to use a parallel grid partitioner such as Parmetis to determine which processor the volume grid in this processor should be assigned to according to certain rules. Traditionally, this process is based on the graph partitioning method. Although graph partitioning has been studied for many years and developed multi-level heuristic methods, covering three stages of coarsening, partitioning, and uncoarsening, its inherent defects are still significant. Since graph partitioning itself is an NP-complete problem, in practical applications, even with existing optimization means, the parallel performance and partitioning quality are still poor. In large-scale parallel computing scenarios, the partitioning time is too long, and the divided sub-regions often cannot meet the requirements of load balancing and communication optimization, and are extremely likely to become the bottleneck of the entire simulation workflow, seriously hindering the improvement of the overall computing efficiency. Summary of the Invention
[0004] In view of this, the present application provides a two-level parallel large-scale grid remeshing method, which determines which processor the volume grid in the processor should be divided into based on the Hiber space-filling curve, avoiding the calculation and memory bottlenecks in traditional partitioning and ensuring its efficiency. The specific solution is as follows:
[0005] A two-level parallel large-scale grid remeshing method, comprising the following steps:
[0006] S1, the main process uses the Message Passing Interface to create several processes, and several processes process in parallel;
[0007] S2, after reading the source file of the data partition, several of the processes respectively create several grid file reading threads, and a signal mechanism is set between several grid file reading threads to achieve synchronous processing;
[0008] Among them, several grid file reading threads respectively include: face threads, volume threads, point threads, and shared point threads;
[0009] S3, after the grid file reading threads finish reading the grid data, the volume thread determines the target processor to which the volume grid in the process is divided based on the Hiber space-filling curve;
[0010] S4, respectively perform grid remeshing of several of the grid file reading threads, and re-divide the grid to the corresponding target processor;
[0011] S5, the source processor calls the solver to calculate the grid divided to the target processor respectively, and outputs the grid after multi-layer mapping area projection.
[0012] Preferably, the signal mechanism includes:
[0013] When the face thread encounters synchronization, it waits to be woken up by the volume thread to execute;
[0014] When the point thread encounters synchronization, it waits for the volume thread to increment the reference count of the four vertices in the tetrahedron in the point element list by 1, and then is woken up to execute;
[0015] When the volume thread encounters synchronization, it waits for the point thread to read the coordinate information of the points, and then is woken up to execute;
[0016] When the shared point thread encounters synchronization, it waits for the point thread to put the point element information received from other processors into the point element list container, and then is woken up to execute.
[0017] Preferably, the grid remeshing process of the face thread includes:
[0018] (1) After being woken up, check whether the volume grid where the face grid is located needs to be divided into other processors;
[0019] (2) Count the number of surface meshes that the current processor needs to send to and receive from other processors respectively, and correspondingly adjust the size of the list of boundary elements of the surface meshes;
[0020] (3) Fill the surface meshes to be sent into the list of boundary elements, and remove the sent surface meshes from the local surface mesh list;
[0021] (4) Send the surface meshes in the list of boundary elements to the target processor, and put the surface meshes received from other processors into the local surface mesh list;
[0022] (5) After the surface mesh re - division is completed on the current processor, write the surface mesh data in the local surface mesh list to the disk for storage, and then exit the current thread.
[0023] Preferably, the process of point - thread re - division includes:
[0024] a. After the volume thread is awakened, traverse the volume meshes in the volume element list, find the volume meshes that need to be divided to other processors. Then, the four vertices of the volume meshes that need to be divided to other processors also need to be divided to other processors. At the same time, decrement the reference count of the four vertex elements corresponding to the volume meshes that need to be divided to other processors in the point element list;
[0025] b. Save the vertex data to be sent into the first mapper map1. The key of the first mapper is the id of the target processor to which the data is to be sent, and the value is the set of global vertex element ids to be sent to the target processor;
[0026] c. According to the global point IDs recorded in the value of the first mapper, find the coordinate information of the point, and send the coordinates of the point to the processor ID saved in the key of the first mapper. Then, put the point element information received from other processors into the vertex element list, and then wake up the shared point element thread to execute;
[0027] d. Add the vertices with a reference count greater than 0 in the point element list and the point data in the vertex element list to the second mapper map2. The key of the second mapper is the global ID of the point, and the value is the coordinate information corresponding to the point. Its main function is to remove duplicates to prevent the same global point from existing in the point element list and the vertex element list, so as to write the point to the nodes file multiple times. Then, write the data in the second mapper map2 to the disk for storage.
[0028] Preferably, the process of volume - thread mesh re - division is as follows:
[0029] Step 1: After being awakened by the point thread, call the parallel grid partitioner to determine the target processor to which the volume meshes in the local processor are divided, and then wake up the surface thread;
[0030] Step 2: Traverse the list of volume elements, increment the reference count of the four vertices of the tetrahedron in the list of point elements, and then wake up the point thread;
[0031] Step 3: Count the number of volume meshes that the current processor needs to send to and receive from other processors, adjust the size of the volume element list container; fill the volume meshes to be sent into the volume element list container, and fill the volume meshes that do not need to be sent into the local volume element list;
[0032] Step 4: Send the volume meshes stored in the volume element list container to the target processor, and put the volume meshes received from other processors into the local volume element list;
[0033] Step 5: Write the data in the local volume element list to the disk for storage.
[0034] Preferably, the determination process of the parallel mesh divider includes:
[0035] 1) Calculate the minimum cuboid region in the current process, call the MPI_Allreduce function to obtain the global minimum cuboid region and the global number of volume meshes, and determine the number of volume meshes in each partition after division according to the number of partitions to be divided;
[0036] 2) Determine the number of coarse bins according to the number of processes. The number of coarse bins is the smallest integer m times greater than or equal to the number of processes and is a multiple of 8. Obtain the corresponding space-filling curve through the m value and the global minimum rectangular region;
[0037] 3) Calculate the coarse bin index corresponding to each volume mesh in the process according to the space-filling curve;
[0038] 4) Calculate the volume meshes being processed in the current stage according to the coarse bin index, and the target processor of the volume meshes being processed, and send the index information of the volume meshes being processed to the target processor;
[0039] The index information is calculated from the global minimum rectangular region and the hiber space-filling curve generated by 21 recursions;
[0040] 5) Each processor calls the MPI_Alltoall function to calculate the number of index information received from other processors in the current process, allocate memory space accordingly, and call MPI_Isend, MPI_Irecv, and MPI_Waitall to send the index information;
[0041] Among them, the process of each processor processing the index information is:
[0042] First, each processor obtains the number of index information received by other processors through the MPI_Allgather function;
[0043] Then, each processor sorts the received index information in ascending order, assigns partition numbers in the sorted order, and the current processor knows which grid of which other processor the received index information comes from, and then sends the partition number to the corresponding other processor;
[0044] Finally, after receiving this information, the other processors store the partition number corresponding to the volume grid locally in the local processor and enter the next round of loop.
[0045] Preferably, the process of the shared point thread for grid re - division includes:
[0046] Step 1: After the point thread is awakened, according to the global id of each point element in the vertex element list, then the reference count of the point element in the point element list is incremented by 1. What is saved in the vertex element list are the points sent by other processors. If the current processor already has this point before division, its reference count will be incremented by 1;
[0047] Step 2: Traverse each point in the key value set of the first mapper map1. The key in map1 is the global ID of the divided - out point, and there are the following four possible situations for this point:
[0048] Situation 1: This point is not a shared point, and the reference count of this point after division is 0;
[0049] Situation 2: This point is not a shared point, and the reference count of this point after division is not 0;
[0050] Situation 3: This point is a shared point, and the reference count of this point after division is 0;
[0051] Situation 4: This point is a shared point, and the reference count of this point after division is not 0;
[0052] Step 3: The current processor determines the owner processors of shared points and non - shared points respectively;
[0053] For shared points, their owner processors are given in the shared file. The shared file will give the list of processors where the shared points are located, and the first processor ID in the processor list is the owner processor of the shared point; the owner processor of non - shared points is defaulted to the current processor;
[0054] Step 4: The current processor sends the information of the partitioned points to the owner of the points in the form of instructions. The owner processor of the points aggregates these instructions and updates the processor list of the points. If the point is a shared point after partitioning, the updated processor list information will be sent to the owner processor of the point. After receiving this information, the owner processor will update the processor list information where the shared point is located locally, and then write this information to the shared file;
[0055] The instructions are as follows:
[0056] When the reference count of a point is 0 after partitioning, the processor commands the owner processor of the point to delete the ID of the current processor and add the ID of the target processor to which the point is partitioned;
[0057] When the reference count of a point is not 0 after partitioning, the processor commands the owner processor of the point to add the ID of the target processor to which the point is partitioned.
[0058] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0059] The present invention uses a space-filling curve to replace the traditional graph-based partitioning scheme, partitions based on the spatial distribution of the centroids of grid cells, and ignores the topological interactions between grid elements. The space-filling curve has good locality characteristics, which maps a multi-dimensional space to a one-dimensional space and tries to maintain the proximity of elements in both spaces.
[0060] The present invention performs multi-level re-partitioning on grid files in Elmer format. There are four types of elements in Elmer-type grids. Compared with sequentially partitioning each element in order, starting four threads to parallelly partition the four types of grid elements can improve the parallel performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flowchart of grid re-partitioning in an embodiment of the present application;
[0062] Figure 2 is a schematic diagram of randomly partitioning the volume grid in a processor to other processors in an embodiment of the present application;
[0063] Figure 3 is a schematic diagram of the thermodynamic solution result before re-partitioning in the test analysis of the robustness of grid re-partitioning of the present application;
[0064] Figure 4 is a schematic diagram of the thermodynamic solution result after re-partitioning in the test analysis of the robustness of grid re-partitioning of the present application;
[0065] Figure 5Schematic diagram of the solution result after randomly re - partitioning the volume meshes in each processor to other processors when re - partitioning the grid for this application;
[0066] Figure 6 Schematic diagram of a grid that has been divided into four parts in an embodiment of this application. Detailed implementation manners
[0067] As Figure 1 shown, a multi - level grid re - partitioning method for Elmer file type is provided, including the following steps:
[0068] S1. The main process uses the Message Passing Interface to create several processes, and these processes process in parallel;
[0069] S2. After reading the source file of the data partition, several of these processes respectively create several grid file reading threads, and a signal mechanism is set among the several grid file reading threads to achieve synchronous processing;
[0070] Among them, the several grid file reading threads respectively include: surface thread, volume thread, point thread, and shared point thread;
[0071] S3. After the grid file reading threads finish reading the grid data, the volume thread determines the target processors divided by the volume meshes within the process based on the Hiber space - filling curve;
[0072] S4. Respectively perform grid re - partitioning of the several grid file reading threads, and re - partition the grid to the corresponding target processors;
[0073] S5. The source processor calls the solver to calculate the grids divided to the target processors respectively, and outputs the grids after multi - layer mapping area projection.
[0074] It should be noted that:
[0075] In this application, the main process uses MPI to create processes;
[0076] The surface thread, volume thread, point thread, and shared point thread respectively read the part.id.boundary file, part.id.nodes file, part.id.elements file, and part.id.shared file in the grid file, where id is the rank + 1 of the current processor in the mpi communication domain. Specifically:
[0077] The main process uses MPI to create P processes. After each process reads the part.id.header file, it creates 4 threads to read the Elmer grid file data respectively. Specifically:
[0078] Face thread, read the boundary file into the list vector <boundaryelement>In boundelems, in the BoundaryElement structure, there is an array variable verts with a length of 3 representing the global IDs of the three points of the surface mesh, and an integer variable parentelement representing the global ID of the volume mesh where the surface mesh is located. Then, the mesh remeshing process is entered.
[0079] Point thread, read the nodes file into the list vector <vertex>In verts, there is an array variable coor with a length of 3 in the Vertex structure to represent the x, y, and z coordinates of the point, an integer variable idx to represent the global id of the point, and an integer variable useCount to represent the reference count of the point. The reference count is initialized to 0. At the same time, <int, int> vertmap stores the mapping from the global id of the point to the local id, and then enters the mesh re - refinement process.
[0080] For the body thread, read the elements file into the list vector <volumeelement>In volelems, in the VolumeElement structure, there is an array variable verts with a length of 4 representing the global IDs of the four points of the volume mesh, and an integer variable idx representing the global ID of the volume mesh. Then, the mesh re - division process is entered.
[0081] The shared - point thread reads the shared file into the list map<int, string> sharedVert. Here, the key represents the global ID of the point, and the value is a string representing the list information of the processors where the point is located. This information includes the total number of processors where the point is located and the IDs of each processor. Then, the mesh re - division process is entered.
[0082] As an open - source finite - element analysis software, Elmer has wide applications in multi - physics simulations such as structural mechanics, electromagnetics, and fluid dynamics. It supports parallel solving and can export the solution results as VTK files. It has an active community and ecosystem. The file formats saved by each process are shown in Table 1.
[0083] Table 1 File Formats Saved by Each Process in Elmer
[0084]
[0085]
[0086] In Table 1, the ID is the number of the current process. Since the Elmer distributed - mesh file mainly consists of four types of elements: face - mesh elements, volume - mesh elements, point elements, and shared - point elements. Compared with sequentially dividing each element in order, in this application, four threads are opened in the current process, namely the face thread, volume thread, point thread, and shared - point thread, to simultaneously divide these four types of mesh elements, which can shorten the mesh - division time. Especially when writing the data of the divided mesh file to the disk. For example, after the face - mesh elements are divided, the face - mesh elements can be immediately written to the disk in the current thread, without having to wait until the point elements, volume elements, and shared - point elements are divided and then write these mesh elements to the disk in sequence, thereby improving its parallel performance.
[0087] Furthermore, the signal mechanism includes:
[0088] When the face thread encounters synchronization, it waits to be woken up and executed after the volume thread finishes executing the parallel mesh partitioner.
[0089] When the point thread encounters synchronization, it waits for the volume thread to traverse the volume - element list (volelems), increment the reference count of the four vertices of the tetrahedron in the point - element list verts, and then be woken up and executed.
[0090] When the in-body thread encounters synchronization, it waits for the point thread to read the coordinate information of the point and then be awakened for execution;
[0091] When the shared point thread encounters synchronization, the waiting point thread puts the point element information received from other processors into the vertex element list container vector <vertex>In migVertex, and then is awakened for execution.
[0092] Further, the mesh re - division process of the face thread includes:
[0093] (1) After being awakened, check whether the volume mesh where the edge - face mesh is located needs to be divided to other processors;
[0094] (2) Statistically calculate the number of face meshes that the current processor needs to send to and receive from other processors respectively, and correspondingly adjust the size of the boundary element list (migBoundary);
[0095] (3) Fill the face meshes to be sent into the boundary element list (migBoundary), and remove the sent face meshes from the local face mesh list (boundelems);
[0096] (4) Send the face meshes in the boundary element list (migBoundary) to the target processor, and put the face meshes received from other processors into the local face mesh list (boundelems);
[0097] (5) After the face mesh re - division of the current processor is completed, write the face mesh data in the local face mesh list (boundelems) to the disk for storage.
[0098] Further, the re - division process of the point thread includes:
[0099] a. After the volume thread is awakened, traverse the volume meshes in the volume element list (volelems), find out the volume meshes that need to be divided to other processors. Then, the four vertices of the volume meshes that need to be divided to other processors also need to be divided to other processors. At the same time, subtract 1 from the reference count of the four vertex elements corresponding to the volume meshes that need to be divided to other processors in the vertex element list (verts);
[0100] b. Save the vertex data to be sent to the first mapper <int, set <int>>In map1, the first mapper key is the ID of the target processor to which data is to be sent, and the value is the set (value) of the IDs of the global vertex elements to be sent to the target processor; among them, Map1[5] = {2, 3, 4} means that the points with global IDs 2, 3, and 4 are sent to processor 5. In map2, the key is the global ID of the point, and the value is the x, y, and z coordinate information of the point;
[0101] c. Place the vertex element information received from other processors into the vertex element list (vector <vertex>in migVertex), and then wake up the shared point element thread for execution;
[0102] d. Add the vertices with a reference count greater than 0 in the point element list (verts) and the vertex element list (migVertex) to the second mapper <int, Vertex> map2. The key of the second mapper is the global ID of the point, and the value is the coordinate information corresponding to the point. Its main function is to remove duplicates to prevent the point element list and the vertex element list from having the same global point, thus writing the point to the nodes file multiple times. Then write the data in the second mapper <int, Vertex> map2 to disk for storage.
[0103] Furthermore, the process of body mesh element thread mesh re - division is as follows:
[0104] Step 1: After the point thread finishes reading the nodes file in Elmer, call the parallel grid partitioner to determine the target processor for body mesh division in the local processor, and then wake up the thread surface thread;
[0105] Step 2: Traverse the body element list (volelems), increment the reference count of the four vertices of the tetrahedron in the point element list (verts), and then wake up the point thread;
[0106] Step 3: Count the number of body meshes that the current processor needs to send to and receive from other processors, and adjust the body element list container vector <volumeelement>The size of migVolume; fill the volume mesh to be sent into the body element list container vector <volumeelement>In migVolume, fill the volume meshes that do not need to be sent into the local volume element list vector <volumeelement>in the localVolume;
[0107] Step 4: Store in the body element list container vector <volumeelement>The volume meshes in migVolume are sent to the target processor, and the volume meshes received from other processors are placed in the local volume element list vector <volumeelement>in the localVolume;
[0108] Step 5: Make the local volume element list vector <volumeelement>The data in localVolum is written to the disk for storage.
[0109] Furthermore, the determination process of the parallel mesh divider includes:
[0110] 1), Calculate the minimum cuboid region (bounding box) in the current process, call the MPI_Allreduce function to obtain the global minimum cuboid (bounding box) and the global volume grid number (total_volnums), and determine the volume grid number (part_size) of each partition after division according to the number of partitions M to be divided;
[0111] Among them, the volume grid number part_size of each partition = total_volnums / M;
[0112] 2) Determine the number of coarse bins according to the number of processes. The number of coarse bins is the smallest integer multiple m of 8 that is greater than or equal to the number of processes. Through the m value and the global minimum cuboid region boundingbox, obtain the space filling curve correspondingly;
[0113] 3) Calculate the associated coarse bin index for each volume grid within the process according to the space filling curve;
[0114] 4) Calculate the volume grid being processed in the current stage according to the coarse bin index (associated_bins), and which target processor should process the volume grid being processed currently, and send the index information of the volume grid to the target processor; Among them, when calculating the index information, the hiber space filling curve generated by 21 - time recursion using the global minimum rectangular region (bounding box) is used;
[0115] 5) Call the MPI_Alltoall function to calculate the number of index information received by the current process from other processors, allocate memory space correspondingly, and call MPI_Isend, MPI_Irecv, MPI_Waitall to send the index information;
[0116] 6) Each processor obtains the number of index information received from other processors through the MPI_Allgather function;
[0117] 7) Each processor sorts the received index information from small to large, and assigns partition numbers in the sorted order. The current processor knows which other processor's which grid the received index information comes from, and then sends the partition number to the corresponding other processor.
[0118] 8) After receiving this information, other processors locally store the partition number corresponding to the volume mesh and enter the next round of the loop.
[0119] It should be noted that:
[0120] In this application, the number of coarse bins is greater than or equal to the number of processes and is an integer multiple of 8 (the smallest integer m);
[0121] Since the number of coarse bins may be greater than the number of processes num_processes, the above steps 4, 5, 6, 7, and 8 are carried out in a for loop. Each round of the loop processes num_processes coarse bin meshes, and each processor processes one coarse bin until all the coarse bins are processed and the loop ends.
[0122] Before step 4 is executed, several variables need to be defined first, namely: int64_t phase_vol, prev_phase_vol = 0: indicating the number of volume meshes processed in the current phase and the number of volume meshes processed in all previous phases respectively. These two variables will be updated at step 6 and step 8. int bin_begin = 0; int bin_end = num_processes: indicating the range of the coarse bin index subscripts processed in the current phase. After step 8 is executed, that is, after this round of the loop ends, both bin_begin and bin_end will be incremented by num_processes;
[0123] In step 4 above, the index information is calculated using the globally minimum cuboid region and the Hilbert space filling curve generated by 21 recursions because when allocating partition numbers in step 7, they are allocated in the order of the space filling curve.
[0124] After the MPI_Allgather function call in step 6 above ends, two variables need to be updated, namely the weight before the coarse grid bin corresponding to the current process: prev_bin_vol = prev_phase_vol; + the sum of the weights of the parallel processes with lower ranks in the current phase. The weight of the current phase: phase_vol = the sum of the weights of all parallel processes in the current phase;
[0125] Note that when assigning partition numbers to index information in step 7 above, there are part_size individual meshes in one partition number. The initial partition number assigned by the forward processor is init_part = prev_bin_vol / part_size, and the number of index information that can be assigned for this partition number is residual = (init_part + 1) * part_size - prev_bin_vol. If the number of index information already assigned for a certain partition number by the current processor reaches part_size, the partition number is incremented by one. Next, the current processor knows the partition number corresponding to each index information it receives, as well as which mesh of which processor the index information comes from, and then tells the source processor the partition number corresponding to the index information. After receiving this information, the source processor stores the partition number corresponding to the volume mesh locally. prev_phase_vol += phase_vol, and enter the next round of loop.
[0126] Further, the process of the shared point thread for mesh re - division includes:
[0127] Step 1: After being woken up by the point thread, obtain the global id of each point element in the vertex element list container (migVertex), find the position of each point element in the point element list (verts) through the mapper (vertmap) according to the global id of each point element, and then perform an operation of incrementing the reference count of each point element in the point element list (verts) correspondingly;
[0128] Step 2: Traverse the value value in the first mapper map1 (set <int>) Each point has the following four possible cases:
[0129] Case 1: The point is not a shared point, and its reference count is 0 after partitioning.
[0130] Case 2: The point is not a shared point, and its reference count is not 0 after partitioning.
[0131] Case 3: The point is a shared point, and its reference count is 0 after partitioning.
[0132] Case 4: The point is a shared point, and its reference count is not 0 after partitioning.
[0133] Step 3: The current processor respectively determines the owner processors of shared points and non-shared points.
[0134] For shared points, their owner processors are given in the shared file. The shared file will give the list of processors where the shared points are located. The first processor ID in the processor list is the owner processor of the shared point. The owner processor of a non-shared point is defaulted to the current processor.
[0135] Step 4: The current processor sends the information of the partitioned points to the owner of the point in the form of instructions. The owner processor of the point collects these instructions and updates the processor list of the point. If the point is a shared point after partitioning, it will send the updated processor list information to the owner processor of the point. After receiving this information, the owner processor will update the processor list information where the shared point is located locally, and then write this information to the shared file.
[0136] The instructions are as follows:
[0137] When the reference count of a point is 0 after migration and partitioning, the processor commands the owner processor of the point to delete the ID of the current processor and add the ID of the target processor to which the point is partitioned.
[0138] When the reference count of a point is not 0 after migration, the processor commands the owner processor of the point to add the ID of the target processor to which the point is partitioned.
[0139] It should be noted that:
[0140] When the mesh is re-partitioned, the update of the volume mesh and the boundary surface mesh structure is simple, but the update of the shared point element structure is more complex because internal points may become boundary shared points and boundary shared points may become internal points. Therefore, when moving a point to another processor, the processor holding the point needs to know which processor the point element is partitioned to. During the mesh partitioning stage, the owner update paradigm is used to achieve distributed notification. Before mesh re-partitioning, each point has an actual owner, and the non-actual owners of the point know this actual owner. The processor where the point is located after partitioning can be collected from the actual owner of this point, and then the non-actual owners of the point are updated after collection. For internal nodes, the processor where the point is located is implicitly the owner of the node. For the actual owner of a shared node, the first processor ID in the processor list in the shared file is the owner of the point.
[0141] In a mesh that has been partitioned into four parts. Although a three-dimensional mesh is generated in the project, a two-dimensional mesh example is used here for simplicity of presentation. The owner of the internal point c is processor 0, the owner of the shared point b is processor 0, and the owner of the vertex a is processor 2. For shared points, each processor stores a list of holders. For example, the list (2,0,1,3) of vertex a is stored on processors 0, 1, 2, and 3. The list (0,1) of vertex b is stored on processors 0 and 1. Point c does not need to store any list because it is an internal vertex. This list is a hash table map, and the key is the global id of the point. During the mesh re-partitioning process, the following several situations may occur:
[0142] Situation 1: The holders of the shared point a (processors 0 and 2) need to notify the owner processor (processor 2) of the target processor (processor 1) to which point a in this processor is partitioned. The owner processor will form a new list of holders (1,3), and then update the holders (processors 1 and 3). Details are as follows: Processor 0 will send the following instructions to processor 2 [(0,0,a),(1,1,a)], which means that processor 0 deletes point a and processor 1 adds point a. Processor 2 will send the following instructions to processor 2 [(2,0,a),(1,1,a)], which means that processor 2 deletes point a and processor 1 adds point a. Processor 2 summarizes these instructions, updates the processor list [2,0,1,3] of the original shared point a to [1,3], and then tells processors 1 and 3 the updated list [a,(1,3)]. After receiving the instructions, processors 1 and 3 update the shared point list of a locally.
[0143] Case 2: The No. 1 processor divides point b to the No. 0 processor. The No. 1 processor sends the following instructions to the actual owner of point b (the No. 0 processor) [(1,0,b),(0,1,b)]. This means that the No. 1 processor deletes point b and the No. 0 processor adds point b. After receiving the instructions, the No. 0 processor will update the processor list that originally shared point b from [0,1] to [0], and then tell the No. 0 processor the updated list [b,(0)]. After receiving the instructions, the No. 0 processor can update the shared point list of b locally. Since only one processor owns point b after the division, point b becomes an internal point.
[0144] Case 3: The No. 0 processor divides the internal point c to the No. 1 and No. 2 processors. The No. 0 processor will send the following instructions to the actual owner of point c (the No. 0 processor) [(0,0,c),(1,1,c)], and [(0,0,c),(2,1,c)]. This means that the No. 0 processor deletes point c and the No. 1 and No. 2 processors add point c. After receiving the instructions, the owner of this point will update the shared point list of point c to [1,2], and then send the instructions [c,(1,2)] to the target processors No. 1 and No. 2. After receiving the instructions, the target processors update the shared point list of local point c to [1,2].
[0145] Case 4: The No. 0 processor divides the internal point d to the No. 2 processor. Point d remains an internal point before and after the division. The instruction sending situation is as follows: The No. 0 processor sends the instruction [(0,0,d),(2,1,d)] to the actual owner of point d (the No. 0 processor), which means that the No. 0 processor deletes point d and the No. 2 processor adds point c. The actual owner updates the list of point d to [2], and then sends the instruction [d,(2)] to the No. 2 processor.
[0146] Role of reference counting: Shared points in this processor, such as point b and point e in Processor 1. Since the reference count of point b is 0 after partitioning, while the reference count of point e is not 0 because there is still a certain triangle in Processor 1 that references point e, Processor 1 will send the instruction [(1,0,b)] to the actual owner of point b, which is Processor 0, meaning that Processor 1 deletes point b, and will not send the instruction [(1,0,e)] to the actual owner of point e, meaning that Processor 1 deletes point e. After the mesh partitioning, the shared point list of point e remains (0,1), and the shared point list of point b becomes (0), turning into an internal point, which conforms to the actual situation. The processing of internal points is the same. For example, point f and point c in Processor 0. Since the reference count of point c becomes 0 after partitioning, the instruction it sends after partitioning has been introduced in the above Case 3. The reference count of point f is not 0 after partitioning, and it will send the instructions [(0,1,f),(1,1,f)] and [(0,1,f),(2,1,f)] to the actual owner of this point, which is Processor 0, meaning that Processors 0, 1, and 2 all add point f. Then Processor 0 will send the shared point list of point f to the holder processor of point f.
[0147] Since the Elmer mesh file consists of four types of elements, namely face mesh elements, volume mesh elements, point elements, and shared point elements, compared with the traditional technology of partitioning each element in sequence, in this application, four threads are opened, namely the face thread, the point thread, the volume thread, and the shared point thread, which respectively and simultaneously correspond to partitioning these four types of mesh elements to improve performance. Especially when writing the data of the Elmer - formatted mesh file after partitioning to the disk. For example, after the face mesh elements are partitioned, the face mesh elements can be immediately written to the disk in the current thread, without having to wait for the point elements, volume elements, and shared point elements to be partitioned and then write these mesh elements to the disk in sequence, because the read - write speed of the disk is much slower than that of the memory. However, when multiple threads update the same data structure, some synchronization means are needed, which brings complexity to programming. For the synchronization problems encountered among multiple threads during the mesh partitioning process, the solution mechanism provided in this application is as follows:
[0148] Synchronization situation encountered by the face thread: The face thread needs to wait for the volume thread to execute the parallel mesh partitioner. The target processor where the face mesh is partitioned is the same as the volume mesh it is in. Only after the volume thread executes the parallel mesh partitioner can the face thread know which processor the face mesh in the current processor needs to be partitioned to.
[0149] Synchronization situations encountered by the point thread: Since the reference count of points needs to be updated, where the reference count indicates how many tetrahedrons a point exists in. After both the point thread and the volume thread have read the mesh data file, the volume thread needs to increment the reference count of the four points in the tetrahedron by 1. And when the tetrahedron is partitioned to other processors, the point thread needs to decrement the reference count of the four partitioned points by 1. To correctly update the reference count of points, it is stipulated that only after the volume thread has updated the reference count of points can the point thread be awakened to execute.
[0150] Synchronization situations encountered by the volume thread: Only after the point thread has read the coordinate information of points from the disk can the volume thread call the parallel mesh partitioner, because the partitioning process is based on the centroid of the tetrahedron, and calculating the centroid of the tetrahedron requires the coordinate information of the four points of the tetrahedron.
[0151] Synchronization situations encountered by the shared point thread: The shared point thread needs to process the points sent from other processors. If the point previously existed in this processor, then the reference count of the point will be incremented. So it needs to wait until the point thread has completed the sending and receiving of point data before it can execute.
[0152] Space Filling Curve (SFC) is a continuous function used to map a multi - dimensional space to a one - dimensional space and has good locality properties, that is, it attempts to maintain the proximity relationship of elements in the two spaces. The idea of using SFC for geometric partitioning is to map the mesh elements to a one - dimensional space and then easily divide the resulting segments into equally weighted sub - segments. A significant advantage of SFC partitioning is that it can be calculated very quickly and is particularly easy to parallelize, especially compared with graph partitioning methods. There are many different definitions of SFC based on different mapping options, including the well - known Peano and Hilbert methods. Therefore, Hilbert SFC is selected in this application.
[0153] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0154] Embodiment
[0155] This embodiment provides a two - level parallel large - scale mesh re - partitioning method for the computational domain during the simulation of an aircraft shape with a complex curved surface structure to achieve its fine mesh partitioning and capture the subtle changes in the flow field. Specifically:
[0156] Build a parallel architecture
[0157] Start the main process on a high-performance computing cluster with an MPI parallel computing environment. According to the cluster resource configuration, the main process uses MPI to create (P = 32) processes, each process corresponding to a computing core in the cluster, and initially constructs a parallel architecture.
[0158] The process reads the configuration file and creates threads
[0159] Each process reads the (part.id.header) file pre-stored in the distributed file system to obtain the basic configuration information for the current simulation grid division, such as grid type, overall scale estimation, etc.
[0160] After the reading is completed, each process creates four threads: a surface thread, a volume thread, a point thread, and a shared point thread.
[0161] The four threads respectively read the corresponding grid data files
[0162] Surface thread: Read the (part.id.boundary) file, which records the surface grid information at the boundary of the simulation area. Taking the external flow field simulation of an aircraft as an example, it contains the surface grid data of key parts such as the aircraft surface boundary and the far-field boundary, and reads it into the (vectorboundelems) list. The (verts) array of the (BoundaryElement) structure stores the global (id) of the three points of the surface grid, which is used to locate the boundary points; the (parentelement) variable indicates the global (id) of the volume grid where the surface grid is located, providing a basis for associating volume grid and boundary grid operations.
[0163] Point thread: Read the (part.id.nodes) file, which covers all node information within the computational domain. In aircraft simulation, it reflects the spatial coordinates and attributes of the discrete points in the flow field, and reads it into the (vectorverts) list. The (coor) array of the (Vertex) structure records the (x), (y), and (z) coordinates of the point, the (idx) is used as the unique global (id) of the point, the (useCount) is initialized to 0 to track the point reference situation, and the (vertmap) saves the mapping from the global (id) of the point to the local (id), facilitating local data management within the process.
[0164] Volume thread: Read the (part.id.elements) file to obtain the volume grid data. For external flow field simulation, the volume grid constructs the basic computational unit of the flow field and reads it into the (vectorvolelems) list. The (verts) array of the (VolumeElement) structure stores the global (id) of the four points of the volume grid, and the (idx) identifies the global (id) of the volume grid, laying a foundation for subsequent volume grid re-meshing and processing.
[0165] Shared point thread: Reads the (part.id.shared) file to process shared point information. In complex flow field simulations, there may be shared points at the boundaries of regions responsible by different processors. These are read into (map<int,string>sharedVert), where the (key) represents the global (id) of the point, and the (value) lists the information of the processor list where the point is located, ensuring cross-processor coordination of shared points.
[0166] Four threads execute the mesh rezoning process respectively
[0167] Face thread rezoning
[0168] After the face thread finishes reading the (boundary) file, it blocks and waits to be woken up by the volume thread. When the volume thread determines based on the volume mesh zoning strategy that some volume meshes need to be zoned and wakes up the face thread, the face thread starts to check whether the volume meshes where the edge face meshes are located need to be zoned to other processors. For example, in aircraft simulations, strong gradient changes such as shock waves may occur in some regions of the flow field (such as the leading edge of the wing), which may cause local volume meshes to need to be refined, and the corresponding boundary meshes also need to be adjusted.
[0169] Then, it counts the number of face meshes that the current processor needs to send to and receive from other processors. Assume that the meshes in the part of the leading edge of the aircraft wing responsible by the current processor need to be refined due to flow field changes, and some boundary meshes need to be sent to adjacent processors for collaborative processing. Through accurate counting, the size of (vectormigBoundary) is adjusted.
[0170] Fill the face meshes to be sent into (migBoundary), and remove them from the local face mesh list (boundelems) to ensure local data consistency.
[0171] Send the face meshes in (migBoundary) to the target processor, and at the same time receive the face meshes sent by other processors and put them into (boundelems) to complete data interaction. Finally, when the face mesh rezoning of this processor is completed, write the face mesh data in (boundelems) to the disk to save the updated boundary mesh information.
[0172] Point thread rezoning
[0173] After the point thread wakes up the volume thread, it blocks and waits to be woken up. Once woken up, it traverses the volume meshes in (volelems). When it is found that a certain volume mesh needs to be zoned to other processors due to flow field characteristic changes (such as the vortex structure at the connection between the aircraft fuselage and the wing), the synchronization operation also zones the four points of this volume mesh to other processors, and at the same time decrements the reference count of the point element in (verts) by 1.
[0174] Classify the data of the points to be sent according to the target processor (id) and save them in (<int, set> map1).
[0175] Receive the point element information transmitted from other processors and put it into the (vector migVertex) list, and then wake up the shared point thread for execution.
[0176] Finally, judge whether the reference count of the points in (verts) is greater than 0. If it is greater than 0, add the point to (<int, Vertex> map2), then add the data in (migVertex) to (map2), and finally write the data in (map2) to the disk to ensure the accurate update and persistence of the node data.
[0177] Volume thread remeshing
[0178] The volume thread blocks and waits to be woken up by the point thread, and then calls the parallel mesh divider to determine the target processor to which the volume mesh in the local processor should be divided. In the aircraft simulation scenario, the parallel mesh divider comprehensively judges the migration direction of the volume mesh based on the physical characteristics of the flow field (such as pressure gradient, velocity change, etc.) and the preset mesh quality index. For example, in high Mach number flight simulations, the volume mesh near the shock wave needs to be more finely divided and may be divided to a processor with more abundant computing resources.
[0179] Traverse (volelems), increment the reference count of the four points of the tetrahedron in (verts) by 1, and then wake up the point thread.
[0180] Count the number of volume meshes that the current processor needs to send to and receive from other processors, adjust the sizes of (vector migVolume) and (vector localVolume), fill the volume meshes to be sent into (migVolume), and fill the volume meshes that do not need to be sent into (localVolume).
[0181] Send the volume meshes in (migVolume) to the target processor, and receive the volume meshes transmitted from other processors and put them into (localVolume).
[0182] Shared point thread remeshing process:
[0183] After thread 4 blocks and waits to be woken up by thread 2, obtain the global id of each point in migVertex, find the position of the point in verts according to vertmap, and then increment the reference count of the point in verts by 1 to ensure the accurate update of the shared point reference count.
[0184] Traverse the value (set) in map1 <int>) For each point, the point is divided into four cases:
[0185] Case 1: The point is not a shared point, and the reference count of the point after migration is 0. For example, in the simulation of the internal flow field of an aircraft, some internal nodes are only used within a single processor after mesh re - division and are no longer referenced by other processors.
[0186] Case 2: The point is not a shared point, and the reference count of the point after migration is not 0. For example, during the refinement of the internal structure mesh of the wing, some nodes, although not shared points, are still referenced multiple times within the current processor.
[0187] Case 3: The point is a shared point, and the reference count of the point after migration is 0. At the boundary between the surface of the aircraft and the external flow field, some shared points are no longer referenced on some processors after re - division.
[0188] Case 4: The point is a shared point, and the reference count of the point after migration is not 0. For example, the shared points at the connection between the fuselage and the wing are still frequently referenced among multiple processors after re - division.
[0189] For shared points, there is a designated owner processor, and other processors know this owner processor. For non - shared points, that is, internal points, the owner processor is the processor that holds them. The processor sends instructions to the owner of the point. After receiving these instructions, the owner of the point will update the processor id that owns the point, and then the owner will send the updated information to the owner. The instructions are as follows:
[0190] For cases 1 and 3, it will tell the owner of the point that this processor deletes the point, and the target processor where the point is partitioned adds the point.
[0191] For cases 2 and 4, it will tell the owner of the point that this processor adds the point, and the target processor where the point is partitioned adds the point. Through a fine - grained shared - point management mechanism, cross - processor point data consistency is ensured.
[0192] Parallel mesh partitioner operation
[0193] The parallel mesh partitioner in the body thread first calculates the bounding box of the current processor, that is, determines the minimum bounding box in three - dimensional space of the volume mesh responsible for the current processor. In aircraft simulation, different processors may be responsible for different parts of the aircraft's volume mesh. For example, one processor is responsible for the leading edge part of the wing, and its bounding box reflects the spatial range of this local area. Call the MPI_Allreduce function to obtain the global bounding box and the total number of volume meshes total_volnums for overall mesh partitioning from a global perspective.
[0194] According to the number of partitions M to be divided, determine the number of volume grids in each partition after division, part_size = total_volnums / M.
[0195] Determine the number of coarse bins (num_bins) according to the number of processes. The number of coarse bins is the smallest integer (m) multiple of 8 that is greater than or equal to the number of processes. Calculate the corresponding coarse bin index associated_bins of each volume grid within the process based on the global bounding box and the space-filling curve obtained by k, to achieve the preliminary classification of volume grids.
[0196] Since the number of coarse bins may be greater than the number of processes num_processes, the following steps are carried out in a for loop. Each round of the loop processes num_processes coarse bin grids. Each processor processes one coarse bin until all coarse bins are processed and the loop ends. Define variables: int64_t phase_vol, prev_phase_vol = 0: representing the number of volume grids processed in the current phase and the number of volume grids processed in the previous phase. int bin_begin = 0; int bin_end = num_processes: representing the range of the coarse bin index subscripts processed in the current phase. After this round of the loop, both bin_begin and bin_end will be incremented by num_processes.
[0197] Calculate the volume grids processed in the current phase according to associated_bins, and which target processor should process this volume grid. Then send the index information generated by this volume grid to the target processor. (The index information of this volume grid uses the hiber space-filling curve generated by the global bounding box and 21 recursions. Because when allocating partition numbers, they should be allocated in the order of the space-filling curve), to ensure the rationality and orderliness of volume grid division.
[0198] Call the MPI_Alltoall function to calculate the number of index information received by this process from other processors, allocate memory space, and call MPI_Isend, MPI_Irecv, and MPI_Waitall to send index information. Each processor needs to know the weight information (the number of index information received) of other processors, which can be obtained through the MPI_Allgather function to achieve efficient cross-processor index information interaction.
[0199] Calculate the weight before the coarse grid bin corresponding to the current process: prev_bin_vol = prev_phase_vol + the sum of the weights of the parallel processes with lower rankings in the current stage. The weight of the current stage: phase_vol = the sum of the weights of all parallel processes in the current stage.
[0200] The current processor assigns a partition number to the received index information. There are part_size individual grids within one partition number. The initial partition number assigned to the current process is init_part = prev_bin_vol / part_size, and the number of index information that can be assigned for this partition number is residual = (init_part + 1) * part_size - prev_bin_vol. If the number of index information already assigned for a certain partition number of the current process reaches part_size, then the partition number is incremented by one.
[0201] The current processor knows the partition number corresponding to each received index information and which grid of which processor the index information comes from, and then tells the source processor the partition number corresponding to the index information. After receiving this information, the source processor will store the partition number corresponding to this body grid locally. Then update the variable prev_phase_vol += phase_vol, enter the next round of loop, and optimize the division and management of the body grid through a fine partition number assignment and feedback mechanism.
[0202] Simulation Results and Advantage Verification
[0203] Through the implementation of the above secondary parallel grid re - division algorithm, the external flow field of the aircraft is simulated and calculated. Compared with the traditional serial grid processing method, under the same simulation accuracy requirements, the calculation time is significantly shortened. Through actual testing, for the flow field simulation of complex aircraft shapes, this algorithm shortens the grid re - division and subsequent processing time by about 60%, significantly improving the simulation efficiency.
[0204] In terms of the accuracy of the simulation results, due to the ability to finely process boundary data, shared point data, and multi - level interactions of body grids during the grid re - division process, the coincidence degree between the simulation results and the experimental data is higher. For example, when predicting the surface pressure distribution of the aircraft, the average error is reduced by about 30% compared with the traditional method, providing more reliable numerical simulation support for related engineering applications such as aircraft design.
[0205] Through this specific embodiment, the high efficiency and high - precision advantages of the patented algorithm in complex scientific computing scenarios are fully demonstrated, providing a strong basis for its popularization and application in more fields.
[0206] The following is a test analysis of the robustness of the grid re - division:
[0207] The test method is to randomly divide the volume meshes in each processor to other processors. Compared with the method of using the Hiber space-filling curve to determine the volume mesh division rule, adopting the random division strategy and performing solution verification on the mesh elements after random division can better prove the robustness of the algorithm. The boundary surface meshes of each processor after random division are as Figure 2 shown. It can be seen that there is a significant dispersion between the boundary surface meshes, and this phenomenon is inevitable. Since the target process to which the boundary surface mesh is divided is the same as the volume mesh it belongs to, and the volume mesh adopts the random division strategy, this results in the boundary surface meshes being randomly divided to different processors as well. Figure 5 shows the results after each processor completes the solution of the volume mesh. It can be clearly seen from the figure that the solution results inside each processor are uneven and discontinuous. This is because we adopted the random division strategy, which makes the volume meshes inside the processor unlikely to be continuous in most cases, and such a result is as expected. After merging the solution results of each processor, the overall result is as Figure 4 shown. For comparison, the solution result before remeshing is presented in Figure 3 . After careful comparison, it can be found that the solution results before and after remeshing are exactly the same. This phenomenon strongly verifies that the remeshing algorithm we adopted has excellent robustness, and even in the complex situation of random division of volume meshes, it can still ensure the accuracy and stability of the solution results.
[0208] Next, the performance of mesh remeshing is tested, and the test results are shown in Table 2. It can be seen that when the number of MPI processes is 64, the remeshing of a grid file with 140 million grids can be completed in less than half a minute, and when there are 128 cores, for a grid file of the 1.4 billion level, the remeshing time is only about 3 minutes. Compared with the subsequent flow field calculation which often takes several days, this little time is really negligible. Moreover, after the mesh remeshing is completed, the grid scales in each processor will tend to be consistent. This optimization effectively avoids the embarrassing situation where some processors quickly complete the calculation due to too small grid scales and then enter the idle waiting state, while some other processors become the bottleneck of the solution performance due to too large grid scales.
[0209] Table 2 Performance test table of mesh remeshing
[0210] < / int> < / int> < / volumeelement> < / volumeelement> < / volumeelement> < / volumeelement> < / volumeelement> < / volumeelement> < / vertex> < / int> < / vertex> < / volumeelement> < / vertex> < / boundaryelement>
Claims
1. A multi-level grid rezoning method for Elmer file types, characterized in that, It includes the following steps: S1. The main process creates several processes using the Message Passing Interface, and these processes process in parallel; S2. After reading the source files of the data partition, several of the processes respectively create several grid file reading threads, and a signal mechanism is set among the several grid file reading threads to achieve synchronous processing; Among them, the several grid file reading threads respectively include: surface thread, volume thread, point thread, and shared point thread; S3. After the grid file reading threads finish reading the grid data, the volume thread determines the target processors divided by the volume grids within the process based on the Hiber space filling curve; S4. Execute the grid re-partitioning of several of the grid file reading threads respectively, and re-partition the grids to the corresponding target processors; S5. The source processor calls the solver to calculate the grids divided to the target processors respectively, and outputs the grids after multi-layer mapping area projection.
2. A multi-level grid re-meshing method for Elmer file types according to claim 1, characterized in that The signal mechanism includes: When the surface thread encounters synchronization, it waits to be woken up by the volume thread to execute; When the point thread encounters synchronization, it waits for the volume thread to increment the reference count of the four vertices in the tetrahedron in the point element list, and then is woken up to execute; When the volume thread encounters synchronization, it waits for the point thread to read the coordinate information of the points, and then is woken up to execute; When the shared point thread encounters synchronization, it waits for the point thread to put the point element information received from other processors into the vertex element list, and then is woken up to execute.
3. A multi-level grid re-meshing method for Elmer file types according to claim 1, characterized in that The grid re-partitioning process of the surface thread includes: (1) After being woken up, check whether the volume grid where the surface grid is located needs to be partitioned to other processors; (2) Respectively count the number of surface grids that the current processor needs to send to and receive from other processors, and correspondingly adjust the size of the boundary element list of the surface grids; (3) Fill the surface grids to be sent into the boundary element list, and remove the sent surface grids from the local surface grid list; (4) Send the surface grids in the boundary element list to the target processors, and put the surface grids received from other processors into the local surface grid list; (5) After the current processor finishes the grid re-partitioning of the surface grids, write the surface grid data in the local surface grid list to the disk for storage, and then exit the current thread.
4. A multi-level grid re-meshing method for Elmer file types according to claim 1, characterized in that The re-partitioning process of the point thread includes: a. After being woken up by the volume thread, traverse the volume grids in the volume element list, find the volume grids that need to be partitioned to other processors, then the four vertices of the volume grids that need to be partitioned to other processors also need to be partitioned to other processors, and at the same time decrement the reference count of the four vertex elements corresponding to the volume grids that need to be partitioned to other processors in the point element list; b. Save the vertex data to be sent to the first mapper map1, where the key of the first mapper is the id of the target processor to be sent, and the value is the set of global vertex element ids to be sent to the target processor; c. According to the global point IDs recorded in the value of the first mapper, find the coordinate information of the points, and send the coordinates of the points to the processor ID saved by the key of the first mapper, put the point element information received from other processors into the vertex element list, and then wake up the shared point element thread to execute; d. Add the vertices with a reference count greater than 0 in the point element list and the point data in the vertex element list to the second mapper map2. The key of the second mapper is the global ID of the point, and the value is the coordinate information corresponding to the point. Its main function is to remove duplicates and prevent the same global point from existing in the point element list and the vertex element list, thus writing the point to the nodes file multiple times. Then, write the data in the second mapper map2 to disk for storage.
5. A multi-level grid re-meshing method for Elmer file types according to claim 1, characterized in that, The re - meshing process of the volume thread grid is as follows: Step 1: After the point thread finishes reading the data of the nodes point file, call the parallel grid partitioner to determine the target processor for volume grid division in the local processor, and then wake up the face thread; Step 2: Traverse the volume element list, increment the reference count of the four vertices of the tetrahedron in the point element list by 1, and then wake up the point thread; Step 3: Statistically calculate the number of volume grids that the current processor needs to send to and receive from other processors, and adjust the size of the volume element list container; fill the volume grids to be sent into the volume element list container, and fill the volume grids that do not need to be sent into the local volume element list; Step 4: Send the volume grids stored in the volume element list container to the target processor, and put the volume grids received from other processors into the local volume element list; Step 5: Write the data in the local volume element list to disk for storage.
6. The multi-level grid redivision method for Elmer file type according to claim 5, characterized in that, The determination process of the parallel grid partitioner includes: 1) Calculate the minimum cuboid area in the current process, call the MPI_Allreduce function to obtain the global minimum cuboid area and the global number of volume grids, and determine the number of volume grids in each partition after division according to the number of partitions to be divided; 2) Determine the number of coarse bins based on the number of processes. The number of coarse bins is the smallest integer m times greater than or equal to the number of processes and is a multiple of 8. Obtain the corresponding space - filling curve through the m value and the global minimum rectangular area; 3) Calculate the coarse bin index corresponding to each volume grid within the process according to the space - filling curve; 4) Calculate the volume grids being processed in the current stage and the target processor of the volume grids being processed according to the coarse bin index, and send the index information of the volume grids being processed to the target processor; The index information is calculated from the global minimum rectangular area and the hiber space - filling curve generated by 21 - level recursion; 5) Each processor calls the MPI_Alltoall function to calculate the number of index information received from other processors by the current process and the corresponding allocated memory space, and calls MPI_Isend, MPI_Irecv, and MPI_Waitall to send the index information; Among them, the process of the source processor processing the index information is as follows: First, each processor obtains the number of index information received from other processors through the MPI_Allgather function; Then, each processor sorts the received index information from smallest to largest, and assigns partition numbers in the sorted order. The current processor knows which other processor's which grid the received index information comes from, and then sends the partition number to the corresponding other processor; Finally, after other processors receive this information, they store the partition number corresponding to the volume mesh locally and enter the next round of loop.
7. A multi-level grid re-meshing method for Elmer file type according to claim 5, characterized in that The process of the shared point thread for mesh re - division includes: Step 1: After the point thread is awakened, according to the global id of each point element in the vertex element list, then the reference count of the point element is incremented by 1 in the point element list correspondingly; what is saved in the vertex element list are the points sent by other processors. If the current processor already has this point before division, its reference count will be incremented by 1. Step 2: Traverse each point in the key - value set of the first mapper map1. The value in map1 is the global ID of the points that have been divided. This point has the following four possible situations: Situation 1: This point is not a shared point, and the reference count of this point after division is 0. Situation 2: This point is not a shared point, and the reference count of this point after division is not 0. Situation 3: This point is a shared point, and the reference count of this point after division is 0. Situation 4: This point is a shared point, and the reference count of this point after division is not 0. Step 3: The current processor determines the owner processors of shared points and non - shared points respectively. For shared points, their owner processors are given in the shared file. The shared file will give the list of processors where the shared points are located. The first processor ID in the processor list is the owner processor of the shared point; the owner processor of non - shared points is defaulted to the current processor. Step 4: The current processor sends the information of the divided points to the owner of this point in the form of instructions. The owner processor of this point collects these instructions and updates the processor list of this point. If this point is a shared point after division, the updated processor list information will be sent to the owner processor of this point. After receiving this information, the owner processor will update the processor list information where the shared point is located locally, and then write this information to the shared file. The instructions are: When the reference count of a certain point after division is 0, the processor commands the owner processor of this point to delete the ID of the current processor and add the ID of the target processor to which this point is divided. When the reference count of a certain point after division is not 0, the processor commands the owner processor of this point to add the ID of the target processor to which this point is divided.