A parallel method for the generation of large-scale body-fitting particles
By parallelizing the sorting and deduplication steps of the stick particle generation method in computer simulation simulation software, using multi-core processing to generate large-scale stick particles, the problems of insufficient particle scale and low generation efficiency in the existing technology are solved, and high-precision and large-scale particle generation are achieved.
Patent Information
- Application Number
- CN202411235860.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-09-04
AI Technical Summary
When generating three-dimensional body-mounted particles, the existing computer simulation software is not large enough to meet the high-precision simulation requirements in complex scenarios, and is limited by the single-machine memory capacity, so large-scale particles cannot be effectively generated.
A parallel generation method for large-scale patch particles is proposed. By further parallelizing the sorting and deduplication steps in the patch particle generation method within a single process, using multiple cores to simultaneously process particle generation tasks, and achieving rapid, patching and uniform generation of large-scale particles.
It realizes a larger-scale particle generation based on ensuring the rapid, body-mounted and uniformity of particle generation, supports the generation of particles above the order of tens of millions, and meets the high-precision simulation needs in complex scenarios.
Smart Images

Figure CN119249835B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a parallel generation method of large-scale body-attached particles, belonging to the technical field of computer simulation and industrial software. Background Art
[0002] Computer simulation software, as an important part of industrial software, is a powerful tool to guide product design, development, and testing. For example, simulation software can simulate various scenarios such as tsunamis and cars wading in water, help respond to natural disasters, and design and improve the appearance of cars.
[0003] The specific technical means of implementing computer simulation software mainly include grid methods and particle methods. Among them, particle methods are an important supplement to traditional grid methods, and are especially suitable for simulation of complex problems such as large deformation and moving boundaries. The spatial distribution of particles has a great influence on the calculation accuracy and stability of particle methods. The ideal particle distribution is uniform and can accurately describe geometric and physical information. Existing particle generation methods mainly include lattice methods, grid methods, physical combination methods, etc., which have problems such as slow speed and non-fitting.
[0004] At present, the simulation method has achieved very good results for simple physical scene problems. However, ensuring high-precision simulation in complex scenes is still a very challenging problem. The intuitive idea is to improve the simulation method at a high-resolution scale. Large-scale particle generation is an important cornerstone for high-precision particle simulation in complex scenes. The existing body-fitting particle generation method of computer simulation software uses strategies such as reasonable hashing to reduce the dimensionality of the problem domain to ensure the body-fitting, uniformity, and speed of particles generated in a single process. To a certain extent, the accuracy and stability of numerical simulations are guaranteed. However, due to hardware factors such as the memory capacity of a single machine, it cannot meet the needs of large-scale particle generation. Summary of the invention
[0005] The purpose of the present invention is to solve the technical problems such as the insufficient magnitude of three-dimensional body-fitting particles generated in computer simulation software, and propose a parallel generation method for large-scale body-fitting particles in computer simulation software. Under the premise of uniform distribution and load balancing, the method further parallelizes the sorting and deduplication steps in the body-fitting particle generation method in a single process, and can further support the generation of body-fitting particles on a larger scale while ensuring the rapid, body-fitting, and uniform characteristics of the generated particles.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme.
[0007] First, the concept of the present invention is explained.
[0008] 1. Particles. They are the basic research unit in particle simulation methods. Each particle carries corresponding geometric and physical information such as position, interface normal vector, velocity, etc. In this method, only the position information of the particle is examined.
[0009] 2. Body-fitting. It is a relationship between a point and a surface, indicating that the point is on the surface. Body-fitting particles refer to particles on the surface. Body-fitting is a positional property of particles.
[0010] 3. Large scale. Describes a large number of particles, at least tens of millions.
[0011] 4. Parallelism. Compared with stand-alone programs, this program framework can use the resources of multiple computers at the same time, supporting and accelerating the calculation of large-scale problems.
[0012] 5. Lattice. It is a periodic geometric structure in three-dimensional space. Common examples include cubes, etc.
[0013] 6. Geometric configuration. The geometric body that describes the internal structure and external surface of an object is composed of points, lines, surfaces, and bodies.
[0014] 7. Mesh. It is the basic unit of industrial entity modeling and the basic calculation unit of finite element method. It is composed of nodes, edges, faces, etc., and other physical parameter information can be added.
[0015] 8.PSRS (Parallel Sorting by Regular Sampling) is a parallel sorting algorithm suitable for large-scale data.
[0016] 9. STL file (STereoLithography). It is a file format that represents the geometric shape of a three-dimensional surface. It consists of a series of triangular patches with three vertex coordinates and three-dimensional normal vector information.
[0017] 10. VTK file (Visualization Toolkit). It is a data set file that can carry a variety of information. It can save particle data with a variety of physical parameter values.
[0018] A parallel generation method for large-scale body-fitting particles includes STL triangle particle generation, lattice hashing and sorting to remove duplicates.
[0019] Step 1: Set the initial parameter values, including the initial particle spacing margin, STL file path, particle generation direction and core number.
[0020] Step 2: Analyze the STL file describing the geometric configuration to obtain the vertex coordinates and normal vector information of each triangle face.
[0021] Step 3: Check the validity of the initial particle spacing.
[0022] Step 4: Subdivide the initial triangle patch.
[0023] Specifically, it iteratively subdivides triangular patches that are too large in area into a series of small patches that can be loaded by a single core. Among them, triangular patches that are too large in area are those in which the number of particles generated at the current resolution will occupy more space than the memory space limit of a single core.
[0024] Step 5: Allocate the subdivided faces to each core evenly according to their area.
[0025] Step 6: Each core generates a single uniform particle from the assigned surface.
[0026] Step 7: Based on the PSRS parallel sorting algorithm, a parallel sorting and deduplication algorithm based on regular sampling is globally executed to obtain globally ordered and non-repetitive body-attached particles.
[0027] Step 8: Each core outputs globally ordered, uniform, and non-repetitive body-fitting particles generated based on the assigned facets.
[0028] Beneficial Effects
[0029] Compared with the prior art, the method of the present invention has the following advantages:
[0030] 1. This method ensures the load balance of parallel particle generation and the legality of parameters. The scale of particle generation in this method depends on the area of the triangles in the geometric configuration. To this end, the initial triangles are reasonably subdivided and evenly distributed to each core to ensure the load balance of particles generated by each core. In addition, combined with parameters such as the initial particle spacing, the overall particle scale can be estimated from the total area. If it exceeds the single-core upper limit, it is illegal.
[0031] 2. This method realizes the rapid generation of large-scale particles. After subdividing the triangular patch, inserting continuous sub-patterns still ensures the spatial continuity in the patch array. Adjacent cores obtain patches in array order, which further ensures spatial locality and reduces the communication overhead between cores. More importantly, on the basis of load balancing, further utilizing the resources of a large number of parallel cores can quickly generate large-scale particles. Of course, it also puts forward certain requirements for large-scale particle file management and post-processing rendering.
[0032] 3. This method ensures that large-scale particles are generated in a close and uniform manner. When a single core runs the stand-alone version of the particle generation method, uniform particles in the single core are obtained. After sorting and removing duplicates between cores, the global multi-core ensures the close, uniform, and non-repetitive properties of the stand-alone version. Therefore, large-scale uniform particles are obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the process of the present invention.
[0034] Figure 2 It is a schematic diagram of subdividing triangular facets in the present invention.
[0035] Figure 3 Schematic diagram of the parallel sorting and deduplication algorithm based on regular sampling in the present invention. DETAILED DESCRIPTION
[0036] The present invention is further described in detail below with reference to the accompanying drawings.
[0037] like Figure 1 As shown, the present invention first parses the initial geometry file. After checking that the parameters are legal, iteratively subdivides the oversized triangular facets and then evenly distributes them to each core, thereby generating independent uniform body-fitting particles inside each core, including spatial hashing, sorting, merging and deduplication. After that, the overall global parallel sorting and deduplication are performed. After executing several parallel communications in the order of regular sampling, sample sorting, selecting the main element for division, dividing by the main element, global exchange in order, sorting and deduplication, global uniform body-fitting particles can be obtained.
[0038] Specifically, the steps include:
[0039] Step 1: Prepare an STL file that describes the geometric configuration. Set parameter values, including the initial particle spacing margin, STL file path, particle generation direction, and core number size.
[0040] Step 2: Parse the STL file describing the geometric configuration to obtain the vertex coordinates and normal vector information of each triangle. Store the three vertices of each triangle in the verts array in sequence.
[0041] Step 3: Check the validity of the initial parameters.
[0042] Specifically, step 3 includes the following steps:
[0043] Step 3.1: Estimate the upper limit Smax of the area that a single core can carry:
[0044] Smax=(Mmax / sizeofParti)*margin^2
[0045] Among them, Mmax is the maximum memory space of a single core, and sizeofParti is the memory space occupied by a single particle.
[0046] Step 3.2: Calculate the area and sum Ssum of all triangles.
[0047] The area of the triangle patch can be calculated by the cross product of the vectors on both sides.
[0048] Step 3.3: If the area allocated to a single core is larger than the upper limit of a single core, that is, Ssum>size*Smax, an overage error is reported.
[0049] Step 4: Traverse all triangles and subdivide those that are too large.
[0050] like Figure 2 As shown in the figure, if the area of a certain patch Si is larger than the recommended load limit Savmax, the patch or the sub-patch obtained by the subdivision is iteratively subdivided until the sub-patch area is less than Savmax. The sub-patch obtained by iterative subdivision is stored in the verts array, covering the storage space of the large patch that was refined in the previous step. Among them, the recommended load limit Savmax is equal to the total area of all triangular patches divided by the number of cores, that is: Si>Savmax=Ssum / size.
[0051] Specifically, step 4 includes the following steps:
[0052] Step 4.1: Traverse the triangle patch vertex array verts[]. If the patch area Si is greater than the recommended load upper limit Savmax, that is, Si>Savmax=Ssum / size, iteratively subdivide the patch or the subdivided sub-patch until the sub-patch area is less than Savmax;
[0053] Step 4.2: The vertex information of the final sub-face is inserted into the position of the original face in the verts array in order, and the vertex information of the original face is deleted.
[0054] Step 5: Evenly distribute the subdivided triangular patches to each core according to their area. Traverse the patches in the verts array and accumulate the area Stemp of the continuous patches from the beginning. When the sum of the temporary patch areas is greater than the load limit, that is, Stemp>Savmax, distribute the continuous patches before this position to the current core and record this position as the starting position of the new temporary continuous patch until all the patches are distributed.
[0055] Step 6: Each core generates a single uniform particle from the assigned surface.
[0056] By reasonably mapping and reducing the dimensionality of the problem domain, the particle generation on the three-dimensional surface is simplified to the one-dimensional linear particle generation, which greatly reduces the particle generation time. At the same time, the two-dimensional plane perspective ensures the body-fitting properties of the generated particles, and the one-dimensional particles obtained by hashing are sorted and deduplicated to ensure the uniformity of the generated particles.
[0057] Step 7: Globally execute the parallel sorting and de-duplication algorithm based on regular sampling to obtain globally ordered and non-repetitive body-fitting particles.
[0058] like Figure 3 As shown in the figure, first, each core obtains ordered, weightless, uniform and body-fitting particles. The array stores the unique hash value of each particle in space, but there may be repeated or globally disordered particles between cores, and deduplication operations are still required globally. Maintaining order globally can greatly speed up the global deduplication process. It can also guarantee the spatial locality of particles to a large extent and speed up various operations based on particle traversal. With the help of the gathering function MPI_Gather(), the core number of sampling particles is first selected from each core particle array. In the figure, the core number size = 4, and the global size cores sample size^2 regular sampling particle elements. Then, after sorting the sampling elements in a single core, continue to sample size-1 partitioning principal elements, and then broadcast the partitioning principal elements to each core with the help of the broadcast function MPI_Bcast(). Specifically, in the figure, the partitioning principal elements are 7, 26, and 64. Each core divides the ordered particle array into size segment sub-arrays according to the principal element value. With the help of the global exchange function MPI_Alltoall(), the new number of particles after the overall sorting of each core is calculated and stored in the particle number array nps[] after the overall sorting of each core. With the help of the global segment exchange function MPI_Alltoallv(), each core exchanges each segment of the particle sub-array after division to the core with the corresponding sequence number according to the segment number and the corresponding segment length array nps[]. In other words, the i-th segment of each core is moved to the i-th segment of the i-th core. Each core then sorts and removes the particles obtained by the exchange and movement to obtain globally non-repetitive and uniform body-fitting particles.
[0059] Specifically, step 7 includes the following steps:
[0060] Step 7.1: From the ordered particles in each core, sample size particles, for a total of size^2 sample particles.
[0061] The implementation of selecting sampling particles is based on the gathering function MPI_Gather() in the MPI framework.
[0062] Step 7.2: Sort the size^2 sample particle array by lattice hash value.
[0063] The specific implementation is based on the library function qsort().
[0064] Step 7.3: Select size-1 from the sample particle array as the partition particle array and broadcast it to all cores.
[0065] The specific implementation is based on the broadcast function MPI_Bcast() in the MPI framework.
[0066] Step 7.4: Each core divides the local particle array on its own core according to the received partitioned particle array into size segments, i.e., size partitioned sub-arrays. The global exchange function MPI_Alltoall() in the MPI framework calculates the new number of particles after the overall sorting of each core and stores it in the particle number array new_partition_size[] after the overall sorting of each core.
[0067] Step 7.5: Each core exchanges each segment of the particle array after division to the core with the corresponding sequence number according to the segment number and the corresponding segment length array new_partition_size[]. That is, the i-th segment of each core is moved to the i-th segment of the i-th core.
[0068] The specific implementation is based on the global segment exchange function MPI_Alltoallv() in the MPI framework.
[0069] Step 7.6: Each core will perform deduplication again on the local single machine after the overall parallel sorting and exchange of the obtained particles, and finally obtain independent global non-repetitive and uniform body-fitting particles.
[0070] Step 8: Each core outputs globally ordered, uniform, and non-repetitive body-fitting particles generated based on the assigned facets.
[0071] Specifically, each core outputs the allocated uniform body-fitting particles into a particle VTK file.
Claims
1. A method for parallel generation of large-scale body-attached particles, characterized in that: The following steps are involved: Step 1: Prepare the STL file describing the geometric configuration; Set parameter values, including initial particle spacing margin, STL file path, particle generation direction and core number size; Step 2: Analyze the STL file describing the geometric configuration to obtain the vertex coordinates and normal vector information of each triangle face; Parse the STL file describing the geometric configuration to obtain the vertex coordinates and normal vector information of each triangle; store the three vertices of each triangle in the verts array in sequence; Step 3: Check the legitimacy of the initial particle spacing; Step 4: Traverse all triangular patches and subdivide patches that are too large; if the area Si of a patch is larger than the recommended load limit Savmax, iteratively subdivide the patch or the subdivided sub-patch until the sub-patch area is smaller than Savmax; the subdivided sub-patch is stored in the verts array, covering the storage space of the large patch that was refined in the previous step; The recommended load limit Savmax is equal to the total area of all triangles divided by the number of cores, that is, Si>Savmax=Ssum / size; Step 5: Evenly distribute the subdivided faces to each core according to their area; Step 6: Each core generates a single uniform particle from the assigned surface; Step 7: Globally execute the parallel sorting and deduplication algorithm based on regular sampling to obtain globally ordered and non-repetitive body-attached particles; Step 8: Each core outputs globally ordered, uniform, and non-repetitive body-fitting particles generated based on the assigned facets.
2. A method for parallel generation of large-scale body-attached particles as claimed in claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: Estimate the upper limit Smax of the area that a single core can carry: Smax=(Mmax / sizeofParti)*margin^2 Among them, Mmax is the maximum memory space of a single core, sizeofParti is the memory space occupied by a single particle; Step 3.2: Calculate the area and sum Ssum of all triangles; Step 3.3: If the area allocated to a single core is larger than the upper limit of a single core, that is, Ssum>size*Smax, an overage error is reported.
3. A method for parallel generation of large-scale body-attached particles as claimed in claim 2, characterized in that: The area of a triangle is calculated by the cross product of the vectors on both sides.
4. The method for parallel generation of large-scale body-attached particles according to claim 1, characterized in that: In step 5, the patches in the verts array are traversed, and the area Stemp of the continuous patches is accumulated from the beginning. When the sum of the temporary patch areas is greater than the load limit, that is, Stemp>Savmax, the continuous patches before this position are allocated to the current core, and this position is recorded as the starting position of the new temporary continuous patches until all the patches are allocated.
5. The method for parallel generation of large-scale body-attached particles according to claim 1, characterized in that: Step 7 includes: Step 7.1: From the ordered particles in each core, sample size particles, a total of size^2 sample particles; Step 7.2: Sort the size^2 sample particle array by lattice hash value; Step 7.3: Select size-1 from the sample particle array as the partition particle array and broadcast it to all cores; Step 7.4: Each core divides the local particle array on its own core according to the received partitioned particle array into size segments, i.e., size sub-arrays after partitioning; the global exchange function MPI_Alltoall() in the MPI framework calculates the new number of particles after the overall sorting of each core, and stores it in the particle number array new_partition_size[] after the overall sorting of each core; Step 7.5: Each core exchanges each segment of the particle array after division to the core with the corresponding sequence number according to the segment number and the corresponding segment length array new_partition_size[]. That is, the i-th segment of each core is moved to the i-th segment of the i-th core; Step 7.6: Each core will perform deduplication again on the local single machine after the overall parallel sorting and exchange of the obtained particles, and finally obtain independent global non-repetitive and uniform body-fitting particles.
Citation Information
Patent Citations
Parallelization acceleration method for a molecular dynamic simulation model
CN109871553A
Parallel implementation method of particle mesh method on ARMv8 processor
CN110275732A