Method, apparatus, and electronic device for obtaining a grid file
By dividing and re-arrangeing unstructured meshes in high-performance computing clusters, the optimized mesh files are generated, which solves the problem of low application performance utilization and improves computing speed and cache efficiency.
Patent Information
- Application Number
- CN202310284470.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-03-22
AI Technical Summary
In actual operation of existing high-performance computing clusters, the performance utilization rate of applications is low, which is far lower than the theoretical maximum performance of the cluster.
By obtaining the unstructured mesh of the target computing area, dividing it into multiple sub-parallel areas, computing the node connection degree in each sub-parallel area, and rearranging the node numbers, generating a grid file to optimize the process of the processor reading data.
It effectively improves the cache hit ratio, optimizes the iteration speed of sparse matrix, and improves the execution speed of CFD applications in the cluster.
Smart Images

Figure CN116303219B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of high-performance integrated computing, and in particular relates to a method, a device and an electronic device for obtaining a grid file. Background Art
[0002] Driven by the rapid expansion of computing needs in different industries and the development of commercial CPUs and high-speed interconnected network equipment, HPC (High-Performance-Computing) clusters have achieved unprecedented development in the past 20 years. By aggregating a large number of microprocessor units, high-performance computing clusters have excellent scalability and extremely high cost-effectiveness, and can achieve rapid solutions to complex problems. As an important means of scientific and technological innovation, high-performance clusters are widely used in many fields such as nuclear explosion simulation, weather forecasting, and engineering calculations, and are the strategic commanding heights of contemporary scientific and technological competition. According to Moore's Law, the computing power of high-performance cluster platforms has increased exponentially every year, but in actual operation, the actual performance of applications is worrying. According to research by NERSC (National-Energy-Research-Scientific-Computing-Center), in the Gordon Bell Award-winning application cases over the years, the proportion of peak performance of large-scale scientific applications to the theoretical performance of the operating platform has dropped from 40%-50% in the 1990s to 5%-10% now. It can be seen that even large-scale high-performance computing applications that have been heavily optimized can only use peak performance far lower than the theoretical maximum performance of HPC clusters. Therefore, improving the performance of HPC in cluster operations and making full use of cluster computing performance are increasingly becoming important issues that need to be urgently addressed in HPC technology. Summary of the invention
[0003] In view of the above problems, an embodiment of the present invention provides a method for obtaining a grid file, so as to overcome the above problems or at least partially solve the above problems.
[0004] According to a first aspect of an embodiment of the present invention, a method for obtaining a grid file is provided, the method comprising:
[0005] Acquire an unstructured grid of a target calculation area; wherein the unstructured grid includes a plurality of nodes, each node corresponds to an initial number, and each node is connected to at least one adjacent node;
[0006] Dividing the unstructured grid to obtain a plurality of sub-parallel regions; wherein the nodes contained in different sub-parallel regions have no intersection;
[0007] Obtaining the node connectivity in each sub-parallel region; wherein the node connectivity is the number of adjacent nodes connected to each node in each sub-parallel region;
[0008] Rearrange the initial numbers of each node based on the node connection degree to obtain the target numbers of each node;
[0009] Generate a grid file for the target calculation area based on multiple sub - parallel areas, and the initial numbers and target numbers corresponding to each node in each sub - parallel area;
[0010] Among them, the grid file is applied in the process of the processor reading the data stored in the nodes of the target calculation area.
[0011] Optionally, the dividing the unstructured grid to obtain multiple sub - parallel areas includes:
[0012] Obtain the node weights and node interconnection information corresponding to each of the multiple nodes; wherein, the node interconnection information is used to characterize the connection relationship between each node and other nodes;
[0013] Based on the node weights and the node interconnection information, obtain the initial calculation undirected graph file corresponding to the unstructured grid;
[0014] Divide the initial calculation undirected graph file to obtain multiple sub - parallel areas.
[0015] Optionally, the dividing the initial calculation undirected graph file to obtain multiple sub - parallel areas includes:
[0016] Determine the target partitioning parameter for partitioning the initial calculation undirected graph file;
[0017] Based on the target partitioning parameter, divide the initial calculation undirected graph file to obtain multiple sub - parallel area undirected graph files;
[0018] Based on multiple sub - parallel area undirected graph files, determine multiple sub - parallel areas.
[0019] Optionally, the determining multiple sub - parallel areas based on multiple sub - parallel area undirected graph files includes:
[0020] Obtain the node interconnection boundary weights between multiple sub - parallel area undirected graphs; wherein, the node interconnection boundary weights are used to characterize the size of the node communication volume of the nodes with interconnection relationships between each sub - parallel area and other sub - parallel areas;
[0021] Update the node interconnection boundary weights;
[0022] Based on the updated multiple sub - parallel area undirected graph files, determine multiple sub - parallel areas.
[0023] Optionally, determining the target partitioning parameter for partitioning the initial computational undirected graph file includes:
[0024] Obtaining the discrete format of the initial computational undirected graph file; wherein, the discrete format includes: finite difference, finite volume, and finite element;
[0025] When the discrete format is finite difference, determining the node partitioning parameter as the target partitioning parameter;
[0026] When the discrete format is finite volume or finite element, determining the element partitioning parameter as the target partitioning parameter.
[0027] Optionally, reordering the initial numbers of each node based on the node connectivity degree to obtain the target number of each node includes:
[0028] Obtaining the target sorting result of each node based on the magnitude of the node connectivity degree;
[0029] Reordering the initial numbers of each node based on the target sorting result to obtain the target number of each node.
[0030] Optionally, obtaining the target sorting result of each node based on the magnitude of the node connectivity degree includes:
[0031] Sequentially adding the initial numbers of each node to the target queue in ascending order of the node connectivity degree and in ascending order of the initial number until the initial numbers of all nodes are added;
[0032] Obtaining the reverse addition order of each node based on the addition order of the initial numbers of each node in the target queue;
[0033] Determining the reverse addition order as the target sorting result.
[0034] Optionally, sequentially adding the initial numbers of each node to the target queue in ascending order of the node connectivity degree and in ascending order of the initial number until the initial numbers of all nodes are added includes:
[0035] Adding the initial number of the first target node with the smallest connectivity degree to the target queue in ascending order of the node connectivity degree; wherein, one target queue is used to store the initial numbers of all nodes within a sub - parallel region;
[0036] Taking the first target node as the starting connection point, determine the second target nodes among all the nodes except the first target node, and the connection relationship between the second target nodes and the first target node;
[0037] According to the connection relationship between the second target nodes and the first target node, and the magnitude of the initial numbers of the second target nodes, add the initial numbers of each of the second target nodes to the target queue in sequence.
[0038] In a second aspect of the embodiments of the present invention, there is provided an apparatus for obtaining a grid file, the apparatus including:
[0039] A first obtaining module, configured to obtain an unstructured grid of a target calculation area; wherein, the unstructured grid includes a plurality of nodes, each node corresponds to an initial number, and each of the nodes is adjacent to at least one node;
[0040] A partitioning module, configured to partition the unstructured grid to obtain a plurality of sub-parallel regions; wherein, the nodes included in different sub-parallel regions do not cross;
[0041] A second obtaining module, configured to obtain the node connection degree in each sub-parallel region; wherein, the node connection degree is the number of adjacent nodes connected by each node in each sub-parallel region;
[0042] A rearrangement module, configured to rearrange the initial numbers of each node based on the node connection degree to obtain the target numbers of each of the nodes;
[0043] A generation module, configured to generate a grid file of the target calculation area based on the plurality of sub-parallel regions, and the initial numbers and the target numbers corresponding to each of the nodes in each sub-parallel region;
[0044] Wherein, the grid file is applied in the process of the processor reading the data stored in the nodes in the target calculation area.
[0045] In a third aspect of the embodiments of the present invention, there is provided an electronic device, the electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the method for obtaining a grid file as described in the first aspect of the embodiments of the present invention.
[0046] A method for obtaining a grid file provided by an embodiment of the present invention first obtains an unstructured grid of a target calculation area; and divides the unstructured grid to obtain a plurality of sub-parallel areas; it can realize parallel calculation of data by using the plurality of sub-parallel areas, improve the operation speed of the unstructured grid, and then obtain the node connectivity in each sub-parallel area, where the node connectivity is the number of adjacent nodes connected to each node in each sub-parallel area; then, according to the node connectivity, re-arrange the initial numbers of each node to obtain the target numbers of each node after re-arrangement, and solve the problem of discontinuous initial numbers after dividing the unstructured grid. Finally, based on the plurality of sub-parallel areas, and the initial numbers and target numbers corresponding to each node in each sub-parallel area, generate a grid file of the target calculation area, and the grid file is used for the processor to read the data stored by the nodes in the target calculation area. The grid file obtained by this method can centralize the storage distribution of data in cells or nodes in the grid in memory when the processor reads the data in the unstructured grid, effectively improve the cache hit ratio, optimize the iteration speed of the sparse matrix, and improve the execution speed of the CFD application in the cluster. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0048] Figure 1 It is a schematic diagram of the change of the mainstream CPU and memory latency in the past 20 years provided by an embodiment of the present invention;
[0049] Figure 2 It is a schematic diagram of a structured grid and an unstructured grid provided by an embodiment of the present invention;
[0050] Figure 3 It is a flowchart of the steps of a method for obtaining a grid file provided by an embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram of an initial undirected graph and a corresponding CSV format file provided by an embodiment of the present invention;
[0052] Figure 5 It is provided by an embodiment of the present invention for Figure 4 A schematic diagram of the result of dividing parallel areas;
[0053] Figure 6 It is a schematic diagram of a device for obtaining a grid file provided by an embodiment of the present invention;
[0054] Figure 7 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0055] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings in the embodiments of the present invention. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0056] In recent years, as Moore's Law has gradually failed, mainstream manufacturers such as Intel and AMD have begun to widely use multi-core architectures to provide higher computing performance. Although the CPU performance has been continuously improved, the performance of the memory access system has remained almost unchanged in recent years. In particular, the gap between metrics such as memory access latency and CPU cycle time has been increasing.
[0057] Refer to Figure 1 , Figure 1 It is a schematic diagram of the changes in the mainstream CPU and memory latency in the past 20 years provided by an embodiment of the present invention; the figure shows the changes in the CPU cycle and memory access latency metrics in the past 25 years. In the past 25 years, the number of CPU execution cycles has been continuously increasing at a rate of 2% to 2.5% per year, but the reduction rate of memory latency does not exceed 1% per year. As the CPU performance continues to improve, more and more software performance bottlenecks have shifted from the computing process to the data access process. Especially for CFD (Computational-Fluid-Dynamic) applications with extremely high requirements for computing performance, CFD computing applications usually use the current most advanced supercomputer clusters to complete the simulation of the motion processes of a series of complex fluids, and have extremely high requirements for the overall performance of the cluster.
[0058] Refer to Figure 2 , Figure 2It is a schematic diagram of a structured grid and an unstructured grid provided by an embodiment of the present invention; in CFD applications, computational grids include two types: structured grids and unstructured grids. The advantage of structured grids is that the data structure is simple, and both the grid generation speed and quality are relatively good. The disadvantage is that it is only applicable to the simulation of regular regions and cannot simulate complex geometric regions. On the contrary, the data structure of unstructured grids is complex and occupies more memory, but it can simulate complex geometric regions. Due to the advantages of unstructured grids such as being able to simulate complex geometric shapes, they are widely used in commercial or open-source CFD software based on the finite element method and the finite volume method. In actual applications, the mainstream is mainly unstructured grids. Since the storage order of grid nodes is related to the numbering during the calculation process. In the calculation process of modern server platforms, due to the discontinuous numbering of grid nodes, a large number of cache misses will occur, making the data reading and writing of the processor gradually become the main bottleneck in the calculation process of structured grids. To solve the above problems, the present invention provides a method for obtaining a grid file to solve the problem of a large number of cache misses caused by the discontinuous numbering of unstructured grid nodes.
[0059] Embodiment 1
[0060] Refer to Figure 3 , Figure 3 It is a flowchart of the steps of a method for obtaining a grid file provided by an embodiment of the present invention; as shown in the figure, the steps of this method include:
[0061] Step S301: Obtain the unstructured grid of the target calculation area; wherein, the unstructured grid includes a plurality of nodes, each node corresponds to an initial number, and each node is adjacent to at least one node.
[0062] In this embodiment, to obtain the unstructured grid of the target calculation area, there are a plurality of nodes in the unstructured grid, and each node stores the data in the target calculation area for the processor to read. Since the data stored in each node is different, each node will correspond to an initial number to distinguish the nodes storing different data. In order to be able to read the data of all nodes in the unstructured grid, each node in the unstructured grid is adjacent to at least one node. When the processor reads the unstructured grid data, it generally reads in the order of the node numbers, and then reads the data stored in all the nodes included in the unstructured grid.
[0063] Step S302: Divide the unstructured grid to obtain a plurality of sub-parallel regions; wherein, the nodes included in different sub-parallel regions do not cross.
[0064] In this embodiment, since unstructured grids are applied to the calculation of high-performance cluster CFD applications, the nodes contained in the unstructured grids and the data they store are also a very large quantity. For example, the number of grid nodes in the unstructured grid is 1000. If a user wants to obtain the node data with the initial number 995 in the unstructured grid, according to the initial node sorting, it is necessary to complete the reading of the previous 994 initial numbers before reading the node data with the initial number 995, which will consume a lot of time to read the data in the unstructured grid, thereby reducing the efficiency of the processor to read the data in the unstructured grid. Therefore, before reading the data stored in the unstructured grid, it is necessary to divide the unstructured grid to obtain multiple sub-parallel regions, and the multiple sub-parallel regions can be read simultaneously, saving the time for the processor to read the data and improving the efficiency of reading the data in the unstructured grid.
[0065] Step S303: Obtain the node connectivity of each sub-parallel region; wherein, the node connectivity is the number of adjacent nodes connected by each node in each sub-parallel region.
[0066] In this embodiment, the node connectivity of each sub-parallel region is obtained. The node connectivity is the number of adjacent nodes connected by each node in each sub-parallel region. For example, if node 3 represents that the initial number of this node is 3, and the nodes connected to node 3 are node 5, node 6, and node 7, then the connectivity of node 3 is 3.
[0067] Step S304: Based on the node connectivity, rearrange the initial numbers of each node to obtain the target numbers of each node.
[0068] In this embodiment, according to the connectivity of the nodes, the initial numbers of the nodes contained in each sub-parallel region are rearranged, and the node numbers within a sub-parallel region are rearranged to be numbers close to each other. The rearranged node numbers are determined as the target numbers of each node.
[0069] For example, taking the nodes contained in a parallel region as an example, assume that the nodes contained in this parallel region are node 100, node 4, and node 70. Among them, the node connectivities corresponding to node 100, node 4, and node 70 are 2, 1, and 1 respectively. According to the node connectivity, the initial numbers 100, 4, and 70 of node 100, node 4, and node 70 are rearranged. Then the target numbers of each node after rearrangement are 2, 3, and 1, solving the problem of discontinuous adjacent node numbers.
[0070] Step S305: Generate a grid file for the target calculation area based on the multiple sub-parallel areas and the initial numbers and target numbers corresponding to each node in each sub-parallel area; wherein, the grid file is applied in the process of the processor reading the data stored in the nodes in the target calculation area.
[0071] In this embodiment, a grid file for the target calculation area is generated according to multiple sub-parallel areas and the initial numbers and target numbers corresponding to each node in each sub-parallel area. Before reading the data stored in the unstructured grid, the grid file of the target calculation area is obtained. The grid file is used to provide data for the processor to read the nodes or cells in the target calculation area. It should be noted here that, first, read the data stored in the nodes in multiple sub-parallel areas in the order of the target numbers. After the reading is completed, according to the mapping relationship between the target numbers and the initial numbers of each node, map the data stored in the nodes read according to the target numbers to solve the problem of discontinuous grid nodes. Obtain the data stored in the unstructured grid according to the mapping relationship between the target number and the initial number, and then read the data stored in the nodes in the target calculation area.
[0072] In the process of obtaining the stored data in the unstructured grid by reading the grid file of the target calculation area, the cache hit ratio is effectively improved, the iteration speed of the sparse matrix is optimized, and the execution speed of the CFD application in the cluster is improved.
[0073] In one embodiment, the dividing the unstructured grid to obtain multiple sub-parallel areas includes: obtaining the node weights and node interconnection information corresponding to each of the multiple nodes; wherein, the node interconnection information is used to represent the connection relationship between each node and other nodes; based on the node weights and the node interconnection information, obtain an initial calculation undirected graph file corresponding to the unstructured grid; divide the initial calculation undirected graph file to obtain multiple sub-parallel areas.
[0074] In this embodiment, the data structure management software METIS and Scotch are used to divide the unstructured grid by the graph balancing method to obtain multiple sub-parallel areas. The data structure management software can only recognize files in a data format. Therefore, the unstructured grid needs to be converted into an undirected graph file that the data structure management software can recognize the information therein. First, identify the node weights in the unstructured grid and the interconnection relationship between each node. The node weight in the grid represents the probability of the node appearing. The node interconnection information represents the connection relationship between each node and other nodes. For example, taking node 3 as an example, the connection relationship of node 3 includes the number of other nodes connected to node 3, and the number of communication or calculation times between other nodes and node 3.
[0075] Exemplarily, a specific embodiment will be constructed and described below:
[0076] Referring to Figure 4 , Figure 4 which is a schematic diagram of an initial undirected graph and a corresponding CSV format file provided by an embodiment of the present invention; Figure 4 The left graph is the initial undirected graph. The numbers in the circles represent node numbers, the numbers in the squares represent the node weights corresponding to the node numbers, and the numbers on the connecting lines between two node numbers represent the number of communications between the two nodes. The connecting lines also represent the existence of a connection relationship between the two nodes; Figure 4 The right graph is the initial undirected graph and its corresponding CSV (Comma-Separated-Values) format file, which can be understood as an initial calculation undirected graph file. Extract the node weights and node interconnection information of the unstructured grid, construct the initial calculation undirected graph file, and convert the initial calculation vertex V 0 and the computational grid topology information E 0 into an undirected graph form G 0 =(V 0 , E 0 ), where the vertex V 0 represents the node number and its corresponding node weight, and the grid topology information E 0 represents the node interconnection information between each node. An initial calculation undirected graph can be constructed. When saving the initial calculation undirected graph, select to save it as a CSV format file, and an initial calculation undirected graph file can be obtained. The first row in the initial calculation undirected graph file represents the number of nodes and interconnected edges in the calculation area; starting from the second row, the (N + 1)-th row has a number corresponding to the node weight of the node, and the subsequent numbers respectively represent the interconnection situation between the node and other nodes and the corresponding edge interconnection throughput. Taking the second row as an example for explanation, the first number in the second row represents the node weight of node 1, the second number represents that the node connected to node 1 is node 5, the third number represents the interconnection communication volume between node 1 and node 5 is 1, the fourth number represents that the node connected to node 1 is node 3, the fifth number represents the interconnection communication volume between node 1 and node 3 is 2, the sixth number represents that the node connected to node 1 is node 2, and the seventh number represents the interconnection communication volume between node 1 and node 2 is 1. Similarly, the explanations for other rows refer to the digital explanations in the second row and will not be elaborated here.
[0077] In one embodiment, the dividing the initial calculation undirected graph file to obtain a plurality of the sub-parallel regions includes: determining a target division parameter for dividing the initial calculation undirected graph file; based on the target division parameter, dividing the initial calculation undirected graph file to obtain a plurality of sub-parallel region undirected graph files; and determining a plurality of the sub-parallel regions based on the plurality of sub-parallel region undirected graph files.
[0078] In this embodiment, before partitioning the initial computational undirected graph file, it is necessary to determine the target partitioning parameters. Since the processor needs to consider the load balancing situation when reading data, load balancing is to divert the access users when the number of user accesses is huge. When a client sends a request, the number of requests sent by the user will be diverted. In order to ensure load balance and improve the parallel computing performance, it is also necessary to determine the target partitioning parameters. The target partitioning parameters are partitioned according to the node weights. For example, the target partitioning parameters are two parallel sub-regions, and the node weights account for 5 / 13 and 8 / 13 respectively. Then the initial computational undirected graph file will be partitioned into two sub-parallel regions, and the node weights of the two sub-parallel regions account for 5 / 13 and 8 / 13 respectively. By partitioning according to the partitioning parameters input by the user, two sub-parallel regions partitioned according to the target partitioning parameters can be obtained. Refer to Figure 5 , Figure 5 is a schematic diagram of the result of partitioning parallel regions provided by an embodiment of the present invention; after obtaining Figure 4 , input the target partitioning parameters. The target partitioning parameters are two parallel sub-regions, and the node weights account for 5 / 13 and 8 / 13 respectively. Then Figure 4 the two sub-parallel regions shown in Figure 5 will be obtained.
[0079] In one embodiment, determining the multiple sub-parallel regions based on the multiple sub-parallel region undirected graph files includes: obtaining the node interconnection boundary weights between the multiple sub-parallel region undirected graphs; wherein, the node interconnection boundary weights are used to characterize the node traffic volume of each sub-parallel region and the interconnection relationship with other sub-parallel regions; updating the node interconnection boundary weights; and determining the multiple sub-parallel regions based on the updated multiple sub-parallel region undirected graph files.
[0080] In this embodiment, in combination with Figure 4 and Figure 5 are described. When using METIS and Scotch software to partition Figure 4 and obtain the multiple sub-parallel region undirected graph files, it is also necessary to modify the node interconnection boundary weights between the obtained multiple sub-parallel region undirected graph files. Refer to Figure 5 , two sub-parallel regions are partitioned. The nodes included in sub-parallel region one are node 1, node 2, and node 3, and the nodes included in sub-parallel region two are node 4, node 5, node 6, and node 7. Refer to Figure 4 , the interconnection traffic volume between node 3 and node 5 is 3. Since in Figure 5Among them, Node 3 and Node 5 are already in different sub-parallel regions, and there is no connection relationship between them. Therefore, the node interconnection boundary weight between Node 3 and Node 5 will be updated to 0. Similarly, other similar node interconnection boundary weights also need to be updated, which will not be elaborated here.
[0081] In one embodiment, the determining the target partitioning parameter for partitioning the initial computational undirected graph file includes: obtaining the discrete format of the initial computational undirected graph file; wherein, the discrete format includes: finite difference, finite volume, and finite element; when the discrete format is finite difference, determining the node partitioning parameter as the target partitioning parameter; when the discrete format is finite volume or finite element, determining the element partitioning parameter as the target partitioning parameter.
[0082] In this embodiment, obtaining the discrete format of the initial computational undirected graph file. Since the equation formats for constructing the initial unstructured grid are finite difference, finite volume, and finite element, there are two ways to store data in the unstructured grid: node storage and element storage. Element storage is to store non-overlapping geometric elements composed of nodes. If data is stored in the form of elements, when the processor reads data from the unstructured grid, it reads according to the element number. When partitioning the initial computational undirected graph file, it is necessary to first determine the discrete format of the initial computational undirected graph file. The discrete format of the initial computational undirected graph file corresponds to the unstructured grid, and then partition according to the storage method corresponding to the discrete format. Determine the target partitioning parameter according to the discrete format. If the discrete format is finite difference, determine the target partitioning parameter according to the node partitioning parameter. When the discrete format is finite volume or finite element, determine the element partitioning parameter as the target partitioning parameter. This embodiment is mainly used to illustrate that the method provided by the present invention for obtaining the grid file can not only be used for partitioning nodes, but also be applicable to element partitioning, which is specifically determined by the actual data storage format of the unstructured grid and will not be limited here.
[0083] In one embodiment, the reordering the initial numbers of each node based on the node connectivity to obtain the target numbers of each node includes: obtaining the target sorting result of each node based on the magnitude of the node connectivity; reordering the initial numbers of each node based on the target sorting result to obtain the target numbers of each node.
[0084] In this embodiment, taking a sub-parallel region as an example, the connection degrees of each node in this region are obtained. According to the magnitudes of the node connection degrees of each node, the target sorting result of each node in a sub-parallel region is obtained, and the initial numbers of all nodes in a sub-parallel region are rearranged according to the target sorting result to obtain the target numbers of all nodes. By this method, the numbering rearrangement of each node in multiple parallel regions is performed to obtain the target numbers of each node.
[0085] In one embodiment, obtaining the target sorting result of each of the nodes based on the magnitude of the node connection degree includes: sequentially adding the initial numbers of each of the nodes to the target queue in ascending order of the node connection degree and in ascending order of the initial number until the initial numbers of each of the nodes are added; based on the adding order of the initial numbers of each of the nodes in the target queue, obtaining the reverse adding order of each of the nodes; and determining the reverse adding order as the target sorting result.
[0086] In this embodiment, since there may be multiple nodes with the same node connection degree in a parallel region, it is also necessary to obtain the initial numbers of each node, and sequentially add the initial numbers of each of the nodes to the target queue in ascending order of the node connection degree and in ascending order of the initial number until the initial numbers of each node are added; based on the adding order of the initial numbers of each node in the target queue, obtaining the reverse adding order of each of the nodes; and determining the reverse adding order as the target sorting result. The node rearrangement can be performed by the CM algorithm for reverse rearrangement.
[0087] In one embodiment, the step of sequentially adding the initial numbers of each of the nodes to the target queue in ascending order of the node connection degree and in ascending order of the initial number until the initial numbers of each of the nodes are added includes: adding the initial number of the first target node with the smallest connection degree to the target queue in ascending order of the node connection degree; wherein, a target queue is used to store the initial numbers of all nodes in a sub-parallel region; taking the first target node as the starting connection point, determining the second target nodes among each of the nodes except the first target node, and the connection relationship between the second target nodes and the first target node; and sequentially adding the initial numbers of each of the second target nodes to the target queue according to the connection relationship between the second target nodes and the first target node and the magnitude of the initial number of the second target node.
[0088] In this embodiment, taking all the nodes in a sub-parallel region as an example, in the order of the node connection degrees from small to large, the initial number of the first target node with the smallest connection degree is added to the target queue; taking the first target node as the starting connection point, the second target nodes among all the nodes except the first target node are determined, as well as the connection relationship between the second target nodes and the first target node. The connection relationship between the second target nodes and the first target node is the relationship that the nodes in the second target nodes are directly or indirectly connected to the first target node. According to the connection relationship between the second target nodes and the first target node, and the size of the initial numbers of the second target nodes, the initial numbers of each second target node are sequentially added to the target queue.
[0089] Exemplarily, in combination with Figure 5 sub-parallel region two, a detailed description of the process of rearranging the node numbers provided by the embodiments of the present invention will be given:
[0090] The nodes included in sub-parallel region two are node 4, node 5, node 6, and node 7. The initial numbers sorted from small to large are 4, 5, 6, 7. The node connection degrees of node 4, node 5, node 6, and node 7 are obtained, and the corresponding node connection degrees are 2, 1, 3, 2 respectively. First, the node with the smallest connection degree is obtained, which is node 5, that is, the first target node. Node 5 is placed in the target queue. Then, the second target nodes that are directly or indirectly connected except node 5 are obtained, that is, node 4, node 6, and node 7. First, the node directly connected to node 5 is obtained, which is node 6. After node 5 has been added to the target queue, node 6 is added to the target queue. At the same time, node 6 is analyzed, and the nodes directly connected to node 6 are obtained, that is, node 4 and node 7. At this time, there are two nodes connected to node 6. According to the order of the initial numbers from small to large, node 4 is added to the target queue first, and then node 7 is added to the target queue. At this time, all the nodes in sub-parallel region two are included in the target queue. At this time, the order of the nodes added to the target queue is obtained, that is, node 5, node 6, node 4, node 7. The initial number order of the node addition order is 5, 6, 4, 7. After reverse sorting, the target node number order is obtained, that is, 4, 3, 2, 1. Node 4 corresponds to the target number 2, node 5 corresponds to the target number 4, node 6 corresponds to the target number 3, and node 7 corresponds to the target number 1.
[0091] For the mesh file rearranged and obtained in this way, when reading the node data in the sub-parallel region, it will be read in the order of the target numbers. During the reading process, the node data related to the first target node will be cached in the processor, the hit rate will be improved, the sparse matrix iteration speed will be optimized, and the computing performance of the CFD application in the cluster using distributed memory will be improved.
[0092] Embodiment Two
[0093] In the second aspect of the embodiments of the present invention, a device for obtaining a grid file is provided. The device includes: a first obtaining module, configured to obtain an unstructured grid of a target calculation area; wherein, the unstructured grid includes a plurality of nodes, each node corresponds to an initial number, and each node is adjacent to at least one node; a partitioning module, configured to partition the unstructured grid to obtain a plurality of sub-parallel regions; wherein, the nodes included in different sub-parallel regions do not cross; a second obtaining module, configured to obtain the node connectivity of each sub-parallel region; wherein, the node connectivity is the number of adjacent nodes connected by each node in each sub-parallel region; a rearrangement module, configured to rearrange the initial numbers of the nodes based on the node connectivity to obtain the target numbers of the nodes; a generating module, configured to generate a grid file of the target calculation area based on the plurality of sub-parallel regions, and the initial numbers and the target numbers corresponding to the nodes in each sub-parallel region; wherein, the grid file is applied in the process of a processor reading data stored in the nodes of the target calculation area.
[0094] In this embodiment, referring to Figure 6 , Figure 6 FIG. is a schematic diagram of a device for obtaining a grid file provided by an embodiment of the present invention; the device includes a first obtaining module 601, a partitioning module 602, a second obtaining module 603, a rearrangement module 604, and a generating module 605.
[0095] The first obtaining module 601 is configured to obtain an unstructured grid of a target calculation area; wherein, the unstructured grid includes a plurality of nodes, each node corresponds to an initial number, and each node is adjacent to at least one node.
[0096] The partitioning module 602 is configured to partition the unstructured grid to obtain a plurality of sub-parallel regions; wherein, the nodes included in different sub-parallel regions do not cross.
[0097] The second obtaining module 603 is configured to obtain the node connectivity of each sub-parallel region; wherein, the node connectivity is the number of adjacent nodes connected by each node in each sub-parallel region.
[0098] The rearrangement module 604 is configured to rearrange the initial numbers of the nodes based on the node connectivity to obtain the target numbers of the nodes.
[0099] A generation module 605, configured to generate a grid file of the target calculation area based on the multiple sub-parallel areas, and the initial numbers and the target numbers corresponding to the respective nodes in each of the sub-parallel areas; wherein, the grid file is applied in the process of the processor reading data stored in the nodes in the target calculation area.
[0100] The target area calculation grid file generated by this device can centralize the storage distribution of data in cells or nodes in the grid when the processor reads data in the unstructured grid, effectively improving the cache hit ratio, optimizing the iteration speed of the sparse matrix, and enhancing the execution speed of the CFD application in the cluster.
[0101] Embodiment III
[0102] In the third aspect of the embodiments of the present invention, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for obtaining a grid file as described in the first aspect of the embodiments of the present invention is implemented.
[0103] In this embodiment, refer to Figure 7 , Figure 7 is a schematic diagram of an electronic device provided by an embodiment of the present invention; as Figure 7 shown, the electronic device 100 includes: a memory 110 and a processor 120. The memory 110 and the processor 120 are communicatively connected via a bus. A computer program is stored in the memory 110, and the computer program can run on the processor 120, thereby implementing the method for obtaining a grid file as described in the first aspect of the embodiments of the present application.
[0104] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0105] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, devices, and electronic devices according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0106] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operational steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.
[0108] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0109] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0110] The above has introduced in detail a method, apparatus and electronic device for obtaining a grid file provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for obtaining a mesh file, characterized in that, the method includes: Obtain an unstructured mesh of a target calculation area; wherein, the unstructured mesh includes a plurality of nodes, each node corresponds to an initial number, and each of the nodes is adjacent to at least one node; Divide the unstructured mesh to obtain a plurality of sub-parallel regions; wherein, the nodes included in different sub-parallel regions do not cross; Obtain the node connectivity of each sub-parallel region; wherein, the node connectivity is the number of adjacent nodes connected by each node in each sub-parallel region; Based on the node connectivity, reorder the initial numbers of each node to obtain the target numbers of each node; Based on a plurality of the sub-parallel regions, and the initial numbers and the target numbers corresponding to each node in each sub-parallel region, generate a mesh file of the target calculation area; wherein, the mesh file is applied in the process of a processor reading data stored in nodes in the target calculation area; wherein, the dividing the unstructured mesh to obtain a plurality of sub-parallel regions includes: Obtain the node weights and node interconnection information corresponding to each of the plurality of nodes; wherein, the node interconnection information is used to characterize the connection relationship between each node and other nodes; Based on the node weights and the node interconnection information, obtain an initial calculation undirected graph file corresponding to the unstructured mesh; Divide the initial calculation undirected graph file to obtain a plurality of the sub-parallel regions; The dividing the initial calculation undirected graph file to obtain a plurality of the sub-parallel regions includes: Determine a target partitioning parameter for partitioning the initial calculation undirected graph file; Based on the target partitioning parameter, divide the initial calculation undirected graph file to obtain a plurality of sub-parallel region undirected graph files; Based on a plurality of the sub-parallel region undirected graph files, determine a plurality of the sub-parallel regions.
2. The method according to claim 1, characterized in that, the determining a plurality of the sub-parallel regions based on a plurality of the sub-parallel region undirected graph files includes: Obtain the node interconnection boundary weights between a plurality of the sub-parallel region undirected graphs; wherein, the node interconnection boundary weights are used to characterize the node traffic volume size of the interconnection relationship between each sub-parallel region and other sub-parallel regions; Update the node interconnection boundary weights; Based on the updated plurality of the sub-parallel region undirected graph files, determine a plurality of the sub-parallel regions.
3. The method according to claim 1, characterized in that, the determining a target partitioning parameter for partitioning the initial calculation undirected graph file includes: Obtain the discrete format of the initial calculation undirected graph file; wherein, the discrete format includes: finite difference, finite volume and finite element; When the discrete format is finite difference, determine the node partitioning parameter as the target partitioning parameter; When the discrete format is finite volume or finite element, determine the element partitioning parameter as the target partitioning parameter.
4. The method according to claim 1, characterized in that, Rearranging the initial numbers of each node based on the node connectivity degree to obtain the target numbers of each node includes: Obtaining the target sorting result of each node based on the magnitude of the node connectivity degree; Rearranging the initial numbers of each node based on the target sorting result to obtain the target numbers of each node.
5. The method according to claim 4, wherein, the obtaining the target sorting result of each node based on the magnitude of the node connectivity degree includes: Adding the initial numbers of each node to the target queue in ascending order of the node connectivity degree and in ascending order of the initial number until the initial numbers of all nodes are added; Obtaining the reverse addition order of each node based on the addition order of the initial numbers of each node in the target queue; Determining the reverse addition order as the target sorting result.
6. The method according to claim 5, wherein, the adding the initial numbers of each node to the target queue in ascending order of the node connectivity degree and in ascending order of the initial number until the initial numbers of all nodes are added includes: Adding the initial number of the first target node with the smallest connectivity degree to the target queue in ascending order of the node connectivity degree; wherein, a target queue is used to store the initial numbers of all nodes in a sub - parallel region; Taking the first target node as the starting connection point, determining the second target nodes among each node except the first target node, and the connection relationship between the second target nodes and the first target node; Adding the initial numbers of each second target node to the target queue in sequence according to the connection relationship between the second target node and the first target node and the magnitude of the initial number of the second target node.
7. A device for obtaining a grid file, wherein, the device includes: A first obtaining module, configured to obtain an unstructured grid of a target calculation region; wherein, the unstructured grid includes multiple nodes, each node corresponds to an initial number, and each node is adjacent to at least one node; A partitioning module, configured to partition the unstructured grid to obtain multiple sub - parallel regions; wherein, the nodes included in different sub - parallel regions do not intersect; A second obtaining module, configured to obtain the node connectivity degree of each sub - parallel region; wherein, the node connectivity degree is the number of adjacent nodes connected by each node in each sub - parallel region; A rearrangement module, configured to rearrange the initial numbers of each node based on the node connectivity degree to obtain the target numbers of each node; A generation module, configured to generate a grid file of the target calculation region based on multiple sub - parallel regions, and the initial numbers and target numbers corresponding to each node in each sub - parallel region; Among them, the grid file is applied in the process of the processor reading the data stored in the nodes in the target calculation area; among them, the dividing the unstructured grid to obtain a plurality of sub-parallel regions includes: Obtaining the node weights and node interconnection information corresponding to each of the plurality of nodes; wherein, the node interconnection information is used to characterize the connection relationship between each node and other nodes; Based on the node weights and the node interconnection information, obtaining an initial calculation undirected graph file corresponding to the unstructured grid; Dividing the initial calculation undirected graph file to obtain a plurality of the sub-parallel regions; The dividing the initial calculation undirected graph file to obtain a plurality of the sub-parallel regions includes: Determining a target partitioning parameter for partitioning the initial calculation undirected graph file; Based on the target partitioning parameter, partitioning the initial calculation undirected graph file to obtain a plurality of sub-parallel region undirected graph files; Based on the plurality of sub-parallel region undirected graph files, determining a plurality of the sub-parallel regions.
8. An electronic device, the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, it implements the method for obtaining a grid file according to any one of claims 1 to 6.
Citation Information
Patent Citations
Mesh generation computing method and device
CN103970715A
Large-scale parallel mesh generation system and method for finite element analysis
CN111125949A