Super computer load balancing strategy determination method and device based on structured grid, equipment and storage medium
By constructing a structural grid and using a graph partitioning algorithm to optimize the supercomputer load balancing strategy, the problem of multi-level communication delay not taken into account in existing technologies is solved, and the load balancing efficiency and user experience are improved.
Patent Information
- Application Number
- CN202511274554.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing technologies do not fully consider the impact of multi-level communication latency in the actual hardware architecture of the Sunway supercomputer in supercomputer load balancing strategies, resulting in insufficient engineering adaptability.
The supercomputer load balancing strategy based on structured grid sets the processors as nodes, constructs a structured grid with target grid parameters, and establishes a load balancing model based on the hardware architecture. The grid is segmented using preset constraints, the vertices, edge weights and degrees are determined, and a directed graph mapping result is constructed. The graph partitioning algorithm is used to optimize load balancing.
It improves the efficiency of the supercomputer load balancing strategy and enhances the user experience.
Smart Images

Figure CN120803745A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computational fluid dynamics, in particular to a supercomputer load balancing strategy determination method, device and equipment based on structured grid and a storage medium. BACKGROUND
[0002] At present, an intelligent multi-dimensional subdivision algorithm is proposed in the prior art to process the grid blocks with large computing load in the original grid blocks corresponding to the supercomputer, and then a genetic algorithm is used to distribute the multi-block structured grid processed by the intelligent multi-dimensional subdivision algorithm to the processor cores. In addition, by referring to the classical greedy load balancing strategy, an improved segmentation strategy for segmenting the original grid blocks corresponding to the supercomputer is also proposed in the prior art, including a new sub-grid block subdivision method and a cycle strategy. The algorithm introduces a grid block evaluation interval and a load lower limit. Furthermore, a block-based recursive binary method is proposed to segment the undirected graph.
[0003] From the prior art analysis, the traditional grid partitioning strategy corresponding to the original grid blocks corresponding to the supercomputer is mostly based on the theoretical load balancing assumption, and only optimizes the task distribution between the computing nodes through the abstract model, but does not fully consider the influence of the multi-level communication delay (such as inter-node, inter-core group and inter-core communication) in the actual hardware architecture of the Sunway supercomputer on the parallel efficiency, resulting in insufficient engineering adaptability.
[0004] As can be seen from the above, how to improve the efficiency of determining the load balancing strategy of the supercomputer based on the structured grid in the process of determining the load balancing strategy of the supercomputer based on the structured grid is a problem to be solved at present. SUMMARY
[0005] Therefore, the purpose of the present application is to provide a supercomputer load balancing strategy determination method, device and equipment based on structured grid and a storage medium, which can improve the efficiency of determining the load balancing strategy of the supercomputer in the process of determining the load balancing strategy of the supercomputer based on the structured grid. The specific scheme is as follows:
[0006] In the first aspect, the present application provides a supercomputer load balancing strategy determination method based on structured grid, comprising:
[0007] The processor in the supercomputer is set as a node, and a structured grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then a load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer; the hardware architecture includes the node, the core group in the node, and the communication structure between the master core and the slave core in the core group; the target grid parameters include the number of grid elements and the number of adjacent surfaces between grid blocks;
[0008] Split the structural grid based on the load balancing model and the preset node calculation time constraint condition, the preset inter-node communication time constraint condition, the preset inter-core group communication time constraint condition and the preset inter-slave core communication time constraint condition to obtain a plurality of split grid blocks;
[0009] Determine vertices and vertex weights based on the split grid blocks and corresponding grid quantities, respectively, determine edges and edge weights based on the association relationship between two split grid blocks and the number of adjacent face grid units, and then set the number of grid blocks adjacent to the two split grid blocks as the degree corresponding to the vertices, to determine a directed graph mapping result based on the vertices, the vertex weights, the edges, the edge weights and the degrees;
[0010] Perform graph partitioning processing on the directed graph mapping result using a preset graph partitioning algorithm to obtain a target graph partitioning result, and then inverse map the target graph partitioning result to a target grid to determine a load balancing strategy corresponding to the supercomputer based on the target grid.
[0011] Optionally, the processors in the supercomputer are set as nodes, and a structural grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then a load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer, including:
[0012] Determine each processor in the supercomputer, the processor including a plurality of core groups, each core group including a master core for managing data and a plurality of slave cores for operating data; the master core and the slave core share the memory of the processor;
[0013] Determine the grid division granularity of the grid based on the computing power of the slave cores in the master-slave core heterogeneous architecture of the supercomputer, and determine the adjacent face connection relationship of the grid based on the communication demand between each core group, to construct a structural grid based on the grid division granularity and the adjacent face connection relationship;
[0014] Determine the hardware architecture corresponding to the supercomputer based on the communication structure between the nodes, the core groups in the nodes and the master cores and the slave cores in the core groups, and determine the corresponding computing level based on the core groups, the master cores, the slave cores and the nodes, to determine the corresponding storage access delay characteristics based on the computing level;
[0015] Establish a load balancing model corresponding to the supercomputer based on the storage access delay characteristics, the hardware architecture and the target grid parameters.
[0016] Optionally, the structure grid is segmented based on the load balancing model and preset node calculation time constraint condition, preset inter-node communication time constraint condition, preset inter-core group communication time constraint condition and preset inter-slave core communication time constraint condition to obtain a plurality of segmented grid blocks, including:
[0017] The preset node calculation time constraint condition is determined based on the number of grid units corresponding to the node, and the preset inter-node communication time constraint condition is determined based on the number of grid surfaces for communication between the node and other nodes;
[0018] The preset inter-core group communication time constraint condition is determined based on the number of grid surfaces for communication between the core groups in the node, and the preset inter-slave core communication time constraint condition is determined based on the number of grid units in the core group;
[0019] The first weighting coefficient corresponding to the preset node calculation time constraint condition, the second weighting coefficient corresponding to the preset inter-node communication time constraint condition, the third weighting coefficient corresponding to the preset inter-core group communication time constraint condition and the fourth weighting coefficient corresponding to the preset inter-slave core communication time constraint condition are determined based on the number of nodes in the structure grid, the number of grid blocks in the node, the number of grid surfaces for communication;
[0020] The total execution time corresponding to the node is determined based on the preset node calculation time constraint condition, the preset inter-node communication time constraint condition, the preset inter-core group communication time constraint condition, the preset inter-slave core communication time constraint condition and the corresponding weighting coefficients;
[0021] The structure grid is segmented based on the load balancing model and the total execution time and a preset grid segmentation algorithm to obtain a plurality of segmented grid blocks with a geometry satisfying a preset shape condition; the preset grid segmentation algorithm includes recursive bisection and multi-block structure grid block segmentation algorithm.
[0022] Optionally, the vertices and vertex weights are determined based on the segmented grid blocks and corresponding grid units, respectively, the edges and edge weights are determined based on the association relationship between two segmented grid blocks and the number of adjacent surface grid units, and the number of grid blocks adjacent to the two segmented grid blocks is set as the degree corresponding to the vertices, so as to determine a directed graph mapping result based on the vertices, vertex weights, edges, edge weights and degrees, including:
[0023] mapping each of the divided grid blocks as a vertex, determining a grid quantity corresponding to the divided grid block, and determining a vertex weight corresponding to the vertex based on the grid quantity and the first weighting coefficient, then determining an association relationship between each two of the divided grid blocks, and mapping the association relationship as an edge;
[0024] determining a number of adjacent face grid cells between each two of the divided grid blocks, and determining an edge weight corresponding to the edge based on the number of adjacent face grid cells, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient, then determining a number of adjacent grid blocks of each two of the divided grid blocks, and mapping the number of adjacent grid blocks as a degree corresponding to the vertex;
[0025] constructing a directed graph mapping result based on the vertex, the vertex weight, the edge, the edge weight, and the degree; the directed graph mapping result is a directed and weighted graph.
[0026] Optionally, after determining the directed graph mapping result based on the vertex, the vertex weight, the edge, the edge weight, and the degree, the method further includes:
[0027] determining each vertex in the directed graph mapping result, and matching each current vertex by using a preset graph coarsening algorithm to obtain a current matched vertex, then determining a current coarsened graph based on the current matched vertex and a corresponding edge weight;
[0028] determining whether a graph size corresponding to the current coarsened graph is smaller than a preset threshold, if the graph size corresponding to the current coarsened graph is smaller than the preset threshold, setting the current coarsened graph as a target coarsened graph, if the graph size corresponding to the current coarsened graph is not smaller than the preset threshold, re-jumping to the step of matching each current vertex by using the preset graph coarsening algorithm;
[0029] or, determining whether the current coarsened graph satisfies a preset condition, if the current coarsened graph satisfies the preset condition, setting the current coarsened graph as the target coarsened graph, if the current coarsened graph does not satisfy the preset condition, re-jumping to the step of matching each current vertex by using the preset graph coarsening algorithm.
[0030] Optionally, the method of performing graph partitioning processing on the directed graph mapping result by using a preset graph partitioning algorithm to obtain a target graph partitioning result includes:
[0031] The preset graph partitioning algorithm includes a graph partitioning algorithm based on an iterative improvement strategy, a graph partitioning algorithm based on a construction method, a graph partitioning algorithm based on a mathematical method, and a graph partitioning algorithm based on an intelligent optimization algorithm.
[0032] If the difference between the current vertex weight results is not less than the preset difference threshold, the step of using the preset graph partitioning algorithm to perform graph partitioning processing on the current directed graph mapping result is re-executed.
[0033] If the difference between the current vertex weight results is less than the preset difference threshold, the edge weight sum corresponding to each current vertex subset is determined, and the current graph partitioning result corresponding to the edge weight sum with the smallest value is set as the target graph partitioning result.
[0034] Optionally, the target graph partitioning result is inversely mapped into a target grid to determine a load balancing strategy corresponding to the supercomputer based on the target grid, including:
[0035] The target graph partitioning result is mapped to the previous layer level by layer, and the mapping result is inversely mapped into a target grid based on the vertex number corresponding to the target graph partitioning result, and a load balancing strategy corresponding to the supercomputer is determined based on the grid partitioning result in the target grid.
[0036] The grid block number in the target grid corresponds to the vertex number one by one; the target execution total time of each node corresponding to the load balancing strategy is not greater than the node execution total time; and the node execution total time is the execution total time of each node in the structure grid.
[0037] In a second aspect, the application provides a supercomputer load balancing strategy determination device based on a structure grid, including:
[0038] A structure grid construction module is configured to set a processor in the supercomputer as a node, construct a structure grid with target grid parameters based on the master-slave core heterogeneous architecture of the supercomputer, and then establish a load balancing model based on the target grid parameters and the hardware architecture of the supercomputer; the hardware architecture includes the node, a core group in the node, and a communication structure between the master core and the slave core in the core group; the target grid parameters include the number of grid units and the number of adjacent surfaces between grid blocks.
[0039] The structural grid segmentation module is configured to segment the structural grid based on the load balancing model, and based on preset node calculation time constraints, preset inter-node communication time constraints, preset inter-core group communication time constraints, and preset inter-slave core communication time constraints, to obtain a plurality of segmented grid blocks.
[0040] The mapping result determination module is configured to determine vertices and vertex weights based on the segmented grid blocks and corresponding grid quantities, respectively, to determine edges and edge weights based on the association between two segmented grid blocks and the number of adjacent face grid units, and to set the number of grid blocks adjacent to the two segmented grid blocks as the degrees corresponding to the vertices, so as to determine a directed graph mapping result based on the vertices, the vertex weights, the edges, the edge weights, and the degrees.
[0041] The load balancing strategy generation module is configured to perform graph partitioning processing on the directed graph mapping result by using a preset graph partitioning algorithm to obtain a target graph partitioning result, to inversely map the target graph partitioning result into a target grid, and to determine a load balancing strategy corresponding to the supercomputer based on the target grid.
[0042] In a third aspect, the present application provides an electronic device, comprising:
[0043] A memory configured to store a computer program.
[0044] A processor configured to execute the computer program to implement the foregoing method for determining a load balancing strategy for a supercomputer based on a structural grid.
[0045] In a fourth aspect, the present application provides a computer-readable storage medium configured to store a computer program, wherein the computer program is executed by a processor to implement the foregoing method for determining a load balancing strategy for a supercomputer based on a structural grid.
[0046] As can be seen from the above, before determining the supercomputer load balancing strategy based on the structural grid, the processor in the supercomputer needs to be set as a node, and the structural grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then the load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer; the structural grid is segmented based on the load balancing model and the preset node calculation time constraint condition, the preset inter-node communication time consumption constraint condition, the preset inter-kernel group communication time consumption constraint condition and the preset inter-slave core communication time consumption constraint condition, to obtain a plurality of segmented grid blocks; the vertices and vertex weights are respectively determined based on the segmented grid blocks and corresponding grid quantities, and the edges and edge weights are respectively determined based on the correlation between the two segmented grid blocks and the number of adjacent face grid units, and then the number of grid blocks adjacent to the two segmented grid blocks is set as the degree corresponding to the vertex, so as to determine the directed graph mapping result based on the preset vertex edge weighted directed graph structure mapping rule, the vertex, the vertex weight, the edge, the edge weight and the degree; the directed graph mapping result is graph partitioned by using the preset graph partitioning algorithm to obtain a target graph partitioning result, and then the target graph partitioning result is inversely mapped to a target grid to determine the load balancing strategy corresponding to the supercomputer based on the target grid.
[0047] As can be seen from the above, before determining the supercomputer load balancing strategy based on the structural grid, the processor in the supercomputer needs to be set as a node, and the structural grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then the load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer; the structural grid is segmented based on the load balancing model and the preset node calculation time constraint condition, the preset inter-node communication time consumption constraint condition, the preset inter-kernel group communication time consumption constraint condition and the preset inter-slave core communication time consumption constraint condition, to obtain a plurality of segmented grid blocks; the vertices and vertex weights are respectively determined based on the segmented grid blocks and corresponding grid quantities, and the edges and edge weights are respectively determined based on the correlation between the two segmented grid blocks and the number of adjacent face grid units, and then the number of grid blocks adjacent to the two segmented grid blocks is set as the degree corresponding to the vertex, so as to determine the directed graph mapping result based on the preset vertex edge weighted directed graph structure mapping rule, the vertex, the vertex weight, the edge, the edge weight and the degree; the directed graph mapping result is graph partitioned by using the preset graph partitioning algorithm to obtain a target graph partitioning result, and then the target graph partitioning result is inversely mapped to a target grid to determine the load balancing strategy corresponding to the supercomputer based on the target grid. In this way, the efficiency of determining the load balancing strategy of the supercomputer is improved based on the structural grid in the process of determining the supercomputer load balancing strategy based on the structural grid, thereby improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only are a part of the present application, and for those skilled in the art, other drawings can be obtained based on the provided drawings without any creative effort.
[0049] Figure 1 A flow chart of a method for determining a load balancing strategy of a supercomputer based on a structured grid according to the present application is disclosed.
[0050] Figure 2 A flow chart of a specific method for determining a load balancing strategy of a supercomputer based on a structured grid according to the present application is disclosed.
[0051] Figure 3 A specific schematic diagram of a processor structure according to the present application is disclosed.
[0052] Figure 4 A flow chart of a specific process of mapping a grid into a point-edge weighted directed graph according to the present application is disclosed.
[0053] Figure 5 A flow chart of a specific process of vertex matching of a coarsening stage according to the present application is disclosed.
[0054] Figure 6 A flow chart of a specific process of inverse mapping a point-edge weighted directed graph into a grid according to the present application is disclosed.
[0055] Figure 7 A schematic diagram of a device structure for determining a load balancing strategy of a supercomputer based on a structured grid according to the present application is disclosed.
[0056] Figure 8 A schematic diagram of an electronic device structure according to the present application is disclosed. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without any creative effort belong to the scope of protection of the present application.
[0058] Currently, the traditional grid partition strategy corresponding to the original grid block corresponding to the supercomputer is based on the theoretical load balancing assumption, and only the task allocation between the computing nodes is optimized through the abstract model, but the influence of the multi-level communication delay (such as inter-node, inter-core group and inter-core communication) in the actual hardware architecture of the Godson supercomputer on the parallel efficiency is not fully considered, resulting in insufficient engineering adaptability. Therefore, the application provides a supercomputer load balancing strategy determination method based on a structure grid, which can improve the efficiency of determining the load balancing strategy of the supercomputer based on the structure grid in the process of determining the load balancing strategy of the supercomputer based on the structure grid.
[0059] Referring to Figure 1 The embodiment of the application discloses a supercomputer load balancing strategy determination method based on a structure grid, comprising:
[0060] Step S11, the processor in the supercomputer is set as a node, and a structure grid with a target grid parameter is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then a load balancing model is established based on the target grid parameter and the hardware architecture of the supercomputer; the hardware architecture includes the node, the core group in the node, and the communication structure between the master core and the slave core in the core group; the target grid parameter includes the number of grid units and the number of adjacent surfaces between grid blocks.
[0061] In this embodiment, the process of determining the supercomputer load balancing strategy based on the structure grid is as shown in Figure 2 Firstly, the application embodiment needs to model the computing load balancing problem, and divide the grid based on the load balancing partition model, and then realize the partition of the grid through graph partitioning.
[0062] In a specific embodiment, as shown in Figure 3 The supercomputer adopts a 64-bit instruction system and a processor designed independently, wherein the chip adopts a unique master-slave core heterogeneous architecture, a single processor (single node) is divided into 6 core groups (Core Groups, CG), each core group contains 1 master core with a frequency of 2.10GHz, called Manage Processing Elements (MPE), and 64 slave cores with a frequency of 2.25GHz, called Computing Processing Elements (CPE), and shares 16G DDR4 main memory.
[0063] Specifically, the processors in the supercomputer are set as nodes, and a structural grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer. Then, a load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer, which may include: determining each processor in the supercomputer, the processor including several core groups, each core group including a master core for managing data and several slave cores for calculating data; the master core and the slave core share the memory of the processor; determining the grid division granularity of the grid based on the computing power of the slave cores in the master-slave core heterogeneous architecture of the supercomputer, and determining the adjacent surface connection relationship of the grid based on the communication requirements between each core group, so as to construct a structural grid based on the grid division granularity and the adjacent surface connection relationship; determining the hardware architecture corresponding to the supercomputer based on the nodes, the core groups in the nodes, and the communication structure between the master core and the slave cores in the core groups, and determining the corresponding computing level based on the core groups, the master core, the slave cores and the nodes, so as to determine the corresponding storage access delay characteristics based on the computing level; establishing a load balancing model corresponding to the supercomputer based on the storage access delay characteristics, the hardware architecture and the target grid parameters.
[0064] Step S12: using the load balancing model, and based on the preset node calculation time constraint, the preset inter-node communication time constraint, the preset node core group communication time constraint and the preset slave core communication time constraint, the structural grid is segmented to obtain a plurality of segmented grid blocks.
[0065] In this embodiment, since the execution time of parallel computing is a complex and comprehensive factor, the total execution time of a single node is affected by both the internal computing time and the communication time with other nodes or within the node. Therefore, it can be expressed by the following formula:
[0066] ;
[0067] in, ,function Indicates the The total execution time of each node, Indicates the The computation time on each node, Indicates the The communication time between a node and other nodes, Indicates the Communication time between core groups within a node, Indicates the Communication time between slave cores in each core group of a node, Indicates the total number of nodes used for calculation.
[0068] It is worth mentioning that the computation time on the node The number of grid cells assigned to the node directly related, communication time between the node and other nodes the number of mesh faces corresponding to each node related to communication with the node and other nodes directly related, communication time between each core group in the node the number of mesh faces related to communication between each core group directly related, communication time between each slave core in each core group in the node the number of mesh cells allocated to the core group directly related.
[0069] Therefore, the above formula can be further written as:
[0070] ;
[0071] wherein, , represents the weighted coefficient of node calculation time, , and respectively represent the weighted coefficients corresponding to the communication time between nodes, between core groups in the node, and between slave cores in the core group.
[0072] It can be understood that each mesh block corresponds to six mesh faces, and each mesh face corresponds to a mesh cell. In addition, it is worth mentioning that the above coefficients , , and are obtained by real machine testing. In one specific embodiment, the embodiment of the present application assumes that the number of nodes used for calculation is , each time nodes in are sampled, and the number of mesh blocks in each node is gradually increased from to during each sampling of the node, and then the number of communication-related mesh faces is gradually increased from to , and the above process is repeated times. By averaging the results of times of execution, the values of the weighted coefficients , , and are finally calculated, wherein is the number of nodes to be sampled after grouping , is the total number of sampling executions, and are multiplied by .
[0073] The objective function of the load balancing problem is to find the minimum value of the maximum time used by all nodes, and the expression is as follows:
[0074] ;
[0075] wherein, .
[0076] Further, in order to reduce the complexity of the grid data to be partitioned, thereby reducing the grid area related to communication, the grid needs to be segmented before partitioning, so that the segmented grid block is "approximately a cube". The grid block segmentation strategy can use recursive bisection method, multi-block structured grid block segmentation algorithm, etc., which is not specifically limited here.
[0077] Specifically, the load balancing model is used to segment the structured grid based on the preset node calculation time constraint condition, the preset inter-node communication time constraint condition, the preset inter-core group communication time constraint condition, and the preset inter-slave core communication time constraint condition, to obtain a plurality of segmented grid blocks, which can include: determining the preset node calculation time constraint condition based on the number of grid elements corresponding to the node, and determining the preset inter-node communication time constraint condition based on the number of grid surfaces for communication between the node and the remaining nodes; determining the preset inter-core group communication time constraint condition based on the number of grid surfaces for communication between the core groups in the node, and determining the preset inter-slave core communication time constraint condition based on the number of grid elements in the core group; determining a first weighting coefficient corresponding to the preset node calculation time constraint condition, a second weighting coefficient corresponding to the preset inter-node communication time constraint condition, a third weighting coefficient corresponding to the preset inter-core group communication time constraint condition, and a fourth weighting coefficient corresponding to the preset inter-slave core communication time constraint condition based on the number of nodes in the structured grid, the number of grid blocks in the node, and the number of grid surfaces for communication; determining the total execution time corresponding to the node based on the preset node calculation time constraint condition, the preset inter-node communication time constraint condition, the preset inter-core group communication time constraint condition, the preset inter-slave core communication time constraint condition, and the corresponding weighting coefficients; using the load balancing model and based on the total execution time and the preset grid segmentation algorithm to segment the structured grid to obtain a plurality of segmented grid blocks whose geometric shapes meet the preset shape condition; the preset grid segmentation algorithm includes recursive bisection method and multi-block structured grid block segmentation algorithm.
[0078] In step S13, vertices and vertex weights are determined based on the two partitioned grid blocks and corresponding grid quantities, edges and edge weights are determined based on the association between the two partitioned grid blocks and the number of adjacent surface grid cells, the number of grid blocks adjacent to the two partitioned grid blocks is set as the degree corresponding to the vertices, and a directed graph mapping result is determined based on the vertices, vertex weights, edges, edge weights and degrees.
[0079] In this embodiment, the partitioned grid blocks need to be partitioned. Because the structural grid scale is large in this embodiment, a multi-layer graph partitioning algorithm in a graph partitioning algorithm is used in the partitioning process. The algorithm mainly includes five small steps: mapping of the grid to a point-edge weighted directed graph, a coarsening stage, an initial partitioning stage, a refinement stage, and inverse mapping of the point-edge weighted directed graph to a grid. The flowchart of mapping of the grid to a point-edge weighted directed graph is shown in FIG. 2, that is, the physical grid is converted into an abstract graph structure. Figure 4
[0080] Specifically, vertices and vertex weights are determined based on the two partitioned grid blocks and corresponding grid quantities, edges and edge weights are determined based on the association between the two partitioned grid blocks and the number of adjacent surface grid cells, the number of grid blocks adjacent to the two partitioned grid blocks is set as the degree corresponding to the vertices, and a directed graph mapping result is determined based on the vertices, vertex weights, edges, edge weights and degrees. This can include: mapping each partitioned grid block to a vertex, determining the grid quantity corresponding to the partitioned grid block, determining the vertex weight corresponding to the vertex based on the grid quantity and a first weighting coefficient, determining the association between each two partitioned grid blocks in the partitioned grid blocks, and mapping the association to an edge; determining the number of adjacent surface grid cells between each two partitioned grid blocks in the partitioned grid blocks, and determining the edge weight corresponding to the edge based on the number of adjacent surface grid cells, a second weighting coefficient, a third weighting coefficient and a fourth weighting coefficient, then determining the number of adjacent grid blocks of each two partitioned grid blocks to map the number of adjacent grid blocks to the degree corresponding to the vertex; constructing a directed graph mapping result based on the vertices, vertex weights, edges, edge weights and degrees; and the directed graph mapping result is a directed and weighted graph.
[0081] Further, in this embodiment, the grid blocks are taken as points in the graph, the grid quantity of the grid block is taken as the weight of the point, the association between two grid blocks is taken as an edge, the number of adjacent surface grid cells between the two grid blocks is taken as the weight of the edge, and the number of adjacent grid blocks of the grid block is taken as the degree of the point. The conversion of the grid quantity to the vertex weight value can be represented based on the product of the weighting coefficient of the node calculation time and the number of grid cells:
[0082]
[0083] wherein, is the conversion result of the number of adjacent face mesh cells between two mesh blocks to the edge weight.
[0084] Further, the conversion of the number of adjacent face mesh cells between two mesh blocks to the edge weight can be based on the weighted coefficient of the node communication time , and the product sum of the number of adjacent face mesh cells between mesh blocks , and the number of core group mesh cells , and the expression is as follows:
[0085] ;
[0086] wherein, is the conversion result of the number of adjacent face mesh cells between two mesh blocks to the edge weight.
[0087] Further, the vertex matching process of the coarsening stage graph is as shown in Figure 5 : in the coarsening stage, the embodiment of the present application needs to match the vertices in the original graph , and combine the matched vertices to obtain the next layer of coarsening graph, and repeat the above process to obtain a series of coarsening graphs , until the coarsening graph is small enough or meets the preset condition, wherein , is the original graph, is the vertex in the original graph, is the edge in the original graph.
[0088] Specifically, after determining the directed graph mapping result based on the vertex, vertex weight, edge, edge weight and degree, it can further include: determining each vertex in the directed graph mapping result, and matching each current vertex by using a preset graph coarsening algorithm to obtain a current matched vertex, and then determining a current coarsening graph based on the current matched vertex and the corresponding edge weight; judging whether the size of the current coarsening graph is smaller than a preset threshold, if the size of the current coarsening graph is smaller than the preset threshold, setting the current coarsening graph as a target coarsening graph, if the size of the current coarsening graph is not smaller than the preset threshold, jumping back to the step of matching each current vertex by using the preset graph coarsening algorithm; or, judging whether the current coarsening graph meets a preset condition, if the current coarsening graph meets the preset condition, setting the current coarsening graph as the target coarsening graph, if the current coarsening graph does not meet the preset condition, jumping back to the step of matching each current vertex by using the preset graph coarsening algorithm.
[0089] In step S14, a preset graph partitioning algorithm is used to perform graph partitioning processing on the directed graph mapping result to obtain a target graph partitioning result, and then the target graph partitioning result is inversely mapped to a target grid to determine a load balancing strategy corresponding to the supercomputer based on the target grid.
[0090] In the initial partitioning stage, the embodiment of the present application can perform partitioning on the coarsest graph obtained in the coarse stage to obtain an initial partitioning. In the process of partitioning the coarsest graph, the target is to minimize the sum of edge weights between vertices in different subsets while meeting the condition that each vertex subset contains vertices whose weight sums are approximately equal, that is, to solve the load balancing objective function: ;
[0091] wherein, .
[0092] It is worth mentioning that the above-mentioned partitioning method can use a graph partitioning algorithm based on an iterative improvement strategy, a graph partitioning algorithm based on a construction method, a graph partitioning algorithm based on a mathematical method, and a graph partitioning algorithm based on an intelligent optimization algorithm, which will not be exemplified one by one.
[0093] Specifically, the graph partitioning processing on the directed graph mapping result using the preset graph partitioning algorithm to obtain the target graph partitioning result can include: performing graph partitioning processing on the current directed graph mapping result using the preset graph partitioning algorithm to obtain a current graph partitioning result, and determining each current vertex subset corresponding to the current graph partitioning result and the current vertex weight sum result corresponding to the current vertex subset; the preset graph partitioning algorithm includes a graph partitioning algorithm based on an iterative improvement strategy, a graph partitioning algorithm based on a construction method, a graph partitioning algorithm based on a mathematical method, and a graph partitioning algorithm based on an intelligent optimization algorithm; it is determined whether the difference between each current vertex weight sum result is less than a preset difference threshold value, if the difference between each current vertex weight sum result is not less than the preset difference threshold value, then jump back to the step of performing graph partitioning processing on the current directed graph mapping result using the preset graph partitioning algorithm; if the difference between each current vertex weight sum result is less than the preset difference threshold value, then determine the edge weight sum corresponding to each current vertex subset, and set the current graph partitioning result corresponding to the edge weight sum with the smallest value in each edge weight sum as the target graph partitioning result.
[0094] Subsequently, the embodiment of the present application needs to refine the target graph partitioning result. In the refinement stage, the primary task is to map the partitioning of the coarsest graph obtained in the initial partitioning stage to the last layer of refined graph, that is, the partitioning of graph is obtained by mapping the partitioning of graph .
[0095] Finally, the embodiment of the present application needs to inversely map the obtained point-edge-weighted directed graph to a grid, and the process of inversely mapping the point-edge-weighted directed graph to the grid is as followsFigure 6 Specifically, the inverse mapping of the target graph partitioning result to the target grid to determine the load balancing strategy corresponding to the supercomputer based on the target grid can include: mapping the target graph partitioning result to the last level layer by layer to obtain a mapping result, then inversely mapping the mapping result to the target grid based on the vertex number corresponding to the target graph partitioning result, and determining the load balancing strategy corresponding to the supercomputer based on the grid partitioning result in the target grid; wherein the grid block number in the target grid corresponds to the vertex number one by one; the target execution total time of each node corresponding to the load balancing strategy is not greater than the node execution total time; the node execution total time is the execution total time of each node in the structure grid.
[0096] It is worth mentioning that after the graph partitioning process is completed, the grid data is always in the memory, the vertex number in the subgraph corresponds to the grid block number in the grid one by one, and the point weight, edge weight and point-to-point association relationship information of the current graph do not need to be recorded again in the process of mapping the directed graph to the grid. Therefore, the embodiment of the present application only needs to inform the vertex number in each subgraph to convert to the grid block number in each partition to obtain the final partitioning result of the grid.
[0097] As can be seen from the above, before the structure grid based supercomputer load balancing strategy determination, the processor in the supercomputer needs to be set as a node, and a structure grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then a load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer; secondly, the structure grid is segmented based on the load balancing model and the preset node calculation time constraint condition, the preset inter-node communication time constraint condition, the preset inter-core group communication time constraint condition and the preset inter-slave core communication time constraint condition to obtain a plurality of segmented grid blocks; then, the vertices and vertex weights are determined based on the segmented grid blocks and the corresponding grid quantities respectively, and the edges and edge weights are determined based on the association relationship between the two segmented grid blocks and the number of adjacent face grid elements respectively, and then the number of grid blocks adjacent to the two segmented grid blocks is set as the degree corresponding to the vertex to determine the directed graph mapping result based on the preset vertex edge weighted directed graph structure mapping rule, the vertex, the vertex weight, the edge, the edge weight and the degree; finally, the graph partitioning algorithm is used to perform graph partitioning processing on the directed graph mapping result to obtain the target graph partitioning result, and then the target graph partitioning result is inversely mapped to the target grid to determine the load balancing strategy corresponding to the supercomputer based on the target grid. In this way, the efficiency of determining the load balancing strategy of the supercomputer is improved based on the structure grid in the process of determining the structure grid based supercomputer load balancing strategy, thereby improving the user experience.
[0098] Correspondingly, referring to Figure 7As shown, the application also provides a structural grid-based supercomputer load balancing strategy determination device, comprising:
[0099] A structural grid construction module 11 is configured to set processors in the supercomputer as nodes, construct a structural grid with target grid parameters based on a master-slave core heterogeneous architecture of the supercomputer, and then establish a load balancing model based on the target grid parameters and a hardware architecture of the supercomputer; the hardware architecture includes the nodes, core groups in the nodes, and a communication structure between master cores and slave cores in the core groups; the target grid parameters include a number of grid cells and a number of adjacent faces between grid blocks;
[0100] A structural grid segmentation module 12 is configured to segment the structural grid based on preset node calculation time constraints, preset inter-node communication time constraints, preset inter-core group communication time constraints, and preset inter-slave core communication time constraints using the load balancing model, to obtain a plurality of segmented grid blocks;
[0101] A mapping result determination module 13 is configured to determine vertices and vertex weights based on the segmented grid blocks and corresponding grid quantities, respectively, determine edges and edge weights based on the association between two segmented grid blocks and the number of adjacent face grid cells, and then set the number of grid blocks adjacent to the two segmented grid blocks as the degrees corresponding to the vertices, to determine a directed graph mapping result based on the vertices, vertex weights, edges, edge weights, and degrees;
[0102] A load balancing strategy generation module 14 is configured to perform graph partitioning processing on the directed graph mapping result using a preset graph partitioning algorithm to obtain a target graph partitioning result, then inversely map the target graph partitioning result to a target grid, and determine a load balancing strategy corresponding to the supercomputer based on the target grid.
[0103] As can be seen from the above, before the supercomputer load balancing strategy based on the structural grid is determined, the processor in the supercomputer needs to be set as a node first, and a structural grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then a load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer; secondly, the structural grid is segmented based on the load balancing model and the preset node calculation time constraint condition, the preset inter-node communication time constraint condition, the preset inter-core group communication time constraint condition and the preset inter-slave core communication time constraint condition, to obtain a plurality of segmented grid blocks; then, the vertices and vertex weights are determined based on the segmented grid blocks and the corresponding grid quantities respectively, and the edges and edge weights are determined based on the correlation between two segmented grid blocks and the number of adjacent face grid units, and then the number of grid blocks adjacent to the two segmented grid blocks is set as the degree corresponding to the vertex, to determine the directed graph mapping result based on the preset vertex edge weighted directed graph structure mapping rule, the vertex, the vertex weight, the edge, the edge weight and the degree; finally, the graph partitioning algorithm is used to perform graph partitioning processing on the directed graph mapping result to obtain a target graph partitioning result, and then the target graph partitioning result is inversely mapped to a target grid to determine the load balancing strategy corresponding to the supercomputer based on the target grid. In this way, the efficiency of determining the load balancing strategy of the supercomputer based on the structural grid is improved in the process of determining the supercomputer load balancing strategy based on the structural grid, thereby improving the user experience.
[0104] In some specific embodiments, the structural grid construction module 11 can specifically include:
[0105] The processor determination unit is configured to determine each processor in the supercomputer, wherein each processor includes a plurality of core groups, each core group includes a master core for managing data and a plurality of slave cores for operating data, and the master core and the slave cores share the memory of the processor.
[0106] The grid division granularity determination unit is configured to determine the grid division granularity of the grid based on the computing power of the slave cores in the master-slave core heterogeneous architecture of the supercomputer, and determine the adjacent face connection relationship of the grid based on the communication demand between each core group, so as to construct a structural grid based on the grid division granularity and the adjacent face connection relationship.
[0107] The hardware architecture determination unit is configured to determine the hardware architecture corresponding to the supercomputer based on the node, the core group in the node, and the communication structure between the master core and the slave core in the core group, and determine the corresponding computing level based on the core group, the master core, the slave core and the node, so as to determine the corresponding storage access delay characteristics based on the computing level.
[0108] A load balancing model construction unit is configured to construct a load balancing model corresponding to the supercomputer based on the storage access delay characteristics, the hardware architecture, and the target grid parameters.
[0109] In some embodiments, the structural grid partitioning module 12 can specifically include:
[0110] A first constraint condition determination unit is configured to determine a preset node calculation time constraint condition based on the number of grid cells corresponding to the node, and determine a preset inter-node communication time consumption constraint condition based on the number of grid faces through which the node communicates with the remaining nodes;
[0111] A second constraint condition determination unit is configured to determine a preset intra-node kernel group communication time consumption constraint condition based on the number of grid faces through which the cores in the node communicate, and determine a preset inter-core communication time consumption constraint condition based on the number of grid cells in the core group;
[0112] A weighting coefficient determination unit is configured to determine a first weighting coefficient corresponding to the preset node calculation time constraint condition, a second weighting coefficient corresponding to the preset inter-node communication time consumption constraint condition, a third weighting coefficient corresponding to the preset intra-node kernel group communication time consumption constraint condition, and a fourth weighting coefficient corresponding to the preset inter-core communication time consumption constraint condition, based on the number of nodes in the structural grid, the number of grid blocks in the node, and the number of grid faces through which communication is performed;
[0113] An execution total time determination unit is configured to determine an execution total time corresponding to the node based on the preset node calculation time constraint condition, the preset inter-node communication time consumption constraint condition, the preset intra-node kernel group communication time consumption constraint condition, the preset inter-core communication time consumption constraint condition, and the corresponding weighting coefficients;
[0114] A grid partitioning unit is configured to partition the structural grid based on the execution total time and a preset grid partitioning algorithm using the load balancing model, to obtain a plurality of partitioned grid blocks with geometrical shapes satisfying a preset shape condition; the preset grid partitioning algorithm includes a recursive bisection method and a multi-block structural grid block partitioning algorithm.
[0115] In some embodiments, the mapping result determination module 13 can specifically include:
[0116] An association relationship determination unit is configured to map each of the partitioned grid blocks to a vertex, determine a grid amount corresponding to the partitioned grid block, determine a vertex weight corresponding to the vertex based on the grid amount and the first weighting coefficient, determine an association relationship between each two of the partitioned grid blocks, and map the association relationship to an edge;
[0117] an edge weight determination unit configured to determine a number of adjacent face grid units between each two of the partitioned grid blocks, and determine an edge weight corresponding to the edge based on the number of adjacent face grid units, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient, and then determine a number of adjacent grid blocks of each two of the partitioned grid blocks, and map the number of adjacent grid blocks as a degree corresponding to the vertex;
[0118] a mapping result determination unit configured to construct a directed graph mapping result based on the vertex, the vertex weight, the edge, the edge weight, and the degree; the directed graph mapping result being a directed and weighted graph.
[0119] In some embodiments, the structure grid based supercomputer load balancing strategy determination apparatus can further include:
[0120] a coarse graph determination unit configured to determine each vertex in the directed graph mapping result, match each current vertex using a preset graph coarsening algorithm to obtain a current matched vertex, and then determine a current coarse graph based on the current matched vertex and a corresponding edge weight;
[0121] a first step jump unit configured to determine whether a graph size corresponding to the current coarse graph is less than a preset threshold, if the graph size corresponding to the current coarse graph is less than the preset threshold, set the current coarse graph as a target coarse graph, and if the graph size corresponding to the current coarse graph is not less than the preset threshold, jump back to the step of matching each current vertex using the preset graph coarsening algorithm;
[0122] a second step jump unit configured to determine whether the current coarse graph satisfies a preset condition, if the current coarse graph satisfies the preset condition, set the current coarse graph as the target coarse graph, and if the current coarse graph does not satisfy the preset condition, jump back to the step of matching each current vertex using the preset graph coarsening algorithm.
[0123] In some embodiments, the load balancing strategy generation module 14 can specifically include:
[0124] a weight and result determination unit configured to perform graph partitioning processing on the current directed graph mapping result using a preset graph partitioning algorithm to obtain a current graph partitioning result, and determine each current vertex subset corresponding to the current graph partitioning result and a current vertex weight and result corresponding to the current vertex subset; the preset graph partitioning algorithm including a graph partitioning algorithm based on an iterative improvement strategy, a graph partitioning algorithm based on a construction method, a graph partitioning algorithm based on a mathematical method, and a graph partitioning algorithm based on an intelligent optimization algorithm;
[0125] The difference judging unit is configured to judge whether the difference between each current vertex weight and the result is less than a preset difference threshold value, and if the difference between each current vertex weight and the result is not less than the preset difference threshold value, the step of performing graph partitioning processing on the current directed graph mapping result by using the preset graph partitioning algorithm is re-executed.
[0126] The edge weight sum determining unit is configured to determine the edge weight sum corresponding to each current vertex subset if the difference between each current vertex weight and the result is less than the preset difference threshold value, and set the current graph partitioning result corresponding to the edge weight sum with the smallest value as the target graph partitioning result.
[0127] In some specific embodiments, the load balancing strategy generation module 14 can specifically include:
[0128] The load balancing strategy generation subunit is configured to map the target graph partitioning result to the previous layer level by layer to obtain a mapping result, then inversely map the mapping result into a target grid based on the vertex number corresponding to the target graph partitioning result, and determine a load balancing strategy corresponding to the supercomputer based on the grid partitioning result in the target grid. The grid block number in the target grid corresponds to the vertex number in a one-to-one manner. The target execution total time of each node corresponding to the load balancing strategy is not greater than the node execution total time. The node execution total time is the execution total time of each node in the structural grid.
[0129] Further, the embodiment of the present application further discloses an electronic device, Figure 8 is an electronic device 20 structure diagram shown according to an exemplary embodiment, the contents in the figure cannot be considered as any limitation on the use range of the present application. The electronic device 20 can specifically include: at least one processor 21, at least one memory 22, power supply 23, communication interface 24, input output interface 25 and communication bus 26. Wherein, the memory 22 is used for storing computer programs, the computer programs are loaded and executed by the processor 21, to realize the related steps in the foregoing any embodiment disclosed structural grid based supercomputer load balancing strategy determination method. In addition, the electronic device 20 in the embodiment of the present application can be an electronic computer.
[0130] In the embodiment, the power supply 23 is used for providing working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the present application, which is not limited specifically herein; the input output interface 25 is used for obtaining external input data or outputting data to the outside world, and the specific interface type can be selected according to the specific application needs, which is not limited specifically herein.
[0131] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage mode can be temporary storage or permanent storage.
[0132] The operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including the computer program capable of completing the structural grid-based supercomputer load balancing strategy determination method performed by the electronic device 20 disclosed in any of the preceding embodiments, the computer program 222 can further include a computer program capable of completing other specific work.
[0133] Further, the present application also discloses a computer readable storage medium for storing a computer program; wherein the computer program is executed by a processor to implement the structural grid-based supercomputer load balancing strategy determination method disclosed in the preceding embodiments. For the specific steps of the method, refer to the corresponding content disclosed in the preceding embodiments, which will not be described here.
[0134] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. For the same or similar parts between the embodiments, refer to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and refer to the method part for the relevant part.
[0135] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of the two. In order to clearly show the interchangeability of hardware and software, the components and steps of each example have been described in the above description. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0136] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0137] Finally, it is to be understood that the phraseology or terminology such as "comprising", "including", "containing", or "consisting of" etc. used herein is merely open-ended, and does not exclude or deny the inherent
[0138] The above detailed description of the technical solutions provided by the present application has been described in detail, and the principles and implementation modes of the present application are described in this paper. The above description of the examples is only used to help understand the method and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in view of the above, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for determining a supercomputer load balancing strategy based on a structured grid, characterized in that: include: The processors in the supercomputer are set as nodes, and a structured grid with target grid parameters is constructed based on the master-slave core heterogeneous architecture of the supercomputer, and then a load balancing model is established based on the target grid parameters and the hardware architecture of the supercomputer; The hardware architecture includes the node, the core group in the node, and the communication structure between the master core and the slave core in the core group; the target grid parameters include the number of grid cells and the number of adjacent faces between grid blocks; Using the load balancing model, and based on preset node calculation time constraints, preset inter-node communication time constraints, preset node inter-core group communication time constraints, and preset inter-slave core communication time constraints, the structured grid is segmented to obtain a plurality of segmented grid blocks; Determining vertices and vertex weights based on the segmented grid blocks and corresponding grid quantities, and determining edges and edge weights based on an association relationship between two segmented grid blocks and a number of adjacent surface grid units, and then setting the number of grid blocks adjacent to the two segmented grid blocks as the degree corresponding to the vertex, so as to determine a directed graph mapping result based on the vertices, the vertex weights, the edges, the edge weights, and the degrees; The directed graph mapping result is subjected to graph partitioning processing using a preset graph partitioning algorithm to obtain a target graph partitioning result, and then the target graph partitioning result is inversely mapped into a target grid to determine a load balancing strategy corresponding to the supercomputer based on the target grid.
2. The method for determining a supercomputer load balancing strategy based on a structured grid according to claim 1, wherein: The method includes setting the processors in the supercomputer as nodes, constructing a structured grid with target grid parameters based on the master-slave core heterogeneous architecture of the supercomputer, and then establishing a load balancing model based on the target grid parameters and the hardware architecture of the supercomputer, including: Determine each processor in the supercomputer, wherein the processor includes a plurality of core groups, each of the core groups includes a master core for managing data and a plurality of slave cores for computing data; the master core and the slave cores share a memory of the processor; Determining a grid partitioning granularity of a grid based on the computing capabilities of the slave cores in the master-slave core heterogeneous architecture of the supercomputer, and determining adjacent surface connection relationships of the grid based on communication requirements between the core groups, so as to construct a structural grid based on the grid partitioning granularity and the adjacent surface connection relationships; determining a hardware architecture corresponding to the supercomputer based on the node, a core group in the node, and a communication structure between a master core and slave cores in the core group, and determining a corresponding computing hierarchy based on the core group, the master core, the slave core, and the node, so as to determine a corresponding storage access latency characteristic based on the computing hierarchy; A load balancing model corresponding to the supercomputer is established based on the storage access delay characteristics, the hardware architecture and the target grid parameters.
3. The method for determining a supercomputer load balancing strategy based on a structured grid according to claim 1, wherein: The load balancing model is used to segment the structured grid based on preset node calculation time constraints, preset inter-node communication time constraints, preset node inter-core group communication time constraints, and preset inter-slave core communication time constraints to obtain a number of segmented grid blocks, including: Determining a preset node computation time constraint based on the number of grid cells corresponding to the node, and determining a preset inter-node communication time constraint based on the number of grid surfaces for communication between the node and other nodes; Determining a preset node inter-core communication time constraint based on the number of grid surfaces communicating between the core groups in the node, and determining a preset slave inter-core communication time constraint based on the number of grid cells in the core group; Determining, based on the number of nodes in the structural grid, the number of grid blocks in the node, and the number of grid surfaces for communication, a first weighting coefficient corresponding to the preset node computation time constraint, a second weighting coefficient corresponding to the preset inter-node communication time constraint, a third weighting coefficient corresponding to the preset inter-node core communication time constraint, and a fourth weighting coefficient corresponding to the preset inter-slave core communication time constraint; Determine the total execution time corresponding to the node based on the preset node calculation time constraint, the preset inter-node communication time constraint, the preset node inter-core communication time constraint, the preset slave inter-core communication time constraint, and corresponding weighting coefficients; The load balancing model is used to segment the structural grid based on the total execution time and a preset grid segmentation algorithm to obtain a number of segmented grid blocks whose geometric shapes meet preset shape conditions; the preset grid segmentation algorithm includes a recursive bisection method and a multi-block structural grid block segmentation algorithm.
4. The method for determining a supercomputer load balancing strategy based on a structured grid according to claim 3, wherein: The steps of determining vertices and vertex weights based on the segmented grid blocks and corresponding grid quantities, and determining edges and edge weights based on an association relationship between two segmented grid blocks and a number of adjacent surface grid units, and then setting the number of grid blocks adjacent to the two segmented grid blocks as the degree corresponding to the vertex to determine a directed graph mapping result based on the vertices, the vertex weights, the edges, the edge weights, and the degrees, include: Mapping each of the segmented grid blocks to a vertex, determining a grid quantity corresponding to the segmented grid block, and determining a vertex weight corresponding to the vertex based on the grid quantity and the first weighting coefficient, then determining an association relationship between every two of the segmented grid blocks in each of the segmented grid blocks, and mapping the association relationship to an edge; Determining the number of adjacent surface mesh units between every two of the segmented mesh blocks in each of the segmented mesh blocks, and determining an edge weight corresponding to the edge based on the number of adjacent surface mesh units, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient, and then determining the number of adjacent mesh blocks between every two of the segmented mesh blocks in each of the segmented mesh blocks, so as to map the number of adjacent mesh blocks to a degree corresponding to the vertex; A directed graph mapping result is constructed based on the vertices, the vertex weights, the edges, the edge weights and the degrees; the directed graph mapping result is a directed weighted graph.
5. The method for determining a supercomputer load balancing strategy based on a structured grid according to claim 1, wherein: After determining the directed graph mapping result based on the vertices, the vertex weights, the edges, the edge weights, and the degrees, the method further includes: Determine each vertex in the directed graph mapping result, and use a preset graph coarsening algorithm to match each current vertex to obtain a current matched vertex, and then determine a current coarsened graph based on the current matched vertex and the corresponding edge weight; Determine whether the graph size corresponding to the current coarsening graph is less than a preset threshold; if the graph size corresponding to the current coarsening graph is less than the preset threshold, set the current coarsening graph as the target coarsening graph; if the graph size corresponding to the current coarsening graph is not less than the preset threshold, jump back to the step of matching each current vertex using the preset graph coarsening algorithm; Alternatively, determine whether the current coarsening graph meets the preset conditions. If the current coarsening graph meets the preset conditions, set the current coarsening graph as the target coarsening graph. If the current coarsening graph does not meet the preset conditions, jump back to the step of matching each current vertex using the preset graph coarsening algorithm.
6. The method for determining a supercomputer load balancing strategy based on a structured grid according to claim 1, wherein: The step of performing graph partitioning on the directed graph mapping result using a preset graph partitioning algorithm to obtain a target graph partitioning result includes: Performing graph partitioning processing on the current directed graph mapping result using a preset graph partitioning algorithm to obtain a current graph partitioning result, and determining each current vertex subset corresponding to the current graph partitioning result and the current vertex weights and results corresponding to the current vertex subset; the preset graph partitioning algorithm includes a graph partitioning algorithm based on an iterative improvement strategy, a graph partitioning algorithm based on a construction method, a graph partitioning algorithm based on a mathematical method, and a graph partitioning algorithm based on an intelligent optimization algorithm; Determine whether the difference between each current vertex weight and the result is less than a preset difference threshold; if the difference between each current vertex weight and the result is not less than the preset difference threshold, jump back to the step of performing graph partitioning processing on the current directed graph mapping result using the preset graph partitioning algorithm; If the difference between the current vertex weights and results is less than the preset difference threshold, the edge weights corresponding to each current vertex subset are determined, and the current graph segmentation result corresponding to the edge weight with the smallest value among the edge weights is set as the target graph segmentation result.
7. The method for determining a supercomputer load balancing strategy based on a structured grid according to any one of claims 1 to 6, characterized in that: The step of inversely mapping the target graph partitioning result into a target grid, and determining a load balancing strategy corresponding to the supercomputer based on the target grid, includes: Mapping the target graph partitioning result to the upper level layer by layer to obtain a mapping result, then inversely mapping the mapping result into a target grid based on vertex numbers corresponding to the target graph partitioning result, and determining a load balancing strategy corresponding to the supercomputer based on a grid partitioning result in the target grid; Among them, the grid block numbers in the target grid correspond one-to-one to the vertex numbers; the target total execution time of each of the nodes corresponding to the load balancing strategy is not greater than the total node execution time; the total node execution time is the total execution time of each of the nodes in the structural grid.
8. A device for determining a supercomputer load balancing strategy based on a structured grid, characterized in that: include: a structured grid construction module, configured to set processors in the supercomputer as nodes, construct a structured grid with target grid parameters based on the master-slave core heterogeneous architecture of the supercomputer, and then establish a load balancing model based on the target grid parameters and the hardware architecture of the supercomputer; The hardware architecture includes the node, the core group in the node, and the communication structure between the master core and the slave core in the core group; the target grid parameters include the number of grid cells and the number of adjacent faces between grid blocks; a structural grid segmentation module, configured to segment the structural grid using the load balancing model and based on preset node computation time constraints, preset inter-node communication time constraints, preset node inter-core group communication time constraints, and preset inter-slave core communication time constraints to obtain a plurality of segmented grid blocks; a mapping result determination module, configured to determine vertices and vertex weights based on the segmented grid blocks and corresponding grid quantities, and to determine edges and edge weights based on an association relationship between two segmented grid blocks and a number of adjacent surface grid units, and then set the number of grid blocks adjacent to the two segmented grid blocks as the degree corresponding to the vertex, so as to determine a directed graph mapping result based on the vertices, the vertex weights, the edges, the edge weights, and the degrees; A load balancing strategy generation module is used to perform graph partitioning on the directed graph mapping result using a preset graph partitioning algorithm to obtain a target graph partitioning result, and then inversely map the target graph partitioning result into a target grid to determine a load balancing strategy corresponding to the supercomputer based on the target grid.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the method for determining a supercomputer load balancing strategy based on a structured grid as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, it implements the method for determining a supercomputer load balancing strategy based on a structured grid according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for accelerating synchronous operation of communication locks between main core and core group based on Shenwei many-core processor
CN110262900A
Underwater three-dimensional sound field model Bellhop3D parallel implementation method based on domestic many-core supercomputing
CN115437782A
System and method for load balancing for parallel computations on structured multi-block meshes in cfd
US20140365186A1
Task processing method and apparatus, many-core system, and computer-readable medium
WO2022171002A1