A method and device for calculating target tasks by supercomputer

By obtaining and analyzing the computing network topology information of supercomputers and dynamically adjusting the allocation of computing resources, the problem of low computing resource utilization in the construction of supercomputers in installments is solved, and load balancing and computing performance improvement is achieved.

CN119576535BActive Publication Date: 2025-05-06COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411618572.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-05-06
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

In supercomputers built in stages, due to changes in hardware topology, the heterogeneity of parallel computing environments is difficult for the existing technology to make full use of computing resources and improve computing efficiency.

Method used

By obtaining the computing network topology information of the supercomputer, determining the computing power weight based on the computing performance of the computing node, establishing an appropriate network topology structure, dynamically adjusting the allocation of network computing resources, and achieving load balancing.

Benefits of technology

It realizes load balancing in heterogeneous hardware environments, improves computing performance, reduces cross-layer communication delays and overhead, and improves the overall communication efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576535B_ABST
    Figure CN119576535B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for calculating a target task through a supercomputer, comprising: obtaining topological information of a computing network corresponding to the supercomputer, the computing network comprising N network nodes, the N network nodes comprising M computing nodes and L switching nodes, and the topological information comprising associations between the network nodes; determining a problem domain represented by a first multivariate array corresponding to a given target computing task; determining a computing power weight corresponding to each computing node according to the computing performance of each computing node, dividing the problem domain into M subdomains according to the topological information and the computing power weight of each computing node, the M subdomains respectively corresponding to the M computing nodes; obtaining a computing result of a subdomain corresponding to each computing node through each computing node, and determining a computing result of the target computing task according to the computing result of each subdomain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high performance computing, and in particular to a method and device for computing a target task by a supercomputer. Background Art

[0002] Computational fluid dynamics (CFD) is a scientific technology that studies and simulates fluid motion. It is widely used in aerospace, automotive design, meteorology, marine engineering and other fields. CFD calculations usually involve solving large-scale differential equations. Such calculations require dividing the computational domain into multiple sub-areas and assigning these sub-areas to different computing nodes for execution to achieve parallel computing.

[0003] With the rapid growth of computing power demand in scientific computing, engineering simulation, climate modeling and other fields, supercomputers have become a key tool for solving complex computing problems. These computers can handle large-scale data and computing tasks in a short period of time through highly parallel processing capabilities. However, since the construction and upgrading of supercomputers requires a lot of capital and technical investment, many countries and scientific research institutions often choose to build in phases, that is, gradually increase computing nodes and hardware resources to cope with the ever-changing computing needs.

[0004] Supercomputers built in phases are usually configured with limited computing resources in their initial stages. As time goes by, new computing nodes and hardware modules are gradually added to the system. Such architectural characteristics cause the hardware topology of supercomputers to change at different times, resulting in the heterogeneity of their parallel computing environment. How to make full use of existing resources and improve computing efficiency in this gradually changing hardware environment is a major challenge for large-scale parallel applications such as computational fluid dynamics (CFD).

[0005] Therefore, a new method and device for calculating target tasks through a supercomputer is needed. Summary of the invention

[0006] The purpose of the present invention is to provide a method, electronic device and computer storage medium for calculating target tasks through a supercomputer, establish an appropriate network topology structure through the applied network nodes, thereby adjusting the allocation of network computing resources according to the computing tasks, realizing load balancing in the supercomputer system, and improving computing performance.

[0007] To achieve the above object, in a first aspect, the present invention provides a method for calculating a target task by a supercomputer, comprising:

[0008] Acquire topological information of a computing network corresponding to a supercomputer, the computing network including N network nodes, the N network nodes including M computing nodes and L switching nodes, the topological information including associations between the network nodes; determine a problem domain represented by a first multivariate array corresponding to a given target computing task;

[0009] Determine the computing power weight corresponding to each computing node according to the computing performance of each computing node, and divide the problem domain into M subdomains according to the topology information and the computing power weight of each computing node, wherein the M subdomains correspond to the M computing nodes respectively;

[0010] The calculation results of the subdomains corresponding to the various computing nodes are obtained through the various computing nodes, and the calculation results of the target computing tasks are determined according to the calculation results of the various subdomains.

[0011] Preferably, according to the topology information and the computing power weight of each computing node, the problem domain is divided into M subdomains, including:

[0012] According to the topology information and the computing power weight of each computing node, the computing power weight of each switching node is determined; according to the computing power weight of each switching node and the computing power weight of each computing node, the problem domain is divided into M subdomains.

[0013] Specifically, the computing network is a tree network of at least 2 layers, the root node of the tree network is a switching node, the leaf nodes in the tree network are computing nodes, and the network nodes on the path between the leaf nodes and the root node in the tree network are switching nodes;

[0014] Determining the computing power weight of each switching node according to the topology information and the computing power weight of each computing node includes:

[0015] Determine the computing power weight of each switching node according to the computing power weight of the lower-layer computing nodes of each switching node;

[0016] According to the computing power weight of each switching node and the computing power weight of each computing node, the problem domain is divided into M subdomains, and the M subdomains correspond to the M computing nodes respectively, including:

[0017] The problem domain is divided multiple times in order from the top layer to the bottom layer of the tree network, wherein the tree network includes, for the i-th layer, the i-th division of the problem domain according to the i-th layer includes determining the part of the problem domain to be divided in this layer, and according to the total number of switching nodes and computing nodes included in the i-th layer, the computing power weight of each switching node in the i-th layer, and the computing nodes, dividing the part to be divided into multiple partition domains corresponding to each switching node and each computing node respectively, the partition domain corresponding to the switching node is used as the part to be divided in the i+1th layer, and the partition domain corresponding to the computing node is used as the subdomain.

[0018] Preferably, the weight of the computing node is determined according to the operating status and computing resources of the computing node.

[0019] Specifically, the weight of the computing node is determined according to the operating status and computing resources of the computing node, including: if the operating status of the computing node is normal, the weight of the computing node is determined according to the computing resources of the computing node; if the operating status of the computing node is a complete failure, the weight of the computing node is set to 0; if the operating status of the computing node is a partial failure, the weight of the computing node is determined according to the actual computing resources of the computing node.

[0020] Specifically, the multi-element array is a ternary array.

[0021] Specifically, the target computing task is a fluid mechanics computing task.

[0022] In a second aspect, the present invention provides an apparatus for calculating a target task by a supercomputer, comprising:

[0023] The acquisition unit is configured to acquire topological information of a computing network corresponding to the supercomputer, wherein the computing network includes N network nodes, wherein the N network nodes include M computing nodes and L switching nodes, and the topological information includes associations between the network nodes; determine a problem domain represented by a first multivariate array corresponding to a given target computing task;

[0024] The processing unit is configured to determine the computing power weight corresponding to each computing node according to the computing performance of each computing node, and divide the problem domain into M subdomains according to the topology information and the computing power weight of each computing node, wherein the M subdomains correspond to the M computing nodes respectively;

[0025] The determination unit is configured to obtain, through each computing node, a calculation result of a subdomain corresponding to each computing node, and determine a calculation result of a target computing task according to the calculation result of each subdomain.

[0026] In a third aspect, the present invention provides an electronic device, comprising:

[0027] at least one memory for storing a program;

[0028] At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method provided in the first aspect.

[0029] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed on an electronic device, enables the electronic device to execute the method provided in the first aspect.

[0030] Compared with the prior art, the present invention has the following advantages: the present invention can balance the workload of computing nodes, dynamically adjust according to the computing power and communication capabilities of different nodes, avoid uneven distribution of computing tasks or the generation of communication bottlenecks, and better adapt to complex supercomputing network structures, effectively reduce the delay and overhead of cross-layer communication, thereby improving the overall communication efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of basic steps for calculating a target task by a supercomputer provided by an embodiment of the present invention;

[0032] Figure 2 A schematic diagram of an overall process of calculating a target task by a supercomputer provided by an embodiment of the present invention;

[0033] Figure 3 A schematic diagram of a tree network with three layers is provided for a computing network according to an embodiment of the present invention;

[0034] Figure 4 A schematic diagram of a structure for dividing a computing domain by network topology provided by an embodiment of the present invention;

[0035] Figure 5 A schematic diagram of a framework of a topology-aware mapping solution provided in an embodiment of the present invention;

[0036] Figure 6 A structural diagram of a device for calculating a target task by a supercomputer provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments.

[0038] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be described below in conjunction with the accompanying drawings. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0039] In the description of the embodiments of the present invention, words such as "exemplary", "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a concrete way.

[0040] Computational fluid dynamics (CFD) is a scientific technology that studies and simulates fluid motion. It is widely used in aerospace, automotive design, meteorology, marine engineering and other fields. CFD calculations usually involve solving large-scale partial differential equations. Such calculations require dividing the computational domain into multiple sub-regions and assigning these sub-regions to different computing nodes for execution to achieve parallel computing. In traditional CFD applications, domain decomposition and process mapping are usually based on the following assumptions:

[0041] Fixed hardware topology: It is assumed that the hardware structure of the supercomputer is static, that is, the number of computing nodes and the topology of the communication network will not change during the entire computing process.

[0042] Uniform computing power: It is assumed that all computing nodes have the same computing power and memory capacity, and there is no heterogeneity between nodes.

[0043] Under this assumption, common domain decomposition methods include geometric decomposition, load-based decomposition, etc. These methods often perform well in static hardware environments, but on supercomputers built in phases, since the number of computing nodes and communication topology may change at any time, traditional methods often cannot adapt to dynamically changing environments, leading to the following problems:

[0044] Low resource utilization: Due to changes in hardware topology, some nodes may be underloaded, while other nodes may be overloaded, making it impossible to effectively utilize all computing resources. Increased communication overhead: As the hardware topology changes, the original process mapping method may lead to increased communication between remote nodes, significantly increasing communication overhead and reducing overall computing efficiency. Degraded computing performance: Due to the inability to dynamically adjust the region decomposition and process mapping strategies, computing performance significantly decreases when the hardware structure changes, making it difficult to meet efficient computing needs.

[0045] Existing domain decomposition and process mapping methods often assume that the topology of supercomputers is static, fail to adapt to changes in hardware computing power after phased construction, fail to fully utilize computing resources, resulting in low computing efficiency, and cannot meet the needs of CFD applications.

[0046] In order to overcome the shortcomings of the topology-aware region decomposition and process mapping strategies in the prior art, a method for computing target tasks by supercomputers was proposed. Figure 1 A schematic diagram of the basic steps of calculating a target task by a supercomputer is provided in an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0047] S102: first, topological information of a computing network corresponding to a supercomputer is obtained, wherein the computing network includes N network nodes, wherein the N network nodes include M computing nodes and L switching nodes, and the topological information includes associations between the network nodes; and a problem domain represented by a first multivariate array corresponding to a given target computing task is determined;

[0048] S104: determining a computing power weight corresponding to each computing node according to the computing performance of each computing node, and dividing the problem domain into M subdomains according to the topology information and the computing power weight of each computing node, wherein the M subdomains correspond to the M computing nodes respectively;

[0049] S106: Obtain the calculation results of the subdomains corresponding to the respective calculation nodes through the respective calculation nodes, and determine the calculation results of the target calculation task according to the calculation results of the respective subdomains.

[0050] First, in step S102, the hardware information of each computing node is obtained through the system administrator interface or by using a hardware information collection tool such as hwloc. The information includes but is not limited to the number of CPU cores, memory size, network bandwidth and storage capacity, and node topology information is generated based on the information.

[0051] For example, Figure 2 FIG. 1 is a schematic diagram of an overall process of calculating a target task by a supercomputer according to an embodiment of the present invention. Figure 2 As shown in the figure, after the user submits a resource application request to the job management system, the system builds a hierarchical network topology of the entire computing network through network topology exploration tools such as InfiniBand Fabric Manager, or by collecting node topology information returned by the job management system. The network node information in the computing network will be recorded and divided into a tree network.

[0052] After obtaining the node topology information, the system will assign a weight to each network node. The weight of a computing node is generally 1. If it is a high-performance computing node, the weight is set according to its computing resources, and the weight of a switching node is assigned according to the number of computing nodes it manages. The algorithm will calculate the weight distribution of the entire computing network based on the connection relationship between computing nodes and switching nodes.

[0053] Secondly, in step S104, exemplarily, the problem domain can be divided into M subdomains according to the topology information and the computing power weight of each computing node, including: determining the computing power weight of each switching node according to the topology information and the computing power weight of each computing node; dividing the problem domain into M subdomains according to the computing power weight of each switching node and the computing power weight of each computing node.

[0054] Figure 3 A computing network provided in an embodiment of the present invention is a tree-like network diagram of three layers, such as Figure 3 As shown, the computing network is a 3-layer tree network, the root node of the tree network is a switching node, the leaf nodes in the tree network are computing nodes, and the network nodes on the path between the leaf nodes and the root node in the tree network are switching nodes;

[0055] Figure 4 A schematic diagram of a structure of dividing a computing domain by network topology provided by an embodiment of the present invention, such as Figure 4 As shown, according to the topology information and the computing power weight of each computing node, the computing power weight of each switching node is determined, including: according to the computing power weight of the lower computing node of each switching node, the computing power weight of each switching node is determined; according to the computing power weight of each switching node and the computing power weight of each computing node, the problem domain is divided into M subdomains, and the M subdomains correspond to the M computing nodes respectively, including: according to the order of the tree network from the top layer to the bottom layer, the problem domain is divided multiple times, wherein the tree network includes for the i-th layer, the i-th division of the problem domain according to the i-th layer includes determining the part of the problem domain to be divided in this layer, according to the total number of switching nodes and computing nodes included in the i-th layer, the computing power weight of each switching node in the i-th layer and the computing node, the part to be divided is divided into multiple division domains corresponding to each switching node and each computing node respectively, the division domain corresponding to the switching node is used as the part to be divided of the i+1th layer, and the division domain corresponding to the computing node is used as the subdomain.

[0056] Exemplarily, the largest unary array in the first multi-element array is selected and divided according to the computing power weight of the top-level network node to generate multiple sub-domains represented by the second multi-element array, one of the sub-domains represented by the second multi-element array is selected, and it is determined whether the network node corresponding to the sub-domain is a computing node. If not, the sub-domain is divided according to the computing power weight of the network node below the network node corresponding to the sub-domain. When the sub-domain is divided to the corresponding computing node, it ends, and the problem domain is divided into M sub-domains. The M computing nodes correspond to the M sub-domains respectively and calculate the target tasks of the sub-domains.

[0057] For example, Figure 5 A schematic diagram of a framework of a topology-aware mapping solution provided by an embodiment of the present invention. Figure 5 As shown in the figure, A is the topology-aware node-to-node process mapping, and B is the top-down problem domain decomposition process. The system obtains the problem domain of the target task input by the user, and then recursively divides the problem domain according to the weight information of the network nodes. The specific division is as follows: The end condition of the recursion is that when the network node to be divided is a computing node, the algorithm directly returns the size of the current subdomain and the corresponding weight, as a separate subdomain, without further division. For each division, it is necessary to first identify the largest dimension of each dimension of the current subdomain to be divided as the target dimension for division. For example, in an initial three-dimensional problem domain with a size of (Global nx ,Global ny ,Global nz )Global nx is the largest, then the x dimension is selected as the partition dimension. Then the partition point is calculated. The system first calculates the sum of the weights W corresponding to all network nodes. total , and then calculate the cumulative weight W of the first half of the weight array half , by calculating The system calculates the appropriate division point and converts the maximum dimension Global nx The division is performed at this point to generate two new subdomain sizes. The other dimensions (Global ny ,Global nz ) remains unchanged, so that the problem domain is divided as evenly as possible according to the network node weight ratio. For each newly generated subdomain, recursively repeat the above steps for further division. Each time the network nodes and weight arrays are divided into two parts in proportion, they are assigned to two new subdomains, and the recursive division continues until each subdomain contains only one computing node.

[0058] Finally, in step S106, after the division is completed, the division results of all subdomains are combined to form the final complete division domain. The size and corresponding weight of each subdomain will be used as the final output. The system will map each subdomain to the corresponding computing node and assign specific computing tasks to each computing node. The target task will be processed based on the data in the problem domain and communicated through switches in the network layer.

[0059] Exemplarily, the weight of the computing node is determined according to the operating state and computing resources of the computing node. Specifically, the weight of the computing node is determined according to the operating state and computing resources of the computing node, including: if the operating state of the computing node is normal, the weight of the computing node is determined according to the computing resources of the computing node; if the operating state of the computing node is a complete failure, the weight of the computing node is set to 0; if the operating state of the computing node is a partial failure, the weight of the computing node is determined according to the actual computing resources of the computing node.

[0060] In different embodiments, the specific structure of the multivariate data may be different. In a specific embodiment, the multivariate array may be a ternary array.

[0061] In different embodiments, the target task may be different specific computing tasks. In a specific embodiment, the target task is a fluid mechanics computing task.

[0062] In order to understand the present invention more clearly, an embodiment is given, and the calculation domain size of this embodiment is (Global nx ,Global ny ,Global nz )=(1024,512,256), it needs to be divided into 8 sub-domains and tasks are allocated according to the load capacity of the computing nodes. First, initialize the input parameters: input the computing domain size (Global nx ,Global ny ,Global nz )=(1024,512,256); Set the number of subdomains to be divided m=8; Input the subdomain weight array W={1,1,1,1,1,1,2,1} obtained from the hardware information, where each weight corresponds to the computing power of a subdomain. Next, identify the maximum dimension: The system identifies the Global nx is the maximum dimension with a length of 1024. This dimension is selected for the initial partition. Then the partition point is calculated: the system calculates the sum of the weights W total =9, and calculate the cumulative weight W of the first half of the weight array half =4; according to W half / W total=4 / 9, the system determines the division point as the 456th unit on the maximum dimension, thereby dividing the computing domain into two new subdomains of size (456,512,256) and (568,512,256). Recursive division is then performed: for each new subdomain, the remaining dimensions and weights are recursively divided according to the above steps until all subdomains are allocated. The final result output: the generated 8 subdomains correspond to computing nodes of size (228,256,256), (228,256,256), (228,256,256), (228,256,256), (228,256,256), (228,256,256), (340,341,256), (340,171,256). Each subdomain is allocated to the computing node with the corresponding load capacity according to the weight and the result is executed to ensure load balancing. Through this embodiment, the system divides the original computing domain into 8 subdomains and distributes them reasonably according to the weights. This method effectively reduces the cross-node communication overhead and improves the overall efficiency of parallel computing. After performance testing, the computing time is reduced by about 15% compared with the traditional uniform division.

[0063] According to yet another embodiment, a device for calculating a target task by a supercomputer is provided. Figure 6 A structural diagram of a device for calculating a target task by a supercomputer provided by an embodiment of the present invention, such as Figure 6 As shown, the device 600 includes:

[0064] The acquisition unit 61 is configured to acquire topology information of a computing network corresponding to the supercomputer, the computing network including N network nodes, the N network nodes including M computing nodes and L switching nodes, and the topology information including associations between the network nodes; determine a problem domain represented by a first multi-element array corresponding to a given target computing task;

[0065] The processing unit 62 is configured to determine the computing power weight corresponding to each computing node according to the computing performance of each computing node, and divide the problem domain into M subdomains according to the topology information and the computing power weight of each computing node, wherein the M subdomains correspond to the M computing nodes respectively;

[0066] The determination unit 63 is configured to obtain, through each computing node, the calculation result of the subdomain corresponding to each computing node, and determine the calculation result of the target computing task according to the calculation result of each subdomain.

[0067] It is understood that the method steps in the embodiments of the present invention can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0068] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)), etc.

[0069] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for calculating a target task by a supercomputer, comprising: Acquire topology information of a computing network corresponding to a supercomputer, wherein the computing network includes N network nodes, wherein the N network nodes include M computing nodes and L switching nodes, and the topology information includes associations between the network nodes; Determine a problem domain represented by a first multi-element array corresponding to a given target computing task; According to the computing performance of each computing node, determine the computing power weight corresponding to each computing node, and according to the topology information and the computing power weight of each computing node, divide the problem domain into M subdomains, and the M subdomains correspond to the M computing nodes respectively; wherein, according to the topology information and the computing power weight of each computing node, divide the problem domain into M subdomains, including: according to the topology information and the computing power weight of each computing node, determine the computing power weight of each switching node; according to the computing power weight of each switching node and the computing power weight of each computing node, divide the problem domain into M subdomains; the computing network is a tree network of at least 2 layers, the root node of the tree network is a switching node, the leaf nodes in the tree network are computing nodes, and the network nodes on the path between the leaf nodes and the root node in the tree network are switching nodes; according to the topology information and the computing power weight of each computing node, determine the computing power weight of each switching node , including: determining the computing power weight of each switching node according to the computing power weight of the lower computing node of each switching node; dividing the problem domain into M subdomains according to the computing power weight of each switching node and the computing power weight of each computing node, wherein the M subdomains correspond to the M computing nodes respectively, including: dividing the problem domain multiple times in the order from the top layer to the bottom layer of the tree network, wherein the tree network includes, for the i-th layer, the i-th division of the problem domain according to the i-th layer includes determining the part of the problem domain to be divided in this layer, and dividing the part to be divided into multiple division domains corresponding to each switching node and each computing node respectively according to the total number of switching nodes and computing nodes included in the i-th layer, the computing power weight of each switching node in the i-th layer, and the computing nodes, wherein the division domain corresponding to the switching node is used as the part to be divided in the i+1th layer, and the division domain corresponding to the computing node is used as the subdomain; The calculation results of the subdomains corresponding to the respective calculation nodes are obtained through the respective calculation nodes, and the calculation results of the target calculation task are determined according to the calculation results of the respective subdomains, wherein the target calculation task is a fluid mechanics calculation task.

2. The method according to claim 1, wherein: The weight of the computing node is determined according to the running state and computing resources of the computing node.

3. The method according to claim 2, wherein: The weight of the computing node is determined according to the operating status and computing resources of the computing node, including: if the operating status of the computing node is normal, the weight of the computing node is determined according to the computing resources of the computing node; if the operating status of the computing node is a complete failure, the weight of the computing node is set to 0; if the operating status of the computing node is a partial failure, the weight of the computing node is determined according to the actual computing resources of the computing node.

4. The method according to claim 1, wherein: The multi-element array is a three-element array.

5. A device for calculating a target task by a supercomputer, comprising: An acquisition unit is configured to acquire topology information of a computing network corresponding to the supercomputer, wherein the computing network includes N network nodes, wherein the N network nodes include M computing nodes and L switching nodes, and the topology information includes associations between the network nodes; Determine a problem domain represented by a first multi-element array corresponding to a given target computing task; The processing unit is configured to determine the computing power weight corresponding to each computing node according to the computing performance of each computing node, and divide the problem domain into M subdomains according to the topology information and the computing power weight of each computing node, and the M subdomains correspond to the M computing nodes respectively; wherein, dividing the problem domain into M subdomains according to the topology information and the computing power weight of each computing node includes: determining the computing power weight of each switching node according to the topology information and the computing power weight of each computing node; dividing the problem domain into M subdomains according to the computing power weight of each switching node and the computing power weight of each computing node; the computing network is a tree network of at least 2 layers, the root node of the tree network is a switching node, the leaf nodes in the tree network are computing nodes, and the network nodes on the path between the leaf nodes and the root node in the tree network are switching nodes; determining the computing power weight of each switching node according to the topology information and the computing power weight of each computing node The computing power weight includes: determining the computing power weight of each switching node according to the computing power weight of the lower computing node of each switching node; dividing the problem domain into M subdomains according to the computing power weight of each switching node and the computing power weight of each computing node, wherein the M subdomains correspond to the M computing nodes respectively, including: dividing the problem domain multiple times in the order from the top layer to the bottom layer of the tree network, wherein the tree network includes, for the i-th layer, the i-th division of the problem domain according to the i-th layer includes determining the part of the problem domain to be divided in this layer, and dividing the part to be divided into multiple division domains corresponding to each switching node and each computing node respectively according to the total number of switching nodes and computing nodes included in the i-th layer, the computing power weight of each switching node in the i-th layer, and the computing nodes, wherein the division domain corresponding to the switching node is used as the part to be divided in the i+1th layer, and the division domain corresponding to the computing node is used as the subdomain; The determination unit is configured to obtain the calculation results of the subdomains corresponding to each computing node through each computing node, and determine the calculation results of the target computing task according to the calculation results of each subdomain, wherein the target computing task is a fluid mechanics computing task.

6. An electronic device comprising: A processor, a memory, and computer program instructions stored in the memory and executable on the processor, wherein the processor is used to implement the method according to any one of claims 1 to 4 when executing the computer program instructions.

7. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 4 when executed by a processor.

Citation Information

Patent Citations

  • Flow field modeling method based on discrete invariance grid convolution operator

    CN114818462A

  • Computing power routing method and device, electronic equipment and storage medium

    CN115412482A