Network architecture, network architecture generation method and related device

By connecting the computing node with the farthest communication distance in the Torus network, the Diagonal Torus network is formed, which solves the problem of limited communication performance of Torus network, and achieves higher throughput, fault tolerance and simplified routing settings.

CN120378311APending Publication Date: 2025-07-25HUAWEI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202410104640.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

As the Torus network scale increases, communication performance is limited, communication delay is long, and the All to all bandwidth is damaged. The network architecture does not have point symmetry, resulting in complex routing settings.

Method used

Based on the Torus network, additionally connect the target computing nodes with the farthest communication distance between computing nodes to form a Diagonal Torus network, reducing communication delay and increasing bandwidth.

Benefits of technology

It improves the throughput and fault tolerance of computing nodes, reduces communication delay, simplifies network routing settings, and improves overall communication performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378311A_ABST
    Figure CN120378311A_ABST
Patent Text Reader

Abstract

The invention discloses a network architecture, a network architecture generation method and a related device, and belongs to the technical field of communication. The network architecture comprises a plurality of computing nodes, the plurality of computing nodes are in communication connection according to a Torus networking mode, and a first computing node in the plurality of computing nodes is also in communication connection with a target computing node; wherein the first computing node is any one computing node in the plurality of computing nodes, and the target computing node is the computing node which has the farthest communication distance with the first computing node under the condition that the plurality of computing nodes are in communication connection according to the Torus networking mode. Therefore, compared with a Torus network, the network architecture provided by the invention not only increases the degree of each computing node, improves the throughput and fault tolerance of the computing nodes, but also reduces the communication delay between the computing nodes, thereby improving the overall communication performance of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a network architecture, a method for generating a network architecture, and related devices. Background Art

[0002] As a direct network architecture, the Torus network shortens the communication distance between computing nodes in the network and improves communication performance by adding loopback links on the basis of a Mesh network.

[0003] As Figure 1 shown, the Torus network can be represented by an N-dimensional grid, and the communication connections between computing nodes in the Torus network are implemented by establishing direct connections between adjacent computing nodes in each dimension (abbreviated as D). For example, in a two-dimensional Torus (abbreviated as 2D Torus) network, each computing node is connected to the four computing nodes adjacent to it above, below, left, and right. Since the Torus network is a network topology with node symmetry, as long as the dimension of the Torus network and the number of nodes in each dimension are determined, its multi-node networking method can be uniquely determined.

[0004] However, as the network scale increases, the communication performance of the Torus network is relatively limited. Therefore, there is an urgent need for a new network architecture to further improve the communication performance such as the node communication delay and network bandwidth of the Torus network. Summary of the Invention

[0005] This application provides a network architecture, a method for generating a network architecture, and related devices, which can significantly improve the communication performance between computing nodes compared with the Torus network. The technical solutions are as follows:

[0006] In a first aspect, a network architecture is provided. The network architecture includes multiple computing nodes. The multiple computing nodes are communicatively connected according to the Torus networking method, and a first computing node among the multiple computing nodes is further communicatively connected to a target computing node; wherein, the first computing node is any one of the multiple computing nodes, and the target computing node is the computing node with the farthest communication distance from the first computing node when the multiple computing nodes are communicatively connected according to the Torus networking method.

[0007] Optionally, the Torus networking mode includes at least one dimension. If the total number of computing nodes in each dimension of the at least one dimension is even, the target computing node includes one computing node, and the coordinate value of the one computing node in the first dimension is the first coordinate value, where the first coordinate value is obtained by taking the modulus of d1 + k1 / 2 with respect to k1; wherein, the first dimension is any one of the at least one dimension, d1 is the coordinate value of the first computing node in the first dimension, and k1 is the total number of computing nodes in the first dimension.

[0008] Optionally, the Torus networking mode includes at least one dimension. If there is a dimension in the at least one dimension where the total number of computing nodes is odd, the target computing node includes at least two computing nodes; when the total number of computing nodes in the first dimension is even, the coordinate values of the at least two computing nodes in the first dimension are both the first coordinate value, where the first coordinate value is obtained by taking the modulus of d1 + k1 / 2 with respect to k1; when the total number of computing nodes in the first dimension is odd, the coordinate values of the first part of the at least two computing nodes in the first dimension are the second coordinate value, and the coordinate values of the second part of the at least two computing nodes in the first dimension are the third coordinate value, and the number of the first part of the computing nodes is equal to the number of the second part of the computing nodes, the second coordinate value is obtained by taking the modulus of d1 + (k1 - 1) / 2 with respect to k1, and the third coordinate value is obtained by taking the modulus of d1 + (k1 + 1) / 2 with respect to k1; wherein, the first dimension is any one of the at least one dimension, d1 is the coordinate value of the first computing node in the first dimension, and k1 is the total number of computing nodes in the first dimension.

[0009] Optionally, if the total number of computing nodes in each dimension of the at least one dimension is odd, the at least two computing nodes include a second computing node and a third computing node; the coordinate value of the second computing node in the first dimension is the second coordinate value, and the coordinate value in the second dimension is the fourth coordinate value, where the fourth coordinate value is obtained by taking the modulus of d2 + (k2 - 1) / 2 with respect to k2; the coordinate value of the third computing node in the first dimension is the third coordinate value, and the coordinate value in the second dimension is the fifth coordinate value, where the fifth coordinate value is obtained by taking the modulus of d2 + (k2 + 1) / 2 with respect to k2; wherein, the second dimension is any one of the at least one dimension other than the first dimension, d2 is the coordinate value of the first computing node in the second dimension, and k2 is the total number of computing nodes in the second dimension.

[0010] It can be seen that after multiple computing nodes in the network architecture of the present application are communicatively connected according to the Torus networking mode, each computing node is further communicatively connected to at least one computing node with the farthest communication distance in the Torus networking mode. Therefore, compared with the Torus network, the network architecture shown in the embodiments of the present application not only increases the degree of each computing node, improves the throughput and fault tolerance of the computing node, but also reduces the communication delay between computing nodes, thereby improving the overall communication performance of the network.

[0011] Optionally, the network architecture further includes at least one network device, and the first computing node and the target computing node are communicatively connected through the at least one network device.

[0012] Optionally, the first computing node and the target computing node are communicatively connected through the same network device.

[0013] Optionally, the multiple computing nodes are formed into multiple Torus networks with the same topology according to the Torus networking mode, and the computing nodes at the same position in the multiple Torus networks are connected to the same network device.

[0014] In summary, when the network architecture includes one network device, the computing nodes in the multiple Torus networks are all connected to the same network device, so as to ensure that all the computing nodes in the multiple Torus networks are fully connected through this network device. When the network architecture includes multiple network devices, some computing nodes in each Torus network are connected to the same network device, so as to ensure that all the computing nodes in the multiple Torus networks are fully connected through the multiple network devices. That is, when the network architecture includes multiple Torus networks formed by multiple computing nodes and at least one network device, the at least one network device can be used to implement the communication connection between the computing nodes with the farthest communication distance in each Torus network, thereby forming a Diagonal Torus network and reducing the communication delay between the computing nodes in the network architecture.

[0015] In a second aspect, a circuit is provided, the circuit includes multiple computing nodes, and the multiple computing nodes form the network architecture as described in the first aspect above.

[0016] In a third aspect, a cabinet is provided, the cabinet includes multiple computing nodes and at least one network device, and the multiple computing nodes and the at least one network device form the network architecture as described in the first aspect above.

[0017] In a fourth aspect, a computing cluster is provided. The computing cluster includes a plurality of racks, each of the racks includes a plurality of computing nodes and at least one network device, and the plurality of computing nodes and the at least one network device form the network architecture as described in the first aspect above.

[0018] In a fifth aspect, a method for generating a network architecture is provided. The network architecture includes a plurality of computing nodes and at least one network device. The plurality of computing nodes are formed into at least one Torus network with the same topology according to the Torus networking method. The method includes: receiving network resource configuration information, where the network resource configuration information includes a target topology type, and the target topology type indicates a target network topology currently expected to be generated. The two computing nodes with the farthest communication distance in the target network topology are communicatively connected through the at least one network device; determining routing information between each of the computing nodes that form the target network topology among the plurality of computing nodes according to the target topology type; and sending the target topology type and the routing information to each of the computing nodes.

[0019] Optionally, when the plurality of computing nodes are formed into a plurality of Torus networks with the same topology according to the Torus networking method and the network architecture includes a plurality of network devices, the computing nodes at the same position in the plurality of Torus networks are connected to the same network device.

[0020] Optionally, the two computing nodes with the farthest communication distance in the same Torus network are communicatively connected through the same network device, and the routing information includes the shortest route between each of the computing nodes.

[0021] Optionally, before determining the routing information between each of the computing nodes that form the target network topology among the plurality of computing nodes according to the target topology type, the method further includes: if the target topology type is a network topology type supported by the network architecture, then performing the step of determining the routing information between each of the computing nodes that form the target network topology among the plurality of computing nodes according to the target topology type.

[0022] Optionally, the method further includes: determining the network topology types supported by the network architecture based on the connection relationship between the plurality of computing nodes and the at least one network device.

[0023] It can be seen that for the connection relationship between multiple computing nodes and at least one network device in a network architecture, when multiple network topology types are supported, the manager determines the communication routes between the computing nodes that make up the target network topology in the network architecture according to the target topology type included in the network resource configuration information. In this way, based on the same network architecture, by generating the target network topology, the application scenarios of the network architecture are improved to meet the requirements of different network sizes and the number of computing nodes required for different data processing tasks. Moreover, since the target network topology is a Diagonal Torus network constructed based on the networking method of the embodiments of the present application, the computing nodes can communicate according to the target network topology and routing information, which can reduce the communication delay during multi-node communication.

[0024] In a sixth aspect, a method for generating a network architecture is provided. The network architecture includes multiple computing nodes and at least one network device. The multiple computing nodes are organized into at least one Torus network with the same topology according to the Torus networking method. The method includes: receiving a target topology type and routing information. The target topology type indicates a target network topology expected to be generated based on the network architecture. The two computing nodes with the farthest communication distance in the target network topology are communicatively connected through the at least one network device. The routing information indicates the communication routes between the computing nodes that make up the target network topology.

[0025] It can be seen that for the computing nodes in the network architecture, after receiving the target topology type and routing information, the target topology type and routing information can be stored, so that when it is necessary to form a network according to the target network topology and process computing tasks, the routing information can be referred to for communication with at least one computing node within the target network topology to transmit relevant data, information, or instructions.

[0026] In a seventh aspect, a device for generating a network architecture is provided. The device for generating a network architecture has the function of implementing the behavior of the method for generating a network architecture in the fifth aspect above. The device for generating a network architecture includes at least one module, and the at least one module is used to implement the method for generating a network architecture provided in the fifth aspect above.

[0027] In an eighth aspect, a device for generating a network architecture is provided. The device for generating a network architecture has the function of implementing the behavior of the method for generating a network architecture in the sixth aspect above. The device for generating a network architecture includes at least one module, and the at least one module is used to implement the method for generating a network architecture provided in the sixth aspect above.

[0028] In a ninth aspect, an electronic device is provided. The electronic device includes a processor and a memory. The memory is used to store a computer program for executing the method for generating the network architecture provided in the fifth aspect or the sixth aspect. The processor is configured to execute the computer program stored in the memory to implement the method for generating the network architecture described in the fifth aspect or the sixth aspect.

[0029] Optionally, the electronic device may further include a communication bus, which is used to establish a connection between the processor and the memory.

[0030] In a tenth aspect, a computer-readable storage medium is provided. Instructions are stored in the storage medium. When the instructions are run on a computer, the computer is caused to execute the method for generating the network architecture described in the fifth aspect or the sixth aspect.

[0031] In an eleventh aspect, a computer program product containing instructions is provided. When the instructions are run on a computer, the computer is caused to execute the method for generating the network architecture described in the fifth aspect or the sixth aspect. Or rather, a computer program is provided. When the computer program is run on a computer, the computer is caused to execute the method for generating the network architecture described in the fifth aspect or the sixth aspect.

[0032] The technical effects obtained in the second aspect, the third aspect, and the fourth aspect are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be elaborated here; similarly, the technical effects obtained in the seventh aspect to the eleventh aspect are similar to the technical effects obtained by the corresponding technical means in the fifth aspect and the sixth aspect, and will not be elaborated here either. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the network architecture of a Torus network provided by an embodiment of the present application;

[0034] Figure 2 is a schematic diagram of the network architecture of a Tofu network provided by an embodiment of the present application;

[0035] Figure 3 is a schematic diagram of the network architecture of a Twisted Torus network provided by an embodiment of the present application;

[0036] Figure 4 is a schematic diagram of the network architecture of a Diagonal 1D Torus network provided by an embodiment of the present application;

[0037] Figure 5 is a schematic diagram of the network architecture of a Diagonal 2D Torus network provided by an embodiment of the present application;

[0038] Figure 6 It is a schematic diagram of the network architecture of a Diagonal 3D Torus network provided by an embodiment of the present application;

[0039] Figure 7 It is a schematic diagram of the network architecture of another Diagonal 1D Torus network provided by an embodiment of the present application;

[0040] Figure 8 It is a schematic diagram of the network architecture of another Diagonal 2D Torus network provided by an embodiment of the present application;

[0041] Figure 9 It is a schematic diagram of the connection method of a computing node of a Diagonal 2D Torus network provided by an embodiment of the present application;

[0042] Figure 10 It is a schematic diagram of the connection method of multi-node networking provided by an embodiment of the present application;

[0043] Figure 11 It is a schematic diagram of another connection method of multi-node networking provided by an embodiment of the present application;

[0044] Figure 12 It is a schematic diagram of yet another connection method of multi-node networking provided by an embodiment of the present application;

[0045] Figure 13 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application;

[0046] Figure 14 It is a schematic diagram of the flowchart of a method for generating a network architecture provided by an embodiment of the present application;

[0047] Figure 15 It is a schematic diagram of the structure of a device for generating a network architecture provided by an embodiment of the present application. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0049] For ease of understanding, before providing a detailed explanation of the network architecture provided by the embodiments of the present application, the terms and application backgrounds involved in the embodiments of the present application will be introduced first.

[0050] First, the terms involved in the embodiments of the present application will be introduced.

[0051] 1. Torus network

[0052] Torus is a directly connected network topology architecture that adds loopback links to a Mesh network to shorten the communication distance between computing nodes in the network and improve communication performance.

[0053] Among them, the number of computing nodes in each dimension (abbreviated as D) in the Torus network is called the radix of the Torus network. For example, if there are M computing nodes in both the X dimension and the Y dimension in the Torus network, then the radix of the Torus network is M, and M is an integer greater than 1.

[0054] 2. Vertex-transitive graph (also known as vertex transitive graph)

[0055] Each point in a vertex-transitive graph has the same local interconnection environment. Based on this, when the overall architecture of a certain network satisfies the vertex-transitive graph, each computing node in this network also has the same local interconnection environment. Therefore, it is relatively easy to design the communication routes between computing nodes.

[0056] 3. Cayley graph (also known as Cayley graph)

[0057] The Cayley graph is also a type of vertex-transitive graph, and the Torus network architecture belongs to the Cayley graph.

[0058] 4. All to all

[0059] A many-to-many collective communication mode that often appears in artificial intelligent (AI) training. In the All to all operation, the data of each computing node is scattered to all computing nodes in the cluster, and at the same time, each computing node also aggregates the data of all computing nodes in the cluster. In other words, All to all is a full exchange operation that allows each computing node to obtain the values of other computing nodes.

[0060] Among them, the specific operation of All to all is: sending the j-th block of data of computing node i to computing node j, and computing node j places the data block received from computing node i in its i-th position.

[0061] 5. All reduce

[0062] A many-to-many collective communication mode that often appears in AI training. All reduce is a collective term for a series of simple arithmetic operations. All reduce applies the same reduce operation on all computing nodes.

[0063] Taking All reduce sum as an example, after execution, each computing node has the sum of the data of all computing nodes.

[0064] 6. Distance

[0065] The minimum number of hops required for communication between two computing nodes in the network is the distance between these two computing nodes.

[0066] 7. Diameter

[0067] The maximum value of the distances between any two computing nodes in the network. Among them, the larger the diameter of the network, the greater the communication delay of the network.

[0068] 8. Average distance

[0069] The average value of the distances of all computing node pairs in the network, and each computing node pair includes any two computing nodes in the network. Among them, the larger the average distance of the network, the greater the overall communication delay of the network.

[0070] 9. Bisection bandwidth

[0071] The bisection bandwidth refers to the communication bandwidth between two identical parts after dividing the network into two identical parts. Among them, the larger the bisection bandwidth, the stronger the communication ability of the network.

[0072] 10. Degree

[0073] In a complex network, the degree is used to measure the number of edges connected to a computing node, that is, the degree of a computing node is defined as the number of edges it has. For example, if a computing node has three edges, then the degree of this computing node is 3.

[0074] 11. A Mod B

[0075] The value of A Mod B is the remainder obtained by taking the modulus of number A with respect to number B, that is, used to calculate the remainder of number A divided by number B. For example, 5 Mod 4 = 1.

[0076] Secondly, the application background of the embodiments of the present application is introduced.

[0077] When building a large-scale network, the Clos architecture is usually adopted to connect and network multiple nodes / devices. However, as a high-performance and scalable non-direct network architecture, the Clos architecture has a high networking cost and is difficult to cope with the rapidly growing network scale. Different from the Clos architecture, the Torus network is a direct network topology architecture that improves the communication performance of the network by adding loopback links on the basis of a mesh network. Since the networking cost / power consumption of the Torus network is low, it has been applied to various large-scale networks.

[0078] See Figure 1 , the Torus network can be represented by an N-dimensional grid. The communication connections between computing nodes in the Torus network are achieved by establishing direct connections between adjacent computing nodes in each dimension. For example, in a 1D Torus network, each computing node is connected to its two adjacent computing nodes; in a 2D Torus network, each computing node is connected to its four adjacent computing nodes above, below, left, and right; in a 3D Torus network, each computing node is connected to its six adjacent computing nodes in front, behind, left, right, above, and below. Since the overall architecture of the Torus network belongs to the Cayley graph, which is a type of point-symmetric network topology, as long as the dimension of the Torus network and the number of nodes in each dimension are determined, the networking method when multiple computing nodes form a Torus network can be uniquely determined.

[0079] However, as the network scale increases, the Torus network has disadvantages such as a large network diameter, long communication latency, and impaired All to all bandwidth, resulting in relatively limited overall communication performance of the Torus network.

[0080] See Figure 2 , the 6D Torus network based on ring fusion (also known as the Torus Fusion network, abbreviated as the Tofu network) uses a 3D structure composed of 12 computing nodes as the computing nodes in the 3D Torus network to construct a 6D network as a whole. The Tofu network uses six coordinate systems to describe each computing node, and the coordinates of each computing node can be denoted as (X, Y, Z, A, B, C). Among them, the coordinate values of the computing node on the A coordinate axis and the C coordinate axis are 0 or 1, and the coordinate value on the B coordinate axis is 0, 1, or 2; the value range of the coordinate values of the computing node on the X coordinate axis, Y coordinate axis, and Z coordinate axis depends on the overall scale of the Tofu network.

[0081] In the Tofu network, as Figure 2 shown, each computing node is connected to four computing nodes in the corresponding 3D structure and is also connected to six computing nodes in the 3D Torus network architecture. Therefore, the degree of each computing node in the Tofu network is 10, which greatly improves the throughput and fault tolerance of the Torus network, and at the same time, the networking cost is relatively low. However, the far-point / cluster communication performance of the Tofu network at the scale of ten thousand cards is still relatively limited, the communication latency still reaches 20+ hops, and the All to all communication performance is poor.

[0082] It should be noted that although the alias of the Tofu network is 6D Torus, its essence is to place a 3D structure at the computing node positions of the 3D Torus. Therefore, the Tofu network does not essentially belong to the Torus network and does not have point symmetry, resulting in a relatively complex communication routing setting for the computing nodes in the network.

[0083] See Figure 3 , the Torus network based on the twisted loop (also known as the Twisted Torus network) is based on the irregular (rectangular) Torus network, and twists the loop of the short side in the Torus network to improve the bisection bandwidth of the overall network. At the same time, it inherits the advantage of the relatively low networking cost of the Torus network. Among them, the degree of each computing node in the Twisted torus network is the same as that of the Torus network.

[0084] It should be noted that since the Twisted torus network is an improved scheme for the irregular Torus network, when networking multiple nodes, the networking method of the Twisted Torus network can be used to connect the computing nodes only when the number of computing nodes in one dimension is twice the number of computing nodes in another dimension.

[0085] That is to say, the Twisted Torus network is only for the irregular Torus network and has the effect of improving the bisection bandwidth, but the irregularity of the network architecture will also lead to a relatively complex communication routing setting for the computing nodes in the entire network.

[0086] It can be seen that although the above-mentioned improved scheme based on the Torus network enhances the throughput, fault tolerance and bisection bandwidth of the network, and inherits the advantage of the relatively low networking cost of the Torus network, its network architecture does not have point symmetry, resulting in a relatively complex communication routing setting for the computing nodes in the network; moreover, there is still a problem of relatively long network communication delay in the case of large-scale node networking in the above-mentioned improved scheme, that is, the defect that the communication performance of the Torus network is relatively limited as the network scale increases is not overcome.

[0087] Based on this, the embodiments of the present application provide a network architecture. On the basis of the Torus network, for any computing node in the network, after communicating and connecting according to the Torus networking method, at least one computing node with the farthest communication distance in the diagonal direction from the computing node is additionally connected. In this way, the network architecture provided by the present application additionally connects the computing nodes with the farthest communication distance on the basis of the Torus network, reducing the communication delay of the network and improving the communication bandwidth; at the same time, the network architecture has point symmetry, making the routing setting of the computing nodes in the network relatively simple.

[0088] It should be understood that the network architecture provided by the embodiments of the present application can be an intra-ACK interconnection mode or a cluster interconnection mode to achieve multi-node networking; it can be applied to, but is not limited to, small-scale general networks, or various service scenarios such as large-scale high performance computing (HPC), artificial intelligence training, and artificial intelligence inference. Of course, with the emergence of new multi-node networking methods and network architectures, the technical concept of connecting the computing node with the farthest communication distance in the additional connection network provided by the embodiments of the present application is also applicable to similar technical problems.

[0089] Next, a detailed explanation of the network architecture provided by the embodiments of the present application will be given.

[0090] The embodiments of the present application provide a network architecture, which includes multiple computing nodes. The multiple computing nodes are communicatively connected according to the Torus networking method, and a first computing node among the multiple computing nodes is also communicatively connected to a target computing node. Among them, the first computing node is any one of the multiple computing nodes, and the target computing node is the computing node with the farthest communication distance from the first computing node when the multiple computing nodes are communicatively connected according to the Torus networking method.

[0091] It should be noted that when implementing the network architecture provided by the embodiments of the present application, the number of the above-mentioned first computing nodes can be all the computing nodes in the network architecture or part of the computing nodes. The embodiments of the present application do not limit the number of the first computing nodes.

[0092] As an example, assuming that the communication distance between a computing node and its adjacent computing node is 1 hop, for the first computing node, the number of hops (or the number of computing nodes) it experiences to reach the target computing node is the most in the Torus networking method. For example, when the multiple computing nodes are communicatively connected according to the Torus networking method, if the first computing node needs to experience 5 hops (5 computing nodes) to communicate with a second computing node in the Torus network, and the number of hops for communicating with other computing nodes is less than 5 hops, then the second computing node is the computing node with the farthest communication distance from the first computing node. At this time, the second computing node is the target computing node that the first computing node needs to additionally connect to.

[0093] It should be understood that other methods can also be used to determine the computing node with the farthest communication distance from the first computing node, not limited to the method in the above example. The embodiments of the present application do not limit this, aiming to illustrate that when the communication distance between the first computing node and the target computing node is the farthest, its communication delay will be relatively large, and the solution provided by the embodiments of the present application is exactly to improve the problem of long communication delay.

[0094] To distinguish it from the Torus network and for ease of description, in the embodiments of the present application, the technical concept of the embodiments of the present application will be adopted to improve the network architecture of the Torus network, so as to form a Diagonal Torus network (abbreviated as Diagonal Torus network) based on diagonals with multiple computing nodes. That is to say, the network obtained by networking and connecting multiple computing nodes according to the network architecture shown in the embodiments of the present application can be called a Diagonal Torus network.

[0095] In a possible implementation manner, during the process of forming a Diagonal Torus network with multiple computing nodes, the multiple computing nodes can be communicatively connected in the Torus networking manner to obtain a Torus network; on the basis of this Torus network, each computing node is further communicatively connected to the target computing node that is farthest from it in terms of communication distance in this Torus network to obtain a Diagonal Torus network.

[0096] Among them, the Torus networking manner includes at least one dimension, and the position of a computing node in the network architecture can be represented by the coordinate value of this computing node. When the Torus networking manner includes one dimension, the coordinates of each computing node in this network architecture are the coordinate values on this one dimension; when the Torus networking manner includes multiple dimensions, the coordinates of each computing node in this network architecture include the coordinate values of this computing node on each dimension.

[0097] As an example, the coordinate of a computing node in a 1D network is a numerical value; the coordinate of a computing node in a 2D network is (x, y), including the coordinate values on the x dimension and the y dimension; the coordinate of a computing node in a 3D network is (x, y, z), including the coordinate values on the x dimension, the y dimension, and the z dimension, and so on. Examples are not given one by one here.

[0098] It should be noted that when multiple computing nodes are communicatively connected in the Torus networking manner in at least one dimension, when determining the target computing node for a first computing node to communicate with, in the case where the total number of computing nodes in each dimension of at least one dimension is an even number, and there is at least one dimension in which the total number of existing computing nodes is an odd number, the method for determining the target computing node for the same first computing node is also slightly different, and will be introduced separately below.

[0099] In the first method for determining the target computing node, multiple computing nodes are communicatively connected in a Torus networking mode, and the Torus networking mode includes at least one dimension. If the total number of computing nodes in each dimension of the at least one dimension is even, for the first computing node among the multiple computing nodes, the target computing node with the farthest communication distance from the first computing node includes one computing node, and the coordinate value of this computing node in the first dimension is the first coordinate value, which is obtained by taking the modulus of d1 + k1 / 2 with respect to k1.

[0100] Wherein, the first dimension is any one of the at least one dimension, d1 is the coordinate value of the first computing node in the first dimension, and k1 is the total number of computing nodes in the first dimension.

[0101] That is to say, if the total number of computing nodes in each dimension of the Diagonal Torus network is even, for the first computing node in the Diagonal Torus network, when communicatively connected to at least two computing nodes in the Torus networking mode, it is also communicatively connected to the target computing node with the farthest communication distance.

[0102] Since common Torus networks include Figure 1 the 1D Torus network, 2D Torus network, and 3D Torus network shown, therefore, after improving the network architectures of these Torus networks in the embodiments of the present application, the Diagonal 1D Torus network, Diagonal 2D Torus network, and Diagonal 3D Torus network are correspondingly obtained. Next, taking the Diagonal 1D Torus network, Diagonal 2D Torus network, and Diagonal 3D Torus network as examples respectively, in the case where the total number of computing nodes in each dimension of the at least one dimension is even, the process of the first computing node in the Diagonal Torus network communicating with the target computing node will be explained.

[0103] (1) Multiple computing nodes form a Diagonal 1D Torus network, and the total number of computing nodes is even.

[0104] When the total number of computing nodes is k and k is even, the set of the k computing nodes in the 1D network can be denoted as {0, 1,..., k - 1}, and each computing node in the 1D Torus network is communicatively connected to two computing nodes.

[0105] Based on this, when k computing nodes are communicatively connected according to the 1D Torus networking mode, the first computing node with a coordinate value of d is communicatively connected to the computing node with a coordinate value of (d + 1) Mod k and the computing node with a coordinate value of (d + k - 1) Mod k respectively. On the basis of the 1D Torus network, when the Diagonal 1D Torus network is formed in the embodiments of the present application, the above-mentioned first computing node with a coordinate value of d is further communicatively connected to a target computing node with a coordinate value of (d + k / 2) Mod k.

[0106] That is to say, if the total number of multiple computing nodes included in the network architecture is an even number, in the Diagonal 1D Torus network formed in the embodiments of the present application, each computing node is communicatively connected to three computing nodes.

[0107] As an example, referring to Figure 4 , assume that the total number of multiple computing nodes included in the network architecture is k = 16, and the coordinate value of the first computing node is d = 1. Based on the above coordinate value calculation method, the first computing node is communicatively connected to the computing node with a coordinate value of 0 and the computing node with a coordinate value of 2 according to the 1D Torus networking mode; at the same time, the first computing node is communicatively connected to the target computing node with a coordinate value of 9. By analogy, when 16 computing nodes are communicatively connected according to the 1D Torus networking mode, each computing node is further communicatively connected to a computing node with a coordinate value of (d + k / 2) Mod k to form a Diagonal 1D Torus network.

[0108] It should be noted that Figure 1 and Figure 4 both use a 1D network with 16 computing nodes included in the network architecture, that is, k = 16 as an example for illustration. However, in practical applications, the network architecture may include more or fewer computing nodes, and the embodiments of the present application do not limit this; moreover, for 16 computing nodes, the embodiments of the present application use numbering from 0 to 15 as an example. In practical applications, other numbering methods may also be used. Without affecting the multi-node networking mode and network architecture, the embodiments of the present application do not limit the setting method of the coordinate values of the computing nodes.

[0109] For the convenience of comparative illustration, in the embodiments of the present application, it is assumed that the single-link bandwidth between two computing nodes is 1. Through calculation and analysis, in a 1D Torus network composed of 16 computing nodes, the diameter of the network is 8, the average distance is 4.3, the bisection bandwidth is 2, the All reduce bandwidth is 2, the All to all bandwidth is 0.5, and the degree of the computing node is 2; in a Diagonal 1D Torus network composed of 16 computing nodes, the diameter of the network is 4, the average distance is 2.6, the bisection bandwidth is 4, the Allreduce bandwidth is 3, the All to all bandwidth is 1, and the degree of the computing node is 3.

[0110] Among them, the degree of the computing node is equivalent to the number of communication connection lines of the computing node in the above 1D network. The more the number of communication connection lines of the computing node, the better the throughput performance and fault tolerance performance of the computing node; the diameter and the average distance can reflect the communication delay situation of the 1D network. The smaller the diameter and the average distance, the smaller the communication delay of the computing node; the bisection bandwidth, the Allreduce bandwidth, and the All to all bandwidth can reflect the data transmission performance between the computing nodes in the 1D network. The larger the bisection bandwidth, the All reduce bandwidth, and the All to all bandwidth, the higher the communication performance of the computing node.

[0111] As the network scale increases, for the Diagonal 1D Torus network obtained by multi-node networking based on the network architecture shown in the embodiments of the present application, and for the 1D Torus network, all network parameters approach a fixed value. Table 1 below gives an exemplary calculation method of the network parameters of the 1D Torus network and the Diagonal 1D Torus network, as well as the comparison of the results of the network parameters.

[0112] Table 1

[0113] k computing nodes, where k is an even number 1D Torus network Diagonal 1D Torus network Comparison situation Degree 2 3 Degree increases by 1 / 2 Diameter k / 2 k / 4 Diameter is halved Average distance k / 4 k / 8 Average distance is halved Bisection bandwidth 2 4 Bisection bandwidth doubles All reduce bandwidth 2 3 All reduce bandwidth increases by 1 / 2 times All to all bandwidth 8 / k 16 / k All to all bandwidth doubles

[0114] As can be seen from Table 1, as the network scale expands and the total number of computing nodes in the network architecture increases, compared with the 1D Torus network, the Diagonal 1D Torus network has greatly improved communication performance. Moreover, from Figure 4 the network structure of the shown Diagonal 1D Torus network, it can be seen that the Diagonal 1D Torus network is the same as the 1D Torus network and also has point symmetry. Its network architecture is easy to implement in engineering, and there is a simple communication route between computing nodes, making the overall routing algorithm setting of this network relatively easy.

[0115] In summary, compared with the 1D Torus network, the Diagonal 1D Torus network obtained by multi-node networking according to the network architecture provided in the embodiments of the present application not only increases the communication bandwidth between computing nodes (including All toall bandwidth, All reduce bandwidth, and bisection bandwidth), but also reduces the communication latency between computing nodes (including maximum communication latency and average communication latency), improving the overall communication performance of the network.

[0116] (2) Multiple computing nodes form a k×k Diagonal 2D Torus network, and k is an even number.

[0117] When the total number of computing nodes is k×k and k is an even number, the set of the k×k computing nodes in the 2D network can be denoted as {(x, y)|0 ≤ x ≤ k - 1, 0 ≤ y ≤ k - 1}, and each computing node in the 2D Torus network is communicatively connected to four computing nodes.

[0118] Based on this, when the k×k computing nodes are communicatively connected according to the 2D Torus networking method, the first computing node with coordinate values (x, y) is communicatively connected to the computing node with coordinate values ((x + 1) Mod k, y), the computing node with coordinate values (x, (y + 1) Mod k), the computing node with coordinate values ((x + k - 1) Mod k, y), and the computing node with coordinate values (x, (y + k - 1) Mod k). On the basis of the 2D Torus network, when the embodiments of the present application form a Diagonal 2D Torus network, the above-mentioned first computing node with coordinate values (x, y) is also communicatively connected to the target computing node with coordinate values ((x + k / 2) Mod k, (y + k / 2) Mod k).

[0119] That is to say, if the total number of multiple computing nodes included in the network architecture is an even number, then in the Diagonal 2D Torus network formed by the embodiments of the present application, each computing node is communicatively connected to five computing nodes.

[0120] As an example, see Figure 5, assume that the total number of multiple computing nodes included in the network architecture is 16, which is used to form a 4×4 2D Torus network or a 4×4 Diagonal 2D Torus network; at the same time, assume that the horizontal direction is the X-axis and the vertical direction is the Y-axis, and the coordinate value of the first computing node is (0, 0). Then the first computing node is communicatively connected to the computing nodes with coordinate values (1, 0), (0, 1), (0, 3), and (3, 0) according to the 2D Torus networking method; at the same time, the first computing node is communicatively connected to the target computing node with the coordinate value (2, 2). And so on, after the 16 computing nodes are communicatively connected according to the 2D Torus networking method, each computing node is also communicatively connected to a computing node with a coordinate value of ((x + k / 2) Mod k, (y + k / 2) Mod k) to form a Diagonal 2D Torus network.

[0121] It should be noted that Figure 1 and Figure 5 in the 2D networks, it is assumed that the network architecture includes 16 computing nodes, and the total number of computing nodes in both dimensions is 4 (i.e., k = 4) for illustration. However, in actual applications, the network architecture may include more or fewer computing nodes, and the number of nodes in the two dimensions of the multiple computing nodes may also be unequal. The embodiments of the present application do not limit this; moreover, for 16 computing nodes, the embodiments of the present application take the value range of the two dimensions as [0, 3] for illustration. In actual applications, other methods may also be used to determine the coordinate values of the computing nodes. Without affecting the multi-node networking method and network architecture, the embodiments of the present application do not limit the setting method of the coordinate values of the computing nodes.

[0122] For the convenience of comparative illustration, it is assumed in the embodiments of the present application that the single-link bandwidth between two computing nodes is 1. Through calculation and analysis, in the 4×4 2D Torus network formed by 16 computing nodes, the diameter of the network is 4, the average distance is 2.13, the bisection bandwidth is 8, the All reduce bandwidth is 4, the All to all bandwidth is 2, and the degree of the computing node is 4; in the 4×4 Diagonal 2D Torus network formed by 16 computing nodes, the diameter of the network is 2, the average distance is 1.67, the bisection bandwidth is 16, the All reduce bandwidth is 5, the All to all bandwidth is 3, and the degree of the computing node is 5.

[0123] Similarly, the degree of a computing node is equivalent to the number of communication connection lines of the computing node in the above 2D network. The more the number of communication connection lines of the computing node, the better the throughput performance and fault tolerance performance of the computing node; the diameter and average distance can reflect the communication delay situation of the 2D network. The smaller the diameter and average distance, the smaller the communication delay of the computing node; the bisection bandwidth, Allreduce bandwidth, and All to all bandwidth can reflect the data transmission performance between computing nodes in the 2D network. The larger the bisection bandwidth, All reduce bandwidth, and All to all bandwidth, the higher the communication performance of the computing node.

[0124] As the network scale increases, for the Diagonal 2D Torus network obtained by multi-node networking based on the network architecture shown in the embodiments of the present application, and for the 2D Torus network, all network parameters approach a fixed value. Table 2 below gives an exemplary calculation method for the network parameters of the 2D Torus network and the Diagonal 2D Torus network, as well as the comparison of the results of the network parameters.

[0125] Table 2

[0126] k×k computing nodes, where k is an even number 2D Torus network Diagonal 2D Torus network Comparison situation Degree 4 3 Degree increases by 1 / 4 Diameter k k / 2 Diameter is halved Average distance k / 2 k / 3 Average distance increases by approximately 0.67 times Bisection bandwidth 2k 4k Bisection bandwidth doubles All reduce bandwidth 4 5 All reduce bandwidth increases by 1 / 4 times All to all bandwidth 8 / k 12 / k All to all bandwidth increases by approximately 1.5 times

[0127] As can be seen from Table 2, as the network scale expands and the total number of computing nodes in the network architecture increases, compared with the 2D Torus network, the Diagonal 2D Torus network has greatly improved its communication performance. Moreover, Figure 5 From the network structure of the Diagonal 2D Torus network, it can be seen that the Diagonal 2D Torus network is the same as the 2D Torus network and also has point symmetry. Its network architecture is easy to implement in engineering, and there is a simple communication route between computing nodes, making it relatively easy to set the overall routing algorithm of the network.

[0128] In summary, compared with the 2D Torus network, the Diagonal 2D Torus network obtained by multi-node networking according to the network architecture provided in the embodiments of the present application not only increases the communication bandwidth between computing nodes (including All to all bandwidth, All reduce bandwidth, and bisection bandwidth), but also reduces the communication delay between computing nodes (including the maximum communication delay and average communication delay), improving the overall communication performance of the network.

[0129] (3) Multiple computing nodes form a k×k×k Diagonal 3D Torus network, and k is an even number.

[0130] When the total number of computing nodes is k×k×k and k is an even number, the set of the k×k×k computing nodes in the 3D network can be denoted as {(x, y, z)|0 ≤ x ≤ k - 1, 0 ≤ y ≤ k - 1, 0 ≤ z ≤ k - 1}, and each computing node in the 3D Torus network is communicatively connected to six computing nodes.

[0131] Based on this, when the k×k×k computing nodes are communicatively connected according to the 3D Torus networking mode, the first computing node with the coordinate value (x, y, z) is communicatively connected to the computing node with the coordinate value ((x + 1)Mod k, y, z), the computing node with the coordinate value (x, (y + 1)Mod k, z), the computing node with the coordinate value (x, y, (z + 1)Mod k), the computing node with the coordinate value ((x + k - 1)Modk, y, z), the computing node with the coordinate value (x, (y + k - 1)Mod k, z), and the computing node with the coordinate value (x, y, (z + k - 1)Modk). On the basis of the 3D Torus network, when the present application embodiment constructs a Diagonal 3D Torus network, the above-mentioned first computing node with the coordinate value (x, y, z) is further communicatively connected to the target computing node with the coordinate value ((x + k / 2)Mod k, (y + k / 2)Mod k, (z + k / 2)Mod k).

[0132] That is to say, if the total number of multiple computing nodes included in the network architecture is an even number, then in the Diagonal 3D Torus network constructed by the present application embodiment, each computing node is communicatively connected to seven computing nodes.

[0133] As an example, such as Figure 6As shown, it is assumed that the total number of nodes of multiple computing nodes included in the network architecture is 64, forming a 3D Torus network of 4×4×4, or a Diagonal 3D Torus network of 4×4×4. At the same time, it is assumed that the horizontal direction is the X-axis, the vertical direction is the Y-axis, and the direction perpendicular to the X-axis and Y-axis respectively is the Z-axis. The coordinate value of the first computing node is (0, 0, 0). Then, the first computing node is communicatively connected to the computing nodes with coordinate values of (1, 0, 0), (0, 1, 0), (0, 0, 1), (3, 0, 0), (0, 3, 0), and (0, 0, 3) according to the 3D Torus networking method. At the same time, the first computing node is also communicatively connected to the target computing node with the coordinate value of (2, 2, 2). By analogy, after the 64 computing nodes are communicatively connected according to the 3D Torus networking method, each computing node is also communicatively connected to a computing node with the coordinate value of ((x + k / 2) Mod k, (y + k / 2) Mod k, (z + k / 2) Mod k) to form a Diagonal 3D Torus network.

[0134] Among them, for the convenience of description, Figure 6 only when the coordinate value of the first computing node is (0, 0, 0), it is shown that the first computing node needs to be communicatively connected to seven computing nodes. Other computing nodes can also refer to the above calculation method to be communicatively connected to the target computing node, which is not shown in the appendix Figure 6 and is not shown in the figure.

[0135] It should be noted that in the embodiments of the present application, only an example is given where the network architecture includes 64 computing nodes and the total number of computing nodes in the three dimensions is 4 (i.e., k = 4). However, in actual applications, the network architecture may also include more or fewer computing nodes, and the number of nodes of multiple computing nodes in the three dimensions may also be different. The embodiments of the present application do not limit this. Moreover, for 64 computing nodes, the embodiments of the present application take the value range of the three dimensions as [0, 3] for example. In actual applications, other methods may also be used to determine the coordinate values of the computing nodes. Without affecting the multi-node networking method and network architecture, the embodiments of the present application do not limit the setting method of the coordinate values of the computing nodes.

[0136] For the convenience of comparison and explanation, in the embodiments of the present application, it is assumed that the bandwidth of a single link (i.e., each communication connection line) between two computing nodes is 1. Through calculation and analysis, in a 4×4×4 3D Torus network composed of 64 computing nodes, the diameter of the network is 6, the average distance is 3, the bisection bandwidth is 32, the All reduce bandwidth is 6, the All to all bandwidth is 2, and the degree of the computing node is 6; in a 4×4×4 Diagonal 3D Torus network composed of 64 computing nodes, the diameter of the network is 3, the average distance is 2.44, the bisection bandwidth is 64, the All reduce bandwidth is 7, the All to all bandwidth is 2.74, and the degree of the computing node is 7.

[0137] Similarly, the height of the computing node is equal to the number of communication connection lines of the computing node in the above 3D network. The more the number of communication connection lines of the computing node, the better the throughput performance and fault tolerance performance of the computing node; the diameter and the average distance can reflect the communication delay situation of the 3D network. The smaller the diameter and the average distance, the smaller the communication delay of the computing node; the bisection bandwidth, the All reduce bandwidth, and the All to all bandwidth can reflect the data transmission performance between the computing nodes in the 3D network. The larger the bisection bandwidth, the All reduce bandwidth, and the All to all bandwidth, the higher the communication performance of the computing node.

[0138] As the network scale increases, for the Diagonal 3D Torus network obtained by multi-node networking based on the network architecture shown in the embodiments of the present application, as well as the 3D Torus network, all network parameters approach a fixed value. Table 3 below gives an exemplary calculation method of the network parameters of the 3D Torus network and the Diagonal 3D Torus network, as well as the result comparison of each network parameter.

[0139] Table 3

[0140]

[0141] As can be seen from Table 3, as the network scale expands and the total number of computing nodes in the network architecture increases, compared with the 3D Torus network, the Diagonal 3D Torus network has greatly improved its communication performance. Moreover, from Figure 6 the network structure of the shown Diagonal 3D Torus network, it can be seen that the Diagonal 3D Torus network, like the 3D Torus network, also has point symmetry. Its network architecture is easy to implement in engineering, and there is a simple communication route between computing nodes, making the overall routing algorithm setting of the network relatively easy.

[0142] In summary, compared with the 3D Torus network, the Diagonal 3D Torus network obtained by multi-node networking according to the network architecture provided in the embodiments of the present application not only increases the communication bandwidth between computing nodes (including All toall bandwidth, All reduce bandwidth, and bisection bandwidth), but also reduces the communication delay between computing nodes (including the maximum communication delay and the average communication delay), improving the overall communication performance of the network.

[0143] In the second method for determining the target computing node, multiple computing nodes are communicatively connected in a Torus networking manner, and the Torus networking manner includes at least one dimension. If there is a dimension in which the total number of computing nodes is odd among the at least one dimension, then for the first computing node among the multiple computing nodes, the target computing nodes with the farthest communication distance from the first computing node include at least two computing nodes. When the total number of computing nodes in the first dimension is even, the coordinate values of the at least two computing nodes in the first dimension are all the first coordinate value, and the first coordinate value is the value obtained by taking the modulus of d1 + k1 / 2 with respect to k1. When the total number of computing nodes in the first dimension is odd, the coordinate values of the first part of the at least two computing nodes in the first dimension are the second coordinate value, and the coordinate values of the second part of the at least two computing nodes in the first dimension are the third coordinate value.

[0144] Among them, the number of the first part of the computing nodes is equal to the number of the second part of the computing nodes. The second coordinate value is the value obtained by taking the modulus of d1 + (k1 - 1) / 2 with respect to k1, and the third coordinate value is the value obtained by taking the modulus of d1 + (k1 + 1) / 2 with respect to k1. The first dimension is any one of the at least one dimension, d1 is the coordinate value of the first computing node in the first dimension, and k1 is the total number of computing nodes in the first dimension.

[0145] Optionally, if the total number of computing nodes in each dimension among the at least one dimension is odd, the at least two computing nodes include a second computing node and a third computing node. The coordinate value of the second computing node in the first dimension is the second coordinate value, and the coordinate value in the second dimension is the fourth coordinate value, and the fourth coordinate value is the value obtained by taking the modulus of d2 + (k2 - 1) / 2 with respect to k2. The coordinate value of the third computing node in the first dimension is the third coordinate value, and the coordinate value in the second dimension is the fifth coordinate value, and the fifth coordinate value is the value obtained by taking the modulus of d2 + (k2 + 1) / 2 with respect to k2. Among them, the second dimension is any one of the at least one dimension other than the first dimension, d2 is the coordinate value of the first computing node in the second dimension, and k2 is the total number of computing nodes in the second dimension.

[0146] That is, when the total number of computing nodes in each dimension is odd, for the first computing node, two coordinate values can be calculated in each dimension in the above manner. Therefore, each first computing node has at most four target computing nodes with the farthest communication distance from it. At this time, communication connections can be made with all four computing nodes, or only with some of the computing nodes. In addition, under the condition that the overall network architecture satisfies point symmetry, connections can also be made only with the two computing nodes in the diagonal direction, that is, the above-mentioned second computing node and the third computing node. The embodiments of the present application do not limit this.

[0147] It can be seen that when multiple computing nodes in a network architecture are communicatively connected in the Torus networking mode in at least one dimension, if there is a dimension in which the total number of computing nodes is odd in the at least one dimension, then when the total number of computing nodes in the first dimension is even, each computing node in the Diagonal Torus network is communicatively connected to one more computing node on the basis of the Torus network; when the total number of computing nodes in the first dimension is odd, each computing node in the Diagonal Torus network is communicatively connected to at least two more computing nodes on the basis of the Torus network.

[0148] When the total number of computing nodes in the first dimension is even, the first computing node located in the first dimension in the Diagonal Torus network is communicatively connected to a target computing node on the basis of the original Torus network. The method for determining the target computing node is similar to the first method for determining the target computing node described above. For the specific implementation process, reference can be made to the relevant description above and will not be elaborated here.

[0149] Next, taking the Diagonal 1D Torus network, the Diagonal 2D Torus network, and the Diagonal 3D Torus network as examples, the process of the first computing node in the Diagonal Torus network being communicatively connected to at least two computing nodes is explained respectively in the case where the total number of computing nodes in each dimension in at least one dimension is odd, and in the case where there are both dimensions with an even total number of computing nodes and dimensions with an odd total number of computing nodes in at least one dimension.

[0150] First, the process of the first computing node in the Diagonal Torus network being communicatively connected to at least two computing nodes is explained in the case where the total number of computing nodes in each dimension in at least one dimension is odd.

[0151] (1) Multiple computing nodes form a Diagonal 1D Torus network, and the total number of computing nodes is odd.

[0152] When the total number of computing nodes is k and k is odd, the set of these k computing nodes in the 1D network can be denoted as {0, 1, …, k−1}. Based on the 1D Torus network, when the embodiments of the present application construct a Diagonal 1D Torus network, the first computing node with a coordinate value of d is also communicatively connected to a target computing node with a coordinate value of (d+(k−1) / 2) Mod k and a target computing node with a coordinate value of (d+(k+1) / 2) Mod k respectively.

[0153] That is to say, if the total number of multiple computing nodes included in the network architecture is odd, in the Diagonal 1D Torus network constructed by the embodiments of the present application, each computing node is communicatively connected to four computing nodes. Among these four computing nodes, two computing nodes are nodes communicatively connected in the 2D Torus networking mode, and the other two computing nodes are nodes additionally connected according to the technical solution of the embodiments of the present application.

[0154] As an example, referring to Figure 7 , assume that the total number of multiple computing nodes included in the network architecture is k = 9, and the coordinate value of the first computing node is d = 1. Then the first computing node is first communicatively connected to the computing nodes with coordinate values of 0 and 2 respectively in the 1D Torus networking mode; meanwhile, the first computing node is communicatively connected to two target computing nodes with coordinate values of 5 and 6 respectively. By analogy, after the 9 computing nodes are communicatively connected in the 1D Torus networking mode, each computing node is also communicatively connected to the computing node with a coordinate value of (d+(k−1) / 2) Mod k and the computing node with a coordinate value of (d+(k+1) / 2) Mod k to construct a Diagonal 1D Torus network.

[0155] It should be noted that Figure 7 takes the 1D network with 9 computing nodes included in the network architecture, i.e., k = 9, as an example for illustration. However, in practical applications, the network architecture may also include more or fewer computing nodes, and the embodiments of the present application do not limit this; moreover, for 9 computing nodes, the embodiments of the present application take the numbering from 0 to 8 as an example. In practical applications, other numbering methods may also be used. Without affecting the multi-node networking mode and network architecture, the embodiments of the present application do not limit the setting method of the coordinate values of the computing nodes.

[0156] It can be seen that when the total number of computing nodes is odd, after multiple computing nodes are communicatively connected in the 1D Torus networking mode, each computing node still needs to be additionally connected to at least two computing nodes. Therefore, compared with the 1D Torus network, the Diagonal 1D Torus network increases the degree of each computing node, reduces the communication delay, and improves the communication bandwidth, resulting in a greater improvement in the overall communication performance of the network.

[0157] (2) Multiple computing nodes form a k×k Diagonal 2D Torus network, where k is odd.

[0158] When the total number of computing nodes is k×k and k is odd, the set of the k×k computing nodes in the 2D network can be denoted as {(x, y)|0≤x≤k - 1, 0≤y≤k - 1}. On the basis of the 2D Torus network, when the present application embodiment forms a Diagonal 2D Torus network, the first computing node with coordinates (x, y) is also communicatively connected to the target computing nodes with coordinates (x+(k - 1) / 2)Mod k and (x+(k + 1) / 2)Mod k in the X-axis direction, and coordinates (y+(k - 1) / 2)Mod k and (y+(k + 1) / 2)Mod k in the Y-axis direction, respectively.

[0159] That is to say, if the total number of multiple computing nodes included in the network architecture is odd, in the Diagonal 2D Torus network formed by the present application embodiment, each computing node communicates with at least six computing nodes. Among the at least six computing nodes, four computing nodes are nodes communicatively connected in the 2D Torus networking mode, and the remaining computing nodes are nodes additionally connected according to the technical solution of the present application embodiment.

[0160] In a possible implementation, when the total number k of computing nodes in each dimension of at least one dimension is odd, the first computing node with the coordinate value (x, y) in the Diagonal 2DTorus network is respectively communicatively connected to all the computing nodes with the farthest communication distance. At this time, on the basis of the 2D Torus network, the first computing node with the coordinate value (x, y) is also respectively communicatively connected to the target computing nodes with the coordinate values ((x+(k - 1) / 2) Mod k, (y+(k - 1) / 2) Mod k), the target computing nodes with the coordinate values ((x+(k - 1) / 2) Mod k, (y+(k + 1) / 2) Mod k), the target computing nodes with the coordinate values ((x+(k + 1) / 2) Mod k, (y+(k - 1) / 2) Mod k), and the target computing nodes with the coordinate values ((x+(k + 1) / 2) Mod k, (y+(k + 1) / 2) Mod k) to form a Diagonal 2D Torus network.

[0161] In this implementation, each computing node is communicatively connected to all the computing nodes with the farthest communication distance. Therefore, the communication delay of the entire Diagonal 2D Torus network is greatly improved. However, since each first computing node needs to additionally connect four target computing nodes on the basis of the 2D Torus network, the number of additional communication connections required for the entire Diagonal 2DTorus network is 4×k×k, which also increases the networking cost to a certain extent.

[0162] In another possible implementation, when the total number k of computing nodes in each dimension of at least one dimension is odd, the first computing node with the coordinate value (x, y) in the Diagonal 2D Torus network is only connected to the two computing nodes with the farthest communication distance in the diagonal direction. At this time, on the basis of the 2D Torus network, the first computing node with the coordinate value (x, y) is also respectively communicatively connected to the target computing node with the coordinate value ((x+(k - 1) / 2) Mod k, (y+(k - 1) / 2) Mod k) and the target computing node with the coordinate value ((x+(k + 1) / 2) Mod k, (y+(k + 1) / 2) Mod k) to form a Diagonal 2D Torus network.

[0163] In this implementation, when connecting the target computing nodes with the farthest communication distance from the first computing node, only the two computing nodes with the farthest communication distance from the first computing node need to be connected in the diagonal direction. In this way, while improving the communication delay of the 2D Torus network, it is ensured that the networking cost will not be too high.

[0164] As an example, refer to Figure 8 , assuming that the total number of multiple computing nodes included in the network architecture is 25, forming a 5×5 2D Torus network, or a 5×5 Diagonal 2D Torus network; at the same time, assuming that the horizontal direction is the X-axis and the vertical direction is the Y-axis, and the coordinate value of the first computing node is (0, 0), then the first computing node is communicatively connected to the computing nodes with coordinate values (1, 0), (0, 1), (0, 4), and (4, 0) according to the 2D Torus networking method; at the same time, the first computing node is communicatively connected to the target computing node with coordinate value (2, 2) and the target computing node with coordinate value (3, 3).

[0165] And so on, after the 25 computing nodes are communicatively connected according to the 2D Torus networking method, each computing node is also respectively connected to the computing node with coordinate value ((x+(k-1) / 2)Mod k, (y+(k-1) / 2)Mod k) and the computing node with coordinate value ((x+(k+1) / 2)Mod k, (y+(k+1) / 2)Mod k) to form a Diagonal 2D Torus network.

[0166] It should be noted that the 2D network introduced in the above example is only illustrated by taking the network architecture including 25 computing nodes and the total number of computing nodes in both dimensions being 5 (i.e., k = 5), but in actual applications, the network architecture may also include more or fewer computing nodes, and the embodiments of the present application do not limit this; moreover, for 25 computing nodes, the embodiments of the present application take the value range of both dimensions as [0, 4] for illustration. In actual applications, other methods can also be used to determine the coordinate values of the computing nodes. Without affecting the multi-node networking method and network architecture, the embodiments of the present application do not limit the setting method of the coordinate values of the computing nodes.

[0167] It can be seen that in the case where the total number of computing nodes in each dimension is odd in at least one dimension, after multiple computing nodes are communicatively connected according to the 2D Torus networking method, each computing node also needs to be additionally connected to at least two computing nodes. Therefore, compared with the 2D Torus network, the Diagonal 2D Torus network increases the degree of each computing node, reduces the communication delay, and improves the communication bandwidth, so that the overall communication performance of the network is further improved.

[0168] (3) Multiple computing nodes form a k×k×k Diagonal 3D Torus network, and k is odd.

[0169] When the total number of computing nodes is k×k×k and k is odd, the set of the k×k×k computing nodes in the 3D network can be denoted as {(x, y, z)|0≤x≤k - 1, 0≤y≤k - 1, 0≤z≤k - 1}. Based on the 3D Torus network, when the embodiments of the present application form a Diagonal 3DTorus network, the first computing node with the coordinate value (x, y, z) is further communicatively connected to the second computing nodes with the coordinate values (x+(k - 1) / 2)Mod k and (x+(k + 1) / 2)Mod k in the X-axis direction, (y+(k - 1) / 2)Mod k and (y+(k + 1) / 2)Mod k in the Y-axis direction, and (z+(k - 1) / 2)Mod k and (z+(k + 1) / 2)Mod k in the Z-axis direction respectively.

[0170] That is to say, if the total number of multiple computing nodes included in the network architecture is odd, in the Diagonal 3DTorus network formed by the embodiments of the present application, each computing node is communicatively connected to at least eight computing nodes. Among the at least eight computing nodes, six computing nodes are nodes communicatively connected in the 3D Torus networking mode, and the remaining computing nodes are nodes additionally connected according to the technical solution of the embodiments of the present application.

[0171] In a possible implementation, when the total number k of computing nodes in each dimension of at least one dimension is odd, the first computing node with the middle coordinate value of (x, y, z) in the Diagonal 3DTorus network is respectively communicatively connected to all computing nodes with the farthest communication distance. At this time, on the basis of the 3D Torus network, the first computing node with the coordinate value of (x, y, z) is also respectively connected to the target computing nodes with the coordinate values of ((x + (k + 1) / 2) Mod k, (y + (k + 1) / 2) Mod k, (z + (k + 1) / 2) Mod k), the target computing nodes with the coordinate values of ((x + (k + 1) / 2) Mod k, (y + (k - 1) / 2) Mod k, (z + (k + 1) / 2) Mod k), the target computing nodes with the coordinate values of ((x + (k + 1) / 2) Mod k, (y + (k + 1) / 2) Mod k, (z + (k - 1) / 2) Mod k), the target computing nodes with the coordinate values of ((x + (k + 1) / 2) Mod k, (y + (k - 1) / 2) Mod k, (z + (k - 1) / 2) Mod k), the target computing nodes with the coordinate values of ((x + (k - 1) / 2) Mod k, (y + (k + 1) / 2) Mod k, (z + (k + 1) / 2) Mod k), the target computing nodes with the coordinate values of ((x + (k - 1) / 2) Mod k, (y + (k - 1) / 2) Mod k, (z + (k + 1) / 2) Mod k), the target computing nodes with the coordinate values of ((x + (k - 1) / 2) Mod k, (y + (k + 1) / 2) Mod k, (z + (k - 1) / 2) Mod k), and the target computing nodes with the coordinate values of ((x + (k - 1) / 2) Mod k, (y + (k - 1) / 2) Mod k, (z + (k - 1) / 2) Mod k) for communication connection to form the Diagonal 3DTorus network.

[0172] In this implementation, each computing node is communicatively connected to all computing nodes with the farthest communication distance. Therefore, the communication delay of the entire Diagonal 3D Torus network is greatly improved. However, since each first computing node needs to additionally connect four target computing nodes on the basis of the 3D Torus network, the number of additional communication connections required for the entire Diagonal 3DTorus network is 8×k×k×k, which also increases the networking cost to a certain extent.

[0173] In another possible implementation, when the total number k of computing nodes in each dimension of at least one dimension is odd, the first computing node with the middle coordinate value (x, y, z) in the Diagonal 3D Torus network is only connected to the two computing nodes with the farthest communication distance in the diagonal direction. At this time, on the basis of the 3D Torus network, the first computing node with the coordinate value (x, y, z) is also respectively connected to the target computing node with the coordinate value ((x+(k-1) / 2)Mod k, (y+(k-1) / 2)Mod k, (z+(k-1) / 2)Mod k), and the target computing node with the coordinate value ((x+(k+1) / 2)Mod k, (y+(k+1) / 2)Mod k, (z+(k+1) / 2)Mod k) for communication connection to form a Diagonal 3D Torus network.

[0174] In this implementation, when connecting the target computing nodes with the farthest communication distance from the first computing node, only connect the two computing nodes with the farthest communication distance from the first computing node in the diagonal direction. In this way, while improving the communication delay of the 3D Torus network, it is ensured that the networking cost will not be too high.

[0175] As an example, in Figure 1 the 3D Torus network of 3×3×3 shown, assuming that the coordinate value of the first computing node is (0, 0, 0), then in the Diagonal 3D Torus network provided in the embodiments of the present application, this first computing node is also respectively connected to the computing node with the coordinate value (1, 1, 1) and the computing node with the coordinate value (2, 2, 2). And so on, on the basis of the above 3D Torus network of 27 computing nodes, each computing node is also respectively connected to the computing node with the coordinate value ((x+(k-1) / 2)Mod k, (y+(k-1) / 2)Mod k, (z+(k-1) / 2)Mod k), and the computing node with the coordinate value ((x+(k+1) / 2)Mod k, (y+(k+1) / 2)Mod k, (z+(k+1) / 2)Mod k) for communication connection to form a Diagonal 3D Torus network.

[0176] It should be noted that the 3D network introduced in the above example only takes the case where the network architecture includes 27 computing nodes and the total number of computing nodes in each of the three dimensions is 3 (i.e., k = 3) for illustration. However, in actual applications, the network architecture may include more or fewer computing nodes, and the embodiments of the present application do not limit this; moreover, for the 27 computing nodes, the embodiments of the present application take the value range of the three dimensions as [0, 2] for illustration. In actual applications, other methods may also be used to determine the coordinate values of the computing nodes. Without affecting the multi-node networking method and network architecture, the embodiments of the present application do not limit the setting method of the coordinate values of the computing nodes.

[0177] It can be seen that when the total number of computing nodes in each dimension is odd in at least one dimension, after multiple computing nodes are communicatively connected according to the 3DTorus networking method, each computing node also needs to be additionally connected to at least two computing nodes. Therefore, compared with the 3DTorus network, the Diagonal 3D Torus network increases the degree of each computing node, reduces the communication delay, and improves the communication bandwidth, resulting in a greater improvement in the overall communication performance of the network.

[0178] Secondly, in the case where there is a dimension in which the total number of computing nodes is odd in at least one dimension, the process of communicatively connecting the first computing node in the Diagonal Torus network to at least two computing nodes is explained.

[0179] When multiple computing nodes are communicatively connected according to the Torus networking method and the Torus networking method includes at least one dimension, if there is both a dimension in which the total number of computing nodes is odd and a dimension in which the total number of computing nodes is even in at least one dimension, assuming that the total number of computing nodes in the first dimension where the first computing node is located is odd and the total number of computing nodes in the second dimension where the first computing node is located is even, then the first computing node needs to be connected to the target computing nodes indicated by two different coordinate values in the first dimension and needs to be connected to the target computing node indicated by one coordinate value in the second dimension.

[0180] As an example, when the total number of computing nodes is k1×k2, where k1 is odd and k2 is even, the set of the k1×k2 computing nodes in the 2D network can be denoted as {(x, y)|0 ≤ x ≤ k1 - 1, 0 ≤ y ≤ k2 - 1}. At this time, based on the 2D Torus network, the first computing node with the coordinate value (x, y) in the Diagonal 2DTorus network is also communicatively connected to the target computing node with the coordinate value ((x+(k1 - 1) / 2) Mod k1, (y + k2 / 2) Mod k2) and the target computing node with the coordinate value ((x+(k1 - 1) / 2) Mod k1, (y + k2 / 2) Mod k2). Similarly, when k1 is even and k2 is odd, based on the 2DTorus network, the first computing node with the coordinate value (x, y) in the Diagonal 2D Torus network is also communicatively connected to the target computing node with the coordinate value ((x + k1 / 2) Mod k1, (y+(k2 - 1) / 2) Mod k2) and the target computing node with the coordinate value ((x + k1 / 2) Mod k1, (y+(k2 + 1) / 2) Mod k2).

[0181] As an example, see Figure 9 , 20 computing nodes are communicatively connected in the 2D Torus networking mode to form a 5×4 2DTorus network. Among them, 5 computing nodes are arranged in each row and 4 computing nodes are arranged in each column in this 2D Torus network. Based on this 2D Torus network, the first computing node with the coordinate value (0, 0) is also communicatively connected to the target computing node with the coordinate value (2, 2) and the target computing node with the coordinate value (3, 2).

[0182] And so on, after 20 computing nodes are communicatively connected in the 2D Torus networking mode, each computing node is also communicatively connected to the two computing nodes with the farthest communication distance to form a Diagonal 2D Torus network.

[0183] It should be understood that for the first computing node in the network architecture, the above two methods for determining the target computing node are only illustrated by forming Diagonal 1D Torus network, Diagonal 2D Torus network, and Diagonal 3D Torus network. With the emergence of larger-scale Torus networks, based on the technical concept of connecting the computing node with the farthest communication distance in the Torus networking method provided in the embodiments of the present application, the network architecture of higher-dimensional Torus networks can also be improved to form higher-dimensional Diagonal Torus networks. The implementation methods are similar and will not be described one by one here.

[0184] Moreover, regardless of the number of dimensions in which multiple computing nodes in the network architecture are networked and connected, for the first computing node, as long as the total number of computing nodes corresponding to each dimension in the network architecture is determined, the coordinate values of the computing nodes that the first computing node needs to communicate and connect in each dimension can be determined in sequence according to the above method, so as to integrate the coordinate values that the first computing node needs to communicate and connect in at least one dimension, and determine the target computing node that communicates and connects with the first computing node.

[0185] It should also be noted that although the calculation method of the target computing node additionally connected by the first computing node is given in the above embodiments, when the number of target computing nodes is multiple, the first computing node in any of the above embodiments can communicate and connect with all target computing nodes with the farthest communication distance, or only communicate and connect with some target computing nodes. The embodiments of the present application do not limit this.

[0186] In summary, in the embodiments of the present application, after multiple computing nodes in the network architecture are communicatively connected according to the Torus networking method, each computing node also communicates and connects with at least one computing node with the farthest communication distance in the Torus networking method. Therefore, compared with the Torus network, the network architecture shown in the embodiments of the present application not only increases the degree of each computing node, improves the throughput and fault tolerance of the computing node, but also reduces the communication delay between computing nodes, thereby improving the overall communication performance of the network.

[0187] In some embodiments, for the network architecture shown in the above embodiments, when the first computing node in the network architecture communicates and connects with the target computing node, the implementation methods of the communication connection may include the following two.

[0188] In the first connection method, the first computing node and the target computing node are connected by a physical wire.

[0189] Regarding this connection method, reference can be made to the above appendix Figures 4 - 9, connect the first computing node to the target computing through physical wiring.

[0190] In the second connection method, the above network architecture further includes at least one network device, and the first computing node and the target computing node are communicatively connected through the at least one network device.

[0191] Among them, the at least one network device may include one network device or multiple network devices, and the embodiments of the present application do not limit this. The at least one network device may be a switch or other routing devices capable of forwarding data. The embodiments of the present application do not limit the specific type of the network device, as long as it can realize the communication connection between multiple computing nodes.

[0192] In a possible implementation, the first computing node and the target computing node are communicatively connected through the same network device. At this time, the shortest communication route exists between the first computing node and the target computing node, reducing the communication delay between the first computing node and the target computing node and improving the communication performance.

[0193] In the case where the network architecture includes multiple computing nodes, after the multiple computing nodes are communicatively connected according to the Torus networking method, there will be multiple pairs of computing nodes with the farthest communication distance. These multiple pairs of computing nodes can be fully connected through the same network device or through multiple network devices. The embodiments of the present application do not limit this.

[0194] It should be understood that if there is a dimension with an odd total number of computing nodes in the Torus network, then there are two target computing nodes with the farthest communication distance for the first computing node in the Torus network. At this time, the first computing can be communicatively connected to these two target computing nodes through the same network device or through two network devices respectively to communicate with these two target computing nodes. The embodiments of the present application do not limit this.

[0195] Among them, when all the computing nodes with the farthest communication distance are fully connected through one network device, the requirement for the number of ports of the network device is relatively high, and the networking cost when multiple nodes are combined with the network device for networking is also relatively high; when all the computing nodes with the farthest communication distance are fully connected through multiple network devices, the requirement for the number of ports of the network device is relatively low, and the networking cost when multiple nodes are combined with the network device for networking is relatively low. Therefore, it can be flexibly selected according to actual needs, and the embodiments of the present application do not limit this.

[0196] It should be noted that when the first computing node and the target computing node can be communicatively connected through at least one network device, the embodiments of the present application do not limit the specific implementation of all computing nodes in the network architecture being communicatively connected through at least one network device to the computing node with the farthest communication distance in the Torus networking mode.

[0197] In some embodiments, multiple computing nodes included in the network architecture can be organized into multiple Torus networks with the same topology according to the Torus networking mode, and the computing nodes in the multiple Torus networks are fully connected through at least one network device.

[0198] Next, taking the network architecture including 64 computing nodes and at least one network device, and the 64 computing nodes forming four 4×4 2D Torus networks as an example, an exemplary description will be given of the specific manner in which the computing nodes in the multiple Torus networks are fully connected through the at least one network device.

[0199] In the first case, the network architecture includes one network device, and all computing nodes in the multiple Torus networks in the network architecture are communicatively connected through the same network device.

[0200] In other words, when the network architecture includes one network device, all computing nodes in the network architecture are connected to the network device, so that each computing node can communicate with at least one computing node with the farthest communication distance through the network device.

[0201] As an example, referring to Figure 10 , the four 4×4 2D Torus networks included in the network architecture, that is, the 64 computing nodes included are all connected to one network device. At this time, the first computing node with the coordinate value (0, 0) in each 2D Torus network can communicate with the target computing node with the coordinate value (2, 2) through the network device; at the same time, the computing nodes in different 2D Torus networks can also communicate through the network device.

[0202] It can be seen that when the network architecture only includes one network device, all computing nodes in the multiple Torus networks are communicatively connected through the network device, that is, all computing nodes in the network architecture are fully connected through one network device. In this way, based on the network device, direct communication can be achieved between any two computing nodes in the network architecture, and the communication route is the shortest.

[0203] It should be noted that when all the computing nodes in four 4×4 2D Torus networks are fully connected through a network device, the requirement for the number of ports of the network device is 64. Therefore, when all the computing nodes in a network architecture are fully connected through a network device, the number of ports of this network device must be greater than or equal to the total number of computing nodes in this network architecture. Such a network architecture has certain requirements for the number of ports of the network device, and the networking cost is relatively high.

[0204] Based on this, to reduce the multi-node networking cost, when the computing nodes in multiple Torus networks can be fully connected through multiple network devices, the sum of the ports of these multiple network devices being greater than or equal to the total number of computing nodes in this network architecture is sufficient. For example, replacing the above-mentioned 64-port network device with 8 16-port network devices, or 4 16-port network devices, etc., increasing the number of network devices while reducing the number of ports of a single network device, thereby controlling the multi-node networking cost.

[0205] It should be understood that different networking architectures of network devices require different numbers of network devices, which can be flexibly set according to actual needs, and the embodiments of the present application do not limit this.

[0206] When the network architecture includes multiple network devices, the computing nodes with the farthest communication distance in multiple Torus networks can all be communicatively connected through these multiple network devices. At this time, for each Torus network, all the computing nodes in this Torus network can be communicatively connected through these multiple network devices, for details, refer to the following second case; or only some of the computing nodes in this Torus network can be communicatively connected through the same network device, for details, refer to the following third case.

[0207] The second case is that the network architecture includes multiple network devices, and in any Torus network, each computing node and the target computing node with the farthest communication distance are communicatively connected through multiple network devices.

[0208] In a possible implementation manner, if the computing nodes in each Torus network are communicatively connected to the target computing node with the farthest communication distance through multiple network devices, then to achieve the full connection of all computing nodes within the network architecture, a solution of using the Torus networking method + Clos networking method can be adopted to achieve the full connection of all computing nodes in multiple Torus networks through multiple network devices.

[0209] As an example, in a Clos network architecture, after the four 16-port network devices in the first layer communicate and connect with the computing nodes in multiple Torus networks, the four 16-port network devices in the second layer are respectively communicatively connected to the four 16-port network devices in the first layer, thereby achieving full connection between all computing nodes within the network architecture.

[0210] See Figure 11 , taking a network architecture including eight 16-port network devices as an example, all computing nodes in four 4×4 2D Torus networks are fully connected through these eight 16-port network devices. In each 2D Torus network, the 8 computing nodes in the first two columns are connected to one network device in the first layer, and the 8 computing nodes in the last two columns are connected to another network device in the first layer. The two network devices connecting the computing nodes in the same Torus network communicate with each other through one network device in the second layer.

[0211] Among them, the computing nodes with coordinate values (0, 0), (1, 0), (0, 1), (1, 1), (0, 2), (1, 2), (0, 3), and (1, 3) in the first 2D Torus network communicate with the first network device; the computing nodes with coordinate values (2, 0), (3, 0), (2, 1), (3, 1), (2, 2), (3, 2), (2, 3), and (3, 3) communicate with the second network device, and the first network device and the second network device communicate with each other through the fifth network device.

[0212] Thus, the computing node with the coordinate value (0, 0) in the first 2D Torus network can communicate with the computing node with the coordinate value (2, 2) according to the communication route of "the first network device - the fifth network device - the second network device"; similarly, the computing nodes with the coordinate values (0, 1), (0, 2), (0, 3), (1, 0), (1, 1), (1, 2), and (1, 3) in the first 2D Torus network also communicate with the computing nodes with the coordinate values (2, 3), (2, 0), (2, 1), (3, 2), (3, 3), (3, 0), and (3, 1) according to the communication route of "the first network device - the fifth network device - the second network device". Thus, all the computing nodes in the first 2D Torus network are sequentially connected to communicate with the computing node with the farthest communication distance in this first 2D Torus network through "the first network device - the fifth network device - the second network device".

[0213] When networking and connecting according to the above second case, it is only necessary to ensure that the total number of ports of the multiple network devices directly connected to all the computing nodes in multiple Torus networks in the first-layer Clos architecture is greater than or equal to the total number of computing nodes in this network architecture. This networking method has relatively low requirements for the port quantity of network devices.

[0214] It should be noted that Figure 11 Only the computing nodes in the first two columns and the last two columns in the first 2D Torus network are connected to different network devices as an example. In specific implementation, the computing nodes in the first two rows can also be connected to the first network device, and the last two rows can be connected to the second network device. Without exceeding the port quantity of the network devices in the first layer, the embodiments of the present application do not limit the quantity of computing nodes connected to each network device in the first layer, nor the 2D Torus network to which the computing nodes connected to each network device belong.

[0215] In the third case, the network architecture includes multiple network devices, and in any Torus network, some computing nodes communicate and connect with at least one computing node with the farthest communication distance through the same network device.

[0216] In a possible implementation, according to the number of network devices, determine the target network device corresponding to the pair of computing nodes with the farthest communication distance in each Torus network, so as to implement the communication connection between some computing nodes with the farthest communication distance in the Torus network through the target network device.

[0217] In other words, for any Torus network, a part of the computing nodes in the Torus network communicate with at least one computing node with the farthest communication distance through a network device; another part of the computing nodes communicate with at least one computing node with the farthest communication distance through another network device. In this way, in this network architecture, each computing node can communicate with the computing node with the farthest communication distance through a network device. Compared with the above second case, the communication delay between the computing nodes with the farthest communication distance is further reduced, and the communication performance is improved.

[0218] Among them, if the number of multiple network devices is P, and the number of pairs of computing nodes with the farthest communication distance in each Torus network is Q, then for each Torus network, connect Q / P pairs of computing nodes to the same network device. For example, if P = 4 and Q = 8, then connect 2 pairs of computing nodes, that is, 4 computing nodes in each Torus network to the same network device.

[0219] As an example, refer to Figure 12 , taking the network architecture including four 16-port network devices as an example, all computing nodes in four 4×4 2D Torus networks are fully connected through the eight 16-port network devices. In each 2D Torus network, the computing nodes with coordinate values (0, 0) and (2, 2), and the computing nodes with coordinate values (2, 0) and (0, 2) communicate through the first network device; the computing nodes with coordinate values (1, 0) and (3, 2), and the computing nodes with coordinate values (3, 0) and (1, 2) communicate through the second network device; the computing nodes with coordinate values (0, 1) and (2, 3), and the computing nodes with coordinate values (2, 1) and (0, 3) communicate through the third network device; the computing nodes with coordinate values (1, 1) and (3, 3), and the computing nodes with coordinate values (3, 1) and (1, 3) communicate through the fourth network device.

[0220] When networking and connecting according to the third case described above, it is only necessary to ensure that the total number of ports of multiple network devices is greater than or equal to the total number of computing nodes in the network architecture. This networking method also has relatively low requirements for the number of ports of network devices. Moreover, compared with the second case described above, this networking and connecting method can achieve full connection of the above 64 computing nodes by using only four network devices, reducing the composition cost and also reducing the communication delay caused by excessive networking layers of network devices.

[0221] It should be noted that the above Figure 10 - Figure 12 only shows the connection method between the computing nodes of the first 2D Torus network and at least one network device. The connection methods between the computing nodes of the other three 2D Torus networks and the at least one network device are similar and not shown in the figure. Moreover, for the sake of convenience of explanation, the above Figure 10 - Figure 12 only takes the example where the total number of computing nodes in two dimensions of the 2D Torus network is the same and both are even numbers. When the multiple Torus networks formed by multiple computing nodes in the network architecture include multiple dimensions and the total number of computing nodes in multiple dimensions is different, or there is at least one dimension where the total number of computing nodes is an odd number, the connection methods shown in the above three cases can still be referred to, and the computing nodes with the farthest communication distance in multiple Torus networks are connected for communication through at least one network device. The networking and connecting methods are similar, so they will not be elaborated here.

[0222] In summary, when the network architecture includes one network device, the above Figure 10 connection method can be referred to, and the computing nodes in multiple Torus networks are all connected to the same network device, so as to ensure that all computing nodes in multiple Torus networks are fully connected through this network device. When the network architecture includes multiple network devices, the above Figure 11 or Figure 12 connection method can be referred to, and some computing nodes in each Torus network are connected to the same network device, so as to ensure that all computing nodes in multiple Torus networks are fully connected through these multiple network devices. That is to say, when the network architecture includes multiple Torus networks formed by multiple computing nodes and at least one network device, the at least one network device can be used to realize the communication connection between the computing nodes with the farthest communication distance in each Torus network, so as to form a Diagonal Torus network and reduce the communication delay between each computing node under the network architecture.

[0223] In some embodiments, multiple Torus networks with the same topology can be logically combined to form a Diagonal Torus network with a larger or smaller scale. When the first computing nodes in each Torus network are connected to the target computing nodes through at least one network device, the computing nodes at the same position in multiple Torus networks can be connected to the same network device. Based on this, when multiple Torus networks are logically combined to form a Diagonal Torus network with a larger or smaller scale, each computing node in the Diagonal Torus network can communicate with the computing node with the farthest communication distance through the same network device, greatly reducing the communication delay between the computing nodes in this logical network.

[0224] In a possible implementation, as Figure 12 shown, the two computing nodes with the farthest communication distance in each Torus network communicate with each other through the same network device. At the same time, the computing nodes at the same position in multiple Torus networks are connected to the same network device.

[0225] Based on Figure 12 the network architecture shown, the Diagonal Torus networks that can be formed logically by 4 Torus networks include: a Diagonal 1D Torus network containing 4 computing nodes, a 4×4 Diagonal 2D Torus network, and a 4×4×4 Diagonal 3D Torus network.

[0226] When forming a Diagonal 1D Torus network containing 4 computing nodes, taking the first row in the first 2D Torus network as an example, the computing nodes with coordinate values (0, 0), (1, 0), (2, 0), and (3, 0) communicate with each other according to the 1D Torus networking method. Moreover, the computing node with coordinate value (0, 0) and the computing node with the farthest communication distance with coordinate value (2, 0) communicate with each other through the first network device, and the computing node with coordinate value (1, 0) and the computing node with the farthest communication distance with coordinate value (3, 0) communicate with each other through the second network device. In this way, in Figure 12 the network architecture shown, the connection method between the computing nodes in each row or each column within a single 4×4 2D Torus network satisfies the actual networking method of the Diagonal 1D Torus network. Therefore, each row or each column of a single 2D Torus network in this network architecture can be logically operated as a Diagonal1DTorus network.

[0227] When constructing a 4×4 Diagonal 2D Torus network, taking the first 2D Torus network as an example, since the computing nodes with coordinate values (0, 0) and (2, 2) in this 2D Torus network, as well as the computing nodes with coordinate values (0, 2) and (2, 0), are interconnected through the first network device, and the other computing nodes with the farthest communication distances in this 2D Torus network are also connected for communication through other network devices. Thus, in Figure 12 the network architecture shown, the network formed by the connection of the computing nodes and network devices of a single 2D Torus network can operate logically as a Diagonal 2D Torus network.

[0228] When constructing a 4×4×4 Diagonal 3D Torus network, since the computing node with coordinate value (0, 0) in the first 2D Torus network and the computing node with coordinate value (2, 2) in the third 2D Torus network are the pair of computing nodes with the farthest communication distances in the 3D Torus network, and in Figure 12 the network architecture shown, these two computing nodes are connected for communication through the first network device, and the other computing nodes with the farthest communication distances in the 3D Torus network are also connected for communication through network devices. Thus, in Figure 12 the network architecture shown, the overall connection relationship of the 64 computing nodes satisfies the actual networking method of the Diagonal 3D Torus network. Therefore, the 4 2D Torus networks in this network architecture can operate logically as a Diagonal 3D Torus network as a whole.

[0229] It should be understood that when multiple computing nodes in a network architecture form a Torus network, a Diagonal Torus network with a dimension equal to or lower than that of the Torus network can also be constructed logically based on this Torus network. For example, in the case where multiple computing nodes form a 4×4 2D Torus network, a Diagonal 1D Torus network with 4 computing nodes or a 4×4 Diagonal 2D Torus network can still be constructed logically according to the above examples.

[0230] Optionally, based on the above technical solution for constructing a logical network topology, when setting communication routes for computing nodes of multiple tenants in a cloud scenario, if full connection has been achieved among multiple computing nodes, different-dimensional logical Diagonal Torus networks can be constructed among the multiple computing nodes to find the shortest communication routes between the computing nodes in the logical Diagonal Torus network, thereby determining the communication routes for mutual communication among the multiple computing nodes. In this way, by leveraging the advantages of shorter communication routes and smaller communication latency in the Diagonal Torus network, other complex fully connected networks can be mapped into the Diagonal Torus network to set the communication routes among multiple computing nodes, thus improving the efficiency of setting node communication routes in complex networks.

[0231] It should be noted that the "logical" Diagonal Torus network here means that full connection has been achieved among multiple computing nodes, but it is not directly constructed into a Diagonal Torus network according to the network architecture provided in the embodiments of the present application. Instead, based on the full connection among multiple computing nodes, the network architecture of multiple computing nodes is converted into the network architecture of a Diagonal Torus network to find the shortest communication route.

[0232] In summary, in the case where the network architecture includes multiple computing nodes and at least one network device, the multiple computing nodes are communicatively connected in accordance with the Torus networking method, and each computing node is further communicatively connected to at least one computing node with the farthest communication distance in the Torus networking method. In this way, the degree of each computing node is increased, the throughput and fault tolerance of the computing nodes, as well as the communication performance between the computing nodes are improved, and at the same time, the communication latency between the computing nodes in the entire network is reduced. Moreover, regardless of whether multiple computing nodes form one or more Torus networks, the computing nodes with the farthest communication distance in each Torus network can be fully connected through network devices, and the communication routes between the computing nodes are shorter.

[0233] Based on the above network architecture, the embodiments of the present application also provide several products including the above network architecture, which will be introduced separately below.

[0234] First, the embodiments of the present application provide a circuit board, which includes multiple computing nodes, and the multiple computing nodes form the network architecture as described above Figure 2 - Figure 9 shown. That is, in the case where the multiple computing nodes are communicatively connected in accordance with the Torus networking method, each computing node is further communicatively connected to a target computing node with the farthest communication distance.

[0235] Among them, each computing node can be a computing chip. After multiple computing chips are fully connected according to the connection manner shown in the above network architecture, the circuit board provided by the embodiment of the present application is obtained.

[0236] Secondly, the embodiment of the present application provides a cabinet, which includes multiple computing nodes and at least one network device, and the multiple computing nodes and at least one network device form the above Figure 10 - Figure 12 shown network architecture. That is, when multiple computing nodes are communicatively connected according to the Torus networking mode to form a Torus network or multiple Torus networks, each computing node is also communicatively connected to the target computing node with the farthest communication distance through at least one network device.

[0237] Among them, each computing node can be a computer or a processor. After multiple computers or processors are fully connected according to the connection manner shown in the above network architecture through at least one network device, the cabinet provided by the embodiment of the present application is obtained.

[0238] In addition, the embodiment of the present application further provides a computing cluster, which includes multiple cabinets, each cabinet includes multiple computing nodes and at least one network device, and the multiple computing nodes and at least one network device form the network architecture as shown above Figures 2 - 12 shown.

[0239] In a possible implementation manner, multiple cabinets can be communicatively connected using the Clos architecture, so as to achieve a full connection of the computing nodes within the computing cluster.

[0240] It should be noted that based on the network architecture provided by the embodiment of the present application and the technical concept of connecting the computing nodes with the farthest communication distance pairs, as the number of computing nodes increases and the networking scale expands, the technical solution of the present application is also applicable to other products, such as data centers, cloud data centers, etc., and the embodiment of the present application does not limit this.

[0241] Based on the above network architecture, the embodiment of the present application further provides a method for generating a network architecture. This method is applied to a manager to form target network topologies with different topological types based on the above network architecture by the manager, and can flexibly switch between multiple network topologies supported by the network architecture. Among them, there is a communication connection between the manager and multiple computing nodes and at least one network device in the network architecture. The manager can receive information or data input by the user and process it; at the same time, the manager can also issue instructions, configuration information, relevant data, etc. to multiple computing nodes and at least one network device in response to the user input information.

[0242] In the embodiment of the present application, the manager is configured to determine the routing information between each computing node that constitutes the target network topology in the network architecture according to the target topology type input by the user and in accordance with the target topology type, and send the target topology type and the routing information to each computing node, so that each computing node communicates according to the target network topology and the routing information.

[0243] In this way, in the case of a known network architecture, combined with the user's requirements, the manager can flexibly switch the target network topology formed by multiple computing nodes in the network architecture to find the shortest communication route between each computing node from the logical network, reducing the communication delay of each computing node in the network architecture and improving the flexibility of the routing algorithm setting at the same time.

[0244] Next, the computer device provided in the embodiment of the present application will be introduced first. The computer device may specifically be a server, a terminal device, etc. When implementing the method for generating the network architecture in the embodiment of the present application, the computing device may be a manager, or a single network device, or any computing node, so as to correspondingly execute the relevant steps in the method for generating the network architecture.

[0245] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a computer device shown according to the embodiment of the present application. The computer device includes at least one processor 1301, a communication bus 1302, a memory 1303, and at least one communication interface 1304.

[0246] The processor 1301 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or may be one or more integrated circuits for implementing the solution of the present application. For example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0247] The communication bus 1302 is used to transmit information between the above components. The communication bus 1302 may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation,Figure 13 It is represented by only one thick line, but it does not mean that there is only one bus or one type of bus.

[0248] The memory 1303 can be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed optical disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1303 can exist independently and be connected to the processor 1301 through the communication bus 1302. The memory 1303 can also be integrated with the processor 1301.

[0249] The communication interface 1304 uses any device such as a transceiver for communicating with other devices or communication networks. The communication interface 1304 includes a wired communication interface and can also include a wireless communication interface. Among them, the wired communication interface can be, for example, an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface can be a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.

[0250] As an example, the processor 1301 can include one or more CPUs, such as Figure 13 CPU0 and CPU1 shown in

[0251] As an example, the computer device can include multiple processors, such as Figure 13 the processor 1301 and the processor 1305 shown in

[0252] In some embodiments, the computer device may further include an output device and an input device (not shown in the figure). The output device communicates with the processor 1301 and can display information in various ways. For example, the output device may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device communicates with the processor 1301 and can receive user input in various ways. For example, the input device may be a mouse, a keyboard, a touch screen device, or a sensing device, etc.

[0253] In some embodiments, the memory 1303 is used to store the program code 1310 for executing the solution of this application, and the processor 1301 can execute the program code 1310 stored in the memory 1303. The program code 1310 may include one or more software modules, and the computer device can implement the network architecture switching method provided in the following Figure 13 embodiments through the processor 1301 and the program code 1310 in the memory 1303.

[0254] Please refer to Figure 14 , Figure 14 which is a flowchart of a method for generating a network architecture provided by an embodiment of this application. As introduced above, the network architecture includes multiple computing nodes and at least one network device. The multiple computing nodes are organized into at least one Torus network with the same topology according to the Torus networking method. The method includes the following steps.

[0255] Step 1401: Receive network resource configuration information, where the network resource configuration information includes a target topology type, and the target topology type indicates the target network topology currently expected to be generated. In the target network topology, the two computing nodes with the farthest communication distance communicate with each other through at least one network device.

[0256] That is to say, the target network topology expected to be generated is the Diagonal Torus network provided by the embodiments of this application.

[0257] Referring to the foregoing introduction, the network architecture includes multiple computing nodes and at least one network device. The multiple computing nodes form one or more Torus networks. Therefore, in the case where the network architecture includes one network device and the multiple computing nodes are organized into one Torus network in the Torus networking mode, if all the computing nodes in one Torus network can be connected to the network device, at this time, the network device acts as a manager to execute steps 1401 - 1403; in the case where the network device includes multiple network devices and the multiple computing nodes are organized into multiple Torus networks with the same topology in the Torus networking mode, the computing nodes with the farthest communication distance in the multiple Torus networks can be communicatively connected through the network device. At this time, a manager independent of the multiple computing nodes and the multiple network devices can execute steps 1401 - 1403.

[0258] In some embodiments, in the case where the multiple computing nodes are organized into multiple Torus networks with the same topology in the Torus networking mode and the network architecture includes multiple network devices, the computing nodes at the same position in the multiple Torus networks are connected to the same network device. In this way, the computing nodes in the multiple Torus networks can all achieve full connection through the multiple network devices.

[0259] That is to say, in the case where the network device includes multiple computing nodes and at least one network device, if the computing nodes in one or more Torus networks achieve full connection through at least one network device, then the one or more Torus networks and the at least one network device can logically form Diagonal Torus networks of multiple dimensions. In other words, based on the connection relationship between the computing nodes in one or more Torus networks and the at least one network device, the network architecture can be mapped to multiple network topologies.

[0260] Among them, based on Figure 10 - Figure 12 the network architecture shown, the target network topology expected to be switched in step 1401 above can be any one of Diagonal 1D Torus network, Diagonal 2D Torus network, and Diagonal 3D Torus network.

[0261] Optionally, the network resource configuration information in step 1401 above can also include the number of networks, and the number of networks indicates the number of target network topologies to be switched. In this way, based on the network resource configuration information, the manager executes the following step 1302 according to the target topology type and the number of networks to determine the routing information between the computing nodes that form one or more target network topologies among the multiple computing nodes.

[0262] As an example, the target network topology in the network resource configuration information may be a Diagonal 2D Torus network, and the number of networks is 5. The network resource configuration information indicates that 5 Diagonal 2D Torus networks are generated based on the network architecture, but the number of computing nodes included in each Diagonal 2D Torus network is not limited.

[0263] Optionally, the network resource configuration information may further include the target number of nodes of the generated target network topology, and the network scale of the target network topology to be generated is described by the target number of nodes. In this way, based on the network resource configuration information, the manager executes step 1402 below according to the target topology type and the target number of nodes, so as to determine the computing nodes constituting the target network topology and the routing information between the respective computing nodes from among the multiple computing nodes included in the network architecture.

[0264] As an example, the target network topology in the network resource configuration information may be a Diagonal 1D Torus network, and the target number of nodes is 5. The network resource configuration information indicates that a Diagonal 1D Torus network including 5 computing nodes is generated based on the network architecture, but the number of Diagonal 1D Torus networks to be formed is not limited.

[0265] Step 1402: Determine the routing information between the respective computing nodes constituting the target network topology among the multiple computing nodes according to the target topology type.

[0266] It should be noted that based on the network architecture, the routing between the computing nodes includes routing information based on the Torus network and / or routing information based on network devices.

[0267] As an example, see Figure 12, the communication route between the computing node with coordinate value (0, 0) and the computing node with coordinate value (1, 1) in the first 2D Torus network is the computing node with coordinate value (0, 0) - the computing node with coordinate value (1, 0) - the computing node with coordinate value (1, 1), or the computing node with coordinate value (0, 0) - the computing node with coordinate value (0, 1) - the computing node with coordinate value (1, 1). Therefore, the routing information between the computing node with coordinate value (0, 0) and the computing node with coordinate value (1, 1) is routing information based on the Torus network. The communication route between the computing node with coordinate value (0, 0) and the computing node with coordinate value (2, 2) in the first 2D Torus network is the computing node with coordinate value (0, 0) - the first network device - the computing node with coordinate value (2, 2). Therefore, the routing information between the computing node with coordinate value (0, 0) and the computing node with coordinate value (2, 2) is routing information based on the network device. The communication route between the computing node with coordinate value (0, 0) and the computing node with coordinate value (2, 1) in the first 2D Torus network can be the computing node with coordinate value (0, 0) - the first network device - the computing node with coordinate value (1, 1) - the computing node with coordinate value (2, 1). Therefore, the routing information between the computing node with coordinate value (0, 0) and the computing node with coordinate value (2, 1) is routing information based on the Torus network and the network device.

[0268] In some embodiments, when two computing nodes with the farthest communication distance in the same Torus network are communicatively connected through the same network device, the above routing information includes the shortest routes between the respective computing nodes.

[0269] In a possible implementation manner, the implementation process of the above step 1402 can be: The manager determines the routing information when the respective computing nodes communicate with each other according to the connection relationship between the respective computing nodes in the target network topology. Among them, for any computing node in the target network topology, the corresponding routing information includes the communication routes between this computing node and all other computing nodes in the target network topology.

[0270] It should be noted that the premise for the manager to execute the above step 1402 is that the multiple computing nodes and at least one network device included in this network architecture can form a network corresponding to the above target topology type, that is, support switching to the above target topology type. Therefore, if the target topology type is a network topology type supported by this network architecture, the manager executes the above step 1402; if the target topology type is not a network topology type supported by this network architecture, that is, this network architecture does not support generating the above target topology type, then the manager does not execute the above step 1402, and the process of the network architecture switching method ends.

[0271] Optionally, when the network architecture does not support generating the above target topology type, the manager can also output topology generation failure information to prompt the user that the network architecture does not support generating the above target topology type. Of course, the topology generation failure information can also carry the network topology types supported by the current network architecture for generation.

[0272] In some embodiments, the implementation process of determining the network topology types supported by the network architecture can be: the manager determines the network topology types supported by the network architecture for generation based on the connection relationships of multiple computing nodes and at least one network device.

[0273] That is, before executing the network architecture generation method provided in the embodiments of the present application, the manager can pre-obtain or receive the connection relationships of multiple computing nodes and at least one network device in the network architecture, so as to determine the network topology types supported by the network architecture for generation. Of course, the user can also directly send the network topology types supported by the network architecture to the manager, and the embodiments of the present application do not limit this.

[0274] Optionally, after determining the network topology types supported by the network architecture for generation, the manager can store the network topology types supported by the network architecture for generation, so as to quickly determine whether the network architecture supports generating the target topology type by comparing the target topology type with the network topology types supported by the network architecture when receiving the network resource configuration information input by the user.

[0275] Step 1403: Send the target topology type and routing information to each computing node.

[0276] Each computing node receives the target topology type and routing information, and stores the target topology type and routing information, so that when it is necessary to form a network according to the target network topology and process computing tasks, it can refer to the routing information and communicate with at least one computing node within the target network topology to transmit relevant data, information or instructions.

[0277] In summary, in the embodiments of the present application, for the connection relationship between multiple computing nodes and at least one network device in a network architecture that supports generating multiple network topology types, the manager determines the communication routes between the computing nodes that form the target network topology in the network architecture according to the target topology type included in the network resource configuration information. In this way, based on the same network architecture, by generating the target network topology, the application scenarios of the network architecture are improved to meet the requirements of different networking scales and the requirements of different data processing tasks for the number of computing nodes. Moreover, since the target network topology is a Diagonal Torus network constructed based on the networking method of the embodiments of the present application, the computing nodes communicate according to the target network topology and routing information, which can reduce the communication delay during multi-node communication.

[0278] Figure 15 FIG. 4 is a schematic structural diagram of a network architecture generation device provided by an embodiment of the present application. The network architecture includes multiple computing nodes and at least one network device. The multiple computing nodes are organized into at least one Torus network with the same topology according to the Torus networking method. The network architecture generation device can be implemented by software, hardware, or a combination of both to be part or all of a management device. Refer to Figure 15 FIG. 4, the network architecture generation device 1500 includes: an information receiving module 1501, a topology generation module 1502, and an information sending module 1503.

[0279] The information receiving module 1501 is configured to receive network resource configuration information, where the network resource configuration information includes a target topology type, and the target topology type indicates the target network topology currently expected to be generated. The two computing nodes with the farthest communication distance in the target network topology are communicatively connected through at least one network device; for the detailed implementation process, refer to the corresponding content in the above method embodiments, and details are not described here again.

[0280] The topology generation module 1502 is configured to determine the routing information between the computing nodes that form the target network topology among the multiple computing nodes according to the target topology type; for the detailed implementation process, refer to the corresponding content in the above method embodiments, and details are not described here again.

[0281] The information sending module 1503 is configured to send the target topology type and the routing information to each computing node. For the detailed implementation process, refer to the corresponding content in the above method embodiments, and details are not described here again.

[0282] Optionally, when the multiple computing nodes are organized into multiple Torus networks with the same topology according to the Torus networking method and the network architecture includes multiple network devices, the computing nodes located at the same position in the multiple Torus networks are connected to the same network device.

[0283] Optionally, two computing nodes with the farthest communication distance in the same Torus network are communicatively connected through the same network device, and the routing information includes the shortest routes between each computing node.

[0284] Optionally, before determining the routing information between each computing node that forms the target network topology among multiple computing nodes according to the target topology type, the apparatus further includes:

[0285] A topology judgment module, configured to, if the target topology type is a network topology type supported by the network architecture to generate, execute the step of determining the routing information between each computing node that forms the target network topology among multiple computing nodes according to the target topology type.

[0286] Optionally, the apparatus further includes:

[0287] A topology determination module, configured to determine the network topology type supported by the network architecture to generate based on the connection relationship between multiple computing nodes and at least one network device.

[0288] In the embodiment of the present application, when the connection relationship between multiple computing nodes and at least one network device in the network architecture supports generating multiple network topology types, the generating apparatus of the network architecture determines the communication routes between each computing node that forms the target network topology in the network architecture according to the target topology type included in the network resource configuration information. In this way, based on the same network architecture, by generating the target network topology, the application scenarios of the network architecture are improved to meet the requirements of different networking scales and the requirements of different data processing tasks for the number of computing nodes. Moreover, since the target network topology is a Diagonal Torus network constructed based on the networking method of the embodiment of the present application, the computing nodes can communicate according to the target network topology and routing information, which can reduce the communication delay during multi-node communication.

[0289] It should be noted that: when the generating apparatus of the network architecture provided in the above embodiment manages the Diagonal Torus network that can be logically generated by multiple Torus in the network architecture, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the generating apparatus of the network architecture provided in the above embodiment and the embodiment of the network architecture generating method belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.

[0290] The embodiment of the present application further provides a schematic structural diagram of a network architecture generation device. The network architecture includes multiple computing nodes and at least one network device. The multiple computing nodes are formed into at least one Torus network with the same topology according to the Torus networking method. The network architecture generation device can be implemented by software, hardware, or a combination of both as part or all of the computing nodes. The network architecture generation device includes: an information receiving module.

[0291] Among them, the information receiving module is used to receive a target topology type and routing information. The target topology type indicates a target network topology expected to be generated based on the network architecture. In the target network topology, the two computing nodes with the farthest communication distance communicate through at least one network device. The routing information indicates the communication routes between the respective computing nodes that make up the target network topology.

[0292] The embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that direct a computing device in a computing cluster to execute the network architecture generation method provided by the embodiment of the present application.

[0293] The embodiment of the present application further provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on the computing devices included in the computing device cluster, it causes the computing device cluster to execute the network architecture generation method provided by the embodiment of the present application.

[0294] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, a non-transitory storage medium.

[0295] It should be understood that the "plurality" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; the "and / or" herein is merely a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit being different.

[0296] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0297] The above are the embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A network architecture, characterized in that, The network architecture includes multiple computing nodes, and the multiple computing nodes are communicatively connected in a Torus networking mode. Moreover, a first computing node among the multiple computing nodes is also communicatively connected to a target computing node; Wherein, the first computing node is any one of the multiple computing nodes, and the target computing node is the computing node with the farthest communication distance from the first computing node when the multiple computing nodes are communicatively connected in the Torus networking mode.

2. The network architecture according to claim 1, characterized in that, The Torus networking mode includes at least one dimension. If the total number of computing nodes in each dimension of the at least one dimension is an even number, the target computing node includes one computing node, and the coordinate value of the one computing node in the first dimension is a first coordinate value, and the first coordinate value is the value obtained by taking the modulus of d1 + k1 / 2 with respect to k1; Wherein, the first dimension is any one of the at least one dimension, d1 is the coordinate value of the first computing node in the first dimension, and k1 is the total number of computing nodes in the first dimension.

3. The network architecture according to claim 1, characterized in that The Torus networking mode includes at least one dimension. If there is a dimension in the at least one dimension where the total number of computing nodes is an odd number, the target computing node includes at least two computing nodes; When the total number of computing nodes in the first dimension is an even number, the coordinate values of the at least two computing nodes in the first dimension are both the first coordinate value, and the first coordinate value is the value obtained by taking the modulus of d1 + k1 / 2 with respect to k1; when the total number of computing nodes in the first dimension is an odd number, the coordinate values of the first part of the at least two computing nodes in the first dimension are the second coordinate value, and the coordinate values of the second part of the at least two computing nodes in the first dimension are the third coordinate value, and the number of the first part of the computing nodes is equal to the number of the second part of the computing nodes, the second coordinate value is the value obtained by taking the modulus of d1 + (k1 - 1) / 2 with respect to k1, and the third coordinate value is the value obtained by taking the modulus of d1 + (k1 + 1) / 2 with respect to k1; Wherein, the first dimension is any one of the at least one dimension, d1 is the coordinate value of the first computing node in the first dimension, and k1 is the total number of computing nodes in the first dimension.

4. The network architecture according to claim 3, characterized in that, If the total number of computing nodes in each dimension of the at least one dimension is an odd number, the at least two computing nodes include a second computing node and a third computing node; The coordinate value of the second computing node in the first dimension is the second coordinate value, and the coordinate value in the second dimension is the fourth coordinate value, and the fourth coordinate value is the value obtained by taking the modulus of d2 + (k2 - 1) / 2 with respect to k2; The coordinate value of the third computing node in the first dimension is the third coordinate value, and the coordinate value in the second dimension is the fifth coordinate value, and the fifth coordinate value is the value obtained by taking the modulus of d2 + (k2 + 1) / 2 with respect to k2; Wherein, the second dimension is any one of the at least one dimension other than the first dimension, d2 is the coordinate value of the first computing node in the second dimension, and k2 is the total number of computing nodes in the second dimension.

5. The network architecture according to any one of claims 1-4, characterized in that The network architecture further includes at least one network device, and the first computing node and the target computing node are communicatively connected through the at least one network device.

6. The network architecture according to claim 5, characterized in that, The first computing node and the target computing node are communicatively connected through the same network device.

7. The network architecture according to claim 5 or 6, characterized in that, The multiple computing nodes are formed into multiple Torus networks with the same topology according to the Torus networking mode, and the computing nodes located at the same position in the multiple Torus networks are connected to the same network device.

8. A circuit, characterized in that, The circuit includes multiple computing nodes, and the multiple computing nodes form the network architecture according to any one of claims 1-4.

9. A cabinet, characterized in that, The cabinet includes multiple computing nodes and at least one network device, and the multiple computing nodes and the at least one network device form the network architecture according to any one of claims 1-7.

10. A computing cluster, characterized in that, The computing cluster includes multiple cabinets, each cabinet includes multiple computing nodes and at least one network device, and the multiple computing nodes and the at least one network device form the network architecture according to any one of claims 1-7.

11. A method for generating a network architecture, characterized in that, The network architecture includes multiple computing nodes and at least one network device, and the multiple computing nodes are formed into at least one Torus network with the same topology according to the Torus networking mode; the method includes: Receiving network resource configuration information, the network resource configuration information includes a target topology type, the target topology type indicates a target network topology currently expected to be generated, and the two computing nodes with the farthest communication distance in the target network topology are communicatively connected through the at least one network device; Determining routing information between each of the computing nodes that form the target network topology among the multiple computing nodes according to the target topology type; Sending the target topology type and the routing information to each of the computing nodes.

12. The method according to claim 11, wherein When the multiple computing nodes are formed into multiple Torus networks with the same topology according to the Torus networking mode and the network architecture includes multiple network devices, the computing nodes located at the same position in the multiple Torus networks are connected to the same network device.

13. The method according to claim 11 or 12, characterized in that The two computing nodes with the farthest communication distance in the same Torus network are communicatively connected through the same network device, and the routing information includes the shortest route between each of the computing nodes.

14. The method according to any one of claims 11-13, characterized in that, Before determining the routing information between each of the computing nodes that form the target network topology among the multiple computing nodes according to the target topology type, the method further includes: If the target topology type is a network topology type supported by the network architecture, then perform the step of determining the routing information between each of the computing nodes that form the target network topology among the multiple computing nodes according to the target topology type.

15. The method according to claim 14, wherein The method further includes: Determine the network topology types supported by the generated network architecture based on the connection relationships among the multiple computing nodes and the at least one network device.

16. A method for generating a network architecture, characterized in that, The network architecture includes multiple computing nodes and at least one network device. The multiple computing nodes are organized into at least one Torus network with the same topology according to the Torus networking method. The method includes: Receiving a target topology type and routing information. The target topology type indicates a target network topology expected to be generated based on the network architecture. In the target network topology, the two computing nodes with the farthest communication distance communicate and connect through the at least one network device. The routing information indicates the communication routes among the respective computing nodes that make up the target network topology.

17. An electronic device, characterized in that, The computer device includes a processor and a memory; The memory is used to store a computer program; The processor is configured to execute the computer program to implement the method according to any one of claims 11-15, or the processor is configured to execute the computer program to implement the method according to claim 16.

18. A computer-readable storage medium, characterized in that, Instructions are stored in the storage medium. When the instructions run on the computer device, the computer device is caused to execute the method according to any one of claims 11-15, or the processor is configured to execute the computer program to implement the method according to claim 16.

19. A computer program product comprising instructions, characterized in that, When the instructions run on the computer device, the computer device is caused to execute the method according to any one of claims 11-15, or the processor is configured to execute the computer program to implement the method according to claim 16.

Citation Information

Cited By

  • All-to-all communication system, communication method and computer equipment

    CN120614255A

  • An all-to-all communication system, a communication method and a computer device

    CN120614255B

  • Network architecture, network architecture generation method, and related device

    WO2025156994A1