Data routing system and data routing method

By employing a hierarchical structure of nested surround network topology, fat tree network topology, and fully interconnected network topology, the high cost and power consumption of AI cluster network topology are addressed. This enables efficient and low-cost data routing, adapts to traffic routing requirements under different parallel modes, and meets the computing needs of large-scale AI clusters.

WO2025185263A9PCT designated stage Publication Date: 2025-12-26HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/137673
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2024-12-09
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively reduce the cost and power consumption of AI cluster network topologies. In particular, the networking scale of fat tree network topologies is limited by the number of switch ports, making it impossible to meet the interconnection requirements of computing cards in large-scale or ultra-large-scale AI clusters.

Method used

A hierarchical network structure is adopted, including nested surrounding network topology, fat tree network topology, and fully interconnected network topology. Large-scale or ultra-large-scale networks are built with network devices with a small number of ports. By combining shortest path routing algorithms and routing strategies for different types of traffic, efficient data routing between computing nodes is achieved.

Benefits of technology

It effectively reduces network costs and power consumption, improves network scalability, meets the computing needs of large-scale or ultra-large-scale AI clusters, reduces expensive global fiber optic links, and improves communication performance and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137673_26122025_PF_FP_ABST
    Figure CN2024137673_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application is a data routing system. The system comprises a first topological network, a second topological network and a third topological network, wherein the network structure of the first topological network is a toroidal network topology structure consisting of a first node group, the first node group comprises a first number of network nodes, and the network nodes in the toroidal network topology structure have connection relationships with upstream nodes and downstream nodes in each dimension; the second topological network is a fat-tree network topology structure consisting of the network nodes in the first topological network and a switch; there is a cascade relationship between the network nodes in the first topological network and the switch; the first topological network is nested in the third topological network, and the network structure of the third topological network is a fully connected network topology structure consisting of a second node group; and the second node group comprises a second number of first node groups, and there is a connection relationship between the first node groups. Thus, large-scale or ultra-large-scale networking can be realized at a relatively low cost.
Need to check novelty before this filing date? Find Prior Art

Description

A data routing system and a data routing method

[0001] This application claims priority to Chinese Patent Application No. 202410276266.2, filed with the State Intellectual Property Office of China on March 7, 2024, entitled "A Data Routing System and Data Routing Method", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer network technology, and in particular to a data routing system, a data routing method, a method for constructing a data routing system, a management device, a computer-readable storage medium, and a computer program product. Background Technology

[0003] As business scale continues to grow, more and more applications require large-scale data center deployments. Taking artificial intelligence (AI) applications as an example, intelligent question-answering systems based on Generative Pre-trained Transformer (GPT) models are booming, and large AI models (also simply called large models or models) represented by GPT are rapidly expanding into various fields. Large AI models can be trained or inferred through AI clusters. As the parameter scale of large AI models continues to evolve, for example from hundreds of billions to trillions, or even tens of trillions, the computing power requirements of AI clusters (such as data centers deploying AI applications) are growing rapidly.

[0004] The rapidly growing demand for computing power poses a significant challenge to the scaling of AI cluster networks, requiring networks exceeding 100,000 computing cards to meet current needs. Designing network topologies for AI clusters and other data centers to effectively reduce network costs and power consumption is a crucial issue that network architecture design must address. Currently, large-scale AI clusters typically employ a fat-tree network topology; however, the scaling of a fat-tree network is limited by the number of switch ports. For example, a Layer 3 fat-tree network built with mainstream 64-port switches can only provide a maximum of 65,536 interconnected ports, which is insufficient to meet the interconnection requirements of hundreds of thousands of computing cards in large AI models.

[0005] How to achieve large-scale or ultra-large-scale networking at a relatively low cost has become a key concern in the industry. Summary of the Invention

[0006] This application provides a data routing system that constructs hierarchical network structures such as nested surround network topologies, fat-tree network topologies, and fully interconnected network topologies using network devices with a small number of ports. This enables the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. This application also provides a data routing method, a method for constructing a data routing system, a management device, a computer-readable storage medium, and a computer program product corresponding to the aforementioned data routing system.

[0007] Firstly, this application provides a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network. The first topology network has a network structure of a surrounding network topology composed of a first group of nodes. The first group of nodes includes a first number of network nodes, and the network nodes in the surrounding network topology are connected to upstream and downstream nodes in each dimension. The second topology network is a fat-tree network topology composed of the network nodes and switches in the first topology network. There is a cascading relationship between the network nodes and switches in the first topology network. The first topology network is nested within the third topology network. The third topology network has a network structure of a fully interconnected network topology composed of a second group of nodes. The second group of nodes includes a second number of first group of nodes, and there is a connection relationship between the second number of first group of nodes.

[0008] This data routing system can build hierarchical network structures such as nested ring network topology, fat tree network topology, and fully interconnected network topology using network devices with a small number of ports (e.g., 9-port network devices, where 9 ports refer to the 9 output ports of the network device, and output ports can also be simply referred to as output ports). This enables the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability.

[0009] In some possible implementations, network nodes belonging to the same first node group in the second topology network can be directly connected to the same switch. Specifically, when the number of ports on a single switch is sufficient to meet the cascading requirements of network nodes in the first topology network, network nodes belonging to the same first node group can be directly connected to the same switch. This allows traffic to hop within the first node group and enables one-hop reachability, thus improving network performance.

[0010] In some possible implementations, the first node group comprises multiple subgroups. The second topology network comprises multiple first switches corresponding to the multiple subgroups and at least one second switch. In the second topology network, network nodes belonging to the same subgroup are directly connected to the first switch corresponding to the subgroup, and the multiple first switches are connected upwards to at least one second switch.

[0011] In this way, a multi-layer fat-tree network topology, including multiple first switches and at least one second switch, can satisfy the cascading requirements of network nodes in a first node group with a large number of network nodes. Even if there are many network nodes in the first node group, traffic can be hopped within the first node group through the first and second switches, and can achieve reachability within a few hops. For example, when the fat-tree network has two layers, traffic can reach the first node group in up to three hops.

[0012] In some possible implementations, subgroups can be obtained by dividing the network into racks, rack groups, or data centers. By dividing the first node groups of different specifications into subgroups according to their size, network nodes belonging to the same subgroup are directly connected to the first switch corresponding to that subgroup, and the first switches corresponding to multiple subgroups are connected upwards to the second switch. This allows the cascading requirements of a large number of network nodes to be met by increasing the number of layers in the fat tree network.

[0013] In some possible implementations, each network node includes n ports. Of these n ports, m ports are used to connect to network nodes surrounding the network topology. These m ports include ports for connecting upstream nodes in each dimension and ports for connecting downstream nodes in each dimension.

[0014] In this system, network nodes can connect to upstream and downstream nodes in various dimensions by providing a small number of ports, thereby enabling interconnection of adjacent nodes within the first node group. This system has good performance for neighbor communication with strong locality, especially for global all-reduce traffic in the model parallel and data parallel modes of AI applications. Communication can be achieved through ring algorithms, with communication mainly occurring between adjacent nodes, making it naturally adapted to ring networks.

[0015] In some possible implementations, h out of n ports are used to connect network nodes in a fully interconnected network topology consisting of a second group of nodes. The second number is less than or equal to... Among them, Len j The number of network nodes in the j-th dimension surrounding the network topology is represented by q, and the number of dimensions surrounding the network topology is represented by q.

[0016] In this way, there can be at least one global link between different first node groups, which can provide direct routes for global communication, alleviate global network congestion, and reduce expensive global fiber optic links compared to fat tree network topology.

[0017] In some possible implementations, the data routing system also includes a fourth topology network. The third topology network is nested within the fourth topology network. The fourth topology network has a fully interconnected network topology consisting of a third group of nodes, where each third group of nodes includes a third number of second group of nodes, and these third group of second group of nodes are interconnected.

[0018] This system compresses the size of the surrounding network by increasing the number of layers in the fully interconnected network topology, such as by using a nested two-level fully interconnected network topology. By increasing the number of network layers, it reduces the network diameter, communication latency, communication performance, and network throughput.

[0019] In some possible implementations, k out of n ports are used to connect network nodes in a fully interconnected network topology consisting of a third node group, where the number of the third group is less than or equal to k * the number of nodes in the second node group + 1.

[0020] In this way, different groups of second nodes can include at least one global link, which can provide a direct route for global communication, alleviate global network congestion, and reduce expensive global fiber optic links compared to fat tree network topology.

[0021] In some possible implementations, the network node is either a switching node independent of the compute node, or a routing engine integrated into the compute node. When the network node uses a routing engine integrated into the compute node, interconnection with switches can be eliminated, further reducing physical costs and power consumption.

[0022] In some possible implementations, the node coordinates of a network node are determined based on the network node's number in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The network node is used to determine data routing paths based on its node coordinates, and to implement data routing between computing nodes according to these routing paths.

[0023] Thus, data routing paths can be determined based on node coordinates using a shortest path routing algorithm. The shortest path routing algorithm provides the shortest communication distance between the source and destination nodes, with the lowest communication latency. Since the data routing system of this application is a hierarchical network structure nested around various network topologies, including thick-tree and fully interconnected network topologies, the deterministic shortest path routing algorithm adapted to the topology features is computationally fast and facilitates hardware implementation.

[0024] In some possible implementations, computing nodes are used to execute AI computing tasks in parallel. The parallelism of these tasks includes at least one of model parallelism, pipelined parallelism, hybrid expert parallelism, or data parallelism. The data routing system of this application, by embedding different types of network topologies to form a hierarchical network, can meet the traffic routing needs under different parallelism methods in AI scenarios.

[0025] In some possible implementations, the first topology network is used to route packets in the global reduction (all-reduce) traffic, the second topology network is used to route packets in many-to-many (all-to-all) or point-to-point (point-to-point) traffic, and the third topology network is used to route packets in point-to-point traffic. Here, the all-reduce traffic can be traffic in model-parallel or data-parallel modes, the all-to-all traffic can be traffic in a hybrid expert-parallel mode, and the point-to-point traffic can be send / receive traffic in a pipelined parallel mode. In this way, routing can be performed through networks with appropriate topologies based on the characteristics of different types of traffic, thus reducing communication latency.

[0026] Secondly, this application provides a data routing method. This method is applied to a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network. The first topology network has a network structure of a surrounding network topology composed of a first group of nodes, which includes a first number of network nodes. The network nodes in the surrounding network topology are connected to upstream and downstream nodes in each dimension. The second topology network has a network structure of a fat-tree network topology composed of the network nodes and switches in the first topology network. There is a cascading relationship between the network nodes and switches in the first topology network. The first topology network is nested within the third topology network. The third topology network has a network structure of a fully interconnected network topology composed of second group of nodes. The second group of nodes includes a second number of first group of nodes, and there is a connection relationship between the second number of first group of nodes.

[0027] In the data routing system, the current node receives the data packet to be routed. The current node is either the source node of the data packet or a node on the path from the source node to the destination node. Then, the current node determines the data routing path based on the coordinates of the destination node. The node coordinates are determined based on the network node's number in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The current node then routes the data packet according to the data routing path.

[0028] This method introduces the node coordinates of the destination node in the data routing system, adapts the data routing system to the topology features of the data routing system to perform data routing, provides the shortest distance communication between the source node and the destination node, achieves low communication latency, and is simple to implement with high availability.

[0029] In some possible implementations, when the number of the first node group to which the current node belongs is equal to the number of the first node group to which the destination node belongs, the current node determines the routing port in each dimension based on the offset values ​​of the numbers of the destination node and the current node in each dimension of the first topology network.

[0030] If the number of the first node group to which the current node belongs is equal to the number of the first node group to which the destination node belongs, it means that the current node and the destination node are in the same first node group. The current node can be routed by surrounding the network topology. When routing by surrounding the network topology, the routing port can be determined based on the offset values ​​of the numbers of the destination node and the current node in each dimension of the first topology network. This allows the routing path to be determined and low communication latency to be achieved.

[0031] In some possible implementations, the number of network nodes in the i-th dimension of the first topology network is L. i When the offset between the destination node and the current node's index in the i-th dimension of the first topology is greater than 0 and less than or equal to L... i / 2, or less than -L i / 2, determine the routing port in the i-th dimension as the port used to connect downstream nodes; when the offset value of the destination node and the current node in the i-th dimension of the first topology network is greater than L i / 2, or greater than or equal to -L i If the value is 2 and less than 0, the routing port in the i-th dimension is determined as the port used to connect to the upstream node. In this way, the shortest path around the network topology can be determined, achieving low communication latency.

[0032] In some possible implementations, the data packets are those in the global reduction all-reduce traffic. All-reduce traffic can communicate using a ring algorithm, where communication mainly occurs between adjacent nodes. This is naturally suited to ring networks, and routing through the network structure to the first topology of the ring network can achieve good performance.

[0033] In some possible implementations, the current node can determine the traffic type based on the coordinates of both the destination and source nodes. The traffic type identifies whether a data packet is intra-group or inter-group traffic. When the traffic type identifies the data packet as inter-group traffic, the current node determines the forwarding port based on the destination node's coordinates. This allows for the redirection of inter-group traffic via a switch, achieving a hierarchical separation of inter-group and intra-group traffic and preventing congestion.

[0034] In some possible implementations, when cross-group traffic is traffic sent from the local node to the external node, the current node can determine the coordinates of network nodes within the group that have a direct link to the destination node, or determine the coordinates of network nodes within the group that have a direct link to a hop node leading to the destination node, based on the node coordinates of the destination node. Then, it determines the port number of the switch to which the network nodes within the group connect. When cross-group traffic is traffic sent from the external node to the local node, the current node determines the forwarding port based on the node coordinates of the destination node and the cascading relationship between the destination node and the switch.

[0035] The current node can perform data forwarding within a fat tree based on the shortest routing path algorithm for the aforementioned cross-group traffic. Specifically, cross-group traffic destined for external nodes can be forwarded to local nodes within the same group that have a direct link to the destination node; for cross-group traffic originating from external nodes and destined for the local node, the data routing path can be determined based on the destination node's coordinates, allowing for hopping within the first node group. Forwarding can be performed via a switch in both the source node's and destination node's node groups, achieving low-latency communication.

[0036] In some possible implementations, data packets can be either all-to-all traffic or point-to-point traffic. For all-to-all traffic in a hybrid expert parallel mode, forwarding can be achieved through switches, thus enabling traffic layering and avoiding congestion. For point-to-point traffic in a pipelined parallel mode, routing is typically performed using a fully interconnected network topology. If the current node is not a node in the same node group as the destination node, traffic can be redirected within the current node's node group via a switch. Specifically, the redirection is to a node in the same node group as the destination node, and then routing is performed to the destination node group via a direct link. This shortens the communication path and achieves low-latency communication.

[0037] In some possible implementations, the data routing system also includes a fourth topology network. A third topology network is nested within the fourth topology network, and the fourth topology network has a fully interconnected network topology consisting of a third group of nodes. The third group of nodes includes a third number of second group of nodes. Connections exist between these third number of second group of nodes. Node coordinates are determined based on the node's number within each dimension of the first topology network, the port number of the switch to which the node is connected, and the numbers of the first and second group of nodes to which the node belongs.

[0038] Accordingly, when the number of the second node group to which the current node belongs is equal to the number of the second node group to which the destination node belongs, and the number of the first node group to which the current node belongs is not equal to the number of the first node group to which the destination node belongs, the current node can determine whether there is a first direct link between the current node and the first node group to which the destination node belongs.

[0039] If yes, the current node can determine that the data routing path includes forwarding data packets from the first direct link to the first node group where the destination node is located; if no, the current node can determine that the data routing path includes forwarding data packets to the second hop node, and the second hop node and the first node group where the destination node is located have a second direct link.

[0040] This method identifies whether the current node is directly connected to the first node group of the destination node when the current node and the destination node are in different first node groups within the same second node group. If so, routing is performed through the direct link; otherwise, routing is first performed to the corresponding jump node, thereby enabling routing through the shortest path.

[0041] In some possible implementations, when the number of the second node group to which the current node belongs is not equal to the number of the second node group to which the destination node belongs, the current node can determine whether there is a third direct link between the second node group to which the current node belongs and the second node group to which the destination node belongs.

[0042] If yes, the current node can determine that the data routing path includes forwarding data packets from the third direct link to the second node group where the destination node is located; if no, the current node can determine that the data routing path includes forwarding data packets to the first hop node. The first hop node and the second node group where the destination node is located have a fourth direct link.

[0043] When the current node and the destination node are in different second node groups, this method identifies whether the current node is directly connected to the second node group where the destination node is located. If so, routing is performed through the direct link; otherwise, routing is first performed to the corresponding jump node. This enables routing through the shortest path.

[0044] In some possible implementations, the current node can also determine whether the number of the first node group to which the first hop node belongs is equal to the number of the first node group to which the current node belongs. If not, the current node can also determine that the data routing path includes forwarding data packets to the ferry node. The ferry node and the first node group to which the first hop node belongs have a fifth direct link.

[0045] In this method, when the current node and the first hop node are in different first node groups, the current node can first route to the ferry node to reach the first node group where the first hop node is located. Then, the data packet can be routed to the first hop node through the intra-group routing of the first node group, and then routing can be performed based on the direct link of the hop node, thereby achieving the shortest path routing.

[0046] In some possible implementations, the data packets are data packets within point-to-point traffic. This point-to-point traffic can be send / receive traffic in a pipelined parallel mode. For the aforementioned point-to-point traffic, routing through a fully interconnected network topology can satisfy the communication requirements of the point-to-point traffic.

[0047] Thirdly, this application provides a method for constructing a data routing system. This method can be applied to management devices, and includes:

[0048] A switch and multiple network nodes are determined for constructing the data routing system, the multiple network nodes being divided into a second number of first node groups, the first node group comprising a first number of network nodes;

[0049] A configuration file is sent to the switch and the plurality of network nodes, so that the network nodes in the first node group establish connections with upstream nodes and downstream nodes in each dimension according to the network structure of the data routing system described in the configuration file, to form a first topology network with a surrounding network topology, and a second topology network cascaded with the switch to form a fat tree network topology, and establish connections between the second number of first node groups to form a third topology network with a fully interconnected network topology. The data routing system includes the first topology network, the second topology network, and the third topology network, with the first topology network nested within the third topology network.

[0050] This method identifies switches and multiple network nodes for constructing the data routing system and instructs these nodes and switches to establish connections according to the network structure of the data routing system in the configuration file. This allows for the construction of nested surround network topologies, fat-tree network topologies, and fully interconnected network topologies, thereby enabling the building of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. Specifically, the surround network topology accelerates the application of localized communication characteristics and supports multi-tenant isolation, accelerating resource pooling; the fat-tree network topology can be used for global traffic forwarding and as a replacement route for fault recovery; for ultra-large-scale supernodes, a Layer 2 fat-tree topology can be used for non-blocking expansion to achieve global traffic forwarding; and the fully interconnected network topology provides direct routes for global communication, alleviating global network congestion and reducing the need for expensive global fiber optic links compared to the fat-tree network topology.

[0051] In some possible implementations, multiple network nodes are divided into a third number of second node groups, which in turn include a second number of first node groups. The configuration file is also used to establish connections between network nodes in the first node groups and the third number of second node groups, forming a fourth topology network with a fully interconnected network topology. The data routing system also includes this fourth topology network, with the third topology network nested within it.

[0052] By nesting multi-layered fully interconnected network topologies, more ports can be provided and more computing nodes can be connected, thereby enabling the construction of ultra-large-scale networks to meet business needs.

[0053] In some possible implementations, the management device can also check the connection relationships of the multiple network nodes based on the network structure of the data routing system described in the configuration file. If the connection relationships of the multiple network nodes are inconsistent with the connection relationships in the network structure of the data routing system, the management device can send a prompt message to the user, suggesting adjustments to the connection relationships of the multiple network nodes. This ensures the accuracy and reliability of the network configuration.

[0054] In some possible implementations, the management device can also generate a routing table based on the node coordinates of multiple network nodes. The node coordinates are determined based on the network node's ID in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the ID of the first node group to which the network node belongs. The management device can then issue entries from the routing table related to at least one of the multiple network nodes.

[0055] In this process, the management device sends the routing table entries related to the network node to at least one network node, which can reduce the storage resources occupied by the routing table entries in the network node and save storage space.

[0056] Fourthly, this application provides a management device. The management device includes at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor is configured to execute the instructions stored in the at least one memory to cause the management device to perform a method for constructing a data routing system as described in the third aspect or any implementation thereof.

[0057] Fifthly, this application provides a computer-readable storage medium storing instructions that instruct a management device to perform the method for constructing a data routing system as described in the first aspect or any implementation thereof.

[0058] In a sixth aspect, this application provides a computer program product containing instructions that, when run on a management device, cause the management device to perform the method for constructing a data routing system as described in the first aspect or any implementation thereof.

[0059] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0060] To more clearly illustrate the technical methods of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below.

[0061] Figure 1 is an architecture diagram of a data routing system provided in this application;

[0062] Figure 2 is a schematic diagram of the ports of a node in a data routing system provided in this application;

[0063] Figure 3 is a schematic diagram of a surround network topology provided in this application;

[0064] Figure 4 is a schematic diagram of a 9-port router provided in this application;

[0065] Figure 5 is a schematic diagram of the structure of an integrated routing engine in a processor provided in this application;

[0066] Figure 6 is a schematic diagram of a ring network topology provided in this application;

[0067] Figure 7 is a schematic diagram of a hierarchical network with a nested surrounding network topology, a fat tree network topology, and a two-level fully interconnected network topology provided in this application.

[0068] Figure 8 is a flowchart of a data routing method provided in this application;

[0069] Figure 9 is a schematic diagram of the data routing process in a data routing system provided in this application;

[0070] Figure 10 is a schematic diagram of data routing via a virtual channel provided in this application;

[0071] Figure 11 is a schematic diagram of a hierarchical network provided in this application, including a nested surrounding network topology, a fat tree network topology, and a first-level fully interconnected network topology.

[0072] Figure 12 is a schematic diagram of a port bandwidth allocation method provided in this application;

[0073] Figure 13 is a schematic diagram of the data routing process in a data routing system provided in this application;

[0074] Figure 14 is a schematic diagram of a fault recovery based on a backup cabinet provided in this application;

[0075] Figure 15 is a flowchart of a method for constructing a data routing system provided in this application;

[0076] Figure 16 is a structural schematic diagram of a management device provided in this application;

[0077] Figure 17 is a schematic diagram of the structure of a computing device provided in this application;

[0078] Figure 18 is a schematic diagram of the structure of a computing device cluster provided in this application;

[0079] Figure 19 is a schematic diagram of another computing device cluster provided in this application;

[0080] Figure 20 is a schematic diagram of another computing device cluster provided in this application. Detailed Implementation

[0081] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.

[0082] First, some technical terms involved in the embodiments of this application will be introduced.

[0083] Network topology refers to the specific arrangement or connection method among the members of a network (such as network devices or communication devices like switches and routers). It is generally divided into physical, real, online structures, or logical, virtual, and programmatic structures. If two networks may differ in their internal physical wiring and the distances between nodes, but their member connection structures are identical, then the two networks can be considered to have the same network topology.

[0084] Networking, short for network construction, typically involves designing a suitable network topology based on network requirements, including scale and performance needs, and then constructing the network according to that topology. In scenarios such as cloud computing, high-performance computing (HPC), or artificial intelligence (AI) computing, it is usually necessary to connect various computing devices (also called computing nodes) through network devices (such as switches, routers, and other switching nodes) to build large-scale or ultra-large-scale interconnected networks. This provides high-bandwidth, low-latency, and high-throughput efficient communication capabilities, effectively maximizing computing performance.

[0085] Cloud computing, also known as network computing, is an internet-based computing method that allows shared hardware and software resources and information to be provided on demand to various terminals or other devices. These devices then use the infrastructure provided by the service provider to perform calculations. High-performance computing refers to providing more powerful computing performance than traditional computers or servers by aggregating computing capabilities. It is typically used in scenarios such as weather forecasting, astrophysical simulations, molecular dynamics simulations, and earthquake simulations. AI computing is a mathematically intensive process that executes machine learning algorithms. It typically uses acceleration systems and software to extract new insights from large datasets and learn new capabilities in the process. AI computing can be applied to model training or model inference, especially large language model (LLM) training or inference.

[0086] For ease of description, this application mainly uses AI computing scenarios as examples.

[0087] AI computing is typically achieved through AI clusters, such as large-scale AI training clusters. An AI cluster can include multiple computing nodes, which are interconnected via network nodes such as routers. It's important to note that computing and network nodes can be independent or integrated; for example, a routing engine can be integrated into a computing node to perform routing functions on a network node. During training, AI clusters can employ distributed parallel strategies to improve efficiency. Depending on the model size or function, distributed parallel strategies can include, but are not limited to, model parallelism (e.g., tensor parallelism), data parallelism, pipelined parallelism, or expert parallelism.

[0088] Model parallelism (MP) divides the entire model (such as the aforementioned large model like LLM) into multiple sub-models (also called model fragments), and each sub-model is sent to different computing nodes for processing. Each computing node is only responsible for processing its own sub-model, calculating local gradients, and transmitting the gradients to the central node (usually a parameter server) through a communication mechanism.

[0089] Data parallelism (DP) involves dividing the complete dataset into multiple subsets, with each computing node processing a different subset, and then aggregating the updated gradients.

[0090] Pipeline parallelism (PP) refers to dividing the model into multiple stages and distributing them to different computing nodes, with each computing node completing the training in a "relay" manner.

[0091] Expert parallelism, also known as mixture of experts (MoE) parallelism, typically decomposes a task into several subtasks. For each subtask, an expert model is trained, and a gating model is developed. The gating model learns which expert model to use based on the input samples and combines the output results. This strategy dynamically activates parts of the neural network, thereby significantly increasing the number of model parameters without increasing computational cost. For example, sparse large models with trillions of parameters can be trained using expert parallelism.

[0092] Model parallelism primarily involves global all-reduce traffic, with data volumes in the gigabyte (GB) range. Furthermore, communication and computation in model parallelism are non-overlapping. Expert parallelism has a significant proportion (e.g., 50% or more) of many-to-many all-to-all (or all2all) global traffic, also in the GB range. Communication and computation in expert parallelism are also non-overlapping. Therefore, for model parallelism and expert parallelism, high bandwidth and high throughput are typically required to accelerate communication performance. Data parallelism primarily involves all-reduce traffic, with data volumes in the GB range. However, communication and computation in data parallelism can overlap (or be masked or hidden), so communication optimization is not a core requirement. Pipeline parallelism mainly involves point-to-point send and receive communication, with data volumes in the megabyte (MB) range. The bandwidth requirement is simply high throughput across nodes. Pipeline parallelism mainly involves point-to-point communication between the sender and receiver, with data volume in the MB range. The bandwidth requirement is simply point-to-point communication across nodes.

[0093] Depending on the model and hardware specifications, different parallel strategies can be combined to train the model. Taking a dense transformer model with hundreds of billions of parameters, such as the Generative Pre-trained Transformer (GPT) model, as an example, GPT3 can include 175 billion (denoted as 175B, where B represents billions) parameters, and the GPT3 network consists of 96 layers. The ratio of computation to communication is approximately 8:2.

[0094] Based on the model and hardware specifications, an efficient parallel deployment strategy can be obtained through parallel strategy search, as shown below:

[0095] For the model-parallel domain, the main traffic is all-reduce traffic, with data volumes in the GB range, and communication and computation cannot overlap. Therefore, a single-node deployment can be used, with cards interconnected via a high-bandwidth bus. For the pipelined parallel domain, the main traffic is send / receive (or send / recv) traffic, with data volumes in the MB range. Communication and computation can overlap, and deployment can be done between nodes via a network. For the data-parallel domain, the main traffic is all-reduce traffic, with data volumes in the GB range. Communication and computation can overlap, so deployment can also be done between nodes via a network.

[0096] For sparse large models with trillions of parameters, communication can account for up to 69.3%. The traffic analysis for the aforementioned sparse large model with trillions of parameters is shown below:

[0097] Table 1: Traffic Analysis of Trillions of Sparse AI Large Models

[0098] Among them, the single-step communication time of the sparse large model with trillions of parameters can be XX12 milliseconds (ms), and the total communication time can be XX15 ms. "XX12" and "XX15" are numbers processed by masking, indicating that the single-step communication time and the total communication time are in the thousands of milliseconds.

[0099] Currently, many AI model training methods, especially for large AI models, typically employ a combination of parallel strategies for distributed parallel training, which generates different types of traffic. For example, training a trillion-parallel AI model can utilize a combination of model parallelism, data parallelism, pipelined parallelism, and hybrid expert parallelism strategies. Correspondingly, the traffic generated during training can include all-reduce traffic, all-to-all traffic, and point-to-point send / recv traffic (also known as peer-to-peer traffic). How to configure the network to meet the routing requirements of different types of traffic has become a key focus in the industry.

[0100] The industry offers various network topologies for networking. Among them, the fat tree network topology (also known as a fat tree network or fat tree topology) is a tree-shaped network topology from leaf nodes (or leaf nodes) to the root node, where network bandwidth does not converge. It is a classic high-performance network topology, boasting advantages such as excellent performance, system stability, and wide adaptability, making it the mainstream network interconnection solution for data centers (such as AI clusters or data centers used for training AI models). The dragonfly network topology (also known as a dragonfly topology or dragonfly network) is a fully interconnected network topology within and between groups. Compared to the fat tree topology, the dragonfly topology can reduce network layers and the number of optical fibers and modules, offering higher cost-effectiveness. The surround network topology (also known as a surround topology or surround network) has low networking costs, strong scalability, and good performance for neighbor communication with strong locality. Using a ring topology as an example of a Torus topology, for all-reduce traffic in model-parallel and data-parallel modes of AI applications, communication can be achieved through a ring algorithm. Communication mainly occurs between adjacent nodes, which is naturally adapted to the Torus topology.

[0101] Among them, the Dragonfly topology is based on switch expansion, and its network scale is limited by the switch port specifications. Based on the current mainstream 64-port high-performance switches, the maximum network scale is only 270,000 ports, which cannot meet the networking requirements of large-scale AI clusters. The communication latency of the Torus topology scales linearly with the network size, which has certain limitations. Moreover, the Torus topology is a congested network. Due to the limitations of the topology and routing algorithm, the network throughput of global traffic is low. Therefore, it is not friendly to all-to-all global traffic in the hybrid expert parallel mode, and its performance is very poor for large-scale all-to-all global traffic.

[0102] Given the advantages of fat-tree topologies such as high network throughput, non-blocking operation, and stable performance, large-scale AI clusters typically employ fat-tree network architectures. Using a mainstream 64-port high-performance switch as an example, a three-layer fat-tree network built with a 64-port high-performance switch can provide a maximum of 65,536 interconnected ports. However, this is still insufficient to meet the interconnection needs of hundreds of thousands of computing cards in large AI models, such as hundreds of thousands of graphics processing units (GPUs). Increasing the number of layers in the fat-tree network to increase the number of interconnected ports significantly reduces the port utilization of the switches. Specifically, the port utilization of switches in a fat-tree network is typically below 33%, and it decreases further with increasing layers. For example, the port utilization of a three-layer fat-tree is only 20%. Furthermore, large-scale fat-tree networks require a large amount of fiber optic cables and modules to connect the various layers, greatly increasing reliability pressure and making it difficult to cope with the networking needs of ultra-large-scale networks. In addition, the port bandwidth of mainstream high-performance switches has reached 400 gigabits per second (Gbps). The power consumption of a 64-port high-performance switch is about 200 to 400 watts (W), while the power consumption of an optical module is about 13 watts. If optical modules are used for connection, the power consumption of the switch can reach as high as 1000 watts.

[0103] In view of this, this application provides a low-cost, highly reliable data routing system. This data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network.

[0104] The first topology network has a surrounding network topology (such as a torus) consisting of a first group of nodes. The first group of nodes includes a first number of network nodes. The network nodes in the surrounding network topology are connected to upstream and downstream nodes in each dimension. The second topology network has a fat-tree network topology consisting of the network nodes and switches in the first topology network. The network nodes and switches in the first topology network are cascaded. The switches can be categorized based on their deployment location, such as top-of-rack (TOR) switches, end-of-row (EOR) switches, and middle-of-row (MOR) switches. The first topology network is also nested within a third topology network. Nesting refers to insertion or embedding; specifically, the first topology network can be embedded as a component or unit of the third topology network. The third topology network can have a full-mesh network topology consisting of a second group of nodes, which includes a second number of first group of nodes, and these second number of first group of nodes are connected to each other.

[0105] This data routing system can construct a hierarchical network structure with nested Torus topology, fat tree topology, and full mesh topology using network devices with a small number of ports (e.g., 9-port network devices; to avoid confusion, this application does not consider access ports). This enables the construction of large-scale or ultra-large-scale networks, such as ultra-large-scale networks with 17.3 million ports, effectively reducing network costs and power consumption and improving network scalability.

[0106] Furthermore, for scenarios with relatively regular and predictable traffic, such as AI computing scenarios, this data routing system can layer traffic based on the communication characteristics of the scenario (or the applications and models within the scenario, such as AI applications or large AI models). The Torus topology is naturally adapted to all-reduce traffic, the fat tree topology can meet the routing requirements of all-to-all traffic within a supernode, or it can redirect global point-to-point traffic (such as send / receive traffic, also known as Send / Recv traffic). The full-mesh topology provides direct routes for global communication, such as point-to-point traffic in a pipelined parallel mode. By integrating these multiple topologies for networking, the number of switches, network links, fiber optic links, and optical modules can be reduced, improving system reliability. This data routing system is load-affine, and by layering traffic routing, it effectively reduces network costs and energy consumption while providing high-performance, highly reliable, and highly available communication.

[0107] To make the technical solution of this application clearer and easier to understand, the system architecture of the data routing system of this application is described below with reference to the accompanying drawings.

[0108] Figure 1 shows a schematic diagram of the architecture of a data routing system. The data routing system 10 is used to implement data routing between computing nodes. The data routing system 10 includes multiple network nodes (hereinafter sometimes collectively referred to as "nodes 110"). The multiple nodes 110 are located in a communication network 150 that includes at least a first topology network (e.g., 120-1, 120-2, ..., 120-Q, hereinafter sometimes collectively referred to as "first topology network 120"), a second topology network 130, and a third topology network 140.

[0109] In embodiments of this application, a first topology network 120 is nested within a third topology network 140. The first topology network 120 has a network structure that is a surrounding network topology composed of a first node group (e.g., 120-1, 120-2, ..., 120-Q, where Q is a positive integer greater than 1). The first node group includes a first number of network nodes. For example, the first node group may include nodes 110-1, 110-2, ..., 110-P, where P is a positive integer greater than 1. The network nodes in the surrounding network topology are connected to upstream and downstream nodes in each dimension. Taking node 110-2 as an example, node 110-2's upstream and downstream nodes in a certain dimension are nodes 110-1 and 110-3, respectively, and node 110-2 is connected to nodes 110-1 and 110-3.

[0110] The second topology network 130 has a fat-tree network topology consisting of network nodes in the first topology network 120 and a switch 132. The switch 132 is used to route traffic from the network nodes in the first topology network 120. Depending on its deployment location, the switch 132 may include, but is not limited to, a TOR switch, an EOR switch, or a MOR switch. There is a cascading relationship between the network nodes in the first topology network 120 and the switch 132. For example, each network node in the first topology network 120 can be connected to at least one port of the switch 132, and different network nodes can be connected to different ports.

[0111] The first topology network 120 is nested within the third topology network 140. The network structure of the third topology network 140 is a fully interconnected network topology composed of a second group of nodes. The second group of nodes includes a second number of first group of nodes, such as 120-1, 120-2, ..., 120-Q. These second number of first group of nodes are interconnected. As illustrated in Figure 1, 120-1, 120-2, ..., 120-Q can be interconnected. It should be noted that the interconnection between node groups can be between nodes in different node groups, for example, a node in 120-1 and a node in 120-2 can be interconnected.

[0112] The data routing system 10 may include a single-layer fully interconnected network or multiple layers of fully interconnected networks. When the data routing system 10 includes multiple layers of fully interconnected networks, these multiple layers of fully interconnected networks may also be nested. For example, a first-layer fully interconnected network may be nested within a second-layer fully interconnected network, and so on, with the (L-1)th layer of fully interconnected network nested within the Lth layer of fully interconnected network.

[0113] This application does not limit the number of layers (or levels) of the fully interconnected network in the data routing system 10. The number of layers of the fully interconnected network can be configured according to the required network scale. The network scale can be represented by the total number of ports. For example, when the total number of ports is less than or equal to a first threshold, configuring a single-layer fully interconnected network topology in the data routing system 10 can meet the requirements. As another example, when the total number of ports is greater than the first threshold and less than or equal to a second threshold, configuring a two-layer fully interconnected network topology in the data routing system 10 can meet the requirements. It should be noted that the aforementioned first threshold can be the maximum number of ports that the data routing system 10 can provide when a single-layer fully interconnected network topology is nested, and the aforementioned second threshold can be the maximum number of ports that the data routing system 10 can provide when a two-layer fully interconnected network topology is used.

[0114] Based on this, the data routing system can also include a fourth topology network, and the third topology network 140 can be nested within the fourth topology network. Similar to the third topology network 140, the network structure of the fourth topology network can also be a fully interconnected network topology. The main difference between the fourth topology network and the third topology network 140 is that the fourth topology network is a fully interconnected network topology based on a third group of nodes. This third group of nodes includes a third number of second group of nodes, and these third group of nodes are interconnected. The third group of nodes is a node group formed by network nodes in the third topology network 140, thus enabling the third topology network 140 to be nested within the fourth topology network.

[0115] Similar to a fully interconnected network topology, a fat-tree network topology can be single-layered or multi-layered. In some possible implementations, the number of ports on a single switch is sufficient to meet the cascading requirements of network nodes in the first topology network. Correspondingly, in the second topology network, network nodes belonging to the same first node group can be directly connected to the same switch. For example, a first node group may include network nodes in a 4x4x4 cabinet. In other words, if the number of network nodes in the first node group is less than or equal to 64, a 64-port switch can meet the cascading requirements of all network nodes in the cabinet, and each network node in the cabinet can be directly connected to at least one port of the aforementioned 64-port switch. In other possible implementations, the first topology network includes a large number of network nodes, and the number of ports on a single switch is insufficient to meet the cascading requirements of the network nodes in the first topology network. Therefore, multiple switches can be used to meet the cascading requirements of the network nodes in the first topology network. Specifically, the first node group includes multiple subgroups, where subgroups can be obtained by dividing according to cabinets, cabinet groups, or data centers. For example, when the first node group includes 1024 network nodes, such as network nodes in a super node formed by 16 8x8 cabinets, subgroups can be obtained by dividing the network according to the cabinets. The second topology network includes multiple first switches corresponding to multiple subgroups and at least one second switch. In the second topology network, network nodes belonging to the same subgroup are directly connected to the first switch corresponding to the subgroup. It should be noted that each subgroup can correspond to one or more first switches. Multiple first switches can be connected upwards to at least one second switch. The first switches can be switches located within cabinets, such as TOR switches, while the second switches can be switches located between cabinets.

[0116] Next, the network topology will be explained in conjunction with the port connections of the network nodes. Referring to Figure 2, each network node can include n ports, such as port 111-1, port 111-2, ..., 111-N. Each network node can provide m ports for connecting network nodes in the surrounding network topology. The surrounding network topology can include one or more dimensions. For example, the surrounding network topology can be a three-dimensional (3D) surrounding network topology (such as a 3D torus). Accordingly, the m ports include ports for connecting upstream nodes in each dimension and ports for connecting downstream nodes in each dimension. Where m is greater than or equal to 2. As shown in Figure 3, taking a three-dimensional surrounding network topology as an example, the X dimension of a standard three-dimensional surrounding network topology can provide X+ ports and X- ports representing positive and negative directions. The X+ port of each upstream node connects to the X- port of the downstream node, forming a surrounding network.

[0117] Each network node can also provide at least one port for connecting to a switch in the fat-tree network topology. For example, when the first node group includes 64 network nodes, the switch can be a 64-port switch, and each network node can provide one port to connect to one port of the 64-port switch.

[0118] Each network node's remaining ports can be used to connect to network nodes in a fully interconnected network topology. The number of remaining ports can be greater than or equal to one. In addition to the ports connecting to network nodes in the surrounding network topology and the ports connecting to switches in the fat-tree network topology, at least one port on a network node can connect to network nodes in other first-node groups to form a fully interconnected network topology. For example, a network node can connect to network nodes in other first-node groups through one port to form a third topology network. Furthermore, a network node can connect to network nodes in other second-node groups through another port to form a fourth topology network.

[0119] In the embodiment of Figure 1, node 110 can be implemented as a physical node or a virtual node. The specific implementations of physical nodes and virtual nodes are described below.

[0120] In some embodiments, node 110 can be a switching node independent of the computing nodes, such as a standalone switch or router. Node 110 can be a low-port switch or low-port router. Expanding using low-port switches or low-port routers results in low networking costs and multi-path characteristics. Multi-path characteristics refer to the use of multiple paths between nodes to avoid single points of failure. For ease of description, a low-port router is used as an example. A low-port router is a router with a small number of ports compared to a high-cardinality router. For example, a low-port router can be a 9-port router, or an 11-port, 12-port, 13-port, or 14-port router, while a high-cardinality router can be a 48-port router or a 64-port router.

[0121] Figure 4 illustrates a schematic diagram of a 9-port router, which includes multiple communication engines, a crossbar matrix, and nine ports. The communication engines are used for communication between the router or routing engine ports and the processor to perform Direct Memory Access (DMA) operations, data read / write, memory copying, and other operations. The crossbar matrix is ​​used for communication between the router or routing engine ports and the communication engines to enable internal data forwarding in low-port mode.

[0122] As shown in Figure 4, the nine ports can include six ports for building a wraparound network topology, one port for building a fat-tree network topology, and two ports for building a fully interconnected network topology. The six ports for building the wraparound network topology can include two ports each in the X, Y, and Z dimensions. Each dimension has two directions: positive and negative. Ports LinkX+ and LinkX- are responsible for interconnecting adjacent nodes in the X dimension; LinkY+ and LinkY- are responsible for interconnecting adjacent nodes in the Y dimension; and LinkZ+ and LinkZ- are responsible for interconnecting adjacent nodes in the Z dimension. It should be noted that in the wraparound network topology, the first and last nodes in the same dimension can also be considered adjacent nodes. Taking the X dimension as an example, when the X dimension includes four network nodes, the first network node is connected not only to the second network node but also to the last network node. The port for building the fat-tree network topology can include Link-SW, used to connect to a switch, such as a 64-port TOR switch within the rack. The two ports used to construct the fully interconnected network topology include L-Link, which implements the first-layer fully interconnected network (such as the aforementioned third topology network), and G-Link, which implements the second-layer fully interconnected network (such as the aforementioned fourth topology network). In L-Link, L can stand for local, and in G-Link, G can stand for global. L-Link enables interconnection between L1-level topology (first topology network) groups (referred to as L1 groups, e.g., the first node group), while G-Link enables interconnection between L3-level topology (third topology network) groups (referred to as L3 groups, e.g., the second node group).

[0123] In other embodiments, node 110 may be a routing engine integrated into the computing node. This routing engine may be a routing engine integrated into a processor. The processor may be a graphics processing unit (GPU), a neural network processing unit (NPU), or other accelerator device, or it may be a central processing unit (CPU).

[0124] Figure 5 illustrates a schematic diagram of a routing engine integrated into a processor. The processor may include multiple cores and a Last Level Cache (LLC). Processors typically employ multi-level caching, with each level of cache being faster than the next. When a core needs to access a block of data or an instruction, it first checks the nearest L1 cache. If the data exists, it indicates a cache hit; otherwise, it indicates a cache miss, requiring further checks of the next L1 cache. In some possible implementations, the processor can be a CPU / GPU memory chip integrating High Bandwidth Memory (HBM). The HBM can be a stack of multiple Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM, or DDR for short), which can then be packaged together with the CPU / GPU die to achieve a large-capacity, high-bandwidth DDR array.

[0125] The structure of the routing engine integrated in the processor can be referenced from the structure of the low-port router in Figure 4. The routing engine can include a communication engine, a crossbar switch matrix, and ports. In the example in Figure 5, the processor integrates or has a built-in 9-port routing engine, requiring only 6 ports to achieve expansion in the X, Y, and Z dimensions, where Kx, Ky, and Kz represent the length of each dimension. The length of each dimension can be characterized by the number of nodes in each dimension. Each dimension can be distinguished into positive and negative directions, with ports Link X+ and Link X- responsible for interconnecting adjacent nodes in the X dimension; Link Y+ and Link Y- responsible for interconnecting adjacent nodes in the Y dimension; and Link Z+ and Link Z- responsible for interconnecting adjacent nodes in the Z dimension. In this embodiment, to control the network diameter, the length of each of the X, Y, and Z dimensions is 4. The 9-port routing engine provides one port, Link SW, for connecting to a rack switch, such as a TOR switch, to form a fat-tree topology network structure. The remaining two ports, L-Link and G-Link, are used to implement a two-layer fully interconnected network topology.

[0126] The processor core can communicate with the routing engine via an internal bus. This internal bus can include, but is not limited to, a Peripheral Component Interconnect Express (PCIe) bus or a Compute eXpress Link (CXL) bus.

[0127] It should be understood that the processor, router, or routing engine of this application may include more or fewer components, functions, or modules. In some embodiments, the processor may also include components such as resonant circuits, and may also be connected to external devices such as storage devices. For example, a processor with an integrated routing engine may also include a memory management unit (MMU) and a memory controller (MC) for managing and controlling external or extended memory such as DDR. The aforementioned memory management unit and memory controller can communicate with the extended DDR via a unified bus (UB). Similarly, the processor may also connect to or extend a solid-state drive (SSD), wherein the processor assembly can communicate with the SSD via the UB bus or Ctrl bus. Furthermore, the processor may connect to or extend devices such as accelerators. In some embodiments, the router or routing engine may be any device that implements data transmission, such as a network device like a switch.

[0128] It should be noted that Figures 4 and 5 illustrate with a 9-port example. In other possible implementations of this application, the routing engine integrated in the low-port router or processor may also include other numbers of ports, such as 8, 10, 11, 12, or 13 ports. Taking an 8-port example, the routing engine integrated in the low-port router or processor can construct a surround network topology using 6 of the 8 ports, a fat-tree network topology using 1 of the 8 ports, and a fully interconnected network topology using the remaining 1 port.

[0129] In this application, computing nodes can be used to execute AI computing tasks in parallel. Computing nodes can be standalone nodes or nodes integrated with a routing engine, such as processors with integrated routing engines. The parallelism of the AI ​​computing tasks includes at least one of model parallelism, pipelined parallelism, hybrid expert parallelism, or data parallelism. Accordingly, the first topology network of the data routing system can be used to route data packets in all-reduce traffic, such as data packets in all-reduce traffic in model parallelism or data parallelism mode; the second topology network is used to route data packets in many-to-many all-to-all traffic or point-to-point traffic, such as routing data packets in all-to-all traffic in hybrid expert parallelism mode, or to enable data packets in point-to-point traffic (such as send / receive traffic in pipelined parallelism mode) to jump within the first node group; and the third topology network is used to route data packets in point-to-point traffic, such as routing send / receive traffic in pipelined parallelism mode.

[0130] To make the technical solution of this application clearer and easier to understand, the following example illustrates the construction of a Torus using a 9-port router or routing engine, and then the construction of a hierarchical network consisting of nested Torus, fat tree, and 2-level full mesh.

[0131] Large-scale AI models with hundreds of billions of parameters primarily employ three parallelism modes: model parallelism, data parallelism, and pipelined parallelism. The main traffic of these large-scale AI models consists of all-reduce traffic from model parallelism and data parallelism, with data volumes typically in the GB range, requiring high bandwidth to accelerate communication. In addition, the large-scale AI models also include pipelined send / receive traffic, with smaller data volumes, typically in the MB range, which does not put much pressure on communication bandwidth.

[0132] All-reduce communication can be transformed into communication between neighboring nodes using the ring algorithm, and the Torus topology is naturally suited to this type of communication. The Torus topology can be expanded based on low-port routers to reduce network costs and energy consumption, and it also features multi-path characteristics. For a 3D Torus network, communication routes can be provided in six directions across the X, Y, and Z dimensions. Direct copper cable connections within the rack can meet the communication requirements of all-reduce communication. Even for cross-rack communication, the Torus architecture can significantly reduce the use of long-distance optical fibers and modules. Therefore, this application can use the Torus topology to construct an L1-level topology to transmit model-parallel and data-parallel all-reduce traffic.

[0133] Large-scale AI models involve extensive parallel computing. If the scale of a supernode formed by multiple computing nodes connected by a network can cover the scale of various parallel computing methods, then the main communication can be concentrated within the supernode, where the large communication bandwidth resources can accelerate communication. However, if the supernode scale is smaller than the scale of various parallel computing methods, a large amount of inter-supernode communication is required. Since the communication bandwidth between supernodes is usually much smaller than the communication bandwidth within a supernode, it will greatly affect communication performance. Therefore, the supernode scale usually needs to be adapted to the scale of parallel computing in large-scale AI models.

[0134] For parallel computing of AI models with hundreds of billions of parameters, a supernode with 64 accelerator cards (e.g., 64 GPUs) can meet the requirements. A standard rack can typically accommodate 64 GPUs; therefore, this application can construct a 4*4*4 3D Torus topology, and a single computer rack can meet the training needs of models with hundreds of billions of parameters.

[0135] Communication between racks for large AI models with hundreds of billions of parameters primarily involves pipelined send / receive traffic, with data volumes in the megabyte range and relatively low bandwidth requirements. Therefore, a full mesh topology can be used for rack networking. Compared to torus topology, full mesh topology offers higher bandwidth splitting. For uniform random traffic, the communication efficiency of full mesh topology is the same as that of fat tree topology, achieving 100% network utilization. When implemented using routers integrated into the processor, direct connections can be made via processor ports without the need for switch expansion. Compared to fat tree topology, this saves 50% on global fiber optic links and modules, effectively reducing network costs and improving network reliability.

[0136] This application allows for nested two-level full mesh topologies for inter-rack expansion. Specifically, building L3 / L4 topologies based on a full mesh, such as the first and second layer fully interconnected networks in a fully interconnected network architecture, can effectively reduce network diameter, communication latency, and alleviate congestion problems caused by global communication. When using a two-level full mesh topology, a network with 17.3 million nodes can be built based on 9-port routers (switches). When using a routing engine integrated into the processor, switch interconnections can be eliminated, further reducing network cost and power consumption.

[0137] It should be noted that since the full mesh interconnect between racks connects each GPU to a different rack, traffic for inter-rack communication needs to be routed between GPU cards within the rack to reach the corresponding direct links. This requires transmission through the rack's Torus topology, which can conflict with all-reduce traffic within the rack. Therefore, to avoid traffic congestion, this application adds a switch, such as a TOR switch within the rack, to handle the routing of pipelined traffic between racks, achieving traffic layering. All-reduce traffic within a rack flows within the Torus topology, while cross-rack traffic flows within the rack via the TOR switch, avoiding congestion. In some possible implementations, the TOR switch can also serve as an alternative route for fault recovery, forwarding traffic. Specifically, the TOR switch within the rack constructs an L2-level topology.

[0138] It should be noted that for a rack containing 64 GPUs, with each GPU connected to one port, a 64-port TOR switch can be configured within the rack to form a fat-tree topology for pipelined traffic routing between racks. In some embodiments, multi-plane switching can be used to increase bandwidth depending on the volume of pipelined data. For example, if 256 ports of switches are needed to connect to the GPUs, four 64-port TOR switches can be used to connect to the GPUs respectively, thus increasing bandwidth at a lower cost.

[0139] The following illustrations, with reference to the accompanying diagrams, demonstrate the construction of a Torus, as well as hierarchical networks consisting of nested Torus, fat trees, and two-level full meshes.

[0140] As shown in Figure 6, the L1-level topology is first constructed using a Torus topology. Specifically, each processor has a built-in 9-port routing engine, requiring only 6 ports (external and external ports), for example, ports 1 to 6, to achieve expansion in the X, Y, and Z dimensions. Here, Kx, Ky, and Kz represent the length of each dimension, indicating the number of network nodes in each dimension. Each dimension is divided into positive and negative directions, with ports Link X+ and Link X- responsible for interconnection in the X dimension; Link Y+ and Link Y- responsible for interconnection in the Y dimension; and Link Z+ and Link Z- responsible for interconnection in the Z dimension. Taking a supernode with 64 compute nodes (e.g., 64 GPUs) as an example, the length of each of the X, Y, and Z dimensions is 4. Therefore, the L1-level topology can connect 4*4*4 = 64 compute nodes.

[0141] Next, referring to Figure 7, the 7th port in the 9-port routing engine, such as Link-SW, can connect to the 64-port TOR switch within the rack to build an L2-level topology. In the L2-level topology, the TOR switch is primarily responsible for routing pipelined send / receive traffic across racks between the 64 compute nodes within the rack, for example, between the 64 GPUs within the rack. Alternatively, when large AI models are trained in a hybrid expert parallel mode, the TOR switch can also be used to route all-to-all traffic in this hybrid expert parallel mode.

[0142] The 8th port of the 9-port routing engine, such as an L-Link, can be used for interconnection between L1-level topology groups. Since there are 64 L-Links within the rack, a full mesh direct connection between up to 65 racks can be achieved. An L3-level topology group can consist of 65 racks, connecting a maximum of 64 * 65 = 4160 compute nodes, primarily responsible for transmitting send / receive traffic between racks.

[0143] The 9th port of a 9-port routing engine, such as a G-Link, can be used for interconnection between L3-level topology groups. Since there are 4160 G-Links within an L3-level topology group, a maximum of 4161 L3-level topology groups can achieve full mesh direct connection. Therefore, an L4-level topology can be composed of 4161 L3-level topologies, and can connect a maximum of 4160 * 4161 = 17,309,760 nodes.

[0144] In this embodiment, if a routing engine is integrated into the processor, 9 ports can meet the networking requirements of 17.3 million GPUs (or compute nodes). By building a direct-connect network, no additional network nodes need to be configured, which can greatly reduce network costs and power consumption. The routing engine integrated in the processor provides 9 external ports. Among them, 6 ports are used for interconnection of neighbor nodes in L1-level topology, the 7th port connects to the TOR switch in the rack to build L2-level topology, the 8th port is used for L3-level full-mesh interconnection, and the 9th port is used for L4-level full-mesh interconnection. The internal 9*9 low-port Cross Bar can meet the internal data forwarding requirements.

[0145] It should be noted that Figures 6 and 7 illustrate examples of using 6 ports out of 9 ports to construct a 3D torus topology, 1 port to construct a fat tree topology, and the remaining 2 ports to construct a two-layer full mesh. In other possible implementations of this application, the routing engine integrated in the low-port router or processor may include other numbers of ports.

[0146] Specifically, each network node includes n ports, of which m ports are used to connect to other network nodes in the surrounding network topology. These m ports include ports for connecting upstream nodes in each dimension and ports for connecting downstream nodes in each dimension. Figures 6 and 7 above illustrate an example of a 3D surrounding network topology constructed with m=6. In practical applications, the surrounding network topology can also be 2D or other dimensional surrounding network topologies.

[0147] In some possible implementations, h out of n ports are used to connect network nodes in the fully interconnected network topology formed by the second node group. The second node group includes a second number of first node groups. Considering the limited number of ports, the maximum value of the second number can be determined based on the number of ports used to connect network nodes in the fully interconnected network topology formed by the second node group and the number of nodes in the first node group. The number of nodes in the first node group can be determined based on the number of nodes around each dimension of the network topology.

[0148] Based on this, the second quantity can satisfy the following formula:

[0149] Where, count 2nd The second quantity is indicated by h, which represents the number of ports used to connect network nodes in the fully interconnected network topology formed by the second group of nodes. j The number of nodes surrounding the j-th dimension of the network topology (also called the length of the j-th dimension) is represented by q, and q represents the number of dimensions surrounding the network topology. This indicates the number of nodes in the first node group, for example, the first number. h can be greater than or equal to 1, so that there is at least one direct global link between different first node groups (or first topology networks).

[0150] For example, when the network topology includes three dimensions (X, Y, Z), the maximum second quantity can be h·Kx·Ky·Kz+1. In the example in Figure 6, Kx, Ky, and Kz can be 4, and h can be 1. Therefore, the maximum second quantity can be 1*64+1=65. It should be noted that in practical applications, h can also be a positive integer greater than 1, such as 2 or 3. This can increase the links between network nodes, reduce the probability of congestion, or connect more groups of first nodes, increasing the network size.

[0151] In some possible implementations, k out of the n ports of a network node are used to connect to network nodes in a fully interconnected network topology consisting of a third node group. The third node group includes a third number of second node groups. Considering the limited number of ports, the maximum value of this third number can be determined based on the number of ports used to connect network nodes in the fully interconnected network topology consisting of the second node groups and the number of nodes in the second node groups. The number of nodes in the second node groups can be determined based on a second number (the number of first node groups in the second node group) and a first number (the number of nodes in the first node group), for example, it can be the product of the first and second numbers.

[0152] Based on this, the third quantity can satisfy the following formula: count 3rd ≤k·count 2nd ·count 1st +1 (2)

[0153] Where, count 1st Represents the first quantity, count 2nd To indicate the second quantity, count 3rd Count represents the third quantity, where k represents the number of ports used to connect network nodes in the fully interconnected network topology formed by the third node group. 2nd ·count 1stThis can be the number of nodes in the second node group. k can be greater than or equal to 1, so that there is at least one direct global link between different second node groups (or second topology networks).

[0154] When the network topology includes three dimensions (X, Y, Z), the maximum second quantity can be h·Kx·Ky·Kz+1, and the maximum third quantity can be k·(h·Kx·Ky·Kz+1)·(Kx·Ky·Kz)+1. In the example in Figure 7, Kx, Ky, and Kz can be 4, and h and k can be 1. Therefore, the maximum third quantity can be 1*(1*64+1)*64+1=4161. Figure 7 illustrates this with h and k as 1. In practical applications, h or k can also be greater than 1, which can increase the links between network nodes and reduce the probability of congestion.

[0155] Based on the aforementioned embodiments, the node coordinates of a network node can be determined based on the network node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. For example, when the data routing system is a hierarchical network with nested surround network topologies, fat tree network topologies, and single-layer fully interconnected network topologies, the node coordinates can be represented by the network node's number within each dimension of the first topology network (such as a surround network topology), the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. Taking a surround network topology as a 3D Torus example, the node coordinates can be represented as {L, P, X, Y, Z}. Here, L represents the number of the first node group to which the network node belongs, P represents the port number of the switch to which the network node is connected, and X, Y, Z represent the network node's number within each dimension of the first topology network.

[0156] Furthermore, the data routing system may also include a fourth topology network, which is a fully interconnected network topology composed of a third node group, nested within the fourth topology network. Accordingly, the node coordinates of a network node can be determined based on the node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the numbers of the first and second node groups to which the network node belongs. For example, when the data routing system is a hierarchical network with nested surround network topologies, fat tree network topologies, and two-layer fully interconnected network topologies, the node coordinates can be represented by the node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the numbers of the first and second node groups to which the network node belongs. For instance, when the surround network topology is a 3D Torus topology, the node coordinates can be represented as {G,L,P,X,Y,Z}. Here, G represents the number of the second node group to which the network node belongs, L represents the number of the first node group to which the network node belongs, P represents the port number of the switch to which the network node belongs, and X, Y, and Z represent the numbers of the network node within each dimension of the first topology network.

[0157] It should be noted that a unified numbering system can be used for routing engines integrated in the processor or for independent low-port routers, low-port switches, and other switching nodes. Correspondingly, network nodes can determine data routing paths based on the node coordinates represented by these numbers, and then implement data routing between computing nodes according to these data routing paths.

[0158] In this network, nodes can determine data routing paths based on node coordinates (such as node numbers) using a shortest path routing algorithm. The shortest path routing algorithm provides the shortest communication distance between the source and destination nodes, has the lowest communication latency, and is the most basic routing algorithm. Since this application uses a hierarchical network nested with Torus, fat tree, and full mesh (e.g., a two-layer full mesh), a deterministic shortest path routing algorithm adapted to the topology features fast computation and ease of hardware implementation.

[0159] The following section describes the specific implementation of the deterministic shortest path routing algorithm based on adaptive topology features for data routing.

[0160] Referring to Figure 8, a flowchart of a data routing method is shown. This method can be applied to the data routing system 10 of the aforementioned embodiment. The data routing system 10 is used to implement data routing between computing nodes. The data routing system 10 includes a first topology network 120, a second topology network 130, and a third topology network 140. The network structure of the first topology network 120 is a surrounding network topology structure composed of a first node group. The first node group includes a first number of network nodes, such as nodes 110-1 to 110-P. The network nodes in the surrounding network topology structure are connected to upstream and downstream nodes in each dimension. The network structure of the second topology network 130 is a fat-tree network topology structure composed of each network node in the first topology network 120 and a switch 132. The network nodes and switches in the first topology network are cascaded. The first topology network 120 is nested within the third topology network 140. The network structure of the third topology network 140 is a fully interconnected network topology structure composed of a second node group. The second node group includes a second number of first node groups. There are connections between the second number of first node groups. In this method, the data routing system 10 may route data packets from the source node to the destination node through at least one network node. The node where the data packet is currently located is called the current node, which can be the source node of the data packet (the source node can also be a network node) or a node on the path from the source node to the destination node. The method includes the following steps:

[0161] S802, The current node receives the data packets to be routed.

[0162] The current node is either the source node of the data packet or a node on the path from the source node to the destination node. The packet header can record information about the source node and the destination node, such as the coordinates of the source node, the coordinates of the destination node, or other information that can be used to determine the coordinates of the source node and the destination node.

[0163] S804. The current node determines the data routing path based on the node coordinates of the destination node.

[0164] Node coordinates are determined based on the network node's number in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. Furthermore, when the data routing topology 10 includes a fourth topology network, the node coordinates can be determined based on the network node's number in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group and the number of the second node group to which the network node belongs.

[0165] The current node can parse data packets to obtain the coordinates of the destination node. In some examples, the coordinates of the destination node can be {Gd,Ld,Pd,Xd,Yd,Zd}, where Gd represents the number of the second node group to which the destination node belongs, Ld represents the number of the first node group to which the destination node belongs, Pd represents the port number of the switch to which the destination node is connected, and Xd, Yd, and Zd represent the numbers of the destination node in each dimension of the first topology network.

[0166] The current node can determine the data routing path based on the coordinates of the destination node and the topology of the data routing system 10 (e.g., a topology graph). It should be noted that the coordinates of the current and destination nodes indicate that when the current and destination nodes are at the same level, the data routing path can be determined based on the topology of that level. For example, if the current and destination nodes are both in the same first-level topology network (or the same first-level node group) at the first level, the current node can determine the data routing path based on the topology of that level, such as the topology of the corresponding node group within that level. Alternatively, if the current and destination nodes are in different node groups at the same level, the current node can determine the data routing path based on the topology of that level.

[0167] Based on this, the current node can first identify whether the node groups (such as the first node group and the second node group) to which the current node and the destination node belong are the same, based on the node coordinates of the current node and the destination node, and then determine the data routing path according to the distribution of the node groups to which the current node and the destination node belong. In some cases, the current node can also determine the data routing path by combining the node coordinates of the source node. For example, the current node can determine the traffic type based on the node coordinates of the destination node and the node coordinates of the source node. The traffic type is used to identify whether the data packet is intra-group traffic or cross-group traffic. Intra-group traffic can be traffic within the same rack, while cross-group traffic can be external traffic across different racks. The current node can determine the corresponding data routing path based on the traffic type.

[0168] S806. The current node routes data packets according to the data routing path.

[0169] Specifically, the data routing path can indicate the routing port or the node coordinates of the next hop node. Accordingly, the current node can route the data packet to the next hop node based on the routing port indicated by the data routing path or the node coordinates of the next hop node.

[0170] Based on the above description, this application provides a data routing method. This method introduces the node coordinates of the destination node in the data routing system, adapts the data routing to the topological characteristics of the data routing system, provides the shortest distance communication between the source node and the destination node, achieves low communication latency, and is simple to implement with high availability.

[0171] Next, the specific implementation of determining the data routing path when the current node and the destination node are in node groups at different levels in the embodiment of Figure 8 will be explained.

[0172] Depending on the location of the source and destination nodes in the network, the source node can route data packets to the destination node either directly or by first routing to a next-hop network node and then continuing the routing from that node until the destination node is reached. Specifically, when the data is at the source node, the source node can determine the next-hop node using a shortest path routing algorithm. If that next-hop node is not the destination node, it can also determine its own next-hop node using the same algorithm until the destination node is reached.

[0173] For ease of description, this application refers to the node where data arrives and is waiting to be processed as the current node. The current node can be the source node or a node on the path from the source node to the destination node. Whether it is the source node or a node on the path from the source node to the destination node, the method of determining the data routing path (next-hop node) using the shortest path routing algorithm is similar. The specific implementation of determining the data routing path is described in detail below.

[0174] Considering that routing methods often differ significantly when the current node and destination node are distributed differently in the network, this application supports identifying the distribution of the current node and destination node, and then performing data routing based on this distribution.

[0175] Specifically, the current node can identify whether the current node and the destination node belong to the same node group based on their node coordinates and the destination node's node coordinates. For example, if the numbers of the first node group to which the current node and the destination node belong are the same, it means that the current node and the destination node are in the same first node group. Similarly, if the numbers of the second node group to which the current node and the destination node belong are the same, it means that the current node and the destination node are in the same second node group. Based on this, data routing can be divided into intra-group routing and inter-group routing between different node groups.

[0176] In the first implementation, when the ID of the current node in the first node group is equal to the ID of the destination node in the first node group, it indicates that the current node and the destination node are in the same first node group (L1 group), and the data routing is an intra-group route within the first node group. Since the current node and the destination node are in the same first node group, they are necessarily in the same second node group as well, and the ID of the current first node in the second node group is equal to the ID of the destination node in the second node group. In this case, the current node can determine the routing port in each dimension based on the offset values ​​of the IDs of the destination node and the current node in each dimension of the first topology network.

[0177] The number of network nodes surrounding the i-th dimension of the network topology can be L. i The current node determines the routing port in the i-th dimension based on the offset value between the destination node and the current node's number in the i-th dimension surrounding the network topology. This can include the following cases:

[0178] In the first case, the offset between the destination node and the current node in the i-th dimension surrounding the network topology is greater than 0 and less than or equal to L. i / 2 indicates that the destination node is in the positive direction, or the offset value is less than -L. i / 2 indicates that the loopback link is the shortest path. Based on this, the current node can determine the routing port in the i-th dimension as the port used to connect to downstream nodes, also known as the forward routing port.

[0179] In the second case, the offset between the destination node and the current node in the i-th dimension surrounding the network topology is greater than L. i / 2 indicates that the loopback link is the shortest path, or the offset value is greater than or equal to -L. i If the value is 2 and less than 0, it indicates that the destination node is in the negative direction. The routing port in the i-th dimension is determined as the port used to connect to the upstream node, also known as the negative routing port.

[0180] The current node can traverse each dimension and perform the above routing operation for each dimension until the destination node is reached.

[0181] It should be noted that the above example uses one primary node group to correspond to one rack, such as a single rack. When the current node (e.g., the source node) and the destination node are located in different racks, the current node can also route data to the destination node through a switch.

[0182] When the data packet is a packet within the all-reduce traffic, the current node can execute the intra-group routing of the first node group mentioned above. For example, in an AI computing scenario, the traffic for training an AI model in parallel using data parallel mode or model parallel mode can be all-reduce traffic. The current node can route the all-reduce traffic through the intra-group routing of the first node group.

[0183] In the second implementation, the current node determines the traffic type based on the coordinates of the destination node and the source node. The traffic type identifies whether the data packet is intra-group traffic or inter-group traffic. Intra-group traffic can be traffic within a rack, while inter-group traffic can be external traffic across racks. For example, if the first node group number of the source node and the first node group number of the destination node are the same, it indicates intra-group traffic, such as traffic within the local rack, which can communicate based on the aforementioned Torus topology. Otherwise, the traffic is inter-group traffic, such as external traffic across racks, which can be redirected through switches in the fat tree network topology. When the traffic type identifies the data packet as inter-group traffic, the current node can determine the forwarding port based on the node coordinates of the destination node.

[0184] Cross-group traffic can be further categorized into local-to-external traffic and external-to-local traffic. Local-to-external traffic can be internal-to-external cross-rack traffic, while external-to-local traffic is external-to-internal cross-rack traffic. The current node can perform fat-tree data forwarding based on the shortest routing path algorithm for the aforementioned cross-group traffic. Specifically, for external cross-rack traffic, it can be forwarded to a local node within the rack that has a direct link to the destination node; for external cross-rack traffic destined for the local node, the data routing path (also known as the forwarding path) can be determined based on the destination node's coordinates, allowing for hopping within the first node group.

[0185] When cross-group traffic is local-to-external traffic, such as cross-rack traffic from inside to outside, the current node can determine the coordinates of network nodes within the group that have a direct link to the destination node (e.g., jump nodes), or network nodes within the group that have a direct link to the jump node leading to the destination node (e.g., shuttle nodes), based on the destination node's coordinates. The current node can then determine the port number of the switch to which the network nodes within the group are connected. Specifically, the jump node's coordinates include its number within the dimensions surrounding the network topology, and the shuttle node's coordinates include its number within the dimensions surrounding the network topology. The current node can then determine the port number of the switch to which these network nodes (e.g., the switch's output port number) are connected, based on the jump node's number within the dimensions surrounding the network topology or the shuttle node's number within the dimensions surrounding the network topology.

[0186] If the destination node and the local node are in the same second node group, then the local node contains a jump node with a direct link to the destination node. The current node can use the jump node's number within the dimension surrounding the network topology to determine the switch's outgoing port number. If the destination node and the local node are in different second node groups, then the coordinates can be mapped to the node coordinates of an internal local node with a direct link to the jump node. This internal local node with a direct link to the jump node is called a ferry node.

[0187] After determining the node coordinates of the aforementioned jump node or ferry node, the port number can be calculated using the following formula: P=X*Ky*Kz+Y*Kz+Z+1 (3)

[0188] Where P represents the port number, Ky and Kz represent the lengths of different dimensions of the first topology network, and X, Y, and Z represent the numbers of the jump nodes with direct links to the destination node in the dimensions surrounding the network topology, and the numbers of the ferry nodes with direct links to the jump nodes in the dimensions surrounding the network topology.

[0189] When cross-group traffic is traffic sent from outside to the local machine, such as cross-rack traffic from outside to inside, for example, when the number of the first node group where the source node is located is not equal to the number of the first node group where the destination node is located, but the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, it indicates that the data has been transferred to the rack where the destination node is located. The current node can determine the forwarding port based on the node coordinates of the destination node and the cascading relationship between the destination node and the switch. Among them, the node coordinates of the destination node include the number of the destination node in the dimension surrounding the network topology. The current node can determine the output port number of the switch based on the number of the destination node in the dimension surrounding the network topology. The calculation of the output port number can refer to the above formula (3). The only difference from the above formula (3) is that X, Y, and Z represent the dimension number of the destination node, such as Xd, Yd, and Zd.

[0190] When data packets are part of all-to-all or point-to-point traffic (send / receive traffic), the current node can route them through a fat-tree network topology, for example, through a switch within the fat-tree network topology. Specifically, for send / receive traffic, if the current node is not a network node directly connected to the destination node's node group (first or second node group), it can first be redirected through a switch in the fat-tree network topology, and then routed through the full mesh network.

[0191] In the third implementation, when the number of the second node group to which the current node belongs is equal to the number of the second node group to which the destination node belongs, and the number of the first node group to which the current node belongs is not equal to the number of the first node group to which the destination node belongs, it means that the current node and the destination node are in different first node groups within the same second node group, and the data routing is intra-group routing within the second node group (inter-group routing within the first node group).

[0192] Accordingly, the current node can determine whether there is a direct link (such as L-LINK) between the current node and the destination node in the first node group. In order to distinguish it from the direct link (such as G-LINK) or other direct links in the second node group of the destination node, this application may refer to the direct link with the destination node in the first node group as the first direct link.

[0193] If yes, the current node can determine that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if not, the current node can determine that the data routing path includes forwarding the data packet to the second hop node. The second hop node and the first node group where the destination node is located have a second direct link. When the data packet arrives at the second hop node, it can reach the first node group where the destination node is located via the second direct link. When the data packet reaches the first node group where the destination node is located, intra-group routing within the first node group can be performed.

[0194] The fourth implementation method is that when the number of the second node group to which the current node belongs is not equal to the number of the second node group to which the destination node belongs, it means that the current node is in a different second node group, and the data routing is an intra-group routing of the third node group (inter-group routing of the second node group).

[0195] Accordingly, the current node can determine whether a third direct link exists between the current node and the destination node in the second node group. If so, the current node can determine that the data routing path includes forwarding the data packet from the third direct link to the destination node's second node group. When the data packet arrives at the destination node's second node group, intra-group routing can be performed. If not, the current node can determine that the data routing path includes forwarding the data packet to the first hop node. The first hop node and the destination node's second node group have a fourth direct link.

[0196] The coordinates of the first jump node can be determined by the following formula:

[0197] Where Xm, Ym, and Zm are the numbers of the first hop node in each dimension of the first topology network, PLm represents the port number of the second node group where the first hop node forwards data to the destination node, Ky and Kz represent the lengths of different dimensions of the first topology network, and Xd, Yd, and Zd are the numbers of the destination node in each dimension of the surrounding network topology.

[0198] Furthermore, the current node can also determine whether the number of the first node group to which the first hop node belongs is equal to the number of the first node group to which the current node belongs. If not, it indicates that the current node and the first hop node are in different first node groups. The current node's determination of the data routing path also includes forwarding data packets to the ferry node. The ferry node and the first hop node's first node group have a fifth direct link. That is, the current node can first route the data packet to the ferry node. Since the ferry node and the aforementioned first hop node are in the same first node group, the data packet can be routed to the aforementioned first hop node through intra-group routing within the first node group. The first hop node and the destination node's second node group have a direct link, such as the aforementioned fourth direct link. Data packets arriving at the first hop node can be routed to the destination node's second node group through the aforementioned fourth direct link. When the data packet reaches the destination node's first node group, intra-group routing within the second node group can be performed.

[0199] In the case of data packets within point-to-point traffic, the current node can perform intra-group routing (inter-group routing of the first node group) or intra-group routing (inter-group routing of the second node group) of the aforementioned second node group, for example, routing the aforementioned point-to-point traffic through a full mesh network (third topology network or fourth topology network). This point-to-point traffic can include send / receive traffic in pipelined parallel mode.

[0200] For ease of description, this application uses the topology example in Figure 7 as an example. In the topology of Figure 7, the node coordinates of the network node can be {G,L,P,X,Y,Z}, where G represents the number of the second node group to which the network node belongs, ranging from 0 to (Kx*Ky*Kz*(Kx*Ky*Kz+1)). When Kx, Ky, and Kz are 4, G can be 0 to 4160 (inclusive of endpoint values). L represents the number of the first node group to which the network node belongs, ranging from 0 to Kx*Ky*Kz. When Kx, Ky, and Kz are 4, L can be 0 to 64 (inclusive of endpoint values).

[0201] P is the port number of the network node (such as a GPU card with routing function) in the L2 topology that connects to the TOR switch in the rack (P = X*Ky*Kz+Y*Kz+Z+1); X, Y, and Z represent the numbers of the network node in the X, Y, and Z dimensions within the L1 Torus, respectively, ranging from 0 to Kx-1, 0 to Ky-1, and 0 to Kz-1. The node number corresponds one-to-one with the topology location of the node.

[0202] The routing engine's ports 1 and 2 serve as routing ports in the forward and reverse directions of the X dimension (port1: LinkX+; port2: LinkX-).

[0203] The routing engine's ports 3 and 4 serve as routing ports in the forward and reverse directions of the Y dimension (port3: LinkY+; port4: LinkY-).

[0204] The routing engine's ports 5 and 6 serve as routing ports in the forward and reverse directions of the Z dimension (port 5: LinkZ+; port 6: LinkZ-).

[0205] The positive direction interface of each adjacent node in each dimension is connected to the negative direction port of the other end. By connecting them in sequence, a Torus ring network is formed, and an L1-level topology 3D Torus can be constructed.

[0206] The routing engine's port 7 connects to the TOR switch inside the rack. As an L2 topology, the rack is interconnected via switches, and GPUs within the rack can reach each other in one hop without blocking.

[0207] Port 8 of the routing engine serves as the interconnection port between various L1-level Torus within the L3-level topology group, and is defined as an L-Link. The L-Link number of each routing engine corresponds one-to-one with the Torus coordinates. PL = X*Ky*Kz + Y*Kz + Z, and the value range of PL is 0 to Kx*Ky*Kz-1.

[0208] The following describes the global link connection relationships within the L3-level topology group.

[0209] The source node is numbered {Gs,Ls,Ps,Xs,Ys,Zs}, the destination node is numbered {Gd,Ld,Pd,Xd,Yd,Zd}, and the current node is numbered {Gc,Lc,Pc,Xc,Yc,Zc}. The source node's L-Link port 8 is connected to the destination node's L-Link port 8. Therefore, the full mesh connection rule between L3 topologies can be: Ls + PLs = Ld && PLd + PLs = Kx * Ky * Kz - 1.

[0210] Where Ls is the number of the first node group to which the source node belongs, Ld is the number of the first node group to which the destination node belongs, PLs is the L-Link port number connecting the source node to the first node group to which the destination node belongs, and PLd is the L-Link port number connecting the destination node to the first node group to which the source node belongs. Based on this connection relationship, the routing relationship between the first node groups can be calculated.

[0211] The global link connectivity within the L4-level topology group is shown below:

[0212] Port 9 of the routing engine serves as the interconnection port between various L3 full mesh topologies within the L4-level topology, defined as a G-Link. Each routing engine's G-Link number corresponds one-to-one with the coordinates of each L3 full mesh and the torus within the group. The G-Link numbering rule is: PG = L*Kx*Ky*Kz + X*Ky*Kz + Y*Kz + Z, where PG values ​​range from 0 to Kl*Kx*Ky*Kz-1. The full mesh connection relationship between L4-level topologies is: Gs + PGs = Gd && PGd + PGs = Kl*Kx*Ky*Kz-1.

[0213] Wherein, Gs is the L4 topology number of the source node, for example, the number of the second node group to which the source node is located; Gd is the L4 topology number of the destination node, for example, the number of the second node group to which the destination node is located; PGs is the G-Link port number connecting the source node to the second node group to which the destination node is located; and PGd is the G-Link port number connecting the destination node to the second node group to which the source node is located. Based on this connection relationship, the routing relationship between the second node groups can be calculated.

[0214] Referring to Figure 9, which illustrates a data routing process in a data routing system, the current node first determines whether the destination node and the current node are in the same second node group. If not, it performs inter-group routing within the second node group; if so, it continues to determine whether the destination node and the current node are in the same first node group. When the destination node and the current node are not in the same first node group, it performs inter-group routing within the first node group. When the destination node and the current node are in the same first node group, the current node can further determine the traffic type of the data packet based on the node coordinates of the destination node and the source node. When the traffic type is local rack-internal traffic, the current node can route the data packet using the L1-level torus topology; when the traffic type is cross-rack external traffic, the current node can route the data packet using the L2-level fat tree topology. Specifically, the traffic type is local rack-internal traffic when the destination node's second node group number is equal to the source node's second node group number, and the destination node's first node group number is equal to the source node's first node group number; otherwise, the traffic is cross-rack external traffic.

[0215] The following sections will explain the specific implementations of inter-group routing for the second node group, inter-group routing for the first node group, fat tree topology routing, and torus topology routing.

[0216] If the destination node and the current node are in different second node groups, i.e., Gd = ! Gc, the first hop node with a direct link to the second node group containing the destination node {Gd, Ld, Xd, Yd, Zd} is defined. The coordinates of the first hop node can be determined by the following formula: Gm = Gd - PGm && PGm = (Kx*Ky*Kz*(Kx*Ky*Kz+1)-1) - PGd, where,

[0217] If the current node has a direct link to the destination node's second node group (denoted as Gd), then the following formula is satisfied:

[0218] In this case, data can be directly forwarded from the current node's G-Link port to reach the destination node's second node group. Otherwise, it is necessary to first route from the current node to the first hop node. The coordinates of the first hop node are {Gc, Lm, Xm, Ym, Zm}, where:

[0219] If the first node group of the first hop node is different from the first node group of the current node (i.e., Lm = ! Lc), the data packet must first be routed to the first node group of the first hop node (denoted as Lm). A node with a direct link to the first node group Lm of the first hop node is defined as a ferry node {Gn, Ln, Xn, Yn, Zn}. Therefore, data packets can first be routed to the ferry node, and then forwarded from the direct link L-Link to the first node group of the first hop node. The physical location of the ferry node can be obtained using the following formula:

[0220] The current node can route data packets to a ferry node via an L1-level torus. The current node can route packets based on coordinate address offsets in the X, Y, and Z dimensions. Upon reaching the ferry node, the data packet can be routed from the direct L-Link port to the first node group containing the first hop node. After arriving at the first node group, the data packet can reach the first hop node via intra-group routing. The first hop node can then reach the second node group containing the destination node via a direct link. Finally, the data packet can reach the destination node via intra-group routing within the second node group.

[0221] The following section provides a detailed explanation of intra-group routing for the second node group (e.g., inter-group routing for the first node group).

[0222] The current node can determine whether it and the destination node are in the same second node group based on their coordinates. If the destination node and the current node are in the same second node group (Gd = Gc, but Ld = !Lc), the current node can determine whether it has a direct link connecting to the destination node in the first node group based on the topology.

[0223] If the current node is a directly connected link to the destination node's first node group, data can be forwarded directly from the directly connected link. See the following formula for details:

[0224] If formula (9) is satisfied, the destination node can be directly reached from the L-Link port by forwarding the link. Otherwise, a node in the first node group connected to the destination node by a direct link is defined as the second hop node, and the coordinates of the second hop node can be:

[0225] Therefore, the first step is to route to the second hop node: Xm = PLm / Ky*Kz; Ym = PLm / Kz; Zm = PLm%Kz, then redirect to the first node group's internal routing, specifically routing along the X, Y, and Z dimensions based on the coordinate address offset. Upon reaching the second hop node, the packet is forwarded from port PLm = (Kx*Ky*Kz-1)-PLd = (Kx*Ky*Kz-1)-Xd*Ky*Kz-Yd*Kz-Zd, thus connecting the direct link to the first node group where the destination node resides. Then, intra-group routing within the first node group can be executed, such as dimension-order routing within the same rack. Dimension-order routing means that each packet is routed on only one dimension at a time; once the appropriate coordinates are reached on that dimension, routing continues on other dimensions in ascending order.

[0226] Specifically, based on the source node number {Gs,Ls,Ps,Xs,Ys,Zs} and the destination node number {Gd,Ld,Pd,Xd,Yd,Zd}, it can be determined whether the data packet is internal traffic within the local cabinet or external traffic across cabinets: if Gs == Gd and Ls == Ld, it indicates that it is local traffic and will communicate through the Torus topology; otherwise, it indicates that it is external traffic across cabinets and needs to be hopped through the TOR switch.

[0227] For cross-rack traffic originating locally and destined for external locations, the local node number within the rack with which the destination node has a direct link can be calculated based on the destination node number (if the destination node and the local node are not in the same first node group, then it corresponds to the internal local node number with a direct link to the intermediate jump node). The port numbering rule from this node to the TOR switch is determined according to P = X*Ky*Kz + Y*Kz + Z + 1.

[0228] For the traffic sent from the external to the local, that is, Gs != Gd or Ls != Ld, but Gc == Gd and Lc == Ld, it indicates that the data has been transferred to the cabinet where the destination node is located. Then, the forwarding port can be determined according to the coordinates of each dimension of the Torus of the destination node: P = Xd * Ky * Kz + Yd * Kz + Zd + 1.

[0229] The intra-group routing of the first node group is described below.

[0230] The destination node and the current node are within the same L4 and L3 topologies, that is, Gd == Gc && Ld == Lc. Therefore, only the coordinate offsets of the three dimensions of the Torus where the destination node is located need to be used to determine the routing:

[0231] First, judge the coordinate offset in the X dimension. Compare the coordinate deviation of the destination node in the X dimension with that of the current node. Define the coordinate deviation DeltX == Xd - Xc. If 0 < DeltX <= Kx / 2, it means the destination node is in the positive direction, and route from the X+ port; if DeltX > Kx / 2, it means that taking the loopback link is the shortest path, and route from the X- port; if 0 > DeltX >= -Kx / 2, it means the destination node is in the negative direction, and route from the X- port; if DeltX < -Kx / 2, it means that taking the loopback link is the shortest path, and route from the X+ port.

[0232] Then, judge the coordinate offset in the Y dimension. Define the coordinate deviation DeltY == Yd - Yc. If 0 < DeltY <= Ky / 2 or DeltY < -Ky / 2, route from the Y+ positive port; if DeltY > Ky / 2 or 0 > DeltY >= -Ky / 2, route from the Y- port.

[0233] Next, judge the coordinate offset in the Z dimension. Define the coordinate deviation DeltZ == Zd - Zc. If 0 < DeltZ <= Kz / 2 or DeltZ < -Kz / 2, route from the Z+ positive port; if DeltZ > Kz / 2 or 0 > DeltZ >= -Kz / 2, route from the Z- port.

[0234] Finally, there is no offset in the Z dimension, reaching the destination node. The DMA communication engine receives the packet and completes the communication.

[0235] Considering that both Torus and full mesh topologies contain loops, deadlock risks can arise. Deadlock refers to the circular occupation of communication resources (such as buffers). For example, in a Torus topology with four nodes along dimension X, each node sends data to its next adjacent node (node ​​0 sends to node 2, node 1 sends to node 3, and so on). If all four nodes send data simultaneously, port resources will be circularly occupied, leading to deadlock, communication congestion, and in severe cases, even network paralysis. Therefore, the data routing system in this application can also prevent deadlock through virtual lanes (VLs).

[0236] Virtual channels can effectively avoid deadlock. In the data routing system of this application, since the first topology network (such as L1 level topology) adopts a standard Torus architecture and the third topology network adopts a full mesh architecture, loops naturally exist in both standard Torus and full mesh topologies. Therefore, multiple virtual channels can be allocated to the ports of the first topology network (such as the first node group) and the third topology network (such as the third node group) to avoid deadlock. Furthermore, when the data routing system also includes a fourth topology network, multiple virtual channels can also be allocated to the ports of the fourth topology network to avoid deadlock.

[0237] Specifically, the ports corresponding to the first topology network are configured with a first virtual channel and a second virtual channel, the ports corresponding to the third topology network are configured with a third virtual channel and a fourth virtual channel, and the ports corresponding to the fourth topology network are configured with a fifth virtual channel and a sixth virtual channel.

[0238] For global traffic in the first topology network, the second virtual channel can be used for routing. Specifically, the first virtual channel can be used for routing the initial global traffic in the first topology network. For global traffic in the third topology network, the fourth virtual channel can be used for routing. Specifically, the third virtual channel can be used for routing the initial global traffic in the third topology network. For global traffic in the fourth topology network, the sixth virtual channel can be used for routing. Specifically, the fifth virtual channel can be used for routing the initial global traffic in the fourth topology network. For ease of understanding, Figure 10 provides an example of a virtual channel allocation scheme. In this example, a port of a network node is configured with multiple virtual channels. This prevents loops between virtual channels, thus avoiding deadlocks between layers.

[0239] The above example uses a large-scale AI model with hundreds of billions of parameters. With the continuous increase in computing power, more and more developers are choosing to train AI models with trillions of parameters through ultra-large-scale clusters. In addition to model parallelism, data parallelism, and pipeline parallelism, trillion-parameter AI models further enhance computational efficiency through hybrid expert parallelism. Hybrid expert parallelism primarily involves global communication, with all-to-all traffic accounting for 50% of the total traffic.

[0240] The communication latency of a Torus topology increases rapidly with network size; the larger the network, the more severe the congestion, especially for all-to-all type globally uniform random traffic networks where throughput is very low. Therefore, the Torus topology is difficult to meet the communication requirements of all-to-all traffic. A fat tree network, on the other hand, is a non-blocking network, highly compatible with all-to-all traffic, and can provide 100% throughput.

[0241] The scale of expert parallelism and the size of supernodes affect the performance of AI applications. Under the same cluster size, the larger the number of supernodes, the better the performance of AI application. Taking GPT-4 as an example, the performance of a 1024-supernode scale is 15% better than that of a 64-supernode scale. However, for large-scale supernodes, due to hardware constraints, multiple racks need to be connected to form a large-scale supernode.

[0242] The following example illustrates the construction of a supernode using 1024 computing cards (e.g., GPUs). A standard rack typically accommodates 64 GPUs, allowing for an 8x8 2D Torus topology. A 3D Torus can be constructed between 16 racks, connecting 16*8*8 = 1024 GPUs. Global traffic across racks (e.g., global jump traffic) can be transmitted via a two-layer fat tree. The rack contains a 2D Torus with 8*8 = 64 GPUs, and the supernode contains a 3D Torus between 16 racks, also containing 16*8*8 = 1024 GPUs. A single-level full mesh network can reach a scale of 1 million GPUs: 1024 * 1025 = 10,496,001,972, as shown in Figure 11. It should be noted that "A" in Figure 11 represents GPUs, and "K" represents CPUs. Each set of 4 CPUs can connect to 8 GPUs. In Figure 11, SW represents a rack switch. The TOR switch located at the top of the rack can be a leaf switch, and SW01 to SW64, which aggregate the leaf switches, can be spine switches. The spine switch in this application can be a 128-port switch.

[0243] The following example illustrates a network configuration using a 12-port general-purpose GPU. A 12-port general-purpose GPU is a GPU with an internal I / O die that provides routing capabilities. This GPU integrates or has a built-in routing engine, thus providing routing forwarding capabilities. A 1-port general-purpose GPU can have a port bandwidth of 200Gbps.

[0244] In practical implementation, as shown in Figure 12, a Hybrid Torus topology can be constructed based on 12-port general-purpose GPUs. Specifically, bandwidth resources are allocated according to the communication type and features of the AI ​​large model. Six 200Gbps ports are used for Torus interconnection. Within a RACK, a 2D Torus consisting of 64 GPUs (arranged in an 8*8 configuration) is constructed using four ports (two ports per dimension). A 3D Torus consisting of 1024 GPUs (arranged in a 16*8*8 configuration) is constructed between the 16 RACKs.

[0245] Experts account for approximately 50% of the traffic in parallel. Therefore, five 200Gbps bandwidth connections can be allocated to the TOR switches within the rack. The racks are connected to form a two-layer fat tree via spine switches, which can provide all-to-all and all-reduce traffic communication between 1024 cards within the supernode.

[0246] Since communication between supernodes is mainly pipelined in parallel, and the pipelined traffic is low-volume send / receive traffic, one 200Gbps port can be allocated for full-mesh direct connection between supernodes, which can effectively reduce network costs and energy consumption while meeting communication requirements.

[0247] Table 2 also shows the bandwidth allocation results, as follows:

[0248] Table 2 Bandwidth Allocation Results

[0249] Each node within a RACK can have bandwidth comprising 1200G of Torus network bandwidth, 1000G of fat tree network bandwidth, and 1024*200G of full mesh network bandwidth. The Torus network transmits only all-reduce traffic, fully utilizing bandwidth resources across all dimensions. The fat tree network transmits all-to-all and all-reduce traffic within a supernode, as well as internal jumps of global traffic across supernodes. The full mesh network transmits send / receive traffic between RACKs (e.g., between supernodes).

[0250] It should be noted that in practical applications, the network size can also be adjusted according to traffic characteristics.

[0251] The following example illustrates data routing based on the aforementioned 12-port general-purpose GPU network.

[0252] Referring to Figure 13, which illustrates a data routing process in a data routing system, the current node first determines whether the destination node and the current node are in the same first node group. If not, inter-group routing within the first node group is executed. If so, the traffic type is identified. If the traffic type is local rack-based traffic, intra-group routing within the first node group is executed, such as a sequential routing within the same rack. If the traffic type is cross-rack external traffic, routing can be performed using a fat tree network topology. Furthermore, if a 12-port general-purpose GPU nested Torus, fat tree, and two-level full mesh are used, it can first determine whether the destination node and the current node are in the same second node group. If they are in different second node groups, inter-group routing within the second node group is executed.

[0253] The following sections will explain the specific implementations of inter-group routing for the second node group, inter-group routing for the first node group, fat tree topology routing, and intra-group routing for the first node group.

[0254] The current node first determines whether the destination node and the current node are in the same second node group. If not, i.e., Gd = ! Gc, then inter-group routing within the second node group can be performed. Specifically, if the current node has a direct link to the destination node's second node group Gd, i.e., satisfying PGc = Lc*Kx*Ky*Kz + Xc*Ky*Kz + Yc*Kz + Zc = (Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)-Ld*Kx*Ky*Kz-Xd*Ky*Kz-Yd*Kz-Zd, then data can be directly forwarded from the current node's G-Link port to reach the destination node's second node group. Otherwise, routing from the current node to the first hop node is required. The first jump node is located at the position number {Gc,Lm,Xm,Ym,Zm}, where PGm=(Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)-Kx*Ky*Kz*Ld-Xd*Ky*Kz-Yd*Kz-Zd, Lm=PGm / Kx*Ky*Kz, Xm=PGm / Ky*Kz, Ym=PGm / Kz, Zm=PGm%Kz.

[0255] If the first node group of the first hop node is different from the first node group of the current node (i.e., Lm = ! Lc), the data packet must first be routed to the first node group Lm of the first hop node. A node with a direct link to the first node group Lm of the first hop node is defined as a ferry node, {Gn, Ln, Xn, Yn, Zn}. Therefore, the data packet must first be routed to the ferry node, and then forwarded from the direct link L-Link to the first node group of the first hop node. Then, it can be routed within the L2 level rack, using a fat-tree network topology to route the data packet to the first hop node. The physical location of the ferry node can be obtained using the following formulas: Gn = Gc, Ln = Lc, PLn = (Kx*Ky*Kz-1) - Xm*Ky*Kz - Ym*Kz - Zm. The coordinates within the Torus of the ferry node are: Xn = PLn / Ky*Kz; Yn = PLn / Kz; Zn = PLn%Kz.

[0256] If the destination node and the current node are in the same second node group (Gd = Gc), the current node can further determine if they are in the same first node group. If not, (Ld = !Lc), then inter-group routing within the first node group is executed. If the current node has a direct link to the destination node's first node group, data can be forwarded directly from the direct link. Specifically, when Lc = Ld - PLc && PLc = (Kx*Ky*Kz - 1) - PLd, or Lc = Ld - (Kx*Ky*Kz - 1) + Xd*Ky*Kz + Yd*Kz + Zd = Xc*Ky*Kz + Yc*Kz + Zc, then data can be directly forwarded from the L-Link port to reach the destination node's first node group. Otherwise, the node in the first node group with a direct link to the destination is defined as the second hop node, and its node coordinates are: Gm = Gc, Lm = Ld - PLm && PLm = (Kx*Ky*Kz - 1) - PLd, PLd = Xd*Ky*Kz + Yd*Kz + Zd. That is, Gm = Gc, Lm = Ld - (Kx*Ky*Kz - 1) + Xd*Ky*Kz + Yd*Kz + Zd && PLm = (Kx*Ky*Kz - 1) - Xd*Ky*Kz - Yd*Kz - Zd. Therefore, the route first goes to the second hop node: Xm = PLm / Ky*Kz; Ym = PLm / Kz; Zm = PLm%Kz, then redirects to the internal routing of the first node group, routing according to the coordinate address offset in the X, Y, and Z dimensions respectively. After reaching the second hop node, the device forwards from port PLm = (Kx*Ky*Kz-1)-PLd = (Kx*Ky*Kz-1)-Xd*Ky*Kz-Yd*Kz-Zd, thus reaching the first node group where the destination node is located via the direct link.

[0257] When the destination node and the current node are in the same first node group, the source node number {Gs, Ls, Ps, Xs, Ys, Zs} and the destination node number {Gd, Ld, Pd, Xd, Yd, Zd} can be used to determine whether the data is internal supernode traffic or external traffic across supernodes. If Gs == Gd and Ls == Ld, it indicates internal supernode traffic. For all-to-all traffic types (the MPI communication library can identify the traffic type based on the application's communication primitives), internal expert parallel traffic can be transmitted through the FatTree topology; for all-reduce traffic types, it can be transmitted through the Torus topology, and some all-reduce traffic can share the bandwidth resources of the FatTree topology (without conflicting with all-to-all traffic); otherwise, it indicates external traffic across supernodes, which can be accessed through a TOR switch to jump to the FatTree topology.

[0258] For traffic within a supernode, such as Gs = Gd and Ls = Ld (the destination node and source node are in the same first node group), or Gc == Gd and Lc == Ld (indicating that the data has flowed to the supernode where the destination node is located), the uplink path can be determined by hash routing based on the node coordinates of the destination node. The spine switches in the Layer 2 FatTree topology then route the data to the leaf switches connected to the destination node based on the topology connection relationship through hash routing. Assuming X represents the rack number (0~15), then whether it is a cross-rack issue is determined based on the X dimension number. If Xd! =Xc indicates that cross-cabinet communication is required. In this case, the traffic needs to enter the Layer 2 fat tree network, and through hash routing, the traffic is forwarded to the corresponding spine switch. The spine switch then selects the output port based on the cabinet number (X dimension number) where the destination node is located and forwards it to the leaf switch in the cabinet where the destination node is located. If Xd == Xc, it means that the traffic has reached the cabinet where the destination node is located. The leaf switch forwards the traffic within the cabinet through the TOR switch according to the encoding rules. The data packet is forwarded from port: P = Y*Kz + Z + 1 to the corresponding node and then transferred to the intra-group routing of the first node group.

[0259] For the cross-supernode traffic sent locally to the external, based on the node coordinates of the destination node, the node coordinates of the local nodes inside the cabinet with direct links to it can be calculated (if the destination node and the local node are not in the same first node group, i.e., Ls!= Ld, then it corresponds to the node coordinates of the local nodes inside that have direct links to the second hop node). For the two-layer FatTree topology, based on the target node number, the upstream path can be determined through hash routing. Then, according to the topology connection relationship, the spine switch routes the data to the leaf switch connected to the target node through hash routing: Assuming X represents the cabinet number (0 - 15), it is judged whether to cross cabinets according to the X dimension number. If Xd!= Xc, it means cross-cabinet communication is required, and the traffic needs to enter the two-layer fat tree network. Through hash routing, the traffic is forwarded to the corresponding spine switch. Then, the spine switch selects the output port according to the cabinet number (X dimension number) where the target node is located and forwards it to the leaf switch in the cabinet where the destination node is located. If Xd == Xc, it means reaching the cabinet where the destination node is located. Then, according to the coding rule, the leaf switch forwards inside the cabinet where the destination node is located through the TOR switch, and forwards the data packet from the port: P = Y * Kz + Z + 1 to the corresponding node. This node forwards the data packet from the port with a direct link to the destination node according to the node coordinates of the destination node and transfers it to the intra-group routing within the first node group.

[0260] The intra-group routing within the first node group is described in detail below.

[0261] If the destination node and the current node are in the same L4 and L3 topology, i.e., Gd == Gc && Ld == Lc, the current node only needs to judge the routing based on the coordinate offsets of the three dimensions of the Torus where the destination node is located:

[0262] First, judge the coordinate offset in the X dimension. Compare the coordinate deviation of the destination node in the X dimension with the current node's coordinates. Define the coordinate deviation DeltX == Xd - Xc. If 0 < DeltX <= Kx / 2, it means the destination node is in the positive direction, and route from the X+ port. If DeltX > Kx / 2, it means taking the loopback link is the shortest path, and route from the X- port. If 0 > DeltX >= -Kx / 2, it means the destination node is in the negative direction, and route from the X- port. If DeltX < -Kx / 2, it means taking the loopback link is the shortest path, and route from the X+ port.

[0263] Then, judge the coordinate offset in the Y dimension. Define the coordinate deviation DeltY == Yd - Yc. If 0 < DeltY <= Ky / 2 or DeltY < -Ky / 2, route from the Y+ positive port. If DeltY > Ky / 2 or 0 > DeltY >= -Ky / 2, route from the Y- port.

[0264] Next, determine the Z - dimension coordinate offset. Define the coordinate deviation DeltZ == Zd - Zc. If 0 < DeltZ <= Kz / 2 or DeltZ < -Kz / 2, route from the Z+ positive - direction port; if DeltZ > Kz / 2 or 0 > DeltZ >= -Kz / 2, route from the Z - port.

[0265] Finally, there is no offset in the Z dimension, reaching the destination node. The DMA communication engine receives the message and completes the communication.

[0266] The operation time of the AI large - model job is long. Usually, it takes several months to complete the training. For large - scale clusters, tens of thousands or hundreds of thousands of computing cards need to communicate concurrently. It is inevitable that equipment failures will occur in hundreds of thousands of system components, resulting in training interruptions. Currently, the average time between failures of the mainstream training clusters in the industry is only about ten hours, and failures occur almost every day, which will greatly affect the training efficiency. Therefore, if the data routing system has backup resources and can quickly replace the faulty nodes, the system availability will be greatly improved. The fusion topology architecture of this application naturally has the characteristics of resource backup. Specifically, taking a cabinet with 64 computing cards as an example, each L3 group can include 65 computer cabinets, and 1 cabinet can be planned as a backup cabinet to back up the other 64 cabinets. Moreover, the algorithm allocation of AI applications is usually allocated in powers of 2. 16 / 64 is a more appropriate allocation. Resource allocation based on the above method will not cause fragmentation.

[0267] In some possible implementation manners, the failure may include different types or granularities of failures. The following gives an example of the rule recovery for different types of failures.

[0268] When the failure is a node failure, as shown in Figure 14, there is a faulty node in a cabinet with 64 cards. The job scheduling system (such as a job scheduler) can allocate a node in the backup RACK that has a direct - connection link with the RACK where the faulty node is located as a replacement node to continue participating in the operation. The data can return to the faulty RACK through the full mesh network, with only 3 additional hop - delays, and the faulty node can be quickly replaced, effectively improving the system availability.

[0269] When the failure is a multi - node failure, the job scheduling system can also use the TOR in the RACK to connect multiple backup nodes in the backup rack for replacement. For example, when there are 2 faulty nodes, the direct - connection links in 6 directions of each of the 2 faulty nodes can be connected, and the backup nodes in the backup RACK connected by the 12 direct - connection links can be used as replacement nodes to continue participating in the operation.

[0270] When the failure is a cabinet failure, the job scheduling system can use the backup RACK to replace the faulty RACK, thus quickly resuming operation.

[0271] The rack-level resource backup solution proposed in this application can support rapid replacement of single-node, multi-node, or even whole-rack failures, which can effectively improve system availability and accelerate the training efficiency of large AI models.

[0272] Based on the aforementioned data routing system 10 and data routing method, this application also provides a method for constructing a data routing system. This method can be executed by a management device. The method for constructing a data routing system according to this application will be described below from the perspective of a management device.

[0273] Referring to Figure 15, a flowchart of a method for constructing a data routing system is shown, which includes the following steps:

[0274] S1502, Management equipment identifies the switches and multiple network nodes used to build the data routing system.

[0275] Switches can include rack-mounted switches, specifically switches located within the rack where network nodes reside. For ease of cabling, rack-mounted switches can be positioned on the top of the rack; that is, rack-mounted switches can be TOR switches. In some examples, rack-mounted switches can include inter-rack switches. Rack-mounted switches can connect different network nodes within different racks. Specifically, inter-rack switches can connect different racks, for example, rack-mounted switches connecting different racks.

[0276] Multiple network nodes are divided into a second number of first node groups, which in turn include a first number of network nodes. After multiple network nodes establish a physical connection, such as through a physical link, the management device can identify the multiple network nodes used to build the data system through node discovery. The management device can discover new network nodes through mechanisms such as heartbeats. Further details are omitted here.

[0277] Furthermore, multiple network nodes can be divided into a third number of second node groups. These second node groups may include a second number of first node groups. The third number of second node groups constitutes a third node group.

[0278] S1504. The management device sends configuration files to the switch and multiple network nodes, so that the network nodes in the first node group can establish connections with upstream and downstream nodes in each dimension according to the network structure of the data routing system in the configuration file, to form a first topology network with a surrounding network topology, and cascade with the switch to form a second topology network with a fat tree network topology, and establish connections between a second number of first node groups to form a third topology network with a fully interconnected network topology.

[0279] In this network topology, network nodes in the first node group can establish connections with upstream and downstream nodes at various dimensions according to the network structure of the data routing system in the configuration file. This connection can be established through network connections, such as a three-way handshake, to facilitate data routing. Similarly, network nodes in the first node group can establish connections between a second number of first node groups according to the network structure of the data routing system in the configuration file. It should be noted that connections between node groups can include connections between network nodes in one node group and network nodes in another node group. In the third topology network, the second node groups can include at least one connection. For example, the connection between second node group A and second node group B can include a connection between node 1 in second node group A and node 2 in second node group B. Further, the connection between second node group A and second node group B can also include a connection between node 11 in second node group A and node 12 in second node group B. Network nodes in the first node group can also establish connections with switches according to the network structure of the data routing system in the configuration file. It should be noted that the connection initiator can be a network node or a switch; this embodiment does not impose any restrictions on this.

[0280] The data routing system includes a first topology network, a second topology network, and a third topology network, with the first topology network nested within the third topology network. Furthermore, when multiple network nodes are divided into a third number of second node groups, the configuration file is also used to establish connections between network nodes in the first node groups and these third number of second node groups, forming a fourth topology network with a fully interconnected network topology. Correspondingly, the data routing system also includes a fourth topology network, with the third topology network nested within it.

[0281] In some possible implementations, the management device can also check the connectivity of multiple network nodes based on the network structure of the data routing system in the configuration file, such as a network structure represented by a topology diagram. Here, the network structure in the configuration file represents the desired network structure, and the connectivity relationships represent the desired connections. The management device can also obtain the actual connectivity relationships of multiple network nodes through probing. The management device can then compare the desired connectivity relationships with the actual connectivity relationships to check the connectivity of multiple network nodes.

[0282] If the connection relationships (e.g., actual connections) of multiple network nodes are inconsistent with the connection relationships in the network structure of the data routing system, the management device can also send a prompt message to the user. This prompt message can be used to suggest adjusting the connection relationships of multiple network nodes. In this way, the accuracy and reliability of the network topology can be guaranteed.

[0283] The management device is also used to distribute routing tables, enabling the data routing system to perform data routing based on these tables. Specifically, the management device can generate routing tables based on the node coordinates of multiple network nodes. The node coordinates are determined based on the network node's number in each dimension of the first topology network and the number of the first node group to which the node belongs. The management device can then distribute entries in the routing table related to at least one of the multiple network nodes to at least one of the network nodes.

[0284] In this embodiment, the management device sends routing table entries related to at least one network node to that network node, which can reduce the storage resources occupied by routing table entries in the network node and save storage space. It should be noted that if the network node has sufficient storage resources, the management device can also send the complete routing table to at least one network node; this embodiment does not impose any restrictions on this.

[0285] Based on the above description, it can be understood that the method for constructing a data routing system in this application determines the switches and multiple network nodes used to construct the data routing system, and instructs the network nodes and switches to establish connections according to the network structure of the data routing system in the configuration file. This constructs a data routing system with nested surround network topology, fat tree network topology, and fully interconnected network topology, thereby enabling the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. Specifically, the surround network topology accelerates the application of localized communication characteristics and supports multi-tenant isolation, accelerating resource pooling; the fat tree network topology can be used for global traffic forwarding and can also serve as a replacement route for fault recovery; for ultra-large-scale supernodes, a Layer 2 fat tree topology can be used for non-blocking expansion to achieve global traffic forwarding; and the fully interconnected network topology provides direct routes for global communication, alleviating global network congestion and reducing expensive global fiber optic links compared to the fat tree network topology architecture.

[0286] As can be seen from the above embodiments, the hierarchical network that integrates Torus, fat tree and full mesh topologies proposed in this application can effectively reduce the network radius and provide a low-cost, low-power, and highly scalable network.

[0287] The processor integrates at least a 9-port routing engine, enabling networking with up to 1700W cards. Specifically, six ports in the 9-port routing engine or 9-port router can build a 4x4x4 3D Torus within the RACK, port 7 can connect to a 64-port TOR switch, and ports 8 / 9 can build a two-tier full mesh. This allows for ultra-large-scale networking with up to 1700W cards and enables traffic stratification, effectively improving the global communication performance of large AI models.

[0288] The table below compares different network topologies in terms of cost, power consumption, and performance. For AI models with trillions of parameters, the network scale is typically 128K GPUs and the supernode scale is 1024 GPUs. The mainstream network architecture in the industry uses a fat-tree topology to construct supernodes, and supernodes are also interconnected using a fat-tree topology. This requires a large number of switches and optical modules, resulting in very high network costs. Using the converged topology proposed in this application, the number of switches is reduced by 80% compared to the fat-tree architecture, and the number of optical modules is reduced by 61%. Therefore, network costs are reduced by 64%, and network energy consumption is reduced by 70%.

[0289] Table 3. Cost Comparison of Topology Networks

[0290] The optical module failure rate is 4%, and using a fat tree topology requires 4,480,000 optical modules, which cannot guarantee the normal operation of large-scale modeling tasks. Adopting this application can reduce the number of optical modules by 61%, effectively improving network system reliability.

[0291] Moreover, this method can achieve RACK-level resource backup, quickly replace faulty nodes, and effectively improve system availability.

[0292] Based on the aforementioned method for constructing a data routing system, this application also provides a management device. The management device will be described below from the perspective of functional modularity.

[0293] Referring to Figure 16, a schematic diagram of a management device is shown. As shown in Figure 16, the management device 1600 includes:

[0294] The determination module 1602 is used to determine the switches and multiple network nodes used to construct the data routing system, the multiple network nodes being divided into a second number of first node groups, the first node groups including the first number of network nodes;

[0295] The sending module 1604 is used to send configuration files to the switch and multiple network nodes, so that the network nodes in the first node group can establish connections with upstream nodes and downstream nodes in various dimensions according to the network structure of the data routing system in the configuration file, to form a first topology network with a surrounding network topology, and a second topology network cascaded with the switch to form a fat tree network topology, and establish connections between a second number of first node groups to form a third topology network with a fully interconnected network topology. The data routing system includes the first topology network, the second topology network, and the third topology network, with the first topology network nested within the third topology network.

[0296] The aforementioned determining module 1602 and sending module 1604 can be implemented in hardware or in software.

[0297] When implemented in software, the determining module 1602 and the sending module 1604 can be applications running on the management device 1600, such as a computing engine. Taking the determining module 1602 as an example, it can be virtualized through a virtualization service and then provided to users. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, and container services. Specifically, a VM service can be a service that uses virtualization technology to create a pool of virtual machine (VM) resources on multiple physical hosts, providing VMs to users on demand. A BMS service is a service that uses virtualization technology to create a pool of BMS resources on multiple physical hosts, providing BMS services to users on demand. A container service is a service that uses virtualization technology to create a pool of container resources on multiple physical hosts, providing containers to users on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines, and features secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services; specific limitations are not discussed here.

[0298] When implemented in hardware, the determining module 1602 may include at least one computing device, such as a server. Alternatively, the determining module 1602 may also be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof. The transmitting module 1604 may be implemented using a transceiver or transceiver unit. The transceiver may include, but is not limited to, an RF transceiver or a fiber optic transceiver.

[0299] In some possible implementations, multiple network nodes are divided into a third number of second node groups, which include the second number of first node groups. The configuration file is also used for network nodes in the first node groups to establish connections between the third number of second node groups, forming a fourth topology network with a fully interconnected network topology. The data routing system also includes the fourth topology network, with the third topology network nested within it.

[0300] In some possible implementations, the management device 1600 may also include a checking module 1606 and a prompting module 1608. The checking module 1606 is used to check the connection relationship of multiple network nodes according to the network structure of the data routing system in the configuration file; the prompting module 1608 is used to send a prompt message to the user if the connection relationship of multiple network nodes is inconsistent with the connection relationship in the network structure of the data routing system, and the prompt message is used to prompt adjustment of the connection relationship of multiple network nodes.

[0301] The inspection module 1606 and the prompting module 1608 can be implemented in software or in hardware.

[0302] When implemented in software, the inspection module 1606 and the prompting module 1608 can be applications running on the management device 1600. This application can also be virtualized and made available to the user via a virtualization service; for example, the inspection module 1606 can be virtualized and made available to the user via a VM service, BMS service, or container service.

[0303] When implemented in hardware, the inspection module 1606 and the prompting module 1608 may include at least one computing device, such as a server. Alternatively, the inspection module 1606 and the prompting module 1608 may also be devices implemented using an ASIC or a programmable logic device (PLD).

[0304] In some possible implementations, the management device 1600 may also include a generation module 1609;

[0305] The generation module 1609 is used to generate a routing table based on the node coordinates of multiple network nodes. The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs.

[0306] The sending module 1604 is further configured to send the entry in the routing table related to the at least one network node to at least one of the plurality of network nodes.

[0307] Similar to the aforementioned inspection module 1606 and prompting module 1608, the generation module 1609 can be implemented in software or in hardware.

[0308] When implemented in software, generation module 1609 can be an application running on management device 1600, such as a VM service, BMS service, or container service running on management device 1600. When implemented in hardware, generation module 1609 can include at least one computing device, such as a server. Alternatively, generation module 1609 can also be a device implemented using an ASIC or a programmable logic device (PLD).

[0309] The management device 1600 of this application will now be described from the perspective of hardware implementation. The management device 1600 can be a computing device, such as a server, or a terminal device. The terminal device can be a desktop computer, a laptop computer, a thin client, or a smartphone. The structure of the computing device of this application will now be described with reference to the accompanying drawings.

[0310] This application also provides a computing device 1700. As shown in FIG17, the computing device 1700 includes: a bus 1702, a processor 1704, a memory 1706, and a communication interface 1708. The processor 1704, the memory 1706, and the communication interface 1708 communicate with each other via the bus 1702. The computing device 1700 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1700.

[0311] Bus 1702 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 17, but this does not imply that there is only one bus or one type of bus. Bus 1702 can include pathways for transmitting information between various components of computing device 1700 (e.g., memory 1706, processor 1704, communication interface 1708).

[0312] Processor 1704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0313] Memory 1706 may include volatile memory, such as random access memory (RAM). Memory 1706 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD). Memory 1706 stores executable program code, which processor 1704 executes to implement the aforementioned method for constructing a data routing system. Specifically, memory 1706 stores instructions for management device 1600 to execute the method for constructing a data routing system.

[0314] The communication interface 1708 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 1700 and other devices or communication networks.

[0315] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0316] As shown in Figure 18, the computing device cluster includes at least one computing device 1700. The memory 1706 of one or more computing devices 1700 in the computing device cluster may store instructions from the same management device 1600 for executing methods to construct a data routing system.

[0317] In some possible implementations, one or more computing devices 1700 in the computing device cluster can also be used to execute some of the instructions used by the management device 1600 to execute the method for constructing the data routing system. In other words, a combination of one or more computing devices 1700 can jointly execute the instructions used by the management device 1600 to execute the method for constructing the data routing system.

[0318] It should be noted that the memory 1706 in different computing devices 1700 in the computing device cluster can store different instructions for executing some functions of the management device 1600.

[0319] Figure 19 illustrates one possible implementation. As shown in Figure 19, two computing devices 1700A and 1700B are connected via a communication interface 1708. The memory in computing device 1700A stores instructions for executing the functions of the determination module 1602. The memory in computing device 1700B stores instructions for executing the functions of the sending module 1604. Furthermore, the memory in computing device 1700A also stores instructions for executing the functions of the checking module 1606 and the prompting module 1608, and the memory in computing device 1700B also stores instructions for executing the functions of the generation module 1609. In other words, the memory 1706 of computing devices 1700A and 1700B jointly stores instructions used by the management device 1600 to execute the method for constructing a data routing system.

[0320] The connection method between the computing device clusters shown in Figure 19 can be considered because the method for constructing the data routing system provided in this application requires a large amount of resources to send configuration files and entries in the routing table to a large number of network nodes. Therefore, it is considered that the functions implemented by the sending module 1604 and the generating module 1609 are performed by the computing device 1700B.

[0321] It should be understood that the functions of computing device 1700A shown in Figure 19 can also be performed by multiple computing devices 1700. Similarly, the functions of computing device 1700B can also be performed by multiple computing devices 1700.

[0322] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 20 illustrates one possible implementation. As shown in Figure 20, two computing devices 1700C and 1700D are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1706 in computing device 1700C stores instructions for executing the function of the determination module 1602. Simultaneously, the memory 1706 in computing device 1700D stores instructions for executing the function of the sending module 1604. The memory in computing device 1700C also stores instructions for executing the functions of the check module 1606 and the prompting module 1608, and the memory in computing device 1700D also stores instructions for executing the function of the generation module 1609.

[0323] It should be understood that the functions of computing device 1700C shown in Figure 20 can also be performed by multiple computing devices 1700. Similarly, the functions of computing device 1700D can also be performed by multiple computing devices 1700.

[0324] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a management device to perform a method for constructing a data routing system.

[0325] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on a management device, it causes the management device to perform the method described above for constructing a data routing system.

[0326] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data routing system, characterized in that, The data routing system is used to implement data routing between computing nodes, and the data routing system includes a first topology network, a second topology network, and a third topology network. The network structure of the first topology network is a surrounding network topology structure composed of a first node group, wherein the first node group includes a first number of network nodes, and the network nodes in the surrounding network topology structure are connected to upstream nodes and downstream nodes in each dimension respectively. The second topology network is a fat tree network topology structure composed of network nodes and switches in the first topology network, and there is a cascading relationship between the network nodes and switches in the first topology network. The first topology network is nested within the third topology network. The network structure of the third topology network is a fully interconnected network topology structure composed of a second node group, wherein the second node group includes a second number of first node groups, and there are connections between the second number of first node groups.

2. The system according to claim 1, characterized in that, In the second topology network, network nodes belonging to the same first node group are directly connected to the same switch.

3. The system according to claim 1, characterized in that, The first node group includes multiple subgroups, and the second topology network includes multiple first switches and at least one second switch corresponding to the multiple subgroups. In the second topology network, network nodes belonging to the same subgroup are directly connected to the first switch corresponding to the subgroup, and the multiple first switches are connected upwards to the at least one second switch.

4. The system according to claim 3, characterized in that, The subgroups are divided according to cabinets, cabinet groups, or computer rooms.

5. The system according to any one of claims 1 to 4, characterized in that, Each network node includes n ports, of which m ports are used to connect to network nodes in the surrounding network topology. The m ports include ports for connecting upstream nodes in each dimension and ports for connecting downstream nodes in each dimension.

6. The system according to claim 5, characterized in that, Of the n ports, h ports are used to connect network nodes in the fully interconnected network topology formed by the second node group, where the second number is less than or equal to The Len j The number of network nodes in the j-th dimension of the surrounding network topology is represented by q, and q represents the number of dimensions of the surrounding network topology.

7. The system according to any one of claims 1 to 6, characterized in that, The data routing system also includes a fourth topology network, and the third topology network is nested within the fourth topology network; The fourth topology network has a fully interconnected network topology consisting of a third group of nodes, wherein the third group of nodes includes a third number of second groups of nodes, and there are connections between the third number of second groups of nodes.

8. The system according to claim 7, characterized in that, K of the n ports are used to connect network nodes in the fully interconnected network topology formed by the third node group, and the third number is less than or equal to k·the number of nodes in the second node group + 1.

9. The system according to any one of claims 1 to 8, characterized in that, The network node is either a switching node independent of the computing node, or a routing engine integrated into the computing node.

10. The system according to any one of claims 1 to 9, characterized in that, The node coordinates of the network node are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The network node is used to determine the data routing path based on the node coordinates, and to realize data routing between the computing nodes according to the data routing path.

11. The system according to any one of claims 1 to 10, characterized in that, The computing nodes are used to execute artificial intelligence (AI) computing tasks in parallel, and the parallelism of the AI ​​computing tasks includes at least one of model parallelism, pipeline parallelism, hybrid expert parallelism, or data parallelism.

12. The system according to any one of claims 1 to 11, characterized in that, The first topology network is used to route data packets in global all-reduce traffic, the second topology network is used to route data packets in many-to-many all-to-all traffic or point-to-point traffic, and the third topology network is used to route data packets in point-to-point traffic.

13. A data routing method, characterized in that, An application is made to a data routing system for implementing data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network. The first topology network has a network structure of a surrounding network topology composed of a first group of nodes, each group comprising a first number of network nodes. These network nodes are connected to upstream and downstream nodes in various dimensions. The second topology network has a network structure of a fat-tree network topology composed of the network nodes and switches in the first topology network. These network nodes and switches in the first topology network are cascaded. The first topology network is nested within the third topology network. The third topology network has a network structure of a fully interconnected network topology composed of second group of nodes, each group comprising a second number of first group of nodes, which are connected to each other. The method includes: The current node in the data routing system receives the data packet to be routed, and the current node is either the source node of the data packet or a node on the path from the source node to the destination node. The current node determines the data routing path based on the node coordinates of the destination node. The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The current node routes the data packet according to the data routing path.

14. The method according to claim 13, characterized in that, The current node determines the data routing path based on the node coordinates of the destination node, including: When the number of the first node group to which the current node belongs is equal to the number of the first node group to which the destination node belongs, the current node determines the routing port in each dimension based on the offset values ​​of the numbers of the destination node and the current node in each dimension of the first topology network.

15. The method according to claim 14, characterized in that, The number of network nodes in the i-th dimension of the first topology network is L i The current node determines the routing port in each dimension based on the offset values ​​of the destination node and the current node's number in each dimension of the first topology network, including: When the offset value of the destination node and the current node in the i-th dimension of the first topology network is greater than 0 and less than or equal to L i / 2, or less than -L i / 2, determine the routing port in the i-th dimension as the port used to connect to the downstream node; When the offset value of the destination node and the current node in the i-th dimension of the first topology network is greater than L i / 2, or greater than or equal to -L i If the value is 2 and less than 0, the routing port in the i-th dimension is determined to be the port used to connect to the upstream node.

16. The method according to claim 14 or 15, characterized in that, The data packets are those in the global reduce (all-reduce) traffic.

17. The method according to claim 13, characterized in that, The current node determines the data routing path based on the node coordinates of the destination node, including: The current node determines the traffic type based on the node coordinates of the destination node and the node coordinates of the source node. The traffic type is used to identify whether the data packet is intra-group traffic or cross-group traffic. When the traffic type identifies the data packet as cross-group traffic, the current node determines the forwarding port based on the node coordinates of the destination node.

18. The method according to claim 17, characterized in that, The current node determines the forwarding port based on the node coordinates of the destination node, including: When the cross-group traffic is traffic sent from the local to the outside, the current node determines the node coordinates of the network nodes in the group that have a direct link to the destination node based on the node coordinates of the destination node, or determines the node coordinates of the network nodes in the group that have a direct link to the jump node to the destination node, and determines the port number of the network nodes in the group that are connected to the switch based on the node coordinates of the network nodes in the group. When the cross-group traffic is traffic sent from outside to the local area, the current node determines the forwarding port based on the node coordinates of the destination node and the cascading relationship between the destination node and the switch.

19. The method according to claim 18, characterized in that, The data packets are data packets in all-to-all traffic or point-to-point traffic.

20. The method according to claim 13, characterized in that, The data routing system further includes a fourth topology network, in which the third topology network is nested. The fourth topology network has a fully interconnected network topology structure consisting of a third node group. The third node group includes a third number of second node groups, and there are connections between the third number of second node groups. The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch connected to the network node, the number of the first node group to which the network node belongs, and the number of the second node group to which the network node belongs. The current node determines the data routing path based on the node coordinates of the destination node, including: When the number of the second node group to which the current node belongs is equal to the number of the second node group to which the destination node belongs, and the number of the first node group to which the current node belongs is not equal to the number of the first node group to which the destination node belongs, the current node determines whether there is a first direct link between the current node and the first node group to which the destination node belongs. If yes, the current node determines that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if no, the current node determines that the data routing path includes forwarding the data packet to the second hop node, where the second hop node has a second direct link with the first node group where the destination node is located.

21. The method according to claim 20, characterized in that, The current node determines the data routing path based on the node coordinates of the destination node, including: When the number of the second node group to which the current node belongs is not equal to the number of the second node group to which the destination node belongs, the current node determines whether there is a third direct link between the second node group to which the current node belongs and the second node group to which the destination node belongs. If yes, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the second node group where the destination node is located; if no, the current node determines that the data routing path includes forwarding the data packet to the first hop node, where the first hop node and the second node group where the destination node is located have a fourth direct link.

22. The method according to claim 21, characterized in that, The method further includes: The current node determines whether the number of the first node group to which the first jump node is located is equal to the number of the first node group to which the current node is located. If not, the current node's determination of the data routing path also includes forwarding the data packet to the ferry node, and the ferry node has a fifth direct link with the first node group to which the first jump node is located.

23. The method according to any one of claims 20 to 22, characterized in that, The data packets are data packets in point-to-point traffic.

24. A method for constructing a data routing system, characterized in that, Applied to the management of equipment, the method includes: A switch and multiple network nodes are determined for constructing the data routing system, the multiple network nodes being divided into a second number of first node groups, the first node group comprising a first number of network nodes; A configuration file is sent to the switch and the plurality of network nodes, so that the network nodes in the first node group establish connections with upstream nodes and downstream nodes in each dimension according to the network structure of the data routing system in the configuration file, to form a first topology network with a surrounding network topology, and cascade with the switch to form a second topology network with a fat tree network topology, and establish connections between the second number of first node groups to form a third topology network with a fully interconnected network topology. The data routing system includes the first topology network, the second topology network, and the third topology network, with the first topology network nested within the third topology network.

25. The method according to claim 24, characterized in that, The plurality of network nodes are divided into a third number of second node groups, the second node groups including the second number of first node groups; The configuration file is also used for network nodes in the first node group to establish connections between the third number of second node groups, forming a fourth topology network with a fully interconnected network topology. The data routing system also includes the fourth topology network, and the third topology network is nested within the fourth topology network.

26. The method according to claim 24 or 25, characterized in that, The method further includes: Based on the network structure of the data routing system described in the configuration file, check the connection relationships of the multiple network nodes; If the connection relationship of the multiple network nodes is inconsistent with the connection relationship in the network structure of the data routing system, a prompt message is sent to the user, which is used to prompt the user to adjust the connection relationship of the multiple network nodes.

27. The method according to any one of claims 24 to 26, characterized in that, The method further includes: A routing table is generated based on the node coordinates of the plurality of network nodes. The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The entry in the routing table related to the at least one network node is sent to at least one of the plurality of network nodes.

28. A management device, characterized in that, The management device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions to cause the management device to perform the method as described in any one of claims 24 to 27.

29. A computer-readable storage medium, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the method of any one of claims 24 to 27.