Data routing system and data routing method
Through the hierarchical structure of nested surround network topology, fat tree network topology and fully interconnected network topology, the problems of high cost and power consumption of AI cluster network topology architecture are solved, and efficient and low-cost data routing is achieved to adapt to the high bandwidth and low-latency communication of AI clusters.
Patent Information
- Application Number
- CN202410276266.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies make it difficult to effectively reduce the cost and power consumption of AI cluster network topology architectures. In particular, the networking scale of the fat-tree network topology is limited by the number of switch ports and cannot meet the computing card interconnection requirements of large AI models.
A hierarchical network structure of nested surround network topology, fat tree network topology and fully interconnected network topology is adopted to build large-scale or ultra-large-scale networks through network devices with a small number of ports. The data routing path is determined by combining the shortest path routing algorithm and node coordinates to achieve efficient data routing between computing nodes.
Effectively reduce network costs and power consumption, improve network scalability, meet the high-bandwidth and low-latency communication requirements of AI clusters, and adapt to traffic routing under different parallel modes.
Smart Images

Figure CN120614291A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer network technology, and in particular to a data routing system, a data routing method, a method for constructing a data routing system, a management device, a computer-readable storage medium, and a computer program product. Background Art
[0002] As the scale of business continues to grow, more and more applications require the use of large-scale data centers to deploy their services. Taking artificial intelligence (AI) applications as an example, intelligent question-answering systems based on generative pre-trained transformers (GPT) are booming, and AI large models represented by GPT (also referred to as large models or models) are rapidly expanding to various fields. AI large models can be trained or inferred through AI clusters. As the parameter scale of AI large models continues to evolve, for example, from hundreds of billions to trillions, or even tens of trillions, the computing power demand of AI clusters (such as data centers where AI applications are deployed) is growing rapidly.
[0003] The rapidly increasing demand for computing power poses a significant challenge to the networking scale of AI clusters. Networks exceeding 100,000 compute cards are required to meet current AI cluster requirements. Designing the network topology for AI clusters and other data centers to effectively reduce network costs and power consumption is a crucial issue in network architecture design. Currently, large-scale AI clusters typically utilize a fat-tree network topology. However, the scale of this topology is limited by the number of switch ports. For example, a three-layer fat-tree network constructed with a mainstream 64-port switch can only interconnect a maximum of 65,536 ports, which is insufficient to meet the interconnection requirements of large AI models with hundreds of thousands of compute cards.
[0004] How to achieve large-scale or ultra-large-scale networking at a lower cost has become a key issue in the industry. Summary of the Invention
[0005] This application provides a data routing system that uses network devices with a small number of ports to construct hierarchical network structures, including nested surround network topologies, fat tree network topologies, and fully interconnected network topologies, thereby building large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption and improving network scalability. This application also provides a data routing method corresponding to the above-mentioned data routing system, a method for constructing the data routing system, a management device, a computer-readable storage medium, and a computer program product.
[0006] In a first aspect, the present application provides a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network. Among them, the network structure of the first topology network is a surround network topology structure composed of a first node group. The first node group includes a first number of network nodes, and the network nodes in the surround network topology structure have a connection relationship with the upstream node and the downstream node in each dimension. The second topology network is a fat tree network topology structure composed of each network node and switch in the first topology network. There is a cascade relationship between each network node and switch in the first topology network. The first topology network is nested in the third topology network. The network structure of the third topology network is a fully interconnected network topology structure composed of a second node group. Among them, the second node group includes a second number of first node groups, and there is a connection relationship between the second number of first node groups.
[0007] The data routing system can form a hierarchical network structure of nested surround network topology, fat tree network topology and fully interconnected network topology through a network device with a small number of ports (for example, a 9-port network device, where 9 ports refer to the 9 output ports of the network device, and the output ports can also be simply referred to as output ports), thereby forming a large-scale or ultra-large-scale network, effectively reducing network costs and power consumption, and improving network scalability.
[0008] In some possible implementations, in the second topology network, network nodes belonging to the same first node group can be directly connected to the same switch. Specifically, when the number of ports on a single switch can meet the cascading requirements of the network nodes in the first topology network, network nodes belonging to the same first node group can be directly connected to the same switch. This allows traffic to be redirected within the first node group and achieve single-hop reachability, thereby improving network performance.
[0009] In some possible implementations, the first node group includes multiple subgroups. The second topology network includes multiple first switches corresponding to the multiple subgroups and at least one second switch. In the second topology network, network nodes belonging to the same subgroup are directly connected to the first switch corresponding to the subgroup, and the multiple first switches are connected to the at least one second switch.
[0010] In this way, a multi-layer fat-tree network topology comprising multiple first switches and at least one second switch can meet the cascading requirements of network nodes in a first node group with a large number of network nodes. Even if the first node group contains a large number of network nodes, traffic can be transferred within the first node group through the first and second switches, and can be reached within a few hops. For example, if the fat-tree network has two layers, traffic can be reached within the first node group within a maximum of three hops.
[0011] In some possible implementations, subgroups can be created based on cabinets, cabinet groups, or computer rooms. By dividing first node groups of different specifications into subgroups based on size, directly connecting network nodes belonging to the same subgroup to the first switch corresponding to that subgroup, and connecting the first switches corresponding to multiple subgroups to a second switch, the number of layers in the fat-tree network can be increased to meet the cascading requirements of more network nodes.
[0012] In some possible implementations, each network node includes n ports, m of which are used to connect to network nodes in a surround network topology, and the m ports include ports for connecting to upstream nodes in each dimension and ports for connecting to downstream nodes in each dimension.
[0013] Among them, network nodes can realize the interconnection of adjacent nodes in the first node group by providing a small number of ports to connect upstream nodes and downstream nodes in various dimensions. It has good performance for neighboring communications with strong locality, especially the global reduction (all-reduce) traffic in the model parallel and data parallel modes of AI applications. Communication can be achieved through a ring algorithm. Communication mainly occurs between adjacent nodes and is naturally adapted to the surround network.
[0014] In some possible implementations, h ports out of the n ports are used to connect network nodes in a fully interconnected network topology structure formed by the second node group. The second number is less than or equal to Among them, Len j represents the number of network nodes in the jth dimension of the surrounding network topology, and q represents the number of dimensions of the surrounding network topology.
[0015] In this way, different first node groups may include at least one global link, which can provide a direct route for global communication, alleviate global network congestion, and reduce expensive global optical fiber links compared to a fat-tree network topology.
[0016] In some possible implementations, the data routing system further includes a fourth topology network. The third topology network is nested within the fourth topology network. The network structure of the fourth topology network is a fully interconnected network topology structure composed of a third node group, wherein the third node group includes a third number of second node groups, and a connection relationship exists between the third number of second node groups.
[0017] The system increases the number of layers of the fully interconnected network topology, such as adopting a nested two-level fully interconnected network topology, to compress the scale of the surrounding network, reduce the network diameter by increasing the network level, reduce communication delay, improve communication performance, and increase network throughput.
[0018] In some possible implementations, k ports out of the n ports are used to connect network nodes in a fully interconnected network topology structure formed by a third node group, and the third number is less than or equal to k·the number of nodes in the second node group+1.
[0019] In this way, different second node groups may include at least one global link, which can provide a direct route for global communication, alleviate global network congestion, and reduce expensive global optical fiber links compared to a fat-tree network topology.
[0020] In some possible implementations, the network nodes are switching nodes independent of the computing nodes, or routing engines integrated into the computing nodes. When the network nodes are integrated into the computing nodes, the routing engines can be eliminated from the interconnection switches, further reducing physical costs and power consumption.
[0021] In some possible implementations, the node coordinates of the network node are determined based on the network node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The network node is configured to determine a data routing path based on the node coordinates and implement data routing between computing nodes based on the data routing path.
[0022] In this way, data routing paths can be determined based on node coordinates using a shortest path routing algorithm. This algorithm provides the shortest distance communication between source and destination nodes, minimizing communication latency. Because the data routing system of this application utilizes a hierarchical network structure comprised of nested wrap-around, fat-tree, and fully interconnected network topologies, a deterministic shortest path routing algorithm adapted to these topological characteristics offers rapid computational speed and facilitates hardware implementation.
[0023] In some possible implementations, computing nodes are used to execute AI computing tasks in parallel, using at least one of model parallelism, pipeline parallelism, hybrid expert parallelism, or data parallelism. The data routing system of this application, by embedding different types of network topologies to form a hierarchical network, can meet the traffic routing requirements of different parallelization methods in AI scenarios.
[0024] In some possible implementations, the first topology network is used to route data packets in global all-reduce traffic, the second topology network is used to route data packets in many-to-many all-to-all traffic or point-to-point traffic, and the third topology network is used to route data packets in point-to-point traffic. All-reduce traffic can be traffic in model parallel or data parallel mode, all-to-all traffic can be traffic in hybrid expert parallel mode, and point-to-point traffic can be send / receive traffic in pipeline parallel mode. This allows routing through networks with corresponding topologies based on the characteristics of different types of traffic, thereby shortening communication latency.
[0025] In a second aspect, the present application provides a data routing method. The method is applied to a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network. The network structure of the first topology network is a ring network topology structure composed of a first node group, the first node group includes a first number of network nodes, and the network nodes in the ring network topology structure have a connection relationship with the upstream node and the downstream node in each dimension. The network structure of the second topology network is a fat tree network topology structure composed of each network node and switch in the first topology network. There is a cascade relationship between each network node and switch in the first topology network. The first topology network is nested in the third topology network. The network structure of the third topology network is a fully interconnected network topology structure composed of a second node group. The second node group includes a second number of first node groups, and there is a connection relationship between the second number of first node groups.
[0026] A current node in a data routing system receives a data packet to be routed. The current node is the source node of the data packet or a node on a path from the source node to a destination node. The current node then determines a data routing path based on the node coordinates of the destination node. The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The current node routes the data packet according to the data routing path.
[0027] This method introduces the node coordinates of the destination node in the data routing system, adapts the topological characteristics of the data routing system to perform data routing, provides the shortest distance communication between the source node and the destination node, and achieves low communication delay. Moreover, this method is simple to implement and has high availability.
[0028] In some possible implementations, when the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, the current node determines the routing port in each dimension based on the offset value of the numbers of the destination node and the current node in each dimension of the first topology network.
[0029] Among them, if the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, it means that the current node and the destination node are in the same first node group, and the current node can be routed through the surrounding network topology structure. When routing through the surrounding network topology structure, the routing port can be determined based on the offset value of the number of the destination node and the current node in each dimension of the first topology network, thereby determining the routing path and achieving low communication latency.
[0030] In some possible implementations, the number of network nodes in the i-th dimension of the first topology network is L i When the offset value of the number of the destination node and the current node in the i-th dimension of the first topology network is greater than 0 and less than or equal to L i / 2, or less than -L i / 2, determine the routing port in the i-th dimension as the port for connecting to the downstream node; when the offset value of the number of the destination node and the current node in the i-th dimension of the first topology network is greater than L i / 2, or greater than or equal to -L i / 2 and is less than 0, the routing port in the i-th dimension is determined to be the port used to connect to the upstream node. In this way, the shortest path within the surrounding network topology can be determined to achieve low communication latency.
[0031] In some possible implementations, the data packet is a data packet in a global all-reduce flow. The all-reduce flow can be communicated using a ring algorithm. Communication primarily occurs between adjacent nodes, making it a natural fit for a ring network. Routing in a first topology network with a ring structure can achieve better performance.
[0032] In some possible implementations, the current node can determine the traffic type based on the node coordinates of the destination node and the source node. The traffic type identifies a packet as intra-group or inter-group traffic. When the traffic type identifies the packet as inter-group traffic, the current node determines the forwarding port based on the node coordinates of the destination node. This allows inter-group traffic to be redirected through the switch, achieving stratification between intra-group and inter-group traffic, thus avoiding congestion.
[0033] In some possible implementations, when cross-group traffic is traffic sent from a local location to an external location, the current node may determine the node coordinates of a network node within the group that has a direct link to the destination node based on the node coordinates of the destination node, or determine the node coordinates of a network node within the group that has a direct link to a jump node to the destination node, and then determine the port number of the network node within the group connected to the switch based on the node coordinates of the network node within the group. When cross-group traffic is traffic sent from an external location to a local location, the current node determines the forwarding port based on the node coordinates of the destination node and the cascade relationship between the destination node and the switch.
[0034] The current node can forward data within the fat tree based on the shortest routing path algorithm for the above-mentioned cross-group traffic. For cross-group traffic sent externally, it can be forwarded to a local node within the group that has a direct link to the destination node. For cross-group traffic sent from the outside to the local, the data routing path can be determined based on the node coordinates of the destination node to jump within the first node group. Whether in the node group where the source node is located or the node group where the destination node is located, it can be forwarded through the switch, achieving low-latency communication.
[0035] In some possible implementations, the data packets are data packets in all-to-all traffic or point-to-point traffic. Among them, the all-to-all traffic in the hybrid expert parallel mode can be forwarded through the switch, so that traffic stratification can be achieved to avoid congestion. For the point-to-point traffic in the pipeline parallel mode, it can usually be routed through a fully interconnected network topology. If the current node is not a node directly connected to the node group where the destination node is located, the traffic can be jumped within the node group where the current node is located through the switch, specifically jumping to a node directly connected to the node group where the destination node is located, and then routing to the node group where the destination node is located through a direct link. In this way, the communication path can be shortened and low-latency communication can be achieved.
[0036] In some possible implementations, the data routing system further includes a fourth topology network. The third topology network is nested within the fourth topology network, and the network structure of the fourth topology network is a fully interconnected network topology structure composed of a third node group. The third node group includes a third number of second node groups. Connection relationships exist between the third number of second node groups. The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs and the number of the second node group to which the network node belongs.
[0037] Correspondingly, when the number of the second node group where the current node is located is equal to the number of the second node group where the destination node is located, and the number of the first node group where the current node is located is not equal to the number of the first node group where the destination node is located, the current node can determine whether there is a first direct link between the current node and the first node group where the destination node is located.
[0038] If so, the current node can determine that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if not, the current node can determine that the data routing path includes forwarding the data packet to the second jump node, and there is a second direct link between the second jump node and the first node group where the destination node is located.
[0039] This method identifies whether the current node is directly connected to the first node group where the destination node is located when the current node and the destination node are in different first node groups within the same second node group. If so, routing is performed through the direct link; if not, routing is first performed to the corresponding jump node, thereby achieving routing through the shortest path.
[0040] In some possible implementations, when the number of the second node group where the current node is located is not equal to the number of the second node group where the destination node is located, the current node may determine whether a third direct link exists between the current node and the second node group where the destination node is located.
[0041] If so, the current node may determine that the data routing path includes forwarding the data packet via the third direct link to the second node group where the destination node is located. If not, the current node may determine that the data routing path includes forwarding the data packet to the first jump node. A fourth direct link exists between the first jump node and the second node group where the destination node is located.
[0042] When the current node and the destination node are in different second node groups, the method identifies whether the current node is directly connected to the second node group where the destination node is located. If so, routing is performed through the direct link; if not, routing is first performed to the corresponding jump node, thereby achieving routing through the shortest path.
[0043] In some possible implementations, the current node may further determine whether the number of the first node group in which the first jump node resides is equal to the number of the first node group in which the current node resides. If not, the current node may further determine that the data routing path further includes forwarding the data packet to a ferry node. A fifth direct link exists between the ferry node and the first node group in which the first jump node resides.
[0044] In this method, when the first jump node and the current node are in different first node groups, the current node can first be routed to the ferry node to reach the first node group where the first jump node is located, and then the data packet can be routed to the first jump node through the intra-group routing of the first node group, and then routed based on the direct link of the jump node, thereby achieving shortest path routing.
[0045] In some possible implementations, the data packets are data packets in point-to-point traffic. The point-to-point traffic may be send / receive traffic in a pipelined parallel mode. Routing such point-to-point traffic through a fully interconnected network topology can meet the communication requirements of the point-to-point traffic.
[0046] In a third aspect, the present application provides a method for constructing a data routing system. The method can be applied to a management device and includes:
[0047] determining a switch and a plurality of network nodes for constructing the data routing system, the plurality of network nodes being divided into a second number of first node groups, the first node groups including a first number of network nodes;
[0048] A configuration file is sent to the switch and the multiple network nodes so that the network nodes in the first node group establish connections with upstream nodes and downstream nodes in various dimensions according to the network structure of the data routing system in the configuration file to form a first topology network with a surround network topology structure, and cascade with the switch to form a second topology network with a fat tree network topology structure, and establish connections between the second number of first node groups to form a third topology network with a fully interconnected network topology structure. The data routing system includes the first topology network, the second topology network, and the third topology network, and the first topology network is nested in the third topology network.
[0049] This method determines switches and multiple network nodes used to construct a data routing system, and instructs the network nodes and switches to establish connections according to the network structure of the data routing system in the configuration file, thereby constructing a data routing system with nested encircling network topologies, fat-tree network topologies, and fully interconnected network topologies. This allows for the establishment of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. The encircling network topology accelerates the application of local communication features, supports multi-tenant isolation, and accelerates resource pooling. The fat-tree network topology can be used for global traffic forwarding and can also be used as an alternative route for fault recovery. For ultra-large-scale supernodes, a two-layer fat-tree topology can be used for non-blocking expansion to achieve global traffic forwarding. The fully interconnected network topology provides direct routing for global communications, alleviates global network congestion, and can reduce expensive global fiber links compared to the fat-tree network topology architecture.
[0050] In some possible implementations, the plurality of network nodes are divided into a third number of second node groups, where the second node groups include a second number of first node groups. The configuration file is further configured to enable the network nodes in the first node group to establish connections between the third number of second node groups, thereby forming a fourth topology network having a fully interconnected network topology. The data routing system further includes a fourth topology network, wherein the third topology network is nested within the fourth topology network.
[0051] By nesting a multi-layer fully interconnected network topology, it is possible to provide more ports and connect more computing nodes, thereby building an ultra-large-scale network to meet business needs.
[0052] In some possible implementations, the management device may also check the connection relationships of the multiple network nodes based on the network structure of the data routing system described in the configuration file. If the connection relationships of the multiple network nodes are inconsistent with the connection relationships in the network structure of the data routing system, the management device may send a prompt to the user, prompting the user to adjust the connection relationships of the multiple network nodes. This ensures the accuracy and reliability of the network.
[0053] In some possible implementations, the management device may further generate a routing table based on node coordinates of multiple network nodes. The node coordinates are determined based on the network node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs. The management device may then distribute an entry in the routing table related to at least one of the multiple network nodes.
[0054] The management device sends the table entries related to the network node in the routing table to at least one network node, which can reduce the storage resources occupied by the routing table entries in the network node and save storage space.
[0055] In a fourth aspect, the present application provides a management device. The management device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor is configured to execute the instructions stored in the at least one memory, so that the management device performs the method for constructing a data routing system as described in the third aspect or any implementation of the third aspect.
[0056] In a fifth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct a management device to execute the method for constructing a data routing system described in the first aspect or any implementation of the first aspect.
[0057] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a management device, enables the management device to execute the method for constructing a data routing system as described in the first aspect or any one of the implementations of the first aspect.
[0058] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical methods of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments.
[0060] Figure 1 An architectural diagram of a data routing system provided for this application;
[0061] Figure 2 A schematic diagram of the ports of a node in a data routing system provided by this application;
[0062] Figure 3 A schematic diagram of a surround network topology provided in this application;
[0063] Figure 4 A schematic diagram of the structure of a 9-port router provided in this application;
[0064] Figure 5 This is a schematic diagram of the structure of a routing engine integrated in a processor provided by this application;
[0065] Figure 6 A schematic diagram of a surround network topology structure provided by this application;
[0066] Figure 7 Schematic diagram of a hierarchical network provided by this application, including a nested surround network topology, a fat tree network topology, and a two-level fully interconnected network topology;
[0067] Figure 8 A flowchart of a data routing method provided by this application;
[0068] Figure 9 A schematic diagram of a data routing process in a data routing system provided by this application;
[0069] Figure 10 A schematic diagram of data routing through a virtual channel provided by this application;
[0070] Figure 11 Schematic diagram of a hierarchical network provided by this application, including a nested surround network topology, a fat tree network topology, and a one-level fully interconnected network topology;
[0071] Figure 12 A schematic diagram of port bandwidth allocation provided in this application;
[0072] Figure 13 A schematic diagram of a data routing process in a data routing system provided by this application;
[0073] Figure 14 A schematic diagram of fault recovery based on a backup cabinet provided in this application;
[0074] Figure 15 A flowchart of a method for constructing a data routing system provided in this application;
[0075] Figure 16 A schematic diagram of the structure of a management device provided for this application;
[0076] Figure 17 A schematic diagram of the structure of a computing device provided in this application;
[0077] Figure 18 A schematic diagram of the structure of a computing device cluster provided in this application;
[0078] Figure 19 A schematic diagram of the structure of another computing device cluster provided in this application;
[0079] Figure 20 This is a structural diagram of another computing device cluster provided in this application. DETAILED DESCRIPTION
[0080] The terms "first" and "second" in the embodiments of this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.
[0081] First, some technical terms involved in the embodiments of this application are introduced.
[0082] Network topology refers to the specific arrangement or connection between network components (e.g., switches, routers, and other network or communication devices). It is generally categorized as either a physical, real, online structure or a logical, virtual, programmatic structure. If two networks have different physical connections and node distances, but the connections between the network components are the same, then the network topologies of the two networks are considered the same.
[0083] Networking, short for network formation, typically involves designing a suitable network topology based on networking requirements, including scale and performance requirements, and then implementing the network according to that topology. In scenarios such as cloud computing, high-performance computing (HPC), and artificial intelligence (AI), it's often necessary to connect computing devices (also known as compute nodes) through network devices (such as switches and routers) to form large-scale or ultra-large-scale interconnected networks. This provides efficient communication capabilities with high bandwidth, low latency, and high throughput, effectively unleashing computing power.
[0084] Among them, cloud computing, also known as network computing, is an Internet-based computing method through which shared software and hardware resources and information can be provided to various terminals or other devices on demand. Terminals and other devices use the infrastructure provided by the service provider for computing. High-performance computing refers to providing more powerful computing performance than traditional computers or servers by aggregating computing power. It can usually be used in scenarios such as weather forecasting, astrophysical simulation, molecular dynamics simulation, and earthquake simulation. AI computing is a mathematically intensive process for executing machine learning algorithms, which can usually use accelerated systems and software. It can extract new insights from large data sets and learn new capabilities in the process. AI computing can be applied to model training or model inference, especially large language model (LLM) training or inference.
[0085] For ease of description, this application mainly uses AI computing scenarios as examples.
[0086] AI computing can usually be achieved through AI clusters, such as large-scale AI training clusters. Among them, the AI cluster may include multiple computing nodes, wherein the multiple computing nodes can be interconnected through network nodes such as routers. It should be noted that the computing nodes and network nodes can be independent of each other, or the computing nodes and network nodes can also be integrated. For example, a routing engine can be integrated into the computing node to realize the routing function of the network node. During training, the AI cluster can adopt a distributed parallel strategy for model training to improve efficiency. Depending on the scale or function of the model, the distributed parallel strategy may include but is not limited to model parallelism (such as tensor parallelism), data parallelism, pipeline parallelism or expert parallelism.
[0087] Model parallelism (MP) involves partitioning an entire model (such as the aforementioned LLM) into multiple sub-models (also called model fragments). Each sub-model is sent to a different compute node for processing. Each compute node is responsible for processing its own sub-model, calculating local gradients and transmitting these gradients to a central node (usually a parameter server) through a communication mechanism.
[0088] Data Parallelism (DP) is to divide the complete dataset into multiple sub-datasets, each computing node processes a different sub-dataset, and then aggregates the updated gradients.
[0089] Pipeline Parallelism (PP) means dividing the model into multiple stages and distributing them to different computing nodes, with each computing node completing the training in a "relay" manner.
[0090] Expert parallelism, also known as mixture of experts (MoE) parallelism, typically decomposes a task into several subtasks, trains an expert model for each subtask, and develops a gating model. The gating model learns the expert models based on the input samples and combines the output results. This strategy dynamically activates parts of the neural network, significantly increasing the number of model parameters without increasing the computational load. For example, large sparse models with trillions of parameters can be trained using the expert parallelism strategy.
[0091] Model parallelism primarily involves global all-reduce traffic, with data volumes in the gigabyte (GB) range. Furthermore, communication and computation in model parallelism are non-overlapping. Expert parallelism involves a significant proportion (e.g., 50% or more) of many-to-many all-to-all (or all2all) global traffic, also in the GB range. Furthermore, communication and computation in expert parallelism are non-overlapping. Therefore, both model and expert parallelism typically require high bandwidth and throughput to accelerate communication performance. Data parallelism primarily involves all-reduce traffic, with data volumes in the GB range. However, communication and computation in data parallelism are overlappable (or hidden), so optimizing communication is not a core requirement. Pipeline parallelism primarily involves point-to-point send and receive communication, with data volumes in the megabyte (MB) range. High cross-node throughput is sufficient for communication bandwidth requirements. Pipeline parallelism mainly involves point-to-point communication between the sender and the receiver. The data volume is also at the MB level, and the communication bandwidth requirement is only point-to-point communication across nodes.
[0092] Model training can be achieved by combining different parallel strategies, depending on the model and hardware specifications. For example, a dense transformer model with hundreds of billions of parameters, such as the Generative Pre-trained Transformer (GPT) model, can include 175 billion (B, with B representing billion) parameters, and its network consists of 96 layers. The ratio of computation to communication is approximately 8:2.
[0093] Based on the specifications of the model and hardware, an efficient parallel deployment strategy can be obtained through parallel strategy search, as shown below:
[0094] For model parallel domains, the primary traffic is all-reduce traffic, with data volumes at the GB level. Communication and computation cannot overlap, so single-node deployment is feasible, with cards interconnected via a high-bandwidth bus. For pipeline parallel domains, the primary traffic is send / receive (or send / recv) traffic, with data volumes at the MB level. Communication and computation can overlap, allowing for inter-node deployment and network interconnection. For data parallel domains, the primary traffic is all-reduce traffic, with data volumes at the GB level. Communication and computation can overlap, allowing for inter-node deployment and network interconnection.
[0095] For large sparse models with trillions of parameters, communication accounts for up to 69.3%. The traffic analysis of the above large sparse models with trillions of parameters is shown below:
[0096] Table 1 Analysis of traffic of trillion-level sparse AI models:
[0097] Parallel mode communication operations traffic Calculation time The need for communication Model Parallelism (MP) all-reduce GB Seconds (non-overlapping) High-speed interconnection within the node Pipeline Parallel (PP) send / recv MB Seconds (overlappable) Cross-node point-to-point communication Mixed-of-Experts Parallel (MOE) all2all GB Seconds (non-overlapping) Cross-node all2all Data Parallelism (DP) all-reduce GB Tens of seconds (partial overlap possible) High throughput across nodes
[0098] For a large, sparse model with trillions of parameters, a single-step communication time can be as low as XX12 milliseconds (ms), and the total communication time can be as low as XX15 ms. The "XX12" and "XX15" numbers are masked, indicating that the single-step and total communication times are in the thousands of milliseconds.
[0099] Currently, many AI model trainings, especially those for large AI models, typically employ a combination of multiple parallel strategies for distributed parallel training, which can generate different types of traffic. For example, when training large AI models with trillions of parameters, a combination of model parallelism, data parallelism, pipeline parallelism, and hybrid expert parallelism can be employed for distributed parallel training. Accordingly, the traffic generated during training can include all-reduce traffic, all-to-all traffic, and point-to-point send / recv traffic (also known as point-to-point traffic). How to configure networks to meet the routing requirements of different types of traffic has become a key concern in the industry.
[0100] The industry offers a variety of network topologies for networking. Among them, the fat tree network topology (also referred to as a fat tree network or fat tree topology) is a tree-shaped network topology in which network bandwidth does not converge from leaf nodes (or leaf nodes) to the root node. It is a classic high-performance network topology with advantages such as excellent performance, system stability, and wide adaptability. It is currently the mainstream network interconnection solution for data centers (such as AI clusters or data centers used to train AI models). The Dragonfly network topology (referred to as the Dragonfly topology or Dragonfly network) is a network topology that is fully interconnected within and between groups. Compared to the fat tree topology, the Dragonfly topology can reduce the network layer and the number of optical fibers and optical modules, resulting in a higher cost-performance ratio. The surround network topology (referred to as the surround topology or surround network) has low networking costs, strong scalability, and good performance for localized neighbor communications. Taking the torus topology as an example, for the all-reduce traffic in the model parallel and data parallel modes of AI applications, communication can be achieved through a ring algorithm. Communication mainly occurs between adjacent nodes and is naturally adapted to the torus topology.
[0101] The Dragonfly topology is based on switch expansion, and its network scale is limited by the switch port specifications. Based on the current mainstream 64-port high-performance switch, the maximum network scale is only 270,000 ports, which cannot meet the networking needs of large-scale AI clusters. The communication latency of the Torus topology scales linearly with network scale, which has certain limitations. Furthermore, the Torus topology is a blocking network, and due to the topology and routing algorithm, the network throughput of global traffic is low. Therefore, it is not friendly to all-to-all global traffic in the hybrid expert parallel mode, and its performance for large-scale all-to-all global traffic is very poor.
[0102] Given the advantages of a fat-tree topology, such as high network throughput, non-blocking performance, and stable performance, large-scale AI clusters often adopt a fat-tree network topology for networking. Using a mainstream 64-port high-performance switch as an example, constructing a three-layer fat-tree network with 64-port high-performance switches can provide a maximum of 65,536 interconnected ports. This still falls short of meeting the interconnection requirements of hundreds of thousands of computing cards in large AI models, such as those for hundreds of thousands of graphics processing units (GPUs). Increasing the number of layers in a fat-tree network to increase the number of port interconnections will significantly reduce the port utilization of the switches in the fat-tree network. Specifically, the port utilization of switches in a fat-tree network is typically less than 33%, and this utilization can decrease as the number of layers increases. For example, the port utilization of a three-layer fat-tree is only 20%. Furthermore, large-scale fat-tree networks require a large amount of optical fiber and optical modules to connect each layer, significantly increasing reliability pressure and making it difficult to meet the networking requirements of ultra-large-scale networks. In addition, the port bandwidth of mainstream high-performance switches has reached 400 Gigabit per second (Gbps). The power consumption of a 64-port high-performance switch is about 200 to 400 watts (the full name is watt, symbol is w), and the power consumption of the optical module is about 13 watts. If an optical module is used for connection, the switch energy consumption can be as high as 1000 watts.
[0103] In view of this, the present application provides a low-cost, highly reliable data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network, a second topology network, and a third topology network.
[0104] The network structure of the first topology network is a ring network topology (e.g., a torus) consisting of a first node group. The first node group includes a first number of network nodes. The network nodes in the ring network topology are connected to upstream nodes and downstream nodes in each dimension. The network structure of the second topology network is a fat tree network topology consisting of the network nodes and switches in the first topology network. A cascade relationship exists between the network nodes and switches in the first topology network. Switches can be classified into different types based on their deployment location. For example, switches can be top of rack (TOR) switches, end of row (EOR) switches, and middle of row (MOR) switches. The first topology network is also nested within a third topology network. Nesting refers to insertion or embedding. Specifically, the first topology network can be embedded within the third topology network as a component or constituent unit of the third topology network. The network structure of the third topology network can be a fully meshed network topology consisting of a second node group. The second node group includes a second number of first node groups, and there is a connection relationship between the second number of first node groups.
[0105] The data routing system can form a hierarchical network structure of nested Torus topology, fat tree topology and full mesh topology through network devices with a small number of ports (for example, a 9-port network device. To avoid confusion, this application does not consider access ports), thereby forming a large-scale or ultra-large-scale network, such as an ultra-large-scale network with 17.3 million ports, effectively reducing network costs and power consumption and improving network scalability.
[0106] Furthermore, for scenarios with relatively regular and predictable traffic, such as AI computing scenarios, the data routing system can perform traffic stratification based on the communication characteristics of the scenario (or the applications and models in the scenario, such as AI applications and AI large models). The Torus topology is naturally adapted to all-reduce traffic, and the fat tree topology can meet the routing requirements of all-to-all traffic within a supernode, or can jump global point-to-point traffic (such as send / receive traffic, also known as Send / Recv traffic). The full-mesh topology provides direct routing for global communications such as point-to-point traffic in a pipeline parallel mode. By integrating the above-mentioned multiple topologies for networking, the number of switches can be reduced, the number of networking links can be reduced, the number of optical fiber links and optical modules can be reduced, and the system reliability can be improved. The data routing system is load-friendly and effectively reduces network costs and energy consumption by performing hierarchical routing of traffic while providing high-performance, high-reliability, and high-availability communications.
[0107] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the data routing system of the present application is introduced below with reference to the accompanying drawings.
[0108] Figure 1 A schematic diagram of the architecture of a data routing system is shown. The data routing system 10 is used to implement data routing between computing nodes. The data routing system 10 includes multiple network nodes (sometimes collectively referred to as "nodes 110" below), wherein the multiple nodes 110 are located in a communication network 150 that includes at least a first topology network (for example, 120-1, 120-2, ..., 120-Q, sometimes collectively referred to as "first topology network 120" below), a second topology network 130, and a third topology network 140.
[0109] In an embodiment of the present application, the first topology network 120 is nested in the third topology network 140. The network structure of the first topology network 120 is a surround network topology structure composed of a first node group (e.g., 120-1, 120-2, ..., 120-Q, where Q is a positive integer greater than 1). The first node group includes a first number of network nodes. For example, the first node group may include node 110-1, node 110-2, ..., node 110-P, where P is a positive integer greater than 1. The network nodes in the surround network topology structure have a connection relationship with the upstream node and the downstream node in each dimension. Taking node 110-2 as an example, the upstream node and downstream node included in a certain dimension of node 110-2 are nodes 110-1 and node 110-3 respectively, and node 110-2 has a connection relationship with node 110-1 and node 110-3.
[0110] The network structure of the second topology network 130 is a fat-tree network topology structure composed of the network nodes in the first topology network 120 and a switch 132. The switch 132 is used to route traffic between the network nodes in the first topology network 120. Depending on the deployment location, the switch 132 may include, but is not limited to, a TOR switch, an EOR switch, or a MOR switch. A cascade relationship exists between the network nodes in the first topology network 120 and the switch 132. For example, each network node in the first topology network 120 can be connected to at least one port of the switch 132, and different network nodes can be connected to different ports.
[0111] The first topology network 120 is nested in the third topology network 140. The network structure of the third topology network 140 is a fully interconnected network topology structure composed of second node groups. The second node groups include a second number of first node groups, for example, 120-1, 120-2, ..., 120-Q. There is a connection relationship between the second number of first node groups. Figure 1For example, a connection relationship may exist between 120-1, 120-2, ..., and 120-Q. It should be noted that the connection relationship between node groups may be a connection relationship between nodes in different node groups, for example, a connection relationship exists between a node in 120-1 and a node in 120-2.
[0112] The data routing system 10 may include a layer of fully interconnected networks, or multiple layers of fully interconnected networks. When the data routing system 10 includes multiple layers of fully interconnected networks, the multiple layers of fully interconnected networks may be nested. For example, the first layer of fully interconnected networks may be nested within the second layer of fully interconnected networks, and so on. The L-1 layer of fully interconnected networks may be nested within the L layer of fully interconnected networks.
[0113] The present application does not impose any restrictions on the number of layers (or levels) of the fully interconnected network in the data routing system 10. The number of layers of the fully interconnected network can be configured according to the required network scale. The network scale can be characterized by the total number of ports. For example, when the total number of ports is less than or equal to the first threshold, the data routing system 10 can be configured with a one-layer fully interconnected network topology to meet the demand. For another example, when the total number of ports is greater than the first threshold and less than or equal to the second threshold, the data routing system 10 can be configured with a two-layer fully interconnected network topology to meet the demand. It should be noted that the above-mentioned first threshold can be the maximum number of ports that the data routing system 10 can provide when a one-layer fully interconnected network topology is nested, and the above-mentioned second threshold can be the maximum number of ports that the data routing system 10 can provide when a two-layer fully interconnected network topology is configured.
[0114] Based on this, the data routing system may further include a fourth topology network, and the third topology network 140 may be nested within the fourth topology network. Similar to the third topology network 140, the network structure of the fourth topology network may also be a fully interconnected network topology structure. The main difference between the fourth topology network and the third topology network 140 is that the network structure of the fourth topology network is a fully interconnected network topology structure based on a third node group. The third node group includes a third number of second node groups, and a connection relationship exists between the third number of second node groups. The third node group is a node group formed by network nodes in the third topology network 140, thereby enabling the third topology network 140 to be nested within the fourth topology network.
[0115] Similar to the fully interconnected network topology, the fat tree network topology can be a single layer or multiple layers. In some possible implementations, the number of ports of a switch can meet the cascading requirements of the network nodes in the first topology network. Accordingly, in the second topology network, the network nodes belonging to the same first node group can be directly connected to the same switch. For example, the first node group can include network nodes of a 4*4*4 cabinet. In other words, the number of network nodes in the first node group is less than or equal to 64. A 64-port switch can meet the cascading requirements of each network node in the cabinet, and each network node in the cabinet can be directly connected to at least one port of the above-mentioned 64-port switch. In other possible implementations, the first topology network includes a large number of network nodes, and the number of ports of a single switch is difficult to meet the cascading requirements of the network nodes in the first topology network. Based on this, the cascading requirements of the network nodes in the first topology network can be met by multiple switches. Specifically, the first node group includes multiple subgroups, wherein the subgroups can be obtained by dividing according to cabinets, cabinet groups or computer rooms. For example, when the first node group includes 1024 network nodes, such as network nodes in a super node formed by 16 8*8 cabinets, the subgroups can be obtained by dividing the cabinets. The second topology network includes multiple first switches corresponding to multiple subgroups and at least one second switch. In the second topology network, the network nodes belonging to the same subgroup are directly connected to the first switch corresponding to the subgroup. It should be noted that each subgroup can correspond to one or more first switches. Multiple first switches can be connected upward to at least one second switch. Among them, the first switch can be a switch set in the cabinet, such as a TOR switch, and the second switch can be a switch set between cabinets.
[0116] Next, the network topology is described in conjunction with the port connections of the network nodes. Figure 2 , each network node may include n ports, such as port 111-1, port 111-2, ..., 111-N. Each network node may provide m ports for connecting network nodes in a surround network topology. The surround network topology may include one or more dimensions. For example, the surround network topology may be a three-dimensional (3D) surround network topology (such as a 3D torus). Accordingly, the m ports include ports for connecting upstream nodes of each dimension and ports for connecting downstream nodes of each dimension. Wherein, m is greater than or equal to 2. Figure 3 As shown, taking a three-dimensional surround network topology as an example, the X dimension of the standard three-dimensional surround network topology can provide X+ ports and X- ports representing positive and negative directions. The X+ port of each upstream node is connected to the X- port of the downstream node, and the ends are connected to form a surround network.
[0117] Each network node may also provide at least one port for connecting to a switch in a fat-tree network topology. For example, when the first node group includes 64 network nodes, the switch may be a 64-port switch, and each network node may provide a port for connecting to a port of the 64-port switch.
[0118] The remaining ports of each network node can be used to connect to network nodes in a fully interconnected network topology. The number of remaining ports can be greater than or equal to one. In addition to the ports connected to network nodes in the surround network topology and the ports connected to switches in the fat-tree network topology, at least one port of the network node, other than the ports connected to network nodes in the surround network topology and the ports connected to switches in the fat-tree network topology, can be connected to network nodes in other first node groups to form a fully interconnected network topology. For example, a network node can connect to network nodes in other first node groups via one port to form a third topology network. Furthermore, a network node can connect to network nodes in other second node groups via another port to form a fourth topology network.
[0119] Figure 1 The node 110 in the embodiment can be implemented as a physical node or a virtual node. The specific implementation of the physical node and the virtual node will be described below.
[0120] In some embodiments, node 110 may be a switching node independent of a computing node, such as an independent switch or router. Node 110 may be a low-port switch or a low-port router. By using a low-port switch or a low-port router for expansion, the networking cost is low and multi-path characteristics are provided. Multi-path characteristics refer to the use of multiple paths to connect nodes to avoid single point failures. For ease of description, a low-port router is used as an example. A low-port router is a router with a small number of ports relative to a high-radix router. For example, a low-port router may be a 9-port router, or an 11-port, 12-port, 13-port, or 14-port router. A high-radix router may be a 48-port router or a 64-port router.
[0121] Figure 4 A schematic diagram of a nine-port router is shown. The nine-port router includes multiple communication engines, a crossbar switch matrix, and nine ports. The communication engines facilitate communication between the ports of a router or routing engine and the processor, enabling operations such as direct memory access (DMA), data read and write, and memory copy. The crossbar switch matrix facilitates communication between the ports of a router or routing engine and the communication engines, enabling internal data forwarding in low-port mode.
[0122] like Figure 4 As shown, the nine ports can include six ports for building a surround network topology, one port for building a fat tree network topology, and two ports for building a fully interconnected network topology. The six ports for building a surround network topology can include two ports each in the X, Y, and Z dimensions. Each dimension distinguishes between positive and negative directions, with ports LinkX+ and LinkX- responsible for interconnecting adjacent nodes in the X dimension, LinkY+ and LinkY- responsible for interconnecting adjacent nodes in the Y dimension, and LinkZ+ and LinkZ- responsible for interconnecting adjacent nodes in the Z dimension. It should be noted that in a surround network topology, the first and last nodes in the same dimension can also be considered adjacent nodes. For example, when the X dimension includes four network nodes, the first network node is connected not only to the second network node but also to the last network node. The one port for building a fat tree network topology can include Link-SW, which is used to connect to a switch, such as a 64-port TOR switch within a rack. The two ports used to construct a fully interconnected network topology include a port L-Link for implementing a first-layer fully interconnected network (such as the aforementioned third topology network) and a port G-Link for implementing a second-layer fully interconnected network (such as the aforementioned fourth topology network). The L in L-Link can represent local, and the G in G-Link can represent global. The L-Link can implement interconnection between L1-level topology (first topology network) groups (abbreviated as L1 group, L1 group, for example, the first node group), and the G-Link can implement interconnection between L3-level topology (such as the third topology network) groups (abbreviated as L3 group, L3 group, for example, the second node group).
[0123] In other embodiments, node 110 may be a routing engine integrated into a computing node. The routing engine may be a routing engine integrated into a processor. The processor may be a graphics processing unit (GPU), a neural network processing unit (NPU), or other accelerator device, or the processor may be a central processing unit (CPU).
[0124] Figure 5A schematic diagram of the architecture of a routing engine integrated into a processor is shown. The processor may include multiple cores and a last-level cache (LLC). Processors typically employ multiple levels of cache, with each level of cache being faster than the next. When a core needs to access a piece of data or instruction, it first checks the nearest L1 cache. If the data exists, it's a cache hit; otherwise, it's a cache miss, requiring the next level of cache to be queried. In some possible implementations, the processor may be a CPU / GPU memory chip integrated with high-bandwidth memory (HBM). The HBM may be a stack of multiple Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM, abbreviated as DDR). After stacking, the stacked chips can be packaged together with the CPU / GPU die to create a large-capacity, high-bit-width DDR combination array.
[0125] The structure of the routing engine integrated in the processor can be referred to Figure 4 In the structure of low- and medium-port routers, the routing engine can include a communication engine, a crossbar switch matrix, and ports. Figure 5 In the example, the processor integrates or has a built-in 9-port routing engine, and only 6 ports are needed to achieve expansion in the three dimensions of X, Y, and Z, where Kx, Ky, and Kz represent the length of each dimension respectively. The length of each dimension can be represented by the number of nodes in each dimension. Each dimension can distinguish between two positive and negative directions, and define ports Link X+ and Link X- to be responsible for interconnecting adjacent nodes in the X dimension; Link Y+ and Link Y- to be responsible for interconnecting adjacent nodes in the Y dimension; and Link Z+ and Link Z- to be responsible for interconnecting adjacent nodes in the Z dimension. In this embodiment, in order to control the network diameter, the lengths of the X, Y, and Z dimensions are all 4. A port Link SW is provided in the 9-port routing engine for connecting to a cabinet switch, such as a TOR switch, thereby forming a fat tree topology network structure. The remaining two ports L-Link and G-Link in the 9 ports are used to implement a two-layer fully interconnected network topology.
[0126] The core of the processor may communicate with the routing engine via an internal bus, which may include but is not limited to a high-speed serial component interconnect bus (Peripheral Component Interconnect Express, PCIe) or a Compute eXpress Link (CXL) bus.
[0127] It should be understood that the processor or router or routing engine of the present application may include more or fewer components, functions or modules. In some embodiments, the processor may also include components such as a resonant circuit, and may also be connected to external devices such as storage devices. For example, a processor with an integrated routing engine may also include a memory management unit (MMU) and a memory controller (MC) for managing and controlling external or extended memory such as DDR. The memory management unit and the memory controller may communicate with the extended DDR via a unified bus (UB). Similarly, the processor may also be connected to or extended with a solid-state drive (SSD), wherein the processor set may communicate with the SSD via a UB bus or a Ctrl bus. In addition, the processor may also be connected to or extended with devices such as an accelerator. In some embodiments, the router or routing engine may be any device that implements data transmission, such as a network device such as a switch.
[0128] Need to explain, Figure 4 、 Figure 5 The examples are all 9 ports. In other possible implementations of the embodiments of the present application, the routing engine integrated in the low-port router or processor may also include other numbers of ports, such as 8 ports, 10 ports, 11 ports, 12 ports, 13 ports, etc. Taking the example of 8 ports, the routing engine integrated in the low-port router or processor can build a surround network topology using 6 of the 8 ports, a fat tree network topology using 1 of the 8 ports, and a fully interconnected network topology using the remaining 1 port.
[0129] In the present application, computing nodes can be used to perform AI computing tasks in parallel. The computing nodes can be independent computing nodes, or computing nodes integrated with a routing engine, such as processors integrated with a routing engine. The parallel methods of AI computing tasks include at least one of model parallelism, pipeline parallelism, hybrid expert parallelism or data parallelism. Accordingly, the first topology network of the data routing system can be used to route data packets in all-reduce traffic, such as data packets in all-reduce traffic in model parallel or data parallel mode, the second topology network is used to route data packets in many-to-many all-to-all traffic or point-to-point traffic, such as routing data packets in all-to-all traffic in hybrid expert parallel mode, or to realize the jump of data packets in point-to-point traffic (such as send / receive traffic in pipeline parallel mode) within the first node group, and the third topology network is used to route data packets in point-to-point traffic, such as routing send / receive traffic in pipeline parallel mode.
[0130] In order to make the technical solution of the present application clearer and easier to understand, the following example illustrates the construction of a Torus using a 9-port router or routing engine, and then constructing a hierarchical network of nested Torus, fat tree, and 2-level full mesh.
[0131] Large AI models with hundreds of billions of parameters primarily utilize model parallelism, data parallelism, and pipeline parallelism for parallel execution. These large AI models primarily utilize all-reduce traffic from model and data parallelism, typically in the GB range, requiring high bandwidth to accelerate communication. Furthermore, large AI models also utilize pipeline parallel send / receive traffic, which typically operates in the MB range, placing less pressure on communication bandwidth.
[0132] Among them, all-reduce communication can be converted into communication between neighboring nodes through the ring algorithm, and the Torus topology is naturally adapted to this type of communication. The Torus topology can be expanded based on low-port routers to reduce network costs and energy consumption, and has multi-path characteristics. For 3D Torus networks, communication routing can be provided in three dimensions: X, Y, and Z, a total of six directions. Direct copper cable connections within the cabinet can meet the communication needs of all-reduce. Even for cross-cabinet communication, the Torus architecture can greatly reduce the use of long-distance optical fibers and optical modules. Therefore, this application can use the Torus topology to build an L1 topology to transmit model-parallel and data-parallel all-reduce traffic.
[0133] Large AI models involve extensive parallel computing. If the scale of a supernode, formed by connecting multiple computing nodes via a network, can accommodate the scale of various parallel computations, primary communication can be concentrated within the supernode. The ultra-large communication bandwidth within the supernode accelerates communication. However, if the supernode scale is smaller than the scale of various parallel computations, extensive inter-supernode communication is required. However, the communication bandwidth between supernodes is typically much smaller than the communication bandwidth within a supernode, significantly impacting communication performance. Therefore, the supernode scale must typically be adapted to the scale of parallel computations required for large AI models.
[0134] For the scale of parallel computing of large AI models with hundreds of billions of parameters, a supernode with 64 accelerator cards (for example, 64 GPUs) can meet the requirements. A standard cabinet can usually accommodate 64 GPUs. Therefore, this application can build a 4*4*4 3D Torus topology, and a single computing cabinet can meet the training needs of large models with hundreds of billions of parameters.
[0135] The communication between cabinets of large AI models with hundreds of billions of parameters mainly consists of parallel send / receive traffic. The data volume is only at the MB level, and the bandwidth requirement is not large. Therefore, a full mesh topology can be used to achieve inter-cabinet networking. Compared with the torus topology, the full mesh topology can provide higher binary bandwidth. For uniform random traffic, the communication efficiency of the full mesh topology is the same as that of the fat tree topology, and 100% network utilization can be achieved. When implemented using the router integrated in the processor, it can be directly connected based on the processor port without the need for switch expansion. Compared with the fat tree topology, it can save 50% of global fiber links and optical modules, effectively reducing network costs and improving network reliability.
[0136] Among them, the present application can nest a two-level full mesh topology to expand between cabinets. Specifically, building an L3 topology / L4 topology based on fullmesh, such as the first-layer fully interconnected network and the second-layer fully interconnected network in the fully interconnected topology network architecture, can effectively reduce the network diameter, reduce communication latency, and effectively alleviate the congestion problem caused by global communication. When a two-level full mesh topology is adopted, a 17.3 million node scale network can be built based on a 9-port router (switch). When a routing engine integrated in the processor is used, switch interconnection can be eliminated, further reducing network costs and power consumption.
[0137] It should be noted that, since the full mesh interconnection between cabinets is connected to different cabinets through each GPU, the traffic for inter-cabinet communication needs to be transmitted between the GPU cards in the cabinet to reach the corresponding direct link, which requires transmission through the Torus topology within the cabinet, which will conflict with the all-reduce traffic inside the cabinet. Therefore, in order to avoid traffic congestion, this application also adds a switch, such as adding a TOR switch in the cabinet, which is responsible for the jump of the parallel flow traffic between cabinets to achieve traffic stratification. The all-reduce traffic inside the cabinet flows within the Torus topology, and the cross-cabinet traffic flows within the cabinet through the TOR switch to avoid congestion. In some possible implementations, the TOR switch can also be used as a replacement route for fault recovery to forward traffic. Among them, the TOR switch in the cabinet constructs an L2 level topology.
[0138] It should be noted that for a cabinet containing 64 GPUs, each GPU is connected to a port. Therefore, a 64-port TOR switch can be set up in the cabinet to form a fat tree topology to jump the parallel flow of traffic between cabinets. In some embodiments, multiple planes can also be used to increase bandwidth based on the amount of parallel flow data. For example, if a 256-port switch is required to connect to each GPU, four 64-port TOR switches can be used to connect to each GPU respectively, thus increasing bandwidth at a lower cost.
[0139] The following diagram illustrates the construction of a Torus, as well as examples of building a hierarchical network of nested Torus, fat trees, and a two-level full mesh.
[0140] like Figure 6 As shown, the Torus topology is first used to build the L1 topology. Specifically, each processor has a built-in 9-port routing engine, and only 6 ports (external ports, outbound ports) are required, such as ports 1 to 6, to achieve expansion in the X, Y, and Z dimensions. Among them, Kx, Ky, and Kz are the lengths of each dimension, indicating the number of network nodes in each dimension. Each dimension distinguishes between positive and negative directions, and defines ports Link X+ and Link X- responsible for X-dimensional interconnection; Link Y+ and Link Y- responsible for Y-dimensional interconnection; Link Z+ and Link Z- responsible for Z-dimensional interconnection. Taking a super node including 64 computing nodes (such as 64 GPUs) as an example, the lengths of the X, Y, and Z dimensions are all 4, so the L1 topology can connect 4*4*4=64 computing nodes.
[0141] Next, see Figure 7The seventh port of the 9-port routing engine, such as a Link-SW, can connect to a 64-port ToR switch within the rack, building an L2 topology. The ToR switch in this L2 topology is primarily responsible for routing parallel send / receive traffic across the racks, for example, between the 64 compute nodes within the rack. Alternatively, when large AI models are trained in parallel using the hybrid expert parallel mode, the ToR switch can also be used to route all-to-all traffic in this mode.
[0142] The eighth port of the 9-port routing engine, such as an L-Link, can be used for interconnection between L1 topology groups. With 64 L-Links within a cabinet, a full mesh of up to 65 cabinets can be directly connected. An L3 topology group can consist of 65 cabinets, connecting up to 64 * 65 = 4160 compute nodes, primarily responsible for transmitting send / receive traffic between cabinets.
[0143] The ninth port of a nine-port routing engine, such as a G-Link, can be used for interconnection between L3 topology groups. Because there are 4160 G-Links within an L3 topology group, a maximum of 4161 L3 topology groups can be connected in a full mesh. Therefore, an L4 topology can be constructed from 4161 L3 topologies, connecting a maximum of 4160 * 4161 = 17,309,760 nodes.
[0144] In this embodiment, if the routing engine is integrated into the processor, 9 ports can meet the networking requirements of 17.3 million GPUs (or computing nodes). By building a direct network, no additional network nodes need to be configured, which can greatly reduce network costs and power consumption. The routing engine integrated in the processor provides 9 external ports, of which 6 ports are used for interconnecting neighbor nodes in the L1 topology, the 7th port connects to the TOR switch in the cabinet to build the L2 topology, the 8th port is used for L3 full-mesh interconnection, and the 9th port is used for L4 full-mesh interconnection. The internal integrated 9*9 low-port Cross Bar can meet internal data forwarding needs.
[0145] It should be noted that Figure 6 、 Figure 7 This example illustrates a 3D torus topology using 6 of the 9 ports, a fat tree topology using 1 port, and a two-layer full mesh topology using the remaining 2 ports. In other possible implementations of the present invention, the routing engine integrated in a low-port router or processor may include other numbers of ports.
[0146] Specifically, each network node includes n ports, m of which are used to connect to network nodes in the surrounding network topology, and the m ports include ports for connecting to upstream nodes in each dimension and ports for connecting to downstream nodes in each dimension. Figure 6 、 Figure 7 An example of constructing a 3D surround network topology structure is given with m=6. In practical applications, the surround network topology structure may also be a 2-dimensional or other dimensional surround network topology structure.
[0147] In some possible implementations, h of the n ports are used to connect to network nodes in a fully interconnected network topology structure formed by a second node group. The second node group includes a second number of first node groups. Given a limited number of ports, a maximum value of the second number can be determined based on the number of ports used to connect to network nodes in the fully interconnected network topology structure formed by the second node group and the number of nodes in the first node group. The number of nodes in the first node group can be determined based on the number of nodes in each dimension of the surrounding network topology structure.
[0148] Based on this, the second quantity can satisfy the following formula:
[0149]
[0150] Among them, count 2nd represents the second quantity, h represents the number of ports used to connect network nodes in the fully interconnected network topology structure formed by the second node group, Len j represents the number of nodes in the jth dimension of the surrounding network topology (also called the length of the jth dimension), and q represents the number of dimensions of the surrounding network topology. represents the number of nodes in the first node group, for example, the first number. h may be greater than or equal to 1, so that different first node groups (or first topology networks) include at least one direct global link.
[0151] For example, when the surround network topology includes three dimensions, namely, X, Y, and Z, the second number can be at most h·Kx·Ky·Kz+1. Figure 6 In the example, Kx, Ky, and Kz can be 4, and h can be 1. Therefore, the maximum second number can be 1*64+1=65. It should be noted that in actual applications, h can also be a positive integer greater than 1, such as 2 or 3. This can increase the number of links between network nodes and reduce the blocking probability, or can connect more first node groups to increase the network scale.
[0152] In some possible implementations, k of the n ports of a network node are used to connect to network nodes in a fully interconnected network topology structure formed by a third node group. The third node group includes a third number of second node groups. Given the limited number of ports, the maximum value of the third number can be determined based on the number of ports used to connect to network nodes in the fully interconnected network topology structure formed by the second node group and the number of nodes in the second node group. The number of nodes in the second node group can be determined based on the second number (the number of the first node group in the second node group) and the first number (the number of nodes in the first node group), for example, the product of the first number and the second number.
[0153] Based on this, the third quantity can satisfy the following formula:
[0154] count 3rd ≤k·count 2nd count 1st +1 (2)
[0155] Among them, count 1st Indicates the first quantity, count 2nd Indicates the second quantity, count 3rd represents the third number, and k represents the number of ports used to connect network nodes in the fully interconnected network topology structure formed by the third node group. 2nd count 1st It can be the number of nodes in the second node group. k can be greater than or equal to 1, so that different second node groups (or second topology networks) include at least one direct global link.
[0156] When the surround network topology includes three dimensions, X, Y, and Z, the second number can be at most h·Kx·Ky·Kz+1, and the third number can be at most k·(h·Kx·Ky·Kz+1)·(Kx·Ky·Kz)+1. Figure 7 In the example, Kx, Ky, and Kz can take the values of 4, and h and k can take the values of 1. Therefore, the third number can be a maximum of 1*(1*64+1)*64+1=4161. Figure 7 Taking h and k as 1 as an example, in actual application, h or k can also be greater than 1, which can increase the links between network nodes and reduce the blocking probability.
[0157] Based on the aforementioned embodiment, the node coordinates of the network node can be determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node is located. For example, when the data routing system is a hierarchical network of nested surround network topology structures, fat tree network topology structures, and one-layer fully interconnected network topology structures, the node coordinates can be represented by the number of the network node in each dimension of the first topology network (such as the surround network topology structure), the port number of the switch to which the network node is connected, and the number of the first node group to which the network node is located. Taking the surround network topology structure as an example of a 3D Torus topology, the node coordinates can be represented as {L, P, X, Y, Z}. Among them, L represents the number of the first node group to which the network node is located, P represents the port number of the switch to which the network node is connected, and X, Y, Z represent the number of the network node in each dimension of the first topology network.
[0158] Furthermore, the data routing system may also include a fourth topology network, which is a fully interconnected network topology structure composed of a third node group, with the third topology network nested within the fourth topology network. Accordingly, the node coordinates of a network node can be determined based on the network node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs and the number of the second node group to which the network node belongs. For example, when the data routing system is a hierarchical network with nested wrap-around network topologies, fat-tree network topologies, and two-layer fully interconnected network topologies, the node coordinates can be represented by the network node's number within each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs and the number of the second node group to which the network node belongs. For example, when the wrap-around network topology is a 3D torus topology, the node coordinates can be represented as {G, L, P, X, Y, Z}. Here, G represents the number of the second node group to which the network node belongs, L represents the number of the first node group to which the network node belongs, P represents the port number of the switch to which the network node is connected, and X, Y, and Z represent the number of the network node within each dimension of the first topology network.
[0159] It should be noted that a unified numbering scheme can be used for switching nodes such as integrated routing engines in processors or independent low-port routers and switches. Accordingly, network nodes can determine data routing paths based on the node coordinates represented by the numbers, and implement data routing between computing nodes based on the data routing paths.
[0160] Among them, network nodes can determine data routing paths using the shortest path routing algorithm based on node coordinates (such as node numbers). The shortest path routing algorithm can provide the shortest distance communication between the source node and the destination node, with the lowest communication latency, and is the most basic routing algorithm. Because this application is a hierarchical network nested with Torus, fat tree, and full mesh (for example, a two-layer full mesh), a deterministic shortest path routing algorithm that adapts to topological characteristics has the advantages of rapid calculation and is conducive to hardware implementation.
[0161] The following describes the specific implementation of data routing using a deterministic shortest path routing algorithm based on adaptive topology features.
[0162] See also Figure 8 The flowchart of a data routing method is shown, which can be applied to the data routing system 10 of the aforementioned embodiment. The data routing system 10 is used to implement data routing between computing nodes. The data routing system 10 includes a first topology network 120, a second topology network 130, and a third topology network 140. The network structure of the first topology network 120 is a ring network topology consisting of a first node group. The first node group includes a first number of network nodes, such as nodes 110-1 through 110-P. The network nodes in the ring network topology are connected to upstream nodes and downstream nodes in each dimension. The network structure of the second topology network 130 is a fat tree network topology consisting of the network nodes in the first topology network 120 and switches 132. The network nodes and switches in the first topology network are in a cascade relationship. The first topology network 120 is nested within the third topology network 140. The network structure of the third topology network 140 is a fully interconnected network topology consisting of a second node group. The second node group includes a second number of first node groups. Connections are established between the second number of first node groups. The data routing system 10 may pass through at least one network node when routing a data packet from a source node to a destination node. The node where the data packet is currently located is called the current node. The current node may be the source node of the data packet (the source node may also be a network node) or a node on the path from the source node to the destination node. The method includes the following steps:
[0163] S802: The current node receives a data packet to be routed.
[0164] The current node is the source node of the data packet or a node on the path from the source node to the destination node. The packet header may record information about the source node and the destination node of the data packet, such as the node coordinates of the source node, the node coordinates of the destination node, or other information that can be used to determine the node coordinates of the source node and the destination node.
[0165] S804: The current node determines the data routing path according to the node coordinates of the destination node.
[0166] The node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node is located. Furthermore, when data routing topology 10 includes a fourth topology network, the node coordinates may be determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, the number of the first node group to which the network node is located, and the number of the second node group to which the network node is located.
[0167] The current node can parse the data packet to obtain the node coordinates of the destination node. In some examples, the node coordinates of the destination node can be {Gd, Ld, Pd, Xd, Yd, Zd}, where Gd represents the number of the second node group in which the destination node is located, Ld represents the number of the first node group in which the destination node is located, Pd represents the port number of the switch to which the destination node is connected, and Xd, Yd, and Zd represent the numbers of the destination node in each dimension of the first topology network.
[0168] The current node can determine the data routing path based on the node coordinates of the destination node, combined with the topology of the data routing system 10 (for example, a topology map). It should be noted that when the node coordinates of the current node and the destination node represent that the current node and the destination node are in the same layer, the data routing path can be determined based on the topology of the layer. For example, when the current node and the destination node are both in the same first topology network (or the same first node group) of the first layer, the current node can determine the data routing path based on the topology of the layer, for example, the topology of the corresponding node group in the layer. For another example, when the current node and the destination node are in different node groups of the same layer, the current node can determine the data routing path based on the topology of the layer.
[0169] Based on this, the current node can first identify whether the node groups (such as the first node group and the second node group) where the current node and the destination node are located are the same based on the node coordinates of the current node and the node coordinates of the destination node, and then determine the data routing path based on the distribution of the node groups where the current node and the destination node are located. In some cases, the current node can also determine the data routing path in combination with the node coordinates of the source node. For example, the current node can determine the traffic type based on the node coordinates of the destination node and the node coordinates of the source node, and the traffic type is used to identify the data packet as intra-group traffic or cross-group traffic. Among them, intra-group traffic can be internal traffic within the cabinet, and cross-group traffic can be external traffic across cabinets. The current node can determine the corresponding data routing path based on the traffic type.
[0170] S806: The current node routes the data packet according to the data routing path.
[0171] Specifically, the data routing path may indicate a routing port or the node coordinates of a next-hop node. Accordingly, the current node may route the data packet to the next-hop node according to the routing port indicated by the data routing path or the node coordinates of the next-hop node.
[0172] Based on the above description, the present application provides a data routing method, which introduces the node coordinates of the destination node in the data routing system, adapts the topological characteristics of the data routing system to perform data routing, provides the shortest distance communication between the source node and the destination node, and achieves low communication latency. Moreover, the method is simple to implement and has high availability.
[0173] Next, Figure 8 In the embodiment, a specific implementation of determining a data routing path when the current node and the destination node are in node groups at different levels is described.
[0174] Depending on the location of the source and destination nodes in the network, the source node can route data packets to the destination node directly, or it can first route to the next-hop network node and then continue routing from the next-hop network node until it reaches the destination node. When the data is at the source node, the source node can use the shortest path routing algorithm to determine the next-hop node. If the next-hop node is not the destination node, the next-hop node can use the shortest path routing algorithm to determine the next-hop node until it reaches the destination node.
[0175] For ease of description, this application refers to the node where data arrives for processing as the current node. The current node can be the source node or a node on the path from the source node to the destination node. Whether it is the source node or a node on the path from the source node to the destination node, the method of determining the data routing path (next hop node) through the shortest path routing algorithm is similar. The specific implementation of determining the data routing path is detailed below.
[0176] Considering that routing methods are usually quite different when the distribution of the current node and the destination node in the network is different, this application supports identifying the distribution of the current node and the destination node, and then routing data according to the distribution.
[0177] Specifically, the current node can identify whether the current node and the destination node are in the same node group based on the node coordinates of the current node and the node coordinates of the destination node. For example, if the number of the first node group in which the current node is located is the same as the number of the first node group in which the destination node is located, it means that the current node and the destination node are in the same first node group. Similarly, if the number of the second node group in which the current node is located is the same as the number of the second node group in which the destination node is located, it means that the current node and the destination node are in the same second node group. Based on this, data routing can be divided into intra-node group routing and inter-node group routing between different node groups.
[0178] In the first implementation, when the number of the first node group in which the current node is located is equal to the number of the first node group in which the destination node is located, it indicates that the current node and the destination node are in the same first node group (L1 group), and the data routing is intra-group routing within the first node group. When the current node and the destination node are in the same first node group, they must also be in the same second node group, and the number of the second node group in which the current first node is located is equal to the number of the second node group in which the destination node is located. In this case, the current node can determine the routing port within each dimension based on the offset value of the number of the destination node and the current node within each dimension of the first topology network.
[0179] The number of network nodes surrounding the i-th dimension of the network topology can be L i The current node determines the routing port in the i-th dimension according to the offset value of the number of the destination node and the current node in the i-th dimension of the surrounding network topology structure, which may include the following situations:
[0180] In the first case, when the offset value of the number between the destination node and the current node in the i-th dimension of the surrounding network topology is greater than 0 and less than or equal to L i / 2, indicating that the destination node is in the positive direction, or the offset value is less than -L i / 2 indicates that the loopback link is the shortest path. Based on this, the current node can determine that the routing port in the i-th dimension is the port used to connect to the downstream node, also known as the forward routing port.
[0181] In the second case, when the offset value of the number between the destination node and the current node in the i-th dimension of the surrounding network topology is greater than L i / 2 indicates that the loopback link is the shortest path, or the offset value is greater than or equal to -L i / 2 and less than 0, indicating that the destination node is in the negative direction, and the routing port in the i-th dimension is determined to be the port used to connect to the upstream node, also known as the negative routing port.
[0182] The current node may traverse each dimension and perform the above routing operation for each dimension until reaching the destination node.
[0183] It should be noted that the above description uses an example where one first node group corresponds to one cabinet, such as one rack. When the current node (eg, source node) and the destination node are located in different cabinets, the current node can also route data to the destination node through a switch.
[0184] If the data packet is part of all-reduce traffic, the current node can perform intra-group routing within the first node group. For example, in AI computing scenarios, traffic for training AI models in parallel using data parallelism or model parallelism can be all-reduce traffic. The current node can route this all-reduce traffic using intra-group routing within the first node group.
[0185] The second implementation method is that the current node determines the traffic type based on the node coordinates of the destination node and the node coordinates of the source node, and the traffic type is used to identify the data packet as intra-group traffic or cross-group traffic. Intra-group traffic can be internal traffic within a cabinet, and cross-group traffic can be external traffic across cabinets. For example, the number of the first node group where the source node is located is equal to the number of the first node group where the destination node is located, which is represented as intra-group traffic, such as local internal traffic within a cabinet, and can be communicated based on the aforementioned Torus topology. Otherwise, the traffic is cross-group traffic, such as external traffic across cabinets, which can be jumped through the switch in the fat tree network topology. In particular, when the traffic type identifies that the data packet is cross-group traffic, the current node can determine the forwarding port based on the node coordinates of the destination node.
[0186] For cross-group traffic, it can be further divided into traffic sent from local to external or traffic sent from external to local. Among them, the traffic sent from local to external can be cross-cabinet traffic from inside to outside, and the traffic sent from external to local can be cross-cabinet traffic from outside to inside. The current node can forward data within the fat tree based on the shortest routing path algorithm for the above cross-group traffic. Among them, for cross-cabinet traffic to the outside, it can be forwarded to the local node inside the cabinet that has a direct link to the destination node; for cross-cabinet traffic sent from the outside to the local, the data routing path (also called forwarding path) can be determined according to the node coordinates of the destination node to jump within the first node group.
[0187] When the cross-group traffic is traffic sent locally to the outside, such as cross-cabinet traffic from the inside to the outside, the current node can determine the node coordinates of the network node (such as the jump node) in the group that has a direct link to the destination node based on the node coordinates of the destination node, or the node coordinates of the network node (such as the ferry node) in the group that has a direct link to the jump node to the destination node. The current node can determine the port number of the network node in the group connected to the switch based on the node coordinates of the network node in the group. Among them, the node coordinates of the jump node include the number of the jump node in the dimension of the surrounding network topology, and the node coordinates of the ferry node with a direct link to the jump node include the number of the ferry node in the dimension of the surrounding network topology. The current node can determine the port number of the network node in the group such as the jump node or the ferry node connected to the switch based on the number of the jump node in the dimension of the surrounding network topology or the number of the ferry node in the dimension of the surrounding network topology, such as the output port number of the switch.
[0188] If the destination node and the local node are in the same second node group, and the local node has a jump node with a direct link to the destination node, the current node can use the jump node's number within the dimension surrounding the network topology to determine the switch's outbound port number. If the destination node and the local node are in different second node groups, the node coordinates of an internal local node with a direct link to the jump node can be used to determine the port number. The internal local node with a direct link to the jump node is a ferry node.
[0189] After determining the node coordinates of the jump node or the ferry node, the port number can be calculated using the following formula:
[0190] P=X*Ky*Kz +Y*Kz+ Z+1 (3)
[0191] Among them, P represents the port number, Ky and Kz represent the lengths of different dimensions of the first topology network, X, Y, and Z represent the numbers of jump nodes with direct links to the destination node in the dimension of the surrounding network topology, and the numbers of ferry nodes with direct links to the jump nodes in the dimension of the surrounding network topology.
[0192] When cross-group traffic is traffic sent from the outside to the local, such as cross-cabinet traffic from the outside to the inside, for example, the number of the first node group where the source node is located is not equal to the number of the first node group where the destination node is located, but the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, it means that the data has flowed to the cabinet where the destination node is located, and the current node can determine the forwarding port based on the node coordinates of the destination node and the cascade relationship between the destination node and the switch. Among them, the node coordinates of the destination node include the number of the destination node in the dimension of the surrounding network topology, and the current node can determine the output port number of the switch based on the number of the destination node in the dimension of the surrounding network topology. Among them, the calculation of the output port number can refer to the above formula (3), and the difference from the above formula (3) is only that X, Y, and Z represent the dimension number of the destination node, for example, Xd, Yd, and Zd.
[0193] For all-to-all or point-to-point (send / receive) traffic, the current node can route the data packet through the fat-tree network topology, for example, through a switch in the fat-tree network topology. For send / receive traffic, if the current node is not directly connected to a node in the node group (the first node group or the second node group) where the destination node resides, the data packet can first be redirected through a switch in the fat-tree network topology and then routed through the full mesh network.
[0194] The third implementation method is that when the number of the second node group where the current node is located is equal to the number of the second node group where the destination node is located, and the number of the first node group where the current node is located is not equal to the number of the first node group where the destination node is located, it means that the current node and the destination node are in different first node groups of the same second node group, and the data routing is the intra-group routing of the second node group (inter-group routing of the first node group).
[0195] Accordingly, the current node can determine whether there is a direct link (such as L-LINK) between the current node and the first node group where the destination node is located. In order to distinguish it from the direct link (such as G-LINK) or other direct links of the second node group where the destination node is located, this application may refer to the direct link with the first node group where the destination node is located as the first direct link.
[0196] If so, the current node can determine that the data routing path includes forwarding the data packet via the first direct link to the first node group where the destination node is located. If not, the current node can determine that the data routing path includes forwarding the data packet to a second jump node. The second jump node has a second direct link with the first node group where the destination node is located. When the data packet arrives at the second jump node, it can reach the first node group where the destination node is located via the second direct link. If the data packet reaches the first node group where the destination node is located, intra-group routing within the first node group can be performed.
[0197] The fourth implementation method is that when the number of the second node group where the current node is located is not equal to the number of the second node group where the destination node is located, it means that the current node is in a different second node group, and the data routing is the intra-group routing of the third node group (inter-group routing of the second node group).
[0198] Accordingly, the current node can determine whether a third direct link exists between the current node and the second node group where the destination node is located. If so, the current node can determine that the data routing path includes forwarding the data packet along the third direct link to the second node group where the destination node is located. When the data packet arrives at the second node group where the destination node is located, intra-group routing within the second node group can be performed. If not, the current node can determine that the data routing path includes forwarding the data packet to the first jump node. A fourth direct link exists between the first jump node and the second node group where the destination node is located.
[0199] The node coordinates of the first jump node can be determined by the following formula:
[0200]
[0201] Among them, Xm, Ym, and Zm are the numbers of the first jump node in each dimension of the first topology network, PLm represents the port number of the second node group where the first jump node forwards data to the destination node, Ky and Kz represent the lengths of different dimensions of the first topology network, and Xd, Yd, and Zd are the numbers of the destination node in each dimension of the surrounding network topology.
[0202] Furthermore, the current node can also determine whether the number of the first node group where the first jump node is located is equal to the number of the first node group where the current node is located. If not, it means that the current node and the first jump node are in different first node groups. The current node determines that the data routing path also includes forwarding the data packet to the ferry node. Among them, there is a fifth direct link between the ferry node and the first node group where the first jump node is located. That is, the current node can first route the data packet to the ferry node. The ferry node and the above-mentioned first jump node are in the same first node group, and the data packet can be routed to the above-mentioned first jump node through the intra-group routing of the first node group. There is a direct link between the first jump node and the second node group where the destination node is located, such as the above-mentioned fourth direct link. When the data packet reaches the first jump node, it can be routed to the second node group where the destination node is located through the above-mentioned fourth direct link. When the data packet reaches the first node group where the destination node is located, intra-group routing of the second node group can be performed.
[0203] If the data packet is a data packet in point-to-point traffic, the current node may perform intra-group routing within the second node group (inter-group routing within the first node group) or intra-group routing within the third node group (inter-group routing within the second node group), for example, routing the point-to-point traffic via a full mesh network (the third topology network or the fourth topology network). The point-to-point traffic may include send / receive traffic in a pipelined parallel mode.
[0204] For the convenience of description, this application uses Figure 7 The topology example is shown in the following figure. Figure 7 In the topology, the node coordinates of the network node can be {G, L, P, X, Y, Z}, where G represents the number of the second node group where the network node is located, ranging from 0 to (Kx*Ky*Kz*(Kx*Ky*Kz+1)), where Kx, Ky, and Kz are 4, G can be 0 to 4160 (including endpoint values), and L represents the number of the first node group where the network node is located, ranging from 0 to Kx*Ky*Kz, where Kx, Ky, and Kz are 4, L can be 0 to 64 (including endpoint values).
[0205] P is the port number of the network node (such as a GPU card with routing function) in the L2 topology connected to the TOR switch in the cabinet (P = X*Ky*Kz+Y*Kz+Z+1); X, Y, Z are the numbers of the network nodes in the X, Y, and Z dimensions within the L1 Torus, ranging from 0 to Kx-1, 0 to Ky-1, and 0 to Kz-1, respectively. The node number corresponds to the topological position of the node.
[0206] Ports 1 and 2 of the routing engine serve as routing ports in the forward and reverse directions of the X dimension (port 1: LinkX+; port 2: LinkX-).
[0207] Ports 3 and 4 of the routing engine serve as routing ports in the forward and reverse directions of the Y dimension (port 3: LinkY+; port 4: LinkY-).
[0208] Ports 5 and 6 of the routing engine serve as routing ports in the forward and reverse directions of the Z dimension (port 5: LinkZ+; port 6: LinkZ-).
[0209] The positive direction interfaces of adjacent nodes in each dimension are connected to the negative direction ports of the opposite end, and the connections are made in sequence to form a Torus ring network, thus constructing a L1 level topology 3D Torus.
[0210] Port 7 of the routing engine is connected to the TOR switch in the cabinet. As a Layer 2 topology, the cabinets are interconnected through the switch, and the GPUs in the cabinet can be reached in one hop without blocking.
[0211] Port 8 of the routing engine serves as the interconnection port between each L1 torus in the L3 topology group and is defined as an L-Link. The L-link number of each routing engine corresponds to the torus coordinate one-to-one, PL = X*Ky*Kz+Y*Kz+Z, and the PL value range is: 0 to Kx*Ky*Kz-1.
[0212] The following describes the global link connection relationship within the L3 topology group.
[0213] The source node is numbered {Gs, Ls, Ps, Xs, Ys, Zs}, the destination node is numbered {Gd, Ld, Pd, Xd, Yd, Zd}, and the current node is numbered {Gc, Lc, Pc, Xc, Yc, Zc}. L-Link port 8 of the source node is connected to L-Link port 8 of the destination node. The full mesh connection rule for the L3 topology is: Ls + PLs = Ld && PLd + PLs = Kx * Ky * Kz - 1.
[0214] Among them, Ls is the number of the first node group where the source node is located, Ld is the number of the first node group where the destination node is located, PLs is the number of the L-Link port connecting the source node to the first node group where the destination node is located, and PLd is the number of the L-Link port connecting the destination node to the first node group where the source node is located. Based on this connection relationship, the routing relationship between the first node groups can be calculated.
[0215] The global link connection relationship within the L4 topology group is as follows:
[0216] Port 9 of the Routing Engine serves as the interconnection port between L3 full mesh topologies within the L4 topology and is defined as a G-Link. The G-Link number of each Routing Engine corresponds to the coordinates of each L3 full mesh and torus within the group. The G-Link numbering rule is: PG = L*Kx*Ky*Kz+X*Ky*Kz+Y*Kz+Z, with the PG value range from 0 to Kl*Kx*Ky*Kz-1. The full mesh connection relationship between L4 topologies is: Gs+PGs = Gd &&PGd+PGs = Kl*Kx*Ky*Kz-1.
[0217] Among them, Gs is the L4 topology number of the source node, for example, the number of the second node group where the source node is located, Gd is the L4 topology number of the destination node, for example, the number of the second node group where the destination node is located, PGs is the G-Link port number of the source node connecting to the second node group where the destination node is located, and PGd is the G-Link port number of the destination node connecting to the second node group where the source node is located. Based on this connection relationship, the routing relationship between the second node groups can be calculated.
[0218] See also Figure 9 The diagram shows a flow chart of data routing in a data routing system. The current node can first determine whether the destination node and the current node are in the same second node group. If not, inter-group routing within the second node group is performed. If so, inter-group routing within the first node group is performed. When the destination node and the current node are not in the same first node group, inter-group routing within the first node group is performed. When the destination node and the current node are in the same first node group, the current node can also determine the traffic type of the data packet based on the node coordinates of the destination node and the node coordinates of the source node. When the traffic type is local cabinet internal traffic, the current node can route the data packet using the L1 torus topology; when the traffic type is cross-cabinet external traffic, the current node can route the data packet using the L2 fat tree topology. When the number of the second node group in which the destination node is located is equal to the number of the second node group in which the source node is located, and the number of the first node group in which the destination node is located is equal to the number of the first node group in which the source node is located, the traffic type is local cabinet internal traffic; otherwise, the above traffic is cross-cabinet external traffic.
[0219] The specific implementations of inter-group routing of the second node group, inter-group routing of the first node group, fat tree topology routing, and torus topology routing are described below.
[0220] If the destination node and the current node are in different second node groups, that is, Gd = !Gc, define the first jump node with a direct link to the second node group where the destination node {Gd, Ld, Xd, Yd, Zd} is located. The node coordinates of the first jump node can be determined by the following formula:
[0221] Gm=Gd-PGm&&PGm=(Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)-PGd, where
[0222] Gm=Gc=Gd-(Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)+Ld*Kx*Ky*Kz+Xd*Ky*Kz+Yd*Kz+Zd && (5)
[0223] PGm=(Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)-Ld*Kx*Ky*Kz-Xd*Ky*Kz-Yd*Kz-Zd.
[0224] If the current node has a direct link to the second node group (denoted as Gd) where the destination node is located, the following formula is satisfied:
[0225]
[0226] Then the data can be forwarded directly from the G-Link port of the current node to the second node group where the destination node is located. Otherwise, it is necessary to first route from the current node to the first jump node. The node coordinates of the first jump node are {Gc, Lm, Xm, Ym, Zm}, where:
[0227]
[0228] If the first node group where the first jump node is located is different from the first node group where the current node is located, that is, Lm=!Lc, it must first be routed to the first node group where the first jump node is located (denoted as Lm). The nodes with direct links to the first node group Lm where the first jump node is located are defined as ferry nodes {Gn, Ln, Xn, Yn, Zn}. Therefore, the data packet can be routed to the ferry node first, and then forwarded from the direct link L-Link to the first node group where the first jump node is located. The physical location of the ferry node can be obtained according to the following formula:
[0229]
[0230] The current node can first route the data packet to the ferry node through the L1 level Torus, where the current node can route according to the coordinate address offset in the three dimensions of X, Y, and Z. After arriving at the ferry node, the data packet can be routed from the direct link L-Link port to the first node group where the first jump node is located. After the data packet arrives at the first node group where the first jump node is located, it can reach the first jump node through the intra-group routing of the first node group. The first jump node can reach the second node group where the destination node is located through the direct link. The data packet arrives at the second node group where the destination node is located and can reach the destination node through the intra-group routing of the second node group.
[0231] The following describes in detail the intra-group routing of the second node group (eg, the inter-group routing of the first node group).
[0232] The current node can determine whether the current node and the destination node are in the same second node group based on the node coordinates. If the destination node and the current node are in the same second node group, that is, Gd = Gc, but Ld = !Lc, the current node can determine whether there is a direct link between the current node and the first node group where the destination node is located based on the topological connection relationship.
[0233] If the current node is a direct link to the first node group where the destination node is located, data can be forwarded directly from the direct link. For details, see the following formula:
[0234]
[0235] If formula (9) is satisfied, the destination node can be directly forwarded from the L-Link port to the first node group. Otherwise, the node with a direct link to the destination node in the first node group is defined as the second jump node. The node coordinates of the second jump node can be:
[0236]
[0237] Therefore, first route to the second jump node: Xm = PLm / Ky*Kz; Ym = PLm / Kz; Zm = PLm%Kz, and then turn to the internal routing of the first node group. Specifically, routing is performed in the three dimensions X, Y, and Z according to the coordinate address offset. After arriving at the second jump node, forwarding from the port PLm = (Kx*Ky*Kz-1)-PLd = (Kx*Ky*Kz-1)-Xd*Ky*Kz-Yd*Kz-Zd can be carried out from the direct link to the first node group where the destination node is located. Then, intra-group routing of the first node group can be performed, such as dimension-ordered routing within the same cabinet. Dimension-ordered routing means that each data packet is routed only on one dimension at a time. After reaching the appropriate coordinates on this dimension, it is routed on other dimensions in the order from low dimension to high dimension.
[0238] Among them, according to the source node numbers {Gs, Ls, Ps, Xs, Ys, Zs} and the destination node numbers {Gd, Ld, Pd, Xd, Yd, Zd}, it can be determined whether the data packet is internal traffic within the local cabinet or external traffic across cabinets: If Gs == Gd and Ls == Ld, it means it is local traffic and will communicate through the Torus topology. Otherwise, it means it is external traffic across cabinets and needs to be redirected through the TOR switch.
[0239] For the cross-cabinet traffic sent from local to external, the local node numbers within the cabinet that have direct links with it can be calculated according to the destination node label (if the destination node and the local node are not in the same first node group, it corresponds to the internal local label that has a direct link with the intermediate redirection node). The port number rule for this node to the TOR switch is determined by P = X * Ky * Kz + Y * Kz + Z + 1.
[0240] For the traffic sent from external to local, that is, Gs!= Gd or Ls!= Ld, but Gc == Gd and Lc == Ld, it means the data has been transferred to the cabinet where the destination node is located. Then, the forwarding port can be determined according to the coordinates of each dimension of the Torus of the destination node: P = Xd * Ky * Kz + Yd * Kz + Zd + 1.
[0241] The intra-group routing of the first node group is described below.
[0242] The destination node and the current node are within the same L4 and L3 topologies, that is, Gd == Gc && Ld == Lc. Therefore, only the coordinate offsets of the three dimensions of the Torus where the destination node is located need to be used to determine the routing:
[0243] First, judge the coordinate offset in the X dimension. Compare the coordinate deviation of the destination node in the X dimension with the current node. Define the coordinate deviation DeltX == Xd - Xc. If 0 < DeltX <= Kx / 2, it means the destination node is in the positive direction and route through the X+ port; if DeltX > Kx / 2, it means taking the loopback link is the shortest path, then route through the X- port; if 0 > DeltX >= -Kx / 2, it means the destination node is in the negative direction and route through the X- port; if DeltX < -Kx / 2, it means taking the loopback link is the shortest path, then route through the X+ port.
[0244] Then, judge the coordinate offset in the Y dimension. Define the coordinate deviation DeltY == Yd - Yc. If 0 < DeltY <= Ky / 2 or DeltY < -Ky / 2, route through the Y+ positive port; if DeltY > Ky / 2 or 0 > DeltY >= -Ky / 2, route through the Y- port.
[0245] Next, determine the Z - dimension coordinate offset. Define the coordinate deviation DeltZ == Zd - Zc. If 0 < DeltZ <= Kz / 2 or DeltZ < -Kz / 2, route from the Z+ positive port; if DeltZ > Kz / 2 or 0 > DeltZ >= -Kz / 2, route from the Z - port.
[0246] Finally, there is no offset in the Z - dimension, reaching the destination node. The DMA communication engine receives the packet to complete the communication.
[0247] Considering that both the Torus topology and the full mesh topology have loops, a deadlock risk can be generated therefrom. Here, deadlock refers to the circular occupation of communication resources (such as cache buffer). For example, for the 4 nodes in the X - dimension of the Torus topology, each node sends data to the adjacent next node. Node 0 sends to node 2, node 1 sends to node 3, and so on. If the 4 nodes send data simultaneously, the port resources will be circularly occupied, which will cause deadlock and communication congestion, and in severe cases, even lead to network paralysis. Based on this, the data routing system of this application can also prevent deadlocks through virtual lanes (VL).
[0248] Based on the fact that virtual lanes can effectively avoid deadlocks, in the data routing system of this application, since the first topology network (such as the L1 - level topology) adopts the standard Torus architecture and the third topology network adopts the full mesh architecture, and the standard Torus and Full mesh topologies naturally have loops, therefore, multiple virtual lanes can be allocated to the ports of the first topology network (such as the first node group) and the third topology network (such as the third node group) to avoid deadlocks. Further, when the data routing system also includes a fourth topology network, multiple virtual lanes can also be allocated to the ports of the fourth topology network to avoid deadlocks.
[0249] Specifically, the port corresponding to the first topology network is configured with a first virtual lane and a second virtual lane, the port corresponding to the third topology network is configured with a third virtual lane and a fourth virtual lane, and the port corresponding to the fourth topology network is configured with a fifth virtual lane and a sixth virtual lane.
[0250] For the global traffic of the first topology network, the second virtual lane can be used for routing the global traffic. Among them, for the first global traffic of the first topology network, the first virtual lane can be used for routing. For the global traffic of the third topology network, the fourth virtual lane can be used for routing the global traffic. Among them, for the first global traffic of the third topology network, the third virtual lane can be used for routing the global traffic. For the global traffic of the fourth topology network, the sixth virtual lane can be used for routing the global traffic. Among them, for the first global traffic in the fourth topology network, the fifth virtual lane can be used for routing the global traffic. For the sake of easy understanding, Figure 10 A virtual channel allocation scheme is provided as an example. In this example, a port of a network node is configured with multiple virtual channels. In this way, no loop is formed between the virtual channels, and therefore no loop deadlock occurs between the layers.
[0251] The above example uses a large AI model with hundreds of billions of parameters. With the continuous increase in computing power, more and more developers are choosing to use ultra-large-scale clusters to train large AI models with trillions of parameters. Trillion-parameter AI models utilize hybrid expert parallelism, in addition to model parallelism, data parallelism, and pipeline parallelism, to improve computational efficiency. Hybrid expert parallelism primarily involves global communication, with all-to-all traffic predominantly accounting for 50%.
[0252] The communication latency of a torus topology increases rapidly with network size. Larger networks experience more severe congestion, especially for all-to-all, globally uniform random traffic, where network throughput is very low. Therefore, a torus topology struggles to meet the communication needs of all-to-all traffic. Fat-tree networks are non-blocking and highly suitable for all-to-all traffic, offering 100% throughput.
[0253] The scale of expert parallelism and the size of supernodes affect the performance of AI applications. Given the same cluster size, larger supernodes result in better application performance. For GPT-4, for example, the performance of a 1024-supernode is 15% higher than that of a 64-supernode. However, due to hardware constraints, larger supernodes require connecting multiple cabinets to form a larger supernode.
[0254] The following example illustrates the construction of a supernode using 1024 computing cards (for example, GPUs). A standard cabinet can usually accommodate 64 GPUs, and an 8*8 2D Torus topology can be constructed. A 3D Torus can be constructed between 16 cabinets, and 16*8*8=1024 GPUs can be connected. Global traffic across cabinets (for example, global jump traffic) can be transmitted through a two-layer fat tree. The RACK is a 2DTorus, which includes a total of 8*8=64 GPUs. A 3D Torus can be constructed between 16 RACKs in a supernode, which includes a total of 16*8*8=1024 GPUs. The single-level full mesh network can reach a scale of 1 million GPUs: 1024*1025=10496001972, such as Figure 11 It should be noted that Figure 11 The "A" in the code represents a GPU. Figure 11 The "K" in the figure represents the CPU, where every 4 CPUs can be connected to 8 GPUs. Figure 11The SW in the figure represents a cabinet switch. The TOR switch at the top of the cabinet can be a leaf switch, and SW01 to SW64 that aggregate the leaf switches can be spine switches. The spine switch in this application can be a 128-port switch.
[0255] The following networking example uses a 12-port general-purpose GPU. A 12-port general-purpose GPU has internal IO Die routing and forwarding capabilities. This GPU has an integrated or built-in routing engine, providing routing and forwarding capabilities. A single-port general-purpose GPU can have a port bandwidth of 200 Gbps.
[0256] When implementing it specifically, Figure 12 As shown, a Hybrid Torus topology can be constructed based on a 12-port general-purpose GPU. Specifically, bandwidth resources are allocated based on the communication type and characteristics of large AI models. Six 200Gbps ports are used for torus interconnection. Within a rack (RACK), four ports (two ports in each dimension) are used to construct a 2D torus consisting of 64 GPUs (arranged in an 8x8 configuration). A 3D torus consisting of 1024 GPUs (arranged in a 16x8x8 configuration) is constructed across 16 racks.
[0257] Expert parallelism accounts for approximately 50% of the communication volume. Therefore, five 200Gbps bandwidths can be allocated to connect the TOR switches within the cabinet. The cabinets are connected through the spine switches to form a two-layer fat tree, which can provide all-to-all traffic and all-reduce traffic communication between 1024 cards in the super node.
[0258] Since the communication between supernodes is mainly pipelined and parallel, and the flow of pipelined and parallel traffic is send / receive traffic with low data volume, one 200Gbps port can be allocated for full-mesh direct connection between supernodes, which can effectively reduce network costs and energy consumption while meeting communication needs.
[0259] Table 2 also shows the bandwidth allocation results, as follows:
[0260] Table 2 Bandwidth allocation results
[0261]
[0262] The bandwidth per node within a RACK can include 1200G of Torus network bandwidth, 1000G of fat-tree network bandwidth, and 1024x200G of full-mesh network bandwidth. The Torus network only transmits all-reduce traffic, fully utilizing bandwidth resources in all dimensions. The fat-tree network transmits all-to-all and all-reduce traffic within supernodes, as well as internal hops for global traffic across supernodes. The full-mesh network transmits send / receive traffic between RACKs (e.g., between supernodes).
[0263] It should be noted that in actual applications, the network scale can also be adjusted according to traffic characteristics.
[0264] The following example illustrates data routing based on the 12-port general-purpose GPU network described above.
[0265] See also Figure 13 The diagram shows a flow chart of data routing in a data routing system. The current node can first determine whether the destination node and the current node are in the same first node group. If not, inter-group routing within the first node group is performed. If so, the traffic type can be identified. If the traffic type is local cabinet internal traffic, intra-group routing within the first node group can be performed, such as dimension-order routing within the same cabinet. If the traffic type is external traffic across cabinets, routing can be performed using a fat-tree network topology. Furthermore, if a 12-port general-purpose GPU is used with nested torus, fat trees, and two-level full meshes, it can first be determined whether the destination node and the current node are in the same second node group. If they are in different second node groups, inter-group routing within the second node group is performed.
[0266] The specific implementations of inter-group routing of the second node group, inter-group routing of the first node group, fat tree topology routing, and intra-group routing of the first node group are described below.
[0267] The current node first determines whether the destination node and the current node are in the same second node group. If not, that is, Gd=!Gc, then inter-group routing of the second node group can be performed. Specifically, if the current node has a direct link to the second node group Gd where the destination node is located, that is, PGc=Lc*Kx*Ky*Kz+Xc*Ky*Kz+Yc*Kz+Zc=(Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)-Ld*Kx*Ky*Kz-Xd*Ky*Kz-Yd*Kz-Zd. Then the data can be forwarded directly from the G-Link port of the current node to the second node group where the destination node is located. Otherwise, it is necessary to first route from the current node to the first jump node. The location of the first jump node is numbered {Gc, Lm, Xm, Ym, Zm}, where PGm = (Kx*Ky*Kz*(Kx*Ky*Kz+1)-1)-Kx*Ky*Kz*Ld-Xd*Ky*Kz-Yd*Kz-Zd, Lm = PGm / Kx*Ky*Kz, Xm = PGm / Ky*Kz, Ym = PGm / Kz, Zm = PGm%Kz.
[0268] If the first node group of the first jump node is different from the first node group of the current node, that is, Lm = !Lc, routing must first be done to the first node group Lm of the first jump node. The ferry node is defined as a node with a direct link to the first node group Lm of the first jump node, {Gn, Ln, Xn, Yn, Zn}. Therefore, the data packet must first be routed to the ferry node and then forwarded via the direct link L-Link to the first node group of the first jump node. L2 intra-cabinet routing can then be used to route the data packet to the first jump node via a fat-tree network topology. The physical location of the ferry node can be determined using the following formulas: Gn = Gc, Ln = Lc, PLn = (Kx*Ky*Kz-1)-Xm*Ky*Kz-Ym*Kz-Zm. The coordinates of the ferry node within the Torus are: Xn = PLn / Ky*Kz; Yn = PLn / Kz; Zn = PLn%Kz.
[0269] When the destination node and the current node are in the same second node group, that is, Gd=Gc, the current node can continue to determine whether they are in the same first node group. If not, that is, Ld=!Lc, then the inter-group routing of the first node group is executed. If the current node is a direct link connecting the first node group where the destination node is located, the data can be forwarded directly from the direct link. Among them, Lc=Ld-PLc&&PLc=(Kx*Ky*Kz-1)-PLd, that is, Lc=Ld-(Kx*Ky*Kz-1)+Xd*Ky*Kz+Yd*Kz+Zd=Xc*Ky*Kz+Yc*Kz+Zc, then directly forwarding from the L-Link port can reach the first node group where the destination node is located. Otherwise, define the node in the first node group with a direct link to the destination as the second hop node. Its node coordinates are: Gm = Gc, Lm = Ld - PLm && PLm = (Kx*Ky*Kz-1) - PLd, PLd = Xd*Ky*Kz + Yd*Kz + Zd. That is, Gm = Gc, Lm = Ld - (Kx*Ky*Kz-1) + Xd*Ky*Kz + Yd*Kz + Zd && PLm = (Kx*Ky*Kz-1) - Xd*Ky*Kz - Yd*Kz - Zd. Therefore, first route to the second hop node: Xm = PLm / Ky*Kz; Ym = PLm / Kz; Zm = PLm%Kz, then switch to routing within the first node group, routing in the X, Y, and Z dimensions based on the coordinate address offset. After reaching the second hop node, forwarding is performed from port PLm=(Kx*Ky*Kz-1)-PLd=(Kx*Ky*Kz-1)-Xd*Ky*Kz-Yd*Kz-Zd, and the first node group where the destination node is located can be reached from the direct link.
[0270] If the destination node and the current node are in the same first node group, the source node number {Gs, Ls, Ps, Xs, Ys, Zs} and the destination node number {Gd, Ld, Pd, Xd, Yd, Zd} can be used to determine whether the data is internal supernode traffic or cross-supernode external traffic. If Gs == Gd and Ls == Ld, it indicates that the traffic is internal supernode traffic. For all-to-all traffic types (the MPI communication library can identify traffic types based on the applied communication primitives), supernode internal expert parallel traffic can be transmitted through the FatTree topology; for all-reduce traffic types, it can be transmitted through the Torus topology, and some all-reduce traffic can share the FatTree topology's bandwidth resources (not conflicting with all-to-all traffic). Otherwise, it indicates that the traffic is cross-supernode external traffic and can be connected to the FatTree topology through a TOR switch.
[0271] For traffic within a supernode, for example, Gs=Gd and Ls=Ld (the target node and the source node are in the same first node group), or Gc==Gd and Lc==Ld (indicating that the data has flowed to the supernode where the destination node is located), the uplink path can be determined by hash routing based on the node coordinates of the destination node. The spine switch in the two-layer FatTree topology then routes the data to the leaf switch connected to the destination node through hash routing based on the topological connection relationship: assuming that X represents the cabinet number (0 to 15), then determine whether it crosses the cabinet based on the X dimension number. If Xd! =Xc, indicating that cross-cabinet communication is required, the traffic needs to enter the Layer 2 fat-tree network, and through hash routing, forward the traffic to the corresponding spine switch. The spine switch then selects the output port according to the cabinet number (X dimension number) where the destination node is located and forwards it to the leaf switch in the cabinet where the destination node is located; if Xd==Xc, indicating that the destination node has reached the cabinet, the leaf switch forwards the data packet inside the cabinet through the TOR switch according to the encoding rules, and forwards the data packet from port: P=Y*Kz+Z+1 to the corresponding node, and then enters the intra-group routing of the first node group.
[0272] For cross-super-node traffic sent from a local node to an external node, the node coordinates of the local node inside the cabinet with a direct link to the destination node can be calculated based on the node coordinates of the destination node (if the destination node and the local node are not in the same first node group, that is, Ls!=Ld, then the node coordinates of the local node inside the cabinet with a direct link to the second jump node are used). For a two-layer FatTree topology, the uplink path can be determined based on the target node number through hash routing. The spine switch then routes the data to the leaf switch connected to the target node through hash routing based on the topological connection relationship: Assuming that X represents the cabinet number (0 to 15), whether it is cross-cabinet is determined based on the X dimension number. If Xd!=Ld, then the uplink path is determined based on the target node number. =Xc, indicating that cross-cabinet communication is required, and the traffic needs to enter the Layer 2 fat-tree network. Through hash routing, the traffic is forwarded to the corresponding spine switch. The spine switch then selects the output port according to the cabinet number (X dimension number) where the target node is located and forwards it to the leaf switch of the cabinet where the destination node is located. If Xd==Xc, indicating that the destination node has been reached, the leaf switch forwards the data packet inside the cabinet where the destination node is located through the TOR switch according to the encoding rules, and forwards the data packet from port: P=Y*Kz+Z+1 to the corresponding node. The node forwards the data packet from the port with a direct link to the destination node according to the node coordinates of the destination node, and transfers it to the intra-group routing of the first node group.
[0273] The following describes in detail the intra-group routing of the first node group.
[0274] If the destination node and the current node are within the same L4 and L3 topologies, i.e., Gd == Gc && Ld == Lc, the current node only needs to determine the routing based on the coordinate offsets of the destination node in the three dimensions of the Torus:
[0275] First, determine the coordinate offset in the X dimension. Compare the coordinate deviation of the destination node in the X dimension with that of the current node. Define the coordinate deviation DeltX == Xd - Xc. If 0 < DeltX <= Kx / 2, it means the destination node is in the positive direction, and route through the X+ port; if DeltX > Kx / 2, it means taking the loopback link is the shortest path, and route through the X- port; if 0 > DeltX >= -Kx / 2, it means the destination node is in the negative direction, and route through the X- port; if DeltX < -Kx / 2, it means taking the loopback link is the shortest path, and route through the X+ port.
[0276] Then, determine the coordinate offset in the Y dimension. Define the coordinate deviation DeltY == Yd - Yc. If 0 < DeltY <= Ky / 2 or DeltY < -Ky / 2, route through the Y+ positive port; if DeltY > Ky / 2 or 0 > DeltY >= -Ky / 2, route through the Y- port.
[0277] Next, determine the coordinate offset in the Z dimension. Define the coordinate deviation DeltZ == Zd - Zc. If 0 < DeltZ <= Kz / 2 or DeltZ < -Kz / 2, route through the Z+ positive port; if DeltZ > Kz / 2 or 0 > DeltZ >= -Kz / 2, route through the Z- port.
[0278] Finally, there is no offset in the Z dimension, reaching the destination node. The DMA communication engine receives the message and completes the communication.
[0279] The operation time of the AI large model is long, usually taking several months to complete training. For large-scale clusters, tens of thousands or hundreds of thousands of computing cards are required for concurrent communication. With hundreds of thousands of system components, equipment failures are inevitable, resulting in interrupted training. Currently, the average time between failures of the mainstream training clusters in the industry is only about a dozen hours, and failures occur almost every day, which will greatly affect the training efficiency. Therefore, if the data routing system has backup resources and can quickly replace faulty nodes, the system availability will be greatly improved. The fusion topology architecture of this application naturally has the characteristic of resource backup. Specifically, taking a cabinet with 64 computing cards as an example, each L3 group can include 65 computer cabinets, and 1 cabinet can be planned as a backup cabinet to back up the other 64 cabinets. Moreover, the algorithm allocation of AI applications is usually allocated in powers of 2, and 16 / 64 is a more appropriate allocation. Resource allocation based on the above method will not cause fragmentation.
[0280] In some possible implementations, the faults may include faults of different types or granularities. The following provides examples of rule recovery for different types of faults.
[0281] When the fault is a node fault, such as Figure 14 As shown, if a 64-card rack contains a faulty node, the job scheduling system (e.g., a job scheduler) can assign a node in the backup rack that has a direct link to the faulty node as a replacement node to continue computing. Data can be returned to the faulty rack via the full mesh network, adding only three hops of latency. This allows for rapid node replacement and effectively improves system availability.
[0282] If the failure involves multiple nodes, the job scheduling system can also use the TOR within the rack to connect multiple backup nodes within the backup rack. For example, if there are two failed nodes, the six direct links between the two failed nodes can be connected, and the backup nodes in the backup rack connected by the 12 direct links can be used as replacement nodes to continue participating in the operation.
[0283] When the fault is a cabinet failure, the job scheduling system can use the backup RACK to replace the faulty RACK, thereby quickly restoring operation.
[0284] The cabinet-level resource backup solution of this application can support the rapid replacement of single-node, multi-node, or even entire cabinet failures, which can effectively improve system availability and accelerate the training efficiency of large AI models.
[0285] Based on the aforementioned data routing system 10 and data routing method, the present application also provides a method for constructing a data routing system. The method can be executed by a management device. The following describes the method for constructing a data routing system from the perspective of a management device.
[0286] See also Figure 15 A flow chart of a method for constructing a data routing system is shown, the method comprising the following steps:
[0287] S1502: The management device determines a switch and multiple network nodes for constructing a data routing system.
[0288] Switches can include in-cabinet switches, specifically switches installed within the cabinet where the network node resides. To facilitate cabling, the cabinet switch can be installed on top of the cabinet, i.e., the cabinet switch can be a TOR switch. In some examples, the cabinet switch can include an inter-cabinet switch. In-cabinet switches can connect different network nodes in different cabinets. Specifically, inter-cabinet switches can connect different cabinets, for example, connecting in-cabinet switches in different cabinets.
[0289] The plurality of network nodes are divided into a second number of first node groups, where the first node group includes the first number of network nodes. After the plurality of network nodes are physically connected, for example, via a physical link, the management device can determine the plurality of network nodes used to construct the data system through node discovery. The management device can discover new network nodes through a heartbeat mechanism, for example. This is not further described here.
[0290] Furthermore, the plurality of network nodes may be divided into a third number of second node groups, wherein the second node groups may include the second number of first node groups, and the third number of second node groups constitute a third node group.
[0291] S1504. The management device sends a configuration file to the switch and multiple network nodes, so that the network nodes in the first node group establish connections with upstream nodes and downstream nodes in various dimensions according to the network structure of the data routing system in the configuration file to form a first topology network with a surround network topology structure, and cascade with the switch to form a second topology network with a fat tree network topology structure, and establish connections between the second number of first node groups to form a third topology network with a fully interconnected network topology structure.
[0292] The network nodes in the first node group can establish connections with upstream nodes and downstream nodes in various dimensions based on the network structure of the data routing system in the configuration file. This can be done by establishing a network connection, such as through a three-way handshake, to facilitate data routing. Similarly, the network nodes in the first node group can establish connections between the second number of first node groups based on the network structure of the data routing system in the configuration file. It should be noted that the connection between the node groups can include a connection between a network node in one node group and a network node in another node group. In the third topology network, the second node groups can include at least one connection. For example, the connection between the second node group A and the second node group B can include a connection between node 1 in the second node group A and node 2 in the second node group B. Furthermore, the connection between the second node group A and the second node group B can also include a connection between node 11 in the second node group A and node 12 in the second node group B. The network nodes in the first node group can also establish a connection with a switch based on the network structure of the data routing system in the configuration file. It should be noted that the connection initiator can be a network node or a switch, and this embodiment does not limit this.
[0293] The data routing system includes a first topology network, a second topology network, and a third topology network, with the first topology network nested within the third topology network. Furthermore, when the plurality of network nodes are divided into a third number of second node groups, the configuration file is used to establish connections between the network nodes in the first node group and the third number of second node groups, thereby forming a fourth topology network having a fully interconnected network topology. Accordingly, the data routing system also includes a fourth topology network, with the third topology network nested within the fourth topology network.
[0294] In some possible implementations, the management device may also check the connection relationships of multiple network nodes based on the network structure of the data routing system in a configuration file, such as a network structure represented by a topology diagram. The network structure in the configuration file is the desired network structure, and the connection relationships are the desired connection relationships. The management device may also obtain the actual connection relationships of multiple network nodes through probing. The management device may compare the desired connection relationships with the actual connection relationships to check the connection relationships of multiple network nodes.
[0295] If the connection relationship between multiple network nodes (e.g., the actual connection relationship) is inconsistent with the connection relationship in the network structure of the data routing system, the management device can also send a prompt message to the user. This prompt message can be used to prompt the user to adjust the connection relationship between the multiple network nodes. In this way, the accuracy and reliability of the network can be guaranteed.
[0296] The management device is further configured to distribute a routing table so that the data routing system can perform data routing based on the routing table. Specifically, the management device may generate a routing table based on the node coordinates of multiple network nodes. The node coordinates are determined based on the number of the network node within each dimension of the first topology network and the number of the first node group to which the node belongs. The management device may then distribute an entry in the routing table related to the at least one network node to at least one of the multiple network nodes.
[0297] The management device sends routing table entries related to at least one network node to reduce storage resources occupied by routing table entries in the network node, thereby saving storage space. It should be noted that if the network node has sufficient storage resources, the management device may also send a complete routing table to at least one network node, and this embodiment does not impose a limitation on this.
[0298] Based on the above description, it can be seen that the method of constructing a data routing system of the present application, by determining the switches and multiple network nodes used to construct the data routing system, and instructing the network nodes and switches to establish connections according to the network structure of the data routing system in the configuration file, thereby constructing a data routing system with nested surround network topology, fat tree network topology, and fully interconnected topology network structure, thereby realizing the formation of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. Among them, the surround network topology has an accelerating effect on the application of local communication characteristics, and supports multi-tenant isolation and accelerates resource pooling; the fat tree network topology can be used for global traffic forwarding, and can also be used as an alternative route for fault recovery to forward traffic. For ultra-large-scale super nodes, a two-layer fat tree topology can be used for non-blocking expansion to achieve global traffic forwarding; the fully interconnected topology network structure provides direct routing for global communication, alleviates the global network congestion level, and can reduce expensive global fiber links compared to the fat tree network topology architecture.
[0299] According to the above embodiments, the hierarchical network proposed in this application that integrates Torus, fat tree and full mesh topologies can effectively reduce the network radius and provide a low-cost, low-power and highly scalable network.
[0300] The processor integrates a minimum 9-port routing engine, enabling networking with 1700W cards. Specifically, 6 ports in the 9-port routing engine or 9-port router can form a 4*4*4 3D Torus within the RACK, port 7 can connect to a 64-port TOR switch, and ports 8 / 9 can form a 2-level full mesh. This enables ultra-large-scale networking with 1700W cards and enables traffic stratification, effectively improving the global communication performance of large AI models.
[0301] The following table compares networks with different topologies in terms of cost, power consumption, and performance. For large AI models with trillions of parameters, the network scale is typically 128K GPUs and the supernode scale is 1024 GPUs. The industry's mainstream network architecture uses a fat tree topology to build supernodes, and supernodes are also interconnected using a fat tree topology. This requires a large number of switches and optical modules, and the network cost is very high. Using the converged topology of this application, the number of switches is reduced by 80% compared to the fat tree architecture, and the number of optical modules is reduced by 61%. As a result, network costs are reduced by 64% and network energy consumption is reduced by 70%.
[0302] Table 3. Topology network cost comparison
[0303]
[0304] Among them, the optical module failure rate is 4%. The fat tree topology requires 4,480,000 optical modules, which cannot guarantee the normal operation of large-scale operations. Using this application can reduce the number of optical modules by 61%, which can effectively improve the reliability of the network system.
[0305] Moreover, this method can achieve RACK-level resource backup, quickly replace failed nodes, and effectively improve system availability.
[0306] Based on the aforementioned method for constructing a data routing system, the present application further provides a management device. The management device will be introduced below from the perspective of functional modularization.
[0307] See also Figure 16 A structural diagram of a management device is shown in FIG. Figure 16 As shown, the management device 1600 includes:
[0308] A determination module 1602 is configured to determine a switch and a plurality of network nodes for constructing a data routing system, wherein the plurality of network nodes are divided into a second number of first node groups, and the first node groups include a first number of network nodes;
[0309] The sending module 1604 is used to send a configuration file to the switch and multiple network nodes, so that the network nodes in the first node group establish connections with upstream nodes and downstream nodes in various dimensions according to the network structure of the data routing system in the configuration file to form a first topology network with a surround network topology structure, and cascade with the switch to form a second topology network with a fat tree network topology structure, and establish connections between the second number of first node groups to form a third topology network with a fully interconnected network topology structure, wherein the data routing system includes the first topology network, the second topology network, and the third topology network, and the first topology network is nested in the third topology network.
[0310] The above-mentioned determination module 1602 and sending module 1604 can be implemented by hardware or software.
[0311] When implemented by software, the determination module 1602 and the sending module 1604 can be applications running on the management device 1600, such as a computing engine. Taking the determination module 1602 as an example, the determination module 1602 can be virtualized through a virtualization service and provided to users for use. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, and container services. Among them, the VM service can be a service that uses virtualization technology to virtualize a virtual machine (VM) resource pool on multiple physical hosts to provide VMs for users to use on demand. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide BMS for users to use on demand. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide containers for users to use on demand. VM is a simulated virtual computer, that is, a logical computer. BMS is a high-performance computing service that is elastically scalable and has computing performance that is no different from that of a traditional physical machine and has the characteristics of secure physical isolation. Containers are a kernel virtualization technology that can provide lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service in the above virtualization services are only specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0312] When implemented in hardware, determination module 1602 may include at least one computing device, such as a server. Alternatively, determination module 1602 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Transmission module 1604 may be implemented using a transceiver or a transceiver unit. The transceiver may include, but is not limited to, a radio frequency transceiver and a fiber optic transceiver.
[0313] In some possible implementations, the plurality of network nodes are divided into a third number of second node groups, where the second node groups include the second number of first node groups. The configuration file is further used to establish connections between the network nodes in the first node group and the third number of second node groups, thereby forming a fourth topology network having a fully interconnected network topology. The data routing system further includes a fourth topology network, in which the third topology network is nested.
[0314] In some possible implementations, the management device 1600 may further include a checking module 1606 and a prompting module 1608. The checking module 1606 is configured to check the connection relationship of the multiple network nodes based on the network structure of the data routing system in the configuration file; the prompting module 1608 is configured to send a prompt message to the user if the connection relationship of the multiple network nodes is inconsistent with the connection relationship in the network structure of the data routing system, and the prompt message is configured to prompt the user to adjust the connection relationship of the multiple network nodes.
[0315] The checking module 1606 and the prompting module 1608 may be implemented by software or hardware.
[0316] When implemented by software, the inspection module 1606 and the prompt module 1608 may be applications running on the management device 1600. The applications may also be virtualized through a virtualization service and provided to the user. For example, the inspection module 1606 may be virtualized through a VM service, a BMS service, or a container service and provided to the user.
[0317] When implemented by hardware, the checking module 1606 and the prompting module 1608 may include at least one computing device, such as a server, etc. Alternatively, the checking module 1606 and the prompting module 1608 may also be implemented using an ASIC or a programmable logic device (PLD).
[0318] In some possible implementations, the management device 1600 may further include a generating module 1609;
[0319] a generating module 1609, configured to generate a routing table based on node coordinates of a plurality of network nodes, where the node coordinates are determined based on numbers of the network nodes in each dimension of the first topology network, port numbers of switches to which the network nodes are connected, and numbers of the first node groups to which the network nodes belong;
[0320] The sending module 1604 is further configured to send an entry in the routing table related to the at least one network node to at least one network node among the multiple network nodes.
[0321] Similar to the aforementioned checking module 1606 and prompting module 1608 , the generating module 1609 may be implemented by software or hardware.
[0322] When implemented via software, generation module 1609 may be an application running on management device 1600, such as a VM service, BMS service, or container service running on management device 1600. When implemented via hardware, generation module 1609 may include at least one computing device, such as a server. Alternatively, generation module 1609 may be implemented using an ASIC or a programmable logic device (PLD).
[0323] The following describes the management device 1600 of the present application from the perspective of hardware implementation. The management device 1600 can be a computing device, such as a server, or a terminal device. The terminal device can be a desktop computer, a laptop computer, a thin client, or a smartphone. The following describes the structure of the computing device of the present application with reference to the accompanying drawings.
[0324] The present application also provides a computing device 1700. Figure 17 As shown, computing device 1700 includes a bus 1702, a processor 1704, a memory 1706, and a communication interface 1708. Processor 1704, memory 1706, and communication interface 1708 communicate with each other via bus 1702. Computing device 1700 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1700.
[0325] The bus 1702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 17 The bus 1702 may include a path for transmitting information between various components of the computing device 1700 (eg, memory 1706, processor 1704, communication interface 1708).
[0326] The processor 1704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0327] The memory 1706 may include volatile memory, such as random access memory (RAM). The memory 1706 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory 1706 stores executable program code, and the processor 1704 executes the executable program code to implement the aforementioned method for constructing a data routing system. Specifically, the memory 1706 stores instructions for the management device 1600 to execute the method for constructing a data routing system.
[0328] The communication interface 1708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1700 and other devices or a communication network.
[0329] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0330] like Figure 18 As shown, the computing device cluster includes at least one computing device 1700. The memory 1706 of one or more computing devices 1700 in the computing device cluster may store instructions for executing the method for building a data routing system by the same management device 1600.
[0331] In some possible implementations, one or more computing devices 1700 in the computing device cluster may also be used to execute some of the instructions of the management device 1600 for executing the method for building a data routing system. In other words, the combination of one or more computing devices 1700 may jointly execute the instructions of the management device 1600 for executing the method for building a data routing system.
[0332] It should be noted that the memory 1706 in different computing devices 1700 in the computing device cluster may store different instructions for executing part of the functions of the management device 1600 .
[0333] Figure 19 A possible implementation is shown. Figure 19 As shown, two computing devices 1700A and 1700B are connected via a communication interface 1708. The memory of computing device 1700A stores instructions for executing the functions of determination module 1602. The memory of computing device 1700B stores instructions for executing the functions of sending module 1604. Furthermore, the memory of computing device 1700A also stores instructions for executing the functions of checking module 1606 and prompting module 1608, and the memory of computing device 1700B also stores instructions for executing the functions of generating module 1609. In other words, the memories 1706 of computing devices 1700A and 1700B jointly store instructions for management device 1600 to execute the method for building a data routing system.
[0334] Figure 19 The connection method between the computing device clusters shown may be based on the fact that the method for building a data routing system provided in this application requires a large amount of resources to send configuration files and routing table entries to a large number of network nodes. Therefore, it is considered that the functions implemented by the sending module 1604 and the generating module 1609 may be performed by the computing device 1700B.
[0335] It should be understood that Figure 19 The functionality of the computing device 1700A shown in FIG. 1 may also be implemented by multiple computing devices 1700. Similarly, the functionality of the computing device 1700B may also be implemented by multiple computing devices 1700.
[0336] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 20 A possible implementation is shown. Figure 20As shown, two computing devices 1700C and 1700D are connected via a network. Specifically, the connection to the network is achieved through a communication interface in each computing device. In this possible implementation, the memory 1706 in computing device 1700C stores instructions for executing the functions of determination module 1602. Simultaneously, the memory 1706 in computing device 1700D stores instructions for executing the functions of sending module 1604. The memory in computing device 1700C also stores instructions for executing the functions of checking module 1606 and prompting module 1608, and the memory in computing device 1700D also stores instructions for executing the functions of generating module 1609.
[0337] It should be understood that Figure 20 The functionality of computing device 1700C shown in FIG. 17 may also be implemented by multiple computing devices 1700. Similarly, the functionality of computing device 1700D may also be implemented by multiple computing devices 1700.
[0338] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct a management device to execute a method for constructing a data routing system.
[0339] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on a management device, the management device executes the method for constructing a data routing system.
[0340] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A data routing system, characterized in that: The data routing system is used to implement data routing between computing nodes, and the data routing system includes a first topology network, a second topology network, and a third topology network; The network structure of the first topology network is a surround network topology structure composed of a first node group, wherein the first node group includes a first number of network nodes, and the network nodes in the surround network topology structure are connected to upstream nodes and downstream nodes in each dimension; The second topology network is a fat tree network topology structure composed of the network nodes and switches in the first topology network, and a cascade relationship exists between the network nodes and switches in the first topology network; The first topology network is nested in the third topology network, and the network structure of the third topology network is a fully interconnected network topology structure composed of second node groups, wherein the second node groups include a second number of first node groups, and there is a connection relationship between the second number of first node groups.
2. The system according to claim 1, wherein: In the second topology network, network nodes belonging to the same first node group are directly connected to the same switch.
3. The system according to claim 1, wherein: The first node group includes multiple subgroups, and the second topology network includes multiple first switches corresponding to the multiple subgroups and at least one second switch. In the second topology network, network nodes belonging to the same subgroup are directly connected to the first switch corresponding to the subgroup, and the multiple first switches are upwardly connected to the at least one second switch.
4. The system according to claim 3, characterized in that The subgroups are divided according to cabinets, cabinet groups or computer rooms.
5. The system according to any one of claims 1 to 4, characterized in that Each network node includes n ports, m of the n ports are used to connect to network nodes in the surround network topology structure, and the m ports include ports for connecting to upstream nodes in each dimension and ports for connecting to downstream nodes in each dimension.
6. The system according to claim 5, characterized in that H ports out of the n ports are used to connect network nodes in the fully interconnected network topology structure formed by the second node group, and the second number is less than or equal to Len j The jth dimension represents the number of network nodes in the surrounding network topology structure, and the q represents the number of dimensions of the surrounding network topology structure.
7. The system according to any one of claims 1 to 6, characterized in that The data routing system further includes a fourth topology network, wherein the third topology network is nested within the fourth topology network; The network structure of the fourth topology network is a fully interconnected network topology structure composed of third node groups, wherein the third node groups include a third number of second node groups, and there is a connection relationship between the third number of second node groups.
8. The system according to claim 7, characterized in that K ports among the n ports are used to connect network nodes in the fully interconnected network topology structure formed by the third node group, and the third number is less than or equal to k·the number of nodes in the second node group+1.
9. The system according to any one of claims 1 to 8, characterized in that The network node is a switching node independent of the computing node, or a routing engine integrated into the computing node.
10. The system according to any one of claims 1 to 9, characterized in that The node coordinates of the network node are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node belongs; The network node is used to determine a data routing path based on the node coordinates, and implement data routing between the computing nodes according to the data routing path.
11. The system according to any one of claims 1 to 10, characterized in that The computing nodes are used to execute artificial intelligence (AI) computing tasks in parallel, and the parallel mode of the AI computing tasks includes at least one of model parallelism, pipeline parallelism, hybrid expert parallelism or data parallelism.
12. The system according to any one of claims 1 to 11, characterized in that The first topology network is used to route data packets in global reduce all-reduce traffic, the second topology network is used to route data packets in many-to-many all-to-all traffic or point-to-point traffic, and the third topology network is used to route data packets in point-to-point traffic.
13. A data routing method, characterized in that: Applied to a data routing system, the data routing system is used to implement data routing between computing nodes, the data routing system includes a first topology network, a second topology network, and a third topology network, the network structure of the first topology network is a surround network topology structure composed of a first node group, the first node group includes a first number of network nodes, and the network nodes in the surround network topology structure have connection relationships with upstream nodes and downstream nodes in each dimension, respectively, the network structure of the second topology network is a fat tree network topology structure composed of each network node and switch in the first topology network, and a cascade relationship exists between each network node and switch in the first topology network, the first topology network is nested in the third topology network, and the network structure of the third topology network is a fully interconnected network topology structure composed of a second node group, the second node group includes a second number of first node groups, and the second number of first node groups are connected. The method includes: A current node in the data routing system receives a data packet to be routed, the current node being a source node of the data packet or a node on a path from the source node to a destination node; The current node determines a data routing path according to node coordinates of the destination node, where the node coordinates are determined based on a number of the network node in each dimension of the first topology network, a port number of the switch to which the network node is connected, and a number of a first node group to which the network node belongs; The current node routes the data packet according to the data routing path.
14. The method according to claim 13, characterized in that The current node determines a data routing path according to the node coordinates of the destination node, including: When the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, the current node determines the routing port in each dimension based on the offset value of the numbers of the destination node and the current node in each dimension of the first topology network.
15. The method according to claim 14, characterized in that The number of network nodes in the i-th dimension of the first topology network is L i The current node determines the routing ports in each dimension according to the offset values of the numbers of the destination node and the current node in each dimension of the first topology network, including: When the offset value of the number of the destination node and the current node in the i-th dimension of the first topology network is greater than 0 and less than or equal to L i / 2, or less than -L i / 2, determining the routing port in the i-th dimension as a port for connecting to a downstream node; When the offset value of the number of the destination node and the current node in the i-th dimension of the first topology network is greater than L i / 2, or greater than or equal to -L i / 2 and is less than 0, it is determined that the routing port in the i-th dimension is a port for connecting to an upstream node.
16. The method according to claim 14 or 15, characterized in that The data packet is a data packet in the global reduction all-reduce traffic.
17. The method according to claim 13, characterized in that The current node determines a data routing path according to the node coordinates of the destination node, including: The current node determines a traffic type according to the node coordinates of the destination node and the node coordinates of the source node, where the traffic type is used to identify whether the data packet is intra-group traffic or inter-group traffic; When the traffic type identifies the data packet as inter-group traffic, the current node determines a forwarding port according to the node coordinates of the destination node.
18. The method according to claim 17, characterized in that The current node determines a forwarding port according to the node coordinates of the destination node, including: When the inter-group traffic is traffic sent locally to the outside, the current node determines, based on the node coordinates of the destination node, the node coordinates of a network node in the group that has a direct link to the destination node, or determines, based on the node coordinates of the network node in the group, the node coordinates of a network node in the group that has a direct link to a jump node to the destination node, and determines, based on the node coordinates of the network node in the group, a port number of the network node in the group connected to the switch; When the cross-group traffic is traffic sent from the outside to the local, the current node determines a forwarding port according to the node coordinates of the destination node and the cascade relationship between the destination node and the switch.
19. The method according to claim 18, characterized in that The data packet is a data packet in all-to-all traffic or point-to-point traffic.
20. The method according to claim 13, wherein The data routing system further includes a fourth topology network, the third topology network is nested within the fourth topology network, the network structure of the fourth topology network is a fully interconnected network topology structure composed of a third node group, the third node group includes a third number of second node groups, a connection relationship exists between the third number of second node groups, and the node coordinates are determined based on the number of the network node in each dimension of the first topology network, the port number of the switch to which the network node is connected, and the number of the first node group to which the network node is located and the number of the second node group to which the network node is located; The current node determines a data routing path according to the node coordinates of the destination node, including: When the number of the second node group where the current node is located is equal to the number of the second node group where the destination node is located, and the number of the first node group where the current node is located is not equal to the number of the first node group where the destination node is located, the current node determines whether a first direct link exists between the current node and the first node group where the destination node is located; If so, the current node determines that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if not, the current node determines that the data routing path includes forwarding the data packet to a second jump node, and the second jump node has a second direct link with the first node group where the destination node is located.
21. The method according to claim 20, characterized in that The current node determines a data routing path according to the node coordinates of the destination node, including: When the number of the second node group where the current node is located is not equal to the number of the second node group where the destination node is located, the current node determines whether a third direct link exists between the current node and the second node group where the destination node is located; If so, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the second node group where the destination node is located; if not, the current node determines that the data routing path includes forwarding the data packet to the first jump node, and there is a fourth direct link between the first jump node and the second node group where the destination node is located.
22. The method according to claim 21, characterized in that The method further comprises: The current node determines whether the number of the first node group where the first jump node is located is equal to the number of the first node group where the current node is located; If not, the current node determines that the data routing path further includes forwarding the data packet to a ferry node, and a fifth direct link exists between the ferry node and the first node group where the first jump node is located.
23. The method according to any one of claims 20 to 22, characterized in that The data packet is a data packet in point-to-point traffic.
24. A method for constructing a data routing system, characterized in that: Applied to managing devices, the method includes: determining a switch and a plurality of network nodes for constructing the data routing system, the plurality of network nodes being divided into a second number of first node groups, the first node groups including a first number of network nodes; A configuration file is sent to the switch and the multiple network nodes so that the network nodes in the first node group establish connections with upstream nodes and downstream nodes in various dimensions according to the network structure of the data routing system in the configuration file to form a first topology network with a surround network topology structure, and cascade with the switch to form a second topology network with a fat tree network topology structure, and establish connections between the second number of first node groups to form a third topology network with a fully interconnected network topology structure. The data routing system includes the first topology network, the second topology network, and the third topology network, and the first topology network is nested in the third topology network.
25. The method according to claim 24, characterized in that the plurality of network nodes being divided into a third number of second node groups, the second node groups including the second number of first node groups; The configuration file is also used for the network nodes in the first node group to establish connections between the third number of second node groups, forming a fourth topology network with a fully interconnected network topology structure. The data routing system also includes the fourth topology network, and the third topology network is nested in the fourth topology network.
26. The method according to claim 24 or 25, characterized in that The method further comprises: Checking the connection relationship of the plurality of network nodes according to the network structure of the data routing system in the configuration file; If the connection relationship between the plurality of network nodes is inconsistent with the connection relationship in the network structure of the data routing system, a prompt message is sent to the user, where the prompt message is used to prompt the user to adjust the connection relationship between the plurality of network nodes.
27. The method according to any one of claims 24 to 26, characterized in that The method further comprises: generating a routing table according to node coordinates of the plurality of network nodes, wherein the node coordinates are determined based on numbers of the network nodes in each dimension of the first topology network, port numbers of the switches to which the network nodes are connected, and numbers of the first node groups to which the network nodes belong; Sending an entry related to the at least one network node in the routing table to at least one network node among the multiple network nodes.
28. A management device, characterized in that: The management device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions, so that the management device performs the method according to any one of claims 24 to 27.
29. A computer-readable storage medium, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 24 to 27.
Citation Information
Cited By
Adaptive weight language model adjustment method based on environment feedback
CN121833274A