Data routing system and data routing method
By adopting a multi-level hierarchical structure of nested fat-tree network topology and fully interconnected network topology in the data center, combined with low-port switching nodes and shortest path routing algorithms, the network cost and power consumption problems of large-scale networking in supercomputing centers are solved, and efficient and reliable data routing is achieved.
Patent Information
- Application Number
- PCT/CN2024/137490
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-09
- Filing Date
- 2024-12-06
- Publication Date
- 2025-12-26
AI Technical Summary
How to design the network topology of a data center to meet the needs of large-scale or ultra-large-scale networking, reduce network costs and power consumption, and improve network scalability and performance, especially in the high-performance computing environment of a supercomputing center.
A multi-level hierarchical network structure with nested fat tree network topology and fully interconnected network topology is adopted. A data routing system is built through low-port switching nodes. Combined with node coordinates and shortest path routing algorithm, efficient data routing between computing nodes is achieved.
It effectively reduces network costs and power consumption, improves network scalability and performance, ensures high reliability and high-performance data routing, and meets the needs of ultra-large-scale networks.
Smart Images

Figure CN2024137490_26122025_PF_FP_ABST
Abstract
Description
A data routing system and a data routing method
[0001] This application claims priority to Chinese Patent Application No. 202410270219.7, filed with the State Intellectual Property Office of China on March 9, 2024, entitled "A Data Routing System and Data Routing Method", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer network technology, and in particular to a data routing system, a data routing method, a method for constructing a data routing system, a management device, a computer-readable storage medium, and a computer program product. Background Technology
[0003] As business scale continues to grow, more and more applications require large-scale data center deployments. Data centers can be categorized into different types based on their application scenarios and requirements. For example, data centers can be divided into cloud data centers and supercomputing centers. Cloud data centers can provide internet services to massive numbers of users, including but not limited to email and database services, while supercomputing centers primarily provide high-performance computing for researchers.
[0004] The following uses a supercomputing center as an example to illustrate data center networking. Supercomputing centers use powerful "supercomputers" to perform calculations, reducing overall solution time and achieving high-performance computing. A supercomputer, also known as a supercomputing system, is a computer capable of performing high-speed calculations that personal computers cannot handle; its specifications and performance are far superior to personal computers. Several countries have successively released exascale supercomputing systems capable of performing 10^18 mathematical operations per second. Post-exascale supercomputing systems face enormous challenges in network scale. Taking a future 10^18 supercomputing system as an example, the scale of the supercomputing system could reach hundreds of thousands of nodes, requiring the network to provide 500,000 port access capabilities.
[0005] How to design the network topology of data centers to meet the needs of large-scale or ultra-large-scale networking has become a key issue of concern in the industry. Summary of the Invention
[0006] This application provides a data routing system that constructs multi-level hierarchical network structures, including nested fat-tree network topologies and fully interconnected network topologies, using small switching nodes with a limited number of ports. This enables the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption while improving network scalability. This application also provides a data routing method, a method for constructing a data routing system, a management device, a computer-readable storage medium, and a computer program product corresponding to the aforementioned data routing system.
[0007] Firstly, this application provides a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network and a second topology network. The first topology network is nested within the second topology network. Nesting refers to insertion or embedding; specifically, the first topology network or components thereof can be inserted into the second topology network. Specifically, the network structure of the first topology network is a fat-tree network topology composed of a first node group and switches. The first node group includes a first number of network nodes, and each network node in the first topology network is cascaded with a switch. The network structure of the second topology network is a fully interconnected network topology composed of second node groups, wherein the second node groups include a second number of first node groups, and there are connections between the second number of first node groups.
[0008] In this application, each network node in the node group is connected to a switch and network nodes in other node groups. Therefore, network nodes can be implemented using small switching nodes with a limited number of ports, such as low-port switching nodes. A large number of low-port switching nodes can be used to construct a multi-level hierarchical network topology that nests fat-tree and fully interconnected network topologies, achieving low-cost, low-power, ultra-large-scale networking. Furthermore, the performance of different topologies in this converged topology complements each other, achieving high reliability, high performance, and high scalability.
[0009] In some possible implementations, network nodes belonging to the same first node group in the first topology network are directly connected to the same switch. Specifically, when the number of ports on a single switch is sufficient to meet the cascading requirements of network nodes in the first topology network, network nodes belonging to the same first node group can be directly connected to the same switch. This allows traffic to hop within the first node group and enables one-hop reachability, thus improving network performance.
[0010] In some possible implementations, the first node group comprises multiple subgroups. The first topology network includes multiple Layer 1 switches corresponding to the multiple subgroups and at least one Layer 2 switch. In the first topology network, network nodes belonging to the same subgroup are directly connected to the Layer 1 switch corresponding to the subgroup, and the multiple Layer 1 switches are connected upwards to at least one Layer 2 switch.
[0011] In this way, the cascading requirements of network nodes in a first node group with a large number of network nodes can be met by a multi-layer fat tree network topology that includes multiple first-layer switches and at least one second-layer switch. Even if there are many network nodes in the first node group, traffic can be hopped within the first node group through first-layer switches and second-layer switches, and multiple hops can be achieved.
[0012] In some possible implementations, the subgroups are obtained by dividing the network into cabinets, cabinet groups, or data centers. By dividing the first node groups of different specifications into subgroups according to their size, network nodes belonging to the same subgroup are directly connected to the first-layer switch corresponding to that subgroup, and the first-layer switches corresponding to multiple subgroups are connected upwards to the second-layer switches, thereby increasing the number of layers in the fat-tree network to meet the cascading requirements of a large number of network nodes.
[0013] In some possible implementations, in the first topology network, network nodes belonging to the same first node group are directly connected to the same first switch and the same second switch. The first switch forwards intra-group traffic, and the second switch forwards inter-group traffic. The source and destination nodes of intra-group traffic are in the same first node group, while the source and destination nodes of inter-group traffic are in different first node groups. This allows for the layering of intra-group and inter-group traffic, avoiding congestion.
[0014] In some possible implementations, each network node includes n ports, of which m ports are used to connect to switches in the first topology network. By providing a small number of ports for cascading with switches, network nodes can achieve traffic routing within the first node group through the switches, thus achieving low communication latency.
[0015] In some possible implementations, h ports out of n ports are used to connect network nodes in a fully interconnected network topology consisting of a second group of nodes. The second number is less than or equal to h * the first number + 1. Thus, different first group of nodes can include at least one global link, providing a direct route for global communication, alleviating global network congestion, and reducing the need for expensive global fiber optic links compared to a fat tree network topology.
[0016] In some possible implementations, the data routing system also includes a third topology network. The second topology network is nested within the third topology network. The third topology network has a fully interconnected network topology consisting of a third group of nodes, where each third group of nodes includes a third number of second group of nodes, and these third group of second group of nodes are interconnected.
[0017] This system compresses the size of the fat tree network by increasing the number of layers in the fully interconnected network topology, such as by using a nested two-level fully interconnected network topology. By increasing the number of network layers, it reduces the network diameter, communication latency, communication performance, and network throughput.
[0018] In some possible implementations, k out of the n ports of each network node are used to connect to network nodes in the fully interconnected network topology formed by the third node group, wherein the third number is less than or equal to k·the number of nodes in the second node group + 1.
[0019] In this way, different groups of second nodes can include at least one global link, which can provide a direct route for global communication, alleviate global network congestion, and reduce expensive global fiber optic links compared to fat tree network topology.
[0020] In some possible implementations, the network node is a switching node independent of the computing node, or a routing engine integrated into the computing node. When the network node uses a routing engine integrated into the computing node, interconnection with switches can be eliminated, further reducing physical costs and power consumption.
[0021] In some possible implementations, the node coordinates of the network node are determined based on the network node's number within a first node group and the number of the first node group to which the network node belongs. The network node's number within the first node group corresponds to the port number of the switch. The network node is used to determine data routing paths based on its node coordinates, and to implement data routing between the computing nodes according to these data routing paths.
[0022] In this way, data routing paths can be determined based on node coordinates using a shortest path routing algorithm. The shortest path routing algorithm provides the shortest communication distance between the source and destination nodes, with the lowest communication latency. Since the data routing system of this application is a hierarchical network structure with nested fat-tree and fully interconnected network topologies, the deterministic shortest path routing algorithm adapted to the topology features is computationally fast and facilitates hardware implementation.
[0023] In some possible implementations, compute nodes are used to perform high-performance computing (HPC) tasks. Traffic generated from executing HPC tasks can include intra-group traffic or cross-group traffic. By employing hierarchical networks with nested fat-tree and fully interconnected network topologies, different traffic can be routed to different topologies, ensuring overall performance.
[0024] Secondly, this application provides a data routing method. This method is applied to a data routing system. The data routing system is used to implement data routing between computing nodes, and the data routing system includes a first topology network and a second topology network. The first topology network is nested within the second topology network. The network structure of the first topology network is a fat-tree network topology composed of a first node group and switches. The first node group includes a first number of network nodes. Each network node in the first topology network is cascaded with a switch. The network structure of the second topology network is a fully interconnected network topology composed of a second node group, and the second node group includes a second number of first node groups, with connections existing between the second number of first node groups.
[0025] Specifically, in the data routing system, the current node receives the data packet to be routed. The current node is either the source node of the data packet or a node on the path from the source node to the destination node. Then, the current node determines the data routing path based on its own coordinates and the coordinates of the destination node. The node coordinates are determined based on the node's number within the first node group and the number of the first node group to which the node belongs. Next, the current node routes the data packet according to the data routing path.
[0026] This method introduces node coordinates from the data routing system, adapts to the topology of the data routing system for data routing, provides the shortest distance communication between the source node and the destination node, achieves low communication latency, and is simple to implement with high availability.
[0027] In some possible implementations, when the number of the current node in the first node group is equal to the number of the destination node in the first node group, the data routing system can determine the routing port based on the number of the destination node in the first node group.
[0028] If the number of the first node group to which the current node belongs is equal to the number of the first node group to which the destination node belongs, it means that the current node and the destination node are in the same first node group. The current node can achieve shortest path routing through the switch. For example, the current node can determine the routing port based on the number of the destination node in the first node group to achieve low communication latency.
[0029] In some possible implementations, when the number of the first node group to which the current node belongs is not equal to the number of the first node group to which the destination node belongs, the current node can determine whether a first direct link exists between the current node and the first node group to which the destination node belongs. If yes, the current node can determine that the data routing path includes forwarding data packets from the first direct link to the first node group to which the destination node belongs; if no, the current node determines that the data routing path includes forwarding data packets to a second hop node, and the second hop node and the first node group to which the destination node belongs have a second direct link.
[0030] In some possible implementations, the data routing system also includes a third topology network, with the second topology network nested within it. The third topology network has a fully interconnected network topology consisting of a third group of nodes, which includes a third number of second node groups. These third number of second node groups are interconnected, and node coordinates are determined based on the node's number within the first node group, the node's number within the first node group, and the node's number within the second node group.
[0031] When the number of the second node group to which the current node belongs is not equal to the number of the second node group to which the destination node belongs, the current node determines whether a third direct link exists between the current node and the second node group to which the destination node belongs. If yes, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the second node group to which the destination node belongs; if no, the current node determines that the data routing path includes forwarding the data packet to the first hop node, and the first hop node and the second node group to which the destination node belongs have a fourth direct link.
[0032] When the current node and the destination node are in different second node groups, this method identifies whether the current node is directly connected to the second node group where the destination node is located. If so, routing is performed through the direct link; otherwise, routing is first performed to the corresponding jump node. This enables routing through the shortest path.
[0033] In some possible implementations, the current node can also determine whether the number of the first node group to which the first hop node belongs is equal to the number of the first node group to which the current node belongs. If not, the current node's determination of the data routing path also includes forwarding data packets to the ferry node, wherein the ferry node and the first node group to which the first hop node belongs have a fifth direct link.
[0034] In this method, when the current node and the first hop node are in different first node groups, the current node can first route to the ferry node to reach the first node group where the first hop node is located. Then, the data packet can be routed to the first hop node through the intra-group routing of the first node group, and then routing can be performed based on the direct link of the hop node, thereby achieving the shortest path routing.
[0035] In some possible implementations, the data packets include those from HPC tasks. By employing hierarchical networks with nested fat-tree and fully interconnected network topologies, different traffic can be routed to different topologies, ensuring overall performance.
[0036] Thirdly, this application provides a method for constructing a data routing system. This method is applied to a management device and includes:
[0037] Multiple network nodes and switches are identified for constructing a data routing system. The multiple network nodes are divided into a second number of first node groups, and the first node groups include a first number of network nodes.
[0038] Configuration files are sent to multiple network nodes and the switches, so that the network nodes in the first node group are cascaded with the switches according to the network structure of the data routing system in the configuration file to form a first topology network with a fat tree network topology. Connections are also established between a second number of first node groups to form a second topology network with a fully interconnected network topology. The data routing system includes the first topology network and the second topology network, with the first topology network nested within the second topology network.
[0039] This method determines the switches and multiple network nodes used to build the data routing system, and instructs the network nodes and switches to establish connections according to the network structure of the data routing system in the configuration file, thereby constructing a data routing system with a nested fat tree network topology and a fully interconnected network topology. This enables the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability.
[0040] In some possible implementations, multiple network nodes are divided into a third number of second node groups, which in turn include a second number of first node groups. Configuration files are also used to establish connections between network nodes in the first node groups and these third number of second node groups, forming a third topology network with a fully interconnected network topology. The data routing system also includes this third topology network, with the second topology networks nested within it.
[0041] By nesting multi-layered fully interconnected network topologies, more ports can be provided and more computing nodes can be connected, thereby enabling the construction of ultra-large-scale networks to meet business needs.
[0042] In some possible implementations, the management device can also check the connectivity of multiple network nodes based on the network structure of the data routing system in the configuration file. If the connectivity of multiple network nodes is inconsistent with the connectivity in the network structure of the data routing system, the management device can send a prompt to the user, suggesting adjustments to the connectivity of the multiple network nodes. This ensures the accuracy and reliability of the network configuration.
[0043] In some possible implementations, the management device can also generate a routing table based on the node coordinates of multiple network nodes, where the node coordinates are determined by the network node's number within a first node group and the number of the first node group to which the node belongs. The management device can then distribute entries from the routing table related to at least one of the multiple network nodes to at least one of those network nodes.
[0044] In this process, the management device sends the entry in the routing table related to the network node to at least one network node, which can reduce the storage resources occupied by the routing table entries in the network node and save storage space.
[0045] Fourthly, this application provides a management device. The management device includes at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor is configured to execute the instructions stored in the at least one memory to cause the management device to perform a method for constructing a data routing system as described in the third aspect or any implementation thereof.
[0046] Fifthly, this application provides a computer-readable storage medium storing instructions that instruct a management device to perform the method for constructing a data routing system as described in the third aspect or any implementation thereof.
[0047] In a sixth aspect, this application provides a computer program product containing instructions that, when run on a management device, causes the management device to perform the method for constructing a data routing system as described in the third aspect or any implementation thereof.
[0048] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0049] To more clearly illustrate the technical methods of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below.
[0050] Figure 1 is a schematic diagram of the architecture of a data routing system provided in this application;
[0051] Figure 2 is a schematic diagram of the ports of a node in a data routing system provided in this application;
[0052] Figure 3 is a schematic diagram of a low-port router provided in this application;
[0053] Figure 4 is a schematic diagram of the structure of an integrated routing engine in a processor provided in this application;
[0054] Figure 5 is a schematic diagram of the structure of a data routing system provided in this application;
[0055] Figure 6 is a flowchart of a data routing method provided in this application;
[0056] Figure 7 is a schematic diagram of the data routing process in a data routing system provided in this application;
[0057] Figure 8 is a schematic diagram of data routing via a virtual channel provided in this application;
[0058] Figure 9 is a schematic diagram of a fault recovery provided in this application;
[0059] Figure 10 is a flowchart of a method for constructing a data routing system provided in this application;
[0060] Figure 11 is a structural schematic diagram of a management device provided in this application;
[0061] Figure 12 is a schematic diagram of the structure of a computing device provided in this application;
[0062] Figure 13 is a schematic diagram of the structure of a computing device cluster provided in this application;
[0063] Figure 14 is a schematic diagram of another computing device cluster provided in this application;
[0064] Figure 15 is a schematic diagram of another computing device cluster provided in this application. Detailed Implementation
[0065] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0066] First, some technical terms involved in the embodiments of this application will be introduced.
[0067] Network topology refers to the specific arrangement or connection method among the members of a network (such as network devices or communication devices like switches and routers). It is generally divided into physical, real, online structures, or logical, virtual, and programmatic structures. If two networks may differ in their internal physical wiring and the distances between nodes, but their member connection structures are identical, then the two networks can be considered to have the same network topology.
[0068] Networking, short for network construction, typically involves designing a suitable network topology based on network requirements, including scale and performance needs, and then constructing the network according to that topology. In scenarios such as cloud computing, high-performance computing (HPC), or artificial intelligence (AI) computing, it is usually necessary to connect various computing devices (also called computing nodes) through network devices (such as switches, routers, and other switching nodes) to build large-scale or ultra-large-scale interconnected networks, such as cloud data centers, supercomputing centers, or AI computing centers, to provide high-bandwidth, low-latency, and high-throughput efficient communication capabilities and effectively leverage computing power.
[0069] Cloud computing, also known as network computing, is an internet-based computing method that allows shared hardware and software resources and information to be provided on demand to various terminals or other devices. These devices then use the infrastructure provided by the service provider to perform calculations. High-performance computing (HPC) refers to providing more powerful computing performance than traditional computers or servers by aggregating computing capabilities. It is commonly used in scenarios such as weather forecasting, astrophysical simulations, molecular dynamics simulations, and earthquake simulations. AI computing is a mathematically intensive process of computing machine learning algorithms. It typically uses acceleration systems and software to extract new insights from large datasets and learn new capabilities in the process. AI computing can be applied to model training or model inference, especially large language model (LLM) inference.
[0070] For ease of description, this application mainly uses the example of building a supercomputing center for high-performance computing as an illustration.
[0071] Supercomputing centers comprise supercomputers. Supercomputers hold strategic importance in the field of science and technology, hence various countries have successively released multiple exascale (E) supercomputing systems. An exascale supercomputing system refers to a supercomputer capable of performing 10^18 mathematical operations per second. Entering the post-exascale era, supercomputing centers face enormous challenges in network scale, scalability, network performance, and reliability and availability. These will be explained in detail below.
[0072] Currently, network equipment in data centers, exemplified by the aforementioned supercomputing centers, accounts for 20%-30% of energy consumption and 10%-20% of network costs. Future 10E-level supercomputing systems could have networks with hundreds of thousands of nodes, requiring 500,000 port access capabilities, and consuming up to 50 megawatts (MW). For 10E-level supercomputing systems, network cost and power consumption present significant challenges. Furthermore, as network scale increases, hop count, latency, and congestion intensify, leading to a sharp decline in network bandwidth utilization. In addition, the technology of Dynamic Random Access Memory (DRAM) lags behind by 2-3 generations, with a failure rate 5 times higher than international standards, severely impacting system reliability and availability. Therefore, designing data center network topologies to effectively reduce network costs and power consumption is a crucial issue that network architecture design must address.
[0073] Many data centers (such as hyperscale supercomputing centers) require interconnect networks to provide high bandwidth, low latency, and high throughput to deliver high-performance communication capabilities. Currently, the mainstream high-performance interconnect network topologies include fat tree network topology, dragonfly network topology, and torus network topology.
[0074] Some exascale supercomputing systems under construction employ the Dragonfly topology. The Dragonfly topology is built using large-port switches. Specifically, it constructs a fully interconnected network structure within and between groups, based on large-port switches. This results in a non-blocking network with a low network diameter (only 4 hops) and low networking costs, reducing valuable global fiber optic links compared to fat-tree networks. However, system scale is limited by switch specifications, requiring a high number of ports. A maximum network size of 270,000 ports based on 64-port switches cannot meet the networking requirements of future 10 exascale supercomputing systems.
[0075] Building upon the Dragonfly topology, the industry has proposed an evolved Dragonfly+ topology. The Dragonfly+ topology is also based on large-port switches. Specifically, it constructs a symmetrical fat-tree structure within groups and a fully interconnected structure between groups, offering strong network scalability, non-blocking operation, and a network diameter of 4 hops. At the same scale, it reduces core switch costs by 20% compared to fat-tree networks and lowers global fiber optic link costs by 50%. However, this solution also requires large-port switches to increase the number of nodes within groups. However, based on 48-port switches, the maximum network size can only accommodate 332,352 ports, which is insufficient to meet the networking requirements of 10E-level supercomputing systems. Furthermore, for adversarial traffic, the throughput can only reach a maximum of 50%.
[0076] Considering the advantages of fat-tree network topology, such as multi-path load balancing, non-blocking performance, stable performance, mature technology, and wide adaptability, it can be used to build ultra-large-scale networks. Building ultra-large-scale networks based on fat-tree topology requires high-cardinality switches (such as 64-port switches and other large-port switches) to meet the scale requirements, reduce network layers, and decrease network diameter. Furthermore, network performance can be improved by increasing port bandwidth and switch forwarding capacity. However, high-cardinality switches face significant challenges in power consumption and manufacturing processes, making it difficult to design and manufacture high-cardinality, high-bandwidth, high-performance switches by improving process specifications.
[0077] Furthermore, increasing port bandwidth and the number of ports can lead to increased network power consumption and costs. Currently, mainstream commercial high-performance switches have port bandwidths reaching 400Gbps, supporting 64 400Gb / s ports. For the 400Gbps bandwidth standard, the energy efficiency of the serializer / deserializer (SerDes) can reach 20pJ / bit. Therefore, the power consumption of SerDes can be 8W / port, while the power consumption of optical modules can be 5W / port. The power consumption of a single-port high-speed SerDes and optical module is at least 13W. Therefore, reducing the number of switch ports, reducing the number of network links, and reducing the number of fiber optic links can effectively reduce network power consumption and costs. In addition, large-scale data centers (such as supercomputing centers) require a large number of fiber optic links to expand network scale. For a fat tree network with 500,000 ports, the number of optical modules is as high as 180,000, and the failure rate of optical modules is as high as 4%, with an average of 20 failures per day, seriously affecting system reliability. Therefore, reducing the number of optical modules can effectively improve system reliability.
[0078] In view of this, this application provides a data routing system. This data routing system can achieve ultra-large-scale networking through a novel converged topology. Specifically, the data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network and a second topology network. The first topology network and the second topology network form a converged topology.
[0079] Specifically, the first topology network has a fat-tree network topology consisting of a first node group and switches. The first node group includes a first number of network nodes. Each network node in the first topology network is cascaded with a switch. The first topology network is nested within a second topology network. Nesting refers to embedding or insertion. Specifically, the first topology network or its components can be embedded into the second topology network as components or constituent units. The second topology network has a fully interconnected network topology consisting of a second node group. The second node group includes a second number of first node groups. Connections exist between the second number of first node groups.
[0080] In this network topology, each network node in a node group connects to a switch and network nodes in other node groups. Therefore, network nodes can be implemented using small switching nodes with a limited number of ports, such as low-port switching nodes. A large number of low-port switching nodes can be used to construct a multi-level hierarchical network topology, enabling low-cost, low-power, ultra-large-scale networking. Furthermore, the performance of different topologies in this converged topology complements each other, achieving high reliability, high performance, and high scalability. Further, computing nodes such as processors can integrate high-speed interconnect chips (IO dies), such as IO ports with routing and forwarding capabilities, thus replacing low-port switching nodes and further reducing network cost and power consumption.
[0081] To make the technical solution of this application clearer and easier to understand, the system architecture of the data routing system of this application is described below with reference to the accompanying drawings.
[0082] Figure 1 shows a schematic diagram of the architecture of a data routing system 10, which includes multiple network nodes (hereinafter sometimes collectively referred to as "nodes 110"), wherein the multiple nodes 110 are located in a communication network 140 that includes at least a first topology network (e.g., 120-1, 120-2, ..., 120-Q, hereinafter sometimes collectively referred to as "first topology network 120") and a second topology network 130.
[0083] In embodiments of this application, a first topology network 120 is nested within a second topology network 130. The network structure of the first topology network 120 is a fat-tree network topology consisting of a first node group and switches. The first node group includes a first number of network nodes; for example, the first node group may include node 110-1, node 110-2, ..., node 110-P, where P is a positive integer greater than 1. Switches can be categorized into different types based on their deployment location; for example, switches can be top-of-rack (TOR) switches, end-of-row (EOR) switches, or middle-of-row (MOR) switches. Each network node in the first topology network 120 (such as node 110-1, node 110-2, ..., node 110-P) is cascaded with a switch. For example, if the first node group may include 64 network nodes and the switch may include 64 ports, then each network node in the first node group can connect to one port of the switch.
[0084] The network structure of the second topology network 130 is a fully interconnected network topology composed of a second group of nodes. The second group of nodes includes a second number of first group of nodes, such as 120-1, 120-2, ..., 120-Q. There are connections between the second number of first group of nodes. As illustrated in Figure 1, there can be connections between 120-1, 120-2, ..., 120-Q. It should be noted that the connections between node groups can be connections between nodes in different node groups, for example, there can be connections between nodes in 120-1 and nodes in 120-2.
[0085] The data routing system 10 may include a single-layer fully interconnected network or multiple layers of fully interconnected networks. When the data routing system 10 includes multiple layers of fully interconnected networks, these networks may be nested. For example, a first-layer fully interconnected network may be nested within a second-layer fully interconnected network, and so on, with the (L-1)th layer fully interconnected network nested within the Lth layer fully interconnected network. Nesting refers to insertion or embedding. Taking the example of a first-layer fully interconnected network nested within a second-layer fully interconnected network, the first-layer fully interconnected network can be embedded into the second-layer fully interconnected network as a component or unit.
[0086] This application does not limit the number of layers (or levels) of the fully interconnected network in the data routing system 10. The number of layers of the fully interconnected network can be configured according to the required network scale. The network scale can be represented by the total number of ports. For example, when the total number of ports is less than or equal to a first threshold, configuring a single-layer fully interconnected network topology in the data routing system 10 can meet the requirements. As another example, when the total number of ports is greater than the first threshold and less than or equal to a second threshold, configuring a two-layer fully interconnected network topology in the data routing system 10 can meet the requirements. It should be noted that the aforementioned first threshold can be the maximum number of ports that the data routing system 10 can provide when a single-layer fully interconnected network topology is nested, and the aforementioned second threshold can be the maximum number of ports that the data routing system 10 can provide when a two-layer fully interconnected network topology is used.
[0087] Based on this, the data routing system can also include a third topology network, and the second topology network 130 can be nested within the third topology network. Similar to the second topology network 130, the network structure of the third topology network can also be a fully interconnected network topology. The main difference between the third topology network and the second topology network 130 is that the third topology network is a fully interconnected network topology based on a third group of nodes. This third group of nodes includes a third number of second node groups, and these third number of second node groups are interconnected. The second node groups are the node groups formed by the network nodes in the second topology network 130, thus enabling the second topology network 130 to be nested within the third topology network.
[0088] Similar to a fully interconnected network topology, a fat-tree network topology can connect to network nodes within a first node group via a single or multiple switch layer. In some possible implementations, the number of ports on a single switch is sufficient to meet the cascading requirements of network nodes in the first topology network. Consequently, network nodes belonging to the same first node group can be directly connected to the same switch. For example, a first node group may include network nodes in a 4x4x4 cabinet. In other words, if the number of network nodes in the first node group is less than or equal to 64, a 64-port switch can meet the cascading requirements of all network nodes in the cabinet, and each network node in the cabinet can be directly connected to at least one port of the aforementioned 64-port switch. In other possible implementations, the first topology network includes a larger number of network nodes, and the number of ports on a single switch is insufficient to meet the cascading requirements. Therefore, multiple switches can be used to meet the cascading requirements of network nodes in the first topology network. Specifically, the first node group includes multiple subgroups, which can be obtained by dividing the network nodes according to cabinets, cabinet groups, or data centers. For example, when the first node group includes 1024 network nodes, such as network nodes in a super node formed by 16 8x8 cabinets, subgroups can be obtained by dividing the network according to the cabinets. The first topology network includes multiple Layer 1 switches corresponding to multiple subgroups and at least one Layer 2 switch. In the first topology network, network nodes belonging to the same subgroup are directly connected to the Layer 1 switch corresponding to the subgroup. It should be noted that each subgroup can correspond to one or more Layer 1 switches. Multiple Layer 1 switches can be connected upwards to at least one Layer 2 switch. The Layer 1 switches can be switches located within cabinets, such as TOR switches, while the Layer 2 switches can be switches located between cabinets.
[0089] Considering that intra-group traffic and inter-group traffic within the first node group can generate resource contention, multiple switches can be configured in the first topology network, such as a first switch and a second switch. In this first topology network, network nodes belonging to the same first node group are directly connected to the same first switch and the same second switch. The first switch forwards intra-group traffic, and the second switch forwards inter-group traffic. The source and destination nodes of intra-group traffic are in the same first node group, while the source and destination nodes of inter-group traffic are in different first node groups. This allows for the layering of intra-group and inter-group traffic, avoiding congestion.
[0090] Next, the network topology will be explained in conjunction with the port connections of the network nodes.
[0091] Specifically, referring to Figure 2, each network node can include n ports, such as port 111-1, port 111-2, ..., 111-n. Each network node can provide m ports for connecting to the switches in the first topology network 120. Here, m can be greater than or equal to 1. Network nodes within the first node group, such as network nodes within a rack, can be interconnected non-blockingly via switches. Communication between nodes requires only one forwarding by the switch, and communication is achieved within a single hop within the rack.
[0092] Each network node has its remaining nm ports used for full interconnection. nm can be greater than or equal to 1. These remaining nm ports can be nested hierarchically according to system scale requirements. Specifically, a network node can connect to other network nodes in the first node group through at least one port other than the aforementioned m ports, forming a fully interconnected network topology. For example, a network node can connect to other network nodes in the first node group through one port other than the m ports to form a second topology network. This ensures that there is at least one global direct link between each first node group. Furthermore, a network node can connect to other network nodes in the second node group through another port other than the m ports to form a third topology network. This ensures that there is at least one global direct link between each second node group.
[0093] In the embodiment of Figure 1, node 110 can be implemented as a physical node or a virtual node. The specific implementations of physical nodes and virtual nodes are described below.
[0094] In some embodiments, node 110 can be a switching node independent of the computing nodes, such as a standalone router. Node 110 can be a low-port switching node, such as a low-port router. Expanding using low-port switching nodes results in low networking costs and multi-path characteristics. Multi-path characteristics refer to the use of multiple paths between nodes to avoid single points of failure and to achieve load balancing based on multiple paths. For ease of description, a low-port router is used as an example. A low-port router is a router with a small number of ports, such as a 3-port router, or a 4-port, 5-port, or 6-port router.
[0095] Figure 3 illustrates a schematic diagram of a low-port router, which includes multiple communication engines, a crossbar matrix, and multiple ports. The communication engines are used for communication between the router or routing engine ports and the processor to perform Direct Memory Access (DMA) operations, data read / write, memory copying, and other operations. The crossbar matrix is used for communication between the router or routing engine ports and the communication engines to enable internal data forwarding in low-port mode.
[0096] As shown in Figure 3, the multiple ports may include at least one port for constructing a fat-tree network topology (such as the first topology network mentioned above) and two ports for constructing a fully interconnected network topology. The at least one port for constructing the fat-tree network topology can be cascaded with a switch, for example, with a TOR switch, to construct the fat-tree network topology. The two ports for constructing the fully interconnected network topology include an L-Link port for implementing a first-layer fully interconnected network (such as the second topology network mentioned above) and a G-Link port for implementing a second-layer fully interconnected network (such as the third topology network mentioned above), where L in L-Link can stand for local and G in G-Link can stand for global.
[0097] In other embodiments, node 110 may be a routing engine integrated into the computing node. In other words, node 110 may be a computing node with an integrated routing engine, which has routing capabilities. The routing engine may be a routing engine integrated into a processor. The processor may be a graphics processing unit (GPU), a neural network processing unit (NPU), or other accelerator devices, or a central processing unit (CPU).
[0098] Figure 4 illustrates a schematic diagram of a routing engine integrated into a processor. The processor may include multiple cores and a Last Level Cache (LLC). Processors typically employ multi-level caching, with each level of cache being faster than the next lower level. When a core needs to access a block of data or an instruction, it first checks the nearest L1 cache. If the data exists, it indicates a cache hit; otherwise, it indicates a cache miss, requiring further checks of the next L1 cache. In some possible implementations, the processor can be a CPU / GPU memory chip integrating High Bandwidth Memory (HBM). The HBM can be a stack of multiple Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), which can then be packaged together with the CPU / GPU die to achieve a large-capacity, high-bandwidth DDR array.
[0099] The structure of the routing engine integrated in the processor can be referenced from the structure of the low-port router in Figure 3. The routing engine can include a communication engine, a crossbar switch matrix, and ports. In the example in Figure 4, the processor integrates or has a built-in multi-port routing engine. At least one port in the multi-port routing engine is used to connect to a switch, such as a TOR switch, thereby constructing a fat-tree network topology. Two ports in the multi-port routing engine, L-Link and G-Link, are used to connect network nodes in other node groups. For example, L-Link is used to connect network nodes in other first node groups, and G-Link is used to connect network nodes in other second node groups, realizing a fully interconnected network topology.
[0100] The processor core can communicate with the routing engine via an internal bus. This internal bus can include, but is not limited to, a Peripheral Component Interconnect Express (PCIe) bus or a Compute eXpress Link (CXL) bus.
[0101] It should be understood that the processor, router, or routing engine of this application may include more or fewer components, functions, or modules. In some embodiments, the processor may also include components such as a bus or interface for interconnection with the routing engine, a processor core, and a resonant circuit, and may also be connected to external devices such as storage devices. For example, a processor integrating a routing engine may also include a memory management unit (MMU) and a memory controller (MC) for managing and controlling external or extended memory such as DDR. The aforementioned memory management unit and memory controller may communicate with the extended DDR via a unified bus (UB). Similarly, the processor may also connect to or extend a solid-state drive (SSD), wherein the processor assembly may communicate with the SSD via the UB bus or the Ctrl bus. Furthermore, the processor may connect to or extend devices such as accelerators. In some embodiments, the router or routing engine may be any device that implements data transmission, such as a network device such as a switch.
[0102] It should be noted that Figures 3 and 4 are examples illustrating at least 3 ports. In other possible implementations of this application, the routing engine integrated in the low-port router or processor may also include other numbers of ports, such as 4 ports, 5 ports, etc.
[0103] In this application, compute nodes can be used to execute HPC tasks. Compute nodes can be standalone compute nodes or compute nodes integrated with a routing engine, such as processors with integrated routing engines. Traffic generated from executing HPC tasks can include intra-group traffic or cross-group traffic. Intra-group traffic has source and destination nodes in the same first node group, such as within the same supernode. Cross-group traffic has source and destination nodes in different first node groups, such as within different supernodes. Accordingly, the first topology network of the data routing system can be used to route intra-group traffic; for example, the first switch in the first topology network can be used to route data packets where the source and destination nodes are in the same supernode, and the second topology network of the data routing system can be used to route cross-group traffic. It should be noted that when routing cross-group traffic, considering the case where the current node is not a network node directly connected to the destination node's first node group, the first topology network can also redirect the cross-group traffic. For example, the second switch in the first topology network can redirect the cross-group traffic to a network node with a direct link to the destination node's first node group, and then route it through that direct link.
[0104] To make the technical solution of this application clearer and easier to understand, the following example illustrates the networking of a processor with an integrated 3-port routing engine and a TOR switch.
[0105] Full mesh is an ideal topology, providing high bandwidth, direct links between any nodes (one-hop reachable), shortest communication distance, lowest latency, non-blocking network, and good performance. However, full mesh requires the number of node ports to be proportional to the scale, making it difficult to meet the needs of large-scale networks. Therefore, this application adopts a layered approach to implement full mesh expansion step by step, enabling ultra-large-scale networking with only a small number of ports and switches. For example, by integrating a processor with a 3-port routing engine (or a 3-port router) and a TOR switch, it can meet the networking requirements of up to 17 million ports, significantly reducing network cost and power consumption.
[0106] Referring to Figure 5, a schematic diagram of a data routing system is shown. A rack can include 64 processors, such as N01 to N64, each of which integrates a 3-port routing engine. The rack also houses a TOR switch, which can be a 64-port box switch. Port 1 of each processor can connect to one port of the TOR switch. Within an L1-level rack, the TOR switch can directly connect to all 64 processors. Communication between processors (e.g., between N01 and N64) requires only one forwarding by the TOR switch, and communication is achieved within a single hop within the rack. Port 2 of each processor is used for full mesh connections between racks. The 64 processors in a rack can provide 64 global direct links, interconnecting nodes across a maximum of 65 racks. Therefore, the L2-level network can reach 64 * 65 = 4160 nodes, and the maximum communication distance between L2-level networks (e.g., the second topology network) is only 3 hops. Port 3 is used for full mesh interconnections between multiple L2-level networks. Each node or processor can provide a global direct link; 65 racks can provide 4160 global direct links. Correspondingly, the L3 network can interconnect a maximum of 4160 * 4161 = 17,309,760 nodes, with a network diameter of only 7 hops. This layered approach enables the creation of a non-blocking network, offering advantages such as ultra-large-scale networking, ultra-low latency, ultra-low cost, and ultra-high performance.
[0107] In some possible implementations, the first topology network can also include additional switches, such as at least one TOR switch. Specifically, the first topology network may include a first switch and a second switch. In the first topology network, network nodes belonging to the same first node group are directly connected to the same first switch and the same second switch. The first switch can be used for intra-group traffic forwarding, and the second switch can be used for cross-group traffic forwarding. Within the group, the source and destination nodes of traffic are in the same first node group, while the source and destination nodes of cross-group traffic are in different first node groups. The added TOR switch can be used for traffic layering, handling global traffic forwarding and avoiding competition for link bandwidth with internal RACK traffic. This achieves a completely non-blocking network, with performance superior to dragonfly topologies and fat tree topologies.
[0108] Figure 5 illustrates an example of three ports: one port for direct connection to the TOR switch within a rack, one port for full mesh connection between racks, and one port for full mesh interconnection between multiple L2-level networks. In practical applications, more ports can be used to connect more nodes or provide more links.
[0109] Specifically, each network node includes n ports, of which m ports are used for cascading with switches in the first network topology. Figure 5 above illustrates this with m=1, constructing a fat-tree network topology. Further, h ports out of the n ports are used to connect network nodes in the fully interconnected network topology formed by the second node group. The second node group includes a second number of first node groups. Considering the limited number of ports, the maximum value of the second number can be determined based on the number of ports used to connect network nodes in the fully interconnected network topology formed by the second node group and the number of nodes in the first node group.
[0110] Based on this, the second quantity can satisfy the following formula: count 2nd ≤h·Q+1 (1)
[0111] Where, count 2ndLet h represent the number of ports used to connect network nodes in the fully interconnected network topology formed by the second node group, and Q represent the number of nodes in the first node group. h can be greater than or equal to 1, ensuring that different first node groups (or first topology networks) are connected by at least one direct global link. Figure 5 illustrates this with h equal to 1. In practical applications, h can also be a positive integer greater than 1, such as 2 or 3, which can increase the number of links between network nodes, reducing the probability of congestion, or connect more first node groups, increasing the network size. Q can be greater than 1; for example, when the first node group consists of network nodes in a single rack, Q can be 4*4*4 = 64.
[0112] In some possible implementations, k out of the n ports of a network node are used to connect to network nodes in a fully interconnected network topology consisting of a third node group. The third node group includes a third number of second node groups. Considering the limited number of ports, the maximum value of this third number can be determined based on the number of ports used to connect network nodes in the fully interconnected network topology consisting of the second node groups and the number of nodes in the second node groups. The maximum number of nodes in the second node groups can be determined based on a second number (the number of first node groups in the second node group) and a first number (the number of nodes in the first node groups), for example, it can be the product of the first and second numbers.
[0113] Based on this, the third quantity can satisfy the following formula: count 3rd ≤k·count 2nd ·count 1st +1 (2)
[0114] Where, count 1st Represents the first quantity, count 2nd To indicate the second quantity, count 3rd Count represents the third quantity, where k represents the number of ports used to connect network nodes in the fully interconnected network topology formed by the third node group. 2nd ·count 1st This can be the number of nodes in the second node group. k can be greater than or equal to 1, so that there is at least one direct global link between different second node groups (or second topology networks). Figure 5 illustrates this with k=1 as an example. In practical applications, k can also be greater than 1, which can increase the number of links between network nodes and reduce the probability of congestion.
[0115] Based on the foregoing embodiments, network nodes can be uniquely identified by their serial numbers in different topology groups. Therefore, these serial numbers or port numbers can be used as the node coordinates of the network node. In some possible implementations, the node coordinates of a network node can be determined based on its serial number within the first node group (e.g., a server rack) and the serial number of the first node group in which the network node belongs. The serial number of the network node within the first node group corresponds to the port number of the switch. It should be noted that when the number of network nodes in the first node group equals the number of switch ports, the serial number of the network node within the first node group and the port number of the switch can have a one-to-one correspondence.
[0116] Furthermore, the data routing system also includes a third topology network. The second topology network is nested within the third topology network, and the third topology network has a fully interconnected network topology structure composed of a third group of nodes. In this case, the node coordinates of a network node can be determined based on the network node's number within the first node group (e.g., its number within a rack), the number of the network node within the first node group, and the number of the network node within the second node group.
[0117] For example, in a hierarchical network with nested fat-tree and two-layer fully interconnected network topologies, node coordinates can be represented by the node's number within the first node group and the numbers of the first and second node groups to which the node belongs. For instance, node coordinates can be represented as {G, L, R}. Here, G represents the node's number in the second node group (e.g., L2 group), L represents the node's number in the first node group (e.g., L1 group), and R represents the node's number within the rack. In some examples, G can range from 0 to 4160 (inclusive), L from 0 to 64 (inclusive), and R from 1 to 64 (inclusive). The coordinates of the switches connecting network nodes within the first node group in the data routing system can be represented as {G, L, 0}.
[0118] It should be noted that a unified numbering system can be used for network nodes such as routing engines integrated in the processor or independent low-port routers. Accordingly, network nodes can determine data routing paths based on the node coordinates represented by these numbers, and then implement data routing between computing nodes according to these paths.
[0119] In this network, network nodes can determine data routing paths based on their coordinates using a shortest path routing algorithm. The shortest path routing algorithm provides the shortest communication distance between the source and destination nodes, has the lowest communication latency, and is the most basic routing algorithm. Since this application uses a hierarchical network with nested fat trees and full meshes (e.g., a two-layer full mesh), a deterministic shortest path routing algorithm adapted to the topology features fast computation and ease of hardware implementation.
[0120] The following section describes the specific implementation of the deterministic shortest path routing algorithm based on adaptive topology features for data routing.
[0121] Referring to Figure 6, a flowchart of a data routing method is shown. This method can be applied to the data routing system 10 of the aforementioned embodiment. The data routing system 10 is used to implement data routing between computing nodes. The data routing system 10 includes a first topology network 120 and a second topology network 130. The first topology network 120 is nested within the second topology network 130. The network structure of the first topology network 120 is a fat-tree network topology composed of a first node group and switches. The first node group includes a first number of network nodes, and each network node in the first topology network 120 is cascaded with a switch. The network structure of the second topology network is a fully interconnected network topology composed of a second node group. The second node group includes a second number of first node groups, and there are connections between the second number of first node groups. The method includes the following steps:
[0122] S602, The current node receives the data packet to be routed.
[0123] Depending on the location of the source and destination nodes in the network, the source node can route data packets to the destination node either directly or by first routing to a next-hop network node and then continuing the routing from that node until the destination node is reached. Specifically, when the data is at the source node, the source node can determine the next-hop node using a shortest path routing algorithm. If that next-hop node is not the destination node, it can also determine its own next-hop node using the same algorithm until the destination node is reached.
[0124] For ease of description, this application refers to the node where data arrives and is waiting to be processed as the current node. The current node can be the source node or a node on the path from the source node to the destination node. Whether it is the source node or a node on the path from the source node to the destination node, the method of determining the data routing path (next hop node) using the shortest path routing algorithm is similar. This application provides a detailed description of data routing from the perspective of the current node.
[0125] Data packets can be data packets for computational tasks, such as data packets for high-performance computing tasks. The packet header can record information about the source node and the destination node, such as the coordinates of the source node, the coordinates of the destination node, or other information that can be used to determine the coordinates of the source node and the destination node.
[0126] S604. The current node determines the data routing path based on the node coordinates of the current node and the node coordinates of the destination node.
[0127] The node coordinates are determined based on the node's number within the first node group (e.g., a cabinet) and the number of the node within the first node group. Furthermore, when the data routing topology 10 includes a third topology network, the node coordinates can be determined based on the node's number within the first node group, the number of the network node within the first node group, and the number of the network node within the second node group.
[0128] The current node can parse data packets to obtain the coordinates of the destination node. In some examples, the destination node's coordinates can be {Gd, Ld, Rd}, where Gd represents the number of the second node group the destination node belongs to, Ld represents the number of the first node group the destination node belongs to, and Rd represents the number of the destination node within the rack. The current node can also obtain its own coordinates; for example, each network node can store its own coordinates after the network is formed, allowing the current node to read its coordinates from storage.
[0129] The current node can determine the data routing path based on its own coordinates and the coordinates of the destination node, combined with the topology of the data routing system 10 (e.g., a topology graph). It should be noted that the coordinates of the current and destination nodes indicate that when they are at the same level, the data routing path can be determined based on the topology of that level. For example, if both the current and destination nodes are in the same first-level topology network (or the same first-level node group), the current node can determine the data routing path based on the topology of that level, such as the topology of the corresponding node group within that level. Alternatively, if the current and destination nodes are in different node groups at the same level, the current node can determine the data routing path based on the topology of that level.
[0130] Based on this, the current node can first identify whether the node groups (such as the first node group and the second node group) to which the current node and the destination node belong are the same based on the node coordinates of the current node and the node coordinates of the destination node. Then, based on the distribution of the node groups to which the current node and the destination node belong, the data routing path can be determined.
[0131] S606. The current node routes data packets according to the data routing path.
[0132] Specifically, the data routing path can indicate the routing port or the node coordinates of the next hop node. Accordingly, the current node can route the data packet to the next hop node based on the routing port indicated by the data routing path or the node coordinates of the next hop node.
[0133] Based on the above description, this application provides a data routing method. This method introduces the node coordinates of the current node and the destination node in the data routing system, adapts the data routing to the topological characteristics of the data routing system, provides the shortest distance communication between the source node and the destination node, achieves low communication latency, and is simple to implement with high availability.
[0134] Next, the specific implementation of determining the data routing path when the current node and the destination node are in node groups at different levels in the embodiment of Figure 6 will be explained.
[0135] In the first implementation, when the ID of the current node in the first node group is equal to the ID of the destination node in the first node group, it indicates that the current node and the destination node are in the same first node group, and the current node can perform intra-group routing. Intra-group routing can be based on a fat-tree topology. Specifically, the current node can determine the routing port based on the destination node's ID within the first node group (e.g., a server rack), and the data routing path can include forwarding data packets from the routing port to the destination node.
[0136] In the second implementation, when the number of the current node in the second node group is equal to the number of the destination node in the second node group, and the number of the current node in the first node group is not equal to the number of the destination node in the first node group, it indicates that the current node and the destination node are in different first node groups within the same second node group. The data routing is an intra-group route within the second node group, also known as an L2 group intra-group route. Considering that the second node group includes different first node groups, the above route is also called an inter-group route within the first node group, also known as an L1 group inter-group route.
[0137] Inter-group routing of the first node group may include: the current node can determine whether there is a direct link (such as L-LINK) between the current node and the destination node in the first node group. In order to distinguish it from the direct link (such as G-LINK) or other direct links in the second node group of the destination node, this application may refer to the direct link with the destination node in the first node group as the first direct link.
[0138] If yes, the current node can determine that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if not, the current node can determine that the data routing path includes forwarding the data packet to the second hop node. The second hop node and the first node group where the destination node is located have a second direct link. When the data packet arrives at the second hop node, it can reach the first node group where the destination node is located via the second direct link. When the data packet reaches the first node group where the destination node is located, L1 Group intra-group routing can be performed.
[0139] The third implementation method involves a second node group number that the current node belongs to, which is not equal to the second node group number of the destination node. This indicates that the current node is in a different second node group, and the data routing is an intra-group route within the third node group, also known as an L3 group intra-group route. The third node group includes multiple second node groups, and this route can also be called an inter-group route within the second node groups, i.e., an L2 group inter-group route.
[0140] Inter-group routing in the second node group may include: the current node determining whether a third direct link exists between the current node and the destination node in the second node group. If so, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the destination node's second node group. When the data packet arrives at the destination node's second node group, intra-group routing (L2 group routing) can be performed. If not, the current node determines that the data routing path includes forwarding the data packet to the first hop node. The first hop node and the destination node's second node group have a fourth direct link.
[0141] Furthermore, the current node can also determine whether the number of the first node group to which the first hop node belongs is equal to the number of the first node group to which the current node belongs. If not, it indicates that the current node and the first hop node are in different first node groups. The current node's determination of the data routing path also includes forwarding data packets to the ferry node. The ferry node and the first hop node's first node group have a fifth direct link. That is, the current node can first route the data packet to the ferry node. Since the ferry node and the aforementioned first hop node are in the same first node group, they can route the data packet to the aforementioned first hop node through L1 group routing. The first hop node and the destination node's second node group have a direct link, such as the aforementioned fourth direct link. Data packets arriving at the first hop node can be routed to the destination node's second node group through the aforementioned fourth direct link. When the data packet reaches the destination node's first node group, L2 group routing can be performed.
[0142] For ease of description, this application uses the topology shown in Figure 5 as an example to illustrate the data routing method. As shown in Figure 5, the node coordinates can be defined as {G, L, R} based on the physical location of the network node in the topology, where G represents the number of the second node group in which the network node is located, L represents the number of the first node group in which the network node is located, and R represents the number of the network node within the rack. The connection relationships of each level of the topology are explained below.
[0143] For L1 topology, port 1 of the routing engine integrated in each processor within the rack is directly connected to the corresponding port of the TOR switch. For example, port 1 of node 1 (e.g., N01) is connected to port 1 of the TOR switch, port 1 of node 2 (e.g., N02) is connected to port 2 of the TOR switch, and port 1 of node 64 (e.g., N64) is connected to port 64 of the TOR switch. In this way, the 64-port TOR switch can satisfy the non-blocking interconnection of 64 nodes within the rack.
[0144] For L2 topologies, port 2 of the routing engine integrated in each processor within the rack serves as the interconnection port between various RACKs within the L2 topology, and is defined as an L-Link. The L-Link number for each routing engine can be the same as the processor's number within the RACK, and the L-Link number PL can range from 0 to 63.
[0145] Specifically, the coordinates of the source node can be represented as {Gs, Ls, Rs}, and the coordinates of the destination node can be {Gd, Ld, Rd}. The L-Link port 2 of the source node is connected to the L-Link port 2 of the destination node. The full mesh connection rule between L2 topologies can be represented as: Ls + PLs = Ld && PLd + PLs = 63 (3)
[0146] Where Ls is the number of the first node group to which the source node belongs, also known as the L2 topology number of the source node; Ld is the number of the first node group to which the destination node belongs, i.e., the L2 topology number of the destination node; PLs is the L-Link port number of the first node group (such as L2 group) to which the source node connects to the destination node; and PLd is the L-Link port number of the first node group (such as L2 group) to which the destination node connects to the source node. Based on this connection relationship, the routing relationship between the first node groups can be calculated.
[0147] For L3 topology, port 3 of the routing engine integrated in each processor within the rack serves as the interconnection port between various L2 full mesh topologies within the L3 topology, and is defined as G-Link. The G-Link number of each routing engine can correspond one-to-one with the L2 full mesh topology and the coordinates of the nodes within the group. The G-Link numbering rule can be: PG = 64 * L + R, and the value range of PG can be: 0-4160. Accordingly, the full mesh connection relationship between L3 topologies can be expressed as: Gs + PGs = Gd && PGd + PGs = 4160 (4)
[0148] Wherein, Gs is the number of the second node group (e.g.) where the source node is located, Gd is the number of the second node group (e.g., L3 topology or L3 group) where the destination node is located, PGs is the G-Link port number connecting the source node to the second node group where the destination node is located, and PGd is the G-Link port number connecting the destination node to the second node group where the source node is located. Based on this connection relationship, the routing relationship between the second node groups can be calculated.
[0149] The data routing method of this application will be described in detail below based on the connection relationship of each level of topology.
[0150] Referring to Figure 7, which illustrates a data routing process in a data routing system, the current node first determines whether the destination node and the current node are in the same second node group. If not, inter-group routing within the second node group is executed. If so, it continues to determine whether the destination node and the current node are in the same first node group. When the destination node and the current node are not in the same first node group, inter-group routing within the first node group is executed. When the destination node and the current node are in the same first node group, intra-group routing within the first node group is executed.
[0151] The specific implementations of inter-group routing for the second node group, inter-group routing for the first node group, and intra-group routing for the first node group are explained below.
[0152] If the destination node and the current node are in different second node groups, i.e., Gd = ! Gc, the first hop node with a direct link to the second node group containing the destination node {Gd, Ld, Rd} is defined. Based on the topology, the coordinates of the first hop node can be determined as follows:
[0153] If the current node has a direct link to the second node group (denoted as Gd) where the destination node is located, then the following formula is satisfied: PGc=64*Lc+Rc=4161-64*Ld-Rd (6)
[0154] In this case, data can be directly forwarded from the current node's G-Link port to reach the destination node in the second node group. Otherwise, it needs to be routed from the current node to the first hop node.
[0155] The above formula is illustrated by the example of a first node group containing 64 nodes, and 65 first node groups interconnected to form a second node group. In practical applications, the maximum number of nodes in a first node group can be represented as Kr, and the maximum number of first node groups in a second node group can be Kl. Accordingly, the current node can determine whether it is the first jump node in the second node group to which the destination node is directly connected by judging whether the following formula holds true: Kr*Lc+Rc==Kl-Kr*Ld-Rd (7)
[0156] Accordingly, the node coordinates of the first jump node can be represented as {Gc, Lm, Rm}, where Lm and Rm satisfy the following formula:
[0157] If the first node group of the first hop node is different from the first node group of the current node, i.e., Lm = ! Lc, the data packet should first be routed to the first node group of the first hop node (denoted as Lm). A node with a direct link to the first node group Lm of the first hop node is defined as a ferry node. The coordinates of the ferry node can be represented as {Gn, Ln, Rn}. Therefore, the data packet can first be routed to the ferry node, and then forwarded from the direct link L-Link to the first node group of the first hop node. The physical location of the ferry node can be obtained according to the following formula: Gn = Gc, Ln = Lc, Rn = PLn = Kr-1-Rm (9)
[0158] Upon reaching the ferry node, data packets can be forwarded to the first hop node via intra-group routing (such as L1-level RACK internal routing) by the TOR switch. After reaching the first hop node, the packets can be forwarded via the direct link L-Link port, thus reaching the destination node's second node group via the direct link. Data packets reaching the destination node's second node group can then reach the destination node via intra-group routing within the second node group.
[0159] The following section provides a detailed explanation of intra-group routing for the second node group (e.g., inter-group routing for the first node group).
[0160] The current node can determine whether it and the destination node are in the same second node group based on their coordinates. If the destination node and the current node are in the same second node group (Gd = Gc, but Ld = !Lc), the current node can determine whether it has a direct link connecting to the destination node in the first node group based on the topology.
[0161] If the current node is a directly connected link to the destination node's first node group, data can be forwarded directly from the directly connected link. See the following formula for details: Ld-Kr+1+Rd==Rc (10)
[0162] If formula (10) is satisfied, it means that the current node is a direct link to the first node group where the destination node is located. Taking Kr = 64 as an example, if Lc = Ld - PLc && PLc = 63 - PLd, that is, Lc = Ld - 63 + Rd = Rc, then it means that the current node is a direct link to the first node group where the destination node is located. The current node can directly forward data packets from the direct link to the first node group where the destination node is located, such as L-Link, to the first node group where the destination node is located. Otherwise, a node with a direct link to the first node group where the destination node is located is defined as the second hop node, and the node coordinates of the second hop node can be determined by the following formula: Gm = Gc, Lm = Ld - PLm && PLm = Kr-1 - PLd, where PLd = Rd (11)
[0163] Taking Kr=64 as an example, the node coordinates of the second hop node can be Gm=Gc, Lm=Ld-PLm&&PLm=63-PLd, PLd=Rd. That is, Gm=Gc, Lm=Ld-63+Rd&&PLm=63-Rd. Therefore, the route first goes to the second hop node: Rm=PLm, and then redirects to the intra-group route of the first node group, where the TOR switch performs routing based on the intra-group number of the first node.
[0164] The following section explains the intra-group routing for the first node group.
[0165] The destination node and the current node are in the same first node group and second node group, i.e., Gd == Gc && Ld == Lc. Therefore, the route can be determined only based on the destination node's number in the first node group, such as its number in the RACK. The TOR switch forwards data packets from port Rd to the destination node. Once the data packet arrives at the destination node, the DMA communication engine receives the data packet, completing the communication.
[0166] Considering the existence of loops in a full mesh topology, deadlock risk can arise. Deadlock refers to the circular occupation of communication resources (such as buffers). For example, if each node sends data to its next adjacent node—node 0 sends to node 2, node 1 sends to node 3, and so on—if all four nodes send data simultaneously, port resources will be circularly occupied, leading to deadlock, communication congestion, and in severe cases, even network paralysis. Based on this, the data routing system in this application can also prevent deadlock through virtual lanes (VLs).
[0167] Virtual channels can effectively avoid deadlock. In the data routing system of this application, since the second topology network (such as L2 level topology) and the third topology network (such as L3 level topology) adopt a full mesh architecture, multiple virtual channels can be allocated to the ports of network nodes in the second topology network and the third topology network to avoid deadlock.
[0168] Specifically, the ports corresponding to the second topology network are configured with a first virtual channel and a second virtual channel, while the ports corresponding to the third topology network are configured with a third virtual channel and a fourth virtual channel. Taking the second topology network as an example, if the destination address is higher than the current address, the first virtual channel can be used, for example, virtual channel 0; if the destination address is lower than the current address, the second virtual channel can be used, for example, virtual channel 1. Communication between nodes is shown in Figure 8. Communication dependencies between nodes do not form loops, thus avoiding deadlock. Similar to the second topology network, in the third topology network, virtual channels do not form loops; therefore, loop deadlocks will not occur between layers.
[0169] Considering the long runtime and large network scale of large-scale applications such as HPC, component failures are inevitable. If the data routing system has backup resources, it can quickly replace faulty nodes, greatly improving system availability. The topology provided in this application can provide RACK-level resource backup. For node-level or RACK-level failures, rapid replacement can be performed based on the backup RACK-level resources, improving system availability.
[0170] In some possible implementations, faults can include different types or granularities. The following examples illustrate rule-based recovery for different types of faults.
[0171] The L2-level topology reserves a primary node group, such as a rack, as a hot backup resource. When a service node in the primary node group other than the hot backup resource fails, as shown in Figure 9, the job scheduling system (such as a job scheduler) can allocate a node within the backup rack that has a direct link to the rack containing the failed node as a replacement node. This can be achieved by routing through a TOR switch to a node directly connected to the backup rack, scheduling the job to a node within the backup rack. This allows for rapid replacement of the failed node without affecting existing jobs, effectively improving system performance. Specifically, the backup rack has direct links to other racks within the L2-level topology, and the corresponding nodes are directly connected to the failed rack via L2 global fiber optic links, enabling rapid replacement of a single node failure.
[0172] When the failure involves multiple nodes, the job scheduling system can also use the TOR within the RACK to connect multiple backup nodes within the backup rack for replacement. For example, if there are two failed nodes, direct links can be established in six directions for each of the two failed nodes, and the backup nodes within the backup RACK connected by these 12 direct links can be used as replacement nodes to continue participating in the computation.
[0173] When the fault is a RACK-level fault, the job scheduling system can use a backup RACK to replace the faulty RACK, thereby quickly restoring operation and greatly improving system availability.
[0174] The rack-level resource backup solution proposed in this application can support rapid replacement of single-node, multi-node, or even whole-rack failures, which can effectively improve system availability and accelerate application operation efficiency.
[0175] Based on the aforementioned data routing system 10 and data routing method, this application also provides a method for constructing a data routing system. This method can be executed by a management device. The method for constructing a data routing system according to this application will be described below from the perspective of a management device.
[0176] Referring to Figure 10, a flowchart of a method for constructing a data routing system is shown, which includes the following steps:
[0177] S1002, The management device identifies multiple network nodes and switches used to build the data routing system.
[0178] Multiple network nodes are divided into a second number of first node groups, which in turn include a first number of network nodes. After multiple network nodes establish a physical connection, such as through a physical link, the management device can identify the multiple network nodes used to build the data system through node discovery. The management device can discover new network nodes through mechanisms such as heartbeats. Further details are omitted here.
[0179] Furthermore, multiple network nodes can be divided into a third number of second node groups. These second node groups may include a second number of first node groups. The third number of second node groups constitutes a third node group.
[0180] Switches can include rack-mounted switches, specifically switches located within the racks where network nodes reside. For ease of cabling, rack-mounted switches can be positioned at the top of the rack; that is, rack-mounted switches can be TOR switches. In some examples, switches may also include inter-rack switches. Switches within different racks can connect upwards to the inter-rack switch. In this way, network nodes within different racks can communicate through the switches.
[0181] S1004. The management device sends configuration files to the switch and multiple network nodes, so that the network nodes in the first node group can be cascaded with the switch according to the network structure of the data routing system in the configuration file to form a first topology network with a fat tree network topology, and establish connections between a second number of first node groups to form a second topology network with a fully interconnected network topology.
[0182] In this network topology, network nodes in the first node group can establish connections with corresponding ports of the switches according to the network structure of the data routing system in the configuration file, thereby achieving cascading with the switches. Switches cascaded with network nodes can then implement data routing between different network nodes. Similarly, network nodes in the first node group can establish connections between a second number of first node groups according to the network structure of the data routing system in the configuration file. It should be noted that connections between node groups can include connections between network nodes in one node group and network nodes in another node group. In the third topology network, the second node groups can include at least one connection. For example, the connection between second node group A and second node group B can include a connection between node 1 in second node group A and node 2 in second node group B. Furthermore, the connection between second node group A and second node group B can also include a connection between node 11 in second node group A and node 12 in second node group B.
[0183] Furthermore, when multiple network nodes are divided into a third number of second node groups, the configuration file is also used to establish connections between network nodes in the first node group and the third number of second node groups, forming a third topology network with a fully interconnected network topology. Correspondingly, the data routing system also includes the third topology network, with the second topology network nested within it.
[0184] In some possible implementations, the management device can also check the connectivity of multiple network nodes based on the network structure of the data routing system in the configuration file, such as a network structure represented by a topology diagram. Here, the network structure in the configuration file represents the desired network structure, and the connectivity relationships represent the desired connections. The management device can also obtain the actual connectivity relationships of multiple network nodes through probing. The management device can then compare the desired connectivity relationships with the actual connectivity relationships to check the connectivity of multiple network nodes.
[0185] If the connection relationships (e.g., actual connections) of multiple network nodes are inconsistent with the connection relationships in the network structure of the data routing system, the management device can also send a prompt message to the user. This prompt message can be used to suggest adjusting the connection relationships of multiple network nodes. In this way, the accuracy and reliability of the network topology can be guaranteed.
[0186] The management device is also used to distribute routing tables, enabling the data routing system to perform data routing based on these tables. Specifically, the management device can generate routing tables based on the node coordinates of multiple network nodes. The node coordinates are determined based on the network node's number within a first node group and the number of the first node group to which the node belongs. The management device can then distribute entries from the routing table related to at least one of the multiple network nodes to at least one of the network nodes.
[0187] In this embodiment, the management device sends routing table entries related to at least one network node to that network node, which can reduce the storage resources occupied by routing table entries in the network node and save storage space. It should be noted that if the network node has sufficient storage resources, the management device can also send the complete routing table to at least one network node; this embodiment does not impose any restrictions on this.
[0188] Based on the above description, the method for constructing a data routing system in this application involves determining the switches and multiple network nodes used to construct the data routing system, and instructing the network nodes to establish connections with the switches and other network nodes in the node group according to the network structure of the data routing system in the configuration file. This constructs a data routing system with a nested fat-tree network topology and a fully interconnected network topology, thereby enabling the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. The converged topology of this application provides direct routes for global communication, alleviating global network congestion, and reduces expensive global fiber optic links compared to the fat-tree topology. Furthermore, network nodes can be implemented by integrating a routing engine into the processor, thus achieving a switchless architecture, effectively reducing communication latency, and further reducing network costs and power consumption.
[0189] The hierarchical network structure with nested fat tree and full mesh topologies proposed in this application can effectively reduce the network radius and provide a low-cost, low-power, high-performance, high-availability, high-reliability, and highly scalable network.
[0190] Table 1 compares the cost and power consumption of networking for a future 10E-class supercomputer. The cost of a switch is approximately $1000, with a power consumption of approximately 220W, while the cost of an optical module is approximately $200, with a power consumption of approximately 12W. A 10E-class supercomputer network can reach a scale of 500,000 nodes. Currently, mainstream high-performance networks primarily use a fat-tree topology, based on the industry's mainstream 64-port high-performance switches, which have a port bandwidth of 400Gbps. Typically, four layers of fat trees are needed to meet the networking requirements of 500,000 nodes.
[0191] Table 1
[0192] As shown in Table 1, a four-layer fat-tree network typically requires 54,880 switches and 3,010,560 optical modules. Using the converged topology proposed in this application, only 7,865 switches of the same specifications are typically needed, representing only 14.33% of the number required for a fat-tree network. The number of optical modules is also only 33.44% of that required for a fat-tree network. Reducing the number of optical modules effectively improves system reliability. The normalized network cost is only 31.74% of that of a fat-tree network, and the network power consumption is only 26.13% of that of a fat-tree network. Since network costs typically account for 10% of cluster costs, using this application can save nearly 70% of network costs.
[0193] This application's multi-level full mesh networking achieves non-blocking connectivity with a network diameter of only 7 hops, enabling low latency and high throughput, significantly improving performance. Furthermore, a minimum of 3-port routing engines integrated within a low-port router or processor can construct a 17 million-port network, demonstrating good scalability. Compared to layer 4 fat-tree networking, this application's networking scheme reduces cost by 68% and power consumption by 73%, achieving low-cost, low-power networking. Moreover, it enables RACK-level resource backup, improving system availability by up to 70%.
[0194] Furthermore, this application can also support the use of optical switches to replace optical fibers. Optical switches have features such as topology adjustment, dynamic programming, and fault tolerance, which can adapt to the needs of large-scale networks.
[0195] Based on the aforementioned method for constructing a data routing system, this application also provides a management device. The management device will be described below from the perspective of functional modularity.
[0196] Referring to Figure 11, a schematic diagram of a management device is shown. As shown in Figure 11, the management device 1100 includes:
[0197] The determination module 1102 is used to determine multiple network nodes and switches for constructing a data routing system, wherein the multiple network nodes are divided into a second number of first node groups, and the first node groups include a first number of network nodes.
[0198] The sending module 1104 is used to send configuration files to multiple network nodes and switches, so that the network nodes in the first node group are cascaded with the switches according to the network structure of the data routing system described in the configuration file to form a first topology network with a fat tree network topology, and to establish connections between the second number of first node groups to form a second topology network with a fully interconnected network topology.
[0199] The data routing system includes a first topology network and a second topology network, with the first topology network nested within the second topology network.
[0200] The aforementioned determining module 1102 and sending module 1104 can be implemented in hardware or in software.
[0201] When implemented in software, the determining module 1102 and the sending module 1104 can be applications running on the management device 1100, such as a computing engine. Taking the determining module 1102 as an example, it can be virtualized through a virtualization service and then provided to users. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, and container services. Specifically, a VM service can be a service that uses virtualization technology to create a pool of virtual machine (VM) resources on multiple physical hosts, providing VMs to users on demand. A BMS service is a service that uses virtualization technology to create a pool of BMS resources on multiple physical hosts, providing BMS services to users on demand. A container service is a service that uses virtualization technology to create a pool of container resources on multiple physical hosts, providing containers to users on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines, and features secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services; specific limitations are not discussed here.
[0202] When implemented in hardware, the determining module 1102 may include at least one computing device, such as a server. Alternatively, the determining module 1102 may also be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The transmitting module 1104 may be implemented using a transceiver or transceiver unit. The transceiver may include, but is not limited to, an RF transceiver or a fiber optic transceiver.
[0203] In some possible implementations, multiple network nodes are divided into a third number of second node groups, which include the second number of first node groups.
[0204] The configuration file is also used to establish connections between network nodes in the first node group and a third number of second node groups to form a third topology network with a fully interconnected network topology. The data routing system also includes the third topology network, with the second topology network nested within it.
[0205] In some possible implementations, the management device 1100 may also include a checking module 1106 and a prompting module 1108. The checking module 1106 is used to check the connection relationship of multiple network nodes according to the network structure of the data routing system in the configuration file; the prompting module 1108 is used to send a prompt message to the user if the connection relationship of multiple network nodes is inconsistent with the connection relationship in the network structure of the data routing system, and the prompt message is used to prompt the user to adjust the connection relationship of multiple network nodes.
[0206] The inspection module 1106 and the prompting module 1108 can be implemented in software or in hardware.
[0207] When implemented in software, the inspection module 1106 and the prompting module 1108 can be applications running on the management device 1100. This application can also be virtualized and made available to the user via a virtualization service; for example, the inspection module 1106 can be virtualized and made available to the user via a VM service, BMS service, or container service.
[0208] When implemented in hardware, the inspection module 1106 and the prompting module 1108 may include at least one computing device, such as a server. Alternatively, the inspection module 1106 and the prompting module 1108 may also be devices implemented using an ASIC or a programmable logic device (PLD).
[0209] In some possible implementations, the management device 1100 may further include a generation module 1109. The generation module 1109 is used to generate a routing table based on the node coordinates of multiple network nodes, where the node coordinates are determined based on the network node's number within a first node group and the number of the first node group to which the node belongs. The sending module 1104 is also used to send entries in the routing table related to at least one network node to at least one of the multiple network nodes.
[0210] The management device 1100 of this application will now be described from the perspective of hardware implementation. The management device 1100 can be a computing device, such as a server, or a terminal device. The terminal device can be a desktop computer, laptop computer, thin client, or smartphone. The structure of the computing device of this application will now be described with reference to the accompanying drawings.
[0211] This application also provides a computing device 1200. As shown in FIG12, the computing device 1200 includes: a bus 1202, a processor 1204, a memory 1206, and a communication interface 1208. The processor 1204, the memory 1206, and the communication interface 1208 communicate with each other via the bus 1202. The computing device 1200 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1200.
[0212] Bus 1202 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 12, but this does not imply that there is only one bus or one type of bus. Bus 1202 can include pathways for transmitting information between various components of computing device 1200 (e.g., memory 1206, processor 1204, communication interface 1208).
[0213] The processor 1204 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0214] Memory 1206 may include volatile memory, such as random access memory (RAM). Memory 1206 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD). Memory 1206 stores executable program code, which processor 1204 executes to implement the aforementioned method for constructing a data routing system. Specifically, memory 1206 stores instructions for management device 1100 to execute the method for constructing a data routing system.
[0215] The communication interface 1208 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1200 and other devices or communication networks.
[0216] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0217] As shown in Figure 13, the computing device cluster includes at least one computing device 1200. The memory 1206 of one or more computing devices 1200 in the computing device cluster may store instructions from the same management device 1100 for executing methods to construct a data routing system.
[0218] In some possible implementations, one or more computing devices 1200 in the computing device cluster can also be used to execute some of the instructions of the management device 1100 for executing the method of constructing a data routing system. In other words, a combination of one or more computing devices 1200 can jointly execute the instructions of the management device 1100 for executing the method of constructing a data routing system.
[0219] It should be noted that the memory 1206 in different computing devices 1200 in the computing device cluster can store different instructions for executing some functions of the management device 1100.
[0220] Figure 14 illustrates one possible implementation. As shown in Figure 14, two computing devices 1200A and 1200B are connected via a communication interface 1208. The memory in computing device 1200A stores instructions for executing the functions of the determination module 1102. The memory in computing device 1200B stores instructions for executing the functions of the sending module 1104. Furthermore, the memory in computing device 1200A also stores instructions for executing the functions of the checking module 1106 and the prompting module 1108, and the memory in computing device 1200B also stores instructions for executing the functions of the generation module 1109. In other words, the memories 1206 of computing devices 1200A and 1200B jointly store instructions used by the management device 1100 to execute the method for constructing a data routing system.
[0221] The connection method between the computing device clusters shown in Figure 14 can be considered because the method for constructing the data routing system provided in this application requires a large amount of resources to send configuration files and entries in the routing table to a large number of network nodes. Therefore, it is considered that the functions implemented by the sending module 1104 and the generating module 1109 are performed by the computing device 1200B.
[0222] It should be understood that the functions of computing device 1200A shown in Figure 14 can also be performed by multiple computing devices 1200. Similarly, the functions of computing device 1200B can also be performed by multiple computing devices 1200.
[0223] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 15 illustrates one possible implementation. As shown in Figure 15, two computing devices 1200C and 1200D are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1206 in computing device 1200C stores instructions for executing the function of the determination module 1102. Simultaneously, the memory 1206 in computing device 1200D stores instructions for executing the function of the sending module 1104. The memory in computing device 1200C also stores instructions for executing the functions of the check module 1106 and the prompting module 1108, and the memory in computing device 1200D also stores instructions for executing the function of the generation module 1109.
[0224] It should be understood that the functions of computing device 1200C shown in Figure 15 can also be performed by multiple computing devices 1200. Similarly, the functions of computing device 1200D can also be performed by multiple computing devices 1200.
[0225] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a management device to perform a method for constructing a data routing system.
[0226] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on a management device, it causes the management device to perform the method described above for constructing a data routing system.
[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data routing system, characterized in that, The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network and a second topology network, with the first topology network nested within the second topology network. The network structure of the first topology network is a fat tree network topology structure consisting of a first node group and a switch. The first node group includes a first number of network nodes, and each network node in the first topology network is cascaded with the switch. The network structure of the second topology network is a fully interconnected network topology structure composed of a second group of nodes, wherein the second group of nodes includes a second number of first groups of nodes, and there are connections between the second number of first groups of nodes.
2. The system according to claim 1, characterized in that, In the first topology network, network nodes belonging to the same first node group are directly connected to the same switch.
3. The system according to claim 1, characterized in that, The first node group includes multiple subgroups, and the first topology network includes multiple first-layer switches and at least one second-layer switch corresponding to the multiple subgroups. In the first topology network, network nodes belonging to the same subgroup are directly connected to the first-layer switch corresponding to the subgroup, and the multiple first-layer switches are connected upward to the at least one second-layer switch.
4. The system according to claim 3, characterized in that, The subgroups are divided according to cabinets, cabinet groups, or computer rooms.
5. The system according to any one of claims 2 to 4, characterized in that, In the first topology network, network nodes belonging to the same first node group are directly connected to the same first switch and the same second switch. The first switch is used to forward intra-group traffic, and the second switch is used to forward inter-group traffic. The source node and destination node of the intra-group traffic are in the same first node group, and the source node and destination node of the inter-group traffic are in different first node groups.
6. The system according to any one of claims 1 to 5, characterized in that, Each network node includes n ports, of which m ports are used to connect to the switches in the first topology network.
7. The system according to claim 6, characterized in that, h of the n ports are used to connect network nodes in the fully interconnected network topology formed by the second node group, where the second number is less than or equal to h·first number+1.
8. The system according to any one of claims 1 to 7, characterized in that, The data routing system further includes a third topology network, in which the second topology network is nested; The network structure of the third topology network is a fully interconnected network topology composed of a third group of nodes, wherein the third group of nodes includes a third number of second group of nodes, and there are connections between the third number of second group of nodes.
9. The system according to claim 8, characterized in that, K ports out of the n ports of each network node are used to connect to network nodes in the fully interconnected network topology formed by the third node group, wherein the third number is less than or equal to k·the number of nodes in the second node group + 1.
10. The system according to any one of claims 1 to 9, characterized in that, The network node is either a switching node independent of the computing node, or a routing engine integrated into the computing node.
11. The system according to any one of claims 1 to 10, characterized in that, The node coordinates of the network node are determined based on the number of the network node in the first node group and the number of the first node group in which the network node is located. The number of the network node in the first node group corresponds to the port number of the switch. The network node is used to determine the data routing path based on the node coordinates, and to realize data routing between the computing nodes according to the data routing path.
12. The system according to any one of claims 1 to 11, characterized in that, The computing nodes are used to execute high-performance computing (HPC) tasks.
13. A data routing method, characterized in that, An application is made to a data routing system for implementing data routing between computing nodes. The data routing system includes a first topology network and a second topology network, with the first topology network nested within the second topology network. The first topology network has a fat-tree network topology structure consisting of a first node group and switches. The first node group includes a first number of network nodes, and each network node in the first topology network is cascaded with the switches. The second topology network has a fully interconnected network topology structure consisting of second node groups, with the second node groups including a second number of first node groups, and these second number of first node groups are interconnected. The method includes: The current node in the data routing system receives the data packet to be routed, and the current node is either the source node of the data packet or a node on the path from the source node to the destination node. The current node determines the data routing path based on the node coordinates of the current node and the node coordinates of the destination node. The node coordinates are determined based on the node's number in the first node group and the number of the first node group to which the node belongs. The current node routes the data packet according to the data routing path.
14. The method according to claim 13, characterized in that, The current node determines the data routing path based on its own node coordinates and the destination node's node coordinates, including: When the number of the current node in the first node group is equal to the number of the destination node in the first node group, the routing port is determined according to the number of the destination node in the first node group.
15. The method according to claim 13, characterized in that, The current node determines the data routing path based on its own node coordinates and the destination node's node coordinates, including: When the number of the first node group to which the current node belongs is not equal to the number of the first node group to which the destination node belongs, the current node determines whether there is a first direct link between the first node group to which the current node belongs and the first node group to which the destination node belongs. If yes, the current node determines that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if no, the current node determines that the data routing path includes forwarding the data packet to the second hop node, where the second hop node has a second direct link with the first node group where the destination node is located.
16. The method according to claim 13, characterized in that, The data routing system further includes a third topology network, in which the second topology network is nested. The network structure of the third topology network is a fully interconnected network topology structure composed of a third node group. The third node group includes a third number of second node groups, and there are connections between the third number of second node groups. The node coordinates are determined based on the node's number in the first node group, the number of the node in the first node group, and the number of the node in the second node group. The current node determines the data routing path based on its own node coordinates and the destination node's node coordinates, including: When the number of the second node group to which the current node belongs is not equal to the number of the second node group to which the destination node belongs, the current node determines whether there is a third direct link between the second node group to which the current node belongs and the second node group to which the destination node belongs. If yes, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the second node group where the destination node is located; if no, the current node determines that the data routing path includes forwarding the data packet to the first hop node, where the first hop node and the second node group where the destination node is located have a fourth direct link.
17. The method according to claim 16, characterized in that, The method further includes: The current node determines whether the number of the first node group to which the first jump node is located is equal to the number of the first node group to which the current node is located. If not, the current node's determination of the data routing path also includes forwarding the data packet to the ferry node, and the ferry node has a fifth direct link with the first node group to which the first jump node is located.
18. The method according to any one of claims 13 to 17, characterized in that, The data packets include data packets for high-performance computing (HPC) tasks.
19. A method for constructing a data routing system, characterized in that, Applied to the management of equipment, the method includes: A plurality of network nodes and switches for constructing the data routing system are determined, the plurality of network nodes being divided into a second number of first node groups, the first node groups comprising a first number of network nodes; A configuration file is sent to the plurality of network nodes and the switch, so that the network nodes in the first node group are respectively cascaded with the switch according to the network structure of the data routing system in the configuration file to form a first topology network with a fat tree network topology, and connections are established between the second number of first node groups to form a second topology network with a fully interconnected network topology. The data routing system includes the first topology network and the second topology network, and the first topology network is nested in the second topology network.
20. The method according to claim 19, characterized in that, The plurality of network nodes are divided into a third number of second node groups, the second node groups including the second number of first node groups; The configuration file is also used for network nodes in the first node group to establish connections between the third number of second node groups to form a third topology network with a fully interconnected network topology. The data routing system also includes the third topology network, and the second topology network is nested within the third topology network.
21. The method according to claim 19 or 20, characterized in that, The method further includes: Based on the network structure of the data routing system described in the configuration file, check the connection relationships of the multiple network nodes; If the connection relationship of the multiple network nodes is inconsistent with the connection relationship in the network structure of the data routing system, a prompt message is sent to the user, which is used to prompt the user to adjust the connection relationship of the multiple network nodes.
22. The method according to any one of claims 19 to 21, characterized in that, The method further includes: A routing table is generated based on the node coordinates of the plurality of network nodes, wherein the node coordinates are determined based on the number of the network node in the first node group and the number of the first node group to which the node belongs; The entry in the routing table related to the at least one network node is sent to at least one of the plurality of network nodes.
23. A management device, characterized in that, The management device includes at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor executes the computer-readable instructions to cause the computing device cluster to perform the method as described in any one of claims 19 to 22.
24. A computer-readable storage medium, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the method of any one of claims 19 to 22.