Data routing system and related method

Through the hierarchical design of nested enhanced surround network topology and fully interconnected network topology, the problems of high interconnection costs, high energy consumption, and linear expansion of communication latency in the cloud data center network topology architecture are solved, and low cost, low power consumption and high scalability of large-scale or hyper-large-scale networks are achieved.

WO2025167061A1PCT designated stage Publication Date: 2025-08-14HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/115333
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-09
Filing Date
2024-08-29
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

The existing cloud data center network topology architecture has problems such as high interconnection costs, high energy consumption, linear expansion of communication latency with scale, and poor global communication performance, making it difficult to form large-scale or super-large-scale networks.

Method used

Using a hierarchical network structure with nested enhanced surround network topology and fully interconnected network topology, large-scale or super-large-scale networks are built through the hierarchical design of nested enhanced surround network topology and fully interconnected networks, reducing network costs and power consumption and improving network scalability.

Benefits of technology

It effectively reduces network costs and power consumption, improves network scalability, supports multi-tenant isolation, accelerates local communication, alleviates global network congestion, and reduces expensive global fiber links.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115333_14082025_PF_FP_ABST
    Figure CN2024115333_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a data routing system. The system is used for realizing data routing between computing nodes. The system comprises a first topological network and a second topological network, wherein the first topological network is nested within the second topological network. The network structure of the first topological network is an enhanced torus network topology structure formed by a first node group, wherein each network node in the enhanced torus network topology structure has a connection relationship separately with an upstream node, a downstream node and a non-adjacent node in each dimension. The network structure of the second topological network is a full mesh network topology structure formed by a second node group, wherein the second node group comprises a second number of first node groups, and there is a connection relationship between the second number of first node groups. In the system, by means of a hierarchical network structure in which an enhanced torus network topology is nested within a full mesh network topology, a large-scale or ultra-large-scale network is established, thereby effectively reducing the network cost and power consumption, and improving the network scalability.
Need to check novelty before this filing date? Find Prior Art

Description

A data routing system and related method

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 9, 2024, with application number 202410179059.5 and invention name “A data routing system and related methods”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer network technology, and in particular to a data routing system, a data routing method, a method for constructing a data routing system, a management device, a computer-readable storage medium, and a computer program product. Background Art

[0003] As business scale continues to grow, more and more applications require large-scale data center deployment. Designing data center network topologies and effectively reducing network costs and power consumption are essential challenges in network architecture design. To facilitate understanding, let's use a cloud data center for cloud computing as an example. Cloud data centers primarily provide services to numerous tenants, and data security is paramount. Therefore, cloud data centers must ensure secure isolation between multiple tenants.

[0004] Currently, mainstream cloud data centers primarily use fat-tree or full-mesh network topologies. Full-mesh network topologies can include spine-leaf network topologies. Some high-performance networks may use dragonfly or torus network topologies.

[0005] However, the aforementioned network topologies, such as the fat-tree and fully interconnected architectures, suffer from high interconnection costs and high network energy consumption. While the torus topology offers flexible networking and low cost, communication latency scales linearly with scale, resulting in poor global communication performance and limited suitability for localized traffic. The dragonfly topology, while saving expensive global fiber links compared to the fat-tree architecture, is limited in system scale by switch specifications and requires a high number of ports. The maximum network size based on a 48-port switch is 90,300 ports, which is insufficient for networking in exascale (100 billion floating-point operations per second) computing systems. Developing large-scale or ultra-large-scale networks at a low cost has become an urgent challenge.

[0006] Summary of the Invention

[0007] This application provides a data routing system that utilizes a hierarchical network structure that combines nested enhanced surround network topologies and fully interconnected network topologies to build large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption and improving network scalability. This application also provides a data routing method corresponding to the aforementioned data routing system, a method for constructing the data routing system, a management device, a computer-readable storage medium, and a computer program product.

[0008] In a first aspect, the present application provides a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network and a second topology network. The first topology network is nested in the second topology network. The network structure of the first topology network is an enhanced surround network topology structure composed of a first node group, wherein the first node group includes a first number of network nodes. The network nodes in the enhanced surround network topology structure have connection relationships with upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. The network structure of the second topology network is a fully interconnected network topology structure composed of a second node group, wherein the second node group includes a second number of first node groups, and there is a connection relationship between the second number of first node groups.

[0009] This system can form a hierarchical network structure consisting of nested enhanced surround network topologies and fully interconnected network topologies through network nodes with a small number of ports, thereby building large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. The enhanced surround network topology accelerates the application of localized communication characteristics, supports multi-tenant isolation, and accelerates resource pooling. The fully interconnected network topology provides direct routing for global communications, alleviates global network congestion, and reduces expensive global fiber links compared to the fat tree network topology architecture.

[0010] In some possible implementations, each network node includes n ports, m of which are used to connect to network nodes in the enhanced surround network topology. The m ports include ports for connecting to upstream nodes, downstream nodes, and non-adjacent nodes in each dimension.

[0011] This system extends the standard surround network topology by adding direct links (also known as cross-links) across nodes by setting up ports that connect to adjacent nodes. This reduces latency within the enhanced surround network topology. In each dimension, any node has a direct link, accessible in a single hop. The maximum intra-cabinet communication distance is only three hops, reducing distance by 50% compared to the standard surround network topology and significantly improving network performance.

[0012] In some possible implementations, h ports out of the n ports are used to connect network nodes in a fully interconnected network topology structure formed by the second node group, and the second number is less than or equal to Len j The jth dimension represents the number of network nodes of the enhanced surround network topology structure, and the q represents the number of dimensions of the enhanced surround network topology structure.

[0013] In the system, different first node groups include at least one global link, which can provide a direct route for global communication, alleviate global network congestion, and reduce expensive global optical fiber links compared to a fat-tree network topology.

[0014] In some possible implementations, the data routing system further includes a third topology network, wherein the second topology network is nested within the third topology network. The network structure of the third topology network is a fully interconnected network topology structure composed of a third node group, wherein the third node group includes a third number of second node groups, and a connection relationship exists between the third number of second node groups.

[0015] The system increases the number of layers of the fully interconnected network topology, such as adopting a nested two-level fully interconnected network topology, to compress the scale of the enhanced surround network, reduce the network diameter by increasing the network level, reduce communication delay, improve communication performance, and increase network throughput.

[0016] In some possible implementations, k ports out of the n ports are used to connect network nodes in the fully interconnected network topology structure formed by the third node group, and the third number is less than or equal to k·the number of nodes in the second node group+1.

[0017] In the system, different second node groups include at least one global link, which can provide a direct route for global communication, alleviate global network congestion, and reduce expensive global optical fiber links compared to a fat-tree network topology.

[0018] In some possible implementations, the network nodes are switching nodes independent of the computing nodes, or routing engines integrated into the computing nodes. When the network nodes are routing engines integrated into the computing nodes, switch interconnection can be eliminated, further reducing physical costs and power consumption.

[0019] In some possible implementations, the node coordinates of the network node are determined based on the number of the network node in each dimension of the enhanced surround network topology and the number of the first node group to which the network node belongs. Accordingly, the network node is configured to determine a data routing path based on the node coordinates, and implement data routing between the computing nodes according to the data routing path.

[0020] In this way, data routing paths can be determined based on node coordinates using a shortest path routing algorithm. This algorithm provides the shortest distance communication between source and destination nodes, minimizing communication latency. Because the data routing system of this application is a hierarchical network structure that nests an enhanced surround network topology and a fully interconnected network topology, a deterministic shortest path routing algorithm that adapts to topological characteristics is computationally scalable and facilitates hardware implementation.

[0021] In some possible implementations, the data routing system further includes a third topology network. The second topology network is nested within the third topology network, and the network structure of the third topology network is a fully interconnected network topology structure composed of a third node group. Node coordinates of network nodes are determined based on the network node's number within each dimension of the enhanced surround network topology structure, as well as the number of the first node group in which the network node is located and the number of the second node group in which the network node is located.

[0022] This enables efficient routing in large-scale networks (e.g., networks connecting more than 17.3 million computing nodes) to meet business needs.

[0023] In a second aspect, the present application provides a data routing method. The method is applied to a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network and a second topology network. The first topology network is nested in the second topology network. The network structure of the first topology network is an enhanced surround network topology structure composed of a first node group, the first node group includes a first number of network nodes, and the network nodes in the enhanced surround network topology structure have connection relationships with upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. The network structure of the second topology network is a fully interconnected network topology structure composed of a second node group. The second node group includes a second number of first node groups, and there is a connection relationship between the second number of first node groups.

[0024] In a specific implementation, a current node in a data routing system receives a data packet to be routed, where the current node is the source node of the data packet or a node on a path from the source node to a destination node. The current node determines a data routing path based on the node coordinates of the current node and the node coordinates of the destination node, where the node coordinates are determined based on the node's number within each dimension of the enhanced surround network topology and the number of the first node group to which the node belongs. The current node then routes the data packet according to the data routing path.

[0025] This method introduces the node coordinates of the current node and the destination node in the data routing system, adapts the topological characteristics of the data routing system to perform data routing, provides the shortest distance communication between the source node and the destination node, and achieves low communication delay. Moreover, this method is simple to implement and has high availability.

[0026] In some possible implementations, when the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, it means that the current node and the destination node are in the same first node group, and the current node can determine the routing port in each dimension based on the offset value of the number of the destination node and the current node in each dimension of the enhanced surround network topology structure.

[0027] In this method, when the current node and the destination node are in the same first node group, the current node can efficiently and quickly determine the routing ports in each dimension in combination with the node coordinates of the destination node, thereby achieving efficient routing.

[0028] In some possible implementations, the number of network nodes in the i-th dimension of the enhanced surround network topology is L i Accordingly, the current node may determine the routing port in the following situations:

[0029] When the offset value of the number between the destination node and the current node in the i-th dimension of the enhanced surround network topology is 1 or -L i +1, determine the routing port in the i-th dimension as the port used to connect to the downstream node;

[0030] When the offset value of the number between the destination node and the current node in the i-th dimension of the enhanced surround network topology is -1 or L i -1, determine the routing port in the i-th dimension as the port used to connect to the upstream node;

[0031] Otherwise, the routing port in the i-th dimension is determined to be a port for connecting to a non-adjacent node.

[0032] In this way, by comparing the offset value of the number of the destination node and the current node in the dimension of the enhanced surround network topology structure with the set value, the port for routing the data packet can be quickly determined, so that the data packet can be routed to the destination node in one hop, greatly shortening the delay.

[0033] In some possible implementations, the data routing system further includes a third topology network, wherein the second topology network is nested within the third topology network, and the network structure of the third topology network is a fully interconnected network topology structure composed of a third number of node groups. The third node group includes a third number of second node groups, and connections exist between the third number of second node groups. Accordingly, the node coordinates are determined based on the node's number within each dimension of the enhanced surround network topology structure, as well as the number of the first node group to which the node belongs and the number of the second node group to which the node belongs.

[0034] Correspondingly, when the number of the second node group where the current node is located is equal to the number of the second node group where the destination node is located, and the number of the first node group where the current node is located is not equal to the number of the first node group where the destination node is located, it means that the current node and the destination node are in different first node groups under the same second node group, and the current node can perform intra-group routing in the second node group.

[0035] The current node may determine whether there is a first direct link between the current node and the first node group where the destination node is located;

[0036] If so, the current node determines the data routing path includes forwarding the data packet along the first direct link to the first node group where the destination node is located. If not, the current node determines the data routing path includes forwarding the data packet to a second jump node. The second jump node has a second direct link to the first node group where the destination node is located.

[0037] The method first checks whether there is a first direct link between the current node and the first node group where the destination node is located. If not, the method determines the jump node with the direct link, which can shorten the delay as much as possible and improve the routing performance.

[0038] In some possible implementations, when the number of the second node group where the current node is located is not equal to the number of the second node group where the destination node is located, it means that the current node and the destination node are in different second node groups, and the current node can perform inter-second node group routing.

[0039] Specifically, the current node may determine whether a third direct link exists between the current node and the second node group where the destination node is located.

[0040] If so, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the second node group where the destination node is located; if not, the current node determines that the data routing path includes forwarding the data packet to the first jump node, and there is a fourth direct link between the first jump node and the second node group where the destination node is located.

[0041] The method first checks whether there is a first direct link between the current node and the second node group where the destination node is located. If not, it determines the jump node with a direct link, which can shorten the delay as much as possible and improve the routing performance.

[0042] In some possible implementations, the current node can further determine whether the number of the first node group where the first jump node is located is equal to the number of the first node group where the current node is located. If not, it means that the first jump node and the current node are in different first node groups. The current node determines that the data routing path also includes forwarding data packets to the ferry node, and there is a fifth direct link between the ferry node and the first node group where the first jump node is located.

[0043] In this way, the first node group where the first jump node is located can be first routed through the ferry node, and the first jump node can be reached through the routing within the first node group, which can greatly shorten the delay.

[0044] In a third aspect, the present application provides a method for building a data routing system. The method is applied to a management device. The management device can be hardware or software. In a specific implementation, the management device can determine multiple network nodes for building a data routing system. The multiple network nodes are divided into a second number of first node groups, and the first node group includes a first number of network nodes. The management device can then send a configuration file to the multiple network nodes so that the network nodes in the first node group, according to the network structure of the data routing system in the configuration file, establish connections with upstream nodes, downstream nodes, and non-adjacent nodes in each dimension, respectively, to form a first topology network whose network structure is an enhanced surround network topology structure, and establish connections between the second number of first node groups to form a second topology network whose network structure is a fully interconnected network topology structure. Wherein, the data routing system includes a first topology network and a second topology network, and the first topology network is nested in the second topology network.

[0045] This method constructs a data routing system by identifying multiple network nodes for building the data routing system and instructing the network nodes to initiate connections according to the network structure of the data routing system in the configuration file. This method then builds a data routing system that nests enhanced surround network topologies and fully interconnected network topologies. This enables the construction of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption and improving network scalability. The enhanced surround network topology accelerates the application of localized communication features, supports multi-tenant isolation, and accelerates resource pooling. The fully interconnected topology provides direct routing for global communications, alleviating global network congestion and reducing expensive global fiber links compared to the fat-tree network topology.

[0046] In some possible implementations, the plurality of network nodes are divided into a third number of second node groups, where the second node groups include a second number of first node groups. The configuration file is further used to establish connections between the network nodes in the first node group, the third number of second node groups, to form a third topology network having a fully interconnected network topology. The data routing system further includes the third topology network, and the second topology network is nested within the third topology network.

[0047] In this way, a super-large-scale network can be built, through which more computing nodes can be connected to meet the growing computing needs.

[0048] In some possible implementations, the management device may also check the connection relationships of multiple network nodes based on the network structure of the data routing system in the configuration file. The network structure in the configuration file is the desired network structure, and the connection relationships are the desired connection relationships. The management device may also obtain the actual connection relationships of multiple network nodes through detection. The management device may compare the desired connection relationships with the actual connection relationships to check the connection relationships of multiple network nodes. If the connection relationships of multiple network nodes are inconsistent with the connection relationships in the network structure of the data routing system, the management device may send a prompt to the user, prompting the user to adjust the connection relationships of the multiple network nodes.

[0049] In this way, real-time monitoring of the network structure of the data routing system can be achieved. When the network structure is inconsistent with the configuration file, the user can be reminded to adjust the connection relationship in time to ensure the accuracy and reliability of the network.

[0050] In some possible implementations, the management device may further generate a routing table based on node coordinates of multiple network nodes. The node coordinates are determined based on the number of the network node in each dimension of the enhanced surround network topology and the number of the first node group to which the node belongs. The management device may then distribute an entry in the routing table related to at least one network node among the multiple network nodes.

[0051] The management device sends routing table entries related to at least one network node to reduce storage resources occupied by routing table entries in the network node, thereby saving storage space. It should be noted that if the network node has sufficient storage resources, the management device may also send the complete routing table to at least one network node, thereby reducing complexity on the management device side.

[0052] In a fourth aspect, the present application provides a management device. The management device can be implemented by at least one computing device, for example, a single computing device or a cluster of computing devices. The management device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions to perform the method described in the third aspect or any implementation of the third aspect.

[0053] In a fifth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct a computing device or a computing device cluster to execute the method described in the third aspect or any implementation of the third aspect.

[0054] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or computing device cluster to execute the method described in the third aspect or any one of the implementations of the third aspect.

[0055] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical methods of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments.

[0057] FIG1 is an architecture diagram of a data routing system provided by the present application;

[0058] FIG2 is a schematic diagram of ports of a node in a data routing system provided by the present application;

[0059] FIG3 is a schematic diagram of an enhanced surround network topology and a standard surround network topology provided by the present application;

[0060] FIG4 is a schematic structural diagram of an 11-port router provided by the present application;

[0061] FIG5 is a schematic diagram of the structure of a routing engine integrated in a processor provided by the present application;

[0062] FIG6 is a schematic diagram of a topology structure for constructing an enhanced surround network provided by the present application;

[0063] FIG7 is a schematic diagram of a nested enhanced surround network topology structure provided by the present application, which is a nested two-level fully interconnected topology network structure;

[0064] FIG8 is a flow chart of a data routing method provided by the present application;

[0065] FIG9 is a schematic diagram of a data routing process in a data routing system provided by the present application;

[0066] FIG10 is a schematic diagram of data routing through a virtual channel provided by the present application;

[0067] FIG11 is a flow chart of a method for constructing a data routing system provided by the present application;

[0068] FIG12 is a comparative schematic diagram of multi-tenant isolation provided by this application;

[0069] FIG13 is a comparative schematic diagram of memory access in a resource pooling situation provided by this application;

[0070] FIG14 is a schematic diagram of the structure of a management device provided by this application;

[0071] FIG15 is a hardware structure diagram of a computing device provided by the present application;

[0072] FIG16 is a hardware structure diagram of a computing device cluster provided by this application;

[0073] FIG17 is a hardware structure diagram of another computing device cluster provided by this application;

[0074] FIG18 is a schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION

[0075] The terms "first" and "second" in the embodiments of this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.

[0076] First, some technical terms involved in the embodiments of this application are introduced.

[0077] Network topology refers to the specific arrangement or connection between network components (e.g., switches, routers, and other network or communication devices). It is generally categorized as either a physical, real, online structure or a logical, virtual, programmatic structure. If two networks have different physical connections and node distances, but the connections between the network components are the same, then the network topologies of the two networks are considered the same.

[0078] Networking, short for network formation, typically involves designing a suitable network topology (also known as a network topology structure or simply "topology") based on networking requirements, including scale and performance requirements, and then constructing the network according to that topology. In cloud computing, high-performance computing (HPC), or artificial intelligence (AI) computing scenarios, it's often necessary to connect computing devices (also known as compute nodes) through network devices (such as switches, routers, and other switching nodes) to form large-scale or ultra-large-scale interconnected networks. This provides efficient communication capabilities with high bandwidth, low latency, and high throughput, effectively unleashing computing power.

[0079] Among them, cloud computing, also known as network computing, is an Internet-based computing method through which shared hardware and software resources and information can be provided to various terminals or other devices on demand, using the infrastructure provided by the service provider for computing. High-performance computing refers to providing more powerful computing performance than traditional computers or servers by aggregating computing power. It can usually be used in scenarios such as weather forecasting, astrophysical simulation, molecular dynamics simulation, and earthquake simulation. AI computing is a mathematically intensive process for calculating machine learning algorithms, which can usually use accelerated systems and software. It can extract new insights from large data sets and learn new capabilities in the process. AI computing can be applied to model training or model reasoning, especially large language model (LLM) reasoning.

[0080] For ease of description, this application mainly uses cloud computing scenarios as examples. Cloud computing services provide services to many tenants, requiring different computing resources to be provided according to the computing power requirements of each tenant, flexible allocation of computing resources, and requiring secure isolation between tenants to ensure user data security. Based on this, the data center network (such as cloud data center) that supports cloud computing services has a large amount of local communication behavior. Related research also shows that the main traffic in the data center network is limited to servers in the same cabinet (such as rack), or servers inside interconnected cabinets. For example, in some cloud data centers, more than 80% of the traffic is concentrated inside the rack, which has obvious local communication characteristics, and the communication mode has a strong spatial locality.

[0081] Currently, mainstream cloud data centers primarily employ fat-tree network topologies (also referred to as fat-tree networks or fat-tree topologies) or full-mesh network topologies (also referred to as full-mesh networks or full-mesh topologies). Full-mesh network topologies may include spine-leaf network topologies. Some high-performance networks may employ dragonfly or wrap-around network topologies. Wrap-around network topologies may include torus network topologies (also referred to as torus networks or torus topologies).

[0082] The fat-tree network topology is a symmetrical topology with high network throughput, non-blocking performance, and stable performance. However, it suffers from issues such as high interconnection costs and high network energy consumption. As network scale expands, the demand for optical modules continues to increase, and in particular, network costs and energy consumption increase linearly with network scale. The torus network topology offers low cost, flexible networking, and strong scalability. However, it is a blocking network, with communication latency scaling linearly with scale. This results in poor global communication performance and makes it suitable for localized traffic. The Dragonfly network topology is a fully interconnected network architecture with both intra-group and inter-group connectivity. It is non-blocking for uniform random traffic, has a low network diameter of only four hops, and offers low networking costs. Compared to the fat-tree network, it can save on expensive global fiber links. Currently, most exascale supercomputers use the Dragonfly topology. However, network scale is limited by switch specifications, requiring a high number of ports. The maximum network size based on a 48-port switch is 90,300 ports, which is insufficient to meet the networking requirements of exascale computing systems.

[0083] In view of this, the present application provides a data routing system. The data routing system is used to implement data routing between computing nodes. The data routing system includes a first topology network and a second topology network. The first topology network is nested within the second topology network. Nesting refers to insertion or embedding. Specifically, the first topology network can be embedded in the second topology network as a component or constituent unit of the second topology network. The second topology network includes the above-mentioned first topology network. For example, the second topology network can be formed by connecting multiple first topology networks.

[0084] The network structure of the first topology network is an enhanced torus network topology structure (Enhanced Torus) composed of a first node group. The first node group includes a first number of network nodes, and the network nodes in the enhanced torus network topology structure are connected to upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. The network structure of the second topology network is a full mesh network topology structure composed of a second node group. The second node group includes a second number of first node groups, and the second number of first node groups are connected to each other.

[0085] This data routing system can form a hierarchical network structure consisting of nested Enhanced Torus and full mesh topologies using network nodes with a small number of ports, thereby building large-scale or ultra-large-scale networks, such as an ultra-large-scale network with 17.3 million ports. This effectively reduces network costs and power consumption, and improves network scalability. The Enhanced Torus topology accelerates applications with localized communication characteristics, such as cloud data center applications, and supports multi-tenant isolation and accelerated resource pooling. The full mesh topology provides direct routing for global communications, alleviating global network congestion and reducing expensive global fiber links compared to the fat tree network topology architecture.

[0086] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the data routing system of the present application is introduced below with reference to the accompanying drawings.

[0087] Figure 1 shows an architectural diagram of a data routing system, where the data routing system 10 includes multiple network nodes (sometimes collectively referred to as "nodes 110" below), wherein the multiple nodes 110 are located in a communication network 140 that includes at least a first topology network (e.g., 120-1, 120-2, ..., 120-Q, sometimes collectively referred to as "first topology network 120" below) and a second topology network 130.

[0088] In an embodiment of the present application, the first topology network 120 is nested in the second topology network 130. The network structure of the first topology network 120 is an enhanced surround network topology structure composed of a first node group, and the first node group includes a first number of network nodes. For example, the first node group may include node 110-1, node 110-2, ..., node 110-P, where P is a positive integer greater than 1. The network nodes in the enhanced surround network topology structure (such as node 110-1, node 110-2, ..., node 110-P) are connected to upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. Taking node 110-2 as an example, the upstream nodes, downstream nodes, and non-adjacent nodes included in a certain dimension of node 110-2 are nodes 110-1, node 110-3, and node 110-4, respectively, and node 110-2 is connected to node 110-1, node 110-3, and node 110-4.

[0089] The network structure of the second topology network 130 is a fully interconnected network topology structure composed of second node groups. The second node groups include a second number of first node groups, such as 120-1, 120-2, ..., and 120-Q. A connection relationship exists between the second number of first node groups. Using FIG1 as an example, a connection relationship can exist between 120-1, 120-2, ..., and 120-Q. It should be noted that the connection relationship between node groups can be a connection relationship between nodes in different node groups, such as a connection relationship between the nodes in 120-1 and the nodes in 120-2.

[0090] The data routing system 10 may include a layer of fully interconnected networks, or multiple layers of fully interconnected networks. When the data routing system 10 includes multiple layers of fully interconnected networks, the multiple layers of fully interconnected networks may be nested. For example, the first layer of fully interconnected networks may be nested within the second layer of fully interconnected networks, and so on. The L-1 layer of fully interconnected networks may be nested within the L layer of fully interconnected networks.

[0091] The present application does not impose any restrictions on the number of layers (or levels) of the fully interconnected network in the data routing system 10. The number of layers of the fully interconnected network can be configured according to the required network scale. The network scale can be characterized by the total number of ports. For example, when the total number of ports is less than or equal to the first threshold, the data routing system 10 can be configured with a one-layer fully interconnected network topology to meet the demand. For another example, when the total number of ports is greater than the first threshold and less than or equal to the second threshold, the data routing system 10 can be configured with a two-layer fully interconnected network topology to meet the demand. It should be noted that the above-mentioned first threshold can be the maximum number of ports that the data routing system 10 can provide when a one-layer fully interconnected network topology is nested, and the above-mentioned second threshold can be the maximum number of ports that the data routing system 10 can provide when a two-layer fully interconnected network topology is configured.

[0092] Based on this, the data routing system may further include a third topology network, and the second topology network 130 may be nested in the third topology network. Similar to the second topology network 130, the network structure of the third topology network may also be a fully interconnected network topology structure. The main difference between the third topology network and the second topology network 130 is that the network structure of the third topology network is a fully interconnected network topology structure based on a third node group. The third node group includes a third number of second node groups, and a connection relationship exists between the third number of second node groups. The second node group is a node group formed by the network nodes in the second topology network 130, thereby enabling the second topology network 130 to be nested in the third topology network. Specifically, referring to Figure 2, each network node may include n ports, such as port 111-1, port 111-2, ..., 111-n. Each network node may provide m ports for connecting network nodes in the enhanced surround network topology structure. The enhanced torus topology may include at least one dimension. For example, the enhanced torus topology may be a three-dimensional (3D) enhanced torus topology. Accordingly, the m ports include ports for connecting upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. Where m is greater than or equal to 3. Network nodes in the enhanced torus topology connect to upstream nodes, downstream nodes, and non-adjacent nodes in each dimension through the m ports, thereby achieving local interconnection in each dimension of the enhanced torus topology.

[0093] As shown in Figure 3, the enhanced torus network topology adds at least one additional port to the standard torus network topology for connecting non-adjacent nodes. Taking a three-dimensional torus network as an example, the X dimension of the standard three-dimensional torus network topology provides X+ and X- ports, representing the positive and negative directions. The X+ port of each upstream node connects to the X- port of the downstream node, forming a torus network. The enhanced three-dimensional torus network topology also provides ports for inter-node connections, such as X-ports. X-ports are used for enhanced inter-node links, connecting non-adjacent nodes. This ensures that any node in a four-dimensional torus has a direct link in each dimension, allowing one-hop reachability. By adding a cross-link, using at least three ports in each dimension, each of the four nodes in the dimension has a direct link, allowing any two nodes to be connected in one hop. This reduces the communication distance within a cabinet to just three hops, reducing the communication distance by 50% compared to a standard torus and significantly improving network performance.

[0094] The remaining nm ports of each network node are used for full interconnection. Where nm can be greater than or equal to 1. A network node can connect to other network nodes in the first node group via at least one port other than the aforementioned m ports, forming a fully interconnected network topology. For example, a network node can connect to other network nodes in the first node group via one port other than the m ports to form a second topology network. Furthermore, a network node can connect to other network nodes in the second node group via another port other than the m ports to form a third topology network.

[0095] The node 110 in the embodiment of Figure 1 can be implemented as a physical node or a virtual node. The specific implementation of the physical node and the virtual node will be described below.

[0096] In some embodiments, node 110 may be a switching node independent of a computing node, such as an independent switch or router. Node 110 may be a low-port switch or a low-port router. By using a low-port switch or a low-port router for expansion, network construction costs are low and multi-path features are provided. Multi-path features refer to the use of multiple paths between nodes to avoid single points of failure. For ease of description, a low-port router is used as an example. A low-port router is a router with a small number of ports relative to a high-radix router. For example, a low-port router may be an 11-port router, or a 12-port, 13-port, or 14-port router. A high-radix router may be a 48-port router or a 64-port router.

[0097] Figure 4 shows a schematic diagram of an 11-port router, which includes multiple communication engines, a crossbar, and 11 ports. The communication engines facilitate communication between the ports of a router or routing engine and the processor, enabling operations such as Direct Memory Access (DMA), data reads and writes, and memory copies. The crossbar facilitates communication between the ports of a router or routing engine and the communication engines, enabling internal data forwarding in low-port mode.

[0098] As shown in Figure 4, the 11 ports may include 9 ports for building an enhanced surround network topology (such as the first topology network described above) and 2 ports for building a fully interconnected network topology. The 9 ports include 3 ports each in the X dimension, Y dimension, and Z dimension. Each dimension distinguishes between positive and negative directions and a crossover direction. Ports LinkX+ and LinkX- are respectively defined to be responsible for interconnecting adjacent nodes in the X dimension, LinkX is responsible for direct connection between non-adjacent nodes in the X dimension (direct connection, not passing through other nodes), LinkY+ and LinkY- are responsible for interconnecting adjacent nodes in the Y dimension, LinkY is responsible for direct connection between non-adjacent nodes in the Y dimension, LinkZ+ and LinkZ- are responsible for interconnecting adjacent nodes in the Z dimension, and LinkZ is responsible for direct connection between non-adjacent nodes in the Z dimension. The bold arrows in Figure 4 are used to identify cross-node connections. The two ports include port L-Link for implementing a first-layer fully interconnected network (such as the aforementioned second topology network) and port G-Link for implementing a second-layer fully interconnected network (such as the aforementioned third topology network), where the L in L-Link can represent local and the G in G-Link can represent global.

[0099] In other embodiments, node 110 may be a routing engine integrated into a computing node. The routing engine may be a routing engine integrated into a processor. The processor may be a graphics processing unit (GPU), a neural network processing unit (NPU), or other accelerator device, or a central processing unit (CPU).

[0100] Figure 5 shows a schematic diagram of the architecture of a routing engine integrated into a processor. The processor may include multiple cores and a last-level cache (LLC). Processors typically employ multiple levels of cache, with each level of cache being faster than the next. When a core needs to access a piece of data or instruction, it first checks the nearest L1 cache. If the data is present, it represents a cache hit; otherwise, it represents a cache miss, requiring the next level of cache to be queried. In some possible implementations, the processor may be a CPU / GPU memory chip integrated with high-bandwidth memory (HBM). The HBM may be a stack of multiple Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM, abbreviated as DDR). After stacking, the stacked chips can be packaged together with the CPU / GPU die to create a large-capacity, high-bit-width DDR array.

[0101] The structure of the routing engine integrated in the processor can refer to the structure of the low-port router in Figure 4. The routing engine may include a communication engine, a crossbar switch matrix, and ports. In the example of Figure 5, the processor integrates or has an 11-port routing engine built-in, requiring only 9 ports to achieve expansion in the three dimensions of X, Y, and Z, where Kx, Ky, and Kz represent the length of each dimension, respectively. Each dimension distinguishes between positive and negative directions and a crossover direction, and defines ports LinkX+ and LinkX- to be responsible for interconnecting adjacent nodes in the X dimension, while LinkX is responsible for direct connections between non-adjacent nodes in the X dimension; LinkY+ and LinkY- to be responsible for interconnecting adjacent nodes in the Y dimension, while LinkY is responsible for direct connections between non-adjacent nodes in the Y dimension; and LinkZ+ and LinkZ- to be responsible for interconnecting adjacent nodes in the Z dimension, while LinkZ is responsible for direct connections between non-adjacent nodes in the Z dimension. In this embodiment, to control the network diameter, the lengths of the X, Y, and Z dimensions are all 4. The remaining two ports, L-Link and G-Link, are used to implement a two-layer fully interconnected network topology.

[0102] The core of the processor may communicate with the routing engine via an internal bus, which may include but is not limited to a high-speed serial component interconnect bus (Peripheral Component Interconnect Express, PCIe) or a Compute eXpress Link (CXL) bus.

[0103] It should be understood that the processor or router or routing engine of the present application may include more or fewer components, functions or modules. In some embodiments, the processor may also include components such as a bus or interface, a processor core, a resonant circuit, etc. for interconnecting with the routing engine, and may also be connected to external devices such as storage devices. For example, the processor of the integrated routing engine may also include a memory management unit (MMU) and a memory controller (MC) for managing and controlling external or extended memory such as DDR. The memory management unit and the memory controller may communicate with the extended DDR via a unified bus (UB). Similarly, the processor may also be connected to or extended with a solid-state drive (SSD), wherein the processor set may communicate with the SSD via a UB bus or a Ctrl bus. In addition, the processor may also be connected to or extended with devices such as accelerators. In some embodiments, the router or routing engine may be any device for implementing data transmission, such as a network device such as a switch.

[0104] It should be noted that Figures 4 and 5 are both illustrated using 11 ports as an example. In other possible implementations of the embodiments of the present application, the routing engine integrated in the low-port router or processor may also include other numbers of ports, such as 10 ports, 12 ports, 13 ports, etc. Taking the 10-port example, the routing engine integrated in the low-port router or processor can build an enhanced surround network topology using 9 of the 10 ports, and the remaining 1 port can build a one-layer fully interconnected network topology.

[0105] In order to make the technical solution of the present application clearer and easier to understand, the following example illustrates how to construct an Enhanced Torus using an 11-port router or routing engine, and then construct a hierarchical network with a 2-level full mesh nested within the Enhanced Torus.

[0106] The Torus topology can be expanded based on low-port routers, offering low networking costs and multipath capabilities. However, for large-scale networks, the network diameter is large, and communication latency increases linearly with network size, resulting in a congested network. Because some applications (such as cloud data centers) are characterized by localized communication, communication traffic is primarily concentrated within the cabinet, with 80% of traffic concentrated within the cabinet. This communication model is a natural fit for the Torus topology. For 3D Torus networks, communication routing can be provided in six directions—X, Y, and Z—effectively supporting near-neighbor communication.

[0107] However, the communication delay of the Torus topology scales linearly with the network scale. The larger the network scale, the more severe the congestion. In particular, it is not friendly to global uniform random traffic of the all-to-all type. Among them, all-to-all can be the data on all XPU cards (a general term for CPU, GPU or NPU) transposed to all XPU cards. For example, in the scenario of model training using model parallelism, all-to-all type global uniform random traffic is usually generated. A standard cabinet can usually accommodate 64 servers and a 4*4*4 3DTorus topology can be constructed. The farthest communication distance within the cabinet is 6 hops. There is a certain amount of congestion and the performance is not optimal.

[0108] Therefore, this embodiment can add a cross-link (a link connecting nodes). With only three ports per dimension, each of the four nodes in the dimension can have a direct link, allowing any two nodes to be connected in a single hop. The maximum intra-cabinet communication distance is only three hops, reducing the communication distance by 50% compared to a standard Torus, effectively improving network performance.

[0109] Furthermore, since the main traffic of a data center (such as a cloud data center) is concentrated inside the cabinet, communication between cabinets does not require a large bandwidth. Moreover, the communication delay of the Torus topology scales linearly with the network scale, so the tours topology is not suitable for interconnection between cabinets. The full mesh topology can provide a higher bisection bandwidth, where the bisection bandwidth refers to the total data rate when each processor is sending, which can usually be the number of processors multiplied by the bandwidth of the connection multiplied by the number of simultaneous transmissions that a processor can execute. For uniform random traffic, the communication efficiency of the full mesh topology is the same as that of the fat tree topology, and can also meet 100% throughput. In addition, it can save 50% of the global optical fiber links compared to the fat tree topology, greatly reducing network costs and improving network reliability. Therefore, this embodiment adopts a nested full mesh topology for expansion between cabinets, which can effectively reduce the network diameter, reduce communication delay, and effectively reduce congestion problems caused by global communication.

[0110] To meet the networking needs of 500,000 nodes in a hyperscale data center, Torus's architecture, nested in a layer of full mesh topology, can provide 570,000 ports, which can meet the networking needs of 500,000 nodes. However, the internal network scale of Torus is too large, and the network diameter is too large. For example, the network diameter can reach 29 hops. To meet the communication needs of some applications (such as cloud data center applications), this embodiment can also use a nested two-level full mesh topology to compress the Torus network scale. By increasing the network level, the network diameter is reduced, communication latency is reduced, communication performance is improved, and network throughput is increased. For example, Torus's nested two-layer full mesh topology can build a 17.3 million node network based on an 11-port router (access port is not considered here to avoid confusion), with a network diameter of only 15 hops. The use of the routing engine integrated in the processor during networking eliminates the need for switch interconnection, further reducing physical costs and power consumption.

[0111] The following describes an example of building an Enhanced Torus and a hierarchical network with a two-level full mesh nested within an Enhanced Torus, with reference to the accompanying figures.

[0112] As shown in Figure 6, the Torus topology is first used to construct the Layer 1 topology. Specifically, each processor has an 11-port routing engine built into it, with nine of these 11 ports used for expansion in the X, Y, and Z dimensions. Kx, Ky, and Kz represent the length of each dimension, indicating the number of network nodes in that dimension. Each dimension distinguishes between positive and negative directions and a spanning direction. Ports LinkX+ and LinkX- are defined to interconnect adjacent nodes in the X dimension, while LinkX directly connects non-adjacent nodes in the X dimension. LinkY+ and LinkY- are defined to interconnect adjacent nodes in the Y dimension, while LinkY directly connects non-adjacent nodes in the Y dimension. LinkZ+ and LinkZ- are defined to interconnect adjacent nodes in the Z dimension, while LinkZ directly connects non-adjacent nodes in the Z dimension. To control the network diameter, the lengths of the X, Y, and Z dimensions are all set to 4. Therefore, the Layer 1 topology can connect 4*4*4=64 nodes. Enhanced spanning links directly connect non-adjacent nodes, making each dimension reachable within a single hop, resulting in a network diameter of only three hops within the cabinet.

[0113] Next, referring to Figure 7 , the 10th port of each routing engine is defined as an L-Link, used for interconnection between L1-level topology groups (referred to as L1 groups, for example, the first node group). Since there are 64 L-Links within a cabinet, a maximum of 65 cabinets can be connected in a full mesh. Therefore, an L2-level topology group (referred to as L2 groups, for example, the second node group) can be composed of 65 L1 groups, connecting 64*65=4160 nodes, with a network diameter of 7 hops. Finally, the 11th port, a G-Link, is defined for interconnection between L2 groups. Since there are 4160 G-Links within an L2 group, a maximum of 4161 L2 groups can be connected in a full mesh. Therefore, an L3-level topology group (referred to as L3 groups, for example, the L3 topology) can be composed of 4161 L2 groups (L2-level topology), connecting 4160*4161=17,309,760 nodes, with a network diameter of only 15 hops.

[0114] In this embodiment, the routing engine only requires 11 external ports to meet the networking requirements of 17.3 million ports. Furthermore, the integrated routing engine in the processor enables the construction of direct-connect networks, eliminating the need for switch interconnection, significantly reducing network costs and power consumption. Nine ports are used for L1 Enhanced Torus interconnection, the tenth port for L2 full mesh interconnection, and the eleventh port for L3 full mesh interconnection. An internally integrated 11x11 low-port crossbar satisfies internal data forwarding requirements, and the communication engine implements Direct Memory Access (DMA) operations, enabling data read / write, memory copy, and other functions.

[0115] It should be noted that Figures 6 and 7 illustrate an example of providing 9 ports out of 11 to construct a 3D torus topology, and the remaining 2 ports to construct a two-layer full mesh. In other possible implementations of the embodiments of the present application, the routing engine integrated in the low-port router or processor may include other numbers of ports.

[0116] Specifically, each network node includes n ports, m of which are used to construct network nodes in the enhanced surround network topology. The m ports include ports for connecting upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. The above Figures 6 and 7 illustrate an example of constructing a 3D enhanced surround network topology using m = 9. In actual applications, the enhanced surround network topology can also be a 2-dimensional or other dimensional enhanced surround network topology. For example, when the enhanced surround network topology is 2-dimensional, m can also be equal to 6.

[0117] In some possible implementations, h of the n ports are used to connect to network nodes in a fully interconnected network topology structure formed by a second node group. The second node group includes a second number of first node groups. Given a limited number of ports, a maximum value of the second number can be determined based on the number of ports used to connect to network nodes in the fully interconnected network topology structure formed by the second node group and the number of nodes in the first node group. The number of nodes in the first node group can be determined based on the number of nodes in each dimension of the enhanced surround network topology structure.

[0118] Based on this, the second quantity can satisfy the following formula:

[0119] Among them, count 2nd represents the second quantity, h represents the number of ports used to connect network nodes in the fully interconnected network topology structure formed by the second node group, Len j The number of nodes in the jth dimension of the enhanced surround network topology (also called the length of the jth dimension) represents the number of dimensions of the enhanced surround network topology. represents the number of nodes in the first node group, for example, the first number. h may be greater than or equal to 1, so that different first node groups (or first topology networks) include at least one direct global link.

[0120] For example, when the enhanced surround network topology includes three dimensions, namely, X, Y, and Z, the maximum second number can be h·Kx·Ky·Kz+1. In the example of Figure 6, Kx, Ky, and Kz can take the values ​​of 4, and h takes the value of 1. Therefore, the maximum second number can be 1*64+1=65. It should be noted that in actual applications, h can also be a positive integer greater than 1, such as 2 or 3. This can increase the number of links between network nodes and reduce the blocking probability, or it can connect more first node groups to increase the network scale.

[0121] In some possible implementations, k of the n ports of a network node are used to connect to network nodes in a fully interconnected network topology structure formed by a third node group. The third node group includes a third number of second node groups. Given the limited number of ports, the maximum value of the third number can be determined based on the number of ports used to connect to network nodes in the fully interconnected network topology structure formed by the second node group and the number of nodes in the second node group. The number of nodes in the second node group can be determined based on the second number (the number of the first node group in the second node group) and the first number (the number of nodes in the first node group), for example, the product of the first number and the second number.

[0122] Based on this, the third quantity can satisfy the following formula: count3rd =k·count 2nd count 1st +1 (2)

[0123] Among them, count 1st Indicates the first quantity, count 2nd Indicates the second quantity, count 3rd represents the third number, and k represents the number of ports used to connect network nodes in the fully interconnected network topology structure formed by the third node group. 2nd count 1st It can be the number of nodes in the second node group. k can be greater than or equal to 1, so that different second node groups (or second topology networks) include at least one direct global link.

[0124] When the enhanced surround network topology includes three dimensions, X, Y, and Z, the second number can be h·Kx·Ky·Kz+1, and the third number can be k·(h·Kx·Ky·Kz+1)·(Kx·Ky·Kz)+1. In the example of Figure 7 , Kx, Ky, and Kz can be 4, and h and k can be 1. Therefore, the maximum third number can be 1*(1*64+1)*64+1=4161. Figure 7 uses h and k as 1 for illustration. In actual applications, h or k can also be greater than 1, which can increase the number of links between network nodes and reduce the blocking probability.

[0125] Based on the aforementioned embodiment, the node coordinates of the network node can be determined based on the number of the network node in each dimension of the enhanced surround network topology and the number of the first node group to which the network node belongs. For example, when the data routing system is a hierarchical network with a nested enhanced surround network topology and a layer of fully interconnected network topology, the node coordinates can be represented by the number in each dimension of the enhanced surround network topology and the number of the first node group to which the network node belongs. For example, when the enhanced surround network topology is a 3D Enhanced Torus topology, the node coordinates can be represented as {L, X, Y, Z}. Among them, L represents the number of the first node group to which the network node belongs, such as the number of the L1 group. X, Y, Z represent the numbers in each dimension of the enhanced surround network topology to which the network node belongs.

[0126] Furthermore, the data routing system also includes a third topology network. In this case, the node coordinates of the network node are determined based on the number of the network node in each dimension of the enhanced surround network topology structure, the number of the first node group to which the network node belongs, and the number of the second node group to which the network node belongs. For example, when the data routing system is a hierarchical network with nested enhanced surround network topology structures and two-layer fully interconnected network topology structures, the node coordinates can be represented by the number in each dimension of the enhanced surround network topology structure and the number of the first node group to which the network node belongs, and the number of the second node group to which the network node belongs. For example, when the enhanced surround network topology structure is a 3D Enhanced Torus topology, the node coordinates can be represented as {G, L, X, Y, Z}. Among them, G represents the number of the second node group to which the network node belongs, such as the number of the L2 group, and L represents the number of the first node group to which the network node belongs, such as the number of the L1 group. X, Y, Z represent the number in each dimension of the enhanced surround network topology structure to which the network node belongs.

[0127] It should be noted that a unified numbering scheme can be used for switching nodes such as integrated routing engines in processors or independent low-port routers and switches. Accordingly, network nodes can determine data routing paths based on the node coordinates represented by the numbers, and implement data routing between computing nodes based on the data routing paths.

[0128] Among them, network nodes can determine data routing paths using the shortest path routing algorithm based on node coordinates. The shortest path routing algorithm provides the shortest distance communication between the source node and the destination node, with the lowest communication latency, and is the most basic routing algorithm. Because this application is a hierarchical network with nested torus and full mesh (for example, a two-layer full mesh), a deterministic shortest path routing algorithm that adapts to topological characteristics has the advantages of rapid computation and is easy to implement in hardware.

[0129] The following describes the specific implementation of data routing using a deterministic shortest path routing algorithm based on adaptive topology features.

[0130] FIG8 shows a flow chart of a data routing method, which can be applied to the data routing system 10 of the aforementioned embodiment. The data routing system 10 is used to implement data routing between computing nodes. The data routing system 10 includes a first topology network 120 and a second topology network 130, wherein the first topology network 120 is nested within the second topology network 130. The network structure of the first topology network 120 is an enhanced surround network topology structure composed of a first node group. The first node group includes a first number of network nodes, such as nodes 110-1 through 110-P. The network nodes in the enhanced surround network topology structure are connected to upstream nodes, downstream nodes, and non-adjacent nodes in each dimension. The network structure of the second topology network 130 is a fully interconnected network topology structure composed of a second node group. The second node group includes a second number of first node groups. There are connections between the second number of first node groups. When routing a data packet from a source node to a destination node, the data routing system 10 may pass through at least one network node. The node where the data packet is currently located is called the current node. The current node can be the source node of the data packet (the source node can also be a network node) or a node on the path from the source node to the destination node. The method includes the following steps:

[0131] S802: The current node receives a data packet to be routed.

[0132] The current node is the source node of the data packet or a node on the path from the source node to the destination node. The packet header may record information about the source node and the destination node of the data packet, such as the node coordinates of the source node, the node coordinates of the destination node, or other information that can be used to determine the node coordinates of the source node and the destination node.

[0133] S804: The current node determines a data routing path according to the node coordinates of the current node and the node coordinates of the destination node.

[0134] The node coordinates are determined based on the node's number within each dimension of the enhanced surround network topology and the number of the first node group to which the node belongs. Furthermore, when data routing topology 10 includes a third topology network, the node coordinates can be determined based on the node's number within each dimension of the enhanced surround network topology, the number of the first node group to which the network node belongs, and the number of the second node group to which the network node belongs.

[0135] The current node can obtain the parsed data packet and obtain the node coordinates of the destination node. In some examples, the node coordinates of the destination node can be {Gd, Ld, Xd, Yd, Zd}, where Gd represents the number of the second node group in which the destination node is located, Ld represents the number of the first node group in which the destination node is located, and Xd, Yd, and Zd represent the numbers of the destination node in each dimension of the enhanced surround network topology. The current node can also obtain the node coordinates of the current node. For example, each network node can store its own node coordinates after the network is completed, so that the current node can read the node coordinates of the current node.

[0136] The current node can determine the data routing path based on the node coordinates of the current node and the node coordinates of the destination node, combined with the topology of the data routing system 10 (for example, a topology map). It should be noted that when the node coordinates of the current node and the destination node represent that the current node and the destination node are in the same layer, the data routing path can be determined based on the topology of the layer. For example, when the current node and the destination node are both in the same first topology network (or the same first node group) of the first layer, the current node can determine the data routing path based on the topology of the layer, for example, the topology of the corresponding node group in the layer. For another example, when the current node and the destination node are in different node groups of the same layer, the current node can determine the data routing path based on the topology of the layer.

[0137] Based on this, the current node can first identify whether the node groups (such as the first node group and the second node group) where the current node and the destination node are located are the same based on the node coordinates of the current node and the node coordinates of the destination node, and then determine the data routing path based on the distribution of the node groups where the current node and the destination node are located.

[0138] S806: The current node routes the data packet according to the data routing path.

[0139] Specifically, the data routing path may indicate a routing port or the node coordinates of a next-hop node. Accordingly, the current node may route the data packet to the next-hop node according to the routing port indicated by the data routing path or the node coordinates of the next-hop node.

[0140] Based on the above description, the present application provides a data routing method, which introduces the node coordinates of the current node and the destination node in the data routing system, adapts the topological characteristics of the data routing system to perform data routing, provides the shortest distance communication between the source node and the destination node, and achieves low communication latency. Moreover, the method is simple to implement and has high availability.

[0141] Next, the specific implementation of determining the data routing path when the current node and the destination node are in node groups at different levels in the embodiment of FIG8 will be described.

[0142] In a first implementation, when the number of the first node group to which the current node belongs is equal to the number of the first node group to which the destination node belongs, the current node and the destination node are in the same first node group (L1 group), and data routing is intra-L1 group routing. Furthermore, when the current node and the destination node are in the same first node group, they are necessarily also in the same second node group, and the number of the second node group to which the current node belongs is equal to the number of the second node group to which the destination node belongs. Based on this, the current node can determine the routing port within each dimension of the enhanced surround network topology based on the offset between the numbers of the destination node and the current node within each dimension.

[0143] The number of network nodes in the i-th dimension of the enhanced surround network topology structure can be L i The current node determines the routing port in the i-th dimension according to the offset value of the numbers of the destination node and the current node in the i-th dimension of the enhanced surround network topology structure, which may include the following situations:

[0144] In the first case, when the offset value of the number of the destination node and the current node in the i-th dimension of the enhanced surround network topology is 1 or -L i +1 indicates that the destination node is the downstream node of the current node. (When the current node is the tail node, the current node can determine that the routing port in the i-th dimension is the port for connecting to the downstream node, also known as the forward routing port.)

[0145] In the second case, when the offset value of the number between the destination node and the current node in the i-th dimension of the enhanced surround network topology is -1 or L i -1, the current node can determine that the routing port in the i-th dimension is the port used to connect to the upstream node, also known as the negative routing port.

[0146] In the third case, when the offset value of the number between the destination node and the current node in the i-th dimension of the enhanced surround network topology is other than 1, -L i or -1 or L i For values ​​other than , the current node may determine that the routing port in the i-th dimension is a port for connecting to non-adjacent nodes, such as a port for cross-node connection (connecting to non-adjacent nodes).

[0147] In the second implementation method, when the number of the second node group where the current node is located is equal to the number of the second node group where the destination node is located, and the number of the first node group where the current node is located is not equal to the number of the first node group where the destination node is located, it means that the current node and the destination node are in different first node groups of the same second node group, and the data routing is L2 group intra-group routing (L1 group inter-group routing).

[0148] Accordingly, the current node can determine whether there is a direct link (such as L-LINK) between the current node and the first node group where the destination node is located. In order to distinguish it from the direct link (such as G-LINK) or other direct links of the second node group where the destination node is located, this application may refer to the direct link with the first node group where the destination node is located as the first direct link.

[0149] If so, the current node can determine that the data routing path includes forwarding the data packet along the first direct link to the first node group where the destination node is located. If not, the current node can determine that the data routing path includes forwarding the data packet to a second jump node. The second jump node has a second direct link to the first node group where the destination node is located. When the data packet arrives at the second jump node, it can reach the first node group where the destination node is located via the second direct link. If the data packet reaches the first node group where the destination node is located, intra-L1 Group routing can be performed.

[0150] In a third implementation, when the number of the second node group where the current node is located is not equal to the number of the second node group where the destination node is located, it indicates that the current node is in a different second node group, and the data routing is L3 intra-group routing (L2 inter-group routing).

[0151] Accordingly, the current node can determine whether a third direct link exists between the current node and the second node group where the destination node resides. If so, the current node can determine that the data routing path includes forwarding the data packet along the third direct link to the second node group where the destination node resides. When the data packet arrives at the second node group where the destination node resides, intra-L2 group routing can be performed. If not, the current node can determine that the data routing path includes forwarding the data packet to the first jump node. A fourth direct link exists between the first jump node and the second node group where the destination node resides.

[0152] Furthermore, the current node can also determine whether the number of the first node group where the first jump node is located is equal to the number of the first node group where the current node is located. If not, it means that the current node and the first jump node are in different first node groups. The current node determines that the data routing path also includes forwarding the data packet to the ferry node. Among them, there is a fifth direct link between the ferry node and the first node group where the first jump node is located. That is, the current node can first route the data packet to the ferry node. The ferry node and the above-mentioned first jump node are in the same first node group, and the data packet can be routed to the above-mentioned first jump node through L1 group routing. There is a direct link between the first jump node and the second node group where the destination node is located, such as the above-mentioned fourth direct link. When the data packet reaches the first jump node, it can be routed to the second node group where the destination node is located through the above-mentioned fourth direct link. When the data packet reaches the first node group where the destination node is located, L2 group routing can be performed.

[0153] For ease of description, this application uses the topology of Figure 7 to illustrate the data routing method. In the topology of Figure 7, the node coordinates of the network node can be {G, L, X, Y, Z}, where G represents the number of the L2 group where the network node is located (for example, the number of the second node group where the network node is located), and the range can be 0 to 4160 (including endpoint values), L represents the number of the L1 group where the network node is located (for example, the number of the second node group where the network node is located), and the range can be 0 to 64 (including endpoint values), and X, Y, and Z respectively represent the numbers of the network node in the three dimensions of Enhanced Torus, and the range can be 0 to 3 (including endpoint values). Among them, the node coordinates (node ​​number) correspond one-to-one to the topological position of the node.

[0154] Ports 1 and 2 of the routing engine serve as routing ports in the forward and reverse directions of the X dimension (port1: LinkX+; port2: LinkX-), port 3 serves as a direct routing port between non-adjacent nodes in the X dimension (port3: LinkX), ports 4 and 5 of the routing engine serve as routing ports in the forward and reverse directions of the Y dimension (port4: LinkY+; port5: LinkY-), port 6 serves as a direct routing port between non-adjacent nodes in the Y dimension (port6: LinkY), ports 7 and 8 of the routing engine serve as routing ports in the forward and reverse directions of the Z dimension (port1: LinkZ+; port2: LinkZ-), and port 9 serves as a direct routing port between non-adjacent nodes in the Z dimension (port9: LinkZ). The positive interfaces of adjacent nodes in each dimension are connected to the negative interfaces of the other end, and these interfaces are connected in sequence to form a Torus ring network. Non-adjacent nodes are directly connected through enhanced ports, such as direct routing ports between non-adjacent nodes, enabling one-hop reach between any nodes in the dimension (node ​​0 connects to node 2; node 1 connects to node 3), thus constructing an L1-level 3D Enhanced Torus. Port 10 of the routing engine serves as the interconnection port between each L1-level Torus in the L2-level topology and is defined as an L-Link. The L-link number (denoted as PL) of each routing engine can correspond one-to-one with the Torus coordinates, for example, PL = X*16+Y*4+Z, with PL values ​​ranging from 0 to 63.

[0155] The following describes the global link connection relationship within the L2 topology. The node coordinates of the source node can be {Gs, Ls, Xs, Ys, Zs}, where s represents source. The node coordinates of the destination node can be {Gd, Ld, Xd, Yd, Zd}, where d represents destination. The node coordinates of the current node can be {Gc, Lc, Xc, Yc, Zc}, where c represents current. The L-Link port of the source node (for example, port 10) is connected to the L-Link port of the destination node. The full mesh connection rule between L2 topologies can be: Ls + PLs = Ld && PLd + PLs = 63 (3)

[0156] Wherein, Ls is the L2 topology number of the source node. The L2 topology can include multiple L1 groups. Based on this, Ls can be the number of the L1 group where the source node is located. Ld is the L2 topology number of the destination node, for example, the number of the L1 group where the destination node is located. PLs is the number of the L-Link port connecting the source node to the L1 group where the destination node is located. PLd is the number of the L-Link port connecting the destination node to the L1 group where the source node is located. Based on this connection relationship, the routing relationship between L1 groups can be calculated. Wherein, the source node and the destination node can be processors with integrated routing engines, or low-port routers or low-port switches, such as low-port routers or low-port switches connected to computing nodes.

[0157] Next, the global link connection relationship within the L3 topology is described. A port of the routing engine (for example, port 11) is used as the interconnection port between each L2 full mesh topology (L2 group) within the L3 group and is defined as a G-Link. The number of each routing engine's G-Link corresponds one-to-one to the coordinates of each L2 full mesh and the Torus within the group. The G-Link numbering rule can be: PG = 64*L+X*16+Y*4+Z, where the PG value range can be from 0 to 4159 (including the endpoint values). The full mesh connection relationship between L3 topologies can be: Gs+PGs=Gd && PGd+PGs=4159 (4)

[0158] Gs is the L3 topology number of the source node. An L3 topology can include multiple L2 groups. Therefore, Gs can be the number of the L2 group in which the source node resides. Gd is the L3 topology number of the destination node, for example, the number of the L2 group in which the destination node resides. PGs is the number of the G-Link port connecting the source node to the L2 group in which the destination node resides, and PGd is the number of the G-Link port connecting the destination node to the L2 group in which the source node resides. Based on this connection relationship, the routing relationship between L2 groups can be calculated.

[0159] Referring to the flowchart of data routing in a data routing system shown in FIG9 , the current node may first determine whether the destination node and the current node are in the same L2 group. If not, L2 group inter-group routing (or L3 group intra-group routing) is performed. If so, continue to determine whether the destination node and the current node are in the same L1 group. When the destination node and the source node are in the same L1 group, for example, the same cabinet, L1 group intra-group routing is performed. When the destination node and the source node are in different L1 groups, specifically different L1 groups of the same L2 group, L1 group inter-group routing (or L2 group intra-group routing) may be performed.

[0160] The following describes the specific implementations of L3 group intra-routing, L2 group intra-routing, and L1 group intra-routing.

[0161] If the destination node and the current node are in different L2 groups, that is, Gd=!Gc, the first jump node with a direct link to the L2 group where the destination node {Gd, Ld, Xd, Yd, Zd} is defined, and the intermediate node (first jump node) to which the port belongs can be calculated based on the direct link port connected to the L2 group where the destination node is located. The node coordinates (node ​​number) of the first jump node are Gm=Gd-PGm&&PGm=4159-PGd, where Gm=Gc=Gd-4159+Ld*64+Xd*16+Yd*4+Zd&&PGm=4159-Ld*64-Xd*16-Yd*4-Zd. It should be noted that this example uses Kx, Ky, and Kz as 4. When Kx, Ky, and Kz take other values, the node coordinates of the first jump node can be determined by the following formula:

[0162] The current node can first determine whether there is a direct link from the current node to the L2 group (group number Gd) where the destination node is located. The current node can determine whether the following formula is satisfied to determine whether there is a direct link from the current node to the L2 group where the destination node is located:

[0163] Assuming Kx, Ky, and Kz are still 4, the current node can determine whether PGc = 64*Lc + Xc*16 + Yc*4 + Zc satisfies 4159 - 64*Ld - Xd*16 - Yd*4 - Zd. If so, a direct link exists from the current node to the destination node's L2 group. If not, no direct link exists from the current node to the destination node's L2 group.

[0164] If there is a direct link from the current node to the destination node's L2 group, the current node can forward the data packet directly from the current node's G-Link port to the destination node's L2 group. Otherwise, the current node can first route the data packet to the first jump node. The node coordinates of the first jump node are {Gc, Lm, Xm, Ym, Zm}, where Lm, Xm, Ym, and Zm can be determined based on PGm. For details, see the above formula (5).

[0165] Furthermore, the current node can also determine whether the L1 group where the first jump node is located is the same as the L1 group where the current node is located. For example, the current node can determine whether Lm is equal to Lc. If the L1 group where the first jump node is located is different from the L1 group where the current node is located, that is, Lm=!Lc, then the route must first be made to the L1 group where the first jump node is located (numbered Lm). The node in the L1 group where the current node is located that has a direct link to the L1 group where the first jump node is located is defined as a ferry node. The node coordinates of the ferry node can be {Gn, Ln, Xn, Yn, Zn}. Therefore, the route must first be made to the ferry node, and then forwarded from the direct link L-Link to the L1 group where the first jump node is located. The physical location of the ferry node can be obtained according to the following formula: Gn=Gc, Ln=Lc, PLn=Kx*Ky*Kz-1-Xm*Ky*Kz-Ym*Kz-Zm (7)

[0166] Still taking Kx, Ky, and Kz as 4 as an example, PLn=63-Xm*16-Ym*4-Zm.

[0167] Based on the above physical location, the coordinates of the ferry node in Torus are: Xn=PLn / (Ky*Kz); Yn=PLn / Ky; Zn=PLn%Kz (8)

[0168] Among them, when Kx, Ky, and Kz are 4, Xn=PLn / 16; Yn=PLn / 4; Zn=PLn%4.

[0169] Based on this, the current node can turn to L1 group intra-group routing, specifically Torus internal routing, and perform routing in the three dimensions of X, Y, and Z according to the offset value of the node coordinates (such as the offset between port numbers). After arriving at the ferry node, the data packet is forwarded from the direct link L-Link port, and then the data packet can be reached from the direct link to the L1 Group where the first jump node is located. After the data packet arrives at the L1 group where the first jump node is located, it can reach the first jump node through the L1 group intra-group routing. The first jump node can reach the L2 group where the destination node is located through the direct link. The data packet that arrives at the L2 group where the destination node is located can reach the destination node through the L2 group intra-group routing.

[0170] The following describes L2 group routing in detail.

[0171] The current node can determine whether it and the destination node are in the same L2 group based on their node coordinates. If the destination node and the current node are in the same L2 group, that is, Gd = Gc, but Ld = !Lc, the current node can determine whether it has a direct link to the destination node's L1 group. If so, the packet is forwarded directly over this direct link.

[0172] From the current node {Gc, Lc, Xc, Yc, Zc} to the destination node {Gd, Ld, Xd, Yd, Zd}, according to the topological connection relationship, if the current node is a direct link connected to the L1 group where the destination node is located, the data packet can be forwarded directly from the direct link. For example, Lc = Ld-PLc&&PLc = 63-PLd. That is, Lc = Ld-63+Xd*16+Yd*4+Zd = Xc*16+Yc*4+Zc, indicating that there is a direct link from the current node to the L1 group where the destination node is located. Then, directly forwarding from the L-Link port can reach the L1 group where the destination node is located. Otherwise, define the node with a direct link to the L1 group where the destination node is located as the second hop node, and its node coordinates can be:

[0173] Among them, when Kx, Ky, and Kz are 4, Gm=Gc, Lm=Ld-PLm&&PLm=63-PLd, and PLd=Xd*16+Yd*4+Zd.

[0174] Therefore, routing is first performed to the second hop node: Xm = PLm / (Ky*Kz); Ym = PLm / Kz; Zm = PLm%Kz. Then, intra-L1 group routing is performed, specifically routing in the X, Y, and Z dimensions based on the offset values ​​of the node coordinates. Upon reaching the second hop node, the packet can be forwarded from port PLm = 63 - PLd = 63 - Xd*16 - Yd*4 - Zd, directly from the link to the L1 group where the destination node resides. Intra-L1 group routing can then be performed, for example, dimension-ordered routing within the same cabinet. Dimension-ordered routing means that each packet is routed one dimension at a time. Once it reaches the appropriate coordinates in that dimension, it is routed in the next dimension in descending order.

[0175] Next, L1 group intra-routing is described in detail.

[0176] If the destination node and the current node are in the same L2 group and L1 group, that is, Gd == Gc && Ld == Lc, the current node only needs to determine the route based on the offset values ​​of the node coordinates of the three dimensions of the Torus where the destination node is located:

[0177] First, the current node determines the coordinate offset value in the X dimension. For example, if the dimension length is 4 and the node coordinates are 0, 1, 2, and 3, the current node calculates the coordinate offset value of the destination node in the X dimension relative to the current node. If it is +1, it indicates an adjacent downstream node and is forwarded from the X+ port; if it is -1, it indicates an adjacent upstream node and is forwarded from the X- port; if it is +3, it is forwarded from the reverse X- port; if it is -3, it is forwarded from the forward X+ port; and if it is 2, it indicates a non-adjacent node and can be forwarded from the X port.

[0178] Next, the current node determines the coordinate offset in the Y dimension. Given a dimension length of 4 and node coordinates of 0, 1, 2, or 3, the current node calculates the Y coordinate offset of the destination node relative to the current node. If the offset is +1, it indicates an adjacent downstream node and is forwarded from the Y+ port. If it is -1, it indicates an adjacent upstream node and is forwarded from the Y- port. If it is +3, it is forwarded from the reverse Y- port. If it is -3, it is forwarded from the forward Y+ port. If it is 2, it indicates a non-adjacent node and is forwarded from the enhanced port Y.

[0179] Next, the current node determines the coordinate offset in the Z dimension. For a dimension length of 4 and node coordinates of 0, 1, 2, or 3, the current node calculates the coordinate offset of the destination node in the Z dimension relative to the current node. If the offset is +1, it indicates an adjacent downstream node and is forwarded from port Z+; if it is -1, it indicates an adjacent upstream node and is forwarded from port Z-; if it is +3, it is forwarded from the reverse Z-port; if it is -3, it is forwarded from the forward Z+ port; and if it is 2, it indicates a non-adjacent node and is forwarded from port Z.

[0180] Finally, if there is no offset in the Z dimension, the message reaches the destination node, and the communication engine in the routing engine receives the message, completing the communication.

[0181] Considering that both the Torus topology and the full mesh topology have loops, they can create a deadlock risk. Deadlock refers to the cyclical occupation of communication resources (such as buffers). For example, in the Torus topology, each of the four nodes in the X dimension sends data to the next adjacent node: node 0 sends data to node 2, node 1 sends data to node 3, and so on. If all four nodes send data simultaneously, the port resources will be cyclically occupied, which will cause deadlock, communication congestion, and even network paralysis in severe cases. Based on this, the data routing system of the present application can also achieve anti-deadlock through virtual lanes (VL).

[0182] Virtual channels can effectively avoid deadlock. In the data routing system of this application, since the first topology network (e.g., L1-level topology) utilizes an Enhanced Torus architecture, there are direct links between any nodes in each dimension, making them reachable in one hop. Deadlock will not occur, and no virtual channels are required. Full mesh topologies naturally have loops, so multiple virtual channels can be allocated to the ports of the second topology network (e.g., L2 group) to avoid deadlock. Furthermore, when the data routing system also includes a third topology network, multiple virtual channels can also be allocated to the ports of the third topology network to avoid deadlock.

[0183] Specifically, a port in a network node used to connect to a node group in a second topology network is configured with multiple virtual channels, such as a first virtual channel and a second virtual channel. For example, an L-LINK in a network node is configured with the first virtual channel and the second virtual channel. Similarly, a port in a network node used to connect to a node group in a third topology network is configured with multiple virtual channels. For example, a G-LINK in a network node is configured with the third virtual channel and the fourth virtual channel.

[0184] For the global routing of the second topology network, the current node can use the first virtual channel of the port to send data and the second virtual channel of the port to receive data. Similarly, for the global routing of the third topology network, the current node can use the third virtual channel of the port to send data and the fourth virtual channel of the port to receive data.

[0185] For example, an L2 group global route uses virtual channel 0 and initially traverses an L2 global link using virtual channel 1. An L3 group global route uses virtual channel 2 and initially traverses an L3 global link using virtual channel 3. As shown in Figure 10, this virtual channel allocation scheme prevents loops between virtual channels and prevents loop deadlocks between the L2 and L3 layers.

[0186] Based on the aforementioned data routing system 10 and data routing method, the present application also provides a method for constructing a data routing system. The method can be executed by a management device. The following describes the method for constructing a data routing system from the perspective of a management device.

[0187] Referring to the flowchart of a method for constructing a data routing system shown in FIG11 , the method includes the following steps:

[0188] S1102: The management device determines a plurality of network nodes for constructing a data routing system.

[0189] The plurality of network nodes are divided into a second number of first node groups, each of which includes the first number of network nodes. After the plurality of network nodes are physically connected, for example, via a physical link, the management device can determine the plurality of network nodes used to construct the data system through node discovery. The management device can discover new network nodes through a heartbeat mechanism, for example. This is not further described here.

[0190] Furthermore, the plurality of network nodes may be divided into a third number of second node groups, wherein the second node groups may include the second number of first node groups, and the third number of second node groups constitute a third node group.

[0191] S1104. The management device sends a configuration file to multiple network nodes, so that the network nodes in the first node group establish connections with upstream nodes, downstream nodes, and non-adjacent nodes in various dimensions according to the network structure of the data routing system in the configuration file, thereby forming a first topology network whose network structure is an enhanced surround network topology structure, and establish connections between a second number of first node groups, thereby forming a second topology network whose network structure is a fully interconnected network topology structure.

[0192] The network nodes in the first node group can establish connections with upstream nodes, downstream nodes, and non-adjacent nodes in various dimensions according to the network structure of the data routing system in the configuration file. The network connection can be established, for example, by establishing a network connection through a three-way handshake to facilitate data routing. Similarly, the network nodes in the first node group can establish connections between the second number of first node groups according to the network structure of the data routing system in the configuration file. It should be noted that the connection between the node groups can include the connection between the network node in one node group and the network node in another node group. In the second topology network, the second node groups can include at least one connection, for example, the connection between the second node group A and the second node group B can include the connection between node 1 in the second node group A and node 2 in the second node group B. Furthermore, the connection between the second node group A and the second node group B can also include the connection between node 11 in the second node group A and node 12 in the second node group B.

[0193] The data routing system includes a first topology network and a second topology network, with the first topology network nested within the second topology network. Furthermore, when the plurality of network nodes are divided into a third number of second node groups, the configuration file is used to establish connections between the network nodes in the first node group and the third number of second node groups, thereby forming a third topology network having a fully interconnected network topology. Accordingly, the data routing system also includes a third topology network, with the second topology network nested within the third topology network.

[0194] In some possible implementations, the management device may also check the connection relationships of multiple network nodes based on the network structure of the data routing system in a configuration file, such as a network structure represented by a topology diagram. The network structure in the configuration file is the desired network structure, and the connection relationships are the desired connection relationships. The management device may also obtain the actual connection relationships of multiple network nodes through probing. The management device may compare the desired connection relationships with the actual connection relationships to check the connection relationships of multiple network nodes.

[0195] If the connection relationship between multiple network nodes (e.g., the actual connection relationship) is inconsistent with the connection relationship in the network structure of the data routing system, the management device can also send a prompt message to the user. This prompt message can be used to prompt the user to adjust the connection relationship between the multiple network nodes. In this way, the accuracy and reliability of the network can be guaranteed.

[0196] The management device is further configured to distribute a routing table so that the data routing system can perform data routing based on the routing table. Specifically, the management device may generate a routing table based on the node coordinates of multiple network nodes. The node coordinates are determined based on the network node's number within each dimension of the enhanced surround network topology and the number of the first node group to which the node belongs. The management device may then distribute an entry in the routing table related to the at least one network node to at least one of the multiple network nodes.

[0197] The management device sends routing table entries related to at least one network node to reduce storage resources occupied by routing table entries in the network node, thereby saving storage space. It should be noted that if the network node has sufficient storage resources, the management device may also send a complete routing table to at least one network node, and this embodiment does not impose a limitation on this.

[0198] Based on the above description, it can be seen that the method for constructing a data routing system in this application, by determining multiple network nodes used to construct the data routing system and instructing the network nodes to initiate connections according to the network structure of the data routing system in the configuration file, thereby constructing a data routing system with a nested enhanced surround network topology structure and a fully interconnected topology network structure, thereby achieving the formation of large-scale or ultra-large-scale networks, effectively reducing network costs and power consumption, and improving network scalability. Among them, the Enhanced Torus topology has an accelerating effect on the application of local communication features, supports multi-tenant isolation, and accelerates resource pooling; the full mesh topology provides direct routing for global communications, alleviates global network congestion, and can reduce expensive global fiber links compared to the fat tree network topology architecture.

[0199] According to the above embodiments, the hierarchical network proposed in this application that integrates Enhanced Torus and full mesh topologies can effectively reduce the network radius, provide a low-cost, low-power, and highly scalable network, and can meet the needs of multi-tenant isolation, memory pooling, etc. in cloud data centers.

[0200] Table 1 shows the topology network feature comparison results between the embodiment of the present application and three comparison examples for a network scale of 500,000 nodes. Among them, the embodiment of the present application can build a 17.3 million node network based on an 11-port switch, with a network diameter of only 15 hops; compared with the standard 3D Torus network, the corresponding network diameter is as high as 384 hops; even if a fat tree is built based on a 12-port switch, a 9-layer fat tree is required, the network diameter is 17 hops, and the latency is higher than the solution of the present application; compared with the previously proposed Hyper Torus topology, although only an 8-port switch is required to achieve the same scale, the network diameter reaches 27 hops, which affects performance.

[0201] Table 1 Comparison of topological network characteristics

[0202] As can be seen from Table 1, the hierarchical network proposed in this application that integrates Enhanced Torus and full mesh topologies can support large-scale networking and reduce network diameter, thereby reducing network latency and improving communication efficiency.

[0203] Table 2 shows the topology network cost comparison results between the embodiment of the present application and two comparison examples under the same network scale. Among them, one comparison example is built based on a 4-layer fat tree topology network structure, and the other comparison example is built based on a 3D Torus.

[0204] Table 2 Comparison of topological network costs

[0205] As can be seen from Table 2, at the same scale, for example, 17.3 million nodes, the network cost of this application is only 28% of that of a fat tree network; and the power consumption is only 20% of that of a fat tree network. If the routing engine integrated in the processor is used to achieve switch-free direct connection, the number of optical modules is 33% of that of a fat tree network, which can further reduce power consumption and greatly improve system reliability. In terms of CrossBar resource consumption, the Enhanced Torus is 27% of that of a fat tree network. The binary bandwidth of the Enhanced Torus is 50% lower than that of the fat tree network. The Enhanced Torus consumes slightly more power than the standard 3D Torus, but the binary bandwidth is 255 times higher, and the network diameter is only 4% of that of the standard 3D Torus.

[0206] It should be noted that the standard Torus network is difficult to meet the demands of multi-tenant security isolation. Multi-tenancy requires that the network architecture can be flexibly divided according to application needs. For example, referring to the comparative diagram of multi-tenant isolation shown in Figure 12, for 16 servers in a 4*4 configuration, Tenant 1 requires 9 servers, Tenant 2 requires 3 servers, and Tenant 3 requires 4 servers. Tenant 1 is allocated 9 servers in a 3*3 configuration, Tenant 2 is allocated 3 servers in a 1*3 configuration, and Tenant 3 is allocated 4 servers in a 1*4 configuration. Once the standard Torus is divided, the torus architecture is destroyed and degenerated into a mesh structure (torus is a mesh with loopback links added to form a surrounding structure), and the performance of the mesh structure is lower than that of the torus structure. For example, since Tenant 1 has no loopback link resources, it is equivalent to a 3*3 mesh topology, and its performance is inferior to that of the standard 3*3 Torus. Since this application adds enhanced direct links between nodes, resources can be flexibly allocated, which can meet the topological requirements of the torus. As shown on the right side of Figure 12, when resources are allocated according to the aforementioned resource allocation scheme, since there are direct links between any nodes in each dimension, Tenant 1 is a 3*3 torus, Tenant 2 is a 1*3 torus, and Tenant 3 is a 1*4 tours, all of which comply with the torus architecture and have no performance loss. Moreover, internal communication between tenants does not need to pass through the communication traffic of other tenants, achieving traffic isolation and meeting the security isolation requirements of multi-tenant scenarios in cloud data centers.

[0207] It can be seen that the hierarchical network of the present application that integrates Enhanced Torus and full mesh topology can dynamically allocate resources according to application requirements, achieve flexible sharding, and support multi-tenant security isolation.

[0208] It's also important to note that Torus topologies typically involve congested networks. Communication between non-adjacent nodes requires forwarding through intermediate nodes, resulting in competition for link bandwidth resources and impacting application performance. This is particularly true for memory pooling scenarios, where increasing the number of network hops increases communication latency, impacting application performance. For ease of illustration, a 4x4 = 16-node resource pool is used as an example. Figure 13 shows a comparative diagram of memory access in a resource pooling scenario. For a standard 4x4 Torus topology, a 3x4 = 12-node compute resource pool and a 1x4 = 4-node memory resource pool can be constructed. A standard Torus topology only has two rows of compute nodes with direct links to the memory pool, enabling direct memory access with minimal latency. To access memory, the compute nodes in the second row of Figure 13 must traverse the links between the nodes in the upper and lower rows. This communication distance is two hops, doubling latency. Furthermore, they conflict with memory access traffic from these two rows, competing for bandwidth resources and reducing concurrent access efficiency by 50%. For the Enhanced Torus topology of this application, since there are direct links between any nodes in each dimension, all computing nodes have direct links to memory nodes, which are reachable in one hop, with low access latency, exclusive link bandwidth, and high communication performance. Compared with the standard Torus, memory resource utilization can be improved by 50%.

[0209] It can be seen that in the resource pooling scenario, the hierarchical network that integrates Enhanced Torus and full mesh topologies in this application can significantly improve memory resource utilization through direct links between non-adjacent nodes.

[0210] Based on the aforementioned method for constructing a data routing system, the present application also provides a management device. The management device will be introduced from the perspective of functional modularization. As shown in FIG14 , the management device 1400 includes:

[0211] a determining module 1402 for determining a plurality of network nodes for constructing the data routing system, wherein the plurality of network nodes are divided into a second number of first node groups, the first node groups including a first number of network nodes;

[0212] The sending module 1404 is used to send a configuration file to the multiple network nodes so that the network nodes in the first node group establish connections with upstream nodes, downstream nodes, and non-adjacent nodes in various dimensions according to the network structure of the data routing system in the configuration file, thereby forming a first topology network whose network structure is an enhanced surround network topology structure, and establish connections between the second number of first node groups to form a second topology network whose network structure is a fully interconnected network topology structure. The data routing system includes the first topology network and the second topology network, and the first topology network is nested in the second topology network.

[0213] Exemplarily, the above-mentioned determination module 1402 and sending module 1404 may be implemented by hardware or software.

[0214] When implemented via software, determination module 1402 and sending module 1404 may be applications or program modules running on a computing device. Applications, etc., may be provided to users via virtualization services. Virtualization services may include virtual machine (VM) services, bare metal server (BMS) services, and container services. VM services may use virtualization technology to create a virtual machine (VM) resource pool on multiple physical hosts to provide VMs to users on demand. BMS services use virtual BMS resource pools on multiple physical hosts to provide BMSs to users on demand. Container services use virtual container resource pools on multiple physical hosts to provide containers to users on demand. A VM is a simulated virtual computer, or logically a single computer. BMS is a scalable, high-performance computing service with computing performance comparable to traditional physical machines and secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service and container service in the above-mentioned virtualization services are only specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.

[0215] When implemented in hardware, determination module 1402 may include at least one computing device, such as a server. Alternatively, determination module 1402 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Transmission module 1404 may be a communication interface, such as a transceiver or a transceiver module.

[0216] In some possible implementations, the plurality of network nodes are divided into a third number of second node groups, the second node groups including the second number of first node groups;

[0217] The configuration file is also used for the network nodes in the first node group to establish connections between the third number of second node groups, forming a third topology network with a fully interconnected network topology structure. The data routing system also includes the third topology network, and the second topology network is nested in the third topology network.

[0218] In some possible implementations, the management device 1400 may further include:

[0219] A checking module 1406, configured to check the connection relationship between the plurality of network nodes according to the network structure of the data routing system in the configuration file;

[0220] The sending module 1404 is further configured to send a prompt message to the user if the connection relationship between the multiple network nodes is inconsistent with the connection relationship in the network structure of the data routing system, wherein the prompt message is used to prompt the user to adjust the connection relationship between the multiple network nodes.

[0221] The checking module 1406 may be implemented by hardware or software.

[0222] When implemented via software, inspection module 1406 may be an application or program module running on a computing device. Applications and the like may be provided to users via virtualization services. Virtualization services may include VM services, BMS services, and container services. When implemented via hardware, inspection module 1406 may include at least one computing device, such as a server. Alternatively, inspection module 1406 may be implemented using an ASIC or a programmable logic device (PLD).

[0223] In some possible implementations, the management device 1400 may further include:

[0224] a generating module 1408 configured to generate a routing table based on node coordinates of the plurality of network nodes, wherein the node coordinates are determined based on numbers of the network nodes in each dimension of the enhanced surround network topology and numbers of the first node group to which the nodes belong;

[0225] The sending module 1404 is further configured to send an entry in the routing table related to the at least one network node to at least one network node among the multiple network nodes.

[0226] The generation module 1408 may be implemented by hardware or software.

[0227] When implemented via software, generation module 1408 may be an application or program module running on a computing device. Applications and the like may be provided to users via virtualization services. Virtualization services may include VM services, BMS services, and container services. When implemented via hardware, generation module 1408 may include at least one computing device, such as a server. Alternatively, generation module 1408 may be implemented using an ASIC or a programmable logic device (PLD).

[0228] The following describes the management device 1400 provided in this application from the perspective of hardware implementation. The management device 1400 can be implemented by a computing device. This application also provides a computing device 1500. As shown in Figure 15, the computing device 1500 includes: a bus 1502, a processor 1504, a memory 1506, and a communication interface 1508. The processor 1504, the memory 1506, and the communication interface 1508 communicate with each other via the bus 1502. The computing device 1500 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1500.

[0229] Bus 1502 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG2 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1502 may include a path for transmitting information between various components of computing device 1500 (e.g., memory 1506, processor 1504, and communication interface 1508).

[0230] The processor 1504 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0231] The memory 1506 may include a volatile memory, such as a random access memory (RAM). The memory 1506 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory 1506 stores executable program code, and the processor 1504 executes the executable program code to implement the aforementioned method of constructing a data routing system. Specifically, the memory 1506 stores instructions for the management device 1400 to execute the method of constructing a data routing system. For example, the memory 1506 may store instructions for the functions of the determination module 1402 and the sending module 1404. Furthermore, the memory 1506 may also store instructions for the functions of the inspection module 1406 and the generation module 1408.

[0232] The communication interface 1508 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1500 and other devices or a communication network.

[0233] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0234] As shown in Figure 16, the computing device cluster includes at least one computing device 1500. The memory 1506 of one or more computing devices 1500 in the computing device cluster may store the same instructions of the management device 1400 for executing the method for building a data routing system.

[0235] In some possible implementations, one or more computing devices 1500 in the computing device cluster may also be used to execute some of the instructions of the management device 1400 for executing the method for building a data routing system. In other words, the combination of one or more computing devices 1500 may jointly execute the instructions of the management device 1400 for executing the method for building a data routing system.

[0236] It should be noted that the memory 1506 in different computing devices 1500 in the computing device cluster may store different instructions for executing part of the functions of the management device 1400 .

[0237] FIG17 illustrates a possible implementation. As shown in FIG17 , two computing devices 1500A and 1500B are connected via a communication interface 1508. The memory in computing device 1500A stores instructions for executing the functions of determination module 1402. The memory in computing device 1500B stores instructions for executing the functions of sending module 1404. In other words, the memories 1506 of computing devices 1500A and 1500B jointly store instructions for management device 1400 to execute the method for constructing a data routing system. The memory 1506 of computing device 1500B may also store instructions for the functions of checking module 1406 and generating module 1408.

[0238] The connection method between the computing device clusters shown in Figure 17 can be considered to be based on the fact that the method for building a data routing system provided by this application requires a large amount of resources for node discovery. To avoid excessive resource consumption by node discovery and thus affecting the distribution of configuration files, it is considered to transfer the functions implemented by the sending module 1404 to the computing device 1500B.

[0239] It should be understood that the functionality of the computing device 1500A shown in FIG17 may also be implemented by multiple computing devices 1500. Similarly, the functionality of the computing device 1500B may also be implemented by multiple computing devices 1500.

[0240] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG. 18 illustrates a possible implementation. As shown in FIG. 18 , two computing devices 1500C and 1500D are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1506 in the computing device 1500C stores instructions for executing the functions of the determination module 1402. Simultaneously, the memory 1506 in the computing device 1500D stores instructions for executing the functions of the sending module 1404. The memory 1506 in the computing device 1500D may also store instructions for executing the functions of the inspection module 1406 and the generation module 1408.

[0241] The connection method between the computing device clusters shown in Figure 18 can be considered to be that the method for building a data routing system provided by this application requires more resources for node discovery, and in order to avoid affecting other businesses, it is considered to entrust the functions implemented by the sending module 1404, the checking module 1406, and the generating module 1408 to the computing device 1500D for execution.

[0242] It should be understood that the functionality of the computing device 1500C shown in FIG18 may also be accomplished by multiple computing devices 1500. Similarly, the functionality of the computing device 1500D may also be accomplished by multiple computing devices 1500.

[0243] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned method applied to the management device 1400 for executing the construction of the data routing system.

[0244] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method for constructing a data routing system.

[0245] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data routing system, characterized in that: The data routing system is used to implement data routing between computing nodes, and the data routing system includes a first topology network and a second topology network, wherein the first topology network is nested in the second topology network; The network structure of the first topology network is an enhanced surround network topology structure composed of a first node group, wherein the first node group includes a first number of network nodes, and the network nodes in the enhanced surround network topology structure have connection relationships with upstream nodes, downstream nodes, and non-adjacent nodes in each dimension; The network structure of the second topology network is a fully interconnected network topology structure composed of second node groups, wherein the second node groups include a second number of first node groups, and a connection relationship exists between the second number of first node groups.

2. The system according to claim 1, wherein: Each network node includes n ports, m of which are used to connect network nodes in the enhanced surround network topology structure, and the m ports include ports for connecting upstream nodes, downstream nodes, and non-adjacent nodes in each dimension.

3. The system according to claim 2, characterized in that H ports out of the n ports are used to connect network nodes in the fully interconnected network topology structure formed by the second node group, and the second number is less than or equal to Len j The jth dimension represents the number of network nodes of the enhanced surround network topology structure, and the q represents the number of dimensions of the enhanced surround network topology structure.

4. The system according to any one of claims 1 to 3, characterized in that The data routing system further includes a third topology network, wherein the second topology network is nested within the third topology network; The network structure of the third topology network is a fully interconnected network topology structure composed of a third node group, wherein the third node group includes a third number of second node groups, and a connection relationship exists between the third number of second node groups.

5. The system according to claim 4, characterized in that K ports among the n ports are used to connect network nodes in the fully interconnected network topology structure formed by the third node group, and the third number is less than or equal to k·the number of nodes in the second node group+1.

6. The system according to any one of claims 1 to 5, characterized in that The network node is a switching node independent of the computing node, or a routing engine integrated into the computing node.

7. The system according to any one of claims 1 to 6, characterized in that The node coordinates of the network node are determined based on the number of the network node in each dimension of the enhanced surround network topology structure and the number of the first node group to which the network node belongs; The network node is used to determine a data routing path based on the node coordinates, and implement data routing between the computing nodes according to the data routing path.

8. The system according to claim 7, characterized in that The data routing system further includes a third topology network, the second topology network is nested in the third topology network, and the network structure of the third topology network is a fully interconnected network topology structure composed of a third node group; The node coordinates of the network node are determined based on the number of the network node in each dimension of the enhanced surround network topology structure, the number of the first node group to which the network node belongs, and the number of the second node group to which the network node belongs.

9. A data routing method, characterized in that: A method for implementing data routing between computing nodes in a data routing system includes a first topology network and a second topology network, wherein the first topology network is nested within the second topology network, wherein the network structure of the first topology network is an enhanced surround network topology structure composed of a first node group, wherein the first node group includes a first number of network nodes, and wherein the network nodes in the enhanced surround network topology structure are connected to upstream nodes, downstream nodes, and non-adjacent nodes in each dimension, respectively; and wherein the network structure of the second topology network is a fully interconnected network topology structure composed of a second node group, wherein the second node group includes a second number of first node groups, and wherein the second number of first node groups are connected to each other. The method includes: A current node in the data routing system receives a data packet to be routed, the current node being a source node of the data packet or a node on a path from the source node to a destination node; The current node determines a data routing path according to node coordinates of the current node and node coordinates of the destination node, wherein the node coordinates are determined based on the number of the node in each dimension of the enhanced surround network topology structure and the number of the first node group to which the node belongs; The current node routes the data packet according to the data routing path.

10. The method according to claim 9, characterized in that The current node determines a data routing path according to the node coordinates of the current node and the node coordinates of the destination node, including: When the number of the first node group where the current node is located is equal to the number of the first node group where the destination node is located, the current node determines the routing port in each dimension based on the offset value of the numbers of the destination node and the current node in each dimension of the enhanced surround network topology structure.

11. The method according to claim 10, characterized in that The number of network nodes in the i-th dimension of the enhanced surround network topology is L i The current node determines the routing ports in each dimension according to the offset values of the numbers of the destination node and the current node in each dimension of the enhanced surround network topology structure, including: When the offset value of the number of the destination node and the current node in the i-th dimension of the enhanced surround network topology structure is 1 or -L i +1, determining that the routing port in the i-th dimension is a port for connecting to a downstream node; When the offset value of the number of the destination node and the current node in the i-th dimension of the enhanced surround network topology structure is -1 or L i -1, determining that the routing port in the i-th dimension is a port for connecting to an upstream node; Otherwise, the routing port in the i-th dimension is determined to be a port for connecting to a non-adjacent node.

12. The method according to claim 9, characterized in that The data routing system further includes a third topology network, wherein the second topology network is nested within the third topology network, wherein the network structure of the third topology network is a fully interconnected network topology structure composed of a third node group, wherein the third node group includes a third number of second node groups, and a connection relationship exists between the third number of second node groups, and the node coordinates are determined based on the number of the node in each dimension of the enhanced surround network topology structure, the number of the first node group in which the node is located, and the number of the second node group in which the node is located; The current node determines a data routing path according to the node coordinates of the current node and the node coordinates of the destination node, including: When the number of the second node group where the current node is located is equal to the number of the second node group where the destination node is located, and the number of the first node group where the current node is located is not equal to the number of the first node group where the destination node is located, the current node determines whether a first direct link exists between the current node and the first node group where the destination node is located; If so, the current node determines that the data routing path includes forwarding the data packet from the first direct link to the first node group where the destination node is located; if not, the current node determines that the data routing path includes forwarding the data packet to a second jump node, and the second jump node has a second direct link with the first node group where the destination node is located.

13. The method according to claim 12, characterized in that The current node determines a data routing path according to the node coordinates of the current node and the node coordinates of the destination node, including: When the number of the second node group where the current node is located is not equal to the number of the second node group where the destination node is located, the current node determines whether a third direct link exists between the current node and the second node group where the destination node is located; If so, the current node determines that the data routing path includes forwarding the data packet from the third direct link to the second node group where the destination node is located; if not, the current node determines that the data routing path includes forwarding the data packet to the first jump node, and there is a fourth direct link between the first jump node and the second node group where the destination node is located.

14. The method according to claim 13, characterized in that The method further comprises: The current node determines whether the number of the first node group where the first jump node is located is equal to the number of the first node group where the current node is located; If not, the current node determines that the data routing path further includes forwarding the data packet to a ferry node, and a fifth direct link exists between the ferry node and the first node group where the first jump node is located.

15. A method for constructing a data routing system, characterized in that: Applied to managing devices, the method includes: determining a plurality of network nodes for constructing the data routing system, the plurality of network nodes being divided into a second number of first node groups, the first node groups including a first number of network nodes; A configuration file is sent to the multiple network nodes so that the network nodes in the first node group establish connections with upstream nodes, downstream nodes, and non-adjacent nodes in various dimensions according to the network structure of the data routing system in the configuration file, thereby forming a first topology network having an enhanced surround network topology structure, and establish connections between the second number of first node groups, thereby forming a second topology network having a fully interconnected network topology structure. The data routing system includes the first topology network and the second topology network, and the first topology network is nested in the second topology network.

16. The method according to claim 15, characterized in that The plurality of network nodes are divided into a third number of second node groups, the second node groups including the second number of first node groups; The configuration file is also used for the network nodes in the first node group to establish connections between the third number of second node groups, forming a third topology network with a fully interconnected network topology structure. The data routing system also includes the third topology network, and the second topology network is nested in the third topology network.

17. The method according to claim 15 or 16, characterized in that The method further comprises: Checking the connection relationship of the plurality of network nodes according to the network structure of the data routing system in the configuration file; If the connection relationship between the plurality of network nodes is inconsistent with the connection relationship in the network structure of the data routing system, a prompt message is sent to the user, where the prompt message is used to prompt the user to adjust the connection relationship between the plurality of network nodes.

18. The method according to any one of claims 15 to 17, characterized in that The method further comprises: generating a routing table according to node coordinates of the plurality of network nodes, wherein the node coordinates are determined based on numbers of the network nodes in each dimension of the enhanced surround network topology structure and numbers of the first node group to which the nodes belong; Sending an entry related to the at least one network node in the routing table to at least one network node among the multiple network nodes.

19. A management device, characterized in that: The management device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions to enable the computing device cluster to execute the method according to any one of claims 15 to 18.

20. A computer-readable storage medium, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 15 to 18.

Citation Information

Patent Citations

  • Method for constructing Torus network, Torus network and routing algorithm

    CN107612746A

  • Hierarchical network-on-chip topology and routing method thereof

    CN109189720A

  • Three-dimensional network topology structure and routing algorithm thereof

    CN109561034A

  • Communicating messages between publishers and subscribers in a mesh routing network

    US20160021043A1

Cited By

  • Multi-layer hexagonal network structure construction method, multi-layer hexagonal network structure path determination method and switch

    CN120880964A

  • Processing method and device of multi-core processor, storage medium and electronic equipment

    CN121579416A