A traffic forwarding method and apparatus

By receiving and parsing routing information in the AI ​​large model network, selecting a reasonable next-hop path, and generating forwarding table entries, the traffic congestion problem is solved, more efficient traffic forwarding is achieved, and the probability of traffic congestion at Spine nodes is reduced.

CN118784573BActive Publication Date: 2026-03-24NEW H3C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traffic congestion is prone to occur in large AI model networks, especially when multiple computing nodes send traffic to the same Leaf node at the same time. Traffic collisions and congestion are likely to occur at the Spine node connection.

Method used

By receiving routing information published by remote leaf nodes, the next-hop list of computing power nodes is determined, and the target next hop is selected according to the routing policy index. A forwarding table entry is generated to ensure that the target next hop of the same computing power node connected to the same remote leaf node is the same, while the target next hop of different computing power nodes connected to the same leaf node is different, thereby optimizing the traffic forwarding path.

Benefits of technology

It effectively reduces the probability of traffic congestion in the AI ​​large model network, avoids traffic congestion on the downlink ports of Spine nodes, and improves the network transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118784573B_ABST
    Figure CN118784573B_ABST
Patent Text Reader

Abstract

The application provides a flow forwarding method and device. In one example, the method comprises: receiving routing information published by a remote leaf node; determining a next hop list corresponding to a computing power node connected to the remote leaf node according to the received routing information; for any computing power node connected to the remote leaf node, selecting a target next hop from the next hop list corresponding to the computing power node according to a routing strategy index of the computing power node; generating a forwarding table entry of the computing power node according to a host route of the computing power node and the target next hop corresponding to the computing power node, and forwarding flow sent to the computing power node according to the forwarding table entry. Application of the embodiment of the application can reduce the probability of flow congestion in an AI large model network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of AI (artificial intelligence) large model and network communication technology, and in particular to a traffic forwarding method and device. BACKGROUND

[0002] AI large model network refers to a computing and communication infrastructure supporting large artificial intelligence model training and operation.

[0003] The AI large model network has the characteristics of periodic fluctuations in traffic and large data traffic, and thus traffic collisions are prone to occur in actual networking.

[0004] For example, when multiple computing power nodes simultaneously send traffic to computing power nodes under the same Leaf node, traffic collisions are prone to occur at the downlink port of the Spine node connected to the Leaf node, resulting in traffic congestion.

[0005] How to reduce the probability of traffic congestion in the AI large model network has become a technical problem to be solved. SUMMARY

[0006] The present application provides a traffic forwarding method and device to solve the problem of traffic congestion prone to occur in the existing AI large model network.

[0007] According to a first aspect of an embodiment of the present application, a traffic forwarding method is provided, comprising:

[0008] receiving routing information published by a remote Leaf node; wherein the routing information includes the host route of the computing power node connected by the remote Leaf node and the routing strategy index of the computing power node;

[0009] determining the next hop list corresponding to the computing power node connected by the remote Leaf node according to the received routing information;

[0010] For any computing power node connected by the remote Leaf node, the target next hop is selected from the next hop list corresponding to the computing power node according to the routing strategy index of the computing power node; wherein for different Leaf nodes, the target next hop corresponding to the same computing power node connected by the same remote Leaf node is the same; for the same Leaf node, the target next hop corresponding to different computing power nodes connected by the same remote Leaf node is different;

[0011] generating the forwarding table entry of the computing power node according to the host route of the computing power node and the target next hop corresponding to the computing power node, and forwarding the traffic sent to the computing power node according to the forwarding table entry.

[0012] According to a second aspect of an embodiment of the present application, a traffic forwarding device is provided, comprising:

[0013] a receiving unit configured to receive routing information published by a remote leaf node; wherein the routing information comprises host routes of computing power nodes connected by the remote leaf node and routing strategy indexes of the computing power nodes;

[0014] a determining unit configured to determine, according to the received routing information, a next hop list corresponding to the computing power nodes connected by the remote leaf node;

[0015] a selecting unit configured to, for any computing power node connected by the remote leaf node, select a target next hop from the next hop list corresponding to the computing power node according to the routing strategy index of the computing power node; wherein for different leaf nodes, the target next hop corresponding to the same computing power node connected by the same remote leaf node is the same; and for the same leaf node, the target next hops corresponding to different computing power nodes connected by the same remote leaf node are different;

[0016] a forwarding control unit configured to generate a forwarding table item of the computing power node according to the host route of the computing power node and the target next hop corresponding to the computing power node, and forward traffic sent to the computing power node according to the forwarding table item.

[0017] According to the technical solution disclosed in the present application, by receiving routing information published by a remote leaf node, determining, according to the received routing information, a next hop list corresponding to computing power nodes connected by the remote leaf node, for any computing power node connected by the remote leaf node, selecting a target next hop from the next hop list corresponding to the computing power node according to the routing strategy index of the computing power node, and generating a forwarding table item of the computing power node according to the host route of the computing power node and the target next hop corresponding to the computing power node, traffic sent to the computing power node can be forwarded according to the forwarding table item. By setting a routing strategy index for the computing power node and selecting a target next hop for the computing power node according to the routing strategy index, for different leaf nodes, the target next hop corresponding to the same computing power node connected by the same remote leaf node is the same, and for the same leaf node, the target next hops corresponding to different computing power nodes connected by the same remote leaf node are different, thereby reducing the probability of traffic congestion in an AI large model network. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is a flowchart of a traffic forwarding method provided by an embodiment of the present application;

[0019] Figure 2 is a schematic diagram of a specific application scenario provided by an embodiment of the present application;

[0020] Figure 3 is a schematic diagram of whole-network topology information maintained by a Leaf node provided by an embodiment of the present application;

[0021] Figure 4 This is a traffic forwarding diagram provided in an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of the structure of a traffic forwarding device provided in an embodiment of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, and to make the above-mentioned objectives, features and advantages of the embodiments of the present invention more apparent and understandable, the technical solutions in the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0024] Please see Figure 1 This is a flowchart illustrating a traffic forwarding method provided in an embodiment of the present invention. This traffic forwarding method can be applied to leaf nodes in a large AI model network based on a Leaf-Spine network architecture, such as... Figure 1 As shown, this traffic forwarding method may include the following steps:

[0025] Step 101: Receive routing information published by the remote leaf node; wherein, the routing information includes the host route of the computing power node connected to the remote leaf node and the routing policy index of the computing power node.

[0026] In this embodiment of the invention, when a Leaf node learns the host routing information of a locally connected computing power node, it can publish the host routing information of that computing power node to a remote Leaf node.

[0027] In this embodiment of the invention, in order to select traffic forwarding links more reasonably, a routing strategy index (also known as an extended strategy index) can be set for any computing power node connected under any Leaf node. This routing strategy index is used to assist in selecting the forwarding link for traffic sent to the computing power node.

[0028] When a Leaf node publishes the host route of a computing node to a remote Leaf node, it can also publish the routing policy index of the computing node to the remote Leaf node.

[0029] Step 102: Based on the received routing information, determine the next-hop list corresponding to the computing power nodes connected to the remote leaf nodes.

[0030] In this embodiment of the invention, the Leaf node can obtain the host route of the computing power node connected to the remote Leaf node, the routing policy index of the computing power node, and the next-hop list corresponding to the computing power node through route resolution based on the routing information published by the remote Leaf node.

[0031] Step 103: For any computing power node connected to a remote leaf node, select the target next hop from the next hop list corresponding to the computing power node according to the routing strategy index of the computing power node; wherein, for different leaf nodes, the target next hop corresponding to the same computing power node connected to the same remote leaf node is the same; for the same leaf node, the target next hop corresponding to different computing power nodes connected to the same remote leaf node is different.

[0032] In this embodiment of the invention, in order to reduce the probability of traffic congestion in the AI ​​large model network, when selecting the next hop (which can be called the target next hop) for the computing power node connected to the remote Leaf node, the selection can be based on the routing strategy index of the computing power node. The next hop selection is based on the principle that for different Leaf nodes, the target next hop corresponding to the same computing power node connected to the same remote Leaf node is the same; and for the same Leaf node, the target next hop corresponding to different computing power nodes connected to the same remote Leaf node is different.

[0033] Since multiple computing nodes typically do not access the same computing node under the same remote Leaf node simultaneously in large AI model networks, ensuring that the target next hop is the same for different Leaf nodes and for different computing nodes under the same remote Leaf node, can effectively avoid downlink port traffic congestion on Spine nodes, provided that the target next hop is different for different computing nodes connected to the same remote Leaf node.

[0034] Step 104: Based on the host route of the computing power node and the target next hop corresponding to the computing power node, generate a forwarding table entry for the computing power node, and forward the traffic sent to the computing power node according to the forwarding table entry.

[0035] In this embodiment of the invention, when a target next hop is selected for the computing power node in the manner described above, a forwarding table entry for the computing power node can be generated based on the host route of the computing power node and the target next hop corresponding to the computing power node.

[0036] Once a forwarding table entry for that computing power node is generated, traffic sent to that computing power node can be forwarded based on that forwarding table entry.

[0037] For example, this forwarding table entry can be sent to the forwarding engine, which can then forward the traffic destined for the computing power node based on this entry.

[0038] It can be seen that, in Figure 1In the illustrated method, routing information published by a remote leaf node is received. Based on the received routing information, the next-hop list corresponding to the computing nodes connected to the remote leaf node is determined. For any computing node connected to the remote leaf node, the target next hop is selected from the next-hop list corresponding to the computing node based on the routing policy index of the computing node. A forwarding table entry for the computing node is generated based on the host route of the computing node and the target next hop corresponding to the computing node. Then, traffic sent to the computing node can be forwarded based on the forwarding table entry. By setting a routing policy index for the computing node and selecting a target next hop for the computing node based on the routing policy index, the target next hop corresponding to the same computing node connected to the same remote leaf node is the same for different leaf nodes, and the target next hop corresponding to different computing nodes connected to the same remote leaf node is different, which reduces the probability of traffic congestion in the AI ​​large model network.

[0039] To enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present invention, the technical solutions provided by the embodiments of the present invention will be described below in conjunction with specific application scenarios.

[0040] Please see Figure 2 This is a schematic diagram illustrating a specific application scenario provided by an embodiment of the present invention, such as... Figure 2 As shown, in this application scenario, Spine nodes 201-203 are all connected via different ports (not in...). Figure 2 (As shown in the image) is connected to Leaf nodes 301-304, and each Leaf node 301-304 is connected via a different port (not shown in the image). Figure 2 (As shown in the diagram) It is connected to Spine nodes 201-203, and each Leaf node is connected to multiple computing nodes through different ports.

[0041] For example, such as Figure 2 As shown, Leaf node 301 is connected to computing nodes 3011 to 3013 through different ports, Leaf node 302 is connected to computing nodes 3021 to 3023 through different ports, and Leaf node 303 is connected to computing nodes 3031 to 3033 through different ports.

[0042] In this embodiment, computing nodes can be represented by GPUs (Graphics Processing Units).

[0043] In this embodiment, a unique global downlink port can be allocated to each GPU on the Spine node for all traffic destined for that GPU, based on the host route corresponding to the destination GPU.

[0044] To achieve the above functionality, for Leaf nodes, the corresponding host routes are generated based on the ARP of the local GPU. When publishing routes, the routing policy index of the local GPU can be carried in the published route information.

[0045] For example, a routing strategy index may include a primary index and sub-indexes.

[0046] The sub-index is used to identify the target port, and the main index is used to identify the Leaf node.

[0047] Optionally, the primary index can be the network device number (the Leaf node number). For example, Leaf node 301 can be numbered 1, Leaf node 302 can be numbered 2, and Leaf node 3 can be numbered 3. The sub-index can be the port number on the Leaf node that connects to the GPU. For example, for Leaf node 301, the port numbers of the ports connecting to GPUs 3011 to 3013 can be 1 to 3 respectively; for Leaf node 302, the port numbers of the ports connecting to GPUs 3021 to 3023 can be 1 to 3 respectively; for Leaf node 303, the port numbers of the ports connecting to GPUs 3031 to 3033 can be 1 to 3 respectively; and for Leaf node 304, the port numbers of the ports connecting to GPUs 3041 to 3043 can be 1 to 3 respectively.

[0048] It should be noted that the GPU's routing strategy index can also be configured manually.

[0049] The host routing and routing policy indexes of GPUs 3011-3013 connected to Leaf node 301 can be shown in Table 1-1.

[0050] Table 1-1

[0051]

[0052]

[0053] The host routing and routing policy indexes of GPUs 3021-3023 connected to Leaf node 302 can be shown in Table 1-2.

[0054] Table 1-2

[0055] Route prefix Route mask Routing policy index 20.0.0.1 32 2:1 20.0.0.2 32 2:2 20.0.0.3 32 2:3

[0056] The host routing and routing policy indexes of GPUs 3031-3033 connected to Leaf node 303 are shown in Table 1-3.

[0057] Table 1-3

[0058] Route prefix Route mask Routing policy index 30.0.0.1 32 3:1 30.0.0.2 32 3:2 30.0.0.3 32 3:3

[0059] The host routing and routing policy indexes of GPUs 3041-3043 connected to Leaf node 304 are shown in Table 1-4.

[0060] Table 1-4

[0061] Route prefix Route mask Routing policy index 40.0.0.1 32 4:1 40.0.0.2 32 4:2 40.0.0.3 32 4:3

[0062] In Tables 1-1 to 1-4, the routing policy index indicates the Leaf node port to which each GPU is connected. A routing index of 2:1 indicates that GPU 3021 is connected to port 2 of leaf node 302.

[0063] When leaf nodes 301-304 publish BGP (Border Gateway Protocol) routes, they can carry the GPU's routing policy index through the BGP private extension community.

[0064] Leaf nodes 301-304 publish OSPF (Open Shortest Path First) routes, which can carry the routing policy index of the GPU through the extended TLV.

[0065] The Leaf node receives routing information published by the remote Leaf node, parses the received routing information, and obtains information such as the host route, routing policy index, and next-hop list of the computing power nodes connected to the remote Leaf node.

[0066] The route information resolved by Leaf node 301 is shown in Table 2.

[0067] Table 2

[0068] Route prefix Route mask Routing policy index Next hop list 20.0.0.1 32 2:1 2.0.0.1,2.0.0.2,2.0.0.3 20.0.0.2 32 2:2 2.0.0.1,2.0.0.2,2.0.0.3 20.0.0.3 32 2:3 2.0.0.1,2.0.0.2,2.0.0.3 30.0.0.1 32 3:1 2.0.0.1,2.0.0.2,2.0.0.3 30.0.0.2 32 3:2 2.0.0.1,2.0.0.2,2.0.0.3 30.0.0.3 32 3:3 2.0.0.1,2.0.0.2,2.0.0.3 40.0.0.1 32 4:1 2.0.0.1,2.0.0.2,2.0.0.3 40.0.0.2 32 4:2 2.0.0.1,2.0.0.2,2.0.0.3 40.0.0.3 32 4:3 2.0.0.1,2.0.0.2,2.0.0.3

[0069] Among them, 30.0.0.1 to 30.0.0.3 are the host routes for GPUs 3031 to 3033 connected to Leaf node 303, respectively. The routing policy index is 3:X, indicating that the GPU is connected through the port numbered X (such as 1, 2, or 3) on Leaf node 3 (i.e., Leaf node 303 mentioned above), which is GPU 3031, GPU 3032, or GPU 3033. The next hops 2.0.0.1, 2.0.0.2, and 2.0.0.3 correspond to Spine nodes 201 to 203, respectively.

[0070] In this embodiment, in addition to maintaining the aforementioned routing information, the Leaf node also maintains its directly connected neighbor information locally.

[0071] For example, taking Leaf node 301 as an example, the information of its directly connected neighbors can be shown in Table 3.

[0072] Table 3

[0073] Neighbor route id Neighbor next hop 1 2.0.0.1 2 2.0.0.2 3 2.0.0.3

[0074] In this embodiment, based on the routing protocol, the Leaf node also maintains the entire network topology information.

[0075] For example, the network topology information maintained by the Leaf node can be as follows: Figure 3 As shown, dashed lines indicate that the link is missing or faulty.

[0076] like Figure 3 As shown, there is a fault in the link between Spine node 201 and Leaf node 302.

[0077] In this embodiment, when routing for the GPU, Leaf nodes 301-304 select routes in the network to reach the GPU accessed by the remote Leaf node based on the same routing strategy.

[0078] For example, taking Leaf node 301 as an example, suppose Leaf node 301 receives a route to 20.0.0.2. The route information includes the prefix 20.0.0.2, the mask 32, the policy index 1:2, and the Leaf1 neighbor list 2.0.0.1, 2.0.0.2, and 2.0.0.3. After sorting according to the preset sorting policy (taking ascending order as an example), the next-hop list is 2.0.0.1, 2.0.0.2, and 2.0.0.3. The sub-index in the policy index is 2, and the corresponding routing policy is to select the second IP address (2.0.0.2) in the sorted neighbor next hop list. If this neighbor IP address exists in the next-hop list, it is selected as the target next hop, a forwarding table entry is generated, and it is sent to the forwarding engine. If this neighbor IP address does not exist in the next-hop list, an alternative target next hop is selected for forwarding. After the route calculation is completed, the relevant table entries are sent to the forwarding engine.

[0079] It should be noted that for scenarios where the port number of the port connecting the computing power node on the Leaf node exceeds the number of neighbor next hops, for example, the number of ports connecting the computing power node on the Leaf node exceeds the number of neighbor next hops, the sorting and sub-index matching can include: taking the sub-index modulo the number of neighbor next hops, and the sorting and modulo results are the same; wherein, when the result of taking the sub-index modulo the number of neighbor next hops is 0, the modulo result is set as the sub-index itself.

[0080] Furthermore, for any computing power node, if the target neighbor next hop is not included in the next hop list corresponding to the computing power node after the target neighbor next hop has been determined in the above manner, a backup target neighbor next hop can be determined and used as the backup target next hop. The specific implementation method can be found in the relevant description of the case where the target next hop is abnormal below. The embodiments of the present invention will not be elaborated here.

[0081] In this embodiment, based on Figure 2 The network topology shown, taking Leaf node 301 as an example, the target next-hop information can be seen in Table 4:

[0082] Table 4

[0083] Route prefix Route mask Routing policy index Next hop 20.0.0.1 32 2:1 2.0.0.1 20.0.0.2 32 2:2 2.0.0.2 20.0.0.3 32 2:3 2.0.0.3 30.0.0.1 32 3:1 2.0.0.1 30.0.0.2 32 3:2 2.0.0.2 30.0.0.3 32 3:3 2.0.0.3 40.0.0.1 32 4:1 2.0.0.1 40.0.0.2 32 4:2 2.0.0.2 40.0.0.3 32 4:3 2.0.0.3

[0084] As shown in Table 4, for the same Leaf node, the target next hops of GPUs connected to the same port number on different remote Leaf nodes are the same. For example, the target next hops of the GPU connected to the port with port number 1 on Leaf node 302 (i.e., GPU3021), the GPU connected to the port with port number 1 on Leaf node 303 (i.e., GPU3031), and the GPU connected to the port with port number 1 on Leaf node 304 (i.e., GPU3041) are the same, all being 2.0.0.1.

[0085] GPUs connected to ports with the same port number on the same remote Leaf node may have different target next hops. For example, the target next hop of the GPU connected to port number 1 on Leaf node 302 (i.e., GPU3021) is 2.0.0.1, while the target next hop of the GPU connected to port number 2 on Leaf node 302 (i.e., GPU3022) is 2.0.0.2.

[0086] based on Figure 2 In the network topology shown, taking Leaf node 302 as an example, the target next-hop information for the remote Leaf node 303 can be seen in Table 5:

[0087] Table 5

[0088] Route prefix Route mask Routing policy index Next hop 30.0.0.1 32 3:1 2.0.0.1 30.0.0.2 32 3:2 2.0.0.2 30.0.0.3 32 3:3 2.0.0.3

[0089] As shown in Tables 4 and 5, for different Leaf nodes, the next hop of the target corresponding to the same GPU connected to the same remote Leaf node is the same.

[0090] For example, for Leaf nodes 301, 302, and 304, the target next hop for GPU 3031 connected to Leaf node 303 is 2.0.0.1.

[0091] For Leaf nodes 301, 302, and 304, the target next hop for GPU 3032 connected to Leaf node 303 is 2.0.0.2.

[0092] For Leaf nodes 301, 302, and 304, the target next hop for GPU 3032 connected to Leaf node 303 is 2.0.0.2.

[0093] For Leaf nodes 301, 302, and 304, the target next hop for GPU 3033 connected to Leaf node 303 is 2.0.0.3.

[0094] Based on the above routing strategy, the target next hops of different computing nodes under the same Leaf node are different on each remote Leaf node side. Since in the AI ​​large model network, there are usually not multiple computing nodes accessing the same computing node under the same remote Leaf node at the same time, the target next hops of GPUs under different Leaf nodes to the same GPU connected to the same remote Leaf node can be set to be the same.

[0095] By implementing the above methods, the probability of traffic congestion at the Spine node can be effectively reduced.

[0096] For example, with Figure 2 Taking the network topology shown as an example, assuming that at a certain moment, GPU 3011 connected to Leaf node 301 sends traffic at line speed to GPU 3031 connected to Leaf node 303, and GPU 3022 connected to Leaf node 302 sends traffic at line speed to GPU 3032 connected to Leaf node 303, then according to the above routing strategy, the next hop for traffic from GPU 3011 to GPU 3031 is 2.0.0.1 (corresponding to Spine node 201), and the next hop for traffic from GPU 3022 to GPU 3032 is 2.0.0.2 (corresponding to Spine node 202). That is, traffic from different remote Leaf nodes accessing different GPUs of the same Leaf node can be forwarded through different Spine nodes, effectively reducing the probability of traffic congestion on the downlink port (the port connecting the Leaf nodes) of the Spine node. The traffic forwarding diagram can be seen as follows: Figure 4 As shown.

[0097] Among them, such as Figure 4As shown, the forwarding path of traffic from GPU3011 to GPU3031 can be represented by the solid arrow in the figure; the forwarding path of traffic from GPU3022 to GPU3032 can be represented by the dashed arrow in the figure.

[0098] In this embodiment, when a link failure occurs, a backup link can be switched over.

[0099] Among them, the sub-index and main index corresponding to the computing power node can be used to select the backup target next hop from the next hop list corresponding to the computing power node.

[0100] If the link connecting Leaf node 301 to Spine node 201 fails, then all traffic sent from Leaf node 301 to the GPUs (such as GPU3021, GPU3031, and GPU3041) connected to port number 1 on each remote Leaf node will need to be switched to a backup link.

[0101] During backup link switching, if all traffic from Leaf node 301 to GPUs connected to port number 1 on each remote Leaf node switches to the same backup link—for example, switching the next hop to the next hop corresponding to Spine node 202—it can easily lead to congestion in the uplink between traffic from Leaf node 301 to GPUs connected to port number 1 on each remote Leaf node and traffic from Leaf node 301 to GPUs connected to port number 2 on each remote Leaf node (such as GPU3022, GPU3032, and GPU3042). Therefore, it is necessary to distribute the traffic from Leaf node 301 to GPUs connected to port number 1 on each remote Leaf node as much as possible.

[0102] Optionally, for the same Leaf node, the next hop of the backup target for the computing power nodes connected to the same port number on different remote Leaf nodes may not be exactly the same.

[0103] When the number of available backup next hops is greater than or equal to the number of computing nodes connected to a single Leaf node, for the same Leaf node, the backup target next hops for computing nodes connected to the same port number on different remote Leaf nodes are different.

[0104] In this embodiment, in the event of a link failure between Leaf node 301 and Spine node 201, the next hop of the backup target can be determined based on the primary and secondary indices in the routing policy index corresponding to the GPU connected to the port with port number 1 on each remote Leaf node.

[0105] Optionally, a new index can be obtained by adding the primary and secondary indexes in the routing strategy index, and the next hop of the alternative target can be determined based on this new index.

[0106] Taking the link failure between Leaf node 301 and Spine node 201 as an example, the information of the GPU of the corresponding remote Leaf node (i.e., the GPU of the target next hop corresponding to Spine node 201) for Leaf node 301 can be shown in Table 6:

[0107] Table 6

[0108] Route prefix Route mask Routing policy index Next hop list 20.0.0.1 32 2:1 2.0.0.2,2.0.0.3 30.0.0.1 32 3:1 2.0.0.2,2.0.0.3 40.0.0.1 32 4:1 2.0.0.2,2.0.0.3

[0109] The directly connected neighbor information maintained by Leaf node 301 can be seen in Table 3.

[0110] For the route prefix 20.0.0.1, the next hop next neighbor that matches the sum of the main index (2) and the sub-index (1) of the routing policy index can be selected from the sorted next hop next neighbors (i.e., 2.0.0.3) as the backup target next hop. Since the backup target next hop next neighbor exists in the next hop list, the backup target next hop next neighbor can be used as the backup target next hop.

[0111] For the route prefix 30.0.0.1, based on the main index (3) and sub-index (1) of the routing policy index, a neighbor next hop that matches the sum of the main index and sub-index (1+3=4) can be selected from the sorted neighbor next hops as the backup target neighbor next hop. Since 4>3 (the number of neighbor next hops), the backup target neighbor next hop can be selected from the sorted neighbor next hops based on the result of taking the modulo of 4 and 3 (i.e., 1), resulting in neighbor next hop 2.0.0.1. This neighbor next hop is the same as the target next hop, and the next neighbor next hop of this neighbor next hop (i.e., 2.0.0.2) is determined as the backup target neighbor next hop. Since this backup target neighbor next hop exists in the next hop list, it can be used as the backup target next hop.

[0112] For the route prefix 30.0.0.1, based on the main index (4) and sub-index (1) of the routing policy index, a neighbor next hop that matches the sum of the main index and sub-index (1+4=5) can be selected from the sorted neighbor next hops as the backup target neighbor next hop. Since 5>3 (the number of neighbor next hops), the backup target neighbor next hop can be obtained from the sorted neighbor next hops based on the result of taking 5 modulo 3 (i.e., 2), resulting in neighbor next hop 2.0.0.2. Since this backup target neighbor next hop exists in the next hop list, it can be used as the backup target neighbor next hop.

[0113] Accordingly, for Leaf node 301, the next hop of the backup target connected to the GPU via port number 1 on each remote Leaf node can be shown in Table 7:

[0114] Table 7

[0115] Route prefix Route mask Routing policy index Next hop 20.0.0.1 32 2:1 2.0.0.1→2.0.0.3 30.0.0.1 32 3:1 2.0.0.1→2.0.0.2 40.0.0.1 32 4:1 2.0.0.1→2.0.0.2

[0116] It should be noted that, in the embodiments of the present invention, when selecting the backup target neighbor next hop, if the selected backup target neighbor next hop is the same as the selected target neighbor next hop (the neighbor next hop corresponding to the abnormal target next hop), the backup target neighbor next hop can be reselected. For example, the next neighbor next hop of the currently selected backup target neighbor next hop can be used as the backup target neighbor next hop.

[0117] Furthermore, if the next hop of the backup target neighbor is not included in the next hop list corresponding to the computing node after the backup target neighbor next hop has been determined in the above manner, the next hop of the backup target neighbor next hop can be selected again. For example, the next hop of the next neighbor of the backup target neighbor next hop can be selected from the next hops of the neighbors sorted according to the preset sorting strategy.

[0118] In this embodiment, for any GPU, if the target next hop is abnormally recovered, a next hop back-cut can be performed.

[0119] Taking Leaf node 301 as an example, assuming that the link between Leaf node 301 and Spine 201 recovers after a failure, the target next hop for the GPU connected to port number 1 on each remote Leaf node after the failure is recovered can be as shown in Table 8:

[0120] Table 8

[0121] Route prefix Route mask Routing policy index Next hop 20.0.0.1 32 2:1 2.0.0.1←2.0.0.3 30.0.0.1 32 3:1 2.0.0.1←2.0.0.2 40.0.0.1 32 4:1 2.0.0.1←2.0.0.2

[0122] Please see Figure 5 This is a schematic diagram of a traffic forwarding device provided in an embodiment of the present invention. The traffic forwarding device can be deployed on a leaf node in a large AI model network based on a leaf-spine network architecture, such as... Figure 5 As shown, the traffic forwarding device may include:

[0123] The receiving unit 510 is used to receive routing information published by the remote leaf node; wherein, the routing information includes the host route of the computing power node connected to the remote leaf node and the routing policy index of the computing power node.

[0124] The determining unit 520 is used to determine the next-hop list corresponding to the computing power node connected to the remote leaf node based on the received routing information.

[0125] Selection unit 530 is used to select a target next hop from the next hop list corresponding to any computing power node connected to a remote leaf node, based on the routing strategy index of the computing power node; wherein, for different leaf nodes, the target next hop corresponding to the same computing power node connected to the same remote leaf node is the same; for the same leaf node, the target next hop corresponding to different computing power nodes connected to the same remote leaf node is different.

[0126] The forwarding control unit 540 is used to generate forwarding table entries for the computing power node based on the host route of the computing power node and the target next hop corresponding to the computing power node, and to forward the traffic sent to the computing power node based on the forwarding table entries.

[0127] In some embodiments, for any computing power node, the routing strategy index corresponding to the computing power node includes a sub-index for identifying the target port, where the target port is the port on the target leaf node that connects to the computing power node, and the target leaf node is the leaf node to which the computing power node is connected.

[0128] Selection unit 530 selects the target next hop from the next hop list corresponding to the computing power node based on the routing strategy index of the computing power node, including:

[0129] Based on the sub-index corresponding to the computing power node, select the target neighbor next hop that matches the sub-index from the neighbor next hops sorted according to the preset sorting strategy;

[0130] If the next-hop list corresponding to the computing node includes the next hop of the target neighbor, then the next hop of the target neighbor is determined as the target next hop;

[0131] For the same leaf node, computing power nodes connected to the same port number on different remote leaf nodes have the same next hop for their target neighbors.

[0132] In some embodiments, the sub-index is the port number of the target port.

[0133] In some embodiments, for any computing power node, the routing strategy index corresponding to the computing power node includes a primary index for identifying the target leaf node;

[0134] Selection unit 530 is also used to select an alternative target next hop from the next hop list corresponding to the computing power node based on the sub-index and main index of the computing power node when the target next hop corresponding to the computing power node is abnormal; wherein, for the same leaf node, the alternative target next hops of computing power nodes connected by the same port number on different remote leaf nodes are not completely the same.

[0135] The forwarding control unit 540 is also used to generate a backup forwarding table entry for the computing power node based on the host route of the computing power node and the backup target next hop corresponding to the computing power node, and to forward the traffic sent to the computing power node based on the backup forwarding table entry.

[0136] In some embodiments, the selection unit 530 selects a backup target next hop from the next hop list corresponding to the computing power node based on the sub-index and the main index corresponding to the computing power node, including:

[0137] Based on the sub-index and main index corresponding to the computing power node, select the backup target neighbor next hop from the neighbor next hops sorted according to the preset sorting strategy, and select the one whose sorting matches the sum of the sub-index and main index.

[0138] If the next-hop list corresponding to the computing node includes a backup target neighbor next hop, then the backup target neighbor next hop is determined as the backup target next hop.

[0139] In some embodiments, the forwarding control unit 540 is further configured to, for any computing power node, in the event that the target next hop corresponding to the computing power node recovers abnormally, generate a forwarding table entry for the computing power node based on the host route of the computing power node and the target next hop corresponding to the computing power node, and forward the traffic sent to the computing power node based on the forwarding table entry.

[0140] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0141] For the apparatus embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to in the description of the method embodiment.

Claims

1. A traffic forwarding method, characterized in that, include: Receive routing information published by a remote leaf node; wherein, the routing information includes the host route of the computing power node connected to the remote leaf node and the routing policy index of the computing power node; Based on the received routing information, determine the next-hop list corresponding to the computing power node connected to the remote leaf node; For any computing power node connected to the far-end leaf node, the target next hop is selected from the next hop list corresponding to the computing power node according to the routing strategy index of the computing power node; wherein, for the same computing power node connected to the same far-end leaf node, different leaf nodes select the same target next hop; for different computing power nodes connected to the same far-end leaf node, the same leaf node selects different target next hops. Based on the host route of the computing power node and the target next hop corresponding to the computing power node, a forwarding table entry for the computing power node is generated, and traffic sent to the computing power node is forwarded according to the forwarding table entry.

2. The method according to claim 1, characterized in that, For any computing power node, the routing strategy index corresponding to the computing power node includes a sub-index for identifying the target port, where the target port is the port on the target leaf node that connects to the computing power node, and the target leaf node is the leaf node to which the computing power node is connected. The step of selecting the target next hop from the next hop list corresponding to the computing power node based on the routing strategy index of the computing power node includes: Based on the sub-index corresponding to the computing power node, select the target neighbor next hop that matches the sub-index from the neighbor next hops sorted according to the preset sorting strategy; If the target neighbor's next hop is included in the next hop list corresponding to the computing power node, the target neighbor's next hop is determined as the target next hop; For the same leaf node, computing power nodes connected to the same port number on different remote leaf nodes have the same next hop for their target neighbors.

3. The method according to claim 2, characterized in that, The sub-index is the port number of the target port.

4. The method according to claim 2, characterized in that, For any computing power node, the routing strategy index corresponding to that computing power node includes a primary index used to identify the target leaf node; The method further includes: For any computing power node, if the target next hop corresponding to the computing power node is abnormal, an alternative target next hop is selected from the next hop list corresponding to the computing power node based on the sub-index and main index corresponding to the computing power node; wherein, for the same leaf node, the alternative target next hops of computing power nodes connected by the same port number on different remote leaf nodes are not completely the same. Based on the host route of the computing node and the next hop of the backup target corresponding to the computing node, a backup forwarding table entry for the computing node is generated, and traffic sent to the computing node is forwarded according to the backup forwarding table entry.

5. The method according to claim 4, characterized in that, The step of selecting a backup target next hop from the next hop list corresponding to the computing power node based on the sub-index and main index of the computing power node includes: Based on the sub-index and main index corresponding to the computing power node, select the backup target neighbor next hop from the neighbor next hops sorted according to the preset sorting strategy, and select the one whose sorting matches the sum of the sub-index and main index. If the next-hop list corresponding to the computing node includes the backup target neighbor next hop, then the backup target neighbor next hop is determined as the backup target next hop.

6. The method according to claim 4, characterized in that, The method further includes: For any computing power node, if the target next hop corresponding to the computing power node recovers abnormally, a forwarding table entry for the computing power node is generated based on the host route of the computing power node and the target next hop corresponding to the computing power node, and traffic sent to the computing power node is forwarded based on the forwarding table entry.

7. A traffic forwarding device, characterized in that, include: The receiving unit is configured to receive routing information published by a remote leaf node; wherein the routing information includes host routes of computing nodes connected to the remote leaf node and routing policy indexes of the computing nodes. The determining unit is used to determine the next-hop list corresponding to the computing power node connected to the remote leaf node based on the received routing information. The selection unit is used to select a target next hop from the next hop list corresponding to any computing power node connected to the far-end leaf node, based on the routing strategy index of the computing power node; wherein, for the same computing power node connected to the same far-end leaf node, different leaf nodes select the same target next hop; for different computing power nodes connected to the same far-end leaf node, the same leaf node selects different target next hops. The forwarding control unit is used to generate forwarding table entries for the computing power node based on the host route of the computing power node and the target next hop corresponding to the computing power node, and to forward the traffic sent to the computing power node according to the forwarding table entries.

8. The apparatus according to claim 7, characterized in that, For any computing power node, the routing strategy index corresponding to the computing power node includes a sub-index for identifying the target port, where the target port is the port on the target leaf node that connects to the computing power node, and the target leaf node is the leaf node to which the computing power node is connected. The selection unit selects the target next hop from the next hop list corresponding to the computing power node based on the routing strategy index of the computing power node, including: Based on the sub-index corresponding to the computing power node, select the target neighbor next hop that matches the sub-index from the neighbor next hops sorted according to the preset sorting strategy; If the target neighbor's next hop is included in the next hop list corresponding to the computing power node, the target neighbor's next hop is determined as the target next hop; Among them, for the same leaf node, the computing power nodes connected to the same port number on different far leaf nodes have the same next hop for their target neighbors; The sub-index is the port number of the target port.

9. The apparatus according to claim 8, characterized in that, For any computing power node, the routing strategy index corresponding to that computing power node includes a primary index used to identify the target leaf node; The selection unit is further configured to, for any computing power node, in the event that the target next hop corresponding to the computing power node is abnormal, select an alternative target next hop from the next hop list corresponding to the computing power node based on the sub-index and main index corresponding to the computing power node; wherein, for the same leaf node, the alternative target next hops of computing power nodes connected by the same port number on different remote leaf nodes are not completely the same. The forwarding control unit is further configured to generate a backup forwarding table entry for the computing power node based on the host route of the computing power node and the backup target next hop corresponding to the computing power node, and forward the traffic sent to the computing power node based on the backup forwarding table entry.

10. The apparatus according to claim 9, characterized in that, The selection unit selects a backup target next hop from the next hop list corresponding to the computing power node based on the sub-index and main index of the computing power node, including: Based on the sub-index and main index corresponding to the computing power node, select the backup target neighbor next hop from the neighbor next hops sorted according to the preset sorting strategy, and select the one whose sorting matches the sum of the sub-index and main index. If the next hop list corresponding to the computing node includes the backup target neighbor next hop, then the backup target neighbor next hop is determined as the backup target next hop; And / or, The forwarding control unit is further configured to, for any computing power node, generate a forwarding table entry for the computing power node based on the host route of the computing power node and the target next hop corresponding to the computing power node in the event of an abnormal recovery of the target next hop of the computing power node, and forward the traffic sent to the computing power node based on the forwarding table entry.

Citation Information

Patent Citations

  • Methods and Apparatuses for Non-Blocking IP Multicast Delivery of Media Data in a Multi-Spine Network

    US20200028774A1

  • Source-initiated distribution of spine node identifiers of preferred spine nodes for use in multicast path selection

    US20200412639A1