Message transmission method, apparatus, Spine device, Leaf device, and computer-readable storage medium

By using different next hops of backup path groups to send fault notifications and data packets in the intelligent computing center network, the faulty link is bypassed, which solves the problem of excessively long fault convergence time caused by link failure in the intelligent computing center network, and realizes fast convergence and normal transmission of data packets.

CN119835212BActive Publication Date: 2025-10-28MAIPU COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411972673.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-10-28
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In intelligent computing center networks, link failures cause excessively long convergence times, affecting the efficiency of AI training tasks. Existing technologies rely on routing protocol processing logic, which also results in excessively long convergence times, making it impossible to quickly restore services.

Method used

When a Spine device detects a link failure, it sends a failure notification message and a data message through different next hops in the backup path group to bypass the failed link. It then uses the first next hop in the backup path group to send the failure notification message to all Leaf devices except the target Leaf device, and sends the data message to the relay Leaf device through the second next hop in the backup path group, and then sends it to the target Leaf device through the second Spine device.

Benefits of technology

It shortens the fault convergence time, improves convergence performance, avoids the long convergence time caused by relying on the control plane to handle faults through routing protocols when the link fails, ensures the normal transmission of data packets, and reduces packet loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119835212B_ABST
    Figure CN119835212B_ABST
Patent Text Reader

Abstract

This invention relates to data communication technology, providing a message sending method, apparatus, Spine device, Leaf device, and computer-readable storage medium. The method includes: receiving a data message, the data message being a message to be forwarded by a target Leaf device among a plurality of Leaf devices; when a direct link failure is detected with the target Leaf device, sending a failure notification message to the remaining Leaf devices (excluding the target Leaf device) through a first next hop in a backup path group, and sending the data message to a relay Leaf device through a second next hop in the backup path group, so that the data message reaches a second Spine device via the relay Leaf device, and is then sent to the target Leaf device for forwarding via the second Spine device. This invention can bypass faulty links, quickly forward data, shorten fault convergence time, and improve network performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data communication technology, and more specifically, to a message sending method, apparatus, Spine device, Leaf device, and computer-readable storage medium. Background Technology

[0002] With the rapid development of artificial intelligence technology, intelligent computing center networks have gradually become a research hotspot in the industry. AI training, due to its characteristics of distributed computing, long-cycle operation, and real-time response, is extremely sensitive to network failures. Since intelligent computing networks typically use a large number of vulnerable components such as optical modules, network failures are inevitable during AI training. Currently, the industry places great emphasis on improving the fault convergence speed in intelligent computing center networks and reducing the impact of network failures on AI training tasks. This involves remote fault convergence, one approach being to combine fault notification with fault handover.

[0003] Intelligent computing centers typically employ relatively regular network topologies, with the Fat-Tree being a common configuration. Fat-Trees in intelligent computing centers can be further categorized into two-level and three-level network configurations. In a typical two-level Fat-Tree network, each Leaf device connects half its ports to servers / GPUs (Graphics Processing Units) and the other half to Spine devices. Each Spine device connects all its ports to Leaf devices. The maximum number of GPUs the network can connect to is limited by the number of ports on the switch.

[0004] For a 64-port box-type secondary box-to-box network, it supports a maximum of 32 Spine devices and 64 Leaf devices, supporting 2048 GPUs. For a 128-port box-type secondary box-to-box network, it supports a maximum of 64 Spine devices and 128 Leaf devices, supporting 8192 GPUs.

[0005] In real-world network deployments, there will be a large number of GPUs, as well as numerous Leaf and Spine devices. Leaf devices have a 1:1 uplink / downlink bandwidth convergence ratio, and each pair of Leaf and Spine devices typically has only one link between them. When a link failure occurs in the network, the devices need to directly or indirectly detect the link change and then notify other devices in the network to update their forwarding table entries to achieve convergence.

[0006] Current solution: According to the standard routing protocol processing logic, fault convergence mainly relies on the device control plane. After the Spine device detects a link failure of the connected Leaf device, it sends a protocol message to other Leaf devices to notify them of the route cancellation. Services can only be restored after other Leaf devices receive the protocol message and update their local routing table entries. In typical scenarios, the convergence time is in the second range, and the service packet loss time is too long. Summary of the Invention

[0007] The present invention aims to provide a message transmission method, apparatus, Spine device, Leaf device, and computer-readable storage medium.

[0008] The embodiments of the present invention can be implemented as follows:

[0009] In a first aspect, the present invention provides a message sending method applied to a switching chip of a first Spine device in an intelligent computing network. The first Spine device is directly connected to multiple Leaf devices. The switching chip is configured with a backup path group for the direct connection links from the first Spine device to each of the Leaf devices. The method includes:

[0010] Receive data packets, wherein the data packets are packets to be forwarded by a target Leaf device among the plurality of Leaf devices;

[0011] When a direct link failure is detected with the target Leaf device, a fault notification message generated based on the data packet is sent to all Leaf devices except the target Leaf device via the first next hop in the backup path group. The data packet is then sent to a relay Leaf device via the second next hop in the backup path group, so that it can reach the second Spine device. Finally, the data packet is sent to the target Leaf device for forwarding via the second Spine device.

[0012] In an optional implementation, the step of sending the fault notification message to the remaining Leaf devices other than the target Leaf device via the first next hop in the backup path group includes:

[0013] Based on the first next hop, the data packet is edited to obtain the fault notification packet carrying fault link information;

[0014] The fault notification message is forwarded to a pre-created multicast group at a preset rate limit so that it can be received by each of the remaining Leaf devices; wherein the multicast group includes each of the plurality of Leaf devices.

[0015] In an optional implementation, the step of editing the data packet to obtain the fault notification packet carrying fault link information includes:

[0016] The source MAC address of the data packet is modified to the MAC address of the first Spine device, the destination MAC address of the data packet is modified to the MAC address of the target Leaf device, and the Ethernet type of the data packet is modified to a custom value, so that any of the other Leaf devices that receive the fault notification message can determine that the direct link between the first Spine device and the target Leaf device is faulty based on the source and destination MAC addresses and the Ethernet type.

[0017] In an optional implementation, the step of sending the data packet to the relay Leaf device via the second next hop in the backup path group includes:

[0018] According to the second next hop, add a preset tag to the data packet;

[0019] The transit Leaf device is determined from all the remaining Leaf devices based on the preset marker;

[0020] The data packet is sent to the second Spine device via the relay Leaf device, and then sent to the target Leaf device via the second Spine device.

[0021] In an optional implementation, the step of determining the transit Leaf device from all remaining Leaf devices based on the preset marker includes:

[0022] By searching for the forwarding action corresponding to the access control list that matches the preset tag, the relay Leaf device is determined from each Leaf device reachable in the pre-created ECMP group according to the preset load balancing policy; the ECMP group includes all other Leaf devices except the target Leaf device among the plurality of Leaf devices.

[0023] In an optional implementation, the step of sending the data packet to the second Spine device via the relay Leaf device, and then sending the data packet to the target Leaf device via the second Spine device, includes:

[0024] Encapsulate the data packet with a tunnel outer IP header, and use the IP address of the relay Leaf device as the destination IP address of the tunnel outer IP header to obtain a tunnel-encapsulated packet;

[0025] The tunnel-encapsulated message is sent to the relay Leaf device, so that the relay Leaf device decapsulates the data packet from the tunnel-encapsulated message and sends it to the target Leaf device through the second Spine device.

[0026] Secondly, the present invention provides a message sending method, applied to any Leaf device in an intelligent computing network that is directly connected to a first Spine device, wherein the first Spine device is directly connected to multiple Leaf devices, and the Leaf device includes a switching chip and a CPU, the method comprising:

[0027] The switching chip receives a fault notification message sent by the first Spine device, wherein the fault notification message is sent by the first Spine device when it detects a fault in the direct link between itself and the target Leaf device;

[0028] The switching chip searches for an access control list that matches the Ethernet type of the fault notification message, and sends the fault notification message to the CPU according to the forwarding action corresponding to the access control list.

[0029] The CPU determines that the direct link between the first Spine device and the target Leaf device is faulty based on the source MAC address and destination MAC address of the fault notification message, and notifies the switching chip to remove the next hop corresponding to the first Spine device from the pre-created ECMP group. The ECMP includes the next hops corresponding to all Spine devices directly connected to this device.

[0030] In an optional implementation, the switching chip receives a tunnel encapsulated packet sent by the first Spine device through an IP tunnel, decapsulates the data packet from the tunnel encapsulated packet, determines a second Spine device from the ECMP group according to a preset load balancing strategy, and sends the data packet to the target Leaf device through the second Spine device.

[0031] Thirdly, the present invention provides a message sending device applied to a switching chip of a first Spine device in an intelligent computing network. The first Spine device is directly connected to multiple Leaf devices. The switching chip is configured with a backup path group for the direct connection links from the first Spine device to each of the Leaf devices. The device includes:

[0032] A receiving module is used to receive data packets, wherein the data packets are packets to be forwarded by a target Leaf device among the plurality of Leaf devices;

[0033] The fault detection module is used to notify the sending module when a fault is detected in the direct link between the target Leaf device;

[0034] The sending module is configured to, according to the notification from the fault detection module, send a fault notification message generated based on the data packet to all Leaf devices except the target Leaf device via the first next hop in the backup path group, and send the data packet to the relay Leaf device via the second next hop in the backup path group, so that the data packet can reach the second Spine device via the relay Leaf device, and then be sent to the target Leaf device for forwarding via the second Spine device.

[0035] Fourthly, the present invention provides a Spine device, including a processor and a switching chip, wherein the switching chip, under the control of the processor, implements the message transmission method as described in the first aspect.

[0036] Fifthly, the present invention provides a Leaf device, which is directly connected to a first Spine device, and the Leaf device includes a switching chip and a CPU;

[0037] The switching chip is used to receive a fault notification message sent by the first Spine device, wherein the fault notification message is sent by the first Spine device when it detects a fault in the direct link between itself and the target Leaf device;

[0038] The switching chip is also used to find an access control list that matches the Ethernet type of the fault notification message, and send the fault notification message to the CPU according to the forwarding action corresponding to the access control list.

[0039] The CPU is configured to determine a direct link failure between the first Spine device and the target Leaf device based on the source MAC address and destination MAC address of the fault notification message, and to notify the switching chip to remove the next hop corresponding to the first Spine device from the pre-created ECMP group, wherein the ECMP includes the next hops corresponding to all Spine devices directly connected to this device.

[0040] In a sixth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the message sending method described in the first aspect.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] When the Spine device detects a direct link failure with the target Leaf device, this invention sends a fault notification message to all Leaf devices except the target Leaf device via the first next hop in the backup path group, and sends a data packet to a relay Leaf device via the second next hop in the backup path group. The data packet then reaches the second Spine device via the relay Leaf device, and is sent to the target Leaf device via the second Spine device. By using different next hops to send the fault notification message and the data packet respectively, and by bypassing the faulty link, the normal transmission of the data packet is achieved. This avoids the long fault convergence time caused by relying on the control plane to handle faults through routing protocols when the link fails, thus shortening the fault convergence time and improving convergence performance. Attached Figure Description

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 This is an example diagram illustrating a fault in a smart computing scenario provided in this embodiment.

[0045] Figure 2 This is a block diagram of a network device provided in this embodiment.

[0046] Figure 3 This is a flowchart illustrating the message sending method provided in this embodiment.

[0047] Figure 4 This is an example diagram illustrating the fault reporting detour in the intelligent computing scenario provided in this embodiment.

[0048] Figure 5 This is an example diagram illustrating the overall fault handling process in the intelligent computing scenario provided in this embodiment.

[0049] Figure 6 This is a flowchart illustrating the message sending device provided in this embodiment.

[0050] Icons: 10-Spine device; 20-Leaf device; 30-Network device; 31-Processor; 32-Switching chip; 100-Message sending device; 110-Receiver module; 120-Fault detection module; 130-Sending module. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0052] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0053] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0054] In the description of this invention, it should be noted that if terms such as "upper," "lower," "inner," or "outer" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed, they are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.

[0055] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0056] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.

[0057] Please refer to Figure 1 , Figure 1 This is an example diagram illustrating a fault in the intelligent computing scenario provided in this embodiment. Figure 1 In a typical Fat-Tree network, there are two Spine devices: Spine1 and Spine2, and two Leaf devices: Leaf1 and Leaf2. Leaf1 communicates with both GPU1 and GPU2 simultaneously, while Leaf2 communicates with both GPU3 and GPU4 simultaneously. Spine1 communicates with both Leaf1 and Leaf2 simultaneously, and Spine2 communicates with both Leaf1 and Leaf2 simultaneously. It should be noted that in actual Fat-Tree networks, there are a large number of Leaf and Spine devices. Figure 1 This is a simplified network diagram created solely to illustrate the technical problem solved by this invention.

[0058] by Figure 1 For example, under normal circumstances, the messages sent by GPU3 are in accordance with... Figure 1 The route is forwarded via path 1. However, when Spine1 detects a link failure with Leaf1, the existing technology sends a protocol message to Leaf2 to notify of the route cancellation. Service can only resume after Leaf2 receives the protocol message and updates its local routing table. Only then will the message sent by GPU3 switch to [the correct path]. Figure 1 The path 2 in the middle is sent, and the fault convergence time is too long.

[0059] In view of this, this embodiment provides a message sending method, apparatus, Spine device, Leaf device, and computer-readable storage medium. The core improvement is that when the Spine device in the intelligent computing network detects a link failure with the target Leaf device, it does not handle the link failure through routing protocol messages to switch the message to the normal path for forwarding. Instead, it sends fault notification messages and data messages separately using different next hops. At the same time, it achieves normal transmission of data messages by bypassing the faulty link. This avoids the long fault convergence time caused by relying on the control plane to handle faults through routing protocols when the link fails, shortens the fault convergence time, and improves the convergence performance. It will be described in detail below.

[0060] Please refer to Figure 2 , Figure 2 This is a block diagram illustrating the network device 30 provided in this embodiment. The network device 30 can be... Figure 1 The Spine device in this embodiment implements the message sending method used to implement the switching chip of the Spine device. Alternatively, it can be... Figure 1 The Leaf device in the embodiment includes a processor 31, which is the CPU of the Leaf device and is used to implement the message transmission method for the Leaf device in this embodiment. The network device 30 includes a processor 31 and a switching chip 32, which are connected via a bus.

[0061] The processor 31 can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the message transmission method in the above embodiments can be completed by the integrated logic circuits in the processor 31 or by software instructions. The processor 31 can be a general-purpose processor, including a CPU (Central Processing Unit), an NP (Network Processor), etc.; it can also be a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Logic Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0062] Switching chip 32 is responsible for high-speed forwarding and processing of data packets. The switching chip has a variety of built-in routing algorithms and can quickly switch paths when a link failure is detected.

[0063] In this embodiment, as an implementation method, the interaction chip 32 of the Spine device also creates multiple eloop ports (engress loop, exit loopback port) and iloop ports (ingress loopback port). The iloop port is the loopback port in the entry process of the switching chip, and the eloop port is the loopback port in the exit process of the switching chip.

[0064] The switching chip 32 of the Spine device implements the message sending method applied to the switching chip 32 of the Spine device in this embodiment under the control of the processor 31 of the Spine device.

[0065] based on Figure 1 and Figure 2 This embodiment provides a message sending method applied to Spine devices. Please refer to [link / reference]. Figure 3 , Figure 3 This is a flowchart illustrating the message sending method provided in this embodiment. The method includes the following steps:

[0066] Step S101: Receive a data packet, which is a packet to be forwarded by the target Leaf device among multiple Leaf devices;

[0067] In this embodiment, the data packet is a service packet that the Spine device needs to send to the target Leaf device. Through the target Leaf device, the data packet can reach the GPU that is directly connected to the target Leaf device.

[0068] Step S102: When a direct link failure is detected between the target Leaf device and the target Leaf device, a fault notification message generated based on the data packet is sent to the remaining Leaf devices other than the target Leaf device via the first next hop in the backup path group, and the data packet is sent to the relay Leaf device via the second next hop in the backup path group, so that the data packet can reach the second Spine device via the relay Leaf device, and then be sent to the target Leaf device for forwarding via the second Spine device.

[0069] In this embodiment, the backup path group includes two paths: one path for sending fault notification messages and the other path for sending data packets. Each path corresponds to a different next hop. Through the first next hop, fault notification messages can be sent to all Leaf devices except the target Leaf device, enabling rapid link convergence. Through the second next hop, data packets can be sent to a relay Leaf device for forwarding. The relay Leaf device is a different Leaf device from the target Leaf device. Through the relay Leaf device, data packets can bypass the faulty link, achieving normal forwarding and avoiding packet loss.

[0070] The method provided in this embodiment enables the normal transmission of data packets by bypassing the faulty link, avoiding the long fault convergence time caused by relying on the control plane to handle faults through routing protocols when the link fails, thus shortening the fault convergence time, improving convergence performance, and avoiding packet loss.

[0071] In an optional implementation, one way to send a fault notification message is as follows:

[0072] First, based on the first next hop, the data packet is edited to obtain a fault notification packet carrying fault link information;

[0073] In this embodiment, to enable other Leaf devices to recognize the received message as a fault notification message, the Ethernet Type field (EtherType) in the data packet can be edited and set to a specific value. To enable other Leaf devices to know the faulty link based on the received fault notification message, information about the faulty link can be set in the data packet. One implementation method is as follows:

[0074] The source MAC address of the data packet is modified to the MAC address of the first Spine device, the destination MAC address of the data packet is modified to the MAC address of the target Leaf device, and the Ethernet type of the data packet is modified to a custom value, so that any other Leaf device that receives the fault notification message can determine that the direct link between the first Spine device and the target Leaf device is faulty based on the source and destination MAC addresses and the Ethernet type.

[0075] Secondly, the fault notification message is forwarded to a pre-created multicast group at a preset rate limit so that it can be received by every other Leaf device; wherein, the multicast group includes each Leaf device among multiple Leaf devices.

[0076] In this embodiment, a multicast group is pre-created in the switching chip of the first Spine device. The multicast group includes each Leaf device among multiple Leaf devices. Through the multicast group, fault notification messages are sent to each Leaf device. However, since the link between the first Spine device and the target Leaf device has failed, the target Leaf device will not receive the fault notification message. That is, the fault notification message will be received by every Leaf device except the target Leaf device.

[0077] To minimize the impact of fault notification messages on bandwidth, this embodiment also performs rate limiting on fault notification messages, forwarding them via multicast groups at a preset rate limit.

[0078] In this embodiment, after receiving a fault notification message, in order to promptly handle the fault based on the fault link information in the fault notification message and prevent the Leaf device from sending messages to the fault link during subsequent message forwarding, this embodiment provides a processing method applied to the Leaf device:

[0079] First, the switching chip of the Leaf device receives a fault notification message sent by the first Spine device, which is sent by the first Spine device when it detects a fault in the direct link between itself and the target Leaf device.

[0080] Secondly, the switching chip of the Leaf device searches for an access control list that matches the Ethernet type of the fault notification message, and sends the fault notification message to the CPU according to the forwarding action corresponding to the access control list.

[0081] Finally, the CPU of the Leaf device determines that the direct link between the first Spine device and the target Leaf device is faulty based on the source MAC address and destination MAC address of the fault notification message. It then notifies the switching chip to remove the next hop corresponding to the first Spine device from the pre-created ECMP group, which includes the next hops corresponding to all Spine devices directly connected to this device.

[0082] In this embodiment, while sending the fault notification message, the data packet is also forwarded normally, bypassing the faulty link. One way to forward the data packet is as follows:

[0083] First, add a preset tag to the data packet based on the second next hop;

[0084] In this embodiment, in order to enable data packets to be forwarded to the target Leaf device through the second Spine device, a preset tag is added to the data packets before they are forwarded to distinguish them from other data packets that are forwarded directly.

[0085] Secondly, the transit Leaf device is determined from all the remaining Leaf devices based on the preset markers;

[0086] In this embodiment, the switching chip of the first Spine device pre-issues an ACL (Access Control List) entry, which is used to match the forwarding action corresponding to the preset tag.

[0087] In this embodiment, the first Spine device is directly connected to multiple Leaf devices. The first Spine device pre-creates an ECMP (Equal-Cost Multi-Path Routing) group. Each ECMP member in the group is associated with a Leaf device, meaning that the Leaf device associated with the first Spine device can be reached through the corresponding ECMP member.

[0088] One method for determining relay Leaf devices is as follows: by searching for the forwarding action corresponding to the access control list that matches the preset tag, the relay Leaf device is determined from each Leaf device reachable in a pre-created ECMP group according to the preset load balancing policy; the ECMP group includes all other Leaf devices in the multiple Leaf devices except the target Leaf device.

[0089] Finally, the data packets are sent to the second Spine device via the relay Leaf device, and then sent to the target Leaf device via the second Spine device.

[0090] In this embodiment, in order to send data packets to the relay Leaf device, the IP address of the relay Leaf device needs to be used as the destination IP address. To avoid affecting the normal transmission of data packets, this embodiment encapsulates the data packets with an additional tunnel IP header. Specifically, this can be achieved by:

[0091] (1) Encapsulate the data packet with the outer IP header of the tunnel and use the IP address of the relay Leaf device as the destination IP address of the outer IP header of the tunnel to obtain the tunnel-encapsulated packet;

[0092] (2) Send the tunnel encapsulated message to the relay Leaf device so that the relay Leaf device can decapsulate the data packet from the tunnel encapsulated message;

[0093] (3) And send it to the target Leaf device via the second Spine device.

[0094] In this embodiment, to enable the relay Leaf device to send data packets to the target Leaf device via the second Spine device, this embodiment provides an implementation method:

[0095] First, the switching chip of the relay Leaf device receives the tunnel encapsulation message. The tunnel encapsulation message is obtained by the first Spine device encapsulating the data packet with the outer IP header of the tunnel and using the IP address of the relay Leaf device as the destination IP address of the outer IP header of the tunnel.

[0096] Secondly, the switching chip of the relay Leaf device decapsulates the data packet from the tunnel-encapsulated message;

[0097] Finally, the switching chip of the relay Leaf device determines the second Spine device from the ECMP group according to the preset load balancing strategy, and sends it to the target Leaf device through the second Spine device.

[0098] In this embodiment, after receiving the fault notification message, the relay Leaf device has already learned that the link between the first Spine device and the target Leaf device is faulty. Therefore, the relay Leaf device will delete the next hop corresponding to the first Spine device in the local ECMP group and send the received data packets out through the second Spine device.

[0099] In this embodiment, the relay Leaf device decapsulates the data packets received from the first Spine device and then sends them to the second Spine device via a pre-created ECMP (Equal-Cost Multi-Path Routing) group. Upon arrival at the second Spine device, the second Spine device can use its routing table to query the destination IP address of the target Leaf device in the data packet and forward the data packet to the target Leaf device.

[0100] To more intuitively illustrate the process of data packet detour and forwarding, please refer to... Figure 4 , Figure 4 This is an example diagram illustrating fault reporting detours in the intelligent computing scenario provided in this embodiment. Figure 4 In this process, GPU2 needs to send data packets to GPU1. The data packets sent by GPU2 pass through Leaf2 and arrive at Spine1. Spine1 detects a link failure between itself and Leaf1, adds a preset tag to the data packet, selects an ECMP member based on the preset tag, encapsulates the data packet with the outermost IP header of the tunnel, and sends it to Leaf2. Leaf2 decapsulates the outermost IP header of the tunnel and sends the data packet to Spine2. Spine2 then sends the data packet to Leaf1, and finally it reaches GPU1.

[0101] In this embodiment, as a specific implementation, to achieve simultaneous forwarding of fault notification messages and data packets, the first Spine device pre-creates an APS (Automatic Protection Switch) group for each Leaf device connected to it. Each Leaf device's APS group includes two paths: a primary path and a backup path group. The primary path is the path formed by the next hop to the corresponding Leaf device; when the link is normal, data packets are sent through the primary path. The backup path group, being a multicast group, is used to handle faults when a link failure occurs on the primary path, utilizing the first next hop in the backup path group and the second next hop in the backup path group for data packet rerouting.

[0102] In this embodiment, two eloop ports (engress loops) and two iloop ports (ingress loops) are created on the first Spine device: the first eloop port and the second eloop port. A first next hop (i.e., the first next hop of the backup path group) is created pointing to the first eloop port, and the first iloop port is associated with the first eloop port. A second next hop (i.e., the second next hop of the backup path group) is created pointing to the second eloop port, and the second iloop port is associated with the second eloop port.

[0103] When a link failure occurs on the main path, data packets will be sent to both the first eloop port and the second eloop port simultaneously.

[0104] To illustrate the overall process described above, please refer to [link / reference]. Figure 5 , Figure 5 This is an example diagram illustrating the overall fault handling process in the intelligent computing scenario provided in this embodiment. Figure 5 In the example of the APS group created by the first Spine device for Leaf-x (the target Leaf device), the processing method for data packets sent to the first eloop interface is as follows:

[0105] (1) The first Spine device edits the data packet to obtain a fault notification packet carrying fault link information;

[0106] (2) Send the fault notification message to the first iloop port through the first eloop port at a preset rate limit;

[0107] This embodiment can bind a rate-limiting hardware entry of a preset rate QoS (Quality of Service) policy to the first eloop port. The preset rate can be set according to the needs of the actual scenario, for example, the preset rate can be set to 1PPS (Packets Per Second).

[0108] (3) The fault notification message is sent to each of the remaining Leaf devices through the multicast group cross-connected through the first iloop port.

[0109] In this embodiment, in order to send the fault notification message to each of the remaining Leaf devices, a multicast group is established on the first Spine device. Each member of the multicast group is associated with each of the remaining Leaf devices. The first iloop port is cross-connected to the multicast group. Specifically, the fault notification message can be forwarded to the multicast group through the first iloop port, so that it can be sent to each of the remaining Leaf devices through each member of the multicast group. Figure 5 In the process, the first Spine device edits the data packet to obtain the fault notification message, and sends it to the first iloop port through rate limiting. The first iloop port sends it to multicast group 2 in a cross-connect manner. Multicast group 2 includes n members from Leaf-1 to Leaf-n. Each member reaches a Leaf device. Thus, the fault notification message is sent to all Leaf devices with normal connection links to the Spine device.

[0110] The processing method for data packets sent to the second eloop port is as follows:

[0111] (1) Add the LogicPort tag as the default tag to the data packet;

[0112] (2) Send the data packet with the preset tag to the second eloop port, and then send it to the second iloop port through the second eloop port;

[0113] (3) By sending an access control list item that matches the preset tag through the second iloop port, the action is to redirect to the ECMP group and determine the relay Leaf device from all the remaining Leaf devices according to the preset load sharing policy;

[0114] In this embodiment, the second iloop port is pre-configured with forwarding actions corresponding to preset tags. As one implementation method, the Spine device can issue an ACL entry on the second iloop port. This entry is used to match the forwarding action corresponding to the logicPort tag, and the forwarding action is to redirect to the ECMP group.

[0115] Based on the pre-configured forwarding action of the matched ACL entry, the target ECMP member is selected from all ECMP members in the ECMP group according to the preset load balancing policy, and the relay Leaf device is reached according to the next hop corresponding to the target ECMP member.

[0116] (4) Send the data packet to the second Spine device through the relay Leaf device, and then send the data packet to the target Leaf device through the second Spine device.

[0117] Figure 5The diagram shows that the traffic received by the first Spine device that needs to be sent to Leaf-x is forwarded relatively evenly through the members of the ECMP group Leaf-1 to Leaf-k. After encapsulating the outer IP header of the tunnel, it is sent to the relay Leaf device. The relay Leaf device decapsulates the data to obtain the original data packet and forwards it to the second Spine device. The second Spine device looks up the routing table and sends the data packet to Leaf-x, thus completing the data packet detour forwarding.

[0118] In this embodiment, besides the Spine device being able to detect link failures with the Leaf device, the Leaf device can also detect failures with the Spine device. In this case, the Leaf device will also take appropriate action. One approach is for the Leaf device to create an ECMP group, where each member of the ECMP group is the next hop for the Spine device connected to the Leaf device. The ECMP group is pre-configured to enable failover, meaning that when a link between the Leaf device and a Spine device fails, the traffic on the failed link will be redirected to one of the remaining Spine devices. For example… Figure 1 In this scenario, when the link between Leaf1 and Spine1 fails, the traffic that was originally loaded onto Spine1 will be loaded onto other Spines (such as Spine2). The outgoing routes for other traffic will not change. If there is still a link between Leaf1 and Spine3, the traffic load on Spine3 will not be affected, and the traffic on Spine1 will only be loaded onto Spine2.

[0119] To perform the corresponding steps applied to the switching chip of the first Spine device in the above embodiments and various possible implementations, an implementation of the message sending device 100 is given below. Please refer to... Figure 6 , Figure 6 This is a block diagram of the message sending device provided in this embodiment. The message sending device 100 is applied to the switching chip of the first Spine device in this embodiment. It should be noted that the message sending device 100 provided by the present invention has the same basic principle and technical effect as the corresponding embodiment described above. For the sake of brevity, it is not mentioned in this embodiment.

[0120] The message sending device 100 includes a receiving module 110 and a fault detection module 120.

[0121] The receiving module 110 is used to receive data packets, which are packets to be forwarded by a target Leaf device among multiple Leaf devices;

[0122] The fault detection module 120 is used to notify the sending module 130 when a fault in the direct link between the target Leaf device is detected.

[0123] The sending module 130 is configured to, according to the notification from the fault detection module 120, send a fault notification message generated based on the data packet to all Leaf devices except the target Leaf device via the first next hop in the backup path group, and send the data packet to the relay Leaf device via the second next hop in the backup path group, so that the data packet can reach the second Spine device via the relay Leaf device, and then be sent to the target Leaf device for forwarding via the second Spine device.

[0124] In an optional implementation, the sending module 130 is specifically used for:

[0125] Based on the first next hop, the data packet is edited to obtain the fault notification packet carrying fault link information;

[0126] The fault notification message is forwarded to a pre-created multicast group at a preset rate limit so that it can be received by each of the remaining Leaf devices; wherein the multicast group includes each of the plurality of Leaf devices.

[0127] In an optional implementation, the sending module 130, when editing a data packet to obtain a fault notification message carrying fault link information, specifically performs the following functions:

[0128] The source MAC address of the data packet is modified to the MAC address of the first Spine device, the destination MAC address of the data packet is modified to the MAC address of the target Leaf device, and the Ethernet type of the data packet is modified to a custom value, so that any other Leaf device that receives the fault notification message can determine that the direct link between the first Spine device and the target Leaf device is faulty based on the source and destination MAC addresses and the Ethernet type value.

[0129] In an optional implementation, the sending module 130, when used to send data packets to the relay Leaf device via the second next hop in the backup path group, is specifically used for:

[0130] Add a preset tag to the data packet based on the second next hop;

[0131] The transit Leaf device is determined from all the remaining Leaf devices based on the preset markers;

[0132] The data packets are sent to the second Spine device via the relay Leaf device, and then sent to the target Leaf device via the second Spine device.

[0133] In an optional implementation, the sending module 130 is specifically used to: determine the relay Leaf device from all remaining Leaf devices based on a preset tag.

[0134] By finding the forwarding action corresponding to the access control list that matches the preset tag, the relay Leaf device is determined from each Leaf device reachable in the pre-created ECMP group according to the preset load balancing policy; the ECMP group includes all other Leaf devices in the multiple Leaf devices except the target Leaf device.

[0135] In an optional implementation, the sending module 130 is specifically used to send data packets to the second Spine device via the relay Leaf device, and to send data packets to the target Leaf device via the second Spine device, in the following ways:

[0136] Encapsulate the data packet with an outer tunnel IP header, and use the IP address of the relay Leaf device as the destination IP address of the outer tunnel IP header to obtain the tunnel-encapsulated packet;

[0137] The tunnel encapsulation message is sent to the relay Leaf device, so that the relay Leaf device can decapsulate the data packet from the tunnel encapsulation message and send it to the target Leaf device through the second Spine device.

[0138] In this embodiment, in order to perform the corresponding steps applied to the Leaf device in the above embodiments and various possible implementations, this embodiment provides an implementation method of the Leaf device:

[0139] The switching chip of the Leaf device is used to receive fault notification messages sent by the first Spine device, wherein the fault notification message is sent by the first Spine device when it detects a fault in the direct link between itself and the target Leaf device;

[0140] The switching chip of the Leaf device is also used to look up the access control list that matches the Ethernet type of the fault notification message, and send the fault notification message to the CPU of the Leaf device according to the forwarding action corresponding to the access control list.

[0141] The CPU of the Leaf device is used to determine the direct link failure between the first Spine device and the target Leaf device based on the source MAC address and destination MAC address of the fault notification message, and to notify the switching chip to remove the next hop corresponding to the first Spine device from the pre-created ECMP group.

[0142] The switching chip of the Leaf device is also used to remove the next hop corresponding to the first Spine device from a pre-created ECMP group, wherein the ECMP includes the next hops corresponding to all Spine devices directly connected to this device.

[0143] In an optional implementation, the switching chip of the Leaf device is also used to receive tunnel encapsulation messages sent by the first Spine device through the IP tunnel, decapsulate data packets from the tunnel encapsulation messages, determine the second Spine device from the ECMP group according to a preset load balancing strategy, and send the data packets to the target Leaf device through the second Spine device.

[0144] This invention provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the message transmission method applied to the first Spine device in the aforementioned embodiments.

[0145] In summary, embodiments of the present invention provide a message sending method, apparatus, Spine device, Leaf device, and computer-readable storage medium, applied to a switching chip of a first Spine device in an intelligent computing network. The first Spine device is directly connected to multiple Leaf devices. The switching chip is configured with a backup path group for the direct connection links from the first Spine device to each Leaf device. The method includes: receiving a data packet, the data packet being a packet to be forwarded by a target Leaf device among the multiple Leaf devices; when a failure of the direct connection link with the target Leaf device is detected, sending a failure notification message to the remaining Leaf devices (excluding the target Leaf device) through a first next hop in the backup path group, and sending the data packet to a relay Leaf device through a second next hop in the backup path group, so that the data packet reaches the second Spine device through the relay Leaf device, and is then forwarded to the target Leaf device by the second Spine device. Compared with the prior art, this embodiment has at least the following advantages: (1) By sending the data packet to the relay Leaf device, and then to the second Spine device through the relay Leaf device, and then sending the data packet to the target Leaf device through the second Spine device, the normal transmission of the data packet is achieved by bypassing the faulty link, which avoids the long fault convergence time caused by relying on the control plane to handle the fault through the routing protocol when the link fails, shortens the fault convergence time, and improves the convergence performance; (2) The timely transmission of the fault notification message and the bypass transmission of the data packet are achieved through two different paths. On the one hand, it can notify other Leaf devices to remove the faulty link in a timely manner, and on the other hand, it ensures the timely and normal forwarding of the data packet, further reducing the fault convergence time; (3) When sending the fault notification message, the rate is limited to avoid the fault notification message occupying too much bandwidth of the data packet.

[0146] The above descriptions are merely various embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A message sending method, characterized in that, A switching chip for a first Spine device in an intelligent computing network, the first Spine device being directly connected to multiple Leaf devices, the switching chip being configured with a backup path group for the direct connection links from the first Spine device to each of the Leaf devices, the method comprising: Receive data packets, wherein the data packets are packets to be forwarded by a target Leaf device among the plurality of Leaf devices; When a direct link failure is detected with the target Leaf device, a fault notification message generated based on the data packet is sent to all Leaf devices except the target Leaf device via the first next hop in the backup path group. The data packet is then sent to a relay Leaf device via the second next hop in the backup path group, so that it can reach the second Spine device. Finally, the data packet is sent to the target Leaf device for forwarding via the second Spine device.

2. The message sending method according to claim 1, characterized in that, The step of sending the fault notification message to the remaining Leaf devices other than the target Leaf device via the first next hop in the backup path group includes: Based on the first next hop, the data packet is edited to obtain the fault notification packet carrying fault link information; The fault notification message is forwarded to a pre-created multicast group at a preset rate limit so that it can be received by each of the remaining Leaf devices; wherein the multicast group includes each of the plurality of Leaf devices.

3. The message sending method according to claim 2, characterized in that, The step of editing the data packet to obtain the fault notification packet carrying fault link information includes: The source MAC address of the data packet is modified to the MAC address of the first Spine device, the destination MAC address of the data packet is modified to the MAC address of the target Leaf device, and the Ethernet type of the data packet is modified to a custom value, so that any of the other Leaf devices that receive the fault notification message can determine that the direct link between the first Spine device and the target Leaf device is faulty based on the source and destination MAC addresses and the Ethernet type.

4. The message sending method according to claim 1, characterized in that, The step of sending the data packet to the relay Leaf device via the second next hop in the backup path group includes: According to the second next hop, add a preset tag to the data packet; The transit Leaf device is determined from all the remaining Leaf devices based on the preset marker; The data packet is sent to the second Spine device via the relay Leaf device, and then sent to the target Leaf device via the second Spine device.

5. The message sending method according to claim 4, characterized in that, The step of determining the transit Leaf device from all remaining Leaf devices based on the preset marker includes: By searching for the forwarding action corresponding to the access control list that matches the preset tag, the relay Leaf device is determined from each Leaf device reachable in the pre-created ECMP group according to the preset load balancing policy; the ECMP group includes all other Leaf devices except the target Leaf device among the plurality of Leaf devices.

6. The message sending method according to claim 4, characterized in that, The step of sending the data packet to the second Spine device via the relay Leaf device, and then sending the data packet to the target Leaf device via the second Spine device, includes: Encapsulate the data packet with a tunnel outer IP header, and use the IP address of the relay Leaf device as the destination IP address of the tunnel outer IP header to obtain a tunnel-encapsulated packet; The tunnel-encapsulated message is sent to the relay Leaf device, so that the relay Leaf device decapsulates the data packet from the tunnel-encapsulated message and sends it to the target Leaf device through the second Spine device.

7. A message sending method, characterized in that, The method is applied to any Leaf device in an intelligent computing network that is directly connected to a first Spine device, wherein the first Spine device is directly connected to multiple Leaf devices, and the Leaf device includes a switching chip and a CPU. The switching chip receives a fault notification message sent by the first Spine device through the first next hop in the backup path group, wherein the fault notification message is sent by the first Spine device when it detects a direct link failure with the target Leaf device; The switching chip searches for an access control list that matches the Ethernet type of the fault notification message, and sends the fault notification message to the CPU according to the forwarding action corresponding to the access control list. The CPU determines that the direct link between the first Spine device and the target Leaf device is faulty based on the source MAC address and destination MAC address of the fault notification message, and notifies the switching chip to remove the next hop corresponding to the first Spine device from the pre-created ECMP group. The ECMP includes the next hops corresponding to all Spine devices directly connected to this device. The switching chip receives a tunnel encapsulation message sent by the first Spine device through an IP tunnel. The tunnel encapsulation message is obtained by encapsulating a data packet sent by the first Spine device through the second next hop in the backup path group. The chip decapsulates the data packet from the tunnel encapsulation message, determines the second Spine device from the ECMP group according to a preset load balancing strategy, and sends the data packet to the target Leaf device through the second Spine device.

8. A message sending device, characterized in that, A switching chip for a first Spine device in an intelligent computing network, the first Spine device being directly connected to multiple Leaf devices, the switching chip being configured with a backup path group for the direct connection links from the first Spine device to each of the Leaf devices, the device comprising: A receiving module is used to receive data packets, wherein the data packets are packets to be forwarded by a target Leaf device among the plurality of Leaf devices; The fault detection module is used to notify the sending module when a fault is detected in the direct link between the target Leaf device; The sending module is configured to, according to the notification from the fault detection module, send a fault notification message generated based on the data packet to all Leaf devices except the target Leaf device via the first next hop in the backup path group, and send the data packet to the relay Leaf device via the second next hop in the backup path group, so that the data packet can reach the second Spine device via the relay Leaf device, and then be sent to the target Leaf device for forwarding via the second Spine device.

9. A Spine device, characterized in that, It includes a processor and a switching chip, wherein the switching chip, under the control of the processor, implements the message transmission method as described in any one of claims 1-6.

10. A Leaf device, characterized in that, The Leaf device is directly connected to the first Spine device, and the Leaf device includes a switching chip and a CPU. The switching chip is used to receive a fault notification message sent by the first Spine device through the first next hop in the backup path group, wherein the fault notification message is sent by the first Spine device when it detects a direct link failure with the target Leaf device; The switching chip is also used to find an access control list that matches the Ethernet type of the fault notification message, and send the fault notification message to the CPU according to the forwarding action corresponding to the access control list. The CPU is configured to determine a direct link failure between the first Spine device and the target Leaf device based on the source MAC address and destination MAC address of the fault notification message, and to notify the switching chip to remove the next hop corresponding to the first Spine device from the pre-created ECMP group, wherein the ECMP includes the next hops corresponding to all Spine devices directly connected to this device. The switching chip is also used to receive tunnel encapsulation messages sent by the first Spine device through an IP tunnel. The tunnel encapsulation messages are obtained by encapsulating data packets sent by the first Spine device through the second next hop in the backup path group. The chip decapsulates the data packets from the tunnel encapsulation messages, determines the second Spine device from the ECMP group according to a preset load balancing strategy, and sends the data packets to the target Leaf device through the second Spine device.

11. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the message sending method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Looped network routing method and looped network node

    CN101272352A

  • Method and system for accelerating convergence of media access control (MAC) address

    CN107483312A