Fault processing method and MLAG system

By having the master and slave devices obtain and send the uplink port status when a peer-link fails, the access device can flexibly adjust the aggregation group port status, solving the problem of traffic forwarding failure in the MLAG system after a peer-link failure and improving the robustness of the system.

CN120639593APending Publication Date: 2025-09-12MAIPU COMM TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410277777.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In existing MLAG systems, when a peer-link fails, especially when the master device's uplink channel fails, traffic cannot be forwarded normally, resulting in poor system robustness.

Method used

When the master and slave devices detect a peer-link failure, they each obtain the uplink port status and send it to the access device. The access device then sends a cancel configuration message based on the uplink port status of the master and slave devices, causing the slave device to cancel the fail-closed state of the aggregation group port, allowing traffic to be sent through the aggregation group port and uplink port of the slave device.

Benefits of technology

This improves the traffic forwarding capability of the MLAG system after a peer-link failure, enhances the robustness of the system, and avoids traffic interruption caused by a master device uplink port failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639593A_ABST
    Figure CN120639593A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data communication, and relates to a fault processing method and an MLAG system, and the method comprises the steps: when a master device detects a Peer-link fault, obtaining the state of an uplink port of the master device, and transmitting the state to an access device; when the slave equipment detects a Peer-link fault, the slave equipment obtains the state of an uplink port of the slave equipment and sends the state to the access equipment; and when the access equipment judges that the state of the uplink port of the master equipment is faulty and the state of the uplink port of the slave equipment is normal, the access equipment sends a configuration cancelling message to the slave equipment so as to indicate the slave equipment to cancel the operation of configuring the own aggregation group port to be in a faulty closed state according to the configuration cancelling message. And the slave device receives the flow of the access device through the own aggregation group port and sends the flow out through the uplink port of the slave device. According to the invention, the robustness of the MLAG system can be improved when the Peer-link fault occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data communications, and in particular to a fault handling method and an MLAG system. Background Art

[0002] MLAG (Multi-chassis Link Aggregation Group) is a mechanism for implementing cross-device link aggregation. It selects one or more ports on each of two adjacent devices to form a logical aggregation port, which is then used for dual-homing access devices, thereby forming an MLAG network. From the perspective of the access devices, the two MLAG devices are logically virtualized into a single device. This allows them to forward traffic together, ensuring the reliability of the MLAG system.

[0003] Please refer to Figure 1 , Figure 1 This is a network example diagram of the MLAG system provided by the present invention. Figure 1 In the example, the MLAG system includes node device 1 and node device 2. A peer link and a keepalive link are established between node device 1 and node device 2. The peer link is the aggregation of direct links between node device 1 and node device 2, used to exchange MLAG protocol messages and transmit data traffic. Node device 1 and node device 2 perform pairing detection, master-slave role election, etc. through MLAG protocol messages. MLAG protocol messages are communicated through the peer link layer 2 link. The keepalive link is the keepalive detection between node device 1 and node device 2, and communication is performed through the layer 3 link. When the peer link fails, node device 1 and node device 2 use the keepalive link to determine whether the other party is still alive. Figure 1 The access devices in the dual-homed MLAG system are connected, and node device 1 and node device 2 can jointly forward traffic from the access devices.

[0004] Figure 1 In the example, node device 1 and node device 2 will determine their respective master-slave roles through master-slave role election, with node device 1 as the master device and node device 2 as the slave device. If the peer-link fails at this time and the keep-alive link is normal, the slave device will enter the suspended state (suspend state) and cut the traffic to the master device. If the uplink port of the master device is in the connected state, the master device will forward the traffic of the access device normally. If all the uplink ports of the master device are in the disconnected state, the master and slave devices will not be able to forward traffic normally. Please refer to Figure 2 , Figure 2 This is an example diagram of an uplink path failure provided by the present invention. Figure 2In a scenario where the peer-link between the master and slave devices fails, the slave device enters a suspended state and cannot forward server traffic. The traffic should be sent to the spine devices through the master's uplink, but the master's uplink fails. Consequently, the server traffic cannot be forwarded by either the master or the slave, even if the slave's uplink is normal. This results in poor MLAG system robustness.

[0005] An existing implementation method is to monitor the situation between the uplink port of the master and slave devices and the CE port of the CE device (Customer Edge, user network edge device), and add it to the decision information. Through the keep-alive link, when the uplink channel between the master device and the CE device fails, the slave device can be selected as the new master device when the peer-link fails one after another, ensuring that the traffic can communicate from the new master device to the CE device. However, in this implementation method, if the peer-link fails first and the uplink channel is not faulty, the role election relationship will still be carried out in the original way. If the uplink channel of the master device fails again at this time, then there is still Figure 2 The traffic in the network cannot be forwarded by either the master device or the slave device. Summary of the Invention

[0006] The object of the present invention is to provide a fault handling method and an MLAG system, which can improve the robustness of the MLAG system.

[0007] The embodiments of the present invention can be implemented as follows:

[0008] In a first aspect, the present invention provides a fault handling method applied to an MLAG system, wherein the MLAG system includes a master device and a slave device, a peer-link is established between the master device and the slave device, and an access device is dual-homed to the MLAG system. The method includes:

[0009] When the master device detects the peer-link failure, it obtains the status of its own uplink port and sends it to the access device;

[0010] When the slave device detects the peer-link failure, it obtains the status of its own uplink port and sends it to the access device;

[0011] When the access device determines that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, the access device sends a cancellation configuration message to the slave device to instruct the slave device to cancel the operation of configuring its own aggregation group port to a fault-closed state according to the cancellation configuration message, and receive the traffic of the access device through its own aggregation group port and send it out through the uplink port of the slave device.

[0012] In an optional implementation manner, the step of obtaining the status of its own uplink port and sending it to the access device includes:

[0013] Get the status of its own uplink port;

[0014] Generate Link Aggregation Control Protocol LACP packets;

[0015] Adding the state of the own uplink port to the LACP message;

[0016] The state of its own uplink port is sent to the access device through the LACP message.

[0017] In an optional implementation manner, when the master device detects the peer-link failure, after obtaining the status of its own uplink port and sending it to the access device, the following steps are included:

[0018] When the master device determines that the state of its own uplink port is faulty, it sets the state of its own aggregation group port to disconnected.

[0019] In an optional embodiment, the master device has multiple aggregation group ports, the uplink port of the master device corresponds to at least one of the multiple aggregation group ports of the master device, and when the master device determines that the status of its own uplink port is faulty, the step of setting the status of its own aggregation group port to disconnected includes:

[0020] When the master device determines that the state of its own uplink port is faulty, it sets the states of all aggregation group ports corresponding to its own uplink port to disconnected.

[0021] In an optional embodiment, the method further comprises:

[0022] When the master device determines that the state of its own uplink port is normal, the master device obtains the state of its own uplink port in real time within a preset time period and synchronizes it to the slave device;

[0023] When the slave device detects that the state of the uplink port of the master device changes from normal to faulty and the state of its own uplink port is normal and the state of the aggregation group port is in a fault-closed state, it clears the fault-closed state of its own aggregation group port to receive the traffic of the access device through its own aggregation group port and send it out through its own uplink port.

[0024] In an optional embodiment, the step of acquiring the status of its own uplink port in real time within a preset time period and synchronizing it to the slave device includes:

[0025] Create a timer with a timing period of the preset duration;

[0026] During the timing period, the state of its own uplink port is acquired in real time and synchronized to the slave device.

[0027] In an optional implementation manner, a keep-alive link exists between the master device and the slave device, and the step of obtaining the status of its own uplink port and synchronizing it to the slave device includes:

[0028] Generate keepalive protocol messages;

[0029] Adding the state of its own uplink port to the keep-alive protocol message;

[0030] The state of its own uplink port is synchronized to the slave device through the keep-alive protocol message and the keep-alive link.

[0031] In an optional embodiment, after the step of the slave device acquiring the status of its own uplink port and sending the status to the access device when detecting the peer-link failure, the method further includes:

[0032] If the deconfiguration message is not received within a preset time period, the aggregation group port itself is configured to be in a fail-closed state.

[0033] In a second aspect, the present invention provides an MLAG system, the MLAG system including a master device and a slave device, a peer-link is established between the master device and the slave device, and an access device is dual-homed to the MLAG system;

[0034] The master device is configured to obtain the status of its own uplink port and send the status to the access device when detecting the peer-link failure;

[0035] The slave device is used to obtain the status of its own uplink port and send it to the access device when detecting the peer-link failure;

[0036] The access device is used to send a cancellation message to the slave device when it is determined that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, so as to instruct the slave device to cancel the operation of configuring its own aggregation group port to a fault-closed state according to the cancellation message, and receive the traffic of the access device through its own aggregation group port and send it out through the uplink port of the slave device.

[0037] In an optional embodiment, the master device is further configured to obtain the status of its own uplink port in real time within a preset time period and synchronize the status to the slave device when determining that the status of its own uplink port is normal;

[0038] The slave device is also used to clear the fault-closed state of its own aggregation group port when it detects that the state of the uplink port of the master device changes from normal to faulty and the state of its own uplink port is normal and the state of the aggregation group port is a fault-closed state, so as to receive the traffic of the access device through its own aggregation group port and send it out through its own uplink port.

[0039] The present invention provides a fault handling method and an MLAG system. When a master and slave device detect a peer-link fault, each obtains the status of its own uplink port and sends the status to an access device. When the access device determines that the status of the master device's uplink port is faulty and the status of the slave device's uplink port is normal, the access device sends a deconfiguration message to the slave device, instructing the slave device to deconfigure its own aggregation group port to the fault-closed state according to the deconfiguration message. As a result, traffic from the access device can be sent through the aggregation group port and uplink port of the slave device, avoiding the problem of traffic from the access device being unable to be forwarded when the peer-link fault occurs and the master device's uplink port fails while the slave device's uplink port is normal, thereby improving the robustness of the MLAG system. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1 This is a network example diagram of the MLAG system provided by the present invention.

[0042] Figure 2 This is an example diagram of an uplink path failure provided by the present invention.

[0043] Figure 3This is an example diagram of the node device role change provided by the present invention.

[0044] Figure 4 This is an example flowchart of the fault handling method provided by the present invention.

[0045] Icons: 10 - Node device 1; 20 - Node device 2; 30 - Access device. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0047] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0048] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0049] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or is the orientation or position relationship in which the product of the invention is usually placed when in use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be understood as a limitation on the present invention.

[0050] In addition, the terms "first", "second", etc., if used, are merely used to distinguish and describe, and should not be understood as indicating or implying relative importance.

[0051] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.

[0052] The inventors conducted an in-depth analysis of the situation after the peer-link failure between the master device and the slave device in the prior art and found that there are mainly two situations: (1) the uplink fails first, then the peer-link fails; (2) the peer-link fails, then the uplink fails again; for situation (1), the above-mentioned solution is feasible, and the traffic can still be forwarded normally; for situation (2), the above-mentioned solution in this scenario still cannot make the traffic forwarded normally. Therefore, the inventors conducted a further in-depth analysis of the processing after the peer-link failure, please refer to Figure 3 , Figure 3 This is an example diagram of the node device role change provided by the present invention. Figure 3 In the MLAG system, when the system is operating normally, two node devices elect a master-slave role, one becoming the master (MASTER) and the other becoming the slave (SLAVE). When the peer-link between the two fails and the keepalive link is normal, the slave device detects the peer-link failure. To avoid a "dual master" situation, the slave device's MLAG port is disconnected, and the slave device enters a suspended state. At this point, the slave device's role becomes SLAVE_SUSPEND (slave suspend) and the master device's role becomes MASTER_ALONE (master single active). The node devices believe that only the master device exists in the active-active system, and only the master device's MLAG port forwards traffic normally. At this time, if the keepalive link fails and the master device's uplink channel is normal, according to the aforementioned solution, traffic will still be forwarded normally through the master device. If the master device's uplink channel fails, traffic cannot be forwarded through either the master device or the slave device. The inventor believes that the reason for the aforementioned forwarding failure is essentially that the fault environment is not flexibly judged based on the current environmental status of the master and slave devices, ultimately finding a path that satisfies traffic forwarding and effectively improving the robustness of the MLAG system.

[0053] In view of this, this embodiment provides a fault handling method and MLAG system. In the event of a peer-link failure, the method comprehensively analyzes the status of the uplink ports of the master and slave devices, flexibly determines the fault environment based on the current status of the uplink ports of the master and slave devices, and ultimately selects a path that meets traffic forwarding requirements. This method is described in detail below.

[0054] Please refer to Figure 4 , Figure 4 This is a flowchart of a troubleshooting method provided by the present invention, which includes the following steps:

[0055] Step S101: When a peer-link failure is detected, the master device obtains the status of its own uplink port and sends the status to the access device.

[0056] Step S102: When a peer-link failure is detected, the slave device obtains the status of its own uplink port and sends the status to the access device.

[0057] In this embodiment, the master device and the slave device can each obtain the status of their own uplink port and send it to the access device in the same way. Since the peer-link is disconnected at this time, normal message interaction can no longer be carried out between the master and the slave devices. However, when the master and the slave devices are performing peer-link processing, they first send the status of their respective uplink ports to the access device. The access device performs a comprehensive analysis of the status of the uplink ports of the master and the slave devices, and then makes a reasonable judgment to set the aggregation group port (MLAG group port) of the master and the slave devices to an appropriate state, thereby ensuring the forwarding of access device traffic to the greatest extent.

[0058] In step S103, when the access device determines that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, the access device sends a cancel configuration message to the slave device to instruct the slave device to cancel the operation of configuring its own aggregation group port to the fault closed state (errdisable state) according to the cancel configuration message, and receives the traffic of the access device through its own aggregation group port and sends it out through the uplink port of the slave device.

[0059] In this embodiment, when the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, at this time, although the master device cannot forward the traffic of the access device, the slave device can. In order to allow the slave device to forward the traffic normally, at this time, the aggregation group port of the slave device cannot be set to the fault closed state according to the logic of the existing technology, but its processing flow needs to be modified, that is, the aggregation group port of the slave device is not set to the fault closed state, so as to ensure that the traffic of the access device can be sent normally to the aggregation group port of the slave device, and then sent out through the uplink port of the slave device.

[0060] The method provided in this embodiment comprehensively analyzes the status of the uplink ports of the master and slave devices, flexibly determines the fault environment based on the current status of the uplink ports of the master and slave devices, and ultimately selects a path that can meet traffic forwarding requirements, thereby improving the robustness of the MLAG system.

[0061] To minimize changes to the current networking and fully reuse existing processing flows, this embodiment further provides a method for obtaining the status of its own uplink port and sending it to the access device. This method can be used for both master devices and slave devices:

[0062] First, obtain the status of its own uplink port;

[0063] Secondly, generate Link Aggregation Control Protocol LACP packets;

[0064] Third, add the status of its own uplink port to the LACP message;

[0065] Finally, the status of its own uplink port is sent to the access device through LACP packets.

[0066] In this embodiment, LACP (Link Aggregation Control Protocol) messages are used. LACP is a protocol that implements dynamic link aggregation and de-aggregation. In the LACP protocol, two devices exchange aggregated link information by sending and receiving LACP messages, thereby aggregating multiple physical links into a single logical link. This aggregation method can improve network bandwidth and reliability while also reducing network management complexity. Because the peer link between the master and slave devices has failed, LACP messages cannot be forwarded via the peer link. However, the master and slave devices can still communicate with the access device. Therefore, the master and slave devices add the status of their respective uplink ports to the LACP messages and send these to the access device via the LACP messages. To avoid disrupting the normal LACP message processing flow, in this embodiment, the master and slave devices add the status of their respective uplink ports to the extension field of the LACP messages, thereby reusing the LACP message forwarding process without impacting the existing forwarding process.

[0067] If the uplink port on the master device fails, to prevent the access device from sending traffic to the master device, the master device needs to perform the following operations after sending the status of its uplink port to the access device:

[0068] When the master device determines that the state of its own uplink port is faulty, it sets the state of its own aggregation group port to disconnected.

[0069] In this embodiment, the master device may have one or more aggregation group ports. When there are multiple aggregation group ports, the uplink port of the master device corresponds to at least one of the multiple aggregation group ports of the master device. To avoid affecting other aggregation group ports of the master device, the following implementation is performed:

[0070] When the master device determines that the state of its own uplink port is faulty, it sets the states of all aggregation group ports corresponding to its own uplink port to disconnected.

[0071] For example, the aggregation group ports of the main device include port 1 and port 2, and the uplink ports include port a and port b. Port a corresponds to port 1, and port b corresponds to port 2. If port a fails, the corresponding port 1 needs to be set to disconnect, while port 2 still maintains the existing logic for traffic forwarding.

[0072] When the master device detects a peer-link failure, if its uplink port is in normal state at the time, this embodiment further provides a processing method to prevent the slave device from being unable to forward traffic when its uplink port is in normal state after the state of the master device's uplink port changes from normal to faulty during the process:

[0073] First, when the master device determines that the status of its own uplink port is normal, it obtains the status of its own uplink port in real time within a preset time period and synchronizes it to the slave device;

[0074] Secondly, when the slave device detects that the status of the uplink port of the master device changes from normal to faulty and the status of its own uplink port is normal and the status of the aggregation group port is fault-disabled, it clears the fault-disabled status of its own aggregation group port to receive the traffic of the access device through its own aggregation group port and send it out through its own uplink port.

[0075] In this embodiment, the preset duration can be set according to actual needs. The master device can obtain the status of its upstream port in real time by periodically obtaining it and synchronizing it to the slave device so that the slave device can promptly know the status changes of the master device's upstream port.

[0076] When the slave device promptly learns that the uplink port of the master device has changed from normal to faulty, and its own uplink port is in normal state, it means that traffic can now be forwarded normally through the slave device. Therefore, if the slave device finds that its own aggregation group port is in the fail-closed state at this time, it will cancel the fail-closed state of its own aggregation group port.

[0077] To ensure that the synchronized uplink port is synchronized within a preset time, this implementation provides a method for obtaining the status of its own uplink port in real time within the preset time and synchronizing it to the slave device:

[0078] First, create a timer with a preset period.

[0079] Secondly, within the timing period, the state of its own upstream port is obtained in real time and synchronized to the slave device.

[0080] In this embodiment, a timer is created to ensure that the status of the uplink port is synchronized within a preset time period, while also ensuring that subsequent processes are processed promptly after the preset time period has expired. Uplink port synchronization is performed when the timer's timing period expires, and is not performed after the timer's timing period expires.

[0081] To minimize changes to the current networking and fully reuse existing processing flows, this embodiment provides a method for synchronizing the status of the master device's uplink port to the slave device:

[0082] First, generate a keep-alive protocol message;

[0083] Secondly, the state of its own uplink port is added to the keep-alive protocol message;

[0084] Finally, the state of its own uplink port is synchronized to the slave device through the keepalive protocol message and the keepalive link.

[0085] In this embodiment, the keep-alive link is a link that already exists during the processing. As an implementation method, the master device adds the status of its own uplink port to the extended field of the keep-alive protocol message, thereby synchronizing the status of the master device's uplink port to the slave device by reusing the keep-alive link without adding additional links between the master and slave devices and without changing the existing fields in the keep-alive protocol message.

[0086] If the slave device finds that the uplink port of the master device has been normal within the preset time, the role of the slave device is suspend state, and the slave device will configure its own aggregation group port to the fault closed state. If the uplink port of the master device fails afterwards, when the timer's timing cycle arrives, the master device will synchronize the state of its uplink port to the slave device. The slave device can still promptly know the changes of the uplink port of the master device, and when the uplink port of the master device changes from normal to faulty, it will promptly cancel the configuration of its own aggregation group port to the fault closed state, so that the traffic of the access device is forwarded from the aggregation group port and the uplink port of the slave device.

[0087] When a slave device detects a peer-link failure, in order to set the correct role of the slave device according to the current processing logic when the status of the master device's uplink port remains normal, this embodiment further provides a processing method after the slave device obtains the status of its own uplink port and sends it to the access device:

[0088] If no cancellation message is received within the preset time, the aggregation group port itself will be configured as a fail-closed state.

[0089] Based on the same inventive concept, this embodiment also provides an MLAGA system. The MLAG system includes a master device and a slave device. A peer-link exists between the master device and the slave device. Access devices are dual-homed to the MLAG system.

[0090] When a peer-link failure is detected, the master device obtains the status of its own uplink port and sends it to the access device.

[0091] When a peer-link failure is detected, the slave device obtains the status of its own uplink port and sends it to the access device.

[0092] The access device is used to send a cancellation message to the slave device when it determines that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, so as to instruct the slave device to cancel the operation of configuring its own aggregation group port to be in a fault-closed state according to the cancellation message, and receive the traffic of the access device through its own aggregation group port and send it out through the uplink port of the slave device.

[0093] In an optional implementation, the master device or the slave device is further used to: obtain the status of its own uplink port; generate a Link Aggregation Control Protocol LACP message; add the status of its own uplink port to the LACP message; and send the status of its own uplink port to the access device via the LACP message.

[0094] In an optional implementation manner, the master device is further configured to: when determining that the status of its own uplink port is faulty, set the status of its own aggregation group port to disconnected.

[0095] In an optional embodiment, the master device has multiple aggregation group ports, and the uplink port of the master device corresponds to at least one of its own multiple aggregation group ports. The master device is also used to: when the master device determines that the status of its own uplink port is faulty, the status of all aggregation group ports corresponding to its own uplink port is set to disconnected.

[0096] In an optional embodiment, the master device is also used to: when determining that the status of its own uplink port is normal, obtain the status of its own uplink port in real time within a preset time period and synchronize it to the slave device; the slave device is also used to: when detecting that the status of the uplink port of the master device changes from normal to faulty and the status of its own uplink port is normal and the status of the aggregation group port is a faulty closed state, clear the faulty closed state of its own aggregation group port, so as to receive the traffic of the access device through its own aggregation group port and send it out through its own uplink port.

[0097] In an optional implementation manner, the master device is further configured to: create a timer with a timing period of a preset duration; and obtain the status of its own uplink port in real time during the timing period and synchronize it to the slave device.

[0098] In an optional embodiment, a keep-alive link exists between the master device and the slave device, and the master device is further used to: generate a keep-alive protocol message; add the status of its own uplink port to the keep-alive protocol message; and synchronize the status of its own uplink port to the slave device via the keep-alive protocol message and the keep-alive link.

[0099] In an optional implementation manner, the slave device is further configured to: if no cancellation configuration message is received within a preset time period, configure its own aggregation group port to a fail-closed state.

[0100] In summary, an embodiment of the present invention provides a fault handling method and an MLAG system, which are applied to an MLAG system of an inter-device link aggregation group. The MLAG system includes a master device and a slave device. A peer link is established between the master device and the slave device. The access device dual-homes to the MLAG system. The method includes: when the master device detects a peer-link fault, it obtains the status of its own uplink port and sends it to the access device; when the slave device detects a peer-link fault, it obtains the status of its own uplink port and sends it to the access device; when the access device determines that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, the access device sends a cancellation message to the slave device to instruct the slave device to cancel the operation of configuring its own aggregation group port to a fault-closed state according to the cancellation message, and receives traffic from the access device through its own aggregation group port and sends it out through the uplink port of the slave device. Compared with the prior art, the embodiments of the present invention have at least the following advantages: (1) by comprehensively analyzing the status of the uplink ports of the master and slave devices, the fault environment can be flexibly judged based on the current status of the uplink ports of the master and slave devices, and finally a path that can meet the traffic forwarding requirements is selected, thereby improving the robustness of the MLAG system; (2) when the uplink port of the master device fails first and the peer-link fails later, the master and slave devices can use LACP messages to send the status of the uplink ports to the access device. When the peer-link fails first and the uplink port of the master device fails later, the keep-alive protocol message is used to enable the slave device to promptly know the status change of the uplink port of the master device, fully reuse the existing processing flow, avoid changing the current networking as much as possible, and reduce the cost of improvement; (3) by using a timer, it can ensure the synchronous processing of the status of the uplink port and ensure that the timely processing of the subsequent process is not affected after the preset time period, fully reuse the existing peer-link fault processing flow.

[0101] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A fault handling method, characterized in that: The method is applied to an MLAG system including a master device and a slave device, wherein a peer-link is established between the master device and the slave device, and an access device is dual-homed to the MLAG system. The method includes: When the master device detects the peer-link failure, it obtains the status of its own uplink port and sends it to the access device; When the slave device detects the peer-link failure, it obtains the status of its own uplink port and sends it to the access device; When the access device determines that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, the access device sends a cancellation configuration message to the slave device to instruct the slave device to cancel the operation of configuring its own aggregation group port to a fault-closed state according to the cancellation configuration message, and receive the traffic of the access device through its own aggregation group port and send it out through the uplink port of the slave device.

2. The fault handling method according to claim 1, wherein: The step of obtaining the status of the upstream port thereof and sending the status to the access device comprises: Get the status of its own uplink port; Generate Link Aggregation Control Protocol LACP packets; Adding the state of the own uplink port to the LACP message; The state of its own uplink port is sent to the access device through the LACP message.

3. The fault handling method according to claim 1, wherein: When the master device detects the peer-link failure, after obtaining the status of its own uplink port and sending it to the access device, the method includes: When the master device determines that the state of its own uplink port is faulty, it sets the state of its own aggregation group port to disconnected.

4. The fault handling method according to claim 3, wherein: The master device has multiple aggregation group ports, the uplink port of the master device corresponds to at least one of the multiple aggregation group ports of the master device, and when the master device determines that the status of its uplink port is faulty, the step of setting the status of its aggregation group port to disconnected includes: When the master device determines that the state of its own uplink port is faulty, it sets the states of all aggregation group ports corresponding to its own uplink port to disconnected.

5. The fault handling method according to claim 1, wherein: The method further comprises: When the master device determines that the state of its own uplink port is normal, the master device obtains the state of its own uplink port in real time within a preset time period and synchronizes it to the slave device; When the slave device detects that the state of the uplink port of the master device changes from normal to faulty and the state of its own uplink port is normal and the state of the aggregation group port is in a fault-closed state, it clears the fault-closed state of its own aggregation group port to receive the traffic of the access device through its own aggregation group port and send it out through its own uplink port.

6. The fault handling method according to claim 5, characterized in that: The step of acquiring the status of the upstream port of the slave device in real time within a preset time period and synchronizing the status to the slave device includes: Create a timer with a timing period of the preset duration; During the timing period, the state of its own uplink port is acquired in real time and synchronized to the slave device.

7. The fault handling method according to claim 5 or 6, characterized in that: There is a keep-alive link between the master device and the slave device, and the step of obtaining the status of its own uplink port and synchronizing it to the slave device includes: Generate keepalive protocol messages; Adding the state of its own uplink port to the keep-alive protocol message; The state of its own uplink port is synchronized to the slave device through the keep-alive protocol message and the keep-alive link.

8. The fault handling method according to claim 1, wherein: After the step of the slave device acquiring the status of its own uplink port and sending the status to the access device when detecting the peer-link failure, the method includes: If the deconfiguration message is not received within a preset time period, the aggregation group port itself is configured to be in a fail-closed state.

9. An MLAG system, characterized in that: The MLAG system includes a master device and a slave device. A peer-link is established between the master device and the slave device. Access devices are dual-homed to the MLAG system. The master device is configured to obtain the status of its own uplink port and send the status to the access device when detecting the peer-link failure; The slave device is used to obtain the status of its own uplink port and send it to the access device when detecting the peer-link failure; The access device is used to send a cancellation message to the slave device when it is determined that the status of the uplink port of the master device is faulty and the status of the uplink port of the slave device is normal, so as to instruct the slave device to cancel the operation of configuring its own aggregation group port to a fault-closed state according to the cancellation message, and receive the traffic of the access device through its own aggregation group port and send it out through the uplink port of the slave device.

10. The MLAG system according to claim 9, wherein: The master device is further configured to obtain the status of its own uplink port in real time within a preset time period and synchronize the status to the slave device when determining that the status of its own uplink port is normal; The slave device is also used to clear the fault-closed state of its own aggregation group port when it detects that the state of the uplink port of the master device changes from normal to faulty and the state of its own uplink port is normal and the state of the aggregation group port is a fault-closed state, so as to receive the traffic of the access device through its own aggregation group port and send it out through its own uplink port.

Citation Information

Cited By

  • Fault processing method and device, network equipment and program product

    CN121418343A

  • Link protocol-based dual-master switch processing method and electronic equipment

    CN121547393A