Loop prevention method and device in a switch MLAG dual-homing failure scenario

CN122741479APending Publication Date: 2026-09-11INSPUR NETWORK TECH (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610909819.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]但这种机制存在明显缺陷:MLAG成员端口的状态震荡与对端成员口的隔离操作为纯异步执行,且无消息确认、无重传、无超时等待,容易出现本端端口已UP、对端隔离未生效的时间窗口

Benefits of technology

1)本申请在MLAG双归故障恢复过程中,本端端口收到UP指令后,先向对端发送开启隔离消息,并对端回复隔离完成确认消息后,才执行端口UP操作。通过严格的先确认隔离、后恢复端口时序控制,从根本上消除端口已UP,但对端尚未隔离的时间窗口,彻底消除环路隐患。并且在故障端口恢复前,对端设备始终保持隔离状态生效,业务流量持续通过正常链路稳定转发,不会出现瞬时环路、MAC地址漂移与流量切换丢包等问题。仅在确认对端已完成隔离配置后,才将故障端口恢复UP并切换回双归模式,能够实现故障恢复全程业务无感知、流量无中断。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122741479A_ABST
    Figure CN122741479A_ABST
Patent Text Reader

Abstract

The application discloses a loop prevention method and equipment in a switch MLAG dual-homing fault scenario, belongs to the technical field of data communication, and is used for solving the loop hidden danger problem in the MLAG dual-homing fault recovery process. The method comprises the following steps: in the MLAG system, when a lower port of a first device receives a UP instruction, it is determined whether the lower port is an MLAG member port; if yes, under the condition that the MLAG system is in a normal state, the first device sends an opening isolation message to a second device at the opposite end, and waits to receive an isolation completion confirmation message returned from the second device; after the first device successfully receives the isolation completion confirmation message, the UP instruction is executed, and the MLAG member port is set to an UP state. The application realizes the sequence control of confirming isolation first and then recovering the port, and can effectively reduce the loop risk in the existing network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data communication technology, and in particular to a method and device for preventing loops in the MLAG dual-homing fault scenario of a switch. Background Technology

[0002] Multi-Chassis Link Aggregation (MLAG) is a cross-device link aggregation technology widely used in data centers and enterprise networks. It can virtualize two physical switches into one logical device, allowing downstream servers, switches, and other terminals to access the network in a dual-homed manner. This improves network reliability from the link level to the device level, effectively avoiding service interruptions caused by single points of failure.

[0003] In the current MLAG dual-homing fault recovery scenario, the recovery of member port status and the isolation configuration of peer member port are executed asynchronously: when a member port fails down, the local end sends a release isolation message to the peer end, and the peer end releases the isolation to ensure traffic bypass; when a member port recovers up, the local end sends an enable isolation message to the peer end, and the peer end reconfigures the isolation.

[0004] However, this mechanism has significant drawbacks: the state oscillations of MLAG member ports and the isolation operations of peer member ports are executed asynchronously, without message confirmation, retransmission, or timeout waiting. This easily leads to a time window where the local port is up, but the peer's isolation has not yet taken effect. For example, when the local port is up, but the peer has not initiated isolation in time due to CPU overload, peer-link congestion, packet loss, or processing delays, a Layer 2 loop will form, causing MAC address drift, broadcast storms, excessive device resource consumption, and packet loss, severely impacting network stability and service continuity. Furthermore, most mainstream vendors currently use similar mechanisms, relying solely on delayed UP, dual-master detection, loop detection, and STP for post-event remediation, failing to eliminate the loop risks inherent in asynchronous isolation operations at their root. Summary of the Invention

[0005] This application provides a loop prevention method and device for MLAG dual-homing fault scenarios of switches, which is used to solve the following technical problem: In the current MLAG dual-homing fault recovery scenario, the recovery of member port status and the isolation configuration of peer member port are executed asynchronously, which leads to the risk of loop in the MLAG system.

[0006] The technical solution adopted in this application is as follows: On the one hand, this application provides a loop prevention method in the MLAG dual-homing fault scenario of a switch. The method includes: in the MLAG system, when the downstream port of the first device receives a UP command, determining whether the downstream port is an MLAG member port; if so, under the condition that the MLAG system is in a normal state, the first device sends an isolation enable message to the second device at the other end and waits to receive an isolation completion confirmation message returned from the second device; after the first device successfully receives the isolation completion confirmation message, it executes the UP command to set the MLAG member port to the UP state.

[0007] In one possible implementation, after determining whether the downstream port is an MLAG member port, the method further includes: determining that the downstream port is not an MLAG member port; the first device directly executes the UP command to set the downstream port to the UP state.

[0008] In one possible implementation, the waiting to receive the isolation completion confirmation message returned from the second device includes: if the first device does not receive the isolation completion confirmation message within a preset waiting time, a retry mechanism is initiated, the timer is reset, and the first device resends the start isolation message to the second device; if the first device successfully receives the isolation completion confirmation message during the retry period, the UP command is executed to set the MLAG member port to the UP state; if the first device still does not receive the isolation completion message during the retry period, the first device actively reports an isolation confirmation timeout alarm, enables a local MAC jitter suppression strategy, and then forcibly executes the UP command to forcibly set the MLAG member port to the UP state.

[0009] In one possible implementation, the retry mechanism includes initiating a three-retry mechanism, wherein the retry waiting times in the three retry mechanisms are 100ms, 200ms and 400ms respectively.

[0010] In one possible implementation, the local MAC jitter suppression strategy includes at least one of MAC address learning restrictions, broadcast packet rate limiting, and unknown unicast rate limiting.

[0011] In one possible implementation, after determining that the downstream port is an MLAG member port, the method further includes: determining that the MLAG system is in an abnormal state, the abnormal state including a dual-master state or a second device failure state; determining the device role of the first device, the device role including at least a master device and a backup device; and determining whether to execute the UP command based on the device role of the first device.

[0012] In one possible implementation, the method further includes: after the first device restarts from a fault, determining whether the negotiation state of the MLAG system has been established normally; if so, when the first device receives the UP instruction, sending an open isolation message to the second device on the other end; if not, determining the device role of the first device, and determining whether to execute the received UP instruction based on the device role of the first device, wherein the device role includes at least a master device and a backup device.

[0013] In one possible implementation, after determining the device role of the first device, the method further includes: if the device role of the first device is a master device, then the first device directly executes the UP instruction to set the MLAG member port to the UP state; if the device role of the first device is a backup device, then the first device determines whether the MLAG member port is set to an Err-Disable state or a Down state; if so, then the first device refuses to execute the UP instruction; if not, then the first device forcibly executes the UP instruction to forcibly set the MLAG member port to the UP state.

[0014] In one possible implementation, when the first device forcibly executes the UP command, the method further includes: the MLAG system enabling loop detection, and if a loop is detected, setting the MLAG member port back to the Down state.

[0015] On the other hand, this application also provides an anti-loop device for a switch MLAG dual-homing failure scenario, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to execute an anti-loop method for a switch MLAG dual-homing failure scenario as described in any of the above implementations.

[0016] This application provides a loop prevention method and device for a switch MLAG dual-homing fault scenario, which has the following beneficial effects: 1) In the MLAG dual-homing fault recovery process, after receiving the UP command, the local port first sends an isolation enable message to the peer, and only executes the port UP operation after the peer replies with an isolation completion confirmation message. Through strict timing control of confirming isolation before restoring the port, the time window where the port is UP but the peer is not yet isolated is fundamentally eliminated, completely eliminating loop risks. Furthermore, before the faulty port is restored, the peer device maintains its isolation state, ensuring continuous and stable forwarding of service traffic through the normal link, preventing issues such as instantaneous loops, MAC address drift, and packet loss during traffic switching. Only after confirming that the peer has completed isolation configuration is the faulty port restored to UP and switched back to dual-homing mode, achieving seamless service operation and uninterrupted traffic throughout the fault recovery process.

[0017] 2) This application introduces a timeout and retransmission mechanism. If the local end does not receive a confirmation message from the peer indicating that isolation is complete within a set time, it will preferably perform three retransmissions to avoid accidental interruption of the process due to a single packet loss or a brief delay, thereby improving reliability. The timeout and retransmission mechanism is configurable and can flexibly set timeout parameters according to device performance and network environment, transforming traditional passive waiting into active retries and intelligent fault tolerance. This effectively addresses abnormal scenarios such as Peer-link congestion, CPU overload, and packet loss, significantly improving the reliability of message interaction and the robustness of the solution.

[0018] 3) This application also includes timeout escape handling. If no confirmation message indicating isolation completion is received after three retransmissions, an alarm message is proactively reported. This facilitates maintenance personnel in quickly locating the cause of the fault and promptly troubleshooting network problems, significantly improving the maintainability and efficiency of the equipment. Simultaneously, a local MAC jitter suppression policy is activated, and the UP command is forcibly executed to set the port to UP state, preventing service interruption issues where the local port can never recover in the event of a complete failure of the peer device or a permanent link failure.

[0019] 4) This application introduces role-based hierarchical decision-making when the MLAG status is abnormal. When the MLAG status is abnormal, such as dual-master status or suspected fault status of the peer, no isolation start message is sent. Instead, the judgment is made based on the device role and port Err-Disable flag to ensure that the port can still be reasonably restored in abnormal scenarios where normal negotiation is not possible, while avoiding dual-master loop.

[0020] 5) This application can reuse the MLAG keepalive link to transmit acknowledgment messages. The ACK messages for enabling isolation and confirming isolation completion are encapsulated in the existing MLAG protocol interaction messages, which use the Keepalive link without consuming additional bandwidth and CPU resources, and inherit the high reliability of the MLAG keepalive mechanism. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart of an anti-loop method for a switch in an MLAG dual-homing fault scenario is provided in this application; Figure 2 Example diagram of the MLAG system provided in this application; Figure 3 Example diagram of Peer-link provided in this application; Figure 4 Example diagram of dual relocation provided for this application; Figure 5 An example diagram of isolation provided for this application; Figure 6 A flowchart of port processing for the MLAG member port failure recovery scenario provided in this application; Figure 7 The interaction diagram between MLAG Peer devices provided in this application; Figure 8 This application provides a schematic diagram of the anti-loop device for a switch in the MLAG dual-homing fault scenario. Detailed Implementation

[0022] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0023] The asynchronous isolation sampling mechanism during MLAG dual-homing fault recovery introduces loop risks. With the large-scale commercial deployment of network equipment, once a Layer 2 loop occurs in the existing network, it will directly cause serious problems such as packet loss, transmission lag, data interruption, excessive CPU and memory usage of switches, and forwarding congestion, thereby affecting service continuity and causing direct economic losses to users.

[0024] This application proposes a proactive loop prevention mechanism to address this issue. By adding isolation message confirmation, timeout retransmission, and anomaly escape mechanisms to the MLAG control plane, it ensures that the peer has completed isolation configuration before the local port goes up. This addresses the loop risks caused by asynchronous operations at the root, significantly reduces the risk of loops in the existing network, and comprehensively improves the operational stability, network reliability, and ease of maintenance of the equipment.

[0025] Optionally, the loop prevention method proposed in this application covers two core scenarios: single-sided link fault recovery and single-sided device fault restart recovery. Unilateral link failure recovery: After receiving the UP command, the local port first sends an isolation enable message to the peer; the peer completes the isolation configuration and replies with an acknowledgment message. The local port then executes the port UP after receiving the acknowledgment; if no reply is received within the timeout period, a timeout alarm is triggered.

[0026] Unilateral device failure recovery: Before restarting the device due to failure and restoring the member port to UP, first send an isolation start message to the peer; after the peer completes the isolation and replies with confirmation, the local end will execute port UP after receiving the confirmation; if no reply is received within the timeout period, a timeout alarm will be triggered.

[0027] The method in this application will be described in detail below with reference to the accompanying drawings.

[0028] Figure 1 A flowchart of an anti-loop method for a switch in an MLAG dual-homing fault scenario provided in this application is shown below. Figure 1 As shown, the method in this application includes at least the following execution steps: Step 101: In the MLAG system, when the downstream port of the first device receives the UP command, determine whether the downstream port is an MLAG member port.

[0029] Example diagram of MLAG system as follows Figure 2 As shown, an example diagram of Peer-link is as follows. Figure 3 As shown in the example diagram of double regression, see below. Figure 4 As shown, an example diagram of isolation is as follows. Figure 5 As shown. Figure 2 In the MLAG system shown, two physical switches are virtualized as a single logical device, with downstream servers, switches, and other terminals accessing via a dual-homed approach. This elevates network reliability from the link level to the device level, effectively preventing service interruptions caused by single points of failure. The peer-link is the core Layer 2 interconnection link between MLAG peers, undertaking functions such as fault traffic backup, MAC entry synchronization, and protocol negotiation. It is primarily used for fault traffic backup forwarding and protocol message exchange. Figure 3In this context, when a single downstream link fails, network traffic can be routed to the peer device via a peer link, ensuring uninterrupted service. Dual-homing refers to a downstream device simultaneously connecting to two peers in the MLAG system via two or more links, such as... Figure 4 The two eth1 paths in the MLAG device belong to the same portchannel1 group, forming a dual-homing aggregation group that includes both local and peer member ports. Isolation is to prevent Layer 2 loops, ensuring unidirectional logical isolation between the peer link and member ports. Figure 5 As shown in the isolation diagram, the MLAG device on the right prohibits the forwarding of packets received from the Peer-link from the member port outwards, i.e., the isolation prohibits forwarding to downstream devices, thereby avoiding the formation of loops.

[0030] In a dual-homed MLAG networking scenario, when the MLAG member port of a device fails, the downstream device will switch from dual-homed forwarding mode to single-homed forwarding mode, and the service traffic will only be transmitted through the normal link.

[0031] Figure 6 This is a flowchart illustrating the port handling process for MLAG member port failure recovery scenarios. Figure 6 As shown, after the downstream port of the first device receives a UP command or an interface UP command, it first determines whether the downstream port is an MLAG member port. If the downstream port receiving the UP command is an MLAG member port, then the fault recovery control process described below is entered.

[0032] As one possible implementation, if the downstream port of the received UP command or interface UP command is not a member port of MLAG, the first device directly executes the UP command, for example, by directly sending the UP command to the underlying driver to set the downstream port to the UP state.

[0033] Step 102: If it is an MLAG member port, then when the MLAG system is in normal condition, the first device sends an isolation start message to the second device on the other end, and waits to receive an isolation completion confirmation message returned from the second device.

[0034] When the downstream port receiving the UP command is an MLAG member port, and the faulty MLAG member port receives the UP command, the system first checks whether the MLAG system is in a normal state, such as whether the MLAG negotiation state has been established normally.

[0035] When the MLAG system is in normal operation, the first device sends an isolation enable request message for the member port to the second device on the other end, and waits to receive an isolation completion confirmation message from the second device. Optionally, after receiving the isolation enable message, the second device immediately configures unidirectional forwarding isolation between the corresponding MLAG member port on its end and the Peer-link, prohibiting the forwarding of packets received from the Peer-link through that member port, and replies to the first device with an isolation completion confirmation message, optionally in the form of an ACK packet.

[0036] Step 103: After the first device successfully receives the isolation completion confirmation message, execute the UP command to set the MLAG member port to the UP state.

[0037] When the first device receives an ACK message containing confirmation that the second device has completed isolation, it executes the UP command to set the local faulty port to the UP state. The downstream device then smoothly resumes dual-homing forwarding from single-homing, and the process ends.

[0038] Figure 7 The diagram shows the interaction between MLAG Peer devices provided in this application. Figure 7 As shown, when device 1 recovers from a local fault, its member port receives a UP command and sends an enable isolation message to device 2. After receiving the enable isolation message, device 2 enables isolation and returns a recovery message ACK to device 1. After receiving the ACK, device 1 executes the UP command to restore port UP, completing the local fault recovery.

[0039] As one possible implementation, such as Figure 6As shown, if the first device does not receive an isolation completion confirmation message within the preset waiting time (i.e., does not receive an ACK from the peer after a timeout), optionally, the preset waiting time is 100ms. This could be due to reasons such as Keepalive link congestion, device CPU overload, or packet transmission delay causing the response to not be returned in time. In this case, the system will initiate a retry mechanism to adapt to short-term abnormal scenarios such as transient network congestion. Optionally, the retry mechanism is a three-retry mechanism, that is, resending the isolation start message to the second device and attempting to resend three times if no isolation completion confirmation message is received, with waiting times of 100ms, 200ms, and 400ms respectively, minimizing port recovery latency and reducing the impact on service forwarding performance while ensuring reliability. If an isolation confirmation message (ACK) is received from the second device during the three retries, the first device executes the UP command, sets the MLAG member port to the UP state, and the process ends. However, if the first device still does not receive the isolation completion confirmation message after three retries, the system will activate the timeout escape mechanism to prevent the local port from failing to recover due to extreme situations such as the complete downtime of the peer device or the interruption of the peer-link. At this time, the first device will proactively report the isolation confirmation timeout alarm, throw the alarm information, and simultaneously enable the local MAC jitter suppression policy. Afterward, the first device will forcibly execute the UP command to force the MLAG member ports to be set to the UP state to ensure service connectivity.

[0040] Furthermore, the aforementioned local MAC jitter suppression strategy includes at least one of MAC address learning restrictions, broadcast packet rate limiting, and unknown unicast rate limiting.

[0041] As one possible implementation, such as Figure 6 As shown, if an abnormal state of the MLAG system is detected after confirming that the downstream port is an MLAG member port, it indicates that the current network may be in a dual-master state or the peer device may be faulty. Since only the primary device's port is in a normal forwarding state in a dual-master scenario, the system will further determine whether the first device is a primary or backup device: If the first device is the master device, the first device will directly execute the UP command to restore the faulty port to the UP state, and the process will end. If the first device is the backup device, the first device first checks whether the member port has been set to Err-Disable or Down state due to dual-master protection: if it has been set to Err-Disable or Down state, there is no need to execute the UP command to restore the port, and the process ends directly; if it has not been set to Err-Disable or Down state, it indicates that the original master device may have failed. In this case, the first device directly forces the UP command to restore the faulty port on this end to the UP state, ensuring that the service is not interrupted.

[0042] Optionally, the above process also applies to recovery scenarios after a complete device failure and restart. After the first device fails and restarts, the MLAG role may switch, and the role determination logic in the process will automatically adapt to the new role; if the peer device has not yet completed MLAG negotiation, the system will directly follow the established procedure. Figure 6 The branch logic handling for MLAG state anomalies shown in the diagram ensures full-scenario coverage and robustness. Specifically, after the first device restarts from a fault, it is determined whether the negotiation state of the MLAG system has been established normally. If it has been established normally, when the first device receives the UP command, it sends an isolation enable message to the second device on the other end. If it has not been established normally, it is determined whether the device role of the first device is the primary device or the backup device, and then it is determined whether to execute the received UP command based on the device role of the first device.

[0043] As one possible implementation, this application, when the first device does not receive an isolation response, or when the first device is a backup device but its port is not set to Err-Disable or Down state, forcibly executes an UP command to force the MLAG member port to the UP state. At this time, the second device on the other end may not have yet set up isolation or completed isolation. To further enhance the reliability of the MLAG system, the following additional safeguards can be adopted: 1) Enable loop detection. If a loop is detected, actively shut down the MLAG member port again. 2) If an isolation completion confirmation message is received from the peer device after sending the isolation start message and no isolation completion confirmation message is received during the forced UP period, the confirmation process will be entered again. However, in order to ensure the stability of the environment, the status of the MLAG member interface will not change and will remain in the UP state.

[0044] 3) If the MLAG system malfunctions and is forced to UP without sending an isolation enable message, and if the MLAG system status returns to normal, the process of sending an isolation enable message to the second device on the other end will continue to be triggered. However, in order to ensure the stability of the environment, the interface status will not change and will remain in the UP state.

[0045] To ensure the reliability of the MLAG system, this application also includes the following additions: The isolation confirmation message in this application is based on the original MLAG system isolation notification message, with the only addition being the ACK response field for the device. The message encapsulation format and transmission channel remain unchanged. It reuses the existing Keepalive link of the upgraded MLAG system for transmission, without consuming additional bandwidth and CPU resources. At the same time, it relies on MLAG's own keepalive mechanism to improve the reliability of message transmission.

[0046] During the port's isolation confirmation period, the downstream device continuously and stably forwards traffic via a single-homed link, without path switching or table entry oscillation, thus preventing handover packet loss. The port only returns to UP state and enables dual-homed forwarding after receiving an isolation completion confirmation message from the peer. If a timeout-forced UP is triggered, the local MAC jitter suppression strategy effectively reduces the risks of broadcast storms, MAC drift, and packet loss caused by loops. Therefore, the loop prevention method in this application eliminates loop risks at their source while maximizing the integrity and continuity of service traffic.

[0047] Furthermore, the anti-loop method in this application ensures the reliability of isolation configuration in normal scenarios while fully compatibility with abnormal scenarios where isolation messages cannot be sent. When a member port receives a UP command but the MLAG system status is abnormal and unable to send isolation messages, the system will decide whether to execute port UP according to the device role hierarchy. In dual-master mode, if the backup device port has been set to non-forwarding state, the process will be terminated directly, perfectly compatible with the existing MLAG dual-master processing mechanism; in other scenarios, the port's UP state will be directly restored to avoid the local port becoming inaccessible due to abnormal peer device.

[0048] By employing the methods described above, this application can fundamentally resolve loop hazards in the MLAG dual-homing fault recovery process, significantly reduce the risk of loops in the existing network, and comprehensively improve the operational stability, network reliability, and ease of maintenance of the equipment.

[0049] Based on the same inventive concept, this application also provides an anti-loop device for MLAG dual-homing fault scenarios of switches, the structure of which is as follows: Figure 8 As shown.

[0050] Figure 8 This application provides a structural schematic diagram of an anti-loop device for a switch operating under an MLAG dual-homing fault scenario. (See attached diagram.) Figure 8 As shown, the anti-loop device 800 for a switch MLAG dual-homing failure scenario in this application specifically includes: at least one processor 801; and a memory 803 communicatively connected to at least one processor 801 (connected via a bus 802); wherein the memory 803 stores instructions executable by at least one processor 801, so that at least one processor 801 can execute an anti-loop method for a switch MLAG dual-homing failure scenario as described in the above embodiments.

[0051] In one possible implementation of this application, the aforementioned processor 801 is configured to perform the following: In the MLAG system, when the downstream port of the first device receives the UP command, it determines whether the downstream port is an MLAG member port; if so, under the condition that the MLAG system is in a normal state, the first device sends an isolation enable message to the second device at the other end and waits to receive an isolation completion confirmation message returned from the second device; after the first device successfully receives the isolation completion confirmation message, it executes the UP command to set the MLAG member port to the UP state.

[0052] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0053] The equipment and method provided in this application are one-to-one correspondences. Therefore, the equipment also has similar beneficial technical effects as its corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the equipment will not be repeated here.

[0054] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0055] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0056] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A loop prevention method for a switch in an MLAG dual-homing fault scenario, characterized in that, The method includes: In the MLAG system, when the downstream port of the first device receives the UP command, it is determined whether the downstream port is an MLAG member port; If so, under normal conditions of the MLAG system, the first device sends an isolation start message to the second device on the other end and waits to receive an isolation completion confirmation message returned from the second device. After the first device successfully receives the isolation completion confirmation message, it executes the UP command to set the MLAG member port to the UP state.

2. The method for preventing loops in a switch MLAG dual-homing fault scenario according to claim 1, characterized in that, After determining whether the downstream port is an MLAG member port, the method further includes: It has been determined that the downstream port is not an MLAG member port; The first device directly executes the UP command to set the downstream port to the UP state.

3. The method for preventing loops in a switch MLAG dual-homing fault scenario according to claim 1, characterized in that, The waiting to receive the isolation completion confirmation message returned from the second device includes: If the first device does not receive the isolation completion confirmation message within the preset waiting time, a retry mechanism is initiated, the timer is reset, and the first device resends the isolation start message to the second device. If the first device successfully receives the isolation completion confirmation message during the retry period, it executes the UP command to set the MLAG member port to the UP state. If the first device still does not receive the isolation completion message during the retry period, the first device will actively report an isolation confirmation timeout alarm, enable the local MAC jitter suppression policy, and then forcibly execute the UP command to force the MLAG member port to be set to UP state.

4. The loop prevention method for a switch in an MLAG dual-homing fault scenario according to claim 3, characterized in that, The startup retry mechanism includes: A three-retry mechanism is initiated, with retry waiting times of 100ms, 200ms, and 400ms respectively.

5. The method for preventing loops in a switch MLAG dual-homing fault scenario according to claim 3, characterized in that, The local MAC jitter suppression strategy includes at least one of MAC address learning restriction, broadcast packet rate limiting, and unknown unicast rate limiting.

6. The method for preventing loops in a switch MLAG dual-homing fault scenario according to claim 1, characterized in that, After determining that the downstream port is an MLAG member port, the method further includes: The MLAG system is determined to be in an abnormal state, which includes a dual master state or a second device failure state. Determine the device role of the first device, wherein the device role includes at least a primary device and a backup device; Whether to execute the UP command is determined based on the device role of the first device.

7. The method for preventing loops in a switch MLAG dual-homing fault scenario according to claim 1, characterized in that, The method further includes: After the first device restarts from a fault, determine whether the negotiation state of the MLAG system has been established normally; If so, when the first device receives the UP command, it sends an isolation enable message to the second device on the other end; If not, the device role of the first device is determined, and the received UP instruction is executed based on the device role of the first device. The device role includes at least a master device and a backup device.

8. A loop prevention method for a switch in an MLAG dual-homing fault scenario according to claim 6 or 7, characterized in that, After determining the device role of the first device, the method further includes: If the device role of the first device is the master device, then the first device directly executes the UP instruction to set the MLAG member port to the UP state; If the device role of the first device is a backup device, the first device determines whether the MLAG member port is set to Err-Disable or Down state; If so, the first device refuses to execute the UP command; If not, the first device will force the UP command to set the MLAG member port to the UP state.

9. The method for preventing loops in a switch MLAG dual-homing fault scenario according to claim 1, characterized in that, When the first device forcibly executes the UP command, the method further includes: The MLAG system enables loop detection, and if a loop is detected, the MLAG member port is set to the Down state again.

10. A loop prevention device for MLAG dual-homing fault scenarios in switches, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, enabling the at least one processor to execute a loop prevention method for a switch MLAG dual-homing fault scenario according to any one of claims 1-9.