Port abnormity automatic recovery method, device, equipment, medium and product
By implementing a tiered recovery strategy and status monitoring for switch ports, switch port faults are automatically repaired, solving the problems of high costs and low efficiency in manual operation and maintenance. This enables fast and accurate fault recovery, improving network availability and service quality.
Patent Information
- Application Number
- CN202511650314.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-01-13
AI Technical Summary
In existing technologies, switch port fault handling relies on manual operation and maintenance, resulting in high costs and low efficiency. In particular, the fault recovery cycle is long in remote unattended sites, making it difficult to meet the needs of enterprise-level networks and data centers for high availability and low-cost operation and maintenance.
By monitoring port link status changes in real time, a tiered recovery strategy is adopted, which involves disabling and re-enabling ports, simulating offline and online optical modules using software, and restarting the device. This strategy combines optical module status signals and DDM parameters to select effective recovery scenarios, differentiates recovery time settings, and limits restart frequency, thereby achieving automated and accurate fault repair.
Faults can be quickly repaired without human intervention, shortening fault recovery time from hours or days to minutes, reducing business interruption time, improving network availability and service quality, and avoiding unnecessary operations and network instability.
Smart Images

Figure CN121333893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, device, medium, and product for automatic recovery of port anomalies. Background Technology
[0002] In network communication systems, switches serve as the infrastructure for data forwarding and link interconnection, and the stable operation of their ports directly determines the continuity and reliability of network transmission. With the continuous expansion of enterprise networks, data centers, and remote office networks, switch ports need to simultaneously connect multiple types of network components such as servers, terminal devices, and edge computing nodes, undertaking the real-time transmission of massive amounts of business data (such as transaction data, office documents, and multimedia streams). Especially in unattended sites in remote areas, the continuous availability of switch ports is directly related to the coverage capability of regional network services. Once a port fails or is interrupted, it may not only cause local network paralysis but also lead to serious consequences such as business stagnation and data transmission interruption, placing extremely high demands on network service quality and business continuity.
[0003] Currently, handling switch port faults mainly relies on traditional manual maintenance. When a port fails due to factors such as physical link fluctuations, protocol negotiation anomalies, electrostatic interference, or momentary port chip failure, maintenance personnel must follow the process of fault reporting, on-site troubleshooting, and manual repair. After arriving at the fault site, they attempt to resolve the problem by restarting the switch, re-plugging and unplugging physical cables, or replacing port hardware modules. For some intermittent faults (such as momentary port unresponsiveness caused by environmental electromagnetic interference), multiple trips to the site may be necessary for repeated troubleshooting to confirm whether the fault has been completely eliminated.
[0004] However, the traditional manual operation and maintenance model has the problem of high cost. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and product for automatic recovery of port anomalies, which can reduce operation and maintenance costs.
[0006] To achieve the above objectives, this application adopts the following technical solution: Firstly, this application provides a method for automatic recovery from port anomalies, including: Acquire the link status change signal of the port; When the link status change signal is detected to change from connected to disconnected, a first recovery is performed, which includes disabling the port and then re-enabling the port; after performing the first recovery, wait for a first duration and obtain the first recovery status of the port; If the first recovery status is not recovered, a second recovery is performed, which includes simulating the optical module being offline and then simulating the optical module being online again; after performing the second recovery, wait for a second duration and obtain the second recovery status of the port; the first duration is less than the second duration; If the second recovery status is not recovered, a third recovery is performed, which includes restarting the device.
[0007] Optionally, performing the first recovery includes: Acquire the optical module status signal of the optical port, the optical module status signal including a first status signal, a second status signal and a third status signal; When the first status signal indicates that the optical module is in place and the second status signal indicates that the optical port is not present and there is a loss of optical signal (LOS) alarm, the first recovery is performed.
[0008] Optionally, the method further includes: Based on the digital diagnostic monitoring (DDM) parameters in the third state signal, determine whether the recovery conditions are met; If the recovery conditions are met, then perform the first recovery.
[0009] Optionally, after performing the third recovery and restart of the device, the method further includes: Restore monitoring of the port; Set a restart time interval. If the port fails again within the restart time interval, the third recovery will be prevented from being executed again, and the first recovery and the second recovery will be executed instead.
[0010] Optionally, the DDM parameters include optical module temperature, operating voltage, bias current, received optical power, and transmitted optical power.
[0011] Secondly, this application provides an apparatus for automatic recovery from port anomalies, comprising: The acquisition module is used to acquire the link status change signal of the port; The recovery module is configured to perform a first recovery when a link status change signal is detected to change from connected to disconnected, the first recovery including disabling and then re-enabling the port; after performing the first recovery, wait for a first duration and obtain the first recovery status of the port; if the first recovery status is not recovered, perform a second recovery, the second recovery including simulating the optical module being offline and then simulating the optical module being online; after performing the second recovery, wait for a second duration and obtain the second recovery status of the port; the first duration is less than the second duration; if the second recovery status is not recovered, perform a third recovery, the third recovery including restarting the device.
[0012] Thirdly, this application provides a computing device, including a memory and a processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the computing device performs the method as described in any one of the first aspects.
[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program for performing the method as described in any one of the first aspects.
[0014] As can be seen from the above technical solution, this application has at least the following beneficial effects: First, through a tiered recovery process—first recovery (disabling and re-enabling the port), second recovery (software simulating offline and online optical modules), and third recovery (rebooting the device)—non-hardware damage faults such as physical link fluctuations and protocol negotiation anomalies can be automatically repaired without on-site human intervention. In particular, it avoids the cost of human intervention at remote unattended sites, reducing fault recovery time from hours to days in the traditional manual mode to minutes, and significantly reducing business interruption time.
[0015] Secondly, by pre-acquiring the optical module's on-site status, optical port LOS alarms, and DDM parameters (temperature, voltage, power, etc.), the recovery process is triggered only when the recovery conditions are met, avoiding invalid operations for physical link failures (such as module missing or signal loss). At the same time, by setting the first and second durations differently, the effective period of different recovery steps is adapted to reduce the risk of process oscillation.
[0016] Finally, by limiting the restart interval, the network fluctuations caused by frequent restarts of core devices are avoided while ensuring the rapid recovery of non-core devices, thus taking into account the business continuity requirements in different scenarios. The full-process status monitoring and tiered recovery design can more effectively deal with occasional and recurring failures, break the vicious cycle of failure, repair, and re-failure in traditional manual operation and maintenance, and improve the overall availability and service quality of the network.
[0017] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0018] Figure 1 A flowchart illustrating a method for automatic port anomaly recovery provided in this application embodiment; Figure 2 A schematic diagram of a device for automatic port fault recovery provided in an embodiment of this application; Figure 3 This is a schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation
[0019] The terms "first," "second," and "third," etc., used in this application specification and accompanying drawings are used to distinguish different objects, not to limit a specific order.
[0020] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0021] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the related technologies is given first: A switch port is a physical interface in a switch used to connect other network devices (such as servers, terminals, and other switches). It is the node through which data enters and exits the switch, and its status directly affects the connectivity of the network link.
[0022] In network communication systems, the stable operation of switch ports is crucial for ensuring continuous network transmission. However, in practical applications, ports often experience abnormal disconnections due to non-physical link failures and are unable to recover autonomously. Specifically, ports may become disconnected due to software or temporary hardware issues such as momentary protocol negotiation failures or chip malfunctions, and cannot recover through their own mechanisms, requiring manual intervention. This ultimately leads to inefficient fault handling and prolonged service interruptions. Among these issues, momentary protocol negotiation failures often stem from mismatches in speed and duplex mode with the connected device, while chip malfunctions are frequently caused by static electricity or electromagnetic interference resulting in brief periods of unresponsiveness.
[0023] The root cause of this problem lies in the limitations of traditional operation and maintenance models and their mismatch with scenario requirements: On the one hand, existing solutions rely on manual on-site handling, which incurs extremely high transportation and time costs for remote unattended sites, and the fault recovery cycle often takes several hours to several days, seriously affecting business continuity; on the other hand, traditional models lack accurate fault screening and tiered recovery strategies, making it easy to perform ineffective operations on physical link failures, and may also cause network instability due to excessive operations. The former involves physical link failures such as missing optical modules and optical port optical signal loss alarms, while the latter often involves frequent restarts of core equipment. These problems together make it difficult for traditional models to adapt to the high availability and low cost operation and maintenance requirements of enterprise-level networks, data centers and other scenarios.
[0024] In view of this, embodiments of this application provide a method for automatic recovery of port anomalies, which can be executed by a switch device.
[0025] This method addresses the issues of low efficiency, high cost, and susceptibility to ineffective operations caused by manual maintenance when existing switch ports are abnormally disconnected due to non-physical link failures (such as protocol negotiation anomalies or momentary chip failures). This application addresses these problems by real-time monitoring of port link status changes, triggering a recovery process only when a link changes from connected to disconnected. It pre-screens effective recovery scenarios using optical module availability status, optical signal loss alarms, and digital diagnostic monitoring parameters, excluding physical link failures to avoid ineffective operations. A tiered recovery strategy is employed, involving disabling and re-enabling the port, simulating optical module offline and then back online, and restarting the device. This strategy progressively processes recovery based on fault severity, while differentiating waiting times to adapt to the effective cycles of different recovery steps and limiting restart intervals to avoid excessive operations. Ultimately, this achieves automated and accurate port anomaly recovery, reducing maintenance costs and improving network availability.
[0026] To make the technical solution of this application clearer and easier to understand, the following describes a method for automatic port anomaly recovery provided by an embodiment of this application, in conjunction with the accompanying drawings. Figure 1As shown, this figure is a flowchart of a method for automatic port anomaly recovery provided in an embodiment of this application. The method includes: S201. The switch device acquires the link status change signal of the port.
[0027] A switch is a device used for network communication. It connects different network nodes (such as servers, terminals, other switches, etc.) through multiple ports to receive, forward, and manage data links. It is an infrastructure that ensures network connectivity.
[0028] A port is an interface on a switch used to physically connect to other network devices, such as an electrical port or an optical port. It is the channel through which data enters and exits the switch, and its status directly determines the connectivity of the corresponding link.
[0029] Link status refers to whether the communication link between the port and the connected device is unobstructed, and is divided into two states: connected (the link is normal and data can be transmitted) and disconnected (the link is interrupted and data cannot be transmitted).
[0030] Link status change signals are signals used to characterize the port link status from connected to disconnected or from disconnected to connected, and are the basis for switches to sense the dynamics of link status.
[0031] Switch devices use built-in monitoring mechanisms to detect the link status of all their ports in real time, continuously capturing information on the switching between connected and disconnected port link statuses. For example, when a port changes from normal data transmission (connected) to being unable to transmit data (disconnected) due to a protocol negotiation anomaly, the switch can recognize this status change and generate a corresponding link status change signal; conversely, when a disconnected link is restored to connectivity, the switch will also capture the corresponding change signal.
[0032] This process is a prerequisite for subsequent determination of whether the port failure recovery process needs to be triggered. By acquiring this signal in real time, the switch can accurately detect whether a port failure has occurred, providing the initial triggering conditions for automated recovery.
[0033] S202. When a link status change signal is detected to change from connected to disconnected, the switch device performs a first recovery, which includes disabling the port and then re-enabling the port; after performing the first recovery, it waits for a first duration and obtains the first recovery status of the port.
[0034] The first recovery is the first operational step in the tiered recovery strategy of this application. It is designed for the basic abnormal scenario of the link changing from connected to disconnected. It attempts to quickly repair the port failure through lightweight configuration adjustments and is a prerequisite for subsequent deeper recovery operations.
[0035] Disabling a port is a software instruction used by a switch to suspend all communication functions of a target port, temporarily stopping the port from receiving and sending data. This is equivalent to cutting off the logical connection between the port and external devices, thereby clearing any abnormal configurations or temporary blocking states that may exist on the port.
[0036] Re-enabling a port is when a port is disabled and then reactivated by the switch via software commands. This restores the port's communication function and allows it to re-negotiate protocols with the connected device (such as speed and duplex mode matching) to attempt to establish a normal link.
[0037] The first duration is the preset time period for waiting for port status feedback after the first recovery is performed. Its duration setting needs to balance fault recovery efficiency and avoid port oscillation. It needs to be sufficient for the port to complete protocol negotiation and present the recovery result, while avoiding waiting too long and prolonging the service interruption time. In this application, it is set to 1 minute by default.
[0038] The first recovery state is the link status result presented by the port after the first recovery is performed and the first duration is waited. It includes only two situations: recovered (the port becomes connected again and can transmit data normally) and unrecovered (the port is still disconnected and cannot transmit data). It is the basis for determining whether the second recovery needs to be triggered.
[0039] The switch only initiates the first recovery process when it detects that the link state of a port has switched from connected to disconnected. This design effectively avoids invalid operations on normal links or scenarios where the first recovery is not required (such as a normal recovery scenario where the link state changes from disconnected to connected), ensuring that the recovery process only applies to ports that are truly abnormal, thus improving operational accuracy.
[0040] In the first recovery process, the switch first disables the target port via software commands. This operation clears minor faults that may exist on the port, such as protocol negotiation anomalies or temporary signal blockages. Subsequently, the switch re-enables the port via software commands, allowing the port to re-negotiate communication parameters such as speed and duplex mode with the connected external device. The entire process is equivalent to performing a logical restart on the port. Fault repair can be attempted solely through software-level configuration adjustments, without involving physical operations such as hardware plugging and unplugging or module replacement, thus minimizing the operational cost and complexity of fault handling.
[0041] After completing all the operations of the first recovery, the switch does not immediately proceed to the subsequent recovery phase. Instead, it sets a one-minute waiting period. This waiting period allows sufficient time for the port to complete protocol negotiation with external devices and for the link status to stabilize. After the waiting period, the switch will obtain the port's first recovery status in real time. If the first recovery status is "recovered," it indicates that the port has reconnected and can transmit data normally, the first recovery process is successful, and no further operations are needed.
[0042] If the first recovery status is not recovered, it means that the current port failure cannot be resolved by lightweight software configuration adjustments. At this time, the switch device will trigger a deeper second recovery process. The progressive recovery strategy ensures the effectiveness of fault handling and avoids the failure from being unable to be resolved for a long time due to the inadequacy of a single recovery method.
[0043] Before performing the first recovery, the method requires the following operations: First, the switch device acquires the optical module status signal of the optical port. The optical module status signal includes a first status signal, a second status signal, and a third status signal.
[0044] An optical port is a physical interface on a switch device used to connect optical modules and transmit data through optical fibers. It must be used with an optical module, and its status is directly related to the operation of the optical module.
[0045] Optical module status signals are data signals sent by the optical module to the switch device that reflect its own operating status. They include information such as whether the optical module exists, whether there is signal loss, and whether the parameters are normal. They are the basis for judging the cause of optical port abnormalities.
[0046] The first status signal is a sub-item of the optical module status signal, used to characterize whether the optical module is physically inserted into the optical port. There are only two states: optical module in place (module has been correctly inserted) and optical module out of place (module is not inserted or is loose).
[0047] The second status signal is a sub-item of the optical module status signal, used to provide feedback on whether there is an optical signal loss problem at the optical port, whether there is an optical signal loss (LOS) alarm at the optical port or an LOS alarm indicating that the optical port does not exist.
[0048] The third status signal is a sub-item of the optical module status signal, containing parameters of the optical module's operation. These are Digital Diagnostic Monitoring (DDM) parameters, used to further determine whether the optical module is within its normal operating range from the perspective of hardware operating status. DDM parameters are a set of operating indicators collected in real time by the optical module through its built-in sensors. DDM parameters include optical module temperature, operating voltage, bias current, received optical power, and transmitted optical power.
[0049] When the first status signal indicates that the optical module is in place and the second status signal indicates that the optical port is not present and there is a loss of optical signal (LOS) alarm, the first recovery is performed.
[0050] The switch device first actively reads the status signals sent by the optical modules connected to the optical ports. These signals are divided into three categories: first status signal, second status signal, and third status signal. The first two categories are used to determine the basic connection and signal reception status of the optical modules, while the third category is used to determine the hardware health status of the optical modules.
[0051] The switching device does not immediately perform the first recovery after receiving the signal. Instead, it first makes a preliminary judgment based on the first and second status signals, and only proceeds to the next stage if both conditions are met simultaneously: The first status signal is that the optical module is in place: This confirms that the optical module has been correctly inserted into the optical port, eliminating physical connection problems such as the module not being inserted or being loose. The second status signal is an optical port LOS alarm: confirming that the optical port can receive optical signals normally, ruling out physical link problems such as fiber breakage or failure of the peer equipment.
[0052] These two conditions serve to first rule out scenarios that cannot be addressed by the first recovery method, such as hardware disconnection or link breakage, and to initially narrow down the scope of non-physical link failures.
[0053] After passing the basic screening criteria, the switch also needs to verify its hardware health using the DDM parameters in the third state signal: The switch device determines whether the recovery conditions are met based on the Digital Diagnostic Monitoring (DDM) parameters in the third status signal; if the recovery conditions are met, the first recovery is performed.
[0054] The switch will compare the real-time collected optical module temperature, operating voltage, bias current, received optical power, and transmitted optical power with the preset normal threshold range one by one. If all five parameters are within the normal threshold (meeting the recovery conditions), it means that the optical module hardware itself is not faulty. The port disconnection is most likely a minor problem such as protocol negotiation abnormality or temporary signal blockage. At this time, performing the first recovery can effectively repair it. If any parameter exceeds the threshold (does not meet the recovery conditions), it indicates that the port abnormality is caused by a hardware failure of the optical module (such as module overheating or insufficient power). If the first recovery cannot solve the problem, the process should be stopped and a manual hardware check should be prompted to avoid wasting system resources on ineffective operations.
[0055] The entire process forms a complete judgment chain of basic screening, hardware verification, and accurate execution, which ensures that the first recovery only applies to repairable scenarios and minimizes invalid operations on physical failure scenarios, thereby improving the accuracy and efficiency of fault handling.
[0056] S203. If the first recovery status is not recovered, the switch device performs a second recovery, which includes simulating the optical module being offline and then simulating the optical module being online. After performing the second recovery, it waits for a second duration and obtains the second recovery status of the port. The first duration is less than the second duration.
[0057] The second recovery is the second operational step of the tiered recovery strategy in this application. It is triggered only when the first recovery fails. It is designed for deeper port anomalies (such as fixed errors in port configuration parameters) that cannot be resolved by the first recovery. It achieves configuration reset by simulating changes in hardware state and is a step that connects the first recovery and the third recovery.
[0058] Software simulation of optical module offline / online means that the switch does not actually plug or unplug the physical optical module, but instead sends a signal to the system that the optical module has been unplugged (offline) or inserted (online) through software commands. This tricks the system into triggering the configuration loading process related to the insertion and removal of the optical module. It is a hardware status simulation operation at the pure software level.
[0059] The second duration is a preset time period for waiting for port status feedback after the second recovery is performed. Since the second recovery involves more steps such as simulating offline-online, reloading configuration, and secondary protocol negotiation, the required effective time is longer than that of the first recovery. Therefore, it is set to 5 minutes to meet the requirement that the first duration is less than the second duration.
[0060] The second recovery status is the link status result presented by the port after the second recovery is executed and the second duration is waited. It is also divided into two types: recovered and not recovered. It is the basis for determining whether the third recovery needs to be triggered.
[0061] The second recovery of a switch device has clear prerequisites: it must first be confirmed that the first recovery status is not recovered, that is, the port is still disconnected after the disable-re-enable operation. This setting strictly follows the step-by-step recovery logic from light to heavy, which can effectively avoid directly executing a more complex second recovery when the light operation has not been attempted or has successfully resolved the problem, thereby reducing unnecessary consumption of system resources.
[0062] In the second recovery process, the switch first enters the simulated optical module offline stage. The switch sends a signal to its own system that the optical module has been unplugged. After receiving the signal, the system will automatically clear the configuration parameters that were originally saved on the port and matched the optical module, such as the port rate parameters set according to the optical module rate. Then, the switch enters the simulated optical module online stage. The switch sends a signal that the optical module has been inserted. The system will then reread the hardware capabilities of the optical module, such as the supported rate range and power threshold, and issue a new adaptation configuration to the port based on this data. This is to repair the port anomalies that could not be resolved by the first recovery due to configuration parameter fixing errors (such as the original configuration not matching the actual capabilities of the optical module).
[0063] After performing the above operations, the switch will not immediately determine the result, but will wait for a second period of 5 minutes. This is because the second recovery involves multiple complex processes such as clearing the configuration, reloading, and secondary protocol negotiation, which require sufficient time for the port and the connected device to complete parameter matching and establish a stable link. After the second period ends, the switch will obtain the second recovery status of the port to provide a basis for determining whether to start the third recovery.
[0064] The reason for specifying that the first recovery time is shorter than the second recovery time is based on the difference in complexity between the two recovery operations: the first recovery is just a simple configuration adjustment of disabling and enabling, and one minute is enough for the port to complete negotiation and stabilize. The second recovery process is more complex and takes longer. The five-minute setting can avoid misjudging failure to recover due to too short a wait, and also avoid unnecessarily prolonging the business interruption time due to too long a wait, ultimately achieving a balance between efficiency and accuracy.
[0065] S204. If the second recovery status is not recovered, the switch device performs a third recovery, which includes restarting the device.
[0066] The third recovery is the final operation step of the tiered recovery strategy of this application. It is triggered only when the first recovery and the second recovery fail. It is designed for deep port anomalies that cannot be resolved by the first two steps. It resets the system state by restarting the device and is the last resort to ensure port recovery.
[0067] Restarting the device enables the switch to complete the entire process of system shutdown, core component initialization, configuration reloading, and port reactivation through software commands or preset mechanisms. It is not a simple power disconnection and reconnection, but an orderly restart that takes into account system stability. It can clear temporary errors accumulated during device operation and resolve cross-module configuration conflicts.
[0068] This final recovery step is only initiated when the switch detects that the second recovery status is not recovered—that is, the port remains disconnected after a software-simulated offline and then online configuration reset of the optical module. This setup continues the recovery logic of first light-weight, then medium-weight, and finally fallback. The first two steps attempt to repair the fault through light-weight configuration adjustment (first recovery) and medium-weight configuration reset (second recovery) of a single port, respectively. Only when the first two steps fail to resolve the issue will a device restart operation, which affects the overall operation of the device, be used to minimize interference with normal network links.
[0069] From an operational perspective, the third recovery process of restarting the device is not simply a matter of power on / off. Rather, the switch completes a systematic process of shutting down the system, initializing core components, reloading basic configurations, and reactivating all ports. The purpose is to clear system-level errors accumulated during device operation, such as temporary core module freezes, cross-port configuration conflicts, and abnormal cached data. These deep-seated faults cannot be addressed by the first two steps, which only target individual ports. Only by restarting and resetting the entire device's operating environment can the problem be solved at its root, allowing the ports to regain normal communication capabilities.
[0070] Meanwhile, in conjunction with the overall solution design, although the third recovery is a fallback measure, it will be paired with a restart time interval limit to avoid network instability caused by frequent restarts of core equipment.
[0071] After the third recovery restart is performed, the switch resumes monitoring of the ports. A restart interval is set. If the port becomes abnormal again within the restart interval, the third recovery is prohibited from being performed again, and the first and second recoveries are performed instead.
[0072] Restoring port monitoring is a fundamental operation after a reboot. During the reboot process, the ports and monitoring modules of the switch device are briefly in an initialization state and cannot detect changes in the link. After the device reboots, the monitoring mechanism needs to be reactivated to track the port link status in real time. This ensures that if a port subsequently changes from connected to disconnected, it can be detected in a timely manner, avoiding missed fault diagnosis due to lack of monitoring.
[0073] Secondly, setting a restart interval is to avoid network interference caused by frequent restarts. While restarting the device can resolve deep-seated faults, it causes the device to go offline briefly and ports to temporarily disconnect. If this is done multiple times in a short period, it may cause network instability. This application sets a default restart interval of 24 hours, specifying the minimum time difference between two third-party recovery attempts, in order to limit the restart frequency.
[0074] Finally, if the port becomes abnormal again during the restart interval, it indicates that the fault may be a minor issue that can be fixed by the first two steps, rather than a system-level fault that requires another restart. In this case, the third recovery should be disabled, and only the first and second recoveries should be triggered. This not only continues the principle of operating from minor to major issues, but also avoids further impacting network stability due to unnecessary restarts.
[0075] Based on the above description, this application has the following beneficial effects: First, through a tiered recovery process—first recovery (disabling and re-enabling the port), second recovery (software simulating offline and online optical modules), and third recovery (rebooting the device)—non-hardware damage faults such as physical link fluctuations and protocol negotiation anomalies can be automatically repaired without on-site human intervention. In particular, it avoids the cost of human intervention at remote unattended sites, reducing fault recovery time from hours to days in the traditional manual mode to minutes, and significantly reducing business interruption time.
[0076] Secondly, by pre-acquiring the optical module's on-site status, optical port LOS alarms, and DDM parameters (temperature, voltage, power, etc.), the recovery process is triggered only when the recovery conditions are met, avoiding invalid operations for physical link failures (such as module missing or signal loss). At the same time, by setting the first and second durations differently, the effective period of different recovery steps is adapted to reduce the risk of process oscillation.
[0077] Finally, by limiting the restart interval, the network fluctuations caused by frequent restarts of core devices are avoided while ensuring the rapid recovery of non-core devices, thus taking into account the business continuity requirements in different scenarios. The full-process status monitoring and tiered recovery design can more effectively deal with occasional and recurring failures, break the vicious cycle of failure, repair, and re-failure in traditional manual operation and maintenance, and improve the overall availability and service quality of the network.
[0078] The above text combined Figure 1 This application provides a detailed description of a method for automatic recovery of port anomalies. The apparatus and devices provided in this application will be described below with reference to the accompanying drawings.
[0079] like Figure 2 As shown in the figure, this is a schematic diagram of a device for automatic port anomaly recovery provided in an embodiment of this application. The device includes: Acquisition module 301 is used to acquire the link status change signal of the port; Recovery module 302 is configured to perform a first recovery when the link status change signal is detected to change from connected to disconnected, the first recovery including disabling the port and then re-enabling the port; after performing the first recovery, wait for a first duration and obtain the first recovery status of the port; if the first recovery status is not recovered, perform a second recovery, the second recovery including simulating the optical module being offline and then simulating the optical module being online; after performing the second recovery, wait for a second duration and obtain the second recovery status of the port; the first duration is less than the second duration; if the second recovery status is not recovered, perform a third recovery, the third recovery including restarting the device.
[0080] Optionally, the acquisition module 301 is specifically used to acquire the optical module status signal of the optical port. The optical module status signal includes a first status signal, a second status signal, and a third status signal. When the first status signal indicates that the optical module is in place and the second status signal indicates that the optical port does not exist and there is a loss of optical signal (LOS) alarm, the first recovery is performed.
[0081] Optionally, the recovery module 302 is further configured to determine whether the recovery conditions are met based on the digital diagnostic monitoring (DDM) parameters in the third state signal; If the recovery conditions are met, then perform the first recovery.
[0082] Optionally, recovery module 302 is also used to restore monitoring of the port; Set a restart time interval. If the port fails again within the restart time interval, the third recovery will be prevented from being executed again, and the first recovery and the second recovery will be executed instead.
[0083] The apparatus for automatic port recovery according to the embodiments of this application can correspond to the execution of the method described in the embodiments of this application, and the other operations and / or functions of each module / unit of the apparatus for automatic port recovery are respectively for implementing Figure 1 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.
[0084] This application also provides a computing device. For example... Figure 3 As shown in the figure, this is a schematic diagram of a computing device provided in an embodiment of this application. The computing device 700 includes a bus 701, a processor 702, a communication interface 703, and a memory 704. The processor 702, the memory 704, and the communication interface 703 communicate with each other via the bus 701.
[0085] The 701 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0086] The processor 702 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0087] The communication interface 703 is used for external communication.
[0088] Memory 704 may include volatile memory, such as random access memory (RAM). Memory 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0089] The memory 704 stores executable code, and the processor 702 executes the executable code to perform the aforementioned method for automatic recovery of port abnormalities.
[0090] Specifically, in achieving Figure 2 In the case of the illustrated embodiment, and Figure 2 When the modules or units of the port anomaly automatic recovery device described in the embodiment are implemented by software, the following steps are performed: Figure 2 The software or program code required for the functions of each module / unit can be partially or entirely stored in memory 704. Processor 702 executes the program code corresponding to each unit stored in memory 704 to perform the aforementioned automatic recovery method for port abnormalities.
[0091] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the aforementioned method for automatic recovery of port anomalies.
[0092] This application also provides a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.
[0093] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0094] When the computer program product is executed by a computer, the computer performs any of the aforementioned methods for automatic port failure recovery. The computer program product can be a software installation package; when any of the aforementioned methods for automatic port failure recovery is required, the computer program product can be downloaded and executed on the computer.
[0095] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the scope of protection of this application.
Claims
1. A method for automatic recovery of port anomalies, characterized by, The method comprises: acquiring a link state change signal of a port; when detecting that the link state change signal changes from being connected to being disconnected, performing a first recovery, the first recovery comprising disabling the port and then re-enabling the port; after performing the first recovery, waiting for a first time length and acquiring a first recovery state of the port; when the first recovery state is not recovered, performing a second recovery, the second recovery comprising simulating that an optical module is offline and then simulating that the optical module is online; after performing the second recovery, waiting for a second time length and acquiring a second recovery state of the port; the first time length is less than the second time length; when the second recovery state is not recovered, performing a third recovery, the third recovery comprising restarting the device.
2. The method of claim 1, wherein, The performing the first recovery comprises: acquiring an optical module state signal of an optical port, the optical module state signal comprising a first state signal, a second state signal and a third state signal; when the first state signal is that the optical module is in place and the second state signal is that there is no loss of signal (LOS) alarm of the optical port, performing the first recovery.
3. The method of claim 2, wherein, The method further comprises: judging whether a recovery condition is met according to a digital diagnostic monitoring (DDM) parameter in the third state signal; if the recovery condition is met, performing the first recovery.
4. The method of claim 1, wherein, After restarting the device in the third recovery, the method further comprises: resuming monitoring of the port; setting a restart time interval, and within the restart time interval, if the port abnormally occurs again, disabling the third recovery from being performed again, and performing the first recovery and the second recovery.
5. The method of claim 3, wherein, The DDM parameter comprises an optical module temperature, an operating voltage, a bias current, a received optical power and a transmitted optical power.
6. An apparatus for automatic recovery from port abnormality, the apparatus comprising: The apparatus comprises: an acquisition module configured to acquire a link state change signal of a port; a recovery module configured to, when detecting that the link state change signal changes from being connected to being disconnected, perform a first recovery, the first recovery comprising disabling the port and then re-enabling the port; after performing the first recovery, waiting for a first time length and acquiring a first recovery state of the port; when the first recovery state is not recovered, performing a second recovery, the second recovery comprising simulating that an optical module is offline and then simulating that the optical module is online; after performing the second recovery, waiting for a second time length and acquiring a second recovery state of the port; the first time length is less than the second time length; when the second recovery state is not recovered, performing a third recovery, the third recovery comprising restarting the device.
7. The apparatus of claim 6, wherein, The acquisition module is specifically configured to acquire an optical module state signal of an optical port, the optical module state signal comprising a first state signal, a second state signal and a third state signal; when the first state signal is that the optical module is in place and the second state signal is that there is no loss of signal (LOS) alarm of the optical port, performing the first recovery.
8. A computing device, comprising: comprise a memory and a processor; wherein one or more computer programs are stored in the memory, the one or more computer programs comprising instructions; when the instructions are executed by the processor, the computing device performs the method in any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store a computer program for performing the method according to any one of claims 1 to 5.
10. A computer program product, characterised in that, The computer program product comprises one or more computer instructions, which, when executed by a computer, perform the method according to any one of claims 1 to 5.
Citation Information
Cited By
Fault self-healing reconnection method based on SRIO communication between FPGA and DSP
CN122086670A