Fault detection method and device, nonvolatile storage medium and electronic equipment

By obtaining the main and backup link switching information and optical module operation status information, and using the status classification model to perform optical module fault detection, the problem of low fault detection efficiency in the existing technology is solved, timely early warning and precise positioning of optical module faults is achieved, and the accuracy and efficiency of fault detection are improved.

CN120474616APending Publication Date: 2025-08-12CHINA TELECOM CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510874652.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing optical module fault detection technology mainly relies on the operating status information of the optical module, resulting in low fault detection efficiency, inability to accurately identify the cause of the fault, and lacks network topology correlation analysis capabilities and preventive maintenance.

Method used

By obtaining the main and backup link switching information and the operating status information of the optical module, fault detection is performed using the status classification model, and combined with the second port status information of the optical fiber connection, the fault positioning and early warning are achieved.

Benefits of technology

It realizes timely early warning and precise positioning of optical module faults, improves fault detection efficiency, can perform predictive maintenance based on historical data, and reduces the dependence of manual empirical judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474616A_ABST
    Figure CN120474616A_ABST
Patent Text Reader

Abstract

The invention discloses a fault detection method and device, a nonvolatile storage medium and electronic equipment. The method comprises the steps that main and standby link switching information of a first port and optical module operation state information of the first port are acquired, and the main and standby link switching information is used for indicating whether main and standby link switching happens to the first port or not; determining whether a first optical module corresponding to the first port is abnormal or not according to the main and standby link switching information and the optical module operation state information of the first port; when it is determined that the first optical module is abnormal, optical module operation state information of a second port is obtained, and the second port is connected with the first port through an optical fiber; and determining a fault detection result according to the optical module operation state information of the first port and the optical module operation state information of the second port. The technical problem of low fault detection efficiency caused by the fact that the related fault detection technology mainly depends on the operation state information of the optical module is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communications, and in particular to a fault detection method, device, non-volatile storage medium, and electronic device. Background Art

[0002] The rapid development of optical communication networks has led to optical modules playing an increasingly important role in data transmission. As a core component of optical communication systems, optical modules are responsible for converting electrical signals into optical signals for transmission, and their performance directly impacts network efficiency and reliability. However, optical modules may experience various faults during operation, leading to network performance degradation or even interruption. Traditional optical module fault monitoring technology primarily relies on the module's built-in DDM information (i.e., optical module operating status information). It determines the module's health status by collecting real-time threshold alarms for parameters such as received optical power, transmitted optical power, operating temperature, and bias current. However, this method can only trigger an alarm based on threshold judgment after a fault occurs, and cannot determine the specific cause of the fault. Determining the cause of the fault still requires manual experience, resulting in low fault detection efficiency.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a fault detection method, apparatus, non-volatile storage medium, and electronic device to at least solve the technical problem of low fault detection efficiency caused by the fact that related fault detection technologies mainly rely on optical module operating status information.

[0005] According to one aspect of an embodiment of the present application, a fault detection method is provided, including: obtaining primary-backup link switching information of a first port and optical module operating status information of the first port, wherein the primary-backup link switching information is used to indicate whether primary-backup link switching occurs at the first port; determining whether there is an abnormality in the first optical module corresponding to the first port based on the primary-backup link switching information and the optical module operating status information of the first port; when it is determined that there is an abnormality in the first optical module, obtaining optical module operating status information of the second port, wherein the second port is connected to the first port via an optical fiber; determining a fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0006] Optionally, the optical module operating status information includes received optical power and transmitted optical power, and the first port fault includes a transmit fault and a receive fault; determining the fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port includes: when the transmitted optical power of the second port is less than a first threshold, determining the fault detection result as a second port fault; when the transmitted optical power of the second port is not less than the first threshold and the received optical power of the first port is less than the second threshold, determining the fault detection result as a receive fault; when the transmitted optical power of the first port is less than the first threshold, determining the fault detection result as a transmit fault.

[0007] Optionally, the type of the first port includes at least one of the following: a user access link port, an uplink port, and a downlink port; when it is determined that the fault detection result is a first port fault, the method further includes: determining whether there is a backup port for the first port; when there is no backup port for the first port, determining that the alarm level corresponding to the first port is a level one alarm; when there is a backup port for the first port, determining that the alarm level corresponding to the first port is a level two alarm; determining a fault handling suggestion based on the fault detection result, the type of the first port, and the alarm level, wherein the fault handling suggestion is used to instruct the target object to handle the fault; and sending alarm information to the target object, wherein the alarm information includes: the alarm level, the fault detection result, and the fault handling suggestion.

[0008] Optionally, obtaining the primary-backup link switching information of the first port and the optical module operating status information of the first port includes: collecting the primary-backup link switching information of the first port and the optical module operating status information of the first port according to a preset time period, wherein the optical module operating status information includes at least one of the following: operating temperature, operating voltage, bias current, received optical power and transmitted optical power; determining the optical module model of the optical module corresponding to the first port, and obtaining threshold information corresponding to the optical module model, wherein the threshold information includes a minimum value and a maximum value corresponding to the optical module operating status information; and normalizing the optical module operating status information of the first port based on the threshold information.

[0009] Optionally, based on the threshold information, normalizing the optical module operating status information of the first port includes: when the value of the optical module operating status information is less than the minimum value corresponding to the optical module operating status information, setting the value of the optical module operating status information to 0; when the value of the optical module operating status information is greater than the maximum value corresponding to the optical module operating status information, setting the value of the optical module operating status information to 1; when the value of the optical module operating status information is between the minimum value and the maximum value corresponding to the optical module operating status information, scaling the value of the optical module operating status information according to the minimum value and the maximum value.

[0010] Optionally, based on the primary-backup link switching information and the optical module operating status information of the first port, determining whether there is an abnormality in the first optical module corresponding to the first port includes: inputting the primary-backup link switching information and the optical module operating status information of the first port into a state classification model, the state classification model is used to classify the health status of the first optical module; obtaining a classification result output by the state classification model, wherein the classification result includes health, warning, and abnormality; when the classification result is abnormal, there is an abnormality in the first optical module corresponding to the first port.

[0011] Optionally, before inputting the primary-backup link switching information and the optical module operating status information of the first port into the state classification model, the method also includes: obtaining historical sampling data, wherein the historical sampling data includes the historical primary-backup link switching information of the port and the historical optical module operating status information; normalizing the historical sampling data; determining a multi-feature dictionary and a sparse coefficient based on the normalized historical sampling data; and constructing a state classification model based on the multi-feature dictionary and the sparse coefficient.

[0012] Optionally, after obtaining the classification result output by the status classification model, the method further includes: when the classification result is a warning, sending warning information to the target object, wherein the warning information is used to instruct the target object to check the first port.

[0013] According to another aspect of an embodiment of the present application, a fault detection device is also provided, including: a first processing module, used to obtain the primary-backup link switching information of the first port and the optical module operating status information of the first port, wherein the primary-backup link switching information is used to indicate whether the primary-backup link switching occurs at the first port; a second processing module, used to determine whether the first optical module corresponding to the first port has an abnormality based on the primary-backup link switching information and the optical module operating status information of the first port; a third processing module, used to obtain the optical module operating status information of the second port when it is determined that the first optical module has an abnormality, wherein the second port is connected to the first port via an optical fiber; a fourth processing module, used to determine the fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0014] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored. When the program is running, a device where the non-volatile storage medium is located is controlled to execute a fault detection method.

[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the fault detection method is executed when the program is run.

[0016] According to another aspect of an embodiment of the present application, a computer program product is further provided, including a computer program, which implements a fault detection method when executed by a processor.

[0017] In an embodiment of the present application, the master-slave link switching information of the first port and the optical module operating status information of the first port are obtained, wherein the master-slave link switching information is used to indicate whether the master-slave link switching occurs at the first port; based on the master-slave link switching information and the optical module operating status information of the first port, it is determined whether the first optical module corresponding to the first port has an abnormality; when it is determined that the first optical module has an abnormality, the optical module operating status information of the second port is obtained, wherein the second port is connected to the first port through an optical fiber; the fault detection result is determined based on the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port failure, second port failure, fault prediction is performed through the optical module DDM information and the master-slave link switching information, and after determining that a fault exists, the specific fault cause is judged and the corresponding alarm information is triggered, thereby achieving the purpose of accurately identifying the optical module status, thereby achieving the technical effect of timely fault warning and precise positioning of the fault cause, and thus solving the technical problem of low fault detection efficiency caused by the fact that the relevant fault detection technology mainly relies on the optical module operating status information. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 is a structural diagram of a computer terminal provided according to an embodiment of the present application;

[0020] Figure 2 This is a flowchart of a fault detection method provided according to an embodiment of the present application;

[0021] Figure 3 This is a schematic diagram of optical module transmitted optical power sampling data before normalization provided in an embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of normalized optical module transmitted optical power sampling data provided in an embodiment of the present application;

[0023] Figure 5 According to the embodiment of the present application, a method based on the norm l is provided. row,0 Schematic diagram of the classification results;

[0024] Figure 6 According to the embodiment of the present application, a method based on the norm l is provided. adaptive,0 Schematic diagram of the classification results;

[0025] Figure 7 This is a schematic diagram of a process for inferring a fault cause according to an embodiment of the present application;

[0026] Figure 8 This is a schematic diagram of an alarm process provided according to an embodiment of the present application;

[0027] Figure 9 This is a schematic diagram of fault alarm information provided according to an embodiment of the present application;

[0028] Figure 10 is a schematic diagram of a framework of a fault detection system provided according to an embodiment of the present application;

[0029] Figure 11 It is a structural diagram of a fault detection device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0033] Optical Module: An optical module is a device used in optical communication systems, responsible for converting electrical signals into optical signals (for transmission) or converting optical signals into electrical signals (for reception).

[0034] Sparse Representation: Sparse representation is a signal processing and data analysis technique that extracts important features by representing a signal as a linear combination of a small number of basis functions. This method has advantages in processing high-dimensional data and can improve the efficiency and accuracy of data analysis.

[0035] DDM (Digital Diagnostic Monitoring): DDM is a monitoring technology used in optical modules and fiber-optic communication systems. The SFF.8472 protocol specifies five digital diagnostic monitoring parameters for evaluating the operating status of optical modules: operating temperature (Temp), operating voltage (Vcc), bias current (Tx_Bias), received optical power (Rx_Power), and transmitted optical power (Tx_Power).

[0036] In related technologies, fault detection for optical modules primarily relies on the module's built-in DDM information. This involves real-time collection of threshold alarms for parameters such as received optical power, transmitted optical power, operating temperature, and bias current to determine the module's health status. The root cause of the fault and the scope of its impact are inferred based on manual experience and local databases, with no intelligent solution available.

[0037] However, this approach has significant drawbacks:

[0038] First, isolated monitoring can lead to misdiagnosis of faults: Abnormal performance of optical modules may be caused by a variety of factors, such as fiber link loss, optical module failure of the other device, or abnormality of the local port. Relying solely on DMM information cannot determine the specific cause of the fault. Currently, the judgment of the fault cause still relies on manual experience.

[0039] Second, there is a lack of network topology correlation analysis capabilities. Existing technologies only monitor the status of individual optical modules, ignoring their logical location and connection relationships in the network, and are unable to automatically correlate information between upstream and downstream devices.

[0040] Third, passive response. Related technologies can only trigger alarms based on threshold judgments after a fault occurs. They lack the ability to predict performance based on historical data, making preventive and proactive maintenance impossible.

[0041] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0042] According to an embodiment of the present application, a method embodiment of a fault detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0043] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal for implementing a fault detection method. Figure 1 As shown, the computer terminal 10 may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0044] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0045] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fault detection method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned fault detection method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0046] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0047] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0048] In the above operating environment, the embodiment of the present application provides a fault detection method, such as Figure 2 As shown, the method includes the following steps:

[0049] Step S202 : acquiring primary / backup link switching information of the first port and optical module operation status information of the first port, wherein the primary / backup link switching information is used to indicate whether primary / backup link switching occurs on the first port.

[0050] Optionally, when the primary and backup links are switched, the load rate of the optical module may change suddenly, thereby accelerating the aging of the device. Therefore, the primary and backup switching information is introduced as an auxiliary feature for determining the health status. Taking device A as an example, if the port configuration contains "lacp priority", the port is the primary link, and its primary and backup switching characteristic value is 1, otherwise it is 0. The primary and backup link switching information includes multiple primary and backup switching characteristic values (0 or 1) arranged in the order of the sampling time points. Through the multiple primary and backup switching characteristic values arranged in time series, the primary and backup link switching information can indicate whether the primary and backup link switching occurs on the first port. For example, when the primary and backup link switching information is {0,0,0,1,1}, it indicates that the primary and backup link switching occurs on the first port at the sampling time point of the fourth primary and backup switching characteristic value.

[0051] In the technical solution provided in step S202, obtaining the primary-backup link switching information of the first port and the optical module operating status information of the first port includes: collecting the primary-backup link switching information of the first port and the optical module operating status information of the first port according to a preset time period, wherein the optical module operating status information includes at least one of the following: operating temperature, operating voltage, bias current, received optical power and transmitted optical power; determining the optical module model of the optical module corresponding to the first port, and obtaining threshold information corresponding to the optical module model, wherein the threshold information includes a minimum value and a maximum value corresponding to the optical module operating status information; and normalizing the optical module operating status information of the first port according to the threshold information.

[0052] Optionally, configure collection tasks through the cloud network acquisition and control platform to manage and collect data from network element devices, including network topology information and DDM data. Configuring collection tasks includes managing network devices such as MSE and DCSW, adding network element MIB databases, configuring collection instructions, and configuring scheduled collection tasks. By adding the SNMP agent's IP address (i.e., the IP address of the cloud network acquisition and control platform server) to the network element device, the network element device can upload information to the next-generation cloud network acquisition and control platform. By adding the SNMP agent's read password to the device, the cloud network acquisition and control platform can read information from the network element device.

[0053] Optionally, collecting the primary and backup link switching information of the first port and the operating status information of the optical module of the first port according to a preset time period includes: using the SNMP GET instruction through the cloud network acquisition and control platform to actively initiate data collection to the network element device every hour. Specifically including module information, port configuration information (extracting dT (indicating that the port is connected to the downlink link), uT (indicating that the port is connected to the uplink) described by the descrition keyword in the configuration), DDM information (operating temperature, operating voltage, bias current, received optical power and transmitted optical power), port status information (up or down), and LLDP neighbor information. The collection instruction can parse the data packet returned by the network element device through the MIB database to make the original data readable. Scheduled collection tasks can be configured on the cloud network acquisition and control platform, and can be set to execute a collection task once an hour. After each collection task is executed, the platform will encapsulate the collected data into an API for subsequent system calls. The system can call the API to read data and store historical data.

[0054] As an optional implementation, normalizing the optical module operating status information of the first port based on the threshold information includes: when the value of the optical module operating status information is less than the minimum value corresponding to the optical module operating status information, setting the value of the optical module operating status information to 0; when the value of the optical module operating status information is greater than the maximum value corresponding to the optical module operating status information, setting the value of the optical module operating status information to 1; when the value of the optical module operating status information is between the minimum value and the maximum value corresponding to the optical module operating status information, scaling the value of the optical module operating status information according to the minimum value and the maximum value.

[0055] Optionally, the data used to determine the health status of the optical module mainly includes DDM information (operation status information) and primary-backup link switching information. Since the thresholds of optical modules of different models are different, it is also necessary to collect the manufacturer information, model information and the corresponding DDM threshold (i.e., threshold information) of the optical module. For example, for each optical module model, DDM information is collected, including operating temperature, operating voltage, bias current, received optical power and transmitted optical power. Since the thresholds of each type and model of optical modules are different, there are certain difficulties in setting the alarm threshold. To solve this problem, this method will normalize the DMM information according to the thresholds of different models (i.e., the threshold information corresponding to the optical module model). For example, for the luminous power Tx of a certain model of optical module, its normal range is [Tx2-Tx1] (where Tx2 is the minimum value corresponding to the operation status information of the optical module, and Tx1 is the maximum value corresponding to the operation status information of the optical module), and its normalized segmentation formula is set as:

[0056]

[0057] That is, if the luminous power is less than Tx1 (that is, when the value of the optical module operation status information is less than the minimum value corresponding to the optical module operation status information), the value of the optical module operation status information is normalized to 0 to indicate an abnormality; if the luminous power is greater than Tx2 (that is, when the value of the optical module operation status information is greater than the maximum value corresponding to the optical module operation status information), the value of the optical module operation status information is normalized to 1 to indicate an abnormality; if the luminous power is within the normal range, linear scaling (1) is performed (that is, the value of the optical module operation status information is scaled according to the minimum value and the maximum value).

[0058] Step S204 : determining whether the first optical module corresponding to the first port is abnormal based on the active / standby link switching information and the operating status information of the optical module of the first port.

[0059] In the technical solution provided in step S204, determining whether there is an abnormality in the first optical module corresponding to the first port based on the primary-backup link switching information and the optical module operating status information of the first port includes: inputting the primary-backup link switching information and the optical module operating status information of the first port into a state classification model, the state classification model is used to classify the health status of the first optical module; obtaining a classification result output by the state classification model, wherein the classification result includes health, warning, and abnormality; when the classification result is abnormal, there is an abnormality in the first optical module corresponding to the first port.

[0060] Optionally, collecting primary / backup link switching information and optical module operating status information for the first port at a preset time period provides records of various performance indicators at different time points, capable of reflecting dynamic changes in the optical module during operation. By inputting the primary / backup link switching information and the optical module operating status information for the first port into a state classification model, a classification result for the health status of the first optical module can be obtained. If the classification result is abnormal, it is determined that the first optical module corresponding to the first port is abnormal, and subsequent root cause determination is required.

[0061] As an optional implementation, before inputting the primary-backup link switching information and the optical module operating status information of the first port into the state classification model, the method also includes: obtaining historical sampling data, wherein the historical sampling data includes the historical primary-backup link switching information of the port and the historical optical module operating status information; normalizing the historical sampling data; determining a multi-feature dictionary and a sparse coefficient based on the normalized historical sampling data; and constructing a state classification model based on the multi-feature dictionary and the sparse coefficient.

[0062] Optionally, the sparse representation classification (SRC) model assumes that the time series features of optical modules in the same state (normal, warning, abnormal) are located in the same low-rank dimensional subspace. By determining a multi-feature dictionary and a sparse coefficient, an optical module health status classification system (i.e., a state classification model) based on multi-factor sparse representation can be constructed.

[0063] Optionally, after acquiring historical sampled data, based on the optical module status (whether the optical module is faulty) corresponding to the historical sampled data, the status of the optical module one week prior to the failure is set to a warning state. For each optical module, historical DDM data and primary / backup link switching information are acquired every hour and are labeled as normal, warning, or abnormal.

[0064] Optionally, after the historical sampling data is annotated, normalization processing is performed, and the normalized segmentation formula is:

[0065]

[0066] That is, if the luminous power is less than Tx1 (that is, when the value of the optical module operation status information is less than the minimum value corresponding to the optical module operation status information), the value of the optical module operation status information is normalized to 0 to indicate abnormality; if the luminous power is greater than Tx2 (that is, when the value of the optical module operation status information is greater than the maximum value corresponding to the optical module operation status information), the value of the optical module operation status information is normalized to 1 to indicate abnormality; if the luminous power is within the normal range, linear scaling (1) is performed (that is, the value of the optical module operation status information is scaled according to the minimum value and the maximum value). Figure 3 shows the sampling data of the luminous power (transmitted optical power) of an optical module before normalization. Figure 4 shows a sample data of normalized light module luminous power, such as Figure 3 and Figure 4 As shown, the sampled data below the minimum value of the threshold information is normalized to 0.

[0067] Optionally, because the data for abnormal and warning states is relatively limited, upsampling is used to expand the dataset. This results in three types of training samples (normal, warning, and abnormal). Each type of sample contains six multi-feature information about the optical module (operating temperature, operating voltage, bias current, received optical power, transmitted optical power, and active / standby switching information).

[0068] The test sample Represented by 6 different performance data in T is the historical time sampling point. Then the corresponding multi-feature dictionary and sparse coefficient can be expressed as and Where C = 1, 2, and 3 represent normal, warning, and alarm states respectively, and k is the kth feature. is the historical value of the first DDM characteristic of the optical module in the recent T period. The c-th state can be obtained through the c-th subset of the dictionary Linear combination representation:

[0069]

[0070] because The status category is unknown, and a combination of C-type sub-dictionaries can also be used express:

[0071]

[0072] The multi-feature sparse coefficient (matrix) can be converted into the following model:

[0073]

[0074] That is, under certain constraints (||α MF || row,0 <=J), the actual value of the characteristic data (k is from 1 to 6) With the feature dictionary matrix D t and the sparse coefficient matrix α t The sum of the errors (measured by the Frobenius norm) between the combined estimates is minimized. Here argmin means finding the α that minimizes the objective function. MF . Where J is the non-zero number of sparse vectors.

[0075] Taking into account the different effects and weights of the performance characteristics of different optical modules on the prediction results of optical modules, we can also consider the adaptability of the characteristics to the prediction results. By introducing the adaptive sparse norm l adaptive,0 Determine the sparsity coefficient:

[0076]

[0077] The adaptive sparse norm allows non-zero values to be distributed in different rows in the sparse coefficients (matrix), and can adaptively select important features, so that the model can perform different degrees of sparse processing on different features according to their actual importance, avoiding the irrationality that may be caused by using a unified sparse standard for all features. Figure 5 shows a method based on the norm l row,0 Schematic diagram of the classification results, Figure 6shows a method based on the norm l adaptive,0 The schematic diagram of the classification results is as follows: Figure 5 and Figure 6 As shown in the figure, the sparse coefficients of different shapes represent different characteristic atoms. It can be seen from the figure that based on l adaptive,0 Important features can be adaptively selected to predict the health status of optical modules ( Figure 5 The sparse coefficients of each row in are relatively evenly distributed, and each feature has a certain number of non-zero coefficients (i.e., colored feature atoms). Figure 6 The sparse coefficients in different rows have different distributions of the number of non-zero elements. Figure 6 It shows that some rows have more non-zero coefficients, while other rows are relatively sparse. This is because the adaptive sparse norm dynamically adjusts the weight of the feature's influence on the classification result. Important features are given greater weight, allowing them to have more non-zero elements in the sparse coefficient matrix to better represent the test sample).

[0078] Optionally, after obtaining the sparse matrix (sparse coefficients), the status of the optical module can be classified according to the minimum reconstruction error principle to obtain the health status of the optical module:

[0079]

[0080] That is, for each category c (c = 1, 2, 3, corresponding to normal, warning, and alarm status respectively), the test sample The multi-feature dictionary (matrix) corresponding to the category and sparse coefficients (matrix) The combined reconstruction results are compared and the errors between them are calculated. By comparing the reconstruction errors of all categories (c = 1, 2, 3), the category with the smallest error is selected as the classification result of the test sample. arg min means finding the category c that minimizes the objective function (reconstruction error) as the final classification result.

[0081] Classification through the sparse representation classification model can take into account the multi-feature historical time performance characteristics and topological information characteristics of the optical module. Compared with other machine learning models, it can better capture the correlation of performance characteristic data and the trend of time changes, which helps to accurately determine the status of the optical module.

[0082] As an optional implementation, after obtaining the classification result output by the status classification model, the method further includes: when the classification result is a warning, sending warning information to the target object, wherein the warning information is used to instruct the target object to check the first port.

[0083] Step S206 : When it is determined that the first optical module is abnormal, the operating status information of the optical module of the second port is obtained, wherein the second port is connected to the first port through an optical fiber.

[0084] Optionally, the link information of the port is obtained through the LLDP protocol (Link Layer Discovery Protocol), and the optical module status of the local end (first port) and the opposite end (second port) and the normalized DDM data are combined to build an optical module database. The specific structure of the optical module database is shown in the following table:

[0085]

[0086]

[0087] The optical module database specifically stores the following information:

[0088] (1) Local port information: device, (first) port number, optical module status (healthy, warning, alarm), link type (uT, dT, user access), normalized DDM information, and fault cause.

[0089] (2) Neighbor device information: user access number, device, (second) port number, optical module status (healthy, warning, alarm), and normalized DDM information.

[0090] By acquiring the operating status information of the optical module of the second port recorded in the database, it can be determined whether the fault occurs at the local end (first port) or the opposite end (second port).

[0091] Step S208 : determining a fault detection result based on the optical module operation status information of the first port and the optical module operation status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0092] In the technical solution provided in step S208, the optical module operating status information includes the received optical power and the transmitted optical power, and the first port fault includes a transmit fault and a receive fault; determining the fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port includes: when the transmitted optical power of the second port is less than a first threshold, determining that the fault detection result is a second port fault; when the transmitted optical power of the second port is not less than the first threshold and the received optical power of the first port is less than the second threshold, determining that the fault detection result is a receive fault; when the transmitted optical power of the first port is less than the first threshold, determining that the fault detection result is a transmit fault.

[0093] Optionally, Figure 7A flow chart of fault cause inference is shown in FIG. Figure 7 As shown, first check whether the optical module (first optical module) at the local end (first port) is abnormal. If abnormal, verify in sequence whether the remote end transmit power (transmitted optical power of the second port), the local end receive power (received optical power of the first port), and the local end transmit power (transmitted optical power of the first port) are lower than the threshold (all are stored normalized values), which correspond to diagnostic results such as "remote end fault", "fiber break / receiving fault", and "transmitter aging", respectively.

[0094] Specifically, first determine whether the local optical module (first optical module) is abnormal. If it is not abnormal, determine whether the classification result of the first optical module is a warning. If it is a warning, send a warning message to the target object, wherein the warning information is used to instruct the target object to check or pay special attention to the first port.

[0095] When the first optical module is abnormal, determine whether the transmitted optical power of the second port (opposite end Tx power) is less than the first threshold. When the transmitted optical power of the second port is less than the first threshold, determine that the fault detection result is a second port fault (opposite end fault). When the transmitted optical power of the second port is not less than the first threshold and the received optical power of the first port is less than the second threshold, determine that the fault detection result is a receiving fault (including a fiber break fault / a local end receiving fault). When the transmitted optical power of the first port is less than the first threshold, determine that the fault detection result is a transmitting fault (for example, transmitter aging). If the transmitted optical power of the second port is not less than the first threshold and the received optical power of the first port is not less than the second threshold and the transmitted optical power of the first port is not less than the first threshold, determine that the detection result is an unknown fault, and determine that the fault handling suggestion is to recommend checking the log.

[0096] Optionally, the type of the first port includes at least one of the following: a user access link port, an uplink port, and a downlink port; when it is determined that the fault detection result is a first port fault, the method further includes: determining whether there is a backup port for the first port; when there is no backup port for the first port, determining that the alarm level corresponding to the first port is a level one alarm; when there is a backup port for the first port, determining that the alarm level corresponding to the first port is a level two alarm; determining a fault handling suggestion based on the fault detection result, the type of the first port, and the alarm level, wherein the fault handling suggestion is used to instruct the target object to handle the fault; and sending alarm information to the target object, wherein the alarm information includes: the alarm level, the fault detection result, and the fault handling suggestion.

[0097] Optionally, Figure 8 A schematic diagram of an alarm process is shown, such as Figure 8As shown, the fault status of the local port is first detected. If a fault is confirmed, the access type of the link is branched and the scope of the optical module fault is inferred: for user access links, the existence of the user_access_code is verified to determine whether a user-level interruption alarm is triggered; for uT uplink links, the device is checked for backup uplink ports (other uT ports associated with device_name) to determine the risk of device disconnection; for dT downlink links, the backup uplink status of the neighbor_device is used to assess the impact of downstream device isolation. The entire process combines neighbor information (neighbor_* fields) and real-time status (module_status) in the database to dynamically generate red alarms and orange warnings, accurately locate the scope of the fault impact (single user, local network, or entire device disconnection), and indicate potential risks (such as traffic limit exceeding).

[0098] Optionally, if there is no backup port, a red alarm (level 1 alarm) is sent; if there is a backup port, an orange alarm (level 2 alarm) is sent. For example, if the link type (type of the first port) is a user access link port, and if the database contains a user_access_code, an orange alarm is sent, including the fault detection result (fault reason), the scope of impact (which may affect the user_access_code), and the corresponding fault handling suggestions (please pay attention to whether the traffic exceeds the limit); if the database does not contain a user_access_code, a red alarm is sent (faultreason, user_access_code user interruption). If the link type is an uplink link port uT, determine whether the first port has a backup port (whether device_name has other uplink ports), and if there is a backup port, send an orange alarm (fault reason, device_name has a backup uplink, and the fault handling suggestion is to pay attention to the risk of disconnection); if there is no backup port, send a red alarm (fault reason, device_name device is disconnected). When the link type is the downlink port dT, determine whether the first port has a backup port (whether the neighbor_device has other uplink ports), and send an orange alarm (fault reason, neighbor_device has a backup uplink, the fault handling suggestion is to pay attention to traffic limit excess) if a backup port exists; if no backup port exists, send a red alarm (fault reason, neighbor_device is offline).

[0099] Optionally, Figure 9 A schematic diagram of a fault warning message is shown, such as Figure 9 As shown, alarm information includes the alarm level, fault detection results, and troubleshooting suggestions (determined by the fault detection results, the type of the first port, and the alarm level). Alarm levels are defined based on the scope and urgency of the fault, providing maintenance personnel with a prioritized approach to troubleshooting. Level 1 alarms indicate widespread impact and require immediate attention; level 2 alarms may have a localized impact or present potential risks, requiring scheduled maintenance. By prioritizing, the system guides the appropriate allocation of maintenance resources, ensuring that critical issues are addressed first and non-urgent issues are properly addressed without impacting network stability. Automatically generated troubleshooting suggestions intelligently recommend the most appropriate maintenance strategy. For example, for a fault on a user access link port, suggestions may include checking the fiber connection, replacing the optical module, or notifying the user to inspect their equipment. For an uplink or downlink port, suggestions may involve link redundancy switching or device-level troubleshooting. Automated suggestion generation significantly reduces manual judgment, making maintenance decisions faster and more accurate.

[0100] Optionally, the alarm generation API can be connected to the message sending system, and the Webhook address provided by the message sending program can be used to create an alarm connection interface to achieve real-time push of alarm information.

[0101] Optionally, Figure 10 A schematic diagram of a fault detection system is shown in FIG. Figure 10 As shown in the figure, the fault detection system includes a data acquisition layer, a data processing layer and a user interaction layer. The data acquisition layer collects data from data network equipment (including switches, routers and OLTs) through the new generation cloud network acquisition and control platform (including optical module DDM data collected by the data acquisition module and network topology data collected by the topology discovery module). The data is used for fault prediction through the data processing layer, and the analysis results are sent to the alarm management platform of the user interaction layer to generate a fault impact report (alarm information). Data processing includes normalization processing, feature selection module (features include time series DDM features and master-slave switching information), sparse prediction module and fault analysis and propagation module.

[0102] Through the above steps, it is possible to achieve early warning of optical module performance degradation. By introducing the analysis technology that integrates network topology perception and multi-dimensional data, it is possible to more effectively analyze the operating status of the optical module, issue an early warning before a fault occurs, and improve the reliability and stability of the optical communication network. The method embodiment of the present application is suitable for evaluating the health status of the optical module and locating the cause of the optical module alarm and inferring the impact range. The optical module status monitoring algorithm of the method embodiment of the present application breaks through the limitations of traditional single DDM parameter analysis, introduces features such as redundant switching records and historical fault modes, constructs a sparse classification prediction model, and provides early warning of potential faults. In addition, topology perception and collaborative diagnosis technology are introduced. By analyzing the logical connection relationship between the upper and lower links, redundant links, etc. of the optical module in the network, combined with the DDM parameters and link status of the opposite module, the optical module itself is distinguished from the optical fiber abnormality and the opposite module abnormality, the misjudgment rate is reduced, and the fault impact range is accurately inferred. The method embodiment of the present application achieves precise, automated, and low-cost health management of the optical module through the deep integration of intelligent algorithms and topology perception, providing core technical support for the stable operation of the optical communication network. It can not only monitor optical module obstacles, but also monitor the aging status of optical modules, and focus on link conditions in advance. In addition, based on the status of the local and opposite optical modules, DDM information, and neighbor information, it can intelligently infer and detect the cause of the optical module failure and the scope of influence, promoting intelligent operation and maintenance. In addition, the implementation threshold of the method embodiment of the present application is low, and no additional detection devices are required. It has the characteristics of real-time, accurate, practical, safe, and zero operation. It can be applied to any operator and has strong practicality and versatility. Specifically, the method embodiment of the present application has the following advantages:

[0103] 1. Considering that the switching of the primary and backup links may cause a sudden change in the load rate of the optical module, accelerating the aging of the device, the method embodiment of the present application breaks through the limitations of single DDM data modeling, introduces redundant switching information, and combines a sparse classification model to train a multi-factor fault prediction model to improve the prediction accuracy of the optical module failure probability.

[0104] 2. Accurately locate the root cause of the fault: By analyzing the logical connection relationship of the optical module in the network topology (uplink and downlink links, redundant links, etc.), combined with collaborative analysis of the module parameters and link status at the opposite end, it can distinguish between optical module faults and external link / device anomalies, reducing the misjudgment rate.

[0105] 3. Establish a topology-driven fault impact assessment mechanism: Calculate the potential impact of optical module failures based on the network topology, providing maintenance personnel with a basis for repair priority decisions and minimizing service losses.

[0106] The present invention provides a fault detection device. Figure 11 is a structural diagram of the device, such as Figure 11 As shown, the device includes: a first processing module 110, used to obtain the primary-backup link switching information of the first port and the optical module operating status information of the first port, wherein the primary-backup link switching information is used to indicate whether the primary-backup link switching occurs at the first port; a second processing module 112, used to determine whether the first optical module corresponding to the first port has an abnormality based on the primary-backup link switching information and the optical module operating status information of the first port; a third processing module 114, used to obtain the optical module operating status information of the second port when it is determined that the first optical module has an abnormality, wherein the second port is connected to the first port through an optical fiber; a fourth processing module 116, used to determine the fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0107] In some embodiments of the present application, the optical module operating status information includes received optical power and transmitted optical power, and the first port fault includes a transmit fault and a receive fault; the fourth processing module 116 determines the fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port, including: when the transmitted optical power of the second port is less than the first threshold, determining that the fault detection result is a second port fault; when the transmitted optical power of the second port is not less than the first threshold and the received optical power of the first port is less than the second threshold, determining that the fault detection result is a receive fault; when the transmitted optical power of the first port is less than the first threshold, determining that the fault detection result is a transmit fault.

[0108] In some embodiments of the present application, the type of the first port includes at least one of the following: a user access link port, an uplink port, and a downlink port; when it is determined that the fault detection result is a first port fault, the fourth processing module 116 is further used to: determine whether there is a backup port for the first port; when there is no backup port for the first port, determine that the alarm level corresponding to the first port is a level one alarm; when there is a backup port for the first port, determine that the alarm level corresponding to the first port is a level two alarm; determine a fault handling suggestion based on the fault detection result, the type of the first port, and the alarm level, wherein the fault handling suggestion is used to instruct the target object to handle the fault; send alarm information to the target object, wherein the alarm information includes: alarm level, fault detection result, and fault handling suggestion.

[0109] In some embodiments of the present application, the first processing module 110 obtains the primary-backup link switching information of the first port and the optical module operating status information of the first port, including: collecting the primary-backup link switching information of the first port and the optical module operating status information of the first port according to a preset time period, wherein the optical module operating status information includes at least one of the following: operating temperature, operating voltage, bias current, received optical power and transmitted optical power; determining the optical module model of the optical module corresponding to the first port, and obtaining threshold information corresponding to the optical module model, wherein the threshold information includes the minimum value and the maximum value corresponding to the optical module operating status information; and normalizing the optical module operating status information of the first port according to the threshold information.

[0110] In some embodiments of the present application, normalizing the optical module operating status information of the first port based on the threshold information includes: when the value of the optical module operating status information is less than the minimum value corresponding to the optical module operating status information, setting the value of the optical module operating status information to 0; when the value of the optical module operating status information is greater than the maximum value corresponding to the optical module operating status information, setting the value of the optical module operating status information to 1; when the value of the optical module operating status information is between the minimum value and the maximum value corresponding to the optical module operating status information, scaling the value of the optical module operating status information according to the minimum value and the maximum value.

[0111] In some embodiments of the present application, the second processing module 112 determines whether there is an abnormality in the first optical module corresponding to the first port based on the primary-backup link switching information and the optical module operating status information of the first port, including: inputting the primary-backup link switching information and the optical module operating status information of the first port into a state classification model, the state classification model is used to classify the health status of the first optical module; obtaining the classification result output by the state classification model, wherein the classification result includes health, warning and abnormality; when the classification result is abnormal, there is an abnormality in the first optical module corresponding to the first port.

[0112] In some embodiments of the present application, before inputting the primary and backup link switching information and the optical module operating status information of the first port into the state classification model, the second processing module 112 is also used to: obtain historical sampling data, wherein the historical sampling data includes the historical primary and backup link switching information of the port and the historical optical module operating status information; normalize the historical sampling data; determine a multi-feature dictionary and a sparse coefficient based on the normalized historical sampling data; and construct a state classification model based on the multi-feature dictionary and the sparse coefficient.

[0113] In some embodiments of the present application, after obtaining the classification result output by the status classification model, the second processing module 112 is further used to: when the classification result is a warning, send a warning message to the target object, wherein the warning message is used to instruct the target object to check the first port.

[0114] It should be noted that the various modules in the above-mentioned fault detection device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0115] An embodiment of the present application provides a non-volatile storage medium, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following fault detection method: obtaining the primary-backup link switching information of the first port and the optical module operation status information of the first port, wherein the primary-backup link switching information is used to indicate whether the primary-backup link switching occurs at the first port; determining whether the first optical module corresponding to the first port has an abnormality based on the primary-backup link switching information and the optical module operation status information of the first port; when it is determined that the first optical module has an abnormality, obtaining the optical module operation status information of the second port, wherein the second port is connected to the first port through an optical fiber; determining a fault detection result based on the optical module operation status information of the first port and the optical module operation status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0116] An embodiment of the present application provides an electronic device, comprising: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the following fault detection method is executed when the program is run: obtaining primary-backup link switching information of a first port and optical module operating status information of the first port, wherein the primary-backup link switching information is used to indicate whether primary-backup link switching occurs at the first port; determining whether an abnormality exists in a first optical module corresponding to the first port based on the primary-backup link switching information and the optical module operating status information of the first port; obtaining optical module operating status information of a second port when it is determined that an abnormality exists in the first optical module, wherein the second port is connected to the first port via an optical fiber; determining a fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0117] An embodiment of the present application provides a computer program product, including a computer program, which implements the following fault detection method when executed by a processor: obtaining primary-backup link switching information of a first port and optical module operating status information of the first port, wherein the primary-backup link switching information is used to indicate whether primary-backup link switching occurs at the first port; determining whether a first optical module corresponding to the first port has an abnormality based on the primary-backup link switching information and the optical module operating status information of the first port; when it is determined that the first optical module has an abnormality, obtaining optical module operating status information of a second port, wherein the second port is connected to the first port via an optical fiber; determining a fault detection result based on the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

[0118] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0119] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0120] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0121] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0122] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0123] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A fault detection method, characterized in that: include: Acquire active / standby link switching information of a first port and operating status information of an optical module of the first port, wherein the active / standby link switching information is used to indicate whether active / standby link switching occurs on the first port; Determining whether a first optical module corresponding to the first port is abnormal based on the primary / backup link switching information and the operating status information of the optical module of the first port; When it is determined that the first optical module is abnormal, obtaining operating status information of the optical module of a second port, wherein the second port is connected to the first port through an optical fiber; A fault detection result is determined according to the optical module operating status information of the first port and the optical module operating status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

2. The fault detection method according to claim 1, characterized in that: The optical module operation status information includes received optical power and transmitted optical power, and the first port fault includes a transmit fault and a receive fault; Determining a fault detection result according to the optical module operating status information of the first port and the optical module operating status information of the second port includes: When the transmitted optical power of the second port is less than a first threshold, determining that the fault detection result is a fault of the second port; When the transmitted optical power of the second port is not less than a first threshold and the received optical power of the first port is less than a second threshold, determining that the fault detection result is the receiving fault; When the transmitted optical power of the first port is less than a first threshold, the fault detection result is determined to be the transmission fault.

3. The fault detection method according to claim 2, characterized in that: The type of the first port includes at least one of the following: a user access link port, an uplink port, and a downlink port; When it is determined that the fault detection result is a fault of the first port, the method further includes: Determining whether a backup port exists for the first port; In a case where the first port does not have a backup port, determining that the alarm level corresponding to the first port is a level one alarm; In a case where a backup port exists for the first port, determining that the alarm level corresponding to the first port is a level 2 alarm; Determining a fault handling suggestion based on the fault detection result, the type of the first port, and the alarm level, wherein the fault handling suggestion is used to instruct a target object to handle the fault; Sending an alarm message to a target object, wherein the alarm message includes: the alarm level, the fault detection result, and the fault handling suggestion.

4. The fault detection method according to claim 1, characterized in that: Acquiring the primary / backup link switching information of the first port and the optical module operating status information of the first port includes: Collecting active / standby link switching information of the first port and operating status information of the optical module of the first port according to a preset time period, wherein the operating status information of the optical module includes at least one of the following: operating temperature, operating voltage, bias current, received optical power, and transmitted optical power; Determine an optical module model of an optical module corresponding to the first port, and obtain threshold information corresponding to the optical module model, wherein the threshold information includes a minimum value and a maximum value corresponding to the optical module operation status information; Normalizing the operating status information of the optical module of the first port according to the threshold information.

5. The fault detection method according to claim 4, characterized in that: Normalizing the optical module operating status information of the first port according to the threshold information includes: When the value of the optical module operation status information is less than the minimum value corresponding to the optical module operation status information, setting the value of the optical module operation status information to 0; When the value of the optical module operation status information is greater than the maximum value corresponding to the optical module operation status information, setting the value of the optical module operation status information to 1; When the value of the optical module operation status information is between a minimum value and a maximum value corresponding to the optical module operation status information, scaling processing is performed on the value of the optical module operation status information according to the minimum value and the maximum value.

6. The fault detection method according to claim 1, characterized in that: Determining whether a first optical module corresponding to the first port is abnormal according to the primary / backup link switching information and the optical module operating status information of the first port includes: Inputting the active / standby link switching information and the operating status information of the optical module of the first port into a state classification model, wherein the state classification model is used to classify the health status of the first optical module; Obtaining a classification result output by the status classification model, wherein the classification result includes healthy, warning, and abnormal; When the classification result is abnormal, the first optical module corresponding to the first port is abnormal.

7. The fault detection method according to claim 6, characterized in that: Before inputting the primary / backup link switching information and the optical module operating status information of the first port into a status classification model, the method further includes: Acquire historical sampling data, wherein the historical sampling data includes historical active / standby link switching information of the port and historical optical module operating status information; performing normalization processing on the historical sampling data; Determining a multi-feature dictionary and a sparse coefficient based on the normalized historical sampling data; The state classification model is constructed based on the multi-feature dictionary and the sparse coefficients.

8. The fault detection method according to claim 6, characterized in that: After obtaining the classification result output by the status classification model, the method further includes: when the classification result is a warning, sending warning information to a target object, wherein the warning information is used to instruct the target object to check the first port.

9. A fault detection device, characterized in that: include: A first processing module is configured to obtain active / standby link switching information of a first port and operating status information of an optical module of the first port, wherein the active / standby link switching information is used to indicate whether active / standby link switching occurs on the first port; a second processing module, configured to determine whether a first optical module corresponding to the first port is abnormal based on the primary / backup link switching information and the operating status information of the optical module of the first port; a third processing module, configured to, when determining that the first optical module is abnormal, obtain operating status information of the optical module of a second port, wherein the second port is connected to the first port via an optical fiber; The fourth processing module is configured to determine a fault detection result based on the optical module operation status information of the first port and the optical module operation status information of the second port, wherein the fault detection result includes at least one of the following: first port fault, second port fault.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the fault detection method according to any one of claims 1 to 8.

11. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the fault detection method according to any one of claims 1 to 8 is executed when the program is run.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the fault detection method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Fault early warning method and electronic equipment

    CN121603394A