Fault detection method, apparatus, system and device for optical fiber communication system
By extracting feature data from the indicator data of optical modules and using fault detection models or feature data matching methods, the problem of difficulty in distinguishing fault types of optical modules is solved, achieving efficient fault detection and rapid service recovery.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2026-04-02
AI Technical Summary
Existing technologies cannot effectively distinguish the types of optical module failures, resulting in low efficiency in optical module failure detection in fiber optic communication systems and prolonging service recovery time.
By extracting feature data from the optical module's performance data, and using a fault detection model or feature data matching method, the specific fault type of the optical module can be determined.
It improves the efficiency and accuracy of identifying optical module fault types, provides more targeted repair measures, and reduces the time required for service recovery.
Smart Images

Figure CN2025096943_02042026_PF_FP_ABST
Abstract
Description
Fault detection method, device, system and equipment of fiber communication system
[0001] The present application claims priority from the Chinese patent application No. 202411360351.3 filed with the State Intellectual Property Office on September 26, 2024 and entitled "Fault detection method, device, system and equipment of fiber communication system", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the field of fault detection, in particular to a fault detection method, device, system and equipment of fiber communication system. BACKGROUND
[0003] In a large-scale data center (such as an artificial intelligence training center), a considerable number of optical fibers are used to improve communication speed, which results in a large number of optical modules in the data center to realize the conversion between optical signals and electrical signals, thereby causing the communication port to be interrupted or even the service to be interrupted due to the fault of the optical module, which becomes one of the common faults of the large-scale data center. Among various faults that may occur in the optical module, the two most common faults are optical module contamination and optical module loosening, but the related art cannot distinguish between the two faults when detecting the fault of the optical module, which seriously reduces the efficiency of solving the fault of the optical module and prolongs the time required for the service to return to normal. SUMMARY
[0004] The present application provides a fault detection method, device, system and equipment of fiber communication system, which improves the efficiency and accuracy of determining the fault type of the optical module according to the feature data extracted from the index data of the optical module, and reduces the time required for the service to return to normal.
[0005] In a first aspect, the present application provides a fault detection method of a fiber communication system, the fiber communication system comprising a first optical module, the method comprising: obtaining index data of the first optical module, the index data being used to indicate the running state of the optical module; performing feature extraction on the index data of the first optical module to obtain first feature data; and determining the fault type of the first optical module according to the first feature data.
[0006] It can be understood that the first feature data obtained by performing feature extraction on the obtained index data of the first optical module can determine the specific fault type of the optical module, which can improve the efficiency of determining the fault type of the optical module, thereby providing more specific fault information for the maintenance engineer, making the repair work more targeted, improving the repair efficiency, and further reducing the time required for the service to return to normal.
[0007] In a possible implementation, the type of fault occurring to the first optical module is determined according to the first feature data, including: obtaining index data of a second optical module; the second optical module is located on the same optical link as the first optical module; performing feature extraction on the index data of the second optical module to obtain second feature data; and determining the type of fault occurring to the first optical module according to the first feature data and the second feature data.
[0008] It can be understood that, according to the feature data of a plurality of optical modules including the first optical module, the accuracy of determining the type of fault occurring to the first optical module can be further improved.
[0009] In a possible implementation, the index data of the second optical module is obtained, including: determining the type of fault occurring to the first optical module and the probability of occurrence of the type of fault according to the first feature data; and if the probability of occurrence of the fault meets a first range, the index data of the second optical module is obtained.
[0010] It can be understood that, by setting a suitable first range, the index data of the second optical module can be further obtained when the probability of occurrence of the type of fault occurring to the first optical module is not high enough according to the first feature data, which helps to perform more accurate fault detection subsequently, and meanwhile, the index data of the second optical module can be avoided to be obtained when the probability of occurrence of the type of fault occurring to the first optical module is high enough according to the first feature data, thereby reducing the occupation of computing resources.
[0011] In a possible implementation, the type of fault occurring to the first optical module is determined according to the first feature data, including: inputting the first feature data into a fault detection model to obtain a fault detection result output by the fault detection model; the fault detection model is trained according to sample feature data corresponding to each type of fault; and the type of fault occurring to the first optical module is determined according to the fault detection result.
[0012] It can be understood that, since the first feature data can include feature data in multiple dimensions, the fault detection model can efficiently and accurately obtain the corresponding fault detection result according to the first feature data, thereby improving the efficiency and accuracy of determining the type of fault occurring to the first optical module.
[0013] In a possible implementation, the type of fault occurring to the first optical module is determined according to the first feature data, including: determining the type of fault occurring to the first optical module according to a matching degree between the first feature data and set feature data corresponding to each type of fault.
[0014] Understandably, determining the fault type by calculating the matching degree between feature data can reduce the investment in the early preparation stage since there is no need to pre-train the model, and it is easier to deploy. Moreover, this method of calculating the matching degree has high stability.
[0015] In one possible implementation, feature extraction is performed on the indicator data of the first optical module to obtain first feature data, including: feature extraction is performed on the indicator data of the first optical module under at least one set dimension to obtain first feature data under the set dimension, wherein the set dimension includes statistical feature dimension, time-series feature dimension and / or frequency domain feature dimension.
[0016] Understandably, extracting the characteristics of the first optical module's index data across different dimensions can help improve the accuracy of determining the type of fault that occurred in the first optical module.
[0017] Secondly, this application provides a fault detection device for performing any of the fault detection methods provided in the first aspect above.
[0018] In one possible implementation, this application can divide the fault detection device into functional modules according to the method provided in the first aspect above. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into a processing module. For example, this application can divide the fault detection device into a display module, a processing module, and an update module according to function. The descriptions of the possible technical solutions and beneficial effects of the various functional modules described above can be found in the technical solutions provided in the first aspect above or its corresponding possible implementations, and will not be repeated here.
[0019] Thirdly, this application provides a fault detection system, comprising: a first device for acquiring indicator data of a first optical module, the indicator data indicating the operating status of the optical module; extracting features from the indicator data of the first optical module to obtain first feature data; determining the fault type and probability of occurrence of the fault type of the first optical module based on the first feature data; sending a trigger signal to a second device in the fault detection system if the probability of occurrence of the fault meets a first range, the trigger signal instructing the second device to perform fault detection on the first optical module; and a second device for acquiring the first feature data and indicator data of a second optical module in response to receiving the trigger signal, the second optical module being located on the same optical link as the first optical module; extracting features from the indicator data of the second optical module to obtain second feature data; and determining the fault type of the first optical module based on the first feature data and the second feature data.
[0020] In a fourth aspect, an embodiment of the present application provides a fault detection method of an optical fiber communication system, applied to a first device and a second device, the optical fiber communication system comprising a first optical module, the method comprising: the first device obtaining index data of the first optical module, the index data being used to indicate an operating state of the optical module; the first device performing feature extraction on the index data of the first optical module to obtain first feature data; the first device determining a fault type of the first optical module and an occurrence probability of the fault type according to the first feature data; if the occurrence probability of the fault meets a first range, the first device sends a trigger signal to the second device in the fault detection system, the trigger signal being used to instruct the second device to perform fault detection on the first optical module; the second device, in response to receiving the trigger signal, obtains the first feature data and index data of a second optical module, the second optical module being located on the same optical link as the first optical module; the second device performs feature extraction on the index data of the second optical module to obtain second feature data; and the second device determines the fault type of the first optical module according to the first feature data and the second feature data.
[0021] In a fifth aspect, an embodiment of the present application provides a computing device, the computing device comprising a processor and a memory, the processor being coupled with the memory; the memory being used to store computer instructions, the computer instructions being loaded and executed by the processor to enable the computing device to implement the fault detection method according to the above-mentioned aspect.
[0022] In a sixth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium storing at least one computer program instruction, the computer program instruction being loaded and executed by a processor to implement the fault detection method according to the above-mentioned aspect.
[0023] In a seventh aspect, an embodiment of the present application provides a computer program product, the computer program product comprising computer instructions stored in a computer readable storage medium. A processor of a computing device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computing device to perform the fault detection method provided in various optional implementation manners of the first aspect.
[0024] The specific description of the second aspect to the seventh aspect and various implementation manners thereof in the present application can refer to the detailed description in the first aspect and various implementation manners thereof; and the beneficial effects of the second aspect to the seventh aspect and various implementation manners thereof can refer to the beneficial effect analysis in the first aspect and various implementation manners thereof, which will not be described herein again.
[0025] These aspects or other aspects of the present application will be more apparent in the following description. BRIEF DESCRIPTION OF DRAWINGS
[0026] FIG. 1 is a common optical module working schematic diagram;
[0027] FIG. 2 is a schematic diagram of a related art;
[0028] FIG. 3 is a flowchart of another related art;
[0029] FIG. 4 is a flowchart of another related art;
[0030] FIG. 5 is a schematic diagram of a network device according to an example embodiment;
[0031] FIG. 6 is a flowchart of a fault detection method according to an example embodiment;
[0032] FIG. 7 is a schematic diagram of a hierarchical fault detection system according to an example embodiment;
[0033] FIG. 8 is a flowchart of a fault detection method related to the architecture shown in FIG. 7;
[0034] FIG. 9 is a schematic diagram of a fault detection apparatus according to an example embodiment. DETAILED DESCRIPTION
[0035] To make the objects, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0036] In the present document, "a plurality of" means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The character " / " generally represents that the associated objects before and after are in an "or" relationship.
[0037] In addition, in the description of the embodiments of the present application, "a plurality of" means two or more, unless otherwise specified. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0038] In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", and the like are used to distinguish the same or similar items with substantially the same functions and effects. Those skilled in the art can understand that the terms "first", "second", and the like do not limit the quantity and execution order, and the terms "first", "second", and the like do not necessarily mean different. Meanwhile, in the embodiments of the present application, the words "exemplary" or "for example" are used to represent as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplary" or "for example" are intended to present the relevant concept in a specific manner, for ease of understanding.
[0039] For the sake of clear and concise description of each embodiment below, a brief introduction of related technologies is given below.
[0040] (1) Optical module
[0041] An optical module is an important component in optical fiber communication, and is an optoelectronic device used to realize the functions of photoelectric conversion and electro-optical conversion in the process of optical signal transmission. The optical module works at the physical layer of the open system interconnect (OSI) model, and is one of the core devices in an optical fiber communication system.
[0042] An optical module mainly consists of optoelectronic devices (such as optical transmitters and optical receivers), functional circuits, and optical interfaces, and its main function is to realize the functions of photoelectric conversion and electro-optical conversion in optical fiber communication. As shown in FIG. 1, FIG. 1 is a schematic diagram of a common optical module.
[0043] In FIG. 1, through the interconnection of a No. 1 optical module and a No. 2 optical module, communication between various network devices (not shown in FIG. 1) connected to a No. 1 Ethernet switch and a No. 2 Ethernet switch can be realized, wherein a transport (Tx) port of the No. 1 optical module and a receive (Rx) port of the No. 2 optical module are connected by an optical fiber, and a receive port of the No. 1 optical module and a transport port of the No. 2 optical module are connected by an optical fiber.
[0044] For example, based on FIG. 1, when communicating, the No. 1 Ethernet switch can first input an electrical signal of a certain code rate to the transmitting port of the No. 1 optical module. After processing by the internal drive chip, the No. 1 optical module can drive the laser diode (LD) or light-emitting diode (LED) to emit a corresponding modulated optical signal, which is transmitted to the receiving interface of the No. 2 optical module through an optical fiber. Then, the No. 2 optical module converts the optical signal into an electrical signal, and outputs a corresponding electrical signal after preamplification, and finally sends the electrical signal to the No. 2 Ethernet switch.
[0045] When the optical module is working, the running state of the optical module can be determined by the temperature of the optical module, the RX power, the TX power, and other index data of the optical module.
[0046] In some possible embodiments, the type of optical module failure can be determined according to the characteristics of the optical module index data.
[0047] The application scenarios of the embodiments of the present application are exemplarily introduced below.
[0048] With the rapid development of artificial intelligence (AI) technology, the scale of deep neural network models has also become larger. The parameter quantity of current many models is moving from hundreds of billions to trillions, and will continue to grow in the future, possibly further reaching 100 trillion or even more. The explosive growth of parameter quantity, on the one hand, improves the ability of these complex models to handle various problems, and even hopes to achieve general artificial intelligence (GAI), but at the same time, the explosive growth of parameter quantity also makes the demand for underlying computing power of these models with large-scale parameters and complex computing structures (referred to as: large models) further upgraded.
[0049] In order to meet the computing power demand of large models (such as shortening the training time of large models, accelerating the iterative upgrade of models, etc.), a large number of acceleration cards (such as graphics cards, tensor processors or other special AI acceleration chips) are usually needed to perform high-speed calculations in a parallel manner. If a data center (DC) that can be used to provide computing power for large models contains ten thousand or more acceleration cards, such a data center is also called a ten-thousand-card cluster.
[0050] A wancai cluster is usually composed of thousands of servers, thousands of switches, thousands of storage devices, and tens of thousands of optical fibers / tens of thousands of optical modules. Based on a cluster of such a scale, a training task of a large model will involve the high-speed operation of millions of components. It can be understood that, since each software / hardware component has a certain failure rate, and the failure modes of software and hardware are complex, fault detection of such a cluster is a great challenge.
[0051] Among the many possible faults that can occur in a wancai cluster, port flashing and service interruption caused by optical module failure are the most common problems in cluster network failure, accounting for more than 80%. However, the current research on this problem takes an average of about 3 days, which significantly affects the normal operation of the business.
[0052] Among optical module failures, the two most common known failures are optical module contamination and optical module loosening. Optical module contamination refers to the presence of dust, oil stains, etc. in the receiving port and / or transmitting port of the optical module. These contaminants can block the transmission of optical signals and affect the normal transmission of optical signals. Optical module loosening refers to the connection between the receiving port and / or transmitting port of the optical module and the optical fiber being not tight enough, which can cause light leakage and other problems, thereby affecting the normal transmission of optical signals.
[0053] In one related technology, as shown in FIG. 2, the method detects the voltage / current of the optical module power supply line through an additional voltage / current detection circuit connected to the optical module, compares it with the pre-set threshold value in the storage module, and determines whether the power supply of the optical module is normal. In the case of abnormal power supply (such as detecting that the power supply line is disconnected), the control module activates the standby optical module power supply to supply power to the optical module. The temperature of the optical module during operation is detected by a temperature sensor, and when the internal temperature of the optical module is too high, the radiator is automatically activated to cool the optical module. Since this method can only simply detect whether the power supply voltage / current or temperature of the optical module is abnormal, it cannot specifically determine whether the optical module has failed due to contamination, loosening, etc.
[0054] In another related technology, as shown in FIG. 3, which is a flowchart of another related technology, the method comprises: S102, obtaining detection data; S104, analyzing according to the detection data to determine whether the optical module is a faulty optical module; S106, if yes, adjusting the running state of the optical module by closing the abnormal channel or adjusting the controllable parameter, so that the faulty optical module is in a low running state; S108, if no, saving the detection data to a server; S110, obtaining detection data of the faulty optical module in a low running state; S112, analyzing according to the detection data of the faulty optical module in the low running state, if the abnormal parameter of the detection data reaches a first preset threshold or is not adjusted to a preset threshold interval within a first preset time threshold, automatically interrupting the transmission interface corresponding to the faulty optical module. Since this method can only determine whether the optical module is a faulty optical module, and a unified processing method is used for the faulty optical module, this method has poor flexibility, and can only detect whether the optical module has failed, but cannot determine the specific fault type, and still has the problem of low efficiency in solving optical module failure.
[0055] In another related technology, as shown in FIG. 4, which is a flowchart of another related technology, the method comprises: step S201, providing a conductor on a connector of an optical fiber, providing two detection positions, a first detection position and a second detection position, at corresponding positions of an optical module, and connecting the first detection position and the second detection position to a microcontroller unit (MCU). Wherein, the corresponding positions refer to that when the connector is inserted into the optical module and the optical fiber is in place, the conductor contacts both detection positions. Step S202, the MCU detects the state of the first detection position and the second detection position respectively. Step S203, the MCU compares the states of the first detection position and the second detection position to determine whether the optical fiber is in place. This method realizes detection of a single optical fiber loosening fault by additionally providing a conductor on the connector of the optical fiber, but when other types of faults occur in the optical module, if this method is used, the type of fault occurring in the optical module cannot be determined, which affects the efficiency of restoring the service.
[0056] In summary, in many related technologies, classification detection of optical module dirt, loosening or other faults cannot be realized, so that when the service is interrupted due to the failure of the optical module in a large data center (such as the aforementioned 10,000 card cluster), it is difficult to quickly detect and locate the problem, which prolongs the time required to restore the service.
[0057] Therefore, the application provides a high-efficiency fault detection method for an optical fiber communication system, which determines the specific fault type of an optical module according to the feature data (such as the standard deviation in the statistical dimension and the change trend slope in the time sequence dimension) of the index data of the optical module in each dimension, realizes the classified detection of the fault of the optical module, improves the efficiency of determining the fault type of the optical module, and thus can provide more specific fault information for the operation and maintenance engineers, make the repair work more targeted, improve the repair efficiency, and further reduce the required time for recovering the service.
[0058] In some possible embodiments, the method comprises: obtaining index data of a first optical module, wherein the index data is used to indicate the running state of the optical module; performing feature extraction on the index data of the first optical module to obtain first feature data; and determining the fault type of the first optical module according to the first feature data. Since the index data of the first optical module can indicate the running state of the optical module, the network device extracts the first feature data from the index data, and then determines the specific fault type of the first optical module according to the first feature data, which significantly improves the efficiency of fault detection, improves the repair efficiency, and further reduces the required time for recovering the service.
[0059] The system architecture of the embodiments of the application is exemplarily introduced below.
[0060] As shown in FIG. 5, FIG. 5 is a structural schematic diagram of a network device according to an exemplary embodiment of the application. The network device 1000 at least comprises a memory 1010, a processor 1020, a communication interface 1030 and a bus 1040.
[0061] The communication interface 1030 can be used to connect with other devices such as optical modules, and the processor 1020 can obtain the index data of the optical module through the communication interface 1030.
[0062] The processor 1020 can be used to obtain the index data of the optical module, perform feature extraction on the index data of the optical module to obtain feature data, and determine the fault type of the optical module according to the feature data.
[0063] The memory 1010 can be used to store the logic code corresponding to the fault detection method provided by the embodiments of the application, or in other words, the memory 1010 can also store the logic code corresponding to the execution of a certain step of the network device 1000 described in the following embodiments.
[0064] Optionally, the network device 1000 can be a switch, a router, an optical fiber transponder, a server, a storage device, a communication device or other devices capable of directly communicating with the optical module, or the network device 1000 can also be an analyzer device, such as a cloud analyzer, a local analyzer or other devices having the capability of obtaining the index data of the optical module.
[0065] Optionally, the memory 1010 can include random access memory (RAM), read-only memory (ROM) and the like, wherein the memory 1010 can run the necessary operating system in the RAM thereof, and the fault detection module, the feature extraction module and the like for executing the fault detection method provided in the present application.
[0066] Optionally, the processor 1020 can be a central processing unit (CPU) or other general-purpose processor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processing (DSP) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the like. The general-purpose processor can be a microprocessor or any conventional processor, and the like.
[0067] Optionally, the communication interface 1030 can be an Ethernet interface, an optical module interface, and the like.
[0068] Optionally, the bus 1040 can be a peripheral component interconnect (PCI) bus or the like, and the type of the bus is not limited in the present application. The bus can be divided into an address bus, a data bus, a control bus, and the like. For the convenience of representation, only one line is shown in FIG. 5, but it does not mean that there is only one bus or only one type of bus. The bus 1040 can include a path for transmitting information between various components (for example, the memory 1010, the processor 1020, the communication interface 1030) of the network device 1000.
[0069] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. It can be known by those skilled in the art that, with the evolution of system architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0070] For the convenience of understanding, the fault detection method provided in the present application is exemplarily introduced below in combination with the drawings, and the fault detection method is applicable to the network device shown in FIG. 5.
[0071] FIG. 6 shows a flow diagram of a fault detection method according to an embodiment of the present application. The fault detection method comprises the following steps:
[0072] S301. The network device acquires index data of the first optical module.
[0073] In the embodiments of the present application, the network device can be the network device 1000 shown in FIG. 5, and the index data is data for indicating the running state of the optical module, which can include but is not limited to the following data: the optical signal intensity (or the transmission power) sent by the optical module, the optical signal intensity (or the reception power) received by the optical module, the temperature of the optical module (which can be the temperature of the transmission port and the reception port of the optical module), the signal-to-noise ratio (SNR) of the optical signal transmitted by the optical module, and the like. In addition, each of the above-mentioned index data can further include corresponding time information.
[0074] In some possible embodiments, if the optical signal transmitted by the first optical module includes two or more optical signals of different wavelengths, the network device can acquire the index data of each optical signal of different wavelengths, or can acquire the index data of only part of the optical signals of different wavelengths when acquiring the index data of the first optical module. In the implementation where the network device acquires the index data of only part of the optical signals of different wavelengths, the network device can determine which index data of the optical signals of different wavelengths to acquire according to the user's settings, which is not limited in the present application.
[0075] For example, the optical signal transmitted by the first optical module includes optical signals of 850 nm and 1310 nm wavelengths, and the network device can acquire the optical signal intensities of the optical signals of the two different wavelengths.
[0076] For another example, the optical signal transmitted by the first optical module includes optical signals of 850 nm and 1310 nm wavelengths, and the network device acquires the SNR data of the optical signal of 850 nm wavelength received by the first optical module according to the user's settings of acquiring only the index data of the optical signal of wavelength less than 1000 nm.
[0077] S302. The network device extracts features from the index data of the first optical module to obtain first feature data.
[0078] In this step, the network device can extract features from the index data of the first optical module in at least one set dimension to obtain the first feature data in the set dimension.
[0079] Specifically, the set dimensions include but are not limited to: a statistical feature dimension, a time sequence feature dimension, and a frequency domain feature dimension.
[0080] For example, based on the determined specific features in each set dimension, the network device obtains the index data of the first optical module as the received power and the transmitted power of the optical signals of 4 different wavelengths in a period of time, and then when the network device extracts features, if in the time sequence feature dimension, the network device can arrange the received power and the transmitted power according to the time information, and then obtain a plurality of time sequence information. The network device can determine the number of mutation points in the time sequence information, the trend slope of fitting each time sequence information, and the autoregressive coefficient of each time sequence information, etc. If in the statistical feature dimension, the network device can calculate the standard deviation of the received power and / or the transmitted power corresponding to the optical signals of each wavelength, and the proportion of the received power (and / or the transmitted power) of the optical signals of each wavelength in the sum of the total received power (and / or the transmitted power) of the optical signals of 4 different wavelengths, or simply the power proportion. If in the frequency domain feature dimension, the network device can calculate the power spectral density (PSD) of the received power and / or the transmitted power corresponding to the optical signals of each wavelength, etc.
[0081] Further, the first feature data obtained by the network device can include one or more feature data in at least one set dimension, wherein which kind or kinds of feature data in each set dimension are obtained by the network device can be determined in advance, which will be described in detail later, and will not be repeated here.
[0082] S303, the network device determines the fault type of the first optical module according to the first feature data.
[0083] When the network device performs the above step S303, at least the following two possible implementation manners can be adopted:
[0084] The first possible implementation manner: the network device inputs the first feature data into a fault detection model to obtain a fault detection result output by the model, and then the network device determines the fault type of the first optical module according to the fault detection result.
[0085] It should be noted that the fault detection model can be pre-trained on the network device or other computing devices according to the sample feature data corresponding to each fault type, and then deployed on the network device, or deployed on other devices that can communicate with the network device. In the embodiments of the present application, the pre-trained fault detection model is mainly deployed on the network device as an example.
[0086] In the embodiments of the present application, the fault detection model can be specifically an extreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or other models based on a boosting algorithm. Such models can train multiple weak classifiers with limited classification detection accuracy, and finally combine them into a strong classifier with high classification detection accuracy through weighting or other methods, and thus perform well in classification detection problems.
[0087] For example, the network device inputs the number of mutation points, the change trend slope of the receiving power and the transmitting power of the first optical module in the time sequence dimension, and the standard deviation in the statistical dimension into the trained XGBoost model, and obtains the fault detection result output by the XGBoost model.
[0088] In some possible embodiments, the fault detection model can directly output the fault type of the first optical module, so that the network device can directly determine the fault type of the first optical module according to the output of the fault detection model, without further judgment, thereby further improving the efficiency of optical module fault detection.
[0089] In some possible embodiments, the fault detection result output by the fault detection model can be the probability of each fault type of the first optical module. For example, the fault detection result output by the aforementioned XGBoost model can be the probability of the dirty fault of the first optical module: 10%, the probability of the loose fault of the first optical module: 80%, and the probability of other types of faults of the first optical module: 5%. Further, the network device can determine the fault type of the first optical module as the fault type with the highest probability.
[0090] It should be noted that the aforementioned model is only an example, and the embodiments of the present application are not limited thereto.
[0091] The second possible implementation manner is that the network device determines the fault type of the first optical module according to the matching degree between the first feature data and the set feature data corresponding to each fault type.
[0092] Specifically, the feature data corresponding to each fault type can be set in advance. For example, the feature data corresponding to the dirty fault of the optical module can include that the number of time sequence mutation points of the receiving power and the transmitting power of the optical module is greater than or equal to a set number, the change trend slope is downward, and the numerical value of the slope is greater than a set value. If the first feature data meets the requirements of the above two feature data, the network device can determine that the first optical module has a dirty fault.
[0093] To further improve the accuracy of fault detection, when the network device performs the above step S303, the embodiment of the application provides a possible implementation manner. When determining the fault type of the first optical module according to the first feature data, the network device can further obtain the index data of a second optical module located on the same optical link as the first optical module. Then, the network device extracts features from the index data of the second optical module to obtain second feature data, and further determines the fault type of the first optical module according to the first feature data and the second feature data. The second optical module can be one or more other optical modules on the optical link where the first optical module is located. For example, the second optical module can be an optical module directly connected to the first optical module through an optical fiber (also referred to as a counter optical module of the first optical module), or the second optical module can be all optical modules on the optical link where the first optical module is located except the first optical module.
[0094] It can be understood that in a complex communication network such as the aforementioned large-scale data center or the Wan card cluster, the optical link where the first optical module is located can include multiple other optical modules. In the embodiment of the application, the second optical module refers to other optical modules on the same optical link as the first optical module. When the first optical module fails, the optical signal transmitted by the first optical module is likely to be abnormal. For example, when the first optical module has a dirty fault, the intensity of the optical signal transmitted by the first optical module can be lower than the normal level, thereby causing the index data of the second optical module receiving the optical signal to deviate from the normal range, for example, the receiving power of the second optical module is lower than the normal level, and so on. Another second optical module receiving the optical signal transmitted by the aforementioned second optical module will also be affected to a certain extent. In other words, when there is a failed optical module in an optical link, the influence can be full-link level, that is, the index data of the second optical module can be used to determine the fault type of the first optical module.
[0095] In some possible embodiments, when the network device extracts features from the index data of the second optical module, the same set dimension and feature data under the set dimension used when extracting features from the index data of the first optical module can be used.
[0096] For example, if the network device extracts the number of mutation points, the trend slope of the receiving power and the transmitting power of the first optical module in the time sequence feature dimension, the network device can also extract the number of mutation points, the trend slope of the receiving power and the transmitting power of the second optical module in the time sequence feature dimension when extracting features from the index data of the second optical module.
[0097] In some possible embodiments, when the network device extracts the index data of the second optical module, the set dimensions and the feature data under the set dimensions used when the network device extracts the index data of the first optical module can be completely different or partially different.
[0098] For example, if the network device extracts the received power and the transmission power of the first optical module, and the first standard deviation of the received power and / or the transmission power corresponding to the optical signals of each wavelength under the statistical feature dimension, when the network device extracts the index data of the second optical module, the network device can also extract the received power and the transmission power of the second optical module, and the second standard deviation of the received power and / or the transmission power corresponding to the optical signals of each wavelength under the statistical feature dimension. Furthermore, the network device can further calculate the difference between the first standard deviation and the second standard deviation.
[0099] In some possible embodiments, when the network device extracts the index data of the second optical module, the set dimensions and the feature data under the set dimensions used when the network device extracts the index data of the first optical module can be completely different or partially different. In this case, the network device can extract the feature data corresponding to the first optical module from the index data of the second optical module.
[0100] For example, if the network device extracts the change trend slope of the transmission power of the first optical module, the network device can further extract the change trend slope of the received power of the second optical module.
[0101] In order to save computing resources and further improve the detection efficiency, the present application provides a possible implementation manner. Before obtaining the index data of the second optical module, the network device determines the fault type of the first optical module and the occurrence probability of the fault type according to the first feature data. If the occurrence probability of the fault meets the first range, the network device obtains the index data of the second optical module.
[0102] For example, the network device determines that the fault type of the first optical module is a dirty fault and the corresponding occurrence probability of the fault is 75% according to the first feature data. The occurrence probability meets the first range: 0-90%. In this case, the network device obtains the index data of the second optical module.
[0103] It can be understood that based on the above setting of the probability range, if the network device determines the fault type of the first optical module and the occurrence probability of the fault type according to the first feature data, and the occurrence probability has exceeded the first range, the network device does not need to obtain the index data of the second optical module. This can save computing resources and improve the detection efficiency.
[0104] In a possible implementation, the network device sets different weights for the first feature data and the second feature data when determining the fault type of the first optical module according to the first feature data and the second feature data.
[0105] The weights can be set in various manners, for example, the farther the second optical module is from the first optical module, the lower the weight of the corresponding feature data.
[0106] In the embodiment, the network device can input the first feature data and the second feature data into a pre-trained fault detection model to obtain a fault detection result output by the model, and then determine the fault type of the first optical module according to the fault detection result, or determine the fault type of the first optical module according to the matching degree between the first feature data and the second feature data and the set feature data corresponding to each fault type. This part is similar to step S303 and will not be described here.
[0107] However, it should be noted that the fault detection model used by the network device to determine the fault type of the first optical module according to the first feature data and the second feature data is different from the fault detection model used by the network device to determine the fault type of the first optical module according to the first feature data. The main difference is the addition of the second feature data. For example, if the network device needs to use a fault detection model to determine the fault type of the first optical module according to the first feature data, the model uses feature data of a single optical module with a known fault as sample data during training. If the network device needs to use a fault detection model to determine the fault type of the first optical module according to the first feature data and the second feature data, the model uses feature data of a single optical module with a known fault and other optical modules on the same optical link as sample data during training.
[0108] Through steps S301-S303, the network device extracts the first feature data from the index data, and then determines the specific fault type of the first optical module according to the first feature data, which significantly improves the efficiency of fault detection, improves the repair efficiency, and reduces the time required for business recovery.
[0109] In the embodiment, the network device needs to determine which feature data or multiple feature data of the optical module in each set dimension. At least the following two possible implementation manners can be used to determine in advance.
[0110] 1) According to human experience and logical inference, the specific feature data under each set dimension is determined.
[0111] Specifically, taking the dirty failure of the optical module as an example, according to experience, there is usually some foreign matter (dirt) at the connection between the optical module and the optical fiber, that is, the dirty failure will cause the transmission of the optical signal to be blocked, and it can be inferred that the temperature of the sending port of the optical module will rise rapidly. Based on such experience and inference, it can be determined that the network device needs to determine whether there are multiple mutation points in the time sequence feature dimension of the temperature data of the optical module.
[0112] Taking the loose failure of the optical module as an example, according to experience, the connection between the optical module and the optical fiber is usually not tight, that is, the loose failure may cause light leakage, and it can be inferred that the receiving power of the optical module is significantly lower than the normal level. Based on such experience and inference, it can be determined that the network device needs to determine whether the change trend slope of the receiving power of the optical module in the time sequence feature dimension is greater than a set value.
[0113] In summary, the specific feature data under each set dimension can be determined in advance according to human experience and logical inference, and the combination and values of multiple specific feature data are written into the network device in advance.
[0114] 2) The specific feature data under each set dimension is determined by using a data-driven method.
[0115] Specifically, there may be some feature data or combinations of multiple feature data that are difficult to predict or explain hidden in the index data of the optical module, and these feature data and combinations may be difficult for relevant personnel to discover. Based on this situation, a certain amount of sample data can be collected in advance and input into a specified machine learning model to assist in determining the specific feature data under each set dimension. The above-mentioned method of using sample data and machine learning models to discover potential feature data and combinations can be referred to as a data-driven method.
[0116] Specifically, the sample data refers to the index data of the optical module within a certain period of time when the optical module has a clear type of failure (such as a dirt or loose failure, etc.), and the specified model refers to, but is not limited to, a gradient boosting decision tree (GBDT) model, an XGBoost model, and a LightGBM model. The aforementioned models can calculate the feature importance of each feature data used to determine the type of failure of the optical module, and output the corresponding ranking results, thereby assisting relevant personnel in discovering potential features, and then screening feature data with high feature importance and filtering out feature data with low feature importance, and performing iterative feature engineering.
[0117] In order to further improve the load balancing of different devices in the fault detection system, the fault detection method provided by the application can also be executed through the following layered architecture fault detection system. In the embodiments of the application, the fault detection system can include one or more first devices and second devices. For ease of distinction, in the following embodiments, the first device is referred to as an end-side device, and the second device is referred to as a cloud-side device.
[0118] As shown in FIG. 7, FIG. 7 is a schematic diagram of the architecture of a layered fault detection system according to an embodiment of the application, which can include a plurality of end-side devices (only two are shown in the figure), a plurality of optical modules (only two are shown in the figure), and a plurality of cloud-side devices (only one is shown in the figure).
[0119] Among them, the first end-side device is connected with the first optical module, the second end-side device is connected with the second optical module, the first optical module and the second optical module are on the same optical link, the first end-side device and the second end-side device can communicate with the cloud-side device, and the first end-side device and the second end-side device can include a data acquisition module, a feature extraction module, a fault detection module, and a reporting module. The cloud-side device at least includes a fault detection module.
[0120] For example, the end-side device (first device) can be a switch, a router, a fiber transponder, a storage device, a communication device, etc. which can directly communicate with the optical module; the cloud-side device (second device) can be a cloud-side analysis server, etc. In some possible embodiments, the computing performance of the cloud-side device is better than that of the end-side device.
[0121] As shown in FIG. 8, FIG. 8 is a flowchart of a fault detection method related to the architecture shown in FIG. 7, which specifically includes:
[0122] S401: The end-side device acquires the index data of the optical module.
[0123] In this step, the data acquisition module in the No. 1 end-side device in FIG. 7 can obtain the index data of the first optical module, or the data acquisition module in the No. 2 end-side device can obtain the index data of the second optical module, wherein the index data includes, but is not limited to, the receiving power, the sending power, the temperature, and the like of the optical module.
[0124] S402: The end-side device obtains the feature data corresponding to the index data.
[0125] In this step, the feature extraction module in the No. 1 end-side device in FIG. 7 can perform feature extraction on the index data obtained in step S401, and then obtain the feature data.
[0126] For example, the feature extraction module in the No. 1 end-side device can extract the number of mutation points and the change trend slope of the receiving power and the sending power of the first optical module in the time sequence feature dimension.
[0127] In some possible embodiments, after obtaining the index data of the optical module, the end-side device can extract a part of the feature data corresponding to the index data, and report the index data to the cloud-side device, and then obtain another part of the feature data of the index data returned by the cloud-side device, wherein the another part of the feature data can be a feature with high calculation complexity. In this way, by cooperation between the end-side device and the cloud-side device, the feature extraction task is shared, and the efficiency of fault detection is further improved.
[0128] Alternatively, in some possible embodiments, after obtaining the index data of the optical module, the end-side device can report the index data to the cloud-side device, and then obtain the feature data of the index data returned by the cloud-side device. In this way, since the end-side device does not need to perform feature extraction on the index data, the calculation resources of the end-side device can be saved.
[0129] Further alternatively, in some possible embodiments, after obtaining the index data of the optical module, the end-side device can extract all the feature data corresponding to the index data, and report the index data and the corresponding feature data to the cloud-side device. In this way, the cloud-side device can obtain the index data of the optical modules in time, without the need to obtain the index data from the end-side device again and perform feature extraction again when fault detection is needed, but can directly perform fault detection according to the feature data of the optical modules, thereby improving the efficiency of fault detection.
[0130] S403: The end-side device performs fault detection according to the feature data.
[0131] In this step, the fault detection module in the No. 1 end-side device in FIG. 7 can perform fault detection according to the feature data.
[0132] Exemplarily, the fault detection module in the No. 1 end-side device inputs the feature data into the pre-trained XGBoost model, and obtains the fault detection result output by the model as: the probability of the dirty fault of the first optical module: 10%, the probability of the loose fault of the first optical module: 70%, and the probability of other types of faults of the first optical module: 5%.
[0133] S404: The end-side device determines whether to report to the cloud-side device according to the result of the fault detection.
[0134] If not, step S405 is performed; if yes, step S406 is performed.
[0135] Exemplarily, based on the fault detection result exemplified in the foregoing step S403, wherein the probability of the loose fault of the first optical module: 70%, meets the set probability range (first range): 50%-80%, in the embodiment of the application, the range can be used not only to indicate whether the index data of the second optical module needs to be collected, but also to indicate whether the cloud-side needs to be reported. The meeting of the range can indicate that the first optical module may have a loose fault, but the fault detection needs to be further determined in combination with the index data of the second optical module, and then the end-side device can report to the cloud-side device, that is, the end-side device performs step S406 to trigger the cloud-side device to perform further fault detection.
[0136] Exemplarily, if the fault detection result is: the probability of the dirty fault of the first optical module: 5%, the probability of the loose fault of the first optical module: 95%, and the probability of other types of faults of the first optical module: 5%, the occurrence probabilities of various faults do not meet the set probability range: 50%-80%, and the cloud-side device does not need to be reported, and then the end-side device performs step S405.
[0137] S405: The end-side device determines the fault type of the optical module.
[0138] In this step, if the fault detection result is: the probability of the dirty fault of the first optical module: 5%, the probability of the loose fault of the first optical module: 95%, and the probability of other types of faults of the first optical module: 5%, wherein, since the probability of the loose fault of the first optical module: 95% is greater than the maximum value in the set probability range: 50%-80%, that is, the probability of the loose fault of the first optical module is high, it can be indicated that the index data of the second optical module does not need to be collected, and it can also be indicated that further fault detection does not need to be performed according to the index data of the two optical modules, but the end-side device can determine that the fault type of the optical module is the loose fault.
[0139] S406: The end-side device sends a trigger signal to the cloud-side device.
[0140] The trigger signal is used to instruct the cloud-side device to perform fault detection on the first optical module.
[0141] In some possible embodiments, when the end-side device sends the trigger signal to the cloud-side device, the end-side device can also send other data, including but not limited to: the index data of the optical module obtained by the end-side device, the feature data corresponding to the index data, and the fault detection result obtained by the end-side device.
[0142] S407: The cloud-side device determines the fault type of the optical module indicated by the trigger signal in response to receiving the trigger signal sent by the end-side device.
[0143] In this step, the cloud-side device can receive the data reported by the plurality of end-side devices, that is, in the embodiment of the present application, the cloud-side device can obtain the feature data corresponding to the index data of the first optical module and the second optical module.
[0144] The trigger signal can include the unique identifier of the first optical module, and then the cloud-side device determines other optical modules (the second optical module) on the optical link where the first optical module is located according to the set link information, and then the cloud-side device obtains the second feature data corresponding to the index data of the second optical module, and determines the fault type of the optical module indicated by the trigger signal according to the first feature data and the second feature data.
[0145] In some possible embodiments, if the cloud-side device receives the trigger signal and also receives the fault detection result sent by the end-side device, for example, the probability of the loosening fault of the first optical module is 75%, then the cloud-side device can preferentially obtain the feature data related to the loosening fault when obtaining the second feature data corresponding to the index data of the second optical module, which can further improve the efficiency of fault detection.
[0146] Through the above steps S401-S407, based on the hierarchical detection architecture as shown in FIG. 7, the fault type of the optical module can be quickly determined on the end-side device, and when necessary, the cloud-side device can perform higher-precision fault detection on the optical module. Through the division of labor and cooperation of the end-side device and the cloud-side device, the efficiency of optical module fault detection is improved, and the time required to restore the service is shortened.
[0147] The above describes the scheme of the embodiments of the present application mainly from the method aspect. It can be understood that, in order to realize the above functions, the fault detection device comprises at least one of the hardware structure and the software module for executing the respective functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be realized in the form of hardware or the combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical scheme. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0148] The embodiments of the present application can divide the functional units of the fault detection device according to the above method examples. For example, each functional unit can be divided according to each function, or two or more functions can be integrated in one processing unit. The integrated unit can be realized in the form of hardware or software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical functional division. In actual implementation, there can be another division manner.
[0149] For example, FIG. 9 shows a structural schematic diagram of a fault detection device 400 provided by an example embodiment of the present application. The fault detection device 400 is applied to a network device. The fault detection device 400 comprises:
[0150] An acquisition module 410 is configured to acquire index data of the first optical module, the index data being used to indicate the running state of the optical module;
[0151] A feature extraction module 420 is configured to perform feature extraction on the index data of the first optical module to obtain first feature data;
[0152] A determination module 430 is configured to determine the fault type of the first optical module according to the first feature data.
[0153] In a possible implementation, the determination module 430 is further configured to:
[0154] acquire the index data of a second optical module; the second optical module is located on the same optical link as the first optical module;
[0155] perform feature extraction on the index data of the second optical module to obtain second feature data;
[0156] determine the fault type of the first optical module according to the first feature data and the second feature data.
[0157] In a possible implementation, the acquisition module 410 is further configured to:
[0158] determine, according to the first feature data, a fault type of the first optical module and a probability of occurrence of the fault type;
[0159] if the probability of occurrence of the fault type meets a first range, acquire the index data of a second optical module.
[0160] In a possible implementation, the determination module 430 is further configured to:
[0161] input the first feature data into a fault detection model to obtain a fault detection result output by the fault detection model, the fault detection model being trained according to sample feature data corresponding to each fault type;
[0162] determine, according to the fault detection result, the fault type of the first optical module.
[0163] In a possible implementation, the determination module 430 is further configured to:
[0164] determine, according to a matching degree between the first feature data and set feature data corresponding to each fault type, the fault type of the first optical module.
[0165] In a possible implementation, the feature extraction module 420 is further configured to:
[0166] extract features of the index data of the first optical module in at least one set dimension to obtain first feature data in the set dimension, the set dimension including a statistical feature dimension, a time series feature dimension and / or a frequency domain feature dimension.
[0167] As a feasible example, the fault detection apparatus 400 provided in the present application is implemented through a software module. For example, the software module can be provided to users through a cloud service subscription mode, and users can select different subscription levels according to needs. For another example, the software module can also provide enterprise-level customized services with professional domain customization, interface personalization and expansion functions according to the needs of users or enterprises.
[0168] In addition, the fault detection apparatus 400 provided in the present application can also be made into a value-added service to provide users, which is not limited in the present application. When the fault detection apparatus 400 is implemented through a software module, the fault detection apparatus 400 can also be embedded into other fault detection software or network management systems.
[0169] In an example embodiment, a computer readable storage medium storing at least one instruction, at least one program, a code set or an instruction set is also provided, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by a processor to implement all or part of the steps of the above exception handling method. For example, the computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk and an optical data storage device, etc.
[0170] In an example embodiment, a computer program product or computer program including computer instructions stored in a computer readable storage medium is also provided. The computer instructions are read by a processor of a computing device from the computer readable storage medium, and the processor executes the computer instructions to cause the computing device to perform all or part of the steps of the method shown in any of the embodiments of FIG. 4.
[0171] In some embodiments, the method shown in the embodiments of the present application can be implemented as computer program instructions encoded in a machine readable format on a computer readable storage medium or on other non-transitory media or articles.
[0172] From the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual applications, the above functions can be completed by different functional modules according to needs, i.e., the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0173] In the several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented by other ways. For example, the device embodiments described above are only schematic, and the division of the modules or units is only a logical function division, and there can be another division way in actual implementation, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0174] The units described as separate components may or may not be physically separate, and the components displayed as units may be a physical unit or multiple physical units, that is, may be located in one place, or also can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0175] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0176] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for making a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0177] The above is only an optional embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method of fault detection for an optical fiber communication system, characterized by, The optical fiber communication system comprises a first optical module, and the method comprises: obtaining index data of the first optical module, the index data being used to indicate an operating state of the optical module; performing feature extraction on the index data of the first optical module to obtain first feature data; determining a fault type of the first optical module according to the first feature data.
2. The method of claim 1, wherein, The determining of the fault type of the first optical module according to the first feature data comprises: obtaining the index data of a second optical module, the second optical module being located on the same optical link as the first optical module; performing feature extraction on the index data of the second optical module to obtain second feature data; determining the fault type of the first optical module according to the first feature data and the second feature data.
3. The method of claim 2, wherein, The obtaining of the index data of the second optical module comprises: determining the fault type of the first optical module and an occurrence probability of the fault type according to the first feature data; if the occurrence probability of the fault type meets a first range, obtaining the index data of the second optical module.
4. The method according to any one of claims 1 to 3, characterized in that, The determining of the fault type of the first optical module according to the first feature data comprises: inputting the first feature data into a fault detection model to obtain a fault detection result output by the fault detection model, the fault detection model being trained according to sample feature data corresponding to each fault type; determining the fault type of the first optical module according to the fault detection result.
5. The method according to any one of claims 1 to 3, characterized in that, The determining of the fault type of the first optical module according to the first feature data comprises: determining the fault type of the first optical module according to a matching degree between the first feature data and set feature data corresponding to each fault type.
6. The method according to any one of claims 1 to 5, characterized in that, The feature extraction on the index data of the first optical module to obtain the first feature data comprises: performing feature extraction on the index data of the first optical module in at least one set dimension to obtain first feature data in the set dimension, the set dimension comprising a statistical feature dimension, a time sequence feature dimension and / or a frequency domain feature dimension.
7. A fault detection apparatus for an optical fiber communication system, characterized by comprising: The apparatus comprises: an obtaining module, configured to obtain index data of the first optical module, the index data being used to indicate an operating state of the optical module; a feature extraction module, configured to perform feature extraction on the index data of the first optical module to obtain first feature data; a determining module, configured to determine a fault type of the first optical module according to the first feature data.
8. The fault detection apparatus of claim 7, wherein, The determining module is further configured to: obtain the index data of a second optical module, the second optical module being located on the same optical link as the first optical module; perform feature extraction on the index data of the second optical module to obtain second feature data; determine the fault type of the first optical module according to the first feature data and the second feature data.
9. The fault detection apparatus of claim 8, wherein, The obtaining module is further configured to: determine the fault type of the first optical module and an occurrence probability of the fault type according to the first feature data; if the occurrence probability of the fault type meets a first range, obtain the index data of the second optical module. If the occurrence probability of the fault meets a first range, the index data of a second optical module is acquired.
10. The fault detection apparatus of any of claims 7-9, wherein, The determination module is further configured to: input the first feature data into a fault detection model to obtain a fault detection result output by the fault detection model, the fault detection model being trained according to sample feature data corresponding to each fault type; determine, according to the fault detection result, a fault type of the first optical module.
11. The fault detection apparatus of any of claims 7-9, wherein, The determination module is further configured to: determine, according to a matching degree between the first feature data and set feature data corresponding to each fault type, the fault type of the first optical module.
12. The fault detection apparatus of any one of claims 7-11, wherein, The feature extraction module is further configured to: extract features of the index data of the first optical module in at least one set dimension to obtain first feature data in the set dimension, the set dimension including a statistical feature dimension, a time sequence feature dimension and / or a frequency domain feature dimension.
13. A fault detection system for an optical fiber communication system, characterized by, The fault detection system comprises: a first device configured to acquire index data of a first optical module, the index data being used to indicate an operating state of the optical module; extract features of the index data of the first optical module to obtain first feature data; determine, according to the first feature data, a fault type of the first optical module and an occurrence probability of the fault type; and if the occurrence probability of the fault meets a first range, send a trigger signal to a second device in the fault detection system, the trigger signal being used to instruct the second device to perform fault detection on the first optical module; the second device is configured to, in response to receiving the trigger signal, acquire the first feature data and index data of a second optical module, the second optical module being located on the same optical link as the first optical module; extract features of the index data of the second optical module to obtain second feature data; and determine, according to the first feature data and the second feature data, the fault type of the first optical module.
14. A method of fault detection for an optical fiber communication system, the method comprising: The method is applied to a first device and a second device, and the optical fiber communication system comprises a first optical module, and the method comprises: the first device acquires index data of the first optical module, the index data being used to indicate an operating state of the optical module; the first device extracts features of the index data of the first optical module to obtain first feature data; the first device determines, according to the first feature data, a fault type of the first optical module and an occurrence probability of the fault type; if the occurrence probability of the fault meets a first range, the first device sends a trigger signal to a second device in the fault detection system, the trigger signal being used to instruct the second device to perform fault detection on the first optical module; the second device, in response to receiving the trigger signal, acquires the first feature data and index data of a second optical module, the second optical module being located on the same optical link as the first optical module; the second device extracts features of the index data of the second optical module to obtain second feature data; and determines, according to the first feature data and the second feature data, the fault type of the first optical module. The second device determines a fault type of the first optical module according to the first feature data and the second feature data.
15. A computing device, comprising: The computing device comprises a processor and a memory; the processor is coupled with the memory; the memory is used to store computer instructions, which are loaded and executed by the processor to enable the computing device to implement the method according to any one of claims 1 to 6.
16. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises computer instructions; when the computer instructions are run in the computing device, the computing device executes the method according to any one of claims 1 to 6.
17. A computer program product, characterised in that, When the computer program product is run in the computing device, the computing device executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
PON fault location method and device
CN112567647A
PON network-based fault diagnosis method, system and device, and storage medium
CN116192245A
Optical module, electronic equipment, communication system and related processing method
CN117176247A
Fault analysis method and apparatus for PON system, and device
WO2024109604A1