A distributed real-time fault monitoring method and system
By deploying monitoring units and intelligent algorithms on a central diagnostic platform in a distributed manner on optical fiber links, the optical power waveform can be monitored in real time and the fault location can be accurately located. This solves the problem that traditional methods cannot capture and accurately locate transient faults in real time, thus improving the reliability and operation and maintenance efficiency of optical networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-06-26
Smart Images

Figure CN121643893B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of optical communication network fault monitoring technology, specifically to a distributed real-time fault monitoring method and system. Background Technology
[0002] With the explosive growth of global data traffic and the arrival of the 5G / 6G era, the reliability and availability of fiber optic communication networks, as the infrastructure carrying information transmission, have become more critical than ever before. However, transient faults in fiber optic links, such as fiber flicker and fiber jitter, have become an increasingly serious challenge facing current optical networks. According to operator reports, these short-term, unexpected transient events account for the vast majority of link failures, incurring significant costs for network operations and severely impacting service continuity and user experience.
[0003] Existing technologies have many limitations in transient fault diagnosis:
[0004] 1. Limitations of Traditional Optical Performance Monitoring (OPM): Most traditional OPM systems are designed to monitor steady-state parameters such as optical power and optical signal-to-noise ratio (OSNR). Their sampling rate and response speed are often insufficient to capture transient power fluctuations at the millisecond or even microsecond level. For sudden and rapidly changing events such as fiber optic flashovers, traditional OPMs often cannot detect them in time or can only record the final result, failing to provide detailed information about the event's process. This makes it difficult to capture complete fault waveform data with high precision in real time at critical moments, hindering the diagnosis of the root cause.
[0005] 2. Limitations of Traditional Optical Time-Domain Reflectometer (OTDR): OTDR is a commonly used tool for fiber optic fault location, capable of accurately detecting physical events such as fiber breaks and loss points. However, it is an offline, intrusive measurement tool, and typically requires a long measurement time to obtain a high-precision link profile. It cannot provide real-time, continuous monitoring capabilities for transient faults, nor can it directly identify the specific type of transient fault. Existing online monitoring methods, even if they can detect faults, often only roughly locate the fault to the optical amplifier section, failing to provide the precise physical location of the fault, resulting in low troubleshooting efficiency.
[0006] Therefore, the industry urgently needs a real-time fault monitoring method that can capture signal waveforms at critical moments, achieve non-intrusive and rapid detection and physical location of millisecond-level transient faults, and effectively handle complex scenarios with multiple faults superimposed, thereby significantly improving the reliability and operation and maintenance efficiency of optical networks. Summary of the Invention
[0007] This application provides a distributed real-time fault monitoring method and system, which can solve the technical problems in the prior art that cannot monitor millisecond-level transient faults in real time and cannot accurately locate the specific physical location of the fault.
[0008] In a first aspect, embodiments of this application provide a distributed real-time fault monitoring method, the distributed real-time fault monitoring method comprising:
[0009] Multiple monitoring units are distributed along the optical fiber link to continuously monitor and collect optical power waveform data. A local intelligent triggering algorithm is used to adaptively determine the trigger based on the instantaneous rate of change of optical power, the power drop amplitude, and dynamic threshold. When a trigger event occurs, the monitoring unit extracts optical power waveform data for a predefined duration before and after the trigger and records the timestamp, generating and reporting event data.
[0010] The central diagnostic platform analyzes event data through a multi-event intelligent decoupling algorithm, identifies and separates multiple fault events that occur simultaneously or overlap, and calculates the location of the fault based on the relative time difference of the fault events on different monitoring units and the propagation delay of the optical signal.
[0011] In conjunction with the first aspect, in one implementation, the local intelligent triggering algorithm includes:
[0012] The optical power waveform data is processed by moving average to obtain smoothed optical power data.
[0013] Maintain smoothed optical power data for a preset time window, and calculate the long-term average and standard deviation of optical power within that preset time window;
[0014] Using the long-term average value as a dynamic baseline, the decrease in optical power, the first-order rate of change, and the second-order rate of change within the preset time window are calculated.
[0015] Determine whether the decrease magnitude, first-order rate of change, and second-order rate of change satisfy the absolute decrease magnitude condition, relative decrease magnitude condition, high-speed rate of change condition, and sudden acceleration condition.
[0016] In conjunction with the first aspect, in one implementation, the dynamic threshold is calculated based on the long-term average and standard deviation.
[0017] In conjunction with the first aspect, in one implementation, the absolute decrease condition refers to the decrease in optical power exceeding a minimum absolute decrease threshold.
[0018] The relative decrease condition refers to the optical power deviating from the dynamic baseline and the deviation exceeding a preset multiple of the standard deviation;
[0019] The high-speed rate of change condition refers to the absolute value of the first-order rate of change exceeding the standard deviation of a preset multiple;
[0020] The sudden acceleration condition refers to the absolute value of the second-order rate of change exceeding the standard deviation of a preset multiple.
[0021] In conjunction with the first aspect, in one implementation, the triggering conditions for fiber breakage or severe fiber bending are either satisfying the absolute descent amplitude condition and the high rate of change condition, or satisfying the relative descent amplitude condition and the sudden acceleration condition.
[0022] The triggering conditions for connector flashover, fiber optic flashover, or poor patch cord contact are to meet the absolute descent amplitude condition and the high rate of change condition, or to meet the relative descent amplitude condition and the sudden acceleration condition.
[0023] The triggering conditions for instantaneous bending of optical fiber, vibration of optical cable or jitter of laser are to meet the sudden acceleration condition and the absolute value of the second-order rate of change exceeds the standard deviation of a preset multiple a specified number of times.
[0024] The triggering conditions for transients caused by channel addition or subtraction operations are the absolute decrease amplitude condition and the high rate of change condition.
[0025] In conjunction with the first aspect, in one implementation, the central diagnostic platform aggregates, preprocesses, and globally aligns multi-source event data, and then analyzes it using a multi-event intelligent decoupling algorithm.
[0026] In conjunction with the first aspect, in one implementation, the multi-event intelligent decoupling algorithm includes:
[0027] Traverse the monitoring unit sequence and identify the power drop start point of the power waveform in each monitoring unit;
[0028] The falling edges identified on different monitoring units are correlated based on the optical signal propagation delay;
[0029] In the current waveform to be processed, find the earliest and largest power drop start point;
[0030] Based on the power drop initiation point, it was determined that the fault occurred between adjacent monitoring units;
[0031] The location of the fault between adjacent monitoring units is calculated based on the time-of-flight principle;
[0032] Calculate the theoretical power drop waveform caused by the fault at all downstream monitoring units, and subtract the theoretical power drop waveform from the current waveform to be processed to generate the residual waveform;
[0033] Add the identified faults to the list of identified faults;
[0034] Repeat the above steps until the unexplained exceptions no longer exist.
[0035] In conjunction with the first aspect, in one implementation, the calculation of the fault location based on the time-of-flight principle includes:
[0036] Based on the time point when the fault event is detected on the adjacent monitoring unit and the time point when the link measurement is received on the adjacent monitoring unit, calculate the link length from the fault point to the upstream and downstream monitoring units.
[0037] In conjunction with the first aspect, in one implementation, the link measurement includes:
[0038] The transceiver periodically sends out link measurement indicators;
[0039] The monitoring unit records the time point at which the link measurement marker is received based on the local time stamp;
[0040] The recorded time point information is uploaded to the central diagnostic platform.
[0041] Secondly, embodiments of this application provide a distributed real-time fault monitoring system, the distributed real-time fault monitoring system comprising:
[0042] Multiple monitoring units are distributed along the optical fiber link to continuously monitor and collect optical power waveform data; they are also used to adaptively determine triggering based on the instantaneous rate of change of optical power, the power drop magnitude, and dynamic thresholds through a local intelligent triggering algorithm; when a triggering event occurs, the monitoring unit extracts optical power waveform data for a predefined duration before and after the trigger and records the timestamp, and generates and reports event data.
[0043] The central diagnostic platform analyzes event data through a multi-event intelligent decoupling algorithm, identifies and separates multiple fault events that occur simultaneously or overlap, and calculates the location of the fault based on the relative time difference of the fault events on different monitoring units and the propagation delay of the optical signal.
[0044] The beneficial effects of the technical solutions provided in this application include:
[0045] By continuously monitoring the optical power waveform through distributed monitoring units, and combining an intelligent triggering algorithm based on instantaneous change rate, power drop magnitude and dynamic threshold adaptation, it can capture millisecond-level transient fault processes in real time, ensuring that key data is not lost, and achieving real-time high-precision detection of millisecond-level transient faults. Compared with traditional OPM systems (which typically have a sampling rate in the second range), the monitoring capability is greatly improved.
[0046] It enables real-time continuous monitoring without interrupting business operations, overcoming the limitations of traditional OTDRs that require business interruption for offline testing, and avoiding business interruptions caused by testing.
[0047] For multiple transient faults that occur simultaneously or overlap, it can identify and isolate the effects of different fault events one by one, solving the problem that traditional fault monitoring systems cannot handle multiple faults and greatly improving fault handling efficiency.
[0048] Utilizing the propagation characteristics of optical signals in optical fibers, when a fault occurs, the fault signal propagates upstream and downstream at the speed of light, and different monitoring units record the arrival time of the fault signal. By calculating the relative time difference between the fault events detected by different monitoring units and combining this with the propagation speed of optical signals in optical fibers, the position of the fault point relative to the monitoring units can be accurately calculated, significantly improving fault handling efficiency. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating an embodiment of the distributed real-time fault monitoring method of this application;
[0050] Figure 2 This is a schematic diagram of a fault location calculation link based on the time-of-flight principle in a specific embodiment of this application;
[0051] Figure 3 This is a schematic diagram of the calculation of the precise location of a fault based on the time-flight principle in a specific embodiment of this application;
[0052] Figure 4 This is a schematic diagram of the architecture of an embodiment of the distributed real-time fault monitoring system of this application. Detailed Implementation
[0053] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0055] In a first aspect, embodiments of this application provide a distributed real-time fault monitoring method.
[0056] In one embodiment, reference is made to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the distributed real-time fault monitoring method of this application. Figure 1 As shown, the distributed real-time fault monitoring method includes:
[0057] Step S1: Distribute multiple monitoring units along the optical fiber link to continuously monitor and collect optical power waveform data;
[0058] Step S2: The trigger is adaptively determined based on the instantaneous rate of change of optical power, the power drop magnitude, and the dynamic threshold through a local intelligent triggering algorithm.
[0059] Step S3: When a trigger event occurs, the monitoring unit extracts optical power waveform data for a predefined duration before and after the trigger and records the timestamp, generating and reporting event data;
[0060] Step S4: The central diagnostic platform analyzes the event data through a multi-event intelligent decoupling algorithm to identify and separate multiple fault events that occur simultaneously or overlap.
[0061] Step S5: Calculate the location of the fault based on the relative time difference of the fault event on different monitoring units and the optical signal propagation delay.
[0062] In this embodiment, multiple ODMUs are deployed along the optical fiber link (e.g., at the optical amplifier) to continuously monitor the optical power waveform in real time at a high sampling rate. Each ODMU integrates an intelligent triggering algorithm based on dynamic threshold adaptation. When a sudden and significant abnormal change occurs in the optical power, the ODMU immediately performs high-precision waveform data capture locally (including waveforms before and after the trigger) and records the timestamp.
[0063] All ODMUs report locally captured event waveform data and timestamps to the central diagnostic platform. The central diagnostic platform aggregates, preprocesses, and globally aligns the received multi-source event data. Then, it analyzes the aggregated waveform data using a multi-event intelligent decoupling algorithm. This algorithm iteratively identifies and isolates the impact of multiple simultaneous or overlapping transient fault events.
[0064] By analyzing the propagation characteristics and relative time differences of the same fault event across different ODMUs, the specific link segment where the fault occurred can be preliminarily identified and located. Specifically, for each decoupled independent fault event, the location of the fault is accurately calculated based on the time-of-flight principle, combining the link segment information where the fault occurred with the timestamps of the fault reported by the ODMUs at both ends of the link segment.
[0065] This application continuously monitors optical power waveforms through a distributed monitoring unit (ODMU). Combined with an intelligent triggering algorithm based on instantaneous rate of change, power drop magnitude, and dynamic threshold adaptation, it can capture millisecond-level transient fault processes in real time, ensuring that critical data is not lost and achieving real-time high-precision detection of millisecond-level transient faults. Compared with traditional OPM systems (which typically have sampling rates on the order of seconds), the monitoring capability is greatly improved.
[0066] This application enables real-time continuous monitoring without interrupting business operations, overcoming the limitations of traditional OTDRs that require business interruption for offline testing, and avoiding business interruptions caused by testing.
[0067] This application addresses the problem of traditional fault monitoring systems being unable to handle multiple transient faults that occur simultaneously or overlap, by identifying and separating the effects of different fault events one by one, thus significantly improving fault handling efficiency.
[0068] This application utilizes the characteristic of optical signals propagating in optical fibers. When a fault occurs, the fault signal propagates upstream and downstream at the speed of light, and different monitoring units record the arrival time of the fault signal. By calculating the relative time difference between the fault events detected by different monitoring units and combining this with the propagation speed of optical signals in optical fibers, the position of the fault point relative to the monitoring units can be accurately calculated, significantly improving fault handling efficiency.
[0069] Furthermore, in one embodiment, the local intelligent triggering algorithm includes:
[0070] The optical power waveform data is processed by moving average to obtain smoothed optical power data.
[0071] Maintain smoothed optical power data for a preset time window, and calculate the long-term average and standard deviation of optical power within that preset time window;
[0072] Using the long-term average value as a dynamic baseline, the decrease in optical power, the first-order rate of change, and the second-order rate of change within the preset time window are calculated.
[0073] Determine whether the decrease magnitude, first-order rate of change, and second-order rate of change satisfy the absolute decrease magnitude condition, relative decrease magnitude condition, high-speed rate of change condition, and sudden acceleration condition.
[0074] In this embodiment, the local fault intelligent triggering algorithm based on dynamic threshold adaptation mainly includes real-time optical power data acquisition and smoothing, dynamic baseline and standard deviation calculation, instantaneous rate of change calculation, and composite triggering condition judgment.
[0075] The real-time optical power data acquisition and smoothing process includes continuously acquiring raw optical power sample values. In the time window Internal Perform a moving average to obtain the smoothed real-time power. The purpose is to filter out high-frequency noise and transient glitches, and reduce false triggering. The moving average calculation formula is shown in the following formula (1):
[0076] (1);
[0077] in, This represents the smoothed optical power data at time t. Indicates a time window. This represents the rate sampling value at time t, which is also the optical power waveform data.
[0078] Dynamic baseline and standard deviation calculations involve maintaining a relatively long time window. History Data. Calculation Inside the window The long-term average value is used to obtain the dynamic baseline. This represents the normal power level of the current link. Dynamic baseline The calculation formula is shown in the following formula (2):
[0079] (2);
[0080] in, This indicates a dynamic baseline. The standard configuration uses a preset time window. 'k' represents the time index.
[0081] in, , For long-term statistical window size, To smooth out window sizes.
[0082] calculate Inside the window Standard deviation This reflects the background noise level of normal power fluctuations. Standard deviation The calculation formula is shown in the following formula (3):
[0083] (3);
[0084] Where, k from t is the time index.
[0085] Subsequently, a dynamic threshold is calculated based on the long-term average and standard deviation, with the aim of enabling the trigger threshold to adaptively adjust according to long-term changes in the network and noise levels.
[0086] The decrease in optical power is expressed as an instantaneous power decrease. This indicates the difference between the current smoothed power and the dynamic baseline, used to detect absolute power drops. Instantaneous power drop. The calculation formula is shown in formula (4) below.
[0087] (4);
[0088] in, .
[0089] instantaneous rate of change That is, the first rate of change, calculated by... The rate of change is obtained and used to capture the speed of power changes. For example, intermittent faults typically exhibit an extremely high rate of change. Instantaneous rate of change The calculation formula is shown in the following formula (5):
[0090] (5);
[0091] in, For weight parameters ( ), usually 0.2 is selected. The larger the value, the faster the response, but the higher the noise sensitivity. A smaller value results in better smoothing, but also increases response latency.
[0092] Instantaneous acceleration That is, the second rate of change, calculated by... The first-order difference is used to capture the suddenness of power changes; extremely fast flashovers will exhibit significant acceleration. Instantaneous acceleration. The calculation formula is shown in the following formula (6):
[0093] (6);
[0094] in, This represents one sampling period.
[0095] The trigger event is determined to have occurred when the magnitude of the descent, the first rate of change, and the second rate of change satisfy one or more of the following conditions: absolute magnitude of descent, relative magnitude of descent, high rate of change, and sudden acceleration.
[0096] Furthermore, in one embodiment, the dynamic threshold is calculated based on the long-term average and standard deviation.
[0097] In this embodiment, based on long-term average values and standard deviation Calculate the dynamic threshold, where the upper threshold is... The lower threshold is , Based on the 3σ principle of normal distribution, it is used to determine whether the power deviates significantly from the dynamic baseline.
[0098] In practical applications, when optical power When the value is below the lower threshold or above the upper threshold, trigger condition B (relative decrease amplitude condition) is met, that is: Furthermore, the dynamic threshold automatically adjusts as network traffic changes without manual intervention. For example, when network traffic increases, the standard deviation... As the network traffic increases, the range of the dynamic threshold also expands; when the network traffic decreases, the standard deviation... As the threshold decreases, the range of the dynamic threshold also shrinks. The method for calculating the dynamic threshold enables the system to adapt to changes in the network environment, effectively reducing the false alarm rate.
[0099] Furthermore, in one embodiment, the absolute decrease condition refers to the optical power decrease value exceeding the minimum absolute decrease threshold;
[0100] The relative decrease condition refers to the optical power deviating from the dynamic baseline and the deviation exceeding a preset multiple of the standard deviation;
[0101] The high-speed rate of change condition refers to the absolute value of the first-order rate of change exceeding the standard deviation of a preset multiple;
[0102] The sudden acceleration condition refers to the absolute value of the second-order rate of change exceeding the standard deviation of a preset multiple.
[0103] The triggering conditions for fiber breakage or severe fiber bending are to satisfy the absolute descent amplitude condition and the high rate of change condition, or to satisfy the relative descent amplitude condition and the sudden acceleration condition.
[0104] The triggering conditions for connector flashover, fiber optic flashover, or poor patch cord contact are to meet the absolute descent amplitude condition and the high rate of change condition, or to meet the relative descent amplitude condition and the sudden acceleration condition.
[0105] The triggering conditions for instantaneous bending of optical fiber, vibration of optical cable or jitter of laser are to meet the sudden acceleration condition and the absolute value of the second-order rate of change exceeds the standard deviation of a preset multiple a specified number of times.
[0106] The triggering conditions for transients caused by channel addition or subtraction operations are the absolute decrease amplitude condition and the high rate of change condition.
[0107] In this embodiment, the trigger signal is generated based on at least one or more of the following combined conditions being met:
[0108] Condition A (Absolute decrease): Current power decrease value Exceeded the minimum absolute descent threshold ,Right now: .
[0109] Condition B (Relative Decrease): The current power deviates significantly from the dynamic baseline, exceeding its normal fluctuation range, i.e.: Among them, the general choice is Based on the 3σ principle of normal distribution.
[0110] Condition C (Rapid Rate of Change): The instantaneous rate of decrease in power exceeds the dynamic threshold, i.e.: ,in, for The dynamic standard deviation, considering first-order derivative noise amplification, Usually more Large, generally choose .
[0111] Condition D (Sudden Acceleration): The power change suddenly accelerates, that is: ,in, for The dynamic standard deviation, considering the further amplification of second-order derivative noise, is generally chosen. 。
[0112] Different fault types have different triggering conditions. Specifically, when an optical fiber breaks or is severely bent, the optical power drops to an extremely low level in a very short time and remains in this low power state. The triggering conditions are either condition A and condition C, or condition B and condition D.
[0113] When a connector breaks, an optical fiber breaks, or a patch cord has poor contact, the optical power drops sharply for a moment, and then usually recovers quickly to near or completely the original level within a short time. This is manifested as a brief negative pulse, triggered by either condition A and condition C, or condition B and condition D.
[0114] When an optical fiber bends instantaneously, the optical cable vibrates, or the laser jitters, the optical power may fluctuate rapidly, but not necessarily very significantly, sometimes exhibiting oscillations or multiple consecutive pulses. The triggering condition is condition D and... The absolute value frequently exceeds .
[0115] Transient events caused by channel addition or subtraction operations refer to brief overshoots or undershoots in the total optical power caused by the transient process of optical power control or gain adjustment during the addition or removal of wavelength channels. These overshoots typically manifest as brief unidirectional pulses or step jumps. The triggering conditions are condition A and condition C.
[0116] Furthermore, in one embodiment, the central diagnostic platform aggregates, preprocesses, and globally aligns multi-source event data, and then analyzes it using a multi-event intelligent decoupling algorithm.
[0117] In this embodiment, once the triggering condition is met, the ODMU immediately generates a trigger signal, records the current time information, and captures raw power waveform data containing a predefined duration before and after the trigger moment, such as 50 milliseconds before the trigger to 50 milliseconds after the trigger. A trigger suppression timer is started, during which no new triggers are accepted to prevent the same fault from causing continuous false alarms.
[0118] Furthermore, in one embodiment, the multi-event intelligent decoupling algorithm includes:
[0119] Traverse the monitoring unit sequence and identify the power drop start point of the power waveform in each monitoring unit;
[0120] The falling edges identified on different monitoring units are correlated based on the optical signal propagation delay;
[0121] In the current waveform to be processed, find the earliest and largest power drop start point;
[0122] Based on the power drop initiation point, it was determined that the fault occurred between adjacent monitoring units;
[0123] The location of the fault between adjacent monitoring units is calculated based on the time-of-flight principle;
[0124] Calculate the theoretical power drop waveform caused by the fault at all downstream monitoring units, and subtract the theoretical power drop waveform from the current waveform to be processed to generate the residual waveform;
[0125] Add the identified faults to the list of identified faults;
[0126] Repeat the above steps until the unexplained exceptions no longer exist.
[0127] In this embodiment, for scenarios where multiple faults may occur simultaneously at multiple locations, the unidirectional propagation characteristics of events in optical fibers are utilized, combined with strong consistency checks in time and amplitude, and the identified fault events are gradually isolated through iterative and residual analysis methods.
[0128] The multi-event intelligent decoupling algorithm mainly includes fault event falling edge detection and candidate fault generation, as well as iterative fault decoupling and localization. Fault event falling edge detection includes traversing the ODMU sequence, falling edge identification, and candidate event association. Iterative fault decoupling and localization includes initialization and iterative iteration.
[0129] The step of traversing the monitoring unit sequence includes starting with the first ODMU upstream of the link and checking the power waveform of each ODMU downstream in sequence.
[0130] The falling edge identification step involves identifying all significant power drop start points (falling edges) in the waveform for each ODMU using methods similar to those in the local triggering algorithm (i.e., instantaneous rate of change, acceleration, amplitude drop, etc.).
[0131] The candidate event association step includes performing an initial association of these falling edges identified on different ODMUs. For example, if exist A falling edge was detected. exist A similar falling edge was detected, and Roughly consistent with optical signals from arrive If the propagation delay is significant, then they likely belong to the same upstream fault.
[0132] The initialization steps include using all raw captured waveforms from the ODMU as the current waveforms to be processed and creating an empty "list of identified faults".
[0133] The iterative steps include:
[0134] a) Find the most significant unexplained anomaly: In the current waveform to be processed, iterate through all ODMUs, looking for the earliest occurring power drop event with the most significant amplitude (or the largest rate of change). This usually indicates the most upstream or most serious unidentified fault. If no new significant anomaly is found, exit the loop.
[0135] b) Preliminary location of the fault:
[0136] Fault segment identification: Determine which ODMU the anomaly originated from (e.g., The number of cases began to decline significantly, while its direct upstream... The waveform remains normal (or experiences only minor fluctuations) during the same period. This will provide an initial indication that the fault occurred... and between.
[0137] Event propagation verification: Check whether the anomaly has propagated to all downstream ODMUs with the required propagation delay.
[0138] c) Calculation of precise fault location: Extraction of upstream monitoring unit Based on the fault detection time, and using a fault precise location calculation method based on the time-of-flight principle, the location of the fault is calculated. and The precise location between them. The fault calculation formula is shown in the following formula (7):
[0139] (7);
[0140] in, This is the speed of light in an optical fiber (approximately 200,000 km / s). and The fault events are respectively in and The moment it was detected and The link measurement is marked in and The moment it is received.
[0141] d) Fault Impact Compensation: Based on the identified fault type (e.g., fiber breakage, fiber flashover, etc.) and its location, calculate the theoretical power drop waveforms caused by the fault at all downstream ODMUs. Subtract these theoretical power drop waveforms from the original current waveform to be processed to generate a new set of residual waveforms.
[0142] e) Add the identified faults to the list of identified faults.
[0143] f) Return to step a) and continue searching for the next fault in the residual waveform.
[0144] Through the aforementioned iterative process, the system can identify and separate multiple fault events that occur simultaneously or overlap, achieving precise decoupling and localization of complex fault scenarios. In actual testing, this algorithm can accurately identify and locate two faults in scenarios where two faults occur simultaneously, significantly improving decoupling accuracy compared to traditional methods.
[0145] Furthermore, in one embodiment, the calculation of the fault location based on the time-of-flight principle includes: calculating the link length from the fault point to the upstream and downstream monitoring units based on the time point at which the fault event is detected on the adjacent monitoring unit and the time point at which the link measurement marker is received on the adjacent monitoring unit.
[0146] The link measurement includes: the transceiver periodically sending link measurement indicators; the monitoring unit recording the time point when the link measurement indicator is received based on the local time stamp; and uploading the recorded time point information to the central diagnostic platform.
[0147] In this embodiment, as Figure 2 As shown, for a bidirectional link, let the local timestamp of node A be . ,node The local timescale is , The local timescale is , The local timescale is .
[0148] Assumption The timescale deviation relative to A is ,have .
[0149] Assumption The timescale deviation relative to A is ,have .
[0150] Assumption relatively The timescale deviation is ,have .
[0151] The reverse process applies relative to node B. The steps for calculating the fault location are as follows:
[0152] Link measurement steps. Specifically, the transceiver periodically sends link measurement markers. When sending a measurement marker, node A records the marker based on its local time stamp and sends a time stamp TAS to node B using a wavelength tag signal; node B records the marker based on its local time stamp and sends a time stamp TBS to node A using a wavelength tag signal. Distributed monitoring units along the route. Upon receiving the measurement marker sent by node A, and recording the received time marker based on the local time stamp, Similarly, for the measurement marker received from node B, the received time marker is also recorded based on the local time stamp. The recorded time stamps are stored and uploaded to the central diagnostic platform. This time stamp information will be updated with new information in the next cycle.
[0153] Link failure event time-stamping measurement steps. Specifically, when a distributed monitoring node detects a fault based on a local intelligent fault triggering algorithm, it records the local timestamp information at the time of triggering and uploads it to the central diagnostic platform. For example, The timestamp recorded when the fault is detected is , The timestamp recorded when the fault is detected is .
[0154] Steps for calculating the precise location of the fault. Specifically, as follows: Figure 3 As shown, the fault point is at location X, and the node is calculated. Link length to fault point X or nodes Link length to fault point X The delay corresponding to each link segment is expressed as follows: , Then there is Assume the link failure occurs relative to the node. The time scale is The relative timing of the link failure The time scale is The relationship between the two is .
[0155] Depend on , We can obtain:
[0156] .
[0157] This formula, after being rearranged, yields: .
[0158] Because the data from the link measurements satisfy the following relationship: Substituting the values, we get:
[0159] .
[0160] Fault point X to node The formula for calculating the link length is:
[0161] ;
[0162] Where v is the speed at which light travels in the optical fiber link.
[0163] Similarly, the distance from fault point X to node The formula for calculating the link length is:
[0164] .
[0165] Using the fault location calculation method based on the time-of-flight principle described above, the system can accurately locate the fault to a specific physical location without interrupting services and achieve high-precision positioning.
[0166] In one specific embodiment, at a certain moment, two transient failure events occurred in the fiber optic link:
[0167] Event A (connector flashover): Occurred at 30km, causing a sudden drop in optical power within a very short time and a rapid recovery (approximately 200 µs), with a maximum drop of approximately 5 dB.
[0168] Event B (instantaneous fiber bending): occurred at 60km, causing a small fluctuation in optical power lasting about 1ms, with a maximum drop of about 1.5dB.
[0169] These two events occurred almost simultaneously and may have superimposed at the downstream ODMU. Assume that before the events occurred, the normal optical power of the link measured by all ODMUs was... Furthermore, the link is stable, and the standard deviation of background noise is extremely small. On the forward path, , and The triggering process is as follows:
[0170] Event A (at 30km) not reached , Only normal link power was detected, and its dynamic baseline was maintained. Stable at The value is around 10.00 dBm, and all rates of change and accelerations are within the normal range. It will not be triggered.
[0171] Located at 40km, it can detect the effects of event A from 30km in the forward direction and the effects of event B from 60km in the reverse direction.
[0172] Dynamic baseline calculation: The average smoothed optical power within a 20ms window is continuously calculated as... and standard deviation Assuming that before the event occurred, its , (Represents background noise).
[0173] Dynamic threshold calculation:
[0174] Power threshold: .
[0175] Rate of change threshold: assuming its current Standard deviation ,but .
[0176] Acceleration threshold: Assuming its current Standard deviation ,but .
[0177] Event A (connector momentary disconnection) arrives :
[0178] At some point , A sharp drop in optical power was detected. Assuming the smoothed power... Within 100µs from The signal strength dropped sharply from 10.00 dBm to 15.00dBm.
[0179] Condition A (Absolute amplitude deviation): .satisfy.
[0180] Condition C (Rapid Rate of Change): Instantaneous Rate of Change .satisfy.
[0181] Triggering result: Since both A and C are satisfied, Trigger immediately and begin caching high-precision power waveform data before and after the trigger. Record the trigger time.
[0182] Located at 80km, the effect will be the superposition of these two events. Assuming the speed of light signal propagation... Event A takes time to travel from 30km to 80km (80 30)km / (2×10⁵km / s)=250µs. The time required for event B to propagate from 60km to 80km is (80)km / (2×10⁵km / s)=250µs. 60)km / (2×10⁵km / s)=100µs. Therefore The monitored waveform will be the result of the superposition of events A and B. However, if the superimposed waveform meets the triggering conditions, ODMU-3 will be triggered and the waveform data will be cached. At this time, the trigger time will also be recorded.
[0183] The waveform in the image contains only event A. The waveform includes events A and B. Based on... The waveform can be used to predict the impact of event A. The effect is that the relative amplitude deviation is 5dB. By subtracting the influence of event A from the waveform, events A and B can be separated.
[0184] Similarly, on the reverse path, The superimposed waveforms of events A and B will be detected. The waveform of event B will be detected. No event will be reported if only normal power is detected. Therefore, in summary, and There is an event in the middle, and There is an event.
[0185] For event A, precise comparison The timescale of the waveform change detected in the positive direction, i.e., the timescale of the event occurrence. ,and The timescale of the waveform change caused by event A, obtained from the reverse-detected stripping event B. Based on periodic link measurements, it can be known that... relatively Time scale deviation is Assuming , Then the location of event A relative to... The distance is
[0186] .
[0187] Secondly, embodiments of this application also provide a distributed real-time fault monitoring system.
[0188] In one embodiment, reference is made to Figure 4 , Figure 4 This is a schematic diagram of the architecture of an embodiment of the distributed real-time fault monitoring device of this application. Figure 4 As shown, the distributed real-time fault monitoring system includes:
[0189] Multiple monitoring units are distributed along the optical fiber link to continuously monitor and collect optical power waveform data; they are also used to adaptively determine triggering based on the instantaneous rate of change of optical power, the power drop magnitude, and dynamic thresholds through a local intelligent triggering algorithm; when a triggering event occurs, the monitoring unit extracts optical power waveform data for a predefined duration before and after the trigger and records the timestamp, and generates and reports event data.
[0190] The central diagnostic platform analyzes event data through a multi-event intelligent decoupling algorithm, identifies and separates multiple fault events that occur simultaneously or overlap, and calculates the location of the fault based on the relative time difference of the fault events on different monitoring units and the propagation delay of the optical signal.
[0191] In this embodiment, multiple ODMUs are deployed along the optical fiber link (e.g., at the optical amplifier) to continuously monitor the optical power waveform in real time at a high sampling rate. Each ODMU integrates an intelligent triggering algorithm based on dynamic threshold adaptation. When a sudden and significant abnormal change occurs in the optical power, the ODMU immediately performs high-precision waveform data capture locally (including waveforms before and after the trigger) and records the timestamp.
[0192] All ODMUs report locally captured event waveform data and timestamps to the central diagnostic platform. The central diagnostic platform aggregates, preprocesses, and globally aligns the received multi-source event data. Then, it analyzes the aggregated waveform data using a multi-event intelligent decoupling algorithm. This algorithm iteratively identifies and isolates the impact of multiple simultaneous or overlapping transient fault events.
[0193] By analyzing the propagation characteristics and relative time differences of the same fault event across different ODMUs, the specific link segment where the fault occurred can be preliminarily identified and located. Specifically, for each decoupled independent fault event, the location of the fault is accurately calculated based on the time-of-flight principle, combining the link segment information where the fault occurred with the timestamps of the fault reported by the ODMUs at both ends of the link segment.
[0194] This application continuously monitors optical power waveforms through a distributed monitoring unit (ODMU). Combined with an intelligent triggering algorithm based on instantaneous rate of change, power drop magnitude, and dynamic threshold adaptation, it can capture millisecond-level transient fault processes in real time, ensuring that critical data is not lost and achieving real-time high-precision detection of millisecond-level transient faults. Compared with traditional OPM systems (which typically have sampling rates on the order of seconds), the monitoring capability is greatly improved.
[0195] This application enables real-time continuous monitoring without interrupting business operations, overcoming the limitations of traditional OTDRs that require business interruption for offline testing, and avoiding business interruptions caused by testing.
[0196] This application addresses the problem of traditional fault monitoring systems being unable to handle multiple transient faults that occur simultaneously or overlap, by identifying and separating the effects of different fault events one by one, thus significantly improving fault handling efficiency.
[0197] This application utilizes the characteristic of optical signals propagating in optical fibers. When a fault occurs, the fault signal propagates upstream and downstream at the speed of light, and different monitoring units record the arrival time of the fault signal. By calculating the relative time difference between the fault events detected by different monitoring units and combining this with the propagation speed of optical signals in optical fibers, the position of the fault point relative to the monitoring units can be accurately calculated, significantly improving fault handling efficiency.
[0198] In one specific embodiment, for a bidirectional shared-cable link, optical modules with wavelength tagging functionality are used in the transceivers at both ends. Distributed ODMUs are deployed at each optical amplifier node on the link. The ODMUs are connected to a central diagnostic platform, uploading locally collected data to the central diagnostic platform for centralized processing to achieve fault location.
[0199] The ODMU contains three modules: an optical power monitoring module, a local event triggering and data capture module, and a data transmission interface. The optical power monitoring module acquires real-time, continuous waveform data of optical power changes over time based on wavelength tags. The local event triggering and data capture module monitors instantaneous changes in optical power in real time. When optical power changes rapidly, such as a sudden drop exceeding a threshold, it immediately triggers a function, saves the original power waveform data for a period before and after the fault occurrence to a cache, and records the current time. The data transmission interface quickly reports the trigger event, timestamp, and captured optical power waveform data to the central diagnostic platform.
[0200] The central diagnostic platform comprises two modules: a time correlation and analysis module and a fault location module. The time correlation and analysis module identifies and separates multiple fault events occurring simultaneously. The fault location module calculates the precise location of the fault by comparing the detection time difference between upstream and downstream ODMUs at the fault location and considering the speed of light propagation in the optical fiber.
[0201] The functions of each module in the above-mentioned distributed real-time fault monitoring system correspond to the steps in the above-mentioned distributed real-time fault monitoring method embodiment, and their functions and implementation processes will not be described in detail here.
[0202] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0203] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.
[0204] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.
[0205] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.
[0206] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.
[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.
[0208] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A distributed real-time fault monitoring method, characterized in that, The distributed real-time fault monitoring method includes: Multiple monitoring units are deployed in a distributed manner along the optical fiber link to continuously monitor and collect optical power waveform data; The system uses a local intelligent triggering algorithm to determine whether preset triggering conditions are met. This includes performing a moving average on the optical power waveform data to obtain smoothed optical power data; calculating the long-term average value of optical power within a preset time window based on the smoothed optical power data; using the long-term average value as a dynamic baseline to calculate and determine whether the instantaneous rate of change of optical power and the power decrease magnitude meet preset triggering conditions; when the triggering conditions are met, the monitoring unit extracts optical power waveform data for a predefined duration before and after the trigger and records the timestamp, generating and reporting event data. The central diagnostic platform determines the fault location through a multi-event intelligent decoupling algorithm, including analyzing event data, traversing the monitoring unit sequence, identifying the power drop start point of the power waveform in each monitoring unit, associating the drop edges identified on different monitoring units based on the optical signal propagation delay, finding the earliest and largest power drop start point in the current waveform to be processed, and determining that the fault occurs between adjacent monitoring units based on the power drop start point.
2. The distributed real-time fault monitoring method as described in claim 1, characterized in that, The local intelligent triggering algorithm also includes: Maintain smoothed optical power data for a preset time window and calculate the standard deviation of optical power within that preset time window; Using the long-term average value as a dynamic baseline, the decrease in optical power, the first-order rate of change, and the second-order rate of change within the preset time window are calculated. The dynamic threshold is used to determine whether the decrease amplitude, first-order rate of change, and second-order rate of change satisfy the absolute decrease amplitude condition, relative decrease amplitude condition, high-speed rate of change condition, and sudden acceleration condition.
3. The distributed real-time fault monitoring method as described in claim 2, characterized in that, The dynamic threshold is calculated based on the long-term average and standard deviation.
4. The distributed real-time fault monitoring method as described in claim 2, characterized in that, The absolute decrease condition refers to the optical power decrease value exceeding the minimum absolute decrease threshold. The relative decrease condition refers to the optical power deviating from the dynamic baseline and the deviation exceeding a preset multiple of the standard deviation; The high-speed rate of change condition refers to the absolute value of the first-order rate of change exceeding the standard deviation of a preset multiple; The sudden acceleration condition refers to the absolute value of the second-order rate of change exceeding the standard deviation of a preset multiple.
5. The distributed real-time fault monitoring method as described in claim 4, characterized in that, The triggering conditions for fiber breakage or severe fiber bending are to satisfy the absolute descent amplitude condition and the high rate of change condition, or to satisfy the relative descent amplitude condition and the sudden acceleration condition. The triggering conditions for connector flashover, fiber optic flashover, or poor patch cord contact are to meet the absolute descent amplitude condition and the high rate of change condition, or to meet the relative descent amplitude condition and the sudden acceleration condition. The triggering conditions for instantaneous bending of optical fiber, vibration of optical cable or jitter of laser are to meet the sudden acceleration condition and the absolute value of the second-order rate of change exceeds the standard deviation of a preset multiple a specified number of times. The triggering conditions for transients caused by channel addition or subtraction operations are the absolute decrease amplitude condition and the high rate of change condition.
6. The distributed real-time fault monitoring method as described in claim 1, characterized in that, The central diagnostic platform aggregates, preprocesses, and aligns multi-source event data globally in time, and then analyzes it using a multi-event intelligent decoupling algorithm.
7. The distributed real-time fault monitoring method as described in claim 1, characterized in that, The multi-event intelligent decoupling algorithm also includes: After calculating the location of the fault between adjacent monitoring units based on the time-of-flight principle; Calculate the theoretical power drop waveform caused by the fault at all downstream monitoring units, and subtract the theoretical power drop waveform from the current waveform to be processed to generate the residual waveform; Add the identified faults to the list of identified faults; Repeat the above steps until the unexplained exceptions no longer exist.
8. The distributed real-time fault monitoring method as described in claim 7, characterized in that, The calculation of the fault location based on the time-of-flight principle includes: Based on the time point when the fault event is detected on the adjacent monitoring unit and the time point when the link measurement is received on the adjacent monitoring unit, calculate the link length from the fault point to the upstream and downstream monitoring units.
9. The distributed real-time fault monitoring method as described in claim 8, characterized in that, The link measurement includes: The transceiver periodically sends out link measurement indicators; The monitoring unit records the time point at which the link measurement marker is received based on the local time stamp; The recorded time point information is uploaded to the central diagnostic platform.
10. A distributed real-time fault monitoring system, characterized in that, The distributed real-time fault monitoring system includes: Multiple monitoring units are distributed along the optical fiber link to continuously monitor and collect optical power waveform data. They are also used to determine whether preset triggering conditions are met through a local intelligent triggering algorithm. This includes performing a moving average on the optical power waveform data to obtain smoothed optical power data; calculating the long-term average value of optical power within a preset time window based on the smoothed optical power data; using the long-term average value as a dynamic baseline to calculate and determine whether the instantaneous rate of change of optical power and the power drop magnitude meet the preset triggering conditions; when the triggering conditions are met, the monitoring unit extracts optical power waveform data for a predefined duration before and after the trigger and records the timestamp, generating and reporting event data. The central diagnostic platform determines the fault location through a multi-event intelligent decoupling algorithm, including analyzing event data, traversing the monitoring unit sequence, identifying the power drop start point of the power waveform in each monitoring unit, associating the drop edges identified on different monitoring units based on the optical signal propagation delay, finding the earliest and largest power drop start point in the current waveform to be processed, and determining that the fault occurs between adjacent monitoring units based on the power drop start point.
Citation Information
Patent Citations
Online and intelligent optical cable monitoring and fault positioning system based on GIS platform
CN106788696A
Optical network intelligent operation and maintenance method and device based on knowledge graph and digital twinning
CN118193751A