A method, system, device and medium for fault detection of PON gateway

By constructing a PON gateway status data sequence and a dynamic window strategy, the problems of high false alarm rate and insufficient identification of related faults in existing PON gateway fault detection technologies are solved, and accurate detection and timely alarm of PON gateway faults are achieved.

CN120786213BActive Publication Date: 2025-12-02SICHUAN TIANYI COMHEART TELECOM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511290198.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-12-02
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing PON gateway fault detection methods suffer from high false alarm rates, inability to capture multi-parameter correlated fault characteristics, and neglect of the temporal correlation of fault characteristics, resulting in insufficient detection accuracy and timeliness.

Method used

By acquiring the status data of the PON gateway, constructing the first and second data sequences, calculating the sub-fluctuation index, generating reference time nodes and windows, and combining fuzzy quantity correction and dynamic window expansion strategies, the system intelligently determines the probability of failure and the detection strategy.

Benefits of technology

Accurately capture transient anomalies in key parameters, reduce false alarm rates, improve the ability to identify complex faults, and ensure real-time alarms for serious faults and early prediction of potential faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120786213B_ABST
    Figure CN120786213B_ABST
Patent Text Reader

Abstract

This invention discloses a PON gateway fault detection method, system, device, and medium, relating to the field of data processing technology. The method includes: acquiring a PON gateway status dataset; acquiring a first data type and a second data type; acquiring first state data and second state data; forming a first data sequence and a second data sequence; acquiring a first sub-fluctuation index; acquiring a second sub-fluctuation index; acquiring a reference time node; acquiring a fuzzy quantity n; acquiring n consecutive acquisition time nodes that include the reference time node as a reference window; determining whether the largest second sub-fluctuation index is located within the reference window; if so, obtaining the fault probability based on the positions of the first state data, second state data, and second state data of the reference time node within the reference window; if not, obtaining a fault detection strategy based on the first state data of the reference time node and a safety threshold. This invention has the advantages of high fault detection efficiency, low false alarm rate, and low false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a PON gateway fault detection method, system, device, and medium. Background Technology

[0002] PON (Passive Optical Network) is a fiber-optic communication technology. Its core feature is that it uses passive optical splitters to provide services to multiple end users through a single optical fiber. In the operation and maintenance management of PON gateways, the accuracy and timeliness of fault detection directly affect user experience and network stability. However, existing fault detection methods have three main drawbacks: First, they rely on static threshold criteria, such as triggering an alarm only when the received optical power exceeds a fixed safety threshold. However, network conditions are dynamic and fluctuating; a momentary exceedance may be a brief interference rather than a real fault, leading to a high false alarm rate. Second, single-dimensional parameter analysis is insufficient. For example, when monitoring only optical power or bit error rate, it is impossible to capture the multi-parameter correlated fault characteristics. When key parameters (such as received optical power) are not analyzed, the problem becomes more prominent. While both the optical power and bit error rate are within the safe threshold range, existing methods cannot capture such correlated anomalies when coordinated abnormal fluctuations occur (e.g., a sudden change in optical power and a sudden increase in bit error rate occur at similar times). Although these latent faults do not trigger alarms for single indicators exceeding limits, they indicate potential risks of equipment performance degradation. Thirdly, the temporal correlation of fault characteristics is ignored. Abnormal fluctuations of different parameters may have a sequential order or time window correlation (e.g., a delay after an abnormal optical power leads to a surge in packet loss rate). However, existing methods only sample independently at each time point and do not establish a correlation between fluctuations across time nodes. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a PON gateway fault detection method, system, device, and medium.

[0004] A PON gateway fault detection method includes: acquiring a state dataset of the PON gateway, and acquiring a first data type and a second data type based on the state dataset; acquiring first state data corresponding to the first data type and second state data corresponding to the second data type at each collection time node in the previous detection cycle within the current detection cycle; forming a first data sequence based on multiple first state data and forming a second data sequence based on multiple second state data; acquiring a first sub-fluctuation index for each first state data based on the first data sequence, and acquiring a second sub-fluctuation index for each second state data based on the second data sequence; acquiring the collection time node corresponding to the largest first sub-fluctuation index and using it as a reference time node; acquiring a fuzzy number n; acquiring n consecutive collection time nodes that include the reference time node and using them as a reference window; and determining whether the largest second sub-fluctuation index is located within the reference window; if so, acquiring a fault probability based on the positions of the first state data, second state data, and second state data of the reference time node within the reference window; if not, acquiring a fault detection strategy based on the first state data of the reference time node and a security threshold.

[0005] Optionally, the fault detection strategy obtained based on the first state data of the reference time node and the safety threshold includes: if the first state data of the reference time node exceeds the safety threshold, the fuzzy quantity n is corrected and a new fuzzy quantity n' is obtained; a new reference window is obtained based on the new fuzzy quantity n', and it is determined whether the largest second sub-fluctuation index is located within the new reference window.

[0006] Optionally, the fuzzy quantity n is corrected and a new fuzzy quantity n' is obtained, which is expressed as: ;in, For the new fuzzy quantity, This represents the number of time points collected within the detection period. For the first state data at the reference time point, The safety threshold for the first state data. For fuzzy quantities.

[0007] Optionally, the number of fuzzy elements n is represented as: ;in, For fuzzy quantities, This represents the number of time points collected within the detection period. This is the splitting coefficient.

[0008] Optionally, the first sub-fluctuation index obtained from each first state data based on the first data sequence is represented as follows: ;in, This is the first sub-fluctuation index for the i-th first state data. This represents the first state data at the i-th data collection time point. This represents the first state data at the (i-1)th acquisition time node. This represents the first state data at the (i+1)th data collection time point. This represents the number of time points collected within the detection period.

[0009] Optionally, the fault probability is obtained based on the positions of the first state data, second state data, and second state data within the reference window at the reference time point, and is expressed as follows: ;in, This represents the probability of failure. To reduce the factor, This represents the number of nodes between the data collection time node and the reference time node corresponding to the largest second-order sub-fluctuation indicator. The safety threshold for the first state data. For the first state data at the reference time point, The safety threshold for the second state data. This is the second state data corresponding to the largest second sub-fluctuation indicator.

[0010] A PON gateway fault detection system is also provided. The system includes: an acquisition module, used to acquire a state dataset of the PON gateway, and acquire a first data type and a second data type based on the state dataset; acquire first state data corresponding to the first data type and second state data corresponding to the second data type at each collection time point within the previous detection cycle; form a first data sequence based on multiple first state data; and form a second data sequence based on multiple second state data; and a data processing module, used to acquire a first sub-fluctuation index for each first state data based on the first data sequence, and acquire a second sub-fluctuation index for each second state data based on the second data sequence. The system employs several modules: a first sampling time node corresponding to the largest first sub-fluctuation indicator and a second detection module; a judgment module to obtain a fuzzy quantity n, n consecutive sampling time nodes containing the reference time node and using them as a reference window, and to determine whether the largest second sub-fluctuation indicator is within the reference window; a first detection module to obtain the fault probability based on the positions of the first state data, second state data, and second state data of the reference time node within the reference window when the largest second sub-fluctuation indicator is within the reference window; and a second detection module to obtain a fault detection strategy based on the first state data of the reference time node and a safety threshold when the largest second sub-fluctuation indicator is not within the reference window.

[0011] Optionally, the second detection module is further configured to: if the first state data of the reference time node exceeds a preset threshold, correct the fuzzy quantity n and obtain a new fuzzy quantity n'; obtain a new reference window based on the new fuzzy quantity n', and determine whether the largest second sub-fluctuation index is located within the new reference window.

[0012] An electronic device is also provided, comprising: a memory storing a computer program thereon; and a processor for executing the computer program in the memory to implement the above-described PON gateway fault detection method.

[0013] A non-transitory computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the above-described PON gateway fault detection method.

[0014] The beneficial effects of this invention are reflected in:

[0015] In the entire PON gateway fault detection method, firstly, cross-node fluctuation quantization technology is used to accurately capture transient anomalies of key parameters (such as a cliff-like drop in optical power), avoiding the omission of normal abnormal signals by existing sampling methods. Secondly, based on the fault propagation timing model, the extreme fluctuation points of the first parameter are used as time anchors to dynamically generate a correlation window covering typical propagation delays (such as a reasonable time period from millisecond-level optical signal degradation to increased bit error rate), intelligently determining whether the extreme fluctuations of the second parameter have temporal coupling. Thirdly, when the fluctuations of the two parameters are spatiotemporally correlated, multi-dimensional safety margin fusion is used, combining the timing delay length and the degree to which the two parameters approach the threshold to generate a continuous probability spectrum, which can effectively identify composite faults (such as a sudden change in optical power accompanied by a gradual increase in bit error rate). Fourthly, when the fluctuations of the two parameters are not temporally correlated, a threshold-driven window expansion strategy is activated. If a single parameter severely exceeds the limit, the system automatically expands the time window to capture potential delay anomalies a second time (such as a delayed spike in bit error rate caused by device reset), ensuring real-time alarm capabilities for serious faults while reducing the false negative rate of correlated faults through an adaptive mechanism. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0017] Figure 1 This is a schematic diagram illustrating the steps of the PON gateway fault detection method of the present invention;

[0018] Figure 2 This is a schematic diagram of a portion of step S5 in the PON gateway fault detection method of the present invention;

[0019] Figure 3 This is a partial flowchart of the PON gateway fault detection method of the present invention;

[0020] Figure 4This is a schematic diagram of another part of the PON gateway fault detection method of the present invention;

[0021] Figure 5 This is a block diagram illustrating an electronic device according to an embodiment of the present invention.

[0022] Figure label:

[0023] 700 - Electronic device; 701 - Processor; 702 - Memory; 703 - Multimedia component; 704 - I / O interface; 705 - Communication component. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0025] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] like Figure 1 , Figure 3 , Figure 4 As shown, a PON gateway fault detection method is provided, including:

[0028] S1. Obtain the status dataset of the PON gateway, and obtain the first data type and the second data type based on the status dataset. At each collection time node in the previous detection cycle, obtain the first status data corresponding to the first data type and the second status data corresponding to the second data type. Form a first data sequence based on multiple first status data and form a second data sequence based on multiple second status data.

[0029] S2. Obtain the first sub-fluctuation index of each first state data according to the first data sequence, and obtain the second sub-fluctuation index of each second state data according to the second data sequence. Obtain the collection time node corresponding to the largest first sub-fluctuation index and use it as the reference time node.

[0030] S3. Obtain the fuzzy quantity n, obtain n consecutive collection time nodes that include the reference time node and use them as the reference window, and determine whether the largest second sub-fluctuation index is located within the reference window.

[0031] S4. If so, the fault probability is obtained based on the position of the first state data, the second state data, and the second state data within the reference window at the reference time node.

[0032] S5. If not, obtain the fault detection strategy based on the first state data of the reference time node and the safety threshold.

[0033] In this embodiment, it should be noted that in S1, the PON gateway fault detection process begins by collecting a full state dataset of the gateway. This dataset contains various key performance indicators recorded during network operation, such as received optical power, uplink bit error rate, or downlink packet loss rate. These indicators comprehensively reflect the real-time status of the gateway. From this dataset, two specific data types are selected as analysis targets. For example, the first data type might be received optical power, and the second data type might be uplink bit error rate. The selection of these data types is based on potential fault correlations, such as the correlation between optical power fluctuations and abnormal bit error rates. Data collection focuses on a detection cycle prior to the current detection time point. This cycle ensures the timeliness of the data and avoids the impact of historical data lag on the accuracy of the analysis. The collection time points are evenly distributed within this cycle, for example, once per second or per minute, to capture the dynamic changing trends of the parameters.

[0034] Furthermore, for each selected data type, corresponding state data is extracted at all acquisition time points within the detection period. For example, the measured values ​​of received optical power at each node are extracted and combined in chronological order to form the first data sequence; similarly, the measured values ​​of uplink bit error rate form the second data sequence. These sequences preserve the integrity of the time dimension, ensuring that each data point corresponds to a specific time point, thereby constructing the respective change trajectories of the two parameters. The serialization process helps transform discrete data into a continuous time series, providing a structural input for subsequent steps to analyze fluctuation patterns and correlations, avoiding the limitations of single-point sampling. The amount of data in the sequence depends on the length of the detection period and the sampling frequency, ensuring coverage of the entire period without omissions.

[0035] In S2, the local fluctuation intensity index corresponding to each acquisition time node is first calculated for the first data sequence (e.g., the received optical power sequence). This index captures abrupt changes by quantifying the sum of differences between each time point and its neighboring data points. For example, for a data point in the middle of the sequence, the absolute values ​​of its deviation from the previous and next data points are calculated simultaneously, and the sum of these two values ​​is used as the fluctuation index value for that node. The first and last nodes of the sequence are treated specially (e.g., the previous neighbor of the first node is considered as itself, and the next neighbor of the last node is considered as itself) to ensure that each node has a calculable fluctuation characterization. This design can effectively amplify the instantaneous jump characteristics of the data and avoid interference from smooth fluctuations—for example, when the optical power suddenly drops at a certain node, the drastic difference in its forward and backward directions will generate significant fluctuation peaks.

[0036] Furthermore, the fluctuation index values ​​of all nodes in the first data sequence are traversed, and the time node with the highest fluctuation intensity is identified and marked as the reference time node. This node represents the core location where the first parameter is most likely to experience an anomaly, providing a time anchor for subsequent multi-parameter correlation analysis. For example, if the received optical power sequence experiences a precipitous drop at a certain moment, the fluctuation index at that moment will inevitably stand out as the maximum value of the sequence, thus being identified as a key reference node. Simultaneously, the same fluctuation index calculation process is performed on the second data sequence (such as the bit error rate sequence) to generate a second sub-fluctuation index sequence for each node, which is used to subsequently determine whether the anomaly event of the second parameter has a temporal correlation with the reference node of the first parameter. This step, by quantifying the fluctuation intensity to establish the coordinates of key events on the time axis, is a prerequisite for capturing coordinated faults.

[0037] In S3, the first step is to determine the width parameter of the time window—the ambiguity number n. This parameter reflects the time tolerance of fault correlation, and its value is determined by the total number of acquisition nodes within the detection period and a preset scaling factor. This ensures that the window width can cover typical fault propagation delays while avoiding the introduction of irrelevant noise due to an excessively large range. For example, if data is acquired at a millisecond-level density within the period, n may correspond to an equivalent time span of tens of milliseconds, used to capture the typical delay process from optical signal degradation to data transmission anomalies.

[0038] Furthermore, a continuous time window is constructed based on a reference time node. This window contains n consecutive acquisition nodes, typically extending symmetrically to both sides of the reference node (automatically adjusting to unilateral expansion when the reference node approaches the period boundary). For example, when detecting a sudden change in received optical power, the system generates a window interval covering a specific time span before and after the event, and then determines whether the maximum fluctuation point of the second parameter (such as the bit error rate) falls exactly within this interval. This design essentially establishes a time channel for fault propagation: if the strongest fluctuation points of the two parameters are within the same associated window, it indicates the existence of a causal chain of "optical power anomaly triggering an increase in bit error rate"; otherwise, it indicates that the fluctuations of the two parameters are not temporally correlated. The window mechanism overcomes the limitations of single-point sampling and can capture hidden fault patterns such as a spike in bit error rate occurring a few milliseconds after an optical power anomaly.

[0039] In S4, when the maximum fluctuation point of the second parameter is confirmed to be within the associated time window centered on the reference node of the first parameter, the multi-parameter collaborative fault analysis mechanism is activated. This mechanism generates dynamic fault probabilities by quantifying three key dimensions: First, the time-series propagation delay effect, i.e., calculating the number of time intervals between the maximum second fluctuation point and the reference node. The smaller the time difference, the stronger the synchronicity of the abnormal event (e.g., a sudden increase in bit error rate within milliseconds after a sudden change in optical power), and the higher the contribution value of the fault probability; when the time difference increases, the probability contribution decays exponentially, but retains tolerance for reasonable propagation delays. Second, a dual assessment of parameter safety margins, simultaneously calculating the closeness of the measured value of the first parameter at the reference time point to the safety threshold, and the closeness of the measured value of the second parameter at its maximum fluctuation point to the safety threshold. The minimum ratio of the two is taken as the safety margin coefficient. This design reflects the weakest link effect; any parameter approaching the threshold boundary will significantly increase the overall risk level. Third, dynamic weight adjustment, using a preset global scaling factor to weight and fuse the above elements to ensure that the probability value is within the standardized range. For example, when the received optical power drops to 90% of the safety threshold and the bit error rate rises to 95% of the threshold after a 3-node interval, a medium probability value will be output; while if the optical power drops to 80% of the threshold and the bit error rate reaches the threshold of 105% in the adjacent node, a high probability alarm will be generated.

[0040] Furthermore, in summary, this approach avoids the mechanical nature of single-point threshold criteria (such as ignoring parameter combinations that do not exceed the threshold but exhibit abnormal mutations) and overcomes the blind spots of traditional methods regarding latent temporal faults. For example, when a slight bend in the optical fiber causes a short-term, drastic fluctuation in optical power (without sustained exceedance), if a continuous jump in the uplink bit error rate (also without exceeding the threshold) is detected within the corresponding time window, the model will output a probability alarm based on the fluctuation intensity and transmission time, indicating "intermittent signal degradation caused by physical damage to the optical fiber." This process does not rely on manual empirical rules. By transforming the hidden correlation of multiple parameters on the time axis into probabilities, maintenance personnel can then handle the situation in stages based on probability values—low-probability events only require recording and observation, while medium- and high-probability events trigger pre-maintenance work orders, providing a quantitative basis for early intervention in sub-optimal conditions and significantly improving the predictive ability for progressive faults.

[0041] In S5, when the maximum fluctuation point of the second parameter is not within the reference time window, a hybrid strategy combining static thresholding and dynamic expansion will be activated. This strategy first checks whether the first state data of the reference time node exceeds the safety threshold: if it does not exceed the limit (e.g., although the received optical power fluctuates abnormally, it is still within the safe range), it is directly determined as a zero-risk event, and the subsequent dynamic expansion mechanism is not triggered; if it exceeds the limit, the tolerance of the time window needs to be corrected. The correction mechanism dynamically adjusts the number of ambiguities based on the severity of the exceedance. When the measured value far exceeds the safety threshold (e.g., a sudden increase in optical power), the node coverage of the time window will be significantly expanded to ensure that the delay fluctuations of the second parameter (e.g., bit error rate) can still be captured. For example, when the optical power exceeds the limit by 5 times, the new window width may be expanded to several times the original window, and then the maximum fluctuation point of the bit error rate is re-evaluated based on the expanded window.

[0042] Furthermore, this mechanism integrates the advantages of single-dimensional threshold criteria and multi-parameter time-series analysis. On the one hand, it directly triggers basic alarms for clearly exceeding standards in severe events; on the other hand, it reserves analysis channels for potential related faults through dynamic window expansion. For example, when an ONU (Optical Network Unit) power supply abnormality causes a sudden change in received optical power above the safety threshold, window expansion is performed first (due to the high degree of exceeding the standard). Subsequently, a spike in bit error rate appearing a few seconds later in the expanded window (due to data errors generated during the device reset process) can still be captured, and a related fault alarm can still be generated. This design ensures immediate response to obvious severe faults while reducing the risk of missed detection of related faults due to independent parameter exceeding the standard through a secondary detection window, forming a layered detection system.

[0043] In summary, the entire PON gateway fault detection method firstly uses cross-node fluctuation quantization technology to accurately capture transient anomalies in key parameters (such as a cliff-like drop in optical power), avoiding the omission of normal abnormal signals by existing sampling methods; furthermore, based on the fault propagation timing model, the extreme fluctuation points of the first parameter are used as time anchors to dynamically generate a correlation window covering typical propagation delays (such as a reasonable period from millisecond-level optical signal degradation to increased bit error rate), intelligently determining whether the extreme fluctuations of the second parameter have temporal coupling; furthermore, when dual-parameter fluctuations exist... In the case of no correlation, a multi-dimensional safety margin fusion is adopted, which combines the timing delay length and the degree to which the two parameters approach the threshold to generate a continuous probability spectrum, which can effectively identify compound faults (such as a sudden change in optical power accompanied by a gradual increase in bit error rate). Furthermore, when the fluctuation of the two parameters has no timing correlation, a threshold-driven window expansion strategy is activated. If a single parameter seriously exceeds the limit, the system automatically expands the time window to capture potential delay anomalies a second time (such as a delayed spike in bit error rate caused by device reset). This not only ensures the real-time alarm capability for serious faults, but also reduces the false negative rate of correlated faults through an adaptive mechanism.

[0044] like Figure 2 As shown, in one embodiment, S5, obtaining the fault detection strategy based on the first state data of the reference time node and the safety threshold includes:

[0045] S51. If the first state data of the reference time node exceeds the safety threshold, then the fuzzy quantity n is corrected and a new fuzzy quantity n' is obtained.

[0046] S52. Obtain a new reference window based on the new fuzzy quantity n', and determine whether the largest second sub-fluctuation index is located within the new reference window.

[0047] In this embodiment, it should be noted that in S51, when the first parameter detection value (such as received optical power) of the reference time node exceeds the safety threshold, a dynamic window expansion mechanism will be activated. Specifically, the tolerance of the time window is intelligently adjusted according to the severity of the exceedance. That is, an expansion coefficient is generated by quantifying the deviation between the actual detection value and the threshold. If the deviation is extremely large, the expansion coefficient drives the number of fuzzy elements to increase exponentially; if the exceedance is slight, only the window width is finely adjusted. An upper limit constraint is set during the correction process (not exceeding the total number of nodes in the detection cycle) to avoid unlimited expansion leading to invalid noise interference. For example, when a loose fiber optic interface causes the optical power to continuously exceed the threshold, an expansion coefficient is automatically generated, making the number of new fuzzy elements several times the original value, thereby expanding the associated window from single-digit seconds to double-digit seconds.

[0048] In S52, a secondary reference window is constructed based on the revised new fuzzy number. The new reference window still contains the original reference nodes, but its width is dynamically adjusted according to the expansion coefficient. If the original window only covers 5 nodes before and after the fault point (e.g., 50 milliseconds), the revised window may expand to 10 nodes before and after (100 milliseconds). The maximum fluctuation point of the second parameter (e.g., bit error rate) is re-evaluated to see if it falls within the expanded window. For example, when an ONU power supply module failure causes a sudden increase in optical power, the original window may miss the associated fluctuation because the device protection mechanism has not yet been triggered; the expanded window can capture the peak bit error rate spike generated during the delayed device reset process (e.g., a data verification error occurs 2 seconds after a power outage). If the maximum second fluctuation point is located within the new window at this time, a composite fault alarm is generated; if it still does not hit, it is determined to be an isolated single-parameter fault. This cascading mechanism of "threshold triggering - window expansion - secondary association" reduces missed detections caused by fault propagation delays and avoids false associations caused by excessive leniency by precisely controlling the expansion amplitude.

[0049] In one implementation, the process of correcting the fuzzy quantity n and obtaining a new fuzzy quantity n' in S51 is represented as follows:

[0050] ;in,

[0051] For the new fuzzy quantity, This represents the number of time points collected within the detection period. For the first state data at the reference time point, The safety threshold for the first state data. For fuzzy quantities.

[0052] In this embodiment, it should be noted that the entire expression is used to address the problem of long delays in associated events caused by sudden severe failures. Specifically, This indicates the degree to which a parameter exceeds the limit. When the optical power significantly exceeds the threshold, at this point... A significant increase in the number of driving fuzzy elements (e.g., a result of 3) indicates that the number of driving fuzzy elements n is increasing. Multiplication expansion. Furthermore, through multiplication... Achieving a non-linear relationship between the degree of exceeding the limit and the window width, for example, when the optical power suddenly increases to 140% of the threshold ( With a scaling factor of 1.4, if the original n=10, then the new number of fuzzy elements n'=min{m, 10×1.4}=14. This design ensures that the more severe the fault, the larger the window expansion, to accommodate associated events with longer delays (such as second-level delay bit errors triggered by device protection mechanisms).

[0053] Furthermore, upper limit constraints The necessity of this is to limit the expanded window to no more than the total number of nodes m in the detection cycle, and to prevent the introduction of irrelevant noise (such as expanding to historical irrelevant data) due to unlimited expansion.

[0054] In summary, when a sudden fault causes an abnormal delay in the second parameter, the original fixed window may not be able to cover it. The formula captures this delay correlation by dynamically expanding the window. For example, if the optical power exceeds the limit by 0.5 times and the scaling factor is 1.5, then n'=min{m, n×1.5}, the window expands by nearly 0.5 times, which can cover the device reset process (such as the peak bit error rate after 2 seconds). Furthermore, to avoid over-expansion causing misjudgment, when the exceedance is slight (such as... Since n'≈n, the window remains almost unchanged, preventing brief disturbances from triggering invalid expansion. In summary, this corrected formula accurately solves the problem of missed detection of severe fault-related events due to the delay of static windows through a three-level mechanism of dynamically quantifying fault severity, adaptively expanding the time window, and constraining the effective range.

[0055] In one implementation, the number of fuzzy elements n obtained in S3 is represented as:

[0056] ;in,

[0057] For fuzzy quantities, This represents the number of time points collected within the detection period. This is the splitting coefficient.

[0058] In this embodiment, it should be noted that in the entire expression, m represents the total number of data collection time points within the detection period, reflecting the monitoring density; for example, if the period is 60 seconds and sampling occurs once per second, then m = 60; if sampling is at the millisecond level, m can reach several thousand. Furthermore, the splitting coefficient... As a mapping coefficient from fault propagation time to sampling density, the typical fault propagation time (e.g., 50 milliseconds for a bit error rate response to an optical power anomaly) is converted into a node number span; for example, when the sampling interval is 10 milliseconds, the typical fault propagation time is 50 milliseconds, corresponding to 5 nodes. If the number of sampling nodes within the monitoring period... If the value is 30, and the system is configured to cover this transmission process, then... We should make n≈5, that is =m / 5=6.

[0059] Furthermore, round down. Ensure that n is an integer number of nodes to avoid time point misalignment caused by non-integer windows and maintain window continuity.

[0060] It should also be noted that the splitting coefficient This can be determined through characteristic calibration and dynamic optimization. First, by simulating faults in the laboratory (such as fiber bending or equipment power failure), the average transmission delay of key parameter pairs is statistically analyzed (e.g., the average delay caused by optical power increasing the bit error rate is 50 milliseconds). Then, the node span baseline value is calculated, and the baseline number of nodes is equal to the transmission delay divided by the sampling interval (e.g., if the sampling interval is 5 ms and the transmission delay is 30 ms, then the baseline number of nodes is 6). Finally, the splitting coefficient is obtained. , This is equal to the base number of nodes multiplied by a safety factor, where the safety factor is between 1.5 and 2. Simultaneously, historical fault data is collected, and the actual propagation time distribution is statistically analyzed. If insufficient coverage of the extended window is found, the splitting factor is increased. .

[0061] In one implementation, the first sub-fluctuation index obtained in S2 based on the first data sequence for each first state data is represented as follows:

[0062] ;in,

[0063] This is the first sub-fluctuation index for the i-th first state data. This represents the first state data at the i-th data collection time point. This represents the first state data at the (i-1)th acquisition time node. This represents the first state data at the (i+1)th data collection time point. This represents the number of time points collected within the detection period.

[0064] In this embodiment, it should be noted that the entire expression is calculated using the summation of the absolute values ​​of the two-way differences (forward difference). Backward difference This method combines several approaches to enhance the detection of sudden anomalies at a single point. For example, in a normal scenario, optical power changes slowly between adjacent nodes (e.g., from -25dBm to -25.1dBm), with both bidirectional difference values ​​close to zero, resulting in a low index value. In a transient fault scenario, the optical power at a certain node changes rapidly (e.g., from -25dBm to -40dBm and then back to -25dBm), with both the forward and backward difference values ​​at 15.0 (drastic change). The combined effect of these two values ​​further highlights the anomaly. Existing methods use a static threshold comparison method, focusing only on whether the instantaneous value exceeds the limit. If the anomaly does not continuously exceed the limit, it will be ignored. This index captures short-term faults through fluctuation amplitude, significantly reducing the false negative rate. Meanwhile, when multiple parameters exhibit coordinated anomalies at similar time points (such as sudden changes in optical power and a sudden increase in bit error rate), even if none of them exceed the threshold, the probability of failure will increase. If the optical power sequence shows high fluctuations at a certain node and the bit error rate sequence shows high fluctuations at another node, S3 and S4 will examine the correlation between the two within the time window and trigger a latent fault warning.

[0065] Furthermore, the first and last nodes of the sequence lack unilateral neighbor nodes. Directly ignoring these will lead to missed detections of critical events (such as sudden failures at the start / end of the detection cycle). To avoid the loss of end-node information, a forced... and It can still reflect forward and backward changes.

[0066] Furthermore, bidirectional differential is more sensitive to transient interference (such as instantaneous noise of a single node), but its response to steady-state drift (such as slow change in the same direction of three consecutive nodes) is weaker than that of the single-point threshold method. It naturally filters low-frequency interference and focuses on high-frequency sudden faults, perfectly matching the dynamic fluctuation scenario of PON.

[0067] In one implementation, the fault probability obtained in S4 based on the positions of the first state data, the second state data, and the second state data within the reference window at the reference time node is expressed as follows:

[0068] ;in,

[0069] This represents the probability of failure. To reduce the factor, This represents the number of nodes between the data collection time node and the reference time node corresponding to the largest second-order sub-fluctuation indicator. The safety threshold for the first state data. For the first state data at the reference time point, The safety threshold for the second state data. This is the second state data corresponding to the largest second sub-fluctuation indicator.

[0070] In this embodiment, it should be noted that the time propagation factor is used throughout the expression. This represents the number of time intervals between the first parameter anomaly event (such as a sudden change in optical power) and the second parameter anomaly event (such as a spike in bit error rate). It implements the addition of timing correlation for fault characteristics; if... (The two parameters fluctuate completely synchronously), and the exponential term exhibits minimal decay. Approaching 1 (highest probability of failure); if As the exponential term increases, its decay accelerates. The reduction is reduced to only the tolerance for reasonable delays (such as a typical 50ms delay from optical power to increased bit error rate).

[0071] Furthermore, It is a two-parameter safety margin factor; where, Assess the criticality of the first parameter (such as optical power) of the reference node: if it is higher than the threshold, the ratio is less than 1 (representing exceeding the standard), the smaller the ratio, the more serious the exceeding of the standard, and the closer it is to the threshold, the closer the ratio is to 1; if it is lower than the threshold, the ratio is greater than 1 (representing the normal situation). Similarly, assess the risk of the second parameter at its point of maximum fluctuation to address the shortcomings of single-dimensional parameter analysis. Taking the minimum of the two parameters ensures that either parameter approaching or exceeding a threshold significantly increases the risk; for example, optical power at 90% of the threshold (…). The bit error rate exceeded the threshold by more than double. The final value is 0.5, reflecting the increased probability of the dominant fault due to a severely excessive bit error rate. Simultaneously, latent fault detection: even if neither parameter exceeds the threshold (e.g., the ratio is greater than 1), the smaller value will be selected for fault probability calculation.

[0072] Furthermore, reduce the coefficient. This is a preset coefficient that adjusts the combined influence of time delay and safety margin on probability; it is typically set to 0.5. This mitigates the high false alarm rate of static thresholds. When the exponent term increases, the exponent term... It is more sensitive to the minimum value and can also react more sensitively to weakly correlated events and weak overshoots. When When decreasing the probability, more stringent associated events and more severe exceedances are required for a significant change in probability. Simultaneously, the reduction coefficient... The acquisition of this data can be accomplished using existing technologies. First, a feature dataset is constructed, collecting multi-parameter time-series data corresponding to historical fault events (such as fiber breakage and optical module aging), and the actual fault time window (i.e., the delay interval from the first parameter anomaly to the second parameter response) is manually marked. Then, the delay and margin features are quantified to obtain the number of node intervals of the actual transmission delay (e.g., a difference of 5 nodes between optical power anomaly and the peak bit error rate). The optimal value is obtained by minimizing the mean square error. , optimal The probability of failure should be minimized. It is positively correlated with the actual failure risk.

[0073] In summary, if the optical power does not exceed the limit (due to environmental interference), the bit error rate delay is extremely long ( Extremely high), high dual-parameter safety margin ( (High value), then Significantly reduces and avoids false alarms; when optical power changes abruptly and exceeds the limit by 20%, and the bit error rate also increases sharply and exceeds the limit by 25%, the minimum value is reduced to 0.8, combined with short latency ( ), to obtain a high probability (such as hour, This identifies latent degradation. Therefore, the entire expression is achieved through... (Time-related strength), Two-parameter safety margin bottleneck effect, The adaptive scaling 3D dynamic coupling captures collaborative anomalies that do not exceed the threshold, breaking through the blind spot of hidden faults; at the same time, it quantifies the propagation delay and establishes multi-parameter time sequence correlation; and at the same time, it uses the exponential decay characteristic to contain transient interference, significantly reducing the false alarm rate.

[0074] A PON gateway fault detection system is also provided, the system including:

[0075] The acquisition module is used to acquire the status dataset of the PON gateway, and acquire a first data type and a second data type based on the status dataset. At each collection time node in the previous detection cycle, the module acquires the first status data corresponding to the first data type and the second status data corresponding to the second data type, and forms a first data sequence based on multiple first status data and a second data sequence based on multiple second status data.

[0076] The data processing module is used to obtain the first sub-fluctuation index of each first state data according to the first data sequence, and to obtain the second sub-fluctuation index of each second state data according to the second data sequence, and to obtain the collection time node corresponding to the largest first sub-fluctuation index as a reference time node.

[0077] The judgment module is used to obtain the fuzzy quantity n, obtain n consecutive collection time nodes that include the reference time node and use them as a reference window, and determine whether the largest second sub-fluctuation index is located within the reference window;

[0078] The first detection module is used to obtain the fault probability based on the positions of the first state data, the second state data, and the second state data within the reference window when the largest second sub-fluctuation index is within the reference window.

[0079] The second detection module is used to obtain a fault detection strategy based on the first state data of the reference time node and the safety threshold when the largest second sub-fluctuation index is not within the reference window.

[0080] In one embodiment, the second detection module is further configured to: if the first state data of the reference time node exceeds a preset threshold, correct the fuzzy quantity n and obtain a new fuzzy quantity n'; obtain a new reference window based on the new fuzzy quantity n', and determine whether the largest second sub-fluctuation index is located within the new reference window.

[0081] In this embodiment, it should be noted that the specific method of performing the operation of the above-mentioned PON gateway fault detection system has been described in detail in the embodiments of the relevant PON gateway fault detection method, and will not be elaborated here.

[0082] Figure 5This is a block diagram of an electronic device illustrating a PON gateway fault detection method according to an exemplary embodiment. Figure 5 As shown, the electronic device 700 may include: a processor 701 and a memory 702. The electronic device 700 may also include one or more of a multimedia component 703, an I / O interface 704 (input / output interface), and a communication component 705.

[0083] The processor 701 controls the overall operation of the electronic device 700 to complete all or part of the steps in the PON gateway fault detection method described above. The memory 702 stores various types of data to support the operation of the electronic device 700. This data may include, for example, instructions for any application or method operating on the electronic device 700, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 702 or transmitted via communication component 705. The audio component also includes at least one speaker for outputting audio signals. I / O interface 704 provides an interface between processor 701 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or a combination thereof, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0084] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the PON gateway fault detection method described above.

[0085] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the PON gateway fault detection method described above. For example, the computer-readable storage medium may be the memory 702 including program instructions described above, which may be executed by the processor 701 of the electronic device 700 to complete the PON gateway fault detection method described above.

[0086] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described PON gateway fault detection method when executed by the programmable device.

[0087] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0088] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0089] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for detecting faults in a PON gateway, characterized in that, include: Obtain the status dataset of the PON gateway, and obtain the first data type and the second data type based on the status dataset. At each collection time node in the previous detection cycle, obtain the first status data corresponding to the first data type and the second status data corresponding to the second data type. Form a first data sequence based on multiple first status data and form a second data sequence based on multiple second status data. The first sub-fluctuation index of each first state data is obtained according to the first data sequence, and the second sub-fluctuation index of each second state data is obtained according to the second data sequence. The collection time node corresponding to the largest first sub-fluctuation index is obtained and used as the reference time node. Get the number of fuzzy items n, get n consecutive collection time nodes that include the reference time node and use them as the reference window, and determine whether the largest second sub-fluctuation index is located within the reference window; If so, the failure probability is obtained based on the positions of the first state data, the second state data, and the second state data within the reference window at the reference time node. The probability of obtaining a fault is expressed as: ;in, This represents the probability of failure. To reduce the factor, This represents the number of nodes between the data collection time node and the reference time node corresponding to the largest second-order sub-fluctuation indicator. The safety threshold for the first state data. For the first state data at the reference time point, The safety threshold for the second state data. The second state data corresponding to the largest second sub-fluctuation indicator; If not, then obtain the fault detection strategy based on the first state data of the reference time node and the safety threshold; The fault detection strategy includes: if the first state data of the reference time node exceeds the safety threshold, the fuzzy quantity n is corrected and a new fuzzy quantity n' is obtained; a new reference window is obtained based on the new fuzzy quantity n', and it is determined whether the largest second sub-fluctuation index is located within the new reference window.

2. The PON gateway fault detection method according to claim 1, characterized in that, The process of correcting the fuzzy quantity n and obtaining a new fuzzy quantity n' is expressed as follows: ;in, For the new fuzzy quantity, This represents the number of time points collected within the detection period. For the first state data at the reference time point, The safety threshold for the first state data. For fuzzy quantities.

3. The PON gateway fault detection method according to claim 1, characterized in that, The number of fuzzy elements n is represented as: ;in, For fuzzy quantities, This represents the number of time points collected within the detection period. This is the splitting coefficient.

4. The PON gateway fault detection method according to claim 1, characterized in that, The first sub-fluctuation index, which obtains each first state data based on the first data sequence, is represented as follows: , , , ;in, This is the first sub-fluctuation index for the i-th first state data. This represents the first state data at the i-th data collection time point. This represents the first state data at the (i-1)th acquisition time node. This represents the first state data at the (i+1)th data collection time point. This represents the number of time points collected within the detection period.

5. A PON gateway fault detection system, characterized in that, The system includes: The acquisition module is used to acquire the status dataset of the PON gateway, and acquire a first data type and a second data type based on the status dataset. At each collection time node in the previous detection cycle, the module acquires the first status data corresponding to the first data type and the second status data corresponding to the second data type, and forms a first data sequence based on multiple first status data and a second data sequence based on multiple second status data. The data processing module is used to obtain the first sub-fluctuation index of each first state data according to the first data sequence, and to obtain the second sub-fluctuation index of each second state data according to the second data sequence, and to obtain the collection time node corresponding to the largest first sub-fluctuation index as a reference time node. The judgment module is used to obtain the fuzzy quantity n, obtain n consecutive collection time nodes that include the reference time node and use them as a reference window, and determine whether the largest second sub-fluctuation index is located within the reference window; The first detection module is used to obtain the fault probability based on the positions of the first state data, the second state data, and the second state data within the reference window when the largest second sub-fluctuation index is within the reference window. The probability of obtaining a fault is expressed as: ;in, This represents the probability of failure. To reduce the factor, This represents the number of nodes between the data collection time node and the reference time node corresponding to the largest second-order sub-fluctuation indicator. The safety threshold for the first state data. For the first state data at the reference time point, The safety threshold for the second state data. The second state data corresponding to the largest second sub-fluctuation indicator; The second detection module is used to obtain a fault detection strategy based on the first state data of the reference time node and the safety threshold when the largest second sub-fluctuation index is not within the reference window. The fault detection strategy includes: if the first state data of the reference time node exceeds the safety threshold, the fuzzy quantity n is corrected and a new fuzzy quantity n' is obtained; a new reference window is obtained based on the new fuzzy quantity n', and it is determined whether the largest second sub-fluctuation index is located within the new reference window.

6. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor is configured to execute the computer program in the memory to implement the PON gateway fault detection method according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the PON gateway fault detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Optical power monitoring system and method applied to optical fiber

    CN118890092A

  • FTTR network data quality optimization method, device, equipment and medium

    CN119996880A