A method, device, equipment and storage medium for tracing causes of laser communication link failure

By screening and analyzing multiple factors during inter-satellite laser communication link failures, calculating the convolution of linear templates with time-series data, and determining the causal correlation of pulse changes, the complex problem of tracing the causes of inter-satellite laser communication link failures was solved, enabling rapid and accurate location of the cause of the failure and reducing manpower costs.

CN122247504APending Publication Date: 2026-06-19CHINA STAR NETWORK SYST RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA STAR NETWORK SYST RES INST CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

The analysis of fault tracing in inter-satellite laser communication links is complex. Existing methods are difficult to effectively diagnose frequent fault scenarios in the early stages, and insufficient historical data makes it difficult to determine the cause of the fault.

Method used

By acquiring multiple factors and original time-series data within a preset time period before and after the fault, a second factor is selected, the convolution of the linear template and the time-series data is calculated, the start and end times of the pulse change are determined, and the cause of the fault is determined based on causal correlation.

Benefits of technology

It improves fault diagnosis efficiency, reduces labor costs, and enables rapid and accurate identification of fault causes in early-stage, frequent fault scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122247504A_ABST
    Figure CN122247504A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, device, and storage medium for tracing the causes of laser communication link failures, comprising: acquiring raw time-series data corresponding to a first factor within a preset time period before and after the failure time; selecting at least one second factor from multiple first factors according to a preset filtering method; calculating the convolution of multiple linear templates with preset slopes with the time-series data of the second factor to obtain the response values ​​of each template at different times; determining the start time, end time, and effective slope of the pulse change corresponding to each second factor based on the response values ​​of the multiple linear templates corresponding to each second factor at different times; determining the causal correlation corresponding to each second factor based on the effective slope and the chronological relationship between the start time, end time, and failure time of the pulse change and the failure time; and determining the cause of the failure based on the causal correlation corresponding to each second factor; thereby improving fault diagnosis efficiency and reducing labor costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of inter-satellite laser communication, and in particular to a method, apparatus, device and storage medium for tracing the cause of laser communication link failures. Background Technology

[0002] In inter-satellite laser communication, when lasers and optical components are in orbit, they are exposed to a complex space environment for a long time. Their reliability is affected by a combination of factors such as temperature changes, micro-vibrations, single-event effects, total dose radiation, vacuum effects, and cosmic background light, which can lead to failures in inter-satellite laser links. These failures are often caused by a combination of factors, making fault analysis extremely complex.

[0003] Currently, fault diagnosis methods for laser link systems are still relatively rudimentary, mostly only capable of detecting the "on" and "off" states of the link. When tracing the cause after a link failure, existing correlation analysis methods can only screen out factors that are statistically strongly correlated with the link failure event, but cannot determine whether the data of these strongly correlated factors are the cause of the failure or whether the failure caused the data changes of these factors. Furthermore, existing data mining-based autonomous fault diagnosis methods are usually designed for more mature scenarios such as power distribution networks, terminals, and wind turbine bearings, requiring a large amount of historical data for automated fault cause analysis. However, laser link failures typically occur frequently in the early stages of their lifecycle, with short initial testing periods, making it difficult to obtain a sufficient sample size. In addition, historical data is insufficient, making it difficult to use artificial intelligence models to solve fault problems. Therefore, existing data mining methods are also difficult to adapt to the fault scenarios of laser links. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for tracing the causes of laser communication link failures, which can improve the efficiency of fault diagnosis and reduce labor costs.

[0005] Firstly, this application provides a method for tracing the causes of laser communication link failures, including: Acquire multiple primary factors and their corresponding original time-series data within a preset time period before and after the fault; At least one second factor is selected from the plurality of first factors according to a preset filtering method; For the original time series data of each second factor, the convolution of multiple linear templates with preset slopes with the time series data of the second factor is calculated to obtain the response value of each template at different times. Based on the response values ​​of multiple linear templates corresponding to each second factor at different times, determine the start time, end time, and effective slope of the pulse change corresponding to each second factor; Based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time and fault time of the pulse change, the causal correlation corresponding to each second factor is determined; wherein, the causal correlation is the quantified value of the change behavior of the original time series data corresponding to the second factor leading to this fault. The cause of this failure was determined based on the causal correlation of each of the second factors.

[0006] In one or more possible embodiments, the step of convolving multiple linear templates with preset slopes with the time-series data of the second factor to obtain the response values ​​of each template at different times includes: For each second factor's original time series data, obtain the corresponding integral time series data and differential time series data for each second factor; Based on multiple preset slope parameters, multiple corresponding linear template vectors are generated, and the non-zero elements in each linear template vector are flipped to obtain the corresponding flipped template vector. Any one of the original time-series data, the integral time-series data, and the differential time-series data is convolved with the flipped template vector corresponding to each slope to obtain a sequence of response values ​​for each slope template at different times; the sequence of response values ​​is used to characterize the degree of shape matching between the original time-series data, the integral time-series data, and the differential time-series data and the multiple linear template vectors.

[0007] In one or more possible embodiments, determining the start time, end time, and effective slope of the pulse change corresponding to each second factor based on the response values ​​of multiple linear templates corresponding to each second factor at different times includes: Obtain the maximum response value for each second factor, and use the slope parameter corresponding to the maximum response value as the effective slope of the pulse change for each second factor. The start time of the pulse change is determined based on the time position corresponding to the maximum response value; The search proceeds from the time position corresponding to the start time in the direction of increasing time, and the time position at which the response value first falls below a preset threshold is determined as the cutoff time of the pulse change.

[0008] In one or more possible embodiments, determining the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time, and fault time of the pulse change includes: Based on the start time of the pulse change, the end time of the pulse change, and the fault time corresponding to each second factor, four time differences are determined respectively; the four time differences include the advance and lag of the start time relative to the fault time, and the advance and lag of the end time relative to the fault time. Based on the preset interval corresponding to the effective slope of each second factor, a preset processing rule is determined; wherein, different preset processing rules are associated with different weight parameter groups, and the weight parameter groups are used to characterize the degree of influence of the four time differences on the fault moment. The causal correlation of each second factor is calculated based on the maximum response value of each second factor, the four time differences, and the weight parameter group associated with the selected preset processing rule.

[0009] In one or more possible embodiments, the fault time is determined in the following manner: Monitor the tracking status, communication status, and error status of the laser terminal; When any of the following states—the tracking state, the communication state, or the bit error state—changes from normal to fault, the moment when the state changes is taken as the fault moment.

[0010] In one or more possible embodiments, before selecting at least one second factor from the plurality of first factors according to a preset filtering method, the method further includes: Determine whether each primary factor is included in a pre-established failure mode library; For each first factor included in the fault mode library, determine whether the time difference between the time when the original time series data corresponding to the first factor becomes abnormal and the time when the fault occurs is less than a preset time difference. If so, determine that the first factor is strongly correlated with the current fault. The original time series data corresponding to the first factor are classified according to the pre-established fault classification model, and the first factor is determined as the cause of the fault based on the classification results.

[0011] In one or more possible embodiments, the step of selecting at least one second factor from the plurality of first factors according to a preset filtering method includes: The original time series data corresponding to multiple first factors within a preset time period before and after the fault time are filtered out, and the first factors that are unrelated to the fault and the first factors whose original time series data have not changed are removed.

[0012] In one or more possible embodiments, it also includes: Based on the causal correlations of each second factor under multiple failure scenarios, a feature matrix of failure scenarios is constructed; wherein, each row of the feature matrix corresponds to a failure scenario, each column corresponds to a second factor, and the matrix elements are the causal correlations of each second factor under the corresponding failure scenario. One fault scenario in the feature matrix is ​​taken as a cluster, and multiple clusters are iteratively merged using an agglomerative hierarchical clustering algorithm to form classification results with different hierarchical structures; Retrieve the keyword description for each cluster in each level; The fault mode library is updated based on the keyword descriptions and classification results.

[0013] Secondly, this application provides a laser communication link fault tracing device, comprising: The data acquisition module is used to acquire multiple primary factors within a preset time period before and after the fault, as well as the original time-series data corresponding to each primary factor. The factor filtering module is used to filter at least one second factor from the plurality of first factors according to a preset filtering method; The response value determination module is used to calculate the convolution of multiple linear templates with preset slopes with the time series data of the second factor for the original time series data of each second factor, so as to obtain the response value of each template at different times. The pulse change determination module is used to determine the start time, end time, and effective slope of the pulse change corresponding to each second factor based on the response values ​​of multiple linear templates corresponding to each second factor at different times. The causal correlation determination module is used to determine the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time and fault time of the pulse change; wherein, the causal correlation is the quantified value of the change behavior of the original time series data corresponding to the second factor leading to this fault. The fault cause determination module is used to determine the cause of the fault based on the causal correlation of each second factor.

[0014] Thirdly, this application provides an electronic device, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform any of the methods in the first aspect.

[0015] Fourthly, this application provides a computer storage medium storing a computer program for causing a computer to perform any of the methods described in the first aspect.

[0016] This application provides a method, apparatus, device, and storage medium for tracing the causes of laser communication link failures, which can improve the efficiency of fault diagnosis and reduce labor costs. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application, and do not constitute an undue limitation of this application.

[0018] Figure 1 This is a flowchart of a laser communication link fault tracing method provided according to an embodiment; Figure 2 This is a flowchart of a clustering method provided according to an embodiment; Figure 3 This is an overall framework diagram of a laser communication link fault tracing method provided according to an embodiment; Figure 4 This is a schematic diagram of a laser communication link fault tracing device according to an embodiment; Figure 5 This is a schematic diagram of an electronic device according to an embodiment; Figure 6 This is a schematic diagram of a computer-readable storage medium provided according to an embodiment. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0020] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0021] Furthermore, in the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0022] Lasers and optical components are greatly affected by the space environment, involving factors such as temperature, micro-vibration, single-event effect, total dose effect, vacuum effect, and cosmic background light. Therefore, there are many factors related to the on / off state of inter-satellite laser links. Current fault diagnosis methods for laser links usually only refer to detecting the on / off state of the link. For tracing the cause after the link is broken, correlation analysis is usually used to extract factors with strong correlations or to establish an artificial intelligence model for a specific problem to identify one or more faults.

[0023] Existing data mining-based fault diagnosis methods are typically designed for mature scenarios such as power distribution networks, terminals, and wind turbine bearings. For example, for one or several fault types, they utilize large amounts of historical data or fault information from multiple terminals to transform the process of manual judgment, analysis, and processing into automated fault cause analysis based on data mining. However, faults in laser links usually occur frequently in the early stages of their lifecycle. The initial testing time is short, making it difficult to obtain a sufficient sample size, and historical data is insufficient, making it difficult to use artificial intelligence models to solve fault problems. Therefore, existing data mining methods are also difficult to adapt to the fault scenarios of laser links.

[0024] To address the aforementioned issues, this application provides a method, apparatus, device, and storage medium for tracing the causes of laser communication link failures, which can improve fault diagnosis efficiency and reduce labor costs.

[0025] First, this application provides a method for tracing the cause of laser communication link failures, specifically as follows: Figure 1 As shown, it includes: Step 101: Obtain multiple first factors within a preset time period before and after the fault time, and the original time series data corresponding to each first factor; In one or more possible embodiments, the fault time is determined by: monitoring the tracking status, communication status, and bit error rate status of the laser terminal; determining when any of the tracking status, communication status, or bit error rate status changes from normal to faulty, and taking the moment of the status change as the fault time; specifically, using the laser terminal's operating status flag, logical judgment is performed; when the laser terminal's operating status is not fine tracking, the laser terminal's tracking status is determined to be disconnected, and the moment when the laser terminal's tracking status is determined to be disconnected is the fault time; using the data frame count sent from the router to the laser terminal. and the number of data frames sent from the laser terminal to the router Theoretical frame count at the current communication rate The moment when the communication status of the laser terminal is disconnected is the moment of the fault. The specific formula is as follows:

[0026] Where Y represents the laser terminal communication working status, 1 indicates normal and 0 indicates disconnection; α represents the set threshold, which is usually set to 0.5; Error rate of data frames sent from the peer laser terminal to the local laser terminal and the theoretical frame count at the current communication rate. To determine the error-prone operating state of the local laser terminal, the moment the error-prone operating state is disconnected is the moment of the fault. The specific formula is as follows:

[0027] in, This indicates the working status of the laser terminal's bit error rate; 1 indicates normal operation, and 0 indicates disconnection. β represents the set threshold, which is usually set to 0.5.

[0028] In one or more possible embodiments, the first factor includes factors such as feed antenna rotation, single event flip, single event latch-up, temperature, motor current, orbital attitude, optical amplifier power, Earth-Moon shadow, and solar-Moon eclipse; taking temperature as an example, assuming sampling once per second, the original time series data corresponding to the temperature within a preset time period before and after the fault time are obtained, resulting in (T1, T2, T3, ..., Tn), indicating that n temperatures have been collected.

[0029] Step 102: Select at least one second factor from the plurality of first factors according to a preset filtering method; In one or more possible embodiments, the original time-series data corresponding to multiple first factors within a preset time period before and after the fault time are filtered out, and first factors unrelated to the fault and first factors whose original time-series data have not changed are removed. For example, the correlation between the time-series data of each first factor and the fault occurrence state can be calculated by the Pearson correlation coefficient method, and the corresponding first factor should be removed by comparing the result of the preset correlation threshold with the calculated correlation. If the correlation of a first factor is less than the preset correlation threshold, it is determined to be a first factor unrelated to the fault and removed. For example, if the time-series data of temperature has not changed within the preset time period before and after the fault time and has always remained at a fixed value or a preset range, it indicates that temperature is a first factor whose original time-series data has not changed.

[0030] Step 103: For the original time series data of each second factor, calculate the convolution of multiple linear templates with preset slopes with the time series data of the second factor to obtain the response value of each template at different times; Step 104: Based on the response values ​​of multiple linear templates corresponding to each second factor at different times, determine the start time, end time, and effective slope of the pulse change corresponding to each second factor. In one or more possible embodiments, the raw time-series data corresponding to temperature is used as an example: Raw temperature time-series data is extracted within a preset time period (e.g., from 1 minute before the failure to 1 minute after the failure) before and after the laser link failure time (t_break); assuming the temperature is collected once per second, (T1, T2, T3, ..., Tn) can be obtained, and each collected temperature corresponds to a time point; then, a preset slope template is generated: according to the formula Y = k Using t + b as a linear template, a set of slope values ​​k to be tested needs to be pre-defined, for example, k = 0.1, 0.5, 1.0, 2.0, representing four change modes: "gradual rise," "medium rise," "rapid rise," and "sudden rise," respectively. For ease of calculation, b is fixed at 1 in this application. For each slope k, a corresponding linear template Y is generated. For example, for k = 0.5, the linear template is Y = 0.5. t + 1; this template is represented graphically as an ascending sloping line; the core of convolution calculation lies in systematically scanning temperature time-series data through a pre-defined linear template to quantify the degree of matching between the original time-series data and linear change patterns at different rates; each linear template has the same time resolution (i.e., the acquisition frequency) as the original time-series data; for example, if a template with a duration of 30 seconds is defined (corresponding to t from 1 to 30), for a slope k=0.5, the corresponding linear template is Y =0.5. 1+1, 0.5 2+1, ..., 0.5 30+1 (essentially a vector) is then used as a sliding window to perform pointwise convolution operations with the original time-series data (T1, T2, ..., Tn) corresponding to the temperature. Essentially, this is equivalent to aligning the template sequentially along the time axis and calculating the dot product between the template sequence and the corresponding data window at each alignment position, generating a continuous "response value sequence". When the actual change trend of the data in a certain time window closely matches the linear rising pattern of the template, a significant peak will appear in the calculated convolution response value. By traversing all preset slope templates (k = 0.1, 0.5, 1.0, 2.0), four response value curves can be obtained. Each curve reveals the matching situation between the temperature data and a specific rate of change. Therefore, by finding the peak value on each response value curve and its corresponding time position, as well as the specific slope template that produces the peak value, it is possible to accurately identify whether there is a pulse-like linear rise in the temperature data, when the rise begins, how long the rise lasts, and its best matching rate of change (effective slope).

[0031] In one or more possible embodiments, the step of convolving multiple linear templates with preset slopes with the time-series data of the second factor to obtain the response values ​​of each template at different times includes: for the original time-series data of each second factor, obtaining the integral time-series data and differential time-series data corresponding to each second factor; generating multiple linear template vectors based on multiple preset slope parameters, and flipping the non-zero elements in each linear template vector to obtain the corresponding flipped template vector; and performing convolution operations on any one of the original time-series data, the integral time-series data, and the differential time-series data with the flipped template vector corresponding to each slope to obtain the response values ​​of each slope template at different times. A sequence of response values ​​at different times; the sequence of response values ​​is used to characterize the shape matching degree between the original time-series data, the integral time-series data, and the differential time-series data and multiple linear template vectors, respectively; specifically, to further analyze the changing trend, the integral time-series data (characterizing the cumulative effect) and the differential time-series data (characterizing the instantaneous rate of change) corresponding to the original time-series data at that temperature can be calculated simultaneously. In subsequent steps, any one of the original time-series data, integral time-series data, or differential time-series data can be used as input data to be matched, and specific examples are not listed one by one; to improve computational efficiency, the non-zero elements in each linear template Y need to be sequentially flipped to obtain a flipped linear template. The specific calculation formula is as follows:

[0032] in, This represents the response value, used to characterize the degree of matching between time-series data and a linear template; This represents the value of the second factor at time point i. This represents the arithmetic mean of the data for the second factor X. This represents the response value of the second factor at time point i; This represents a linear template (i.e., the linear template Y mentioned above). This represents the arithmetic mean of the data in the linear template. This represents the new vector obtained by reversing the order of the non-zero elements in the linear template Y. `std` represents convolution; `std` represents standard deviation; `C` is the constant term.

[0033] Based on the above formula, multiple corresponding response value sequences can be obtained from the original time series data corresponding to temperature and the corresponding multiple flipped linear templates. At the same time, multiple corresponding response value sequences can also be obtained from the integral time series data corresponding to temperature, and multiple corresponding response value sequences can also be obtained from the differential time series data corresponding to temperature. In one or more possible embodiments, determining the start time, end time, and effective slope of the pulse change corresponding to each second factor based on the response values ​​of multiple linear templates corresponding to each second factor at different times includes: obtaining the maximum response value of each second factor; using the slope parameter corresponding to the maximum response value as the effective slope of the pulse change for each second factor; determining the start time of the pulse change based on the time position corresponding to the maximum response value; searching from the time position corresponding to the start time in the direction of increasing time to determine the time position where the response value first falls below a preset threshold as the end time of the pulse change; specifically, finding the maximum response value from the entire response value sequence corresponding to temperature, where the maximum response value is obtained by traversing all response value sequences generated by all preset slope templates. The global peak value selected is used as the reference. The slope of the linear template associated with this peak value is defined as the effective slope of the temperature pulse change. At the same time, the precise position of the peak value on the time axis is determined as the start time of this pulse change. After determining the start time, the algorithm uses this time point as a reference and searches in the sequence of response values ​​corresponding to the specific slope template that generated the maximum response value in the direction of increasing time. The search goal is to find the time point when the response value first drops below the preset response threshold. When the search finds the time point when the response value first drops below the preset response threshold, the corresponding time position is determined as the cutoff time of the pulse change. Alternatively, a simpler cutoff time calculation can be performed, for example, the cutoff time of the pulse change can be determined after the time period of the preset cutoff threshold from the determined start time of the pulse change.

[0034] Step 105: Determine the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time and fault time of the pulse change; wherein, the causal correlation is the quantified value of the change behavior of the original time series data corresponding to the second factor leading to this fault. In one or more possible embodiments, determining the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time, and fault time of the pulse change includes: determining four time differences based on the start time, end time, and fault time of the pulse change corresponding to each second factor; the four time differences include the advance and lag of the start time relative to the fault time, and the advance and lag of the end time relative to the fault time; determining preset processing rules based on preset intervals corresponding to the effective slope of each second factor; wherein different preset processing rules are associated with different weight parameter sets, the weight parameter sets being used to characterize the degree of influence of the four time differences on the fault time; and calculating the causal correlation of each second factor based on the maximum response value of each second factor, the four time differences, and the weight parameter sets associated with the selected preset processing rules.

[0035] Specifically, the calculation using the formula above will yield a maximum response value for each second factor. There will be a corresponding effective slope. The preset processing rules are determined based on the effective slope of each second factor. The specific formula is as follows:

[0036] in, This represents the causal correlation of the j-th second factor in the i-th fault scenario. Represents the minimum meaningful time resolution, for example = 1 second, Represents the weight parameters. dt represents the maximum response value, k is the effective slope, and dt represents the time difference. The following explanation uses the "temperature" factor as an example, where the effective slope k=1 and the pulse start time have been obtained. = Pulse cutoff time 8 seconds before fault = Maximum response value 2 seconds after failure When the fault occurred = 0 seconds, detailed explanation Calculation process: First, calculate the four time differences, using the following formulas:

[0037] in, (Start with advance measurement): At the moment of failure Previously, therefore = |0 - (-8)| = 8 seconds; (Initial lag) Not at the time of the failure After that, therefore = inf(infinity); (Lead time): (+2 seconds) Not at the moment the fault occurred Previously, therefore = inf; (Cumulative lag): At the moment of failure After that, so =|2 - 0| = 2 seconds; Based on the effective slope k=1, the first branch is selected, and the weight parameter set is known. , , and Four terms are calculated, and the largest one is selected as the final causal correlation corresponding to temperature.

[0038] Step 106: Determine the cause of this failure based on the causal correlation of each second factor.

[0039] In one or more possible embodiments, the second factor with the highest score can be directly identified as the single primary cause of the failure; or, a threshold can be set to identify several factors with the highest causal correlation scores (such as the top A factors) as a group of related causes leading to the failure.

[0040] In one or more possible embodiments, before selecting at least one second factor from the plurality of first factors according to a preset screening method, the method further includes: determining whether each first factor is included in a pre-established fault mode library; for each first factor included in the fault mode library, determining whether the time difference between the time when the original time series data corresponding to the first factor becomes abnormal and the time when the fault occurs is less than a preset time difference; if so, determining that the first factor is strongly correlated with the current fault; classifying the original time series data corresponding to the first factor according to a pre-established fault classification model, and determining whether the first factor is the cause of the current fault based on the classification result; specifically, by comparing real-time fault data with a prior knowledge base and model, efficient and structured initial screening of causes can be achieved; firstly, all identified first factors are compared with a pre-established fault mode library, which is based on historical fault cases and domain knowledge. The system establishes a set of primary factors that typically cause anomalies under various typical faults. If a primary factor exists in the pattern library corresponding to the current fault type, it becomes a target for further investigation. For the primary factor to be investigated, the system determines the chronological relationship and time difference between the time of the anomaly and the time of the fault. For example, in a laser link fault, if a sudden change in motor current data occurs within 5 seconds before the fault (assuming a preset time difference threshold of 10 seconds), then the time difference is less than the threshold, indicating that the change in motor current data is closely coupled with the fault in time, and can be identified as a strongly correlated factor. Finally, a pre-trained fault classification model (such as a convolutional neural network classification model) is called, and the original time-series data or extracted features of these strongly correlated primary factors are input into the model, with the motor stall / non-stall state as the output, to determine whether the fault belongs to the motor stall type. Based on this, it can be confirmed whether the primary factor is the root or direct cause of this fault.

[0041] In one or more possible embodiments, the method further includes: constructing a feature matrix of fault scenarios based on the causal correlations of each second factor under multiple fault scenarios; wherein each row of the feature matrix corresponds to a fault scenario, each column corresponds to a second factor, and the matrix elements are the causal correlations of each second factor under the corresponding fault scenario; taking a fault scenario in the feature matrix as a cluster, and using an agglomerative hierarchical clustering algorithm to iteratively merge multiple clusters to form classification results with different hierarchical structures; obtaining the keyword description of each cluster in each level; and updating the fault mode library based on the keyword descriptions and classification results; specifically, taking "temperature", "motor current", and "vibration amplitude" as three second factors as examples; first, constructing a feature matrix, where each fault scenario is described by a 3-dimensional feature vector, for example, including four fault scenarios, constructing a feature matrix of fault scenarios, arranging the feature vectors of all fault scenarios into a matrix, where each row represents a fault scenario, and each column represents the causal correlation of a second factor, as follows:

[0042] Each failure scenario is treated as an independent cluster: cluster 1 = {A}, cluster 2 = {B}, cluster 3 = {C}, cluster 4 = {D}; and the Euclidean distance between any two clusters is calculated according to the following formula:

[0043] Where D represents the Euclidean distance between the two clusters. Indicates the i-th cluster, Indicates the j-th cluster. This indicates that there are m causal correlations in the i-th cluster. This indicates that there are m causal correlations in the j-th cluster;

[0044] Find the two closest clusters (with the smallest Euclidean distance D) among all current cluster pairs. Assuming that cluster 3 and cluster 4 have the smallest distance (because their characteristics are both "high temperature, low other conditions", they are very similar), merge cluster 3 and cluster 4 into a new cluster, denoted as cluster 34. Then continue to merge cluster 1, cluster 2 and cluster 34 until there is only one cluster left, thus completing the clustering. Essentially, it is to classify all historical failure scenarios that failed to match known causes based on the overall feature pattern formed by the causal correlation of each second factor in each scenario, thereby automatically discovering potential, common failure types.

[0045] In one or more possible embodiments, during the clustering process, the term frequency of the name of the second factor is extracted according to the TF-IDF algorithm (Term Frequency–Inverse Document Frequency algorithm) to provide keyword descriptions for the classification results of different layers. The specific formula is as follows:

[0046] in, Term frequency measures the relative frequency or importance of word i in document j. This indicates the number of times word i appears in document j; This represents the total number of occurrences of all words in document j. The symbol k here is the summation index, which means traversing and accumulating all words.

[0047] In one or more possible embodiments, the specific process is as follows: Figure 2 As shown, it includes: Step 201: Initialize the name text of the second factor; In one or more possible embodiments, initializing the name text of the second factor is a preprocessing process that transforms the collected, unstructured raw name text into a standardized term list that can be used for subsequent TF-IDF algorithm calculation. The initialization process includes text cleaning, text segmentation, stop word filtering, etc.

[0048] Step 202: Divide the initialization samples into independent clusters according to the fault scenario; In one or more possible embodiments, initially, each fault scenario (i.e. each sample) is regarded as an independent cluster. If there are N historical fault scenarios or fault cases to be analyzed, N clusters are initialized. Each cluster contains only one fault scenario and the corresponding set of second factors. This step is the starting point of agglomerative hierarchical clustering and lays the foundation for subsequent iterative merging of similar clusters to form a hierarchical structure.

[0049] Step 203: Calculate the distance between clusters; Step 204: Merge the most similar clusters; Step 205: Obtain the name of the second factor in each cluster whose causal correlation is greater than a preset causal threshold; In one or more possible embodiments, for each cluster (e.g., the "motor-related faults" cluster), all historical fault scenarios (e.g., fault A and fault B) contained in the cluster are traversed. For each fault scenario, the names of the second factors whose causal correlation is greater than a preset causal threshold (e.g., 0.7) are extracted based on the causal correlation of each second factor.

[0050] Step 206: Label the word frequencies of the names of the first p second factors in each layer to automatically provide interpretable keyword descriptions for each layer of classification results generated during the agglomerative hierarchical clustering process. In one or more possible embodiments, "the classification result of each layer" refers to a series of nested, hierarchical classification results, from the finest to the coarsest granularity, generated by the agglomerative hierarchical clustering algorithm during the merging process. It is not a single, planar classification, but a multi-level classification system that can be expanded or merged layer by layer like a tree. The agglomerative hierarchical clustering process is bottom-up and progressively merged, with each merge producing a new, more macroscopic "layer." For example, the bottom layer (layer 0): Initially, each fault scenario forms its own category. If there are 5 fault scenarios (A, B, C, D, E), there are 5 clusters; this is the finest layer. The intermediate layers (layer 1, layer 2...): The algorithm finds A and B to be most similar and merges them into a new cluster AB; C and D are most similar and merge them into CD. At this point, the classification result becomes: AB, CD, E (3 clusters), which constitutes a new, more macroscopic layer. Then, the algorithm may find CD and E to be most similar and merge them into CDE, resulting in AB. CDE (2 clusters), which is a new, more macroscopic layer; the top layer (Nth layer): finally, all faults are merged into a single cluster ABCDE, which is the coarsest layer.

[0051] "Automatically providing interpretable keyword descriptions" refers to automatically generating corresponding keyword descriptions for each cluster contained in each level (or each possible merging stage) of the tree diagram. Using the TF-IDF algorithm, keywords are generated for each cluster at each level. For example, Level 0: Cluster "Fault A": keywords are "motor current"; Cluster "Fault B": keywords are "motor current" and "bearing vibration"; Cluster "Fault C": keywords are "optical temperature"; Cluster "Fault D": keywords are "optical temperature" and "heat sink temperature"; Level 1: Cluster "Motor-related faults (Fault AB)": collects the names of all secondary factors of faults A and B. After TF-IDF calculation, the representative keywords for this level are "motor" and "vibration"; Cluster "Temperature-related faults (Fault CD)": collects the names of the secondary factors of faults C and D, and the representative keywords for this level are "optical" and "temperature"; Top level: Cluster "All faults (ABCD)": keywords may be more generalized, such as "motor". "Temperature" indicates that both of these types of problems have appeared in the overall data; in this way, each layer and each cluster is given a clear, data-driven semantic label, which greatly enhances the understandability and engineering practicality of the clustering results, and makes it easier for operation and maintenance personnel to quickly grasp the common characteristics of fault modes at different granularities.

[0052] Step 207: Determine if the number of clusters is 1; if yes, proceed to step 208; otherwise, proceed to step 203.

[0053] Step 208: Output the clustering results.

[0054] In one or more possible embodiments, the relevant data and keyword descriptions generated by data mining are printed, and after expert diagnosis, a list of second factors and a fault log are formed. The list of second factors and the fault log are sorted out, and the second factor items of the laser terminal and the laser terminal fault classification model are supplemented and improved for updating the fault mode library and classification model.

[0055] The overall framework of this application is as follows: Figure 3 As shown, based on the working status of the laser link, considering the status of dual-end tracking, communication, bit error rate, etc., the time point of link failure is quickly screened. Then, using the fault mode library built with existing experience, combined with the fault classification model and the logical relationship between the event and the time of link failure, the fault is quickly diagnosed. The first factor is used for judgment. During the judgment, the first factor can be judged sequentially according to the preset priority. The time relationship between the trigger time of the first factor (the time when the data of the first factor is abnormal) and the time of failure is judged. When the time relationship is less than the time resolution (i.e., the preset time difference) corresponding to the first factor, it indicates that the interruption is mainly affected by this factor. In addition, the existing classification model is used to judge whether the fault belongs to a known category. For unidentified factors, data mining algorithms are used to sort out the main factors of failure. After expert judgment, a second factor or fault log is formed to supplement the fault mode library and update the classification model.

[0056] According to the laser communication link fault tracing method provided in this application, based on expert logs and second-factor analysis, the system can automatically iteratively update the fault model, significantly reducing manual maintenance costs and improving diagnostic efficiency. Through a novel correlation algorithm, while improving the analysis speed, it can accurately extract key features such as the start and end times and slope of data changes. The algorithm judges causal relationships based on the chronological order and can effectively handle complex scenarios with low temporal resolution. Finally, clustering and TF-IDF methods are used to classify common parameters and label keywords, enhancing the interpretability of fault mode induction.

[0057] This application also provides a laser communication link fault tracing device, such as... Figure 4 As shown, it includes: The data acquisition module 401 is used to acquire multiple first factors within a preset time period before and after the fault time, as well as the original time series data corresponding to each first factor. Factor filtering module 402 is used to filter at least one second factor from the plurality of first factors according to a preset filtering method; The response value determination module 403 is used to calculate the convolution of multiple linear templates with preset slopes with the time series data of the second factor for the original time series data of each second factor, so as to obtain the response value of each template at different times. The pulse change determination module 404 is used to determine the start time, end time and effective slope of the pulse change corresponding to each second factor based on the response values ​​of multiple linear templates corresponding to each second factor at different times. The causal correlation determination module 405 is used to determine the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time and fault time of the pulse change; wherein, the causal correlation is the quantified value of the change behavior of the original time series data corresponding to the second factor leading to the current fault. The fault cause determination module 406 is used to determine the cause of the fault based on the causal correlation of each second factor.

[0058] Since the device embodiments of the present invention correspond to the method embodiments described above, details not disclosed in the device embodiments can be referred to in the method embodiments described above, and will not be repeated in the present invention.

[0059] This application also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method for tracing the cause of laser communication link failure.

[0060] like Figure 5 As shown, the device includes a processor 501, a memory 502, a communication interface 503, and a bus 504. The processor 501, memory 502, and communication interface 503 are interconnected via the bus 504.

[0061] Processor 501 is configured to read instructions from memory 502 and execute them, so that at least one processor can execute the laser communication link fault tracing method provided in the above embodiments.

[0062] The memory 502 is used to store various instructions and programs for the laser communication link fault tracing method provided in the above embodiments.

[0063] Bus 504 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0064] Processor 501 can be a central processing unit (CPU), a network processor (NP), a graphics processing unit (GPU), or any combination of CPU, NP, and GPU. It can also be a hardware chip. The aforementioned hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0065] In addition, this application also provides a computer-readable storage medium, such as Figure 6 As shown, the computer storage medium stores a computer program that is used to cause the computer to perform any of the methods described in the above embodiments.

[0066] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) 601 and / or cache memory 602, and may further include read-only memory (ROM) 603.

[0067] The memory may also include a program / utility 605 having a set (at least one) of program modules 604, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0068] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0070] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0071] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0072] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method of laser communication link fault traceback, the method comprising: include: Acquire multiple primary factors and their corresponding original time-series data within a preset time period before and after the fault; At least one second factor is selected from the plurality of first factors according to a preset filtering method; For the original time series data of each second factor, the convolution of multiple linear templates with preset slopes with the time series data of the second factor is calculated to obtain the response value of each template at different times. Based on the response values ​​of multiple linear templates corresponding to each second factor at different times, determine the start time, end time, and effective slope of the pulse change corresponding to each second factor; Based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time and fault time of the pulse change, the causal correlation corresponding to each second factor is determined; wherein, the causal correlation is the quantified value of the change behavior of the original time series data corresponding to the second factor leading to this fault. The cause of this failure was determined based on the causal correlation of each of the second factors.

2. The method of claim 1, wherein, The step of convolving multiple linear templates with preset slopes with the time-series data of the second factor to obtain the response values ​​of each template at different times includes: For each second factor's original time series data, obtain the corresponding integral time series data and differential time series data for each second factor; Based on multiple preset slope parameters, multiple corresponding linear template vectors are generated, and the non-zero elements in each linear template vector are flipped to obtain the corresponding flipped template vector. Any one of the original time-series data, the integral time-series data, and the differential time-series data is convolved with the flipped template vector corresponding to each slope to obtain a sequence of response values ​​for each slope template at different times; the sequence of response values ​​is used to characterize the degree of shape matching between the original time-series data, the integral time-series data, and the differential time-series data and the multiple linear template vectors.

3. The method according to claim 1 or 2, characterized in that, The step of determining the start time, end time, and effective slope of the pulse change corresponding to each second factor based on the response values ​​of multiple linear templates corresponding to each second factor at different times includes: Obtain the maximum response value for each second factor, and use the slope parameter corresponding to the maximum response value as the effective slope of the pulse change for each second factor. The start time of the pulse change is determined based on the time position corresponding to the maximum response value; The search proceeds from the time position corresponding to the start time in the direction of increasing time, and the time position at which the response value first falls below a preset threshold is determined as the cutoff time of the pulse change.

4. The method according to claim 3, characterized in that, The determination of the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time, and fault time of the pulse change includes: Based on the start time of the pulse change, the end time of the pulse change, and the fault time corresponding to each second factor, four time differences are determined respectively; the four time differences include the advance and lag of the start time relative to the fault time, and the advance and lag of the end time relative to the fault time. Based on the preset interval corresponding to the effective slope of each second factor, a preset processing rule is determined; wherein, different preset processing rules are associated with different weight parameter groups, and the weight parameter groups are used to characterize the degree of influence of the four time differences on the fault moment. The causal correlation of each second factor is calculated based on the maximum response value of each second factor, the four time differences, and the weight parameter group associated with the selected preset processing rule.

5. The method according to claim 1, characterized in that, The time of the fault is determined in the following way: Monitor the tracking status, communication status, and error status of the laser terminal; When any of the following states—the tracking state, the communication state, or the bit error state—changes from normal to fault, the moment when the state changes is taken as the fault moment.

6. The method according to claim 1, characterized in that, Before selecting at least one second factor from the plurality of first factors according to a preset filtering method, the method further includes: Determine whether each primary factor is included in a pre-established failure mode library; For each first factor included in the fault mode library, determine whether the time difference between the time when the original time series data corresponding to the first factor becomes abnormal and the time when the fault occurs is less than a preset time difference. If so, determine that the first factor is strongly correlated with the current fault. The original time series data corresponding to the first factor are classified according to the pre-established fault classification model, and the first factor is determined as the cause of the fault based on the classification results.

7. The method according to claim 1, characterized in that, The step of selecting at least one second factor from the plurality of first factors according to a preset filtering method includes: The original time series data corresponding to multiple first factors within a preset time period before and after the fault time are filtered out, and the first factors that are unrelated to the fault and the first factors whose original time series data have not changed are removed.

8. The method according to claim 6, characterized in that, Also includes: Based on the causal correlations of each second factor under multiple failure scenarios, a feature matrix of failure scenarios is constructed; wherein, each row of the feature matrix corresponds to a failure scenario, each column corresponds to a second factor, and the matrix elements are the causal correlations of each second factor under the corresponding failure scenario. One fault scenario in the feature matrix is ​​taken as a cluster, and multiple clusters are iteratively merged using an agglomerative hierarchical clustering algorithm to form classification results with different hierarchical structures; Retrieve the keyword description for each cluster in each level; The fault mode library is updated based on the keyword descriptions and classification results.

9. A laser communication link fault tracing device, characterized in that, include: The data acquisition module is used to acquire multiple primary factors within a preset time period before and after the fault, as well as the original time-series data corresponding to each primary factor. The factor filtering module is used to filter at least one second factor from the plurality of first factors according to a preset filtering method; The response value determination module is used to calculate the convolution of multiple linear templates with preset slopes with the time series data of the second factor for the original time series data of each second factor, so as to obtain the response value of each template at different times. The pulse change determination module is used to determine the start time, end time, and effective slope of the pulse change corresponding to each second factor based on the response values ​​of multiple linear templates corresponding to each second factor at different times. The causal correlation determination module is used to determine the causal correlation of each second factor based on the effective slope corresponding to each second factor and the chronological relationship between the start time, end time and fault time of the pulse change; wherein, the causal correlation is the quantified value of the change behavior of the original time series data corresponding to the second factor leading to this fault. The fault cause determination module is used to determine the cause of the fault based on the causal correlation of each second factor.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

11. A computer storage medium, characterized in that, The computer storage medium stores a computer program that causes the computer to perform any one of the methods claimed in claims 1-8.