Navigation equipment fault anomaly detection method based on multilayer dynamic causal relationship network

By building a multi-layer dynamic causal relationship network, combining distributed stream processing and message queue technology, real-time automatic detection and positioning of general navigation equipment failures is achieved, solving the problems of difficulty in fault detection and insufficient human resources in the existing technology, and improving the fault processing efficiency and the safety and stability of general navigation equipment.

CN120067948AActive Publication Date: 2025-05-30THREE GORNAVIGATION AUTHORITY

Patent Information

Application Number
CN202510212005.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and locate the failure of navigation equipment, especially in the environment of complex changes in multi-source data, resulting in false alarms, underreports and insufficient human resources.

Method used

Using a method based on a multi-layer dynamic causal relationship network, a dynamic causal network of the macro layer and device layer of the system is constructed through the NOTEARS algorithm, combining distributed stream processing technology and the real-time message transmission mechanism of message queues, multi-source data is analyzed in real time, and automatic detection and positioning of faults is realized.

Benefits of technology

It improves the efficiency of fault handling, reduces false alarms and missed reports, reduces the need for manual intervention, and ensures the safe and stable operation of general navigation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067948A_ABST
    Figure CN120067948A_ABST
Patent Text Reader

Abstract

The invention relates to a navigation equipment fault anomaly detection method based on a multi-layer dynamic causal relationship network, and the method comprises the steps: collecting multi-source data from a navigation system, and constructing the multi-layer dynamic causal relationship network comprising a system macroscopic layer and an equipment layer; processing a real-time data stream of the navigation system by adopting a message queue, and extracting feature data used for real-time fault detection in the real-time data stream; analyzing the feature data by using a multi-layer dynamic causal relationship network, judging whether equipment abnormity exists or not, and if the judgment result is yes, further positioning the reason of the equipment abnormity by using the multi-layer dynamic causal relationship network; and if the equipment is abnormal, a fault alarm is given out. According to the invention, real-time automatic detection and positioning of equipment faults are realized, engineers can process the equipment faults conveniently, fault processing efficiency is improved, and navigation safety and stability are effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing and fault detection, and particularly relates to a method for detecting faults and anomalies of navigation equipment based on a multi-layer dynamic causality network. Background Art

[0002] At present, the national shipping demand is increasing continuously, and the dam ship locks with shipping functions are increasing day by day. As an important navigation dam in the Yangtze River Basin, the Three Gorges Dam has about 100 daily navigation ship trips. In the navigation scenario, it involves ship locks, control centers, hydrological monitoring equipment, communication equipment, etc. For the complex and large navigation system to complete the navigation task, various equipment needs to be highly interdependent and collaborative, so faults will inevitably occur. The rapid detection and location of faults need to be solved urgently.

[0003] Traditional fault detection methods mainly rely on manual work or partially rely on manual work. However, in the current situation of diversified equipment and intensive shipping, it is difficult to meet the requirements. A single navigation equipment fault may lead to system shutdown or serious economic losses. For example, the fault of the ship lock's pulling and releasing gates may cause ships to be stranded, delay the shipping plan, and affect the operation of the upper and lower ship locks. In recent years, in order to realize digital navigation management and assist the efficient and safe operation of navigation, the Three Gorges Navigation Administration has established a CCTV monitoring system, a network monitoring system, a dispatching system, etc., and uses these to monitor and manage navigation equipment in real time. Although the above systems can monitor and capture a large amount of operation data of navigation equipment in real time, they lack effective analysis methods and are difficult to extract key information and perform other tasks. It has become an inevitable requirement to make full use of the above multi-source data and combine dynamic causality network analysis to provide a new intelligent fault detection method.

[0004] The existing fault detection technical solutions mainly rely on preset rules and simple data analysis methods. In the Three Gorges navigation scenario, many fault detections still rely on the fault warning rules provided by equipment manufacturers. For example, based on the operating thresholds of equipment current, water level, container pressure, etc. The equipment alarms only when the threshold is exceeded. However, this threshold-based method is too dependent on manual setting and cannot cope with the complex performance of equipment under different working conditions, and is prone to false alarms and missed alarms. At the same time, it is impossible to accurately judge the faulty components on some monitoring data. In addition, the fault detection methods based on statistical analysis, such as simple statistical indicators like average value and standard deviation, although they can detect abnormal changes of some equipment, are difficult to capture the complex dynamic changes of multi-source data. Especially when facing various data types involved in navigation equipment, the detection effect of such methods is not comprehensive enough.

[0005] Existing fault location methods usually rely on the alarm signals provided by the above-mentioned fault detection, and infer the location where the fault occurs through the operation navigation equipment server logs, monitoring data of the equipment, and the experience and knowledge of technicians. Most fault location technologies require experienced technicians to check one by one. For example, when an abnormal alarm appears in the navigation declaration database, the technician first checks the alarm content, and then checks the status of each server in the navigation declaration system. When it is detected that the CPU of the navigation management server has a high utilization rate, relevant CPU index data will be checked one by one to further determine the problem. Through these relevant checks, it is finally concluded that it is caused by a resource-intensive query requested by the navigation control center. In addition, technicians need to record the problem situation and generate reports. This kind of fault location technology requires a huge amount of human resources and time consumption. In the relatively complex navigation system, if a large-scale equipment failure occurs, the human resources required will not be sufficient to support fault monitoring and location, resulting in shutdown or serious economic losses.

[0006] In recent years, machine learning algorithms have been gradually applied to the field of equipment fault detection. The algorithms use supervised learning or unsupervised learning methods to train the operation data of the equipment, learn the data performance of different fault types, and are used to detect faults after maturity. However, the application of such methods faces problems such as difficult data annotation, long model training time, and difficult to explain complex faults. Due to the dynamic nature of the navigation equipment operation environment, there are potential dynamic causal relationships between the multi-source data of the equipment, and the existing algorithms are relatively difficult to handle the dynamic causal relationships, and the accuracy and real-time performance of fault detection still need to be further improved. Summary of the Invention

[0007] The purpose of the present invention is to address the above problems, and provide a navigation equipment fault anomaly detection method based on a multi-layer dynamic causal relationship network. A multi-layer dynamic causal relationship network including a system macro layer and an equipment layer is constructed through the NOTEARS algorithm. The dynamic causal network of the system macro layer is used to identify the dependency relationships between the navigation subsystems, and the dynamic causal network of the equipment layer is used to identify the potential fault chains between the equipment; combined with the distributed stream processing technology and the real-time message transmission mechanism of the message queue, real-time analysis of the multi-source data of the navigation system is carried out to realize real-time detection and location of equipment faults, facilitate engineers to handle equipment faults, improve the efficiency of fault handling, and ensure the safe and stable navigation.

[0008] In order to achieve the above object, the technical solution provided by the present invention is as follows: A navigation equipment fault anomaly detection method based on a multi-layer dynamic causal relationship network, comprising the following steps: Step S1: Collect multi-source data from the navigation system, and the multi-source data includes navigation equipment data, navigation equipment server logs, navigation monitoring data, navigation execution work orders, and navigation ship information data; Step S2: preprocessing the multi-source data, wherein the preprocessing includes data standardization, conversion into time series, and synchronization of data timestamps; Step S3: construct a multi-layer dynamic causal relationship network, including constructing a dynamic causal network at the system macro layer and a dynamic causal network at the equipment layer; the dynamic causal network at the system macro layer is used to identify the dependencies between the various navigation subsystems, and the dynamic causal network at the equipment layer is used to process the nonlinear dynamic causal relationship of the navigation equipment and identify the potential fault chain between the equipment; Step S4: using a message queue to process the real-time data stream of the navigation system and extracting characteristic data therefrom for real-time fault detection; Step S5: using the dynamic causal relationship network obtained in step S3, analyzing the characteristic data obtained in step S4 to determine whether there is an equipment abnormality. If the determination result is yes, further using the dynamic causal relationship network to locate the cause of the equipment abnormality; Step S6: According to the judgment result of step S5, if there is any equipment abnormality, a fault alarm is issued.

[0009] Furthermore, the step S2 specifically includes the following sub-steps: S201: Use the z-score method to standardize the indicator data of different devices; S202: Use the triple standard deviation method to eliminate abnormal data to ensure data accuracy; S203: using the drain algorithm to extract the log data of the general aviation equipment server into events and convert them into time series, and using the TraceExtract algorithm to extract the key path of the call chain data and convert it into time series; S204: Synchronize the timestamps of each category of data so that the indicators, work orders, server logs, navigation monitoring and asset data of the navigation equipment can be synchronously analyzed for multivariate time series.

[0010] Preferably, the step S3 uses an improved NOTEARS algorithm to construct a multi-layer dynamic causal relationship network, and the improved NOTEARS algorithm improves the existing NOTEARS algorithm by including the following steps: 1) Time series data modeling: The core of dynamic causal relationship modeling is to process data with changing time dimensions and learn the causal structure that changes over time on the time axis; divide the time series data into time slices, each of which corresponds to observations over a certain period of time; introduce past observations as lagged variables and establish the causal relationship between the lagged variables and the current variables; 2) Expand the graph structure of the NOTEARS algorithm into a temporal structure. In the original optimization problem of the NOTEARS algorithm, the dynamic causal relationship is a static graph. Change the edges in the static graph to time-dependent edges to obtain a dynamic graph. Correspondingly, add temporal dimension constraints to ensure the acyclicity of the dynamic causal graph. Add cross-time acyclicity constraint conditions on the basis of the original constraints of the NOTEARS algorithm. Through the Notears algorithm, learn the dynamic causal relationships within each time slice, and then use a dynamic Bayesian network to add directed edges between time slices. The directed edges are used to represent the dependency relationships between lag variables and current variables. 3) Dynamic optimization. For each time slice, still use the original optimization strategy of Notears, that is, minimize the squared error or other loss functions.

[0011] Preferably, step 3 specifically includes the following sub-steps: Step S301: Feature selection and data transformation; At the macro level of the navigation system, according to business knowledge and data exploration results, divide time slices according to the time dimension, and select key features at the navigation subsystem level within each time slice; represent the feature data of each navigation subsystem as a feature matrix X (t) _subsystems; At the device level within the navigation subsystem, for each navigation subsystem, divide time slices according to the time dimension, and select relevant data of its internal devices within each time slice; construct a feature matrix X (t) _devices for each navigation subsystem respectively; Step S302: Initialize the weighted adjacency matrix and Lagrange multipliers of the dynamic causal network, and construct a smoothing function; At the macro level of the navigation system, select an initial weighted adjacency matrix W (t) 0 _subsystems and Lagrange multipliers α 0 ; construct a smoothing function h(W (t) a) within the time slice at the system level to enable it to encode acyclicity constraints; construct an acyclicity constraint function ht(W (t) a) between time slices at the system level to ensure the acyclicity constraint of the dynamic causal graph from the time dimension; At the device level within the navigation subsystem, initialize the weighted adjacency matrix W (t) 0 _devices and Lagrange multipliers β 0 ; construct a smoothing function h(W (t)b) to enable it to encode acyclicity constraints; construct an acyclic constraint function ht(W between time slices at the device level (t) b) for ensuring acyclicity constraints of the dynamic causal graph from the time dimension; Step S303: Optimize the weighted adjacency matrix to obtain a dynamic directed acyclic graph, i.e., a dynamic causal network; Use a numerical optimization method to minimize the NOTEARS objective function at the macro level of the navigation system and the device level within the subsystem, satisfying the acyclicity constraints within the time slices where h(W (t) a)=0 and h(W (t) b)=0 and the acyclicity constraint ht(W in the time dimension (t) a)=0 and ht(W (t) b)=0; at the same time, iteratively update the weighted adjacency matrix W (t) _devices at the system level and the weighted adjacency matrix W (t) _subsystems at the device level, and threshold the optimized matrices W (t) _devices and W (t) _subsystems to obtain a dynamic directed acyclic graph as the dynamic causal network; Step S304: Verify and correct the dynamic causal network; At the macro level of the navigation system, use the test data of the navigation subsystem to verify the dynamic causal network between the learned navigation subsystems. If there is an incorrect causal chain, perform manual intervention and correction on the incorrect causal chain to finally determine the dynamic causal network at the system level; At the device level within the navigation subsystem, use the device-level test data to verify the dynamic causal network between the devices within each navigation subsystem. By manually intervening and correcting unreasonable dynamic causal relationship chains, finally determine the dynamic causal network structure at the device level.

[0012] Preferably, in step S4, use the message queue Kafka to transmit the real-time data stream of the navigation system, use the stream processing engine Flink to process the real-time data stream, and use a sliding window to analyze and extract the feature data in the real-time data stream.

[0013] Furthermore, in step S5, based on the multi-layer dynamic causal relationship network obtained in step S3, analyze the feature data of the real-time data stream provided by the stream processing engine Flink, calculate the health score of the device according to the feature data, compare it with a predetermined threshold to determine whether the device state is abnormal, and trace the causal chain of the device anomaly through the directed acyclic graph of the multi-layer dynamic causal network to find the root cause of the device anomaly.

[0014] As another object of the present invention, the present invention provides a navigation equipment fault and anomaly detection system, including: Multi-source data acquisition and processing module: Real-time collect the operation data of navigation equipment through multiple sensors, CCTV systems, network monitoring systems, and dispatching systems, and perform standardized processing, outlier removal, and timestamp synchronization on the data to ensure the consistency and availability of the data; Dynamic causal relationship network construction module: Based on the improved NOTEARS algorithm, construct a multi-layer dynamic causal network including the system macro layer and the equipment layer, generate a directed acyclic graph, optimize and learn the causal chain through the gradient descent algorithm, and analyze the potential non-linear causal relationship between equipment; Real-time data stream processing module: Transmit real-time data through a message queue and perform real-time data stream processing using a stream processing engine. Extract and preprocess the feature data of the equipment sensor data stream and the log data stream, and extract important time series features; Fault detection and early warning module: Analyze the real-time data based on the multi-layer dynamic causal network, judge whether the current equipment state is abnormal, and infer the root cause of the fault through the causal chain. When an anomaly is detected, trigger an alarm and generate a fault report.

[0015] Compared with the prior art, the beneficial effects of the present invention include: 1) The present invention constructs a multi-layer dynamic causal relationship network including the system macro layer and the equipment layer. The dynamic causal network of the system macro layer is used to identify the dependence relationship between each navigation subsystem, and the dynamic causal network of the equipment layer is used to identify the potential fault chain between equipment; combined with the distributed stream processing technology of the Apache Flink stream processing engine and the real-time message transmission mechanism of the message queue Kafka, real-time analyze the multi-source data from different navigation equipment, and extract the feature data therein, realizing the real-time automatic detection and positioning of equipment faults, improving the fault handling efficiency, and effectively ensuring the safe and stable navigation.

[0016] 2) The improved NOTEARS algorithm, combined with the dynamic Bayesian network DBN, divides the multi-source data of the navigation system into time slices, introduces past observations as lag variables, establishes the causal relationship between lag variables and current variables, facilitates the processing of data with changing time dimensions, and learns the causal relationship changing with time; expands the graph structure of the existing NOTEARS algorithm into a time series structure, improves the static graph into a dynamic graph, and adds cross-time acyclicity constraints and within-time-slice acyclicity constraints, facilitating the processing of non-linear dynamic causal relationships of navigation equipment, identifying potential fault chains between equipment, realizing the causal chain reasoning of equipment faults, and being beneficial to quickly locating the root cause of equipment fault anomalies.

[0017] 3) By integrating multi-source data such as general aviation system sensors, logs, and network monitoring, and leveraging Kafka and Flink for real-time data transmission and processing, the present invention effectively solves the complexity and latency problems in multi-source data analysis for general aviation scenarios. The multi-layer dynamic causal network model is adopted to achieve real-time detection of equipment anomalies and accurate inference of the root causes of faults, significantly reducing false alarms and missed alarms and improving the accuracy of fault location.

[0018] 4) The general aviation equipment fault and anomaly detection system of the present invention relies on the multi-source data of the general aviation system and, based on the multi-layer dynamic causal relationship network, realizes the automated causal chain reasoning of equipment faults and anomalies and the dynamic adjustment of detection indicators, reducing the need for manual intervention and enhancing the fault response speed and processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below in conjunction with the drawings and embodiments.

[0020] Figure 1 It is a schematic diagram of the general aviation equipment fault detection system according to an embodiment of the present invention.

[0021] Figure 2 It is a schematic flowchart of the multi-layer dynamic causal relationship network for processing multi-source data streams according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] As Figure 1 shown, the general aviation equipment fault and anomaly detection method based on the multi-layer dynamic causal relationship network includes: Step S1: Collect multi-source data from the general aviation system; The general aviation equipment data includes real-time operation indicators of the equipment (such as temperature, pressure, speed, etc.), general aviation equipment server logs (such as equipment startup, shutdown, abnormal events, etc.), general aviation monitoring data (such as sensor data, CCTV system, environmental information, etc.), general aviation execution work order data (such as equipment ID, maintenance records, processing status, etc.), and general aviation ship information (such as ship number, declaration time, arrival at anchor time, etc.), and store them as CSV files for easy reading by the algorithm.

[0023] At the same time, since the general aviation equipment server logs contain non-numerical data, they can be mapped to different numerical data for storage, as shown in Table 1.

[0024] Table 1 General aviation equipment server data

[0025] The sample of the index data is shown in Table 2.

[0026] Table 2 Index data

[0027] Step S2: preprocessing the multi-source data, wherein the preprocessing includes data standardization, conversion into time series, and synchronization of data timestamps; The present invention uses the z-score method to standardize the indicators of different devices. The z-score method formula is: is the average value, is the standard deviation.

[0028] ; For example, if a hydrological monitoring device monitors a water level of 70m, the average water level in historical data is 65m, and the standard deviation is 5m, then the z-score is calculated as: ; The three-times standard deviation method is used to eliminate abnormal data to ensure data accuracy. For example, water levels exceeding three times the standard deviation will be considered abnormal and will be eliminated before further processing.

[0029] The drain algorithm is used to extract the log data of the general aviation equipment server into events and convert them into time series. The TraceExtract algorithm is used to extract the key path of the call chain data and convert it into a time series to facilitate the subsequent dynamic causal network learning of the call relationship between services.

[0030] The TraceExtract algorithm takes the root span of the trace as input and computes the critical path starting from its end. All child spans of the root span are sorted in descending order according to their end time.

[0031] The pseudo code is: 1traceextract(root) 2 If root.children is None or root.kind is Producer, then 3 return[root] 4 end if 5 if root.kind is client, then 6 return traceextract(root.chlid) 7 end if 8 path ← [root.end] 9 children ← Sort root.children in descending order by end time 10 lastChild ← children[0] 11 path ← traceextract (lastChild).extend(path) 12 For each span in children[1 :], do 13 If span.endTime < lastChild.startTime, then 14 path ← traceextract (span).extend(path) 15 lastChild ← span 16 End if 17 end for 18 path ← [root.start].extend(path) 19 return path 20 end Meanwhile, align the timestamps of various data sources so that the indicators, work orders, server logs, navigation monitoring, and asset data of navigation equipment can be synchronously analyzed for multivariate time series.

[0032] Step S3: Use the improved NOTEARS algorithm to construct a multi-layer dynamic causal relationship network, including constructing a dynamic causal network at the system macro level and a dynamic causal network at the equipment level; the dynamic causal network at the system macro level is used to identify the dependency relationships between navigation subsystems, and the dynamic causal network at the equipment level is used to process the non-linear dynamic causal relationships of navigation equipment and identify potential fault chains between equipment, as Figure 2 shown.

[0033] The improvements of the improved NOTEARS algorithm over the existing NOTEARS algorithm include the following steps: 1) Time series data modeling. The core of dynamic causal relationship modeling lies in processing data with a changing time dimension and learning the causal structure that changes over time on the time axis; divide the time series data into time slices t 1 , t 2 ……t T , and each time slice corresponds to the observations for a certain period of time; introduce past observations as lag variables, such as X t-3 , X t-2 that may have a causal impact on the current X t period.

[0034] 2) Expand the graph structure of the NOTEARS algorithm into a temporal structure. In the original optimization problem of the NOTEARS algorithm, the dynamic causal relationship is a static graph. To change it into a dynamic graph, the edges in the causal graph need to be changed into time-dependent edges, that is, the edges from time t−1 to time t; correspondingly, add temporal dimension constraints to ensure the acyclicity of the dynamic causal graph; on the basis of the original constraints, add the following cross-time acyclicity constraint conditions: ; where represents the dynamic causal relationship of the current time slice, represents the dynamic causal relationship of the previous time slice; d represents the number of variables in the causal relationship network; ht( ) represents the dynamic acyclicity constraint function at time t; Tr( ) represents the matrix trace operation function.

[0035] Through the Notears algorithm, learn the dynamic causal relationship within each time slice, and then use the framework of the Dynamic Bayesian Network (DBN) to add directed edges between time slices. These directed edges represent the dependence relationship between lag variables and current variables.

[0036] 3) Dynamic optimization. For each time slice, still use the original optimization strategy of Notears, that is, minimize the squared error or other loss functions to learn the dynamic causal relationship matrix W (t) The optimization objective is: ; where is the lag variable matrix containing multiple time slices; is the regularization term; loss( ) represents the loss function; ( ) represents the regularization function; represents the regularization coefficient.

[0037] In the embodiment, step 3 specifically includes the following sub-steps: Step S301: Feature selection and data transformation; At the macro level of the navigation system, according to business knowledge and data exploration results, divide time slices according to the time dimension, and select key features at the navigation subsystem level within each time slice (such as subsystem performance indicators, interdependent data traffic, etc.); represent the feature data of each navigation subsystem (such as the response time, error rate, data throughput, etc.) of the subsystem as a feature matrix X (t) _subsystems, where each row represents a time point and each column represents the features of a subsystem. Example features include the declaration subsystem status, navigation monitoring subsystem status, ship monitoring data, lock execution status, etc.

[0038] At the device level within the navigation subsystem, for each navigation subsystem, time slices are divided according to the time dimension, and relevant data of its internal devices are selected within each time slice, such as navigation device sensor data, server operation logs, fault records, server internal service call chains, etc.; a feature matrix X (t) _devices is constructed for each navigation subsystem, where each row represents a time point and each column represents a feature of a device. Example features include: real-time device metrics such as temperature, pressure, network uplink and downlink rates, packet loss rate, memory pressure, etc., work order data such as maintenance records and fault types, and navigation device server logs such as device exception events and device start / stop records.

[0039] Data transformation: Convert the selected features into the feature matrix X (t) , for example, server-related features: X (t) = ; Step S302: Initialize the weighted adjacency matrix and Lagrange multipliers of the dynamic causal network, and construct a smoothing function.

[0040] At the macro level of the navigation system, select the initial weighted adjacency matrix W (t) 0 _subsystems and Lagrange multipliers α 0 ; Construct a smoothing function h(W (t) a) such that it can encode the acyclicity constraint; construct ht(W (t) a) to ensure the acyclicity constraint of the dynamic causal graph in the time dimension; At the device level within the navigation subsystem, initialize the device-level weighted adjacency matrix W (t) 0 _devices and Lagrange multipliers β 0 ; Construct a smoothing function h(W (t) b) such that it can encode the acyclicity constraint; construct ht(W (t) b) to ensure the acyclicity constraint of the dynamic causal graph in the time dimension; Step S303: Optimize the weighted adjacency matrix to obtain a dynamic directed acyclic graph, i.e., a dynamic causal network.

[0041] Use a numerical optimization method such as L-BFGS to minimize the NOTEARS objective function at the macro level of the navigation system and at the device level within the subsystem, satisfying the acyclicity constraint within the time slices where h(W (t) a)=0, h(W (t) b)=0 and the acyclicity constraint in the time dimension ht(W (t) a)=0, ht(W(t) b) = 0; At the same time, iteratively update W (t) _devices and W (t) _subsystems, and threshold the optimized matrix W (t) _devices and W (t) _subsystems to obtain a dynamic directed acyclic graph DAG as the dynamic causal network; Step S304: Verify and correct the dynamic causal network.

[0042] At the macro level of the navigation system, use the test data of the navigation subsystems to verify the dynamic causal network W (t) _subsystems between the learned subsystems. If there are error chains, perform manual intervention and correction on the wrong causal chains to finally determine the dynamic causal network.

[0043] At the device level within the navigation subsystems, use device-level test data to verify the dynamic causal network W (t) _devices between the devices within each subsystem. Through manual intervention and correction of unreasonable dynamic causal relationship chains, finally confirm the dynamic causal network structure at the device level.

[0044] Give an example of W (t) : If there are three feature variables, network delay (delay) of the navigation declaration system, bandwidth usage (bandwidth) of the navigation declaration system, and packet loss rate (packet loss) of the navigation declaration system.

[0045] W (t) = ; W (t) 12 = 0.3 indicates that the network delay of the navigation declaration system has a positive causal impact on the bandwidth usage of the navigation declaration system.

[0046] W (t) 23 = 0.4 indicates that the bandwidth usage of the navigation declaration system has a positive causal impact on the packet loss rate of the navigation declaration system.

[0047] W (t) 31 = -0.3 indicates that the packet loss rate of the navigation declaration system has a negative causal impact on the network delay of the navigation declaration system.

[0048] Directed acyclic graph DAG: Navigation declaration system delay → Navigation declaration system bandwidth → Navigation declaration system packet loss rate.

[0049] Step S4: Use a message queue to process the real-time data stream of the navigation system and extract the feature data for real-time fault detection; Each subsystem of the navigation system transmits real-time data through Kafka. For example, Kafka topics such as network_metrics and device_metrics are set for transmitting network monitoring data and device metric data respectively. As a message queue, Kafka can ensure that real-time data collected from different data sources (such as the navigation network monitoring system and hydrological monitoring device sensors) can be efficiently and timely transmitted to the processing system. Kafka producers send these real-time data streams from the data sources to the consumer side, and Flink, as a Kafka consumer, will subscribe to these topics to receive real-time data streams.

[0050] Flink processes the received real-time data stream. It performs preprocessing, feature extraction, and real-time analysis of the data that conforms to the steps in S2. The z-score is used to standardize the network latency and bandwidth usage data to ensure data consistency and accuracy. Subsequently, key features are calculated through the sliding window technique, such as the maximum and average values of network latency, and the change in packet loss rate, and these features are extracted for subsequent analysis.

[0051] Step S5: Use the dynamic causal relationship network obtained in step S3 to analyze the feature data obtained in step S4 to determine whether there is a device anomaly. If the judgment result is yes, further use the multi-layer dynamic causal relationship network to locate the cause of the device anomaly.

[0052] In an embodiment, based on feature data such as network latency, bandwidth usage, and device metrics, it is evaluated whether the current system state is abnormal. If the features in the real-time data stream exceed the predetermined threshold, the system will perform further analysis through the dynamic causal network model, generate a judgment result and be confirmed by humans. If there are relevant changes, the threshold will be dynamically adjusted. This process includes calculating the health score of the device, judging whether there is an anomaly in the device according to the score, thereby triggering an alarm and generating a detailed fault report. The device score is calculated through the following formula: ; where, W (t) is the weighted adjacency matrix, X (t) is the feature matrix, is the activation function.

[0053] According to the DAG graph, trace the causal chain of the abnormal variable to find the possible upstream causes. For example, if the packet loss rate of the navigation declaration system increases abnormally, the causal chain in the DAG graph will indicate that the packet loss rate of the navigation declaration system is affected by network latency and bandwidth usage. Trace back from the packet loss rate of the navigation declaration system, check the status of the network latency and bandwidth usage of the navigation declaration system, and further find the possible causes of the anomaly. Combine the actual data to determine the location of the fault and perform further processing.

[0054] Step S6: According to the judgment result of step S5, if there is equipment abnormality, a fault alarm is issued.

[0055] The implementation effect shows that the present invention realizes real-time and efficient detection of equipment faults and root cause location by constructing a multi-layer dynamic causal relationship network including the system macro layer and the equipment layer, improves the fault handling efficiency, and can escort the safe operation of the shipping system.

[0056] In another embodiment of the present invention, a navigation equipment fault detection system based on dynamic causal network analysis is provided, as Figure 1 shown. The system includes: Multi-source data acquisition and processing module: Collect equipment operation data through multiple sensors and systems, including sensor data of navigation equipment such as hydrological monitors, navigation monitoring data (such as CCTV systems), data of the navigation declaration system server, network monitoring data, log data of the dispatching system, manually input data by staff, etc. The data types are complex, including structured time series data, unstructured log data, and call chain data between systems. Given the data input format and perform data preprocessing according to the format. Among them, log data can be extracted as events using the drain algorithm and sorted in time series; call chain data uses the TraceExtract algorithm to extract the critical path and perform time sorting to facilitate the subsequent learning of the call relationship between services in the dynamic causal network.

[0057] Dynamic multi-layer causal relationship network construction module: Construct a dynamic causal network between subsystems at the macro level, identify the dependency relationships between subsystems, and then establish a dynamic causal network for the equipment inside each subsystem. Use the improved dynamic NOTEARS algorithm to construct the dynamic causal relationship network. The improved dynamic NOTEARS algorithm on a single time slice transforms the causal structure learning problem into a differentiable optimization problem and ensures that the learned graph is a directed acyclic graph through an acyclic constraint. The graph network generated by the improved dynamic NOTEARS algorithm may have errors, and a dynamic feedback mechanism is introduced to improve the accuracy of the graph network construction. In the navigation equipment management scenario, the dynamic causal network is used to analyze the causal chain of equipment faults. The input data of the improved dynamic NOTEARS algorithm is the sensor data of the equipment, the time-serialized log data, and the call chain data. The algorithm optimizes and learns the dynamic causal relationship through the gradient descent algorithm, thereby generating a dynamic causal relationship model between the equipment. The generated dynamic causal network can not only reflect the correlation between devices, but also adjust the network structure through intervention operations, and finally be used for the accurate positioning of the fault source.

[0058] Real-time Streaming Data Processing Module: To perform real-time detection of navigation equipment failures, the system transmits real-time data through Kafka, and Flink performs real-time processing and analysis of the data stream. In practical applications, the navigation subsystem server transmits the collected device sensor data stream, log data stream, and call chain data stream through Kafka. Flink preprocesses the above data streams, extracts important time series features, and the anomaly detection server receives the above data streams.

[0059] Fault Detection and Early Warning Module: The navigation equipment anomaly detection server combines the dynamic causal network model to perform real-time fault detection of navigation equipment and determine whether the current system state is abnormal. When an anomaly is detected, the dynamic causal network is used to infer the root cause of the fault, and an alarm is triggered through the system, and the results are returned to each subsystem device to prompt the maintenance personnel to handle it promptly.

Claims

1. A method for detecting abnormal faults in navigation equipment based on a multi-layer dynamic causal relationship network, characterized in that: The following steps are involved: Step S1: Collect multi-source data from the navigation system, wherein the multi-source data includes navigation equipment data, navigation equipment server log, navigation monitoring data, navigation execution work order and navigation ship information data; Step S2: preprocessing the multi-source data, wherein the preprocessing includes data standardization, conversion into time series, and synchronization of data timestamps; Step S3: construct a multi-layer dynamic causal relationship network, including constructing a dynamic causal network at the system macro layer and a dynamic causal network at the equipment layer; the dynamic causal network at the system macro layer is used to identify the dependencies between the various navigation subsystems, and the dynamic causal network at the equipment layer is used to process the nonlinear dynamic causal relationship of the navigation equipment and identify the potential fault chain between the equipment; Step S4: using a message queue to process the real-time data stream of the navigation system and extracting characteristic data therefrom for real-time fault detection; Step S5: using the dynamic causal relationship network obtained in step S3, analyzing the characteristic data obtained in step S4 to determine whether there is an equipment abnormality. If the determination result is yes, further using the dynamic causal relationship network to locate the cause of the equipment abnormality; Step S6: According to the judgment result of step S5, if there is any equipment abnormality, a fault alarm is issued.

2. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 1 is characterized in that: The step S2 specifically includes the following sub-steps: S201: Use the z-score method to standardize the indicator data of different devices; S202: Use the triple standard deviation method to eliminate abnormal data to ensure data accuracy; S203: Use the drain algorithm to extract the log data of the general aviation equipment server into events and convert them into time series, extract the key path of the call chain data, and convert it into time series; S204: Synchronize the timestamps of each category of data so that the indicators, work orders, server logs, navigation monitoring and asset data of the navigation equipment can be synchronously analyzed for multivariate time series.

3. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 2 is characterized in that: The step S3 uses the improved NOTEARS algorithm to construct a multi-layer dynamic causal relationship network. The improved NOTEARS algorithm improves the existing NOTEARS algorithm by including the following steps: 1) Time series data modeling: The core of dynamic causal relationship modeling is to process data with changing time dimensions and learn the causal structure that changes over time on the time axis; divide the time series data into time slices, each of which corresponds to observations over a certain period of time; introduce past observations as lagged variables and establish the causal relationship between the lagged variables and the current variables; 2) Expand the graph structure of the NOTEARS algorithm to a time-series structure. The dynamic causal relationship in the original optimization problem of the NOTEARS algorithm is a static graph. The edges in the static graph are changed to time-dependent edges to obtain a dynamic graph. Accordingly, add time dimension constraints to ensure the acyclicity of the dynamic causal graph. Add cross-time acyclic constraints on the basis of the original NOTEARS algorithm constraints. The dynamic causal relationship within each time slice is learned through the Notears algorithm, and then a dynamic Bayesian network is used to add directed edges between time slices. The directed edges are used to represent the dependency relationship between the lagged variables and the current variables. 3) Dynamic optimization: For each time slice, Notears' original optimization strategy is still used, that is, minimizing the square error or other loss function.

4. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 3 is characterized in that: In step 2), the cross-time acyclic constraint condition is: ; in Represents the dynamic causal relationship of the current time slice, represents the dynamic causal relationship of the previous time slice; d represents the number of variables in the causal network; ht( ) represents the dynamic acyclic constraint function at time t; Tr( ) represents the matrix trace operation function.

5. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 3 or 4, characterized in that: In step 3), the optimization objective for learning the dynamic causal relationship matrix for each time slice is: ; in is a lagged variable matrix containing multiple time slices; is the regularization term; loss( ) represents the loss function; ( ) represents the regularization function; represents the regularization coefficient.

6. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 5 is characterized in that: The step 3 specifically includes the following sub-steps: Step S301: feature selection and data conversion; At the macro level of the general aviation system, based on business knowledge and data exploration results, time slices are divided according to the time dimension, and key features at the general aviation subsystem level are selected in each time slice; The characteristic data of each navigation subsystem is represented as the characteristic matrix X (t) _subsystems; At the equipment level within the general aviation subsystem, for each general aviation subsystem, time slices are divided according to the time dimension, and relevant data of its internal equipment is selected in each time slice; the equipment-level feature matrix X is constructed for each general aviation subsystem. (t) _devices; Step S302: Initialize the weighted adjacency matrix and Lagrange multipliers of the dynamic causal network, and construct a smoothing function; At the macro level of the navigation system, the initial weighted adjacency matrix W is selected for the optimization problem of the improved NOTEARS algorithm. (t) 0_subsystems and Lagrange multiplier α0; construct the smooth function h(W (t) a), so that it can encode acyclic constraints; construct the acyclic constraint function ht(W between system-level time slices) (t) a) It is used to ensure the acyclic constraint of the dynamic causal graph from the time dimension; At the equipment level in the navigation subsystem, initialize the equipment-level weighted adjacency matrix W (t) 0_devices and Lagrange multiplier β0; construct the smooth function h(W (t) b) to encode the acyclic constraint; construct the acyclic constraint function ht(W between device-level time slices) (t) b) It is used to ensure the acyclicity constraint of the dynamic causal graph from the time dimension; Step S303: Optimize the weighted adjacency matrix to obtain a dynamic directed acyclic graph, i.e., a dynamic causal network; Numerical optimization methods are used to minimize the NOTEARS objective function at the macro level of the general aviation system and the equipment level within the subsystem to meet h(W (t) a)=0, h(W (t) b)=0 in the time slice and the acyclic constraint in the time dimension ht(W (t) a)=0、ht(W (t) b)=0; at the same time, iteratively update the system-level weighted adjacency matrix W (t) _devices and the device-level weighted adjacency matrix W (t) _subsystems, the optimized matrix W (t) _devices and W (t) _subsystems is thresholded to obtain a dynamic directed acyclic graph as a dynamic causal network; Step S304: verifying and correcting the dynamic causal network; At the macro level of the general aviation system, the dynamic causal network between the learned general aviation subsystems is verified using the test data of the general aviation subsystems. If an incorrect causal chain occurs, manual intervention and correction are performed on the incorrect causal chain to ultimately determine the dynamic causal network at the system level. At the equipment level within the general aviation subsystem, device-level test data is used to verify the dynamic causal network between the internal devices of each general aviation subsystem. Through manual intervention and correction of unreasonable dynamic causal relationship chains, the dynamic causal network structure at the equipment level is ultimately determined.

7. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 1, 2, 3, 4 or 6, characterized in that: In step S4, the message queue Kafka is used to transmit the real-time data stream of the navigation system, the stream processing engine Flink is used to process the real-time data stream, and the sliding window is used to analyze and extract the characteristic data in the real-time data stream.

8. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 7 is characterized in that: In step S5, based on the multi-layer dynamic causal network obtained in step S3, the characteristic data of the real-time data stream provided by the stream processing engine Flink is analyzed, the health score of the device is calculated according to the characteristic data, and compared with the predetermined threshold to determine whether the device status is abnormal, and the causal chain of the device abnormality is tracked through the directed acyclic graph of the multi-layer dynamic causal network to find the root cause of the device abnormality.

9. The method for detecting abnormal faults of navigation equipment based on a multi-layer dynamic causal relationship network according to claim 8 is characterized in that: The step S5 calculates the health score of the device by the following formula: ; in, represents the device health score, W (t) is the weighted adjacency matrix, X (t) is the feature matrix, is the activation function.

10. The system of the navigation equipment fault anomaly detection method according to claim 1 or 2 or 3 or 4 or 6 or 8 or 9, characterized in that: The system comprises: Multi-source data acquisition and processing module: Through multiple sensors, CCTV systems, network monitoring systems and dispatching systems, the operation data of navigation equipment is collected in real time, and the data is standardized, outliers are removed, and timestamps are synchronized to ensure data consistency and availability; Dynamic causal network construction module: Based on the improved NOTEARS algorithm, a multi-layer dynamic causal network including the system macro layer and the device layer is constructed to generate a directed acyclic graph. The causal chain is optimized and learned through the gradient descent algorithm to analyze the potential nonlinear causal relationship between devices. Real-time data stream processing module: transmits real-time data through message queues, and uses stream processing engines to process real-time data streams, extracts features and preprocesses device sensor data streams and log data streams, and extracts important time series features; Fault detection and early warning module: Analyzes real-time data based on a multi-layer dynamic causal network to determine whether the current equipment status is abnormal, and infers the root cause of the fault through the causal chain. When an abnormality is detected, an alarm is triggered and a fault report is generated.

Citation Information

Patent Citations

  • Ship equipment fault diagnosis method and equipment based on multi-dimensional feature knowledge extraction

    CN112810772A

  • Micro-service fault diagnosis method and system

    CN115640159A

  • KPIs abnormal root cause positioning method based on LSTM-CNN causal discovery model

    CN116302647A

  • Fault detection method based on iterative depth time sequence causal discovery

    CN117290786A

  • Network fault adaptive detection system based on machine learning

    CN119420627A

Cited By

  • Deep well multistage intelligent ventilation control method and system

    CN120315292A