Causally driven counterfactual data generation and its application method, device and storage medium in anomaly detection
Through the causal-driven counterfactual data generation method, the causal relationship of system monitoring variables is mined and counterfactual data that conforms to the causal relationship is generated, which solves the adaptability problem of the anomaly detection model in the absence of historical data and improves the model's generalization and recognition capabilities.
Patent Information
- Application Number
- CN202411546090.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Existing technologies for anomaly detection in industrial systems and equipment lack sufficient historical monitoring data and fault data, which causes the anomaly detection model to fail in new environments and cannot accurately identify abnormal conditions.
Through a causal-driven approach, the causal relationship between system monitoring variables is mined, the graph deconvolution network model is used to extract component-level degradation state representation, and the causal adversarial generative network is used to generate counterfactual data that conforms to the causal relationship for data enhancement of the anomaly detection model.
The generalization ability and stability of the anomaly detection model have been improved, and it can adapt to environmental changes and system upgrades, identify abnormal conditions not covered in historical monitoring data, reduce false alarm rates and missed alarm rates, and improve the recognition ability and robustness of the model.
Smart Images

Figure CN119442104B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of anomaly detection technology, and in particular to a method, device, and storage medium for generating causal-driven counterfactual data and its application in anomaly detection. Background Art
[0002] With the rapid development of modern information technology, the process of industrial digitization has accelerated, enabling the effective collection of monitoring data reflecting the operating status of industrial systems and equipment. This has led to the widespread application of data-driven methods in anomaly detection in industrial systems and equipment. Among these, supervised learning models with excellent fitting performance have attracted much attention and achieved significant results.
[0003] However, building high-performance anomaly detection models relies on massive amounts of clean, complete historical monitoring data, a requirement that is difficult to meet in real-world applications. For high-reliability industrial systems and equipment, the collected historical monitoring data is often patchy and failure data is scarce. Directly applying current data-driven methods often fails to extract true failure mechanisms and characteristics from historical monitoring data, causing anomaly detection models to fail in new environments and operating conditions. Summary of the Invention
[0004] In view of this, the present disclosure proposes a causal-driven counterfactual data generation method, device, and storage medium for its application in anomaly detection.
[0005] According to one aspect of the present disclosure, a method for generating causal-driven counterfactual data and applying the same to anomaly detection is provided, the method comprising:
[0006] determining a causal relationship between system monitoring variables based on the system monitoring variables, wherein the system monitoring variables include a plurality of monitored monitoring variables related to the health status of the system;
[0007] Determining a component-level degradation state representation based on the system monitoring variables and the causal relationship through a preset neural network model, wherein the component-level degradation state representation is used to indicate the health state of each of the plurality of components of the system;
[0008] According to the component-level degradation state representation, counterfactual data that conforms to the causal relationship is generated, and the counterfactual data is used for data enhancement of an anomaly detection model, and the anomaly detection model is used to perform anomaly detection on the system.
[0009] In one possible implementation, the type of the system monitoring variable includes at least one of a system monitoring parameter, a system operating condition variable, and a system operating mode. The system monitoring parameter includes variables collected by a data collection device set at a specified location of the system. The system operating condition variable includes external condition variables and / or environmental variables of the system. The system operating mode is the operating mode adopted by the system when implementing its functions.
[0010] In another possible implementation, determining the causal relationship between the system monitoring variables based on the system monitoring variables includes:
[0011] According to the system monitoring variables, a fast causal inference (FCI) algorithm based on a priori constraints is adopted to determine the causal relationship between the system monitoring variables.
[0012] In another possible implementation, the prior constraints include:
[0013] There is no direct causal relationship between any one of the system operating condition variables and any one of the system operating modes, and / or there is no direct causal relationship between any two of the system operating condition variables.
[0014] In another possible implementation, the neural network model is a graph deconvolution network model, and determining the component-level degradation state representation through a preset neural network model based on the system monitoring variables and the causal relationship includes:
[0015] Inputting the system monitoring variables and the causal relationship into the preset graph deconvolution network model, and outputting the component-level degradation state representation;
[0016] The graph deconvolution network model is used to extract the component-level degradation state representation from the system monitoring variables and the causal relationship.
[0017] In another possible implementation, generating counterfactual data that conforms to the causal relationship based on the component-level degradation state representation includes:
[0018] Generate counterfactual data that conforms to the causal relationship using a preset Causal Generative Adversarial Network (CGAN) model based on the component-level degradation state representation, the system operating condition variables, and the system operating mode;
[0019] The counterfactual data includes simulated system monitoring parameters, actual system operating condition variables, actual system operating mode and simulated system fault label data.
[0020] In another possible implementation, the CGAN model includes a generator and a discriminator. The generation of counterfactual data that conforms to the causal relationship through a preset causal adversarial generative network (CGAN) model based on the component-level degradation state representation, the system operating condition variables, and the system operating mode includes:
[0021] generating candidate counterfactual data by the generator according to the component-level degradation state representation, the system operating condition variables, and the system operating mode;
[0022] In a case where the discriminator determines that the candidate counterfactual data satisfies the causal relationship, the candidate counterfactual data is output as the counterfactual data.
[0023] In another possible implementation, the method further includes:
[0024] Adding the counterfactual data to a training dataset of the anomaly detection model;
[0025] The anomaly detection model is trained according to the added training data set to obtain the data-enhanced anomaly detection model.
[0026] According to another aspect of the present disclosure, a device for generating causal-driven counterfactual data and applying the same to anomaly detection is provided, the device comprising:
[0027] a first determining module, configured to determine a causal relationship between system monitoring variables based on system monitoring variables, wherein the system monitoring variables include a plurality of monitored monitoring variables related to a health status of the system;
[0028] a second determining module, configured to determine, based on the system monitoring variables and the causal relationship, a component-level degradation state representation using a preset neural network model, wherein the component-level degradation state representation is used to indicate the health state of each of the plurality of components of the system;
[0029] A generation module is used to generate counterfactual data that conforms to the causal relationship based on the component-level degradation state representation, and the counterfactual data is used for data enhancement of anomaly detection model, and the anomaly detection model is used to perform anomaly detection on the system.
[0030] In one possible implementation, the type of the system monitoring variable includes at least one of a system monitoring parameter, a system operating condition variable, and a system operating mode. The system monitoring parameter includes variables collected by a data collection device set at a specified location of the system. The system operating condition variable includes external condition variables and / or environmental variables of the system. The system operating mode is the operating mode adopted by the system when implementing its functions.
[0031] In another possible implementation, the first determining module is further configured to:
[0032] According to the system monitoring variables, a FCI algorithm based on prior constraints is adopted to determine the causal relationship between the system monitoring variables.
[0033] In another possible implementation, the prior constraints include:
[0034] There is no direct causal relationship between any one of the system operating condition variables and any one of the system operating modes, and / or there is no direct causal relationship between any two of the system operating condition variables.
[0035] In another possible implementation, the neural network model is a graph deconvolution network model, and the second determination module is further configured to:
[0036] Inputting the system monitoring variables and the causal relationship into the preset graph deconvolution network model, and outputting the component-level degradation state representation;
[0037] The graph deconvolution network model is used to extract the component-level degradation state representation from the system monitoring variables and the causal relationship.
[0038] In another possible implementation, the generating module is further configured to:
[0039] Generate counterfactual data that conforms to the causal relationship through a preset CGAN model based on the component-level degradation state representation, the system operating condition variables, and the system operating mode;
[0040] The counterfactual data includes simulated system monitoring parameters, actual system operating condition variables, actual system operating mode and simulated system fault label data.
[0041] In another possible implementation, the CGAN model includes a generator and a discriminator, and the generation module is further configured to:
[0042] generating candidate counterfactual data by the generator according to the component-level degradation state representation, the system operating condition variables, and the system operating mode;
[0043] In a case where the discriminator determines that the candidate counterfactual data satisfies the causal relationship, the candidate counterfactual data is output as the counterfactual data.
[0044] In another possible implementation, the apparatus further includes a training module configured to:
[0045] Adding the counterfactual data to a training dataset of the anomaly detection model;
[0046] The anomaly detection model is trained according to the added training data set to obtain the data-enhanced anomaly detection model.
[0047] According to another aspect of the present disclosure, a device for causally driven counterfactual data generation and its application in anomaly detection is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-mentioned method when executing the instructions stored in the memory.
[0048] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.
[0049] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of a computing device, the processor in the computing device executes the above method.
[0050] The embodiments of the present disclosure provide a method for generating causal-driven counterfactual data and its application in anomaly detection. The method determines the causal relationship between system monitoring variables based on system monitoring variables, where the system monitoring variables include multiple monitoring variables monitored and related to the health status of the system. Based on the system monitoring variables and the causal relationship, a component-level degradation state representation is determined through a preset neural network model, where the component-level degradation state representation is used to indicate the health status of multiple components of the system. Based on the component-level degradation state representation, counterfactual data that conforms to the causal relationship is generated, and the counterfactual data is used to enhance the data of an anomaly detection model, which is used to perform anomaly detection on the system. That is, by determining the causal relationship, extracting the component-level degradation state representation based on the causal relationship, and generating counterfactual data that conforms to the causal relationship, the technical problem of being unable to generate causal-driven counterfactual data when causal information is inaccurate is solved. Data enhancement of the anomaly detection model based on the causal-driven counterfactual data enables the model to adapt to the influence of environmental changes, system upgrades, etc., and identify anomaly states not covered in historical monitoring data, thereby improving the generalization ability and stability of the anomaly detection model.
[0051] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0053] Figure 1 A flowchart of a method for generating causal-driven counterfactual data and applying the same to anomaly detection is shown in an exemplary embodiment of the present disclosure.
[0054] Figure 2 A schematic diagram of a PAG involved in a method for causally driven counterfactual data generation and its application in anomaly detection provided by an exemplary embodiment of the present disclosure is shown.
[0055] Figure 3 A schematic diagram showing a causal relationship between any two nodes in a PAG provided by an exemplary embodiment of the present disclosure is shown.
[0056] Figure 4 A schematic diagram of pseudo code for determining a skeleton based on path constraints provided by an exemplary embodiment of the present disclosure is shown.
[0057] Figure 5 A schematic diagram of a pseudo code for determining a PAG based on a direction constraint provided by an exemplary embodiment of the present disclosure is shown.
[0058] Figure 6 A schematic diagram illustrating the causal relationship between different types of system monitoring variables of a system provided by an exemplary embodiment of the present disclosure is shown.
[0059] Figure 7 A schematic structural diagram of a graph deconvolution network model provided by an exemplary embodiment of the present disclosure is shown.
[0060] Figure 8 A flowchart of a process for generating counterfactual data through CGAN provided by an exemplary embodiment of the present disclosure is shown.
[0061] Figure 9 A flowchart of a comparative experiment of different data enhancement methods provided by an exemplary embodiment of the present disclosure is shown.
[0062] Figure 10 A schematic structural diagram of a device for generating causal-driven counterfactual data and applying it in anomaly detection provided by an exemplary embodiment of the present disclosure is shown.
[0063] Figure 11 is a block diagram of a device provided according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0064] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0065] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0066] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0067] In related technologies, evaluating the effectiveness of improvement strategies at the model level (such as adjusting model parameters or using more powerful algorithms) is often challenging. However, at the data level, when generating simulated data (i.e., counterfactual data) to improve the distribution of historical monitoring data, the counterfactual data follows the meaning of the real monitoring data and can therefore be verified based on variable ranges, experience, and other factors. Currently, several data generation methods for improving model training data have been proposed and applied in various fields, including anomaly detection. For example, the Synthetic Minority Over-sampling Technique (SMOTE) can be used to generate counterfactual fault samples based on the Euclidean space distribution of monitored real data. This improved SMOTE method only oversamples the fault sample boundaries to enhance the classification decision boundary. Furthermore, Adaptive Synthetic (ADASYN) adaptively shifts the classification decision boundary by using weighted distributions for different fault samples. The emerging Generative Adversarial Network (GAN) can generate counterfactual data consistent with the distribution of real fault data through adversarial learning. However, most current research focuses on generating synthetic samples that are similar to the collected fault samples without considering the underlying causal mechanisms within the system (data generation mechanisms), so the performance of anomaly detection models heavily depends on the quality of historical monitoring data.
[0068] In recent years, some researchers have begun exploring how causality can facilitate the generation of counterfactual data. However, current research primarily focuses on image recognition and relies on known, established causal information. In practical applications, due to a lack of prior knowledge, it is often difficult to accurately mine and verify causal information in complex industrial systems and equipment. Methods for generating counterfactual data based on established causal relationships cannot be extended to anomaly detection tasks. Furthermore, the advantages of current data-driven approaches are not fully considered, which undoubtedly abandons the research heights and achievements of data-driven approaches. Therefore, given the premise that causal information is inaccurate and cannot be verified, the relevant technologies have yet to provide a reasonable and effective technical solution for how to enable causality to facilitate the generation of counterfactual data.
[0069] To address the technical issue in related technologies where current methods cannot generate causally driven counterfactual data when causal relationships are inaccurate, the present disclosure provides a method for generating causally driven counterfactual data and applying it to anomaly detection. This method reduces the false positive and false negative rates of anomaly detection and improves the safety of industrial systems and equipment. The proposed method can be applied to small-sample anomaly detection and is suitable for anomaly detection tasks in high-reliability systems such as electronic systems, electromechanical systems, and software systems. It can also be extended to other small-sample classification tasks.
[0070] The causal-driven counterfactual data generation method provided by the embodiments of the present disclosure includes three parts: causal information mining, causal mechanism decoupling, and counterfactual data generation. Causal information mining: Based on the FCI algorithm of prior constraints, the causal relationship between system monitoring variables is mined. Causal effect decoupling: The graph deconvolution network is designed to decouple the influence (causal effect) between system monitoring variables and extract component-level degradation state representation. Counterfactual data generation: CGAN is designed to generate counterfactual data that conforms to the causal relationship between system monitoring variables.
[0071] Methods for applying counterfactual data to anomaly detection include data augmentation for anomaly detection models. This involves incorporating counterfactual data into training datasets to enhance anomaly detection models. This allows the models to adapt to environmental changes, system upgrades, and other factors, and to identify anomalies not captured by historical monitoring data, thereby improving the model's generalization and stability. Furthermore, the generated counterfactual monitoring data corresponds to actual monitoring characteristics, making it practically useful. This data can not only be compared with real monitoring data to verify its effectiveness, but also be analyzed based on expert knowledge to analyze its own distribution characteristics.
[0072] Below, several exemplary embodiments are used to introduce the causal-driven counterfactual data generation and its application method in anomaly detection provided by the embodiments of the present disclosure.
[0073] Please refer to Figure 1 , which shows a flowchart of a method for generating causal-driven counterfactual data and its application in anomaly detection provided by an exemplary embodiment of the present disclosure. This embodiment uses the method applied to a computing device as an example. The method includes the following steps.
[0074] Step 101 : determining a causal relationship between system monitoring variables based on system monitoring variables, where the system monitoring variables include a plurality of monitored monitoring variables related to the health status of the system.
[0075] The computing device obtains system monitoring variables, which include multiple monitoring variables collected from the system, such as temperature, pressure, vibration, etc. These data are actually observed and reflect the health status of the system. System monitoring variables can be collected in real time by a data acquisition device (such as a sensor), and data preprocessing techniques such as wavelet analysis, median filtering, Kalman filtering, etc. are used to improve data quality and reduce the impact of noise and interference. Among them, the system can be an electronic system, an electromechanical system, a software system, or other systems, and the embodiments of the present disclosure are not limited to this.
[0076] In some embodiments, the types of system monitoring variables include at least one of system monitoring parameters, system operating condition variables, and system operating modes. The system monitoring parameters include variables collected by a data collection device set at a designated location of the system. The system operating condition variables include the system's external condition variables and / or environmental variables. The system operating mode is the operating mode adopted by the system when implementing its functions.
[0077] Based on the acquired system monitoring variables, the computing device determines the causal relationships between them. Causality refers to a logical relationship between variables, where changes in one variable cause changes in another. In system monitoring, determining causal relationships between variables helps understand system behavior and predict system responses.
[0078] In some embodiments, the computing device can determine the causal relationship between the system monitoring variables using an FCI algorithm based on prior constraints based on the system monitoring variables. The FCI algorithm is a constraint-based causal discovery algorithm, and the core concept is that different causal structures mean different independent relationships. The FCI algorithm can discover potential (unobserved) confounding factors, and when a "Y" structure is established in the graph, the causal relationship from X to Y is guaranteed not to be confused. In optimization problems, prior constraints refer to providing meaningful initial constraints for certain variables in the absence of sufficient observation data, making the optimization problem solvable. In the FCI algorithm, prior constraints can enhance the robustness of the optimization problem and help the optimizer obtain more accurate estimation results by providing known prior information for the variables in the graph.
[0079] In some embodiments, the a priori constraints include: there is no direct causal relationship between any system operating condition variable and any system operating mode, and / or there is no direct causal relationship between any two system operating condition variables.
[0080] It should be noted that the details of using the FCI algorithm based on prior constraints to determine the causal relationship between system monitoring variables can be found in the relevant description in the following embodiments and will not be introduced here.
[0081] Step 102 : Determine component-level degradation status representations based on system monitoring variables and causal relationships through a preset neural network model. The component-level degradation status representations are used to indicate the health status of each of the plurality of components of the system.
[0082] The computing device decouples the causal mechanism based on the system monitoring variables and the causal relationship between the system monitoring variables through a preset neural network model to extract the component-level degradation state representation.
[0083] In some embodiments, the computing device inputs the system monitoring variables and causal relationships into a preset neural network model and outputs a component-level degradation state representation. The neural network model is pre-trained and is used to extract component-level degradation state representations from system monitoring variables and causal relationships. Causal mechanism decoupling refers to analyzing and understanding the causal relationships between system monitoring variables, separating these relationships from the data, in order to better understand the behavior and dynamics of the system. The component-level degradation state representation is used to indicate the health status of multiple components of the system. The health status of multiple components of the system cannot be observed. The health status of multiple components of the system directly determines the system fault state and directly affects the system monitoring variables. The neural network model can be trained based on system monitoring variable samples and causal relationship samples, as well as the component-level degradation state representation labels corresponding to the samples, based on the training method in the prior art.
[0084] In some embodiments, the neural network model can be a graph deconvolution network model, which is a deep learning model based on a graph structure and is used to handle deconvolution problems in signals or data. Deconvolution is a mathematical operation used to recover the original signal from observed data, particularly in the fields of signal processing and image processing. The graph deconvolution network model solves the deconvolution problem by learning the deep features of the data, thereby improving the accuracy of signal recovery. The computing device determines the component-level degradation state representation based on the system monitoring variables and causal relationships using a preset neural network model. This may include: inputting the system monitoring variables and causal relationships into the preset graph deconvolution network model, and outputting the component-level degradation state representation. The graph deconvolution network model is used to extract the component-level degradation state representation from the system monitoring variables and causal relationships. That is, the input parameters of the graph deconvolution network are: the system monitoring variables and the causal relationships between the system monitoring variables, and the output parameter is: the component-level degradation state representation.
[0085] It should be noted that, for the sake of convenience, the following only takes the use of a graph deconvolution network model to determine the component-level degradation state representation as an example. The relevant details of this step can be referred to the relevant description in the following embodiment and will not be introduced here.
[0086] Step 103 : Generate counterfactual data that conforms to the causal relationship based on the component-level degradation state representation. The counterfactual data is used for data enhancement of the anomaly detection model, and the anomaly detection model is used to detect anomalies in the system.
[0087] The computing device can generate causally consistent counterfactual data based on the component-level degradation state representation using a preset neural network model. The neural network model is configured to generate causally consistent counterfactual data based on the component-level degradation state representation. In some embodiments, the causally consistent counterfactual data is generated based on the component-level degradation state representation, system operating condition variables, and system operating mode using a preset neural network model.
[0088] In some embodiments, the neural network model can be a CGAN model. The CGAN model is a deep learning model that can generate condition-dependent data. In anomaly detection, the CGAN can be used to generate simulated data, or counterfactual data, that is similar to actual observed data but contains anomalous features. This data can be used to enhance the training dataset.
[0089] In some embodiments, the computing device can generate counterfactual data that conforms to the causal relationship based on the component-level degradation state representation, system operating condition variables, and system operating mode through a preset CGAN model. The basic idea of the CGAN model is to train two networks, the generator and the discriminator, so that the generator obtains candidate counterfactual data by sparsely operating the component-level degradation state representation, while the discriminator is responsible for judging whether the candidate counterfactual data conforms to the causal relationship. That is, the computing device can input the component-level degradation state representation, the system operating condition variables, and the system operating mode into the preset CGAN model, and generate candidate counterfactual data through the generator; when the discriminator determines that the candidate counterfactual data conforms to the causal relationship, the candidate counterfactual data is output as counterfactual data.
[0090] Schematically, CGAN can include a generator and a discriminator. The generator includes a graph decoder and a causal mechanism model. The graph decoder is used to generate simulated system monitoring parameters based on component-level degradation state representation. The causal mechanism model is used to generate simulated system fault label data based on component-level degradation state representation, system operating condition variables and system operating mode. The system fault label data is used to indicate whether the system is in a faulty state. The system fault label data can be recorded data of whether the system is in a faulty state, thereby generating simulated counterfactual data, namely candidate counterfactual data. The candidate counterfactual data includes simulated system monitoring parameters, actual system operating condition variables, actual system operating mode and simulated system fault label data. The discriminator is used to determine whether the candidate counterfactual data conforms to the causal relationship. If it is determined that the candidate counterfactual data conforms to the causal relationship, the candidate counterfactual data is output as counterfactual data.
[0091] It should be noted that, for the sake of convenience, the following only uses the CGAN model to generate counterfactual data that conforms to the causal relationship as an example. The relevant details of this step can be referred to the relevant description in the following embodiment and will not be introduced here.
[0092] Counterfactual data refers to data that does not exist in the training dataset of an anomaly detection model but conforms to causal relationships. This type of data predicts possible outcomes by assuming certain conditions change. It can help understand how other variables will respond if one variable changes. In anomaly detection, counterfactual data can be data that could have occurred under specific conditions but did not actually occur. This data is used to simulate system behavior under different circumstances to enhance the model's ability to identify anomalies. Counterfactual data can include simulated system monitoring parameters, system operating condition variables, system operating modes, and simulated system fault label data. Specifically, the system monitoring parameters and system fault label data in the counterfactual data are simulated data, while the system operating condition variables and system operating modes in the counterfactual data are actually observed data.
[0093] Counterfactual data is used for data augmentation in anomaly detection models. Data augmentation is a technique used to improve model performance by creating additional training examples. In anomaly detection models, data augmentation can help the model learn a more diverse data distribution, thereby improving the model's ability to identify anomalies. Anomaly detection models are machine learning models used to identify anomalies or unusual patterns in data. Anomaly detection models can be either unsupervised or semi-supervised and are typically used to identify outliers in data that may indicate errors, fraudulent behavior, or system failures.
[0094] In some embodiments, the computing device may add counterfactual data to a training data set of an anomaly detection model; and train the anomaly detection model based on the added training data set to obtain a data-enhanced anomaly detection model.
[0095] A training dataset is used to train anomaly detection models. In anomaly detection models, the training dataset typically contains both normal and abnormal data. Incorporating counterfactual data into the training dataset increases data diversity and improves the model's generalization capabilities. Data-augmented anomaly detection models are anomaly detection models that have been augmented with counterfactual data. By incorporating counterfactual data into the training process, the model learns a wider range of abnormal patterns, thereby improving its performance and robustness in real-world applications.
[0096] After training the anomaly detection model, the application process of the anomaly detection model may include: obtaining real-time monitoring data from the system during operation; invoking the pre-trained anomaly detection model based on the real-time monitoring data, and outputting an anomaly detection result. The anomaly detection result is used to indicate whether the system is in a fault state. Invoking the pre-trained anomaly detection model based on the real-time monitoring data and outputting the anomaly detection result may further include: pre-processing the real-time monitoring data to obtain pre-processed real-time monitoring data; inputting the pre-processed real-time monitoring data into the pre-trained anomaly detection model; and outputting the anomaly detection result. The anomaly detection result is typically a probability value or classification label indicating whether the system is in a normal or faulty state.
[0097] In some embodiments, the anomaly detection model can be used to interpret and make decisions based on the anomaly detection results. If an anomaly is detected, appropriate measures may need to be taken, such as issuing an alarm, automatically shutting down the system, or performing maintenance. In some embodiments, the anomaly detection model can be evaluated and iteratively improved based on actual system performance and maintenance records. This may include retraining the model to improve its accuracy and robustness. In some embodiments, the anomaly detection model is integrated into the system's monitoring system to achieve automated anomaly detection and response.
[0098] In summary, the embodiments of the present disclosure provide a method for generating causal-driven counterfactual data and its application in anomaly detection, by determining the causal relationship between system monitoring variables based on system monitoring variables, where the system monitoring variables include multiple monitoring variables monitored and related to the health status of the system; based on the system monitoring variables and the causal relationship, a component-level degradation state representation is determined through a preset neural network model, and the component-level degradation state representation is used to indicate the health status of multiple components of the system; based on the component-level degradation state representation, counterfactual data that conforms to the causal relationship is generated, and the counterfactual data is used for data enhancement of the anomaly detection model, and the anomaly detection model is used to detect anomalies in the system; this solution brings the following beneficial effects: 1. Generating causal-driven counterfactual data: By determining the causal relationship, extracting the component-level degradation state representation based on the causal relationship, and generating counterfactual data that conforms to the causal relationship, the technical problem of being unable to generate causal-driven counterfactual data when causal information is inaccurate is solved; 2. Improving the accuracy of anomaly detection: by generating counterfactual data that conforms to the causal relationship, the anomaly detection model can learn a more comprehensive data distribution, including data under normal and abnormal conditions, thereby improving the model's ability to recognize anomalies. 3. Enhance the generalization ability of the model: The addition of counterfactual data makes the training data set more diverse, helps the model maintain good performance when facing unseen data, and enhances the generalization ability of the model. 4. Improve the robustness of the model: The introduction of counterfactual data can help the model maintain stable performance when facing noise and outliers in the data, and improve the robustness of the model. 5. Optimize the model training process: The counterfactual data generated by the CGAN model can simulate the behavior of the system under different circumstances, which helps to optimize the model training process and make model training more efficient. 6. Improve the speed of anomaly detection: By optimizing the model training process, the speed of anomaly detection can be improved, allowing the model to quickly respond to system anomalies and take timely measures.
[0099] The causal-driven counterfactual data generation method provided in this disclosure consists of three main components: 1. An FCI algorithm that mines prior constraints on system causal information; 2. A graph deconvolutional network model that decouples causal mechanisms; and 3. A CGAN model for generating counterfactual data. These three main components are further described below.
[0100] Part 1: FCI algorithm for mining system causal information.
[0101] The FCI algorithm takes as input the system's monitored variables and outputs a partial ancestral graph (PAG), which indicates the causal relationships between these variables. The FCI algorithm is effective in mining causal information even in the presence of biased or confounding data, and is therefore used as the benchmark causal discovery algorithm in this disclosure.
[0102] The output parameter of the FCI algorithm is usually a PAG. The schematic diagram of PAG is as follows Figure 2 As shown. PAG is a graph object used to represent a set of causal Bayesian networks (CBNs) that cannot be distinguished by the algorithm. In PAG, x and y represent variable nodes in the graph, and the relationship between them is represented by different types of edges, which reveal the causal relationship between system monitoring variables. The causal relationship between system monitoring variables includes: the causal relationship structure and causal relationship strength between system monitoring variables. The causal relationship structure between any two system monitoring variables can be one of a plurality of predefined causal relationship structures to identify the causal relationship between them; the causal relationship strength between any two system monitoring variables is used to indicate the size of the causal influence between any two system monitoring variables, reflecting the strength of the influence between them. The causal relationship between any two nodes in a PAG may exist as follows. Figure 3 There are five types of relationships shown (taking nodes i and j as an example, where i and j are both positive integers). There is no causal relationship between nodes i and j: This means there is no causal connection between nodes i and j. Node i is an ancestor of node j: This means node i precedes node j in the causal chain. The causal direction between nodes i and j cannot be determined: This means the direction of the causal relationship between nodes i and j cannot be determined. Node i is the direct cause of node j: This means node i directly affects node j. There is a confounding factor between nodes i and j: This means there are confounding factors between nodes i and j that may affect the causal relationship.
[0103] A PAG containing K nodes can be represented by a K×K dimensional matrix (called the causal adjacency matrix A):
[0104]
[0105] Where K is a positive integer, a ij (i≠j) is the indicative value of the causal relationship from node i to node j (different values are used to represent different types of causal relationships). All a in the causal adjacency matrix A ii is defined as 0.
[0106] In addition, the embodiment of the present disclosure uses average causal effects (ACE) to quantitatively describe the strength of the causal relationship and constructs the causal strength matrix A S as follows:
[0107]
[0108] where a S ij (i≠j) represents the strength of the causal relationship from node i to node j (the size of the causal influence), a S ii is defined as 0.
[0109] In the embodiment of the present disclosure, for Figure 3 The five possible causal relationship structures in the PAG, the elements a of the causal adjacency matrix and the causal strength matrix ij and a S ij The definition of is shown in Table 1. ACE is calculated using the Python Causal Inference Toolkit. The backdoor criterion and linear regression methods are selected in pre-defined functions (such as the model.estimate_effect function). In this example, "backdoor.linear_regression" specifies the use of the backdoor criterion and linear regression methods to estimate ACE. This method effectively controls confounding bias and produces more accurate causal effect estimates.
[0110] Table 1
[0111]
[0112] Considering that prior knowledge is usually limited and it is often difficult to determine that there is absolutely no causal relationship (including direct and indirect causal relationships) between two variables, the prior constraints designed and added in causal discovery in the embodiments of the present disclosure are divided into the following two types (taking nodes i and j as an example):
[0113] Path constraints: Also known as "direct causal relationship constraints," in this disclosure, there are two types of path constraints: a direct causal relationship between node i and node j (denoted as a positive path constraint); and no direct causal relationship between node i and node j (denoted as a negative path constraint). The set of all node groups (i, j) with positive path constraints is denoted as the positive path constraint set L+, and the set of node groups (i, j) with negative path constraints is denoted as the negative path constraint set L-.
[0114] Directional constraint: Constrain node i to be the ancestor of node j, that is, node i is the direct or indirect cause of node j. In the causal network, there is only a causal path from node i to node j. Here, a causal path refers to the combination of links from the cause node to the result node, including direct causal paths and indirect causal paths. The set of all node groups (i, j) with directional constraints is represented as a directional constraint set.
[0115] By limiting the search space of the FCI algorithm and adding prior constraints, we can not only obtain a more realistic causal network, but also reduce the computational complexity to a certain extent. The causal discovery process by limiting the search space and adding prior constraints is also divided into two steps, as shown below:
[0116] The pseudo code for determining the skeleton based on path constraints is as follows Figure 4 As shown, this pseudocode describes an algorithm for determining a skeleton C and a separation set S based on path constraints. The following is a natural language description of the various steps of the pseudocode:
[0117] 1. Input the node set (variable set) V, the positive path constraint set L+, and the negative path constraint set L-. These sets define the relationship between nodes. A positive path constraint indicates that a direct path exists between two nodes, while a negative path constraint indicates that no direct path exists between two nodes.
[0118] 2. Output the skeleton C and the separation set S. Skeleton C is the processed graph. Separation set S means that in skeleton C, for each pair of directly connected nodes Vi and Vj (denoted as edge (i, j)), there exists a specific node set that can sever or block the direct connection between Vi and Vj. In other words, separation set S contains all such node sets that can make the originally directly connected node pairs no longer directly connected after considering the influence of these node sets. In short, separation set S is the sum of the node sets corresponding to all connecting edges in skeleton C. These node sets can separate the directly connected node pairs (i, j).
[0119] 3. Create a completely undirected graph that contains all nodes in the node set V, and there is an edge between any two nodes.
[0120] 4. Delete all corresponding edges in the positive path constraint set L+ and the negative path constraint set L-. This means that if there is a positive path constraint between two nodes, the edge between them will be deleted; if there is a negative path constraint, the edge between them will also be deleted.
[0121] 5. Initialize variable l to -1 and C to the empty set.
[0122] 6. Start a loop that will repeat until a certain condition is met.
[0123] 7. At the beginning of each loop, increase the value of l by 1.
[0124] 8. Start another loop, which will be repeated until all undirected edge node groups (i, j) in C are traversed.
[0125] 9. In this loop, select a new pair of undirected edge node groups (i, j) in the skeleton C, and the number of connected nodes of node i except node j is greater than or equal to l.
[0126] 10. Start an inner loop that will be executed repeatedly until all possible subsets k of the node set V are traversed.
[0127] 11. Select a new subset k, where the node set k contains l elements.
[0128] 12. If nodes i and j are independent under the given conditions (i.e., the edge between them can be deleted), then proceed to the next step.
[0129] 13. Delete undirected edges And define the new graph as skeleton C. This means that node i and node j are no longer directly connected.
[0130] 14. Save the node set k into the separation set S(i,j) from i to j and the separation set S(j,i) from j to i. This means that k is the node set that makes i and j independent.
[0131] 15. If nodes i and j are not independent, continue checking the next subset k.
[0132] 16. Until the undirected edge is deleted or all possible subsets k are traversed.
[0133] 17. Until all undirected edge node groups (i, j) in C are traversed.
[0134] 18. If the conditions are met for all undirected edge node groups (i, j) in C, the loop ends.
[0135] 19. Add all the edges corresponding to the positive path constraint set L+ to the skeleton C. This means adding the edges corresponding to the previously deleted positive path constraints back to the graph.
[0136] The pseudo code for determining PAG based on directional constraints is as follows Figure 5 As shown, this pseudocode describes an algorithm for determining PAG based on directional constraints. The following is a natural language description of the various steps of the pseudocode:
[0137] 1. Input: The algorithm receives three inputs: skeleton C, separation set S, and orientation constraint set The skeleton C is the graph structure determined based on the path constraints. The separation set S contains the node set that can separate certain node pairs. The direction constraint set Contains information about the directionality between nodes.
[0138] 2. Output: The output of the algorithm is PAG G, which is a directed graph in which the edges represent the causal relationship between variables.
[0139] 3. Traverse the groups of nodes that are not directly connected: The algorithm traverses all pairs of nodes (i, j) that are not directly connected in the skeleton C, as well as their common neighbor nodes k.
[0140] 4. Check conditions: For each such node group (i, j) and k, the algorithm checks whether there is a condition that determines that there is a directional relationship between nodes i and j.
[0141] 5. Replace edge: If the conditions are met, the algorithm replaces the undirected edge in the skeleton C Replace with directed edges And define the resulting graph as G.
[0142] 6. End condition check: This step is the end of the previous step, indicating that the condition check and edge replacement of all indirect node groups are completed.
[0143] 7. End traversal: Complete the traversal of all pairs of nodes that are not directly connected.
[0144] 8. Traverse the node group corresponding to the direction constraint: the algorithm traverses all the nodes in the direction constraint set The node pairs (i, j) specified in , and the path p between them.
[0145] 9. Path check: If the path p contains only undirected edges Directed edges and “→”, then all undirected edges in path p are oriented as
[0146] 10. Processing includes Path: If the path p contains type of edge, the algorithm will Targeted And direct all undirected edges in path p to
[0147] 11. Handling paths containing “←”: If the path p contains an edge of type “←”, the algorithm needs to return to the first step, adjust the parameters of the conditional independence test, or delete the edge with the smallest causal strength in the path p to block the path p.
[0148] 12. End path check: complete the check of all paths.
[0149] 13. End traversal of direction constraints: complete the traversal of all direction constraints.
[0150] 14. Apply the FCI algorithm: In the resulting graph G, orient as many undirected edges as possible by repeatedly applying the 10 orientation rules R1-R10 used in the FCI algorithm to determine the PAG.
[0151] Typically, prior constraints are derived from prior knowledge such as expert experience and physical knowledge. Considering that in practical applications, it is possible that insufficient prior knowledge of the target system makes it impossible to obtain an effective prior constraint set, the present disclosure proposes a method for constructing a prior constraint set based on data.
[0152] Assume that the target system observation data D = [V, Y], V∈R N×K is a system monitoring variable related to the health status of the system, Y∈R N×1 is the system fault label data, that is, the record data of whether the system is in a fault state. The target system observation data includes N samples, and the system monitoring variables include K variables, where N and K are both positive integers. In this embodiment, the system monitoring variables include multiple variables V=(V1, V2, ..., V K ) is divided into three categories: system monitoring parameters η, system operating condition variables C and system operating mode M. Among them, system monitoring parameters refer to variables collected by data collection devices (such as sensors) set at designated locations of the system based on expert experience. System operating condition variables refer to the collected system exogenous variables, that is, the external condition variables and / or environmental variables of the system. The system operating mode refers to the operating mode (that is, operation or setting mode) when it realizes different functions, which is usually determined by relevant staff based on actual needs and experience. Multiple variables V can be arranged in the order of system monitoring parameters η, system operating condition variables C and system operating mode M, that is, η=(η1,η2,...,η H )=(V1,V2,...,V H ), C=(C1,C2,...,C Q )=(V H+1 ,V H+2 ,...,V H+Q ), M=V H+Q+1 =V K, where the system monitoring parameters η include H, and the system operating condition variables C include Q, and both H and Q are positive integers.
[0153] Schematically, based on the classification of variable categories, the causal relationship between different types of system monitoring variables can be described as Figure 6 The general structure shown. The variables in the solid box (i.e., system monitoring parameters η, system operating condition variables C, and system operating mode M) are usually observable and collected, while the variables in the dotted box (i.e., the health status U of each component) cannot be observed. The health status U of each component includes H', where H' is a positive integer. Among them, the health status of each component directly determines the fault state of the complex system and directly affects the system monitoring variables. The system operating mode and the system operating condition variables affect the system monitoring parameters by affecting the interaction relationship between components (i.e., the causal relationship structure and strength). Different system operating modes often correspond to different combinations and connection relationships of components within the system, so that the system can achieve different functions. Different system operating condition variables will affect the frequency and intensity of interactions between components. In addition, it is often impossible to determine whether the system operating mode and system operating condition variables will necessarily affect all system monitoring parameters.
[0154] Taking the system as a high-speed rail electromechanical control system as an example, the health status of each component cannot be observed but directly affects the system fault state. System operating condition variables, such as external temperature, humidity, required speed, input voltage, etc., represent the application environment instantaneously monitored by the electromechanical system, and therefore will not affect the health status of each component, but will affect the system monitoring parameters (such as line current, internal temperature). The system operating mode corresponds to the functions implemented by the electromechanical system. Usually, different operating modes are selected by the train driver according to mileage, stations, emergencies, etc., so there is no causal relationship between it and the system application environment, and it will not affect the health status of each component, but it will affect the system monitoring parameters. In some embodiments, the system monitoring parameters include multiple state monitoring values, and there is a one-to-one correspondence between the multiple state monitoring values and the multiple components, that is, each state monitoring value is used to uniquely indicate the health status of a component.
[0155] In the disclosed embodiment, the data-based prior constraints only include negative path constraints, which can only determine which system monitoring variables do not have direct causal relationships. Therefore, based on the causal relationships between different types of system monitoring variables in the above system, the following prior constraints can be constructed:
[0156] 1. There is no direct causal relationship between any system operating condition variable and the system operating mode, that is, the node pairs (H+1,K), (H+2,K),..., (H+Q,K) are all elements of the negative path constraint L-; and / or, 2. There is no direct causal relationship between any two system operating condition variables, that is, for any node pair (s,t), as long as H+Q≥s,t≥H+Q and s≠t are satisfied, these node pairs are all elements of the negative path constraint L-, where s and t are both positive integers.
[0157] Part 2: Graph deconvolution network model for decoupling causal mechanisms.
[0158] Under a single system operation mode and system operating conditions, when it is assumed that the electromechanical system contains H'=H components and each component has only one condition monitoring value, H is a positive integer, and it is assumed that the structural causal equations between components are all linear regression equations, and the component-level degradation state representation U'=[U1',U2',...,U' H ] to simplify the causal mechanism, the causal relationship of the system (i.e., the generation mechanism of the system monitoring parameter η) can be expressed as:
[0159]
[0160] Among them, A S It is the causal strength matrix of the whole system. The causal strength matrix is used to indicate the causal relationship (i.e., causal relationship structure and causal relationship strength) between multiple system monitoring parameters of the system. I is the unit matrix.
[0161] When the causal relationships among multiple system monitoring parameters are linear, the process of decoupling component-level degradation state representation is as follows based on the causal relationship (i.e., the generation mechanism of the system monitoring parameter η shown in the above formula):
[0162] U′=η(Ι-A S )
[0163] In order to prevent the dimension change from affecting the fault information in the data, it is generally hoped that the representation learning process does not change the data dimension. Therefore, the embodiment of the present disclosure normalizes the above formula as shown below:
[0164]
[0165] in is a diagonal matrix whose i-th row and i-th column elements are (i.e. A S The sum of the elements in the jth row of ), and all other elements are 0. It is called the causal intensity matrix A S The degree matrix of .
[0166] The disclosed embodiments use linear causal models multiple times to fit the nonlinear causal relationships that are prevalent in anomaly detection tasks, thereby decoupling component-level degradation state representation from system monitoring parameters. In some embodiments, the structural diagram of the graph deconvolution network model is as follows: Figure 7 As shown in the figure, the core idea of the graph deconvolution network model is to expand the graph convolution operation in the frequency domain, and then process the frequency domain representation of the graph through a series of causal decoupling layers, where the jth causal decoupling layer is j The operation process is as follows:
[0167]
[0168] Among them, L j and L j+1 They are the jth causal decoupling layer of the graph deconvolution network model. j Input and output parameters of W j ∈R H×H It's Layer j The trainable weight matrix of .
[0169] The optimization goal of the graph deconvolution network is to make the decoupled causal representation vector explainable to the system operation mode and system fault state. Explainability here means that the obtained component-level degradation state representation vector can distinguish different system operation modes and system fault states. The loss function of the proposed graph deconvolution network is DA for:
[0170] Loss DA =-silhouette_score(U′,YM)
[0171] Among them, YM is a variable obtained by label encoding the combination of system fault label data Y and system operation mode M, indicating the number of possible combinations of Y and M. silhouette_score() is a function that calculates the average silhouette coefficient, which can be directly calculated by the silhouette_score() function in the Python toolkit sklearn.metrics.
[0172] By optimizing the number of graph deconvolution network layers and the feature weight matrix of each layer to minimize the loss function, a component-level degradation state representation U′ can be obtained. This process involves adjusting the network structure and parameters to ensure that the representation vector output by the model is highly interpretable in distinguishing different system states.
[0173] Part III, CGAN for generating counterfactual data.
[0174] The CGAN consists of a generator G and a discriminator D. The generator includes a graph decoder η′ = g(U′) and a causal mechanism model Y = f′(U′, C, M). The discriminator uses a pre-defined FCI algorithm to determine whether the generated data conforms to the system's causal relationships. Generally speaking, counterfactual data should not differ significantly from the actual data, otherwise its credibility is difficult to assess. Considering that instantaneous monitoring values of environmental factors have little impact on system fault states and that the system's operating mode is fixed, the generator generates counterfactual data by sparsifying the component-level degradation state representation U′. This sparsification operation involves randomly selecting one or more features in the component-level degradation state representation U′ and replacing them with white noise. This approach aims to simulate various abnormal conditions that may be encountered in real applications, thereby enhancing the model's adaptability and robustness to different system states. If the discriminator determines that the generated data does not conform to the system's causal relationships, the discriminator improves the sparsification operation by restarting the random mapping mechanism. This approach increases the diversity and uncertainty of the data generation process by replacing the selected features and injecting white noise. In this way, richer counterfactual data that is closer to actual application scenarios can be generated.
[0175] In some embodiments, a graph decoder is designed and trained to convert the component-level degradation state representation U′ into simulated system monitoring parameters. The graph decoder has the same number of layers as the trained graph deconvolution network to ensure structural consistency during the decoding process. The i-th layer of the graph decoder can be expressed as:
[0176]
[0177] Among them, L i and L i+1 are the input parameters and output parameters of the i-th layer of the graph decoder, W i ′ is the trainable weight matrix of layer i.
[0178] The graph decoder optimizes W i ′ is trained to minimize the gap between the actual system monitoring parameter η and the system monitoring parameter η′ output by the decoder. The loss function Loss2 of the decoder is:
[0179] Loss2=||η-η′||2
[0180] Where || ||2 is the L2-norm of the matrix, representing the sum of the squares of the nonzero elements in the matrix. This loss function measures the error between the actual system monitoring parameters and the system monitoring parameters output by the decoder, and the graph decoder is trained by minimizing this error.
[0181] After training, the resulting graph decoder is denoted as η′=g(U′). This trained graph decoder can accurately convert component-level degradation state representations back to simulated system monitoring parameters, providing a powerful tool for system health monitoring and fault diagnosis.
[0182] In some embodiments, a causal mechanism model may be defined as:
[0183] Y=f′(U′,C,M)
[0184] Considering that the optimization goal of the component-level degradation state representation U′ is to cluster samples in the data [U′, YM] according to different YM, the resulting U′ has high separability, eliminating the need for a complex machine learning model to fit its mapping relationship with Y. The disclosed embodiment uses a single-layer perceptron training to obtain the causal mechanism model Y = f′(U′, C, M).
[0185] In some embodiments, the loss function Loss3 of CGAN is as follows:
[0186] Loss3=||A″-A||2
[0187] Where A is the causal adjacency matrix of the system monitoring variables, A″ is the causal adjacency matrix obtained using the FCI algorithm based on counterfactual data, and || ||2 is the L2-norm of the matrix, which represents the sum of the squares of the nonzero elements in the matrix.
[0188] To minimize the impact of data size on the discrimination process (i.e., determining whether counterfactual data fits the causal network), the sample size of the counterfactual data generated by the GAN is set to be the same as the sample size of the original data. When a large amount of counterfactual data is needed, it can be obtained in batches by running the causal GAN multiple times.
[0189] In some embodiments, a flowchart of the process of generating counterfactual data by CGAN is shown as follows: Figure 8The process includes the following steps: 1. Obtain input parameters (U′, C, M), which include the component-level degradation state representation U′, the system operating condition variable C, and the system operating mode M. 2. Sparsely operate on the features in U′. This operation involves selecting one or more features in U′ and replacing them with white noise to simulate uncertainty or anomalies in the data, thereby obtaining data (U″, C, M), where U″ represents the data after the sparse operation. 3. Input the data (U″, C, M) into the generator G and output the counterfactual data (η″, C, M, Y″), where the generator includes a graph decoder and a causal mechanism model. The graph decoder is used to generate the simulated system monitoring parameter η″ based on the component-level degradation state representation U′, and the causal mechanism model is used to generate the simulated system fault label data Y″ based on the component-level degradation state representation U′, the system operating condition variable C and the system operating mode M. 4. Input the counterfactual data (η″, C, M, Y″) into the discriminator D. The discriminator D judges whether the generated counterfactual data conforms to the causal relationship of the system based on the classic FCI algorithm. 5. If the generated counterfactual data conforms to the causal relationship of the system relationship, then these data are considered to be valid counterfactual data and can be output for further analysis or application. 6. If the generated counterfactual data does not conform to the causal relationship of the system, these data will be rejected, and it is necessary to improve the operation mode, adjust its parameters and regenerate data until the generated data can pass the evaluation of the discriminator. The whole process is an iterative process. The generator and the discriminator constantly compete with each other during the training process. The generator tries to generate more and more realistic data, while the discriminator tries to more accurately identify which data is real and which is generated. In this way, CGAN is able to learn to generate high-quality counterfactual data that meets specific conditions and system causal relationships.
[0190] In an illustrative example, based on a simulation dataset D of an aircraft-borne electromechanical system with imbalanced categories, ( ' 10) 、D ( ' 20) 、D ( ' 30) 、D ( ' 50) 、D ( ' 100) , verifying the effectiveness of the causal-driven counterfactual data generation method provided by the embodiment of the present disclosure, the flowchart of the comparative experiment of different data enhancement methods is shown in FIG. Figure 9As shown. The process includes the following steps: 1. Obtain observation data of complex electromechanical systems. 2. Necessary data preprocessing: including cleaning data, handling missing values, standardizing or normalizing features, etc., to ensure that the data is suitable for training models. 3. Training set preparation: The data is divided into multiple training sets, which are listed here as training set 1, training set 2, and up to training set 10. These training sets can be used for different model training or for cross-validation. 4. Generate data: Generate additional data to enhance the training set, which are listed here as generated data 1, generated data 2, and up to generated data 10. These data can be obtained through data augmentation methods to improve the generalization ability of the model. 5. Test set preparation: Similar to the training set, the test set is also divided into multiple, including test set 1, test set 2, and up to test set 10. These test sets are used to evaluate the performance of the model. 6. Processing of different fault diagnosis models: Use different fault diagnosis models to train and test the data. Different fault diagnosis models can include: Logistic Regression (LR), K-Nearest Neighbors (KNN) algorithm and Support Vector Machine (SVM). 7. Model evaluation: Evaluate the performance of these models on the test set. Evaluation indicators include the mean precision, the variance of precision, the mean recall and the variance of recall. The mean precision refers to the average value of the precision calculated in multiple experiments; the variance of precision refers to the variance of the precision values calculated in multiple experiments, which reflects the fluctuation of the precision values; the mean recall refers to the average value of the recall values calculated in multiple experiments; the variance of recall refers to the variance of the recall values calculated in multiple experiments, which reflects the fluctuation of the recall values. The entire process starts with data preprocessing, goes through model training and testing, and finally evaluates the performance of the model through statistical analysis of precision and recall.
[0191] To validate the effectiveness of data augmentation with correlated counterfactual data, we used common data augmentation methods: SMOTE, Borderline-SMOTE1, Borderline-SMOTE2, ADASYN, and the classic GAN for comparison. When comparing the impact of adding different data to the training process (i.e., incorporating them into the training set) on model generalization performance, it is often necessary to use an off-the-shelf classification algorithm as an anomaly detection model to avoid the impact of the algorithm's flexible and uncertain fitting relationships on the results. Therefore, we used the LR, KNN, and SVM algorithms as baseline anomaly detection models.
[0192] Table 2, Table 3, Table 4, Table 5, Table 6 are the results of different data augmentation methods on the aviation airborne electromechanical system simulation dataset D( ' 10) 、D ( ' 20) 、D ( ' 30) 、D ( ' 50) 、D ( ' 100) The models shown in the table are combinations of different data augmentation methods and different anomaly detection models. For example, SMOTE-LR is an anomaly detection model constructed using the LR algorithm after generating data based on SMOTE. LR, KNN, and SVM are anomaly detection models that do not use data augmentation methods. LR*, KNN*, and SVM* are anomaly detection models constructed using the LR, KNN, and SVM algorithms after data augmentation using associated counterfactual data. The generalization effect is evaluated using multiple evaluation metrics, including mean precision, precision variance, mean recall, and recall variance.
[0193] Table 2
[0194]
[0195]
[0196] Table 3
[0197] Model Mean precision Precision Variance Mean recall Recall variance LR 0.6324 0.0033 0.5854 0.0080 LR* 0.7884 0.0003 0.7085 0.0009 SMOTE-LR 0.6306 0.0041 0.5860 0.0056 BorderlineSMOTE1-LR 0.6229 0.0052 0.5690 0.0078 BorderlineSMOTE2-LR 0.6263 0.0052 0.5566 0.0078 ADASYN-LR 0.6329 0.0084 0.5437 0.0086 GAN-LR 0.6268 0.0094 0.5710 0.0076 KNN 0.6209 0.0085 0.5747 0.0066 KNN* 0.7740 0.0009 0.7066 0.0007 SMOTE-KNN 0.6330 0.0092 0.5982 0.0029 BorderlineSMOTE1-KNN 0.6246 0.0093 0.6084 0.0063 BorderlineSMOTE2-KNN 0.6106 0.0020 0.6341 0.0046 ADASYN-KNN 0.6269 0.0022 0.6062 0.0038 GAN-KNN 0.6234 0.0018 0.5987 0.0039 Support Vector Machine 0.6606 0.0020 0.5321 0.0049 SVM* 0.7518 0.0007 0.7286 0.0008 SMOTE-SVM 0.5959 0.0014 0.6134 0.0056 BorderlineSMOTE1-SVM 0.6068 0.0013 0.5962 0.0093 BorderlineSMOTE2-SVM 0.5856 0.0014 0.6247 0.0095 ADASYN-SVM 0.5982 0.0070 0.6214 0.0098 GAN-SVM 0.5995 0.0069 0.6082 0.0127
[0198] Table 4
[0199] Model Mean precision Precision Variance Mean recall Recall variance LR 0.5976 0.0082 0.6407 0.0075 LR* 0.7625 0.0003 0.7597 0.0004 SMOTE-LR 0.6144 0.0021 0.6533 0.0101 BorderlineSMOTE1-LR 0.6097 0.0009 0.6592 0.0122 BorderlineSMOTE2-LR 0.6148 0.0013 0.6484 0.0123 ADASYN-LR 0.6184 0.0050 0.6207 0.0114 GAN-LR 0.5985 0.0053 0.6055 0.0127 KNN 0.5858 0.0088 0.6022 0.0149 KNN* 0.7561 0.0014 0.7030 0.0010 SMOTE-KNN 0.5990 0.0109 0.5749 0.0092 BorderlineSMOTE1-KNN 0.6099 0.0101 0.5813 0.0079 BorderlineSMOTE2-KNN 0.6163 0.0103 0.6077 0.0043 ADASYN-KNN 0.6363 0.0072 0.6236 0.0041 GAN-KNN 0.6459 0.0068 0.6189 0.0036 Support Vector Machine 0.7146 0.0046 0.6078 0.0021 SVM* 0.7765 0.0006 0.7559 0.0003 SMOTE-SVM 0.6143 0.0067 0.6521 0.0072 BorderlineSMOTE1-SVM 0.6111 0.0068 0.6422 0.0102 BorderlineSMOTE2-SVM 0.5875 0.0068 0.6263 0.0126 ADASYN-SVM 0.5832 0.0068 0.6308 0.0130 GAN-SVM 0.6009 0.0124 0.6093 0.0122
[0200] Table 5
[0201]
[0202]
[0203] Table 6
[0204]
[0205]
[0206] As can be seen, current simple classifiers (LR, KNN, and SVM) have poor generalization performance for class-imbalanced simulation data. Common data augmentation methods such as SMOTE, Borderline-SMOTE1, Borderline-SMOTE2, ADASYN, and classic GAN can all improve model generalization performance, but it's difficult to conclude based on this case that these methods improve model stability. In fact, the stability of these commonly used data augmentation methods often depends on the distribution differences between the training and test sets.
[0207] However, when using models built based on the associated counterfactual data generation method (i.e., LR*, KNN*, SVM*), it can be found that these models have the highest average precision and recall rates, and the smallest variance. This shows that the data enhancement technology combined with the method provided by the embodiment of the present disclosure can not only significantly improve the generalization performance of the anomaly detection model, but also has high stability. This finding has been verified on simulation data sets of different complexity, thereby confirming the effectiveness of the solution provided by the embodiment of the present disclosure, and further showing that the data enhancement method based on counterfactual data generation has wide applicability.
[0208] The following is an apparatus embodiment of the present disclosure. For parts not described in detail in the apparatus embodiment, reference may be made to the technical details disclosed in the above method embodiment.
[0209] Please refer to Figure 10 , which shows a schematic diagram of the structure of an apparatus for causally driven counterfactual data generation and its application in anomaly detection, provided by an exemplary embodiment of the present disclosure. The apparatus can be implemented in whole or in part as a computing device using software, hardware, or a combination of both. The apparatus includes a first determination module 1010, a second determination module 1020, and a generation module 1030.
[0210] A first determining module 1010 is configured to determine a causal relationship between system monitoring variables based on system monitoring variables, where the system monitoring variables include a plurality of monitored monitoring variables related to the health status of the system;
[0211] A second determination module 1020 is configured to determine a component-level degradation state representation based on system monitoring variables and causal relationships using a preset neural network model, wherein the component-level degradation state representation is used to indicate the health status of each of the plurality of components of the system;
[0212] The generation module 1030 is used to generate counterfactual data that conforms to the causal relationship based on the component-level degradation state representation. The counterfactual data is used for data enhancement of the anomaly detection model, and the anomaly detection model is used to detect anomalies in the system.
[0213] In one possible implementation, the type of system monitoring variables includes at least one of system monitoring parameters, system operating condition variables and system operating modes. The system monitoring parameters include variables collected by a data collection device set at a designated location of the system. The system operating condition variables include the system's external condition variables and / or environmental variables. The system operating mode is the operating mode adopted by the system when implementing its functions.
[0214] In another possible implementation, the first determining module 1010 is further configured to:
[0215] According to the system monitoring variables, the FCI algorithm based on prior constraints is used to determine the causal relationship between the system monitoring variables.
[0216] In another possible implementation, the prior constraints include:
[0217] There is no direct causal relationship between any system operating condition variable and any system operating mode, and / or there is no direct causal relationship between any two system operating condition variables.
[0218] In another possible implementation, the neural network model is a graph deconvolution network model, and the second determining module 1020 is further configured to:
[0219] Input system monitoring variables and causal relationships into the preset graph deconvolution network model, and output component-level degradation state representation;
[0220] Among them, the graph deconvolution network model is used to extract component-level degradation state representation from system monitoring variables and causal relationships.
[0221] In another possible implementation, the generating module 1030 is further configured to:
[0222] Based on the component-level degradation state representation, system operating condition variables, and system operation mode, the preset CGAN model is used to generate counterfactual data that conforms to the causal relationship.
[0223] Among them, the counterfactual data includes simulated system monitoring parameters, actual system operating condition variables, actual system operating mode and simulated system fault label data.
[0224] In another possible implementation, the CGAN model includes a generator and a discriminator, and the generation module 1030 is further configured to:
[0225] Generate candidate counterfactual data through a generator based on component-level degradation state representation, system operating condition variables, and system operating mode;
[0226] When the discriminator determines that the candidate counterfactual data conforms to the causal relationship, the candidate counterfactual data is output as the counterfactual data.
[0227] In another possible implementation, the apparatus further includes a training module configured to:
[0228] Add counterfactual data to the training dataset of the anomaly detection model;
[0229] The anomaly detection model is trained based on the added training data set to obtain a data-enhanced anomaly detection model.
[0230] It should be noted that, when the device provided in the above embodiment realizes its function, it only uses the division of the above-mentioned functional modules as an example. In actual application, the above-mentioned functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0231] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0232] An embodiment of the present disclosure further provides a computing device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0233] The embodiment of the present disclosure further provides a non-volatile computer-readable storage medium having computer program instructions stored thereon, which implement the above method when the computer program instructions are executed by a processor.
[0234] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of a computing device, the processor in the computing device executes the above method.
[0235] Figure 11 1 is a block diagram of an apparatus 1900 provided according to an exemplary embodiment of the present disclosure. For example, the apparatus 1900 may be provided as a server or a terminal device, and may include the above-mentioned causal-driven counterfactual data generation and its application in anomaly detection. Figure 11The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.
[0236] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, MacOS X™, Unix™, Linux™, FreeBSD™, or the like.
[0237] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the apparatus 1900 to perform the above-described method.
[0238] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0239] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0240] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0241] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0242] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0243] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0244] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0245] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0246] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A causal-driven counterfactual data generation method and its application in anomaly detection, characterized in that: The method comprises: determining a causal relationship between system monitoring variables based on the system monitoring variables, wherein the system monitoring variables include a plurality of monitored monitoring variables related to the health status of the system; Determining a component-level degradation state representation based on the system monitoring variables and the causal relationship through a preset neural network model, wherein the component-level degradation state representation is used to indicate the health state of each of the plurality of components of the system; generating counterfactual data that conforms to the causal relationship based on the component-level degradation state representation, wherein the counterfactual data is used for data enhancement of an anomaly detection model, and the anomaly detection model is used to perform anomaly detection on the system; The neural network model is a graph deconvolution network model. The component-level degradation state representation is determined by a preset neural network model based on the system monitoring variables and the causal relationship, including: Inputting the system monitoring variables and the causal relationship into the preset graph deconvolution network model, and outputting the component-level degradation state representation; wherein the graph deconvolution network model is used to extract the component-level degradation state representation from the system monitoring variables and the causal relationship; Generating counterfactual data that conforms to the causal relationship based on the component-level degradation state representation includes: Generate counterfactual data that conforms to the causal relationship through a preset causal adversarial generative network (CGAN) model based on the component-level degradation state representation, the system operating condition variables, and the system operating mode; The counterfactual data includes simulated system monitoring parameters, actual system operating condition variables, actual system operating mode and simulated system fault label data.
2. The method according to claim 1, characterized in that The types of the system monitoring variables include at least one of system monitoring parameters, system operating condition variables and system operating modes. The system monitoring parameters include variables collected by a data collection device set at a designated location of the system. The system operating condition variables include external condition variables and / or environmental variables of the system. The system operating mode is the operating mode adopted by the system when implementing its functions.
3. The method according to claim 2, characterized in that Determining the causal relationship between the system monitoring variables according to the system monitoring variables includes: According to the system monitoring variables, a fast causal inference (FCI) algorithm based on prior constraints is adopted to determine the causal relationship between the system monitoring variables.
4. The method according to claim 3, characterized in that The prior constraints include: There is no direct causal relationship between any one of the system operating condition variables and any one of the system operating modes, and / or there is no direct causal relationship between any two of the system operating condition variables.
5. The method according to claim 1, wherein The CGAN model includes a generator and a discriminator. The CGAN model generates counterfactual data that conforms to the causal relationship based on the component-level degradation state representation, the system operating condition variables, and the system operating mode through a preset causal adversarial generative network (CGAN) model, including: generating candidate counterfactual data by the generator according to the component-level degradation state representation, the system operating condition variables, and the system operating mode; In a case where the discriminator determines that the candidate counterfactual data satisfies the causal relationship, the candidate counterfactual data is output as the counterfactual data.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Adding the counterfactual data to a training dataset of the anomaly detection model; The anomaly detection model is trained according to the added training data set to obtain the data-enhanced anomaly detection model.
7. A causal-driven counterfactual data generation and its application in anomaly detection, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 6 when executing the instructions stored in the memory.
8. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Digital media advertisement effect evaluation system
CN117829914A
Traffic network abnormal situation spatio-temporal evolution method fusing causal inference
CN118197047A