Industrial control anomaly explanation method and system based on generative counterfactual sample differences

By using the conditional variational autoencoder in the industrial control system to generate counterfactual samples and combining the abnormal score information, the problems of insufficient interpretability and dissatisfaction with real-time requirements in the existing industrial control abnormal detection methods are solved, and higher quality and correlation abnormal interpretation is achieved, and more effective decision-making is supported.

CN119067225BActive Publication Date: 2025-05-16QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411569749.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-05-16
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

The existing industrial control anomaly detection methods lack interpretability when explaining the causes of abnormalities, are difficult to meet the needs of real-time monitoring, and are difficult to generate diverse and meaningful counterfactual samples, resulting in limited comprehensiveness and accuracy of abnormal explanations.

Method used

By leveraging the anomaly score information output from the anomaly detection model, a diverse and meaningful counterfactual sample is generated using the conditional generation method, and a variability measurement mechanism is introduced to ensure the diversity and representativeness of the generated samples. The specific method is to use the conditional variational autoencoder (CVAE), use the exception score as a condition for the generation process, generate counterfactual samples, and make a reasonable explanation by the decision information of the feature importance score as the anomaly detection model.

Benefits of technology

It improves the quality and correlation of generated samples, can better capture the essential characteristics of exceptions, provide more comprehensive and reliable exception interpretation, meet real-time monitoring needs, and provide stronger decision-making support for system administrators and operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119067225B_ABST
    Figure CN119067225B_ABST
Patent Text Reader

Abstract

The present invention relates to an industrial control anomaly interpretation method and system based on generative counterfactual sample differences, which belongs to the technical field of industrial control system anomaly detection research, including: predicting anomaly score results according to an industrial control anomaly detection model, obtaining an industrial control mixed data set through anomaly score results and multi-sensor time series data set, and preprocessing; taking the original time series data set as input, and the anomaly score output by the industrial control anomaly detection model as a condition, inputting into a conditional variational autoencoder for training; collecting the anomaly threshold output by the industrial control anomaly detection model when predicting the data set, and generating counterfactual samples by changing the threshold size in the conditional variational autoencoder; and obtaining feature importance scores by comparing counterfactual samples with the originally collected multi-sensor time series samples. The present invention improves the practicality of anomaly detection and interpretation in industrial control systems, and provides a more powerful decision support tool for system administrators and operators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial control system anomaly detection research, and specifically relates to an industrial control anomaly interpretation method and system based on generative counterfactual sample differences. Background Art

[0002] As the industrial control environment becomes increasingly complex, a large number of sensors and actuators in industrial control systems continuously generate multi-dimensional time series data, which contains important information about the system's operating status. Anomaly detection, as a key technology to ensure the safe and stable operation of industrial control systems, has always been a hot topic in academia and industry. However, with the increasing complexity of industrial control systems, simply detecting anomalies is no longer sufficient to meet actual needs. Engineers and operators not only need to know whether the system is abnormal, but also need to understand the causes and effects of the anomalies in order to take appropriate response measures.

[0003] Traditional industrial control anomaly detection methods, such as those based on statistical analysis, machine learning, and deep learning, have achieved certain results in anomaly identification, but they are still significantly insufficient in anomaly interpretability. These methods are usually regarded as "black box" models and are difficult to provide clear and interpretable anomaly cause analysis. The lack of interpretability not only affects the system administrator's understanding and handling of abnormal situations, but also increases the risk of false positives and false negatives.

[0004] In recent years, Explainable Artificial Intelligence (XAI) technology has made significant progress in many fields, providing new ideas for industrial control anomaly explanation. Among them, counterfactual explanation, as an emerging interpretability method, provides intuitive explanations for decision-making by generating hypothetical scenarios of "if...then...". However, the application of counterfactual explanation to the field of industrial control anomaly detection still faces many challenges, such as: 1) Data complexity: Industrial control systems generate high-dimensional, multivariate time series data. How to generate meaningful counterfactual samples on such complex data is a difficult problem; 2) Real-time requirements: Industrial control systems usually require real-time response. How to meet real-time requirements while ensuring the quality of explanation is a key challenge; 3) Anomaly score utilization: Existing counterfactual generation methods often do not fully utilize the anomaly score information output by the anomaly detection model. The anomaly score not only reflects the degree to which the data point deviates from the normal pattern, but also contains potential anomaly cause information. How to effectively integrate the anomaly score into the counterfactual sample generation process to improve the quality and relevance of the generated samples is a direction worth exploring. Summary of the invention

[0005] Aiming at the shortcomings of existing industrial control anomaly explanation technology, the present invention proposes an industrial control anomaly explanation method based on generative counterfactual sample differences;

[0006] The present invention aims to solve several key problems existing in the current anomaly detection and interpretation methods of industrial control systems: first, the interpretability of anomaly interpretation is insufficient, making it difficult to provide operators with clear and intuitive analysis of the causes of anomalies; second, existing methods are inefficient in processing high-dimensional, multivariate industrial control time series data, making it difficult to meet the needs of real-time monitoring; third, there is a lack of effective mechanisms to generate diverse and meaningful counterfactual samples, resulting in limited comprehensiveness and accuracy of anomaly interpretation. The purpose is to improve the practicality of anomaly detection and interpretation in industrial control systems and provide system administrators and operators with a more powerful decision support tool.

[0007] The present invention makes full use of the anomaly score information output by the anomaly detection model to generate diverse and meaningful counterfactual samples through conditional generation. At the same time, the present method introduces a new difference measurement mechanism to ensure the diversity and representativeness of the generated samples, thereby providing a more comprehensive and reliable anomaly explanation. In addition, the present invention also takes into account practical needs such as computational efficiency and operability, providing decision makers with reliable and feasible decision support.

[0008] Taking into account the particularity of industrial control data, especially its strong temporal correlation and multi-source heterogeneity, the present invention innovatively proposes a generative counterfactual sample method based on anomaly score conditions. The method first uses the output score of the anomaly detection model as conditional information to guide the conditional variational autoencoder (CVAE) to generate counterfactual samples, and generates importance scores for features based on the feature differences between the counterfactual samples and the original samples, and makes reasonable explanations for the decision information of the industrial control anomaly detection model from both the sample level and the feature level. The present invention can effectively meet the needs of improving the quality and relevance of generated samples and better capturing the essential characteristics of anomalies, and provide a traceable and operational solution for the safety monitoring and anomaly handling of industrial control systems.

[0009] The present invention also proposes an industrial control anomaly explanation system based on generative counterfactual sample differences.

[0010] Terminology explanation:

[0011] 1. Explainable AI: It is a technical method that aims to make the decision-making process and results of AI systems transparent and understandable. It solves the problem that traditional "black box" AI models are difficult to explain their decision-making reasons. It can reveal how AI systems reach specific conclusions through technical means such as feature importance analysis, decision tree visualization, and local interpretability models. In industrial control systems, complex process control, real-time monitoring, and key decisions are usually involved, which directly affect production efficiency, safety, and economic benefits. Explainable AI is a key technology to ensure the safe, reliable, and efficient operation of the system.

[0012] 2. Counterfactual explanation: It is a method to explain the decision of AI model by generating hypothetical samples that are slightly different from the original input but lead to different outputs. Counterfactual explanation answers the question "how would the result change if the input was slightly different". Its core idea is to find the smallest input change that significantly changes the model output.

[0013] 3. Anomaly score: It is a numerical indicator used in anomaly detection tasks to quantify the degree to which a data point deviates from the normal pattern. It reflects the degree of difference between a specific data instance and the expected normal behavior. The anomaly score is usually calculated by an anomaly detection algorithm, and the higher the score, the more likely the data point is to be an anomaly. In industrial control systems, anomaly scores can help operators quickly identify potential system failures or security threats and prioritize the most serious anomalies.

[0014] 4. Conditional Variational Autoencoder (CVAE): is a widely used generative artificial intelligence model. Conditional Variational Autoencoder is an extended version of Variational Autoencoder (VAE), which introduces additional conditional information into the generative model. CVAE not only learns the potential representation of data, but also learns how to generate data based on given conditions.

[0015] 5. Industrial control anomaly detection model: It is an anomaly detection technology specifically used for industrial control systems (ICS), which aims to monitor and analyze data streams from various sensors, actuators, and controllers in real time to identify abnormal patterns that may indicate system failures, security threats, or performance anomalies. Industrial control anomaly detection models usually use machine learning or deep learning techniques, such as long short-term memory networks (LSTM), autoencoders, or convolutional neural networks (CNNs), to capture normal behavior patterns in complex time series data and detect anomalies that deviate from these patterns.

[0016] The technical solution of the present invention is:

[0017] The industrial control anomaly explanation method based on generative counterfactual sample differences includes:

[0018] The anomaly score results are predicted according to the industrial control anomaly detection model. The industrial control mixed data set is obtained through the anomaly score results and the originally collected multi-sensor time series data set, and preprocessed. After preprocessing, it is divided into a training set and a test set.

[0019] The original time series dataset is used as input, and the anomaly score output by the industrial control anomaly detection model is used as a condition and input into the conditional variational autoencoder (CVAE) for training;

[0020] Collect the anomaly thresholds output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder;

[0021] The feature importance score is obtained by comparing the counterfactual samples with the originally collected multi-sensor time series samples. The size of the feature importance score represents the contribution of the feature to the prediction of the detection results by the anomaly detection model.

[0022] Further preferably, continuous data is collected from each sensor in the industrial control system at fixed time intervals as an original time series data set, namely, a multi-sensor time series data set.

[0023] Further preferably, the original time series data set is merged with the anomaly score result as an industrial control mixed data set, that is, at each time point, the content of the industrial control mixed data set is the operating status data collected at the current time point and the corresponding anomaly score.

[0024] Further preferably, the pretreatment comprises:

[0025] 1) Data cleaning: remove missing values;

[0026] 2) Data standardization.

[0027] Preferably, according to the present invention, the original multi-sensor time series is processed by the industrial control anomaly detection model and stored according to the preset time window sampling, including:

[0028] Assume that in an industrial control system, the data set is defined express Devices and time steps The measured device values ​​obtained on the device include sensors and actuators. Indicates Devices, Yes Equipment In time Defining the industrial control anomaly detection model , then the industrial control anomaly detection model is used for the preprocessed multi-sensor time series dataset The predicted anomaly score is , .

[0029] Preferably, according to the present invention, an encoder, a decoder, an anomaly score embedding module, and a reparameterized sampling module of a conditional variational autoencoder are constructed; conditional parameters are input into the encoder of the conditional variational autoencoder, that is, the anomaly score is predicted to obtain the mean and variance; conditional parameters are also input into the decoder to guide the reconstruction training of the conditional variational autoencoder; including:

[0030] Build a conditional variational autoencoder, including building an encoder, a decoder, anomaly score embedding module, and a reparameterized sampling module, and introduce conditional information in the encoding and decoding process. The conditional information is: anomaly score, which represents the numerical value of the degree of anomaly of data in each time window, so as to realize data generation under specific conditions;

[0031] In the anomaly score embedding module, a linear layer is used to perform an embedding operation on the input anomaly score row, and the score is encoded so that the conditional information is embedded into a continuous vector space. The specific method is shown in formula (I):

[0032] (I);

[0033] In formula (I), is the anomaly score result after embedding transformation in the linear layer, represents the ReLU activation function, Represents an operation through a linear layer. Represents the anomaly score predicted by the industrial control anomaly detection model. The transformed anomaly score dimension is changed from Transformed to , define the final output dimension as 2;

[0034] The embedded anomaly score is merged with the encoder input tensor. The encoder includes three hidden layers. The specific method is shown in formula (II):

[0035] (II);

[0036] In formula (II), A tensor representing the merged raw time series dataset, used as the input to the encoder, Represents a merge operation. Represents an operation through a linear layer. represents the ReLU activation function, Represents the latent space vector, which is the output of the encoder;

[0037] Similarly, in the decoder, the original time series data set and conditional information are mapped to the mean and logarithmic variance of the latent space; the decoder and the encoder are symmetric results, and the specific implementation is shown in formula (III):

[0038] (III);

[0039] In formula (III), are the mean and logarithmic standard deviation of the encoder output, represents the sampling operation, The operation in formula (II) A tensor representing the merged original time series dataset, used as the input to the decoder, Represents a merge operation. Represents an operation through a linear layer. represents the ReLU activation function, represents the latent space vector, i.e., the output of the encoder, is the result of the decoder output, that is, the reconstruction result of the conditional variational autoencoder;

[0040] The output of the decoder is input into the reparameterized sampling module, which transforms the random sampling process into a deterministic function and accepts a random noise as input. The specific method is shown in formula (IV):

[0041] (IV);

[0042] In formula (IV), are the mean and logarithmic standard deviation of the encoder output, represents the tensor splitting function, represents the latent space sampling point, represents random noise sampled from a standard normal distribution N(0, 1).

[0043] Preferably, according to the present invention, training a conditional variational autoencoder comprises:

[0044] Initialize the encoder, decoder, and anomaly score embedding module; the encoder and decoder use a multi-layer neural network structure, using the Linear layer and ReLU activation function; the anomaly score embedding module converts conditional information into a continuous vector representation;

[0045] In each training iteration, the conditional variational autoencoder receives input data x and conditional information score; the anomaly score embedding module processes the conditional information through formula (I) to generate an embedding vector; the encoder receives input data and embedding vector and outputs the mean of the latent space and log variance ; Through the reparameterized sampling module, the latent variables are sampled from the latent space distribution through formula (IV) ; The decoder uses formula (III) to use hidden variables and conditional embedding to reconstruct input data;

[0046] After each iteration, the gradients are calculated through back-propagation and the parameters of the conditional variational autoencoder are updated using the Adam optimizer.

[0047] According to the preferred embodiment of the present invention, the loss function of the conditional variational autoencoder is Including: reconstruction error and KL divergence, that is: ,in, represents the total loss, represents the reconstruction loss, Represents KL divergence loss; reconstruction loss is used to evaluate the reconstruction quality, and KL divergence is used to regularize the latent space distribution; the calculation method of reconstruction loss is shown in formula (V):

[0048] (V);

[0049] In formula (IV), represents the mean square error, represents the mean absolute error, is the weight coefficient, represents the original input data, Represents the reconstructed data;

[0050] The calculation method of KL divergence loss is shown in formula (VI):

[0051] (VI);

[0052] In formula (VI), represents the mean value of the encoder output, Represents the variance of the encoder output.

[0053] Preferably, according to the present invention, the abnormal threshold output by the industrial control abnormality detection model when predicting the data set is collected, and the counterfactual sample is generated by changing the threshold value in the conditional variational autoencoder; including:

[0054] Based on the trained conditional variational autoencoder, it is assumed that the conditional variational autoencoder can accept the original sample and the anomaly score as input and generate a reconstructed sample; the conditional variational autoencoder receives the test data loader, which includes the original sample , Raw Anomaly Score and anomaly threshold ;

[0055] Set a new threshold , the original anomaly score Make adjustments; the specific operations are:

[0056] Increase the anomaly threshold of each sample by a fixed value ; The specific method is shown in formula (VII):

[0057] (VII);

[0058] The adjusted anomaly threshold With the original sample are input into the trained conditional variational autoencoder together; in the generation mode, the trained conditional variational autoencoder is based on the adjusted anomaly threshold With the original sample , generate counterfactual samples , ; Make the smallest change in the features to flip the prediction results of the conditional variational autoencoder. Make the smallest change in the features, which means: Transformed into .

[0059] Preferably, according to the present invention, a multi-dimensional feature difference mapping mechanism is constructed to reveal the influence of key features by comparing the positions of original samples and generated counterfactual samples in feature space, explain the decision of each feature on the anomaly detection model, and use high-dimensional feature extraction technology to obtain comprehensive feature representation from the original samples, form a difference spectrum, and construct a feature difference matrix; including:

[0060] Step 1: After obtaining the counterfactual sample, compare the data changes between the original sample and the generated counterfactual sample to obtain the difference between the features; use the independent sample T test to compare each feature of the original sample and the generated counterfactual sample, and calculate the significant difference between the original time series and the generated counterfactual time series;

[0061] The assumptions for the independent sample T test are defined as follows: Null hypothesis: the means of the two groups of samples are equal; Alternative hypothesis: the means of the two groups of samples are not equal; By calculating the T statistic and querying the corresponding P value, decide whether to reject the null hypothesis; If the P value is less than the significance level, reject the null hypothesis and believe that there is a significant difference in the means of the two groups of samples; The significance level is 0.05 or 0.01, otherwise, if the P value is greater than or equal to the significance level, the null hypothesis cannot be rejected and it is believed that the difference in the means of the two groups of samples is not statistically significant;

[0062] Step 2: Obtain the importance score of each feature based on the difference between the original sample and the generated counterfactual sample, including: for each feature, and counterfactual samples Extract the value of this feature from the sample to form two independent samples. The sample size is ;

[0063] First, calculate the combined standard deviation S;

[0064] Then, the T statistic of the two samples is calculated, and the specific implementation is shown in formula (VIII):

[0065] (VIII);

[0066] In formula (VIII), Represents the operation of calculating T statistics, Represents the calculation expectation operation;

[0067] According to the calculated T statistic, the counterfactual difference of each feature is defined as shown in formula (IX):

[0068] (IX);

[0069] In formula (IX), represents the calculation of counterfactual difference value, Represents the operation of calculating T statistics;

[0070] Find the corresponding P value through the T statistic distribution table; the T statistic distribution table is also called the Student's t distribution table, which is a pre-calculated statistical table that lists the critical t values ​​under different degrees of freedom and significance levels;

[0071] The feature prediction importance calculates the influence of each feature on the prediction result. Specifically, after the training of the industrial control anomaly detection model is completed, each feature is converted into a random noise in turn to erase the effect of the feature on the prediction. After obtaining the true prediction value and the actual prediction value, the influence of each feature on the prediction is normalized. The specific implementation method is shown in formula (X):

[0072] (X);

[0073] In formula (X), Represents the feature prediction importance value, represents the normalization operation, represents the prediction result of the industrial control anomaly detection model, I represents the set of all feature sequences, Representative in the entire series The operation of erasing features;

[0074] Finally, the weighted sum is taken as the final feature importance value to form a difference spectrum. The specific method is shown in formula (XI):

[0075] (XI);

[0076] In formula (XI), represents the final feature importance value, Represents the feature prediction importance value, represents the calculation of counterfactual difference value, Used to align values;

[0077] Construct a feature importance matrix. The feature importance matrix is ​​a two-dimensional array, in which rows represent different abnormal samples and columns represent different important features. Each element in the feature importance matrix represents the importance of a specific abnormal sample on a specific important feature. Score value; according to the preset threshold, the features that significantly affect the anomaly detection results are screened out, and for each significant feature, an explanation statement is generated to describe its degree and direction of anomaly.

[0078] The industrial control anomaly explanation system based on generative counterfactual sample differences includes:

[0079] The industrial control mixed data set acquisition and preprocessing module is configured to: predict the anomaly score result according to the industrial control anomaly detection model, obtain the industrial control mixed data set through the anomaly score result and the originally collected multi-sensor time series data set, and perform preprocessing, and divide it into a training set and a test set after preprocessing;

[0080] The conditional variational autoencoder training module is configured to: take the original time series data set as input, and the anomaly score output by the industrial control anomaly detection model as a condition, and input it into the conditional variational autoencoder for training;

[0081] The anomaly threshold prediction module is configured to: collect the anomaly threshold output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder;

[0082] The industrial control anomaly interpretation module is configured to obtain the feature importance score by comparing the counterfactual samples with the originally collected multi-sensor time series samples. The size of the feature importance score represents the contribution of the feature to the prediction of the detection results by the anomaly detection model.

[0083] The beneficial effects of the present invention are:

[0084] In the field of anomaly detection and interpretation of industrial control systems, there have long been problems such as insufficient accuracy, lack of interpretability, and limited practicality. These challenges have seriously affected the application effect of anomaly detection technology in actual industrial control environments. Compared with the prior art, the industrial control anomaly interpretation method based on generative counterfactual sample differences under anomaly score conditions proposed in this invention has the following significant advantages:

[0085] 1. Accurately locate the root cause of anomalies: The present invention cleverly constructs a "hypothesis-verification" framework by generating counterfactual samples. This method can not only identify anomalies, but also accurately locate the key factors that cause anomalies through comparative analysis. By generating counterfactual samples that are highly similar to abnormal samples but do not trigger alarms, this method can provide technicians with an intuitive and quantifiable abnormality explanation mechanism, which can help improve the accuracy and credibility of abnormal diagnosis and provide more reliable decision support for troubleshooting of industrial control systems.

[0086] 2. Deeply capture data complexity: The conditional variational autoencoder (CVAE) structure adopted by this invention breaks through the limitations of traditional methods in processing high-dimensional nonlinear data. By taking the anomaly score as a condition for the generation process, this method can accurately adjust key features while maintaining the overall distribution characteristics of the data. Compared with traditional interpretation methods, this structural design takes into account the continuity of time series data and can effectively capture the complex interactive relationship between multiple sensors.

[0087] 3. Multi-dimensional difference evaluation: The present invention introduces a feature difference mapping mechanism to achieve a comprehensive comparison of counterfactual samples and original abnormal samples. Through multi-angle and multi-scale analysis methods, on the one hand, sample differences are evaluated from the perspective of overall distribution, and on the other hand, the degree of change of each feature is quantified at the specific feature level. By comprehensively using a variety of advanced difference measurement indicators, accurate quantification of the magnitude of feature changes is achieved, and the weight of the impact of these changes on the system state can be evaluated. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 This is the overall framework diagram of the conditional variational autoencoder (CVAE) of the present invention;

[0089] Figure 2 This is a schematic diagram of the principal component analysis PCA visualization result of the conditional variational autoencoder reconstructing the samples of the present invention;

[0090] Figure 3 This is a schematic diagram of the T-SNE visualization result of the conditional variational autoencoder reconstructing the sample of the present invention;

[0091] Figure 4 A schematic diagram showing the p-value results of calculating the importance score of time series features in the present invention. DETAILED DESCRIPTION

[0092] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.

[0093] Example 1

[0094] The industrial control anomaly explanation method based on generative counterfactual sample differences includes:

[0095] The anomaly score results are predicted according to the industrial control anomaly detection model. The industrial control mixed data set is obtained through the anomaly score results and the originally collected multi-sensor time series data set, and preprocessed. After preprocessing, it is divided into a training set and a test set.

[0096] The original time series dataset is used as input, and the anomaly score output by the industrial control anomaly detection model is used as a condition and input into the conditional variational autoencoder (CVAE) for training;

[0097] Collect the anomaly thresholds output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder;

[0098] The feature importance score is obtained by comparing the counterfactual samples with the originally collected multi-sensor time series samples. The size of the feature importance score represents the contribution of the feature to the prediction of the detection results (such as anomalies) by the anomaly detection model.

[0099] Based on the industrial control anomaly detection model, such as but not limited to the long short-term memory network (LSTM) structure, Transformer structure, information bottleneck (Information Bottleneck) structure and other deep learning architectures, the abnormal pattern is identified by analyzing the time dependency of time series data. The model outputs the anomaly score at each time point, and the higher the score, the greater the possibility of anomaly.

[0100] Example 2

[0101] The difference between the industrial control anomaly interpretation method based on generative counterfactual sample differences described in Example 1 is that:

[0102] Continuous data is collected from various sensors (such as temperature, pressure, flow, etc.) in the industrial control system at fixed time intervals as the original time series data set, namely the multi-sensor time series data set. These data reflect the operating status of the industrial control system at different time points.

[0103] The original time series dataset is combined with the anomaly score result as an industrial control mixed dataset. That is, at each time point, the content of the industrial control mixed dataset is the operating status data collected at the current time point and the corresponding anomaly score.

[0104] Pre-processing; including:

[0105] 1) Data cleaning: remove missing values;

[0106] 2) Data normalization: Normalize features of different scales to the same range, such as [0, 1].

[0107] After the preprocessing operation, the industrial control mixed dataset is divided into training and testing sets;

[0108] The original multi-sensor time series is processed by the industrial control anomaly detection model and stored according to the preset time window sampling, including:

[0109] Assume that in an industrial control system, the data set is defined express Devices and time steps The measured device values ​​obtained on the device include sensors and actuators. Indicates Devices, Yes Equipment In time Defining the industrial control anomaly detection model , then the industrial control anomaly detection model is used for the preprocessed multi-sensor time series dataset The predicted anomaly score is , .

[0110] like Figure 1 As shown, the encoder, decoder, anomaly score embedding module, and reparameterized sampling module of the conditional variational autoencoder are constructed; conditional parameters are input into the encoder of the conditional variational autoencoder, that is, the anomaly score is predicted to obtain the mean and variance; conditional parameters are also input into the decoder to guide the reconstruction training of the conditional variational autoencoder; including:

[0111] Build a conditional variational autoencoder (CVAE), including building an encoder, a decoder, anomaly score embedding module, and a reparameterized sampling module, and introduce conditional information in the encoding and decoding process. The conditional information is: anomaly score, which represents the numerical value of the degree of abnormality of the data in each time window, so as to realize data generation under specific conditions;

[0112] In the anomaly score embedding module, a linear layer is used to embed the input anomaly score once, encode the score, and embed the conditional information into a continuous vector space for easy merging. The specific method is shown in formula (I):

[0113] (I);

[0114] In formula (I), is the anomaly score result after embedding transformation in the linear layer, represents the ReLU activation function, Represents an operation through a linear layer. Represents the anomaly score predicted by the industrial control anomaly detection model. The transformed anomaly score dimension is changed from Transformed to , define the final output dimension as 2; satisfying the size of the expected and contrast concatenated vectors in subsequent operations.

[0115] The embedded anomaly score is merged with the encoder input tensor. The purpose is to guide the model to learn through the anomaly score results and reconstruct the sampling points and conditional information in the latent space into the output of the original data space. The encoder includes three hidden layers. The specific method is shown in formula (II):

[0116] (II);

[0117] In formula (II), A tensor representing the merged raw time series dataset, used as the input to the encoder, Represents a merge operation. Represents an operation through a linear layer. represents the ReLU activation function, Represents the latent space vector, which is the output of the encoder;

[0118] Similarly, in the decoder, the anomaly score is used to guide model learning, mapping the original time series data set and conditional information to the mean and logarithmic variance of the latent space; the decoder and the encoder are symmetric results, and the specific implementation is shown in formula (III):

[0119] (III);

[0120] In formula (III), are the mean and logarithmic standard deviation of the encoder output, represents the sampling operation, The operation in formula (II) A tensor representing the merged original time series dataset, used as the input to the decoder, Represents a merge operation. Represents an operation through a linear layer. represents the ReLU activation function, represents the latent space vector, i.e., the output of the encoder, is the result of the decoder output, that is, the reconstruction result of the conditional variational autoencoder;

[0121] The output of the decoder is input into the reparameterized sampling module to ensure the differentiability of the CVAE model. The random sampling process is converted into a deterministic function that accepts a random noise as input. The purpose is to ensure the differentiability of the CVAE model sampling so that back propagation and gradient calculation can be performed. The specific method is shown in formula (IV):

[0122] (IV);

[0123] In formula (IV), are the mean and logarithmic standard deviation of the encoder output, represents the tensor splitting function, represents the latent space sampling point, represents random noise sampled from a standard normal distribution N(0, 1).

[0124] Training a conditional variational autoencoder; including:

[0125] Initialize the encoder, decoder, and anomaly score embedding module; the encoder and decoder use a multi-layer neural network structure, using the Linear layer and ReLU activation function; the anomaly score embedding module converts conditional information into a continuous vector representation;

[0126] In each training iteration, the conditional variational autoencoder receives input data x and conditional information score; the anomaly score embedding module processes the conditional information through formula (I) to generate an embedding vector; the encoder receives input data and embedding vector and outputs the mean of the latent space and log variance ; Through the reparameterized sampling module, the latent variables are sampled from the latent space distribution through formula (IV) ; The decoder uses formula (III) to use hidden variables and conditional embedding to reconstruct input data;

[0127] After each iteration, the gradients are calculated through back-propagation and the parameters of the conditional variational autoencoder are updated using the Adam optimizer.

[0128] Loss function of conditional variational autoencoder Including: reconstruction error and KL divergence, that is: ,in, represents the total loss, represents the reconstruction loss, Represents KL divergence loss; reconstruction loss is used to evaluate the reconstruction quality, and KL divergence is used to regularize the latent space distribution; by constraining the latent space distribution, the generalization ability of the model is enhanced.

[0129] The calculation method of reconstruction loss is shown in formula (V):

[0130] (V);

[0131] In formula (IV), represents the mean square error, represents the mean absolute error, is the weight coefficient, represents the original input data, Represents the reconstructed data;

[0132] The calculation method of KL divergence loss is shown in formula (VI):

[0133] (VI);

[0134] In formula (VI), represents the mean value of the encoder output, Represents the variance of the encoder output.

[0135] The original time series data set is used as the input of the conditional variational autoencoder. The anomaly detection model outputs the anomaly score result as the conditional input to the CVAE for training. The original anomaly score is adjusted to simulate the performance of the sample under abnormal conditions. A "counterfactual" feature vector different from the original feature vector is determined in the latent space for sampling to generate counterfactual samples.

[0136] Collect the anomaly thresholds output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder; including:

[0137] Based on the trained conditional variational autoencoder, it is assumed that the conditional variational autoencoder can accept the original sample and the anomaly score as input and generate the reconstructed sample; Figure 2 This is a schematic diagram of the principal component analysis PCA visualization result of the conditional variational autoencoder reconstructing the samples of the present invention; Figure 3 This is a schematic diagram of the T-SNE visualization result of the conditional variational autoencoder reconstructing the sample of the present invention; the conditional variational autoencoder receives a test data loader, and the test data loader includes the original sample , Raw Anomaly Score and anomaly threshold ;

[0138] Set a new threshold , the original anomaly score Make adjustments; the specific operations are:

[0139] Increase the anomaly threshold of each sample by a fixed value ; By improving the anomaly score, simulate the performance of the sample under abnormal conditions. The specific method is shown in formula (VII):

[0140] (VII);

[0141] The adjusted anomaly threshold With the original sample are input into the trained conditional variational autoencoder together; in the generation mode, the trained conditional variational autoencoder is based on the adjusted anomaly threshold With the original sample , generate counterfactual samples , ; This process is performed in the generative mode of the model to ensure that the model parameters are not changed. In the generative mode, the conditional variational autoencoder generates a set of , making the smallest change in the features to flip the prediction results of the conditional variational autoencoder, making the smallest change in the features means: Transformed into .

[0142] Construct a multi-dimensional feature difference mapping mechanism to reveal the influence of key features by comparing the positions of the original sample and the generated counterfactual sample in the feature space, and explain the decision of each feature on the anomaly detection model. This process involves analyzing the changes of each feature between the original sample and the counterfactual sample, and how these changes affect the output of the anomaly detection model. By comparing the difference of each feature in the original sample and the counterfactual sample, the contribution of each feature to the decision of the anomaly detection model can be quantified; use high-dimensional feature extraction technology to obtain a comprehensive feature representation from the original sample to form a difference spectrum. The difference spectrum is a multi-dimensional data structure used to represent and analyze the difference distribution between the original sample and the corresponding counterfactual sample in each feature dimension, including the size, direction, and correlation of the difference; construct a feature difference matrix; the feature difference matrix is ​​a two-dimensional array, in which rows represent different sample pairs (original samples and corresponding counterfactual samples), columns represent different features, and each element in the matrix represents the difference value of a specific sample pair on a specific feature, reflecting different degrees of feature changes. Including:

[0143] Step 1: After obtaining the counterfactual sample, compare the data changes between the original sample and the generated counterfactual sample to obtain the difference between the features; use the independent sample T test to compare each feature of the original sample and the generated counterfactual sample, and calculate the significant difference between the original time series and the generated counterfactual time series;

[0144] T test is a statistical method used to compare whether there is a significant difference between the means of two independent samples. By calculating the T statistic and the corresponding P value, it can be determined whether the difference between the two groups of samples on a certain feature is statistically significant. The assumptions for defining the independent sample T test are as follows: Null hypothesis (H0): The means of the two groups of samples are equal; Alternative hypothesis (H1): The means of the two groups of samples are not equal; By calculating the T statistic and querying the corresponding P value, decide whether to reject the null hypothesis; If the P value is less than the significance level, reject the null hypothesis and believe that there is a significant difference in the means of the two groups of samples; The smaller the P value, the more significant the difference in the feature; The significance level is 0.05 or 0.01, which is a commonly used standard in statistics, otherwise, if the P value is greater than or equal to the significance level, the null hypothesis cannot be rejected and the difference in the means of the two groups of samples is considered not statistically significant; In this case, it can be considered that there is no difference in the feature between the original sample and the counterfactual sample. Figure 4 A schematic diagram showing the p-value results of calculating the importance score of time series features in the present invention.

[0145] Step 2: Obtain the importance score of each feature based on the difference between the original sample and the generated counterfactual sample, including: for each feature, and counterfactual samples Extract the value of this feature from the sample to form two independent samples. The sample size is ;

[0146] First, calculate the combined standard deviation S;

[0147] Then, the T statistic of the two samples is calculated, and the specific implementation is shown in formula (VIII):

[0148] (VIII);

[0149] In formula (VIII), Represents the operation of calculating T statistics, Represents the calculation expectation operation;

[0150] According to the calculated T statistic, the counterfactual difference of each feature is defined as shown in formula (IX):

[0151] (IX);

[0152] In formula (IX), represents the calculation of counterfactual difference value, Represents the operation of calculating T statistics;

[0153] Find the corresponding P value through the T statistic distribution table; the smaller the P value, the more significant the difference between the two sets of data on the feature, which is used to evaluate the importance of each feature by comparing the difference between the original sample and the generated sample. The T statistic distribution table, also known as the Student's t distribution table, is a pre-calculated statistical table that lists the critical t values ​​under different degrees of freedom and significance levels;

[0154] The counterfactual generator can generate corresponding counterfactual samples according to the specified anomaly score, capturing the relationship between time series features. However, when changing the feature value, no relevant constraints are made, which may lead to too many unexpected explanation results in the generated time series (for example, the model only tends to change individual features or some features lose their discreteness in the process of change). Therefore, an additional feature prediction importance is added. Different from the feature importance obtained from the difference between the counterfactual sample and the original sample, the feature prediction importance calculates the influence of each feature on the prediction result. Specifically, after the training of the industrial control anomaly detection model is completed, each feature is converted into a random noise in turn to erase the effect of the feature on the prediction. After obtaining the true prediction value and the actual prediction value, the influence of each feature on the prediction is normalized. The specific implementation method is shown in formula (X):

[0155] (X);

[0156] In formula (X), Represents the feature prediction importance value, represents the normalization operation, represents the prediction result of the industrial control anomaly detection model, I represents the set of all feature sequences, Representative in the whole set The operation of erasing features;

[0157] Finally, the weighted sum is taken as the final feature importance value to form a difference spectrum. The specific method is shown in formula (XI):

[0158] (XI);

[0159] In formula (XI), represents the final feature importance value, Represents the feature prediction importance value, represents the calculation of counterfactual difference value, Used to align values; preferred value is 0.001.

[0160] Feature importance reflects the suspicious features that cause anomalies and provides feasible actions on how to change the features to cause anomalies.

[0161] Construct a feature importance matrix. The feature importance matrix is ​​a two-dimensional array, in which rows represent different abnormal samples and columns represent different important features. Each element in the feature importance matrix represents the importance of a specific abnormal sample on a specific important feature. Score value; reflects the degree and direction of anomaly. When explaining the decision of each feature to the anomaly detection model, The higher the score, the greater the contribution of the feature to anomaly detection. Then, based on the preset threshold (for example, the absolute value is greater than 2), the features that significantly affect the anomaly detection results are screened out. For each significant feature, an explanation statement is generated to describe its degree and direction of abnormality (for example, "the pressure value is abnormally higher than the normal level, with a Z score of 3.5").

[0162] Example 3

[0163] The industrial control anomaly explanation system based on generative counterfactual sample differences includes:

[0164] The industrial control mixed data set acquisition and preprocessing module is configured to: predict the anomaly score result according to the industrial control anomaly detection model, obtain the industrial control mixed data set through the anomaly score result and the originally collected multi-sensor time series data set, and perform preprocessing, and divide it into a training set and a test set after preprocessing;

[0165] The conditional variational autoencoder training module is configured to: take the original time series data set as input, and the anomaly score output by the industrial control anomaly detection model as a condition, and input it into the conditional variational autoencoder for training;

[0166] The anomaly threshold prediction module is configured to: collect the anomaly threshold output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder;

[0167] The industrial control anomaly interpretation module is configured to obtain the feature importance score by comparing the counterfactual samples with the originally collected multi-sensor time series samples. The size of the feature importance score represents the contribution of the feature to the prediction of the detection results by the anomaly detection model.

Claims

1. An industrial control anomaly explanation method based on generative counterfactual sample differences, characterized by: include: The anomaly score results are predicted according to the industrial control anomaly detection model. The industrial control mixed data set is obtained through the anomaly score results and the originally collected multi-sensor time series data set, and preprocessed. After preprocessing, it is divided into a training set and a test set. The original time series dataset is used as input, and the anomaly score output by the industrial control anomaly detection model is used as a condition and input into the conditional variational autoencoder for training; Collect the anomaly thresholds output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder; The feature importance score is obtained by comparing the counterfactual samples with the originally collected multi-sensor time series samples. The feature importance score represents the contribution of the feature to the prediction of the detection results by the anomaly detection model. Construct the encoder, decoder, anomaly score embedding module, and reparameterized sampling module of the conditional variational autoencoder; Input conditional parameters into the encoder of the conditional variational autoencoder, that is, predict the anomaly score and obtain the mean and variance; input conditional parameters into the decoder to guide the reconstruction training of the conditional variational autoencoder; including: Build a conditional variational autoencoder, including building an encoder, a decoder, anomaly score embedding module, and a reparameterized sampling module, and introduce conditional information in the encoding and decoding process. The conditional information is: anomaly score, which represents the numerical value of the degree of anomaly of data in each time window, so as to realize data generation under specific conditions; In the anomaly score embedding module, a linear layer is used to perform an embedding operation on the input anomaly score row, and the score is encoded so that the conditional information is embedded into a continuous vector space. The specific method is shown in formula (I): Score(x i ) * =(ReLU(Linear(Score(x i )))) (Ⅰ) In formula (I), Score(x i ) * is the abnormal score result after the linear layer embedding transformation. ReLU(·) represents the ReLU activation function, Linear(·) represents the operation after a linear layer, Score(x i ) represents the anomaly score predicted by the industrial control anomaly detection model. The transformed anomaly score dimension changes from R T×D Transform to R T×2 , define the final output dimension as 2; The embedded anomaly score is merged with the encoder input tensor. The encoder includes three hidden layers. The specific method is shown in formula (II): In formula (II), Combined E Represents the tensor after merging the original time series dataset, which is used as the input of the encoder. concate(·) represents the merging operation, Linear(·) represents the operation after passing through a linear layer, ReLU(·) represents the ReLU activation function, and Z represents the latent space vector, that is, the output result of the encoder. Similarly, in the decoder, the original time series data set and conditional information are mapped to the mean and logarithmic variance of the latent space; the decoder and the encoder are symmetric results, and the specific implementation is shown in formula (III): In formula (III), μ, σ are the mean and logarithmic standard deviation of the encoder output, Sample represents the sampling operation, and Encoder(·,·) represents the operation Combined in formula (II). D represents the tensor after merging the original time series dataset, which is used as the input of the decoder. concate(·) represents the merging operation. Linear(·) represents the operation after passing through a linear layer. ReLU(·) represents the ReLU activation function. Z represents the latent space vector, i.e., the output result of the encoder. x i ' is the result of the decoder output, that is, the reconstruction result of the conditional variational autoencoder; The output of the decoder is input into the reparameterized sampling module, which transforms the random sampling process into a deterministic function and accepts a random noise as input. The specific method is shown in formula (IV): In formula (IV), μ, σ are the mean and logarithmic standard deviation of the encoder output, chunk(·) represents the tensor segmentation function, Z represents the latent space sampling point, and ε represents the random noise sampled from the standard normal distribution N(0,1).

2. The method for explaining industrial control anomalies based on generative counterfactual sample differences according to claim 1, characterized in that: The original multi-sensor time series is processed by the industrial control anomaly detection model and stored according to the preset time window sampling, including: Assume that in an industrial control system, the data set x is defined i ∈R T×D represents D devices and the measured device values ​​obtained at time step T. The devices include sensors and actuators. D Indicates the Dth device, x i Is device X D The measured value at time t; define the industrial control anomaly detection model AD(·), then the industrial control anomaly detection model is used for the preprocessed multi-sensor time series data set x i The predicted anomaly score is Score(x i )∈R T×D ,Score(x i )=AD(x i ).

3. The industrial control anomaly explanation method based on generative counterfactual sample differences according to claim 1 is characterized in that: Training a conditional variational autoencoder; including: Initialize the encoder, decoder, and anomaly score embedding module; the encoder and decoder use a multi-layer neural network structure, using the Linear layer and ReLU activation function; the anomaly score embedding module converts conditional information into a continuous vector representation; In each training iteration, the conditional variational autoencoder receives input data x and conditional information score; the anomaly score embedding module processes the conditional information through formula (Ⅰ) to generate an embedding vector; the encoder receives the input data and the embedding vector, and outputs the mean μ and logarithmic variance σ of the latent space; through the reparameterized sampling module, the latent variable Z is sampled from the latent space distribution through formula (Ⅳ); the decoder reconstructs the input data using the latent variable Z and the conditional embedding through formula (Ⅲ); After each iteration, the gradients are calculated through back-propagation and the parameters of the conditional variational autoencoder are updated using the Adam optimizer.

4. The industrial control anomaly interpretation method based on generative counterfactual sample differences according to claim 1 is characterized in that: The loss function L of the conditional variational autoencoder includes: reconstruction error and KL divergence, that is: L = L rec +L KL , where L represents the total loss, L rec represents the reconstruction loss, L KL Represents KL divergence loss; reconstruction loss is used to evaluate the reconstruction quality, and KL divergence is used to regularize the latent space distribution; the calculation method of reconstruction loss is shown in formula (V): L rec =||x i '-x i ||2+λ||x i '-x i ||1 (Ⅴ) In formula (IV), ||x i '-x i ||2 represents the mean square error, ||x i '-x i ||1 represents the mean absolute error, λ is the weight coefficient, x i Represents the original input data, x i ' represents the reconstructed data; The calculation method of KL divergence loss is shown in formula (VI): L KL =-0.5·∑(1+σ-μ 2 -s 2 ) (Ⅵ) In formula (VI), μ represents the mean of the encoder output, and σ represents the variance of the encoder output.

5. The industrial control anomaly interpretation method based on generative counterfactual sample differences according to claim 1 is characterized in that: Collect the anomaly thresholds output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder; include: Based on the trained conditional variational autoencoder, it is assumed that the conditional variational autoencoder can accept the original sample and the anomaly score as input and generate a reconstructed sample; the conditional variational autoencoder receives the test data loader, which includes the original sample x i 、Original anomaly score Score(x i ) and abnormal threshold thr; Set a new threshold thr * , the original anomaly score Score(x i ) to make adjustments; the specific operations are: The abnormal threshold of each sample is increased by a fixed value α; the specific method is shown in formula (VII): thr * =thr+α (Ⅶ) The adjusted abnormal threshold thr * With the original sample x i are input into the trained conditional variational autoencoder together; in the generation mode, the trained conditional variational autoencoder is based on the adjusted anomaly threshold thr * With the original sample x i , generate counterfactual samples Making the smallest change in features flips the prediction results of the conditional variational autoencoder. Making the smallest change in features means: Transformed into 6. The method for explaining industrial control anomalies based on generative counterfactual sample differences according to claim 1, characterized in that: Construct a multi-dimensional feature difference mapping mechanism to reveal the influence of key features by comparing the positions of original samples and generated counterfactual samples in the feature space, explain the decision of each feature on the anomaly detection model, and use high-dimensional feature extraction technology to obtain comprehensive feature representation from the original samples, form a difference spectrum, and construct a feature difference matrix; including: Step 1: After obtaining the counterfactual sample, compare the data changes between the original sample and the generated counterfactual sample to obtain the difference between the features; use the independent sample T test to compare each feature of the original sample and the generated counterfactual sample, and calculate the significant difference between the original time series and the generated counterfactual time series; The assumptions of the independent sample T test are defined as follows: Null hypothesis: the means of the two groups of samples are equal; Alternative hypothesis: the means of the two groups of samples are not equal; By calculating the T statistic and querying the corresponding P value, decide whether to reject the null hypothesis; If the P value is less than the significance level, reject the null hypothesis and believe that there is a significant difference in the means of the two groups of samples; The significance level is 0.05 or 0.01, otherwise, if the P value is greater than or equal to the significance level, the null hypothesis cannot be rejected and it is believed that the difference in the means of the two groups of samples is not statistically significant; Step 2: Obtain the importance score of each feature based on the difference between the original sample and the generated counterfactual sample, including: for each feature, respectively from the original sample x i and counterfactual samples Extract the value of this feature from the sample to form two independent samples. The sample size is n. T ; First, calculate the combined standard deviation S; Then, the T statistic of the two samples is calculated, and the specific implementation is shown in formula (VIII): In formula (VIII), T(·,·) represents the operation of calculating T statistics, and E(·) represents the operation of calculating expectation; According to the calculated T statistic, the counterfactual difference of each feature is defined as shown in formula (IX): In formula (IX), val cf (·) represents the calculation of counterfactual difference value, T(·,·) represents the operation of calculating T statistic; Find the corresponding P value through the T statistic distribution table; the T statistic distribution table is also called the Student's t distribution table, which is a pre-calculated statistical table that lists the critical t values ​​under different degrees of freedom and significance levels; The feature prediction importance calculates the influence of each feature on the prediction result. Specifically, after the training of the industrial control anomaly detection model is completed, each feature is converted into a random noise in turn to erase the effect of the feature on the prediction. After obtaining the true prediction value and the actual prediction value, the influence of each feature on the prediction is normalized. The specific implementation method is shown in formula (X): val pre (i)=Norm(D(X I / {i} ))(X) In formula (X), val pre (·) represents the feature prediction importance value, Norm(·) represents the normalization operation, D(·) represents the prediction result of the industrial control anomaly detection model, I represents the set of all feature sequences, and I / {·} represents the operation of erasing features in the entire set I; Finally, the weighted sum is taken as the final feature importance value to form a difference spectrum. The specific method is shown in formula (XI): choice(i)=choice pre (i)-α·val cf (i)(XI) In formula (XI), val(·) represents the final feature importance value, val pre (·) represents the feature prediction importance value, val cf (·) represents the calculation of counterfactual difference values, and α is used to align the values; Construct a feature importance matrix, which is a two-dimensional array, in which rows represent different abnormal samples and columns represent different important features. Each element in the feature importance matrix represents the val score of a specific abnormal sample on a specific important feature. According to the preset threshold, filter out the features that significantly affect the anomaly detection results. For each significant feature, generate an explanatory statement to describe its degree and direction of abnormality.

7. The method for explaining industrial control anomalies based on generative counterfactual sample differences according to claim 1, characterized in that: Continuous data is collected from each sensor in the industrial control system at fixed time intervals as the original time series data set, namely the multi-sensor time series data set.

8. The industrial control anomaly interpretation method based on generative counterfactual sample differences according to any one of claims 1 to 7, characterized in that: The original time series data set is combined with the anomaly score result as an industrial control mixed data set. That is, at each time point, the content of the industrial control mixed data set is the operating status data collected at the current time point and the corresponding anomaly score; Pre-processing; including: 1) Data cleaning: remove missing values; 2) Data standardization.

9. An industrial control anomaly explanation system based on generative counterfactual sample differences, characterized by: include: The industrial control mixed data set acquisition and preprocessing module is configured to: predict the anomaly score result according to the industrial control anomaly detection model, obtain the industrial control mixed data set through the anomaly score result and the originally collected multi-sensor time series data set, and perform preprocessing, and divide it into a training set and a test set after preprocessing; The conditional variational autoencoder training module is configured to: take the original time series data set as input, and the anomaly score output by the industrial control anomaly detection model as a condition, and input it into the conditional variational autoencoder for training; The anomaly threshold prediction module is configured to: collect the anomaly threshold output by the industrial control anomaly detection model when predicting the data set, and generate counterfactual samples by changing the threshold size in the conditional variational autoencoder; The industrial control anomaly interpretation module is configured to: obtain feature importance scores by comparing counterfactual samples with the originally collected multi-sensor time series samples. The feature importance scores represent the contribution of the features to the prediction of the detection results by the anomaly detection model; Construct the encoder, decoder, anomaly score embedding module, and reparameterized sampling module of the conditional variational autoencoder; Input conditional parameters into the encoder of the conditional variational autoencoder, that is, predict the anomaly score and obtain the mean and variance; input conditional parameters into the decoder to guide the reconstruction training of the conditional variational autoencoder; including: Build a conditional variational autoencoder, including building an encoder, a decoder, anomaly score embedding module, and a reparameterized sampling module, and introduce conditional information in the encoding and decoding process. The conditional information is: anomaly score, which represents the numerical value of the degree of anomaly of data in each time window, so as to realize data generation under specific conditions; In the anomaly score embedding module, a linear layer is used to perform an embedding operation on the input anomaly score row, and the score is encoded so that the conditional information is embedded into a continuous vector space. The specific method is shown in formula (I): Score(x i ) * =(ReLU(Linear(Score(x i )))) (Ⅰ) In formula (I), Score(x i ) * is the abnormal score result after the linear layer embedding transformation. ReLU(·) represents the ReLU activation function, Linear(·) represents the operation after a linear layer, Score(x i ) represents the anomaly score predicted by the industrial control anomaly detection model. The transformed anomaly score dimension changes from R T×D Transform to R T×2 , define the final output dimension as 2; The embedded anomaly score is merged with the encoder input tensor. The encoder includes three hidden layers. The specific method is shown in formula (II): In formula (II), Combined E Represents the tensor after merging the original time series dataset, which is used as the input of the encoder. concate(·) represents the merging operation, Linear(·) represents the operation after passing through a linear layer, ReLU(·) represents the ReLU activation function, and Z represents the latent space vector, that is, the output result of the encoder. Similarly, in the decoder, the original time series data set and conditional information are mapped to the mean and logarithmic variance of the latent space; the decoder and the encoder are symmetric results, and the specific implementation is shown in formula (III): In formula (III), μ, σ are the mean and logarithmic standard deviation of the encoder output, Sample represents the sampling operation, and Encoder(·,·) represents the operation Combined in formula (II). D represents the tensor after merging the original time series dataset, which is used as the input of the decoder. concate(·) represents the merging operation. Linear(·) represents the operation after passing through a linear layer. ReLU(·) represents the ReLU activation function. Z represents the latent space vector, i.e., the output result of the encoder. x i ' is the result of the decoder output, that is, the reconstruction result of the conditional variational autoencoder; The output of the decoder is input into the reparameterized sampling module, which transforms the random sampling process into a deterministic function and accepts a random noise as input. The specific method is shown in formula (IV): In formula (IV), μ, σ are the mean and logarithmic standard deviation of the encoder output, chunk(·) represents the tensor segmentation function, Z represents the latent space sampling point, and ε represents the random noise sampled from the standard normal distribution N(0,1).

Citation Information

Patent Citations

  • Industrial control system-oriented anomaly detection system and method

    CN115484102A

  • Transformer substation safe operation monitoring method and system based on OpenCV

    CN118485973A