Industrial Data Feature Dimensionality Reduction Method Based on Causal Inference
Through the characteristic dimensionality reduction method based on causal inference, using PCMCI+ algorithm and autoencoder, the problems of poor information loss and interpretability in traditional methods are solved, and high accuracy and low cost industrial fault detection are achieved.
Patent Information
- Application Number
- CN202310178391.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-28
AI Technical Summary
The correlation-based feature dimensionality reduction method in the prior art has poor interpretability and low robustness in industrial fault detection, and traditional feature selection methods may lead to information loss, affecting the accuracy of fault detection.
The characteristic dimensionality reduction method based on causal inference is adopted, and the causal relationship between monitoring variables is determined through the PCMCI+ algorithm, a causal relationship matrix is constructed, and the data dimensionality reduction is used using principal component analysis and autoencoder, key information is extracted, and fault detection is performed.
It improves the accuracy of industrial system fault diagnosis, reduces calculation costs, and can perform effective fault detection when only normal operating status data is used, which is highly interpretable and robust.
Smart Images

Figure CN116226648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of feature dimensionality reduction methods, and particularly to an industrial data feature dimensionality reduction method based on causal inference. Background Art
[0002] Fault detection refers to using an equipment model or equipment monitoring data to evaluate the health status of industrial equipment, determine whether the equipment has failed, so as to take timely maintenance measures to avoid production safety accidents caused by equipment failures. Fault detection technology is of great significance for reducing production costs and ensuring the safety of the production process. Fault detection methods mainly include model-based methods and data-driven methods. With the complication of modern industrial production processes, it has become increasingly difficult to build an accurate industrial system model. With the development of intelligent sensors, industrial Internet of Things, and data storage technologies, manufacturing enterprises can obtain a large amount of monitoring data from industrial equipment and use the data for equipment fault detection. At the same time, the improvement of computer operation speed and the development of machine learning algorithms have simplified the process of applying detection data to fault detection. Therefore, data-driven industrial equipment fault diagnosis has become a popular research topic.
[0003] However, a large amount of monitoring data also brings problems such as large computational amount, data redundancy, and noise pollution. As the dimensionality of industrial monitoring data increases, the computational burden of fault monitoring grows exponentially. The noise in the original data also affects the accuracy of fault detection. Therefore, it is necessary to perform data dimensionality reduction processing on the original industrial monitoring data before fault detection to extract the key information of equipment operation from the original data. Feature dimensionality reduction refers to mapping the data in a high-dimensional vector space to a low-dimensional vector space through a certain data processing rule. Feature dimensionality reduction includes feature selection and feature extraction. Feature selection refers to selecting a feature subset from the data that can retain most of the information of the original data. Feature extraction is to combine the useful information in the original data into some new features. Traditional feature dimensionality reduction methods include correlation ranking, principal component analysis, independent component analysis, linear discriminant analysis, isometric mapping, etc., all of which are correlation-based methods. In recent years, research and practice have shown that correlation-based feature dimensionality reduction methods have problems such as poor interpretability, unstable performance, and poor dimensionality reduction effect on high-coupling systems.
[0004] Currently, researchers have proposed a large number of feature selection methods based on causal inference, also known as Markov boundary discovery algorithms. However, feature selection ignores the information contained in features outside the feature subset, which may lead to information loss and thus affect the accuracy of industrial system fault detection. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the present invention provides an industrial data feature dimensionality reduction method based on causal inference, so as to explore the essential relationships among industrial monitoring variables, extract information that can reflect the actual operating state of the system, filter out useless noise, and thereby improve the accuracy of industrial system fault diagnosis and reduce the calculation cost.
[0006] To achieve the above object, the present invention provides an industrial data feature dimensionality reduction method based on causal inference, including the steps of:
[0007] S1: Obtain the industrial system monitoring data set X;
[0008] S2: Perform data preprocessing operations on the obtained industrial system monitoring data set X;
[0009] S3: Use the PCMCI+ causal relationship inference method to obtain the causal relationships among the monitoring variables;
[0010] S4: Convert the causal relationships into a causal relationship matrix;
[0011] S5: Use principal component analysis to reduce the dimensionality of the causal relationship matrix, extract the key information in the causal relationship matrix, and form a feature vector matrix V2;
[0012] S6: Multiply the industrial system monitoring data set X by the feature vector matrix V2 to construct a data set X2 containing k variables, X2 = XV2;
[0013] S7: Use an autoencoder to perform fault detection on the data set X2 after data dimensionality reduction.
[0014] Preferably, in the step S1, the expression of the industrial system monitoring data set X is:
[0015]
[0016] where m represents the number of variables in the data, n represents the number of data points for each variable, and the data point x ij represents the monitoring value of the i-th variable at the j-th moment; where i = 1, 2,..., n, j = 1, 2,..., m; the data set is divided into data collected under the normal operating state of the industrial system and data collected during a fault.
[0017] Preferably, the step S2 further includes the steps of:
[0018] S21: Perform missing value filling operations on the industrial system monitoring data set X to prevent the missing values in the data set from affecting the accuracy of the causal inference and fault detection operation results;
[0019] S22: Perform linear normalization on the industrial system monitoring dataset X to prevent the test results from being dominated by variables with a large order of magnitude, and at the same time alleviate the slowdown of the model convergence speed caused by the excessive difference in the order of magnitude of the monitoring data.
[0020] Preferably, in the S3 step:
[0021] The expression of the causal relationship is (x (t-τ)i , x tj ); where the cause is x (t-τ)i , the result is x tj , and the direction of the causal relationship is from x (t-τ)i to x tj ; The causal relationship specifically includes the object variables of the causal relationship, the direction, the causal relationship strength S(x (t-τ)i , x tj ) and the time delay τ of the causal relationship; the value range of the causal relationship strength is from -1 to 1; the time delay of the causal relationship is a non-negative integer; t represents time.
[0022] Preferably, in the S4 step:
[0023] The causal relationship matrix is an m×m matrix; for the element in the i-th row and j-th column of the causal relationship matrix, if there is a causal relationship between the i-th feature variable and the j-th feature variable in the causal network, then the elements a ij and a ji have the value of the causal relationship strength S(x (t-τ)i , x tj ); if there are multiple causal relationships with different time delays between two variables, then select the causal relationship strength value with the largest absolute value among them; the expression of the causal relationship matrix is:
[0024]
[0025] Where,
[0026]
[0027] Preferably, in the S5 step:
[0028] Construct the covariance matrix of the causal relationship matrix and perform eigenvalue decomposition on the covariance matrix; select the first k eigenvalues with the largest eigenvalues or retain the eigenvalues with a preset proportional difference, and the eigenvectors corresponding to these eigenvalues form the eigenvector matrix V2.
[0029] Preferably, in the S6 step, the fault detection step is divided into an offline training stage and an online detection stage;
[0030] S61: In the offline training stage, use the data collected under the normal operating state of the industrial system after dimensionality reduction through steps S1 - S6 for training; the training objective is to minimize the error between the input and output of the autoencoder; determine the error threshold for fault detection in the offline training stage;
[0031] S62: In the online detection stage, input the monitored data of the industrial system after dimensionality reduction through steps S1 - S6 into the autoencoder; if the error between the input and output is greater than the error threshold determined in step S61, it is determined that the industrial system is in a fault state; if the error is less than the error threshold, it is determined that the industrial system is in a normal operating state.
[0032] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects:
[0033] Compared with the existing industrial data dimensionality reduction technologies, the present invention has the following innovative points:
[0034] 1. Aiming at the disadvantages of the current correlation - based data dimensionality reduction methods, such as weak interpretability and low robustness, a feature extraction method based on causal inference is proposed.
[0035] 2. Use the causal relationship matrix to quantitatively represent the strength of the causal relationship between variables.
[0036] 3. Traditional fault detection methods require the input of datasets under both normal operating conditions and fault states at the same time. However, it is difficult to obtain data samples in the fault state. The present invention only needs to input the data under the normal operating state of the industry during the training process, reducing the cost and difficulty of obtaining monitoring data.
[0037] Compared with the existing industrial data dimensionality reduction technologies, the present invention has the following advantages:
[0038] 1. The industrial data feature dimensionality reduction method based on causal inference proposed by the present invention can reduce the dimension of the original monitoring data, retain the useful information related to the system operating state, and remove the irrelevant noise. This feature dimensionality reduction method can improve the accuracy of industrial system fault diagnosis and reduce the computational cost of fault diagnosis.
[0039] 2. The method proposed by the present invention performs data feature dimensionality reduction based on causal relationships, and has the advantages of strong interpretability and high robustness compared with the traditional correlation - based causal dimensionality reduction methods. It has the advantage of retaining as much effective information as possible compared with the feature selection methods.
[0040] 3. The method proposed by the present invention does not require the intervention of industry experts and only needs the monitored data of the industrial system to complete the task of data feature dimensionality reduction.
[0041] 4. The method proposed by the present invention is not limited to the application scenario of industrial system fault detection, and can also be used in various data analysis tasks, such as analyzing the causes of climate change, investigating customer needs, assisting in decision-making, etc. It helps relevant industry personnel to mine feature information closely related to the target task, so as to perform data analysis, information mining and other operations more efficiently and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of the industrial data feature dimensionality reduction method based on causal inference according to an embodiment of the present invention;
[0043] Figure 2 is a curve graph of the fault detection results of the original data set according to an embodiment of the present invention;
[0044] Figure 3 is a curve graph of the fault detection results of the data set after dimensionality reduction by principal component analysis according to an embodiment of the present invention;
[0045] Figure 4 is a curve graph of the fault detection results of the data set after dimensionality reduction by the dimensionality reduction method based on causal inference according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following is based on the accompanying drawings Figures 1 to 4 , and the preferred embodiments of the present invention are given and described in detail to better understand the functions and features of the present invention.
[0047] The present invention uses the Tennessee Eastman dataset and is programmed in Python 3. The Tennessee Eastman (TE) dataset is a chemical process simulation dataset developed by Eastman Chemical Company of the United States according to the actual industrial production process. The simulation platform for generating the Tennessee Eastman dataset mainly consists of multiple operating units such as a continuous stirred tank reactor, a condenser, a gas-liquid separation column, a stripping column, and a centrifugal compressor. When performing production process simulation, three gas raw materials A, D, and E and a small amount of inert gas B enter the reactor. Raw material C and a certain amount of raw material A enter the simulation process through a partial condenser. The condenser cools the product stream of the reactor and sends it to the gas-liquid separation column for separation. The separated steam is sent to the reactor after being processed by the compressor and recycled. In order to prevent the large accumulation of inert component B and reaction by-products during the reaction process, continuous discharge and recycling treatment are required. The condensed components output from the gas-liquid separation column are pumped to the stripping column for treatment. The remaining raw materials that have not reacted are combined with the recycle stream and participate in the subsequent production process. Finally, products G and H output from the bottom of the stripping column are sent to the downstream process. The Tennessee Eastman dataset has time-varying, strongly coupled, and nonlinear characteristics and is widely used to test the practical application effects of complex industrial process fault diagnosis methods. The Tennessee Eastman dataset has a total of 52 variables, including 41 process variables and 11 control variables. The sampling interval for all variables is 3 minutes. The Tennessee Eastman dataset consists of a training set and a test set, and each of the training set and the test set has 22 data samples. The Tennessee Eastman simulation system has a total of 22 operating states, including 1 normal operating state and 21 fault states. Each operating state corresponds to a data sample in the training set and the test set respectively. The simulation time of the normal operating data sample in the training set is 25 hours, and the number of data points for each variable is 500. The simulation time of the data sample with a fault in the training set is 24 hours, and the number of data points for each variable is 480, and the fault is introduced from the start of the simulation system operation. The simulation time of the data samples in the test set is 48 hours, and the number of data points for each variable is 960. In the test set, the data sample with a fault starts to introduce the fault from the 161st data point, that is, the simulation system is in a normal operating state for the first 160 data points (the first 8 hours), and the simulation system is in a fault state for the next 800 data points (the next 40 hours).
[0048] Please refer to Figures 1 to 4 , a method for reducing the dimensionality of industrial data features based on causal inference according to an embodiment of the present invention, includes the following steps:
[0049] Step S1, collect the normal operating data of the Tennessee Eastman simulation process, that is, the data of the control variables and process variables when the simulation process is operating normally;
[0050] Step S2: Perform data preprocessing on the collected bearing data. Since there are no missing values in the Tennessee - Eastman process dataset, only normalization is carried out. The normalization of data is one of the key steps in causal inference and fault detection, which can avoid the dominant position of variables with large magnitudes caused by different orders of magnitude and alleviate the slowdown of the iterative convergence speed caused by different orders of magnitude. The present invention adopts a linear normalization method, and its function expression is:
[0051]
[0052] where the maximum value of the column where the data point x ij is located is denoted as max{x j}, and the minimum value of the column where the data point x ij is located is denoted as min{x j}.
[0053] Step S3: Use the PCMCI+ causal relationship inference method to obtain the causal relationships among the monitoring variables under normal operating conditions. The causal relationship test process of PCMCI+ is divided into a lag causal relationship test stage and an instantaneous causal relationship test stage: PCMCI+ is a causal inference algorithm that can detect lag and instantaneous causal relationships in high - dimensional non - linear data. PCMCI+ is a two - stage method. First, it detects the lag causal relationship, and then it detects the instantaneous causal relationship and tests whether the causal relationship is true.
[0054] Step S3.1: In the lag causal relationship test stage, detect the lag cause x tj of the data point x (t-τ)i , where τ is a positive integer. To simplify the calculation, a maximum time delay τ max is often set and τ is made not greater than τ max . In the present invention, τ max = 60.
[0055] Step S3.1.1: Assume that all data points lagging behind x tj are the direct causes of x tj , that is, all initial points in the set are candidate direct causes of x tj :
[0056]
[0057] Step S3.1.2: Use the conditional independence test to test the causal relationship (x (t-τ)i , x tj ). is a set of data points, and its initial value is an empty set. If x tj and x (t-τ)iIf they are conditionally independent, then the lagged causality (x (t-τ)i , x tj ) does not hold, and the algorithm terminates. If x tj and x (t-τ)i are conditionally correlated, go to step S3.1.3;
[0058] Step S3.1.3: Add a data point to the set , and go to step S3.1.2 to continue the iteration until no more data points can be added or the number of data points in tj and x (t-τ)i are still conditionally correlated, then the lagged causality (x (t-τ)i , x tj ) holds, and the algorithm terminates.
[0059] Step S3.2: In the instantaneous causality test phase, detect the instantaneous cause x tj of the data point x ti , and verify whether the lagged causality output in step 3.1 is correct. The specific method is to perform a conditional independence test CI(x (t-τ)i , x tj , Z). Among them, Z is the union of the sets and . Among them, the set is a subset of all elements with a count of p in the set {x ti |i ≠ j}. The output causal relationship diagram is as Figure 2 shown.
[0060] Step S4: Convert the causal relationship into a causal relationship matrix. The causal relationship matrix is a 52-row and 52-column matrix. For the element in the i-th row and j-th column of the causal relationship matrix, if there is a causal relationship between the i-th feature variable and the j-th feature variable in the causal network, then the value of the element a ij and a ji is the causal relationship strength S(x (t-τ)i , x tj ). If there are multiple causal relationships with different time delays between two variables, then select the strength value with the largest absolute value. The expression of the causal relationship matrix is:
[0061]
[0062] Among them,
[0063]
[0064] Step S5: Use principal component analysis to reduce the dimension of the causal relationship matrix and extract the key information in the causal relationship matrix. Construct the covariance matrix of the causal relationship matrix and perform eigenvalue decomposition on the covariance matrix. In the present invention, eigenvalues that retain 99% of the differences in the original causal relationship matrix are selected, and 43 eigenvalues are recombined. The eigenvectors corresponding to these eigenvalues are V2, which is a matrix with 52 rows and 43 columns.
[0065] Step S6: Multiply the monitoring data set X by the eigenvector matrix V2 to construct a data set X2 = XV2 containing 43 variables. The monitoring data set mentioned here includes the training set under normal working conditions and the data set to be detected (test set).
[0066] Step S7: Use an autoencoder to perform fault detection on the data set X2 after data dimension reduction. Fault detection is divided into an offline training stage and an online detection stage:
[0067] Step S7.0: In the offline training stage, use the data set X2 for training. The training objective is to minimize the error between the input and output of the autoencoder. Determine the error threshold for fault detection in the offline training stage. In the present invention, the squared prediction error (SPE) is used to determine the error threshold. Denote the squared prediction error of the i-th variable as:
[0068]
[0069] where x i and are respectively the input and output values of the i-th variable in the data set X2. The probability density of the squared prediction error is predicted using the kernel density estimation method:
[0070]
[0071] Let ε be the preset threshold. In the present invention, following the practices of most practical applications, let ε = 0.01. Let to obtain the warning upper limit SPE of the squared prediction error thr .
[0072] Step S7.1: In the online detection stage, multiply the data set X' to be detected by the eigenvector matrix V2:
[0073] X ′ 2 = X'V2
[0074] Input the data set X ′ 2 after data dimension reduction processing into the autoencoder. If the squared prediction error SPE new between the input and output is greater than or equal to the fault alarm upper limit threshold SPE thr, it is determined that the industrial system is in a fault state. If the error is less than the alarm threshold, it is determined that the industrial system is in a normal operating state.
[0075] The present invention performs fault detection based on the original data set, the data after dimensionality reduction by principal component analysis, and the data after dimensionality reduction by the method proposed by the present invention, respectively. Figures 2 to 4 are the prediction results of these three methods. In the figure, the abscissa represents the data points, and the ordinate represents the squared prediction error between the input data and the output data of the autoencoder. The first 160 data points represent the data points collected under normal operating conditions of the simulation system, and the last 800 data points represent the data points collected under fault conditions. The black dashed line is the alarm threshold SPE thr . Figure 2 is the fault detection result of the original data set. The squared prediction error fluctuates around the alarm threshold, and fault detection cannot be performed. Figure 3 is the fault detection result of the data after dimensionality reduction by principal component analysis. From Figure 3 it can be seen that the curve of the first 160 data points fluctuates around the alarm threshold, indicating a high false alarm rate of faults. Figure 4 is the fault detection result of the data after dimensionality reduction by the method proposed by the present invention. From Figure 4 it can be seen that there is a significant difference in the squared prediction error between when the system is operating normally and when it is faulty, and it can basically be distinguished by the alarm threshold. Table 1 shows the comparison of the fault detection results of the three data sets. Fault prediction using the data after dimensionality reduction by the causal inference method achieves both high recall and high precision:
[0076] Table 1 Comparison table of fault detection results of three data sets
[0077] Original data Data after dimensionality reduction by principal component analysis Data after dimensionality reduction by causal inference method Recall rate 0.1820 0.9835 0.9425 Precision rate 0.9625 0.7560 0.9937 F1-score 0.3061 0.8548 0.9674
[0078] The present invention has been described in detail above in conjunction with the embodiments with reference to the drawings. Those of ordinary skill in the art can make various variations of the present invention according to the above description. Therefore, certain details in the embodiments should not constitute a limitation to the present invention, and the protection scope of the present invention will be defined by the scope of the appended claims.
Claims
1. An industrial data feature dimensionality reduction method based on causal inference, comprising the steps: S1: Obtain the industrial system monitoring data set X; S2: Perform data preprocessing operations on the obtained industrial system monitoring data set X; S3: Use the PCMCI+ causal relationship inference method to obtain the causal relationships between monitoring variables; S4: Convert the causal relationships into a causal relationship matrix; S5: Use principal component analysis to reduce the dimensionality of the causal relationship matrix, extract the key information in the causal relationship matrix, and form a feature vector matrix V2; S6: Multiply the industrial system monitoring data set X by the feature vector matrix V2 to construct a data set X2 containing k variables, X2 = XV2; S7: Use an autoencoder to perform fault detection on the data set X2 after dimensionality reduction.
2. The method for dimensionality reduction of industrial data features based on causal inference according to claim 1, wherein In the S1 step, the expression of the industrial system monitoring data set X is: where m represents the number of variables in the data, n represents the number of data points for each variable, and the data point x ij represents the monitored value of the i-th variable at the j-th moment; where i = 1, 2, …, n, j = 1, 2, …, m; the data set is divided into data collected under normal operating conditions of the industrial system and data collected during a fault.
3. The method for dimensionality reduction of industrial data features based on causal inference according to claim 2, wherein, The S2 step further includes the steps: S21: Perform missing value filling operations on the industrial system monitoring data set X; S22: Perform linear normalization processing on the industrial system monitoring data set X.
4. The method for reducing the dimensionality of industrial data features based on causal inference according to claim 3, wherein In the S3 step: The expression of the causal relationship is (x (t-τ)i , x tj ); where the cause is x (t-τ)i , the result is x tj , and the direction of the causal relationship is from x (t-τ)i to x tj ; The causal relationship specifically includes the object variables of the causal relationship, direction, causal relationship strength S(x (t-τ)i , x tj ) and the time delay τ of the causal relationship; The value range of the causal relationship strength is from -1 to 1; The time delay of the causal relationship is a non-negative integer; t represents time.
5. The industrial data feature dimensionality reduction method based on causal inference according to claim 4, wherein In the S4 step: The causal relationship matrix is an m×m matrix; for the element in the i-th row and j-th column of the causal relationship matrix, if there is a causal relationship between the i-th feature variable and the j-th feature variable in the causal network, then the element a ij in the causal relationship matrix is the same as a ji , and its value is the causal relationship strength S(x (t-τ)i , x tj ); if there are multiple causal relationships with different time delays between two variables, then the causal relationship strength value with the largest absolute value is selected; the expression of the causal relationship matrix is: Wherein, 6. The method for dimensionality reduction of industrial data features based on causal inference according to claim 5, wherein In the S5 step: Construct the covariance matrix of the causal relationship matrix and perform eigenvalue decomposition on the covariance matrix; Select the first k eigenvalues with the largest eigenvalues or retain the eigenvalues with a preset proportional difference, and the eigenvectors corresponding to these eigenvalues form the feature vector matrix V2.
7. The method for dimensionality reduction of industrial data features based on causal inference according to claim 6, wherein, In the S6 step, the fault detection step is divided into an offline training stage and an online detection stage; S61: In the offline training stage, use the data collected under the normal operation state of the industrial system after being dimensionally reduced through steps S1 to S6 for training; The training objective is to minimize the error between the input and output of the autoencoder; Determine the error threshold for fault detection in the offline training stage; S62: In the online detection stage, input the industrial system monitoring data after being dimensionally reduced through steps S1 to S6 into the autoencoder; If the error between the input and output is greater than the error threshold determined in step S61, it is determined that the industrial system is in a fault state; If the error is less than the error threshold, it is determined that the industrial system is in a normal operation state.
Citation Information
Patent Citations
Detection method for causal connection strength of magnetic resonance brain imaging based on PCA (Principal component analysis) and GCA (Granger causality analysis)
CN102366323A
Causal relationship adjacency matrix feature extraction method for fault diagnosis
CN111860686A