JB-OCCA-PCA method for quality-related fault detection
By applying the JB-OCCA-PCA method in industrial processes, the problem that the prior art is difficult to capture fault characteristics when processing high-dimensional, dynamic and noise-free data is solved, and the accurate distinction and detection of quality-related faults is achieved, and the accuracy and reliability of detection are improved.
Patent Information
- Application Number
- CN202510351845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively capture the fault characteristics of data when processing high-dimensional, dynamic and noise-free data in industrial processes, especially in distinguishing between quality-related and irrelevant faults.
A JB-OCCA-PCA method is proposed. By standardizing the data and subspace division, quality-related features are extracted using Jarque-Bera statistics and orthogonal typical correlation analysis (OCCA), and dimensionality reduction is reduced by principal component analysis (PCA), and finally the detection results are integrated through Bayesian method.
This method can effectively process Gaussian non-Gaussian mixed data, accurately distinguish between quality-related and irrelevant faults, improve the accuracy and reliability of fault detection, and reduce the false alarm rate and missed alarm rate.
Smart Images

Figure CN120180274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting quality-related faults, and more particularly to a JB-OCCA-PCA method for quality-related fault detection. Background Art
[0002] In modern industrial processes, real-time monitoring and fault detection are crucial for ensuring production safety and product quality. Due to the complex characteristics of industrial process data, such as high dimensionality, dynamics, and noise, there may be significant differences in the data distribution characteristics of variables in industrial processes. For example, some data conforms to the Gaussian distribution, while some data exhibits non-Gaussianity. This makes it difficult for a single fault detection method to effectively capture all the fault characteristics of the data. In addition, quality-related fault detection requires more precise methods to distinguish whether a fault is related to quality or not. Existing fault detection methods mainly include principal component analysis (PCA), partial least squares (PLS), and canonical correlation analysis (CCA), etc. However, these methods have certain limitations in dealing with non-Gaussian data and quality-related fault detection. To address these problems, multi-subspace partitioning and feature extraction techniques have received extensive attention. This method decomposes the data into different subspaces so that data analysis methods suitable for each subspace can be applied separately.
[0003] Orthogonal canonical correlation analysis (OCCA) is a technique for multivariate statistical analysis used to find the correlation between two sets of variables. OCCA extends the traditional canonical correlation analysis (CCA) by maximizing the correlation between two sets of variables while preserving their orthogonality, thereby reducing redundant information and multicollinearity. This method is particularly suitable for fault detection in multi-dimensional data, and can effectively extract features related to fault characteristics, improving the accuracy and robustness of detection results, and has important application value in fault detection of complex industrial processes. Principal component analysis (PCA) is a classical dimensionality reduction method, aiming to extract features with large variances in the data through linear transformation and map the data to a new low-dimensional space. PCA finds the principal components of the data, enabling most of the data changes to be represented by a few principal components, thereby reducing the data dimension, simplifying the data structure, reducing data redundancy, and at the same time retaining the main information of the data. In fault detection, PCA determines whether a fault has occurred in the process by calculating the threshold of the principal component space and based on whether the test set statistic exceeds the threshold.
[0004] The present invention proposes a new fault detection method based on JB - OCCA (Jarque - Bera - Oriented Canonical Correlation Analysis) and PCA (Principal Component Analysis). It divides data into different feature sub - spaces, while reducing redundancy, retains the main information of various types of data, facilitates further fault detection, can effectively process Gaussian - non - Gaussian mixed data, and distinguish whether faults are quality - related. Summary of the Invention
[0005] The present invention proposes a JB - OCCA - PCA method for quality - related fault detection, which can effectively process Gaussian - non - Gaussian mixed data and accurately distinguish quality - related and non - quality - related faults, thereby improving the accuracy and reliability of fault detection. First, the method of the present invention standardizes the original data to ensure that the data is analyzed under the same dimension. Secondly, the data is partitioned into molecular spaces according to Gaussian and non - Gaussian, quality - related and quality - unrelated features. In the Gaussian - related variable subspace and the non - Gaussian - related variable subspace, the Jarque - Bera (JB) statistic is used to replace the original data, and then OCCA is used for correlation analysis to extract quality - related features. For the Gaussian - unrelated variable subspace and the non - Gaussian - unrelated variable subspace, the standard deviation and skewness within the time window are calculated as the feature matrix, and the principal components with the largest amount of information in the feature matrix are extracted based on PCA to replace the original data. Then, a fault detection model is established using PCA. Finally, the statistics with the same physical meaning in each subspace are integrated through the Bayesian method, a threshold is set for fault detection, and an alarm signal is issued when a fault is detected. This method can not only improve the accuracy and robustness of fault detection for high - dimensional complex data, but also effectively reduce the false alarm rate and missed alarm rate, and is a more suitable fault detection solution for complex industrial processes.
[0006] The technical solution adopted by the present invention to solve the above - mentioned technical problems is: a JB - OCCA - PCA method for quality - related fault detection, characterized by including the following steps:
[0007] The implementation process of the offline modeling stage is as follows:
[0008] Step (1): Data acquisition and data pre - processing, and the specific implementation process is as follows:
[0009] Step (1.1): Collect the monitoring data of m sensors at n sampling moments under normal operating conditions to form a process data set X = [x1, x2,..., x m ∈ R n×m , where n is the number of data set samples, m is the number of process variables, and R is the set of real numbers;
[0010] Step (1.2): Collect the monitoring data of l sensors at n sampling moments under normal operating conditions to form a quality data set Y = [y1, y2, …, y l ∈ R n×l , where n is the number of data set samples, m is the number of process variables, and R is the set of real numbers;
[0011] Step (1.3): For the process data set X = [x1, x2, …, x m ∈ R n×m and the quality data set Y = [y1, y2, …, y l ∈ R n×l , use the z - score normalization method for data pre - processing to obtain the normalized process data set X = [x1, x2, …, x m ∈ R n×m and the quality data set Y = [y1, y2, …, y l ∈ R n×l . Taking the process variable x i as an example, the specific calculation formula is as follows:
[0012] i = 1, 2, 3,..., m (1)
[0013] where, μ i represents the mean of x i , and σ i represents the standard deviation.
[0014] Step (2): Perform the Jarque - Bera test on the normalized data set X = [x1, x2, …, x m ∈ R m×m , and divide the process data set into a Gaussian - related subspace X1, a non - Gaussian - related subspace X2, a Gaussian - unrelated subspace X3, and a non - Gaussian - unrelated subspace X4 according to the mutual information (MI) value between the process variable and the quality variable.
[0015] Step (2.1): The basis for dividing the Gaussian and non - Gaussian subspaces is whether the variable follows a Gaussian distribution. If it does, it is a Gaussian variable; otherwise, it is a non - Gaussian variable. Use the Jarque - Bera test to determine whether the variable x i ∈ R n follows a Gaussian distribution. The JB statistic is based on skewness and kurtosis to measure the normality of the data. The smaller the JB statistic, the more the variable x i follows a Gaussian distribution. When skew = 0 and kurt = 3, the variable x i strictly follows a Gaussian distribution. The JB statistic is calculated as follows
[0016]
[0017] Since the JB statistic follows a chi-square distribution with 2 degrees of freedom, the probability p that the statistic falls in the right tail of the current observed value can be obtained by looking up the table. Considering the statistical significance, the confidence level is set to 0.05. When p > 0.05, the variable is classified into the Gaussian subspace; otherwise, it is classified into the non-Gaussian subspace. 2 Since the JB statistic follows a chi-square distribution with 2 degrees of freedom, the probability p that the statistic falls in the right tail of the current observed value can be obtained by looking up the table. Considering the statistical significance, the confidence level is set to 0.05. When p > 0.05, the variable is classified into the Gaussian subspace; otherwise, it is classified into the non-Gaussian subspace.
[0018] Step (2.2): The basis for dividing the quality-related and quality-unrelated subspaces is whether the MI value of the process variable and the quality variable exceeds the threshold η. If MI > η, that is, it exceeds the threshold, it is a quality-related variable; otherwise, it is a quality-unrelated variable. For the variable x in the process variable set X = [x1, x2, …, x m ∈ R n×m within, its joint distribution probability with the quality variable y in the quality variable set Y = [y1, y2, …, y i ∈ R l is denoted as P(x n×l , y j ), then the specific calculation formula for the mutual information value between the process variable x i and the quality variable y j is as follows: i j m n×m
[0019]
[0020] Step (2.3): According to the above steps, the process variable set X = [x1, x2, …, x m ∈ R n×m is divided into four subspaces, namely the Gaussian-related subspace X1, the non-Gaussian-related subspace X2, the Gaussian-unrelated subspace X3, and the non-Gaussian-unrelated subspace X4.
[0021] Step (3): For the Gaussian-related subspace X1 and the quality data set Y, the JB statistic is used to replace the original data in the subspace, and then the quality-related features are extracted through orthogonal canonical correlation analysis (OCCA) for fault detection. The steps are as follows:
[0022] Step (3.1): The sliding window can analyze the dynamic changes of the data in a finer granularity, which helps to detect instantaneous faults or anomalies. Select the window size s, convert the subspace data into a data matrix in window form, and perform z-score standardization. Calculate the JB statistic of the data within the window, replace the original data, and obtain the new data matrices X JB and Y JB .
[0023] Step (3.2): Use the OCCA algorithm to decompose the data matrix X JB into quality-related and -unrelated subspaces: For the data matrices X JB and Y JB , calculate the variable covariance matrix and perform canonical correlation analysis (CCA) to obtain the weight matrices A and B in the direction of the maximum correlation, and obtain the quality-related subspace projection V1 and the quality-unrelated subspace projection V2 through singular value decomposition (SVD). The specific calculations are as follows:
[0024] Define the canonical variables U = A T ·X JB and V = B T ·Y JB . The goal is to maximize the canonical correlation coefficient ρ, where the constraint conditions are A T cov(X JB )A = I and B T cov(Y JB )B = I:
[0025]
[0026] Solve the above optimization problem through generalized eigenvalue decomposition. Let then the canonical vectors and
[0027] Perform CCA modeling on the matrix X JB and the quality variable Y JB to obtain the correlation coefficient matrix H:
[0028] Y JB B = X JB AΛ (5)
[0029] Y JB = X JB AΛB -1 = X JB H (6)
[0030] Perform SVD decomposition on the correlation coefficient matrix H:
[0031] H = U[Φ 0][V1 V2] T (7)
[0032] where V1 corresponds to the quality-related subspace and V2 corresponds to the quality-unrelated subspace.
[0033] Step (3.3): Use the OCCA decomposition to obtain the quality-related and -unrelated subspace matrices V1 and V2, and construct two corresponding detection statistics: the quality-related statistic and the quality-unrelated statistic Used to detect different types of faults. The specific calculation formula is as follows:
[0034]
[0035]
[0036] Step (3.4): To determine whether a sample is in a normal state or a fault state, it is usually necessary to set a threshold according to the distribution of historical normal data statistics. Calculate T based on kernel density estimation (KDE) 2 the distribution of statistics, set the confidence level r = 0.99, and determine the threshold that satisfies the cumulative distribution function CDF. For a set of T 2 statistic data, the kernel density estimation formula and CDF conditions are as follows:
[0037]
[0038]
[0039] where K is the kernel function and h is the bandwidth parameter. Through the KDE method, and the thresholds of are respectively and
[0040] Step (4): For the non-Gaussian correlated subspace X2 and the quality data set Y, according to Step (3), calculate its statistic and threshold
[0041] Step (5): For the Gaussian uncorrelated subspace X3, calculate the standard deviation and skewness within the time window as the feature matrix, extract the principal component with the largest amount of information in the feature matrix based on PCA to replace the original data, and then establish a fault detection model using PCA. It includes the following steps:
[0042] Step (5.1): According to the industrial process requirements, select the window size s, convert the subspace data into a window-form data matrix, and perform z-score standardization. For each variable in the subspace, calculate the standard deviation st(i) and skewness sk(i) within the window using a sliding window at each time point, and use them as the feature f(i) = [st(i), sk(i)] at that moment. Combine the standard deviation and skewness of the variables within each window into the feature matrix F (n-s+1)×2 。 By this strategy, the one-dimensional variable is transformed into a two-dimensional feature matrix. For this two-dimensional matrix F (n-s+1)×2 , use the PCA method to extract the first principal component, that is, its most important feature, to replace the one-dimensional original variable. For each variable, use the above method to obtain the new data matrix XP 。
[0043] Step (5.2): For the new data matrix X P , perform dimensionality reduction through PCA, project the data onto the principal component direction, and remove noise and irrelevant information. For the new data matrix X P Calculate the covariance matrix cov(X P ), and perform eigenvalue decomposition, where Q is the eigenvector matrix and Λ is the diagonal matrix of eigenvalues:
[0044] QΛQ T = cov(X P ) (12)
[0045] Select the first k principal components whose cumulative variance ratio reaches the set threshold of 85% to form the projection matrix P pca , project X P onto the low-dimensional space to obtain the dimensionality-reduced feature matrix X pca . The specific calculation formula is as follows:
[0046] X pca = X P ·P pca (13)
[0047] Step (5.3): Calculate the T 2 statistic of the dimensionality-reduced dataset for fault detection.
[0048]
[0049] Step (5.4): Obtain the threshold in the manner of step (3.4)
[0050] Step (6): For the non-Gaussian independent subspace X4, calculate its statistic and threshold
[0051] Step (7): Through the Bayesian integration method, fuse the quality-related statistics and quality-unrelated statistics of the same type in the four subspaces to obtain the comprehensive quality-related fault detection index BIC(Q) and quality-unrelated fault detection index BIC(NQ) for judging different types of faults. The specific implementation process is as follows:
[0052] Step (7.1): Calculate the probability density of the current sample in the normal and faulty states according to the fault detection statistic.
[0053] Step (7.2): Calculate the posterior probability according to the prior probability and the probability density of each subspace, and calculate the comprehensive quality-related fault detection index BIC(Q) and the quality-unrelated fault detection index BIC(NQ) using the Bayesian formula. The thresholds of both detection indexes are the confidence level L.
[0054] The above steps (1) to (7) are the offline modeling stage of the method of the present invention. The following steps (8) to (17) are the implementation process of the online analysis of the method of the present invention.
[0055] Step (8): Collect the sample data x at the new sampling moment t ;
[0056] Step (9): Process x t using the mean and standard deviation obtained in step (1.3) to obtain the standardized test sample data;
[0057] Step (10): According to the judgment results of Gaussian and non-Gaussian, quality-related and quality-unrelated in the offline modeling, divide x t into four subspaces;
[0058] Step (11): Based on the variable selection criterion and projection in step (3), find the corresponding statistic and
[0059] Step (12): According to the variable selection criterion and projection in step (4), find the corresponding statistic and
[0060] Step (14): According to the variable selection criterion and projection in step (5), find the corresponding statistic
[0061] Step (15): According to the variable selection criterion and projection in step (6), find the corresponding statistic
[0062] Step (16): Calculate the comprehensive quality-related fault detection index BICt(Q) and the quality-unrelated fault detection index BICt(NQ) in the manner of step (7).
[0063] Step (17): Compare the quality-related fault detection index BICt(Q) and the quality-unrelated fault detection index BICt(NQ) with the thresholds (confidence levels) respectively to determine whether quality-related faults and quality-unrelated faults occur in the process.
[0064] Compared with the traditional method, the advantage of the method of the present invention lies in that the data is divided into different feature subspaces, and the data within a certain window is used to replace the original data respectively based on the JB statistic and the maximum principal component strategy to extract the dynamic features of the data. In addition, by establishing the OCCA and PCA models, while reducing redundancy, the main information of various types of data is retained, and the detection accuracy is improved. Description of the Drawings
[0065] Figure 1 It is a detection result diagram of the fault 7 working condition under the implementation case of the present invention.
[0066] Figure 2 It is a fault detection flow chart of the method of the present invention. Detailed Embodiment
[0067] The method of the present invention will be described in detail below in conjunction with the drawings and specific implementation cases.
[0068] The present invention discloses a JB-OCCA-PCA method for quality-related fault detection. The specific implementation process of the method of the present invention will be described below in conjunction with a specific industrial process case.
[0069] The Tennessee Eastman (TE) process is a reliable chemical plant simulation program based on the actual chemical reaction process and is widely used as a benchmark. The TE process includes 41 process variables and 11 operating variables. In addition, the TE process has 21 different types of faults.
[0070] First, a detection model is established using the sample data sampled under the normal operating conditions of the TE process, including the following steps:
[0071] Step (1): Under the normal operating conditions of the industrial process, the number of samples is 960. The first 22 columns of the normal operating condition data set are extracted to form the process variable data set X = [x1, x2,..., x 22 ∈ R 960×22 , and the 35th and 36th columns of the normal operating condition data set are extracted to form the quality variable data set Y = [x 35 , x 36 ∈ R 960×2 . The data sets X = [x1, x2,..., x 22 ∈ R 960×22 and the data set Y = [x 35 , x 36 ∈ R 960×2 are standardized according to the z-score method to obtain the standardized process data set X = [x1, x2,..., x 22 ∈ R 960×22 and the quality data set Y = [x 35 , x36 ∈ R 960×2 ;
[0072] Step (2): According to Gaussian and non-Gaussian, quality-related and quality-unrelated, divide the standardized process data set X = [x1, x2, …, x 22 ∈ R 960×22 into four subspaces: Gaussian-related subspace X1, non-Gaussian-related subspace X2, Gaussian-unrelated subspace X3, and non-Gaussian-unrelated subspace X4.
[0073] Step (3): For the Gaussian-related subspace X1, the characteristic column is c1 = [2, 3, 7, 10, 11, 13, 16, 18, 19, 20, 21], and the corresponding data set X1 is obtained. Use the JB statistic to replace the current data, and then perform data analysis using orthogonal canonical correlation analysis (OCCA) to calculate the T 2 statistic, including the following steps:
[0074] Step (3.1): According to the industrial process requirements, select the window size s = 20, calculate the JB statistic and assign it to the data set to obtain the new data matrices X JB and Y JB ;
[0075] Step (3.2): Use the OCCA algorithm to decompose the data matrix X JB into quality-related and quality-unrelated subspaces: For the data matrices X JB and Y JB , calculate the variable covariance matrix and perform canonical correlation analysis (CCA) to obtain the weight matrices A and B in the direction of the maximum correlation, and obtain the quality-related subspace projection V1 and quality-unrelated subspace projection V2 through singular value decomposition (SVD);
[0076] Step (3.3): Use the quality-related and quality-unrelated subspace matrices V1 and V2 obtained by OCCA decomposition to construct two corresponding detection statistics: the quality-related statistic and the quality-unrelated statistic to detect different types of faults;
[0077] Step (3.4): Through the KDE method, estimate the thresholds of and to be and
[0078] Step (4): For the non-Gaussian correlated subspace X2, the characteristic column is c2 = [1, 4, 8, 15, 22], and the corresponding characteristic datasets X2 and XT2 are obtained. Replace the current data in the way of step (3) with the JB statistic, and then perform data analysis using orthogonal canonical correlation analysis (OCCA) to calculate its statistic and the threshold
[0079] Step (5): For the Gaussian uncorrelated subspace X3, the characteristic column is c3 = [9, 12]. Calculate the standard deviation and skewness within the time window as the feature matrix, and extract the principal component with the maximum information content in the feature matrix based on PCA to replace the original data. Then, establish a fault detection model using PCA. The specific implementation process is as follows:
[0080] Step (5.1): According to the industrial process requirements, select the window size s = 20, convert the subspace data into a window-form data matrix, and perform z-score standardization. For each variable in the subspace, calculate the standard deviation st(i) and skewness sk(i) within the window using a sliding window at each time point, and use them as the feature f(i) = [st(i), sk(i)] at that moment. Combine the standard deviation and skewness of the variables within each window into the feature matrix F (n-s+1)×2 . Through this strategy, the one-dimensional variables are converted into a two-dimensional feature matrix. For this two-dimensional matrix F (n-s+1)×2 , use the PCA method to extract the first principal component, that is, its most important feature, to replace the one-dimensional original variable. For each variable, use the above method to obtain the new data matrix X P ;
[0081] Step (5.2): For the new data matrix X P , through PCA dimensionality reduction, project the data onto the principal component direction to remove noise and irrelevant information;
[0082] Step (5.3): Calculate the statistic of the dimensionality-reduced dataset, and obtain the threshold in the way of step (3.4)
[0083] Step (6): For the non-Gaussian uncorrelated subspace X4, the characteristic column is c4 = [5, 6, 14, 17]. According to step (5), calculate its statistic and the threshold
[0084] Step (7): Through the Bayesian integration method, fuse the quality-related statistics and quality-unrelated statistics of the same type in the four subspaces to obtain the comprehensive quality-related fault detection index BIC(Q) and quality-unrelated fault detection index BIC(NQ) for judging different types of faults.
[0085] The offline modeling phase is now complete, and the next step is to analyze the online data and conduct testing using sample data.
[0086] Step (8): Collect sample data x under TE process fault 7 t ;
[0087] Step (9): For x t Using the mean and standard deviation obtained in step (1) to process, thus obtaining standardized test sample data;
[0088] Step (10): According to the judgment results of Gaussian and non-Gaussian, quality-related and quality-independent in offline modeling, x t Divided into four subspaces;
[0089] Step (11): Calculate the corresponding statistics based on the variable selection criteria and projection in step (3) and
[0090] Step (12): Calculate the corresponding statistics according to the variable selection criteria and projection in step (4) and
[0091] Step (14): Calculate the corresponding statistics according to the variable selection criteria and projection in step (5)
[0092] Step (15): Calculate the corresponding statistics according to the variable selection criteria and projection in step (6)
[0093] Step (16): Calculate the comprehensive quality-related fault detection index BICt(Q) and the quality-independent fault detection index BICt(NQ) in the same manner as in step (7).
[0094] Step (17): The quality-related fault detection index BICt(Q) and the quality-independent fault detection index BICt(NQ) are compared with the threshold (confidence level) respectively to determine whether quality-related faults and quality-independent faults occur in the process.
[0095] It can be found that the method of the present invention realizes targeted modeling by dividing data into different feature subspaces, ensures the accuracy of model detection, and at the same time reduces the computational complexity of high-dimensional data, making the algorithm more suitable for the requirements of online real-time fault detection. The present invention uses JB-OCCA to extract key features of quality-related variables, uses PCA to reduce the dimension of quality-unrelated variables, and combines Bayes to integrate the detection results with specific significance in multiple subspaces, which not only improves the sensitivity of detection, but also enhances the resistance to noise, and reduces the false positive rate and false negative rate. Figure 1 This is the detection result overrun situation diagram of the fault 7 working condition in the implementation case of the present invention, which further demonstrates the feasibility and superiority of the method of the present invention.
[0096] The above implementation cases are only used to explain the specific implementation of the present invention, rather than to limit the present invention. Any modification made to the present invention within the scope of the claims of the present invention falls within the protection scope of the present invention.
Claims
1. A JB-OCCA-PCA method for quality-related fault detection, characterized in that: The following steps are involved: Step (1): Data collection and data preprocessing. The specific implementation process is as follows: Step (1.1): Collect monitoring data of m sensors at n sampling moments under normal operating conditions to form a process data set X = [x1, x2, ..., x m ]∈R n×m , where n is the number of data set samples, m is the number of process variables, and R is a real number set; Step (1.2): Collect monitoring data of l sensors at n sampling moments under normal operating conditions to form a quality data set Y = [y1, y2, ..., y l ]∈R n×l , where n is the number of data set samples, m is the number of process variables, and R is a real number set; Step (1.3): Substitute the process data set X = [x1, x2, …, x m ]∈R n×m And the quality data set Y = [y1,y2,…,y l ]∈R n×l The z-score standardization method is used for data preprocessing to obtain the standardized process data set X = [x1, x2, ..., x m ]∈R n×m And the quality data set Y = [y1,y2,…,y l ]∈R n×l . With process variable x i For example, the specific calculation formula is as follows: Among them, μ i Represents x i The mean value, σ i Represents standard deviation. Step (2): Standardize the dataset X = [x1, x2, …, x m ]∈R n×m The Jarque-Bera test was performed, and according to the mutual information (MI) values between process variables and quality variables, the process data set was divided into Gaussian correlation subspace X1, non-Gaussian correlation subspace X2, Gaussian independent subspace X3, and non-Gaussian independent subspace X4. Step (2.1): The division of Gaussian and non-Gaussian subspaces is based on whether the variable obeys Gaussian distribution. If it obeys, it is a Gaussian variable, otherwise it is a non-Gaussian variable. The Jarque-Bera test is used to determine whether the variable x i ∈R n Whether it obeys Gaussian distribution. JB statistic is based on skewness and Kurtosis to measure the normality of the data. The smaller the JB statistic, the smaller the variable x i The more it obeys the Gaussian distribution. When skew = 0 and kurt = 3, the variable x i Strictly obey the Gaussian distribution. The JB statistic is calculated as follows Since the JB statistic follows the chi-squared relationship with 2 degrees of freedom 2 Distribution, you can look up the table to get the probability p that the statistic falls in the right tail of the current observation. Considering the statistical significance, the confidence level is set to 0.
05. When p>0.05, the variable is classified into the Gaussian subspace; otherwise, it is classified into the non-Gaussian subspace. Step (2.2): The basis for dividing the quality-related and quality-independent subspaces is whether the process variable and the quality variable MI value exceed the threshold η. If MI>η, that is, it exceeds the threshold, it is a quality-related variable, otherwise it is a quality-independent variable. For the process variable set X = [x1, x2, …, x m ]∈R n×m The variable x in i , which is related to the quality variable set Y = [y1,y2,…,y l ]∈R n×l Medium quality variable y j The joint distribution probability of i ,y j ), then the process variable x i and the quality variable y j The specific calculation formula of the mutual information value is as follows: Step (2.3): According to the above steps, the process variable set X = [x1, x2, ..., x m ]∈R n×m It is divided into four subspaces, namely, Gaussian related subspace X1, non-Gaussian related subspace X2, Gaussian independent subspace X3, and non-Gaussian independent subspace X4. Step (3): For the Gaussian correlation subspace X1 and the quality data set Y, the JB statistic is used to replace the original data in the subspace, and then the quality-related features are extracted through orthogonal canonical correlation analysis (OCCA) for fault detection. Step (4): For the non-Gaussian correlation subspace X2 and the quality data set Y, calculate its statistics according to step (3): and threshold Step (5): For the Gaussian independent subspace X3, the standard deviation and skewness in the time window are calculated as the feature matrix, and the principal component with the largest amount of information in the feature matrix is extracted based on PCA to replace the original data, and then the fault detection model is established using PCA. Step (6): For the Gaussian independent subspace X3, the standard deviation and skewness in the time window are calculated as the feature matrix, and the principal component with the largest amount of information in the feature matrix is extracted based on PCA to replace the original data, and then the fault detection model is established using PCA. Step (7): By using the Bayesian integration method, the same type of quality-related statistics and quality-independent statistics in the four subspaces are integrated to obtain the comprehensive quality-related fault detection index BIC(Q) and the quality-independent fault detection index BIC(NQ) for judging different types of faults. The specific implementation process is as follows: Step (7.1): Based on the statistics of fault detection, calculate the probability density of the current sample in normal and fault states. Step (7.2): Based on the prior probability and the probability density of each subspace, the posterior probability is calculated using the Bayesian formula, and the comprehensive quality-related fault detection index BIC(Q) and the quality-independent fault detection index BIC(NQ) are calculated. The thresholds of the two detection indicators are both the confidence level L.
2. The JB-OCCA-PCA method for quality-related fault detection according to claim 1, characterized in that: The step (3) uses the sliding window JB statistic to replace the subspace original data, and then extracts quality-related features through orthogonal canonical correlation analysis (OCCA). The specific implementation process is as follows: Step (3.1): The sliding window can analyze the dynamic changes of data in a more fine-grained manner, which helps to detect instantaneous faults or anomalies. Select the window size s, convert the subspace data into a windowed data matrix, and perform z-score standardization. Calculate the JB statistic of the data in the window, replace the original data, and get the new data matrix X JB and Y JB . Step (3.2): Use the OCCA algorithm to transform the data matrix X JB Decomposition into quality-dependent and quality-independent subspaces: For the data matrix X JB and Y JB , calculate the variable covariance matrix, and perform canonical correlation analysis (CCA) to obtain the weight matrices A and B in the direction of maximum correlation, and obtain the quality-related subspace projection V1 and the quality-independent subspace projection V2 through singular value decomposition (SVD). The specific calculation is as follows: Define the canonical variable I = A T ·X JB and V = B T ·Y JB , the goal is to maximize the canonical correlation coefficient ρ, where the constraints are A T cov(X JB )A=I and B T cov(Y JB )B=I: The above optimization problem is solved by generalized eigenvalue decomposition. Let The typical vector and For the matrix X JB and quality variable Y JB Perform CCA modeling and obtain the correlation coefficient matrix H: Y JB B=X JB AL (5) Y JB =X JB AΛB -1 =X JB H (6) Perform SVD decomposition on the correlation coefficient matrix H: H=U[Φ 0][V1 V2] T (7) Among them, V1 corresponds to the mass-dependent subspace, and V2 corresponds to the mass-independent subspace. Step (3.3): Use OCCA decomposition to obtain the quality-related and quality-independent subspace matrices V1 and V2, and construct the corresponding two detection statistics: the quality-related statistic Statistics not related to quality Used to detect different types of faults. The specific calculation formula is as follows: Step (3.4): In order to determine whether a sample is in a normal state or a faulty state, it is usually necessary to set a threshold based on the distribution of historical normal data statistics. Calculate T based on kernel density estimation (KDE) 2 The distribution of the statistic, set the reliability r = 0.99, and determine the threshold that satisfies the cumulative distribution function CDF. 2 Statistics, Kernel Density Estimation The formula and CDF conditions are as follows: Among them, K is the kernel function and h is the bandwidth parameter. Through the KDE method, and The thresholds are as well as 3. The JB-OCCA-PCA method for quality-related fault detection according to claim 1, characterized in that: In the step (5), the standard deviation and skewness in the time window are calculated as the feature matrix, and the principal element with the largest amount of information in the feature matrix is extracted based on PCA to replace the original data, and then the fault detection model is established by using PCA. The specific implementation process is as follows: Step (5.1): According to the industrial process requirements, select the window size s, convert the subspace data into a windowed data matrix, and perform z-score standardization. For each variable in the subspace, use a sliding window to calculate the standard deviation st(i) and skewness sk(i) in the window at each time point, and use it as the feature f(i) = [st(i), sk(i)] at that moment. The standard deviation and skewness of the variables in each window are combined into the feature matrix F (n-s+1)×2 Through this strategy, the one-dimensional variable is transformed into a two-dimensional feature matrix. For this two-dimensional matrix F (n-s+1)×2 , the PCA method is used to extract the first principal component, that is, its most important feature, to replace the original single-dimensional variable. For each variable, the above method is used to obtain a new data matrix X P . Step (5.2): For the new data matrix X P , through PCA dimensionality reduction, the data is projected onto the principal component direction to remove noise and irrelevant information. P Calculate the covariance matrix cov(X P ), and perform eigendecomposition, where Q is the eigenvector matrix and Λ is the eigenvalue diagonal matrix: QΛQ T =cov(X P ) (12) Select the first k principal components whose cumulative variance ratio reaches the set threshold of 85% to form the projection matrix P pca , X P Project to low-dimensional space to get the reduced-dimensional feature matrix X pca The specific calculation formula is as follows: X pca =X P ·P pca (13) Step (5.3): Calculate T of the reduced-dimensional dataset 2 Statistics used for fault detection. Step (5.4): Calculate the threshold value as in step (3.4)
Citation Information
Cited By
Fault detection method and system for industrial equipment
CN121881278A
A method and system for failure detection of industrial equipment
CN121881278B