Non-linear and non-Gaussian system fault detection method based on quality correlation virtual variable
By constructing the quality-associated virtual variable Qav and using mutual information MI to divide process variables, a KICA model was established, and a strong dependence on quality variables in nonlinear systems was solved, and efficient fault detection was achieved in complex industrial processes.
Patent Information
- Application Number
- CN202510522980.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art has a strong dependence on mass variables in nonlinear system fault detection, and when the mass variables are unmeasurable or the measured values are unreliable, fault detection is difficult, especially in complex industrial processes, the detection effect of the existing methods is not good.
By constructing the quality-associated virtual variable Qav, using mutual information MI to calculate the correlation between process variables and Qav, the process variables are divided into quality-associated groups and irrelevant groups, and a KICA model is established to calculate the relevant and irrelevant statistics and control limits, and perform fault detection.
Reliance on quality variables is reduced, fault detection rate is improved, false alarm rate is reduced, and accurate fault identification can be achieved especially when the quality variable is unpredictable.
Smart Images

Figure CN120406393A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fault detection, and in particular, to a fault detection method for nonlinear non-Gaussian systems based on quality-associated virtual variables. Background Art
[0002] Quality-related fault monitoring, as a key means to ensure the safe or effective operation of large-scale equipment and industrial processes, has formed a relatively mature theoretical system in the scenario where quality variables are measurable. Currently, quality-related fault detection methods can be divided into two mainstream methods: projection space decomposition method and process variable partitioning method. The projection space decomposition method first establishes a model for all process variables to form a projection space; subsequently, the projection space is orthogonally decomposed to form a quality-related subspace and a quality-unrelated subspace for quality-related fault monitoring. Typical methods include Total Partial Least Squares (TPLS), Improved Partial Least Squares (IPLS), and Total Principal Component Regression (T-PCR), etc. The core lies in constructing a unified projection space containing all process variables, decoupling the space into quality-related and quality-unrelated subspaces through an orthogonal decomposition strategy, and then achieving accurate identification of quality-related faults. With the development of deep learning technology, scholars have attempted to construct a neural network-based nonlinear feature extraction architecture, but its essence still follows the correlation modeling paradigm between process variables and quality variables. The process variable partitioning method divides process variables into quality-related and quality-unrelated groups according to the correlation between process variables and quality variables (such as mutual information, minimum redundancy maximum correlation, transfer entropy, variable importance, etc.), and models these two groups of variables separately to achieve quality-related fault detection.
[0003] However, the actual industrial production scenario is often more complex. During the operation of actual industrial processes or large-scale equipment, due to the harsh working environment and measurement technology limitations, some quality variables are difficult or even impossible to obtain through measuring instruments, which makes the monitoring methods relying on the premise of measurable quality variables lose their application basis. In addition, the current theoretical research on the fault detectability of nonlinear systems is not yet perfect, which stems from the mathematical complexity brought by kernel methods. When using kernel methods to extend fault monitoring from the linear space to the kernel space, since fault detectability analysis needs to establish the numerical relationship between variables and statistics in the original data space, and the pre-image problem of kernel functions is difficult to accurately reconstruct the original data, it seriously increases the difficulty of fault detectability analysis of nonlinear systems. Summary of the Invention
[0004] In view of the above deficiencies in the prior art, the fault detection method for nonlinear non-Gaussian systems based on quality-associated virtual variables provided in this application realizes the rapid identification of quality-related variables and reduces the dependence on quality variables in the process of quality-related fault monitoring.
[0005] In order to achieve the above invention purpose, the technical solution adopted in this application is as follows:
[0006] The fault detection method for nonlinear non-Gaussian systems based on quality-associated virtual variables includes:
[0007] S2: Based on the Qav, calculate the correlation between the process variables and the Qav using mutual information MI, and divide the process variables into a quality-related variable group and a quality-unrelated variable group;
[0008] S3: According to the quality-related variable group and the quality-unrelated variable group, establish KICA models, denoted as KICAr and KICAu respectively;
[0009] S4: Based on the KICAr model and the KICAu model, calculate the quality-related statistics and quality-unrelated statistics of the training set samples, and obtain the quality-related control limit and the quality-unrelated control limit according to the kernel density estimation method;
[0010] S5: Based on the division result of the process variables, divide the monitoring samples into a quality-related part and a quality-unrelated part, calculate the quality-related joint statistic and the quality-unrelated joint statistic of the monitoring samples, and determine the quality-related joint control limit and the quality-unrelated joint control limit;
[0011] S6: Based on the quality-related joint statistic, the quality-unrelated joint statistic, the quality-related joint control limit, and the quality-unrelated joint control limit, perform fault detection to obtain the fault detection result.
[0012] Further, the S1 specifically includes:
[0013] S101: Construct a process variable matrix and perform standardization processing on the process variable matrix;
[0014] where N and m represent the number of samples and the number of variables respectively, i = [1, 2,..., N], x i is a column vector representing the sampling sample at the i-th moment, j = [1, 2,..., m], x j is a row vector representing the time series data of the j-th variable;
[0015] S102: Pass the standardized process variables through the nonlinear mapping φ * (.), project the x j onto the linear space x j →φ* (x j );
[0016] S103: Construct the objective function to maximize the sum of the squares of the inner products of the quality - related dummy variable α and all variables:
[0017]
[0018] Among them, is the quality - related dummy variable, α is the transpose of the quality - related dummy variable, N and m represent the number of samples and the number of variables respectively, x j is the row vector representing the time - series data of the j - th variable, αα T is the outer - product operation of the quality - related dummy variable, N - 1 is the sample standard quantity, φ * (x j ) is the projection of the variable x j in the linear space;
[0019] S104: According to the Lagrange multiplier method, transform the objective function into:
[0020]
[0021] Take the partial derivative with respect to α:
[0022]
[0023] Among them, is the partial derivative of L(α) with respect to α;
[0024] S105: Perform eigenvalue decomposition on
[0025]
[0026] Among them, Φ * is the matrix that the process variable is mapped to the high - dimensional linear space through the non - linear transformation φ * (.), Φ *T is the transpose matrix corresponding to Φ * , λ is the eigenvalue;
[0027] S106: Select the eigenvector α corresponding to the largest eigenvalue as Qav.
[0028] Furthermore, the specific steps of S2 include:
[0029] S201: Calculate the mutual - information MI value between the process variable and Qav;
[0030] S202: Divide the process variable according to the MI value and the preset threshold to obtain a quality - related variable group and a quality - unrelated variable group.
[0031] Further, calculating the mutual information MI value between the process variable and the Qav in S201 includes:
[0032] Using a sliding window with a window width of L to intercept the time series data into m - L + 1 sub - blocks. The i - th sub - block can be expressed as:
[0033]
[0034] where, is the i - th sub - block, represents the matrix of quality - related dummy variable matrix, is the quality - related dummy variable matrix, α i is the quality - related dummy variable, x i represents the i - th scalar, and c is the kernel parameter;
[0035] Taking the MI between the j - th variable in and as denoted as
[0036] Calculating the MI values of each sub - block. The MI vector between the j - th variable and Qav is expressed as:
[0037]
[0038] where, I j is the MI vector between the j - th variable and Qav, is the MI between the j - th variable in and denoted as is the j - th variable in;
[0039] Calculating the mean value j of the vector I and taking the mean value as the mutual information MI value between the j - th variable and the Qav.
[0040] Further, partitioning the process variable according to the MI value and a preset threshold to obtain a quality - related variable group and a quality - unrelated variable group includes:
[0041] If the mean value of the MI value is greater than the preset threshold, the variable corresponding to the MI value is a quality - related variable in the quality - related variable group; otherwise, the variable corresponding to the MI value is a quality - unrelated variable in the quality - unrelated variable group.
[0042] Further, the preset threshold includes:
[0043] Generate a set of random samples that follow a normal distribution using the Monte Carlo method;
[0044] Use a sliding window with a fixed window width to block the random samples and the Qav, and calculate the MI values between each variable of the random samples and the Qav;
[0045] Arrange the MI values in ascending order, and select the quantile with a confidence level of 95% as the preset threshold.
[0046] Further, in step S5, based on the partitioning result of the process variables, the monitoring samples are divided into a quality-related part and a quality-unrelated part, and the quality-related joint statistic and the quality-unrelated joint statistic are calculated, specifically including:
[0047] Project the quality-related part and the quality-unrelated part onto the KICAr model and the KICAu model respectively, and calculate the quality-related joint statistic and the quality-unrelated joint statistic according to the calculation formula. The calculation formula:
[0048]
[0049] where x new is the monitoring sample, x new,r is the quality-related part, x new,u is the quality-unrelated part, ψ r (x new ) and ψ u (x new ) are the statistics of the quality-related part and the quality-unrelated part, δ Ir and δ Qr are the quality-related control limits, δ Iu and δ Qu are the quality-unrelated control limits, I 2 (x new,r ) and Q(x new,r ) are the statistics of the quality-related part, I 2 (x new,u ) and Q(x new,u ) are the statistics of the quality-unrelated part.
[0050] Further, step S6 specifically includes:
[0051] According to the quality-related joint statistic and the quality-related joint control limit, determine whether the quality-related joint statistic of the quality-related part is greater than or equal to the quality-related joint control limit, and obtain a first judgment result;
[0052] If the first judgment result is negative, determine whether the quality-independent combined statistic of the quality-independent part is greater than or equal to the quality-independent combined control limit according to the quality-independent combined statistic of the quality-independent part and the quality-independent combined control limit, to obtain a second judgment result; if the first judgment result is positive, a quality-related fault occurs in the system.
[0053] If the second judgment result is positive, a quality-independent fault occurs in the system; if the second judgment result is negative, it is determined that the system is operating normally.
[0054] The beneficial effects of this application are as follows:
[0055] This application constructs a quality-associated virtual variable Qav to characterize the association between the quality variable and the process variable, and through mutual information analysis, divides the process variables into a quality-related variable group and a quality-independent variable group, reducing the direct dependence on the quality variable. Especially in complex situations where the quality variable cannot be directly obtained, this method can accurately identify the quality variable by constructing Qav, effectively reducing the false alarm rate in quality-related fault monitoring and improving the fault detection rate. Description of the Drawings
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.
[0057] Figure 1 It is a schematic flowchart of a method for fault detection of a nonlinear non-Gaussian system based on a quality-associated virtual variable provided by an embodiment of the present application;
[0058] Figure 2 It is a schematic mutual information diagram between each variable and the quality variable and the quality-associated virtual variable provided by an embodiment of the present application;
[0059] Figure 3 It is a schematic diagram of a division result of process variables provided by an embodiment of the present application;
[0060] Figure 4 It is a schematic diagram of the change trend of the quality variable XMEAS(35) over time in IDV(12) provided by an embodiment of the present application;
[0061] Figure 5 It is a schematic diagram of the fault monitoring results of IDV(12) by different methods provided by an embodiment of the present application;
[0062] Figure 6Schematic diagram of the change trend of the quality variable XMEAS(35) over time in the IDV(11) provided by the embodiment of the present application;
[0063] Figure 7 Schematic diagram of the fault monitoring results of the IDV(11) by different methods provided by the embodiment of the present application. Detailed implementation manners
[0064] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0065] To more clearly illustrate the fault detection method for a nonlinear non-Gaussian system based on quality-associated virtual variables provided by the embodiments of the present application, an exemplary description will be given of a possible application scenario of the fault detection method for a nonlinear non-Gaussian system based on quality-associated virtual variables provided by the embodiments of the present application. It can be understood that the following examples are only a possible application scenario of the fault detection method for a nonlinear non-Gaussian system based on quality-associated virtual variables provided by the embodiments of the present application. In other possible embodiments, the fault detection method for a nonlinear non-Gaussian system based on quality-associated virtual variables provided by the embodiments of the present application can also be applied to other possible application scenarios, and the following examples do not impose any restrictions on this.
[0066] In actual large-scale equipment or industrial processes, due to the action of the internal transfer and feedback control mechanisms in the system, intricate interactions and dependencies are formed among the quality-related variables. These variables often exhibit collaborative change trends during the operation of the system, reflecting their inherent connections and mutual influences, while the quality-unrelated variables lack such co-variation trends.
[0067] Based on the fact that the quality variables are unmeasurable or the measured values are unreliable, to reduce the dependence on the quality variables, a virtual variable can be constructed to reflect the quality characteristics of the system, that is, the quality-associated virtual variable Qav.
[0068] It can be understood that, essentially, the quality-associated virtual variable Qav is a mathematically constructed virtual index, which is a quantitative representation of the internal connections and comprehensive characteristics of multiple quality-related variables. This variable has a strong correlation with all quality-related variables and a weak correlation with quality-unrelated variables. This characteristic enables us to use the Pearson correlation coefficient or MI to calculate the correlation strength between each process variable and Qav, accurately identify the quality-related variables, which will greatly reduce the dependence on the quality variables and provide a basis for subsequent quality-related fault monitoring.
[0069] Based on this, the embodiments of the present application provide a fault detection method for a non - linear non - Gaussian system based on quality - associated virtual variables. This method can be seen in Figure 1 , Figure 1 As shown in the following is a schematic flowchart of a fault detection method for a non - linear non - Gaussian system based on quality - associated virtual variables provided by the embodiments of the present application, including:
[0070] S1: Construct the quality - associated virtual variable Qav.
[0071] Further, the specific steps of S1 include:
[0072] S101: Construct the process variable matrix and standardize the process variable matrix;
[0073] where N and m represent the number of samples and the number of variables respectively, i = [1, 2,..., N], x i is a column vector representing the sampling sample at the i - th moment, j = [1, 2,..., m], x j is a row vector representing the time - series data of the j - th variable;
[0074] S102: Project the standardized process variables through the non - linear mapping φ * (.), project the variable x j onto the linear space x j →φ * (x j );
[0075] S103: Construct the objective function to maximize the sum of the squares of the inner products of the quality - associated virtual variable α and all variables:
[0076]
[0077] where, is the quality - associated virtual variable, α T is the transpose of the quality - associated virtual variable, αα T is the outer - product operation of the quality - associated virtual variable, N - 1 is the sample standard quantity, φ * (x j ) is the projection of the variable x j in the linear space;
[0078] S104: According to the Lagrange multiplier method, transform the objective function into:
[0079]
[0080] Take the partial derivative of α:
[0081]
[0082] Among them, is the partial derivative of L(α) with respect to α;
[0083] S105: Let Perform eigenvalue decomposition:
[0084]
[0085] Among them, Φ * is the matrix that maps the process variable to a high-dimensional linear space through the non-linear transformation φ * (.), Φ *T is the transpose matrix corresponding to Φ * , and λ is the eigenvalue;
[0086] S106: Select the eigenvector α corresponding to the largest eigenvalue as Qav.
[0087] It can be understood that the process variable matrix: is the process variable matrix of a linear system under normal operating conditions. Normalizing the process variable matrix means that the matrix has zero mean and unit variance.
[0088] S2: Based on the Qav, use the mutual information MI to calculate the correlation between the process variable and the Qav, and divide the process variable into a quality-related variable group and a quality-unrelated variable group.
[0089] Furthermore, the S2 specifically includes:
[0090] S201: Calculate the mutual information MI value between the process variable and the Qav;
[0091] S202: Divide the process variable according to the MI value and a preset threshold to obtain a quality-related variable group and a quality-unrelated variable group.
[0092] It can be understood that the mutual information (MI) is a useful measurement method in information theory. It can quantitatively evaluate the index of the mutual dependence degree between two random variables. The larger the MI value, the stronger the correlation between variables; conversely, the smaller the MI value, the weaker the correlation.
[0093] For the variable pair x i and x j , MI can be expressed as the relative entropy of the inner product of the joint probability distribution p(x i , x j ) and the marginal probability distributions p(x i )p(x j ):
[0094]
[0095] Among them, x i is a column vector representing the sampling sample at time i, and x j is a row vector representing the time series data of the j-th variable. I(x i , x j ) is the mutual information between the variables x i and x j . p(x i , x j ) is the joint probability distribution between the variables x i and x j . p(x i ) is the probability distribution of the variable x i . p(x j ) is the probability distribution of the variable x j . p(x i )p(x j ) is the marginal probability distribution between the variables x i and x j .
[0096] Introduce the information entropy and transform the I(x i , x j ) as follows:
[0097] I(x i , x j ) = H(x i ) + H(x i ) - H(x i , x j )
[0098]
[0099] Among them, H(x i , x j ) is the relative entropy of the variable pair x i and x j . H(x i ) is the relative entropy of x i . The larger the MI value I(x i , x j ), the stronger the correlation between the variables. On the contrary, the smaller the I(x i , x j ), the weaker the correlation.
[0100] Furthermore, calculating the MI value of the mutual information between the process variable and the Qav includes:
[0101] The time series data is intercepted into m - L + 1 sub - blocks using a sliding window with a window width of L. The i - th sub - block can be expressed as:
[0102]
[0103] Where, is the i - th sub - block, represents a matrix of quality - related dummy variable matrices, is the quality - related dummy variable matrix, α i is the quality - related dummy variable, x i represents the i - th scalar, and c is the kernel parameter;
[0104] Let the MI between the j - th variable and in be denoted as
[0105]
[0106] where, I j is the MI vector between the j - th variable and Qav, is the MI between the j - th variable and in is the j - th variable in
[0107] Calculate the mean value j of the vector I If is greater than the preset threshold Th, then the j - th variable x j is determined as a quality - related variable; otherwise, it is a quality - unrelated variable.
[0108] Where, the specific steps for determining the threshold Th are as follows:
[0109] (1) Data generation: Generate a set of random samples obeying the normal distribution through the Monte Carlo method to simulate the statistical characteristics of quality - unrelated variables;
[0110] (2) Mutual information calculation: Use a sliding window with a window width of L to block the random samples and Qav, and calculate the mutual information MI between each variable and Qav.
[0111] (3) Threshold determination: Arrange the above - mentioned MI values in ascending order, and select the quantile with a confidence level of 95% (i.e., 95% of the MI values are lower than this threshold) as the judgment threshold Th.
[0112] S3: Establish a KICA model based on the quality-related variable group and the quality-unrelated variable group, denoted as KICAr and KICAu respectively.
[0113] It can be understood that the process of establishing the KICAr model and the KICAu model for the quality-related group and the quality-unrelated group is similar to the process of establishing the KICA model for the non-linear data X. Taking the construction of the KICA model for X as an example for illustration.
[0114] To construct the KICA model for the non-linear data X, X is mapped to a high-dimensional linear space F through a non-linear transformation φ(.):
[0115]
[0116] The covariance matrix of the matrix Φ is:
[0117]
[0118] Among them, C F is the covariance matrix, which can also be denoted as C F = ΦΦ T / N, Φ is the matrix that maps the non-linear data X to the high-dimensional linear space, φ(x i ) is the variable that maps the variable x i to the high-dimensional linear space, and N is the number of variables.
[0119] Since φ(.) is difficult to be explicitly represented, the kernel trick is introduced to calculate the kernel matrix K. Using the radial basis function, the element in the i-th row and j-th column can be expressed as:
[0120] [K] i,j = k(x i , x j ) = exp(-||x i - x j || 2 / σ)
[0121] Among them, σ is the kernel parameter, x i represents the i-th row vector, x j is the j-th column vector, k(x i , x j ) is the element in the i-th row and j-th column of the kernel matrix K, K is the kernel matrix and K = Φ T Φ.
[0122] Perform the following centering on the kernel matrix K:
[0123] K = K - 1 N K - K1 N + 1 N K1 N
[0124]
[0125] Among them, 1 N is an N-dimensional square matrix with all elements being 1 / N.
[0126] To simplify the independent component extraction problem, whiten the data matrix X. Specifically, perform eigenvalue decomposition on the centered matrix K:
[0127] Ku i = λ i u i
[0128] Among them, λ i and u i represent the eigenvalue and eigenvector respectively. Usually, the eigenvectors corresponding to the first p largest eigenvalues are selected to whiten the data, and the value of p is the number of all eigenvalues that satisfy the condition . Left-multiply the above formula by Φ to get:
[0129] ΦKu i = Φλ i u i
[0130] Substitute K = Φ T Φ into the above formula to get:
[0131]
[0132] It can be seen from this that the eigenvector of the covariance matrix C F is v i = Φu i , and the corresponding eigenvalue is λ i / n. To make the whitened data have zero mean and unit variance, normalize v i . Then the eigenvector matrix of matrix C F can be expressed as:
[0133] V = ΦUΛ -1 / 2
[0134] Among them, U = [u1, u2,... u p , is the eigenvector matrix, Λ = diag(λ1, λ2,... λ p ), is the diagonal matrix corresponding to all eigenvalues, λ(i = 1, 2, LN) represents the i-th eigenvalue, V represents the corresponding eigenvector matrix. Therefore, the whitening matrix and the whitening score matrix can be expressed as:
[0135]
[0136] Among them, Φ(x i) is the sample x i The high-dimensional vector after non-linear mapping, P is the whitening matrix, and Z is the whitening score matrix.
[0137] After whitening, it is necessary to extract the feature W so that the linear combination of the whitened data Z has the maximum non-Gaussianity, and the mixing matrix A and the demixing matrix Q can be further obtained:
[0138]
[0139] For the given sample x r and the corresponding kernel vector k, its whitening score vector and kernel independent components (KICs) can be expressed as:
[0140]
[0141] s = W T z = Qk
[0142] S4: Based on the KICAr model and the KICAu model, calculate the quality-related statistics and quality-unrelated statistics of the training set samples, and obtain the quality-related control limit and the quality-unrelated control limit according to the kernel density estimation method.
[0143] In a possible embodiment, the quality-related x r statistics can be calculated as:
[0144]
[0145] where s is the kernel independent component, s T is the transpose of the kernel independent component, k is the kernel vector, and W and W T are the demixing matrix and the transpose matrix of the demixing matrix.
[0146] Obtain all quality-related and Q r statistics, and determine its corresponding control limits δ Ir and δ Qr .
[0147] Similarly, obtain the quality-unrelated statistics and Q u and determine the control limits δ Iu and δ Qu .
[0148] S5: Based on the division result of the process variables, divide the monitoring sample into a quality-related part and a quality-unrelated part, calculate the quality-related joint statistic and the quality-unrelated joint statistic of the monitoring sample, and determine the quality-related joint control limit and the quality-unrelated joint control limit.
[0149] Further, project the quality-related part and the quality-unrelated part onto the KICAr model and the KICAu model respectively, and calculate the quality-related joint statistic and the quality-unrelated joint statistic according to the calculation formula. The calculation formula is as follows:
[0150]
[0151] where x new is the monitoring sample, x new,r is the quality-related part, x new,u is the quality-unrelated part, ψ r (x new ) and ψ u (x new ) are the statistics of the quality-related part and the quality-unrelated part, δ Ir and δ Qr are the quality-related control limits, δ Iu and δ Qu are the quality-unrelated control limits, I 2 (x new,r ) and Q(x new,r ) are the statistics of the quality-related part, I 2 (x new,u ) and Q(x new,u ) are the statistics of the quality-unrelated part.
[0152] S6: Based on the quality-related joint statistic, the quality-unrelated joint statistic, the quality-related joint control limit, and the quality-unrelated joint control limit, perform fault detection to obtain a fault detection result.
[0153] In a possible embodiment, the fault detection result includes a quality-related fault, a quality-unrelated fault, and normal system operation.
[0154] Further, S6 is specifically as follows:
[0155] According to the quality-related joint statistic and the quality-related joint control limit, determine whether the quality-related joint statistic of the quality-related part is greater than or equal to the quality-related joint control limit to obtain a first judgment result;
[0156] If the first judgment result is negative, according to the quality-unrelated joint statistic of the quality-unrelated part and the quality-unrelated joint control limit, determine whether the quality-unrelated joint statistic of the quality-unrelated part is greater than or equal to the quality-unrelated joint control limit to obtain a second judgment result; if the first judgment result is positive, a quality-related fault occurs in the system;
[0157] If the second judgment result is yes, a quality-independent fault occurs in the system; if the second judgment result is no, it is determined that the system is operating normally.
[0158] It can be understood that the above system fault judgment is based on the following judgment logic:
[0159]
[0160] For the fault detectability proposed in this application, the following theorem exists:
[0161] Theorem: For a nonlinear system, the sample x = x * +ξ j f j The sufficient condition for it to be detectable by the ψ statistic is:
[0162] or
[0163] where,
[0164]
[0165] Σ = δ I Σ Q -δ Q Σ I
[0166] δ = δ I -δ ψ δ I δ Q
[0167] In the formula, the operator (.) p represents the p-th element of the vector, x * and ξ j f j respectively refer to the faulty part and the fault-free part, δ I is the control limit of the statistic I, δ Q is the control limit of the statistic Q, and δ Ψ is the control limit of the statistic Ψ;
[0168] For the above theorem, the proof is as follows:
[0169] Given the sample x to be measured, its I 2 and Q statistics can be expressed as:
[0170]
[0171] where, I 2 (x) and Q(x) are the statistics of the sample x, and k(x) represents the kernel vector of the sample x.
[0172] The ψ statistic of sample x can be simplified to:
[0173]
[0174] where ψ ∈ {ψ r , ψ u}, δ I ∈ {δ Ir , δ Iu}, δ Q ∈ {δ Qr , δ Qu}. According to the diagnostic logic, a sufficient condition for sample x to be detected by the ψ statistic is min{ψ(x)} > δ ψ , which is equivalent to
[0175] min{δ I - k(x) T Σk(x)} > δ ψ δ I δ Q
[0176] The above equation can be further transformed into:
[0177] max{k(x) T Σk(x)} < δ
[0178] Analyze the relationship between the kernel vector k(x) and the jth variable. Substitute x = x * + ξ j f j into [K] i,j = k(x i , x j ) = exp(-||x i - x j || 2 / σ) to get:
[0179]
[0180] Then,
[0181]
[0182] Substitute the above into max{k(x) T Σk(x)} < δ to get:
[0183]
[0184] According to the Rayleigh entropy theorem, there is the following inequality:
[0185]
[0186] In this way, the formula Can be transformed into:
[0187]
[0188] In the above formula, when |(x * ) j + f j -(x t ) j | is the smallest it has the maximum value; when the elements of the sample x and x t are as close as possible (except for the j-th element) it has the maximum value. Therefore, it is necessary to find the sample x that satisfies the formula t* from the training set X. At this point, the above formula can be further transformed into:
[0189]
[0190] Substitute Δ j and k -j (x * , x i ) into the above formula, and take the logarithm of both sides of the inequality to obtain:
[0191]
[0192] Therefore, the sufficient condition for the sample x to be detected by the ψ statistic is: or
[0193] At this point, the theorem is proven.
[0194] Considering that large and complex equipment systems have characteristics such as multiple indicators, high dimensions, large samples, process variables, and key performance indicators, similar to the Tennessee - Eastman Process (TEP), this application uses TEP to verify the effectiveness of the proposed method. TEP is a complex nonlinear dynamic process and has been widely used to evaluate monitoring and diagnostic methods. The TEP process includes five main units: reactor, condenser, compressor, stripper, and separator. The method proposed in this application is verified below through the data collected in the TEP experiment.
[0195] There are two types of data in the TEP: 41 measurement variables, denoted as XMEAS(1 - 41), and 11 manipulated variables, denoted as XMV(1 - 11). The measurement variables XMEAS(1 - 22) and the manipulated variables XMV(1 - 11) together form the process variable set X, and XMEAS(35) is selected as the quality variable y characterizing the product quality. It should be specifically noted that in the methods proposed in this application, it is assumed that the quality variable is unmeasurable during verification, that is, the verification is carried out in the case of missing y.
[0196] (1) Experiment on correlation analysis between variables
[0197] Similar to numerical simulation, the MI between each variable and the quality variable y and the quality - related virtual variable α is calculated respectively. The results are as Figure 2 shown. The elements in the figure represent the MI values between the corresponding variables. NaN indicates immeasurable. The lighter the color in the figure, the stronger the correlation between variables, while the darker the color, the weaker the correlation.
[0198] From Figure 3 it can be seen that the MI between the process variables XMEAS(7, 13, 16) and the quality variable y are all greater than 1, and the MI between XMEAS(7, 13, 16) and the quality - related virtual variable α are also all greater than 1. This indicates that the variables strongly correlated with y are also strongly correlated with α; at the same time, the MI value between y and α is 1.32. These findings indicate that there is a strong correlation between y and α.
[0199] (2) Experiment on variable partitioning based on Qav
[0200] Considering the non - linear characteristics among the variables in the TEP, the method described in this application performs quality - related variable identification, and sets the sliding window width as L = 12 through cross - validation. The variable partitioning results are as Figure 3 shown, where the quality - related variables are marked in orange, the quality - unrelated variables are marked in blue, and the red dashed line represents the control limit. Specifically, Figure 3 (a) in shows the MI values between the true quality variable y and the process variables, while Figure 3 (b) in presents the MI values between Qav and the process variables. Through comparison, it can be seen that the variable partitioning results using Qav are exactly the same as those directly using the quality variable: the variables XMEAS(1, 3, 7, 10, 11, 13, 16, 18 - 20) and XMV(2, 3, 5, 6, 9) are identified as quality - related variables, and the remaining variables are quality - unrelated variables. This result further confirms the effectiveness of the proposed method. Even in complex situations where the quality variable cannot be directly obtained, the method can accurately identify the quality variable by constructing Qav.
[0201] (3) Experiment on quality - related fault monitoring based on Qav
[0202] There are 15 known faults in the TEP, denoted as IDV(1)-IDV(15). Among them, IDV(3, 9, 15) are minor faults and will not be discussed here. The remaining faults can be roughly divided into the following three categories according to their impacts on quality variables: The first category is strongly quality-related faults. The variables corresponding to these faults have crucial impacts on quality variables. For example, IDV(2), IDV(6), IDV(8), and IDV(12), etc. They will significantly change the state of quality variables throughout the process. The second category is weakly quality-related faults. Such faults will have certain impacts on quality variables in the initial period, but with the intervention of the fault-tolerant controller, the fault effects will gradually be eliminated, such as IDV(1), IDV(5), and IDV(7). The third category is quality-unrelated faults. These faults hardly have any impact on quality variables. For example, IDV(4), IDV(11), and IDV(14).
[0203] In this application, for comparative analysis, three non-linear methods with good performance are selected: MI-KPCA, KPI-KPLS, and KPCA-CCA for comparative experiments. The first 160 samples of each type of fault in the test dataset are normal samples, and the latter 800 are fault samples.
[0204] (i) Monitoring performance of quality-related faults. The MI-KPCA, KPI-KPLS, KPCA-CCA, and the proposed Qav-KICA methods are respectively used to monitor quality-related faults. The FDR of quality-related statistics is shown in Table 1, where the best monitoring results are shown in boldface.
[0205] Table 1 FDR of quality-related faults in TEP
[0206]
[0207] As can be seen from Table 1, in most fault modes, the proposed Qav-KICA method can achieve the highest FDR. Generally speaking, the detection rate of Qav-KICA is similar to that of MI-KPCA, but significantly better than KPI-KPLS and KPCA-CCA. The main reason is that a large number of false negatives occur when these two methods process the IDV(8) and IDV(12) fault modes. The average FDR of Qav-KICA is increased by 0.33, 13.00, and 23.83 percentage points compared with MI-KPCA, KPI-KPLS, and KPCA-CCA respectively.
[0208] Taking IDV(12) as an example, the changing trend of the quality variable XMEAS(35) over time is as Figure 4As shown, starting from the 161st sampling point, this variable shows a continuous trend of deviating from the normal operating condition and is accompanied by significant fluctuation characteristics, indicating that IDV(12) will affect product quality. Figure 5 shows the monitoring results of four monitoring methods, where the blue line represents the change trajectory of the statistic, and the red dashed line represents the control limit. In Figures 6(b) and Figure 5 (c), we can see that the statistics of a large number of samples are below the control limit, indicating that KPI-KPLS and KPCA-CCA cannot continuously and stably detect IDV(12) faults. On the contrary, in Figures 5(a) and Figure 5 (d), the statistics of the samples are all above the control limit, indicating that MI-KPCA and Qav-KICA can effectively detect IDV(12) faults.
[0209] (ii) Monitoring performance of quality-related faults. Similarly, four different methods are used to monitor IDV(4, 11, 14) for faults respectively, and the FAR (False Alarm Rates) and FDR (False Detection Rates) of each method are recorded in Table 2.
[0210] Table 2 FAR and FDR of quality-unrelated faults in TEP
[0211]
[0212] Analyzing the data in Table 2, it can be seen that the proposed Qav-KICA method is significantly higher than other methods in terms of FDR. At the same time, its FAR is significantly lower than the other three methods; compared with MI-KPCA, KPI-KPLS and KPCA-CCA are reduced by 1.13, 18.21 and 4.33 percentage points respectively. This result shows that the Qav-KICA method effectively reduces unnecessary false alarms while maintaining high-efficient fault detection ability, thus demonstrating more superior performance in the field of fault monitoring.
[0213] Taking IDV(11) as an example, the change trend over time is as Figure 6 shown. After the fault occurs, the quality variable does not change significantly, indicating that IDV(11) has no effect on this quality variable, so it is determined as a quality-unrelated fault. Figure 7 shows the fault monitoring results of the four methods. From Figure 7 (a), it can be seen that there are many undetected cases in the quality-unrelated statistics of MI-KPCA; while in Figure 7 (b) and Figure 7As shown in (c), during the monitoring process of KPI-KPLS and KPCA-CCA, the quality-related statistics of some samples are higher than the control limit, that is, there are some false alarms; while the proposed Qav-KICA method has a 79.50% FDR while maintaining a low FAR of 1.25%.
[0214] Aiming at the problem that traditional quality-related fault monitoring is difficult to implement due to the unmeasurable or unreliable quality-related variables in industrial processes, a fault monitoring method based on quality-associated virtual variables is proposed. This method is significantly different from traditional methods and gets rid of the dependence on quality variables. By constructing a quality-associated virtual variable Qav and applying it to quality-related fault monitoring. The quality-related fault monitoring method based on quality-associated virtual variables is applied to the Tennessee-Eastman industrial process data to illustrate and verify the proposed method. The proposed method effectively reduces the FAR and improves the FDR, verifying the feasibility and reliability of the proposed Qav-based fault monitoring method, as well as the sufficient conditions for fault detectability. The proposed method is especially suitable for industrial process fault monitoring where quality variables are difficult to obtain or the measured values are unreliable; in complex scenarios involving multiple quality variables, this method avoids the cumbersome one-by-one evaluation process between process variables and quality variables, greatly simplifying the computational workload. In addition, this Qav-based fault monitoring method has wide applicability. It can not only be effectively applied to quality-related fault monitoring based on variable partitioning methods, but also be applicable to quality-related fault monitoring based on space projection methods and other multivariate statistical methods.
[0215] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. The above description is only a preferred embodiment of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. A fault detection method for a nonlinear non-Gaussian system based on a quality-related virtual variable, characterized in that Including: S1: Construct the quality-associated virtual variable Qav; S2: Based on the Qav, use the mutual information MI calculation process to calculate the correlation between the process variables and the Qav, and divide the process variables into a quality-related variable group and a quality-unrelated variable group; S3: According to the quality-related variable group and the quality-unrelated variable group, establish KICA models, denoted as KICAr and KICAu respectively; S4: Based on the KICAr model and the KICAu model, calculate the quality-related statistics and quality-unrelated statistics of the training set samples, and obtain the quality-related control limit and the quality-unrelated control limit according to the kernel density estimation method; S5: Based on the division result of the process variables, divide the monitoring samples into a quality-related part and a quality-unrelated part, calculate the quality-related joint statistic and the quality-unrelated joint statistic of the monitoring samples, and determine the quality-related joint control limit and the quality-unrelated joint control limit; S6: Based on the quality-related joint statistic, the quality-unrelated joint statistic, the quality-related joint control limit and the quality-unrelated joint control limit, perform fault detection to obtain a fault detection result.
2. The fault detection method for a non-linear non-Gaussian system based on a quality-related virtual variable according to claim 1, wherein The specific content of S1 includes: S101: Construct a process variable matrix And perform normalization processing on the process variable matrix; where N and m represent the number of samples and the number of variables respectively, i = [1, 2,..., N], and x i is a column vector representing the sampling sample at time i, j = [1, 2,..., m], and x j is a row vector representing the time series data of the j-th variable; S102: Project the standardized process variable into a linear space \(x\to\varphi(x)\) through a non - linear mapping \(\varphi(\cdot)\). * (.) j by projecting the \(x\) j →φ * (x j ) S103: Construct an objective function to maximize the sum of the squares of the inner products of the quality-associated virtual variable α and all variables: Among them, is a quality-related dummy variable, α T is the transpose of the quality-related dummy variable, α T α is the outer product operation of the quality-related dummy variable, N is the number of samples, φ * (x j ) is the projection of the variable x j in the linear space; S104: According to the Lagrange multiplier method, transform the objective function into: Take the partial derivative of α: wherein, is the partial derivative of L(α) with respect to α; S105: Perform eigenvalue decomposition: Among them, Φ * is the matrix that maps the process variable to a high-dimensional linear space through the non-linear transformation φ * (.), Φ *T is the transposed matrix corresponding to Φ * , and λ is the eigenvalue; S106: Select the eigenvector α corresponding to the largest eigenvalue as Qav.
3. The fault detection method for a non-linear non-Gaussian system based on quality-related virtual variables according to claim 1, wherein The specific content of S2 includes: S201: Calculate the mutual information MI value between the process variables and the Qav; S202: According to the MI value and a preset threshold, divide the process variables to obtain a quality-related variable group and a quality-unrelated variable group.
4. The fault detection method for a non-linear non-Gaussian system based on quality-related virtual variables according to claim 3, characterized in that, Calculating the mutual information MI value between the process variables and the Qav in S201 includes: Use a sliding window with a window width of L to intercept the time series data into m - L + 1 sub-blocks. The i-th sub-block can be expressed as: Among them, is the i-th sub-block, represents the matrix of quality-related dummy variable matrix, is the quality-related dummy variable matrix, α i is the quality-related dummy variable, x i represents the i-th scalar, and c is the kernel parameter; Denote the j-th variable in and as the MI between them as Calculate the MI value of each sub-block. The MI vector between the j-th variable and Qav is expressed as: where I j is the MI vector between the j-th variable and Qav, is the MI between the j-th variable and in is the j-th variable in Calculate the vector I j mean value and use the mean value as the mutual information MI value between the j-th variable and Qav.
5. The fault detection method for a nonlinear non-Gaussian system based on quality-associated virtual variables according to claim 4, characterized in that, The dividing the process variables according to the MI value and a preset threshold to obtain a quality-related variable group and a quality-unrelated variable group includes: If the mean value of the MI value is greater than the preset threshold, the variable corresponding to the MI value is a quality-related variable in the quality-related variable group; otherwise, the variable corresponding to the MI value is a quality-unrelated variable in the quality-unrelated variable group.
6. The fault detection method for a non-linear non-Gaussian system based on a quality-correlated virtual variable according to claim 3, characterized in that The preset threshold includes: Use the Monte Carlo method to generate a group of random samples that follow a normal distribution; Use a sliding window with a fixed window width to block the random samples and the Qav, and calculate the MI value between each variable of the random samples and the Qav; Arrange the MI values in ascending order, and select the quantile with a confidence level of 95% as the preset threshold.
7. The fault detection method for a non-linear and non-Gaussian system based on quality-associated virtual variables according to claim 1, wherein In S5, based on the division result of the process variables, divide the monitoring samples into a quality-related part and a quality-unrelated part, and calculate the quality-related joint statistic and the quality-unrelated joint statistic, specifically including: Project the quality-related part and the quality-unrelated part onto the KICAr model and the KICAu model respectively, and calculate the quality-related joint statistic and the quality-unrelated joint statistic according to the calculation formula. The calculation formula is as follows: Among them, x new is the monitoring sample, x new,r is the quality-related part, x new,u is the quality-unrelated part, ψ r (x new ) and ψ u (x new ) are the quality-related joint statistic and the quality-unrelated joint statistic, δ Ir and δ Qr are the quality-related control limits, δ Iu and δ Qu are the quality-unrelated control limits, I 2 (x new,r ) and Q(x new,r ) are the statistics of the quality-related part, I 2 (x new,u ) and Q(x new,u ) are the statistics of the quality-unrelated part.
8. The fault detection method for a non-linear non-Gaussian system based on quality-related virtual variables according to claim 1, characterized in that The specific content of S6 includes: According to the quality-related joint statistic and the quality-related joint control limit, determine whether the quality-related joint statistic of the quality-related part is greater than or equal to the quality-related joint control limit, and obtain a first judgment result; If the first judgment result is negative, according to the quality-unrelated joint statistic of the quality-unrelated part and the quality-unrelated joint control limit, determine whether the quality-unrelated joint statistic of the quality-unrelated part is greater than or equal to the quality-unrelated joint control limit, and obtain a second judgment result; if the first judgment result is positive, a quality-related fault occurs in the system; If the second judgment result is positive, a quality-unrelated fault occurs in the system; if the second judgment result is negative, it is determined that the system is operating normally.