Industrial process fault detection method based on deep multi-sampling probability principal component analysis

By building an end-to-end equivalent model through deep multi-sampling probabilistic principal component analysis, the feature extraction problem of multi-sampling rate data is solved, efficient fault detection and real-time monitoring are achieved, and the detection accuracy and sensitivity of the wastewater treatment process are improved.

CN120705695APending Publication Date: 2025-09-26ZHEJIANG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510736462.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing multivariate statistical methods and deep learning methods cannot effectively capture feature distribution when processing multi-sampling rate industrial process data, resulting in inaccurate fault detection or poor interpretability, and unable to achieve timely fault monitoring.

Method used

A method based on deep multi-sampling probabilistic principal component analysis is adopted to build an end-to-end equivalent model. A multi-layer cascade structure and expectation maximization algorithm are used to extract features and optimize model parameters. T2 and SPE statistics are established to achieve effective processing of multi-sampling rate data.

Benefits of technology

It significantly improves the accuracy and sensitivity of multi-sampling rate process fault detection, provides a more reliable real-time monitoring method, can effectively identify abnormal conditions, reduce computing pressure and retain the interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005433444690000031
    Figure BDA0005433444690000031
  • Figure BDA0005433444690000037
    Figure BDA0005433444690000037
  • Figure BDA0005433444690000039
    Figure BDA0005433444690000039
Patent Text Reader

Abstract

The invention discloses an industrial process fault detection method based on a deep multi-sampling probability principal component analysis model, and the method comprises the steps: carrying out the online collection of industrial process online measurement data, obtaining a potential factor corresponding to the online measurement data through a constructed end-to-end equivalent model, calculating the T2 and SPE statistics at a moment, comparing the T2 and SPE statistics with a statistical threshold value, and carrying out the fault detection. Judging whether a fault occurs or not; in the model construction stage, firstly, a deep multi-sampling probability principal component analysis model is constructed by utilizing a training sample set, initial values of parameters of an end-to-end equivalent model are obtained according to model parameters of each layer of the deep multi-sampling probability principal component analysis model, and finally, the end-to-end equivalent model is obtained through optimization. Compared with a traditional multi-sampling-rate model, the model provided by the invention has higher detection sensitivity and accuracy, and provides a more reliable real-time monitoring means for industrial processes such as a wastewater treatment process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial process monitoring, and in particular relates to an industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis. Background Art

[0002] Industrial process fault detection is a key technology for ensuring production safety, improving equipment reliability, and optimizing production efficiency. Take wastewater treatment, for example. With the acceleration of modern industrialization and urbanization, wastewater discharge continues to grow, and water pollution has become a major global environmental issue. Wastewater treatment processes treat wastewater and return it to the water cycle with minimal environmental impact or direct reuse. Due to the complexity of operating conditions and variability in raw materials, wastewater treatment processes can sometimes experience multiple faults, placing higher demands on process monitoring. Within wastewater treatment processes, data sampling frequencies vary significantly due to differences in the dynamic characteristics of the measurement objects and different acquisition methods, exhibiting a significant multi-rate characteristic. Fast dynamic process parameters (such as water temperature and pressure) can be sampled at high frequencies in milliseconds; equipment status parameters (such as flow rate and speed) are typically collected at sub-second or graded frequencies; quality indicators (such as concentration) that require offline analysis and manual inspection data may only be acquired once every few hours or even weeks. Traditional multivariate statistical methods (such as principal component analysis, partial least squares method and its extended methods) are all based on the assumption of a single sampling rate. They cannot effectively process data with multiple sampling rate characteristics and make timely and accurate judgments on the status of industrial processes.

[0003] Currently, most fault monitoring methods for multi-rate processes rely on probabilistic frameworks or deep learning for process modeling. While probabilistic process modeling considers the global dynamics of multi-rate processes and can directly process multi-rate data, it struggles to fully capture the characteristic distribution of multi-rate data through a single feature extraction. Deep learning-based multi-rate process modeling can achieve high accuracy, but the black-box models used suffer from poor interpretability and destroy the original data structure. To effectively utilize information from multi-rate processes and improve process monitoring performance, an efficient process data modeling and fault monitoring approach is needed. Summary of the Invention

[0004] In order to overcome the shortcomings and deficiencies of the existing technology, the present invention provides an industrial process fault monitoring method based on a deep multi-sampling probabilistic principal component analysis model, which performs multi-level extraction of features of multi-sampling rate data, thereby effectively improving the accuracy of multi-sampling rate process fault detection.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] A method for industrial process fault detection based on deep multi-sampling probabilistic principal component analysis is proposed. The online measurement data of the wastewater treatment process is collected online. The latent factors corresponding to the online measurement data are obtained by using the constructed end-to-end equivalent model. The T at that moment is calculated. 2 and SPE statistics, and compare them with the statistical threshold to determine whether a fault has occurred; in the model construction stage of the end-to-end equivalent model, the deep multi-sampling probabilistic principal component analysis model is first constructed using the training sample set, and the initial values ​​of the end-to-end equivalent model parameters are obtained according to the model parameters of each layer of the deep multi-sampling probabilistic principal component analysis model, and finally the end-to-end equivalent model is optimized.

[0007] Furthermore, the end-to-end equivalent model structure adopts a multi-sampling factor analysis (MFA) model structure.

[0008] Furthermore, in the deep multi-sampling probabilistic principal component analysis model, the first layer adopts a multi-sampling probabilistic principal component analysis (MPPCA) model structure, and the remaining layers adopt a probabilistic principal component analysis (PPCA) model structure.

[0009] Furthermore, the deep multi-sampling probabilistic principal component analysis model and / or the end-to-end equivalent model adopts an expectation maximization (EM) algorithm to achieve estimation of latent variables and optimization of model parameters.

[0010] Furthermore, a wastewater treatment process fault detection method based on deep multi-sampling probabilistic principal component analysis includes the following steps:

[0011] (1) Collect data samples online from industrial processes (e.g., wastewater treatment processes) to form a training sample set for modeling;

[0012] (2) Perform normalization preprocessing on the training sample set to eliminate the dimensional differences between data samples;

[0013] (3) For the normalized training sample set, a multi-sampling probability principal component analysis model is established, and the expectation maximization algorithm is used to obtain the model parameters and latent variables corresponding to the first layer of the deep multi-sampling probability principal component analysis model;

[0014] (4) The latent variables extracted from the first layer are used as the input of the second layer, and the expectation maximization algorithm is used to establish the next layer of probabilistic principal component analysis model to obtain the corresponding model parameters and deep feature information; this process is repeated L-1 times to complete the layer-by-layer feature extraction of the deep multi-sampled probabilistic principal component analysis of the L layer, and obtain the deep features corresponding to the original multi-sampled input;

[0015] (5) Based on the relationship between the original multi-sampled input and the deep features, the initial values ​​of the equivalent model parameters are calculated, and the expectation maximization algorithm is used to construct an end-to-end multi-sampled factor analysis model to complete the construction of the end-to-end equivalent model;

[0016] (6) At the same time, in the end-to-end equivalent model training phase, calculate T 2 and SPE statistics, and calculate the statistical thresholds respectively;

[0017] (7) Collect new online measurement data of industrial processes and preprocess them using modeling parameters to obtain test data;

[0018] (8) Input the normalized test data into the end-to-end equivalent model, extract the latent factors in the new sampled data, and calculate the residual matrix to calculate T at that moment. 2 and SPE statistics, and compare them with the statistical threshold to determine the process status.

[0019] Furthermore, the structure of the end-to-end equivalent model is:

[0020]

[0021] The subscript i corresponds to the i-th sampling rate, is an equivalent data sample, is the load matrix is the bias vector, is noise; t (L) is the latent variable of the model;

[0022] The initial values ​​of are obtained by the following formula:

[0023]

[0024] By the noise variance matrix Σ i By singular value decomposition, Σ i It is obtained from the following formula:

[0025]

[0026] in, is the noise variance corresponding to the first layer of the deep multi-sampling probability principal component analysis model; σ (2) , σ (l+1) are the noise variances corresponding to the 2nd layer and the l+1 layer respectively, is the load matrix corresponding to the first layer model of the deep multi-sampling probability principal component analysis model, W j is the load matrix corresponding to the j-th layer model of the deep multi-sampling probability principal component analysis model, and L is the total number of layers of the deep multi-sampling probability principal component analysis model;

[0027] W' i Obtained by the following formula:

[0028]

[0029] Among them, W (l) is the loading matrix corresponding to the l-th layer model of the deep multi-sampling probabilistic principal component analysis model;

[0030] X i is the data sample;

[0031] μ (2) It is the bias vector corresponding to the second layer model of the deep multi-sampling probabilistic principal component analysis model.

[0032] Furthermore, in the stage of training the end-to-end equivalent model, T is obtained through F distribution and chi-square distribution. 2 and the threshold value of the SPE statistic.

[0033] In step (1), an industrial process (such as a wastewater treatment process) is sampled online until a sample sequence with T sampling moments is collected to form a training sample set for modeling. The sampling rate of the collected data samples is generally different.

[0034] In step (2), the training sample set is preprocessed so that the mean of each variable is 0 and the variance is 1, eliminating the influence of dimensional differences between variables.

[0035] In step (3), a deep multi-sampled probabilistic principal component analysis (DeMPPCA) model is constructed. For the normalized training sample set, a multi-sampled probabilistic principal component analysis (MPPCA) model is established, and the model parameters are estimated using the expectation maximization (EM) algorithm.

[0036] In step (4), the next layer of probabilistic principal component analysis (PPCA) model is established for the latent variable matrix extracted in step (3) to deeply extract the potential information in the features. This process is repeated L-1 times, and the layer-by-layer feature extraction stage of the deep multi-sampling probabilistic principal component analysis of L layers is completed.

[0037] In step (5), based on the relationship between the original multi-sampled input and the deep features, the input and output are connected to construct an end-to-end equivalent model, which is equivalent to the multi-sampled factor analysis (MFA) model. The initial values ​​of the equivalent model parameters are calculated based on the model parameters of each layer, and the parameters are fine-tuned based on the parameter update method of MFA. The DeMPPCA model training is completed.

[0038] In step (6), a fault detection model is constructed based on DeMPPCA. During offline training, the model is built based on the training samples under normal working conditions, and T is calculated. 2 and SPE statistics, and the statistical thresholds were estimated and calculated using F distribution and chi-square distribution respectively.

[0039] In step (7), new online measurement data of the industrial process are collected and preprocessed and normalized using the modeling parameters, that is, the mean of each variable during modeling is subtracted and divided by the corresponding standard deviation, so that the new data is unified to a dimensional scale that is suitable for the model.

[0040] In step (8), the normalized test data is input into the end-to-end equivalent model, the latent factors in the new sampled data are extracted, and the residual matrix is ​​calculated. 2 and SPE statistics, and compare them with the statistical threshold to determine the process status.

[0041] As an embodiment, the industrial process is a wastewater treatment process.

[0042] This paper introduces deep learning concepts into a multi-sampled probabilistic principal component analysis model, constructing a multi-layer cascade deep architecture. Through a deep feature extraction mechanism, this model can more effectively mine deep-level feature information from multi-sampled process data. By connecting the input and output terminals, the model strengthens the characterization of deep features, significantly improving the online monitoring performance of abnormal conditions in wastewater treatment processes. Compared with traditional multi-sampled models, the proposed model has higher detection sensitivity and accuracy, providing a more reliable real-time monitoring method for wastewater treatment processes.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] (1) This method is a multi-sampling rate data processing method based on probabilistic modeling theory, which can effectively handle random noise in industrial processes and significantly improve the system's robustness to outliers. It uses the expectation maximization (EM) algorithm to estimate latent variables and optimize model parameters, effectively alleviating the computational pressure of industrial high-dimensional data.

[0045] (2) This method uses a hierarchical cascade structure to increase the complexity of the multi-sampling rate model and improve the feature representation and extraction capabilities. Since traditional machine learning methods are transparent and have good interpretability, through effective deepening, it can not only enhance the performance of traditional models but also retain the advantages of the original models.

[0046] (3) This method, based on the construction of a deep model, refers to the end-to-end form of deep learning, establishes an end-to-end connection between the original input and the deep features to construct an equivalent model. The equivalent model is a single-layer multi-sampling rate model, which means that the deep structure collapses to a single-layer form, simplifying the model structure. In addition, the shallow and deep knowledge in the layer-by-layer feature extraction data is transferred to the equivalent model through parameters. The equivalent model can obtain good initialization parameters, reducing the possibility of falling into a poor local optimal value. The present invention significantly improves the accuracy and timeliness of fault detection through feature extraction and deep information mining of the multi-sampling rate process. The fault detection based on this model can effectively identify abnormal conditions in the multi-sampling rate process. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flow chart of the multi-sampling rate process fault detection method based on the deep multi-sampling probability principal component analysis model of the present invention.

[0048] Figure 2 Flow chart of the wastewater treatment process.

[0049] Figure 3 This is a diagram of the architecture of the deep multi-sampling probabilistic principal component analysis model of the present invention.

[0050] Figure 4 This is the result of the deep multi-sampling probability principal component analysis model of the present invention and other models in detecting the wastewater treatment process. DETAILED DESCRIPTION

[0051] The invention will be further described below with reference to specific embodiments, but the scope of protection of the invention is not limited thereto. Those skilled in the art can and should know that any simple changes or substitutions based on the essence of the invention should fall within the scope of protection claimed by the invention.

[0052] This invention addresses the problem of online real-time detection of multi-sampling process cases, such as wastewater treatment processes. Based on the data collected during the process and a deep multi-sampling probabilistic principal component analysis model, it can promptly capture process anomalies and detect faults.

[0053] See also Figure 1 The main steps of the technical solution adopted by the present invention are as follows:

[0054] The first step is to perform online sampling of the wastewater treatment process through various measuring instruments installed in the wastewater treatment process until the sample sequence of the full T sampling time is collected, which can be recorded as The training sample set for modeling training is composed of is the sample data of the Sth sampling rate, N S is the number of data in the data sample with sampling rate S, MS is the number of variables with a sampling rate of S. In this embodiment, there are S sampling rates in total. Since the sampling frequencies of variables in industrial processes vary, for industrial processes with S different sampling rates, the number and dimensions of variables collected at the same time are different, resulting in S sets of data subsets with different dimensions.

[0055] The second step is to normalize the training sample set so that the mean of each variable is 0 and the variance is 1, eliminating the influence of the dimensional difference between variables, and obtain the normalized data sequence X=[X1 X2...X S ].

[0056] The third step is to use the normalized data sequence X=[X1 X2...X S ], training a deep multi-sampled probabilistic principal component analysis model (see Figure 3 ), the EM algorithm is used to optimize the model parameters Θ of each layer. Specifically, the first layer of the deep multi-sampled probabilistic principal component analysis (DeMPPCA) model is the original multi-sampled probabilistic principal component analysis (MPPCA), and the basic form is as follows:

[0057]

[0058] in represents one of the data sample observations at sampling time k (corresponding to the i-th sampling rate), is the loading matrix corresponding to the sampling rate, and To measure noise, the Gaussian distribution is is the variance of the sample noise corresponding to the i-th sampling rate corresponding to the first layer. Latent variable t (1) ∈R D The constraints between measurements at different sampling rates are extracted and stored, and shared across all sampling rates at sampling time k. The model parameters of the first MPPCA layer are denoted as The model parameters are updated using the EM algorithm. In the E step of model parameter estimation, the posterior distribution estimate of the latent variable is obtained based on the current model parameters. The specific formula is:

[0059]

[0060] in is the latent variable expectation based on the k sampling moment X(k). ki It is a sampling index, which is used to record the sampling status of the i-th sampling rate at the current sampling moment. When the variable is collected at this sampling moment, the sampling index φ ki =1, otherwise 0. In step M, the updated values ​​of the model parameters are obtained according to the updated results of step E, where N is the number of samples corresponding to the maximum sampling rate. The parameter update method is:

[0061]

[0062] Where trace(·) is the matrix trace operator.

[0063] From the third step, the model parameters and features t of the first layer model of the deep multi-sampling probability principal component analysis model can be obtained (1) .

[0064] The fourth step is to extract the features t from the first layer (1) As the input of the second layer, the extracted features are complete data structures, and PPCA (Probabilistic Principal Component Analysis) is used to extract deeper information. The model form of the second layer is:

[0065] t (1) =W (2) t (2) +μ (2) +e (2) (7)

[0066] where μ (2) is the bias vector corresponding to the second layer, W (2) is the loading matrix corresponding to the second layer, e (2) is the measurement noise corresponding to the second layer. The subsequent layers are similar to the second layer and are also modeled using PPCA in the form of (taking the l-1 layer as an example, l = 2, 3, 4,,…L, L is the total number of layers in the deep multi-sampling probabilistic principal component analysis model):

[0067] t (l-1) =W (l) t (l) +μ (l) +e (l) (8)

[0068] where μ (l) is the bias vector corresponding to the lth layer, W (l) is the loading matrix corresponding to the lth layer, e (l) is the measurement noise corresponding to the lth layer, σ (l) is the variance of the Gaussian distribution of the sample noise corresponding to the lth layer. The model parameters of each layer are estimated using the EM algorithm. In the E step, the posterior distribution estimate of the latent variable corresponding to the lth layer at time k is obtained based on the current model parameters.

[0069]

[0070] G=(W (l)T W (l) +σ (l) I) -1 (10)

[0071] The model parameters are updated in the M step as follows:

[0072]

[0073] The fifth step is to calculate the probability distribution estimate between the original multi-sampled input and the deep features based on the relationship between the input and output in each layer model to construct an end-to-end equivalent model. First, for each layer of the model, the prior distribution of the latent variable is defined as:

[0074] p(t)=N(0,I) (13)

[0075] For the first layer of the deep model (deep multi-sampling probability principal component analysis model), the conditional distribution of the multi-sampling data can be obtained as:

[0076]

[0077] For the subsequent l-layer PPCA model, the following conditional distribution is obtained:

[0078] p(t (l-1) |t (l) )=N(W (l) t (l) +μ (l) ,σ (l) I) (15)

[0079] For the L-layer DeMPPCA model, its end-to-end multi-sampled input X and deep output t (L) The relationship between can be expressed in multiple integral forms:

[0080]

[0081] Since all conditional probabilities in DeMPPCA are multivariate Gaussian distributions, Equation (16) can be solved to obtain a closed-form solution, resulting in the following end-to-end conditional distribution:

[0082]

[0083] According to the characteristics of multivariate Gaussian distribution, we construct the (L) The end-to-end generation process to X is:

[0084] X i =W' i t (L) +μ' i +ε' i (18)

[0085]

[0086] ε' i ~N(0,Σ i ) (twenty one)

[0087]

[0088] Since the bias of each layer of PPCA is 0 after the second layer, the bias vector is simplified. i Is a real symmetric matrix, through the singular value decomposition to obtain the orthogonal matrix P and diagonal matrix Λ, the decomposition form is:

[0089]

[0090] Through the orthogonal matrix P i Perform a linear transformation on the model parameters in the form of:

[0091]

[0092]

[0093] After linear transformation, the equivalent model (18) is rewritten as:

[0094]

[0095] The measurement noise Formula (27) is equivalent to the multi-sampling factor analysis (MFA) model. Based on the parameter update method of MFA, the parameters are fine-tuned. The model parameters obtained by the DeMPPCA model are used to calculate the model parameters of the multi-sampling factor analysis (MFA) model. The parameters are used as the initial values ​​and the EM algorithm is also used to estimate the model parameters. In the E step, we get:

[0096]

[0097] The model parameters are updated in the M step as follows:

[0098]

[0099] The subscript i corresponds to the i-th sampling rate, is an equivalent data sample, is the load matrix is the bias vector, is noise; t (L) is the latent variable of the model, and its initial value is the deep output of the deep multi-sampling probability principal component analysis model; is the noise variance corresponding to the first layer of the deep multi-sampling probability principal component analysis model; σ (2) , σ (l+1) are the noise variances corresponding to the 2nd layer and the l+1 layer respectively, is the load matrix corresponding to the first layer model of the deep multi-sampling probability principal component analysis model, W j is the load matrix corresponding to the j-th layer model of the deep multi-sampling probability principal component analysis model, L is the total number of layers of the deep multi-sampling probability principal component analysis model; W (l) is the load matrix corresponding to the l-th layer model of the deep multi-sampling probability principal component analysis model; X i is the data sample; μ (2) It is the bias vector corresponding to the second layer model of the deep multi-sampling probabilistic principal component analysis model.

[0100] Step 6: Use T 2 The process is monitored by using the SPE statistic. The threshold of the monitoring statistic is estimated during the training phase using the F distribution and the chi-square distribution (refer to the method mentioned in the publication number CN119644832A for calculating T 2 and the SPE statistic threshold (or control limit) T 2 lim and SPE lim ).

[0101] The seventh step is to collect new wastewater treatment process operation data and normalize it using modeling parameters, that is, subtract the mean of each variable during modeling and divide it by the corresponding standard deviation, so that the new data is unified to a dimensional scale that is suitable for the model.

[0102] The eighth step is to input the normalized test data into the end-to-end equivalent model, extract the latent factors in the new sampled data and obtain the residual matrix, and calculate T 2 And SPE statistics. When a new online sampling sample X is obtained o (k new ), calculate T 2 And the SPE statistic:

[0103]

[0104] SPE i (k new )=e i (k new ) T e i (k new ) (36)

[0105] The obtained T 2 Compared with the SPE statistics and the statistical threshold, if the statistics at that moment exceed the threshold, it is determined to be a fault condition, otherwise the process operation status is normal.

[0106] To demonstrate the fault detection performance of the deep multi-sampled probabilistic principal component analysis (MPPCA) of the present invention, multi-rate data from a wastewater treatment process was used for validation. The MPPCA and MFA methods, relevant to the present invention, were compared with the relatively effective multi-sampled Gaussian linear state model (MLGSS). Fifteen process and quality variables from the R2S anaerobic reactor in the wastewater treatment process were selected for modeling. The variable descriptions are shown in Table 1.

[0107] Table 1

[0108]

[0109] Table 2

[0110]

[0111]

[0112] Table 2 reports the false alarm rate (FAR) under normal conditions and the false alarm rate and missed detection rate (MDR) for two selected fault scenarios. The bold numbers indicate the optimal missed detection rate for each fault. As can be seen from Table 2, for the two selected fault types, the fault detection performance of the present invention achieves the optimal missed detection rate, while the false alarm rate is within an acceptable range, validating the advantages of the deep multi-sampling probabilistic principal component analysis model of the present invention. Figure 4 The results of fault 1 in the wastewater treatment process are shown in Figure 2. This fault occurred between 480-820 and 900-1200. Comparing the results of DeMPPCA with those of MPPCA, MFA, and MLGSS, MLGSS has a period of false positives under normal operating conditions and cannot obtain reliable detection results. The statistics obtained by MPPCA and MFA fluctuate greatly and have a high missed detection rate. DeMPPCA shows excellent detection capabilities, especially at T 2 The performance of statistics.

Claims

1. An industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis, characterized in that: Collect online measurement data of industrial processes online, use the constructed end-to-end equivalent model to obtain the potential factors corresponding to the online measurement data, and calculate T at that moment. 2 and SPE statistics, and compare them with the statistical threshold to determine whether a fault has occurred; in the model construction stage, the deep multi-sampling probabilistic principal component analysis model is first constructed using the training sample set, and the initial values ​​of the end-to-end equivalent model parameters are obtained according to the model parameters of each layer of the deep multi-sampling probabilistic principal component analysis model, and finally the end-to-end equivalent model is optimized.

2. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: The end-to-end equivalent model structure adopts a multi-sampling factor analysis model structure.

3. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: In the deep multi-sampling probabilistic principal component analysis model, the first layer adopts a multi-sampling probabilistic principal component analysis model structure, and the remaining layers adopt a probabilistic principal component analysis model structure.

4. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: The deep multi-sampling probabilistic principal component analysis model and / or the end-to-end equivalent model adopts the expectation maximization algorithm to achieve the estimation of latent variables and the optimization of model parameters.

5. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: The steps include: (1) Collect data samples from industrial processes online to form a training sample set for modeling; (2) Perform normalization preprocessing on the training sample set to eliminate the dimensional differences between data samples; (3) For the normalized training sample set, a multi-sampling probability principal component analysis model is established, and the expectation maximization algorithm is used to obtain the model parameters and latent variables corresponding to the first layer of the deep multi-sampling probability principal component analysis model; (4) The latent variables extracted from the first layer are used as the input of the second layer, and the expectation maximization algorithm is used to establish the next layer of probabilistic principal component analysis model to obtain the corresponding model parameters and deep feature information; this process is repeated L-1 times to complete the layer-by-layer feature extraction of the deep multi-sampled probabilistic principal component analysis of the L layer, and obtain the deep features corresponding to the original multi-sampled input; (5) Based on the relationship between the original multi-sampled input and the deep features, the initial values ​​of the equivalent model parameters are calculated, and the expectation maximization algorithm is used to construct an end-to-end multi-sampled factor analysis model to complete the construction of the end-to-end equivalent model; (6) At the same time, in the end-to-end equivalent model training phase, calculate T 2 and SPE statistics, and calculate the statistical thresholds respectively; (7) Collect new online measurement data of industrial processes and preprocess them using modeling parameters to obtain test data; (8) Input the normalized test data into the end-to-end equivalent model, extract the latent factors in the new sampled data, and calculate the residual matrix to calculate T at that moment. 2 and SPE statistics, and compare them with the statistical threshold to determine the process status.

6. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: The structure of the end-to-end equivalent model is: The subscript i corresponds to the i-th sampling rate, is an equivalent data sample, is the load matrix is the bias vector, is the noise, t (L) is the latent variable of the model; The initial values ​​of are obtained by the following formula: By the noise variance matrix Σ i By singular value decomposition, Σ i It is obtained from the following formula: in, is the noise variance corresponding to the first layer of the deep multi-sampling probability principal component analysis model; σ (2) , σ (l+1) are the noise variances corresponding to the 2nd layer and the l+1 layer respectively, is the load matrix corresponding to the first layer model of the deep multi-sampling probability principal component analysis model, W j is the load matrix corresponding to the j-th layer model of the deep multi-sampling probability principal component analysis model, and L is the total number of layers of the deep multi-sampling probability principal component analysis model; W i 'Obtained by the following formula: Among them, W (l) is the loading matrix corresponding to the l-th layer model of the deep multi-sampling probabilistic principal component analysis model; X i is the data sample; μ (2) It is the bias vector corresponding to the second layer model of the deep multi-sampling probabilistic principal component analysis model.

7. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: In the stage of training the end-to-end equivalent model, T is obtained through F distribution and chi-square distribution. 2 and the threshold value of the SPE statistic.

8. The industrial process fault detection method based on deep multi-sampling probabilistic principal component analysis according to claim 1 is characterized in that: The industrial process is a wastewater treatment process.

Citation Information

Patent Citations

  • Industrial process monitoring method based on lightweight deep principal component analysis-auto-encoder model

    CN119644832A

Cited By

  • Industrial system fault monitoring method based on physical information enhanced probability principal component analysis

    CN120974165A