An Adaptive Weighted Semi-Supervised Probabilistic CPLS Quality-Related Monitoring Method
Through the adaptive weighted semi-supervised probability concurrent latent structure projection method, the problem of scarcity and high uncertainty of quality data in the process industry is solved, and high-precision quality-related monitoring and fault detection are realized, which is suitable for immediate learning and fault monitoring of continuous processes.
Patent Information
- Application Number
- CN202011575920.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-12-28
AI Technical Summary
The prior art is difficult to effectively monitor quality-related information in the process industry, especially when quality data is scarce and uncertain, it is impossible to accurately predict and soft measurement modeling, and traditional methods ignore the local spatial and temporal dynamics of the data.
Adaptive weighted semi-supervised probability concurrent latent structural projection (AWS-CPLS) method is adopted to improve model accuracy and fault detection capabilities through data preprocessing, model hidden variable initialization, expectation maximization EM algorithm and adaptive weighting calculation, combined with local weighting principles and probability graph theory.
Effectively utilize local spatiotemporal information, solve the problems of sparse mark samples and noise uncertainty, realize automatic fault detection of quality data and process data, improve the accuracy and fault monitoring performance of soft measurement models, and is suitable for immediate learning and fault monitoring of continuous processes.
Smart Images

Figure CN114692708B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial process monitoring, and particularly relates to a quality-related monitoring method based on adaptive weighted semi-supervised probabilistic CPLS. Background Art
[0002] Data-driven modeling methods for process industrial processes must consider the characteristics of high-dimensional data, hidden information, and complex noise. In the past 20-odd years, data-driven modeling methods have made great progress in theoretical research and practical applications. Multivariate statistical process control (MSPC) methods are widely used in the monitoring and fault diagnosis of process industrial processes, such as principal component analysis (PCA), partial least squares (PLS), slow feature analysis (SFA), linear discriminant analysis (LDA), and independent component analysis (ICA) methods. MSPC methods extract information from process data through dimensionality reduction projection methods and construct effective fault monitoring statistics. In recent years, data-driven regression modeling methods for process key quality indicators have been successfully used for the estimation of key performance of process product quality and fault monitoring, not only saving a large amount of manpower and material resources and improving the monitoring of quality-sensitive faults. Moreover, typical regression modeling methods mainly include PLS, artificial neural networks, and support vector regression methods. Partial least squares (PLS) is a regression modeling method based on data-driven that can handle multiple dependent variables against multiple independent variables. Because of its characteristic of extracting quality-related information, it has been widely used in quality-related complex industrial process monitoring and has become a research hotspot in the field of complex industrial process fault detection and diagnosis in recent decades [References 1-5].
[0003] In the actual operation of the process industry, it is often affected by various uncertain factors, making the process observation variables highly random. Therefore, process data described by probability models and monitoring models are easier to interpret. Probability modeling methods have the advantages of being easy to handle the problem of missing latent variables in process data and having good scalability, which is beneficial to solving the problems of excessive dimensions of large-scale data and modeling of uncertain data [References 6-9]. In addition, the variational Bayesian inference method can be used to automatically determine the parameters of the probability model, avoid overfitting problems, and make better use of the implicit information in the process data. Currently, data-driven probability modeling methods mainly include (mixture) probabilistic principal component analysis (PPCA), probabilistic partial least squares ((M)PPLS), probabilistic linear state space (PLSS), Dirichlet process Gaussian mixture model (DPGMM), probabilistic Fisher discriminant analysis (PFDA), variational autoencoder (VAE), etc. Theoretical analysis and simulation results show that probability modeling methods can effectively solve the problems of sample uncertainty and insufficient sample quantity, and achieve ideal results in soft sensor modeling and process industry monitoring.
[0004] To overcome the shortcomings of PLS, Zhou D (2010) proposed the total projection to latent space (T-PLS) algorithm for process monitoring [Reference 10] based on the idea that different types of information belong to different subspaces, so as to achieve accurate monitoring of different types of information. T-PLS divides the input X into a subspace X y related to the output Y and a subspace X o orthogonal to the output. The remaining part is further divided into a subspace X r with larger variation and a residual subspace E r . The output Y is divided into a part related to the subspace X yRelated predictions and prediction errors. Obviously, this model can achieve quality-related fault prediction or soft sensor modeling. In practice, the influence of irrelevant process variables on the prediction accuracy of the PLS model is considered to improve the explanatory ability of process quality variables for process variables and the monitoring performance. Zhao C in 2013 extended T-PLS to multiple spaces and constructed multiple process statistics [Reference 11]. Although T-PLS has the ability to explain process variables and the ability of multiple monitoring statistics to improve process monitoring accuracy, however, this model cannot monitor unpredictable quality changes, and dividing the input space into too many subspaces is not conducive to improving the monitoring accuracy. To address the above problems, Qin S J (2013) proposed the concurrent projection to latent structures (CPLS) model [Reference 2]. This model divides the output space into a predictable subspace, an unpredictable subspace, and a residual subspace, and similarly divides the input space into a prediction-related subspace, a process variable change subspace, and a residual space. CPLS projects process variables and quality variables simultaneously into five subspaces: the covariance subspace of the combined input and output, the output principal subspace, the output residual subspace, the input principal subspace, and the input residual subspace, which better explains fault pairs. CPLS only further decomposes process variables and quality variables from a global monitoring perspective, without using the local space information and dynamic information of the sample set, and ignores the impact of data uncertainty on the model accuracy.
[0005] In addition, in the process of industrial process quality forecasting and quality-related process monitoring modeling, general quality data is often scarcer than unlabeled data, and obtaining such quality data often requires manpower and material resources. Not only is its cycle much larger than that of unlabeled data, but the number of quality data samples is also very scarce [Reference 12]. Figure 1 represents the relationship between quality data and unlabeled data. In Figure 1 , the quality data acquisition cycle is N'T s , where T sLet \(\tau\) be the data sampling period. In actual industrial processes, general quality data is often scarcer than unlabeled data, and obtaining such quality data usually requires a significant amount of manpower and material resources. Not only is its period much larger than that of unlabeled data, but the number of quality data samples is also extremely scarce. During the quality prediction modeling process, considering the scarcity of quality data, people hope to fit the quality data as accurately as possible. Generally speaking, process data has obvious dynamics and locality. It is necessary to consider not only the dynamic correlation between quality data and its previous process data but also the spatial similarity between quality data and previous process data [References 1, 13 - 17]. The reason is that for the distance between quality samples and each historical sample in the sample space, it is assumed that the smaller the spatial distance between samples, the greater their correlation, and the sample weights are suitable for process mutation modeling. However, historical samples with a small spatial distance may be far from the query sample of the current operating state on the time axis. Therefore, historical samples with large weights do not necessarily represent the latest process state, ignoring the trends of samples and the dynamic relationship of the sample time series. Additionally, for some industrial processes with small fluctuations, adjacent samples have a strong temporal relationship. Using past adjacent samples to predict query samples should be of higher importance. Therefore, the temporal distance between samples can well consider the relationship between adjacent samples. In summary, simply using spatial distance or temporal distance cannot simultaneously describe the mutation and slow change properties of the process and the dynamics of the system. Only the weighted combination of temporal weights and spatial weights can solve the above problems.
[0006] Based on the above considerations, the present invention proposes a quality - related measurement method of adaptive weighted semi - supervised probabilistic concurrent latent structure projection (Adaptive weighted semi - supervised concurrent projection to latent space, AWS - CPLS). Summary of the Invention
[0007] The object of the present invention is to design a quality - related measurement method of adaptive weighted semi - supervised probabilistic CPLS for the deficiencies existing in the above - mentioned prior art.
[0008] To achieve the above object, the technical solution of the present invention is as follows:
[0009] A quality - related monitoring method of adaptive weighted semi - supervised probabilistic CPLS includes the following steps:
[0010] S1. Data pre - processing - Collect a dataset from the stable operating conditions of the industrial process and a quality index dataset Then perform normalization pre - processing on the quality data and input data;
[0011] S2. Initialization of model latent variables and model parameters - A probabilistic partial least squares (PPLS) model is established using normalized input and output data, and the total latent variables t of the model are determined. c Dimension q c And the corresponding sparse matrices are used as the initial values of matrices P and C; then the input data and data are reconstructed by the PPLS model, and the reconstruction errors ex of the output and input data are calculated. n' ey n' For the data set The private latent variables t in the input space are determined by using PCA to calculate 85% of the variance contribution rate. x Dimension q x The private latent variables t in the output space y Dimension q y And their corresponding loading matrices are used as the initial values of matrices Q and D respectively, and their reconstruction error covariance matrices are used as the initial values of Σ x And Σ y ;
[0012] S3. Initialization of the expectation-maximization (EM) algorithm. Set the number of iteration steps as t, the stopping threshold ε of the EM algorithm is 10 -5 , set the maximum number of iterations Maxiter for the algorithm to run, set the confidence level c of the control limit to determine the size of the control limit of the monitoring statistic, generally take 95%-99%, and the initial value of the algorithm iteration step is 0, that is, t = 0;
[0013] S4. Calculate the weights of the unlabeled process data samples;
[0014] S5. Increment the iteration step by 1, and use the E-step of the EM algorithm to calculate the expectations of the latent variables and the M-step to update the model parameters until the algorithm reaches the stopping threshold or the number of iterations reaches the maximum number of iterations;
[0015] S6. Calculate the covariance matrix of the latent variables;
[0016] S7. Calculate the control limits of the monitoring statistic;
[0017] S8. Calculate the monitoring statistic of the quality data;
[0018] S9. Calculate the monitoring statistic of the process data;
[0019] S10. Determine whether each monitoring statistic calculated in steps S8 and S9 satisfies and SPE ≤ SPE lim . If not satisfied, it is determined that the process is abnormal, and then the space where the abnormal variable is located is obtained.
[0020] The quality-related soft sensor method of the adaptive weighted semi-supervised probabilistic CPLS of the present invention is used for industrial process monitoring. This method includes data preprocessing, initialization of model latent variables and model parameters, calculation of the weights of unlabeled process data samples, the covariance matrix of latent variables, the control limits of fault monitoring statistics, the monitoring statistics of quality data, the monitoring statistics of process data, and judging whether a fault occurs by comparing the monitoring statistics with the control limits according to the above calculation results. The present invention provides a feasible way to solve the problem of variable-rate data modeling, which is easy to be extended to online learning and fault monitoring of continuous processes and batch processes, and can better improve the accuracy of the soft sensor model and the performance of fault monitoring.
[0021] The further refined technical solution of the present invention is as follows:
[0022] Preferably, in step S1, equations (39) and (40) are used to perform normalization preprocessing on quality data and input data.
[0023]
[0024]
[0025] where x n' is the data after x n normalization processing, mean(X) is the average value of all training data in the input space, std(X) is the standard deviation of all training data in the input space, y m' is the data after y m normalization processing, mean(Y') is the average value of all training data in the output space, std(Y') is the standard deviation of all training data in the output space.
[0026] Preferably, in step S2, the initial values of the means μ x and μ y of the process input data and quality output data in the model are both 0.
[0027] Preferably, in step S4, it is assumed that the sampling period of the labeled samples is N'T s , the sampling period of the unlabeled data is T s , there are N'-1 unlabeled samples x n(N'-1) , x n(N'-2) , …, x N'(n-1)+1 , and the quality samples are denoted as (x nN' , y nN' ). According to equations (5) and (6), the spatial weights and temporal weights of the unlabeled data samples are calculated.
[0028]
[0029]
[0030] wherein, is the spatial weight of the unlabeled data sample, T is the transpose of a matrix or vector, and σ s is an adjustable parameter for controlling the rate of change of the spatial weight with respect to the Euclidean distance, is the temporal weight of the unlabeled data sample, and σ t is an adjustable parameter for controlling the decay rate of the temporal weight with respect to the time of the unlabeled sample and the time of the distance-quality sample, i = 1, 2, …, N'-1, n = 1, 2, …, N L ;
[0031] Then, according to the quality data sample, the final weight of the unlabeled data sample is calculated by Equation (7)
[0032]
[0033] wherein, w nN'-i is the final weight of the unlabeled data sample, λ nN' is the weighting coefficient and ||y nN' -y (n-1)N' || represents the Euclidean distance between y nN' and y (n-1)N' and, y (n-1)N' is the output of the (n - 1)-th sample, and its sampling point position is (n - 1)N', and y nN' represents the output of the n-th sample, and σ w is an adjustable parameter for controlling the rate of change of the weighting coefficient with respect to the change in adjacent quality data.
[0034] Preferably, in step 5, the E-step of the EM algorithm calculates the expectation of the latent variable, and the specific process is as follows: The first-order and second-order statistics of the latent variables t x , t c , t y of the training data are calculated using Equations (15) to (18),
[0035]
[0036]
[0037]
[0038]
[0039] wherein, E(·) represents the expectation of a random variable, and the superscript T represents the transpose of a vector or matrix, Denote the m-th unlabeled sample, and respectively denote the private latent variable and the public latent variable, P and Q are respectively the coefficient matrices of and in the model, Θ old denotes the known model parameters, Σ x denotes the noise covariance matrix of the input space, C and D are respectively the coefficient matrices of the public latent variable and the private latent variable of the output space in the model private latent variable in the model, x n , y n are respectively the input and output of the n-th labeled sample, Σ y is the noise covariance matrix of the output space;
[0040] The specific process of updating the model parameters in the M-step of the EM algorithm is as follows: Use equations (19) to (26) to update the model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y ,
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049] In the formula, L'(X, Y, X u , T', Θ) is the complete likelihood function of the model, X and Y respectively represent the sets composed of the input data and output data of all labeled samples, X u represents the set composed of all unlabeled samples, and T' represents the set composed of all latent variables. The superscript new represents the updated value of the model parameters, N L represents the number of labeled samples, N u represents the number of unlabeled samples, μ x , μ y respectively represent the means related to the input data and output data in the model, wm Indicates the weight.
[0050] Preferably, in the step S6, based on the optimal model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y , according to the relationship between covariance and first-order and second-order statistics cov(t) = E(ttT)-E(t)E(t) T , use equations (27)–(28), (35)–(36) to calculate the covariance matrices cov(Τ y ), cov(Τ c ), and cov(Τ x ) of the latent variables of the training data,
[0051]
[0052]
[0053]
[0054]
[0055] In the formula, I represents the identity matrix, and the other symbols are the same as before. Obtained from the second-order statistics shown in equation (28), and the covariance calculation methods of other latent variables are similar.
[0056] Preferably, in the step S7, calculate the control limits of the monitoring statistics and SPE lim ,
[0057]
[0058]
[0059] In the formula, is the control limit of the statistic T 2 , SPE lim is the control limit of the statistic SPE, is the degree of freedom distribution under the given confidence level c, n d is the dimension of the variable corresponding to the T 2 statistic, n e is the dimension of the variable corresponding to the SPE statistic.
[0060] Preferably, in the step S8, for the new quality data pair (x new , y new ), first normalize the quality data pair using equations (39′)–(40′),
[0061]
[0062]
[0063] Then, calculate the process data x according to equations (30′) to (34′). new and the quality data y new for the latent variables and the reconstruction error monitoring statistics T 2 and SPE.
[0064]
[0065]
[0066]
[0067]
[0068]
[0069] In the formula, is the statistic T for the latent variable 2 value, is the monitoring statistic T for the latent variable 2 value. The remaining monitoring statistics have similar meanings. is the monitoring statistic value for the latent variable . SPE x is the value of the monitoring statistic SPE constructed from the reconstruction residuals of x new . is the reconstruction of the model with respect to x new . are the mathematical expectations of the common latent variable and the private latent variable of x new , y new corresponding to (x new ), respectively. cov(Τ c ) and cov(Τ x ) represent the covariance matrices of the latent variables of all labeled and unlabeled training samples, respectively. is the value of the monitoring statistic T 2 constructed for the output space. is the monitoring statistic T new constructed from the expectation of the private latent variable of the output y value. SPE 2 is the value of the monitoring statistic SPE constructed for the output space. SPE y is the value of the monitoring statistic SPE constructed from the output residuals. new,y is the value of the monitoring statistic SPE constructed from the output residuals. is y new reconstruction of is the output y new private latent variable expect to construct the monitoring statistic T 2 value, SPE new,y construct the monitoring statistic SPE value for the output residual is y new reconstruction of, cov(Τ y ) represents the covariance matrix of the outputs of all labeled training samples corresponding to the private latent variables.
[0070] Preferably, in the step S9, for the new process data x new , first normalize it using Equation (39′)
[0071]
[0072] Calculate the process data x according to Equations (30) to (32) new latent variable and the reconstruction error monitoring statistic T 2 and SPE
[0073]
[0074]
[0075]
[0076] Compared with the prior art, the present invention has the following advantages:
[0077] (1) Drawing on the principle of local weighting, adaptively weight the time-dynamic weighting coefficient and the space-similarity weighting coefficient of the local data set according to the quality change, make better use of the local spatio-temporal information of the sample set, and better consider the global and local information of the training samples;
[0078] (2) Introduce the probability graph theory, make full use of the information of unlabeled samples and labeled samples, solve the problem of scarce labeled samples, solve the problems of insufficient sample quantity and sample uncertainty caused by measurement noise, and propose a semi-supervised probabilistic concurrent latent structure projection method to improve the accuracy of the model;
[0079] (3) Construct monitoring statistics in the input space and the output space to realize the automatic fault detection of quality data and process data without quality indicators;
[0080] (4) This method provides a feasible way to solve the problem of variable-rate data modeling, is easy to be extended to online learning and fault monitoring of continuous processes and batch processes, and effectively improves the accuracy of soft measurement models and the performance of fault monitoring. Brief Description of the Drawings
[0081] The present invention will be further described below in conjunction with the drawings.
[0082] Figure 1 It is a sampling diagram of quality data in the present invention. Detailed Embodiments
[0083] The present invention will be described in detail below in conjunction with the embodiments and the accompanying drawings of the specification.
[0084] Embodiment 1 Adaptive Weighted Semi-Supervised Probabilistic Concurrent Latent Structure Projection
[0085] Based on standard PLS, CPLS will output divided into a prediction part, an unpredictable part, and a measurement noise part, and the input consists of an explanatory output part, a part of the independent variable variation of the unexplained output, and a measurement noise part. Therefore, it is necessary to introduce the probability of probabilistic concurrent latent structure projection (PCPLS) of two types of latent variables, and its model can be expressed as
[0086]
[0087] Wherein, and are both weight matrices; is the common latent variable of the input data and the output data, is the private latent variable of the input data that has nothing to do with the output data, is the latent variable of the non-predicted output; and are the mean vectors of the input data and the output data respectively; and are the input data noise and the output data noise respectively. Due to the randomness of the process data and the measurement of the key performance indicators, the latent variables and the noise variables in the PCPLS model are all random data and the noise is independent of each other. Here, it is assumed that these random variables all follow a Gaussian distribution, that is, t c ~N(0, I), t x ~N(0, I), t y ~N(0, I), ε y ~N(0, Σ y ), ε x ~N(0, Σ x ), I is the identity matrix of the corresponding dimension. Here, the variances of the process variable noises are different, that is and are both diagonal matrices, is the variance of the i-th variable in the input space, is the variance of the j-th variable in the output space, and D' and E' are the dimensions of the input and output spaces respectively.
[0088] Based on the Gaussian distribution assumption of the latent variables and noise variables in the probabilistic C-PLS model, the conditional probability distributions of the input data and the output data are
[0089] p(x|μ x ,t c ,t x ,P,Q,Σ x )=N(μ x +Pt c +Qt x ,Σ x ) (2)
[0090] p(y|μ y ,t c ,t y ,C,D,Σ y )=N(μ y +Ct c +Dt y ,Σ y ) (3)
[0091] Then, the marginal distributions of the labeled samples x and y can be calculated in the following form, that is
[0092]
[0093] To effectively improve the CPLS modeling accuracy, this embodiment draws on the local weighted principle, makes full use of the local space information and the sample time dynamic information, solves the problems of sample uncertainty, scarce labeled samples and a large number of unlabeled samples, and proposes an adaptive weighted semi-supervised probabilistic concurrent latent structure projection AWS-CPLS quality-related prediction method. For simplicity, assume that the sampling period of the labeled samples is N'T s , and the sampling period of the unlabeled data is T s , assume that there are N = N L N' labeled data and unlabeled data samples, then there are N L as the number of labeled samples, and the number of unlabeled samples is N'(N L - 1). Based on the dynamics and locality of the process data, the quality sample (x nN' ,y nN' ) has a high correlation with the N'-1 unlabeled samples x n(N'-1) ,x n(N'-2) ,…,x N'(n-1)+1 in front of it. Making full use of these unlabeled samples is beneficial to improving the quality sample (x nN' ,y nN') Modeling accuracy. Assume that the value of N' is large enough. Then, the quality sample (x nN' , y nN' ) is greatly affected by the unlabeled samples in front of it, and little affected by other samples. Here, n = 1, 2, …, N . First, use Equation (5) to calculate the spatial distance weight between the quality sample (x L , y nN' , y nN' ) and the samples in the sample set as
[0094]
[0095] where is the spatial weight between the quality sample (x nN' , y nN' ) and the samples in the sample set . represents the Euclidean distance between x nN' and x nN'-i . T is the transpose of a vector or matrix. σ s is an adjustable parameter used to control the rate of change of the spatial weight with the Euclidean distance. i = 1, 2, …, N' - 1, n = 1, 2, …, N L .
[0096] Generally speaking, the quality data (x nN' , y nN' ) is more affected by the samples adjacent in time and less affected by the samples with a longer time interval. Therefore, the time weight of the samples in the sample set is calculated by the following method
[0097]
[0098] where is the time weight of the samples in the sample set . σ t is an adjustable parameter used to control the decay rate of the time weight with the time of the unlabeled sample and the distance from the time of the quality sample. i = 1, 2, …, N' - 1, n = 1, 2, …, N L .
[0099] Note that using spatial distance information can help model mutation quality data, and time distance information can help model smoothly changing quality data. Combine the spatial weight and the time weight to construct a sample weight that can reflect the local dynamic information and spatial information of the data, that is
[0100]
[0101] where w nN'-iis the final weight of the sample, λ nN'-i is the weighting coefficient and is used to adjust the relative importance of the time weight and the space weight; ||y nN' -y (n-1)N' || represents the Euclidean distance between y nN' and y (n-1)N' and, and y (n-1)N' is the output of the (n - 1)-th sample, and its sampling point position is (n - 1)N', y nN' represents the output of the n-th sample, σ w is an adjustable parameter used to control the rate at which the weighting coefficient changes with adjacent quality data. Obviously, if the adjacent quality values change greatly, more spatial information is required to implement model modeling, that is, it is required that the space weight plays a dominant role in the overall sample weight; otherwise, it is required that the time weight plays a dominant role in the overall sample weight. In this way, the sample weight will adaptively adjust the ratio of the space weight and the time weight according to the severity of the quality change, effectively improving the model accuracy.
[0102] Suppose the process data set collected in sequence from the industrial site and the quality index data set Among them, the subscripts of the process data and the quality index data represent the sampling time. Calculate the weight w' of each unlabeled sample according to formulas (5) to (7) n' , n' = 1, 2, … N' - 1, N' + 1, …, 2N' - 1, 2N' + 1, …, N - 1. Then the collected sample data set is re-divided into a labeled sample set and an unlabeled sample set Let the weight of the unlabeled data be w m . The model parameters Θ = {μ x , μ y , P, Q, C, D, Σ x , Σ y} of the adaptive weighted semi-supervised probabilistic C-PLS can be obtained by maximizing the log-likelihood function in the following form, that is
[0103]
[0104] To maximize the above-formulated likelihood function, similar to the solution method of the PPCR optimization problem, the EM algorithm is used to perform iterative optimization on it to obtain the optimal solution of the model parameters. To use the EM algorithm, first calculate the complete likelihood function including the latent variable and the complete likelihood function including the latent variable, labeled data, and unlabeled data,
[0105] L'(X, Y, X u, T', Θ) = L(X, Y, T, Θ) + L(X u , T u , Θ) (9)
[0106] where T' = {T, T u}, T u is the set of model latent variables corresponding to all unlabeled sample sets X u , and T is the set of latent variables corresponding to all labeled samples. The expectation of the latent variables in Equation (9), that is
[0107]
[0108] where
[0109]
[0110]
[0111] Here are respectively the common latent variable corresponding to the labeled sample (x n , y n ), the private latent variable of x n , and the private latent variable corresponding to y n . are respectively the common latent variable and the private latent variable of the unlabeled sample . N u is the number of unlabeled samples, and Θ old is the known model parameter. The EM algorithm is an iterative optimization algorithm that includes the expectation E-step and the maximization M-step, and can theoretically guarantee the convergence of the EM algorithm. When the EM algorithm converges iteratively, the parameter values of the adaptive weighted semi-supervised probability C-PLS model reach the optimal.
[0112] Note that in the E-step of the EM algorithm, Θ old is the determined parameter obtained in the previous step, so that the expected value of the complete likelihood function can be obtained. According to the Bayesian formula, noting that the process variable distributions and are constants, then the posterior distribution of the latent variable is
[0113]
[0114]
[0115] Since both terms on the right side of the formula are Gaussian distributions, the posterior distribution of the latent variable must also be a Gaussian distribution. Therefore, its first-order and second-order statistics are calculated as follows:
[0116]
[0117]
[0118]
[0119]
[0120] In the formula, E(·) represents the mathematical expectation of a random variable, and the superscript T represents the transpose of a vector or matrix. represents the m-th unlabeled sample. and respectively represent 's private latent variable and public latent variable. P and Q are the coefficient matrices of and in the model respectively. Θ old represents the known model parameters, and Σ x represents the noise covariance matrix of the input space. C and D are the coefficient matrices of the public latent variable and the private latent variable in the output space of the model respectively. x n , y n are the input and output of the n-th labeled sample respectively. Σ y is the noise covariance matrix of the output space.
[0121] In the M-step of the EM algorithm, the optimal solution of the model parameters is obtained by maximizing the likelihood function. This step is achieved by calculating the partial derivative of E[L(X,Y,T',Θ)] with respect to Θ and setting it equal to 0, that is
[0122] Let the first-order and second-order statistics of the latent variables obtained from the labeled dataset be denoted as For and , the labeling methods of the first-order and second-order statistics are similar. The first-order and second-order statistics of the latent variables obtained from the unlabeled dataset are denoted as, For the remaining latent variables , the notations of the first-order and second-order statistics are similar. Then, for the M-step of the EM algorithm for all data, the method for obtaining the optimal model parameter values is in the following form
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131] wherein, L'(X, Y, X u , T', Θ) is the complete likelihood function of the model, X and Y respectively represent the sets composed of the input data and output data of all labeled samples, X u represents the set composed of all unlabeled samples, and T' represents the set of all latent variables. The superscript new represents the updated value of the model parameters, N L represents the number of labeled samples, and N u represents the number of unlabeled samples, μ x and μ y respectively represent the means of the input data and output data in the model, w m represents the weight, δE[L'(X, Y, X u , T', Θ)] / δQ represents the derivative of the function E[L'(X, Y, X u , T', Θ)] with respect to Q, and E[L'(X, Y, X u , T', Θ)] represents the mathematical expectation of the complete likelihood function L'(X, Y, X u , T', Θ) with respect to the latent variable T'.
[0132] Example 2 Quality-related Monitoring Method Based on AWS-CPLS
[0133] For the adaptive weighted semi-supervised probabilistic C-PLS model, given a new quality data pair (x new , y new ), first calculate the first-order and second-order statistics of the latent variables corresponding to the input data and output data according to the optimal model parameters Θ * :
[0134]
[0135]
[0136] wherein, μ x and μ y are respectively the mean vectors of the process input and output. The output corresponding to the obtained latent variable estimate is
[0137]
[0138] Construct the monitoring statistics T 2 and SPE in the input space as
[0139]
[0140]
[0141]
[0142] Among them, cov(Τ c ) and cov(Τ x ) represent the covariance matrices of the latent variables of all the labeled and unlabeled training samples. The monitoring statistics T 2 and SPE are
[0143]
[0144]
[0145] Among them cov(Τ y ) represents the covariance matrix of the latent variables of all the labeled training samples.
[0146] If there is only new input data then the corresponding latent variables are
[0147]
[0148]
[0149] The monitoring statistics of its input space, such as T 2 and SPE, are calculated as in Eqs. (29) to (31). P and Q are the coefficient matrices of the latent variables and respectively, Σ x , μ x are the noise covariance of the input space and the mean of the input data respectively, T is the transpose of a matrix or vector, are the expectations of the public and private latent variables respectively, and I is the identity matrix of the corresponding dimension. Obtained from the second-order statistics shown in Eq. (28), and the covariance calculation methods of other latent variables are similar.
[0150] T 2 and the two monitoring statistics of SPE are both constructed by the Mahalanobis distance, and the χ 2 distribution with the corresponding degrees of freedom can be used for estimation. Given the confidence level c, and the control limits of the statistics T 2 and SPE are as follows:
[0151]
[0152]
[0153] where n d and n e correspond to the dimensions of the variables corresponding to the statistics respectively.
[0154] A feasible idea is to use probabilistic partial least squares (PPLS) and principal component analysis (PPCA) to determine the appropriate latent variable dimensions in the input space and output space. First, use the probabilistic PLS model to determine the total number of latent variable dimensions of the model by the cross-validation method; then, based on the reconstruction errors of the input data and output data, use PCA to calculate the variance contribution rate to determine the appropriate private latent variable dimensions of the input space and output space. In this embodiment, the contribution rate is taken as 85%.
[0155] In summary, a quality-related monitoring method for adaptive weighted semi-supervised probabilistic CPLS is as follows:
[0156] S1. Data preprocessing - Collect the data set from the stable working conditions of the industrial process and the quality index data set Then, use equations (39) and (40) to perform normalization preprocessing on the quality data and input data.
[0157]
[0158]
[0159] where x n' is the data after x n normalization processing, mean(X) is the average value of all training data in the input space, std(X) is the standard deviation of all training data in the input space, y m' is the data after y m normalization processing, mean(Y') is the average value of all training data in the output space, and std(Y') is the standard deviation of all training data in the output space.
[0160] S2. Initialization of model latent variables and model parameters - Use the normalized input data and output data to establish a probabilistic partial least squares (PPLS) model and determine the total number of latent variables t c dimension q c of the model and the corresponding sparse matrices as the initial values of matrices P and C; then, reconstruct the input data and data by the PPLS model, and calculate the reconstruction errors of the output and input data ex n' , ey n' . Respectively, for the data set use PCA to calculate 85% of the variance contribution rate to determine the private latent variable t x dimension q x, output the spatially private latent variable t y Dimension q y , and their corresponding loading matrices are used as the initial values of matrices Q and D respectively, and their reconstruction error covariance matrix is used as Σ x and Σ y The initial values of the means μ x and μ y of the process input data and quality output data in the model are both 0.
[0161] S3. Set the iteration step t to 0, the expected maximum EM algorithm stop threshold ε = 10 -5 , the maximum number of iterations Maxiter for the algorithm to run, and the confidence level c of the control limit. That is:
[0162] Iteration step t = 0, algorithm stop threshold ε = 10 -5 , maximum number of iterations Maxiter, confidence level c.
[0163] S4. Calculate the weights of the unlabeled process data samples - assume the sampling period of the labeled samples is N'T s , the sampling period of the unlabeled data is T s , there are N' - 1 unlabeled samples x n(N'-1) , x n(N'-2) , …, x N'(n-1)+1 , and denote the quality samples as (x nN' , y nN' ). Calculate the spatial weight and temporal weight of the unlabeled data samples according to equations (5) and (6),
[0164]
[0165]
[0166] In the formula, is the spatial weight of the unlabeled data sample, T is the transpose of a matrix or vector, and σ s is an adjustable parameter used to control the rate of change of the spatial weight with the Euclidean distance, is the temporal weight of the unlabeled data sample, and σ t is an adjustable parameter used to control the decay rate of the temporal weight with the time of the unlabeled sample and the time distance from the quality sample time, i = 1, 2, …, N' - 1, n = 1, 2, …, N L ;
[0167] Then, calculate the final weight of the unlabeled data sample according to the quality data sample by equation (7)
[0168]
[0169] where w nN'-i is the final weight of the unlabeled data sample, λ nN'-i is the weighting coefficient and represents the Euclidean distance between y nN' and y (n-1)N' , y (n-1)N' is the output of the (n - 1)-th labeled sample, and σ w is an adjustable parameter used to control the rate at which the weighting coefficient changes with adjacent quality data.
[0170] S5. Increment the iteration step by 1, and use the E-step of the EM algorithm to calculate the expectation of the latent variable and the M-step to update the model parameters until the algorithm reaches the algorithm stop threshold or the number of iterations reaches the maximum number of iterations. That is:
[0171] repeat
[0172] t = t + 1
[0173] E-step of the EM algorithm (calculate the expectation of the latent variable): Use equations (15) to (18) to calculate the first- and second-order statistics of the latent variables t x , t c , t y of the training data.
[0174] M-step of the EM algorithm (update the model parameters): Use equations (19) to (26) to update the model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y .
[0175] until |L(X, Y, X u , Θ (t) ) - L(X, Y, X u , Θ (t-1) )| ≤ ε or t ≤ Maxiter.
[0176] The E-step of the EM algorithm calculates the expectation of the latent variable. The specific process is as follows: Use equations (15) to (18) to calculate the first- and second-order statistics of the latent variables t x , t c , t y of the training data,
[0177]
[0178]
[0179]
[0180]
[0181] In the formula, E(·) represents the expectation of a random variable, and the superscript T represents the transpose of a vector or matrix. represents the m-th unlabeled sample. and represent respectively the private latent variable and the public latent variable of and respectively. P and Q are the coefficient matrices of old and x in the model. Θ the private latent variable in the output space of the model. x n , y n are the input and output of the n-th labeled sample respectively. Σ y is the noise covariance matrix of the output space.
[0182] The specific process of updating the model parameters in the M-step of the EM algorithm is as follows: Use equations (19) to (26) to update the model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y ,
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191] In the formula, L'(X, Y, X u , T', Θ) is the complete likelihood function of the model. X and Y represent the sets composed of the input data and output data of all labeled samples respectively. X u represents the set composed of all unlabeled samples. T' represents the set composed of all latent variables. The superscript new represents the updated value of the model parameters. N L represents the number of labeled samples. Nu denotes the number of unlabeled samples, μ x , μ y respectively denote the means of the input data and the output data in the model, w m denotes the weight.
[0192] S6. Calculate the covariance matrix of the latent variables — Based on the optimal model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y , according to the relationship between covariance and first - order and second - order statistics cov(t) = E(tt T ) - E(t)E(t) T , use equations (27) - (28), (35) - (36) to calculate the covariance matrices cov(Τ y ), cov(Τ c ) and cov(Τ x ) of the latent variables of the training data,
[0193]
[0194]
[0195]
[0196]
[0197] In the formula, I denotes the identity matrix, and the other symbols are the same as before.
[0198] S7. Calculate the control limits of the monitoring statistics — Calculate the control limits of the monitoring statistics according to (37) - (38) and SPE lim ,
[0199]
[0200]
[0201] In the formula, is the control limit of the statistic T 2 , SPE lim is the control limit of the statistic SPE, is the degree - of - freedom distribution under the given confidence level c, n d is the dimension of the variables corresponding to the T 2 statistic, n e is the dimension of the variables corresponding to the SPE statistic.
[0202] S8. Calculate the monitoring statistic of quality data — For the new quality data pair (x new , y new ), first normalize the quality data pair using equations (39′) to (40′).
[0203]
[0204]
[0205] Then, calculate the monitoring statistics T new and SPE of the latent variables and reconstruction errors of process data x new and quality data y 2 according to equations (30′) to (34′).
[0206]
[0207]
[0208]
[0209]
[0210]
[0211] In the formula, is the value of the statistic T of the latent variable 2 . is the monitoring statistic of the latent variable The remaining monitoring statistics have similar meanings. are the mathematical expectations of the common latent variable corresponding to (x new , y new ) and the private latent variable of x new respectively. cov(Τ c ) and cov(Τ x ) represent the covariance matrices of the latent variables of all labeled and unlabeled training samples respectively. is the value of the monitoring statistic T new constructed from the expectation of the private latent variable of the output y 2 . SPE new,y is the value of the monitoring statistic SPE constructed from the output residual. is the reconstruction of y new . cov(Τ y ) represents the covariance matrix of the private latent variables corresponding to the outputs of all labeled training samples.
[0212] S9. Monitoring Statistic of Process Data Calculation — For new process data x new , first, normalize it using Equation (39).
[0213]
[0214] Calculate process data x according to Equations (30′) to (32′). new Monitoring Statistic T of Latent Variable and Reconstruction Error 2 and SPE
[0215]
[0216]
[0217]
[0218] S10. Judging Whether a Fault Occurs — Judge whether each monitoring statistic calculated in Steps S8 and S9 satisfies and SPE ≤ SPE lim , if not satisfied, judge that the process is abnormal, and then obtain the subspace where the abnormal variable is located, so as to obtain information on whether the fault is related to quality.
[0219] The implementation process of the above method includes two stages: offline modeling and online monitoring. Among them, Steps S1 to S7 establish a monitoring model offline, and Steps 8 and 9 use quality data pairs and process data pairs to monitor the process online respectively.
[0220] References
[0221] 1. Bei Wang, Zhichao Li, Zhenwen Dai, Neil Lawrence, Xuefeng Yan. Data-driven mode identification and unsupervised fault detection for nonlinearmultimode processes. IEEE Transactions on Industrial informatics, 16(6):3651 - 3660, 2020
[0222] 2. Qin S J, Zheng Y. Quality-relevant and process-relevant fault monitoring with concurrent projection to latent structures. AIChE Journal, 59(2):496-504, 2013.
[0223] 3. Xiaofeng Yuan, Zhiqiang Ge, Biao Huang, Zhihuan Song. A Probabilistic Just-in-Time Learning Framework for Soft Sensor Development With Missing Data. IEEE TRANSACTIONS ON CONTROL SYSTEMS TECHNOLOGY, 25(3), 2017.
[0224] 4. Kong Xiangyu, Cao Zehao, Du Baiyang, Luo Jiayu. Quality-related multimodal fault detection technology based on partial least squares. Control and Decision, 34(12):2547-2557, 2019.
[0225] 5. Kaixiang Peng, Kai Zhang, Bo You, Jie Dong, Zidong Wang. A Quality-based nonlinear fault diagnosis framework focusing on industrial multimode batch processes. IEEE Transactions on Industrial electronics, 63(4):2615-22623, 2019.
[0226] 6. Zheng J, Song Z, Ge Z. Probabilistic learning of partial least squares regression model: Theory and industrial applications. Chemometrics & Intelligent Laboratory Systems, 158:80-90, 2016.
[0227] 7. Weiming Shao, Zhiqiang Ge, Zhihuan Song. Quality variable prediction for chemical process based on semi-supervised Dirichlet process mixture Gaussians. Chemical engineering science, 193:394 - 410, 2019.
[0228] 8. Reza Sharifi, Reza Langari. Nonlinear sensor fault diagnosis using mixture of probabilistic PCA models. Mechanical Systems and Signal Processing Volume 85:638 - 650, 2017.
[0229] 9. Jinlin Zhu, Zhiqiang Ge, Zhihuan Song. Distributed Gaussian mixture model for monitoring plant-wide processes with multiple operating modes. Journal of Systems and Control, 6:1 - 15, 2018.
[0230] 10. Zhou D, Li G, Qin S J. Total projection to latent structures for process monitoring[J]. AIChE Journal, 56(1):168 - 178, 2010.
[0231] 11. Zhao C, Sun Y. The multi - space generalization of total projection to latent structures(MsT - PLS) and its appli - cation to online process monitoring[C]. The 10th Conf on Int Control and Automation. Hangzhou: IEEE, 2013:1441 - 1446.
[0232] 12. Shao Weiming, Ge Zhiqiang, Li Hao, Song Zhihuan. Semi-supervised dynamic soft sensor modeling method based on recurrent neural network. 33(11):2277 - 13, 2019.
[0233] 13. Xiaofeng Yuan, Jiao Zhou, Yalin Wang. A spatial-temporal LWPLS for adaptive soft sensor modeling and its application for an industrial hydrocracking process. Chemometrics and intelligent laboratory systems, 197:103921, 2020.
[0234] 14. Jingxiang Liu, Tao Liu, Junghui Chen. Sequential local-based Gaussian mixture model for monitoring multiphase batch processes. Chemical Engineering Science Volume, 181:101 - 113, 2018.
[0235] 15. Liu Jie, Valeria Vitelli, Enrico Zio, Redouane Seraoui. A novel dynamic-weighted probabilistic support vector regression-based ensemble for prognostics of time series data. IEEE transactions on Reliability, 64(4):1203 - 1214, 2015.
[0236] 16. Li Yuan, Ma Yuhan, Zhang Cheng, Feng Liwei. Fault detection of multimodal batch processes based on locally proximal standardized partial least squares. Control Theory & Applications. 37(5): 1109 - 1117, 2019.
[0237] 17. Junhua Zheng, Jinlin Zhu, Guangjie Chen, Zhihuan Song, Zhiqiang Ge. Dynamic Bayesian network for robust latent variable modeling and fault classification. Engineering Applications of Artificial Intelligence, 8, 2020.
Claims
1. An adaptive weighted semi-supervised probabilistic CPLS quality-related monitoring method, characterized in that, Including the following steps: S1. Data preprocessing - Collecting datasets from the stable operating conditions of industrial processes and the quality index datasets Then, perform normalization preprocessing on the quality data and input data; S2, Model Latent Variables and Model Parameter Initialization - A Probabilistic Partial Least Squares (PPLS) model is established using normalized input and output data, and the total number of latent variables t of the model is determined c Dimension q c And the corresponding sparse matrices are used as the initial values of matrices P and C; then the PPLS model is used to reconstruct the input data and the data, and the reconstruction errors of the output and input data, ex n' and ey n' are calculated respectively for the dataset The private latent variable t of the input space is determined by using PCA to calculate 85% of the variance contribution rate x Dimension q x and the private latent variable t of the output space y Dimension q y The corresponding loading matrices are used as the initial values of matrices Q and D respectively, and the reconstruction error covariance matrix is used as Σ x and Σ y as their initial values; S3. Set the number of iteration steps as t, and the stopping threshold ε of the expectation-maximization EM algorithm is 10 -5 , set the maximum number of iterations Maxiter for the algorithm to run, set the confidence level c of the control limit, and the initial value of the algorithm iteration step is 0, that is, t = 0; S4. Calculate the weights of the unlabeled process data samples; Assume that the sampling period of the labeled samples is \(N'T\). s , and the sampling period of the unlabeled data is \(T\). s , there are \(N' - 1\) unlabeled samples \(x\). n(N'-1) , \(x\). n(N'-2) , …, \(x\). N'(n-1)+1 , and denote the quality samples as \((x\). nN' , \(y\). nN' ). Calculate the spatial weight and temporal weight of the unlabeled data samples according to Equations (5) and (6). In the formula, is the spatial weight of the unlabeled data sample, The superscript T is the transpose of a vector or matrix, and σ s is an adjustable parameter used to control the rate of change of the spatial weight with respect to the Euclidean distance, is the temporal weight of the unlabeled data sample, and σ t is an adjustable parameter used to control the decay rate of the temporal weight with respect to the time of the unlabeled sample and the time of the distance-quality sample, where i = 1, 2, …, N' - 1 and n = 1, 2, …, N L ; Then, calculate the final weights of the unlabeled data samples from the quality data samples according to Equation (7). where w nN'-i is the final weight of the unlabeled data sample, λ nN' is the weighting coefficient and ||y nN' -y (n -1 )N '|| represents the Euclidean distance between y nN' and y (n-1)N' and y (n-1)N' is the output of the (n - 1)-th sample, and its sampling point position is (n - 1)N', y nN' represents the output of the n-th sample, and σ w is an adjustable parameter used to control the rate at which the weighting coefficient changes with adjacent quality data; S5. Increment the iteration step by 1, and use the E-step of the EM algorithm to calculate the expectations of the latent variables and the M-step to update the model parameters until the algorithm reaches the algorithm stop threshold or the number of iterations reaches the maximum number of iterations; S6. Calculate the covariance matrix of the latent variables; S7. Calculate the control limits of the measurement statistics; S8. Calculate the monitoring statistics of the quality data; S9. Calculate the monitoring statistics of the process data; S10. Determine whether each monitoring statistic calculated in steps S8 and S9 satisfies and SPE ≤ SPE lim . If not satisfied, it is determined that an abnormality has occurred in the process, and then the space where the abnormal variable is located is obtained.
2. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 1, wherein In step S1, preprocess the quality data and the input data using Equations (39) and (40). where x n' is the data after x n normalization, mean(X) is the mean of all training data in the input space, std(X) is the standard deviation of all training data in the input space, and y m' is the data after y m normalization, mean(Y') is the mean of all training data in the output space, and std(Y') is the standard deviation of all training data in the output space.
3. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 2, characterized in that In the step S2, the initial values of the mean μ of the process input data and the quality output data in the model x and μ y are both 0.
4. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 3, wherein In the said step 5, the E-step of the EM algorithm calculates the expectations of the latent variables, and the specific process is as follows: Use equations (15) to (18) to calculate the latent variables t of the training data respectively x , t c , t y for the first-order and second-order statistics where \(E(\cdot)\) represents the expectation of a random variable, the superscript \(T\) represents the transpose of a vector or matrix, denotes the \(m\)-th unlabeled sample, and denote respectively the private latent variable and the public latent variable of and are the coefficient matrices of old in the model, \(\Theta\) x represents the known model parameters, \(\Sigma\) represents the noise covariance matrix of the input space, \(C\) and \(D\) are the coefficient matrices of the public latent variable and the private latent variable n in the output space of the model, \(x\) n and \(y\) y are respectively the input and output of the \(n\)-th labeled sample, \(\Sigma\) The specific process of updating the model parameters in the M-step of the EM algorithm is as follows: Use equations (19) to (26) to update the model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y , where \(L'(X, Y, X u , T', \Theta)\) is the complete likelihood function of the model, \(X\) and \(Y\) respectively represent the sets composed of the input data and output data of all labeled samples, \(X u represents the set composed of all unlabeled samples, \(T'\) represents the set of all latent variables, the superscript new represents the updated value of the model parameters, \(N L represents the number of labeled samples, \(N u represents the number of unlabeled samples, \(\mu x \), \(\mu y respectively represent the means related to the input data and output data in the model, \(w m represents the weight.
5. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 4, wherein In the step S6, based on the optimal model parameters Q, D, P, C, μ x , μ y , Σ x , Σ y , according to the relationship between covariance and first- and second-order statistics cov(t) = E(tt T ) - E(t)E(t) T , use equations (27)-(28), (35)-(36) to calculate the covariance matrices cov(Τ y ), cov(Τ c ) and cov(Τ x ) of the hidden variables of the training data. where I represents the identity matrix, 6. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 5, characterized in that In the step S7, the control limits of the monitoring statistic are calculated according to (37) to (38). and SPE lim , where is the control limit of the statistic T 2 , SPE lim is the control limit of the statistic SPE, is the degree of freedom distribution under the given confidence level c, n d is the dimension of the variable corresponding to the T 2 statistic, n e is the dimension of the variable corresponding to the SPE statistic.
7. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 6, characterized in that, In the said step S8, for the new quality data pair (x new , y new ), first, normalize the quality data pair by using equations (39′) to (40′). Then, calculate the process data x according to formulas (30′) to (34′) new and the quality data y new for the latent variable and the reconstruction error monitoring statistics T 2 and SPE In the formula, is the latent variable statistic T 2 value, is the latent variable monitoring statistic are respectively the mathematical expectations of the common latent variable corresponding to (x new , y new ) and the private latent variable of x new . cov(Τ c ) and cov(Τ x ) respectively represent the covariance matrices of the latent variables of all labeled and unlabeled training samples. is the output y new private latent variable expectation to construct the monitoring statistic T 2 value, SPE new,y is the value of the monitoring statistic SPE constructed by the output residual. is the reconstruction of y new . cov(Τ y ) represents the covariance matrix of the output corresponding private latent variables of all labeled training samples.
8. The quality-related monitoring method of an adaptive weighted semi-supervised probabilistic CPLS according to claim 7, wherein In the step S9, for the new process data x new , first, normalize it using the formula (39′). Calculate the process data x according to equations (30') to (32') new Latent variable and reconstruction error monitoring statistic T 2 and SPE
Citation Information
Patent Citations
Method for detecting industrial process faults basing on multi-sampling rate factor analysis model
CN109085805A
Competitive adaptive reweighting key data extraction method for Raman spectrum analysis of insulating oil
CN110567937A