Deep extension VAE soft measurement method based on instant learning
By introducing a deep extended structure and an instant learning framework into the VAE network, the problem of insufficient reconstruction error accumulation and time-variability adaptation in the multi-layer VAE network is solved, and the prediction accuracy and adaptability of the soft measurement model are improved.
Patent Information
- Application Number
- CN202510137370.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing soft measurement technology based on multi-layer VAE networks has the accumulation of reconstruction errors and insufficient adaptation to time-variability of industrial processes, resulting in a degradation of predictive performance.
A deep-scaling VAE soft measurement method based on real-time learning is proposed. By constructing a deep-scaling DE-VAE network model, multiplexing prior information to enhance feature extraction, and introducing an instant learning framework in the online prediction stage, and using similar data to update model parameters.
The problems of reconstruction error accumulation and insufficient adaptation of time-variability are effectively overcome, the prediction accuracy of the deep-scaling variational autoencoder model is improved, and the adaptability to industrial processes is enhanced.
Smart Images

Figure CN120108537A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a deep extended VAE soft measurement method based on instant learning, belonging to the technical field of industrial process control. Background Art
[0002] Soft sensing methods use mathematical modeling and data processing technology to achieve real-time prediction of difficult-to-measure key variables based on process variables that are easy to measure during the production process. They can be divided into two categories: mechanism modeling methods and data-driven modeling methods. Compared with mechanism modeling methods, data-driven modeling methods rely less on the mechanism knowledge of the process and can be flexibly adjusted and adapted according to different production environments and data characteristics, thereby capturing complex nonlinear relationships and dynamic changes. Therefore, they have been increasingly widely used in actual industrial processes.
[0003] As an unsupervised deep learning network, autoencoder (AE) has good nonlinear feature extraction ability and scalability. However, the soft sensor modeling method based on AE is easily affected by measurement noise, which affects the prediction performance of the model. To this end, a variational autoencoder (VAE) is proposed. By introducing the concept of Bayesian variational inference between the encoder and the decoder, it can better capture the potential structure of the data when generating samples. VAE can also further enhance the feature extraction and generation capabilities of the model by stacking multiple encoding layers and decoding layers. Shen et al. proposed a nonlinear probabilistic latent variable regression model (NPLVR) based on variational autoencoder in "Nonlinear probabilistic latent variable regression models for softsensor application: From shallow to deep structure. Control Engineering Practice 94 (2020): 104198.", and expanded the NPLVR model from a shallow structure to a deep structure by stacking, thereby extracting deeper nonlinear features for soft sensor modeling. Chinese patent publication number CN118444641A discloses a soft measurement method based on weighted gated stacked quality supervised VAE, which improves the prediction accuracy of the soft measurement model by introducing quality information and a feature integration mechanism of gated linear units.
[0004] However, the above model does not take into account that during the multi-layer VAE pre-training process, as the number of layers increases, the reconstruction error may accumulate, thereby affecting the model's feature extraction ability, resulting in the inability to effectively map the correlation between input and output variables, and reducing the model's prediction accuracy. In addition, the time-varying nature of process data may also lead to a decrease in the model's prediction accuracy. Summary of the invention
[0005] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a deep extended VAE soft measurement method based on just-in-time learning. By constructing a deep extended DE-VAE network model, the correlation between the extracted features and the original input and output is enhanced by reusing prior information; then the idea of just-in-time learning is introduced to propose a deep extended variational autoencoder (just-in-time learning-Deep extended-Variational Autoencoder, JITL-DE-VAE) based on just-in-time learning for the prediction of butane concentration at the bottom of the debutanizer. In the pre-training stage, the deep extended structure is used to fully extract effective features to complete the training of the basic model; in the fine-tuning stage, based on the idea of just-in-time learning, similar data is used to update the model parameters to improve the prediction accuracy of the deep extended variational autoencoder model; so as to overcome the problems of reconstruction error accumulation and poor prediction performance caused by insufficient adaptation of the model to the time-varying nature of the industrial process in the existing soft measurement technology based on the multi-layer VAE network.
[0006] A deep extended VAE soft sensing method based on instant learning, comprising:
[0007] Step 1: Select the input variables of the model based on theoretical analysis, collect data, including input variables and quality variables, and standardize the data;
[0008] The labeled training dataset is Where x is the input variable, y is the quality variable, N l is the number of samples; the standardized set of input variables in the training data set is X s , the sample to be predicted is X new ;
[0009] Step 2: A single-layer VAE extracts high-level features of the data. To ensure the relevance of the extracted features to the quality variables, the prediction error of the quality variables is added to the reconstruction error of the VAE. In this way, the features extracted at each layer are not only feature representations of the input variables, but also effective explanations of the quality variables.
[0010] Step 3: Build a deep extended DE-VAE network model, and use the hidden layer and input variables of the previous E-VAE as the input of the next E-VAE, making full use of the effective information of the features and preventing excessive loss of effective information;
[0011] Step 4: Place the X s As the input of the DE-VAE model, calculate the reconstruction value of the input sample and the predicted target value for pre-training;
[0012] 4.1: Xs As the input of DE-VAE encoder, the sample is calculated to obey the distribution of features in the latent space The mean and variance [μ,σ 2 ];
[0013] 4.2: Using the reparameterization technique, according to the mean and variance [μ,σ 2 ] Sampling obtains the latent variable z;
[0014] 4.3: Use the hidden variable z as the input of the decoder to obtain the reconstructed value of the sample And the predicted value of the sample
[0015] Step 5: Repeat step 4 to calculate the loss function LOSS of each pre-training layer, and use the Adam optimization method to update the model parameters;
[0016] Step 6: After pre-training is completed, use the hidden variable z to calculate the linear mapping y of the original output k , and get the final output form Fine-tune the global by minimizing the loss function minΓ; save network parameters for online prediction;
[0017] Step 7: Introduce a real-time learning framework in the online prediction stage. When a test sample appears, a small batch of samples that are most similar to the test sample will be searched in the historical data (including the training set and the predicted test sample). This batch of data is then used to update the DE-VAE model parameters to obtain the updated JITL-DE-VAE model;
[0018] Step 8: The sample to be predicted X new After the same standardization, the data is input into the updated JITL-DE-VAE model to obtain the final prediction value.
[0019] The DE-VAE network in step 3 is different from the traditional AE. VAE has both the data mining and nonlinear modeling capabilities of deep learning, and can also model the uncertainty of the process and data noise like a probabilistic model, which can greatly improve the model's ability in probabilistic data description and feature extraction. Assume that the input variable x and the quality variable y are generated by random continuous hidden variables z.
[0020]
[0021] Assuming that x and y are conditionally independent of each other, the latent variable z is sampled from the latent space and is passed through two parameters δ x and δ y The neural network is used to obtain x and y. According to the above generation process, the joint probability distribution of the generation model is:
[0022] p δ (x,y,z)=p δx (x|z)p δy (y|z)p(z)
[0023] The log-likelihood function of the marginal probability distribution p(x,y) of the data sample point is:
[0024]
[0025] in To infer the model, an additional variational posterior is used to approximate the true complex posterior probability p(z|x).
[0026] According to Jensen inequality:
[0027]
[0028] in, is the lower bound of the marginal probability likelihood function. In order to calculate and δ, maximizing the evidence lower bound:
[0029]
[0030] The lower bound of evidence can be understood in two parts. The first two items are related to the generative model p of x and y. δx (x|z), p δy (y|z) is related to z Subordinate to the approximate variational posterior Assumptions It follows a normal distribution.
[0031] Since the generative models all obey the normal distribution, logp δx (x|z) and logp δy (y|z) can be expressed as:
[0032]
[0033] For real-valued samples, these two terms represent the mean squared error between x and y. is the approximate variational posterior and the KL divergence between the prior p(z). Because is a normal distribution, and p(z) is a standard normal distribution, so the term can be written as:
[0034]
[0035] The mean μ and variance σ in the formula are 2 Obtained by two encoders respectively.
[0036] The loss function of quality supervised VAE pre-training is negative For real-valued data, it can be expressed as:
[0037]
[0038] Where N l is the number of training samples, x n represents the original input variable, represents the reconstructed value of the original input variable, y n represents the true quality variable, represents the predicted value of the quality variable, σ 2 represents the variance of the latent variable, and μ represents the mean of the latent variable; because the loss function involves information related to the target value, labeled data is required for supervised training.
[0039] After training the parameters in the network by minimizing the loss function, they can be directly stacked into the deep neural network of DE-VAE, and the hidden variables and inputs of the previous E-VAE layer are used as the inputs of the next E-VAE layer. The loss function of the k-th layer E-VAE is:
[0040]
[0041] Where k>1, μ k is the hidden variable z k The mean value, σ k is the hidden variable z k The standard deviation of x i represents the input variable, Represents the reconstructed value of the input variable, y i represents the true quality variable, represents the predicted value of the quality variable, and They represent the i-th hidden variable and reconstructed value output in the j-th layer E-VAE respectively, and m is the hidden variable dimension of the k-th layer E-VAE.
[0042] Fine-tuning the parameters in the network through the BP back-propagation algorithm can be defined as minimizing the following unconstrained optimization problem:
[0043]
[0044] Where N l is the number of training samples, represents the predicted value of the quality variable of the nth sample, and y(n) represents the true quality variable of the nth sample. It is equivalent to minimizing the mean square error (MSE) between the predicted value and the actual target value.
[0045] The instant learning framework in step 7 takes into account the different effects of different process variables on quality variables when measuring the similarity between the test sample and the historical data. Therefore, each process variable needs to be weighted when calculating the sample similarity. To this end, the present invention introduces the maximum mutual information coefficient (MIC) to compare the correlation between process variables and quality variables. MIC measures the degree of association between variables by comprehensively analyzing the linear or nonlinear relationship between them. Compared with the Pearson correlation coefficient, MIC is more comprehensive in reflecting the connection between variables. By traversing all historical data, a small batch of samples that are most similar to the query sample are selected. Each historical marked sample is similar to the query sample x. q The weighted Euclidean distance between is calculated as:
[0046]
[0047] where x h represents historical data, Q represents an n×n diagonal matrix, where the diagonal elements are the maximum mutual information coefficients between process variables and quality variables.
[0048] The beneficial effects of this application are:
[0049] 1. Aiming at the problem of error accumulation in reconstruction of variational autoencoder (VAE), a deep extended structure DE-VAE is proposed. This structure takes the original input data, quality variable data and the latent variables of the previous E-VAE as the input of the next layer of E-VAE. By reusing the prior information, the correlation between the extracted features and the original input and output is enhanced, and the feature information is fully utilized to filter out information irrelevant to the output.
[0050] 2. In order to cope with the time-varying nature of industrial processes, an adaptive framework is introduced. The weighted Euclidean distance and cosine distance are used to find a small number of samples that are most similar to the query samples in the historical database in the online prediction stage. This batch of similar samples is used to fine-tune the model and update the model parameters to enhance the model's ability to handle time-varying process data and improve the accuracy of the model.
[0051] 3. In the fine-tuning stage, based on the idea of just-in-time learning, similar data is used to update model parameters and improve the prediction accuracy of the deep extended variational autoencoder model; in order to overcome the existing soft measurement technology based on multi-layer VAE networks, which suffers from the accumulation of reconstruction errors and the lack of adaptation of the model to the time-varying nature of industrial processes, resulting in a decline in prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1 A flow chart of a deep extended VAE soft sensing method based on real-time learning in the second embodiment provided by the present invention;
[0054] Figure 2 The process flow chart of debutanizer tower involved in the second embodiment provided by the present invention;
[0055] Figure 3 A schematic diagram of the quality supervision VAE part in the JITL-DE-VAE model in a deep extended VAE soft-sensing method based on real-time learning in the second embodiment provided by the present invention;
[0056] Figure 4 The overall structure diagram of the DE-VAE model in the deep extended VAE soft sensing method based on real-time learning in the second embodiment provided by the present invention;
[0057] Figure 5A The SQAE model is used in the second embodiment of the present invention to predict the butane concentration of the debutanizer.
[0058] Figure 5B The SVAE model is used in the second embodiment of the present invention to predict the butane concentration of the debutanizer.
[0059] Figure 5C A fitting curve diagram for predicting butane concentration in debutanizer using the H-NPLVR model in Example 2 provided by the present invention;
[0060] Figure 5D The W-GSTVAE model is used in the second embodiment of the present invention to predict the butane concentration of the debutanizer.
[0061] Figure 5E A fitting curve diagram for predicting butane concentration in debutanizer using the DE-VAE model in Example 2 provided by the present invention;
[0062] Fig. 5F The JITL-DE-VAE model proposed in this application is used in Example 2 of the present invention to predict the butane concentration of the debutanizer.
[0063] Figure 6 This is a box plot of the prediction errors of the six models in Example 2 provided by the present invention. DETAILED DESCRIPTION
[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0065] Embodiment 1:
[0066] This embodiment provides a deep extended VAE soft sensing method based on real-time learning, the method comprising:
[0067] Step 1: Soft sensor modeling includes offline modeling and online prediction. The variables of the model are selected according to theoretical analysis, and then data is collected to obtain the historical data set. The historical data set is divided into training set and test set, and the data is standardized. The labeled training data set is Where x is the input variable, y is the quality variable, N l is the number of samples. The set of standardized input variables in the training data set is X s , the sample to be predicted is X new ;
[0068] Step 2: A single-layer VAE extracts high-level features of the data. To ensure the relevance of the extracted features to the quality variables, the prediction error of the quality variables is added to the reconstruction error of the VAE. In this way, the features extracted at each layer are not only feature representations of the input variables, but also effective explanations of the quality variables.
[0069] Step 3: Build a deep extended DE-VAE network model. First, build a deep DE-VAE by stacking multiple E-VAEs. Then, use the hidden layer and input variables of the previous E-VAE as the input of the next E-VAE, making full use of the effective information of the features and preventing excessive loss of effective information.
[0070] Step 4: Place the X s As the input of the DE-VAE model, calculate the reconstruction value of the input sample and the predicted target value for pre-training;
[0071] 4.1: X s As the input of DE-VAE encoder, the sample is calculated to obey the distribution of features in the latent space The mean and variance of
[0072] 4.2: Use the reparameterization technique to calculate the mean and variance Sampling to get hidden variables z ;
[0073] 4.3: Using Hidden Variables z As the input of the decoder, the reconstructed value of the sample is obtained And the predicted value of the sample
[0074] Step 5: Repeat step 4 to calculate the loss function LOSS of each pre-training layer, and use the Adam optimization method to update the model parameters;
[0075] Step 6: After pre-training is completed, use the hidden variable z to calculate the linear mapping y of the original output k , and get the final output form Fine-tune the global network by minimizing the loss function minΓ. Save the network parameters for online prediction.
[0076] Step 7: Introduce a real-time learning framework in the online prediction stage. When a test sample appears, a small batch of samples that are most similar to the test sample will be searched in the historical data (including the training set and the predicted test sample). This batch of data is then used to update the DE-VAE model parameters.
[0077] Step 8: The sample to be predicted X new After standardization, it is input into the updated model (JITL-DE-VAE) to obtain the final prediction value.
[0078] Embodiment 2:
[0079] This embodiment provides a method for predicting the butane concentration of a debutanizer. The method uses a deep extended VAE soft sensing method based on real-time learning given in the first embodiment to predict the butane concentration of a debutanizer. Figure 1 , the method comprising:
[0080] Step 1: Collect input and output data to form a historical training sample database.
[0081] This example uses a public data set of a real debutanizer industrial process. Figure 2 , mainly composed of heat exchanger, condenser, reboiler, reflux pump, separator and accumulator. Table 1 shows the 7 variables measured by the sensor. Since the main purpose of the debutanizer is to control the butane content to a minimum, the butane concentration needs to be measured in real time. However, in practice, the butane content is obtained by sampling and analysis by a gas chromatograph, and there is a large measurement delay in this process. In addition, there are complex nonlinear relationships between variables. The data set includes a total of 7 input variables and 1 quality variable, namely the butane concentration at the bottom of the debutanizer, with a total of 2394 sample data.
[0082] Considering the process dynamics, 13-dimensional augmented data variables are obtained based on expert knowledge and physical analysis:
[0083] [u 1 (k),u 2 (k),u3 (k),u 4 (k),u 5 (k),u 5 (k-1),u 5 (k-2),u 5 (k-3),(u 6 (k)+u 7 (k)) / 2,y(k-1),y(k-2),y(k-3),y(k-4)] T Among them, k represents the time and T represents the transpose.
[0084] Table 1 Process variables of debutanizer process
[0085]
[0086] Step 2: After augmenting the original sample data set, 2390 sample data are finally obtained. The data is standardized and the standardized data set is divided into a training set and a test set. The first 1600 samples are used as the training data set, and the remaining samples are used as the test data set to build and train the JITL-DE-VAE network. The standardized set of input variables in the training data set is X s , the sample to be predicted is X new .
[0087] Step 3: Deep Extended VAE (DE-VAE) design.
[0088] Considering that only compressing the number of hidden neurons in the deep layer of stacked VAE can effectively process high-dimensional data, it may cause loss of effective information in the features. Therefore, quality supervision information is added to the training process of each layer of VAE to guide feature learning. For a single-layer VAE with quality supervision, see Figure 3 In order to reduce the reconstruction error accumulation caused by multi-layer VAE, the hidden layer and input variables of the previous E-VAE are further used as the input of the next E-VAE to improve the prediction accuracy of the model. By stacking multiple E-VAEs, a deep extended VAE, namely DE-VAE, is obtained. See the structure for details. Figure 4 .
[0089] For a single-layer VAE model with quality variables, assuming that the input variable x and the output target value y in the VAE are generated by random continuous hidden variables z, the generation process can be expressed as:
[0090]
[0091] Assuming that x and y are conditionally independent of each other, the latent variable z is sampled from the latent space and is passed through two parameters δ x and δ yThe neural network is used to obtain x and y. According to the above generation process, the joint probability distribution of the generation model is:
[0092] p δ (x,y,z)=p δx (x|z)p δy (y|z)p(z)
[0093] The log-likelihood function of the marginal probability distribution p(x,y) of the data sample point is:
[0094]
[0095] in To infer the model, an additional variational posterior is used to approximate the true complex posterior probability p(z|x).
[0096] According to Jensen inequality:
[0097]
[0098] is the lower bound of the marginal probability likelihood function. In order to calculate and δ, maximizing the evidence lower bound:
[0099]
[0100] The lower bound of evidence can be understood in two parts. The first two items are related to the generative model p of x and y. δx (x|z), p δy (y|z), where z follows the approximate variational posterior Assumptions Obeys normal distribution. And because the generative models all obey normal distribution, logp δx (x|z) and logp δy (y|z) can be expressed as:
[0101]
[0102] For real-valued samples, these two terms represent the mean squared error between x and y. is the approximate variational posterior and the KL divergence between the prior p(z). Because is a normal distribution, and p(z) is a standard normal distribution, so the term can be written as:
[0103]
[0104] The mean and variance in the formula are obtained by two encoders respectively.
[0105] For single-layer quality supervised VAE pre-training, the loss function is negative For real-valued data, it can be expressed as:
[0106]
[0107] in, represents the approximate variational posterior KL divergence with the prior p(z), N l is the number of training samples, x n represents the pre-training input variable, represents the reconstructed value of the pre-trained input variable, y n represents the true quality variable, represents the predicted value of the quality variable, σ 2 represents the variance of the latent variables of this layer, and μ represents the mean of the latent variables of this layer;
[0108] After training the parameters in the network by minimizing the loss function, they can be directly put into the deep neural network of DE-VAE, and the hidden variables and inputs of the previous E-VAE are used as the input of the next E-VAE. The loss function of the k-th layer E-VAE is:
[0109]
[0110] Where k>1, μ k is the hidden variable z k The mean value, σ k Hidden variable z k The standard deviation of and They represent the i-th latent variable and reconstruction value in the j-th E-VAE respectively, and m is the latent variable dimension of the k-th E-VAE.
[0111] Step 4: Place the X s As the input of the DE-VAE model, calculate the reconstruction value of the input sample and the predicted target value for pre-training;
[0112] 4.1: X s As the input of DE-VAE encoder, the sample is calculated to obey the distribution of features in the latent space The mean and variance [μ,σ 2 ];
[0113] 4.2: Using the reparameterization technique, according to the mean and variance [μ,σ 2 ] is sampled to hidden variable z;
[0114] 4.3: Use the hidden variable z as the input of the decoder to obtain the reconstructed value of the sample And the predicted value of the sample
[0115] Step 5: Repeat step 4 to calculate the loss function LOSS of each pre-training layer, and use the trial and error method to determine the network structure. The final DE-VAE model stacks three E-VAE models, and the number of neuron nodes in the input layer, hidden layer, and output layer are 13-7-14, 20-7-21, and 27-7-28 respectively. The Adam optimization method is used to update the model parameters. The number of training rounds in the pre-training stage is set to 100, the batch size is set to 64, and the learning rate is set to 0.01.
[0116] Step 6: After pre-training is completed, use the hidden variable z to calculate the linear mapping y of the original output k , and get the final output form Fine-tuning the parameters in the network through the BP back-propagation algorithm can be defined as minimizing the following unconstrained optimization problem:
[0117]
[0118] Where N l is the number of training samples, represents the predicted value of the quality variable of the nth sample, y(n) represents the true quality variable of the nth sample, ||·|| 2 Represents the L2 norm.
[0119] It is equivalent to minimizing the mean square error (MSE) between the predicted value and the actual target value.
[0120] The number of training rounds in the fine-tuning phase is set to 100, the batch size is set to 64, and the learning rate is set to 0.01. The network parameters are saved for online prediction.
[0121] Step 7: Introduce a just-in-time learning framework in the online prediction stage. When a test sample appears, a small batch of samples that are most similar to the test sample will be searched in the historical data (including the training set and the predicted test sample). This batch of data is then used to update the DE-VAE model parameters. When measuring the similarity between the test sample and the historical data, the effects of different process variables on the quality variables are different. Therefore, when calculating the sample similarity, each process variable needs to be weighted. To this end, this application introduces the maximum mutual information coefficient (MIC) to compare the correlation between process variables and quality variables. MIC measures the degree of association between variables by comprehensively analyzing the linear or nonlinear relationship between them. Compared with the Pearson correlation coefficient, MIC is more comprehensive in reflecting the connection between variables. By traversing all historical data, a small batch of samples that are most similar to the query sample are selected. Each historical marked sample is related to the query sample x. q The weighted Euclidean distance between is calculated as:
[0122]
[0123] where x h represents historical data, Q represents an n×n diagonal matrix, where the diagonal elements are the weight coefficients of the MIC of process variables and quality variables.
[0124] The calculation process of MIC can be described as follows: if there is a relationship between two variables, a grid G can be drawn on a scatter plot with these two variables as the X-axis and the Y-axis to divide the data to encapsulate this relationship. There is a data set D containing two variables with a total of n samples. According to the above method, the data set is divided into a and b parts along the X-axis and the Y-axis using the grid G. The probability distribution corresponding to each point on the grid G is represented by D G Indicates that mutual information (MI) can be calculated under this grid division. Different grid divisions will produce different probability distributions and mutual information values. The maximum MI value obtained by dividing the grid G is defined as:
[0125] I(D,a,b)=maxI(D G )
[0126] The MI values under different grids are standardized to obtain modified values between 0 and 1 to ensure fair comparison, so the MIC can be defined as:
[0127]
[0128] In order to further enhance the correlation between features and target values, extract features that are highly correlated with output prediction, calculate the MIC between input variables and quality variables when looking for similar samples, assign more weight to variables related to output, and give less attention to variables with low correlation with output. l The labeled training dataset E l , the importance of a variable is determined by its MIC with the target variable. The weight of the d-th dimension input variable is calculated as follows:
[0129]
[0130] Therefore, the diagonal elements of Q can be expressed as λ 1 ... n .
[0131] Step 8: The sample to be predicted X new After the same standardization, it is input into the updated model (JITL-DE-VAE) to obtain the final predicted value.
[0132] This application proposes a time-varying process adaptive model update framework based on a deep extended variational autoencoder, which can fully extract effective information of process variables and instantly update model parameters and improve performance. In order to solve the problem that traditional deep learning methods cannot update models in real time, resulting in reduced model performance, this application scheme first expands the input layer and output layer of each layer of E-VAE to retain as much effective information as possible, thereby improving the model feature extraction capability. Then a real-time fine-tuning strategy is used to update the DE-VAE model, so that deep learning can quickly adapt to the process operation status and further enhance the model prediction accuracy.
[0133] In order to compare the superiority of JITL-DE-VAE proposed in this application, this application uses 5 models for comparison, namely SQAE, SVAE, H-NPLVR, W-GSTVAE and DE-VAE. Among them, SQAE is an improved method based on SAE in the above derivation process; SVAE is a supervised variational autoencoder model method; H-NPLVR is a hierarchical nonlinear probability latent variable regression model method; W-GSTVAE is an improved method based on VAE, and specific reference can be made to the Chinese patent with publication number CN118444641A; DE-VAE is a VAE method with an extended structure introduced in this application. In order to ensure the rationality of the comparison, the model structures of the first three regression models are set to 13-10-7-4-1 respectively. Among them, 10, 7, and 4 are the number of hidden layer neuron nodes in the three-layer network structure, and the model structure of H-NPLVR is set to 13-10-10-10-1. At the same time, this application uses the following two parameters as the prediction model performance evaluation indicators in each prediction method, and their calculation formulas are as follows:
[0134] (1) Root mean square error (RMSE):
[0135]
[0136] Where N is the number of test set samples, and y(n) is the true output value of the nth sample in the test set. is the predicted value of the nth sample in the test set. RMSE focuses on the prediction error of the sample. The smaller the RMSE, the better the model performance.
[0137] (2) Coefficient of Determination (R 2 ):
[0138]
[0139] in It is the average value of the true output value of all samples in the test set. 2It represents the square correlation between the true output value and the predicted value of the test set, reflecting the ability of the model to explain the variance of the output data. 2 The closer it is to 1, the better the model performance.
[0140] All experiments in the examples of this application were carried out on the same computer: Intel(R) Core(TM) i7-9750H CPU@2.60GHz 2.60GHz processor, 16GB memory, 64-bit Windows11 operating system under Pytorch3.8SQAE, SVAE, H-NPLVR, W-GSTVAE, DE-VAE and the proposed solution JITL-DE-VAE for butane concentration prediction results are shown in Figure 2. Figure 5A , Figure 5B , Figure 5C , Figure 5D , Figure 5E , Fig. 5F and Figure 6 The detailed experimental comparison results of the test set are shown in Table 2. The evaluation indicators are RMSE and R 2 . As can be seen from Figure 5, these six methods can well reflect the changing trend of butane content. However, in the prediction of peak values and trough values, the prediction accuracy of the JITL-DE-VAE of the present application scheme is better than that of other models. From the perspective of the model's peak prediction effect, the SQAE model prediction error is about 0.2, which is significantly weaker than the other five models, which proves that the denoising ability of VAE is better than that of SAE; from the perspective of the prediction of trough values, the prediction error of the JITL-DE-VAE model of the present application scheme is less than 0.02, which is better than other models. This proves the effectiveness of adding expansion layers, introducing output constraints, and adopting an adaptive update strategy. Therefore, compared with the other five models, the input-output related features extracted by the JITL-DE-VAE model of the present application scheme are more suitable for process modeling.
[0141] Table 2 Evaluation indicators of six methods on DC dataset
[0142]
[0143] exist Figure 6It can be clearly seen from the error box plot that the JITL-DE-VAE model proposed in this application has a narrower error range than the other five methods, which has obvious advantages. The JITL-DE-VAE proposed in this application pays more attention to those samples that cause larger prediction errors, rather than just those samples that account for a larger proportion of the distribution. It can more accurately predict the butane concentration, which once again proves the superiority of the proposed method. From the perspective of industrial processes, it is very important to predict the butane content as accurately as possible when the butane content is relatively high, because this helps to immediately detect abnormalities in the process state.
[0144] The present application discloses a soft measurement modeling method of deep extended VAE based on real-time learning, which can fully extract effective information of process variables and update model parameters in real time to improve performance. In view of the problem that traditional deep learning methods cannot update models in real time, resulting in reduced model performance, the present application scheme first expands the input layer and output layer of each E-VAE to retain as much effective information as possible, thereby improving the ability of model feature extraction. Then a real-time fine-tuning strategy is used to update the DE-VAE model, so that deep learning can quickly adapt to the process operation status and further enhance the model prediction accuracy. Finally, the superiority of the proposed method is verified on the actual industrial data set of the debutanizer. It can provide strong technical support for the butane concentration monitoring of the debutanizer.
[0145] Some of the steps in the embodiments of the present application may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.
[0146] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A deep extended VAE soft sensing method based on real-time learning, characterized in that: The method comprises: Step 1: Select the input variables of the model, collect data, including input variables and quality variables, and standardize the data; Step 2: When the single-layer VAE extracts high-level features of the data, the prediction error of the quality variable is also added to the reconstruction error of the VAE; Step 3: Build a deep extended DE-VAE network model. Build a deep VAE by stacking multiple single-layer VAEs, and use the hidden layer and input variables of the previous E-VAE as the input of the next E-VAE. Step 4: Place the X s As the input of the DE-VAE model, calculate the reconstruction value of the input sample and the predicted target value, and pre-train the DE-VAE model; Step 5: Repeat step 4 to calculate the loss function LOSS of each pre-training layer, and use the Adam optimization method to update the model parameters; Step 6: After pre-training is completed, use the hidden variable z to calculate the linear mapping y of the original output k , and get the final output form Fine-tune the global by minimizing the loss function minΓ; Step 7: Introduce a just-in-time learning framework in the online prediction stage. When a test sample appears, search for a small batch of samples that are most similar to the test sample in the historical data, and then use this batch of data to update the DE-VAE model parameters to obtain the updated model JITL-DE-VAE. The historical data includes the training set and the predicted test sample. Step 8: The sample to be predicted X new After standardization, it is input into the updated model JITL-DE-VAE to obtain the final prediction value.
2. The method according to claim 1, characterized in that The data collected in step 1 is recorded as Where x is the input variable, y is the quality variable, N l is the number of samples; the standardized set of input variables in the training data set is X s , the sample to be predicted is X new .
3. The method according to claim 2, characterized in that Each historical tag sample and query sample x q The weighted Euclidean distance between is calculated as: where x h represents historical data, Q represents an n×n diagonal matrix, and the diagonal elements are the maximum mutual information coefficients between process variables and quality variables, and T represents transpose.
4. The method according to claim 3, characterized in that The step 4 comprises: 4.1: X s As the input of DE-VAE encoder, the sample is calculated to obey the distribution of features in the latent space The mean and variance [μ,σ 2 ]; 4.2: Using the reparameterization technique, according to the mean and variance [μ,σ 2 ] Sampling obtains the latent variable z; 4.3: Using Hidden Variables z As the input of the decoder, the reconstructed value of the sample is obtained And the predicted value of the sample 5. The method according to claim 4, characterized in that In the model JITL-DE-VAE, the loss function of the first layer of E-VAE pre-training is: in, represents the approximate variational posterior KL divergence with the prior p(z), N l is the number of training samples, x n represents the pre-training input variable, represents the reconstructed value of the pre-trained input variable, y n represents the true quality variable, represents the predicted value of the quality variable, σ1 2 represents the variance of the first layer latent variable, μ1 represents the mean of the first layer latent variable; The loss function of the k-th layer E-VAE is: Where k>1, μ k is the hidden variable z k The mean value, σ k is the hidden variable z k The standard deviation, N l is the number of training samples, and They represent the i-th latent variable and reconstruction value in the j-th E-VAE respectively, and m is the latent variable dimension of the k-th E-VAE.
6. The method according to claim 5, characterized in that The step 6 comprises: The minimization loss function minΓ is: Where N l is the number of training samples, represents the predicted value of the quality variable of the nth sample, y(n) represents the true quality variable of the nth sample, and ||·||2 represents the L2 norm.
7. The method according to claim 1, characterized in that The industrial processes include petrochemicals, blast furnace refining and fermentation processes.
8. A method for predicting the bottom concentration of a debutanizer tower based on deep extended VAE with real-time learning, characterized in that: The method is implemented based on the method described in any one of claims 1 to 6, and the method comprises: Step 1: Collect historical data from the debutanizer process, including 7 input variables and 1 quality variable; Step 2: Augment the collected historical data and use the augmented data as training samples to train the JITL-DE-VAE model; Step 3: Collect the input variables in the debutanizer process in real time, input them into the JITL-DE-VAE model trained in step 2, and obtain the predicted values of the quality variables.
9. The method according to claim 8, characterized in that The seven input variables are top temperature u1, top pressure u2, top reflux u3, top discharge u4, plate 6 temperature u5, first tower bottom temperature u6, second tower bottom temperature u7; the mass variable is butane concentration at the bottom of the debutanizer tower.
10. The method according to claim 9, characterized in that The data after the augmentation processing in step 2 includes: [u1(k), u2(k), u3(k), u4(k), u5(k), u5(k-1), u5(k-2), u5(k-3), (u6(k)+u7(k)) / 2, y(k-1), y(k-2), y(k-3), y(k-4)] T Among them, k represents the time and T represents the transpose.
Citation Information
Patent Citations
Counting data soft measurement modeling method based on stacked Poisson auto-encoder network
CN114692507A
Soft measurement method based on weighted gating stacking quality supervision VAE
CN118444641A
DEEP HIERARCHICAL VARIATIONAL AUTOCODER
DE102021206286A1
Cited By
A soft measurement method and device based on a stack feature expansion network
CN122757845A