Chemical process key parameter component prediction method based on transfer learning and sparse self-attention
By adopting a prediction method based on transfer learning and sparse self-attention in the chemical process, combined with confrontation learning and sparse self-attention mechanism, the problems of incomplete data and noise interference in the chemical process are solved, and the accurate prediction of key parameter components of the chemical process is achieved and the calculation efficiency is improved.
Patent Information
- Application Number
- CN202510345260.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
There are problems of incomplete data, complex noise interference, uncertainty, dynamicity and lack of labels in the chemical process, resulting in low robustness, weak generalization and difficulty in migration in the prediction of key parameter components in traditional deep learning models.
The key parameter component prediction method of chemical process based on transfer learning and sparse self-attention is adopted. The model architecture is constructed through the sparse self-attention mechanism, combined with adversarial learning and transfer learning, the source domain data set is constructed using DCS data and online analyzer data, and the target domain data set is constructed using DCS data and offline assay data to achieve accurate prediction of key parameter components of chemical process.
It improves the accuracy and calculation efficiency of predicting key parameter components in chemical process, enhances the anti-interference ability of the model and the ability to capture complex data distribution, solves the problem of data scarcity, and improves the effect of offline test data prediction.
Smart Images

Figure CN120199352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial soft sensing, and particularly but not limited to a method for predicting the composition of key parameters in a chemical process based on transfer learning and sparse self-attention. Background Art
[0002] Chemical processes are characterized by complex production processes, harsh production conditions, variable operating conditions, and strict product requirements. The increasing market demand and the continuous improvement of the automation management level have led more and more petrochemical enterprises to expect to reduce production costs and improve product quality through digital and intelligent means. In order to better monitor the operating state of chemical processes, ensure the safe operation of chemical processes, and improve the product quality and economic benefits of enterprises, it is necessary to monitor the key variables of chemical processes, such as important process parameters or product quality parameters, in real time. However, the harsh production environment, expensive on-line monitoring instruments, and time-consuming off-line tests pose difficulties for the real-time monitoring of key variables. Some scholars have proposed a soft sensing technology of "using auxiliary variables to infer main variables", that is, using directly measurable process variables (i.e., auxiliary variables) to calculate or predict difficult-to-measure key variables (i.e., main variables). With the rapid development of distributed control systems (DCS), laboratory information management systems (Lims), various sensing devices, and computer technologies, the chemical industry has accumulated a large amount of process data, providing a guarantee for data-driven soft sensing modeling.
[0003] As a powerful data-driven method for dealing with nonlinear problems, deep learning has been widely applied to the field of soft sensor modeling in chemical processes because it can effectively extract useful information hidden in massive process data. However, when traditional deep learning-based soft sensor modeling methods are applied to actual chemical production processes, there are still many engineering problems to be solved, such as incomplete process data, complex noise interference, uncertainty, dynamics, lack of labels, etc., as well as scientific problems brought about by engineering problems, such as sparse data distribution inference, low robustness of models, weak generalization ability, and difficulty in migration. These factors make the prediction of key parameter components a complex sequence prediction problem. Deep learning models for solving time series prediction problems are mainly divided into three types: recurrent neural networks (RNNs), convolutional neural networks (CNNs), and Transformer-based models. Among them, RNNs can effectively process sequential data, but have poor gradient stability and complex training. CNNs can cleverly process the detection of local temporal patterns, but it is difficult to handle long-distance dependencies. In contrast, Transformer-based models overcome these limitations and achieve excellent performance using the self-attention mechanism. Transformer has three main advantages: unparalleled parallelism, enhanced interpretability provided by the self-attention mechanism, and a detailed representation of the input data. To alleviate the computational requirements and running efficiency of the traditional self-attention mechanism, the probabilistic sparse self-attention mechanism has emerged, achieving a more efficient analysis process by probabilistically isolating key queries.
[0004] Abnormal situations in industrial environments (such as inaccurate measurements and system operation interruptions) introduce complex perturbations and outliers into the operation dataset, posing challenges to traditional deep learning frameworks. Deep learning models usually operate under the assumptions of data independence and identical distribution, resulting in sensitivity to anomalies and distribution changes. To address this challenge, adversarial learning has emerged, integrating adversarial games into the design of network architectures and loss functions, enhancing the model's ability to identify resilient features and accurately approximate the data distribution.
[0005] In addition, the chemical process also faces the important obstacle of the lack of accurate measurement data in multivariable systems, which may lead to overfitting or underfitting of traditional deep learning models. Although online analyzers can monitor the fluctuations of key parameter components, due to the harsh production environment and limitations of measurement accuracy, the data of online analyzers are prone to deviation. On the other hand, accurate parameter component content values are usually obtained through offline assays, which results in significant time delays and limited sample availability. Due to the high cost of offline experiments and insufficient samples for obtaining accurate quality variables, this problem becomes more prominent. Deep transfer learning has become an excellent method for solving few-shot quality prediction, improving the performance in the target domain through insights obtained from the source domain.
[0006] In view of this, a new structure or control method is needed to solve at least some of the above problems. Summary of the Invention
[0007] Aiming at one or more problems in the prior art, the present invention proposes a method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention, realizing accurate prediction of the key parameter components of the chemical process.
[0008] The technical solution to achieve the object of the present invention is as follows:
[0009] A method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention, comprising:
[0010] S1. Collect DCS data, on-line analyzer data, and off-line laboratory data in the chemical process. Among them, the DCS data and on-line analyzer data constitute the source domain data set, and the DCS data and off-line laboratory data constitute the target domain data set;
[0011] S2. Based on the auxiliary variables and their corresponding key parameter component data in the source domain data set, construct a training set and a validation set, and use the sparse self-attention mechanism to construct a model architecture for describing the inertia and non-linear characteristics of the chemical process data;
[0012] S3. Through adversarial learning, use the discriminator to perform a minimax game with the model architecture to obtain a pre-trained model;
[0013] S4. Adopt transfer learning, freeze some network layers of the pre-trained model, and use the target domain data set to train its unfrozen network layers to obtain a soft sensor model for predicting the key parameter component data for off-line laboratory analysis;
[0014] S5. Use the soft sensor model to predict the key parameter components of the chemical process.
[0015] Furthermore, in the method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention of the present invention, the specific steps of using the sparse self-attention mechanism to construct the model architecture in S2 include:
[0016] S2-1. In the source domain data set, given the t-th sequence input as and After being processed by the encoder and decoder, the output is where seq (S) is the encoder sequence length, lab (S) is the decoder sequence length, is the total dimension of the data set, is the total dimension of y, pre (S)Let \(seq\) be the predicted sequence length, the subscripts \(e\) and \(d\) represent the encoder and decoder respectively, \(concat()\) represents the concatenation of two sets of data, and the superscript \(t\) represents the \(t\)-th sequence. At the same time, for the input sequence perform reconstruction to obtain
[0017] S2 - 2. The \(t\)-th input sequence of the encoder after positional encoding is After passing through the linear layer, the query vector is obtained the key vector the value vector
[0018]
[0019] where are learnable parameters, \(\tau(\cdot)\) represents the linear layer, and the subscripts \(q\), \(k\), \(v\) represent query, key, and value respectively;
[0020] S2 - 3. Perform head splitting. The \(h\)-th head is The number of heads is \(H\), the single - head dimension is \(D'=D / H\), \(h = \{1,\cdots,H\}\). Introduce the sparse self - attention mechanism. For randomly sample \(L1\) key vectors to obtain
[0021] 1) The distribution of sampled under the condition of the \(i\)-th query vector:
[0022]
[0023] where \(i\in\{1,\cdots,seq (S) \}\), \(j\in\{1,\cdots,L1\}\), \(l\in\{1,\cdots,L1\}\);
[0024] 2) Define a uniform distribution:
[0025]
[0026] 3) Use the KL - divergence to calculate the similarity between the two distributions. The calculation process is as follows:
[0027]
[0028] The upper and lower bounds of the above formula are:
[0029]
[0030] 4) Substitute the upper bound into the KL - divergence to obtain the approximate result of the sparsity metric criterion:
[0031]
[0032] Among them, denotes the i-th query vector in and calculate the sparsity score. For the first m e query vectors with the highest scores, denote them as Use and to calculate. The calculation process is as follows:
[0033]
[0034] Among them, The corresponding value is The remaining part is
[0035] 5) The first part of the required attention value is filled by multiplying with , and the second part is filled by taking the average. Concatenate the results of the two parts to obtain the result of the h-th head. The calculation process is as follows:
[0036]
[0037] Among them, Finally, concatenate the results of (11) to obtain the final result
[0038] S2-4. The t-th input sequence of the decoder after positional encoding is Obtain The two parts respectively pass through a linear layer and a residual connection layer, and then through a cross-attention layer to obtain
[0039] Pass through a linear layer to obtain the final prediction result This process is expressed as:
[0040] Finally, obtain the reconstruction result through a feed-forward neural network:
[0041] Furthermore, for the key parameter component prediction method of the chemical process based on transfer learning and sparse self-attention of the present invention, S3 to obtain the pre-trained model specifically includes;
[0042] S3-1. Introduce a discriminator network. The discriminator network classifies samples into two categories. Real samples are denoted as "true", and generated samples are denoted as "false". Use the reconstruction result obtained in S2 as the generated sample of the discriminator network. The sample classification in the discriminator network is expressed as:
[0043]
[0044] Among them, Z (S) is the input of the discriminator network, and Ω(·) is a feed-forward neural network. are the training parameters of the discriminator in the source model. represents the output of the discriminator network, and δ(·) represents the Sigmoid function;
[0045] S3-2. Based on the output of the discriminator network, the pre-trained model is iteratively updated to minimize the difference between the real samples and the generated samples;
[0046] S3-3. The loss function of the pre-trained model A-informer is Among them, is the prediction error term, is the adversarial loss term, is the input reconstruction term;
[0047] Furthermore, in the chemical process key parameter composition prediction method based on transfer learning and sparse self-attention of the present invention, in S3-3:
[0048] The prediction error term is: Among them, t represents the t-th sequence input, and i represents the i-th y,seq (S) is the sequence length, and pre (S) is the predicted sequence length. represents the predicted value, is the true value;
[0049] The input reconstruction term Among them, lab (S) is the sequence length, is the true value of the input and output, is the reconstructed value of the input and output;
[0050] The adversarial loss term Among them is a type of kernel function of, is a reproducing kernel Hilbert space, and MMD is the maximum mean discrepancy.
[0051] Furthermore, in the chemical process key parameter composition prediction method based on transfer learning and sparse self-attention of the present invention, the soft measurement model described in S4 includes an encoder, a decoder, and a discriminator, among which:
[0052] The encoder includes an encoding layer and a probabilistic sparse attention layer. The encoding layer is used to enhance the model's perception ability of position information, and the probabilistic sparse attention layer is used to process the information interaction of different positions in the input sequence;
[0053] The decoder includes an encoding layer, a probabilistic sparse attention layer and a cross-attention layer, and is used to obtain the reconstructed samples;
[0054] The discriminator is a feed-forward neural network, which is used to minimize the difference between the real samples and the reconstructed samples through adversarial learning.
[0055] Finally, the output result of the decoder is input into a fully connected neural network layer FNN to obtain the prediction result.
[0056] Furthermore, for the key parameter component prediction method of chemical process based on transfer learning and sparse self-attention of the present invention, the soft measurement model obtained in S4 for predicting the key parameter component data for offline laboratory tests specifically includes:
[0057] S4-1. In the target domain, given the t-th sequence input as and where seq (T) is the encoder sequence length, lab (T) is the decoder sequence length, pre (T) is the prediction sequence length, is the total dimension of the data set, is the total dimension of y, and the superscript T represents the target domain. The target domain data is encoded and decoded to obtain
[0058] S4-2. Freeze the FNN layer of the final output of the pre-trained model, randomly initialize the corresponding parameters, and obtain the prediction result as follows:
[0059]
[0060] where, is initialized in the target domain, represents the FNN layer;
[0061] S4-3. Select the MSE loss function to update the model parameters of the target domain data to obtain the final soft measurement model:
[0062] where t represents the t-th sequence input, i represents the i-th y, seq (T) is the sequence length, pre (T) is the prediction sequence length, represents the predicted value, is the true value.
[0063] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0064] 1. The method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention of the present invention adopts an encoder-decoder framework based on sparse self-attention, improving the accuracy and computational efficiency of the method for predicting the key parameter components of a chemical process.
[0065] 2. The method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention of the present invention integrates adversarial learning technology based on game theory principles, enhancing the anti-interference ability of the model and the ability to capture complex data distributions.
[0066] 3. The method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention of the present invention adopts a transfer learning architecture for scarce data, learns the essential process information from the rich data of DCS and on-line analyzers, and transfers it to the data modeling task of very scarce off-line laboratory tests, enhancing the prediction effect of the model for limited off-line laboratory test data. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings are used to provide a further understanding of the present invention, and together with the description are used to explain the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0068] Figure 1 The sampling data flow chart of the method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention of the present invention is shown.
[0069] Figure 2 The simplified flow chart of the dimethyl oxalate synthesis process according to an embodiment of the present invention is shown.
[0070] Figure 3 The flow chart of the method for predicting the key parameter components of a chemical process based on transfer learning and sparse self-attention of the present invention is shown.
[0071] Figure 4 The prediction curve chart of the dimethyl oxalate synthesis process according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To further understand the present invention, the preferred embodiments of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.
[0073] The description of this part only targets typical embodiments, and the present invention is not limited to the scope described in the embodiments. Combinations of different embodiments, mutual replacement of some technical features in different embodiments, and mutual replacement of the same or similar prior art means with some technical features in the embodiments are also within the scope of description and protection of the present invention.
[0074] As Figure 2 shown, the auxiliary variables in the dimethyl oxalate synthesis process include methanol inlet flow rate, methanol recovery flow rate, syngas flow rate to the nitric acid reduction tower, methanol MN solution flow rate, methanol recovery flow rate, nitric acid addition flow rate, fresh CO gas pressure entering the boundary area, mixed gas circulation pressure, outlet gas pressure, bottom pressure of the tower, internal temperature of the reactor, bottom temperature of the tower, bottom temperature of the tank, top temperature of the tower, etc. The key process parameter component is the CO concentration. Taking the monitoring and control system for the dimethyl oxalate synthesis process as an example in this embodiment, the mass of carbon monoxide (CO) in the recycle gas is estimated online, and the soft measurement model TA-informer based on transfer learning and sparse self-attention proposed in this scheme is used to perform online estimation of the CO mass prediction in the recycle gas during the dimethyl oxalate synthesis process. The specific steps of the technical solution are as follows:
[0075] S1. Collect three data sources: DCS data, online analyzer data, and offline laboratory data, with sampling frequencies of seconds, minutes, and days respectively. As Figure 1 shown, at the same time, the DCS data and the online analyzer data constitute the source domain dataset G (S) ={x (S) , y (S)}, and the DCS data and the offline laboratory data constitute the target domain dataset G (T) ={x (T) , y (T)}.
[0076] S2. Construct a training set and a validation set based on the auxiliary variables and their corresponding key parameter component data of the source domain dataset collected in S1, and construct a model architecture Informer based on the sparse self-attention mechanism to describe the inertia and nonlinear characteristics of the chemical process. S2-1. In the source domain, given the t-th sequence input The final output is where seq (S) and lab (S) are the sequence lengths, pre (S) is the prediction sequence length, the subscripts e and d represent the encoder and decoder, and the superscript t represents the t-th sequence. Also, is reconstructed to obtain
[0077] S2-2. To handle the information interaction at different positions in the input sequence, A-informer adopts sparse self-attention to find important query vectors, and then only calculates the attention values of these query vectors to optimize the computational efficiency, enabling the model to better capture the long-range dependencies and context information in the sequence.
[0078] The t-th input sequence of the encoder after positional encoding is The query vector is obtained after passing through the linear layer Key vector Value vector
[0079]
[0080] Among them, are learnable parameters, τ(·) represents the linear layer, and the subscripts q, k, v represent query, key, and value respectively.
[0081] S2-3. Perform head splitting. The h-th head is The number of heads is H, the single-head dimension is D′ = D / H, h = {1,..., H}, and a sparse self-attention mechanism is introduced. Only a partial value vector is sampled to calculate the sparsity with the query vector. For Randomly sample L1 key vectors to obtain
[0082] 1) Represent the distribution of sampled under the condition of the i-th query vector in probability form:
[0083]
[0084] where i ∈ {1,..., seq (S)}, j ∈ {1,..., L1}, l ∈ {1,..., L1}.
[0085] 2) Define a uniform distribution:
[0086]
[0087] 3) Use the KL divergence to calculate the similarity between the two distributions. The calculation process is as follows:
[0088]
[0089] The upper and lower bounds of the above formula can be expressed as:
[0090]
[0091] 4) Substitute the upper bound into the KL divergence formula to obtain the approximate result of the final sparsity metric criterion:
[0092]
[0093] denote the $i$-th query vector in Calculate the sparsity score. After the calculation is completed, find the top $m$ e query vectors with the highest scores, denoted as Next, use and to calculate. The calculation process is as follows:
[0094]
[0095] where the corresponding value is The remaining part is
[0096] 5) The first part of the desired attention value is filled by multiplying with and the second part is filled by taking the mean. Concatenate the results of the two parts to obtain the result of the $h$-th head. The calculation process is as follows:
[0097]
[0098] where Finally, concatenate the results of (11) to obtain the final result
[0099] S2-4. The $t$-th input sequence of the decoder after position encoding is Similar to step S2-3, we get The two parts respectively pass through a linear layer and a residual connection layer. The obtained results are then subjected to a cross-attention to obtain the final result
[0100] Pass through a linear layer to obtain the final prediction result This process can be expressed as: It also passes through a feed-forward neural network to obtain the reconstruction result:
[0101] S3. Adopt adversarial learning, use the discriminator to play a minimax game with Informer to obtain the final pre-trained model A-Informer.
[0102] S3-1. Introduce a discriminator network to play a min-max game with Informer;
[0103] S3-2. The discriminator network classifies the samples into two categories. The real samples are denoted as "true", and the generated samples are denoted as "false". The generated samples obtained in S2-4 serve as the generated samples for the discriminator. The sample classification in the discriminator network can be expressed as:
[0104]
[0105] where Z (S) is the input of the discriminator network; Ω(·) is a feedforward neural network; are the training parameters of the discriminator in the source model; represents the output of the discriminator network, and δ(·) represents the Sigmoid function. Based on the output of the discriminator, the entire A-informer is iteratively updated to minimize the difference between the real samples and the generated samples.
[0106] S3-3. The loss function of the A-informer is where, is the prediction error term, is the adversarial loss term, is the input reconstruction term.
[0107] The prediction error term is: where t represents the t-th sequence input, i represents the i-th y,seq (S) is the sequence length, pre (S) is the predicted sequence length, represents the predicted value, is the true value;
[0108] The input reconstruction term where lab (S) is the sequence length, is the true value of the input-output, is the reconstructed value of the input-output;
[0109] The adversarial loss term where is a type of kernel function of, is a reproducing kernel Hilbert space, and MMD is the maximum mean discrepancy.
[0110] The above three terms together constitute the final objective function of A-Informer, and the stochastic gradient descent SGD algorithm is used for optimization to obtain the final pre-trained model A-Informer.
[0111] S4. Adopt transfer learning to freeze some layers of the pre-trained model A-Informer obtained in S3, and then use the target domain dataset to train the unfrozen layers to fine-tune the network parameters, obtaining the finally trained soft measurement model TA-Informer for predicting the key parameter components data in off-line chemical analysis. Figure 3 The overall structure diagram of TA-Informer is given;
[0112] S4-1. In the target domain, given the t-th sequence input as where the superscript T represents the target domain. The target domain data passes through two parts: an encoder and a decoder, obtaining This process is the same as that in S3.
[0113] S4-2. Freeze the FNN layer of the final output of the model, randomly initialize the corresponding parameters, and obtain the prediction result as follows:
[0114]
[0115] where is initialized in the target domain, represents the FNN layer.
[0116] S4-3. Select the MSE loss function to update the model parameters for the target domain data, obtaining the final soft measurement model: where t represents the t-th sequence input, i represents the i-th y,seq (T) is the sequence length, pre (T) is the predicted sequence length, represents the predicted value, is the true value.
[0117] S5. Use the soft measurement model TA-Informer trained in S4 to predict the CO concentration, a key parameter component, in the dimethyl oxalate synthesis process. The curve of the final prediction result is as Figure 4 shown.
[0118] The description and application of the present invention herein are illustrative and are not intended to limit the scope of the present invention to the above embodiments. The relevant descriptions of effects or advantages involved in the specification may not be reflected in actual experimental examples due to uncertainties in specific condition parameters or other factors, and the relevant descriptions of effects or advantages are not used to limit the scope of the invention. Modifications and changes to the disclosed embodiments herein are possible, and various substitutions and equivalent components of the embodiments are known to those of ordinary skill in the art. Those skilled in the art should clearly understand that the present invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the present invention. Other modifications and changes can be made to the disclosed embodiments herein without departing from the scope and spirit of the present invention.
Claims
1. A method for predicting key parameter components of chemical processes based on transfer learning and sparse self-attention, characterized in that: include: S1, collecting DCS data, online analyzer data, and offline test data in the chemical process, wherein DCS data and online analyzer data constitute the source domain data set, and DCS data and offline test data constitute the target domain data set; S2. constructing a training set and a validation set based on the auxiliary variables in the source domain data set and their corresponding key parameter component data, and using a sparse self-attention mechanism to construct a model architecture for describing the inertia and nonlinear characteristics of chemical process data; S3. Through adversarial learning, a discriminator is used to perform a minimum-maximum game with the model architecture to obtain a pre-trained model; S4, using transfer learning to freeze part of the network layers of the pre-trained model, and using the target domain data set to train the unfrozen network layers, to obtain a soft measurement model for predicting key parameter component data for offline testing; S5. Utilize the soft sensor model to predict key parameter components of the chemical process.
2. The method for predicting key parameters and components of a chemical process based on transfer learning and sparse self-attention according to claim 1, characterized in that: The model architecture constructed using the sparse self-attention mechanism in S2 specifically includes: S2-1. In the source domain dataset, given the tth sequence input is and After being processed by the encoder and decoder, the output is in, seq (S) is the encoder sequence length, lab (S) is the decoder sequence length, is the total dimension of the dataset, is the total dimension of y, pre (S) To predict the length of the sequence, the subscripts e and d represent the encoder and decoder respectively, concat() represents the connection of two sets of data, and the superscript t represents the tth sequence. Reconstruction to obtain S2-2, the tth input sequence of the encoder after position encoding is After the linear layer, we get the query vector Key Vector Value vector in, is a learnable parameter, τ(·) denotes a linear layer, and subscripts q, k, and v represent query, key, and value, respectively; S2-3, split the heads, the hth head is The number of heads is H, the dimension of a single head is D′=D / H, h={1,...,H}, and the sparse self-attention mechanism is introduced. Randomly sample L1 key vectors and get 1) Sampled under the condition of the i-th query vector Distribution: Where i∈{1,...,seq (S) }, j∈{1,...,L1}, l∈{1,...,L1}; 2) Define a uniform distribution: 3) Use KL divergence to calculate the similarity between two distributions. The calculation process is as follows: The upper and lower bounds of the above formula are: 4) Substituting the upper bound into the KL divergence yields an approximate result for the sparsity metric: in, express The i-th query vector in Calculate the sparsity score and convert the first m e The query vector with the highest score is denoted as use and To calculate, the calculation process is as follows: in, The corresponding value is The rest is 5) The first part of the attention value is used and Multiply them to fill, take the mean of the second part to fill, concatenate the two parts to get the result of the hth head. The calculation process is as follows: in, Finally, the results of (11) are concatenated to obtain the final result. S2-4, the t-th input sequence of the decoder after position encoding is get The two parts pass through the linear layer and the residual connection layer respectively, and then pass through the cross attention layer to obtain After a linear layer, the final prediction result is obtained This process is represented as: Finally, the reconstruction result is obtained through the feedforward neural network:
3. The method for predicting key parameters and components of a chemical process based on transfer learning and sparse self-attention according to claim 1, characterized in that: S3 obtains the pre-trained model specifically including: S3-1, introduce the discriminator network, the discriminator network divides the samples into two categories, the real samples are represented as "true", and the generated samples are represented as "false". As the generated samples of the discriminator network, the sample classification in the discriminator network is expressed as: Among them, Z (S) is the input of the discriminator network, Ω(·) is the feedforward neural network, are the training parameters of the discriminator in the source model, represents the output of the discriminator network, δ(·) is represented as the Sigmoid function; S3-2. Based on the output of the discriminator network, the pre-trained model is iteratively updated to minimize the difference between real samples and generated samples; S3-3. The loss function of the pre-trained model A-informer is in, is the prediction error term, To combat the loss term, Reconstruct the term for the input.
4. The method for predicting key parameters and components of a chemical process based on transfer learning and sparse self-attention according to claim 3 is characterized in that: In S3-3: Prediction Error Term for: Among them, t represents the tth sequence input, i represents the i-th y,seq (S) is the sequence length, pre (S) To predict the sequence length, represents the predicted value, is the true value; Enter the reconstruction item Among them, lab (S) Sequence length, Output the real value for the input, Reconstruct values for input and output; Adversarial Loss Term in is the kernel function A category of is a reproducing kernel Hilbert space, and MMD is the maximum mean difference.
5. The method for predicting key parameters and components of a chemical process based on transfer learning and sparse self-attention according to claim 1, characterized in that: The soft measurement model in S4 includes an encoder, a decoder and a discriminator, wherein: The encoder includes an encoding layer and a probabilistic sparse attention layer. The encoding layer is used to enhance the model's perception of position information, and the probabilistic sparse attention layer is used to process information interactions at different positions in the input sequence. The decoder includes an encoding layer, a probabilistic sparse attention layer, and a cross attention layer to obtain reconstructed samples; The discriminator is a feed-forward neural network that is used to minimize the difference between real and reconstructed samples through adversarial learning. Finally, the output of the decoder is input into a fully connected neural network layer FNN to obtain the prediction result.
6. The method for predicting key parameters and components of a chemical process based on transfer learning and sparse self-attention according to claim 1, characterized in that: The soft sensor model for predicting key parameter component data for offline analysis obtained in S4 specifically includes: S4-1. In the target domain, given the t-th sequence input is and where seq (T) is the encoder sequence length, lab (T) is the decoder sequence length, pre (T) To predict the sequence length, is the total dimension of the dataset, is the total dimension of y, the superscript T represents the target domain, and the target domain data is obtained after encoding and decoding S4-2. Freeze the FNN layer of the final output of the pre-trained model, randomly initialize the corresponding parameters, and obtain the prediction results as follows: in, It is initialized in the target domain. represents the FNN layer; S4-3. Select the MSE loss function to update the model parameters of the target domain data to obtain the final soft sensor model: Among them, t represents the tth sequence input, i represents the i-th y,seq (T) is the sequence length, pre (T) To predict the sequence length, represents the predicted value, is the true value.
Citation Information
Cited By
Soft measurement method and device for industrial multi-rate acquisition and medium
CN121352633A