A soft measurement modeling method based on a VMD-BiLSTMA-GPR model
Patent Information
- Application Number
- CN202511613618.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-11-06
AI Technical Summary
尽管现有技术中存在一些将深度学习与概率模型结合的尝试,但面对复杂信号中混叠的多种频率成分和强烈非线性特征,模型的预测精度和区间可靠性仍面临巨大挑战
提升了对复杂非线性信号的刻画能力与预测精度:通过引入变分模态分解(VMD)作为前置处理单元,本发明能够将原始非平稳、非线性的目标变量序列自适应地分解为一系列相对平稳、窄带的子序列(本征模态函数)。这一过程有效解耦了原始信号中的复杂特征,显著降低了后续建模的难度,为获得高精度的点预测奠定了坚实基础。
Smart Images

Figure CN121435190B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a soft measurement modeling method based on the VMD-BiLSTMA-GPR model. Background Technology
[0002] Currently, in many complex processes such as industrial production and environmental monitoring, there are numerous key variables (such as chemical component concentration and pollutant indicators) that are closely related to product quality and system status. However, due to technological or economic limitations, these key variables are often difficult or impossible to measure directly in real time and online using sensors (i.e., "hard measurement"). They usually have to rely on time-consuming and labor-intensive offline laboratory analysis, resulting in a serious lag in monitoring information and failing to meet the needs of real-time control and optimization decision-making.
[0003] To address this issue, soft measurement technology has emerged. It is an effective method for indirectly estimating target variables by establishing a mathematical model between easily measurable process variables (auxiliary variables) and difficult-to-measurable key variables (dominant variables). Traditional soft measurement modeling methods, such as those based on partial least squares or support vector machines, often fall short in handling highly nonlinear and strongly coupled industrial processes, exhibiting limited generalization ability. In recent years, deep learning models, particularly recurrent neural networks (RNNs) and their variant, long short-term memory networks (LSTMs), have been applied to soft measurement due to their powerful ability to capture temporal features. However, single LSTM models may still suffer from information forgetting or gradient problems when processing long sequences, and their predictions are typically single-point predictions, failing to provide quantitative information about the uncertainty of the prediction.
[0004] In practical engineering applications, simply providing a single-point prediction is far from sufficient. Decision-makers are more interested in understanding the reliability of the prediction, i.e., the range of fluctuation in the prediction result. Although there are some attempts in existing technologies to combine deep learning with probabilistic models, the prediction accuracy and interval reliability of models still face significant challenges when dealing with the aliased multiple frequency components and strong nonlinear characteristics of complex signals. Specifically, existing technologies lack an integrated solution that can first adaptively decompose the non-stationary target variable signal to simplify modeling, then comprehensively utilize deep networks for accurate point prediction, and finally provide highly reliable interval prediction. Summary of the Invention
[0005] The purpose of this invention is to provide a soft measurement modeling method based on the VMD-BiLSTMA-GPR model to solve the problems mentioned in the background art.
[0006] In a first aspect, embodiments of the present invention provide a soft measurement modeling method based on the VMD-BiLSTMA-GPR model, the method comprising the following steps: S1. Data acquisition steps: Acquire direct measurement data of process variables and offline detection data of target variables; S2, VMD decomposition steps: Perform variational mode decomposition on the data of the target variable to obtain a series of intrinsic mode function subsequences and a residual sequence; S3. BiLSTMA modeling steps: A fusion model of bidirectional long short-term memory network and self-attention mechanism is used to train and predict each of the intrinsic mode function subsequences and residual sequences obtained in step S2; S4. Result reconstruction step: The prediction results of each sequence obtained in step S3 are superimposed to reconstruct the point prediction value of the target variable. S5, GPR Interval Prediction Steps: Based on the point prediction values obtained in step S4, the Gaussian process regression algorithm is used to probabilistically model the prediction residuals and output the interval prediction results with specific confidence intervals.
[0007] Optionally, in step S2, the variational mode decomposition is used to decompose the non-stationary target variable data into multiple sequences of intrinsic mode functions with stationary modes.
[0008] Optionally, in step S3, the bidirectional long short-term memory network and self-attention mechanism fusion model includes a bidirectional long short-term memory network layer that processes temporal information bidirectionally, and a self-attention mechanism layer that weights the output of the bidirectional long short-term memory network layer with key information.
[0009] Optionally, the bidirectional long short-term memory network layer models sequence information from both forward and backward directions through the structure of forget gate, input gate, cell state, and output gate.
[0010] Optionally, the self-attention mechanism layer calculates the query matrix, key matrix, and value matrix, and generates attention weights based on the similarity between the query matrix and the key matrix, so as to obtain the output by weighted summation of the value matrix.
[0011] Optionally, in step S4, the reconstruction is achieved by directly adding the predicted values of each of the intrinsic mode function subsequences to the predicted values of the residual sequence to restore the original point prediction values of the target variable.
[0012] Optionally, in step S5, the Gaussian process regression algorithm assumes that the predicted residuals follow a Gaussian distribution, and determines the confidence interval by calculating the posterior distribution mean and variance of the predicted point values.
[0013] Optionally, the confidence interval is the range of predicted value fluctuations calculated based on the mean and standard deviation of the posterior distribution at different confidence levels.
[0014] Optionally, the method is used for soft measurement of variables that are difficult to measure directly in real time in industrial processes or environmental monitoring.
[0015] Optionally, the method can handle high-dimensional complex system data with nonlinear relationships between variables, and simultaneously provide high-precision point predictions and interval predictions with probabilistic interpretations.
[0016] The present invention has achieved the following beneficial effects: This invention enhances the ability to characterize and predict complex nonlinear signals: By introducing Variational Mode Decomposition (VMD) as a pre-processing unit, it adaptively decomposes the original nonstationary, nonlinear target variable sequence into a series of relatively stationary, narrow-band subsequences (eigenmode functions). This process effectively decouples the complex features in the original signal, significantly reduces the difficulty of subsequent modeling, and lays a solid foundation for obtaining high-precision point predictions.
[0017] This approach achieves deep mining and efficient utilization of temporal contextual information: a fusion model of bidirectional long short-term memory network and self-attention mechanism (BiLSTMA) is used as the core predictor. The BiLSTM layer can fully capture the long-term dependencies of the sequence in both forward and backward directions, while the self-attention mechanism can dynamically assign different weights to features at different time steps, highlighting key information and suppressing redundant noise. The fusion of the two makes the model's learning of temporal dynamic characteristics more comprehensive and accurate. As shown in Table 1, its point prediction accuracy is significantly better than that of a single LSTM or BiLSTM model.
[0018] This model provides interval prediction results with clear probabilistic significance, enhancing its reliability and practicality. Innovatively, in the Gaussian Process Regression (GPR) layer, it probabilistically models the residuals of BiLSTMA point predictions, not the original target variable. This fully leverages the inherent advantages of GPR in handling uncertainty, ultimately outputting prediction intervals with specific confidence levels (e.g., 90%, 95%). As shown in Table 3, this model achieves high interval coverage while maintaining reasonable interval width, enabling users to not only obtain predicted values but also clearly understand the range of uncertainty in the predictions, providing crucial information for risk assessment and decision-making.
[0019] This creates a synergistic technology chain of "decomposition-point prediction-interval calibration": VMD, BiLSTMA, and GPR are not simply stacked together, but form an organic whole. VMD decomposition creates the conditions for high-precision prediction by BiLSTMA; the high-precision point prediction results of BiLSTMA, in turn, make its residual sequence closer to stationary random noise, thus meeting the ideal application prerequisites of GPR, making the final interval prediction more accurate and reliable. This synergistic effect makes this invention particularly suitable for soft measurement tasks in complex industrial processes that are high-dimensional, nonlinear, and strongly coupled.
[0020] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0022] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a soft measurement modeling method based on the VMD-BiLSTMA-GPR model according to the present invention. Figure 2 This is a schematic diagram of the VMD decomposition results of the COD index. Figure 3 This is a schematic diagram of the BiLSTM model structure; Figure 4 A schematic diagram of the COD point prediction results; Figure 5 This is a schematic diagram of the COD interval prediction results. Detailed Implementation
[0023] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0024] Figure 1 A flowchart of a soft measurement modeling method based on the VMD-BiLSTMA-GPR model is provided as an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps: S1. Data acquisition steps: Acquire direct measurement data of process variables and offline detection data of target variables; S2, VMD decomposition steps: Perform variational mode decomposition on the data of the target variable to obtain a series of intrinsic mode function subsequences and a residual sequence; S3. BiLSTMA modeling steps: A fusion model of bidirectional long short-term memory network and self-attention mechanism is used to train and predict each of the intrinsic mode function subsequences and residual sequences obtained in step S2; S4. Result reconstruction step: The prediction results of each sequence obtained in step S3 are superimposed to reconstruct the point prediction value of the target variable. S5, GPR Interval Prediction Steps: Based on the point prediction values obtained in step S4, the Gaussian process regression algorithm is used to probabilistically model the prediction residuals and output the interval prediction results with specific confidence intervals.
[0025] In step S2, the variational mode decomposition is used to decompose the non-stationary target variable data into multiple intrinsic mode function subsequences with stationary modes.
[0026] In step S3, the bidirectional long short-term memory network and self-attention mechanism fusion model includes a bidirectional long short-term memory network layer that processes temporal information bidirectionally, and a self-attention mechanism layer that weights the output of the bidirectional long short-term memory network layer with key information.
[0027] The bidirectional long short-term memory network layer models sequence information from both forward and backward directions through the structure of forget gate, input gate, cell state, and output gate.
[0028] The self-attention mechanism layer calculates the query matrix, key matrix, and value matrix, and generates attention weights based on the similarity between the query matrix and the key matrix, and then performs a weighted summation on the value matrix to obtain the output.
[0029] In step S4, the reconstruction is achieved by directly adding the predicted values of each intrinsic mode function subsequence to the predicted values of the residual sequence to restore the original point prediction values of the target variable.
[0030] In step S5, the Gaussian process regression algorithm assumes that the predicted residuals follow a Gaussian distribution and determines the confidence interval by calculating the posterior distribution mean and variance of the predicted point values.
[0031] The confidence interval is the range of predicted value fluctuations calculated based on the mean and standard deviation of the posterior distribution at different confidence levels.
[0032] The method is used for soft measurement of variables that are difficult to measure directly in real time in industrial processes or environmental monitoring.
[0033] The method can handle high-dimensional complex system data with nonlinear relationships between variables, and simultaneously provides high-precision point predictions and interval predictions with probabilistic interpretations.
[0034] The working principle of the technical solution of this invention is as follows: First, the system acquires directly and quickly measurable process variable data (such as total phosphorus, pH, temperature, and pressure) from the data acquisition system at the production or monitoring site. Simultaneously, it collects target variable data (such as chemical oxygen demand, COD) through offline laboratory analysis, together forming the historical dataset required for modeling. Then, the core model of this invention begins operation: First, the Variational Mode Decomposition (VMD) module performs adaptive signal processing on the non-stationary target variable time series data. By solving a constrained variational problem, it accurately decomposes the originally complex and fluctuating signal, which is difficult to model directly, into multiple relatively stationary intrinsic mode function (IMF) subsequences with different center frequencies and a residual sequence representing the trend. This decomposition process effectively removes aliasing features from the signal, transforming the complex global modeling problem into multiple simpler sub-problems, laying a solid foundation for subsequent accurate prediction. The second step involves the BiLSTMA prediction module, which performs parallel modeling on each VMD decomposition subsequence (including all IMFs and residual sequences). Within this module, a Bidirectional Long Short-Term Memory (BiLSTM) network operates first. Its internal gating mechanism (forget gate, input gate, and output gate) uses a sophisticated combination of sigmoid and tanh functions to deeply learn and memorize the temporal dynamics of each subsequence from both forward and backward time dimensions, thus comprehensively capturing its contextual dependencies. Then, a self-attention mechanism processes the hidden state sequence output by BiLSTM, dynamically generating attention weights by calculating the similarity between the query, key, and value matrices. This assigns different importance to features at different time points, focusing on and strengthening key time-point information while suppressing noise interference, ultimately outputting a high-precision predicted value for each subsequence. The third step involves the result reconstruction module linearly superimposing the BiLSTMA prediction results of all the subsequences to restore the final predicted value for the original target variable. The fourth step involves the Gaussian Process Regression (GPR) interval calibration module, which probabilistically models the residual between the BiLSTMA point predictions and the true values. GPR assumes that this residual follows a Gaussian process, calculates the covariance between training and test samples using a kernel function, and then derives the posterior probability distribution of the predicted values at a given confidence level. The center (mean) of this distribution can be used for fine-tuning the point prediction, while its standard deviation is directly used to calculate the prediction interval. Ultimately, the model outputs a prediction interval centered on the point prediction with certain upper and lower bounds (e.g., a 95% confidence interval), thus clearly and quantitatively demonstrating the uncertainty of the prediction results. This entire workflow forms a complete technical loop from "signal preprocessing -> deep feature learning and accurate point prediction -> residual probability analysis and interval generation," giving this invention unique advantages of high accuracy, high reliability, and strong interpretability when dealing with complex industrial data.
[0035] Specifically, the technical solution of the present invention includes the following steps when implemented: The example used in this invention is the prediction of the chemical oxygen demand (COD) index in river water quality monitoring.
[0036] S101: Process variable data that can be directly obtained using a water quality testing system The process variables include total phosphorus (TP), total nitrogen (TN), dissolved oxygen (DO), and pH. Data for the target variable (COD) to be modeled are then collected using offline detection methods. m is the number of samples, n is the number of process variables, and R is a real number. Usually, the index only has positive numbers.
[0037] S102: The COD data is decomposed using VMD (Variational Mode Decomposition). The objective variable is decomposed into multiple stationary mode subsequences IMF1, IMF2, IMF3 and residual res sequences. The res sequence does not have stationary modes. The decomposition results are as follows: Figure 2 As shown. The VMD decomposition method in step S102 is shown in formula (1).
[0038] in represent A set of modalities. represent A set of center frequencies It is the exponential term that shifts the modal spectrum to the fundamental frequency. It is the original input signal. is the penalty factor used to balance bandwidth minimization and reconstruction error, K is the number of modes in variational mode decomposition, i.e., the number of intrinsic mode functions (IMFs) to be decomposed, δ(t) is the Dirac delta function, i.e., the unit impulse function, j is the imaginary unit, satisfying j² = -1, and t is the time variable.
[0039] S103: Each decomposed subsequence (including the res sequence) is modeled using a two-layer bidirectional long short-term memory network and self-attention mechanism fusion model (BiLSTMA).
[0040] In step S103, the BiLSTMA model is a fusion of BiLSTM and an attention mechanism. The BiLSTM module consists of a forget gate, an input gate, a cell state gate, and an output gate. The BiLSTM module has two LSTM networks with opposite timings, processing sequence information in the forward and backward directions respectively. The result of the BiLSTM is as follows... Figure 3 As shown. Step S103 can be further divided into five sub-steps, S201-S205.
[0041] S201: Forget Gate: Calculates the information to be forgotten based on the cell state of the previous layer. Some information in the data is partially forgotten and retained with a certain probability, and the output of the previous unit is used for further processing. and the current input sequence Through the sigmoid activation function The forgetting factor was calculated. The probability, with a value between 0 and 1, represents the probability of forgetting previous information; 1 represents complete retention, and 0 represents complete discarding. Its mathematical expression is: in It is the weight matrix of the forget gate. Offset of the forget gate S202: Input gate: Calculates the information to be retained, divided into two parts: the sigmoid function activation part and the tanh function learning part, whose mathematical expression is as follows: in and This determines the new information the network learns at any given moment. and Here is the weight matrix of the input gate. and The bias term represents the input gate.
[0042] S203 cell state update involves multiplying the old cell state by a forgetting factor to obtain the retained information. In addition, the new information learned by the network at this moment This gives us the current cell state. As shown in the following formula: Where ⊙ represents the Hadamard product, which is the element-wise multiplication of a matrix or vector; S204: Output gate: Acts on the current hidden state ht of the output cell, determining the current cell state through a sigmoid layer and a tanh layer. The information that needs to be output is expressed mathematically as follows: in This is the weight matrix of the output gate. This is the bias term for the output gate.
[0043] S205: Attention mechanism, which processes the output of BiLSTM. The BiLSTM and attention mechanism can be combined and nested multiple times.
[0044] in It is the input sequence matrix. These are the weight matrices for the query, key, and value matrices, respectively.
[0045] in It is the first of the attention weight matrix. Line 1 Column elements, softmax is the normalization exponential function used to transform a set of values into a probability distribution, d k is the dimension of the key vector K, used to scale the dot product and prevent gradient vanishing. Value matrix The The line represents the input sequence. A vector of values at each position, This is the final output value.
[0046] Table 1 Prediction results of different benchmark models The evaluation metrics in Table 1 show that BiLSTMA has higher point prediction accuracy than both BiLSTM and LSTM, indicating that bidirectional LSTM and the attention mechanism significantly improve the accuracy of point prediction.
[0047] S104: BiLSTM predicts all subsequences separately, and then reconstructs the results as shown in formula (11). The reconstructed results of each subsequence prediction are as follows: Figure 4 As shown.
[0048] Table 2. Impact of VMD Sequence Decomposition on Model Prediction Performance The evaluation metrics in Table 2 show that the use of VMD sequence decomposition significantly improves the accuracy of point prediction.
[0049] S105: The specific method for probabilistically modeling the prediction residuals using the GPR algorithm is as follows: assuming that the sample data is affected by additive noise. Pollution, observations in the training set y As shown in formula (12): f(x) is a latent, noise-free true function; Further assumptions After a series of changes, the prior distribution of the observed value y and the observed value can be obtained. and predicted value The joint prior distribution is shown in equations (13)-(16): K(x i , x j ) is the kernel function used to calculate the input x. i and x j Covariance between; I n X is an n×n identity matrix, where n is the number of training samples; * x is the input data matrix for the test set; i Let be the i-th training input data point, where i = 1, 2, ..., N, and N is the total number of training samples.
[0050] The covariance of the test set itself, the predicted value The posterior distribution is shown in formulas (17)-(19): As the point prediction results of the GPR algorithm, the interval prediction results are predicted here at 90%, 60%, and 30% confidence levels, respectively. . 、 and . No. i The probability density function of the predicted value in each period is shown in equation (20). The final prediction results for different confidence intervals are as follows: Figure 5 As shown.
[0051] Table 3 Interval Assessment at Different Confidence Levels The indicators in Table 3 show that the model can effectively capture the distribution of true values while controlling the width of the interval.
[0052] This example further illustrates that the present invention has high prediction accuracy and can adapt to high-dimensional and complex data. Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of the present invention and their equivalents, the present invention also intends to include these modifications and variations.
Claims
1. A soft sensor modeling method based on the VMD-BiLSTMA-GPR model, characterized in that, The method includes the following steps: S1. Data acquisition steps: Acquire direct measurement data of process variables and offline detection data of target variables; S2, VMD decomposition steps: Perform variational mode decomposition on the data of the target variable to obtain a series of intrinsic mode function subsequences and a residual sequence; S3. BiLSTMA modeling steps: A fusion model of bidirectional long short-term memory network and self-attention mechanism is adopted to train and predict each of the intrinsic mode function subsequences and residual sequences obtained in step S2; the fusion model of bidirectional long short-term memory network and self-attention mechanism includes a bidirectional long short-term memory network layer that processes temporal information bidirectionally, and a self-attention mechanism layer that weights the output of the bidirectional long short-term memory network layer with key information. S4. Result reconstruction step: The prediction results of each sequence obtained in step S3 are superimposed to reconstruct the point prediction value of the target variable. S5, GPR Interval Prediction Step: Based on the point prediction values obtained in step S4, the Gaussian process regression algorithm is used to probabilistically model the prediction residuals and output the interval prediction results with specific confidence intervals; the Gaussian process regression algorithm assumes that the prediction residuals follow a Gaussian distribution and determines the confidence interval by calculating the posterior distribution mean and variance of the point prediction values.
2. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, In step S2, the variational mode decomposition is used to decompose the non-stationary target variable data into multiple intrinsic mode function subsequences with stationary modes.
3. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, The bidirectional long short-term memory network layer models sequence information from both forward and backward directions through the structure of forget gate, input gate, cell state, and output gate.
4. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, The self-attention mechanism layer calculates the query matrix, key matrix, and value matrix, and generates attention weights based on the similarity between the query matrix and the key matrix, and then performs a weighted summation on the value matrix to obtain the output.
5. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, In step S4, the reconstruction is achieved by directly adding the predicted values of each intrinsic mode function subsequence to the predicted values of the residual sequence to restore the original point prediction values of the target variable.
6. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, The confidence interval is the range of predicted value fluctuations calculated based on the mean and standard deviation of the posterior distribution at different confidence levels.
7. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, The method is used for soft measurement of variables that are difficult to measure directly in real time in industrial processes or environmental monitoring.
8. The soft sensor modeling method based on the VMD-BiLSTMA-GPR model as described in claim 1, characterized in that, The method can handle high-dimensional complex system data with nonlinear relationships between variables, and simultaneously provides high-precision point predictions and interval predictions with probabilistic interpretations.