A Dynamic Data Calibration Method for Improving the Prediction of Distributed Output in the Process
Through the DDR method combined with Bayesian formula and MAP estimation, the prediction error of distributed products in the chemical process is corrected, which solves the problem of prediction inaccuracy under the influence of measurement noise, and achieves higher prediction accuracy and model performance improvement.
Patent Information
- Application Number
- CN202210618056.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The online measurement of distributed product output during chemical industry has measurement noise interference, which affects the prediction accuracy, and existing filtering technologies cannot effectively correct the model prediction error.
The dynamic data correction method DDR is used, combining Bayesian formula and maximum posterior probability estimation MAP, and iterative calculation is used to correct the model prediction error and improve the accuracy of prediction output.
In the complex measurement environment, the accuracy of the prediction output of distributed products is improved, the measurement noise interference is reduced, and the performance of the prediction model is improved.
Smart Images

Figure CN114996938B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of chemical process quality modeling and prediction for distributed products, and particularly relates to a dynamic data correction method for improving the accuracy of predicted output of a distributed product model. Background Art
[0002] With the increasing demand for product diversification, the product outputs of some chemical processes exhibit distributed characteristics, such as the crystal size distribution in the crystallization process, the particle size distribution in the powder industry, the molecular weight distribution (MWD) in the polymerization process, etc. These output variables with distributed characteristics are directly related to the quality characteristics of products, such as strength characteristics, hardness characteristics, stress-strain characteristics, etc. At present, the on-line measurement of the distributed product output of chemical processes is a challenging problem. Therefore, to achieve efficient prediction of product output, data-driven modeling methods have been gradually applied to chemical production processes with distributed output characteristics.
[0003] In recent years, due to the rapid development of intelligent control technology, the neural network (NN) has also achieved good results in the process model learning of distributed outputs. However, the large training set and the generalization ability of the NN under a given modeling task are still the main problems faced by the NN. Compared with the NN, kernel learning (KL) based on support vector regression / least squares support vector regression (LSSVR) can use a limited training set to improve the generalization ability of the regression model, and is an effective solution for non-linear chemical process modeling. Based on the KL model, the proposal of just-in-time kernel learning (JKL) solves the problem of the decline in the model prediction performance of the KL single global model when the operating conditions, production batches and other working conditions change. In addition, the idea of ensemble learning is also used in the JKL modeling process. Based on the model error to define the weights of individual candidate models, the EJKL (Ensemble JKL) modeling method can adjust the model weights according to the accuracy of each model during the process of integrating candidate models. This integrated comprehensive model has more accurate prediction performance than a single regression model.
[0004] In the process of predicting the distributed product output based on the EJKL idea, even though the EJKL model is already quite accurate, the existence of prediction errors is still inevitable. In order to further correct the prediction results, it is necessary to collect output measurement information under offline conditions. However, considering the complexity of the environment, when the measurement conditions change, the measurement output may introduce a kind of measurement noise. The existence of measurement noise will affect the process monitoring of the system and also interfere with the control process of the system. Some classic filtering techniques, such as the Kalman Filter (KF), also have a certain effect on correcting the predicted output under the influence of measurement noise, but KF is only applicable to the state space model. Traditional digital filtering techniques are even more limited to the suppression of measurement noise and do not have a correction effect on the model predicted output.
[0005] Based on the above analysis, the present invention proposes a Dynamic Data Reconciliation (DDR) method to improve the distributed product output prediction performance of the EJKL model for the model prediction error of the distributed product prediction output under the influence of measurement noise. DDR can use inaccurate measurement information and prediction information to improve the accuracy of the predicted output and also reduce the interference of measurement noise on the measurement output. DDR is not limited to a specific model and can synthesize measurement information and model prediction information to obtain the optimal estimated value of the process output according to the Maximum a Posteriori Estimation (MAP) idea, thereby realizing the improvement of the accuracy of the distributed product prediction output. Summary of the Invention
[0006] The object of the present invention is to propose a dynamic data correction method for improving the product output prediction performance in the case of inaccurate measurement information for the prediction error existing in the soft sensor model of the distributed product output prediction based on EJKL.
[0007] The object of the present invention is achieved through the following technical solutions:
[0008] A dynamic data correction method for improving the distributed output prediction of a process, comprising the following steps:
[0009] (1) Establish a soft sensor prediction model:
[0010] Construct a prediction model, collect measurable variables that determine the process operating conditions as the input of the soft sensor model, and the predicted output of the model is the unmeasured distributed product output variable; divide the data composed of output variables and input variables measured offline into a training set and a test set;
[0011] (2) Prediction of the distributed product output:
[0012] Train the established prediction model. After training is completed, input measurable variables to predict the distribution shape of the product output online. Compare the model prediction results with the true output measured offline and analyze the reliability of the model;
[0013] (3) Iteratively calculate the model prediction error using the measurement information:
[0014] When the measurement environment changes and causes inaccurate measurement, iteratively calculate the model prediction error. After multiple iterations, when the prediction error converges, select the stabilized model prediction error value for the final prediction output correction;
[0015] (4) Online correct the prediction results according to the measurement information and prediction information:
[0016] Use the Bayesian formula for the measurement information and model prediction information to obtain the posterior distribution of the actual output based on the measurement information and prediction information; Solve the posterior distribution expression according to the maximum a posteriori probability estimation MAP idea to obtain the estimated value of the true output of the product; Use the true output estimated value as the dynamic data correction DDR for the correction result of the distributed product prediction output.
[0017] Furthermore, in the step (1), the least squares support vector regression method is used to construct a prediction model based on just-in-time kernel learning JKL. The specific process is as follows:
[0018] Step 1.1, establish M candidate JKL prediction models for the i-th sampling point obtained based on the training set. The expression formula is as follows:
[0019]
[0020] where, u q and y i respectively represent the measurable variables that determine the operation status of the process and the output variables measured offline, that is, the test set of the prediction model; x q,i and S qi are the query sample and the data set similar to x q,i in the database respectively; [C m , σ m are M pairs of candidate parameters;
[0021] Step 1.2: Apply the FLOO criterion to evaluate the feasibility of each candidate model. For x q,i , the FLOO error of each candidate model is expressed as:
[0022]
[0023] In the formula, l qi represents the number of the most similar samples defined according to the cumulative similarity factor. denotes the prediction error of the j-th sample calculated based on the FLOO criterion;
[0024] Step 1.3: Combine each JKL model using an ensemble strategy. The model weights defined by the ensemble strategy are:
[0025]
[0026] As can be seen from the above formula, the larger the FLOO error of a single model, the smaller the weight of the model; finally, in the case of N q sampling points, the final predicted distribution shape EJKL model is given:
[0027]
[0028] Furthermore, the process of step (3) is as follows:
[0029] Considering the influence of measurement noise, in the case of inaccurate measurement, the iterative method is used to correct the model prediction error;
[0030] Step 3.1: Take the difference between the actual measurement output y mea (t) containing measurement noise and the predicted output y ejkl (t) in step (2) as the initial value of the model prediction error, and set its covariance to Σ δ0 ;
[0031] Step 3.2: For the first iteration, the iterative process is as follows:
[0032] y ddr1 (t) = y ejkl (t) + (Σ ρ -1 + Σ δ0 -1 ) -1 Σ ρ -1 (y mea (t) - y ejkl (t))
[0033] In the formula, Σ ρ represents the covariance of the measurement noise; take the difference between y ddr1 (t) and y ejkl (t) as the first iteration result of the model prediction error, and set its covariance to Σ δ1 ;
[0034] Step 3.3: Assume that after n iterations, the model prediction error starts to converge, and set the covariance of the model prediction error convergence value to Σ δn , and use Σ δn for the calculation of the final predicted output correction.
[0035] Further, the process of step (4) is as follows:
[0036] Step 4.1: Use Bayes' formula for the measured output y mea (t) and the model predicted output y ejkl (t) to obtain the posterior distribution of the true output based on y ejkl (t), y mea (t):
[0037]
[0038] where N represents the vector dimension, y(t) represents the true output, Σ ρ and Σ δ are the covariance of the measurement noise and the model prediction error, respectively.
[0039] Step 4.2: According to the maximum a posteriori probability estimation MAP, the y(t) when the posterior distribution expression takes the maximum value is the estimated value y ddr (t) of the actual output; take the extreme value of the expression of the posterior distribution in step 4.1 to obtain the expression of the estimated value of y(t) as:
[0040] y ddr (t) = y ejkl (t) + (Σ ρ -1 + Σ δ -1 ) -1 Σ ρ -1 (y mea (t) - y ejkl (t));
[0041] Step 4.3: Let the error between the estimated value y ddr (t) and the actual value y(t) be ξ(t). It can be proved through mathematical derivation that:
[0042] tr{[cov[ξ(t)]]} < min{tr[Σ ρ , Σ δ}
[0043] The above formula shows that the variance of the estimated value error is less than the minimum of the measurement noise variance and the model prediction error variance, that is, the dynamic data correction DDR can achieve the correction of the model predicted output.
[0044] The beneficial effects of the present invention are as follows. Based on the measurement information contaminated by measurement noise and the prediction information of the EJKL soft sensor model, the DDR method is used to improve the accuracy of the predicted output of distributed products. First, the Bayesian formula is used for the measurement information and the model prediction information to obtain the posterior distribution of the actual output, and then the estimated value of the actual output is obtained based on the MAP idea. Compared with the measurement information and the model prediction information, the corrected output information is closer to the true state of the process. In the case of a complex measurement environment, the present invention improves the accuracy of the model predicted output while reducing the interference of measurement noise and improving the performance of the prediction model for the distributed product output process. Description of the Drawings
[0045] Figure 1 FIG. is a flowchart of the implementation of a dynamic data correction method for improving the distributed output prediction of a process proposed by the present invention;
[0046] Figure 2 is the iterative process of the model prediction error;
[0047] Figure 3 is the correction process of the DDR for the predicted output;
[0048] Figure 4 is the correction result of the MWD prediction for the polymerization process in the embodiment of the present invention (the output distribution and the output distribution error under the 23rd operating condition);
[0049] Figure 5 is the correction result of the MWD prediction for the polymerization process in the embodiment of the present invention (the output distribution and the output distribution error under the 27th operating condition). Detailed Embodiment
[0050] The present invention will be described in more detail with reference to the accompanying drawings of the present invention.
[0051] Embodiment: An embodiment for the MWD prediction of the styrene radical polymerization process is as Figure 1 follows. The present invention will be described in detail below:
[0052] (1) Establish a soft sensor prediction model
[0053] The monomer concentration of the inlet monomer feed stream, the concentration of the solvent in the solvent feed, the initial concentration of the initial feed stream, the volume flow rate, and the inlet feed temperature are collected as the inputs of the EJKL soft sensor model. The steady-state concentration of the polymer is the unmeasured distributed product output variable. 3000 sets of input and output variables measured offline, with every 100 sets of data as one operating condition. Among them, 2000 sets of data under 20 operating conditions are used for the training of the EJKL model, and another 1000 sets of data under 10 operating conditions are used for the testing of the model.
[0054] (2) Prediction of Distributed Product Output
[0055] After the prediction model is trained, the input variables of 10 operating conditions (21st - 30th) are used to predict the distribution shape of the product output online. The model prediction results under the 23rd and 27th operating conditions are randomly selected and compared with the true output measured offline to analyze the reliability of the model.
[0056] (3) Iterative Calculation of Model Prediction Error Using Measurement Information
[0057] Considering the influence of measurement noise, an iterative method is used to correct the model prediction error when the measurement is inaccurate.
[0058] Step 3.1:
[0059] Figure 2 is the algorithm flow for iterative prediction error. Here, var(x) represents taking the variance of x, that is, the trace of the covariance matrix. Since the true output is unknown, the measurement information is used to represent the true information during the iteration process. The prerequisite for the measurement information to have reference value is that the prediction error can converge during the iterative calculation. Therefore, on the premise that the iterative calculation does not diverge, the measurement noise variances under the 23rd and 27th operating conditions are set to 5×10 -11 and 9×10 -10 . The difference between the actual measured output y mea (t) containing measurement noise and the predicted output y ejkl (t) in step (2) is used as the initial value of the model prediction error, and its covariance is set to Σ δ0 .
[0060] Step 3.2:
[0061] The first iteration process is as follows:
[0062] y ddr1 (t) = y ejkl (t) + (Σ ρ -1 + Σ δ0 -1 ) -1 Σ ρ -1 (y mea (t) - y ejkl (t))
[0063] In the formula, Σ ρ represents the covariance of the measurement noise. The difference between y ddr1 (t) and y ejkl (t) is taken as the first iteration result of the model prediction error, and its covariance is set to Σ δ1 .
[0064] Step 3.3:
[0065] During the iteration process, if the prediction errors of 5 iterations do not change, the iterative calculation is regarded as having converged. After 14 and 24 iterations respectively, the variances of the model prediction errors under the 23rd and 27th operating conditions converge to 2.9225×10 -10 , 2.1380×10 -9 , and the converged prediction errors are used for the calculation of the final prediction output correction.
[0066] The algorithm of the above iterative process is exactly based on the DDR principle, and the principle of DDR is described in detail in step (4).
[0067] (4) Online correct the prediction result according to the measurement information and the prediction information
[0068] Figure 3 This is the correction process of DDR for the prediction output. ε(t) and δ(t) represent the measurement noise and the model prediction error respectively. L(y(t)|y mea (t)) and L(y(t)|y ejkl (t)) represent the likelihood functions of y(t) based on the measurement information and the model prediction information respectively. p(y(t)) is the prior distribution of y(t). Using the Bayesian formula for the measurement information and the model prediction information, the posterior distribution of the actual output based on the measurement information and the prediction information is obtained; according to the MAP idea, solving the posterior distribution expression can obtain the estimated value y ddr (t) of the true output of the product, where K = (Σ ρ -1 + Σ δ -1 ) -1 Σ ρ -1 ; The estimated value of the true output is used as the correction result of DDR for the distributed product prediction output.
[0069] Step 4.1:
[0070] Using the Bayesian formula for the measurement output y mea (t) and the model prediction output y ejkl (t) to obtain the posterior distribution of the true output based on y ejkl (t), y mea (t):
[0071]
[0072] where N represents the vector dimension, y(t) represents the true output, Σ ρ is the covariance of the measurement noise, Σ δRepresents the covariance of the model prediction error after convergence of iterative calculation. Under the 23rd operating condition, Σ ρ and Σ δ have traces of 5×10 -11 , 2.9225×10 -10 respectively; under the 27th operating condition, Σ ρ and Σ δ have traces of 9×10 -10 , 2.1380×10 -9 .
[0073] Step 4.2:
[0074] According to MAP, the y(t) when the posterior distribution expression takes the maximum value is the estimated value ŷ ddr (t) of the actual output. Taking the extreme value of the above formula, the expression of y(t) is obtained as:
[0075] ŷ ddr (t) = ŷ ejkl (t) + (Σ ρ -1 + Σ δ -1 ) -1 Σ ρ -1 (y mea (t) - ŷ ejkl (t))
[0076] Step 4.3:
[0077] Let the error between the estimated value ŷ ddr (t) and the actual value y(t) be ξ(t). It can be proved through mathematical derivation that:
[0078] tr{cov[ξ(t)]} < min{tr[Σ ρ , Σ δ}
[0079] The above formula indicates that the variance of the estimated value error (the trace of the covariance matrix) is less than the minimum of the measurement noise variance and the model prediction error variance, that is, DDR can achieve the correction of the model prediction output.
[0080] Table 1 Comparison of RMSE of distribution shape error before and after DDR correction
[0081]
[0082] Figure 4 and Figure 5They are the correction results of the DDR for the predicted output under the 23rd and 27th operating conditions respectively. Table 1 shows the calculation results of the Root Mean Squared Error (RMSE) of the output distribution shape error. From Figure 4 , Figure 5 and the comparison with the results in Table 1, it can be found that the predicted output distribution shape of the product based on EJKL is roughly the same as the true distribution shape, but there are still errors in the predicted shape in some areas. After correcting the predicted output using the DDR method proposed in the present invention, the prediction error is significantly reduced, and the corrected predicted distribution is closer to the true distribution shape. This shows that DDR can effectively improve the prediction error caused by inaccurate prediction models, making the predicted product distribution closer to the true state of the process, and has application value in process monitoring and product quality analysis.
[0083] Based on the measurement information contaminated by measurement noise and the prediction information of the EJKL soft measurement model, the present invention uses the DDR method to improve the accuracy of the distributed product predicted output. This method first uses the Bayesian formula for the measurement information and the model prediction information to obtain the posterior distribution of the actual output, and then obtains the estimated value of the actual output based on the MAP idea. Compared with the measurement information and the model prediction information, the corrected output information is closer to the true state of the process. In the case of a relatively complex measurement environment, the present invention improves the accuracy of the model predicted output while reducing the interference of measurement noise and improving the performance of the distributed product output prediction model.
[0084] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that those skilled in the art can think of according to the inventive concept.
Claims
1. A dynamic data correction method for improving the distributed output prediction of a process, characterized in that Including the following steps: (1) Establish a soft-sensing prediction model: Construct a prediction model, collect measurable variables that determine the process operating conditions as the input of the soft-sensing model, and the predicted output of the model is the unmeasured distributed product output variable; divide the data composed of the output variable and the input variable measured offline into a training set and a test set; (2) Prediction of distributed product output: Train the established prediction model. After the training is completed, input the measurable variables to predict the distribution shape of the product output online, compare the model prediction result with the true output measured offline, and analyze the reliability of the model; (3) Iteratively calculate the model prediction error using measurement information: When the measurement environment changes and causes inaccurate measurement, iteratively calculate the model prediction error. After multiple iterations, when the prediction error converges, select the stabilized model prediction error value for the final prediction output correction; (4) Online correct the prediction result according to the measurement information and the prediction information: Use the Bayesian formula for the measurement information and the model prediction information to obtain the posterior distribution of the actual output based on the measurement information and the prediction information; solve the posterior distribution expression according to the maximum a posteriori probability estimation (MAP) idea to obtain the estimated value of the true product output; use the estimated value of the true output as the dynamic data correction (DDR) for the correction result of the distributed product prediction output; In step (1), the least squares support vector regression method is used to construct a prediction model based on just-in-time kernel learning (JKL). The specific process is as follows: Step 1.1, establish M candidate JKL prediction models for the i-th sampling point obtained based on the training set, and the expression formula is as follows: wherein, u q and y i respectively represent the measurable variables determining the operation status of the process and the output variables measured offline, i.e., the test set of the prediction model; x q,i and S qi are respectively the query sample and the data set similar to x q,i in the database; [C m , σ m is the M pairs of candidate parameters; Step 1.2: Apply the FLOO criterion to evaluate the feasibility of each candidate model. For x q,i , the FLOO error of each candidate model is expressed as: where l qi represents the number of the most similar samples defined according to the cumulative similarity factor, represents the prediction error of the j-th sample calculated based on the FLOO criterion; Step 1.3: Combine each JKL model using an ensemble strategy, and the model weights defined based on the ensemble strategy are: As can be seen from the above formula, the larger the FLOO error of a single model, the smaller the weight of the model; finally, in the case of having N q sampling points, the final distribution shape prediction EJKL model is given:
2. The dynamic data correction method for distributed output prediction in the lifting process according to claim 1, wherein The process of step (3) is: Considering the influence of measurement noise, in the case of inaccurate measurement, use the iterative method to correct the model prediction error; Step 3.
1. Take the difference between the actual measurement output y mea (t) containing measurement noise and the predicted output y ejkl (t) in step (2) as the initial value of the model prediction error, and set its covariance to Σ δ0 ; Step 3.2, the first iteration, the iteration process is: where, Σ ρ represents the covariance of the measurement noise; take the difference between y ddr1 (t) and y ejkl (t) as the first iteration result of the model prediction error, and set its covariance to Σ δ1 ; Step 3.3: Assume that after n iterations, the model prediction error begins to converge, and assume that the covariance of the model prediction error convergence value is Σ δn , and use Σ δn for the calculation of the final prediction output correction.
3. The dynamic data correction method for distributed output prediction in the lifting process according to claim 1, characterized in that The process of step (4) is: Step 4.
1. For the measurement output y mea (t) and the model prediction output y ejkl (t), use Bayes' formula to obtain the posterior distribution of the true output based on y ejkl (t), y mea (t): where N represents the vector dimension, y(t) represents the true output, Σ ρ and Σ δ are the covariance of the measurement noise and the model prediction error, respectively; Step 4.2: According to the maximum a posteriori probability estimation MAP, the y(t) when the posterior distribution expression takes the maximum value is the estimated value y of the actual output ddr (t); Taking the extreme value of the expression of the posterior distribution in Step 4.1, the expression for the estimated value of y(t) is obtained as follows: y ddr y(t) = y ejkl (t)+(Σ ρ -1 +Σ δ -1 ) -1 Σ ρ -1 (y mea (t)-y ejkl (t)); Step 4.3: Let the error between the estimated value y ddr (t) and the actual value y(t) be ξ(t). It can be proven through mathematical derivation that: tr{[cov[ξ(t)]]} < min{tr[Σ ρ , Σ δ} The above formula shows that the variance of the estimated value error is less than the minimum of the measurement noise variance and the model prediction error variance, that is, the dynamic data correction (DDR) can achieve the correction of the model prediction output.
Citation Information
Patent Citations
Coal mill primary air volume soft measurement system based on LSSVM (least square support vector machine) and self-adaptive recursion
CN113343178A
Polymerization process molecular weight distribution prediction method based on integrated probability modeling
CN113658646A