Sulfur recovery unit soft-sensing method based on bayesian regularization
By using a Bayesian regularized probabilistic principal component regression modeling method, the number of principal components is automatically determined, which solves the problem of inappropriate selection of the number of principal components in traditional principal component regression analysis in sulfur recovery devices and improves the online detection effect of key indicators.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-26
- Publication Date
- 2026-03-03
AI Technical Summary
Traditional principal component regression analysis lacks the ability to automatically select the number of principal components in sulfur recovery units, which affects the stability and detection performance of soft measurement models.
A probabilistic principal regression modeling method based on Bayesian regularization is adopted. The number of principal components in the principal regression model is automatically determined by the Bayesian regularization method, and a probabilistic principal regression soft measurement model based on Bayesian regularization is established.
This method improves the online detection effect and performance of key indicators in sulfur recovery units, enabling online estimation of hydrogen sulfide and sulfur dioxide content, and overcoming the shortcomings of traditional methods.
Smart Images

Figure CN116935992B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of soft measurement modeling and application in chemical production processes, and specifically relates to a soft measurement modeling and online detection method for sulfur recovery devices based on Bayesian regularization. Background Technology
[0002] As a traditional recovery device, sulfur recovery units can recover toxic and harmful substances from acidic gas streams, preventing pollution of the surrounding environment due to large-scale emissions. Typically, sulfur recovery units receive two different gas streams: one rich in hydrogen sulfide, and the other rich in both hydrogen sulfide and ammonia. These gases are then incinerated in the sulfur recovery unit and converted into the final product, sulfur, through subsequent cooling and other processes.
[0003] Throughout the process, effective control of the sulfur conversion requires the measurement of two key indicators: hydrogen sulfide and sulfur dioxide content. In the absence of analytical instruments, a data-driven soft sensor model needs to be established based on the relationship between other process variables and these two key indicators to perform online measurements. Principal component regression analysis is a widely used soft sensor modeling method; however, this method lacks the ability to automatically select the number of principal components. In practical applications, the number of principal variables usually needs to be determined manually, which significantly affects the stability and detection performance of the soft sensor model. Summary of the Invention
[0004] The purpose of this invention is to address the difficulty of real-time detection of key variables in sulfur recovery devices by providing a probabilistic principal component regression modeling and online detection method based on Bayesian regularization.
[0005] The objective of this invention is achieved through the following technical solution:
[0006] A soft measurement method for a sulfur recovery device based on Bayesian regularization includes: online detection of process variable data of the sulfur recovery device, and preprocessing and normalizing the data; inputting the processed process variable data into a constructed probabilistic principal component regression soft measurement model based on Bayesian regularization to obtain the key indicator values corresponding to the online process variable data.
[0007] During the training process, the probabilistic principal component regression soft measurement model based on Bayesian regularization automatically determines the number of principal components in the principal component regression model using the Bayesian regularization method.
[0008] As a technical solution, the Bayesian regularization-based probabilistic principal component regression soft measurement model is constructed using the following method:
[0009] (1) Obtain process variable data during the normal operation of the sulfur recovery device by using a distributed control system, obtain key indicator values corresponding to the process variable data by offline means, and form a training data sample set for modeling.
[0010] (2) Preprocess and normalize the process variables and key indicator samples in the training data sample set respectively;
[0011] (3) Using the preprocessed and normalized process variable data as the input dataset and the preprocessed and normalized key indicator values as the output dataset, a probabilistic principal component regression soft measurement model based on Bayesian regularization is established.
[0012] In step (1), the process variable data of the sulfur recovery unit are collected using the distributed control system to form the input training data sample set for modeling: X∈R n×m Where n is the number of samples and m is the number of process variables, the dataset is stored in a database for later use. Then, by sampling on-site and conducting offline laboratory analysis, the key indicator values corresponding to the process variable data samples used for modeling in the historical database are obtained, which serve as the training sample set Y∈R for the soft sensor model output. n×r Where n is the number of samples and r is the number of key indicators in the sulfonium recovery unit, the dataset is stored in a database for later use. Using the above method, the original datasets X and Y for modeling are obtained.
[0013] In step (2), the process variable data and key indicator value samples are preprocessed and normalized respectively, so that the mean of each process variable and key indicator is zero and the variance is 1, resulting in a new data matrix set. and This step yields the dataset used for modeling. and
[0014] For the input and output datasets of the normalized soft measurement model, a probabilistic principal component regression soft measurement model based on Bayesian regularization is established, and the parameters of the model are stored in the model database for later use.
[0015] In step (3), the normalized process variable matrix is... As input to the soft measurement model, the key indicator data matrix As the output of the soft measurement model, the following probabilistic principal component regression soft measurement model is established, namely, the Bayesian regularized probabilistic principal component regression soft measurement model:
[0016] x = Pt + e
[0017] y = Ct + f
[0018] Where x and y correspond to the preprocessed and normalized process variables and key indicator samples, respectively; P∈R m×k and C∈R r×k Let m be the loading matrix for the process variables and key indicator samples, respectively, where m is the number of process variables, k is the number of principal components, and r is the number of key indicators; t∈R k×1 The extracted principal components follow a normal distribution with a mean of 0 and a variance of 1, i.e., p(t) = N(0, I); e ∈ R m×1 and f∈R r×1 The noise corresponding to the process variables and key indicator samples, respectively, both follow a normal distribution with zero mean. in, and This represents the corresponding variance value.
[0019] During model training, in order to obtain the optimal parameter set in the principal component regression soft sensor model... We need to maximize the following likelihood function, namely...
[0020]
[0021] During model training, the expectation-maximization algorithm can be used to obtain the optimal model parameters of a Bayesian regularized probabilistic principal regression soft sensor model. In the expectation step of the expectation-maximization algorithm, the posterior probability density function of the principal variables in the model is estimated. In the maximization step, the optimal model parameter values are obtained by maximizing the model optimization function. By iterating through the expectation and maximization steps repeatedly until the model parameters converge, the optimal model parameters are obtained.
[0022] In the training process of the principal component regression model, a Bayesian regularization method is introduced to automatically determine the number of principal variables. First, the prior distributions of the process variables and key indicator loading matrices can be given by the following probability density function:
[0023]
[0024]
[0025] Where, p i Let c be the i-th column vector of the load matrix P. j Let α be the j-th column vector of the load matrix C. α = {α1, α2, ..., α...} d} and β={β1,β2,…,β d} are the parameter variables of the two probability density functions, where d is the dimension of the latent variable, and α is the parameter variable of the probability density function. i and β iThese are the i-th and i-th columns of the corresponding parameter matrix, respectively. Introducing these two probability density function distributions into the optimization function of the principal component regression model yields the following new optimization function.
[0026]
[0027] Here, const represents a constant. Detailed algorithms and main reasoning steps are given in the implementation details.
[0028] Furthermore, in the expectation step of the EM algorithm, the posterior probability density function of the principal variables in the principal component regression model is estimated, i.e.:
[0029]
[0030] The estimates of its first and second-order statistics are as follows:
[0031]
[0032]
[0033] in, and The posterior conditional probabilities of x and y, where p(t) is the prior probability of the latent variables. Let be the joint probability distribution of the observed values and latent variables. Using the posterior probability estimates of the principal variables mentioned above, and based on Bayesian regularization theory, the effective number of principal components is determined, i.e., the magnitude of d is determined.
[0034] In the maxima step, the current optimal model parameter values are as follows:
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041] in, Let be the estimated value of the principal variable corresponding to the i-th sample. Let be the estimated value of the second-order statistic of the posterior probability density function of the principal variable corresponding to the i-th sample. Let A = diag(α) be the transpose matrix of the estimate of the first-order statistic of the posterior probability density function of the principal variable corresponding to the nth sample. i(i = 1, 2, ..., d) and B = diag(β) j ,j=1,2,…,d) are diagonal matrices, and Trace(·) is the trace operator of the matrix.
[0042] The modeling method based on Bayesian regularization theory can determine the number of effective principal variables within the framework of a probabilistic model by judging the core direction of the load matrix, and then rank them according to their importance. Therefore, the principal regression model based on Bayesian regularization in this invention can automatically determine the number of principal variables during the modeling process, which greatly improves the online soft measurement performance of traditional principal regression models.
[0043] Meanwhile, this invention provides another soft-sensor modeling method for sulfur recovery devices based on Bayesian regularization, comprising the following steps:
[0044] (1) Use the distributed control system to collect the operating data of the sulfur recovery unit to form a training data sample set for modeling: X∈R n×m Where n is the number of sample datasets and m is the number of process variables, the datasets are stored in the database for later use.
[0045] (2) Key indicator values corresponding to samples used for modeling in historical databases are obtained through on-site sampling and offline laboratory analysis, and used as the training sample set Y∈R for the soft sensor model output. n×r Where n is the number of sample datasets and r is the number of key indicators in the sulfonium recovery unit, the datasets are stored in the database for later use.
[0046] (3) Preprocess and normalize the process variables and key indicator samples respectively, so that the mean of each process variable and key indicator is zero and the variance is 1, resulting in a new data matrix set. and
[0047] (4) For the input and output datasets of the normalized soft measurement model, establish a probabilistic principal component regression soft measurement model based on Bayesian regularization, and store the parameters of the model in the model database for later use.
[0048] A soft measurement method for a sulfur recovery device based on Bayesian regularization includes the following steps:
[0049] (a-1) Collect process variable data of the new sulfur recovery unit online and preprocess and normalize them.
[0050] (a-2) Input the normalized new data directly into the pre-constructed Bayesian regularization-based probabilistic principal component regression soft measurement model to calculate the key indicator values corresponding to the real-time data.
[0051] In step (a-1), the same preprocessing and normalization methods as in the modeling process are used for the preprocessing and normalization.
[0052] For the new data after normalization Inputting this data into a Bayesian regularized principal component regression soft sensor model, the key indicator values corresponding to this real-time data are calculated online as follows: First, the values of the principal variables corresponding to the new data are calculated as follows:
[0053]
[0054] Based on this, the values of the key variables corresponding to the new data are calculated as follows:
[0055]
[0056] As further verification, if the measured value obtained by the process through laboratory testing is y new The real-time measurement error of the soft measurement model can be obtained as follows:
[0057] This invention introduces a probabilistic modeling method on the basis of the traditional principal component regression model, and automatically determines the number of principal components in the principal component regression model through Bayesian regularization, overcoming the shortcomings of the traditional principal component regression model and greatly improving the online detection effect and performance of key indicators in sulfur recovery devices.
[0058] The beneficial effects of this invention are:
[0059] This invention models the correlation between process variables and two key indicators in a sulfur recovery unit using principal component regression, and automatically determines the number of principal variables in the model using Bayesian regularization. In practical applications, easily measurable variables in the process are used to perform online soft measurement of the difficult-to-measure indicators, thereby achieving online estimation of hydrogen sulfide and sulfur dioxide content in the sulfur recovery unit. Attached Figure Description
[0060] Figure 1 It measures the importance of each vector in the process variable loading matrix;
[0061] Figure 2 It measures the importance of each vector in the load matrix of key variables;
[0062] Figure 3 The results are based on an online soft measurement of a principal component regression model with Bayesian regularization. Detailed Implementation
[0063] This invention addresses the problem of detecting key indicators in sulfur recovery devices. By utilizing easily measurable variables during the process, a probabilistic principal component regression soft measurement model based on Bayesian regularization is employed to perform online soft measurement of the hydrogen sulfide and sulfur dioxide content in the process.
[0064] The following steps one through four are the modeling steps, using which a Bayesian regularized probabilistic principal component regression soft measurement model is obtained. During online soft measurement, the constructed Bayesian regularized probabilistic principal component regression soft measurement model is directly used through steps five and six to achieve the soft measurement:
[0065] The first step is to collect data X of various process variables in the sulfur recovery unit through a distributed control system and a real-time database system: X = {x i ∈R m} i=1,2,…,n This data serves as the input to the soft measurement model. Here, n is the number of samples, and m is the number of process variables. These data are stored in a historical database, and a subset of the data is selected as samples for modeling.
[0066] The second step involves obtaining key indicator values corresponding to process variable samples used for modeling from the historical database through on-site sampling and offline laboratory analysis. These values serve as the output Y∈R of the soft sensor model. n×r .
[0067] The third step involves preprocessing and normalizing the data samples for process variables and key indicators, ensuring that the mean of each process variable and key variable is zero and the variance is 1, resulting in a new data matrix set. and in The normalized process variable matrix, This is the normalized key indicator data matrix.
[0068] The collected process data in the historical database is preprocessed to remove outliers and obvious rough error data. To ensure that the scale of the process data does not affect the soft measurement results, the data for different variables are normalized separately, so that the mean of each variable is zero and the variance is 1. In this way, the data of different process variables are at the same scale, which will not affect the subsequent modeling and soft measurement results.
[0069] After obtaining the normalized process variables and key variables in the fourth step, a probabilistic principal component regression soft sensor model based on Bayesian regularization is established, and the parameters of the soft sensor model are stored in the database for later use.
[0070] Specifically, the normalized process variable matrix As input to the soft measurement model, the key indicator data matrix As the output of the soft sensor model, the following probabilistic principal component regression soft sensor model is established:
[0071] x = Pt + e
[0072] y = Ct + f
[0073] Where x and y are a set of process variables and key indicators corresponding to each other in the process variable matrix and key indicator data matrix, respectively, P∈R m×k and C∈R r×k Let R be the loading matrix of process variables and key indicators, t∈R k×1 The extracted principal components follow a normal distribution with a mean of 0 and a variance of 1, i.e., p(t) = N(0, I), and k is the number of principal components. e∈R m×1 and f∈R r ×1 The noise corresponding to the process variables and key indicators, respectively, both follow a normal distribution with zero mean. in, and denoted as the corresponding variance value, and r as the number of key indicators.
[0074] In principal component regression models, Bayesian regularization is introduced to automatically determine the number of principal variables. First, the prior distributions of the process variables and key indicator loading matrices can be given by the following probability density function:
[0075]
[0076]
[0077] Where, p i Let c be the i-th column vector of the load matrix P. j Let α be the j-th column vector of the load matrix C. α = {α1, α2, ..., α...} d} and β={β1,β2,…,β d} are the parameter variables of two probability density functions, where α i and β j Let be the i-th and j-th columns of α and β in the corresponding parameter matrices, respectively, where d is the dimension of the latent variables, and |||| denotes the L2 norm. Introducing these two probability density function distributions into the optimization function of the principal component regression model yields the following new optimization function.
[0078]
[0079] Where const represents a constant; p(X,Y|P,C) is the posterior conditional probability of X and Y; p(P) is the prior probability of P; p(C) is the prior probability of C; and p(X,Y,P,C) is the joint probability distribution of the observed values and latent variables. Let be the likelihood function.
[0080] Based on the above optimization function, in order to obtain the optimal model parameter values, the expectation-maximization algorithm (i.e., the EM algorithm) is adopted. This algorithm consists of two steps: the expectation step and the maximization step, as detailed below:
[0081] In the expected step of this algorithm, the posterior probability density function of the principal variables in the principal component regression model is estimated, i.e.
[0082]
[0083] and The posterior conditional probabilities of x and y, where p(t) is the prior probability of the latent variables. This is the joint probability distribution of the observed values and latent variables.
[0084] Since all options on the right-hand side of the above equation are normally distributed, the posterior probability density function of the principal variable is also normally distributed. Therefore, the estimates of its first and second-order statistics are as follows:
[0085]
[0086]
[0087] In the maximal step of the algorithm, by taking the partial derivatives of the optimization function with respect to each different model parameter and setting them equal to zero, the optimal parameter values (i.e., model parameters P, C, ...) can be obtained. α i and β j Current updated value and ).Right now:
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094] in, Let be the estimated value of the principal variable corresponding to the i-th sample. Let be the estimated value of the second-order statistic of the posterior probability density function of the principal variable corresponding to the i-th sample. Let A = diag(α) be the transpose matrix of the estimate of the first-order statistic of the posterior probability density function of the principal variable corresponding to the nth sample. i (i = 1, 2, ..., d) and B = diag(β) j Let j = 1, 2, ..., d be diagonal matrices, and Trace(·) be the trace operator of the matrix. By iterating repeatedly through the expectation step and the maximization step, the optimal parameter values can be obtained when the model parameters converge.
[0095] Using steps one through four, a Bayesian regularized probabilistic principal component regression soft measurement model can be constructed, which can be used for actual detection in steps five and six below.
[0096] Step 5: Collect new process data x new And preprocess and normalize it.
[0097] For newly collected data samples during the operation of the sulfur recovery unit, in addition to preprocessing them, the data points are normalized using the model parameters used in modeling, i.e., the modeling mean is subtracted and the modeling standard deviation is divided.
[0098] The sixth step is to normalize the new data. The data is directly input into the soft sensor model to calculate the key indicator values corresponding to the real-time data.
[0099] For the new data after normalization Inputting this data into a Bayesian regularized probabilistic principal component regression soft sensor model, the key indicator values corresponding to this real-time data are calculated online, as follows: First, the values of the principal variables corresponding to the new data are calculated as follows:
[0100]
[0101] Based on this, the values of the key variables corresponding to the new data are calculated as follows:
[0102]
[0103] As an alternative, after obtaining the key variable values corresponding to the new data, the measurement values y obtained through laboratory testing are used. new The real-time measurement error of the soft measurement model can be obtained as follows:
[0104] Steps five and six can be used to obtain the output variables in the soft-sensor modeling, namely the content of hydrogen sulfide and sulfur dioxide in the gas. Typically, obtaining gas content values through offline laboratory analysis often takes several hours, leading to a lag in the control of key variables in the sulfur recovery unit. This invention estimates the difficult-to-measure gas content values using easily measurable variables in the process, greatly improving the real-time performance of key indicator control and significantly benefiting product quality control in this process.
[0105] The effectiveness of this invention will be illustrated below using a specific example of a sulfur recovery device. For this device, a total of 2000 data points were collected, of which 1000 were used for modeling (n=1000), and their corresponding key variable values were analyzed and labeled offline. The remaining 1000 data samples were used to verify the effectiveness of the soft sensor model. In this process, five process variables (m=5) were selected for soft sensor modeling of the key indicators of the process. These five process variables are: feed gas flow rate, first air pipe flow rate, second air pipe flow rate, sulfur-containing area gas flow rate, and sulfur-containing area air flow rate. The implementation steps of this invention will now be described in detail using this specific process as an example:
[0106] 1. Preprocess and normalize the process variables and key variables in the 1000 modeling samples respectively, so that the mean of each process variable and key variable is zero and the variance is 1, to obtain a new modeling data matrix.
[0107] 2. Bayesian Regularized Principal Component Regression Soft Sensor Modeling
[0108] The data matrix consisting of the five selected process variables is used as the input of the soft measurement model, and the data matrix of hydrogen sulfide and sulfur dioxide content (r=2) is used as the output of the soft measurement model. Following the detailed method given in the above implementation steps, a Bayesian regularized principal component regression analysis soft measurement model (Bayesian regularized probabilistic principal component regression soft measurement model) is established.
[0109] 3. Acquire real-time measurement data during the process, and preprocess and normalize it.
[0110] To test the effectiveness of the new method, we tested it on 1000 validation samples and processed them using the normalization parameters used in modeling.
[0111] 4. Online soft measurement of hydrogen sulfide and sulfur dioxide content
[0112] Online soft measurement was performed on 1000 validation samples to obtain the corresponding online estimates. Figure 1 and Figure 2The importance ranking of each column vector in the two loading matrices obtained using the Bayesian regularization method is presented. Based on this result, we can easily select the number of principal variables in the principal component regression model; that is, the number of principal variables should be selected as 3 (d=3). Figure 3 The online soft sensor results for 1000 validation samples using the method of this invention are presented. Here, "*" represents the online estimate of the soft sensor model, and "o" represents the offline analysis value for each sample. Figure 3 It can be seen that the measurement results obtained by the soft measurement method of the present invention have a smaller error compared with the offline analysis values of each sample.
[0113] The above embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
Claims
1. A soft measurement method for a sulfur recovery device based on Bayesian regularization, characterized in that, include: The process variable data of the sulfur recovery unit are monitored online, and then preprocessed and normalized. The processed process variable data is input into a Bayesian regularized probabilistic principal component regression soft measurement model to obtain the key indicator values corresponding to the online process variable data; During the training process, the number of principal components in the probabilistic principal regression soft measurement model based on Bayesian regularization is automatically determined using the Bayesian regularization method. The structure of the Bayesian regularized probabilistic principal component regression soft sensor model is as follows: ; x and y correspond to the preprocessed and normalized process variables and key indicator samples, respectively; and The loading matrices are the process variable and key indicator samples, respectively, where m is the number of process variables, k is the number of principal components, and r is the number of key indicators. The extracted principal components follow a normal distribution with a mean of 0 and a variance of 1; and The noise corresponding to the process variables and key indicator samples, respectively, both follow a normal distribution with zero mean. and This represents the corresponding variance value; The optimal model parameters of a Bayesian regularized probabilistic principal component regression soft sensor model are obtained using the expectation-maximization algorithm. In the expectation step of the expectation-maximization algorithm, the posterior distribution density function of the principal variables in the model is estimated; in the maximization step of the expectation-maximization algorithm, the optimal model parameter values are obtained by maximizing the model optimization function; by iterating the expectation step and the maximization step repeatedly until the model parameters converge, the optimal model parameters are obtained. The optimization function is as follows: ; Input training data sample set; To output the training sample set, Represents a constant; Let be the likelihood function. It is a 2-norm; For load matrix The i-th column vector, For load matrix The j-th column vector; and Let be the parameter matrices of the two probability density functions corresponding to P and C, respectively, where and These are the i-th columns of the corresponding parameter matrices; the two probability density functions are respectively: ; ; Where d is the dimension of the latent variable.
2. The soft measurement method for sulfur recovery devices based on Bayesian regularization according to claim 1, characterized in that, The Bayesian regularization-based probabilistic principal component regression soft sensor model is constructed using the following method: (1) Obtain process variable data collected during the normal operation of the sulfur recovery device using the distributed control system, obtain the key index values corresponding to the process variable data using offline means, and form a training data sample set for modeling. (2) Preprocess and normalize the process variables and key indicator samples in the training data sample set respectively; (3) Using the preprocessed and normalized process variable data as the input dataset and the preprocessed and normalized key indicator values as the output dataset, construct a probabilistic principal component regression soft measurement model based on Bayesian regularization.
3. The soft measurement method for sulfur recovery devices based on Bayesian regularization according to claim 2, characterized in that, In the expected step, the posterior probability density function of the principal variables in the principal component regression model is estimated, i.e.: ; The estimates of its first and second-order statistics are as follows: ; ; in, and Let x and y be the posterior conditional probabilities. For the prior probability of the latent variable. Let I be the joint probability distribution of the observed values and latent variables, and let I be the unit vector.
4. The soft measurement method for sulfur recovery devices based on Bayesian regularization according to claim 3, characterized in that, During the maximization step, the current optimal model parameter values are as follows: ; Among them, (x) i ,y i Let be the i-th sample. Let be the estimated value of the principal variable corresponding to the i-th sample. Let be the estimated value of the second-order statistic of the posterior probability density function of the principal variable corresponding to the i-th sample. , Let be the transpose matrix of the estimate of the first-order statistic of the posterior probability density function of the principal variable corresponding to the i-th sample. and They are diagonal matrices, Let be the trace operator of the matrix.
Citation Information
Patent Citations
Probabilistic principal component regression model-based method for soft sensing of butane content of debutanizer
CN103389360A
Dynamic sulfur recovery soft measurement modeling method based on parameterized FIR model
CN110442991A