A Carbon Dioxide Concentration Prediction Method Based on Semi-Supervised Deep Probability Model

By applying a semi-supervised depth probability model in the carbon dioxide absorption tower, using labelless data and deep learning technology, the problem of real-time measurement of carbon dioxide concentration is solved, and the prediction performance is significantly improved.

CN119538205BActive Publication Date: 2025-05-30SOUTHEAST UNIV

Patent Information

Application Number
CN202411553277.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-05-30
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

The prior art is difficult to measure the concentration of carbon dioxide in the CO2 absorption tower in a simple and economical way in real time, and the predictive performance of the main element regression model is limited by the scarcity of shallow architecture and label data.

Method used

A semi-supervised depth probability model is used to introduce a large number of label-free data samples, and an online prediction model of carbon dioxide concentration is established through deep learning to extract process data information.

Benefits of technology

Effectively utilize large-scale label-free data information to improve the online prediction effect of carbon dioxide concentration and improve prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538205B_ABST
    Figure CN119538205B_ABST
Patent Text Reader

Abstract

The present invention discloses a carbon dioxide concentration prediction method based on a semi-supervised deep probability model for on-line detection of the carbon dioxide content in a carbon dioxide absorption tower. Based on the conventional probabilistic principal component regression model, the present invention simultaneously introduces the ideas of semi-supervised learning and deep learning, expands the basic probabilistic principal component regression model into the structure of a semi-supervised deep learning model, that is, proposes a new soft sensing method based on a semi-supervised deep probabilistic principal component regression model for on-line detection of the carbon dioxide concentration in a carbon dioxide absorption tower. Compared with the conventional principal component regression model, the method of the present invention can not only effectively utilize a large amount of cheap unlabeled sample information, but also deeply extract the information hidden in the process data, thereby improving the actual prediction effect of the carbon dioxide concentration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial process data-driven modeling and applications, and particularly relates to a method for predicting carbon dioxide concentration based on a semi-supervised deep probability model. Background Art

[0002] As an indispensable device in the large-scale ammonia synthesis production process, the carbon dioxide absorption tower plays an important role. In this device, the concentration of carbon dioxide is a core variable. How to obtain the concentration value of carbon dioxide in the tower in real time is crucial for the overall operation quality control and quality optimization of the carbon dioxide absorption tower. However, limited by the current measurement means, it is still a difficult problem to measure the carbon dioxide content in the tower in a relatively simple and economical way in real time. With the continuous development of data-driven modeling technology, it has become possible to indirectly achieve real-time measurement by using a model to predict the carbon dioxide concentration. Although the principal component regression model is one of the most commonly used methods in data-driven modeling, on the one hand, the prediction performance of the model is limited by the shallow architecture, and on the other hand, due to the scarcity of labeled data samples, it is difficult for the model to obtain satisfactory practical application effects. The present invention effectively combines semi-supervised learning and deep learning ideas, expands the probabilistic principal component regression model into a deep model, and at the same time introduces a large number of cheap unlabeled data samples through the semi-supervised model structure to establish a semi-supervised deep probability model and apply it to the online prediction of carbon dioxide concentration in the carbon dioxide absorption tower. Compared with traditional methods, the method of the present invention can not only utilize the information of large-scale unlabeled data, but also deeply extract the process data information, thereby effectively improving the online prediction effect of the carbon dioxide concentration content. Summary of the Invention

[0003] The object of the present invention is to provide a new prediction method based on a semi-supervised deep probability model for the problem that the prediction performance of the carbon dioxide concentration in the carbon dioxide absorption tower is limited.

[0004] The object of the present invention is achieved by the following technical solutions:

[0005] A method for predicting carbon dioxide concentration based on a semi-supervised deep probability model, characterized by including the following steps:

[0006] (1) Use a conventional instrument measurement system to collect normal operating condition data during the operation of the carbon dioxide absorption tower as the input data training sample set when constructing the prediction model: X ∈ R N×m . Wherein, N is the number of training samples in the model input data set, and m is the number of variables in the input data. The input data training sample set is stored in the modeling database for standby.

[0007] (2) Obtain the carbon dioxide content value in the carbon dioxide absorption tower for modeling offline through manual sampling at the industrial production site and chemical analysis in the analysis cabin, and use it as the output data training sample set y ∈ R when constructing the prediction model. n , where n is the number of sample data sets, and this data set is also stored in the modeling database for later use.

[0008] (3) Extract the training sample set from the modeling database and divide it into a labeled data set and an unlabeled data set Perform normalization processing on different variables of the input data and output data respectively, so that the mean value of each variable is zero and the variance is 1, and unify the dimension of the variables to the same scale.

[0009] (4) Based on the standardized data sample set, first establish the first layer structure of the deep probabilistic regression model, that is, the semi-supervised probabilistic principal component regression model, and store the parameters of this model in the model library for later use.

[0010] (5) Starting from the principal component variables extracted from the upper layer semi-supervised probabilistic principal component regression model, re-call the output variable information to establish the second layer structure of the deep probabilistic model, that is, the second layer semi-supervised form of the probabilistic principal component regression model, and also store the model parameters in the model library for later use.

[0011] (6) And so on, we establish a semi-supervised deep probabilistic model including L hidden layers, and store all the model parameters of the hidden layers in the model library for later use. At the same time, integrate the principal component variables extracted from each hidden layer together to form a new comprehensive principal component variable.

[0012] (7) Based on the data samples corresponding to the comprehensive principal component variables, establish the regression relationship with the output variables, that is, construct a final semi-supervised probabilistic principal component regression model, and also store the parameters of the model in the model database for later use.

[0013] (8) Collect new real-time measurement data from the industrial site in the carbon dioxide absorption tower, and use the parameters stored in the semi-supervised probabilistic model library to perform normalization processing on the new data as a preparation for subsequent model prediction. (9) Use the normalized new data as the input of the semi-supervised deep probabilistic model, calculate the model output value corresponding to the current data, and complete the online prediction of the carbon dioxide concentration value.

[0014] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the carbon dioxide concentration prediction method based on the semi-supervised deep probabilistic model.

[0015] A computer-readable storage medium stores computer instructions thereon, and when the computer instructions are executed by a processor, the carbon dioxide concentration prediction method based on the semi-supervised deep probability model is implemented.

[0016] Advantages of the present invention:

[0017] Based on the basic probabilistic principal component regression model, by integrating semi-supervised learning and deep learning ideas, the present invention establishes a regression relationship model between the easily measurable variables in the carbon dioxide absorption tower and the carbon dioxide concentration, and realizes the online prediction of the carbon dioxide concentration in the tower. Description of the drawings

[0018] Figure 1 It is a schematic diagram of the carbon dioxide absorption tower process;

[0019] Figure 2 It is the online prediction result of the method of the present invention in an actual carbon dioxide absorption tower. Specific embodiments

[0020] Aiming at the problem that the performance of the online prediction model for the carbon dioxide concentration in the carbon dioxide absorption tower is limited, the present invention establishes a new semi-supervised deep probability model to perform online prediction on the carbon dioxide content in this process.

[0021] Embodiment: A carbon dioxide concentration prediction method based on a semi-supervised deep probability model, the method comprising the following steps:

[0022] First step: Use a conventional instrument measurement system to collect normal operating condition data during the operation of the carbon dioxide absorption tower as the input data training sample set when constructing the prediction model: X ∈ R N×m . Where N is the number of training samples in the model input data set, m is the number of variables in the input data, and the input data training sample set is stored in the modeling database for later use.

[0023] Second step: Offline obtain the carbon dioxide content value in the modeling carbon dioxide absorption tower through manual sampling at the industrial production site and analysis in the analysis cabin, as the output data training sample set y ∈ R n , where n is the number of sample data sets, and the data set is also stored in the modeling database for later use.

[0024] Third step: Extract the training sample set from the modeling database and divide it into a labeled data set and an unlabeled data set Normalize different variables of the input data and output data respectively, so that the mean of each variable is zero and the variance is 1, and unify the dimension of the variables to the same scale.

[0025] Step 4: Based on the standardized data sample set, first establish the first-layer structure of the deep probabilistic regression model, that is, the semi-supervised probabilistic principal component regression model, and store the parameters of this model in the model library for later use. The structure of the semi-supervised probabilistic principal component regression model in the first layer is as follows:

[0026] x = P (1) t (1) + e (1)

[0027] y = C (1) t (1) + f (1)

[0028] Where P (1) ∈ R m×k and C (1) ∈ R 1×k are the load matrices of the input variables and carbon dioxide concentration for the first-layer model, t (1) ∈ R k×1 is the principal component feature extracted by the first-layer model, and k is the number of principal components. e (1) ∈ R m×1 and f (1) ∈ R are the noises corresponding to the input variables and carbon dioxide content respectively, and both follow a normal distribution with zero mean, that is Where and are the corresponding variance values. Generally, we need to maximize the likelihood functions corresponding to the labeled data and unlabeled data to obtain the optimal parameters

[0029]

[0030] In the process of cyclic iterative optimization, since the principal component variable is a Gaussian distribution variable, therefore, we only need to determine its first-order and second-order statistics. For the labeled data and unlabeled data, the corresponding statistics are calculated as follows respectively:

[0031]

[0032] On this basis, we can obtain the following optimal parameters of the first-layer model:

[0033]

[0034] Where n lb and n ul are the numbers of labeled samples and unlabeled samples respectively, trace represents calculating the trace of the matrix, and E represents the expected value.

[0035] Step 5: Starting from the principal component variables extracted from the upper-layer semi-supervised probabilistic principal component regression model, re-call the output variable information Build the second-layer structure of the deep probabilistic model, that is, the second-layer semi-supervised probabilistic principal component regression model, and also store the model parameters in the model library for future use. Among them, the second-layer probabilistic principal component regression model is as follows:

[0036] t (1) = P (2) t (2) + e (2)

[0037] y = C (2) t (2) + f (2)

[0038] By optimizing the following likelihood function, we can obtain the optimal parameters of the second-layer probabilistic principal component regression model

[0039]

[0040] Step 6: And so on, we build a semi-supervised deep probabilistic model including L hidden layers, and store all the model parameters of the hidden layers in the model library for future use. At the same time, integrate the principal component variables extracted from each hidden layer together to form a new comprehensive principal component variable. Among them, the likelihood function of each layer of the semi-supervised deep probabilistic model is given by the following formula:

[0041]

[0042] where l = 2,..., L. On the basis of completing the construction of the hidden layer models of the deep model, we integrate the principal component variable sets obtained from all L hidden layers into a comprehensive principal component variable, as follows:

[0043] g = [t (1) t (2) ...t (L) T

[0044] Step 7: Based on the data samples corresponding to the comprehensive principal component variable, establish the regression relationship with the output variable, that is, construct a final semi-supervised probabilistic principal component regression model, and also store the parameters of the model in the model database for future use. Among them, the structure of the final semi-supervised probabilistic principal component regression model is as follows:

[0045] g = P (g) t (g) + e (g)

[0046] y = C (g) t (g) + f (g) ​

[0047] Similarly, to obtain the optimal parameter set we maximize the following likelihood function through an iterative loop:

[0048]

[0049] Step 8: Collect new real-time measurement data at the industrial production site and normalize it using the parameters in the model library to unify the scale of the new data to that of the modeling data samples.

[0050] Step 9: Take the normalized new data as the input of the semi-supervised deep probability model, calculate the output variable values corresponding to this real-time data, and complete the online prediction of the carbon dioxide concentration. First, using the model parameters of each hidden layer in the deep model, obtain the posterior probability distribution of the principal component variables corresponding to the new data x new as follows:

[0051]

[0052] Based on this, calculate the mean value of each principal component variable as follows:

[0053]

[0054] Then, by integrating the principal component variables obtained from different hidden layers of the deep model, obtain the comprehensive principal component variable information, and then input it into the final semi-supervised probabilistic principal component regression model to calculate the predicted value of the carbon dioxide concentration of the new data as follows:

[0055]

[0056] Next, we specifically combine an actual carbon dioxide absorption tower case to illustrate the effectiveness of the present invention. Figure 1 The process flow diagram of the carbon dioxide absorption tower is given. To predict the carbon dioxide concentration in the tower online, based on the analysis of variable correlations, a total of 11 highly correlated process variables are selected as the input variables of the model, including various temperature, pressure, flow rate and other data. The specific information is shown in Table 1. Limited by the measurement of carbon dioxide concentration in this production process, we can only collect 50 labeled data samples, that is, including both the input variables and output variable information of the model. In addition, 1000 unlabeled data samples are collected under normal operating conditions to participate in the construction of the semi-supervised model, that is, only including the information of the input variables. In addition, to test the prediction performance of the established model, we independently collected 2000 data samples. Next, the implementation steps of the present invention will be elaborated in detail in combination with this specific process:

[0057] 1. For the labeled data sample set and the unlabeled data sample set, we perform normalization on different process variables and carbon dioxide concentration variables respectively to unify the dimension of the data and obtain a standardized training data sample set.

[0058] 2. Construction of semi-supervised deep probability model

[0059] Take the selected 11 process variables as the input of the semi-supervised deep probability model, and the carbon dioxide concentration variable as the output of the model. According to the detailed steps given in the implementation steps of the present invention, establish an on-line prediction method for carbon dioxide concentration based on the semi-supervised deep probability model.

[0060] 3. During the normal production process of carbon dioxide, 11 process variable data in the absorption tower are obtained in real time and normalized using the parameters in the model library.

[0061] To test the effectiveness of the model we built, perform on-line prediction on the carbon dioxide concentration values corresponding to 2000 samples collected in real time. First, normalize them using the model parameters to unify the data dimension scale.

[0062] 4. On-line prediction of carbon dioxide concentration in test sample data

[0063] Use the established semi-supervised deep probability model to perform on-line prediction on 2000 test data collected in real time and calculate their corresponding carbon dioxide concentration values. Figure 2 The on-line prediction results of the method of the present invention are given. It can be seen from the figure that the new model gives relatively satisfactory results, and the prediction error is within the range of operation in actual applications.

[0064] Table 1: Explanation of variables in the carbon dioxide absorption tower

[0065] Label Name U1 Process gas pressure entering E3 U2 Liquid level of separator 2 U3 Lean liquid temperature at the outlet of E1 U4 Lean liquid flow rate to the CO2 absorption tower U5 Semi-lean liquid flow rate to the CO2 absorption tower U6 Process gas temperature at the outlet of separator 2 U7 Process gas pressure difference at the inlet of the CO2 absorption tower U8 Rich liquid temperature at the outlet of the CO2 absorption tower U9 Liquid level of the CO2 absorption tower U10 Liquid level of separator 1 U11 Outlet process gas pressure Y Residual CO2 content in the process gas

[0066] The above embodiments are used to explain the present invention, rather than to limit the present invention. Any modification and change made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A method for predicting carbon dioxide concentration based on a semi-supervised deep probability model, characterized in that: The method comprises the following steps: (1) Use the instrument measurement system to collect normal operating data of the carbon dioxide absorption tower during operation, which is used as the input data training sample set for building the prediction model: X∈R N×m , where N is the number of training samples of the model input data set, m is the number of variables of the input data, and the input data training sample set is stored in the modeling database for standby use. (2) Through manual sampling at the industrial production site and chemical analysis in the analysis room, the CO2 content value in the modeled CO2 absorption tower is obtained offline as the output data training sample set y∈R when building the prediction model n , where n is the number of sample data sets, and the data sets are also stored in the modeling database for backup. (3) Extract training sample sets from the modeling database and divide them into labeled data sets and unlabeled datasets Normalize the different variables of input data and output data respectively, so that the mean of each variable is zero and the variance is 1, and unify the dimensions of the variables to the same scale. (4) Based on the standardized data sample set, the first layer structure of the deep probabilistic regression model, namely the semi-supervised probabilistic principal component regression model, is first established, and the parameters of the model are stored in the model library for future use. (5) Starting from the principal component variables extracted from the previous semi-supervised probability principal component regression model, the output variable information is re-called The second layer structure of the deep probability model is established, that is, the second layer semi-supervised probability principal component regression model, and the model parameters are also stored in the model library for backup. (6) Similarly, a semi-supervised deep probability model with L hidden layers is established. The model parameters of all hidden layers are stored in the model library for backup. At the same time, the principal component variables extracted from each hidden layer are integrated together to form a new comprehensive principal component variable. (7) Based on the data samples corresponding to the comprehensive principal component variable, the regression relationship between the principal component variable and the output variable is established, that is, a final semi-supervised probabilistic principal component regression model is constructed, and the parameters of the model are also stored in the model database for future use. (8) Collect new industrial field real-time measurement data in the carbon dioxide absorption tower, and use the parameters stored in the semi-supervised probability model library to normalize the new data as a preparation for subsequent model prediction. (9) The normalized new data is used as the input of the semi-supervised deep probability model, the model output value corresponding to the current data is calculated, and the online prediction of the carbon dioxide concentration value is completed.

2. The method for predicting carbon dioxide concentration based on a semi-supervised deep probability model according to claim 1, characterized in that: The step (4) is specifically as follows: based on the standardized data sample set, firstly, the first layer structure of the deep probabilistic regression model, that is, the semi-supervised probabilistic principal component regression model, is established, and the parameters of the model are stored in the model library for standby use, wherein the structure of the first layer of the semi-supervised probabilistic principal component regression model is as follows: x=P (1) t (1) +e (1) y=C (1) t (1) +f (1) , Among them, P (1) ∈R m×k and C (1) ∈R 1×k is the loading matrix of the first-layer model for the input variables and carbon dioxide concentration, t (1) ∈R k×1 is the principal component feature extracted by the first layer model, k is the number of principal components, e (1) ∈R m×1 and f (1) ∈R are the noises corresponding to the input variables and carbon dioxide content, and both obey the normal distribution with zero mean, that is, in, and For the corresponding variance value, it is necessary to maximize the likelihood function corresponding to the labeled data and the unlabeled data to obtain the optimal parameter In the iterative optimization process, since the principal variable is a Gaussian distribution variable, it is only necessary to determine its first-order and second-order statistics. For labeled data and unlabeled data, the corresponding statistics are calculated as follows: On this basis, the following first-layer model optimization parameters are obtained: Among them, n lb and n ul are the number of labeled samples and unlabeled samples respectively, trace represents the trace of the calculation matrix, and E represents the expected value.

3. The method for predicting carbon dioxide concentration based on a semi-supervised deep probability model according to claim 2, characterized in that: The steps (5) and (6) are specifically as follows: starting from the principal component variables extracted from the previous layer of semi-supervised probability principal component regression model, re-calling the output variable information The second layer structure of the deep probability model, that is, the second layer semi-supervised probabilistic principal component regression model, is established. The model parameters are also stored in the model library for standby. Similarly, a semi-supervised deep probability model including L hidden layers is established. The model parameters of all hidden layers are stored in the model library for standby. At the same time, the principal component variables extracted from each hidden layer are integrated together to form a new comprehensive principal component variable. The likelihood function of each layer of the semi-supervised deep probability model is given by the following formula: Where l = 2, ..., L. After completing the construction of each hidden layer model of the deep model, the principal variable obtained from all L hidden layers is integrated into a comprehensive principal variable, as shown below: g=[t (1) t (2) …t (L) ] T 。 4. The method for predicting carbon dioxide concentration based on a semi-supervised deep probability model according to claim 3, characterized in that: The step (7) is specifically as follows: based on the data sample corresponding to the comprehensive principal component variable, a regression relationship between the principal component variable and the output variable is established, that is, a final semi-supervised probabilistic principal component regression model is constructed, and the parameters of the model are also stored in the model database for standby use. The final semi-supervised probabilistic principal component regression model structure is as follows: g=P (g) t (g) +e (g) y=C (g) t (g) +f (g) In order to obtain the optimal parameter set Maximize the following likelihood function by looping and iterating:

5. The method for predicting carbon dioxide concentration based on a semi-supervised deep probability model according to claim 4, characterized in that: The step (9) is specifically as follows: using the standardized new data as the input of the semi-supervised deep probability model, calculating the output variable value corresponding to the real-time data, and completing the online prediction of the carbon dioxide concentration. First, using the model parameters of each hidden layer in the deep model, the new data x new The corresponding posterior probability distribution of the main variable is as follows: On this basis, the mean of each main variable is calculated as follows: Then, by integrating the principal component variables obtained from the hidden layers of different depth models, the comprehensive principal component variable information is obtained, and then it is input into the final semi-supervised probabilistic principal component regression model to calculate the predicted value of carbon dioxide concentration of the new data as follows:

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the carbon dioxide concentration prediction method based on the semi-supervised deep probability model as described in any one of claims 1 to 5 above is implemented.

7. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, the carbon dioxide concentration prediction method based on the semi-supervised deep probability model as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Probabilistic principal component regression model-based method for soft sensing of butane content of debutanizer

    CN103389360A

  • Semi-supervised soft measurement method based on improved Tri-training GPR

    CN118468021A

Cited By

  • Soft measurement method and device based on deep KPLSR network

    CN121525314A