A Soft Sensor Modeling Method for Fermentation Process Based on Fast Component Transfer Learning
By using TCA to align feature distributions and RELM for model transfer, the method addresses the challenge of modeling no-label data in penicillin fermentation processes, achieving improved prediction accuracy and model generalization.
Patent Information
- Application Number
- CN202111633686.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-28
AI Technical Summary
During penicillin fermentation, soft measurement modeling with no label data in multiple operating conditions is difficult, and traditional methods cannot be directly applied to different operating conditions, resulting in insufficient model generalization capabilities.
The rapid component transfer learning method is used to reduce the edge probability distribution distance between operating conditions through transfer component analysis (TCA), and a regularized limit learning machine (RELM) is used to build a model in the feature-aligned space to predict the penicillin concentration in the operating conditions without label data.
It effectively reduces the difference in feature distribution between operating conditions, improves the prediction accuracy and generalization ability of the model under the label-free data conditions, and reduces the model training time.
Smart Images

Figure CN114334024B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soft sensor modeling for fermentation processes, and particularly to a soft sensor modeling method for fermentation processes based on fast component transfer learning. Background Art
[0002] In the chemical production process, the stable operation of production equipment and the guarantee of product quality are highly related to specific key variables. Currently, the prediction tasks of key indicators in industrial processes mostly adopt data-driven soft sensor modeling methods. Soft sensor modeling relies on methods such as statistical analysis or machine learning to mine potential information in data. It reduces the dependence on the internal mechanism or mathematical model of the industrial process, and at the same time greatly reduces the requirements for process prior knowledge. Among many soft sensor modeling methods, the extreme learning machine, as a typical neural network algorithm, has been widely used due to its advantages such as fast training, simple structure, and strong generalization ability.
[0003] The penicillin fermentation process is a typical batch process, that is, it has the characteristics of a multi-condition process. For the actual penicillin fermentation process, soft sensor technology is mostly used to effectively monitor key variables such as penicillin concentration. The establishment of the soft sensor model is based on the auxiliary variables and key variables. In the actual multi-condition process, collecting sufficient key variables, that is, label values, for soft sensor modeling is time-consuming and laborious, and there may even be cases where there are no labels for some conditions. At the same time, due to the different non-linear relationships and data distributions between different conditions, the models established in specific conditions cannot be directly used for other conditions. Therefore, how to model the conditions of these unlabeled data is a problem worthy of discussion. Summary of the Invention
[0004] To solve the problem of difficult modeling of unlabeled data in a certain condition of the penicillin process, the present invention proposes a soft sensor modeling method for fermentation processes based on fast component transfer learning. By using the Transfer Component Analysis (TCA) method to reduce the distance of the marginal probability distributions between two conditions, making the feature distributions of the two conditions similar, and then using the Regularized Extreme Learning Machine (RELM) method to quickly establish a model in the space where the data features are aligned to predict the penicillin concentration in the unlabeled data condition.
[0005] The technical solution of the present invention is as follows:
[0006] A soft sensor modeling method for fermentation processes based on fast component transfer learning, comprising the following steps:
[0007] 1) Acquisition and preprocessing of penicillin data:
[0008] The penicillin data used in the present invention is obtained through simulation on the Benchmark simulation platform (Pensim). To accelerate the model convergence speed and reduce the influence between data with different dimensions, the original data is normalized. In the stage of building the soft sensor model for multi-condition processes, one condition is the source domain condition with auxiliary variables, namely characteristic variables and key variables, and the other condition is the target domain condition with only characteristic variables and no key variables.
[0009] 2) Establish a model for fast component transfer learning:
[0010] Based on the TCA method, the marginal probability distribution distance between the source domain and the target domain is reduced, thereby completing feature transfer. In the source domain and target domain spaces aligned by TCA, a regularized extreme learning machine model is established based on the mapped source domain data to quickly predict the penicillin concentration in the target domain condition.
[0011] 3) Model performance evaluation:
[0012] To more objectively evaluate the method proposed in the present invention, evaluation indicators Root Mean Square Error (RMSE) and Mean Absolute Error (MAE) are introduced.
[0013] Furthermore, the process of step 1) is as follows:
[0014] Step 1.1) Obtain penicillin data:
[0015] The penicillin data used in the present invention is obtained through simulation on the Benchmark simulation platform (Pensim). The penicillin concentration is the key variable to be predicted in the process, and six correlation variables are used as auxiliary variables, namely carbon dioxide concentration, ventilation rate, substrate feeding temperature, stirrer power, culture dish volume, and pH value.
[0016] Step 1.2) Data normalization:
[0017] To accelerate the model convergence speed, reduce the model training time, and at the same time reduce the influence between data with different dimensions, the data is normalized. The formula is as follows:
[0018]
[0019] In the formula, x is the data after normalization; a is the original data collected; a min is the minimum value in the original data; a max is the maximum value in the original data.
[0020] Step 1.3) Determine the data of the source domain and target domain conditions:
[0021] Arbitrarily select one working condition from the normalized datasets of different working conditions as the source domain working condition dataset, and arbitrarily select one from the remaining working conditions as the target domain working condition dataset. The source domain data has both feature variables and key variables, denoted as {X S , Y S}, and the target domain data only has feature variables, denoted as {X T}.
[0022] Furthermore, the process of step 2) is as follows:
[0023] Step 2.1) Feature transfer between the source domain and the target domain:
[0024] Assume that there exists a feature mapping ψ such that the distributions of the source domain and the target domain after mapping satisfy P(ψ(X S )) = P(ψ(X T ))). The distance between the source domain and the target domain after mapping is measured using the Maximum Mean Discrepancy (MMD), expressed as:
[0025]
[0026] where m is the number of source domain samples, n is the number of target domain samples, the above formula represents measuring the distance between two distributions in the Reproducing Kernel Hilbert Space, and H represents the Reproducing Kernel Hilbert Space.
[0027] The process of solving the MMD distance is transformed into a learning process of the kernel function by using the kernel function. D(X S , X T ) is transformed into the following form:
[0028] tr(KL) - λtr(K)
[0029] where tr(.) is to find the inverse of the matrix, λ is the introduced parameter, K is the kernel matrix obtained by mapping using the kernel function, and L is the matrix introduced by the MMD algorithm. The calculation method of each element of it is:
[0030]
[0031] where D S represents the source domain, D T represents the target domain, x i and x j represent the feature variables of any two samples in the source domain or the target domain. Then, the problem of solving tr(KL) - λtr(K) is transformed into the following optimization problem:
[0032]
[0033] s.t. WT KMKW = I m
[0034] Wherein, M is the central matrix, μ is the introduced parameter, and I m is the m-dimensional identity matrix. W is a matrix with a lower dimension than K, and its solution is (KLK + μI) -1 The first p eigenvalues of KMK. Using the TCA algorithm, map the source domain feature variable X S and the target domain feature variable X T to a new space to obtain the new source domain data feature X Snew and the target domain data feature X Tnew . The number of rows of matrix X is the number of samples, and the number of columns is the total number of features.
[0035] Step 2.2) Establish a soft sensor model based on RELM:
[0036] After the above TCA method operation, X Snew and X Tnew are similarly distributed in space. Establish a regularized extreme learning machine soft sensor model with {X Snew , Y S}, and its optimization objective function is:
[0037]
[0038] s.t. j h(x i )β = y i - ξ i , i = 1, 2,..., m
[0039] Wherein, x Si is the feature variable of the source domain sample, y Si is the true key variable of the source domain sample, h(·) is the function for solving the hidden layer matrix, β is the output weight, and ξ i represents the prediction error of the i-th sample, and γ is the regularization coefficient of the model. Substitute the constraint term into the objective loss function formula and transform it into:
[0040]
[0041] Adopt the Moore-Penrose (generalized inverse of Moore-Penrose) method to obtain the output matrix of the model as H is the hidden layer matrix of the RELM model. Therefore, for the unlabeled working condition X Tnew , the label value obtained through the regularized extreme learning machine is:
[0042]
[0043] Furthermore, the process of the said step 3) is:
[0044] Step 3.1) Root Mean Square Error (RMSE) evaluation:
[0045] The root mean square error is defined as follows:
[0046]
[0047] where: r is the total number of samples in the test set; y t represents the true label value of the input sample x t ; represents the predicted value of the input sample x t . The smaller the RMSE, the better the prediction performance of the model;
[0048] Step 3.2) Mean Absolute Error (MAE) value evaluation:
[0049] The mean absolute error can be expressed as:
[0050]
[0051] The smaller the MAE, the smaller the prediction error of the model, and the more it can illustrate the superiority of this method.
[0052] The beneficial effects of the present invention are mainly manifested in that: based on the transfer component analysis method, the present invention reduces the feature distribution distance between the source domain and the target domain, making the distributions of the two groups of data similar. Then, the regularized extreme learning machine method is used to quickly establish a soft sensor model on the mapped penicillin data in the source domain to predict the penicillin concentration in the target domain. This method solves the problem of difficult establishment of the soft sensor model when there is no labeled data under a certain working condition for penicillin data. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 (a) and (b) are respectively the results of the RELM model established under working condition 1 predicting the penicillin concentrations of working condition 1 and working condition 2;
[0054] Figure 2 (a) and (b) are respectively the results of the RELM model established under working condition 1 predicting the penicillin concentrations of working condition 1 and working condition 2 after being processed by the TCA method;
[0055] Figure 3 (a) and (b) are respectively the results of the RELM model established under working condition 1 predicting the penicillin concentrations of working condition 1 and working condition 3;
[0056] Figure 4 (a) and (b) are respectively the results of the RELM model established under working condition 1 predicting the penicillin concentrations of working condition 1 and working condition 3 after being processed by the TCA method;
[0057] Figure 5(a) and (b) are the results of the RELM model established for operating condition 2 predicting the penicillin concentrations in operating conditions 2 and 3 respectively;
[0058] Figure 6 (a) and (b) are the results of the RELM model established for operating condition 2 predicting the penicillin concentrations in operating conditions 2 and 3 respectively after being processed by the TCA method;
[0059] Figure 7 (a), (b), and (c) are the scatter plots of data distributions before and after adopting the TCA method when operating condition 1 migrates to operating condition 2, operating condition 1 migrates to operating condition 3, and operating condition 2 migrates to operating condition 3 respectively;
[0060] Figure 8 is the flow chart of the present invention. Detailed implementation manners
[0061] The present invention will be further described below with reference to the accompanying drawings.
[0062] Refer to Figures 1 to 8 , a soft sensor modeling method for fermentation process based on fast component transfer learning, the specific steps are as follows:
[0063] (1) Acquisition and preprocessing of penicillin data:
[0064] Step 1.1: Acquisition of penicillin data
[0065] The penicillin data adopted in the present invention is obtained through simulation on the Benchmark simulation platform (Pensim). The penicillin concentration is the key variable to be predicted in the process, and six correlation variables are used as auxiliary variables, namely carbon dioxide concentration, ventilation rate, substrate feeding temperature, stirrer power, culture dish volume, and pH value.
[0066] Step 1.2: Data normalization processing
[0067] To accelerate the model convergence speed, reduce the model training time, and at the same time reduce the influence between data with different dimensions, the data is normalized, and the formula is as follows:
[0068]
[0069] In the formula, x is the data after normalization processing; a is the collected original data; a min is the minimum value in the original data; a max is the maximum value in the original data.
[0070] Step 1.3: Determine the source domain and target domain operating condition data
[0071] Arbitrarily select one working condition from the normalized datasets of different working conditions as the source domain working condition dataset, and arbitrarily select one from the remaining working conditions as the target domain working condition dataset. The source domain data has both feature variables and key variables, denoted as {X S ,Y S}, and the target domain data only has feature variables, denoted as {X T}.
[0072] (2) Establish a model for fast component transfer learning:
[0073] Step 2.1: Feature transfer between the source domain and the target domain
[0074] Assume that there exists a feature mapping ψ such that the distributions of the source domain and the target domain after mapping P(ψ(X S )) = P(ψ(X T )) are equal. The maximum mean discrepancy (MMD) is used to measure the distance between the source domain and the target domain after mapping, which is expressed as:
[0075]
[0076] where m is the number of samples in the source domain, n is the number of samples in the target domain, the above formula represents measuring the distance between two distributions in the reproducing kernel Hilbert space, and H represents the reproducing kernel Hilbert space.
[0077] The kernel function is used to transform the process of solving the MMD distance into the learning process of the kernel function, and D(X S ,X T ) is transformed into the following form:
[0078] tr(KL) - λtr(K)
[0079] where tr(.) is to find the inverse of the matrix, λ is the introduced parameter, K is the kernel matrix obtained by mapping using the kernel function, L is the matrix introduced by the MMD algorithm, and the calculation method of each of its elements is:
[0080]
[0081] where D S represents the source domain, D T represents the target domain, x i and x j represent the feature variables of any two samples in the source domain or the target domain. Then, the problem of solving tr(KL) - λtr(K) is transformed into the following optimization problem:
[0082]
[0083] s.t. W TKMKW = I m
[0084] where M is the central matrix, μ is the introduced parameter, and I m is the m-dimensional identity matrix. W is a matrix with a lower dimension than K, and its solution is (KLK + μI) -1 the first p eigenvalues of KMK. Using the TCA algorithm, map the source domain feature variable X S and the target domain feature variable X T to a new space to obtain the new source domain data feature X Snew and the target domain data feature X Tnew . The number of rows of matrix X is the number of samples, and the number of columns is the total number of features.
[0085] Step 2.2: Establish a soft sensor model based on RELM
[0086] After the above TCA method operation, X Snew and X Tnew are similar in distribution in the space. Establish a regularized extreme learning machine soft sensor model with {X Snew , Y S}, and its optimization objective function is:
[0087]
[0088] s.t. j h(x i )β = y i - ξ i , i = 1, 2,..., m
[0089] where x Si is the feature variable of the source domain sample, y Si is the true key variable of the source domain sample, h(·) is the function to solve the hidden layer matrix, β is the output weight, and ξ i represents the prediction error of the i-th sample, and γ is the regularization coefficient of the model. Substitute the constraint term into the objective loss function formula and transform it into:
[0090]
[0091] Adopt the Moore - Penrose (generalized inverse) method to obtain the output matrix of the model as H is the hidden layer matrix of the RELM model. Therefore, for the unlabeled working condition X Tnew , the label value obtained through the regularized extreme learning machine is:
[0092]
[0093] (3) Model performance evaluation:
[0094] Step 3.1: Root Mean Square Error (RMSE) Evaluation
[0095] The root mean square error is defined as follows:
[0096]
[0097] where: r is the total number of samples in the test set; y t represents the true label value of the input sample x t ; represents the predicted value of the input sample x t . The smaller the RMSE, the better the prediction performance of the model;
[0098] Step 3.2: Mean Absolute Error (MAE) Value Evaluation
[0099] The mean absolute error can be expressed as:
[0100]
[0101] The smaller the MAE, the smaller the prediction error of the model, and the more it can illustrate the superiority of the method.
[0102] (4) Prediction Results of Penicillin Data:
[0103] A total of 3 sets of penicillin data for different working conditions were obtained, namely working condition 1, working condition 2, and working condition 3, with 200 samples for each working condition. In this paper, the effectiveness of the proposed method was verified through 3 migrations between working conditions, namely from working condition 1 to working condition 2, from working condition 1 to working condition 3, and from working condition 2 to working condition 3. Figure 1 (a) and (b) are the results of establishing a model using only the RELM method for working condition 1 to predict the penicillin concentrations of working condition 1 and working condition 2. Through the TCA method, the feature distributions of working condition 1 and working condition 2 were aligned, and then a RELM model was established based on the mapped data of working condition 1, Figure 2 (a) and (b) are the prediction results of the penicillin concentrations of working condition 1 and working condition 2 respectively. It can be seen from the figure that the penicillin concentration of working condition 2 predicted by the TCA and RELM methods better fits the original data value, and at the same time, the accuracy of the prediction result of the penicillin concentration of working condition 1 is also ensured. Similarly, Figure 3 (a) and (b) are the results of the RELM model established for working condition 1 to predict the penicillin concentrations of working condition 1 and working condition 3 respectively, Figure 4 (a) and (b) are the results of the RELM model established for working condition 1 to predict the penicillin concentrations of working condition 1 and working condition 3 after the TCA method. Figure 5 (a) and (b) are the results of the RELM model established for working condition 2 to predict the penicillin concentrations of working condition 2 and working condition 3 respectively, Figure 6(a) and (b) are respectively the results of predicting the penicillin concentrations of operating condition 2 and operating condition 3 by the RELM model established under operating condition 2 after the TCA method. Meanwhile, the RMSE and MAE values of the prediction results of each operating condition in the target domain are recorded in Table 1. It can be obtained from Table 1 that the TCA+RELM method can significantly reduce the RMSE and MAE values of the target domain prediction, further verifying the effectiveness of the TCA+RELM method in predicting the penicillin data concentration without labels.
[0104] Figure 7 (a), (b), and (c) are respectively the scatter plots of data distribution before and after the TCA method for the migration of operating condition 1 to operating condition 2, operating condition 1 to operating condition 3, and operating condition 2 to operating condition 3. It can be seen from the figure that the data distributions between the original two operating conditions are far apart, and the TCA method narrows the distance between the two operating conditions, which well explains why TCA+RELM can improve the accuracy of penicillin concentration prediction in the target domain.
[0105] Table 1
[0106]
[0107] The method of the present invention is based on the knowledge of transfer learning to solve the problem of difficult modeling of unlabeled data in a certain operating condition of the penicillin process. The distance between the source domain and the target domain feature distributions is reduced by the TCA method, and then a RELM model is established in the space with similar feature distributions to predict the penicillin concentration of the target domain operating condition with only feature data. This method has high accuracy and is universal and general.
[0108] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also covers equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A soft-sensing modeling method for fermentation process based on fast component transfer learning, characterized in that, Including the following steps: 1) Acquisition and preprocessing of penicillin data: The penicillin data is obtained through simulation on the Benchmark simulation platform Pensim. To accelerate the model convergence speed and reduce the influence between data with different dimensions, the original data is normalized. In the stage of building the soft sensor model for multi-condition processes, one condition is the source domain condition with auxiliary variables, namely characteristic variables and key variables, and the other condition is the target domain condition with only characteristic variables and no key variables. 2) Establish a model for fast component transfer learning: Based on the Transfer Component Analysis (TCA) method, reduce the marginal probability distribution distance between the source domain and the target domain, thereby completing feature transfer. In the source domain and target domain spaces aligned by TCA, based on the mapped source domain data, establish a regularized extreme learning machine model to quickly predict the penicillin concentration in the target domain condition. The process of step 2) is as follows: Step 2.1) Feature transfer between the source domain and the target domain: Suppose there exists a feature mapping ψ such that the distributions of the source domain and the target domain after mapping, P(ψ(X S )) = P(ψ(X T )); The maximum mean discrepancy MMD is used to measure the distance between the source domain and the target domain after mapping, which is expressed as: Among them, m is the number of source domain samples, and n is the number of target domain samples. The above formula represents the distance between two distributions in the reproducing kernel Hilbert space. represents the reproducing kernel Hilbert space; The process of solving the MMD distance is transformed into the learning process of the kernel function by using the kernel function, and D(X S , X T ) is transformed into the following form: tr(KL)-λtr(K) where tr(.) is to find the inverse of the matrix, λ is the introduced parameter, K is the kernel matrix obtained by mapping with the kernel function, and L is the matrix introduced by the MMD algorithm, and the calculation method of each of its elements is: Among them, represents the source domain, represents the target domain, x i and x j represent the characteristic variables of any two samples in the source domain or the target domain; then, the problem of solving tr(KL)-λtr(K) is transformed into the following optimization problem: s.t.W T KMKW=I m Among them, M is the central matrix, μ is the introduced parameter, and I m is the m-dimensional identity matrix, and W is a matrix with a lower dimension than K. Its solution is (KLK + μI) -1 for the first p eigenvalues of KMK; using the TCA algorithm, map the source domain feature variable X S and the target domain feature variable X T to a new space to obtain the new source domain data feature X Snew and the target domain data feature X Tnew . The number of rows of matrix X is the number of samples, and the number of columns is the total number of features; Step 2.2) Establish a soft sensor model based on the Regularized Extreme Learning Machine (RELM): After the above TCA method operation, X Snew and X Tnew have similar distributions in space; a regularized extreme learning machine soft sensor model is established with {X Snew , Y S}, where Y S represents the key variable in the source domain, and its optimization objective function is: s.t. h(x Si )β = y Si -ξ i , i = 1, 2,..., m where x Si is the feature variable of the source domain sample, y Si is the true key variable of the source domain sample, h(·) is the function for solving the hidden layer matrix, β is the output weight, and ξ i represents the prediction error of the i-th sample, and γ is the regularization coefficient of the model; substituting the constraint term into the formula of the objective loss function, it is transformed into: Using the Moore-Penrose generalized inverse method, the output matrix of the model is obtained as H is the hidden layer matrix of the RELM model; thus, for the unlabeled working condition X Tnew , the label value obtained by the regularized extreme learning machine is: 3) Model performance evaluation: Introduce the evaluation indicators Root Mean Square Error (RMSE) and Mean Absolute Error (MAE).
2. The soft sensor modeling method for fermentation process based on fast component transfer learning according to claim 1, characterized in that The process of step 1) is as follows: Step 1.1) Obtain penicillin data: The penicillin concentration is the key variable to be predicted in the process, and six correlation variables are used as auxiliary variables, namely carbon dioxide concentration, ventilation rate, substrate feeding temperature, stirrer power, culture dish volume, and pH value. Step 1.2) Data normalization processing: To accelerate the model convergence speed, reduce the model training time, and at the same time reduce the influence between data with different dimensions, the data is normalized, and the formula is as follows: In the formula, x is the data after normalization processing. a is the original data collected; a min is the minimum value in the original data; a max is the maximum value in the original data; Step 1.3) Determine the source domain and target domain condition data: Arbitrarily select one working condition from the normalized datasets of different working conditions as the source domain working condition dataset, and arbitrarily select one from the remaining working conditions as the target domain working condition dataset; the source domain data has both feature variables and key variables, denoted as {X S , Y S}, and the target domain data has only feature variables, denoted as {X T}.
3. The soft sensor modeling method for fermentation process based on fast component transfer learning according to claim 1, characterized in that The process of step 3) is as follows: Step 3.1) Root Mean Square Error (RMSE) evaluation: The definition of the root mean square error is as follows: where: r is the total number of samples in the test set; y t represents the true label value of the input sample x t ; represents the predicted value of the input sample x t ; The smaller the RMSE, the better the prediction performance of the model; Step 3.2) Mean Absolute Error (MAE) value evaluation: The mean absolute error can be expressed as: The smaller the MAE, the smaller the prediction error of the model, and the more it can illustrate the superiority of this method.