D-pantothenic acid fermentation multi-objective dynamic optimization method based on biological-digital twinning

By constructing a bio-digital twin model and combining it with machine learning and optimization algorithms, the shortcomings of traditional fermentation process monitoring and control methods have been addressed. This has enabled accurate prediction and optimization of the fermentation process, thereby improving D-pantothenic acid yield and fermentation efficiency.

CN121963918APending Publication Date: 2026-05-01ZHEJIANG UNIV OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV OF TECH
Filing Date
2025-12-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional fermentation process monitoring and control methods rely on human experience, making it difficult to achieve real-time and accurate multivariate fermentation process regulation. Existing models lack accuracy, robustness, and adaptability in complex fermentation data processing.

Method used

A multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins was adopted. Through data acquisition, multivariate prediction model construction, model optimization and validation, interpretability analysis and fermentation process optimization, combined with machine learning and optimization algorithms, a bio-soft sensor model was constructed to achieve accurate prediction and control of the fermentation process.

Benefits of technology

It improves the predictability and controllability of the fermentation process, enables accurate prediction of D-pantothenic acid yield and optimization of fermentation process parameters, and enhances fermentation efficiency and yield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963918A_ABST
    Figure CN121963918A_ABST
Patent Text Reader

Abstract

The invention provides a D-pantothenic acid fermentation multi-target dynamic optimization method based on biological-digital twinning, which comprises the following steps: constructing a biological soft sensor through multiple machine learning model algorithms, and combining the constructed biological soft sensor with a neural network. Meanwhile, fermentation process conditions are optimized by applying a particle swarm algorithm and a genetic algorithm, so that the whole fermentation process simulation monitoring under data driving is realized, the problems of insufficient accuracy and poor robustness when a traditional prediction technology is used for processing complex fermentation data are effectively solved, and the predictability and controllability of the fermentation process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microbial fermentation technology, specifically relating to a multi-objective dynamic optimization method for D-pantothenic acid fermentation based on biological-digital twins. Background Technology

[0002] In the field of modern biotechnology, the optimization and control of fermentation processes are crucial. Fermentation involves multiple biochemical reactions, which are influenced by various environmental factors, such as ammonia quality, sugar supplementation, OD growth, pH, temperature, and dissolved oxygen levels. To achieve precise control of the fermentation process, real-time monitoring and prediction of these key parameters are necessary. Traditional fermentation monitoring and control methods often rely on manual experience and offline analysis, which are not only cumbersome and untimely but also fail to achieve real-time, precise control. In recent years, with the rapid development of machine learning technology, its application in bioengineering has become increasingly widespread. Machine learning, a subfield of artificial intelligence, has the advantage of learning knowledge from data. It can train and learn from data to build internal mathematical models, which helps in processing new data and discovering potential connections hidden within large amounts of data.

[0003] Traditional fermentation data prediction methods mostly rely on single mathematical models or empirical formulas. These methods often have limitations when dealing with complex and variable fermentation data. Fermentation is a typical nonlinear, nonstationary, high-dimensional, and slowly time-varying complex system. The dynamic behavior of biological reactions and the relationship between input variables (such as substrate concentration, dissolved oxygen, pH, and temperature) are not simple linear relationships, but rather exhibit complex and variable patterns. Although the Monod equation and logistic equation have been proposed in fermentation kinetics to describe the cell growth process, these theories are only applicable to batch fermentation models, while industrial production methods mostly use fed-batch fermentation. Therefore, modeling and simulating the entire bio-fermentation process often suffers from low accuracy and poor applicability. When faced with large amounts of high-dimensional fermentation data, traditional methods often fail to meet practical needs in terms of computational efficiency and prediction accuracy. Furthermore, traditional methods lack the ability to adapt to new data and handle uncertainties, which limits their application in complex fermentation environments.

[0004] Therefore, a multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins is developed to overcome the limitations of traditional methods and achieve accurate prediction and control of the fermentation process. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins. This method addresses the problems of insufficient accuracy and poor robustness of traditional prediction techniques when processing complex fermentation data, thereby improving the predictability and controllability of the fermentation process.

[0006] To solve the above problems, the technical solution adopted in this application is:

[0007] The first objective of this invention is to provide a multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins, comprising the following steps:

[0008] Step 1: Data Acquisition: Collect multi-dimensional fermentation parameters during the fermentation of D-pantothenic acid by genetically engineered strains under different fermentation process conditions, and construct a fermentation process database for the genetically engineered strains.

[0009] Step 2, Model Construction: Using the preprocessed multidimensional fermentation parameters as input feature variables and D-pantothenic acid yield as the model output value, a multivariate prediction model is established to achieve accurate prediction of D-pantothenic acid yield by the initial biosoft sensor.

[0010] Step 3, Model Optimization and Validation: The hyperparameters of the multivariate prediction model are tuned using the grid search method, and the generalization ability of the model is verified on the independent validation set to construct a biosoft sensor model.

[0011] Step 4: Model interpretability analysis: The SHAP value analysis method (SHapley Additive ex Planations) is used to quantify the contribution of each feature variable value to the prediction results. Through feature importance ranking and local interpretation visualization, the interpretability analysis of the biosoft sensor model is carried out to explain the prediction results of the biosoft sensor model.

[0012] Step 5: Fermentation process optimization: The constructed bio-soft sensor model is coupled with a multi-objective optimization algorithm to form a fermentation process optimization system. The fermentation process parameters are adjusted with the objective functions of maximizing D-pantothenic acid yield and minimizing energy consumption.

[0013] The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins protected in this application constructs a bio-soft sensor model using machine learning model algorithms, combines the constructed bio-soft sensor model with a neural network, and optimizes the fermentation process conditions using particle swarm optimization and genetic algorithms, thereby realizing data-driven simulation and monitoring of the entire fermentation process.

[0014] As a preferred embodiment of this application, the multi-dimensional fermentation parameters include fermentation time (Time), OD (October Spectrum), and other parameters. 600One or more of the following: fermentation temperature (Tem), pH, dissolved oxygen (DO), rotational speed, ammonia concentration, feed supplement, and residual sugar.

[0015] As a preferred embodiment of this application, the genetically engineered strain (hereinafter referred to as strain W1) is *E. coli* ZJUTDPAP16 (E. coli W3110, Trc-EcilvD* / BspanBA* / CgpanC* / alsS* / BspanB* / nac) GTG / ΔP 345ptsH / gltA GTG / gltA TTG / Trc- PpfkB, already disclosed in CN116590209A).

[0016] Preferably, the sample data in the fermentation process database is randomly divided into a training set, a validation set, and a test set in a 6:2:2 ratio. The training set is used for training the machine learning model, the validation set is used to evaluate the performance of the trained model, and the test set is used for the final performance evaluation of the model. Specifically, the fermentation process database contains a total of 207 samples.

[0017] As a preferred embodiment of this application, in step 2, the candidate model used to establish the multivariate prediction model is any one of the following: elastic network, support vector machine, K-nearest neighbor, random forest, and extreme gradient boosting tree.

[0018] As a preferred embodiment of this application, the preprocessing of the multi-dimensional fermentation parameters in step 2 includes:

[0019] (a) Normalization processing: When using support vector machine and K-nearest neighbor model, the multi-dimensional fermentation parameters in the fermentation process database are normalized to eliminate the influence of dimensional differences on model training.

[0020] (b) Feature interaction processing: When using elastic networks, random forests and extreme gradient boosting tree models, interaction terms between feature variables are constructed through multinomial feature generation methods.

[0021] (c) Feature selection: Key feature variables are selected from normalized features and interaction features using the recursive feature elimination method.

[0022] As a preferred embodiment of this application, the normalization process is performed according to formula (1):

[0023] (1);

[0024] Where: xnorm The normalized value; x i x represents the original data value. min x is the minimum value in the data. max This represents the maximum value in the data.

[0025] As a preferred embodiment of this application, all computations, model building, and plotting in the feature interaction processing of step (b) are implemented using the tidymodels framework in R and the tensorflow framework in Python to generate interaction items.

[0026] As a preferred embodiment of this application, in step 3, the hyperparameters of five model algorithms are tuned on the training set using a grid search method, 50 sets of hyperparameter combinations are randomly generated, and finally the coefficient of determination R is used. 2 The root mean square error (RMSE) is used as an evaluation metric, and its calculation formulas are as follows:

[0027] (2);

[0028] (3); Where: R 2 The coefficient of determination represents the proportion of variability explained by the model to the total variability; the closer it is to 1, the better the model explains the variability of the observations. RMSE (Root Mean Square Error) is the square root of the average of the squares of the differences between the actual values ​​and the model's predicted values. It measures the deviation between the model's predicted values ​​and the actual values; the smaller the RMSE, the smaller the model's prediction error and the better its predictive performance. RES It is the sum of squared residuals; y is the total sum of squares; n is the sample size; i For the i-th observation; Let be the i-th predicted value.

[0029] As a preferred embodiment of this application, in step 3, based on the grid search results, Bayesian optimization and simulated annealing are used to iteratively optimize the hyperparameters within each model. The maximum number of iterations for Bayesian optimization and simulated annealing is set to 100, and the iteration stops after 10 iterations without improvement.

[0030] As a preferred embodiment of this application, in step 3, the final constructed bio-soft sensor model (*Sim_SVM_rbf) is a support vector machine (Sim_SVM_rbf) using a radial basis kernel function. This model is obtained by retraining on a dataset that combines the training set and the validation set.

[0031] As a preferred embodiment of this application, the bio-soft sensor model (*Sim_SVM_rbf) calculates the average of the absolute values ​​of the SHAP values ​​of the feature variables for each sample during training (Mean |SHAP Value|). This value is used to reflect the contribution of each feature variable to the overall model; the larger the value, the greater the influence.

[0032] As a preferred embodiment of this application, in step 5, the coupling of the constructed bio-soft sensor model with the multi-objective optimization algorithm includes:

[0033] (1) Construct a front-end neural network architecture with 3 inputs, 2 hidden layers, and 6 outputs. Its inputs are time, fermentation temperature and pH, and its outputs are dissolved oxygen, stirring speed, ammonia consumption volume, feed volume, residual sugar and OD600.

[0034] (2) The output of the pre-neural network is used as the input to the biosoft sensor;

[0035] (3) Particle Swarm Optimization (PSO) and Genetic Algorithm (GA) were used to optimize the fermentation temperature (Tem) and pH to find the optimal fermentation process conditions for a specific time period in order to improve the yield of D-pantothenic acid.

[0036] (4) Based on the results of the optimization algorithm, conduct actual fermentation experiments using the corresponding fermentation process conditions.

[0037] As a preferred embodiment of this application, the hidden layer structure of the pre-neural network is as follows: the first layer has 128 neurons and the second layer has 64 neurons.

[0038] Compared with the prior art, the beneficial effects of this application are:

[0039] (1) Multidimensional fermentation parameters under different process conditions were collected through multiple fermentation experiments, and a fermentation process database containing multiple feature variables was constructed. The training set, validation set and test set were reasonably divided to carry out the machine learning process.

[0040] (2) This application uses five machine learning models, including elastic networks and support vector machines, and performs hyperparameter tuning through grid search, Bayesian optimization, and simulated annealing. Combined with 5-fold cross-validation, the support vector machine based on the radial basis kernel function is selected as the best performing model, with an RMSE of 4.93 and an R-value of 1.5% on the validation set. 2 The score reached 0.982, with an RMSE of 5.33 on the test set. 2 With a value of 0.977, it can accurately predict the yield of D-pantothenic acid, effectively overcoming the drawbacks of traditional detection methods.

[0041] (3) A front-end neural network was constructed using PSO and GA and combined with a bio-soft sensor to optimize fermentation process conditions. GA and PSO respectively provided optimal solutions, and actual fermentation experiments were conducted near their solution spaces. In the actual fermentation experiment, the highest yield of D-pantothenic acid reached 68h-112.6g / L, which is close to the result of the optimization algorithm, confirming the accuracy and feasibility of the bio-soft sensor model-coupled optimization algorithm. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the development process for the DPA bio-soft sensor model.

[0043] Figure 2 This is a schematic diagram of the DPA bio-soft sensor coupling optimization algorithm.

[0044] Figure 3 This is one of the hyperparameter tuning methods in Example 1; where (A) is the grid search method and (B) is the Bayesian optimization method.

[0045] Figure 4 This is the second hyperparameter tuning example in Example 1; where (C) is the simulated annealing method; and (D) is the hyperparameter tuning result of the grid search method.

[0046] Figure 5 This is the third example of hyperparameter tuning in Example 1; where (E) is the hyperparameter tuning result of Bayesian optimization method; and (F) is the hyperparameter tuning result of simulated annealing method.

[0047] Figure 6 This is an evaluation of the prediction performance of *Sim_SVM_rbf in Example 1.

[0048] Figure 7 The image shows the fermentation results of DPA in a 5 L bioreactor in Example 1.

[0049] Figure 8 This represents the iterative process of the GA algorithm; the red line represents the maximum value, the blue line represents the mean value, and the green line represents the minimum value. Detailed Implementation

[0050] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application.

[0051] It should be noted that the process equipment or apparatus not specifically mentioned in the following embodiments are all conventional equipment or apparatus in the art.

[0052] Furthermore, it should be understood that the existence of other method steps before or after the combined steps, or the insertion of other method steps between these explicitly mentioned steps, does not preclude the existence of other method steps before or after the combined steps, or the insertion of other method steps between these explicitly mentioned steps, unless otherwise stated. It should also be understood that the combined connection relationship between one or more devices / apparatus mentioned in this invention does not preclude the existence of other devices / apparatus before or after the combined devices / apparatus, or the insertion of other devices / apparatus between these explicitly mentioned devices / apparatus, unless otherwise stated. Moreover, unless otherwise stated, the numbering of each method step is merely a convenient tool for identifying each method step, and not for limiting the order of the method steps or limiting the scope of the invention. Changes or adjustments to their relative relationships, without substantially altering the technical content, should also be considered within the scope of the invention.

[0053] The parent strain E. coli W3110 described in this invention was deposited at the Coli Genetic Stock Center of Yale University on August 5, 1975, with accession number CGSC#4474, and has been disclosed in patents US 2009 / 0298135A1 and US 2010 / 0248311 A1.

[0054] The genetically engineered bacterium of this invention (hereinafter referred to as strain W1) is the sclerotium ZJUTDPAP16 (E. coli W3110, Trc-EcilvD* / BspanBA* / CgpanC* / alsS* / BspanB* / nac GTG / ΔP 345ptsH / gltA GTG / gltA TTG / Trc-PpfkB, already disclosed in CN116590209A).

[0055] The HPLC method for determining D-pantothenic acid content is as follows: Chromatographic conditions: C18 column (250 × 4.6 mm, particle size 5 μm, Agilent Technologies Co., Santa Clara, CA, USA), detection wavelength: 200 nm, column temperature: 30℃; Sample preparation: The sample was diluted with ultrapure water to maintain the D-pantothenic acid content between 0.05 g / L and 0.40 g / L; Mobile phase: acetonitrile / water / phosphoric acid: (50 / 949 / 1); Data acquisition time: 25 min.

[0056] Example 1

[0057] The present invention provides a multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins, comprising:

[0058] S1, Collection of training data: Multi-dimensional fermentation parameters of strain W1 for the production of D-pantothenic acid were collected in a 5L fermenter at different temperatures and pH levels. A fermentation process database of strain W1 was constructed as the basis for building a biosoft sensor for strain W1.

[0059] The multidimensional fermentation parameters include fermentation time (Time), OD (Oxygen Demand), and so on. 600 Fermentation temperature (Tem), pH, dissolved oxygen (DO), rotational speed, ammonia concentration, feed supplement, and residual sugar.

[0060] By using normalization, the collected data is mapped to the [0,1] interval to eliminate the influence of different parameter units, which facilitates subsequent model training.

[0061] Specifically, when using moving average filtering, let x be the real-time data collected within the window. i ;

[0062] x i Let x be the i-th original data value, i∈{1,2,...,n}; n is the window size, and x is the minimum value among the real-time data collected within the window. min The maximum value among the real-time data collected within the window is x. max The normalized value is:

[0063] (1);

[0064] Where: x norm The normalized value; x min x is the minimum value in the data. max This represents the maximum value in the data.

[0065] Through the above steps, methods, and formulas, real-time and accurate acquisition and preprocessing of fermentation parameters were achieved, providing a reliable data foundation for subsequent optimization of D-pantothenic acid yield prediction models and control strategies.

[0066] S2, Model Construction: Using preprocessed multidimensional fermentation parameters as input feature variables and D-pantothenic acid yield as model output, a multivariate prediction model is established to achieve accurate prediction of D-pantothenic acid yield by the initial biosoft sensor.

[0067] Specifically, fermentation time (Time) and OD 600Nine characteristic variables—fermentation temperature (Tem), pH, dissolved oxygen (DO), rotational speed, ammonia concentration, feed supplementation, and residual sugar—were used as inputs to the model, with D-pantothenic acid yield as the output. A multivariate prediction model was established using candidate models; these candidate models could be any one of elastic network, support vector machine, K-nearest neighbor, random forest, and extreme gradient boosting tree.

[0068] The specific process is as follows:

[0069] Hyperparameter tuning of five model algorithms was performed on the training set using a grid search method. This process was combined by randomly generating 50 sets of hyperparameter combinations, and finally, the coefficient of determination R was used. 2 The root mean square error (RMSE) is used as an evaluation metric, and its calculation formulas are as follows:

[0070] (2);

[0071] (3); Where: R 2 The coefficient of determination represents the proportion of variability explained by the model to the total variability; the closer it is to 1, the better the model explains the variability of the observations. RMSE (Root Mean Square Error) is the square root of the average of the squares of the differences between the actual values ​​and the model's predicted values. It measures the deviation between the model's predicted values ​​and the actual values; the smaller the RMSE, the smaller the model's prediction error and the better its predictive performance. RES It is the sum of squared residuals; y is the total sum of squares; n is the sample size; i For the i-th observation; Let be the i-th predicted value.

[0072] The results show that the support vector machine model algorithm based on radial basis function kernel function performs best in 5-fold cross-validation, with the lowest RMSE of 7.38 and R0. 2 It reached 0.95.

[0073] Based on the grid search results, Bayesian optimization and simulated annealing were used to iteratively optimize the hyperparameters within each model. The maximum number of iterations for Bayesian optimization and simulated annealing was set to 100, and iteration was stopped after 10 iterations without improvement.

[0074] More specifically: The Bayesian optimization process consists of four steps:

[0075] 1. Initially sample some parameter combinations for evaluation and train an initial agent model;

[0076] 2. The acquisition function is used to select the next hyperparameter combination to be evaluated based on the prediction results of the surrogate model;

[0077] 3. Repeat the iteration until the preset number of iterations is reached or the stopping condition is met;

[0078] 4. Select the optimal hyperparameters.

[0079] The main processes of simulated annealing include:

[0080] 1. Randomly select a set of hyperparameters as the initial solution;

[0081] 2. Randomly generate a new solution within the neighborhood of the current solution;

[0082] 3. Calculate the objective function values ​​corresponding to the new solution and the current solution. If the objective function value of the new solution is better, then accept the new solution; otherwise, accept the new solution with a certain probability. This probability is related to the current "temperature". The higher the temperature, the greater the probability of accepting the worse solution.

[0083] 4. As the number of iterations increases, the temperature is gradually reduced, which gradually decreases the probability of accepting a poor solution, thus finding the optimal solution.

[0084] After hyperparameter tuning using the two methods described above, the results show that the support vector machine with radial basis function kernel is the optimal choice. Following simulated annealing optimization, its RMSE further decreased from 7.38 to 7.26.

[0085] Step 3, Model Optimization and Validation: The hyperparameters of the initial biosoft sensor model are tuned using a grid search method, and the generalization ability of the model is verified on an independent validation set to construct the biosoft sensor model.

[0086] Specifically, this step involves evaluating the bio-soft sensor model based on a machine learning strategy: The hyperparameter-tuned model is validated on a validation set. The Sim_SVM_rbf radial basis function support vector machine, originally optimized by simulated annealing, showed the best performance on both the training and validation sets, with the lowest RMSE of 4.93. 2 The value was 0.981. Therefore, we chose to retrain Sim_SVM_rbf on the dataset that combined the training and validation sets to obtain *Sim_SVM_rbf, which will become the final biosensor we constructed for the high-yield D-pantothenic acid strain W1.

[0087] Step 4: Model interpretability analysis: The SHAP value analysis method is used to quantify the contribution of each feature variable value to the prediction results. Through feature importance ranking and local interpretation visualization, the interpretability analysis of the biosoft sensor model is carried out to explain the prediction results of the biosoft sensor model.

[0088] Interpretability analysis was performed on the bio-soft sensor model based on a machine learning strategy. Specifically, the mean (Mean |SHAP Value|) of the absolute values ​​of the feature variables for each sample during training of the previously constructed bio-soft sensor model *Sim_SVM_rbf was calculated. The results showed that the two most influential feature variables were Feed Supplement and Time, with Mean |SHAP Value| values ​​of 18.42 and 11.60, respectively, which is consistent with the actual fermentation process.

[0089] Step 5: Fermentation process optimization: The constructed bio-soft sensor model is coupled with a multi-objective optimization algorithm to form a fermentation process optimization system. The fermentation process parameters are adjusted with the objective functions of maximizing D-pantothenic acid yield and minimizing energy consumption.

[0090] Specifically, Tem and pH were optimized during the fermentation process of strain W1. First, a pre-neighbor neural network architecture with 3 inputs, 2 hidden layers (128 neurons in the first layer and 64 neurons in the second layer) and 6 outputs was constructed. Using Time, Tem and pH as inputs, other fermentation process data automatically controlled by the fermenter instrument were fitted. Then, the output of the pre-neighbor neural network was used as input to the constructed bio-soft sensor model.

[0091] In this step, GA and PSO were used to optimize Tem and pH, and the optimal fermentation process conditions were sought for a specific time period to further increase the yield of D-pantothenic acid.

[0092] The GA algorithm yielded the optimal solution at generation 645: Time = 63.9h, Tem = 30.1℃, and pH = 6.77; the optimal solution for PSO was: Time = 65.8h, Tem = 30.0℃, and pH = 6.78.

[0093] Based on the results of the optimization algorithm, actual fermentation experiments were conducted using corresponding fermentation process conditions. The experiments were performed at a temperature of 30℃ and pH values ​​of 6.7 and 6.8. The results showed that when Tem was set at 30℃ and pH at 6.8, the yield of D-pantothenic acid reached a maximum of 112.6 g / L after 68 h, representing a 4.9% increase compared to the maximum value in the database (107.3 g / L). This result demonstrates the feasibility of using the bio-soft sensor coupling optimization algorithm to optimize fermentation process conditions.

[0094] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.

Claims

1. A multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins, characterized in that, Includes the following steps: Step 1: Data Acquisition: Collect multi-dimensional fermentation parameters during the fermentation of D-pantothenic acid by genetically engineered strains under different fermentation process conditions, and construct a fermentation process database for the genetically engineered strains. Step 2, Model Construction: Using the pre-processed multi-dimensional fermentation parameters as input feature variable values ​​and D-pantothenic acid yield as the model output value, a multivariate prediction model is established. Step 3, Model Optimization and Validation: The hyperparameters of the multivariate prediction model are tuned using the grid search method, and the generalization ability of the model is verified on the independent validation set to construct a biosoft sensor model. Step 4: Model interpretability analysis: The SHAP value analysis method is used to quantify the contribution of each feature variable value to the prediction results. Through feature importance ranking and local interpretation visualization, the interpretability analysis of the biosoft sensor model is carried out to explain the prediction results of the biosoft sensor model. Step 5: Fermentation process optimization: The constructed bio-soft sensor model is coupled with a multi-objective optimization algorithm to form a fermentation process optimization system. The fermentation process parameters are adjusted with the objective functions of maximizing D-pantothenic acid yield and minimizing energy consumption.

2. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 1, characterized in that, The sample data in the fermentation process database are randomly divided into a training set, a validation set, and a test set in a 6:2:2 ratio. The training set is used to train the machine learning model, the validation set is used to evaluate the performance of the model after training, and the test set is used for the final performance evaluation of the model.

3. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 1, characterized in that, In step 2, the candidate model used to establish the multivariate prediction model is any one of the following: elastic network, support vector machine, K-nearest neighbor, random forest, and extreme gradient boosting tree.

4. The method as described in claim 3, characterized in that, Step 2 involves preprocessing the multidimensional fermentation parameters, including: (a) normalization. When using support vector machines and K-nearest neighbors models, the multidimensional fermentation parameters in the fermentation process database are normalized to eliminate the influence of dimensional differences on model training. (b) Feature interaction processing: When using elastic networks, random forests, and extreme gradient boosting tree models, interaction terms between feature variables are constructed using a multinomial feature generation method; and (c) Feature selection: Key feature variables are selected from normalized features and interaction features using the recursive feature elimination method.

5. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 4, characterized in that, Normalization is performed according to formula (1): (1); Where: x norm The normalized value; x i Let x be the i-th original data value, i∈{1,2,...,n}; n is the window size; min x is the minimum value in the data. max This represents the maximum value in the data.

6. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 1, characterized in that, In step 3, the hyperparameters of the five model algorithms are tuned on the training set using a grid search method, 50 sets of hyperparameter combinations are randomly generated, and finally the coefficient of determination R is used. 2 The root mean square error (RMSE) is used as an evaluation metric, and its calculation formulas are as follows: (2); (3); Where: R 2 The coefficient of determination is RMSE; the root mean square error is SS. RES It is the sum of squared residuals; y is the total sum of squares; n is the sample size; i For the i-th observation; Let be the i-th predicted value.

7. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 6, characterized in that, In step 3, based on the grid search results, Bayesian optimization and simulated annealing are used to iteratively optimize the hyperparameters within each model. The maximum number of iterations for Bayesian optimization and simulated annealing is set to 100, and the iteration stops after 10 iterations without improvement.

8. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 1, characterized in that, In step 3, the final biological soft sensor model is a support vector machine using radial basis function kernel function. This model is obtained by retraining on a dataset that combines the training set and the validation set.

9. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 1, characterized in that, In step 5, the coupling of the constructed bio-soft sensor model with the multi-objective optimization algorithm includes: (1) A front-end neural network architecture with 3 inputs, 2 hidden layers, and 6 outputs was constructed. The inputs are time, fermentation temperature, and pH, and the outputs are dissolved oxygen, stirring speed, ammonia consumption volume, feed volume, residual sugar, and OD. 600 ; (2) The output of the pre-neural network is used as the input to the biosoft sensor; (3) The fermentation temperature and pH were optimized using particle swarm optimization and genetic algorithm to find the optimal fermentation process conditions under a specific time period in order to improve the yield of D-pantothenic acid. (4) Based on the results of the optimization algorithm, conduct actual fermentation experiments using the corresponding fermentation process conditions.

10. The multi-objective dynamic optimization method for D-pantothenic acid fermentation based on bio-digital twins as described in claim 9, characterized in that, The hidden layer structure of the aforementioned front-end neural network is as follows: the first layer has 128 neurons, and the second layer has 64 neurons.

Citation Information

Patent Citations

  • Genetically engineered bacterium for producing D-pantothenic acid, construction method and application

    CN116590209A

  • Method for fermentative production of L-methionine

    US20090298135A1

  • Increasing methionine yield

    US20100248311A1