An artificial intelligence method for adaptive dynamic precision fermentation control

CN121325786BActive Publication Date: 2026-08-11ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

虽然这些方法在氨基酸、有机酸等传统发酵产品生产中积累了丰富经验,但在应对合成生物学构建的新型细胞工厂时暴露出显著不足:(1)适应性缺陷:固定控制策略无法适应原料批次差异和环境扰动

Benefits of technology

[0041] 1. During the production process, the control model is continuously optimized through the "learning-verification-rollback-replacement-iteration" model to achieve production goals using production equipment. It can accurately grasp the characteristics of specific production equipment and their interaction with the cell factory, thereby maximizing the production potential of the cell factory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121325786B_ABST
    Figure CN121325786B_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence method for adaptive, dynamic, and precise fermentation control. During the production process, it continuously optimizes the control model's ability to achieve production goals using production equipment through a "learn-try-rollback-substitution-iteration" model. This method accurately grasps the characteristics of specific production equipment and their interaction with the cell factory, maximizing the cell factory's production potential. Employing a control model independent of value functions significantly improves the model's control accuracy and enhances its effectiveness and stability in achieving production goals. The control model uses a two-level configuration mode, supporting both precise parameter adjustments at individual time points and setting constant strategies for time intervals. It iteratively converges the optimal parameter combination using a gradient descent algorithm. This flexibility is of significant value in actual production; for example, a fixed strategy can be used in the early stages of fermentation to ensure microbial activity, while dynamic adjustments can be made later to improve product synthesis efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically an artificial intelligence method for adaptive, dynamic, and precise fermentation control. Background Technology

[0002] The biomanufacturing industry is entering a golden age of rapid development. According to the OECD, by 2030, its scale will reach 35% of global industrial output. In this field, China's fermentation industry occupies a pivotal position, ranking first in the world in scale, with its amino acid and vitamin products accounting for 60%-80% of the global market.

[0003] With the continuous advancement of synthetic biology technology, microbial cell factories—biomanufacturing systems that engineer microbial cells (such as bacteria and yeast) to optimize metabolic pathways for the efficient synthesis of high-value products such as drugs and chemicals—have become a core platform for green production. However, the current industry development faces a critical bottleneck: how to establish a precise fermentation process control system to ensure that laboratory-constructed cell factories maintain optimal performance during industrial scale-up. This technological breakthrough will directly impact the high-quality development level of the biomanufacturing industry.

[0004] In terms of control methods, the current industrial fermentation field still widely adopts static control strategies based on preset parameters, including classical control modes such as proportional-integral-derivative (PID) control and fixed-value feeding batch fermentation. Although these methods have accumulated rich experience in the production of traditional fermentation products such as amino acids and organic acids, they have exposed significant shortcomings when dealing with novel cell factories constructed by synthetic biology: (1) Adaptability defects: Fixed control strategies cannot adapt to batch differences in raw materials and environmental disturbances. For example, fluctuations in the composition of composite raw materials can lead to differences in cell growth rates, and traditional control strategies lack the ability to adjust in real time. (2) Insufficient dynamic response: Microbial metabolism has typical nonlinear time-varying characteristics, and this dynamic change makes parameter tuning of traditional PID control extremely difficult.

[0005] Advances in modern monitoring technologies have enabled us to comprehensively characterize the state of fermentation systems through multimodal observation data, including external environmental parameters (such as temperature, pH, and dissolved oxygen) and internal omics data (such as transcriptomics, proteomics, and metabolomics). However, current online monitoring technologies still have limitations. The lag in offline detection leads to data gaps, and the system's insufficient robustness to noise interference further contributes to incomplete observation data. In practical applications, interpolation is usually used to fill in missing values ​​or replace outliers, but this process introduces errors, ultimately resulting in insufficient modeling accuracy.

[0006] Based on this monitoring data, we can dynamically regulate the fermentation process through control signals. Specific operational methods include environmental parameter adjustment (such as temperature and dissolved oxygen control), chemical intervention (such as adding small metabolic molecules), and nutritional strategies (such as fed-batch feeding). To overcome the limitations of traditional methods, advanced technologies such as Model Predictive Control (MPC) and expert systems have gradually developed in recent years, but they still face key technical bottlenecks in practical applications. First, reliance on mechanistic models: While MPC technology can achieve multivariate coordinated control, its performance is highly dependent on accurate mechanistic models. This is clearly insufficient to meet the timeliness requirements for industrialization of rapidly iterating synthetic biology strains. Second, lack of dynamic control capabilities: On the one hand, there is a lack of dynamic control handover mechanisms for production fluctuations; on the other hand, there is a lack of closed-loop iterative architecture to support continuous strategy optimization, leading to rigid and lagging control strategies. Third, low-quality modeling data further exacerbates the above problems, creating a vicious cycle.

[0007] As the complexity of biomanufactured products increases, the industry has put forward new requirements for fermentation control technology: (1) possessing universal control capabilities for a variety of complex products; (2) achieving dynamic and precise regulation based on real-time monitoring of cell metabolic state; and (3) supporting an iterative mechanism for continuously optimized intelligent control strategies. These requirements are driving the development of fermentation process control technology towards intelligence and self-adaptation.

[0008] To achieve the aforementioned requirements, this patent proposes an artificial intelligence method for adaptive, dynamic, and precise fermentation control. By acquiring multi-dimensional parameters in real time through a multi-modal sensor array and utilizing an intelligent model capable of dynamically optimizing control signals, a closed-loop iterative mode of "learning-trial-substitution-iteration" is employed to achieve precise control of the fermentation process. This control mode possesses adaptive, automated, and continuous optimization capabilities, significantly reducing human intervention. This solution is particularly suitable for optimizing industrial fermentation processes for high-value-added products such as antibiotics and enzyme preparations, offering significant economic benefits and application value. Summary of the Invention

[0009] This invention discloses an artificial intelligence method for adaptive dynamic precision fermentation control, which achieves optimized control through a closed-loop cycle of "learning-trial-rollback-substitution-iteration". The technical solution provided by this invention is as follows:

[0010] 1) Set the baseline control model A and the learning control model B for the current loop;

[0011] 2) During the learning phase, the baseline control model A controls the fermentation process, synchronously records and standardizes the observation data; the learning control model B learns and trains based on this observation data.

[0012] 3) During the trial phase, baseline control model A is suspended, and learning control model B takes over fermentation control;

[0013] 4) During the rollback phase, the baseline control model A resumes control, while the learning control model B continues to learn based on the observation data;

[0014] 5) In the replacement phase, the learning control model B replaces the reference control model A and becomes the reference control model A for the next cycle;

[0015] 6) The control model is iterated through the loop of steps 2)–5) above;

[0016] The control model in step 1) has the following characteristics: ① It completes time-series training using the collected time-series trajectory data of the coordinated changes in the fermentation system state and the fermentation control strategy; ② It can predict the fermentation system state at the next time point based on the historical trajectory of the fermentation system state and the fermentation control strategy of the current fermentation batch, as well as the current fermentation control strategy; ③ It can calculate the optimal control strategy scheme from the current state to the target state based on the historical trajectory of the fermentation system state and the fermentation control strategy of the current fermentation batch, with the expected system state at the target time point as the control objective.

[0017] As a further improvement, the control model described in step 1) has the following characteristics: ① It adopts a control model that does not depend on a value function; ② It adopts a control model that does not assume that the initial states of batches are the same; ③ It adopts a neural network control model that can be trained without requiring standardized training data; ④ It adopts a two-level configuration method to set the control trajectory, specifically including: configuring the control strategy: defining the set of all controllable parameters at any specific time point; configuring the selection trajectory: establishing the correspondence between all time points in the fermentation process and the control strategy; ⑤ When configuring the control strategy, all or some of the controllable parameters are set as unknown parameters, and the unknown parameters are estimated by the neural network model.

[0018] As a further improvement, the conditions for switching from the learning stage in step 2) to the trial stage in step 3) are as follows: During the period when the baseline control model A is in control, the learning control model B predicts the state changes of the fermentation system caused by the control strategy of the baseline control model A, and verifies the difference between the prediction results and the actual observed state changes; a fixed-length sliding monitoring window is used, and if the average prediction error of the learning control model B is lower than the preset threshold within the time window, the transition from the learning stage to the trial stage will be triggered.

[0019] As a further improvement, the conditions for switching from the trial phase (step 3) to the rollback phase (step 4) are as follows: During the period when the learning control model B is in control, the learning control model B is continuously used to predict the state changes of the fermentation system caused by its own control strategy, and the difference between the prediction results and the actual observed state changes is verified; a fixed-length sliding monitoring window is used, and if the average prediction error of the learning control model B within the time window is higher than a preset threshold, the rollback from the trial phase to the learning phase will be triggered.

[0020] As a further improvement, the condition for switching from the trial stage in step 3) to the replacement stage in step 5) is that the learning control model B maintains control at no less than 90% of the time points throughout the entire fermentation process of the current batch.

[0021] As a further improvement, in step 5), the condition for the learned control model B to replace the baseline control model A and become the baseline control model A for the next cycle is: when the learned control model B can trigger the replacement phase in five consecutive fermentation batches.

[0022] As a further improvement, the control model described in step 1) includes four cooperating sub-modules: a feature encoding module F, a forward generation module G, a backward generation module G', and a decoding module D. Each module works collaboratively to achieve the functions of encoding, predicting, backtracking, and decoding time-series features. Specifically, module F adopts a single-layer neural network structure, with an independent F module at each time point. Its input is a constant value of 1, and its output is a k-dimensional feature vector corresponding to the time point. To improve training efficiency, all F modules at all time points are integrated into a parallel processing module, with its input layer having T neurons corresponding to the total number of observation time points, and its output layer having k neurons corresponding to the feature vector dimension. Modules G and G' both adopt a multi-layer neural network architecture, where module G uses the current time... The system uses the k-dimensional feature vector and m-dimensional control strategy of the previous time point as input to predict the k-dimensional feature vector of the next time point. The G' module realizes state backtracking by constructing a reverse time correlation. It reconstructs the k-dimensional feature vector of the previous time point using the k-dimensional feature vector of the current time point and the m-dimensional control strategy of the previous time point as input. The input layers of both modules G and G' are configured with k+m neurons, and the output layers are configured with k neurons. The D module adopts a multi-layer neural network and is responsible for mapping the k-dimensional feature vector to the system state value in the d-dimensional observation space. Its input layer is set with k neurons, and its output layer is set with d neurons corresponding to the number of system state indicators. All modules adopt a fully connected structure and initialize the network weight parameters using a Xavier normal distribution.

[0023] As a further improvement, the process of synchronously recording and standardizing observation data described in step 2) is as follows: b batches of fermentation process data are synchronously collected using online monitoring and offline detection equipment. Each batch contains t observation time points, with a total number of time points T = b * t. The raw data is organized into two structured datasets: a system state dataset and a control strategy dataset. The system state dataset is a T × d dimensional matrix containing d observation indicators characterizing the fermentation system state. This dataset allows for missing values ​​due to detection errors or sampling intervals. The control strategy dataset is a T × m dimensional matrix containing m controllable control signals. A complete data record for each time point consists of a set of d-dimensional system state vectors and a set of m-dimensional control strategy vectors. Data standardization is performed using Min-Max standardization or Z-Score standardization based on parameter characteristics. Each feature dimension is normalized separately to ensure comparability of parameters with different dimensions. The resulting standardized dataset is used to construct training data samples.

[0024] As a further improvement, the time-series training described in step 6) using the collected time-series trajectory data of the coordinated changes in the fermentation system state and the fermentation control strategy includes the following steps:

[0025] 1) Generate a pair of positive and negative training samples for any two time points in each batch, for a total of s pairs.

[0026] To learn the dynamic relationship between state and control across different time spans;

[0027] 2) Construct a T×T diagonal matrix A, where column p of row p (p = (i-1)*t+j, representing the j-th time point of the i-th batch) is 1, and the remaining columns are 0; generate the following for two time points j1 and j2 (j2 = j1+g) in the i-th batch that are g time units apart: ① Forward training samples: The time identifier vector is the p1-th row of matrix A (p1 = (i-1)*t+j1), denoted as i f The counting step size is g; the observation vector is the p2th row of the system state dataset (p2 = (i-1)*t + j2), denoted as x. f Control strategy matrix c f (Dimension is g×m), recording g groups of continuous control strategies from time point j1 to j2-1; ② Reverse training samples: the time label vector is the p2th row of matrix A, denoted as i b The counting step size is g; the observation vector is row p1 of the system state dataset, denoted as x. b Control strategy matrix c b (Dimension is g×m), record the g groups of reverse control strategies from time point j2-1 to j1;

[0028] 3) Positive training: i f Inputting the F module yields a k-dimensional initial feature vector. If g equals 0, directly inputting the D module yields the predicted value, denoted as... If g is not 0, then the eigenvector is compared with c. f The first row is integrated into a (k+m) dimensional vector, which is input into the G module to generate the k-dimensional feature vector for the next time point, while g is decremented by 1; this process is repeated until g reaches zero; the final feature vector is input into the D module to obtain the predicted value. Calculate the loss value The formula is

[0029] 4) Reverse training: i b Inputting the F module yields a k-dimensional initial feature vector; if g equals 0, directly inputting the D module yields the predicted value, denoted as... If g is not 0, then the eigenvector is compared with c. b The first row is integrated into a (k+m) dimensional vector, which is input into the G' module to generate the k-dimensional feature vector of the previous time point, while g is decremented by 1; this process is repeated until g returns to zero; the final feature vector is input into the D module to obtain the predicted value. Calculate the loss value The formula is

[0030] 5) Model optimization: Total loss value After aggregating the loss of all samples, the model gradient is calculated through backpropagation. The Adam optimizer is used to update the network weight parameters of the F, G, G' and D modules. When there are missing observations, the corresponding dimension error is masked out and does not participate in the gradient calculation.

[0031] As a further improvement, the prediction of the fermentation system state at the next time point based on the historical trajectory of the fermentation system state and fermentation control strategy of the current fermentation batch, as well as the current fermentation control strategy, as described in step 6) includes the following steps:

[0032] 1) Based on the n×d system state data and n×m control strategy data of the current fermentation batch at n time points, an independent F module is constructed for the current time point, with 1 input layer node and k output layer nodes; the input value 1 is given to the F module to generate a k-dimensional feature vector, which is then processed by the D module to obtain a d-dimensional state prediction value. At the same time, the feature vectors of the previous n-1 time points are reconstructed sequentially through the G' module and transformed into state prediction values ​​of historical time points through the D module; the mean square error between the predicted values ​​and the actual observed values ​​at the current time point and all historical time points is calculated; the G, G' and D modules are fixed, and the network weight parameters of the F module are updated using the gradient descent method.

[0033] 2) The optimized F module is used to obtain the feature vector at the current time point. The current feature vector is integrated with the current control strategy and input into the G module to obtain the feature vector at the next time point. The final predicted value is output through the D module to achieve single-step prediction. By connecting the G module in series, the feature vector sequence of subsequent time points is generated in sequence, and the final predicted value is output through the D module to achieve multi-step prediction.

[0034] The optimal control strategy for transitioning from the current state to the target state, as described in step 6), includes the following steps:

[0035] 1) Construct the control configuration matrix: Generate an h×h one-hot encoding matrix B, with rows corresponding to h time points to be optimized and columns representing control strategy groups. Control strategies are bound by setting the same non-zero column index to achieve constant control within the time interval.

[0036] 2) Create a control generation module: Construct a control generation module C with a single-layer fully connected neural network structure. The number of input layer neurons is h, and the number of output layer neurons is m. The module is randomly initialized using a Xavier normal distribution. Input the control configuration matrix B into module C to obtain an initial control policy matrix with dimensions h×m.

[0037] 3) Target-oriented strategy optimization: First, obtain the feature vector at the current time point, integrate it with the first row of the control strategy matrix output by module C, and input it into module G to generate the feature vector at the next time point; iterate this process h times to obtain the feature vector at the target time point, convert it into the system state prediction value by module D, and calculate the mean square error loss with the preset target state value.

[0038] During the optimization process, the network weight parameters of all other modules are fixed, and the network weight parameters of only module C are optimized using the gradient descent algorithm; after the loss function converges, matrix B is input into the optimization.

[0039] The C module then outputs the optimal control strategy matrix containing h time points.

[0040] Beneficial effects

[0041] 1. During the production process, the control model is continuously optimized through the "learning-verification-rollback-replacement-iteration" model to achieve production goals using production equipment. It can accurately grasp the characteristics of specific production equipment and their interaction with the cell factory, thereby maximizing the production potential of the cell factory.

[0042] 2. By adopting a control model that does not rely on value functions, the errors caused by the use of inaccurate value functions in traditional control models are avoided, significantly improving the control accuracy of the model and enhancing the effectiveness and stability of the model in achieving production goals;

[0043] 3. A control model that does not assume the same initial state for each batch is adopted. During the control process, the model's estimation of the system state at past time points for this batch is constantly revised based on the observed system state and control signal changes. This enables adaptive and precise control for each fermentation batch, significantly improving the effectiveness and stability of the model in achieving production goals.

[0044] 4. The neural network control model, which can be trained without requiring standardized training data, avoids the errors introduced by interpolation to fill in missing values ​​or replace outliers in application scenarios where the observation of system state often has missing values ​​and outliers, thus significantly improving the accuracy and robustness of the model.

[0045] 5. A control model with two configuration modes is adopted, supporting both precise parameter adjustment at a single time point and setting a constant strategy for a time interval; and the optimal parameter combination is iteratively converged through a gradient descent algorithm. This flexibility is of great value in actual production, for example, using a fixed strategy to ensure microbial activity in the early stages of fermentation, and then dynamically adjusting to improve product synthesis efficiency in the later stages. Attached Figure Description

[0046] Figure 1 A flowchart for the "learn-try-rollback-replace-iterate" control pattern;

[0047] Figure 2 A schematic diagram of an artificial intelligence control model framework;

[0048] Figure 3 A schematic diagram of a precision fermentation control system;

[0049] Figure 4 This is a diagram of the RelA-IκB dynamic system controlled by dual signals. Detailed Implementation

[0050] This invention proposes an intelligent control method applicable to bio-fermentation processes. In the fermentation system, state variables are constructed by real-time monitoring of multimodal indicators (such as temperature, dissolved oxygen concentration, cell density, and key metabolic concentrations); simultaneously, control variables are constructed based on the control signals from control equipment (such as operating parameters of acid / alkali pumps and aeration valves). By using an artificial neural network to learn the patterns of coordinated changes between the system's state indicators and control signals in the bioreactor, goal-oriented dynamic control is implemented.

[0051] To clarify the implementation process of this invention, a specific implementation method is described using the RelA-IκB dynamic expression and control system as an example. RelA-IκB is a classic gene transcriptional regulatory network: when the RelA transcription factor enters the cell nucleus, it can activate the transcription of target genes, including IκB; newly synthesized IκB-mRNA is transferred to the cytoplasm and translated into protein, then binds to RelA in the cytoplasm to form an inactive complex that remains in the cytoplasm, thus realizing a negative feedback regulatory loop. Simultaneously, nuclear RelA can also regulate the expression of other downstream target genes, with its transcription product ds-mRNA and translation product ds-Protein serving as response indicators for downstream transcription and translation levels, respectively. In this system, by introducing α-factor mating pheromones (α-factors) and ethanol as synergistic control signals, different control strategies can be constructed based on different concentration combinations, achieving dual control of the RelA-IκB dynamic behavior. Figure 4 ).

[0052] The ReLa-IκB theoretical model described above is used to simulate the internal biological mechanisms of the fermentation strain. Based on this model, given the initial state and the control strategy (α-factor / ethanol) at each time point, the trajectory of the fermentation system state evolution over time can be accurately calculated by solving the ReLa-IκB equations. This trajectory is considered as the reference trajectory of the actually observed system state in this example. The ReLa-IκB equations, as shown in Eq.1-7, define a total of 6 state variables, corresponding to 6 observation indicators of the system, including the cytoplasmic ReLa protein (S... RelA ), RelA-IκB binding protein (Sx), IκB protein (S I ), IκB-mRNA (S mI ), ds-mRNA(S m ) and ds-Protein(S p ).

[0053]

[0054]

[0055] R N =R tot -S RelA -Sx#(Eq.7)

[0056] α-factor (C α-factor ) and ethanol (C ethanol As a control signal, K is adjusted respectively. m and V h This alters the evolutionary trajectory of the RelA-IκB system, and its specific mechanism is shown in equation Eq.8-9.

[0057]

[0058] The parameter settings for the above equation system are based on the reference (Cell Systems, 2023, 14(5):382-391), and the specific values ​​are shown in the table below:

[0059]

[0060] 1. Learning phase: Model A implements control, Model B initializes.

[0061] First, model A is set as the baseline control model. Model A can be an artificial neural network model constructed using the method of this invention, or a control model obtained based on other methods, including traditional control models that determine fixed parameters through process optimization and adjust the control strategy accordingly. In this example, model A adopts a predefined linear control strategy: the α-factor concentration is 2*10 -5 The constant rate of increase was linear from 0 μM to 0.1 μM, while the ethanol concentration increased at a rate of 3*10 μM / min. -5 The constant rate of μM / min was linearly increased from 0 μM to 0.15 μM, and the entire control process lasted for 500 min.

[0062] Secondly, model B is constructed as a learning control model. In this example, an implementation of the artificial intelligence method described in claim 6 (Tac-BTSTN method, see paper mSystems, 2024, 9(8):e00697-24. "Learning metabolic dynamics from partial observations by bidirectional time-series mechanical network") is used to construct model B. This implementation provides a complete function interface, which can conveniently realize the training, prediction and control functions of the control model. The Tac-BTSTN framework mainly includes four cooperating sub-modules: feature encoding module (F), forward generation module (G), backward generation module (G') and decoding module (D). Specifically: Module F employs a single-layer neural network structure, with an independent F module at each time point; its input is a constant (e.g., the value 1), and its output is the feature vector of the corresponding time point; both modules G and G' employ multi-layer neural network architectures. Module G takes the feature vector of the current time point and the control policy as input to predict the feature vector of the next time point; module G' achieves state backtracking by constructing inverse time correlations, reconstructing the feature vector of the previous time point using the feature vector of the current time point and the control policy of the previous time point as input; module D employs a multi-layer neural network and is responsible for mapping the feature vectors to system state values ​​in the observation space. These modules work collaboratively to achieve the functions of encoding, predicting, backtracking, and decoding time-series features.

[0063] This example directly calls the Tac-BTSTN package. <btstn>Model B(obj) is created using the following class: The number of output layer neurons (o_hid_dims) in module F is set to 6; the number of hidden layer neurons (g_inner) in modules G and G' is set to 128; the number of hidden layers (g_layer) is set to 2; the number of hidden layer neurons (d_inner) in module D is set to 32; the number of hidden layers (d_layer) is set to 2; ReLU activation functions are used in modules G, G', and D; network weights are initialized using a Xavier normal distribution. To ensure data comparability, the Min-Max normalization method is used to normalize each index of the system state. The creation code is as follows:

[0064]

[0065] 2. Learning Phase: The Evolution of Model B's Online Learning System

[0066] Fermentation process data are obtained using the process described in claim 7 for synchronous recording and normalization of observation data. In this example, x is used. t Let u represent the system state observed at time point t. t Let S represent the control strategy implemented at time point t. The fermentation start time is set to t = 1, and the initial state x1 of the system is set to: S RelA =100μM, S C =100μM, S I =25μM, S mI =0μM, S m =0μM, S p =0μM, denoted as x1 = [100,100,25,0,0,0]; the initial control strategy is C α-factor =0μM, C ethanol =0μM, denoted as u1=[0,0].

[0067] During fermentation, the system state was dynamically monitored at 5-minute intervals (the system state was obtained by calculating the system of equations consisting of Eq. 1–9), and the following two datasets were updated in real time: ① System state dataset (Odata): Each row contains 8 columns of data, where the first column is the batch name (assigned the string 'batch 1'), the second column is the time point (in minutes, assigned values ​​of 5, 10, 15… according to the actual sampling time); columns 3-8 record the system state x at the corresponding sampling time point. t ② Control policy dataset (Cdata): Each row contains 4 columns of data. Columns 1-2 are consistent with Odata, and columns 3-4 record the control policy u at the corresponding sampling time point. t .

[0068] During the learning phase, Model A outputs a control policy and calculates the trajectory of system state changes using the RelA-IκB model. Every 5 minutes, the observed system state is updated to Odata, and the corresponding control policy is updated to Cdata. Upon receiving new data, Model B (obj) completes model training using the method described in claim 9. In this example, training can be started by calling its built-in `fit` method.

[0069]

[0070] 3. Learning Phase: Real-time evaluation of the prediction accuracy of Model B

[0071] After Model B completes training, Model B uses the method for predicting the state of the fermentation system at the next time point as described in claim 10 to predict the changes in the system state caused by the control strategy of Model A, and to verify the difference between the predicted results and the actual observed state changes.

[0072] In this example, the specific process is as follows: Assume that at the sampling time point t = 5 (the 25th minute), the control strategy output by model A is u5 = [5 * 10 -4 7.5*10 -4 After sorting, we get Sdata = ['batch 1', 25, 5*10]. -4 7.5*10 -4 Model B, based on the current Sdata, calls its built-in forecast method to predict the system state at the next time point. When t=6 (the 30th minute) is reached, the actual system state x6 is observed and compared with the predicted value. Perform error calculation: Where i represents six indicators of the system state.

[0073] The prediction performance of model B is monitored using a sliding window, with the following settings: the monitoring window length is set to 5 time points, and the prediction error threshold is ∈1; if the average prediction error δ of model B within the current monitoring window is... B If the value is below the threshold ∈1, then model A suspends control, and model B takes over the control of the fermentation process, and the system switches from the learning phase to the trial phase.

[0074]

[0075] 3. Trial Phase: Model B Implements Control

[0076] After Model B takes over the control of the fermentation process, the optimal control strategy calculation method described in claim 10 is used to execute the control strategy for the next moment based on real-time observation data.

[0077] In this example, the specific process is as follows: Assume the current sampling time point is t = 10 (the 50th minute), and the target time point is set to 450 minutes after the current time point (i.e., Δt = 90). The target state is that the downstream protein ds-Proten reaches 35 μM. Since the control target is only for ds-Proten, the other state indicators can be set to the missing value NA. Therefore, the target state data can be represented as rdata = [NA, NA, NA, NA, NA, 35].

[0078] Call the built-in `search_scheme` method of Model B to output the sequence of all control strategies s = [u 10 ,u 11 ,...,u 99 The system first uses u. 10 As the current control strategy, model B is based on u 10 Call the forecast method to predict the system state at time point 11. When t = 11 (the 55th minute) is reached, the observed system state x 11 and the predicted value Perform error calculation: Where i represents six indicators of the system state.

[0079] 4. Rollback Phase: Model A regains control.

[0080] During the control period of Model B, a sliding window is used to continuously monitor the control effect of Model B and verify the difference between the predicted results and the actual observed state changes. The specific settings are as follows: the monitoring window length is set to 5 time points, and the prediction error threshold is ∈2. If the average prediction error δ of Model B within the current monitoring window... B If the value exceeds a threshold ∈ 2, a rollback mechanism is triggered, meaning Model A regains control of the fermentation process, while Model B continues learning to optimize the network weights. If Model B maintains control for at least 90% of the time points during the entire fermentation process of the current batch, the system transitions from the trial phase to the replacement phase.

[0081]

[0082] 5. Replacement Phase: Model B replaces Model A as the new benchmark, initiating a new round of iterations.

[0083] If model B maintains stable control throughout five consecutive batch runs (i.e., model B holds control for at least 90% of the fermentation time in each batch), it signifies that model B has passed stability verification and meets the conditions for model replacement. At this point, the system will automatically trigger the baseline control model update process: first, the better-performing model B replaces the original model A, making it the new baseline control model; simultaneously, the system will initialize the new model B, with its initial network weight parameters being the same as the previous version of model B, thus initiating the next round of training. Through the iteration of the aforementioned steps, the control model is iterated.

[0084] The above is not intended to limit the specific embodiments of this patent. It should be noted that those skilled in the art can make various changes, modifications, additions, or substitutions without departing from the essential scope of this invention, and these improvements and refinements should also be considered within the scope of protection of this invention.< / btstn>

Claims

1. An artificial intelligence method for adaptive dynamic precision fermentation control, characterized in that, Optimized control is achieved through a closed-loop process of "learning-trying-rollback-substitution-iteration," including the following steps and features: 1) Set the baseline control model A and the learning control model B for the current loop; 2) During the learning phase, the baseline control model A controls the fermentation process, synchronously recording and standardizing the observation data; The learning control model B is trained based on this observation data; 3) During the trial phase, baseline control model A is suspended, and learning control model B takes over fermentation control; 4) During the rollback phase, the baseline control model A resumes control, while the learning control model B continues to learn based on the observation data; 5) In the replacement phase, the learning control model B replaces the reference control model A and becomes the reference control model A for the next cycle; 6) The control model is iterated through the loop of steps 2) – 5) above; The control model in step 1) has the following characteristics: ① Time-series training was completed using the collected time-series trajectory data of the coordinated changes in the state of the fermentation system and the fermentation control strategy; ② It can predict the state of the fermentation system at the next time point based on the historical trajectory of the fermentation system state and fermentation control strategy of the current fermentation batch, as well as the current fermentation control strategy; ③ Based on the current fermentation batch's fermentation system state and the historical trajectory of the fermentation control strategy, and with the expected system state at the target time point as the control objective, the system can calculate the optimal control strategy scheme from the current state to the target state. The conditions for switching from the learning phase (step 2) to the trial phase (step 3) are as follows: During the period when the baseline control model A is in control, the learning control model B predicts the state changes of the fermentation system caused by the control strategy of the baseline control model A, and verifies the difference between the prediction results and the actual observed state changes; a fixed-length sliding monitoring window is used, and if the average prediction error of the learning control model B is lower than a preset threshold within the time window, the transition from the learning phase to the trial phase will be triggered. The condition for switching from the trial phase (step 3) to the replacement phase (step 5) is that the learning control model B maintains control at no less than 90% of the time points throughout the entire fermentation process of the current batch.

2. The artificial intelligence method for adaptive dynamic precision fermentation control according to claim 1, wherein, The control model described in step 1) has the following characteristics: 1) Adopt a control model that does not depend on the value function; 2) A control model that does not assume that the initial state of each batch is the same is adopted; 3) Employ a neural network control model that can be trained without requiring standardized training data; 4) A two-level configuration method is used to set the control trajectory, specifically including: ① Configure control strategy: Define the set of all controllable parameters at any given time point; ② Configure and select the trajectory: Establish the correspondence between all time points in the fermentation process and the control strategy; 5) When configuring the control strategy, all or some of the controllable parameters are set as unknown parameters, which are estimated by a neural network model.

3. Artificial intelligence method for adaptive dynamic precision fermentation control according to claim 1 or 2, characterized in that, The conditions for switching from the trial phase (step 3) to the rollback phase (step 4) are as follows: During the period when the learning control model B is in control, the learning control model B is continuously used to predict the state changes of the fermentation system caused by its own control strategy, and the difference between the prediction results and the actual observed state changes is verified; a fixed-length sliding monitoring window is used, and if the average prediction error of the learning control model B within the time window is higher than a preset threshold, the rollback from the trial phase to the learning phase will be triggered.

4. The artificial intelligence method for adaptive dynamic precision fermentation control according to claim 1 or 2 or 3, characterized in that, In step 5), the condition for learning control model B to replace the baseline control model A and become the baseline control model A for the next cycle is: when learning control model B can trigger the replacement phase in five consecutive fermentation batch operation cycles.

5. The artificial intelligence method for adaptive dynamic precision fermentation control according to claim 4, characterized in that, The control model described in step 1) comprises four cooperating sub-modules: a feature encoding module F, a forward generation module G, a backward generation module G', and a decoding module D. These modules work together to achieve the functions of encoding, predicting, backtracking, and decoding time-series features. Specifically, module F employs a single-layer neural network structure, with an independent F module at each time point. Its input is a constant value of 1, and its output is the value at the corresponding time point. 3D feature vectors; to improve training efficiency, the F module at all time points is integrated into a parallel processing module, whose input layer is set to... Each neuron node corresponds to a total number of observation time points, and the output layer is configured accordingly. Each neuron node corresponds to a feature vector dimension; both the G module and the G' module employ a multi-layer neural network architecture, where the G module uses the current time point as the feature vector dimension. 3D eigenvectors and The control strategy is used as input to predict the next time point. 3D eigenvectors; G The module then achieves state backtracking by constructing a reverse time correlation, using the current time point... 3D feature vector and the previous time point The control strategy is used as input to reconstruct the previous time point. 3D feature vectors; the input layers of both the G module and the G' module are configured with Each neuron node is configured in the output layer. One neuron node; the D module uses a multi-layer neural network, responsible for... 3D feature vectors are mapped to The system state values ​​in the dimensional observation space, and its input layer settings Each neuron node, output layer settings Each neuron node corresponds to a number of system state indicators. Each module adopts a fully connected structure, and the network weight parameters are initialized using a Xavier normal distribution.

6. The artificial intelligence method for adaptive dynamic precision fermentation control according to claim 5, characterized in that, The process of synchronously recording and standardizing observation data described in step 2) is as follows: Data is collected synchronously through online monitoring and offline detection equipment. Fermentation process data for each batch, each batch containing Number of observation time points, total number of time points ; The raw data was organized into two structured datasets: a system state dataset and a control policy dataset; the system state dataset is as follows: 3D matrix, containing The dataset contains several observational indicators characterizing the state of the fermentation system, and allows for missing values ​​due to detection errors or sampling intervals; the control strategy dataset is... 3D matrix, containing One adjustable control signal; complete data recording at each time point is composed of a set of... A dimensional system state vector and a set of The control strategy vectors are jointly constructed; the data normalization process adopts Min-Max standardization or Z-Score standardization according to the parameter characteristics, and normalizes each feature dimension separately to ensure that parameters of different scales are comparable. The final normalized dataset is used to construct training data samples.

7. The artificial intelligence method for adaptive dynamic precision fermentation control according to claim 5 or 6, characterized in that, Step 6) describes using the collected time-series trajectory data of the coordinated changes in the fermentation system state and fermentation control strategy to complete the time-series training, which includes the following steps: 1) Generate a pair of positive and negative training samples for any two time points in each batch, for a total of s pairs. To learn the dynamic relationship between state and control across different time spans; 2) Construction diagonal matrix , No. OK Column 1 is set to 1, and the rest of the columns are set to 0. , indicating the first The first batch At the 1st time point, for the 1st Batch interval Two time points within a time unit and , Generation: ① Forward training samples: The time label vector is a matrix The OK , recorded as The counting step size is The observation vector is the first one in the system state dataset. OK , recorded as Control strategy matrix , dimension Record from arrive time point Group continuous control strategy; ② Reverse training samples: time label vector is a matrix The Line, denoted as The counting step size is The observation vector is the system state dataset. Line, denoted as Control strategy matrix , dimension Record from arrive time point Group reverse control strategy; 3) Positive training: Input module F to get An initial eigenvector of dimension, if If the value is 0, directly input it into the D module to obtain the predicted value, denoted as . ;like When the value is not 0, the feature vector is compared with... The first line is integrated into A dimensional vector is input into the G module to generate the next time point. 3D feature vector, at the same time Decrease by 1; repeat this process until... Zero out; input the final feature vector into the D module to obtain the predicted value. ; Calculate the loss value The formula is ; 4) Reverse training: Input F module to obtain Initial eigenvectors; if If the value is 0, directly input it into the D module to obtain the predicted value, denoted as . ;like When the value is not 0, the feature vector is compared with... The first line is integrated into A dimensional vector is input into the G' module to generate the previous time point. 3D feature vector, at the same time Decrease by 1; repeat this process until... Zero out; input the final feature vector into the D module to obtain the predicted value. ; Calculate the loss value The formula is ; 5) Model optimization: Total loss value After aggregating the loss of all samples, the model gradient is calculated through backpropagation. The Adam optimizer is used to update the network weight parameters of modules F, G, G' and D. When there are missing observations, the corresponding dimensional errors are masked and do not participate in the gradient calculation.

8. The artificial intelligence method for adaptive dynamic precision fermentation control according to claim 7, characterized in that, The process described in step 6) of predicting the fermentation system state at the next time point based on the historical trajectory of the fermentation system state and fermentation control strategy of the current fermentation batch, as well as the current fermentation control strategy, includes the following steps: 1) Based on the current fermentation batch Each time point System status data and The control strategy data is used to construct an independent F module for the current time point, with 1 input layer node and 1 output layer node. Input the value 1 into module F to generate The 3D feature vector is obtained through the D module. The predicted state value is then reconstructed sequentially using the G' module. The feature vectors at each time point are transformed into predicted state values ​​at historical time points through the D module; the mean square error between the predicted values ​​and the actual observed values ​​at the current time point and all historical time points is calculated; the network weight parameters of the F module are updated by fixing the G, G' and D modules and using the gradient descent method. 2) The optimized F module is used to obtain the feature vector at the current time point. The current feature vector is integrated with the current control strategy and input into the G module to obtain the feature vector at the next time point. The final predicted value is output through the D module to achieve single-step prediction. By connecting the G module in series, the feature vector sequence of subsequent time points is generated in sequence, and the final predicted value is output through the D module to achieve multi-step prediction. Step 6) describes calculating the optimal control strategy to transition from the current state to the target state, using the expected system state at the target time point as the control objective. This includes the following steps: 1) Construct the control configuration matrix: Generate One-hot encoding matrix , row corresponding Each time point to be optimized is represented by a column that represents a control strategy group. Control strategies are bound together by setting the same non-zero column index to achieve constant control within the time interval. 2) Create the control generation module: Construct a control generation module C with a single-layer fully connected neural network structure, where the number of neurons in the input layer is [number missing]. The number of neurons in the output layer is The control configuration matrix is ​​initialized using a Xavier normal distribution. Input the C module to obtain the dimension. Initial control strategy matrix; 3) Goal-oriented strategy optimization: First, obtain the feature vector at the current time point, integrate it with the first row of the control strategy matrix output by module C, and input it into module G to generate the feature vector at the next time point; iteratively execute... This process obtains the feature vector at the target time point, which is then converted into a predicted system state value by module D and compared with the preset target state value for mean square error loss calculation. During optimization, the network weight parameters of all other modules are fixed, and only the network weight parameters of module C are optimized using the gradient descent algorithm. Once the loss function converges, the matrix... The input is the optimized C module, i.e., the output contains... The optimal control strategy matrix at each time point.

Citation Information

Patent Citations

  • Anaerobic fermentation process soft measurement modeling method based on TSSA-CNN-LSTM

    CN120388639A

  • Method, computer system, and program for predicting characteristics of target compound

    US20210319853A1