Virus vector production control method based on virtual cells

By using an intelligent production control method based on virtual cell models, the production process parameters of viral vectors were optimized, solving the problems of titer decline and quality detection lag in the large-scale production of viral vectors. This enabled efficient and stable production of viral vectors, promoting the large-scale application of gene therapy technology.

CN122024813APending Publication Date: 2026-05-12SHENTUO BIOTECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENTUO BIOTECHNOLOGY (HANGZHOU) CO LTD
Filing Date
2026-02-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Large-scale production of viral vectors faces challenges such as reduced titer, increased impurities, and decreased recovery rates due to process scale-up effects. Furthermore, lagging quality testing leads to high production costs and poor batch stability, hindering the large-scale application of gene therapy technology.

Method used

Based on virtual cell models, combined with transfection or infection kinetics, cell population and bioreactor delivery models, a digital twin of the upstream process is established. Process parameters are optimized through optimization algorithms, and real-time prediction of quality attributes is achieved by combining deep neural networks, thus constructing an intelligent production control system.

Benefits of technology

It improved virus titer and transfection efficiency, ensured virus recovery rate and product purity, enabled forward quality control, shortened process development cycle, reduced production costs, and improved product quality stability and predictability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024813A_ABST
    Figure CN122024813A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medicine production intelligent control, and discloses a virtual cell-based virus vector production control method, which comprises the following steps: constructing a virtual cell (HEK293) model integrated with multiple sub-models based on bioreactor real-time data, laboratory material quality data and equipment static data; coupling transfection or infection kinetics, a cell population and a bioreactor transfer model, and establishing an upstream process digital twin body; taking a process optimization target as guidance, and obtaining upstream culture and transfection parameters through an optimization algorithm; and finally, generating a control instruction or parameter setting of actual virus vector production based on the optimization parameters. According to the method, digital and intelligent regulation and control of the production process are realized, the product quality stability and the production efficiency are improved, the cost is reduced, and the risk is amplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control technology for drug production, specifically a method for controlling the production of viral vectors based on virtual cells. Background Technology

[0002] With the rapid development of gene therapy technology, lentiviruses and adeno-associated viruses have become core tools for gene delivery, and their market demand continues to surge with the progress of clinical translation and commercialization. However, the large-scale production of viral vectors still faces many technical bottlenecks that urgently need to be solved, which seriously restricts the accessibility of gene therapy drugs.

[0003] Traditional viral vector production process development relies heavily on trial and error at the laboratory scale. When the process is scaled up from the laboratory to pilot-scale and commercial production, problems such as decreased titer, increased impurities, and reduced recovery rate often occur, leading to the failure of the scale-up effect. As the core of production, HEK293 cells are easily affected by multiple factors such as substrate concentration, plasmid quality, and culture conditions, resulting in significant fluctuations between production batches.

[0004] Meanwhile, quality testing methods often lag behind the production process, making it difficult to achieve real-time early warning and precise control during the process, and thus unable to avoid production risks in a timely manner. In addition, the high cost of key raw materials such as plasmid DNA, specialized culture media, and chromatography purification consumables further increases the overall production cost of viral vectors.

[0005] The aforementioned problems, such as difficulty in scaling up processes, poor batch stability, lagging quality control, and high production costs, have become key obstacles to the large-scale application of gene therapy technology. There is an urgent need to develop an intelligent and digital production control scheme to achieve accurate prediction and efficient regulation of the entire viral vector production process. Summary of the Invention

[0006] The purpose of this application is to provide a method for controlling the production of viral vectors based on virtual cells, so as to solve the problems mentioned in the background art.

[0007] According to a first aspect of this application, a method for controlling the production of viral vectors based on virtual cells is provided, comprising the following steps: S1. A virtual cell model is constructed based on historical physiological parameter data. The virtual cell model integrates a genome-scale metabolic network model, a virus assembly kinetics sub-model, and a cell stress response sub-model. The historical physiological parameter data comes from real-time process time-series data of the bioreactor, material and quality attribute data from the laboratory information management system, and static configuration data of the production equipment. S2. Based on the virtual cell model, couple the transfection or infection kinetics model, cell population model, and bioreactor delivery model to establish an upstream process digital twin; S3. Based on the upstream process digital twin, guided by the preset process optimization target, an optimization algorithm is executed to obtain optimized process parameters, which include upstream culture parameters and transfection parameters; S4. Based on the optimized process parameters, generate instructions or parameter settings for controlling the actual virus vector production process.

[0008] Preferably, the construction of the virtual cell model based on historical physiological parameter data specifically involves loading a pre-set genome-scale metabolic network model; Based on the historical cell physiological parameter data, the metabolic exchange reaction flux constraints of the preset genome-scale metabolic network model are calibrated to obtain the calibrated metabolic network model. The calibrated metabolic network model is logically connected with the pre-set virus assembly kinetics sub-model and cell stress response sub-model to form the virtual cell model.

[0009] Preferably, establishing the upstream process digital twin specifically involves instantiating the virtual cell model into multiple instances to form the cell population model; The cell population model is coupled with the transfection or infection kinetic model and the bioreactor delivery model. The environmental parameters calculated by the bioreactor delivery model are applied to the cell population model. The metabolic and production states summarized by the cell population model are used to update the material field of the bioreactor delivery model. In the actual production process of the viral vector, the internal state parameters of the upstream process digital twin are dynamically adjusted based on real-time process time-series data using a data assimilation algorithm.

[0010] Preferably, the data assimilation algorithm is an extended Kalman filter algorithm, the state vector of the upstream process digital twin includes cell density, nutrient concentration and metabolic byproduct concentration, and the real-time process time series data is used as an observation vector to update the estimated value of the state vector.

[0011] Preferably, the execution optimization algorithm is specifically a Bayesian optimization algorithm, which includes: Based on historical experimental datasets, a Gaussian process regression model is trained as a surrogate model for the optimization objective; Within the parameter search space defined by process constraints, the next combination of process parameters to be evaluated is selected based on the surrogate model and the acquisition function. The process parameters to be evaluated are input into the upstream process digital twin for virtual production simulation to obtain the predicted process results. The combination of process parameters to be evaluated and its corresponding predicted process results are added to the historical experimental dataset, and the surrogate model is updated. The selection, simulation, and update steps are executed iteratively until the stopping condition is met, and the optimized process parameters are selected from all evaluated combinations of process parameters.

[0012] Preferably, after step S3 and before step S4, the method further includes: Based on the downstream unit operation digital model and multi-objective optimization algorithm, the downstream purification process parameters are optimized to obtain the downstream optimized process parameters; The downstream unit operation digital model includes a chromatography model based on multi-component adsorption kinetics equations and a membrane filtration model based on clogging mechanisms.

[0013] Preferably, the multi-objective optimization algorithm is a non-dominated sorting genetic algorithm, and its optimization objectives include at least two of the following: virus recovery rate, impurity residue level, and production cost.

[0014] Preferably, after step S3 and before step S4, the method further includes: Using the optimized process parameters as input, and based on the scaling criteria and the target production scale equipment parameters, a virtual scale-up production simulation is performed in the scaled upstream process digital twin. By comparing the results of virtual scale-up production simulation with those of small-scale simulation, the predicted changes in key performance indicators are identified. Based on a preset risk threshold, it is determined whether the predicted change constitutes a risk of process scale-up, and a corresponding risk assessment result is generated.

[0015] Preferably, a quality attribute prediction model is constructed, wherein the quality attribute prediction model takes production process data and / or process parameters as input and outputs the predicted values ​​of key quality attributes of the viral vector; During the actual production process, real-time or phased production process data is input into the quality attribute prediction model to obtain real-time predicted values ​​of key quality attributes. The real-time predicted values ​​of key quality attributes are compared with preset product release standards to generate real-time release test conclusions.

[0016] Preferably, the viral vector production cell is a HEK293 cell, or other mammalian cells or their derived stable production cell lines used for viral vector production.

[0017] Preferably, the other mammalian cells include at least one of the following: HEK293T cells, CHO cells, PER.C6 cells, Vero cells, BHK cells, CAP cells, or MDCK cells.

[0018] In a second aspect, this application also provides a virus vector production control system based on virtual cells, comprising: The virtual cell model construction module is used to construct a virtual cell model based on historical physiological parameter data. The virtual cell model integrates a genome-scale metabolic network model, a virus assembly kinetics sub-model, and a cell stress response sub-model. The historical physiological parameter data comes from real-time process time-series data of the bioreactor, material and quality attribute data from the laboratory information management system, and static configuration data of the production equipment. The upstream twin construction module is used to establish an upstream process digital twin based on the virtual cell model, coupled with a transfection or infection kinetics model, a cell population model, and a bioreactor delivery model. The process optimization module is used to execute an optimization algorithm based on the upstream process digital twin, guided by a preset process optimization target, to obtain optimized process parameters, including upstream culture parameters and transfection parameters. The production control instruction generation module is used to generate instructions or parameter settings for controlling the actual virus vector production process based on the optimized process parameters.

[0019] This application significantly improves virus titer and transfection efficiency through intelligent optimization of upstream process parameters; ensures virus recovery rate and product purity through digital simulation and parameter optimization of downstream processes; achieves forward quality control and real-time release testing through real-time prediction of quality attributes; and improves the first-time success rate of process scale-up by identifying and mitigating scale-up risks in advance through virtual simulation of process scale-up. It transforms viral vector production from "experience-driven" to "data-driven," significantly shortening the process development cycle, reducing production and trial-and-error costs, improving the stability and predictability of product quality, and promoting the development of viral vector production towards intelligence, standardization, and large-scale production. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A schematic diagram of a virus vector production control method based on virtual cells provided in this application embodiment; Figure 2 This is a schematic diagram of the process parameter optimization process provided in the embodiments of this application; Figure 3 This is a schematic diagram of the real-time prediction process for quality attributes provided in an embodiment of this application; Figure 4 This is a schematic diagram of a virus vector production control system based on virtual cells, provided as an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] It should be noted that all user information (including but not limited to user device information, user personal information, object information corresponding to device usage data, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, device usage data, etc.) involved in all embodiments of this application are information and data authorized by the user or fully authorized by all parties.

[0024] This method is applicable to scenarios involving the development, parameter optimization, scale-up, and intelligent management of the entire production process for viral vectors such as lentiviruses and adeno-associated viruses for gene therapy, using mammalian cells (HEK293 cells as an example) as the production substrate. The method is deployed and run on servers or cloud computing platforms capable of multi-source data acquisition and processing, high-performance parallel computing, and model storage and retrieval. The virtual cell model refers to a composite digital model that uses cells as the modeling object and integrates mechanistic and data-driven models; its internal states include at least metabolic flux, viral assembly rate, and stress response parameters. Viral vector production refers to a process and quality control system centered on production process data, quality attributes, and scale-up control. Implementing this method typically requires support from both physical and software levels. Physically, it necessitates establishing standardized data interfaces for the bioreactor process control system, online sensors, and laboratory information management system to achieve interconnectivity between real-time process parameters and offline quality testing data. Software-wise, it requires a pre-built library of basic mathematical models and a process database. The model library covers HEK293 cell genome-scale metabolic network models, chromatographic adsorption kinetic models, and bioreactor fluid dynamics calculation correlations, while the database stores structured data such as process parameters, material properties, and quality testing standards. Based on these conditions, a runnable virtual simulation environment can be constructed to conduct subsequent intelligent analysis and production decision-making.

[0025] The following detailed description, using HEK293 cells as an example, illustrates the implementation process of the virtual cell-based viral vector production control method described in this application. It is understood that the method of this application can also be applied to other mammalian cells. It should be noted that this embodiment is only for explaining this application and not for limiting the scope of protection of this application. Conventional adjustments or substitutions made by those skilled in the art to each step without departing from the concept of this application should be included within the scope of protection of this application.

[0026] like Figure 1 As shown in the figure, this application discloses a schematic diagram of a viral vector production control method based on virtual cells, including the following method steps: S1. A virtual cell (HEK293) model is constructed based on historical physiological parameter data. The virtual cell (HEK293) model integrates a genome-scale metabolic network model, a virus assembly kinetics sub-model, and a cell stress response sub-model. The historical physiological parameter data is derived from real-time process time-series data of the bioreactor, material and quality attribute data from the laboratory information management system, and static configuration data of the production equipment. S2. Based on the virtual cell (HEK293) model, couple the transfection or infection kinetics model, cell population model, and bioreactor delivery model to establish an upstream process digital twin; S3. Based on the upstream process digital twin, guided by the preset process optimization target, an optimization algorithm is executed to obtain optimized process parameters, which include upstream culture parameters and transfection parameters; S4. Based on the optimized process parameters, generate instructions or parameter settings for controlling the actual virus vector production process.

[0027] In some embodiments, the data processing flow of this method is triggered when a viral vector production batch is started or a process development experiment is conducted. The data acquisition engine of the computing system establishes a stable connection with the distributed control system or programmable logic controller of the bioreactor, and collects key process parameters such as pH, dissolved oxygen concentration, temperature, agitator speed, and viable cell density within the reactor at a fixed frequency to form a time-series process dataset. Simultaneously, the data acquisition engine accesses the laboratory information management system through an application programming interface to retrieve raw material attribute data such as plasmid DNA concentration, purity, and supercoiling ratio, as well as quality detection data such as viral titer, host cell protein residue, and host cell DNA residue, forming a quality detection dataset; and reads static parameters such as bioreactor geometry, culture medium formulation, and HEK293 cell subtype information from the system configuration database.

[0028] After obtaining the original data from multiple sources, the system smooths and handles outliers in the process dataset. A sliding window averaging method is used to smooth the parameter sequence, and outliers are identified by calculating the difference between adjacent data points. Linear interpolation is then used to replace these outliers. Subsequently, the quality inspection dataset is correlated to a unified time base according to the sampling time. After cleaning, all data is structured and stored in a relational or time-series database according to a predefined data model with the process batch as the root entity, generating batch data ready records to provide high-quality data support for subsequent model building and simulation.

[0029] In some embodiments, the digital twin is a simulation model that precisely digitally maps the physiological state of HEK293 cells and the virus production process. Integrating mechanistic modeling and artificial intelligence technologies, it can achieve full-dimensional dynamic simulation of cell metabolism, virus assembly, and stress response. Based on the aforementioned structured storage of historical successful batch data, the modeling engine constructs a personalized virtual cell (HEK293) model adapted to the current process scenario. First, it loads the pre-built genome-scale metabolic network model iHEK293, which is stored in systems biology markup language format and contains 2356 metabolic reactions and 1805 metabolites, defining a metabolite stoichiometric matrix. ,in For the amount of metabolites, The number of metabolic reactions is defined, and upper and lower flux limits are preset for each reaction. , This is the lower bound of the reaction flux. This is the upper bound of the reaction flux.

[0030] To improve the model's prediction accuracy, the iHEK293 model was individually calibrated based on historical data. The calibration process was based on the principle of flux balance analysis, extracting the cell specific growth rate from the historical data. glucose consumption rate The optimal baseline flux distribution is obtained by solving linear programming problems using time-averaged values ​​of key physiological parameters to maximize cell growth rate. The default flux constraints for metabolic exchange reactions such as glucose input and lactate output are adjusted using the least squares method. This makes the adjusted constraint center value approximate the flux value of the corresponding exchange reaction in the optimal flux distribution. For example, the lower bound after calibration of glucose exchange reaction. ,in To pre-set slack variables, The theoretical minimum flux of this reaction was used to obtain the calibrated metabolic network model iHEK293_calibrated.

[0031] The iHEK293_calibrated model is logically integrated with the viral assembly kinetics submodule and the cellular stress response submodule. The viral assembly kinetics submodule, based on the law of mass action and stochastic processes, uses the flux of metabolic precursors such as nucleotides and amino acids output by iHEK293_calibrated as the raw material supply rate for viral capsid synthesis and genome replication, simulating the generation kinetics of functional and empty-shell viral particles. The cellular stress response submodule takes the reactive oxygen species (ROS) concentration and viral protein overexpression rate simulated by the model as inputs; when the ROS concentration exceeds a threshold... At that time, it outputs the ratio of cell growth to virus production inhibition factor. When the viral protein synthesis rate exceeds the endoplasmic reticulum's processing capacity, it triggers the unfolded protein response and enhances apoptosis signals. All of these stress response signals are fed back to the iHEK293_calibrated model, dynamically adjusting the upper limit of the relevant metabolic response rate, forming a cell physiological state simulation system, and completing the construction of the personalized virtual cell (HEK293) model.

[0032] In some embodiments, for step S2, based on the aforementioned personalized virtual cell (HEK293) model, a digital twin of the upstream multi-scale process covering molecular, cellular, and reactor scales is constructed to achieve full-process simulation of the upstream production process of the viral vector. At the molecular scale, a transfection kinetic model is instantiated, and ordinary differential equations are used to simulate the entire process of complex formation, endocytosis, endosome escape, and plasmid nuclear transport. The model rate parameters are related to plasmid mass, N / P ratio, and cell type. At the cellular scale, the virtual cell (HEK293) model is instantiated into multiple individuals, and random differences in cell cycle stages are introduced to simulate the heterogeneity of biological populations, forming a cell population model. At the reactor scale, a fluid dynamics-mass transfer model is established, and based on the reactor geometric parameters and operating parameters, a correlation equation is used to simulate the process. Calculate the volumetric oxygen mass transfer coefficient ,in This is the power input for aeration and stirring. To represent apparent air velocity, To determine the fitting constants related to the reactor and medium properties, and simultaneously estimate the mean shear stress. With mixing time .

[0033] The three-scale models are bidirectionally coupled through data flow; the reactor model calculations... Environmental parameters such as nutrient concentration are applied to each virtual cell in the cell population model. The cell population model summarizes the metabolic consumption and virus production rate of all individuals and feeds back to update the material concentration field of the reactor model. The number of successfully transported nuclear plasmids output by the transfection kinetics model is used as the trigger input for the virus assembly submodule.

[0034] In one embodiment, to achieve synchronous evolution between the digital twin and the physical production process, the system employs an extended Kalman filter algorithm for online data assimilation, defining the state vector of the digital twin. ,in For cell density, This refers to the glucose concentration. This refers to the lactic acid concentration. This refers to the concentration of ammonium ions. The concentration of virus particles is given by the equation of state. , For control inputs such as feeding rate, Process noise; real-time sensor data as the observation vector. The observation equation is , To address observation noise, the Extended Kalman Filter algorithm advances the state based on the state equation in the prediction step, and utilizes the Kalman gain in the update step. Corrected state estimate By dynamically adjusting hidden states such as metabolic flux within the model, the predicted trajectory of the digital twin closely tracks the physical production process data, enabling it to reflect and predict the production status in real time.

[0035] In some embodiments, for step S3, this step addresses the technical problems of traditional viral vector upstream process parameter optimization relying on laboratory-based trial and error, high experimental costs, long development cycles, and susceptibility to local optima. By combining a Gaussian process regression artificial intelligence model with a Bayesian optimization algorithm, and using a high-fidelity upstream process digital twin as a simulation carrier, global and efficient optimization of process parameters is achieved. The principle is to use a Gaussian process regression model to fit and predict the nonlinear relationship between process parameters and production effects, combined with the exploration and utilization of Bayesian optimization's acquisition function to balance parameter search, and to replace physical experiments with virtual experiments, locking in the optimal combination of process parameters within the fewest evaluation times. The technical effect is reflected in a process development cycle shortened by more than 60%, material consumption reduced by more than 50%, and a significant improvement in the efficiency and accuracy of parameter optimization.

[0036] like Figure 2 As shown, Figure 2 This is a schematic diagram of the process parameter optimization process provided in an embodiment of this application. In S201, a Gaussian process regression model is trained as a surrogate model for the optimization objective based on historical experimental datasets. When upstream process parameter optimization is performed, the system's upstream intelligent surrogate module extracts relevant experimental datasets from the historical database. ,in This is a vector combining process parameters such as DNA dosage, N / P ratio, culture temperature, and feed rate. This represents the vector of experimental results such as viral titer, cell viability, and transfection efficiency. The number of data samples is specified. For each optimization objective, the system trains a Gaussian process regression model as a surrogate model. The Gaussian process regression model is a probability-based nonlinear regression artificial intelligence model that can capture the complex nonlinear characteristics of the data through kernel functions, while outputting the uncertainty of the predicted values, providing a basis for subsequent parameter selection.

[0037] Among them, the Gaussian process regression model assumes an optimization objective It follows a Gaussian process distribution, and its expression is: ,in The mean function is used to describe the overall trend of the optimization objective. In this embodiment, it is set as a constant function to simplify model calculation while ensuring the fitting effect. The covariance function, also known as the kernel function, is used to measure the combination of two process parameters. and To assess the similarity between them, this embodiment uses a radial basis function kernel, the expression of which is: ,in The signal variance reflects the model's ability to fit the overall fluctuations of the data. Using length as the scale, the sensitivity of the model to changes in process parameters is controlled. The Euclidean distance is the combination of process parameters. The noise variance is the random noise used to fit the experimental data. Let Kronecker function be used when hour ,when hour .

[0038] The input layer of the model is a vector of process parameter combinations. It covers all adjustable parameters of upstream culture and transfection. These parameters are standardized before being input into the model to eliminate the influence of dimensional differences on the fitting effect. The core computational layer of the model is based on kernel function-based covariance matrix calculation. For the input... One sample, construct order covariance matrix The first in the matrix Line number The elements of the column are The augmented covariance matrix is ​​obtained by combining the noise variance. , The identity matrix is ​​used; the output layer of the model is the predicted mean of the optimization objective. With prediction variance For any new combination of process parameters The formulas for calculating the predicted mean and variance are as follows: , ,in The covariance vector of the new parameter combination and the historical samples. The target value is the optimization value for historical samples.

[0039] The core of model training is hyperparameters. The optimization is achieved by maximizing the marginal likelihood function to find the optimal hyperparameters. The expression for the marginal likelihood function is: ,in To the determinant of the augmented covariance matrix, The sample size is given. The system uses gradient descent to maximize the marginal likelihood function. First, it calculates the partial derivatives with respect to the hyperparameters to obtain the gradient of the marginal likelihood function with respect to each hyperparameter. Then, it updates the hyperparameters along the gradient ascent direction, setting the learning rate and the number of iterations until the marginal likelihood function converges or reaches the preset number of iterations, thus obtaining the optimal hyperparameters. By substituting the optimal hyperparameters into the model, the Gaussian process regression surrogate model is trained. The trained model can accurately predict the optimization objective for any combination of process parameters and provide the uncertainty range of the prediction results.

[0040] In S202, within the parameter search space defined by process constraints, the next combination of process parameters to be evaluated is selected based on the surrogate model and the acquisition function. After completing the training of the Gaussian process regression surrogate model, the system starts a Bayesian optimization loop, guided by a preset process optimization objective, to select the combination of process parameters to be evaluated within the parameter search space defined by process constraints.

[0041] Define the currently evaluated dataset The initial dataset was a historical experimental dataset. In each iteration, the system selects the next parameter point to be evaluated through a data acquisition function. In this embodiment, the desired improvement function is used as the acquisition function, balancing the exploratory and exploitative aspects of parameter search. The expression for the desired improvement function is: ,in for The optimal optimization objective value is... This is a preset small positive number used to balance exploration and utilization. , The cumulative distribution function of the standard normal distribution. It is the probability density function of the standard normal distribution.

[0042] In S203, the combination of process parameters to be evaluated is input into the upstream process digital twin for virtual production simulation to obtain the predicted process results. The system solves for the maximum value of the desired improvement function in the parameter search space using numerical calculation methods, and uses the corresponding parameter points as... The data is then input into an upstream process digital twin that has undergone data assimilation to perform a complete virtual production simulation from cell inoculation to virus harvesting. After the simulation, the predicted process results are output. This includes optimization targets related to indicators such as viral titer and transfection efficiency.

[0043] In S204, the combination of process parameters to be evaluated and its corresponding predicted process results are added to the historical experimental dataset, and the surrogate model is updated. Add to The Gaussian process regression surrogate model was retrained using the updated dataset, and the model's covariance matrix and hyperparameters were updated to improve the model's fitting accuracy to the parameter search space.

[0044] In S205, the selection, simulation, and update steps are executed iteratively until a stopping condition is met. The optimized process parameters are selected from all evaluated combinations of process parameters. The iterative process of selecting acquisition function points, digital twin virtual simulation, and model updating is repeated until a preset stopping condition is met. The stopping condition can be set to 50 iterations or the improvement value of the optimization target being less than 1 for 10 consecutive rounds. Or, the maximum value of the expected improvement function is less than a threshold. After the iteration terminates, the system selects the optimal or second-best combination from all evaluated process parameter combinations based on the optimization objective. If it is a multi-objective optimization, the system selects the parameter combination that meets the production requirements through Pareto front analysis as the optimized upstream process parameters, including upstream culture parameters and transfection parameters, to provide a precise basis for parameter setting in actual production.

[0045] In some embodiments, for the downstream purification process of viral vectors, a machine learning-based digital model of unit operations is constructed, covering chromatography, membrane filtration, and viral stability prediction models, to achieve accurate simulation and parameter optimization of the downstream purification process. For the chromatography unit, a multi-component adsorption kinetic model based on the generalized Langmuir isotherm is used to describe the competitive adsorption behavior of viral particles, host cell proteins, and host cell DNA. The adsorption rate equation is as follows: ,in Components The amount of adsorption, The concentration of the mobile phase. , The adsorption and desorption rate constants are, The maximum adsorption capacity was determined. The chromatography model parameters were calibrated using a machine learning algorithm. Using breakthrough and elution curves from historical chromatography experiments as the training set, a random forest regression model was employed to fit the relationship between the model parameters and experimental results. The model hyperparameters were then optimized using a grid search method to obtain the calibrated set of packing material-component characteristic parameters. This improves the model's prediction accuracy for actual tomography processes.

[0046] For the membrane filtration unit, a permeate mass transfer model based on the clogging mechanism is established, and the membrane flux decay law is as follows: ,in This is the initial membrane flux. To determine the clogging constant, the system employs a linear regression machine learning algorithm to fit the membrane filtration model parameters based on historical tangential flow filtration experimental data. , The parameters related to virus retention rate are used; the virus stability prediction model takes process stresses such as shear force and gas-liquid interface exposure as inputs, and uses machine learning algorithms to fit the relationship between stress and virus activity, predicting the stability changes of virus particles during the process, and providing a basis for optimizing downstream process parameters. The above unit operation models are logically connected in sequence according to the purification process, with the elution product information of the previous unit as the input of the next unit, constructing a complete digital twin of the downstream purification process flow, realizing digital simulation of the entire downstream process.

[0047] In one embodiment, when downstream purification processes require optimization, the system activates the downstream intelligent agent module. It employs a non-dominated sorting genetic algorithm with an elitist strategy (NSGA-II) to perform multi-objective optimization of downstream process parameters. Optimization objectives include virus recovery rate, impurity residue level, production cost, and processing time. Decision variables include adjustable parameters such as chromatographic loading capacity, elution flow rate, buffer pH, and membrane filtration transmembrane pressure. The system randomly generates a scale of... The initial population is established, with each individual representing a set of downstream process parameters. Individual parameters are input into a digital twin of the downstream process flow for virtual simulation, yielding the corresponding objective function value. The population undergoes non-dominated sorting and crowding calculation. Based on Pareto dominance, individuals are divided into different frontiers, and the crowding distance within the same frontier is calculated to measure their distribution density in the target space. A binary tournament selection method is used to select parent individuals. Offspring are generated through simulated binary crossover and polynomial mutation. The parent and offspring populations are merged, with elites retained to form a new generation. This process is iteratively executed until a preset number of generations is reached. The individuals at the first frontier constitute the Pareto optimal solution set, from which the system selects parameter combinations that meet production requirements as downstream optimized process parameters.

[0048] In one embodiment, when scaling up an optimized process from a laboratory or pilot-scale operation to commercial production, the system's cost and risk proxy module performs a quantitative prediction of the process scale-up effect. This involves obtaining a digital twin of the small-scale process and the design parameters of the target large-scale production equipment, based on constant... , or mixed time The first-principles scaling criteria are used to scale the process parameters. For example, if a constant is chosen... The stirring speed is obtained by inversely solving the criterion using the geometric parameters of a large-scale reactor. With ventilation rate Based on the scaled parameters and the parameters of large-scale equipment, a scaled-up digital twin of the upstream process is constructed, the feeding strategy is adjusted to intelligent control based on metabolite concentration, and a complete virtual scale-up production simulation is executed.

[0049] After the simulation is completed, the key performance indicators of the scaled-up and small-scale simulations are compared, and the predicted changes are calculated. , These are the performance index values ​​for large-scale simulations. These are small-scale simulation metrics, with key performance indicators including virus titer, cell density, lactate concentration, and empty-shell virus proportion. The system makes judgments based on preset risk thresholds. To determine whether a process scale-up risk exists, for example, if the change in virus titer is less than -20%, it is marked as high risk; if the change in lactic acid concentration is greater than 30%, it is marked as medium risk. For the identified risks, the system retrieves the cause explanations and mitigation strategies from the knowledge base and generates a process scale-up prediction and risk assessment report to provide data support for actual process scale-up.

[0050] In some embodiments, the method further includes real-time prediction of quality attributes based on deep neural networks. This addresses the technical problem in traditional viral vector production where quality inspection lags behind the production process, making real-time early warning and control difficult. By constructing a deep neural network artificial intelligence model that integrates process trajectory features, real-time prediction of key quality attributes of viral vectors is achieved. The principle is to utilize the multi-layer nonlinear mapping capability of deep neural networks to uncover the complex correlation between production process data and key quality attributes. Static process parameters and dynamic process trajectory features are used as model inputs to achieve accurate real-time prediction of key quality attributes. The technical effect is to realize the transformation of quality control from "post-event inspection" to "process prediction," provide early warning of potential quality risks, provide a data foundation for real-time release testing, and shorten the product release cycle by more than 70%.

[0051] In some embodiments, the viral vector production cells are specifically HEK293 cells, or other mammalian cells or their derived stable production cell lines for viral vector production. Specifically, the other mammalian cells include at least one of the following: HEK293T cells, CHO cells, PER.C6 cells, Vero cells, BHK cells, CAP cells, or MDCK cells.

[0052] like Figure 3 As shown, Figure 3 This is a schematic diagram of the real-time prediction process for quality attributes provided in an embodiment of this application. In S301, a quality attribute prediction model is constructed, which takes production process data and / or process parameters as input and outputs the predicted values ​​of key quality attributes of the viral vector.

[0053] The deep neural network model is a feedforward network structure that adopts an end-to-end learning approach. It directly extracts features from production process data and predicts key quality attributes. The model consists of three parts: an input layer, a hidden layer, and an output layer. Each layer is fully connected, and the connection weights between neurons are obtained through training. There are no connections between neurons within a layer, which enables the layer-by-layer extraction and nonlinear transformation of features.

[0054] The input layer receives high-dimensional feature vectors. The model can have hundreds of dimensions, and its feature vectors contain four types of features, all of which are standardized before being input into the model. Specifically, these include: static / formulation features, such as culture medium type, cell passage, and plasmid batch number; categorical features are converted into numerical features through one-time thermal encoding or embedding layer processing; key operational parameters, such as upstream core process parameters like transfection N / P ratio, DNA usage, transfection time, and culture temperature; process trajectory features, which are the core input features of the model, extracted from real-time process data of upstream process digital twins or physical production, including time-series data of parameters such as pH, dissolved oxygen concentration, cell density, and lactate concentration during the culture cycle, and extracted through feature engineering to include statistical features such as mean, variance, and slope; time-domain features such as the time to reach maximum cell density and the time point of rapid lactate accumulation; and frequency-domain features such as the dominant frequency extracted through Fourier transform; and downstream process parameter features, including downstream core process parameters such as chromatography loading capacity, elution conductivity, and membrane filtration concentration factor.

[0055] The hidden layer comprises three fully connected layers, employing the ReLU activation function to introduce non-linearity and alleviate overfitting. Dropout and batch normalization layers are also included to further enhance the model's generalization ability. For example, the hidden layer structure is: input layer, 256-dimensional fully connected layer, ReLU activation function, Dropout, 128-dimensional fully connected layer, ReLU activation function, 64-dimensional fully connected layer, ReLU activation function. Dropout represents randomly discarding 30% of the neurons to avoid excessive dependence between neurons. Each neuron in the hidden layer receives the outputs of all neurons in the previous layer, which are then weighted, summed, and transformed by the activation function before being passed to the next layer's neurons. This process achieves layer-by-layer abstraction and extraction of input features, uncovering the non-linear patterns behind the data.

[0056] The output layer, designed according to the quality prediction task, employs a multi-task learning approach to simultaneously predict multiple key quality attributes, including viral titer, purity, host cell protein residue, and host cell DNA residue. The number of neurons in the output layer matches the number of key quality attributes, and a linear activation function is used to directly output the predicted values ​​of these attributes, meeting the needs of quantitative prediction. If product qualification / non-qualification classification is required, the output layer can use a sigmoid activation function to output predicted probability values.

[0057] In step S302, the deep neural network model is trained using historical production batch data as the training set. The model uses structured historical production batch data as the training set, and each batch of data is processed into a feature vector. With key quality attributes true labels The key quality attribute labels are the actual values ​​from offline laboratory testing. For numerical labels such as virus titer, logarithmic transformation is performed before use in model training to improve training performance. The dataset is randomly divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used for model parameter learning, the validation set is used for hyperparameter optimization and early stopping detection, and the test set is used for evaluating the model's generalization ability.

[0058] The model uses mean squared error as the loss function, and the expression for the loss function is as follows: ,in For the sample size, These are the model's predicted values. The model uses the true label values; the Adam optimizer is used to update the model parameters. The Adam optimizer combines momentum and adaptive learning rate, resulting in fast convergence and good stability. The initial learning rate is set to... The weight decays to This is to prevent the model from overfitting.

[0059] The model training employs mini-batch stochastic gradient descent with a batch size of 32 and 200 iterations. During training, the loss value is monitored in real-time on the validation set. If the validation set loss value fails to decrease for 10 consecutive iterations, an early stopping mechanism is triggered, terminating model training and preventing overfitting. After training, the model is evaluated on the test set using the coefficient of determination. The mean absolute percentage error (MAPE) is used to evaluate the model's prediction accuracy. When MAPE < 5%, the model is considered to meet the needs of practical applications, and a trained deep neural network quality attribute prediction model is obtained.

[0060] In S303, during the actual production process control, real-time or phased production process data is input into the quality attribute prediction model to obtain real-time predicted values ​​of key quality attributes. During the actual production of the viral vector, the system collects real-time production process data at a fixed frequency, extracts static features, key operating parameter features, process trajectory features, and downstream process parameter features, and combines them into a current feature vector. The input is then fed into a trained deep neural network model, which performs feedforward propagation calculations and outputs real-time predicted values ​​of key quality attributes. Meanwhile, the system uses the Monte Carlo Dropout method to estimate the uncertainty interval of the predicted value. By randomly dropping neurons multiple times and making predictions, the mean of the prediction results is used as the final predicted value, and twice the standard deviation is used as the uncertainty interval to reflect the reliability of the prediction results.

[0061] In S304, the real-time predicted value of the key quality attribute is compared with the preset product release standard to generate a real-time release test conclusion.

[0062] The system compares real-time predicted values ​​with preset product release standards and generates real-time release test conclusions. If all predicted values ​​of key quality attributes meet the release standards, the product is deemed qualified and can be released in advance. If there are cases where predicted values ​​do not meet the release standards, the system automatically triggers a quality risk warning. Combining the simulation results of the upstream process digital twin, the system analyzes the root cause of the quality risk, provides suggestions for adjusting process parameters, assists operators in timely intervention, avoids batch failures, and achieves real-time and forward-looking quality control.

[0063] In some embodiments, for step S4, the system integrates the optimized upstream process parameters, downstream process parameters, and parameter adjustment suggestions in the process scale-up risk assessment report to generate instructions and parameter settings for controlling the actual viral vector production process. These instructions and settings cover real-time control parameters such as temperature, pH, dissolved oxygen concentration, stirring speed, and feeding rate of the bioreactor; operational parameters such as DNA usage, N / P ratio, and transfection reagent addition for the transfection operation; operational parameters of the downstream purification chromatography and membrane filtration units; and abnormal handling strategies and parameter adjustment thresholds during the production process.

[0064] The production control system distributes the aforementioned instructions and parameter settings to each production device through a standardized data interface, achieving automated and precise control of the virus vector production process. Simultaneously, the system collects actual production data in real time and dynamically compares it with the simulation results of the upstream process digital twin. If deviations occur, the extended Kalman filter algorithm dynamically adjusts the digital twin parameters, and production control instructions are optimized in real time based on the deviations. This forms a production control system encompassing model prediction, parameter control, data feedback, and model correction, ensuring the stability of the production process and the consistency of product quality.

[0065] Thus, intelligent optimization of upstream process parameters significantly improved virus titer and transfection efficiency; digital simulation and parameter optimization of downstream processes ensured virus recovery rate and product purity; real-time prediction of quality attributes enabled proactive quality control and real-time release testing; and virtual simulation of process scale-up allowed for early identification and mitigation of scale-up risks, improving the first-time success rate of process scale-up. This shift from "experience-driven" to "data-driven" viral vector production significantly shortened the process development cycle, reduced production and trial-and-error costs, improved product quality stability and predictability, and propelled viral vector production towards intelligent, standardized, and large-scale development.

[0066] It should be noted that although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0067] Please see Figure 4 , Figure 4 This application provides a block diagram of a virus vector production and control system based on virtual cells. The system specifically includes: The virtual cell model construction module 401 is used to construct a virtual cell (HEK293) model based on historical physiological parameter data. The virtual cell (HEK293) model integrates a genome-scale metabolic network model, a virus assembly kinetics sub-model, and a cell stress response sub-model. The historical physiological parameter data comes from real-time process time-series data of the bioreactor, material and quality attribute data from the laboratory information management system, and static configuration data of the production equipment. The upstream twin construction module 402 is used to establish an upstream process digital twin based on the virtual cell (HEK293) model, coupled with a transfection or infection kinetics model, a cell population model, and a bioreactor delivery model. The process optimization module 403 is used to execute an optimization algorithm based on the upstream process digital twin, guided by a preset process optimization target, to obtain optimized process parameters, which include upstream culture parameters and transfection parameters. The production control instruction generation module 404 is used to generate instructions or parameter settings for controlling the actual virus vector production process based on the optimized process parameters.

[0068] It should be noted that the working process of each module in the virtual cell-based viral vector production control system described in this embodiment can refer to the working process of the virtual cell-based viral vector production control method described in the above embodiments, and the technical effect achieved is the same as that of the virtual cell-based viral vector production control method described in the above embodiments, so it will not be repeated here.

[0069] The above description represents the preferred embodiments of the present invention. It should be noted that, for those skilled in the art, various improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for controlling the production of viral vectors based on virtual cells, characterized in that, Includes the following steps: S1. Construct a virtual cell model corresponding to the virus vector production cells based on historical physiological parameter data. The virtual cell model integrates a genome-scale metabolic network model, a virus assembly kinetics sub-model, and a cell stress response sub-model. The historical physiological parameter data comes from real-time process time-series data of the bioreactor, material and quality attribute data from the laboratory information management system, and static configuration data of the production equipment. S2. Based on the virtual cell model, couple the transfection or infection kinetics model, cell population model, and bioreactor delivery model to establish an upstream process digital twin; S3. Based on the upstream process digital twin, guided by the preset process optimization target, an optimization algorithm is executed to obtain optimized process parameters, which include upstream culture parameters and transfection parameters; S4. Based on the optimized process parameters, generate instructions or parameter settings for controlling the actual virus vector production process.

2. The method for controlling the production of viral vectors based on virtual cells according to claim 1, characterized in that, The construction of the virtual cell model based on historical physiological parameter data specifically involves loading a pre-set genome-scale metabolic network model. Based on the historical cell physiological parameter data, the metabolic exchange reaction flux constraints of the preset genome-scale metabolic network model are calibrated to obtain the calibrated metabolic network model. The calibrated metabolic network model is logically connected with the pre-set virus assembly kinetics sub-model and cell stress response sub-model to form the virtual cell model.

3. The method for controlling the production of viral vectors based on virtual cells according to claim 2, characterized in that, The establishment of the upstream process digital twin specifically involves instantiating the virtual cell model into multiple instances to form the cell population model. The cell population model is coupled with the transfection or infection kinetic model and the bioreactor delivery model. The environmental parameters calculated by the bioreactor delivery model are applied to the cell population model. The metabolic and production states summarized by the cell population model are used to update the material field of the bioreactor delivery model. In the actual production process of the viral vector, the internal state parameters of the upstream process digital twin are dynamically adjusted based on real-time process time-series data using a data assimilation algorithm.

4. The method for controlling the production of viral vectors based on virtual cells according to claim 3, characterized in that, The data assimilation algorithm is specifically an extended Kalman filter algorithm. The state vector of the upstream process digital twin includes cell density, nutrient concentration and metabolic byproduct concentration. The real-time process time series data is used as an observation vector to update the estimated value of the state vector.

5. The method for controlling the production of viral vectors based on virtual cells according to claim 1, characterized in that, The execution optimization algorithm is specifically the execution of a Bayesian optimization algorithm, which includes: Based on historical experimental datasets, a Gaussian process regression model is trained as a surrogate model for the optimization objective; Within the parameter search space defined by process constraints, the next combination of process parameters to be evaluated is selected based on the surrogate model and the acquisition function. The process parameters to be evaluated are input into the upstream process digital twin for virtual production simulation to obtain the predicted process results. The combination of process parameters to be evaluated and its corresponding predicted process results are added to the historical experimental dataset, and the surrogate model is updated. The selection, simulation, and update steps are executed iteratively until the stopping condition is met, and the optimized process parameters are selected from all evaluated combinations of process parameters.

6. The method for controlling the production of viral vectors based on virtual cells according to claim 1, characterized in that, After step S3 and before step S4, the method further includes: Based on the downstream unit operation digital model and multi-objective optimization algorithm, the downstream purification process parameters are optimized to obtain the downstream optimized process parameters; The downstream unit operation digital model includes a chromatography model based on multi-component adsorption kinetics equations and a membrane filtration model based on clogging mechanisms.

7. The method for controlling the production of viral vectors based on virtual cells according to claim 6, characterized in that, The multi-objective optimization algorithm is a non-dominated sorting genetic algorithm, and its optimization objectives include at least two of the following: virus recovery rate, impurity residue level, and production cost.

8. The method according to claim 1, characterized in that, After step S3 and before step S4, the method further includes: Using the optimized process parameters as input, and based on the scaling criteria and the target production scale equipment parameters, a virtual scale-up production simulation is performed in the scaled upstream process digital twin. By comparing the results of virtual scale-up production simulation with those of small-scale simulation, the predicted changes in key performance indicators are identified. Based on a preset risk threshold, it is determined whether the predicted change constitutes a risk of process scale-up, and a corresponding risk assessment result is generated.

9. The method for controlling the production of viral vectors based on virtual cells according to claim 1, characterized in that, Also includes: A quality attribute prediction model is constructed, which takes production process data and / or process parameters as input and outputs the predicted values ​​of key quality attributes of the viral vector. During the actual production process, real-time or phased production process data is input into the quality attribute prediction model to obtain real-time predicted values ​​of key quality attributes. The real-time predicted values ​​of key quality attributes are compared with preset product release standards to generate real-time release test conclusions.

10. The method for controlling the production of viral vectors based on virtual cells according to claim 1, characterized in that, The specific cell used for producing the viral vector is HEK293 cell, or other mammalian cells or their derived stable production cell lines used for viral vector production.

11. The method for controlling the production of viral vectors based on virtual cells according to claim 10, characterized in that, The other mammalian cells include at least one of the following: HEK293T cells, CHO cells, PER.C6 cells, Vero cells, BHK cells, CAP cells, or MDCK cells.

12. A virus vector production control system based on virtual cells, characterized in that, include: The virtual cell model construction module is used to construct a virtual cell model based on historical physiological parameter data. The virtual cell model integrates a genome-scale metabolic network model, a virus assembly kinetics sub-model, and a cell stress response sub-model. The historical physiological parameter data comes from real-time process time-series data of the bioreactor, material and quality attribute data from the laboratory information management system, and static configuration data of the production equipment. The upstream twin construction module is used to establish an upstream process digital twin based on the virtual cell model, coupled with a transfection or infection kinetics model, a cell population model, and a bioreactor delivery model. The process optimization module is used to execute an optimization algorithm based on the upstream process digital twin, guided by a preset process optimization target, to obtain optimized process parameters, including upstream culture parameters and transfection parameters. The production control instruction generation module is used to generate instructions or parameter settings for controlling the actual virus vector production process based on the optimized process parameters.