Monitoring, simulation, and control of biological processes
Through machine learning, predicting the unit transport rate of cellular metabolites in bioreactors, the problem of real-time monitoring and optimization of biological processes in the prior art is solved, and efficient and accurate biomaterial production and process control are achieved.
Patent Information
- Application Number
- CN202180029655.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-20
- Filing Date
- 2021-01-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-01-14
AI Technical Summary
The prior art is difficult to monitor, simulate and control cellular metabolic conditions in biological processes in real time, making it difficult to determine the conditions for producing high-quality biological materials, and lacks direct measurement of metabolic activity, resulting in difficulty in process optimization and fault diagnosis.
Using machine learning methods, the metabolite unit transport rate of cell culture in the bioreactor is predicted by training the model, and the characteristics of biological processes are predicted based on this rate, thereby real-time monitoring and optimization of process conditions are achieved.
Real-time monitoring and prediction of biological processes is achieved, and cell metabolism conditions can be determined more accurately, the quality and efficiency of biological material production can be improved, and the process is timely diagnosed and adjusted, and resource waste can be avoided.
Smart Images

Figure CN115427898B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to computer-implemented methods, computer programs, and systems for monitoring, simulating, and controlling biological processes. Certain methods, programs, and systems of the present disclosure use machine learning to predict one or more variables representing cellular metabolic conditions in a biological process. Background Art
[0002] Bioprocesses use biological systems to produce specific biomaterials, such as biomolecules with therapeutic effects. The process typically involves placing cells and / or microorganisms in a bioreactor with a culture medium containing nutrients under controlled atmospheric conditions. The culture medium is consumed by the cells and used for growth and other metabolic functions, including the production of specific biomaterials and byproducts.
[0003] Bioreactors typically contain or are associated with instrumentation that continuously (e.g., once per second, per minute, per hour) measures process conditions (e.g., temperature, pH, and dissolved oxygen) as well as the addition of nutrients and gases and the flow rates and contents of streams leaving the bioreactor. Typically, culture samples are taken periodically (e.g., once or twice per day, or more than twice) to measure the contents of the bulk fluid, including the concentration of one or more metabolites (e.g., glucose, glutamine, lactate, NH4, etc.), the concentration of product biomaterial (also known as titer), cell metrics (e.g., total cell and viable cell density (VCD)), and quality metrics of product biomaterial (sometimes also referred to as "quality attributes" or "critical quality attributes" (CQA), e.g., glycosylation profile or activity of the product).
[0004] Statistical process analysis methods can be used to assess the good performance of bioprocesses. In particular, multivariate statistical models (including principal component analysis (PCA) and (orthogonal) partial least squares ((O)PLS)) regression have become popular tools for identifying process conditions that are important for ensuring that CQAs are within specifications (collectively referred to as “critical process parameters” (CPPs)) and determining the acceptable range of these process conditions as the bioprocess progresses to completion. Such tools have been used in The results are implemented in the Sartorius Stedim Data Analytics software suite, a leading data analytics software for modeling and optimizing biopharmaceutical development and manufacturing processes.
[0005] In a typical bioprocess analysis, a series of process variables (e.g., dozens of process variables, including temperature, concentrations of key nutrients and metabolites, pH, volume, gas concentrations, viable cell density, etc.) are measured during the completion of the bioprocess. Together, these process variables represent the "process conditions." Many of these variables are highly correlated, so methods such as PCA and PLS can be used to identify summary variables that capture the correlation structure in the data. These variables (usually relatively few) can then be extracted, and the range of values of these variables that define "normal" process conditions can be estimated. In addition, models such as (O)PLS can be used to predict product titer from these process variables.
[0006] All of these approaches simulate the effects of process parameters on the product produced by the cell, but do not provide an understanding of how process parameters affect the function of the cell and how this ultimately leads to changes in CQAs and product titers. During process development, the lack of direct measurement of metabolic activity creates a situation where process operation is designed using trial and error. Using such approaches, experiments are performed in an ad hoc manner to determine the conditions that produce high-quality and relatively high-yield products. In addition, the lack of knowledge of the cellular metabolic conditions (the process conditions that determine the production of specific CQAs and titers) means that diagnosing / predicting failures (e.g., abnormal metabolic operation) or optimizing the process requires subject matter experts (SMEs) to come up with biologically relevant hypotheses for which metabolic conditions can underlie the observed behavior.
[0007] Furthermore, all of these approaches rely on a posteriori analysis of biological processes and have limited ability to predict how biological processes will evolve. This means that these approaches are not suitable for implementing preemptive actions to correct the course of a biological process (rather than diagnosing after the fact that the course is wrong, which can waste time and resources), nor for building models that mimic the course of a biological process (e.g. for process optimization).
[0008] Therefore, there is a need for systems and methods for improved methods of monitoring, modeling, and controlling biological processes. Summary of the invention
[0009] According to a first aspect of the present disclosure, there is provided a computer-implemented method for monitoring a bioprocess, the bioprocess comprising a cell culture in a bioreactor, the method comprising the following steps: obtaining values of one or more process conditions, the one or more process conditions comprising one or more process parameters, one or more metabolite concentrations and / or one or more biomass-related metrics of the bioprocess at one or more maturities; determining a specific transport rate of one or more metabolites in the cell culture using the obtained values as input to a machine learning model, the machine learning model being trained to predict the specific transport rate of one or more metabolites at the latest maturity or later maturity among the one or more maturities based at least in part on the values of the one or more process conditions of the bioprocess at one or more maturities; and predicting one or more characteristics of the bioprocess based at least in part on the determined specific transport rates.
[0010] The method of the first aspect may have any one or any combination of the following optional features.
[0011] Predicting one or more characteristics of a biological process may include: comparing the unit transport rate or a value derived from the unit transport rate to one or more predetermined values; and determining whether the process is operating normally based on the comparison.
[0012] The one or more predetermined values may be specified to vary with maturity, and comparing the unit transport rate, or a value derived from the unit transport rate, to the one or more predetermined values may include comparing the unit transport rate, or a value derived from the unit transport rate, to one or more predetermined values associated with maturity corresponding to later maturity values.
[0013] In the context of the present disclosure, maturity can be expressed in units of time. For example, maturity can refer to the amount of time since the start of a biological process or any other reference time. Therefore, references to a specific "maturity" and "maturity function" should be interpreted as including "time point", "time" and "time function".
[0014] Obtaining values of one or more process conditions may include directly or indirectly measuring the values of the one or more process conditions. Obtaining values of one or more process conditions may include receiving the values, for example, from a user, a computing device, or a memory.
[0015] The unit transport rate of a metabolite may be a unit consumption rate or a unit production rate. The unit transport rate of a metabolite i may be the net amount of the metabolite transported between the cell and the culture medium per cell and per unit of maturity.
[0016] As will be further described below, predicting one or more characteristics of a biological process may include classifying internal metabolic conditions of the process between a plurality of classifications, wherein the internal metabolic conditions of the process include a unit transport rate or a value derived from the unit transport rate. Classifying the internal metabolic conditions of the process may include classifying the internal metabolic conditions of the process into an optimal category or a suboptimal category for biomaterial production.
[0017] As will be further described below, predicting one or more characteristics of a biological process can include predicting a state trajectory of the process using one or more models of process evolution that include a unit transport rate or a variable derived from the unit transport rate. The one or more models can include a kinetic growth model and a material balance equation. The state of the process can include the value of any variable in such a model, such as the value of one or more cell culture parameters and / or metabolite concentrations. The state trajectory of the process can include the values of state variables at multiple time points / maturity.
[0018] As will be further described below, predicting one or more characteristics of a biological process includes using one or more models of process evolution that include a unit transport rate or variables derived from the unit transport rate to predict a future state of the process after one or more process parameters are changed.
[0019] As will be further described below, predicting one or more characteristics of a biological process includes predicting current or future critical quality attributes (CDQs) associated with the process using one or more predictive models, the predictive models being trained to predict values of one or more CQAs based at least in part on a unit transport rate or a variable derived from the unit transport rate.
[0020] The one or more process conditions may include one or more process parameters selected from the group consisting of dissolved oxygen, dissolved CO 2 , pH, temperature, osmotic pressure, stirring speed, stirring power, headspace gas composition (e.g. CO 2The process conditions may include one or more biomass-related metrics selected from the group consisting of live cell density, total cell density, cell viability, dead cell density, and lysed cell density. The one or more metabolite concentrations may include the concentration of one or more metabolites in the culture medium chamber. Unless otherwise indicated by the context, the metabolite concentration generally refers to the metabolite concentration in the culture medium. The values of one or more process conditions for predicting the unit transport rate of one or more metabolites may include at least one metabolite concentration value. The values of one or more process conditions may also include at least one other value of the process condition, preferably at least two other values. The values of one or more metabolite concentrations may include the concentration of one or more metabolites for which the unit transport rate has been determined. The one or more metabolites may alternatively or additionally include the concentration of one or more metabolites, which are precursors of the one or more metabolites for which the unit transport rate has been determined. The one or more metabolites may alternatively or additionally include the concentration of one or more metabolites, which are reaction products that produce or consume the one or more metabolites for which the unit transport rate has been determined.
[0021] Predicting one or more characteristics of a bioprocess includes predicting the value of one or more critical quality attributes of the bioprocess using a predictive model that has been trained to predict CQAs using a set of predictor variables including one or more unit transport rates. The predictive model can be a machine learning model, such as a multivariate model. The predictive model can be trained using data from multiple training runs.
[0022] The method may also include determining values of the one or more latent variables using a pre-trained predictive model, wherein the predictive model is a linear model using process variables including the unit transport rate as predictor variables. In such an embodiment, the step of comparing the value derived from the unit transport rate to a predetermined value may include comparing the value of the one or more latent variables to a corresponding predetermined value.
[0023] The pre-trained prediction model may have been trained to determine the value of one or more latent variables that vary with maturity. The prediction model may be a linear model that uses a unit transport rate and an optional process condition as predictor variables and uses maturity as a response variable. The prediction model may be a multivariate model such as a PLS or OPLS model. The multivariate model may be a PLS model defined in equations (1) and (2), wherein: X is an m×n matrix of process variables at maturity m, Y is an m×1 matrix of maturity values, and T is an m×l matrix of score values that describe aspects of the process variables that are most relevant to maturity, including one or more latent variables.
[0024] The prediction model may be a principal component regression (PCR). The prediction model may be a PCR model, wherein PCA is applied to a matrix X of process variables at maturity m, and a matrix Y of maturity values is regressed on the principal components obtained thereby to identify the principal components most correlated with maturity, and the PCA scores represent latent variables describing aspects of the process variables most correlated with maturity.
[0025] The predictive model may have been pre-trained using data from a plurality of similar biological processes that are believed to be operating normally, wherein a similar biological process is a biological process that uses the same cells for the same purpose. The predictive model may have been pre-trained using data from a plurality of similar biological processes, wherein at least some of the biological processes differ from one another in one or more process conditions that vary with maturity.
[0026] The machine learning model may be a regression model. The machine learning model may be selected from a linear regression model, a random forest regressor, an artificial neural network (ANN), and a combination thereof.
[0027] Advantageously, the machine learning model is an ANN or a collection of ANNs. The inventors have found that ANNs are particularly well suited for the task at hand. Without wishing to be bound by theory, the inventors believe that this is at least in part because ANNs are well suited for predicting values using input data with complex correlation structures. Therefore, ANNs perform well in jointly predicting multiple unit transport rates.
[0028] The linear regression model may be a partial least squares or an orthogonal partial least squares model.The machine learning model may include a plurality of machine learning models, wherein each machine learning model has been trained to predict unit transport rates for a separately selected subset of one or more metabolites.
[0029] The machine learning model may have been trained to jointly predict the unit transport rates of one or more metabolites at later maturity based at least in part on the values of one or more process conditions of a biological process at one or more previous maturity. Jointly predicting the unit transport rates of multiple metabolites (i.e., all or a subset thereof) may advantageously improve the accuracy of the prediction, wherein the unit transport rates of the jointly predicted metabolites are correlated with each other. Without wishing to be bound by theory, many metabolites will be associated with unit transport rates that are at least somewhat correlated due to the metabolites' participation in related metabolic pathways within the cell.
[0030] Obtaining values of one or more process conditions at one or more maturities may include obtaining values of one or more process conditions at multiple maturities; and the machine learning model has been trained to predict the unit transport rate of one or more metabolites at the latest or later maturity of the multiple maturities based at least in part on the values of the one or more process conditions of the biological process at multiple maturities. The inventors have found that predicting the unit transport rate using predictor variables at multiple previous maturities advantageously improves the accuracy of the prediction compared to using predictor variables at a single previous maturity. Without wishing to be bound by theory, the accuracy of machine learning predictions can be improved by using predictor variables at multiple time points because such data can capture information about the dynamics of the biological process.
[0031] Advantageously, the machine learning model may have been trained to predict the unit transport rate of one or more metabolites at the latest or later maturity of two different degrees of maturity based at least in part on the values of one or more process conditions of the biological process at the two different degrees of maturity. The inventors have found that predicting the unit transport rate using predictor variables at two previous degrees of maturity advantageously improves the accuracy of the prediction compared to using predictor variables at a single previous degree of maturity, but the degree to which the accuracy is further improved by including additional degrees of maturity is not as significant. In other words, the inventors have found that many of the benefits associated with using multiple degrees of maturity can be obtained by using only two degrees of maturity.
[0032] The values of one or more process conditions used as input to the machine learning model can be associated with multiple maturity levels that are separated by a difference in maturity that is approximately equal to the maturity difference between the values used to train the machine learning model.
[0033] Predicting one or more characteristics of a biological process may include determining values of one or more variables derived from a unit transport rate by using the unit transport rate to determine a concentration of corresponding one or more metabolites at a later maturity, optionally wherein determining the concentration of corresponding one or more metabolites at a later maturity includes solving a corresponding material balance equation.
[0034] In embodiments, one or more metabolites comprise a desired product. Also described herein are methods of predicting the concentration of a desired product in a biological process.
[0035] Determining the concentration of metabolite i at maturity k can include the i Integrating any of equations (4), (4a)-(4f), and (28) between a known previous maturity and a maturity k, wherein k is the maturity associated with the predicted unit transport rate. The method may also include determining values of one or more variables derived from the unit transport rate by using the unit transport rate to determine concentrations of corresponding one or more metabolites at a later maturity, and using one or more of the above concentrations to determine a value of a biomass-related metric at the later maturity, optionally wherein determining the value of the biomass-related metric at the later maturity includes solving a kinetic growth model. Solving the kinetic growth model at maturity k may include determining the value of x at x. v 、x l 、x d and / or x t Integrate any of equations (14) to (17) between the known previous maturity and the maturity k.
[0036] In embodiments, the biomass-related metric comprises the amount of biomass. In some such embodiments, the biomass is a desired product of a biological process.
[0037] The method may also include using one or more metabolite concentrations and / or biomass-related metrics as inputs to the machine learning model to predict unit transport rates at other maturity levels. In such embodiments, the method may be used to simulate a biological process by iteratively determining values of process conditions used as inputs to the machine learning model to predict unit transport rates at subsequent maturity levels, and determining metabolite concentrations and biomass-related metrics from the subsequent maturity levels.
[0038] Similarly, a method for simulating a bioprocess is also described according to the present aspect, the bioprocess comprising a cell culture in a bioreactor, the method comprising: obtaining initial values of one or more process conditions, the one or more process conditions comprising one or more process parameters, one or more metabolite concentrations and / or one or more biomass-related metrics of the bioprocess at one or more initial maturity levels; determining a unit transport rate of one or more metabolites in the cell culture using the obtained values as input to a machine learning model, the machine learning model being trained to predict the unit transport rate of one or more metabolites at the latest maturity or later maturity in one or more maturity levels based at least in part on the values of the one or more process conditions of the bioprocess at one or more maturity levels; predicting the values of one or more process conditions of the bioprocess based at least in part on the determined unit transport rate, the one or more process conditions comprising one or more process parameters, one or more metabolite concentrations and / or one or more biomass-related metrics; and repeating the above-mentioned determining step and predicting step using the predicted values of the one or more process conditions.
[0039] The values of any process parameters used as inputs to the machine learning model may be set to corresponding initial values (i.e., assuming that the process parameters remain constant). Alternatively or in addition, the values of any process parameters used as inputs to the machine learning model may be obtained from a user interface, a computing device, or a memory. For example, trajectories of one or more process parameters may be provided as inputs to the method, wherein the trajectories of the process parameters include values of the process parameters at multiple maturity levels.
[0040] The various methods of this aspect can be performed using process condition values received from a user, a computing device, or a memory. The various methods of this aspect can be performed using process condition values obtained in real time during the bioprocess operation. Such process condition values can include operation settings (i.e., parameters set by an operator), and / or values measured online or offline during the bioprocess operation. Therefore, the various methods of monitoring bioprocesses described herein can be implemented in real time, i.e., implemented in real time during the operation of the bioprocess. In such an embodiment, the step of obtaining the value of one or more process conditions can include receiving the value of one or more process conditions at the latest maturity (such values having been measured or determined at the latest maturity), and optionally including obtaining the value of one or more process conditions at one or more previous maturity from a data storage.
[0041] Any of the methods of this aspect may include recording in a data store the value of one or more process conditions, the determined unit transport rate, and / or a value derived from the unit transport rate.
[0042] The method may further comprise: outputting a signal to a user if the comparing step indicates that the biological process is not operating normally. The signal may be output via a user interface such as a screen or via any other means such as audio or tactile signalling.
[0043] Comparing the value of one or more unit transport rates or variables derived from one or more unit transport rates with one or more predetermined values may include comparing the value of one or more unit transport rates or variables derived from one or more unit transport rates with an average value of the corresponding variable in a set of biological processes that are considered to be operating normally. If the value of one or more unit transport rates or variables derived from one or more unit transport rates is within a predetermined range of the average value of the corresponding variable in a set of biological processes that are considered to be operating normally, then the biological process may be considered to be operating normally. The predetermined range may be defined as a function of the standard deviation associated with the average value of the corresponding variable in a set of biological processes that are considered to be operating normally. If the value of one or more variables t is within a range defined as average(t)±n*SD(t), then the biological process may be considered to be operating normally, where average(t) is the average value of the variable t in a set of biological processes that are considered to be operating normally, SD(t) is the standard deviation associated with average(t), and n is a predetermined constant (n may be the same for the subrange average(t)+n*SD(t) and the subrange average(t)-n*SD(t), or may be different between these subranges). In an embodiment, n is 1, 2, 3, or a value that achieves a selected confidence interval (e.g., a 95% confidence interval). In an embodiment, if the value of one or more variables t is within a range defined as a confidence interval (e.g., a 95% confidence interval around average(t) based on an assumed distribution of t), the biological process may be considered to be operating normally. The assumed distribution may be a Gaussian (normal) distribution, a chi-squared distribution, etc. In the case where the assumed distribution is a normal distribution, a p% confidence interval (where p may be, for example, 95) may be equivalent to a range of average(t)±n*SD(t), where n is a single value that achieves a p% confidence interval (e.g., for a 95% confidence interval, n may be approximately 1.96).
[0044] Unless the context indicates otherwise, all steps of the methods described herein are computer-implemented. In particular, any step of the method may be implemented by a computing device, optionally in operable communication with one or more sensors, other computing devices, and / or user interfaces.
[0045] The above method may also include predicting the impact of a specific value of a process parameter at a later maturity by including the specific value in an input value used in a material balance equation, a kinetic growth model, and / or a machine learning model to predict unit transport rates at other maturity levels.
[0046] According to a second aspect, there is provided a computer-implemented method for controlling a biological process comprising a cell culture in a bioreactor, the method comprising: implementing the steps of any embodiment of the above-mentioned first aspect; comparing the unit transport rate or a value derived from the unit transport rate with one or more predetermined values; determining whether to implement a corrective action based on the above-mentioned comparison; and if the above-mentioned determination step indicates that a corrective action is to be implemented, sending a signal to one or more effector devices to implement the corrective action.
[0047] The method according to this aspect may have any feature disclosed in relation to the first aspect. The method according to this aspect may also have any feature or any combination of the following optional features.
[0048] Determining whether to implement corrective action based on the comparison may include determining whether the biological process is operating normally based on the comparison, and sending a signal to one or more effector devices to implement corrective action if the comparison indicates that the biological process is not operating normally.
[0049] Determining whether to implement corrective action based on the comparison may include determining whether the biological process is operating optimally based on the comparison, and sending a signal to one or more effector devices to implement corrective action if the comparing step indicates that the biological process is not operating optimally.
[0050] Determining whether the biological process is operating optimally can include determining whether different sets of one or more process conditions are associated with an improved unit transport rate or a value derived from the unit transport rate. In such embodiments, the one or more predetermined values can include values associated with one or more different sets of process conditions.
[0051] The above method may further comprise repeating the steps of the above method of monitoring a biological process after a predetermined period of time has elapsed since a previous measurement was obtained.
[0052] The corrective action may be associated with a change in the value of one or more process conditions. The method may also include predicting the effect of the corrective action by including the above values as inputs to a material balance equation, a kinetic growth model, and / or a machine learning model for predicting unit transport rates at other maturity levels to determine the corrective action to be implemented.
[0053] An effector device may be any device coupled to a bioreactor that is used to alter one or more physical or chemical conditions in the bioreactor.
[0054] According to a third aspect, a method for optimizing a biological process is provided, the process comprising a cell culture in a bioreactor, the method comprising performing the method of any embodiment of the preceding aspects using a first set of process conditions and at least another set of process conditions, and determining whether the another set of process conditions is better than the first set of process conditions by comparing unit transport rates associated with each set of process conditions or values derived from the unit transport rates.
[0055] The method may also include selecting another set of process conditions and comparing the unit transport rate associated with the other set of process conditions, or a value derived therefrom, with the unit transport rate associated with one or more previously used sets of conditions, or a value derived therefrom.
[0056] The step of selecting another set of process conditions may include receiving another set of conditions from a user interface, obtaining another set of conditions from a database or a computing device, determining another set of conditions using an optimization algorithm, or a combination thereof.
[0057] According to a fourth aspect, a method of providing a tool for monitoring a bioprocess comprising a cell culture in a bioreactor is also disclosed, the method comprising the following steps: obtaining measurements of one or more process conditions, the one or more process conditions comprising one or more process parameters, one or more metabolite concentrations, and / or one or more biomass-related metrics of the bioprocess at multiple maturities; determining unit transport rates of one or more metabolites in the cell culture at multiple maturities using a material balance model and the above measurements; using the above measurements and the corresponding unit transport rates determined to train a machine learning model, the machine learning model being trained to predict the unit transport rates of one or more metabolites at the latest or later maturity in the one or more maturities based at least in part on the values of the one or more process conditions of the bioprocess at the one or more maturities. The method may include any of the features described with respect to the first aspect.
[0058] According to a fifth aspect, a system for monitoring a bioprocess is provided, the bioprocess comprising a cell culture in a bioreactor, the system comprising: at least one processor; at least one non-transitory computer-readable medium containing instructions, which, when executed by the at least one processor, cause the at least one processor to perform the following operations: obtain values of one or more process conditions, the one or more process conditions comprising one or more process parameters, one or more metabolite concentrations, and / or one or more biomass-related metrics of the bioprocess at one or more maturities; determine a unit transport rate of one or more metabolites in the cell culture using the obtained values as input to a machine learning model, the machine learning model being trained to predict the unit transport rate of one or more metabolites at the latest maturity or later maturity in one or more maturities based at least in part on the values of the one or more process conditions of the bioprocess at one or more maturities; and predict one or more characteristics of the bioprocess based at least in part on the determined unit transport rate.
[0059] The system according to this aspect can be used to implement the method according to any embodiment of the first aspect. In particular, the at least one non-transitory computer-readable medium may contain instructions, which, when executed by at least one processor, cause the at least one processor to perform operations including any operations described with respect to the first aspect.
[0060] The above system may also include one or more of the following operably connected to the processor: a user interface, wherein the instructions further cause the processor to provide to the user interface one or more of the following for output to the user: values of one or more unit transport rates or variables derived from one or more unit transport rates, the results of the above comparison step, and a signal indicating that the biological process has been determined to be operating normally or abnormally; one or more biomass sensors; one or more metabolite sensors; one or more process condition sensors; one or more effector devices.
[0061] According to a sixth aspect of the present disclosure, there is provided a system for controlling a biological process, the system comprising:
[0062] A system for monitoring a biological process according to the above aspects; and
[0063] At least one effector device operably connected to a processor of the above-described system for monitoring a biological process.
[0064] The system according to this aspect can be used to implement the method of any embodiment of the second aspect. In particular, the at least one non-transitory computer-readable medium may contain instructions, which, when executed by at least one processor, cause the at least one processor to perform operations including any operations described with respect to the second aspect.
[0065] According to a seventh aspect, there is provided a system for providing a tool for monitoring a biological process, the biological process comprising a cell culture in a bioreactor, the system comprising: at least one processor; at least one non-transitory computer readable medium containing instructions, which when executed by the at least one processor, causes the at least one processor to perform the following operations:
[0066] Obtaining measurements of one or more process conditions, the one or more process conditions including one or more process parameters, one or more metabolite concentrations, and / or one or more biomass-related metrics of a biological process at multiple degrees of maturity; determining unit transport rates of the one or more metabolites in a cell culture at multiple degrees of maturity using a material balance model and the above measurements; and using the above measurements and the determined corresponding unit transport rates to train a machine learning model, the machine learning model being trained to predict the unit transport rate of the one or more metabolites at the latest maturity or later maturity among the one or more degrees of maturity based at least in part on the values of the one or more process conditions of the biological process at the one or more degrees of maturity.
[0067] The system according to this aspect can be used to implement the method of any embodiment of the fourth aspect. In particular, the at least one non-transitory computer-readable medium may contain instructions that, when executed by at least one processor, cause at least one processor to perform operations including any operations described with respect to the fourth aspect. According to another aspect, a computer-implemented method is provided that uses a cell metabolic state observer to predict the amount of at least one biological material produced or consumed by a biological system in a bioreactor. The method comprises:
[0068] measuring process conditions, the process conditions comprising one or more process parameters, one or more metabolite concentrations, and one or more biomass-related metrics of the biological system that vary over time;
[0069] determining a unit transport rate for the biological system, the unit transport rate comprising a unit consumption rate of a metabolite and a unit production rate of a metabolite, using the measured process conditions as input to a machine learning model trained to predict the unit transport rate of one or more metabolites at a latest time or later times among the one or more times based at least in part on the values of one or more process conditions of the biological process at the one or more times;
[0070] The above process conditions and the above unit transport rate are provided to a hybrid system model for predicting the production of biomaterials, the hybrid system model comprising:
[0071] Kinetic growth models, used to estimate cell growth over time; and
[0072] a metabolic condition model that selects process conditions based on a specific consumption or secretion rate of a metabolite, wherein the metabolic condition model is used to classify the biological system into internal metabolic conditions; and
[0073] Predict the amount of biomaterial based on a mixed system model.
[0074] The kinetic growth model can be used to estimate the viable cell density. The kinetic growth model can be used to account for lysed cells. In an embodiment, the hybrid model is used as a state observer and provides an estimate of the internal metabolic state of the biological system and an estimate of the cellular state of the biological system. In an embodiment, the kinetic growth model is used as a state observer and provides an estimate of the cellular state of the biological system. The method may include: obtaining a current measurement of a metabolite; determining a consumption rate of the metabolite using a machine learning model; predicting a future concentration of the metabolite using a metabolic state observer and current measurements.
[0075] The method may further include: classifying the internal metabolic condition into an optimal category or a suboptimal category for biomaterial production; and sending a notification to a user when the internal metabolic condition is classified into a suboptimal category.
[0076] The kinetic growth model may include a Monod kinetic model or a saturation kinetic model. In an embodiment, the cell density or cell viability of a biological system that varies over time may be measured. The kinetic growth model may also be used to estimate the growth of microbial cells that varies over time.
[0077] The metabolic condition model may include one or more of a machine learning model, a deep learning model, a principal component analysis (PCA) model, a partial least squares (PLS) model, a partial least squares discriminant analysis (PLS-DA) model, or an orthogonal partial least squares discriminant analysis (OPLS-DA) model.
[0078] In an embodiment, a test sample can be obtained from a bioreactor, and the method can determine whether the amount of biological material in the test sample is within the range predicted by the hybrid system model.
[0079] When the hybrid system model is running, the parameters of the hybrid system model may be updated. The parameters of the hybrid system may include coefficients associated with the hybrid system model.
[0080] According to any aspect, process conditions can include one or more of pH, temperature, dissolved oxygen, osmotic pressure, process stream leaving the bioreactor, growth medium, byproducts, amino acids, metabolites, oxygen flow rate, nitrogen flow rate, carbon dioxide flow rate, air flow rate, and agitation rate. The growth medium or feed can include nutrients, including amino acids, sugars, or organic acids. Byproducts of the bioprocess can include amino acids, sugars, organic acids, or ammonia.
[0081] The method of the present invention may also include: determining optimal process conditions for the bioreactor based on a hybrid system model; using one or more sensors to measure experimental process conditions of the bioreactor that change over time; monitoring the measured experimental process conditions to detect deviations from the optimal process conditions; and sending a notification to a user when a deviation is detected.
[0082] The above method may also include: determining optimal process conditions for the bioreactor based on the hybrid system model; using one or more sensors to measure the experimental process conditions of the bioreactor that change over time; monitoring the measured experimental process conditions to detect deviations from the optimal process conditions; providing feedback to a controller that controls the bioreactor to automatically adjust the experimental process conditions to minimize deviations from the optimal process conditions.
[0083] The above method may also include: determining optimal process conditions for the bioreactor based on the hybrid system model; using one or more sensors to measure the experimental process conditions of the bioreactor that change over time; monitoring the measured experimental process conditions to detect deviations from the optimal process conditions; sending a notification to a user when a deviation is detected; providing feedback to a controller that controls the bioreactor to automatically adjust the experimental process conditions to minimize deviations from the optimal process conditions.
[0084] The method may further include: simulating the predicted amount of at least one biological material using a hybrid system model, wherein the hybrid system model is initialized by the process conditions; and determining one or more states of the biological system based on the simulation.
[0085] The above method may also include: adjusting the process conditions based on the optimization method to determine a set of process conditions that optimize the predicted trajectory, product quantity (titer) and / or product quality.
[0086] The above method may include calibrating a hybrid system model for predicting a biomaterial produced by a biological system in a bioreactor, comprising: obtaining experimental data including measurements of one or more process conditions, the one or more process conditions including one or more process parameters, one or more metabolite concentrations, and one or more biomass-related metrics (e.g., cell mass) for a plurality of bioreactor batches, each batch being associated with a specific set of process conditions; determining a growth rate under ideal conditions using a kinetic model of the hybrid system model based on the experimental data; determining a cell lysis parameter using the kinetic model based on the growth rate under ideal conditions and the growth rate from the experimental data; determining a specific production rate or a specific consumption rate of a metabolite using a material balance equation and the experimental data; determining a kinetic parameter of a factor that inhibits growth to minimize a difference between the growth rate under ideal conditions and the growth rate in the experimental data; training a machine learning model to predict the specific production rate / specific consumption rate; and training a metabolic condition model of the hybrid system model based on the measured specific consumption rate or the measured secretion rate of the metabolite, the metabolic condition model being used to classify the biological system into a metabolic state associated with a specific productivity of the biomaterial produced by the biological system.
[0087] A set of parameters may be provided to the optimization module to determine process conditions that optimize biomaterial production, wherein the set of parameters includes growth rate, specific consumption rate and specific production rate of metabolites under ideal conditions, and new process conditions.
[0088] The method may include monitoring the bulk properties of the bioreactor using principal component analysis (PCA). The output of the bioreactor may be predicted using partial least squares (PLS) regression. The output may be an amount of biological material. "Biological material" may include metabolites, cells, desired proteins, antibodies, immunoglobulins, toxins, one or more byproducts, target molecules, or any other type of molecule produced using the bioreactor. There may be more than one biological material of interest, including products, target organisms.
[0089] As used herein, "growth-inhibiting factors" may include substrate limitation, temperature or pH changes, or growth-inhibiting metabolites. The term "specific productivity" may refer to the amount of product produced on a per cell basis. The term "process optimization" may refer to determining the best adjustments or settings for a process. Metabolites may include any suitable analyte, including, but not limited to, amino acids (e.g., alanine, arginine, aspartic acid, asparagine, cysteine, cysteine, glutamic acid, glutamine, glycine, histidine, hydroxyproline, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, etc.), sugars (e.g., fucose, galactose, glucose, glucose-1-phosphoglucose, lactose, mannose, raffinose, sucrose, xylose, etc.), organic acids (e.g., acetic acid, butyric acid and 2-hydroxybutyric acid, 3-hydroxybutyric acid, citric acid, formic acid, fumaric acid, isovaleric acid, lactic acid, maleic acid, propionic acid, pyruvic acid, succinic acid, etc.), other organic compounds (e.g., acetone, ethanol, pyroglutamic acid, etc.).
[0090] According to another aspect, there is provided a non-transitory computer readable medium comprising instructions which, when executed by at least one processor, cause the at least one processor to perform a method of any embodiment of any aspect described herein.
[0091] According to another aspect, there is provided a computer program comprising codes which, when executed on a computer, cause the computer to perform the method of any embodiment of any aspect described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] As examples, embodiments of the present disclosure will now be described with reference to the accompanying drawings, in which:
[0093] Figure 1 A simplified process diagram showing a general biological process according to an embodiment of the present disclosure;
[0094] Figure 2 is a flowchart illustrating a method for providing a tool according to an embodiment of the present disclosure; in particular, the flowchart illustrates a model calibration process, which results in a calibrated model that can be used to predict one or more metabolic condition variables of a biological process;
[0095] Figure 3 is a flow chart illustrating a method for predicting one or more metabolic condition variables according to an embodiment of the present invention; in particular, the flow chart illustrates a model deployment process by which metabolic condition variables of a biological process can be predicted;
[0096] FIG4 illustrates the effects of selected parameters on variables of a kinetic growth model according to an embodiment of the present disclosure; Figure 4A It is shown that for the parameter θ i,n, a graph showing the value of a correction factor (y-axis) that captures the effect of a growth inhibitory variable on growth rate as a function of the value of a growth inhibitory variable (x-axis, e.g., the concentration of a cytostatic metabolite), in this example wherein the parameter θ i,n Represents the variable z n When the approximate value is greater than this approximate value, the variable z n Start to inhibit growth; Figure 4B It is shown that for the parameter θ s,n , a graph of the value of a correction factor (y-axis) that captures the effect of a substrate limiting variable on growth rate as a function of the value of a substrate limiting variable (x-axis, e.g., concentration of a metabolite such as a nutrient), in this example, the parameter θ s,n Represents the variable z n When the value is smaller than the approximate value, the variable z n Begins to have a restrictive effect on growth; Figure 4C and Figure 4D It is shown that for the parameter θ q,n Three exemplary values of (μ q,n The value of (D) and the parameter μ q,n Three exemplary values of (θ q,n ) (C), a plot of the value of the correction factor (y-axis) that captures the effect of a variable with a quadratic effect on growth as a function of the value of the quadratic effect variable (x-axis, e.g., temperature, pH), where, in these examples, the parameter θ q,n represents the spread of the effect, and the parameter μ q,n indicates the value at which maximum growth occurs;
[0097] Figure 5 A system according to an embodiment of the present disclosure is schematically shown;
[0098] Figure 6 Schematically illustrates a computing architecture for implementing a method according to an embodiment of the present disclosure;
[0099] Figure 7 A flow chart showing an exemplary method for monitoring a biological process using a hybrid model as described herein;
[0100] Figure 8 A flow chart showing an exemplary method for simulating a biological process using a hybrid model as described herein;
[0101] Figure 9A-9D An example of using a hybrid model as described herein is shown; Figure 9A-9CExamples of the effects of independent variable adjustments on growth profiles according to the present disclosure are shown: for each figure, the profile of the predicted growth state is shown as a solid line, and the associated data of the measured state is included for comparison; (A) is an example of measured and predicted trajectories for a batch with temperature change (inhibited growth), (B) is a measured and predicted trajectory for a batch with pH change and low feed rate (glucose depletion), (C) is a measured and predicted trajectory for a batch with pH and temperature change; (D) shows an example output of cell state classification by a hybrid model according to the present disclosure; the figure shows a PCA score scatter plot, providing observability of metabolic disturbances caused by glucose depletion (e.g., predictions from a state observer using PCA); glucose concentrations within a circle centered at the origin represent normal operation - glucose values outside this range indicate reduced titer or increased risk of product quality issues;
[0102] FIG10 shows the results of a bioprocess monitoring / control process according to an exemplary embodiment of the present disclosure (A), and shows a comparison between predictions obtained using various embodiments according to the present disclosure (BD); (A) for each metabolite, each graph compares: (i) on the left, (a) the unit consumption rate / secretion rate of the metabolite predicted by the metabolic neural network and (b) the mean square error (MSE) between the corresponding unit consumption rate / secretion rate calculated retrospectively based on the measurement results of the corresponding metabolite concentration (wherein the height of the bar represents the average MSE of the 6-fold model cross-validation process, and the error bar ( bar) indicates 6-fold standard deviation); and (ii) on the right, (a) the MSE between the unit consumption rate / secretion rate calculated for the corresponding metabolite (calculated as the average calculated back-tested from 12 batches of metabolic concentration measurements) and (b) the corresponding unit consumption rate / secretion rate calculated from the corresponding back-tested metabolite concentration measurements (wherein the height of the bar indicates the average MSE of the 12 batches and the error bars indicate the standard deviation around this average); note that the "benchmark" data are calculated using metrics derived from the same measurements that were used as ground truth for the MSE calculations and are therefore artificially depressed; for the corresponding metabolite, each panel (BD) compares: (a) the unit consumption rate / secretion rate of the metabolite predicted by the metabolic neural network and (b) the corresponding unit consumption rate / secretion rate calculated back-tested from the measurements of the corresponding metabolite concentration (wherein the height of the bar indicates the average MSE of the 6-fold model cross-validation process and the error bars indicate the 6-fold standard deviation); (B) the bars on the left show data using the metabolic neural network with lag=0, and the bars in the middle show data using the metabolic network with lag=1 (e.g. Fig. 10A The bars on the right show the data of the metabolic neural network using lag = 2; (C) The bars on the left show the data of the metabolic neural network using only metabolite concentrations as input values, and the bars on the right show the data of the metabolic neural network using all available variables as input values (such as Fig. 10A (D) The left bar shows the metabolic neural network using only metabolite concentrations as input values (e.g. Fig. 10C The bars on the right show data of a metabolic neural network using four of the five metabolite concentrations (excluding glucose concentration) used by the network generating the data for the bars on the left as input values;
[0103] Fig.11 Results of a bioprocess monitoring / control process according to an exemplary embodiment of the present disclosure are shown: the figure compares: (i) on the left, the MSE between the specific production rate (SPR) predicted by the titer neural network and the corresponding SPR calculated from the retrospectively measured product concentration (where the height of the bar represents the average MSE of the 6-fold network cross-validation process and the error bars indicate the standard deviation of the 6-fold); and (ii) on the right, the MSE between the SPR calculated as the average of 12 batches and the SPR calculated from the retrospectively measured product concentration (where the height of the bar represents the average MSE of the 12 batches and the error bars indicate the standard deviation around the average); again, the "benchmark" data is calculated using metrics derived from the same measurements, which are used as ground truth for the MSE calculation and are therefore artificially low;
[0104] Fig.12Results of a bioprocess simulation according to an exemplary embodiment of the present disclosure are shown; each graph is associated with a metabolite represented in a kinetic growth model used for the simulation and shows: (i) on the left: (a) the mean squared error (MSE) between the corresponding metabolite concentration predicted by the kinetic growth model using a metabolic neural network (for predicting the unit transport rate of each of glucose, glutamine, lactate, glutamate and ammonia at each time point) and (b) the corresponding (back-tested) measured concentration (where the height of the bar represents the average MSE of a 6-fold model cross-validation process and the error bars indicate the standard deviation of the 6-fold); and (ii) on the right: (a) the MSE between the corresponding metabolite concentration calculated by the same kinetic growth model using the average unit transport rate of 12 batches calculated from the (back-tested) measured concentration and (b) the corresponding (back-tested) measured concentration (where the height of the bar represents the average MSE of the 12 batches and the error bars indicate the standard deviation around the average); note that the "benchmark" data are calculated using metrics derived from the same measurements, which are used as ground truth for the MSE calculation and are therefore artificially depressed.
[0105] The accompanying drawings shown herein illustrate embodiments of the present invention, and these accompanying drawings should not be interpreted as limiting the scope of the present invention. Where appropriate, the same reference numerals will be used in different figures to represent the same structural features in the illustrated embodiments. DETAILED DESCRIPTION
[0106] Specific embodiments of the present invention will be described below with reference to the accompanying drawings.
[0107] Biological Process
[0108] As used herein, the term "bioprocess" (also referred to herein as a "biomanufacturing process") refers to a process in which biological components (e.g., cells, cell parts (such as organelles), or multicellular structures (such as organisms or spheroids)) are maintained in a liquid culture medium in an artificial environment (such as a bioreactor). In an embodiment, the bioprocess refers to a cell culture. A bioprocess typically produces a product, which may include biomass and / or one or more compounds produced as a result of the activity of the biological component. A bioreactor can be a disposable container or a reusable container that can contain a liquid culture medium suitable for conducting the bioprocess. Example bioreactor systems suitable for bioprocesses are described in US2016 / 0152936 and WO2014 / 020327. For example, the bioreactor can be selected from: an advanced microbial reactor (e.g., The Automation Partnership Ltd's 250 or 15 bioreactors), disposable bioreactors (e.g. bag bioreactors, such as those of Sartorius Stedim Biotech GmbH STR bioreactors), stainless steel bioreactors (e.g. The present invention is applicable to any type of bioreactor, and is particularly applicable to any supplier and any scale of bioreactor from a benchtop system to a manufacturing scale system.
[0109] Cell culture refers to a biological process in which living cells are stored in an artificial environment (such as a bioreactor). The methods, tools and systems described herein are applicable to biological processes using any type of cells (whether eukaryotic or prokaryotic) that can be stored in a culture. In particular, the present invention can be used to monitor and / or control biological processes using cell types, including but not limited to mammalian cells (such as Chinese hamster ovary (CHO) cells, human embryonic kidney (HEK) cells, Vero cells, etc.), non-mammalian animal cells (such as chicken embryo fibroblast (CEF) cells), insect cells (such as Drosophila melanogaster (D. melanogaster) cells, silkworm (B. mori) cells, etc.), bacterial cells (such as Escherichia coli (E. coli) cells), fungal cells (such as Saccharomyces cerevisiae (S. cerevisiae) cells) and plant cells (such as Arabidopsis thaliana (A. thaliana) cells). Bioprocesses typically produce products, which can be cells themselves (e.g., cell populations for further bioprocessing, cell populations for cell therapy, cell populations for use as products (e.g., probiotics, raw materials, etc.)), macromolecules or macromolecular structures (e.g., proteins, peptides, nucleic acids or viral particles (e.g., monoclonal antibodies, immunogenic proteins or peptides, viral or non-viral vectors for gene therapy, enzymes for example for use in the food industry and environmental applications such as water purification, depollution, etc.)), or small molecules (e.g., alcohols, sugars, amino acids, etc.).
[0110] Figure 1A simplified process diagram of a general bioprocess is shown. The bioprocess is implemented in a reactor 2, which in the embodiment shown is equipped with a stirring device 22. Four flows (also referred to herein as "streams") are depicted, but any or all of these flows may not be present depending on the particular situation. The first flow 24 is a feed flow F containing any material added to the culture in the bioreactor. F The second stream 26 is a bleed flow F having the same composition as the culture in the bioreactor. B The third flow 28A is a harvest flow F obtained by processing an auxiliary harvest flow 28C using the cell separation device 28. H The cell separation device 28 is used to produce a third (harvest) stream and a fourth stream 28B, which is a recycle flow F including cells and any culture medium that has not been completely separated in the cell separation device 28. R In the embodiment, since only the harvesting flow F is considered H is sufficient to capture the flow effectively output from the bioreactor through the harvesting and cell separation processes, so the recycle flow F can be neglected R Thus, reference to the presence or absence of a harvest stream may refer to the presence or absence of an auxiliary harvest stream 28C (as well as derived harvest streams and recycle streams - F H and F R ). It can be assumed that the harvest flow F H The feed stream, effluent stream and harvest stream (F F 、F B 、F H and F R ) may not exist, in which case the bioprocess is called an "unfed batch process" or simply a "batch process". F and harvest flow F H When the feed flow F F and discharge flow F B A bioprocess may be referred to as a "continuous" culture when it is operated in a (pseudo) steady state (from the point of view of process conditions, ie, in particular, keeping the volume of the culture constant). FBut there is no output flow (discharge flow and harvest flow, F B and F H ), the biological process can be called a "fed-batch" process. The present invention is applicable to all the above-mentioned operation modes.
[0111] The product of a bioprocess may have one or more critical quality attributes (CQA). As used herein, "critical quality attributes" are any attributes (especially including any chemical, physical, biological and microbiological attributes) of a product that can be defined and measured to characterize the quality of a product. The quality characteristics of a product may be defined to ensure that the safety and effectiveness of the product remain within a predetermined boundary. CQA may especially include the molecular structure of a small molecule or a macromolecule (especially including any one of the primary, secondary and tertiary structures of a peptide or protein), the glycosylation spectrum of a protein or a peptide, etc. The product may be associated with a "norm", which provides the value or range of the value of one or more CQA that the product needs to meet. If all CQA of a product meet the norm, the product may be referred to as a "norm" (or "meeting the norm", "within the norm", etc.), otherwise it may be referred to as "non-norm" (or "non-normative"). CQA may be associated with a set of key process parameters (CPP) and the range (alternatively, maturity-related range) of the CPP value that realizes an acceptable CQA. A bioprocess run (i.e., a specific instance of execution of a bioprocess) can be referred to as “normal” or “on specification” if the CPP is within a predetermined range considered to achieve acceptable CQAs, and “off specification” or “out of specification” otherwise.
[0112] As used herein, the term "maturity" refers to a measure of the completion of a bioprocess. Maturity is usually measured in terms of the time from the start of a bioprocess to the end of the bioprocess. Therefore, the term "maturity" or "bioprocess maturity" may refer to the amount of time starting from a reference time point (e.g., the start of a bioprocess). Therefore, the phrase "varies with bioprocess maturity" (e.g., quantifying a variable as "varies with bioprocess maturity") may refer to "varies over time" (e.g., quantifying a variable as "varies over time, e.g., since the start of a bioprocess") in some embodiments. Conversely, unless otherwise indicated by the context, reference to time-dependent variables (whether in the text or in an equation) should be understood to be applicable to any maturity measure (including but not limited to time). In particular, any measure that increases monotonically over time may be used, for example, the amount of the desired product (or unwanted byproduct) accumulated or extracted in the culture medium since the start of the bioprocess, the integrated cell density, etc. may be used. Maturity can be expressed as a percentage (or other fractional measure) or as an absolute value that tapers to a certain value (usually a maximum or minimum) at which the biological process is considered complete.
[0113] As used herein, the term "process conditions" refers to any measurable physicochemical parameter of a bioprocess operation. Process conditions may include, among others, parameters of the culture medium and bioreactor operation, such as pH, temperature, culture medium density, volume / mass flow rate of material into and out of the bioreactor, volume of the reactor, agitation rate, etc. Process conditions may also include measurements of biomass (e.g., total cell density, viable cell density, etc.) in a bioreactor or measurements of the amount of metabolites in the entire chamber of a bioprocess (including, among others, the amount of metabolites in any cell chamber, including a cell chamber, a culture chamber including culture medium and cells, and a culture chamber).
[0114] As used herein, the term "process output" refers to a value or set of values that quantifies the desired outcome of a process. The desired outcome of a process can be the production of biomass itself, the production of one or more metabolites, the degradation of one or more metabolites, or a combination of these outcomes.
[0115] The term "metabolite" refers to any molecule consumed or produced by a cell during a biological process. Metabolites include, among others, nutrients (such as glucose, amino acids, etc.), byproducts (such as lactate and ammonia), desired products (such as recombinant proteins or peptides), complex molecules involved in biomass production (such as lipids and nucleic acids), and any other molecules consumed or produced by a cell, such as oxygen (O 2 As will be appreciated by those skilled in the art, the same molecule may be considered a nutrient, a byproduct, or a desired product, depending on the particular circumstances, and this may even change with the operation of a biological process. However, all molecules that participate in cellular metabolism (whether as input or output of a reaction performed by the cellular machinery) are referred to herein as "metabolites".
[0116] Cell metabolic conditions
[0117] The term "cellular metabolic condition" (also referred to herein as "metabolic condition" or "cellular condition") refers to the value of one or more variables that characterize the metabolism of a cell in a biological process (i.e., the metabolic activity of a cell in a biological process). Cellular metabolic conditions may include, among other things, a unit transport rate of a metabolite into or out of a cell, also referred to herein as a unit consumption rate (e.g., when a metabolite (e.g., a nutrient) is primarily consumed by a cell) or a unit secretion rate (when a metabolite is primarily produced by a cell; in particular, when the metabolite is a desired product, this may also be referred to as a unit production rate (SPR)), or any variable derived from a set of variables including one or more of the following variables (e.g., variables using multivariate analysis techniques as will be further described below). For example, in some embodiments, the metabolic condition of a cell culture may be represented as [metaboliccondition] = f(δ, u, m, s), where f is a function that transforms a set of original variables into one or more variables that capture the relationship between the original variables, δ is a unit secretion rate / consumption rate of one or more metabolites, u is a process variable such as temperature, pH, etc., m is a metabolite concentration, and s is a variable representing the state of a cell culture system. As will be further described below, a state can be a variable modeled by a system of differential equations that together form a kinetic growth model. For example, the state variables of a cell culture system can include one or more of a live cell density, a lysed cell density, a total cell density, a dead cell density, or a related variable (e.g., cell viability). As will be further described below, the function f can be obtained, for example, using PCA, PLS, or OPLS. The cell consumption rate or secretion rate / production rate of a metabolite (i.e., the unit transport rate of a metabolite in and out of a cell) and the concentration of an intracellular metabolite (which can be expressed in units of mass per volume or per cell) can be considered to represent a metabolic variable (because these variables characterize the metabolism of a cell). Note that a metabolite can be transported into a cell or moved outside of a cell (e.g., a metabolite can be consumed or produced), in which case the unit transport rate quantifies the combined effects of movement in both directions. In other words, the unit transport rate of a metabolite quantifies the net amount of a metabolite transported between a cell and a liquid culture medium (e.g., as a change in the amount of a metabolite in a culture medium, reflecting movement from a culture medium to a cell, and vice versa). Furthermore, the concentration of the same metabolite in a compartment of a bioprocess (e.g., in a bulk component or liquid culture medium, which may be expressed in units of mass per volume) may be considered to represent a process variable (because the concentration characterizes a macroscopic process variable). For example, the oxygen or glucose concentration in a liquid culture medium (e.g., in units of mass / volume) may be considered a process variable (also referred to herein as a "process parameter") that describes the process (process condition) at a macroscopic level, while the oxygen or glucose concentration in a cell (e.g., in units of mass / cell) may be considered a metabolic variable that describes the metabolic condition of the cell.
[0118] The term "multivariate statistical model" refers to a mathematical model designed to capture the relationship between multiple variables. Commonly used multivariate statistical models include principal component analysis (PCA), partial least squares regression (PLS) and orthogonal PLS (OPLS). The term "multivariate statistical analysis" refers to the establishment (including but not limited to design and parameterization) and / or use of multivariate statistical models.
[0119] Principal component analysis (PCA) is used to identify a set of orthogonal axes (called "principal components") that capture decreasing amounts of variance in the data. The first principal component (PC1) is the direction (axis) that maximizes the variance of the projection of a set of data onto the PC1 axes. The second principal component (PC2) is the direction (axis) orthogonal to PC1 that maximizes the variance of the projection of the data onto the PC1 and PC2 axes. The coordinates of a data point in the new space defined by one or more principal components are sometimes called "scores." PCA, as a dimensionality reduction method, obtains scores for each data point that capture the contributions of multiple underlying variables to the diversity of the data. PCA can be applied to a set of historical data from a run of a bioprocess to characterize and distinguish between good (normal) and bad (abnormal) process conditions. This enables retrospective identification of when historical batches deviated from acceptable process conditions and explains which of the individual process variables had the greatest impact on the observed deviations in global process conditions. This can be used to investigate how such deviations can be avoided in the future.
[0120] PLS is a regression tool that identifies a linear regression model by projecting a set of predictor variables and corresponding observable variables into a new space. In other words, PLS identifies the relationship between the prediction matrix X (dimension mxn) and the response matrix Y (dimension mxp) as:
[0121] X=TP t +E (1)
[0122] Y=UQ t +F (2)
[0123] Where T and U are matrices of dimension mxl, T and U are the scores of X (the projection of X onto the new space of "latent variables") and Y (the projection of Y onto the new space), respectively; P and Q are orthogonal loading matrices (defining the new space and having dimensions nxl and pxl, respectively); and matrices E and F are error terms (assuming that both E and F are independent and identically distributed (IID) random normal variables). The score matrix T summarizes the variation in the predictors in X, and the score matrix U summarizes the variation in the responses in Y. The matrix P represents the correlation between X and U, and the matrix Q represents the correlation between Y and T. X and Y are decomposed into matrices of scores and corresponding loadings to maximize the covariance between T and U. OPLS is a variation of PLS in which the variation in X is split into three parts: the predictive part that is correlated with Y (such as the TP in the PLS model), t ), the orthogonal part (capturing system variations T that are unrelated to Y orth P orth t ) and the noise portion (such as E in the PLS model, capturing the residual). Partial least squares (PLS) and orthogonal PLS (OPLS) regression can be used to characterize the effect of process conditions on the desired process output (product concentration, quality attributes, etc.). This can be performed by fitting an (O)PLS model as described above, where X includes one or more process variables that are believed to have an effect on the process output and Y includes the corresponding measure of the process output. This can be used to determine which process variables can be controlled and how these variables should be controlled to improve or control the desired output.
[0124] The software suite (Sartorius Stedim Data Analytics) also includes a so-called "batch evolution model (BEM)", which describes the time series evolution of process conditions, called process "paths". The process path is obtained by fitting an (O)PLS model as described above, but in this model, X includes one or more process variables that are considered to be potentially relevant, measured at multiple times (maturity values) in the evolution of the process, and Y includes the corresponding maturity values. For example, a set of n process variables can be measured at m maturity values, and these nxm values can be included as coefficients in the matrix X. The corresponding matrix Y is an mx1 matrix of maturity values (i.e. a vector of length m). Therefore, the T matrix includes score values for each of the m maturity values and each of the l identified latent variables that describe various aspects of the process variables that are most relevant to maturity. By using the score values in T to train a BEM about the process path to achieve the desired product quality at the end of the process, a "golden BEM" can be defined that describes the range of acceptable process paths for future batches (achieving CQAs within specification). This allows the batches to be monitored to know that the ongoing batches are within specification. It also means that if an ongoing batch looks like it will deviate from the accepted path range, an alert can be issued to the operator to let the operator know that corrective action needs to be taken to prevent product loss. In addition, the process measurements that caused the deviation in process conditions (by analyzing the variables in X that contributed most to the score in T that was observed to have deviated from the expected process) can be highlighted to the operator to help diagnose the problem and identify the appropriate corrective action process. This can all be done in real time. In addition, the operator only needs to consider a small set of high-level parameters in normal batch processing operations, and only when problems arise does the operator choose to discuss them in depth with the appropriate subject matter experts.
[0125] Calculate the specific transport rate using the material balance equation
[0126] The unit transport rate of metabolite i (hereinafter referred to as δ m,i ) (i.e., specific consumption rate or production rate / secretion rate, depending on whether the rate is positive or negative from the perspective of the reactor) is an important variable that captures aspects of cellular metabolic conditions during a biological process. In addition, where a metabolite is a desired product, the specific secretion rate / production rate of that metabolite provides a useful indication of the productivity of the cell culture. These variables, δ, can be calculated using the material balance equations at a specific time point as described below m,i (One variable δ for each metabolite of interest m,i ), while measuring the concentration of the corresponding metabolites and the viable cell density at two consecutive time points.
[0127] The equation representing the evolution of the bulk concentration of metabolites (including in particular nutrients, byproducts and desired products) can be expressed based on the material balance equation (such as the following equation (3)):
[0128] [Total change in the amount of metabolites in the reactor] = [Total flow of metabolites entering the reactor] - Total flow of metabolites leaving the reactor] + [Metabolites secreted by cells in the reactor] - Metabolites consumed by cells in the reactor] (3)
[0129] Equation (3) represents in mathematical form the conservation of mass of the metabolite being studied in the system (reactor). Equation (3) needs to be satisfied at every time point t. The flow of metabolites in equation (3) can be expressed as mass flow or molar flow (because molar flow can be converted to mass flow by molar mass and vice versa, so that the conservation of mass expressed in the equation can be verified regardless of the unit selected), and a person skilled in the art will be able to convert one to the other. Therefore, reference to mass flow is intended to include the use of the corresponding molar flow with corresponding adjustments to the consistency of the units within the equation. Similarly, reference to concentration can refer to mass concentration or molar concentration. The flow of metabolites entering the bioreactor depends on the feed flow F F (If such a flow exists, that is, F F ≠0) and the value of the concentration of the metabolite in this flow. The flow of metabolites leaving the bioreactor depends on the harvest flow F H (if any) and the discharge flow F B The values of the transport rates (if any) and the concentrations of the metabolites in these corresponding streams. The consumption and secretion of metabolites by cells in a bioreactor depends on the viable cell density in the reactor and a variable called the "specific transport rate" (sometimes also called the "metabolic rate"), which may also be referred to as the "specific consumption rate" of the metabolite by the cells (generally, if the "specific transport rate" is negative, the cells are consuming the metabolite) or the "specific secretion rate / production rate" of the metabolite by the cells (generally, if the "specific transport rate" is positive, the cells are producing the metabolite). Thus, for a general system (e.g., Figure 1 As shown), for metabolite i, the material balance described in equation (3) can be written as the following equation (4):
[0130]
[0131] Among them, δ m,i The unit transport rate of metabolite i by cells in culture, m i is the concentration of metabolite i in the reactor, V is the volume of the culture in the bioreactor, m F,i is the concentration of metabolite i in the feed stream, mH,i is the concentration of metabolite i in the harvest stream, m B,i is the concentration of metabolite i in the effluent stream, x v is the viable cell density in the reactor, and F F 、F H and F B are the volumetric feed flow rate, volumetric harvest flow rate, and volumetric discharge flow rate, respectively (although mass flow rates may be used equivalently with appropriate coefficients for the densities of the respective flows). m,i *x v Can be , where ε is a constant selected to ensure that m is below the detection limit of the metabolite i The value of ε does not lead to errors in the estimate of the unit transport rate. Typically, ε is chosen to be approximately equal to the detection limit of the metabolite (eg, where the metabolite concentrations are normalized, ε may be chosen to be 0.05).
[0132] Equation (4) assumes that harvest stream 28A contains only material that leaves the system through auxiliary harvest stream 28C (i.e., since metabolites only leave the system through the harvest stream, auxiliary harvest streams and return flows do not need to be included in the model), and that the action of cell separation device 28 makes it possible to assume that harvest stream 28A does not contain cells. Equation (4) can be applied to include auxiliary harvest stream 28C (if mass flow rate is used, the corresponding m A,i and density ρ A ) and reflux 28B (and the corresponding m R,i and ρ R ). In addition, equation (4) can be modified to model the removal of some cells by the harvest stream. In other words, depending on the bioprocess setup and the assumptions made, additional terms can be added to equation (4) and some terms can be removed. Some examples of common bioprocess setups and their corresponding simplified equations are provided below.
[0133] As will be appreciated by those skilled in the art, the general equation in (3) may be expressed differently depending on the mode of operation (e.g., fed-batch, non-fed-batch, etc.) and the assumptions made (e.g., variable volumes, variable concentrations in various streams and bioreactors, etc.). Based on the teachings provided herein, one skilled in the art will be able to express and solve equation (3) accordingly. In addition, whether a particular assumption is reasonable may depend on the circumstances, and one skilled in the art will be able to verify whether this is the case using well-known techniques. For example, one skilled in the art will be able to verify whether the volume of the culture is constant (e.g., by checking the amount of material flowing into and out of the bioreactor or using a level sensor), whether the density of the culture medium is constant (e.g., using a hydrometer), whether the concentration of one or more metabolites is the same in one or more chambers and / or streams (e.g., using one or more metabolite sensors to measure the concentration of the metabolites in these chambers and / or streams, respectively), etc. One skilled in the art will also recognize that a particular assumption may be reasonable in one situation but not in another. For example, the concentration of small molecule metabolites in the culture medium may be the same in the bioreactor and the effluent streams (harvest stream and / or discharge stream), but the concentration of large molecules may be different between the bioreactor and one or more effluent streams if the large molecules may be intercepted by filters or other structures.
[0134] The unit transport (consumption / secretion) rate of a metabolite at a particular time point can be calculated based on equation (4) (or a simplified variation thereof as described below), using known (i.e., measured or simulated) values of metabolite concentration and viable cell density, and using a first-order finite difference approximation. For example, using such an approximation, equation (4) can be solved according to equation (5) to yield δ at time k m,i (expressed as δ m,i (t k ) or δ m,i,k ):
[0135]
[0136] where the subscripts k and k+1 represent the values at the kth time point and k+1th time point at which the values of metabolite concentration and viable cell density are available, IVCD k is the integrated viable cell density between time point k and time point k + 1. Note that the kth observed consumption rate is prospective, meaning that it represents the consumption rate for the time interval k → k + 1.
[0137] As will be further described below, the unit transport (consumption / secretion) rate of one or more metabolites of interest can be used as a variable in a metabolic condition model that classifies the metabolic condition of a cell into an optimal or suboptimal state or category, e.g., for biomaterial production.
[0138] For perfusion cultures (where there is a feed flow F F , discharge flow F B and harvest flow F H ), equation (4) can be simplified by making some assumptions. For example, it is assumed that the metabolite concentration is the same everywhere in the medium of the bioreactor, and therefore the metabolite concentration in the harvest stream and the effluent stream is also the same (in other words, it is assumed that the concentration gradient in the reactor can be ignored, so that m B,i =m H,i =m i ), and the number of cells lost in the discharge and harvest flows can be ignored, then equation (4) can be written as:
[0139]
[0140] Assume further that the volume of the culture is constant (i.e., F F =F H +F B ), the flow is constant, and using the first-order finite difference approximation of the derivative, equation (18a) can be solved to obtain at time t k The unit consumption rate / secretion rate of the metabolite is:
[0141]
[0142] For fed-batch cultures (where there is a feed flow but no effluent or harvest flow, i.e., F H =F B =0), equation (18) can be written as:
[0143]
[0144] Solving equation (4b) using the first-order finite difference approximation of the derivative yields k The unit consumption rate / secretion rate of the metabolite is:
[0145]
[0146] In embodiments where the feed stream is continuous or semi-continuous (e.g., for a trickle feed stream), the method in Eq. (5b) may be particularly useful. In embodiments where a bolus feeding strategy is implemented (i.e., the instantaneous addition of a relatively large feed stream), the pseudo-metabolite concentration pm may be used. iRewriting equation (4b), the pseudo-metabolite concentration pm i This allows the feed flow to be eliminated from equation (4b), namely:
[0147]
[0148] For a metabolite provided in the feed stream, the pseudo-metabolite concentration pm can be obtained by i : (i) using the measured (or otherwise determined, e.g., based on the volume of the initial reactor and the volume provided by one or more feed batches) reactor volume and the known feed concentration to determine how much metabolite was added to the reactor in each feed, and (ii) subtracting the value in (i) from all measurements of metabolite concentration after the feed. For metabolites that were not present in the feed (or that can be assumed not to be present in the feed), the pseudo-metabolite concentration pm can be obtained by i : (i) using the measured (or otherwise determined, e.g., based on the volume of the initial reactor and the volume provided by one or more feed batches) reactor volume to determine the change in concentration due to dilution with each feed, and (ii) adding the value in (i) from all measurements of metabolite concentration after the feed. Equation (4d) can be solved using a first-order finite difference approximation of the derivative to give the unit transport rate of the metabolite at time k:
[0149]
[0150] Equation (5d) can also be written as:
[0151]
[0152] Among them, m i,k is the concentration of metabolite i at time k, mAdd i,k is the amount of metabolite i in the bolus addition of metabolite i at time k, V k is the total volume in the bioreactor and iVCD is the total viable cell density. In addition, if the metabolite is a product of the cell (i.e., a metabolite not expected to be present in the feed, such as a desired product), the specific production rate of the metabolite can be written as:
[0153]
[0154] Among them, δ m,i (t k ) can also be written as q IgG (t k ), and m i,k+1 、m i,k It can also be written as C IgG,k+1 , C IgG,k, to indicate that the metabolite is the desired product, such as a recombinant antibody (IgG).
[0155] For unfed batch cultures (no feed, effluent or harvest flow, i.e., F F =F H =F B =0), equation (4) can be written as:
[0156]
[0157] Solving this equation, we can obtain the unit consumption rate / secretion rate at time k as:
[0158]
[0159] In an embodiment, IVCD may be calculated using equation (6): k :
[0160] IVCD k =(αx v,k +βx v,k+1 )*(t k+1 -t k ) (6)
[0161] Wherein, coefficient α and coefficient β weight the relative influence of two viable cell density values, and make α+β=1. For example, the weight of these two values can be the same, that is, a=β=0.5. In an embodiment, α and β can be selected so that α>β (for example, α=0.6 and β=0.4). These coefficients can be selected to reflect the observed cell growth behavior, and these coefficients can be selected independently for each time point. If exponential growth can be assumed, the total viable cell density can be calculated using a logarithmic transformation. For example, the total cell density can be calculated using equation (6a).
[0162]
[0163] in,
[0164] Alternatively, any method for calculating total viable cell density can be used in the methods described herein.
[0165] The above δ at time k can be solved using the known (usually measured) biomass concentration and metabolite concentration at time / maturity k and k+1 m,iThe equation (or any corresponding equation defined taking into account the configuration of the process and the set of assumptions made) is used to obtain the metabolite transport rate at each time point / maturity value for which the above measurements are available. In addition, this can be done separately for each measured metabolite. The resulting metabolite transport rate characterizes the metabolic condition of the cells in the culture as it changes with maturity and is expressed as the amount (mass or moles) of metabolite per cell per unit of maturity (i.e., typically per unit of time). This represents very valuable information about the metabolic condition of the cells, which can be used by metabolic condition models to monitor cell cultures as will be further described below. Note that in all equations for unit transport rates, the signs of all terms can be reversed, depending on whether negative rates are used to represent cell consumption of metabolites, positive rates are used to represent cell production of metabolites (i.e., the rates are described from the perspective of the culture medium), or the opposite (i.e., positive rates are used to represent cell consumption of metabolites, negative rates are used to represent cell production of metabolites, in other words, the rates are described from the perspective of the cell compartment).
[0166] Predicting unit transport rates using machine learning
[0167] Using the above methods, unit transport rates can only be accurately calculated at time points where metabolite concentration and viable cell density data are available. This limits our ability to predict the behavior of biological processes, whether for predictive control, process optimization or simulation (digital twin). The present invention solves this problem by using machine learning methods (i.e., machine learning algorithms that train and / or deploy specific machine learning models) to predict the unit transport rates of one or more metabolites at future time points based on known (measured or calculated) values of one or more variables that characterize the biological process at one or more previous time points.
[0168] In other words, according to the present invention, the unit consumption rate / secretion rate of metabolites δ can be obtained. i =f ML,i (u,m,s), where f ML,i is a unit consumption rate / secretion rate δ that has been trained to predict the unit consumption rate / secretion rate associated with i (alone or with other δ i The model f is a model of a cell culture system comprising: one or more variables selected from a process variable u (e.g., temperature, pH, etc.), one or more metabolite concentrations m (wherein the metabolite concentration can also be considered as part of the process variable, i.e., the symbol u can refer to a physico-chemical process variable and / or a metabolite concentration in the bulk culture medium), and one or more variables s representing the state of the cell culture system (e.g., viable cell density, lysed cell density, total cell density, cell viability). For the avoidance of any doubt, any of the above categories of variables may not be present, i.e., the ...ML,i (u, m, s) may not include variables selected from process variables u, may not include variables selected from metabolite concentrations m, and / or may not include variables selected from cell culture state variables s (provided that at least one variable from at least one of these categories is included). In an embodiment, the one or more variables include at least metabolite concentration m. Therefore, in some embodiments, process variables and / or cell state variables may not exist, and the trained model may predict the unit consumption rate / secretion rate δ based at least (or only) on one or more metabolite concentrations. i . Preferably, the one or more metabolite concentrations include the concentration of a metabolite for which a unit consumption rate has been predicted, and / or a metabolite whose concentration is highly correlated with the concentration of a metabolite (e.g., a direct product or a precursor thereof) for which a unit consumption rate has been predicted. Without wishing to be bound by theory, because metabolite concentrations are naturally correlated with unit transport rates, machine learning that performs satisfactorily in the task of predicting unit transport rates can be obtained using only these metabolite concentrations. However, process variables and cell culture state variables can carry complementary information, and machine learning models can advantageously learn to use such information to improve their prediction accuracy. Therefore, including a greater number and / or a greater variety of predictor variables (e.g., including variables from one or more or each of the u, m, and s categories) can obtain a model with improved prediction accuracy and / or the ability to achieve good and / or improved accuracy in more cases.
[0169] The term "machine learning model" refers to a mathematical model that has been trained to predict one or more output values based on input data, wherein "training" refers to the process of learning using training data and parameters of the mathematical model to obtain a model that can predict output values with minimal error compared to comparison (known) data associated with the training data, wherein these comparison values are generally referred to as "labels". The term "machine learning algorithm" or "machine learning method" refers to an algorithm or method for training and / or deploying a machine learning model. The machine learning models used in the present invention can be considered regression models because these machine learning models capture the relationship between the dependent variable (the unit transport rate being predicted) and a set of independent variables (also called predictors). According to the present invention, any machine learning regression model can be used. In the context of the present invention, a machine learning model is trained by using a learning algorithm to identify the function F:u,m,s→δ i , where F is a function parameterized by a set of parameters θ such that:
[0170]
[0171] in, is the predicted unit transport (consumption / secretion) rate, and θ is a set of parameters identified to satisfy equation (8):
[0172]
[0173] Wherein, L is a loss function that quantifies the model prediction error based on the observed and predicted unit consumption rate. The specific selection of function F, parameter θ and function L and the specific algorithm (learning algorithm) used to obtain θ depend on the specific machine learning method used. Any method that satisfies the above equations can be used in the context of the present invention, including any selection of loss functions, model types and architectures. In an embodiment, the machine learning model is a linear regression model. The linear regression model is a model in the form of equation (9), which can also be written as follows according to equation (9b):
[0174] Y=Xβ+ε (9)
[0175] y i =β 0 +β 1 x i1 +..β p x ip +ε i i=1,...,n (9b)
[0176] where Y is a vector with n elements yi (one for each dependent variable) and X is a vector with elements x for each of the p predictor variables and each of the n dependent variables. i1 ..x ip is a matrix with n elements 1 as the intercept values, β is a vector of p+1 parameters, and ε is a vector of n error terms (one for each dependent variable).
[0177] In an embodiment, the machine learning model is a random forest regression (random forest regressor). Random forest regression is described in, for example, Breiman, Leo. "Random forests." Machine learning 45.1 (2001): 5-32. Random forest regression is a model that includes a set of decision trees and outputs a classification that is the average prediction of each tree. The decision tree recursively partitions the feature space until each leaf (final partition set) is associated with a single value of the target. The regression tree has leaves (predicted results) that can be considered to form a set of continuous numbers. Random forest regression is usually parameterized by obtaining a set of shallow decision trees (shallow decision trees). In an embodiment, the machine learning model is an artificial neural network (ANN, also referred to as "neural network (neural network, NN)"). ANN is usually parameterized by a set of weights, which are applied to the input of each connected neuron in a plurality of connected neurons to obtain a weighted sum that is fed to an activation function to produce a neuron output. The parameters of a neural network can be trained using a method called backpropagation (see, e.g., Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. "Learning representations by back-propagating errors." Nature 323.6088 (1986): 533-536), in which the connection weights are adjusted to compensate for errors discovered during the learning process, combined with a weight update process such as stochastic gradient descent (see, e.g., Kiefer, Jack, and Jacob Wolfowitz. "Stochastic estimation of the maximum of a regression function." The Annals of Mathematical Statistics 23.3 (1952): 462-466).
[0178] Suitable loss functions used in regression problems (such as those described herein) include mean squared error, mean absolute error, and Huber loss. Any of these loss functions can be used in accordance with the present invention. Mean squared error (MSE) can be expressed as:
[0179]
[0180] The mean absolute error (MAE) can be expressed as:
[0181]
[0182] MAE is considered to be more robust than MSE for outlier observations. Huber loss (see, e.g., Huber, Peter J. "Robust estimation of a location parameter." Breakthroughs in statistics. Springer, New York, NY, 1992. 492-518) can be expressed as:
[0183]
[0184] Where α is a parameter. Huber loss is considered to be more robust to outliers than MSE and is strongly convex near its minimum. However, because MSE can solve optimization problems more easily, it is still a very commonly used loss function, especially when strong outlier effects are not expected.
[0185] In an embodiment, the machine learning model comprises a collection of multiple models whose predictions are combined. Alternatively, the machine learning model may comprise a single model. In an embodiment, the machine learning model may be trained to predict the unit secretion rate / consumption rate of a single metabolite. Alternatively, the machine learning model may be trained to jointly predict the unit secretion rate / consumption rate of multiple metabolites. In this case, the loss function used may be modified to be the average (optionally, a weighted average) of all predictor variables, as described in equation (13):
[0186]
[0187] Among them, α i are optional weights that can be selected individually for each metabolite i, δ and is a vector of actual and predicted unit consumption rates for all metabolites. Optionally, δ can be calculated before including it in the loss function. i ) (e.g., by normalizing so that the labels of all jointly predicted variables have equal variance), e.g., to reduce the variance of some jointly predicted The risks of leading training.
[0188] In an embodiment, a machine learning model is trained based on input values to predict the unit transport rate of one or more metabolites, the input values comprising known (measured or calculated) values of one or more variables characterizing a biological process at time point k. In an embodiment, a machine learning model is trained based on input values to predict the unit transport rate of one or more metabolites, the input values comprising known (measured or calculated) values of one or more variables characterizing a biological process at multiple time points (k, k-1, ...). For example, a machine learning model can be trained based on input values to predict the unit transport rate of one or more metabolites, the input values comprising known (measured or calculated) values of one or more variables characterizing a biological process at a first time point k and a second time point k-1. The above-mentioned input variables may include the values of one or more variables at each time point in a plurality of time points, wherein one or more variables may be selected independently for each time point. For example, one or more variables at a time point may partially or completely overlap with one or more variables at another time point. In an embodiment, for each time point in a plurality of time points, the above-mentioned one or more variables are the same. In some embodiments, there may be missing values in the training data (data used to train the model) or the data provided to the model in use. Training and / or using a machine learning model may include inputting one or more missing values. Imputation methods are known in the art. Imputation methods suitable for use in this context include, for example, linear interpolation, mean substitution, etc.
[0189] In an embodiment, a machine learning model is trained based on input values to predict the unit transport rate of one or more metabolites at a future time point, the input values comprising known (measured or calculated) values of one or more variables characterizing a biological process at one or more time points k, k-1, etc. In other words, the training data used may be such that: the model prediction based on the data at one or more time points k, k-1, etc. is evaluated according to the known corresponding values at time j>k, k-1, ... In an embodiment, multiple time points in the input data used for training are separated by a predetermined time period (e.g., 1 hour, 2 hours, 3 hours, 12 hours, 1 day, 2 days, etc.). For example, the input data used for training may include values at one time point and at a second time point separated by a fixed time period from the first time point. In an embodiment, the model prediction based on the data at one or more time points k, k-1, etc. is evaluated according to the known corresponding values (label values) at time j>k, k-1, ..., wherein the time point j is separated from one or more time points in k, k-1, ... by a predetermined time period. In a simple example, the machine learning model is trained to predict the unit transport rate of one or more metabolites on the second day (i.e., within 1 day) from the day when the latest input value among the above input values is input. In this example, if the input values include two time points separated by a fixed time period of one day, the machine learning model is trained to predict the unit transport rate of one or more metabolites on the second day based on the values of one or more variables on the current day and the previous day.
[0190] For the entire training data set, the above time periods (whether between input values or between input values and predicted values) can be approximately the same. Alternatively, the training data can include a set of input values that are not separated by the same time difference and / or a set of input values and corresponding known (label) values. For example, the training data can include measurements of multiple biological processes, wherein data is obtained every day in some of the multiple biological processes and every half day in other biological processes. Preferably, the training data used includes a set of input values separated by approximately the same time (or maturity, as the case may be) difference and a set of input values and corresponding label values. In the above example, this can be achieved by using only measurements associated with multiple consecutive days of the training data (including more than daily measurements), or conversely by inputting measurements at time points where measurements are not obtained. The trained machine learning model can be advantageously used to predict a unit transport rate associated with a time difference, which is the time difference between the prediction and the latest input value and / or the time difference between multiple input values similar to the corresponding time difference in the training data. The reference to time and time difference can refer to the corresponding maturity and maturity difference. In addition, a machine learning model that has been trained to predict a unit transport rate based on input values included at multiple time points can be used to predict a unit transport rate based on input values including missing values (e.g., fewer time points than the time points for training the machine learning model). For example, when the value of a single time point is available (e.g., when the machine learning model is used to predict based on initial conditions), this may be the case. Therefore, for example, a machine learning model that has been trained to predict a unit transport rate based on input values included at two consecutive time points can be used to predict a unit transport rate based on input values of zero or estimated values of another time point that the machine learning model expects as input at one time point and two consecutive time points. In general, various methods can be used to estimate missing data, such as replacing missing values with the mean, median or mode of a set of corresponding values, values randomly extracted from a set of corresponding values, etc. Alternatively, a machine learning algorithm that supports missing values can be used. Such an algorithm may include k-nearestneighbour or classification and regression trees. For example, a machine learning model can be a random forest.
[0191] Kinetic growth model
[0192] The term "kinetic growth model" refers to any model that captures the dynamics of a cell population in a bioprocess. Thus, a kinetic growth model can be used to monitor or simulate the number of viable cells (and other culture-related parameters) in a bioreactor and predict the number of cells in the bioreactor at a future time point. For example, a kinetic growth model may include one or more differential equations that model the maturity-related (usually time-related) behavior of one or more cell population variables. A cell population variable is a specific type of process condition that characterizes the viable cell population, dead cell population, lysed cell population, and / or total cell population in a bioprocess. Common cell population variables include viable cell density (VCD), dead cell density, and lysed cell density, which capture the concentrations of viable cells, dead cells, and lysed cells in a bioreactor, respectively. In an embodiment, a kinetic growth model uses the Monod equation to capture the concentration of a limiting nutrient as a function of cell growth rate. An example of a kinetic growth model is provided in the following equations (14) to (17), which describe the viable cell density x v , dead cell density x d , total cell density x t and lysed cell density x l Changes over time / maturity:
[0193]
[0194]
[0195]
[0196] x t =x v +x d +x l (17)
[0197] In equations (14)-(17), F b and F h are the discharge rate and the harvest rate (see above and Figure 1 ), V is the reactor volume, μ eff , μ d and k l are the effective growth rate, effective death rate, and lysis rate, respectively. Equation (14) contains the following assumptions: (i) living cells grow at an effective growth rate μ eff (ii) the living cells can be formed through the discharge flow F b (iii) living cells can die at an effective rate μ d (iv) no live cells are harvested through the F hexiting the reactor (in other words, when a harvest stream is present, comprising a perfect cell retaining filter - this can be a valid assumption when the amount of any cells exiting the bioreactor via the harvest stream is negligible). Calculation of the effective growth rate is critical to the operation of the model and is described below.
[0198] Equation (15) contains the following assumptions: (i) living cells die at an effective rate μ d (ii) Dead cells are formed by a first-order process at a rate k l (iii) dead cells can be converted into lysed cells through the outlet flow F b (iv) no dead cells pass through the harvest stream F h Leave the reactor. The calculation of the effective death rate is detailed below. Equation (16) contains the following assumptions: (i) The dead cells are dying at a rate k l (ii) the lysed cells can be formed through the discharge flow F b (iv) Lysed cells can be harvested by the harvesting flow F h (if any) leave the reactor. Equation (17) includes the assumption that cells are either alive, dead, or lysed. This can be used to calculate one variable from the other, such as lysed cells (which are usually not directly measurable), by breaking the equilibrium between live and dead cells.
[0199] In an embodiment, for example, as provided in the following equation (18), (by parameter μ d The captured) mortality process is modeled as a combination of the basic mortality rate and a toxicity factor.
[0200] μ d =k d +k t φ t (18)
[0201] In equation (18), k d is the primary (basic) mortality rate, k t is the “toxicity rate”, φ t is a measure of toxicity. The variable φ t It can represent the concentration of one or more components that are toxic to cells. In the embodiment, it is assumed that φ t Equal to the concentration of lysed cells x l (i.e. φ t =x l ). In other embodiments, assuming φ tis equal to the concentration of unknown biological material that accumulates in the culture as a result of cell growth and is inhibitory to cell growth and / or toxic to cells, which is designated as φ b (so that φ t =φ b ). The evolution of this variable can be captured, for example, using the following equation (19):
[0202]
[0203] This equation assumes that the production rate of the unknown product is related to the viable cell density x v proportional to the product, and the product can be obtained through the discharge flow and the harvest flow (F b 、F h ) (if present) exits the bioreactor.
[0204] In an embodiment, (by parameter μ eff The growth process is modeled as the growth rate μ under ideal conditions. max The product of (assuming that growth rate is maximal under ideal conditions) and one or more factors describing the effects of other system variables on growth. These factors can in some cases take one of three functional forms: substrate limitation η s , quadratic η q , or inhibit η i The relationship between these factors and the effective growth rate achieved can be captured in equation (20):
[0205] μ eff =μ max η s η q η i (20)
[0206] Among them, the correction factor η s Capture N s The contribution of the substrate-limiting variable, η q Capture N q The contribution of the secondary influencing variables, η i Capture N i The contribution of the suppressor variable. s Can be equal to 0 (no substrate limiting variable), 1 (one substrate limiting variable), or any natural number. N q Can be equal to 0 (no quadratic influence variable), 1 (one quadratic influence variable), or any natural number. N i Can be equal to 0 (no inhibitory variable), 1 (one inhibitory variable) or any natural number. The above correction factor can be calculated as the product of multiple correction factors, each of which captures the contribution of substrate-limited variables / secondary effect variables / inhibitory variables, as shown in the following equations (21)-(23):
[0207]
[0208]
[0209]
[0210] In an embodiment, the inhibition effect is captured using a correction factor having the form provided by equation (24) below:
[0211]
[0212] Among them, z n is the concentration of the substance with growth inhibitory effect, θ i,n It means z n A parameter of a level above which an inhibitory effect occurs. Figure 4A It is shown that for the parameter θ i,n Example of how the value of the growth inhibitory factor varies with the concentration of the growth inhibitory substance (whose effect is simulated by the correction factor) for different values of . The variable φ in the above equations (18) and (19) is t can be considered a special case of inhibition, where the concentration of the inhibitory substance depends on the viable cell density (e.g. because the inhibitory substance is produced by or due to the presence of cells). t Sometimes, the formula provided in equation (24) can also be used for modeling. Therefore, in some embodiments, equation (20) can be written as:
[0213]
[0214] Among them, θ t (or θ b , as the case may be) is a biological material The coefficient of inhibition of growth by accumulation of Can be equal to x l or As will be appreciated by those skilled in the art, in some embodiments, more than one substance may be present. Each of these substances can be modeled using equations (18) and (25) and the corresponding terms in equation (19). Substances that have growth inhibitory effects and substances that are produced by cells in culture or due to the presence of cells (i.e., substances Examples of substances with growth inhibitory effects that are not necessarily related to viable cell density include substances (e.g. antibiotics) that are included in the culture medium to have the desired effect but may also have a (ideally, slightly) growth inhibitory effect on the cell culture.
[0215] In an embodiment, a correction factor having the form provided by equation (26) below is used to capture the effects of substrate limitation:
[0216]
[0217] Among them, θ s,n is a parameter that represents the approximate level below which the variable z n (growth-limiting substrate) begins to have a limiting effect on growth. Figure 4B It is shown that for the parameter θ s,n Examples of how the value of the substrate limitation factor varies with the concentration of the limiting substrate (whose effect is modeled by the correction factor) for different values of . In equation (26), the factor 2 is used to assign the parameter θ s,n An intuitive biological meaning: approximate concentration, when z n Below this approximate concentration, limiting effects can be seen (e.g., values such that when z n The value is lower than θ s,n , the inhibition exceeds the value of the ∼0.95 threshold). This factor can be any value to achieve the same behavior while making the parameter θ s,n is adjusted accordingly and no longer has the same intuitive interpretation. In particular, this factor can be omitted entirely (i.e. set to 1). Similarly, in equations (24) and (25), the cubic term (i.e. (z n / θ i,n )) can include a coefficient that assigns the parameter θ i,n An intuitive biological meaning: approximate concentration, when z n Above this approximate concentration, limiting effects can be seen (e.g., values such that when z n The value reaches θ i,n (eg, 0.37), the inhibition exceeds the threshold value of ˜0.95). Examples of substances with substrate-limiting effects include nutrients such as glucose, amino acids, and the like.
[0218] In an embodiment, the quadratic effect is captured using a correction factor having the form provided by equation (27):
[0219]
[0220] Among them, μ q,n is a parameter representing the target value (i.e., the value at which maximum growth occurs), θ q,n is a parameter that represents the “spread” of the influence. A factor of 1 / 25 is used to assign the parameter θ q,n An intuitive meaning is that, q,n A value equal to 1 means that the quadratic effect is at the target value μ q,n±1 exceeds the 95% threshold. This factor can be any value to achieve the same behavior while making the parameter θ q,n are adjusted accordingly and no longer have the same intuitive interpretation. In particular, the factor can be omitted entirely (i.e. set to 1). Because it can make it easier for the user to set these parameters (using, for example, biological knowledge or assumptions) or actual bounds for these parameters, it is useful to use a parameter that provides these parameters (e.g., θ s,n and θ q,n ) factors with an intuitive interpretation may be advantageous. Figure 4C It is shown that for the parameter μ q,n Different values of (θ q,n An example of how the value of the quadratic influence factor varies with the concentration of a substance whose influence is modeled by the correction factor, with the value of being fixed equal to 1. Figure 4D It is shown that for the parameter θ q,n Different values of (μ q,n Example of how the value of the quadratic impact factor varies with the concentration of a substance whose impact is modeled by the correction factor, with the value of being fixed equal to 5.
[0221] As will be appreciated by those skilled in the art, other equations may be used in place of or in addition to the above equations to model cell population variables and factors that affect these variables. Together, these equations may form a state space model (sometimes also referred to as a state observer). In practice, these equations are based on a series of state variables x v 、x d 、x t 、x l The evolution of the cell culture state (also referred to herein as "cell state") is described as a function of changes in inputs to the system, including, for example, the concentrations of various substances that affect the cell state.
[0222] As described above, the state space model (including the kinetic growth model) can be extended to include one or more equations that capture the evolution of the concentration of one or more metabolites (especially nutrients, byproducts and desired products) in the main culture. In addition, these equations can be used to calculate the unit secretion rate / consumption rate. However, these equations only allow the calculation of the unit consumption rate / secretion rate at a specific time point where the metabolite concentration can be measured. Therefore, the unit transport rate calculated using the material balance equation can be used to identify that a failure has occurred, but these equations cannot predict the value of the unit consumption rate / secretion rate at a future time point. Therefore, the concentration of the metabolite that can be expected at a future time point cannot be calculated, because this requires including the unit transport rate that varies with time in equation (4) (or an equivalent equation). Because the concentration of substrate-limited, inhibited or secondary-effect metabolites will affect the cell state variables (through the kinetic growth model described above), this in turn limits the ability to calculate the cell state variables that can be expected at such a future time point. The present invention solves these problems.
[0223] According to the present invention, the unit transport rate of one or more metabolites of interest is predicted by a machine learning model. These predictions can be integrated into the above state model by including these predictions in equation (4) and equivalent equations, i.e., by extending the above kinetic growth model with the following equations for one or more metabolites:
[0224]
[0225] where ε can be set to 0 or a value reflecting the detection limit of the above metabolites, and f ML,i where (u,m,s) are the values of the corresponding variables at one or more previous time points. Such an extended state model can be solved at each time point t to predict m i 、x v 、x d and x l These new values can be used to predict the new value of the unit transport rate at the next time point using a machine learning model, which can be plugged into the state model to predict m i 、x v 、x d and x l The new value of , etc.
[0226] Figure 22 is a flow chart showing a model calibration process, which can be used to obtain a calibrated machine learning model for predicting one or more metabolic condition variables of a biological process. In step 200, the values of multiple process variables 200a and metabolite concentrations 200b (wherein process variables and metabolite concentrations can be collectively referred to as process condition variables) associated with the biological process are received. The values of these variables may have been measured using sensors (as will be further described below), and / or have been inferred from the measurements. Optionally, the values of one or more cell state variables 200c are also received. The values of these variables can be obtained by measuring and / or by inferring values from the measurement results using, for example, a kinetic growth model as described herein. The values received in step 200 include input data for the machine learning model and data (i.e., labels or data from which labels can be calculated) that can be used to verify the output of the machine learning model. In step 210, the one or more values received are provided as input to the machine learning model, and the machine learning model provides a set of predictions of unit transport rates as output in step 220. In step 230, the predicted unit transport rate is compared with the corresponding consumption rate obtained directly or indirectly from the measurement of step 200. For example, the predicted rate can be compared with the corresponding rate calculated using the material balance equation as described above. Alternatively, the predicted rate can be indirectly compared with the corresponding measured value by calculating one or more other values based on the above-mentioned predicted rate and comparing the one or more values with the corresponding measured value (or the value derived from the unit transport rate). For example, the predicted unit consumption rate / secretion rate can be used to calculate the concentration of one or more metabolites, and these concentrations can be compared with the corresponding known values (i.e., the measured concentration or the concentration derived from the measured value). Comparing the predicted value and the corresponding known value generally includes calculating the loss function value as described above. In step 240, the optimization algorithm uses the loss value calculated in step 230 to modify the machine learning model. Steps 200, 210, 220, 230 and 240 are usually repeated using different or partially overlapping input value sets received in step 200, all of which together constitute the training data. The process can be repeated until one or more stopping criteria are met. For example, the stopping criteria can include the maximum number of iterations, a threshold value of the amount of change in the value of the loss function compared to one or more previous iterations, a target loss function value that has been reached in one or more iterations, etc.
[0227] Figure 3is a flow chart showing a model deployment process by which metabolic condition variables of a biological process can be predicted. In step 300, values of a plurality of process variables 300a and metabolite concentrations 300b (wherein process variables and metabolite concentrations can be collectively referred to as process condition variables) associated with the biological process are received. Optionally, values of one or more cell state variables 200c are also received. The machine learning model trained at step 210 with the values received at step 300 is used to calculate one or more predictions of the unit transport rate output at step 220. The above predictions can then be used at step 330 to monitor, control, optimize or simulate the biological process. All values received at step 300 may have been measured using sensors, inferred from measurements, predicted by a model, or set by a user (or a combination of these).
[0228] For example, in the context of process monitoring, some or all of the values may have been measured. Machine learning models can use these measurements to predict unit transport rates, which provide indications of future metabolic conditions of cells in the bioprocess being monitored. These predictions can also be used to calculate future metabolite concentrations and / or cell states (e.g., using material balance equations and / or kinetic growth models as described herein). All of these features can be used to determine whether the bioprocess is operating normally, such as using parameterized multivariate models as described above to define the range of behaviors that are considered to be "normal". This in turn can be used to determine whether corrective actions should be taken, and can also be used to select appropriate corrective actions. Importantly, although there are multivariate methods for identifying bioprocesses that are "out of specification" operating, the current implementation of these methods only allows determination of whether the bioprocess is currently operating normally. These methods do not allow prediction of whether the bioprocess is operating normally at a future time point, or whether the bioprocess is operating normally if one or more process conditions change.
[0229] Figure 5An embodiment of a system for monitoring and / or controlling a bioprocess according to an embodiment of the present disclosure is shown. The system includes a computing device 1, which includes a processor 101 and a computer-readable memory 102. In the illustrated embodiment, the computing device 1 also includes a user interface 103, which is shown as a screen, but may include any other device for transmitting information to a user, such as by sound or visual signals. The computing device 1 is operably connected to a bioprocess control system, such as via a network 6, which includes a bioreactor 2, one or more sensors 3, and one or more actuators 4. The computing device may be a smart phone, a tablet computer, a personal computer, or other computing device. The computing device is used to implement a method for monitoring a bioprocess as described herein. In an alternative embodiment, the computing device 1 is used to communicate with a remote computing device (not shown), which itself is used to implement a method for monitoring a bioprocess as described herein. In this case, the remote computing device may also be used to send the results of the method for monitoring a bioprocess to the computing device. The communication between the computing device 1 and the remote computing device may be via a wired or wireless connection, and may be performed on a local network or a public network (e.g., on the public Internet). Each of the sensors 3 and optional actuators 4 may be wired to the computing device 1 or may be capable of communicating via a wireless connection (e.g., via WiFi, as shown). The connection between the computing device 1 and the actuators 4 and sensors may be direct or indirect (e.g., via a remote computer). One or more sensors 3 are used to acquire data related to the biological process performed in the bioreactor 2, which may be as follows: Figure 1 The bioprocess is implemented as shown. One or more actuators 4 are used to control one or more process parameters of the bioprocess performed in the bioreactor 2.
[0230] One or more sensors 3 can be online sensors (sometimes also referred to as "inline sensors") that automatically measure properties of a bioprocess as it proceeds (with or without taking a sample of the culture), or offline sensors (whether manually or automatically, samples are obtained and subsequently processed to obtain measurements). Each measurement from a sensor (or a value derived from such a measurement) represents a data point that is associated with a time (or corresponding maturity) value. One or more sensors 3 include sensors for recording biomass in the bioreactor 2, referred to herein as "biomass sensors". Biomass is typically in the form of viable cell density or a parameter from which viable cell density can be estimated. The biomass sensor can record a physical parameter from which the biomass in the bioreactor can be estimated (typically in the form of total cell density or viable cell density). For example, biomass sensors based on optical density or capacitance are known in the art. One or more sensors also include one or more sensors that measure one or more metabolite concentrations, referred to herein as "metabolite sensors". Metabolite sensor can measure the concentration of single or multiple metabolites (e.g., from several metabolites to hundreds or even thousands of metabolites) in the whole culture, culture medium chamber, biomass chamber (i.e., whole cell) or unit cell chamber. Examples of metabolite sensors are known in the art, and these examples include NMR spectrometers, mass spectrometers, enzyme-based sensors (sometimes referred to as "biosensors", such as for monitoring glucose, lactic acid, etc.), etc. Most commonly used metabolite sensors measure the concentration of metabolites in culture medium. As used herein, sensor 3 (e.g., metabolite sensor and biomass sensor) can also refer to a system that estimates metabolite concentration or biomass from one or more measured variables (e.g., measured variables provided by other sensors). For example, a metabolite sensor can actually be implemented as a processor (e.g., processor 101), which receives information from one or more sensors (e.g., measuring the physical / chemical properties of the system), and uses one or more mathematical models to estimate metabolite concentrations according to such information. For example, a metabolite sensor can be implemented as a processor that receives a spectrum from a near infrared spectrometer and estimates the concentration of metabolites from these spectra. Such sensors may be referred to as "soft sensors" (referring to the use of software to obtain the "measurements" of these sensors, rather than obtaining the "measurements" by direct measurement). The one or more sensors 3 also include one or more sensors that measure other process conditions, such as pH, volume of culture, volume / mass flow rate of material into and out of the bioreactor, culture medium density, temperature, etc. Such sensors are known in the art. Whether one or more sensors 3 that measure other process conditions are necessary or advantageous may depend at least on the mode of operation and the assumptions made by the material balance module as will be further described below.For example, where the bioprocess is not operated in a fed-batch manner, it may be advantageous to include one or more sensors for measuring the amount and / or composition of flows entering and / or leaving the bioreactor. Additionally, where the material balance module does not assume a constant volume in the bioreactor, it may be advantageous to include a sensor for measuring the volume of liquid in the bioreactor (e.g., a level sensor). Measurements from the sensor 3 are transmitted to the computing device 1, which may store the data permanently or temporarily in the memory 102. The memory 102 of the computing device may store a trained machine learning algorithm as described herein. The processor 101 may execute instructions to generate a computer program as described herein (e.g., by reference). Figure 3 ) uses the trained machine learning model and data from one or more sensors 3 to predict one or more unit transport rates. Note that as referenced Figure 5 The system can also be used as described herein (e.g., by reference to Figure 2 ) to train machine learning models.
[0231] In the context of process simulation, some values may have been measured by the user, while other values may have been predicted or set by the user. For example, the user may wish to study the effect of specifically changing one or more process conditions on one or more metrics of bioprocess performance. To this end, one or more values may be set based on the time course of the measurement, and other values may be set to represent the expected changes. Alternatively, all values may be set to represent a set of conditions intended to be used. The above values include at least the initial conditions sufficient for the machine learning model to predict the first set of unit transport rates. These predictions can be used to calculate future metabolite concentrations and / or cell states (e.g., using material balance equations and / or kinetic growth models as described herein). These values can be fed back to the machine learning model, optionally in combination with other process parameters that the user intends to set. Then, a new set of predictions can be obtained as described above, and the process can be repeated as many times as desired. As described above, these predictions themselves can provide an indication of the metabolic conditions of the cells in the bioprocess being simulated. In addition, metabolite concentrations and / or cell states calculated as part of the process simulation process (e.g., using material balance equations and / or kinetic growth models as described herein) can also provide information indicating the performance of the cell culture. Because the above simulations replicate real process conditions in bioinformatics to understand the impact of these process components on cell culture performance, such applications can be called "digital twins". For example, such simulation processes can be used to study the impact of process conditions (such as temperature profile, pH, dissolved oxygen profile, agitation, medium composition, flow parameters, etc.) on biological processes. In this case, the initial conditions can include all starting process conditions used by the machine learning model, including metabolite concentrations, and the process conditions at each subsequent time point (for example, used as input to the machine learning model and / or as parameters of the material balance / kinetic growth model) other than the metabolite concentrations (to be calculated by the model) can be set to those process conditions being studied.
[0232] In the context of process optimization, the above-described simulation process can be integrated as part of an optimization process by which multiple process conditions or combinations thereof can be studied and compared based on one or more desired criteria. For example, the desired criteria can include the concentration of a desired product or the total amount of a desired product produced within a predetermined amount of time.
[0233] As will be appreciated by those skilled in the art, the accuracy of predictions from any machine learning model depends on a combination of the data used to train the model and the specific purpose of the model. For example, a machine learning model may provide more accurate predictions when the input data has features based on time differences between multiple time points (similar to at least some of the time differences in the training data used to train the model). As another example, a machine learning model that has been trained using training data reflecting various process conditions may provide more accurate predictions than a machine learning model that has been trained using training data that captures a narrow set of conditions. Thus, without wishing to be bound by theory, for the purpose of predictive process monitoring (where the process is expected to operate under normal conditions), it may be sufficient to use data representing a set (usually, relatively narrow range) of normal conditions (conditions of the process known to be implemented within specifications). In contrast, for simulation or optimization, training machine learning using data representing various conditions may perform better. Note, however, that such simulation or optimization is not possible in the absence of predictions, so even an imperfect machine learning model may provide an advantage.
[0234] A specific embodiment of using the above predictions of specific consumption rate / secretion rate as part of a hybrid model for bioprocess monitoring, control, simulation and optimization will now be described.
[0235] Hybrid models for predictive bioprocess monitoring, control, simulation and optimization
[0236] Figure 6An exemplary embodiment of a computational architecture for a hybrid model as described herein is shown. The kinetic and metabolic state observation system 5 includes a plurality of specific processing modules, including a kinetic growth model 50, a metabolic condition model 60, a state correction model 75, a process monitoring engine 80, a tagging and alarm engine 85, and a consumption rate and secretion rate module 90. The metabolic condition model 60 may include additional modules, including a multivariate statistical modeling engine 65 (also referred to herein as a PCA and PLS statistical modeling engine) and a data-driven machine learning engine 70. The kinetic growth model 50 and the metabolic condition model 60 together form a model called a "hybrid model". The kinetic model is a model for determining the amount of living cells, lysed cells, living cells, and cell density, etc. (as described above). The metabolic condition model 60 is a model for providing product titers about a biological process and quality control information about the titers and properties (e.g., byproducts, etc.). The output of the metabolic condition model 60 can be fed to the kinetic growth model 50, i.e., some values calculated using the metabolic condition model 60 can be used as input variables in the kinetic growth model 50. The state correction model 75 can be used to update the estimate of the state of the kinetic growth model 50 based on the experimental data. An error can be derived based on the difference between the measured experimental output and the estimated hybrid model output, and the parameters of the hybrid model can be adjusted based on the error signal to drive the error signal to zero over time.
[0237] The kinetic growth model 50 is a state space model based on the Monod growth equation (see equations (14)-(27)) and the metabolite material balance equation (see equations (3)-(6) and (27)). The Monod growth equation and the metabolite material balance equation are a series of differential equations that can be used to describe cell growth (e.g., microbial cell growth), cell density and viable cell density, total cells (e.g., viable cells, dead cells, and lysed cells), etc. The inputs to the kinetic growth model may include temperature, feed conditions, pH, etc., and the outputs include state estimates. The kinetic growth model can be viewed as a typical state space model. In this context, the parameters that are modeled internally are called states. For example, these parameters may include x v 、x d 、x l 、m i . Kinetic growth models can be used to monitor the number of viable cells (and other parameters) in a bioreactor and predict the number of cells in the bioreactor at a future time point. In addition to the kinetic growth equation (Eqs. (14)-(27) above) and the material balance equation (e.g., any of Eq. (28) or Eq. (4) and equivalent equations), the kinetic growth model can include a separate equation for describing the rate at which the desired biomaterial changes over time. This can be expressed as:
[0238]
[0239] in, is the rate of change of the biomaterial over time, Qp is a function whose inputs include the metabolite concentration m, δ (m,i) (t) is the unit production rate or unit consumption rate of the metabolite at the current time, u is a set of independent variables (i.e., variables that are provided as input to the model and for which the model does not calculate new values; for example, these variables may include process parameters such as temperature, pH, etc.), x v is the viable cell density. In the case where the biological material itself is a metabolite, equation (29) can take the form of a material balance equation as described above (and therefore the equation (29) can be included as part of a set of material balance equations). Similarly, in the case where the biological material itself is biomass, equation (29) can take the form of a related kinetic growth equation (e.g., an equation that captures the viable cell density). In other words, if the desired biological material is biomass or a metabolite that has been captured in the material balance equation included in the model, equation (29) can already form part of the model as described above. The output of the kinetic growth model includes product titer, metabolite concentration, viable cell density, and vitality. The kinetic growth model can also be used to calculate the unit consumption (or secretion) rate of one or more metabolites at a given time based on the above-mentioned measurement data. According to the present invention, the trained machine learning model as described herein can be used to predict the unit consumption (or secretion) rate of one or more metabolites at a future time. In addition, some unit transport rates can be calculated using the kinetic growth model and measurement data, and other unit transport rates can be predicted using the machine learning model.
[0240] The metabolic condition model 60 classifies the internal metabolic state of the system into the best or suboptimal state or category, for example, for biomaterial production. Optionally, the input to the metabolic condition model may include temperature, feed conditions, etc. Typically, the metabolic condition model is independent of the kinetic growth model, and metabolic condition monitoring may be performed independently of kinetic growth. In an embodiment, the metabolic condition model may be used to enhance the kinetic growth model by providing an estimate of titer and / or quality to improve titer prediction. The metabolic condition model uses the unit consumption / unit production of metabolites as the minimum input. As described above, these rates may be calculated from metabolite concentration and viable cell density data, or may be predicted using a machine learning model. The metabolic condition model may optionally use additional measurement parameters and / or unmeasured states as input to improve the prediction of product titer and / or quality. In certain embodiments, the metabolic condition model includes a statistical modeling engine 65, which is used to construct a principal component analysis (PCA) or partial least squares (PLS) or orthogonal partial least squares (OPLS) model using unit consumption rate / production rate (and optionally, additional parameters or states measured from the process) as input, and generates one or more multivariate scores as output. The statistical modeling engine may include any suitable engine, including a (PCA) model, a partial least squares (PLS) model, a partial least squares discriminant analysis (PLS-DA) model, and / or an orthogonal partial least squares discriminant analysis (OPLS-DA) model. PCA can be used to characterize the changes in the data set, i.e., in this case, for characterizing metabolic changes. PLS can be used to correlate metabolic changes with the productivity (titer) or production of an important quality metric (product quality). For example, techniques for performing PLS can be found in Wold et al., PLS-regression: a basic tool of chemometrics, Chemometrics and Intelligent Laboratory Systems 58 (2001) 109-130. For example, techniques for performing PCA can be found in Wold et al., Principal Component Analysis, Chemometrics and Intelligent Laboratory Systems 2 (1987) 37-52. Both PCA and PLS can be used to reduce the dimensionality of the data set, i.e., extract a set of summary variables that capture as much variability in the data as possible. PCA is an unsupervised dimensionality reduction technique that can summarize data through linear combinations of variables without losing much information. PLS is a supervised dimensionality reduction technique that is applied based on the correlation between the dependent and independent variables. These techniques are considered to be within the scope of the state of the art in the art.
[0241] In some embodiments, the output of the metabolic condition model 60 is fed to the kinetic growth model 50. The metabolic condition model 60 can also allow the metabolic conditions of cells in the biological process to be visualized. Therefore, the output of the metabolic condition model 60 can be used as the input of the kinetic growth model 50, and can also facilitate the monitoring and visualization of the metabolic conditions of cell metabolism. The metabolic condition model can contain a data-driven machine learning engine 70 or be linked to a data-driven machine learning engine 70. The machine learning engine 70 can include one or more machine learning models, such as a neural network, a deep learning model, or other machine learning models. The machine learning engine 70 can be trained using known techniques to classify the state of the biological system / bioreactor as an optimal state or a suboptimal state. In addition, the machine learning engine 70 can be trained using known techniques to classify the state of the metabolite as an optimal state or a suboptimal state. In an embodiment, the machine learning engine 70 performs a classification comparable to the statistical modeling engine 65, and in some cases, the machine learning engine 70 performs classification with higher accuracy and precision than the statistical modeling engine 65. The machine learning engine 70 also includes one or more machine learning models, which are trained as described herein for predicting unit transport rates.
[0242] The state correction model 75 reduces the error in the system. The difference between the output of the hybrid model and the measured data can be determined, and the difference is provided to the state correction model 75 as an error signal. The state correction model is used to minimize the error and to drive the error signal to zero over time. State correction can correspond to a technique associated with an extended Kalman filter. In an embodiment, the state correction model 75 can be a separate module, or alternatively, it can be integrated into a kinetic growth model. The process monitoring engine 80 can be connected to a plurality of sensors for measuring one or more parameters associated with a bioreactor. These parameters can include temperature, oxygen level, feed conditions, pH value, or other aspects of a bioprocess that can be monitored in real time or nearly real time. These measurements can be provided to a hybrid model that simulates biological reactions. These measurements can also be provided to a metabolic condition module 60 to monitor cell metabolism and generate state estimates. The marking and alarm engine 85 monitors the process deviation of the system. If the output of the bioreactor deviates from its expected / predicted output, an alarm is provided to the user. In some aspects, the bioprocess may be suspended by the system until the process is corrected. In other aspects, the system can compensate for the above deviations (e.g., adjusting feed or process conditions to achieve a desired state (e.g., an optimal state)). The tagging and alarm engine 85 can also send notifications to the user about the state of the bioreactor. In other aspects, a notification can be sent to the user when the internal metabolic state is classified as a suboptimal category. The consumption rate and secretion rate module 90 determines the unit consumption rate and unit production rate of the metabolite / analyte in the bioprocess reactor. The unit consumption data is used by the metabolic condition model 60. The output of the metabolic condition model 60 (e.g., an estimate of unit production / unit consumption) can be provided to the kinetic growth model to predict product titer. The controller 95 can receive feedback (e.g., the output of the bioreactor) to control the bioreactor to automatically adjust the experimental process conditions to minimize the deviation from the optimal process conditions.
[0243] The database 30 contains various types of data for the kinetic and metabolic state observation system 5. The training data 32 corresponds to data for identifying kinetic model coefficients, calculating the unit consumption rate / production rate of metabolites, and / or training the metabolic condition model 60 to classify the state of the cell as an optimal state or a suboptimal state or determine the production amount of the estimated biological material (e.g., unit productivity), and / or training a machine learning model for predicting the unit consumption rate / secretion rate. The process conditions 34 correspond to the process conditions of the current bioprocess reaction. The process conditions 34 may also include ideal process conditions that have been experimentally determined. These conditions may be provided to the hybrid model (50, 60) to facilitate the process monitoring of the current bioprocess operation or the simulation / prediction of the bioreactor. The output 36 is the output of the kinetic and metabolic state observation system 5, and the output may be subtracted from the output of the experimental system to generate an error signal that is fed back into the input of the hybrid model.
[0244] The measurement of metabolites and cell density can be obtained by the process monitoring engine 80, and the measurement is provided to the consumption rate and secretion rate module 90 to determine the unit consumption rate or unit secretion rate of the metabolite by using the material balance equation as described above or by calling the machine learning engine 70. The unit consumption rate and unit secretion rate can be provided as the input of the metabolic condition model. The unit consumption rate and unit secretion rate allow the process conditions to be converted into the amount (i.e., metabolic condition variables) consumed or produced by each cell for each metabolite. In an embodiment, the unit consumption rate and unit secretion rate can be provided to the metabolic condition model 60, wherein the PCA and PLS statistical modeling engine 65 and / or the data-driven machine learning engine 70 classify the state of the cell to determine whether the system is in, for example, the optimal condition or suboptimal condition with respect to process parameters (e.g., temperature, feed concentration, pH, etc.).
[0245] like Figure 7As shown, the above architecture can be used to monitor biological processes. Using such an architecture instead of using only measurable features of the biological process for monitoring means that the internal state of the biological process (i.e., characteristics of cell culture and cell metabolism) can be estimated, thereby providing a richer depiction of the state of the biological process. For example, a set of process parameters (also called "inputs" or "outputs" of the bioreactor, depending on whether these parameters are set as operating conditions or whether these parameters are generated by cell activity in the biological process) can be measured in step 705 and used as inputs to the hybrid model. These parameters may include metabolite concentration, viable cell density (VCD), product titer, product quality, cell viability, product quality, temperature, pH, dissolved oxygen (DO), etc. The above process parameters are usually measured over time. In an embodiment, zero-order hold is used to estimate the value of the sampling interval. The measured process parameters are provided to the kinetic growth model 50. In step 710, the kinetic growth model can use some of such data (particularly initial conditions) to initialize state values (e.g., metabolite values), and the state includes x v 、x d 、x l 、m i The kinetic growth model 50 may also be initialized at step 715 with parameters that may be provided by the user, retrieved as default values, or received with the measurement data. The kinetic growth model 50 may then be used to determine the state x at step 725 simply by solving the equations in the model. v 、x d 、x l 、m iThe value of . This can be achieved by integrating the kinetic equations in the kinetic growth model from the time when the reaction starts to the current time given the measured values of the process parameters. Solving the kinetic growth model requires understanding the unit consumption rate / production rate at each modeling time point. In step 720, the unit consumption rate / production rate at each modeling time point can be understood from previous experiments, that is, by calculating the rate from the metabolite concentration measurement and the viable cell density measurement in the previous experiment, and assuming that these rates are also applicable to the present process (especially if the values obtained at the corresponding time are used). Alternatively, in step 720, as described herein, a machine learning model can be used to predict one or more unit consumption rates / secretion rates based on process parameters and / or kinetic growth model state estimates at previous time points. The output of the kinetic growth model 50 may include product titer, metabolite concentration, viable cell density, and viability. As understood by those skilled in the art, the kinetic growth model may include equations that capture changes in the concentration of one or more metabolites. Similarly, the unit consumption rate / secretion rate of one or more metabolites can be predicted and used in a metabolic condition model.
[0246] The kinetic growth model does not reflect the internal state of the cells in the bioreactor. If the cells are subject to fluctuations or deviations in process conditions (e.g., changes in temperature, increases or decreases in metabolites, changes in feed conditions, etc.), the cells may enter a suboptimal state, and the output (e.g., biological titer in production) may be suboptimal. Therefore, in an embodiment, the kinetic growth model 50 can be supplemented with a metabolic condition model 60. This provides a method to estimate the internal state of the cell and associate output production with environmental variables to optimize production. In a multidimensional system with a large number of states (including some states that are related to each other), this may be particularly useful because it is difficult to determine the process conditions that should be adjusted in this case. Therefore, the metabolic condition model 60 provides a method to estimate the internal state of the cell and associate output production with environmental conditions (variables) to optimize titer production. In addition, if the system deviates from the optimal range, the user can receive a notification that prompts the user to correct the bioprocess to return the reaction to the optimal conditions. In some embodiments, the system can automatically correct the feed or environmental conditions to return the system to the optimal conditions. For example, by providing feedback to a controller that controls the bioreactor, the experimental process conditions can be automatically adjusted to minimize deviations from the optimal process conditions. As described below, you can also use the optimization process to determine process adjustments.
[0247] The function of the metabolic condition model 60 will now be described. In step 750, based on the unit consumption rate and unit production rate calculated and / or predicted in step 745 as described above, the metabolic condition is determined using the metabolic condition model. In step 755, alternatively, the metabolic condition can also be used to calculate specific productivity or one or more quality attributes. For example, this can be completed using a (O) PLS model, which has been trained to predict these values according to the metabolic condition. The product titer can be estimated using the specific productivity (production of the target protein (target protein)) in the kinetic growth model. In order to carry out process monitoring, in step 760, one or more classification methods (for example, machine learning classifiers or by comparing the current metabolic condition with the metabolic condition or condition range considered to correspond to the normal state and / or the optimal state) can be used to classify the current metabolic condition. In addition, the data-driven method provided herein (for example, PCA and PLS statistical modeling engine 65 and data-driven machine learning engine 70) can be used to reduce the dimensionality of the system, and / or allow identification of conditions that affect titer. For example, when observing a system with a large number of state variables including correlated variables (e.g., including multiple process conditions, metabolite concentrations, growth variables, and specific consumption rates / secretion rates), it can be difficult to even identify deviations from an optimal situation, let alone determine the process conditions that should be adjusted to affect titer. The present technology provides an improvement over the prior art because it provides detailed and specific control of a bioreactor by identifying process variables that are outside of an optimally determined range. In an embodiment, "optimal" refers to a range of process conditions or feed conditions that correspond to optimal titer production. However, "optimal" can be defined arbitrarily.
[0248] At step 730, the output of the hybrid model (e.g., state estimate, metabolic state, etc.) or the output of a separate kinetic growth model (in embodiments that do not include a metabolic condition model) can be compared to the output of the bioreactor, and the difference between the measured parameters and the estimated parameters can be fed back to the input of the hybrid model 220 through the state correction model 75 to improve the model. The state correction model 75 seeks to modify the parameters to minimize the difference between the measured bioreactor output and the hybrid model output to drive the error signal to zero over time. In general, when the state of the system cannot be measured directly, a state estimator can be used to estimate the internal state of the system. In particular, at step 740, a Kalman filter can be used to determine the best estimate of the internal system state based on indirect measurements in a noisy environment. That is, the Kalman filter can be used to best estimate the internal state of the system based on process conditions and a kinetic model. The Kalman filter is particularly suitable for achieving the best estimate of the system state in a noisy system. In this example, the state correction model 75 can include a Kalman filter or an extended Kalman filter. A Kalman filter or an extended Kalman filter (EKF) can be used to determine the optimal state value, wherein the error is combined with the uncertainty in the model state estimate, the uncertainty in the state measurement, and the covariance of the error. The Kalman filter can be applied to the kinetic growth model to improve the accuracy of the cell state estimate provided by the model. As is known in the art, historical training data can be used to calibrate (the process refers to identifying the appropriate parameters in the model) the kinetic growth model and / or the metabolic condition model.
[0249] The hybrid models described above can also be used for optimization. Optimization can include searching various input variables, typically using an optimization algorithm, to obtain a set that maximizes titer, quality, or other desired outcomes. Optimization can be performed at run time (e.g., when monitoring a bioprocess as described above) to predict process conditions that can be used to restore a process that has been identified as performing suboptimally using a hybrid model to an optimal state. Optimization can also be performed independent of any particular run, for example, to identify optimal process conditions for future bioprocesses. Input variables typically include nutrient additions and independent process parameters such as temperature and pH. By reference Figure 8 , the optimization process can be performed as follows. In step 810, the input variable (u i ) , the input variable may include nutrient addition. In step 815, the state value (x v 、x d 、x l 、m i) (i.e., the state of the kinetic growth model). Then, in step 820, the equations of the kinetic growth model are integrated to determine the growth trajectory and metabolite trajectory over the appropriate time range. In step 825, the cellular metabolic conditions are determined using the metabolic condition model, and the metabolic conditions are classified using data-driven classification (via a multivariate statistical modeling engine and / or a data-driven machine learning engine). In operation 830, the specific productivity and product quality are predicted using a metabolic condition model (e.g., using a multivariate model trained to predict titer based on metabolic condition variables) and / or a kinetic growth model (e.g., by integrating equations capturing product titer and / or mass changes). In operation 835, the titer and quality of the biomaterial are predicted. This can be repeated using a set of different input variable trajectories and / or initial state values to identify the values of these variables that meet one or more optimal criteria. The optimization algorithm implements various steps to explore more possible values to identify a set of one or more values that meet these optimal criteria.
[0250] Finding the best adjustment or setting for a process can be defined as finding the set of control variables (u or independent variables, in this case process conditions and feed) that minimizes (or maximizes) the mathematical objective function j. For example, (t+k) The objective function of the titer can take the following form:
[0251]
[0252] in, is the predicted titer at a future time point, x| [t] is the current value of the state, u| [t] is the set of manipulated variables to be implemented from the current time point [t] to the future time point [t+k]. In an embodiment, multiple objectives are optimized simultaneously. This can be achieved by weighting these objectives according to their importance. For example, in order to maximize titer and keep the quality metric consistent with the target, the objective function can take the following form:
[0253]
[0254] Among them, θ is the relative weight of each parameter to be optimized, q sp is the target or set point for the quality parameter q. There is a function similar to that used for IgG (f q ), which is used to predict the future quality variable at a future time point. In addition, constraints can be added to the function. For example, the optimization goal of maximizing titer while keeping the quality within the operating specification can be controlled by equation (31) subject to the following constraints:
[0255]
[0256] In addition, to prevent the optimization algorithm from selecting a new set of infeasible inputs, constraints may also be placed on u. In addition, when exploring the space of possible values or providing instructions to the controller to change the current operating conditions (e.g., where a suboptimal state has been identified and the model has been used with the optimization algorithm to identify changes in process conditions that can correct the suboptimal state), the degree to which the optimization algorithm modifies the condition u can be adjusted by imposing a penalty on changes in u from the recipe or current settings. This prevents the controller from making unstable or large changes to the process conditions that fail to improve the target parameters. The entire objective function is then described as obtaining the best set of feasible u that maximizes titer (IgG) and keeps the quality consistent with the target and within specified limits. This may be achieved by subjecting the constraints in equation (32) and satisfying u min ≤u| [t] ≤u max The following equation governs:
[0257]
[0258] Among them, θ u is the penalty weight of u (note that there may be more than one u), u sp is the target value of u, which is usually the set value or the current value.
[0259] The technology provided herein provides a model that accurately simulates cell behavior in a bioreactor. Using experimental data, the predicted VCD and viability profiles are shown to match the experimental values measured under different feed, pH and temperature profiles. As shown in Figure 9, the above hybrid model is able to replicate the experimentally measured behavior and identify different functional cell states. In particular, Figure 9A-9C As shown, temperature changes predicted by the kinetic growth model to inhibit growth were indeed confirmed to inhibit growth. Changes in pH appear to slightly increase the rate of cell death, but do not appear to inhibit the rate of cell growth. The cells appear to adapt well to and recover from glucose depletion and glutamine depletion. Therefore, because the cells may metabolize other carbon sources, no obvious changes in growth are observed. Fig.9D As shown, cell state classification using a data-driven approach (PCA in this example) on unit consumption rates calculated based on measured data can identify biological processes operating in a suboptimal range.
[0260] In addition, the hybrid model effectively acts as a soft sensor for cell metabolism and metabolites, allowing the unit consumption and unit production of metabolites to be monitored and characterized, as well as the changes in cell state and metabolic activity to be monitored and characterized. Unlike other models that do not estimate lysed cells or otherwise consider lysed cells, the hybrid model considers the number of lysed cells, which affects the toxicity of the main fluid. Such a method makes the hybrid model more accurate than other models that do not consider this feature, and the time dimension is longer than other models. Other advantages of the hybrid model include being able to learn more about cell metabolism and drive cell growth, cell death, vitality, titer and product quality factors. The hybrid model can also simulate the performance of new process conditions (e.g., feed, temperature, pH profile, etc.) to maximize productivity and observe cell state (e.g., metabolic activity, etc.) or its changes. In other aspects, perfusion performance can be predicted from fed-batch operation. These technologies improve prediction, improve titer prediction and product quality based on monitoring and prediction. This technology can be applied to various application fields, including simple univariate metabolite state estimator, comprehensive multivariate metabolite state estimator, self-generated system (e.g., digital twin simulation), etc. Thus, the present technology provides improvements in the fields of bioreactor control and biologics manufacturing.
[0261] Example
[0262] Exemplary methods for calibrating machine learning models and exemplary methods for simulating biological processes will now be described.
[0263] Materials and methods
[0264] data
[0265] The data used in these examples were collected from 12 batches of cell culture in a microbioreactor. Each individual batch was allowed to run for 12 days. The cells in the bioreactor were Chinese hamster ovary (CHO) cells that produce recombinant antibodies (IgG for short). The concentration of such a product is called "titer". Over a 12-day period, these batches were active, and the state of the bioreactor was measured 2-3 times a day (irregularly), resulting in more than 300 observations. A total of 11 independent variables were measured: viable cell density (VCD), cell viability (which can be obtained as VCD / TCD, where TCD is the total cell density and can be equal to VCD+DCD, where DCD is the dead cell density - in this example, the above cell viability and VCD are obtained by measuring TCD and then counting cells in the presence of a dye (the dye is, for example, a fluorescent dye, used to stain live cells to obtain VCD)), dissolved oxygen (DO), pH, temperature, volume, and concentrations of glucose, glutamine, lactate, glutamate, and ammonia. Furthermore, titers were measured once a day, making the titer dataset roughly one-third the size of the metabolic dataset.
[0266] Missing values were linearly interpolated to keep as many observations as possible. Some titer measurements were poorly performed, resulting in negative unit production rates, which are biologically implausible, so these titer measurements were removed.
[0267] The data for each measurement used for training is normalized (each observation X is scaled to Where μ is the mean of the variable and σ is the standard deviation). Standardization is expected to increase the speed and stability of training because it reduces the risk that variables with large values will have a disproportionate impact on training (because of their values rather than because of their importance to prediction). Standardization causes all variables Z to be distributed with mean = 0 and standard deviation = 1. Normalization is performed only on the training data (see below) to reduce bias on the validation dataset.
[0268] Tag calculation
[0269] The term “label” refers to the (hypothesized) true value of a variable that is predicted by a machine learning model. In these examples, the machine learning model is trained to predict the value of a variable at multiple time points (t k ) m,i ) specific consumption rate (SCR) and product (q IgG) of the unit production rate (SPR). Therefore, the label of each of the above multiple time points is the (hypothetical) true value of these SCR and SPR. For each time point, these parameters are calculated using the following equations:
[0270]
[0271] q IgG (t k )=(C IgG,k -C IgG,k+1 )iVCD -1
[0272] iVCD=(0.6x v,k -0.4x v,k+1 )(t k+1 -t k )
[0273] Among them, m i,k is the concentration of metabolite i at time k, mAdd i,k is the batch addition of metabolite i at time k (liquid added to the reactor at discrete time intervals to provide more nutrients), V k is the total volume in the bioreactor, iVCD is the total viable cell density, x v,k is the viable cell density at time k, C IgG,k is the concentration of the product in the bioreactor at time k.
[0274] All labels on the training dataset were normalized as described above. The final titer labels were calculated after filtering the data to remove unreasonable values.
[0275] Machine Learning Models
[0276] In these examples, two machine learning models were trained independently: one machine learning model was used to predict the unit production rate of the product, and the other machine learning model was used to (jointly) predict the unit consumption rate / production rate of each measured metabolite. The former machine learning model is called the "titer network" and the latter machine learning model is called the "metabolic network". Both networks are feedforward neural networks (neural networks, NN). Various architectures were tested (data not shown) before selecting a fully connected neural network with three hidden layers. Multiple lag values (no lag, lag = 1, lag = 2) and multiple input variables were tested.
[0277] The final metabolic network uses lag=1 and one-step-ahead prediction (i.e., the metabolic network predicts the value of one or more variables at time t - i.e., the unit consumption rate between time point t and time point t+1 - as a function of the value of one or more variables at time t and t-1). In fact, the performance of the network with lag=2 (two lag values) was not found to be significantly improved over the performance of the network with lag=1, taking into account the increase in computational time caused. In contrast, lag=1 was associated with a significant improvement in performance compared to lag=0 (see Fig. 10B ). The final metabolic network had 22 input nodes (11 variables, each provided for two time points), 22 nodes in each hidden layer, and 5 output values. Different sets of input variables were also tested to investigate the effect of using all available variables, the effect of using only metabolite concentrations, or the effect of using a subset of metabolite concentrations. Results (see Figure 10C-10D ) showed that using metabolite concentrations alone (or even a subset of metabolite concentrations) was sufficient to obtain acceptable predictions. However, since additional variables were expected to improve the predictions or at least the robustness of these variables, all available variables were used in the final network.
[0278] The final titer NN does not use hysteresis values and predicts one step ahead (i.e., the metabolic network predicts the value of one or more variables at time t - i.e., the unit consumption rate between time point t and time point t+1 - as a function of the value of one or more variables at time t). In fact, comparing the titer NN with lag = 1 to the titer NN with lag = 0, no significant difference was found between the two (MES = 123 ± 116 for the titer NN without lag, MES = 122 ± 109 for the titer NN with lag 1). Since the introduction of hysteresis did not significantly improve performance, a simpler architecture was chosen for the bioreactor simulator for computational efficiency and ease of implementation.
[0279] ReLU is used as the activation function for the above two networks.
[0280] Training of machine learning models
[0281] Training or fitting a network is done by using a training dataset. This dataset contains input values and the corresponding true output values or labels. The input values are fed through the network and the output of the network is compared to the labels through a loss function.
[0282] The loss function determines how close the network's prediction is to the label, if the loss is 0 then the network's prediction is the same as the label. A common loss function used for regression problems is the mean squared error (MSE). This loss or error is then fed back through the network to calculate the gradient of the loss function with respect to the weights. This is called backpropagation of the error. The optimizer then adjusts the weights to minimize the loss function by using the aforementioned gradients (usually based on gradient descent).
[0283] Typically, the network is trained on the entire dataset multiple times, with each iteration of the dataset being called an epoch. Computing the loss and gradients for the entire dataset is expensive in terms of computational time and storage, so the loss is usually computed on a small sample of the dataset, called a batch. Computing gradients and updating weights on a small sample of the dataset will result in lower accuracy for the overall minimized loss function, but doing so multiple times will give similar results while using less computational power resources and time.
[0284] By training the network for enough epochs, it is often possible to achieve 0 training loss. However, this is often a case of overfitting. An overfitted model does not generalize well to unknown data. To avoid overfitting, it is common to calculate the loss of the model on a validation dataset, which is called the validation loss. The weights are not updated when checking the validation dataset. When the validation loss stops decreasing, it is often best to stop training the network to avoid overfitting.
[0285] In this example, Adam (Diederik P. Kingma, Jimmy Ba, arXiv:1412.6980, Dec 2014) is used as the optimizer, MSE is used as the loss function, and the batch size of both networks is 5. -3 The learning rate is 10, and then the training is repeated for 10 epochs. -4 The learning rate is used to train for 10 cycles. The learning rate determines how fast the optimization process converges by affecting how much the weights in the neural network are updated during training. The learning rate is usually set by trial and error and / or based on experience. When the above optimization has converged to a region of parameter space, reducing the learning rate as the optimization proceeds can allow fine-tuning of the network's weights. Titrate the NN with 10 -3 The training was performed at a learning rate of 100 epochs at most, or until no decrease in validation loss was observed, after which the model with the smallest validation loss was used. Training was stopped after 30 consecutive epochs with no decrease in validation loss. Note that this “early stopping” approach also applies to the training of metabolic NNs.
[0286] Cross Validation
[0287] This example uses k-fold cross-validation, which is based on repeating the training and test calculations on different randomly selected subsets of the original dataset. The process makes it possible to form partitions of the dataset by dividing it into k non-overlapping subsets. The test loss is then estimated by calculating the average test loss over the k trials. In the i-th trial, the i-th subset of the data is used as the test set, and the remaining subsets of the data are used for training.
[0288] In this example, 6-fold cross validation was used. The final labeled dataset for the metabolic network included 310 values, which were split into a training set of size 258 and a validation set of size 52 at each iteration of the cross validation process. The final labeled dataset for the titer network included 131 values, which were split into a training set of size 109 and a validation set of size 22 at each iteration of the cross validation process. To assign the validation set, instead of the standard practice of randomly assigning observations, observations from two full batches were assigned as validation sets. Since the observations in a batch are correlated, this strategy reduces the risk of overly optimistic validation performance.
[0289] Assessment-Monitoring
[0290] The predictions of the metabolic network were compared to the average SCR values at each time point across 12 batches (29 values per metabolite over 12 days). This reflects the situation that, without the present method, the only way to predict values for the specific consumption rate / production rate of a metabolite would be to use values derived from a set of batches that were assumed to follow a similar course.
[0291] In particular, the MSE between the predictions of the metabolic NN and the labels (see above) was calculated (for each validation dataset in the cross-validation process) and the average MSE per metabolite was used for comparison. This is called the network MSE. Similarly, the MSE between the average SCR / SPR and the labels was calculated for the 12 batches. This is called the baseline MSE.
[0292] Assessment - Simulation
[0293] In addition, the metabolic network is used to predict the specific consumption rate / production rate using the simulated values obtained from the bioreactor model, which includes a kinetic growth model as described above and a material balance equation for each of the 5 metabolites for which data is available. The model is initialized with measured metabolite concentrations and can then be run for 12 days using measured values of VCD, pH, temperature, DO, and activity but only using predicted values of metabolite concentrations. Using the kinetic growth model and the material balance equation, the above values are predicted based on the corresponding SCR / SPR in the material balance equation provided by the metabolic NN (conversely, the metabolite concentrations obtained by solving the kinetic growth equation and the material balance equation are used by the metabolic NN to predict the SCR / SPR in the next step).
[0294] The predictions were evaluated by comparing the metabolite concentrations predicted by the simulator with the average SCR / SPR values calculated from 12 batches of data (data not shown) used by a similar simulator. For each time point, the simulator used the average SCR / SPR values before and near the integration time limit.
[0295] Note that the above comparison does not fully reflect the performance of the NN method because the NN method is theoretically capable of predicting SPR / SCR values for a variety of settings (e.g., different process conditions), and the use of averages from previously acquired data assumes that the process conditions used to acquire the data are at least very similar (preferably, identical) to the process conditions being simulated. In other words, the method using averages from previously acquired data is limited to simulating known conditions and is likely to perform significantly worse when attempting to simulate unknown conditions.
[0296] Example 1 - Monitoring
[0297] In this example, the metabolic network and the titer network were trained as described in the methods section, and the predictions of the networks were compared to corresponding benchmark values calculated from previously acquired data as described above.
[0298] The results are as follows Fig. 10A and Fig.11 As shown. Fig. 10A , for each figure, the following are compared: on the left, the MSE between the SCR of metabolites predicted by the metabolic neural network and the corresponding labels (wherein the height of the bar represents the average MSE of 6-fold cross validation and the error bar represents the standard deviation of the 6-fold); and on the right, the MSE between the SCR calculated as the average of 12 batches of metabolites and the corresponding labels (wherein the height of the bar represents the average MSE of 12 batches and the error bar indicates the standard deviation around the average).
[0299] Similarly, Fig.11 The following are compared: on the left, the MSE between the SPR predicted by the titer NN and the corresponding label (where the height of the bar represents the average MSE of the 6-fold cross validation and the error bar represents the standard deviation of the 6-fold); on the right, the MSE between the SPR calculated as the average of 12 batches and the corresponding label (where the height of the bar represents the average MSE of the 12 batches and the error bar indicates the standard deviation around the average).
[0300] like Fig. 10A As shown, for each metabolite, the MSE of the metabolic network is significantly lower than the benchmark MSE. Fig.11 The mean MSE of the titer SPR predictions for the titer network is shown to be slightly lower than that of the benchmark. Note that the titer network uses less data as input (due to hysteresis = 0) and is only trained to predict a single value (the SPR of the product). In contrast, the metabolic network uses a hysteresis of 1 and is trained to jointly predict 5 values (the SCR of each of the 5 metabolites).
[0301] The SCR / SPR of metabolites are biologically related through the metabolism of cells. Therefore, if the network architecture can capture the correlations between the values of multiple metabolites (these correlations reflect the biological properties of these metabolites), then jointly predicting the values of multiple metabolites can improve the performance of the network. In addition, the above products can be regarded as metabolites, so that the SPR of the product and the SCR of the metabolite can be predicted together using a single network. Assuming that the network architecture can capture the correlations between the SCR / SPR of metabolites that reflect the basic biological properties, this can improve the accuracy of the prediction.
[0302] Fig. 10B The benefit of increasing the lag value in the metabolic network is shown in terms of the MSE between the SCR of the metabolites predicted by the metabolic NN and the corresponding labels (where the height of the bar represents the average MSE of the 6-fold cross validation and the error bar indicates the standard deviation of the 6-fold). In each figure, the bar on the left shows the MSE of the predictions from the network trained with lag=0, and the bar in the middle shows the MSE from the network trained with lag=1 (as shown in Figure 2). Fig. 10A The MSE of predictions from the network trained with lag=2 is shown on the right. Fig. 10B The data suggest that increasing the lag from 0 to 1 may be beneficial in this case, whereas increasing from lag = 1 to lag = 2 does not significantly improve the prediction. Fig. 10CThe performance of a metabolic network using all available variables as input (bars on the right in each figure) is compared to the performance of a metabolic network using only 5 metabolite concentrations as input (bars on the left in each figure) in terms of the MSE calculated as described above, namely viable cell density (VCD), cell viability, dissolved oxygen (DO), pH, temperature, volume, and concentrations of glucose, glutamine, lactate, glutamate, and ammonia. The data suggest that in this case, metabolite concentrations will be sufficient to obtain useful predictions. Fig. 10D The MSE calculated as described above is shown using only five metabolite concentrations (the left bar in each figure - corresponding to Fig. 10C The performance of the metabolic network using the 5 available variables (the bar on the left in each figure) is compared to the performance of the metabolic network using 4 of the 5 metabolite concentrations (except glucose, the bar on the right in each figure). The data suggest that using concentrations of metabolites for which SCRs are predicted (or closely related metabolites) can improve the accuracy of predictions for specific metabolites. Indeed, glucose SCR predictions appear to be most affected by the lack of glucose data, while other SCR predictions remain acceptable. Since there does not seem to be a clear disadvantage to using all available variables, and the data suggest that the MSE tends to decrease with the addition of additional variables, the final network (used to generate the 5 available variables) was optimized for the 5 available variables. Fig. 10A All of these variables were used in the data.
[0303] Example 2 - Simulation
[0304] In this example, a metabolic network was trained as described in the Materials and Methods sections, and the network's predictions were used to simulate a bioreactor with initial conditions and process conditions, where the initial conditions corresponded to the initial conditions in the data shown in the Materials and Methods sections, and the process conditions were the process conditions in the data (i.e., pH, temperature, DO, volume) except for the metabolite concentrations and viable cell density (predicted by the model). The predictions for each simulation based on the metabolite concentrations were then compared to the corresponding measured concentrations by calculating the MSE. The results are shown in Figure 2. Fig.12 shown.
[0305] exist Fig.12 , each figure compares the following: on the left, the MSE between the metabolite concentrations predicted by the above model using the metabolic NN used to provide the metabolite SCR and the corresponding measured concentrations (where the height of the bar represents the average MSE of 12 batches and the error bars indicate the standard deviation around the mean); on the right, the MSE between the metabolite concentrations calculated by the same model but using the average SCR (which is calculated based on 12 batches of metabolites) and the corresponding labels (where the height of the bar represents the average MSE of the metabolites for 12 batches and the error bars indicate the standard deviation around the mean).
[0306] like Fig.12 As shown, at least for some of these metabolites, the MSE of the metabolic network is significantly lower than the baseline MSE. For the metabolites for which the network simulation performs worse than the data-based simulation, the difference in performance between the network simulation and the data-based simulation is relatively small, given that the data-based simulation compares the measured metabolite concentrations with those already obtained using the model (which uses SCRs that are calculated using metabolite concentrations that are compared to the simulation results). Moreover, the MSE of the network simulation is always an order of magnitude smaller than the data-based MSE (if, that is, the network simulation is worse than the data-based simulation). As mentioned above, the above error is an acceptable level of error, given that the network simulation opens up new avenues for simulation and optimization that are not possible with data-based simulations.
[0307] Equivalents and scope
[0308] All documents mentioned in this specification are incorporated herein by reference in their entirety.
[0309] The term "computer system" includes hardware, software and data storage devices for implementing the system according to the above-described embodiments or executing the method according to the above-described embodiments. For example, the computer system may include a central processing unit (CPU), an input device, an output device and a data storage device, and the computer system may be implemented as one or more connected computing devices. Preferably, the computer system has a display or includes a computing device with a display to provide a visual output display (for example, in the design of a business process). The data storage device may include RAM, a disk drive or other computer-readable media. The computer system may include multiple computing devices connected by a network and capable of communicating with each other through the network.
[0310] The methods of the above embodiments may be provided as a computer program or a computer program product or a computer-readable medium carrying a computer program. When the computer program is run on a computer, the computer program is used to perform the above methods.
[0311] The term "computer-readable medium" includes, but is not limited to, any non-transitory medium or medium that can be directly read and accessed by a computer or computer system. The above-mentioned medium may include, but is not limited to, magnetic storage media (such as floppy disks, hard disk storage media, and magnetic tapes); optical storage media (such as optical disks or CD-ROMs); electronic storage media (such as memories, including RAM, ROM, and flash memory); and hybrids and combinations of the above-mentioned media, such as magnetic / optical storage media.
[0312] Unless the context dictates otherwise, the descriptions and definitions of the features described above are not limited to any particular aspect or embodiment of the invention, but apply equally to all aspects and embodiments described.
[0313] As used herein, "and / or" should be understood to specifically disclose each of two specific features or components, covering situations where additional features or components are included or excluded. For example, "A and / or B" means that each of (i) A, (ii) B, and (iii) A and B is specifically disclosed, just as if each situation were listed separately here.
[0314] Note that the singular forms "a", "an", and "the" as used in the specification and the appended claims include plural references unless the context clearly dictates otherwise. As used herein, a range may be expressed as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when a value is expressed as an approximation by using the antecedent "about" or "approximately", it is understood that the particular value forms another embodiment. The terms "about" or "approximately" in relation to a numerical value are optional and mean, for example, + / - 10%.
[0315] Throughout the specification and claims, unless the context requires otherwise, the words "comprises" and "comprising" and variations thereof, will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.
[0316] Unless the context dictates otherwise, further aspects and embodiments of the present invention may be provided as described above by replacing the term "comprising" with the term "consisting of" or "consisting essentially of."
[0317] Features disclosed in the foregoing description or claims or in the drawings may be represented in their specific form, or may be represented by a device for performing the disclosed function or a method or process for obtaining the disclosed result, may be represented individually or in any combination of these features to implement the invention in different forms.
[0318] Although the present invention has been described in conjunction with the above exemplary embodiments, many equivalent modifications and variations will be apparent to those skilled in the art based on this disclosure. Therefore, the above exemplary embodiments of the present invention are considered to be illustrative rather than restrictive. Various changes may be made to the described embodiments without departing from the spirit and scope of the present invention.
[0319] In order to avoid any doubt, any theoretical explanations provided herein are used to facilitate the reader's understanding. The inventors do not wish to be bound by these theoretical explanations.
[0320] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
Claims
1. A method for monitoring a bioprocess comprising a cell culture in a bioreactor, the method comprising: include: obtaining values of one or more process conditions, the one or more process conditions comprising one or more process parameters of the biological process at one or more maturity levels, one or more metabolite concentrations, and one or more biomass-related metrics; determining a unit transport rate of one or more metabolites in the cell culture using the obtained values as input to a machine learning model trained to predict the unit transport rate of the one or more metabolites at the latest or later maturity of the one or more degrees of maturity based at least in part on the values of the one or more process conditions of the biological process at the one or more degrees of maturity; as well as One or more characteristics of the biological process are predicted based at least in part on the determined specific transport rate.
2. The method according to claim 1, in, Predicting one or more characteristics of the biological process includes: comparing the unit transport rate, or a value derived from the unit transport rate, to one or more predetermined values; and A determination is made based on the comparison whether the process is operating normally.
3. The method according to claim 1, in, The specific transport rate of metabolite i is the net amount of metabolite transported between the cell and the medium per cell and per unit of maturity.
4. The method according to claim 1, in, The one or more process conditions include one or more process parameters selected from the group consisting of dissolved oxygen, dissolved CO2, pH, temperature, osmotic pressure, stirring speed, stirring power, CO 2 pressure, feed rate, discharge rate, harvest rate, feed medium composition and culture volume; wherein the one or more process conditions include one or more biomass-related metrics, and the one or more biomass-related metrics are selected from the group consisting of live cell density, total cell density, cell viability, dead cell density and lysed cell density; and / or wherein the one or more metabolite concentrations include the concentration of one or more metabolites in a cell compartment, a culture medium compartment or the entire cell culture.
5. The method according to claim 1, in, The one or more process condition values include at least one metabolite concentration value and at least two other values, and / or wherein the one or more metabolite concentration values include concentrations of one or more metabolites for which a unit transport rate has been determined.
6. The method according to claim 1, in, Predicting one or more characteristics of the bioprocess includes predicting values of one or more critical quality attributes (CQAs) of the bioprocess using a prediction model that has been trained to predict CQAs using a set of predictor variables that includes one or more of the unit transport rates.
7. The method according to claim 1, in, The machine learning model is a regression model, or wherein, The machine learning model is selected from a linear regression model, a random forest regressor, an artificial neural network ANN, and combinations thereof; and / or wherein the machine learning model comprises a plurality of machine learning models, wherein each machine learning model has been trained to predict the unit transport rate of a separately selected subset of the one or more metabolites.
8. The method according to claim 1, in, The machine learning model has been trained to jointly predict the specific transport rates of the one or more metabolites at later maturity based at least in part on the values of the one or more process conditions of the biological process at one or more previous maturity.
9. The method according to claim 1, in, Obtaining values of one or more process conditions at one or more degrees of maturity includes obtaining values of the one or more process conditions at multiple degrees of maturity; and the machine learning model has been trained to predict the unit transport rate of the one or more metabolites at the latest maturity or later maturity among the multiple degrees of maturity based at least in part on the values of the one or more process conditions of the biological process at the multiple degrees of maturity.
10. The method according to claim 9, in, The machine learning model has been trained to predict the unit transport rate of the one or more metabolites at the latest or later maturity of the two different maturity levels based at least in part on the values of the one or more process conditions of the biological process at the two different maturity levels.
11. The method according to claim 1, in, The values of the one or more process conditions used as input to the machine learning model are associated with multiple maturity levels that are separated from each other by a maturity difference that is approximately equal to the maturity difference between the values used to train the machine learning model.
12. The method according to claim 1, in, Predicting one or more characteristics of the biological process includes determining values of one or more variables derived from the unit transport rate by using the unit transport rate to determine a concentration of the corresponding one or more metabolites at the later maturity.
13. The method according to claim 12, in, Determining the concentration of the corresponding one or more metabolites at the later maturity includes solving corresponding material balance equations.
14. The method according to claim 12, in, Determine the concentration of metabolite i at maturity k included in m i Integrate any of equations (4), (4a)-(4d) and (28) between the known previous maturity and the maturity k, where k is the maturity associated with the predicted unit transport rate: Among them, δ m,i is the unit transport rate of metabolite i by cells in culture, m i is the concentration of metabolite i in the reactor, pm i is the pseudo-metabolite concentration when metabolite i is provided in the feed stream using a batch feeding strategy, V is the volume of the culture in the bioreactor, m F,i is the concentration of metabolite i in the feed stream, m H,i is the concentration of metabolite i in the harvest stream, m B,i is the concentration of metabolite i in the effluent stream, x v is the viable cell density in the reactor, and F F 、F H and F B are the volumetric feed flow rate, volumetric harvest flow rate and volumetric discharge flow rate, respectively, and f ML,i (u,m,s) is the predicted value of the unit transport rate of metabolite i by cells in culture confirmed using the machine learning model, and ε can be set to 0 or a value reflecting the detection limit of the above metabolite.
15. The method according to claim 12, further comprising: include: The values of the one or more variables derived from the unit transport rate are determined by using the unit transport rate to determine the concentration of the corresponding one or more metabolites at the later maturity, and using one or more of the concentrations to determine the value of the biomass-related metric at the later maturity.
16. The method according to claim 15, in, Determining the value of the biomass-related metric at the later maturity includes solving a kinetic growth model.
17. The method of claim 16, further comprising using one or more of the metabolite concentrations and / or the biomass-related metrics as inputs to the machine learning model to predict unit transport rates at other maturity levels.
18. The method according to claim 16, further comprising: include: The impact of specific values of process parameters at later maturity are predicted by including the specific values in the material balance equation, the kinetic growth model and / or the machine learning model as inputs for predicting unit transport rates at other maturity levels.
19. A system for monitoring and / or controlling a biological process, the system include: at least one processor; as well as At least one non-transitory computer readable medium comprising instructions which, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of the preceding claims.
20. The system according to claim 19, in, The system further includes one or more of the following operably connected to the processor: a user interface, wherein the instructions further cause the processor to provide to the user interface for output to a user one or more of: values of the one or more unit transport rates or variables derived from the one or more unit transport rates, a result of the comparing step, and a signal indicating that the biological process has been determined to be operating normally or not operating normally; one or more biomass sensors; one or more metabolite sensors; one or more process condition sensors; and One or more effector devices.
Citation Information
Patent Citations
Bioreactor consumable units
US20160152936A1
Bioreactor vessels and associated bioreactor systems
WO2014020327A1
Systems and methods for automated biologic development determinations
WO2019070517A1
Predicting the metabolic condition of a cell culture
WO2019129891A1