Method for measuring process variables in a cell culture process - Patent Application 20070122997

A data-driven model using machine learning estimates key cell culture parameters in real-time, addressing the inefficiencies of offline sampling in mammalian cell cultures, improving process control and yield.

JP7763228B2Active Publication Date: 2025-10-31F HOFFMANN LA ROCHE & CO AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023215592
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-14
Filing Date
2023-12-21
Publication Date
2025-10-31
Estimated Expiration
2040-08-12

AI Technical Summary

Technical Problem

Existing mammalian cell culture processes face challenges in efficiently measuring critical process variables in real-time due to the reliance on offline sampling, which is labor-intensive, error-prone, and limited in small-scale systems, leading to suboptimal process control and product yield.

Method used

A data-driven model using a feature matrix of process variables like 'Time', 'CHT.PV', 'ACOT.PV', etc., generates online estimates of viable cell density, viable cell volume, glucose, and lactate concentrations through machine learning, enabling real-time monitoring without sampling.

Benefits of technology

This approach provides accurate, real-time monitoring of key culture parameters, enhancing process control and product yield in small-scale mammalian cell cultures, particularly for CHO cells expressing antibodies, by reducing time offsets and analytical effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763228000048
    Figure 0007763228000048
  • Figure 0007763228000049
    Figure 0007763228000049
  • Figure 0007763228000050
    Figure 0007763228000050
Patent Text Reader

Abstract

To provide a method for adjusting the glucose concentration to a target value during the mammalian cell cultivation.SOLUTION: The method according to the invention comprises: (a) determining current values at least for process variables in cultivation; (b) determining current glucose concentration in a cultivation medium using the measured values of (a) by means of a data-driven model for mammalian cell cultivation, which is generated using a feature matrix comprising the process variables; and (c) adding glucose until a target value is reached if the current glucose concentration of (b) is lower than a target value, and thus adjusting the glucose concentration to the target value. The method is characterized in that: the one or more process variables are selected from process variables viable cell density, viable cell volume, the glucose concentration in the cultivation medium, and lactate concentration in the cultivation medium; and that the method is carried out without sampling and exclusively using on-line measured values from this cultivation.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is in the field of mammalian cell culture processes. More specifically, the object of the present invention is a method for online measurement of a process target parameter based on historical online and offline values ​​of a set of process variables. [Background technology]

[0002] Technology background Quality and reproducibility are paramount requirements for the production of therapeutic drugs in the pharmaceutical industry. Therefore, empirical standards (GMP guidelines, Good Manufacturing Practices for Drugs and Quasi-drugs) are established to define target values, process limits, and deviations to meet these requirements. Recently, the U.S. Food and Drug Administration (FDA) initiated a PAT (Process Analytical Technology) initiative, calling on the pharmaceutical industry to develop a better understanding of the processes implemented to improve product quality [1]. In recent years, new technologies, such as computer-based models, have been used to advance understanding of cell culture processes, such as CHO cells, used to produce therapeutic proteins.

[0003] Bioreactors are most commonly used for culturing cells. In bioreactors, various process variables are recorded during the culture. These allow process monitoring and control and help maintain control of environmental conditions. A distinction is made between online and offline values. Both provide important information about the process. Online values ​​are collected by appropriate sensors used for direct online control. However, offline values ​​are measured by manual sampling with subsequent external analytical methods. Such offline parameters are, for example, viable cell density, glucose concentration, and lactate concentration. They can be used to evaluate the current culture conditions and, if necessary, intervene to regulate the process.

[0004] Analyzing samples requires increased manual labor, especially for high-throughput culture systems. These external methods can also lead to errors and device failure in some situations. To make the process more efficient and robust, it is possible to obtain information online using online values ​​already recorded during culture. In this way, existing measured parameters and their relationships can be analyzed to explain them using appropriate mathematical models of machine learning.

[0005] An artificial neural network (ANN) has been described for monitoring biomass in fed-batch processes [8], and Kroll et al. have described a model-based soft sensor for measuring CHO cell biomass subpopulations [9].

[0006] Hutter, S. et al. disclose glycosylation flux analysis of immunoglobulin G in Chinese hamster ovary perfused cell cultures (Process 6 (2018) 176). The authors describe a metabolic flux analysis-based approach to generate insights into glycosylation pathways. Hutter et al. focus on metabolic flux analysis in perfused cell culture experiments. Using only offline measured parameters, they fit a mechanistic (linear) model using a random forest model to rank the influence of input parameters on glycosylation outcomes. Thus, Hutter et al. disclose a statistical analysis based on offline data, i.e., a modeling tool for understanding the (biological) meaning of historical data, performed after culture. No predictive or online algorithms are disclosed.

[0007] The white paper "Biopharma PAT - Quality Attributes, Critical Process Parameters and Key Performance Indicators at the Bioreactor" (available at https: / / www.researchgate.net / publication / 326804832_Biopharma_PAT_-_Quality_Attributes_Criticcal_Process_Parameters_Key_Performance_Indicators_at_the_Bioreactor) provides a high-level overview of process analytical technology. The white paper discloses cultivation principles (e.g., batch, fed-batch, and perfusion, monitoring methods), where the impact of measurements such as dissolved oxygen is used to gain process understanding. Prediction of output parameters or machine learning techniques is not disclosed.

[0008] Rubin, J. et al. reported that pH deviations affect CHO cell culture performance and antibody N-linked glycosylation (Bioprocess. Biosys. Eng., 41 (2018) 1731-1741). The authors present a study on the effect of cell culture pH on antibody glycosylation using typical offline measurements of process parameters performed in a given culture, as well as the effect of pH fluctuations.

[0009] Downey, BJ et al. report a novel approach for using dielectric spectroscopy to predict viable cell volume (VCV) in early process development (Biotechnol. Prog. 30 (2014) 479-487).

[0010] Xiao, P. et al. reported metabolic characterization of CHO cell size expansion phase in fed-batch culture (Appl. Microbiol. Biotechnol. 101 (2017) 8101-8113).

[0011] Kroll, P. et al. report a soft sensor for monitoring biomass subpopulations in mammalian cell culture processes (Biotechnol. Lett. 39 (2017) 1667-1673). The authors used a turbidity physical sensor to measure viable cell count (VCC, equivalent to VCD) based on a linear model. Summary of the Invention

[0012] The present invention is based, at least in part, on the discovery that by selecting specific process variables from a historical data set, a useful data-driven model can be obtained that includes key parameters for the cultivation of CHO cells in real time, such as VCD (viable cell density), VCV (viable cell volume), glucose, and lactate. The method according to the present invention makes it possible to provide accurate online-like values ​​of target variables throughout the course of the cultivation without sampling.

[0013] A method for culturing CHO cells expressing an antibody and for measuring the viable cell density and / or viable cell volume and / or the glucose concentration in the culture medium and / or the lactate concentration in the culture medium during the culture using only online measurements from said culture by a model for the culturing of said CHO cells, characterized in that the method generates a model based on a feature matrix including the features "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV".

[0014] The characteristics indicate the following parameters: TIFF0007763228000001.tif101129

[0015] In one embodiment, the model is generated using a random forest method.

[0016] In one embodiment, the training dataset comprises at least 10 culture runs, preferably at least 60 culture runs.

[0017] In one embodiment, the model is derived using a training dataset that includes culture runs of mammalian cells expressing composite IgG, i.e., antibodies that include a morphology that differs from the wild-type, Y-shaped full-length antibody by including additional domains, e.g., one or more Fabs. In one embodiment, the training dataset also includes culture runs of mammalian cells expressing standard IgG, i.e., a Y-shaped, wild-type-like antibody with no domains added or deleted.

[0018] In one embodiment, approximately 80% of the dataset available for model generation is used as the training dataset, and the remaining dataset is used as the test dataset.

[0019] In one embodiment, a) randomly dividing the dataset available for modeling into a training dataset and a test dataset in an 80:20 ratio; b) forming a model; c) determining a mean and standard deviation for measuring the target parameter of a dataset from the training dataset and a mean and standard deviation for measuring the target parameter of a record from the test dataset; d) Steps a) to c) are repeated until parity is achieved for the split between the test and training datasets, ie, means and standard deviations within at most 10%, preferably at most 5%, of each other.

[0020] In one embodiment, missing data points in the dataset are filled in by interpolation.

[0021] In one embodiment, the data set includes at least 60 minutes of data points, preferably about every 5-10 minutes.

[0022] Specific Embodiments of the Invention 1. A method for measuring one or more process variables during mammalian cell culture, comprising: The process variable(s) may simply be i) by a data-driven model of mammalian cell culture generated using a feature matrix including the process variables "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV"; and ii) A method in which the measurement is made by using only online measurements from the culture.

[0023] 2. The method of embodiment 1, wherein the online measurements use at least the process variables of the culture: "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV".

[0024] 3. A method for adjusting glucose concentration to a target value during mammalian cell culture, comprising: a) measuring the current values ​​of at least the process variables "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV" of the culture; b) measuring the current glucose concentration in the culture medium using the values ​​measured in a) by a data-driven model for mammalian cell culture generated using a feature matrix including the process variables "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV", and c) if the current glucose concentration measured in b) is lower than the target value, adding glucose until the target value is reached, thereby adjusting the glucose concentration to the target value.

[0025] 4. The method according to any one of the preceding claims, characterized in that the process variables are selected from the process variables viable cell density, viable cell volume, glucose concentration in the culture medium, and lactate concentration in the culture medium.

[0026] 5. The method according to any one of embodiments 1 to 4, characterized in that the method is carried out without sampling and only online measured values ​​from the culture are used.

[0027] 6. The method according to any one of embodiments 1 to 5, wherein the data-driven model is generated by machine learning.

[0028] 7. The method according to any one of the preceding claims, wherein the data-driven model is generated using a method selected from the group comprising artificial neural networks and ensemble learning.

[0029] 8. The method according to any one of the preceding claims, wherein the data-driven model is generated using a random forest method.

[0030] 9. The method according to any one of embodiments 1 to 7, wherein the data-driven model is generated using an MLP Regressor method.

[0031] 10. The method according to any one of embodiments 1 to 7, wherein the data-driven model is generated using the XGBoost method.

[0032] 11. The method of any one of embodiments 1 to 10, wherein the data-driven model is generated through supervised learning.

[0033] 12. The method according to any one of embodiments 1 to 11, wherein the data-driven model is validated by cross-validation.

[0034] 13. The method of embodiment 12, wherein the cross-validation is 10-fold cross-validation.

[0035] 14. The method of any one of embodiments 1 to 13, wherein the data-driven model is generated using a training dataset comprising at least 10 culture runs.

[0036] 15. The method of embodiment 14, wherein the training dataset comprises at least 60 culture runs.

[0037] 16. The method according to any one of embodiments 1 to 15, characterized in that approximately 80% of the dataset available for model generation is used as a training dataset, and the remaining dataset is used as a test dataset.

[0038] 17. The method according to any one of embodiments 1 to 16, a) the dataset available for modeling is randomly split into training and testing datasets in a ratio of 70:30 to 80:20; b) forming a model; c) determining means and standard deviations for measuring the process variables of a dataset from the training dataset and determining means and standard deviations for measuring the process variables of a dataset from the test dataset; Repeating steps a) to c) until comparable means and standard deviations are achieved for the test and training data sets, i.e., within 10% of each other, preferably within 5% of each other, and the partitioning obtained in a) is different for each new run.

[0039] 18. The method of any one of embodiments 1 to 17, wherein the data sets used to generate the data-driven model each contain the same number of data points.

[0040] 19. The method of any one of embodiments 1 to 18, wherein the data points in the dataset used to generate the data-driven model are each for the same time point in culture.

[0041] 20. The method of any one of embodiments 1 to 19, wherein missing data points in the dataset are obtained by interpolation.

[0042] 21. The method according to embodiment 20, wherein missing data points of glucose concentration and / or viable cell volume are obtained by third-order polynomial fitting.

[0043] 22. The method according to embodiment 20 or 21, wherein missing data points of lactate concentration are obtained by univariate spline fitting.

[0044] 23. The method according to any one of embodiments 20 to 22, wherein the missing data points of viable cell density are obtained by Peleg fitting.

[0045] 24. The method of any one of embodiments 1 to 23, wherein each data set comprises a data point at least every 144 minutes.

[0046] 25. The method of any one of embodiments 1 to 24, wherein each data set includes a data point at least every 60 minutes.

[0047] 26. The method of any one of embodiments 1 to 25, wherein each data set includes a data point approximately every 5 to 10 minutes.

[0048] 27. The method according to any one of embodiments 1 to 26, characterized in that the mammalian cells are CHO cells.

[0049] 28. The method of any one of embodiments 1 to 27, wherein the mammalian cells are CHO-K1 cells.

[0050] 29. The method according to any one of embodiments 1 to 28, wherein the mammalian cells express and secrete a therapeutic protein.

[0051] 30. The method according to any one of embodiments 1 to 29, wherein the mammalian cells express and secrete the antibody.

[0052] 31. The method according to embodiment 30, wherein the antibody is a monoclonal antibody and / or a therapeutic antibody.

[0053] 32. The method according to embodiment 30 or 31, characterized in that the antibody is not a standard IgG antibody, i.e., is a wild-type four-chain full-length antibody, or is a composite antibody, i.e., an antibody comprising additional antibody and / or non-antibody domains compared to a standard antibody.

[0054] 33. The method of any one of embodiments 1 to 32, wherein the data-driven model is generated using a training dataset that includes only culture runs of composite IgG.

[0055] 34. The method according to any one of embodiments 1 to 33, wherein the data-driven model is generated using a training dataset that also includes a standard IgG culture run.

[0056] 35. The method according to any one of embodiments 1 to 34, characterized in that the mammalian cells express and secrete composite IgG or standard IgG.

[0057] 36. The method according to any one of embodiments 1 to 35, characterized in that the culture volume is 300 mL or less.

[0058] 37. The method according to any one of embodiments 1 to 36, characterized in that the culture volume is 250 mL or less, 200 mL or less, 100 mL or less, 75 mL or less, 200 to 250 mL, or 50 to 100 mL.

[0059] 38. The method according to any one of embodiments 1 to 37, wherein the culture is a fed-batch culture.

[0060] 39. The method according to any one of embodiments 1 to 38, wherein the culturing is carried out in a stirred tank reactor.

[0061] 40. The method according to any one of embodiments 1 to 39, characterized in that the water is gassed during the culture.

[0062] 41. The method according to any one of embodiments 1 to 40, characterized in that the culturing is carried out in a disposable bioreactor (SUB).

[0063] 42. The method according to any one of embodiments 1 to 41, characterized in that the mammalian cells are cultured in suspension or are mammalian cells growing in suspension.

[0064] 43. The method of any one of embodiments 1 to 42, wherein the data-driven model is generated by regression analysis.

[0065] 44. Use of viable cell volume as a target parameter in generating data-driven models to measure process variables for culturing mammalian cells in volumes of 300 mL or less.

[0066] 45. Use according to embodiment 44, characterized in that the process variables are selected from the group comprising the process variables viable cell density, viable cell volume, glucose concentration in the culture medium, and lactate concentration in the culture medium.

[0067] 46. ​​Use according to embodiment 44 or 45, characterized in that the culture is carried out without sampling.

[0068] 47. Use according to any one of embodiments 44 to 46, characterized in that the mammalian cells are CHO cells.

[0069] 48. Use according to any one of embodiments 44 to 47, characterized in that the mammalian cells are CHO-K1 cells.

[0070] 49. Use according to any one of embodiments 44 to 48, characterized in that the mammalian cells express and secrete a therapeutic protein.

[0071] 50. Use according to any one of embodiments 44 to 49, characterized in that the mammalian cells express and secrete the antibody.

[0072] 51. The use according to embodiment 50, characterized in that the antibody is a monoclonal antibody and / or a therapeutic antibody.

[0073] 52. Use according to embodiment 50 or 51, characterized in that the antibody is not a standard IgG antibody or is a composite antibody.

[0074] 53. Use according to any one of embodiments 44 to 52, characterized in that the data-driven model has been generated using a training dataset comprising only culture runs of complex IgG.

[0075] 54. Use according to any one of embodiments 44 to 53, characterized in that the data-driven model has been generated using a training dataset that also includes a culture run of a standard IgG.

[0076] 55. Use according to any one of embodiments 44 to 54, characterized in that the mammalian cells express and secrete composite IgG or standard IgG. DETAILED DESCRIPTION OF THE INVENTION

[0077] Detailed Description of Embodiments of the Invention To achieve high throughput of test cultures, especially for complex molecules and molecular formats, the size of the culture vessels must be reduced and the culture must be automated. The success of a culture depends on controlled process variables; only when optimal culture conditions are provided can the desired molecule be produced in high yield. Therefore, rapid and efficient control of the relevant process variables is required to set each process variable and maintain optimal culture conditions. This control is particularly necessary for small-scale parallel cultures, since each culture must be monitored separately. In particular, so-called offline process variables are problematic here because, on the one hand, the required sampling and separate analysis results are time-offset, i.e., the culture continues, and the process variable measured offline differs from the actual process variable. On the other hand, the number of sampling points is significantly smaller than the process variable available online, resulting in poor temporal control of this process variable.

[0078] It is therefore an object of the present invention to make process variables that cannot be measured online, but which are only measured offline, especially due to the size of the culture vessels used, available as well as process variables available online at the culture scale to be used in real time based on data-driven models.

[0079] To produce recombinant proteins, bioreactors are most often operated using a fed-batch process.[4] In addition to the fed-batch process, there are other modes of operation, such as batch processes and continuous culture modes.

[0080] The fed-batch or fed process is one of the partially open systems. The advantage of this process is that nutrients such as glucose, glutamine, and other amino acids can be added to the culture during the process. This can avoid the resulting substrate limitation and allow for longer process times. The substrate can be added continuously or in the form of one or more concentrated masses. To better control the inhibitory effects and accumulation of toxic by-products, an appropriate feeding strategy can be used. However, this requires thorough knowledge of the process as well as control of the process.

[0081] To provide and maintain optimal conditions during the cultivation of mammalian cells, such as CHO cells, bioreactors are used almost exclusively [2]. The bioreactors used are mostly stirred-tank reactors. The cultivation is carried out with cells growing in suspension, i.e., in suspension.

[0082] Aerobic mammalian cells, such as CHO cells, require oxygen to maintain their cellular metabolism. Cells are typically supplied with oxygen by submerged gassing of the culture broth. The dissolved oxygen concentration in the reactor is one of the most important parameters for aerobic cell cultivation. The concentration of dissolved oxygen in the medium is measured by several transport resistances. Diffusion transports oxygen from bubbles to cells, where it can ultimately be metabolized by the cells. While the transport mechanism can be measured using the oxygen transport rate (oxygen transfer rate, abbreviated as OTR), it has been shown that the oxygen consumption rate by the cells themselves can be measured using the oxygen consumption rate (oxygen uptake rate, abbreviated as OUR) [2]. Appropriate exhaust gas analysis can provide the data necessary to calculate OUR and OTR. Process variables such as temperature, pH, and dissolved oxygen concentration are monitored with appropriate sensors and are among the parameters controlled during cultivation. These process variables significantly affect the effective productivity of mammalian cell lines [3].

[0083] To shorten the development and setup time of bioreactors, research and development is increasingly focused on single-use technologies (single-use bioreactors; abbreviated as SUB). A major advantage of these systems is that they do not require complex cleaning processes and the necessary complex and expensive cleaning methods such as CIP (cleaning in place) and SIP (sterilization in place).

[0084] Automated high-throughput culture systems, such as the ambr250 system (automated microscale bioreactor), can help expedite drug development. Twelve single-use bioreactors, each with a volume of 250 mL, are available within this system. An automated liquid handler is used for pipetting and sampling. The operation is controlled by central processing software. To ensure a sterile environment during operation, the entire ambr250 system is placed under a laminar flow box.

[0085] Soft sensors have been increasingly used industrially over the past two decades to monitor process variables [6]. These variables can usually only be measured externally, i.e., offline, with high analytical effort. Especially when using small-scale, single-use systems, it is often not possible to install the necessary additional sensors (space and availability, or connectivity to disposable bioreactors, and in some cases, gamma irradiation is not possible, etc.). Therefore, there is a lack of continuous data on these process variables, i.e., important process variables that can be used for process monitoring and allow adjustment of process target parameters, especially at small cultivation scales. The term "soft sensor" combines the terms "software" and "sensor." The term "software" refers to the computer-assisted programming of models. The output of these models provides information about the cultivation, especially real-time values ​​of process variables that are unavailable due to the lack of respective physical sensors [5].

[0086] Basically, soft sensors can be divided into two classes: model-driven soft sensors and data-driven soft sensors.

[0087] Model-driven soft sensors are influenced by theoretical process models. They require detailed knowledge of the ongoing process and describe it using differential equations of state. This means that the dynamic behavior of the process must be represented using mechanistic models. Such models are primarily developed for the planning and design of manufacturing plants and focus on describing ideal equilibrium states.

[0088] Data-driven soft sensors (called black-box models) use models based on machine learning. These include empirical models that use historical data to represent process variable correlations. Biological processes are complex, and not all aspects of the metabolism of cultured mammalian cells are yet fully understood.

[0089] The application field of data-driven soft sensors within the pharmaceutical industry is broad: typically, to monitor and record cultures.

[0090] It has now been discovered by the present inventors that such historical data can be used to generate data-driven models for online estimation of offline process variables.

[0091] Process variables are primarily measured, i.e., made available, in real time. They can usually only be measured with difficulty and with increased analytical effort and associated time offsets. Furthermore, robust, long-term stable online sensor systems are not always available for online monitoring of some process variables, such as biomass or specific substrate and product concentrations [7]. These parameters contain important information about the cultivation process, but are only available at limited time points during cultivation, i.e., when samples are taken and analyzed offline.

[0092] In small systems such as the ambr250 system, it is not possible to measure certain process variables such as turbidity and / or conductivity due to the lack of a probe port. Furthermore, due to their design, some common probes require a relatively large amount of space, which is not available in these small volume systems.

[0093] Machine learning is the application of algorithms to represent the underlying structure of a dataset. Machine learning can be divided into two parts: supervised learning and unsupervised learning.

[0094] Supervised learning is used when a model is prepared to make predictions for future or unknown data based on training data. The training dataset is controlled because it already contains information about the desired output value. An example is spam email screening

[10] . Thus, the algorithm receives a dataset consisting of spam and non-spam messages and already contains information about spam / non-spam that goes through a learning phase. For new, unmarked emails, the algorithm tries to predict what type of message it is. This is the target variable in a classification (spam / non-spam), hence the term "classification."

[0095] Unsupervised learning attempts to capture relationships within a dataset without presenting a target variable to the algorithm. The focus is on exploring the underlying structure of the data in order to extract meaningful information from it. The simplest example of this group is clustering. This exploratory data analysis attempts to divide a dataset into meaningful subgroups without prior knowledge of the actual population membership.

[0096] When the target variable is continuous, it is called regression or regression analysis. The variables used to explain the regression model are called independent or explanatory variables. Based on this, an attempt is made to find a mathematical relationship between the input variables and the target parameter in order to be able to predict the outcome.

[0097] The method according to the invention uses supervised learning, where the target variable is represented by a regression.

[0098] The modeling can be schematically arranged in the steps of pre-processing, training, evaluation and estimation of the target variable.

[0099] Data preprocessing is necessary to ensure that the model can correctly interpret the information on which it is based. The dataset is prepared in the form of a feature matrix x, containing m features (columns) and n rows, thus representing the explanatory variables. Each row n contains the feature specification for a particular data point. TIFF0007763228000002.tif32128

[0100] The target variables are arranged in the vector y. Therefore, the feature matrix x (n) Each row of (n) Contains information about the associated value of

[0101] Statistical analysis is used to identify suitable features. Once suitable features have been identified and a corresponding feature matrix created, a subset (70-80% of the entire dataset) is made available for the model to learn from. This subset is called the training dataset.

[0102] Typical data preprocessing may involve providing the dataset to the model in a standardized format, so that the data for each feature is given the properties of a standard normal distribution with a mean of 0 and a standard deviation of 1. This increases the comparability of features with each other and allows the learning algorithm to achieve its optimal performance

[10] .

[0103] Learning is the core part of model building. During learning, the model attempts to understand and recognize relationships between data. Each model follows a mathematical formula with specific parameters, which are adapted during the training process to represent the relationships between data as well as possible.

[0104] Some models, such as neural networks, have other parameters that do not change during the learning process. These are called hyperparameters. They affect the complexity of the model or the speed of the learning process and are measured before the training process. There is no set method for selecting the correct hyperparameters. Therefore, different models are trained with different hyperparameters and then tested. Only then can we determine which model is the most suitable.

[0105] A randomized and raster-based algorithm is used to search for the optimal combination of hyperparameters. Each hyperparameter is represented by a list with different values. The model is trained with a grid search on all possible combinations from the respective lists. The required computational effort can be reduced by the randomized search. Various random parameter combinations can be used, and the computational effort can be measured in advance. In one embodiment, the model is first run through a randomized search for rough estimates of the hyperparameters, and then a grid search is performed to fine-tune the hyperparameters. The goal of the training is to train the model so that bias and variance are kept as low as possible.

[0106] A model often learns the relationships between training data better than subsequent predictions using unknown data sets. This behavior is called overfitting. Thus, the model memorizes the training data set and represents the relationships with new data with insufficient accuracy. Similar behavior can also be caused by over-dispersion. Here, the model uses too many input parameters for the data set being trained, resulting in a complex model that only fits this data set with high data variance. Therefore, the model cannot map the actual relationships and has learned the noise in the data.

[0107] On the other hand, if the model is not complex enough to react to changes in the test dataset, this is called undertraining: the bias is too large and the model can only inaccurately map relationships in the training data to the test data.

[0108] Already during learning, k-fold cross-validation of the training dataset offers the possibility of avoiding overfitting the model

[11] . The training dataset is divided into k subsets. Then, k-1 subsets are used to train the model, and the remaining subsets are used as the test dataset. This procedure is repeated k times. In this way, k models are trained and k approximations of the target variable are obtained.

[0109] Model performance estimate E i is generated for each run. For example, the mean squared deviation, a measure of error, is used as a performance approximation for regression. In practice, 10-fold cross-validation has proven to be a good compromise between bias and variance in most cases

[12] : TIFF0007763228000003.tif14128

[0110] Artificial neural networks (ANNs) were first described in 1943 by Warren McCulloch and Walter Pitts with a mathematical model of the neuron. In this way, information transmission in biological systems could be understood.

[13] Frank Rosenblatt then linked the McCulloch-Pitts model of the artificial neuron with a learning rule, and thus was able to describe the perceptron.

[14] Perceptrons still form the basis of ANNs.

[0111] A simple perceptron has n inputs x1,....,x n ∈IR, with weights w1,....,w n ∈IR. The output is denoted by o∈IR. The processing of the input signal with appropriate weighting is the propagation function (input function) σ, TIFF0007763228000004.tif13128This describes the network input of a neuron via an activation function φ: The output of the perceptron, o, is then measured. Various functions can be used for φ, which can cause the activation of the perceptron.

[0112] Thus, an activation function calculates how strongly a neuron is activated depending on the threshold and the network input

[15] . If several of these neurons are interconnected in an appropriate structure, they can map complex relationships between the input and output layers. The simplest form of such structural interconnections of simple neurons is a feedforward network. These are arranged in layers and consist of an input layer, an output layer, and, depending on the structure, several hidden layers.

[0113] In feedforward networks (so-called multilayer perceptrons), every neuron in one layer is connected to every other neuron in the next layer. Thus, these networks propagate the information content created through the network in a forward direction. Each neuron first weights its input signal with a randomly selected weight and adds a bias term. The output of this neuron corresponds to the sum of all weighted input data. Depending on the number of neurons in a layer and the number of hidden layers, the complexity of a neural network can be measured.

[0114] Multilayer feedforward networks with error feedback (backpropagation) are primarily used for supervised learning with ANNs

[16] .

[0115] Training such a neural network can be divided into three steps: · Step 1: Feed forward; ·Step 2: Error calculation; Step 3: Backpropagation

[0116] In the first step, an input is provided to the network's input layer, and this input is propagated layer by layer through the network until there is an output from the network. In the second step, the network's output is compared to an expected value, and the network error is calculated using an error function. Depending on the current weighting, each neuron in the hidden layer contributes to the calculated error to a different extent. In the third step, the error is propagated backward through the network, and the weights are adjusted depending on the contribution of each individual neuron's weight to the error. The goal of the backpropagation algorithm is to minimize the error, and it usually uses gradient descent

[17] . According to this method, the quadratic distance between the network's output and the expected output is calculated as the error function. TIFF0007763228000006.tif9128

[0117] To calculate the error contribution of each neuron weight, the weights w ij We must derive the error function Err from . Therefore, only continuous and differentiable activation functions can be used here

[17] . This determines the weight adjustment delta used in the next iteration. This relationship can be mathematically described as follows: TIFF0007763228000007.tif12128

[0118] The learning coefficient η, along with the number of iterations, is a hyperparameter established before training the model. The two steps are repeated until a maximum number of iterations or a defined error value is reached, achieving good results for unknown inputs.

[0119] Furthermore, the Random Forest (RF) algorithm can be used in machine learning for regression problems

[18] . RF learns through a large number of decision trees and therefore belongs to the category of ensemble learners. Decision trees can spread from a root (superordinate nodes, no predecessors). Each node divides the dataset into two groups based on its features. The successors of the root can be leaves (no successors) or nodes (at least one successor). Nodes and leaves are connected by edges. For regression problems

[19] · Each inner node (including the root) is assigned a feature; A specific value of the target variable to be predicted is assigned to each leaf of the decision tree; · For each edge, a relation is assigned to the threshold.

[0120] In a preferred embodiment, RF uses the bagging principle (bootstrap aggregation principle) proposed by Breiman

[18] to create appropriate training sets. The training sets are created by sampling from the entire training dataset with replacement. Some data may be selected multiple times, while other data are not selected as training data. The number of training sets always corresponds to the number of data in the entire training dataset. Each selected training set is used to make a decision using a decision tree (classifier). The decisions from all training sets are then averaged, and the final classification is determined by majority vote. Therefore, generating bootstrap samples reduces the correlation between individual classifiers. Furthermore, it can reduce the variance of individual classifiers, improving overall classification performance

[18] .

[0121] In a preferred embodiment, features are used to determine splits (node ​​divisions) during the creation of a decision tree, and the features make the most clear decisions regarding a random selection of features in the dataset. The selected split is not selected as the best split for all features, but only as the best split within a random selection of features. As a result of this randomization, bias (distortion, systematic error) of the decision tree increases during the creation process. Variance decreases because the average value of all decision trees included in the RF is formed. The reduction in variance is of greater value than the increase in bias, resulting in an increase in the accuracy of the model

[20] .

[0122] Furthermore, RF prediction largely prevents model overfitting, as the average of all individual decisions is always considered

[18] .

[0123] XGBoost (eXtreme Gradient Boosting) uses an ensemble of regression trees as the basis for model formation. It uses both the bagging principle already described and a specialized boosting technique to train the ensemble for the most accurate predictions possible. Simply put, boosting techniques can be viewed as a combination of gradient descent with many weak learners

[21] . These weak learners are typically less accurate than random guesses and are grouped together as strong learners in the process of creating the ensemble. A typical example of such weak learners is a simple regression tree with only one node. The principle of the boosting algorithm is to use these weak learners to select difficult-to-classify training data in order to learn from these poorly classified objects, thereby improving the performance of the ensemble. Due to the complexity of XGBoost, the algorithm is considered a black box. However, due to its scalability and speed of problem solving, the algorithm has been used very successfully in direct comparisons of different machine learning models

[22] .

[0124] The method implemented by XGBoost combines gradient descent and boosting techniques and is described below using the original paper by Tianqi Chen, “XGBoost: A Scalable Tree Boosting System”

[22] .

[0125] Using an ensemble of k decision trees, the model can be expressed as follows: TIFF0007763228000008.tif14128In formula, f k is the prediction of a single decision tree. Looking across all decision trees, we can make the following predictions: TIFF0007763228000009.tif14128, x i is the feature vector of the i-th data point. To train the model, we optimize the loss function L. For regression problems, the RMSE (Root Mean Square Error) is used: TIFF0007763228000010.tif14128

[0126] Regularization is a key part of preventing a model from overfitting: TIFF0007763228000011.tif14128 where T is the number of leaves and w 2 j is the achieved scoring of the j-th leaf. When the regularization and loss function are taken together, the basic objective function of the model can be formulated as: TIFF0007763228000012.tif5128, where the loss function determines the predictive power and regularization controls the complexity of the model. The objective function is optimized using gradient descent. Given TIFF0007763228000013.tif4128, the gradient descent is calculated at each iteration: TIFF0007763228000014.tif6128 and TIFF0007763228000015.tif4128 is changed along a downward gradient so that the objective function Obj is minimized.

[0127] To create a regression tree, internal nodes are split based on the features of the dataset. The resulting edges define the range of values ​​that allow splitting the dataset. Leaves in the regression tree are weighted, with the weights corresponding to the predicted values. The number of iterations indicates how often the bagging and boosting process is repeated. The XGBoost algorithm provides a very large list of hyperparameters that significantly contribute to the formation of a good model.

[0128] Regardless of the model used, correlation can be used to assess and represent the linear relationship between two variables. The Pearson correlation coefficient r (or r 2 ) provides a common measure for assessing this relationship. It is dimensionless and is calculated according to: TIFF0007763228000016.tif12128 and varies within the range -1≦r≦+1. The counter is the empirical covariance s xy represents the sum of the deviation products of two variables x and y from the mean corresponding to the x and s y The mean value of the quantity to be correlated is The linear relationship according to Fahrmeir

[23] can be interpreted as follows: r<0.5: Weak linear relationship 0.5≦r<0.8: Moderate linear relationship 0.8≦r: Strong linear relationship

[0129] Note that correlation analysis can only show linear relationships. Therefore, the Bravais-Pearson correlation coefficient is not suitable for representing nonlinear relationships. This may mean that there is a strong nonlinear dependence of variables, even though the correlation coefficient is 0.0≦r≦0.2.

[0130] Through mutual information, the nonlinear dependence of two random variables can be measured. It is used in information theory

[24] . Probability is used to represent the information content of a random variable compared to a second random variable. The basic formal relationship is: TIFF0007763228000018.tif12128

[0131] Therefore, this approach was developed by Kraskov et al. and Ross et al.

[25]

[26] so that it can be used to select appropriate continuous variables.

[0132] Appropriate metrics should be used to compare different models, which can help to express the accuracy with which the model can represent the target variable.

[0133] Measurement coefficient R 2 indicates what proportion of the variance in the target variable y is accounted for by the model. The measurement coefficients can be calculated according to: TIFF0007763228000019.tif14128 where, TIFF0007763228000020.tif4128 is the estimated value of the goal variable for the i-th example, and y i is the associated truth value. TIFF0007763228000021.tif4128 is the average. The measurement coefficient can take a value between 0 and 1. The closer the measurement coefficient is to 1, the better the model can fit the target variable.

[0134] Root mean square error (RMSE) is another statistical measure that can be used to measure model quality, where the root mean square of the actual distance to the estimate is calculated: TIFF0007763228000022.tif12128

[0135] By squaring the error and then forming the root, the RMSE can be interpreted as the standard deviation of the variable being estimated, where n is the number of observations and TIFF0007763228000023.tif4128 is the estimated value of the target variable y. The RMSE error is an absolute error value that can yield values ​​of different sizes depending on the target parameter being examined. Therefore, it makes sense to relate the RMSE to the mean. TIFF0007763228000024.tif11128

[0136] Therefore, the RMSE is the average true value TIFF0007763228000025.tif4128. This allows for a better assessment of the error for target variables of different sizes.

[0137] method According to the method of the present invention, it is possible to measure the timeline of cell growth, i.e., cell density, as well as the timeline of certain metabolites, particularly glucose and lactate, from online process variables in real time during cultivation, especially at a small cultivation scale. Thus, the method of the present invention can provide real-time values ​​of process variables that were previously unavailable in real time but only available offline. This is an improvement over conventional methods for measuring the timeline of cell growth and certain metabolites, particularly glucose and lactate, insofar as the method of the present invention does not require sampling from the culture medium.

[0138] In a preferred embodiment, the method of the invention is used to measure cell density, glucose concentration and lactate concentration from online process variables in fed-batch cultures of mammalian cells having a culture volume of 300 mL or less, and the method is performed without sampling, i.e., with feedback-controlled sampling.

[0139] The method of the invention allows the culture to be carried out on a small scale, i.e. with a culture volume of 300 mL or less, fully automatically, i.e. without sampling, and relevant process variables such as cell density cannot be measured online but only offline.

[0140] The methods of the present invention are particularly suitable for monitoring and controlling mammalian cell cultures on a small scale.

[0141] The method according to the present invention provides a method for measuring viable cell density, glucose and lactate concentrations as target parameters in CHO cell cultures using database-based soft sensors, where machine learning models are used to represent the various target variables.

[0142] The present invention is based, at least in part, on the discovery that the selection of process variables used to generate a model has a significant impact on the quality of the measured target process variable.

[0143] Furthermore, the present invention is based, at least in part, on the discovery that the type of division of an existing dataset, ie, its allocation into training and test datasets, affects the quality of the model.

[0144] Furthermore, the present invention is based, at least in part, on the discovery that the type of antibody being produced influences the selection of optimal target parameters.

[0145] The method of the present invention is described below using 155 exemplary data sets obtained from cultures on the ambr250 system. This should not be understood as a limitation of the teachings or the method of the present invention, but rather as an exemplary application of the teachings of the present invention. Other data sets generated in the same or different culture systems may also be used in the method of the present invention.

[0146] 155 data sets were analyzed and examined for relevant features. The target parameters were mapped using a corresponding interpolation strategy so that the selected model could provide values ​​for all target parameters at discrete time points. The models were evaluated in terms of error and model quality. The method based on it allowed providing a robust and accurate model for each target / process variable.

[0147] The molecular formats of the antibodies produced in the cultures in the dataset varied, and an overview of the various projects and molecular formats and the respective culture numbers is shown in Table 1 below.

[0148] (Table 1) Data summary TIFF0007763228000026.tif75143

[0149] Data related to the entire culture process, i.e., online parameter sets and associated date and time stamps, were used for each culture. The data density of various process values ​​varied over the timeline. These data density deviations can be attributed to the fact that the system recorded new data points for online parameters only when the measurements changed by a delta specifically defined for each measurement. To ensure that continuous process data was available and runs could be compared with each other, corresponding online parameters were interpolated for all missing time stamps.

[0150] For online process variables, keep in mind that if the data is heavily smoothed, the fluctuations in the measurements will be lost. However, this noise also represents process-related changes that are occurring and is included as information in the process values. Therefore, it is important not to over-smooth the process values ​​and to allow for process course changes even after interpolation.

[0151] The offline data contains a variable number of analytical values ​​depending on the number of samples in the culture (8-13). Each data set contains a date / time stamp for each data point and the associated analytical values ​​of the offline parameters.

[0152] Preprocessing the online and offline data by interpolation results in a data set containing the same number of data points for all process variables simultaneously, regardless of whether they are online or offline process variables. The analysis was based on the interpolated data set. If data points were simultaneously available at the same frequency for all online and offline process variables, such interpolation would not be necessary.

[0153] By pre-processing the available online and offline data, the different profiles of the individual process variables resulting from different measurement frequencies are standardized into a uniform time profile, i.e., a single timeline. Faulty values ​​caused by technical and process controls are identified, deselected or corrected, and existing time gaps are closed, so that all process variables in one data set for a culture and all data sets for all cultures are uniform in terms of time and number of process variables.

[0154] Data collected during the first and last 12 hours of the culture were not used to ensure that fluctuations in the measurement signal caused by turning the control on at the beginning of the culture or turning it off at the end of the culture would not falsify the model formation. In the specific example, this means that a time range from 0.5 days to 13.5 days was used. This ensures that changes in the process variables can only be attributed to processes in the cell culture. Interpolation of the online data was performed on the entire data set. Figure 1 shows an example of linear interpolation of the process value "AO.PV."

[0155] As shown in Figure 1, the online signal progression is well described by linear interpolation. At the beginning (<0.5 days>), it is possible to understand how the measurement value fluctuated when the control was started. Peaks (larger process value changes in a short time) can also be well mapped with this type of interpolation.

[0156] For offline data, the analytical values ​​(VCD, VCV, glucose, and lactate) obtained were fitted with three different interpolations. Figure 2 shows an example of the interpolation of VCD using various fitting methods.

[0157] Each measurement coefficient R 2 was calculated to evaluate the individual interpolations of the VCD. The univariate spline was used to evaluate the individual interpolations of the VCD, where the maximum R 2 Although it achieved a value of 0.01, there was a tendency toward significant overfitting. Thus, while the univariate spline accurately represents almost all measurements, it does not represent the typical growth curve of a biological system. On the other hand, the difference between Peleg fitting and polynomial fitting is smaller. However, Peleg fitting can represent the various growth stages of a biological system sufficiently well and is therefore used to interpolate the target variable in VCD

[27] .

[0158] Interpolation of lactate and glucose profiles showed that univariate splines fit offline data more satisfactorily. 2 The results showed that the profile for lactate was sufficiently well represented. Since polynomial fitting interpolates negative values ​​of lactate from day 10 onwards, a univariate spline interpolation was defined as the target vector y for lactate. However, for glucose, a polynomial fitting (third order) was used to fit the target variable (glucose:univariate spline (R 2 =0.999) and polynomial fitting (R 2 =0.958); lactate: univariate spline (R 2 =0.999) and polynomial fitting (R 2 =0.959).

[0159] Furthermore, data sets with too few offline data points for preprocessing (three or fewer) were no longer used for analysis. This was the case for two data sets. Thus, the entire interpolated and adjusted data set contained 153 cultures.

[0160] Because the interpolated data set with a maximum resolution of 5 minutes contains a large number of data points, to reduce the computational effort, the analysis was performed at a resolution of 1 / 10 day, which can be done using the JMP program.

[0161] Figure 3 shows the data set from Project 2 (12 cultures). As can be seen, the various interpolation methods (Peleg fitting, univariate spline and polynomial fitting) have very little effect on the strength of the correlation.

[0162] In the scatter plot in Figure 3, the online parameters are shown as features (lines). The columns represent different interpolations of the VCD. The ellipses in the scatter plot always contain 95% of the data. The closer the ellipses are, the stronger the linear relationship between the variables. The calculated Bravais-Pearson correlation coefficients are shown in Table 2 below.

[0163] (Table 2) Numerical values ​​of Pearson correlation coefficients for a sample dataset from Project B corresponding to the values ​​in Figure 3. TIFF0007763228000027.tif77141

[0164] Looking at the "O2.PV" values ​​as an example, the calculated coefficients for the interpolation are very close to each other (0.9547; 0.9490; 0.9490).

[0165] Therefore, a correlation analysis was performed on the entire data set. Table 3 below shows the Bravais-Pearson correlation coefficients thus determined.

[0166] (Table 3) Pearson correlation coefficient calculated for the entire data set (153 cultures), target variable VCD fitted with Peleg fitting. TIFF0007763228000028.tif97128

[0167] Compared to the correlation analysis on a single ambr250 run (see Table 3 and Figure 3 above), the correlation analysis showed a significantly weaker linear relationship across the entire dataset. Apart from the strength of the correlation, the analysis of the entire dataset also produced other online parameters as the best candidates. The independent variables were also found to be correlated with each other. Table 4 below shows some of the correlations between the parameters "O2.PV" and "N2.PV" and the other independent variables.

[0168] (Table 4) Correlations of the independent variables with each other, shown using the example of O2.PV and N2.PV, which had the highest correlation values ​​in the previously performed correlation analysis of Project B. TIFF0007763228000029.tif68133

[0169] If independent variables are correlated with each other, it means multicollinearity. As shown using the example of "O2.PV", there is a clear linear relationship between the two best correlation coefficients of "N2.PV" and "O2.PV" in Figure 3 and the remaining independent parameters.

[0170] Figure 4 shows the calculated information content (mutual information) for all features for the target variable VCD for the entire data set. Figure 4 shows that some of the available features have a high level of information about the VCD target variable. Thus, with respect to VCD, the mutual information may have the highest index for "Time," "CHT.PV," "ACOT.PV," "FED2T.PV," "GEW.PV," "CO2T.PV," "ACO.PV," "AO.PV," "O2.PV," "N2.PV," and "LGE.PV."

[0171] Based on the results of the information content calculation and correlation analysis, the best 10 process variables (CHT.PV, ACOT.PV, FED2T.PV, GEW.PV, CO2T.PV, ACO.PV, AO.PV, LGE.PV, O2.PV and N2.PV) are selected and the corresponding feature matrix X is created. The matrix contains interpolated data of the available data set. Features (f1...f 10) with a resolution of 5 minutes and the duration of the culture (hours) were selected as additional columns in the matrix: TIFF0007763228000030.tif66143

[0172] The division into training and test datasets was done in such a way that these were only datasets from the cultures of Project 2. The target variable "VCD" was divided according to the distribution of the feature matrix.

[0173] To confirm the quality of the obtained models, we calculated the relative frequency density of errors across the entire test dataset. Histograms of predictions across the test dataset for models measured using MLPRegressor (a), Random Forest (b), and XGBoost (c) for the target variable VCD are shown on the X-axis, while the relative frequency of errors is shown on the Y-axis, representing the error of the estimated VCD value compared to the predicted value. All three distributions exhibit a left-skewed trend, indicating that the VCD is underestimated. Furthermore, examination of all histograms shows that the estimates from all three models yielded comparable results. XGBoost exhibits the most uniform distribution of calculated errors, although it can be observed that the target variable is overestimated here.

[0174] For each model, the RMSE and R 2 was calculated based on the entire test data set. Both values ​​relate to a Peleg fit of the target variable VCD. The results for the three models are summarized in Table 5 below.

[0175] (Table 5) Approximation results of VCD for MLPRegressor, Random Forest and XGBoost. TIFF0007763228000031.tif21128

[0176] All models achieved comparable results in terms of RMSE and coefficient of measurement.

[0177] Examining some specific datasets (best models) measured with random forests, it is clear that it is not possible to accurately map the Pelegic fit of the VCD over the entire culture period (see Figure 5). The model in the upper part of the figure fails to correctly represent the relationship of the data to the VCD from day 5 onwards. The lower part of the figure shows the opposite behaviour: the model estimates the VCD too high from the very beginning and therefore fails to achieve a sufficiently accurate description of the VCD.

[0178] Surprisingly, it has been found that swapping features within a feature matrix that have significantly less information content, but still have measurable information content, can significantly increase the quality of the prediction.

[0179] It has been found that the expansion of the matrix by the features “CO2.PV”, “FED3T.PV”, “OUR”, and “PH.PV”, as well as the removal of the redundant feature “O2.PV” (gassing with N2 and O2) leads to an improvement in the quality of the predictions.

[0180] The improved feature matrix includes the following 14 features: "Time", "ACO.PV", "ACOT.PV", "AO.PV", "CHT.PV", "CO2.PV", "CO2T.PV", "FED2T.PV", "FED3T.PV", "GEW.PV", "PH.PV", "N2.PV", "LGE.PV" and "OUR.PV".

[0181] Furthermore, the selection or division into training and testing datasets was found to affect the quality of the predictions.

[0182] Comparing the already selected training and test datasets with respect to the target variable, the distribution of VCDs for the training dataset consisting of the cultures from Project 2 has a mean μ Train = 84.60, and σ Train = 48.62, while the test data set has a mean μ Test = 64.22, and σ Test It was found to have a standard deviation of 38.02.

[0183] When making predictions about cells expressing structurally distinct proteins, we find that obtaining a training dataset from only one project is disadvantageous, and that randomly distributing the training dataset across existing datasets is advantageous.

[0184] In this example, to distribute the datasets more evenly, 30 random numbers between 0 and 152 were generated (since there were 153 datasets). Each number represented one culture run. Random numbers were generated iteratively until a comparable mean and standard deviation for the split between the test and training datasets could be achieved in the trained model. The final split was determined by σ Train μ at =47.11 Train = 80.72 and σ Test μ at =48.70 Test = 80.11, which was used as the split ratio for the two data sets in further courses.

[0185] Thus, in one embodiment of the method according to the invention, an existing, preferably pre-processed, dataset is divided into a training dataset and a test dataset, where the training dataset is 70-80% of the total dataset (in this example 80%, thus 123 culture runs) and the test dataset contains 20-30% of the data of the total dataset (in this example 30 randomly selected cultures of the total dataset validated as above were available for model validation).

[0186] The models were then trained and tested on the new distribution of the augmented feature matrix and dataset. The strategy for optimizing the hyperparameters outlined above was retained for this. The corresponding histograms of the VCD estimates with the newly split training and test datasets show that the distribution of errors for all three models is significantly narrower, which can be attributed to more accurate estimates of the target parameters (Figure 6).

[0187] All three models may achieve error distributions that vary more clearly around the true value of the target variable (x-axis at 0). Again, the XGBoost histogram shows the most uniform error distribution. The Random Forest histogram shows small errors across the entire range. When comparing the two histograms (a) and (c) with each other, XGBoost often estimates the target value more accurately than the MLPRegressor. However, due to the width of the MLPRegressor's distribution, which results in a lower degree of error, we can infer that the accuracy is roughly the same for both models.

[0188] Table 6. Approximation results of VCD for MLPRegressor, Random Forest, and XGBoost, and the new distributions for test and training data. TIFF0007763228000032.tif21128

[0189] All three models can achieve closely related results. Figure 7 is an illustration of an estimate of the best model using individual cultures.

[0190] Therefore, a near-ideal approximation of the target variable based on a Peleg fit of the raw data is achieved. Looking at the entire test dataset, all models achieve an R 2 and good results in terms of RMSE can be achieved.

[0191] The glucose values ​​fitted by a third-order polynomial fit were used as the target parameters for estimating glucose concentrations. The feature matrix used for training contained the same features as the VCD. The same division into training and test datasets was also used.

[0192] Similar to VCD, the histograms show comparable results in terms of error. Again, XGBoost can result in small errors between the actual and estimated values ​​in most cases. The Random Forest histogram also shows small errors between the interpolated and estimated values ​​of the target variable, which are uniformly distributed around the actual value of glucose. The MLP Regressor shows the largest error compared to the other two histograms.

[0193] Table 7. Estimation results of glucose values ​​for MLPRegressor, Random Forest and XGBoost. TIFF0007763228000033.tif21128

[0194] Figure 9 shows two typical cultures obtained using random forests. The dependent variable was well described with a coefficient of determination of 0.93.

[0195] For the estimation of lactate concentration, lactate values ​​fitted with a univariate spline method were used as the target parameter. The feature matrix used for training contained the same features as VCD and glucose. The same division into training and test datasets was also used. The histogram shows the different results in terms of error (Figure 11).

[0196] Considering the histogram of the MLPRegressor, it is not possible to approximate with small errors as frequently as the other two models. On the other hand, Random Forest and XGBoost have very narrow distributions. While it seems possible to make very good predictions with almost no error for some estimates of the target variable, these quickly lead to larger errors in the entire test dataset. The neural network has the most uniform error distribution here.

[0197] Table 8 below shows the RMSE and R for all models. 2 The results of lactic acid evaluation are shown below. (Table 8) Estimation results of lactate values ​​for MLPRegressor, Random Forest and XGBoost. TIFF0007763228000034.tif21128

[0198] Figure 12 shows XGBoost predictions for lactate for an example culture from the test dataset. A near-ideal description of the fitted lactate course can be seen in the upper subimage. In the lower image, the course is represented by an R2 of 0.98.

[0199] For validation, we first studied which model could most efficiently represent the interrelationships of features on the test dataset. To this end, the model was initially provided with only 10 datasets for training. As the process progressed, the number of datasets was increased by 10 each. This resulted in 12 training sessions in which the model received between 10 and 120 datasets. After each training session, the target variable was estimated based on the test dataset. The RMSE for each was calculated. The test dataset also consisted of 30 randomly selected datasets validated as described above. VCD was selected as the target variable. This resulted in the learning response depicted in Figure 13.

[0200] As shown in Figure 13, both Random Forest and XGBoost can achieve smaller errors in predicting the test dataset than neural networks with fewer datasets. However, this effect seems to decrease as the number of training datasets increases, so that comparable errors compared to the other two models can be achieved from about 80 datasets onwards. Up to 120 datasets, Random Forest achieves the lowest RMSE. However, the errors of all models are within a very narrow range.

[0201] A detailed evaluation of the model's estimates for predicting VCD for 30 cultures of the test dataset was performed. Despite good results (histograms, coefficients of measurement, RMSE) across the entire dataset, some predictions were found to still exhibit significantly large deviations. Figure 14 shows a culture run where the estimated VCD course clearly outperforms the actual distribution.

[0202] The cultures from projects 1 and 3 were more likely to be observed to have poor approximation accuracy. In both projects, the cultured cells produced complex molecular formats.

[0203] It was found that the VCDs of IgG-based formats (projects 2 and 4), which have or largely retain the characteristic Y-shape of natural IgG antibodies, are on average higher than those of cells with complex molecular formats as target products (projects 1 and 3), and the calculated cell diameters have higher values ​​than those of projects with complex molecular formats.

[0204] Figure 15 shows the average cell diameters for each sample, grouped by Y-shaped IgG (IgG, Projects 2 and 4) and complex IgG (Conjugate, Projects 1 and 3), as well as the standard deviation in the form of a box plot. The figure shows that the green box plots (complex protein formats; left at each time point) are located above the blue box plots (Y-shaped IgG antibodies; right at each time point). At the beginning of the culture period, both molecular formats are still relatively close together. Cells with the complex molecular format as the target product only become significantly larger as the culture period progresses. In contrast, cells with the standard antibody grow larger until day 7, after which the cell diameter does not increase further.

[0205] The relationship between higher VCD and smaller cell diameter for IgG formats, and smaller VCD and larger cells in complex protein formats, was found to not allow accurate prediction of VCD.

[0206] It was found that viable cell volume (VCV) is a more suitable target variable than VCD not only for cultures producing composite antibody formats but also for cultures producing Y-shaped IgG antibodies.

[0207] VCV is calculated using the following formula: TIFF0007763228000035.tif8128

[0208] Therefore, VCV is a better approximation for describing the living biomass in a culture than VCD.

[0209] Because the VCV calculations, like all other offline parameters, only included the time of sampling, the new target parameters were fitted with a third-order polynomial fit. The model was then trained and evaluated for the new target size as previously described for the other target parameters above.

[0210] Individual models were evaluated using RMSE and coefficient of measurement. In summary, the best model with 14 features achieved the following results:

[0211] Table 9. Comparison of RMSE and coefficient of determination of the best models for the target variable VCD TIFF0007763228000036.tif21128

[0212] For the target variable VCV, the calculated errors and measurement coefficients of the individual models are summarized in Table 10 below.

[0213] Table 10: Comparison of RMSE and coefficient of determination of the best model for the target variable VCV TIFF0007763228000037.tif21128

[0214] By using the target variable VCV instead of VCD, all models were able to achieve a measurement coefficient above 0.9. The model improvement was due to a lower RMSE and a higher R 2 Both values ​​are acceptable.

[0215] To demonstrate the improved results in comparing viable cell density and cell volume, scatter plots were generated showing both the estimates for the entire training set and the test data set. Random forests estimate the best results for VCD and VCV. The two scatter plots are shown in Figure 16.

[0216] Comparing the two scatterplots, we see that VCV predictions are close to ideal approximations and have significantly less spread across the test and training datasets than VCD predictions. When considering only the training data (blue dots), the model learns a better relationship of features to cell volume than to viable cell density. These features therefore enable a more accurate approximation of cell volume across the entire test dataset for all trained models.

[0217] The extent to which splitting the antibodies into different groups and using only a limited dataset for training the method affects quality was investigated as follows.

[0218] If all four projects are considered separately in terms of the course of the target parameter VCV, the box plot shown in Figure 17 is obtained. As can be seen, the VCV of project 4 behaves between projects 1 and 3 on the one hand and project 2 on the other. This means that the datasets from projects 1, 3, and 4 can also be classified as complex IgG antibody formats (classification 2). Therefore, the calculations were repeated with this classification. Various combinations of training and test datasets were also tested. The results are shown in Table 11 and Figures 18 and 19.

[0219] (Table 11) RMSE for various combinations of training and testing datasets. TIFF0007763228000038.tif165143

[0220] It has been shown that, among the various combinations, prediction using the random forest method achieved the best results, i.e., the lowest RMSE.

[0221] The RMSE showed a significant improvement (decrease) in all combinations of the training or test datasets when VCV was used as the target parameter compared to VCD.

[0222] Various combinations of training and test datasets showed that the selection of datasets according to molecular format affected the RMSE of the target parameters. When training a model using a dataset with a standard format and estimating VCD or VCV with a complex format, this combination achieved the highest RMSE. Training using a dataset with a complex molecular format and predicting VCD or VCV resulted in a smaller RMSE. When a mixed dataset was used for standard Y-IgG and complex molecular formats, the smallest RMSE could be achieved.

[0223] Furthermore, the model was evaluated on the approximation of the training data set and the test data set to check whether the already trained model was overfitted. The trained model of the target variable VCV was estimated on the test data set and the training data set. The approximation value was evaluated according to RMSE, and then the difference between the test data set and the training data set was shown in the form of a bar graph (Figure 20).

[0224] Figure 20 shows that the MLPRegressor achieves a lower error on the test dataset than on the training dataset. Therefore, the calculated difference is negative. Random Forest and XGBoost make larger errors on the test dataset, which causes the difference shown here to be positive. Therefore, both decision tree-based models tend to overfit.

[0225] Conventional technology Prior art uses parameters such as glucose, lactate, ammonia, and VCD (all of which are offline parameters) as input variables for random forest regression analysis to explain the dynamic behavior of intracellular activity, but does not use them to predict or model offline parameters.

[0226] In contrast to the prior art, in the present invention the parameters used in the machine learning model are exclusively online parameters (used to control the fermentation conditions).

[0227] Thus, the present invention utilizes typical online measured parameters that have been generated through culture and statistical models to estimate parameters such as VCV, glucose, etc., without the need for additional sensors or sampling.

[0228] Summary and Overview By interpolating existing online and offline culture datasets, a standardized, uniform dataset was obtained, which was used for model generation to predict target parameters that were only available offline.

[0229] For offline data that were considered as target variables for further courses, it was essential to find an interpolation that could representatively describe the course of the respective target parameter. Since the viable cell density is related to the growth process of biological systems, traditional interpolations such as polynomial fitting or univariate spline fitting often describe this target parameter with insufficient accuracy. An incorrect extrapolation leads to an incorrect description of the target variable. The chosen interpolation allows for a more accurate and accurate analysis of R. 2 Although comparable results have been obtained for , the selected interpolation by M. Peleg

[27] may best describe the growth process of cell culture processes. The background to the interpolation strategy is the combination of a continuous logistic equation to describe cell growth and a mirrored logistic equation (Fermi equation) to describe cell death behavior.

[0230] The results of the correlation analysis are only slightly affected by the choice of interpolation strategy.

[0231] The accuracy of the VCD target variable estimate can be increased by the split ratio used to fit the dataset to the training and test datasets. To this end, a validation dataset was selected with respect to the distribution of the target variable so that the mean and standard deviation were as small as possible relative to each other. The goal was not to artificially generate a more suitable dataset for prediction. Rather, it was assumed that the previously generated test dataset could not be used to describe the entire dataset with sufficient accuracy. This gives rise to the corresponding method known as cross-validation.

[0232] Calculation of cell volume and its associated relationship to cell size may represent a better approximation of biomass than VCD, and therefore VCV was obtained as the new target parameter.

[0233] Calculated cell volume, as an approximation of biomass description, provided more information about process characteristics than previously used viable cell densities of cultures measured by sample analysis. The average volume of a cell culture can be determined from the average diameter of the measured cells. Cell size, especially that of cells carrying complex target molecules as products, can be shown to increase continuously with increasing culture time. However, viable cell density cannot map this relationship. Finally, the metabolic activity of cultured cells may be better described by viable cell volume than by viable cell density.

[0234] To measure target parameters in real time, estimations should be made at predetermined intervals, for example 10 minutes. For CHO cells, which have a doubling time of approximately 24 hours, this interval is an acceptable resolution.

[0235] [The present invention 1001] 1. A method for regulating glucose concentration to a target value during mammalian cell culture, comprising: (a) measuring, during incubation, current values ​​of at least the process variables "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV"; (b) measuring the current glucose concentration in the culture medium using the measurements of (a) by a data-driven model for mammalian cell culture generated using a feature matrix including the process variables "Time", "CHT.PV", "ACOT.PV", "FED2T.PV", "GEW.PV", "CO2T.PV", "ACO.PV", "AO.PV", "N2.PV", "LGE.PV", "CO2.PV", "FED3T.PV", "OUR", and "PH.PV"; and (c) if the current glucose concentration of (b) is lower than the target value, adding glucose until the target value is reached, thereby adjusting the glucose concentration to the target value. A method comprising: [The present invention 1002] 1001. The method of claim 1001, wherein said process variables are selected from the process variables viable cell density, viable cell volume, glucose concentration in the culture medium, and lactate concentration in the culture medium. [The present invention 1003] 1001 or 1002, characterized in that the method is performed without sampling, using only online measurements from this culture. [The present invention 1004] The method according to any one of claims 1001 to 1003, wherein the data-driven model is generated by machine learning. [The present invention 1005] The method according to any one of claims 1001 to 1004, wherein the data-driven model is generated using a random forest method. [The present invention 1006] 1006. The method of any one of claims 1001 to 1005, wherein the data-driven model is generated using a training dataset comprising at least 10 culture runs. [The present invention 1007] (a) The dataset available for modeling is randomly divided into training and testing datasets in a ratio of 70:30 to 80:20; (b) a model is generated; (c) means and standard deviations for measuring process variables of a dataset are measured from the training dataset, and means and standard deviations for measuring process variables of a dataset are measured from the test dataset; (d) steps (a)-(c) are repeated until comparable means and standard deviations are achieved for the partitions between the test and training data sets, the partitions obtained under (a) being different for each new run; The method according to any one of claims 1001 to 1006 of the present invention, [The present invention 1008] 8. The method of any one of claims 1001 to 1007, wherein the data sets used to generate the data-driven model each contain the same number of data points. [The present invention 1009] 9. The method of any one of claims 1001 to 1008, wherein the data points in the dataset used to generate the data-driven model are each for the same incubation time. [The present invention 1010] 1009. The method of any of claims 1001 to 1009, wherein missing data points in the data set are completed by interpolation. [The present invention 1011] The method of the present invention 1010, characterized in that missing data points for glucose concentration and / or viable cell volume can be obtained by cubic polynomial fitting, missing data points for lactate concentration can be obtained by univariate spline fitting, and / or missing data points for viable cell density can be obtained by Peleg fitting. [The present invention 1012] 1012. The method of any one of claims 1001 to 1011, wherein the dataset comprises a data point at least every 144 minutes. [The present invention 1013] The method of any one of claims 1001 to 1012, wherein the mammalian cells are CHO-K1 cells. [The present invention 1014] The method of any one of claims 1001 to 1013, wherein the mammalian cell expresses and secretes an antibody. [The present invention 1015] 1015. The method of any one of claims 1001 to 1014, wherein the data-driven model is generated using a training dataset comprising a composite IgG culture run and a standard IgG culture run. [The present invention 1016] The method according to any one of claims 1001 to 1015, wherein the culture volume is 300 mL or less. The following examples and figures serve only to illustrate the present invention. The scope of protection is defined by the pending claims. However, modifications to the disclosed embodiments can be made without departing from the principles according to the invention. [Brief explanation of the drawings]

[0236] [Figure 1] Linearly interpolated measurements using the ACO.PV example range from day 0.5 to day 13.5. [Figure 2] Interpolated measurement curve of viable cell density for a typical culture. Interpolation and measurement coefficients: Peleg fitting (R2 = 0.957), univariate spline (R2 = 0.998), and cubic polyfit (R2 = 0.864). [Figure 3] An example correlation analysis of the ambr250 dataset performed from Project 2. Comparison of correlation coefficients for different interpolation strategies. The figure shows a scatter plot of the individual online parameters of the VCD. [Figure 4] Information content calculated according to mutual information for the target variable VCD for the entire data set. [Figure 5] Random Forest VCD estimates for two separate runs. In the top of the figure, an R2 estimate of 0.20317 was achieved. In the bottom of the figure, an R2 estimate of 0.54896 was achieved. [Figure 6] Histograms of predictions of the newly created test dataset for the models MLPRegressor (a), Random Forest (b), and XGBoost (c) for the target variable "VCD." The error of the fitted VCD values ​​against the predicted values ​​is shown on the X-axis. The Y-axis shows the relative frequency of the error. [Figure 7] Random forest VCD estimates for two example runs of the test dataset. In the top of the figure, an R2 estimate of 0.98944 was achieved. In the bottom of the figure, an R2 estimate of 0.99837 was achieved. [Figure 8] Information content calculated for the entire dataset according to mutual information for the target variable glucose. [Figure 9] Glucose estimation from random forest for two example runs of the test dataset. In the top of the figure, an R2 estimate of 0.99 was achieved. In the bottom of the figure, an R2 estimate of 0.97 was achieved. [Figure 10] Information content calculated for the entire dataset according to mutual information for the target variable lactate. [Figure 11] Histograms of predictions for the test dataset for MLPRegressor (a), Random Forest (b), and XGBoost (c) for the target variable lactate. The error of lactate values ​​added to the predicted value is shown on the x-axis. The y-axis shows the relative frequency of the error. [Figure 12] XGBoost lactate estimation for two example runs of the test dataset. In the top of the figure, an R2 estimate of 0.99 was achieved. In the bottom of the figure, an R2 estimate of 0.98 was achieved. [Figure 13] RMSE calculated for MLPRegressor, Random Forest and XGBoost with different numbers of training datasets. [Figure 14] Random forest VCD estimation for monocultures. Peleg fitting of VCD is shown in blue, and VCD estimates are shown in orange. [Figure 15] Representation of the average diameter of each sample over the entire culture period. Projects 1 and 3 have a complex molecular format as the product (shown here in blue, left). Projects 2 and 4 have a Y-shaped IgG format as the product of interest (shown here in green, right). Box plots include means; units are shown normalized. [Figure 16]Left part of the figure: Random forest approximation for VCD. In red, the test dataset approximation to the true value. In blue, the training dataset approximation to the true value. Ideal approximations for the test and training datasets are shown in black. Right part of the figure: Random forest approximation for VCV. In red, the test dataset approximation to the true value. In blue, the training dataset approximation to the true value. Ideal approximations for the test and training datasets are shown in black. [Figure 17] Representation of the mean diameter of each sample over the entire culture period for each project. Project 1 = purple, Project 2 = red, Project 3 = green, Project 4 = blue. Box plots include means. [Figure 18] Comparison of VCD / VCV using random forest model (best model). [Figure 19] Behavior of RMSE considering all models (MLPRegressor, Random Forest, XGBoost) with training dataset depending on the target parameter VCV. [Figure 20] Bar chart of the difference in RMSE between the test and training datasets, which is the best model for the target variable VCV.

[0237] References TIFF0007763228000039.tif189142TIFF0007763228000040.tif43137

[0238] List of Abbreviations TIFF0007763228000041.tif113128

[0239] List of symbols TIFF0007763228000042.tif34128TIFF0007763228000043.tif232128TIFF0007763228000044.tif118148 [Example]

[0240] material software: For the entire work, the programming language Python was used in the Spyder development environment. The implementation was carried out using object-oriented programming. Several classes were written that implement individual tasks within the project. TIFF0007763228000045.tif50137

[0241] method Data Processing The total dataset included 155 culture runs, which were divided into online and offline data. Data processing was performed using Spyder, a Python programming language. The data were available as CSV files. The data were read with the "csv" program library, which allows for quick and easy data reading and conversion into new data structures within the development environment. A "PIFileParser" class for online data and an "OfflineDataParser" class for offline data were implemented.

[0242] interpolation As the data was available at different data densities, it had to be interpolated accordingly. For this purpose, linear interpolation and interpolation using the moving average method were used. Both functions were implemented in the "scipy" library: "linear-interpolation-interval 1d" and "moving-average-convolve". This ensured that the interpolated value was always between the two raw measurements. Thus, the interpolation was always within the range of the natural fluctuation of the measured signal of the process variable. As each process variable had a different timestamp within the file, a separate CSV file had to be created. The "timeline mapping", which included all start and end times of each incubation, was created by a separate database query. Three different intervals were chosen for the data resolution: -Timestamp of the relevant sampling time for offline data 1 / 10 days ·5 minutes

[0243] Linear interpolation was not applied to the offline data due to the rather low data density and nonlinear data progression. Here, three different interpolation strategies were used for fitting: Peregrine fitting Polynomial fitting ·spline

[0244] Interpolation by M. Peleg can map biological growth through additional functional terms and therefore fully explain the growth process

[27] . Therefore, the raw data for viable cell density were fitted with all three interpolations. For glucose and lactate, we did not assume biological behavior, so we used polynomial and spline methods for interpolation. The online and offline datasets were merged at different intervals and saved as CSV files for each culture. Correlation analyses were then performed based on these datasets.

[0245] Correlation analysis The correlation analysis was performed using JMP®, which allows statistical analysis to be applied to the data set. Multivariate statistics of the online data (features) for each target variable (lactate, glucose, VCD, VCV) was applied. The data were analyzed for both statistical significance and linear relationships in the description of the target variables. The correlation analysis shows the linear relationship between the independent and dependent variables in the form of a Bravais-Pearson correlation coefficient.

[0246] mutual information Another method for identifying suitable features is in the form of mutual information. Mutual information measures the information content contained in an independent variable X to describe a target variable Y. Dependence was calculated and implemented using sklearn via "mutual information regression." Based on the size of the dataset with a 5-minute resolution, the information content was calculated for each culture separately, and then an average of the values ​​obtained across all cultures was generated.

[0247] Creating the feature matrix / the resulting vector The creation of the feature matrix was based on the results of correlation analysis and statistical evaluation based on the information content. It can be represented as a matrix, with one feature per column and one time point for each version of the feature. The feature matrix was saved as a Panda DataFrame. Therefore, a suitable file format was available for training and testing the model.

[0248] Modeling and Evaluation With the help of the correlation analysis, a separate data set was created for each target variable. To train the model, it was necessary to separate the feature matrix into a training data set and a test data set. Subsequent use for online prediction required the reservation of a complete validation project. The training data set included 80% of the entire data set, thus 123 culture runs.

[0249] Since all the target variables were constant target parameters, only regressors were used as models. Several hyperparameters were available for the models, which varied from model to model. Therefore, training the models served to adapt the hyperparameters to map the target variables as accurately as possible.

[0250] For the training itself, the entire feature matrix was normalized with the standard scalar from the Scikit-Learning library.

[0251] Hyperparameter optimization Hyperparameters were optimized using randomized search (RandomizedSearchCV) and grid-based search (GridSearchCV) from the Scikit-Learn library. All models were trained using the randomized search from the Scikit-Learning library in combination with 10-fold cross-validation of the training dataset. Various regions of hyperparameters were explored for the minimum RMSE. The randomized search was performed 30 times. Thus, a different set of randomly selected hyperparameters was used in each iteration. The hyperparameters of the 10 models with the minimum RMSE were output. The hyperparameters of the grid search were then further refined based on the hyperparameters from the randomized search. The grid search was then performed again using 10-fold cross-validation of the dataset. The model with the minimum error (minimum RMSE) was saved and then used to estimate the target variable from the test dataset.

[0252] Multilayer Perceptron We implemented a multi-layer perceptron (MLP) using the Scikit-Learning library. The following list contains the hyperparameters used to train the model: Number of neurons in the input layer Number of neurons in the hidden layer Solver algorithms for setting weights (adam, lbfgs, sgd) ·Activation functions (identity, logistic, tanh, relu) Learning Rate Maximum number of repetitions

[0253] Random Forest Random forest was also implemented using the Scikit-Learn library. The following candidates were available as hyperparameters in this optimization: Number of decision trees Number of features per decision tree Maximum depth of decision tree The minimum number of datasets to create a new node · Method for selecting the dataset (bootstrap = true / false)

[0254] XGBoost The XGBoost algorithm was integrated into the project structure via the XGBoost library, covering the following hyperparameter spaces: The number of regression trees in the ensemble Maximum depth of decision tree Learning rate η Number of datasets per decision tree -Minimum weight of child nodes in decision tree γ error evaluation As the hyperparameters used.

[0255] Model evaluation Model evaluation was primarily performed by displaying error histograms, which show the errors (residuals) that the model has in predicting the test data set relative to the actual values ​​of the target parameters.

[0256] The RMSE was calculated for the estimation accuracy of the target parameters and compared with the mean value of the target parameters.

[0257] To check the model for overfitting, the RMSE was calculated for the entire training and test datasets, and the difference between the two errors was used as an indicator of model overfitting. Overfitting = RMSE 試験 -RMSE 訓練

[0258] Measurement coefficients for the entire test dataset and for each culture considered individually were used to further describe the quality of the model.

[0259] Example 1 Ambr250-Culture 155 datasets based on cultivation in the ambr250 system were collected. The eukaryotic cells used were CHO cells extracellularly expressing the target molecule. The cultivation was carried out using the fed-batch method. The ambr system used allows for 12 simultaneous cultivations. The cultivation time for the main cultivation was 13-14 days. A single-use bioreactor (250 mL) provided the reaction space for this. Pre-cultures were carried out in shake flasks and lasted for 3 weeks. The starting conditions in terms of cell volume and number at inoculation were comparable for each reactor. The medium used was only a chemically defined medium. Only one medium batch was used per cultivation.

[0260] Several process variables were available to provide optimal culture conditions within this system. The controlled parameters were pH, temperature, and dissolved oxygen concentration in the medium. The table below contains a complete list of all process variables used in this work.

[0261] (Table 12) Online measurement parameters. TIFF0007763228000046.tif97128

[0262] All measured variables were recorded over the entire culture period by the so-called PI system, which only includes variables measured online.

[0263] The parameters listed here were available to monitor optimal culture conditions. For each reactor, exhaust gas analysis from BlueSens was also available. This detects the O2 and CO2 content in the exhaust gas stream from the bioreactor, thereby providing another important component in process control. These two measured variables of the exhaust gas stream can be used to measure OUR and OTR.

[0264] Samples were taken daily during the culture and then analyzed for various concentrations of metabolites and product titers using a Cedex Bio HAT® (Roche Diagnostics GmbH, Mannheim, Germany).

[0265] Additionally, cell count measurements were performed. This measurement provides information on viable cell density, total cell density, viability, aggregation rate, and cell diameter. These parameters can be used to infer the growth behavior of the cultures. Offline size was measured with a Cedex HiRes® (oche Diagnostics GmbH, Mannheim, Germany) cell counter. The error from these cell counting and cell analysis systems is in the range of 10%. All offline measurements used are listed in the table below.

[0266] (Table 13) Variables measured offline. TIFF0007763228000047.tif39128

Claims

1. 1. A method for determining viable cell density or viable cell volume in a culture medium for mammalian cell culture during mammalian cell culturing, the method comprising: (a) During the incubation, at least the process variables "time", "temperature of the cooling element (°C)" ("CHT.PV"), "CO in the exhaust gas stream" 2 Total (vol %)" ("ACOT.PV"), "Cumulative Feed 2 (ml)" ("FED2T.PV"), "Fermenter Weight (g)" ("GEW.PV"), "CO 2 Cumulative inflow volume (ml)" ("CO2T.PV"), "CO in exhaust gas flow 2 Concentration (volume %)" ("ACO.PV"), "O in the exhaust gas stream 2 Concentration (volume%)" ("AO.PV"), "N 2 Inflow (ml / min)" ("N2.PV"), "Cumulative amount of base added (g)" ("LGE.PV"), "CO 2 "Inflow (ml / ml)" ("CO2.PV"), "Feed 3 Cumulative (ml)" ("FED3T.PV"), "Oxygen Utilization (mol / (l * measuring the current values ​​of "OUR" and "pH" ("PH.PV"); and (b) determining viable cell density or viable cell volume in the culture medium using the measurements of (a) with a data-driven model for mammalian cell culture generated using a feature matrix including the process variables "Time," "CHT.PV," "ACOT.PV," "FED2T.PV," "GEW.PV," "CO2T.PV," "ACO.PV," "AO.PV," "N2.PV," "LGE.PV," "CO2.PV," "FED3T.PV," "OUR," and "PH.PV." Including, the method is performed without sampling, using only online measurements from the culture, and the mammalian cells are CHO-K1 cells; method.

2. The method of claim 1 , wherein the data-driven model is generated by machine learning.

3. The method of claim 1 or 2, wherein the data-driven model is generated using a random forest method.

4. 4. The method of claim 1, wherein the data-driven model is generated using a training dataset comprising at least 10 culture runs.

5. (a) the dataset available for modeling is randomly split into training and testing datasets in a ratio of 70:30 to 80:20; (b) a model is generated; (c) a mean and standard deviation for determining a process variable of a dataset is determined from the training dataset, and a mean and standard deviation for determining a process variable of a dataset is determined from the test dataset; (d) steps (a)-(c) are repeated until comparable means and standard deviations are achieved for the partitions between the test and training data sets, the partitions obtained under (a) being different in each new run; The method according to claim 4, characterized in that

6. 6. The method of claim 4 or 5, wherein the data sets used to generate the data-driven model each contain the same number of data points.

7. 7. The method according to claim 4, wherein the data points in the dataset used to generate the data-driven model are each for the same incubation time.

8. A method according to any one of claims 4 to 7, characterized in that missing data points in the data set are filled by interpolation.

9. 9. The method of claim 8, wherein missing data points of viable cell volume can be obtained by third order polynomial fitting and / or missing data points of viable cell density can be obtained by Peleg fitting.

10. The method of any one of claims 4 to 9, wherein the data set comprises a data point at least every 144 minutes.

11. The method according to any one of claims 1 to 10, characterized in that the mammalian cells express and secrete an antibody.

12. 12. The method of any one of claims 1 to 11, wherein the data-driven model is generated using a training dataset comprising a composite IgG culture run and a standard IgG culture run.

13. The method according to any one of claims 1 to 12, characterized in that the culture volume is 300 mL or less.

Citation Information

Patent Citations

  • Regulator for culture vessel and culture device

    JP2006296423A

  • Operation-regulating device of culturing vessel

    JP2007202500A

  • Cell culture device

    JP2018117567A