Method for determining process variables in a cell culture process

CN114223034BActive Publication Date: 2026-08-11F HOFFMANN LA ROCHE & CO AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

没有公开预测或在线算法

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114223034B_ABST
    Figure CN114223034B_ABST
Patent Text Reader

Abstract

High-throughput culture systems are used in drug research and development. In this process, samples are collected and analyzed using external analytical techniques for key parameters. The results of these analyses are used to evaluate the culture process and provide important information about it. Particularly in the case of parallel cultures, sample preparation is labor-intensive and prone to error. To avoid the need for sampling and thus minimize errors, this patent application describes a method that allows access to desired target parameters via previously recorded process variables in the form of soft sensors. This document describes a method for determining process-related parameters, particularly glucose, lactate, and viable cell density or volume, in CHO (Chinese hamster ovary) processes during high-throughput culture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention belongs to the field of mammalian cell culture. More specifically, the object of this invention is a method for determining process target parameters online based on historical online and offline values ​​of a set of process variables. Background Technology

[0002] For the production of therapeutic agents in the pharmaceutical industry, the highest requirements are placed on quality and reproducibility. To this end, empirical standards (GMP guidelines, good manufacturing practices) have been developed to define target values, process limits and deviations to meet these requirements. The U.S. Food and Drug Administration (FDA) recently required the pharmaceutical industry to better understand the processes performed in order to improve the quality of its products through the PAT (Process Analytical Technology) program [1]. In recent years, new technologies (such as computer-based models) have been used to facilitate the understanding of cell culture processes, such as those used to produce therapeutic proteins from CHO cells.

[0003] Bioreactors are most commonly used for cell culture. During culture, various process variables are recorded. These enable process monitoring and control, and are used to maintain controlled environmental conditions. A distinction is made between online and offline values. Both values ​​provide important information about the process. Online values ​​are collected by appropriate sensors for direct online control. However, offline values ​​are determined through manual sampling and subsequent external analysis. These offline parameters are, for example, viable cell density, glucose, and lactate concentrations. These can be used to assess current culture conditions and, if necessary, to intervene in the regulation of the process.

[0004] Sample analysis requires more manual intervention, especially in high-throughput culture systems. In some cases, these external methods can also lead to errors and equipment malfunctions. To make the process more efficient and robust, online information can be obtained using online values ​​recorded during culture. This allows for the analysis of existing measured parameters and their relationships, enabling the use of appropriate machine learning mathematical models for description.

[0005] Artificial neural networks (ANNs) for monitoring biomass during fed-batch processes have been described [8]. Kroll et al. described a model-based soft sensor for identifying subpopulations of CHO cell biomass [9].

[0006] Hutter, S. et al. published a glycosylation flux analysis of immunoglobulin G in perfused cell cultures of Chinese hamster ovaries (Process 6 (2018) 176). They described a method based on metabolic flux analysis to gain insights into the glycosylation pathway. Hutter et al. focused on metabolic flux analysis in perfused cell culture experiments. Only offline determined parameters were used to fit a mechanical (linear) model, and a random forest model was used to rank the influence of input parameters on glycosylation results. Therefore, Hutter et al. disclosed a modeling tool based on offline data and statistical analysis performed post-culture, even if historical data had (biological) significance. No predictions or online algorithms were disclosed.

[0007] The white paper “Biopharma PAT – Quality Attributes, Critical Process Parameters & Key Performance Indicators at the Bioreactor” (available at https: / / www.researchgate.net / publication / 326804832_Biopharma_PAT_-_Quality_Attributes_Criticcal_Process_Parameters_Key_Performance_Indicators_at_the_Bioreactor) provides a high-level overview of process analysis techniques. This paper describes cultivation principles (e.g., batch, fed-batch, and perfusion, monitoring methods). The impact of measurements such as dissolved oxygen is used to gain process understanding. No predictions of output parameters or any machine learning methods are disclosed.

[0008] Rubin, J. et al. reported that pH shifts affect CHO cell culture performance and antibody N-linked glycosylation (Bioprocess. Biosys. Eng., 41(2018) 1731-1741). The authors published a study on the effect of cell culture pH on antibody glycosylation, which used typical offline measurements of process parameters in any culture and the effect of pH changes on them.

[0009] Downey, BJ et al. reported a novel method for predicting living cell volume (VCV) using dielectric spectroscopy in early process development (Biotechnol. Prog. 30 (2014) 479-487).

[0010] Xiao, P. et al. reported the metabolic characteristics of the CHO cell size increase phase in fed-batch culture (Appl. Microbiol. Biotechnol. 101(2017) 8101-8113).

[0011] Kroll, P. et al. reported a soft sensor for monitoring biomass subpopulations during mammalian cell culture (Biotechnol. Lett. 39 (2017) 1667-1673). The authors used a turbidity physics sensor to measure live cell counts (VCC, equivalent to VCD) based on a linear model. Summary of the Invention

[0012] This invention is based, at least in part, on the discovery that useful data-driven models containing important parameters for culturing CHO cells can be obtained by selecting certain specific process variables from historical datasets. These parameters include real-time VCD (vital cell density), VCV (vital cell volume), glucose, and lactate. Using the method according to the invention, accurate, near-online values ​​for the target variables can be provided throughout the culture process without sampling.

[0013] The object of this invention is a method for determining viable cell density and / or viable cell volume and / or glucose concentration and / or lactate concentration in the culture medium using only online measurements from the culture, during and in the culture of antibody-expressing CHO cells, characterized in that the model is generated based on a feature matrix comprising the following terms: 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'.

[0014] The feature representation parameters are as follows:

[0015]

[0016] In one embodiment, the model is generated using the random forest method.

[0017] In one embodiment, the training dataset contains at least 10 culture runs, preferably at least 60 culture runs.

[0018] In one embodiment, the model is obtained using a training dataset that includes culture runs of mammalian cells expressing complex IgG, i.e., containing a form of antibody different from the wild-type Y-shaped full-length antibody, for example by including additional structural domains, such as one or more Fabs. In one embodiment, the training dataset also contains culture runs of mammalian cells expressing standard IgG, i.e., Y-shaped wild-type-like antibodies without additional or missing structural domains.

[0019] In one embodiment, approximately 80% of the dataset available for model formation is used as the training dataset, while the remaining dataset is used as the test dataset.

[0020] In one embodiment

[0021] a) Randomly divide the dataset that can be used for modeling into training and testing datasets at a ratio of 80:20.

[0022] b) Form a model,

[0023] c) Determine the mean and standard deviation of the target parameters for the dataset from the training dataset, and determine the mean and standard deviation of the target parameters for the records from the test dataset.

[0024] d) Repeat steps a) to c) until comparable mean and standard deviation are achieved for the split between the test dataset and the training dataset, i.e., within 10% of each other, preferably within 5%.

[0025] In one embodiment, missing data points in the dataset are supplemented by interpolation.

[0026] In one embodiment, the dataset contains data points for at least 60 minutes, preferably approximately every 5 to 10 minutes. Attached Figure Description

[0027] Figure 1 : Linear interpolation values ​​obtained using ACO.PV as an example. The interpolation range is from day 0.5 to day 13.5.

[0028] Figure 2 Interpolation curves for the density of live cells in an exemplary culture. Interpolation and coefficients of determination: Peleg fit (R² = 0.957), univariate spline (R² = 0.998), and third-order polynomial fit (R² = 0.864).

[0029] Figure 3 Example correlation analysis of the dataset from the ambr250 run in Project 2. Comparison of correlation coefficients for different interpolation strategies. This figure shows a scatter plot of a single online parameter of the VCD.

[0030] Figure 4 The information content is calculated based on the mutual information of the target variable VCD across the entire dataset.

[0031] Figure 5 The random forest VCD estimation was performed twice independently. An estimate with R² of 0.20317 was achieved in the upper part of the graph. An estimate with R² of 0.54896 was achieved in the lower part of the graph.

[0032] Figure 6 Histograms of predictions for newly created test datasets for models MLPRegressor (a), Random Forest (b), and XGBoost (c) targeting the variable 'VCD'. The error in fitting the predicted VCD values ​​is shown on the X-axis. The Y-axis represents the relative frequency of the error.

[0033] Figure 7 : Estimation of VCD for random forest using two exemplary runs on the test dataset.

[0034] In the upper part of the graph, an estimate of R² of 0.98944 is achieved. In the lower part of the graph, an estimate of R² of 0.99837 is achieved.

[0035] Figure 8 : Calculate the information content of the entire dataset based on the mutual information of the target variable glucose.

[0036] Figure 9 : Estimate glucose based on two exemplary random forests run on the test dataset.

[0037] In the upper part of the graph, an estimate of R² of 0.99 can be obtained. In the lower part of the graph, an estimate of R² of 0.97 can be obtained.

[0038] Figure 10 Calculate the information content of the entire dataset based on the mutual information of the target variable lactic acid.

[0039] Figure 11 : Histograms of predictions for the target variable lactate on the test dataset using MLPRegressor (a), Random Forest (b), and XGBoost (c). The error added to the predicted lactate values ​​is shown on the X-axis. The Y-axis represents the relative frequency of the error.

[0040] Figure 12 For two exemplary runs on the test dataset, lactate is estimated using XGBoost. In the upper part of the graph, an estimate with R² of 0.99 is achieved. In the lower part of the graph, an estimate with R² of 0.98 is achieved.

[0041] Figure 13: Calculate the RMSE of MLPRegressor, Random Forest and XGBoost using different amounts of training datasets.

[0042] Figure 14 : Estimation of VCD for a single culture. Peleg fits of VCD are shown in blue, and VCD estimates are shown in orange.

[0043] Figure 15 : Indicates the average diameter of each sample taken throughout the entire culture cycle. Items 1 and 3 have complex molecular formats (shown in blue here, on the left) as products. Items 2 and 4 have Y-shaped Ig-G formats (shown in green here, on the right) as target products. Box plots include the mean; units are displayed in a normalized manner.

[0044] Figure 16 Left side of the graph: Random forest estimation of VCD. The red part is the estimate used against the real test dataset. The blue part is the estimate used against the real training dataset. Ideal estimates for both the test and training datasets are shown in black. Right side of the graph: Random forest estimation of VCV. The red part is the estimate used against the real test dataset. The blue part is the estimate used against the real training dataset. Ideal estimates for both the test and training datasets are shown in black.

[0045] Figure 17 : Represents the average diameter of each sample for each item throughout the entire culture period. Item 1 = purple, Item 2 = red, Item 3 = green, and Item 4 = blue. The box plot includes the mean.

[0046] Figure 18 Compare VCD / VCV using the Random Forest model (the best model).

[0047] Figure 19 Considering all models (MLPRegressor, Random Forest, XGBoost) and the training dataset, the behavior of RMSE depends on the target parameter VCV.

[0048] Figure 20 Bar chart showing the difference in RMSE between the test and training datasets, and the best model for the target variable VCV.

[0049] Specific embodiments of the present invention

[0050] 1. A method for determining one or more process variables during mammalian cell culture, characterized in that,

[0051] Process variables are determined individually.

[0052] i) A data-driven model of mammalian cell culture was generated using a feature matrix containing process variables 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'.

[0053] as well as

[0054] ii) Use only / only online measurements from the culture.

[0055] 2. The method according to Example 1, characterized in that the online measured values ​​are used for at least the process variables 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'.

[0056] 3. A method for adjusting glucose concentration to a target value during mammalian cell culture, comprising the following steps:

[0057] a) Determine at least the current values ​​of the following process variables for cultivation: 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'.

[0058] b) Using the values ​​determined in a), determine the current glucose concentration in the culture medium using the data-driven model for mammalian cell culture, which is generated using a feature matrix containing process variables 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'.

[0059] as well as

[0060] c) If the current glucose concentration determined in b) is lower than the target value, add glucose until the target value is reached, thereby adjusting the glucose concentration to the target value.

[0061] 4. The method according to any one of Examples 1 to 3, characterized in that the process variables are selected from the group consisting of: live cell density, live cell volume, glucose concentration in the culture medium, and lactate concentration in the culture medium.

[0062] 5. The method according to any one of Examples 1 to 4, characterized in that the method is performed without sampling and uses only online determined / measured values ​​from the culture.

[0063] 6. The method according to any one of embodiments 1 to 5, characterized in that the data-driven model is generated through machine learning.

[0064] 7. The method according to any one of embodiments 1 to 6, characterized in that the data-driven model is generated using a method selected from the group including artificial neural networks and ensemble learning.

[0065] 8. The method according to any one of embodiments 1 to 7, characterized in that the data-driven model is generated using the random forest method.

[0066] 9. The method according to any one of embodiments 1 to 7, characterized in that the data-driven model is generated using the MLPRegressor method.

[0067] 10. The method according to any one of embodiments 1 to 7, characterized in that the data-driven model is generated using the XGBoost method.

[0068] 11. The method according to any one of embodiments 1 to 10, characterized in that the data-driven model is generated through supervised learning.

[0069] 12. The method according to any one of Embodiments 1 to 11, characterized in that the data-driven model is validated through cross-validation.

[0070] 13. The method according to Example 12, wherein the cross-validation is 10x cross-validation.

[0071] 14. The method according to any one of Examples 1 to 13, characterized in that the data-driven model is generated using a training dataset containing at least 10 culture runs.

[0072] 15. The method according to Example 14, wherein the training dataset contains at least 60 training runs.

[0073] 16. The method according to any one of Embodiments 1 to 15, characterized in that approximately 80% of the dataset that can be used to generate the model is used as a training dataset, and the remaining dataset is used as a test dataset.

[0074] 17. The method according to any one of Examples 1 to 16, characterized in that,

[0075] a) Randomly divide the datasets available for modeling into training and testing datasets at a ratio between 70:30 and 80:20.

[0076] b) Form a model,

[0077] c) Determine the mean and standard deviation of the process variables used to determine the dataset from the training dataset, and determine the mean and standard deviation of the process variables used to determine the dataset from the test dataset.

[0078] d) Repeat steps a) to c) until comparable mean and standard deviation are reached, i.e., within 10% of each other, preferably within 5% of each other, relative to the test dataset and the training dataset, wherein the split obtained under a) is different for / with each new run.

[0079] 18. The method according to any one of Examples 1 to 17, characterized in that the datasets used to generate the data-driven model each contain the same number of data points.

[0080] 19. The method according to any one of Examples 1 to 18, characterized in that the data points in the dataset used to generate the data-driven model each point are for the same cultivation time point.

[0081] 20. The method according to any one of embodiments 1 to 19, characterized in that missing data points in the dataset are obtained by interpolation.

[0082] 21. The method according to Example 20, characterized in that missing data points of glucose concentration and / or live cell volume are obtained by fitting a third-order polynomial.

[0083] 22. The method according to any one of Examples 20 to 21, characterized in that missing data points of lactic acid concentration are obtained by univariate spline fitting.

[0084] 23. The method according to any one of Examples 20 to 22, characterized in that missing data points of live cell density are obtained by Peleg fitting.

[0085] 24. The method according to any one of Examples 1 to 23, characterized in that each dataset contains at least one data point every 144 minutes.

[0086] 25. The method according to any one of Examples 1 to 24, characterized in that each dataset contains at least one data point every 60 minutes.

[0087] 26. The method according to any one of Examples 1 to 25, characterized in that each dataset contains a data point approximately every 5 to 10 minutes.

[0088] 27. The method according to any one of Examples 1 to 26, characterized in that the mammalian cell is a CHO cell.

[0089] 28. The method according to any one of Examples 1 to 27, characterized in that the mammalian cell is a CHO-K1 cell.

[0090] 29. The method according to any one of Examples 1 to 28, characterized in that mammalian cells express and secrete therapeutic proteins.

[0091] 30. The method according to any one of Examples 1 to 29, characterized in that mammalian cells express and secrete antibodies.

[0092] 31. The method according to Example 30, wherein the antibody is a monoclonal and / or therapeutic antibody.

[0093] 32. The method according to any one of Examples 30 to 31, characterized in that the antibody is not a standard IgG antibody, i.e., a wild-type full-length four-chain antibody or a complex antibody, i.e., an antibody containing additional antibody and / or an antibody with non-antibody domains compared to a standard antibody.

[0094] 33. The method according to any one of Examples 1 to 32, characterized in that the data-driven model is generated using a training dataset that contains only complex IgG culture runs.

[0095] 34. The method according to any one of Examples 1 to 33, characterized in that the data-driven model is generated using a training dataset, which further includes a standard IgG culture run.

[0096] 35. The method according to any one of Examples 1 to 34, characterized in that mammalian cells express and secrete complex or standard IgG.

[0097] 36. The method according to any one of Examples 1 to 35, characterized in that the culture volume is 300 mL or less.

[0098] 37. The method according to any one of Examples 1 to 36, characterized in that the culture volume is 250 mL or less, 200 mL or less, 100 mL or less, 75 mL or less, between 200 and 250 mL, or between 50 and 100 mL.

[0099] 38. The method according to any one of Examples 1 to 37, characterized in that the culture is a fed-batch culture.

[0100] 39. The method according to any one of Examples 1 to 38, characterized in that the cultivation is carried out in a stirred tank reactor.

[0101] 40. The method according to any one of Examples 1 to 39, characterized in that immersion aeration is present during cultivation.

[0102] 41. The method according to any one of Examples 1 to 40, characterized in that the cultivation is carried out in a disposable bioreactor (SUB).

[0103] 42. The method according to any one of Examples 1 to 41, characterized in that the mammalian cells are cultured in suspension or are mammalian cells grown in suspension.

[0104] 43. The method according to any one of Examples 1 to 42, characterized in that the data-driven model is generated through regression analysis.

[0105] 44. The use of live cell volume as a target parameter in generating data-driven models for determining process variables of culturing mammalian cells in volumes of 300 mL or less.

[0106] 45. The use according to Example 44, characterized in that the process variable is selected from the group consisting of: live cell density, live cell volume, glucose concentration in the culture medium, and lactate concentration in the culture medium.

[0107] 46. ​​The use according to any one of Examples 44 to 45, characterized in that the culture is carried out without sampling.

[0108] 47. The use according to any one of Examples 44 to 46, characterized in that the mammalian cell is a CHO cell.

[0109] 48. The use according to any one of Examples 44 to 47, characterized in that the mammalian cell is a CHO-K1 cell.

[0110] 49. The use according to any one of Examples 44 to 48, characterized in that the mammalian cell expresses and secretes a therapeutic protein.

[0111] 50. The use according to any one of Examples 44 to 49, characterized in that the mammalian cells express and secrete antibodies.

[0112] 51. The use according to Example 50, characterized in that the antibody is a monoclonal and / or therapeutic antibody.

[0113] 52. The use according to any one of Examples 50 to 51, characterized in that the antibody is not a standard IgG antibody or a complex antibody.

[0114] 53. The use according to any one of Examples 44 to 52, characterized in that the data-driven model is generated using a training dataset that contains only complex IgG culture runs.

[0115] 54. The use according to any one of Examples 44 to 53, characterized in that the data-driven model is generated using a training dataset, which further includes a standard IgG culture run.

[0116] 55. The use according to any one of Examples 44 to 54, characterized in that the mammalian cells express and secrete complex or standard IgG. Detailed Implementation

[0117] To achieve high-throughput testing cultures, especially for complex molecules and molecular forms, the size of culture vessels must be reduced and the cultures must be automated. The success of a culture depends on controlled process variables, and the desired molecules can only be produced in high yields when optimal culture conditions are provided. Therefore, rapid and efficient control of relevant process variables is required to set the appropriate process variables and maintain optimal culture conditions. This control is particularly necessary for small-scale parallel cultures, as each culture must be monitored individually. Offline process variables are a problem here because, on the one hand, the required sampling and individual analysis lead to time skew—that is, as the culture continues, the process variable determined offline differs from the actual process variable; on the other hand, the number of sampling points is much smaller compared to online available process variables, resulting in poor time control of that process variable.

[0118] Therefore, one object of the present invention is to make process variables that cannot be determined online but can only be determined offline (especially due to the size of the culture vessel used) available in real time, based on a data-driven model, at the culture scale used, just like process variables that are available online.

[0119] To produce recombinant proteins, most bioreactors use fed-batch processes [4]. In addition to fed-batch processes, there are other operating modes, such as batch processes and continuous culture modes.

[0120] Fed-batch or fed-batch processes are partly open systems. The advantage of this process is that nutrients such as glucose, glutamine, and other amino acids can be added to the culture during the process. Substrate limitations can be avoided, and longer treatment times can be ensured. Substrate can be added continuously or in the form of concentrated pellets (one or more). Appropriate feeding strategies can be used to better control inhibition and the accumulation of toxic byproducts. However, this requires sufficient knowledge and process control.

[0121] To provide and maintain optimal conditions during the culture of mammalian cells, such as CHO cells, almost exclusively bioreactors are used [2]. Most of the bioreactors used are stirred tank reactors. Cultures are carried out in suspension, i.e., cells with suspension growth.

[0122] Aerobic mammalian cells, such as CHO cells, require oxygen to maintain their cellular metabolism. Oxygen is typically supplied to cells by submerging the culture medium. The concentration of dissolved oxygen in the reactor is one of the most important parameters for aerobic cell culture. The concentration of dissolved oxygen in the culture medium is determined by a number of transport resistances. Diffusion causes oxygen to be transported from the bubbles into the cells, where it is eventually metabolized. Transport mechanisms can be described using the oxygen transport rate (OTR), while the oxygen consumption of the cells themselves can be determined using the oxygen consumption rate (OUR) [2]. Appropriate exhaust gas analysis can provide the data needed to calculate OUR and OTR. Process variables such as temperature, pH, and dissolved oxygen concentration are monitored by appropriate sensors and included in the parameters to be controlled within the culture. These process variables have a significant impact on the effective productivity of mammalian cell lines [3].

[0123] To shorten the development and setup time of bioreactors, research and development are increasingly focused on single-use technologies (single-use bioreactors; abbreviation: SUB). A significant advantage of these systems is that they eliminate the need for complex cleaning processes and the necessary, complex, and expensive cleaning methods such as CIP (clean-in-situ) and SIP (sterilize-in-situ).

[0124] Automated high-throughput culture systems such as the ambr250 system (automated microbioreactor) help accelerate drug development. This system contains twelve disposable bioreactors, each with a volume of 250 mL. An automated liquid processor is used for pipetting and sampling. Operation is controlled by central process software. The entire ambr250 system is located below a laminar flow chamber to ensure a sterile environment during operations.

[0125] Over the past two decades, soft sensors have been increasingly used in industry for monitoring process variables[6]. These process variables are typically determined only through high-resolution analysis or externally (i.e. offline). In particular, when using disposable systems on a small scale, it is often impossible to install the additional sensors required (space and availability or connectivity with disposable bioreactors, the inability to perform gamma irradiation, etc.). As a result, there is a lack of continuous data on important process variables, especially at small-scale cultivation, which can be used for process monitoring and allow for adjustment of these process variables, i.e., process target parameters. The name “soft sensor” combines the terms “software” and “sensor”. The term “software” refers to the computer-aided programming of the model. The outputs of these models provide information about the cultivation, particularly real-time values ​​of process variables that would otherwise be unavailable due to the lack of corresponding physical sensors[5].

[0126] Basically, soft sensors can be divided into two categories: model-driven soft sensors and data-driven soft sensors.

[0127] Model-driven soft sensors are constrained by theoretical process models. These require a detailed understanding of the ongoing process and its description using state-differential equations. This means that mechanical models must be used to describe the dynamic behavior of the process. Such models are primarily developed for the planning and design of process plants and focus on describing ideal equilibrium states.

[0128] In data-driven soft sensors (so-called black-box models), machine learning-based models are used. These include empirical models, which use historical data to describe the correlations of process variables. Biological processes are complex, and every single aspect of the metabolism of cultured mammalian cells is not yet fully understood.

[0129] Data-driven soft sensors have wide applications in the pharmaceutical industry. They are typically used for monitoring and recording culture data.

[0130] The inventors have now discovered that such historical data can be used to generate data-driven models for online estimation of offline process variables.

[0131] Process variables are primarily determined in real time, i.e., available; they are typically determined only by increasing the analytical workload and the associated time offset. Furthermore, for online monitoring of certain process variables (such as biomass or certain substrate and product concentrations), robust and long-term stable online sensor systems are not always available [7]. Although these parameters contain important information about the culture process, they are only available at limited time points during the culture period, i.e., at the time points of sampling and offline analysis.

[0132] Using small systems, such as the ambr250 system, certain process variables, such as turbidity and / or conductivity, cannot be measured due to the lack of probe ports. Furthermore, some common probes require relatively large spaces due to their design, which is unavailable in these small-capacity systems.

[0133] Machine learning is the application of algorithms to describe the basic structure of a dataset. Machine learning can be divided into two parts: supervised learning and unsupervised learning.

[0134] Supervised learning is used when the model is ready to make predictions about future or unknown data based on the training data. It is supervised because the training dataset already contains information about the desired output value. An example is spam filtering

[10] : Thus, the algorithm receives a dataset consisting of spam and non-spam emails, which already contains information about spam / non-spam emails it passed through the learning phase. Using unlabeled new emails, the algorithm now tries to predict what type of message it is. Since this is a classification target variable (spam / non-spam), the term "classification" is used.

[0135] In unsupervised learning, the aim is to extract relationships within a dataset without providing the algorithm with a target variable. The focus is on exploring the underlying structure of the data to extract meaningful information. The simplest example of this is clustering. In this exploratory data analysis, the goal is to divide the dataset into meaningful subgroups without knowing the actual group members beforehand.

[0136] If the target variable is a continuous variable, it is called regression or regression analysis. The variables used to describe the regression model are called independent variables or explanatory variables. Based on this, the aim is to find the mathematical relationship between the input variables and the target parameter in order to predict the results.

[0137] The method according to the invention uses supervised learning, wherein the target variable is described by regression.

[0138] Modeling can be schematically aligned in the following steps: preprocessing, learning, evaluation, and estimation of the target variable.

[0139] Data preprocessing is required to ensure the model can correctly interpret the information it is based on. The dataset is prepared as a feature matrix x, containing m features (columns) and n rows, thus representing explanatory variables. Each row n contains feature descriptions for a specific data point.

[0140]

[0141] The target variables are arranged in the vector y. Therefore, the feature matrix x (n) Each row contains the target variable y (n) Information on relevant values.

[0142] Statistical analysis is used to identify suitable features. Once the appropriate feature locations are determined and the corresponding feature matrix is ​​created, the model can be trained using a subset (70%-80% of the entire dataset). This subset is called the training dataset.

[0143] Typical data preprocessing may involve providing the model with a standardized dataset. Thus, the data for each feature has the property of a standard normal distribution with a mean of 0 and a standard deviation of 1. This increases the comparability of features with each other and enables the learning algorithm to achieve its optimal performance

[10] .

[0144] Learning is the core part of model building. During learning, the model attempts to understand and identify relationships between data. Each model follows a mathematical formula with specific parameters. These are tuned during training to describe the relationships between data as well as possible.

[0145] Some models, such as neural networks, have other parameters that do not change during the learning process. These are called hyperparameters. They affect the complexity of the model or the speed of the learning process and are determined before training. There is no fixed formula for choosing the right hyperparameters. Therefore, different models are trained with different hyperparameters and then tested. Only in this way can it be determined which model is most suitable.

[0146] Randomization and a grid-based algorithm are used to search for the optimal combination of hyperparameters. Each hyperparameter is represented by a list with different values. The model is trained using a grid search with every possible combination from the corresponding list. Random search reduces the computational effort required. Different random parameter combinations are used, where the computational effort can be predetermined. In one embodiment, a random search is first used to perform the model to coarsely estimate the hyperparameters, followed by a grid search to fine-tune the hyperparameters. The goal of learning is to train the model to keep bias and variance as low as possible.

[0147] Models often learn relationships between training data better than they do with subsequent predictions using unknown datasets. This behavior is called overfitting. Thus, the model has memorized the training dataset and uses new data to describe relationships with insufficient accuracy. Similar behavior can be attributed to excessive variance. Here, the model uses too many input parameters to train on the dataset, causing a complex model to only fit datasets with high data variance. Therefore, the model learns noise from the data without being able to map actual relationships.

[0148] On the other hand, if the model is not complex enough to react to changes in the test dataset, this is called underfitting. In this case, the bias is too large, and the model can only imprecisely map the relationship between the training and test data.

[0149] During the learning process, k-fold cross-validation of the training dataset provides the possibility of avoiding model overfitting

[11] . The training dataset is divided into k subsets. Then, the k-1 subsets are used to train the model, and the remaining subsets are used as the test dataset. This process is repeated k times. In this way, k models are trained and estimates of k target variables are obtained.

[0150] Each run generates a performance estimate E for the model. i For example, mean squared deviation is used as a measure of error for performance estimation of regression. In practice, 10-fold cross-validation has proven to be a good trade-off between bias and variance in most cases

[12] :

[0151]

[0152] Artificial neural networks (ANNs) were described in 1943 by Warren McCulloch and Walter Pitts using mathematical models of neurons. This made it easier to understand the transmission of information in biological systems

[13] . Frank Rosenblatt was then able to connect the McCulloch-Pitts model of artificial neurons with learning rules, thus describing the perceptron

[14] . The perceptron remains the foundation of ANNs.

[0153] A simple perceptron has n inputs x1,...,x n ∈IR, each input has a weighted sum w1,...,w n ∈IR. The output is represented by o∈IR. The processing of the input signal with appropriate weighting is the propagation function (input function) σ.

[0154]

[0155] It describes the network input of neurons. Through the activation function φ,

[0156] o=φ(σ),

[0157] Then determine the output o of the perceptron. Various functions can be used with respect to φ, which can lead to the activation of the perceptron.

[0158] Therefore, the activation function calculates the intensity of neuron activation based on the threshold and network input

[15] . If several of these neurons are interconnected in a suitable structure, a complex relationship between the input and output layers can be mapped. The simplest form of this structural interconnection of simple neurons is the feedforward network. They are arranged in layers and consist of an input layer, an output layer, and several hidden layers, depending on the structure.

[0159] In a feedforward network (so-called a multilayer perceptron), each neuron in one layer is connected to all other neurons in the next layer. Thus, these networks propagate forward the information content created by the network. Each neuron weights the input signal using initially randomly chosen weights and adds a bias term. The neuron's output corresponds to the sum of all weighted input data. The complexity of the neural network can be determined based on the number of neurons in one layer and the number of hidden layers.

[0160] Multilayer feedforward networks containing error feedback (backpropagation) are mainly used for supervised learning of ANNs

[16] .

[0161] The training of this neural network can be divided into three steps:

[0162] Step 1: Feedforward;

[0163] Step 2: Error Calculation;

[0164] Step 3: Backpropagation.

[0165] In the first step, the input layer of the network is fed into the network, which propagates through the network layer by layer until there is an output from the network. The output of the network is compared with the expected value in the second step, and the error of the network is calculated using an error function. Each neuron in the hidden layer contributes differently to the calculation of the error depending on the current weights. In the third step, the error is propagated back through the network, where the weights are adjusted according to the contribution of the weights of individual neurons to the error. The goal of the backpropagation algorithm is to minimize the error, and gradient descent is typically used

[17] . Using this method, the quadratic distance between the network output and the expected output is calculated as the error function:

[0166]

[0167] To calculate the contribution of each neuron's weight to the error, the error function Err must be derived from the weights w considered. ij Therefore, only continuous and differentiable activation functions can be used here

[17] . This determines the weight adjustment increment to be used in the next iteration step. This relationship can be described mathematically as follows:

[0168]

[0169] The learning rate η, along with the number of iterations, is a hyperparameter established before training the model. These two steps are repeated until the maximum number of iterations or a defined error value is reached, and good results are obtained for unknown inputs.

[0170] Furthermore, the Random Forest (RF) algorithm can be used for machine learning of regression problems

[18] . RF learns through a large number of decision trees, thus belonging to the category of ensemble learners. A decision tree can be expanded from the root (top node, which has no predecessor). Each node divides the dataset into two groups based on features. The successor of the root can be a leaf (without a successor) or a node (with at least one successor). Nodes and leaves are connected by an edge. In the case of regression problems,

[19]

[0171] • Assign a feature to each internal node (including the root);

[0172] • Assign a specific value of the target variable to be predicted to each leaf of the decision tree;

[0173] • For each edge, assign a relation to a threshold.

[0174] In a preferred embodiment, RF uses the packing principle (bootstrapping aggregation principle) of Breiman

[18] to create a suitable training set, where the training set is created by substitution sampling from the entire training dataset. Some data may be selected multiple times, while others will not be selected as training data. The number of training sets always corresponds to the number of the entire training dataset. Each selected training set is used to generate a decision using a decision tree (classifier). The decisions of all training sets are then averaged, where the majority decision determines the final classification. Thus, the generation of bootstrap samples produces low correlation between individual classifiers. Furthermore, the variance of individual classifiers can be reduced and the overall classification performance can be improved

[18] .

[0175] In a preferred embodiment, the feature used to determine the split (the partitioning of nodes) during tree creation makes the clearest decision about the random selection of features in the dataset. The selected split is no longer chosen as the best split among all features, but rather the best split among the randomly selected features. Due to this randomization, the bias (distortion, systematic error) of the tree increases during creation. The variance decreases due to the formation of the average of all trees contained in the RF. The reduced variance has a greater added value than the increase in bias, thereby improving the accuracy of the model

[20] .

[0176] Furthermore, since the average of all individual decisions is always taken into account, overfitting of the model is almost prevented in RF prediction

[18] .

[0177] XGBoost (eXtreme Gradient BOOSTing) uses an ensemble of regression trees as the basis for model formation. The packing principle and special augmentation techniques described above are used to train the ensemble to achieve the most accurate predictions. Simply put, the augmentation technique can be seen as a combination of gradient descent methods consisting of many weak learners

[21] . These weak learners are usually no more accurate than random guessing and are grouped into strong learners during the creation of the ensemble. A typical example of such a weak learner is a simple regression tree with only one node. The principle of the augmentation algorithm is to select training data that is difficult to classify so that the weak learners can learn from these poorly classified objects, thereby improving the performance of the ensemble. Due to the complexity of XGBoost, the algorithm is considered a black box. However, due to its scalability and speed of problem-solving, the algorithm has been very successful in direct comparisons of different machine learning models

[22] .

[0178] The method implemented by XGBoost combines gradient descent and boosting techniques, and is described below using Tianqi Chen’s original paper “XGBoost: A Scalable Tree Boosting System”

[22] .

[0179] Using an ensemble of K decision trees, the model can be described under the following conditions:

[0180]

[0181] Where f k This is the prediction from a single decision tree. It can be seen across all decision trees that predictions can be made.

[0182]

[0183] Where x i Let L be the feature vector of the i-th data point. To train the model, optimize the loss function L. In the case of a regression problem, use RMSE (Root Mean Square Error):

[0184]

[0185] Regularization is an important part of preventing model overfitting:

[0186]

[0187] Where T is the number of leaves, and w 2 j This is the achievement score of the j-th leaf. If we combine regularization and the loss function, the model's basic objective function can be formulated as follows:

[0188] Obj = L + Ω,

[0189] The loss function determines the predictive power, and regularization controls the model's complexity. Gradient descent is used to optimize the objective function. Given the objective function to be optimized... Calculate gradient descent in each iteration

[0190]

[0191] and The objective function Obj is minimized by changing along the descending gradient.

[0192] To create a regression tree, internal nodes are partitioned based on features of the dataset. The resulting edges define the range of values ​​that allow for partitioning the dataset. Leaves within the regression tree are weighted, with weights corresponding to predicted values. The number of iterations represents the frequency with which the packing and augmentation processes are repeated. The XGBoost algorithm provides a very extensive list of hyperparameters that greatly contribute to forming a good model.

[0193] Regardless of the model used, correlation can be used to assess and represent the linear relationship between two variables. The Pearson correlation coefficient r (or r0) is a key indicator of this relationship. 2 This provides a commonly used metric for evaluating this relationship. It is dimensionless and calculated according to the following formula:

[0194]

[0195] And it varies within the range of -1 ≤ r ≤ +1. The counter describes the sum of the products of the deviations of the two variables x and y as the mean, which corresponds to the empirical covariance s. xy The denominator is a single empirical standard deviation s. x and s y The root of the product. The average of the related quantities is described as... and According to the linear relationship of Fahrmeir

[23] , it can be interpreted as follows:

[0196] • r < 0.5: Weak linear relationship

[0197] • 0.5 ≤ r < 0.8: Moderate linear relationship

[0198] • 0.8 ≤ r: Strong linear relationship

[0199] In correlation analysis, it's important to note that only linear relationships can be shown. Therefore, the Bravais-Pearson correlation coefficient is not suitable for describing non-linear relationships. This might mean that even though the correlation coefficient is 0.0 ≤ r ≤ 0.2, there is a strong non-linear correlation between the variables.

[0200] Mutual information can be used to determine the nonlinear correlation between two random variables. It is used in information theory

[24] . With the help of probability, it describes the information content of a random variable compared to a second random variable. The basic formal relationship is:

[0201]

[0202] Kraskov et al. and Ross et al. extended this approach accordingly so that it could be used to select appropriate continuous variables

[25]

[26] .

[0203] Appropriate metrics must be used to compare different models. With the help of these, statements can be made regarding the accuracy with which the model describes the target variable.

[0204] Coefficient of determination R 2 This indicates which proportion of the variance of the target variable y the model can describe. The coefficient of determination can be calculated as follows:

[0205]

[0206] in Let y be the estimated target variable for the i-th instance, and y i These are the associated real values. It is the mean. The coefficient of determination can take values ​​between 0 and 1. The closer the coefficient of determination is to 1, the better the model fits the target variable.

[0207] Root mean square error (RMSE) is another statistical measure that can be used to determine the quality of a model. Here, the root mean square distance between the actual and estimated values ​​is calculated:

[0208]

[0209] By squaring the error and then taking its root, RMSE can be interpreted as the standard deviation of the variable being estimated. Here, n is the number of observations, and... This is an estimate of the target variable y. RMSE represents the absolute error, which provides different values ​​depending on the target parameter being examined. Therefore, it makes sense to associate RMSE with the mean:

[0210]

[0211] Therefore, it can be relative to the average true value Calculate the RMSE. This allows for a better assessment of the error of the target variable at different sizes.

[0212] method

[0213] Using the method according to the invention, cell growth, i.e., the timeline of cell density, and certain metabolites, particularly glucose and lactate, can be determined in real time from online process variables during culture, especially on a small culture scale. Therefore, the method according to the invention provides real-time values ​​for process variables that were previously only obtainable offline and not in real time. This represents an improvement over conventional methods for determining the timelines of cell growth and certain metabolites (especially glucose and lactate), because the method according to the invention does not require sampling from the culture medium.

[0214] In a preferred embodiment, the method according to the invention is used to determine cell density, glucose concentration, and lactate concentration in fed-batch cultures of mammalian cells with a culture volume of 300 mL or less from online process variables, wherein the method is performed without sampling, i.e., feedback control sampling.

[0215] The method according to the invention allows for fully automated culture on a small scale, i.e., with a culture volume of 300 mL or less, without the need for sampling, where relevant process variables, such as cell density, cannot be determined online but can only be determined offline.

[0216] The method of this invention is particularly suitable for small-scale monitoring and control of mammalian cell culture.

[0217] According to the method of the present invention, a method is provided for determining viable cell density, glucose, and lactate concentration as target parameters in CHO cell culture, wherein the method employs data-based soft sensors. Machine learning models are used to describe the different target variables.

[0218] This invention is based, in at least part, on the finding that the selection of process variables used for model generation has a significant impact on the quality of the determined target process variables.

[0219] Furthermore, this invention is based, at least in part, on the discovery that the type of partitioning (i.e., allocation) of an existing dataset into training and testing datasets affects model quality.

[0220] Furthermore, the present invention is at least in part based on the discovery that the type of antibody produced influences the selection of optimal target parameters.

[0221] The method according to the invention is described below using 155 exemplary datasets obtained from cultures in an ambr250 system. This should not be construed as limiting the teachings or methods according to the invention, but rather as an exemplary application of the teachings. Other datasets already generated using the same or different culture systems can also be readily used and employed in the method according to the invention.

[0222] Suitable features for 155 datasets were analyzed and examined. Corresponding interpolation strategies were used to map the target parameters, ensuring that the selected model provides values ​​for all target parameters at discrete time points. Model error and model quality were evaluated. The method based on this approach allows for the provision of robust and accurate models for the corresponding target / process variables.

[0223] The antibodies produced by the centralized culture of the dataset have different molecular forms. Table 1 below shows an overview of the various items and molecular forms, as well as the corresponding culture quantities.

[0224] Table 1: Data Overview.

[0225]

[0226] Data relevant to the entire culture process, namely the online parameter set, along with associated dates and timestamps, are used for each culture. The data density for different process values ​​varies over time. These deviations in data density can be attributed to the fact that, due to system limitations, if a measured value is changed in an increment specifically defined for each measurement, only the new data point for the online parameter is recorded. To provide continuous process data and ensure that runs can be compared with each other, the corresponding online parameters are interpolated for any missing timestamps.

[0227] It should be noted that for online process variables, excessive data smoothing can lead to a loss of measured value fluctuations. However, this noise also represents any process-related changes that are occurring and is included as information in the process values. Therefore, it is important not to over-smooth process values ​​and to ensure that changes in the process are accessible even after interpolation.

[0228] The offline data contain varying numbers of analytical values, depending on the number of samples during the culture period (between 8 and 13). Each dataset contains the date and timestamp for each data point, along with the relevant analytical values ​​for the offline parameters.

[0229] A dataset is generated by preprocessing online and offline data through interpolation. This dataset contains the same number of data points for all process variables, regardless of whether they are online or offline. The analysis is based on the interpolated dataset. This interpolation is unnecessary if data points are available at the same frequency and time for all online and offline process variables.

[0230] Due to the preprocessing of available online and offline data, the different time curves of individual process variables caused by different measurement frequencies are standardized into a unified time curve, i.e., a single timeline. Undesirable values ​​caused by technology and process management are identified and deselected or corrected, and existing time gaps are closed, ensuring that the time and quantity of all process variables in one dataset from one culture are consistent with those in all datasets from all cultures.

[0231] To prevent fluctuations in measurement signals caused by turning controls on at the start of culture or off at the end of culture from spoofing model formation, data collected during the first and last 12 hours of culture were not used. In this specific instance, this meant using a time range from day 0.5 to day 13.5. This ensured that changes in process variables could only be attributed to processes within the cell culture. Online data interpolation was performed on the entire dataset. Figure 1 This shows an example of linear interpolation of the process value 'AO.PV'.

[0232] from Figure 1 As can be seen, the process of an online signal with linear interpolation is well described. At the beginning (<0.5 days), we can see how the values ​​measured when control begins fluctuate. Peaks (large changes in process values ​​over a short period) can also be well mapped by this type of interpolation.

[0233] For offline data, the obtained analytical values ​​(VCD, VCV, glucose, lactate) were fitted using three different interpolation methods. Figure 2 Examples of VCD interpolation using different fitting methods are shown.

[0234] Calculate the corresponding coefficient of determination R 2 To evaluate a single interpolation of the VCD. Univariate splines achieve the highest R-value here. 2 The values ​​are high, but tend to overfit significantly. Therefore, univariate splines describe almost every measured value, but do not depict the typical growth curve of the biological system. On the other hand, there is a small difference between Peleg fitting and polynomial fitting. However, Peleg fitting can better describe the different growth stages of the biological system, and is therefore used for interpolation of the VCD target variable

[27] .

[0235] Interpolation of the lactate and glucose curves indicates that univariate splines achieve better R-squared. 2 Mapping offline data and better describing the curves under lactate conditions. Since the polynomial fitting interpolates negative values ​​for lactate starting from day 10, the interpolation of the univariate spline is defined as the target vector y for lactate. However, for glucose, a polynomial fitting (third order) is used to describe the target variable (glucose: univariate spline (R... 2 =0.999) and polynomial fitting (R2 =0.958); Lactic acid: Univariate spline (R 2 =0.999) and polynomial fitting (R 2 =0.959).

[0236] Furthermore, datasets with too few offline data points (three or fewer) used for preprocessing were no longer used for analysis. This was the case for two datasets. Therefore, the entire interpolation and adjustment dataset contained 153 cultures.

[0237] Since the interpolated dataset with a maximum resolution of five minutes contains a large number of data points, the analysis is performed at a resolution of only 1 / 10 of a day to reduce the computational workload. The program can be used for this purpose.

[0238] Figure 3 The dataset from Project 2 (12 cultures) is shown. It can be seen that different interpolation methods (Peleg fitting, univariate spline, and multinomial fitting) have little impact on the correlation strength.

[0239] exist Figure 3 In the scatter plot, online parameters are displayed as features (lines). Columns represent different interpolations of VCD. The ellipses in the scatter plot always contain 95% of the data. The closer the ellipses are, the stronger the linear relationship between the variables. The calculated Bravais-Pearson correlation coefficients are shown in Table 2 below.

[0240] Table 2: Pearson correlation coefficient values ​​for the sample dataset from Project B corresponding to... Figure 3 The value.

[0241]

[0242] Taking the value of "O2.PV" as an example, the calculated interpolation coefficients are very close to each other (0.9547; 0.9490; 0.9490).

[0243] Correlation analysis was then performed on the entire dataset accordingly. Table 3 below shows the Bravais-Pearson correlation coefficients determined in this manner.

[0244] Table 3: Calculate the Pearson correlation coefficient for the entire dataset (153 cultures), with the target variable VCD being a good fit for Peleg.

[0245]

[0246] Correlation analysis with a single ambr250 run (see Table 3 above and...) Figure 3In comparison, correlation analysis showed a significantly weaker linear relationship across the entire dataset. Besides the strength of the correlation, the analysis of the entire dataset also yielded other online parameters as optimal candidates. Correlation between the independent variables was also found. Table 4 below shows the correlations between the parameters “O2.PV” and “N2.PV” and other independent variables.

[0247] Table 4: Correlation between independent variables. Taking O2.PV and N2.PV as examples, they have the highest correlation values ​​in the previous correlation analysis of project B.

[0248]

[0249] When independent variables are correlated with each other, one refers to multicollinearity. This is illustrated in the example using 'O2.PV' here. Figure 3 There is a clear linear relationship between the two best correlation coefficients of 'N2.PV' and 'O2.PV' and the remaining independent parameters.

[0250] Figure 4 Displays the computational information (mutual information) of all features of the target variable VCD in the entire dataset. Figure 4 This displays some available features with high-level information about the VCD target variable. For VCD, the mutual information can therefore have the highest indices as 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'O2.PV', 'N2.PV', and 'LGE.PV'.

[0251] Based on the results of information content calculations and correlation analysis, the ten optimal process variables (CHT.PV, ACOT.PV, FED2T.PV, GEB.PV, CO2T.PV, ACO.PV, AO.PV, LGE.PV, O2.PV, and N2.PV) were selected, and a corresponding feature matrix X was created. This matrix contains interpolated data from the available dataset. A five-minute resolution (f1...f) was selected for the features. 10 The culture duration (in hours) and other information are added as additional columns in the matrix:

[0252]

[0253] The training and test datasets are partitioned in such a way that they are entirely derived from the training in Project 2. The target variable 'VCD' is partitioned based on the distribution of the feature matrix.

[0254] To examine the quality of the obtained models, the relative frequency density of errors was calculated across the entire test dataset. Histograms of predictions for the entire test dataset using MLPRegressor (a), Random Forest (b), and XGBoost (c) models for the target variable VCD are shown on the X-axis as the error between the estimated VCD value and the predicted value, and on the Y-axis as the relative frequency of the error. All three distributions show a left-skewed trend, indicating that VCD is underestimated. Furthermore, examination of all histograms shows that the estimates from all three models produce comparable results. XGBoost shows the most uniform distribution of computational errors; however, an overestimation of the target variable can also be observed here.

[0255] For each model, calculate RMSE and R based on the entire test dataset. 2 Both of these values ​​are related to the Peleg fit of the target variable VCD. The results of the three models are summarized in Table 5 below.

[0256] Table 5: VCD estimation results for MLPRegressor, Random Forest, and XGBoost.

[0257] Model RMSE <![CDATA[R 2 ]]> MLPRegressor 0.38 0.58 Random Forest 0.33 0.67 XGBoost 0.34 0.66

[0258] All models achieved comparable results in terms of RMSE and coefficient of determination.

[0259] If you examine some specific datasets determined using random forests (the optimal model), it becomes apparent that the Peleg fit to the VCD cannot accurately map the entire training period (see...). Figure 5 Starting from day 5, the model in the upper part of the graph fails to accurately depict the relationship between the data and VCDs. The lower part of the graph shows the opposite behavior. The model overestimated VCDs from the beginning and therefore could not provide a sufficiently accurate description of VCDs.

[0260] It has now been surprisingly found that exchanging features in a feature matrix with high information content with those features with significantly lower information content but still having deterministic information content can significantly improve the quality of predictions.

[0261] It has been found that expanding the matrix by using features 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV', as well as removing the redundant feature 'O2.PV' (filled with N2 and O2), leads to improved prediction quality.

[0262] The improved feature matrix contains the following 14 features: 'Time', 'ACO.PV', 'ACOT.PV', 'AO.PV', 'CHT.PV', 'CO2.PV', 'CO2T.PV', 'FED2T.PV', 'FED3T.PV', 'GEW.PV', 'PH.PV', 'N2.PV', 'LGE.PV', and 'OUR.PV'.

[0263] Furthermore, it has been found that the selection or splitting of training and test datasets affects the quality of predictions.

[0264] When comparing the training and test datasets selected for the target variable, it was found that the VCD distribution of the training dataset consisting of the cultivation data from Project 2 has a mean μ. 训练 = 84.60, standard deviation is σ 训练 =48.62, while the mean of the test dataset is μ. 测试 =64.22, standard deviation is σ 测试 =38.02.

[0265] It has been found that using a training dataset from only one project is disadvantageous when making predictions about cells expressing proteins with different structures. It has been found to be advantageous to randomly distribute the training dataset across the entire existing dataset.

[0266] In this example, to distribute the dataset more evenly, 30 random numbers between 0 and 152 were generated (since there are 153 datasets). Each number represents one training run. Random numbers were generated repeatedly until a comparable mean and standard deviation between the test and training datasets could be achieved using the trained model. The final split resulted in μ. 训练 =80.72, σ 训练 =47.11 and μ 测试 =80.11, σ 测试 =48.70, and is used as the split ratio between the two datasets in the further process.

[0267] Therefore, in one embodiment of the method according to the invention, the existing preferred preprocessed dataset is divided into a training dataset and a test dataset, wherein the training dataset is 70%-80% of the total dataset (80% in this example, hence 123 culture runs), and the test dataset contains 20%-30% of the data in the entire dataset (in this example, 30 randomly selected cultures of the entire dataset validated as described above can be used to validate the model).

[0268] The models were then trained and tested using the expanded feature matrix and the new distribution of the dataset. The hyperparameter optimization strategy described above was retained for this purpose. Histograms of the corresponding VCD estimates using the newly partitioned training and test datasets show that the error distributions of all three models narrowed significantly, which can be attributed to more accurate estimations of the target parameters. Figure 6 ).

[0269] All three models achieve a clearer error distribution that fluctuates around the true value of the target variable (the X-axis at 0). Here, the XGBoost histogram also shows the most uniform error distribution. The Random Forest histogram shows small errors across the entire region. Comparing the two histograms (a) and (c), XGBoost estimates the accurate target value more often than MLPRegressor. However, since the MLPRegressor distribution has a lower error width, it can be inferred that the accuracy of the two models is roughly the same.

[0270] Table 6: Results of MLPRegressor, Random Forest, and XGBoost estimations of VCD using the new distributions of the test and training data.

[0271] Model RMSE <![CDATA[R 2 ]]> MLPRegressor 0.23 0.86 Random Forest 0.20 0.89 XGBoost 0.22 0.85

[0272] All three models are able to achieve closely related results. Figure 7 An exemplary estimate using the optimal model with a single culture is shown.

[0273] Therefore, Peleg fitting based on the original data achieved a near-ideal estimate of the target variable. Looking at the entire test dataset, all models were able to utilize the dataset's segmentation ratio as described above in R... 2 Good results were achieved in terms of RMSE.

[0274] The glucose values ​​fitted using a third-order polynomial were used as the target parameter for estimating glucose concentration. The feature matrix used for training contained the same features as the VCD. The same partitioning of the training and test datasets was also used.

[0275] Just like with VCDs, histograms show comparable results in terms of error. Here, XGBoost can typically achieve small errors between actual and estimated values. The Random Forest histogram also shows small errors between the interpolated and estimated values ​​of the target variable, which are evenly distributed around the actual values ​​of glucose. Compared to the other two histograms, MLPRegressor shows the largest error.

[0276] Table 7: Glucose estimation results for MLPRegressor, Random Forest, and XGBoost.

[0277] Model RMSE <![CDATA[R 2 ]]> MLPRegressor 0.13 0.94 Random Forest 0.12 0.95 XGBoost 0.14 0.93

[0278] Figure 9 Two exemplary cultivations obtained using random forest are shown. The target variable is appropriately described with a coefficient of determination of 0.93.

[0279] To estimate lactate concentration, lactate values ​​fitted using a univariate spline method were used as the target parameter. The feature matrix used for training contained the same features as for VCD and glucose. The same partitioning of the training and test datasets was also used. Histograms show the different results regarding the error ( Figure 11 ).

[0280] If we consider the histogram of the MLPRegressor, it's impossible to make estimates with small errors as frequently as the other two models. On the other hand, the distributions of Random Forest and XGBoost are very narrow. It seems that for some estimates of the target variable, very good predictions can be made with almost no error, but these quickly lead to larger errors across the entire test dataset. The neural network has the most uniform error distribution here.

[0281] Table 8 below shows the RMSE and R for all models. 2 The results of lactic acid evaluation.

[0282] Table 8: Lactate estimation results for MLPRegressor, Random Forest, and XGBoost.

[0283] Model RMSE <![CDATA[R 2 <!-- 21 -->]]> MLPRegressor 0.18 0.95 Random Forest 0.35 0.81 XGBoost 0.27 0.88

[0284] Figure 12 The image shows XGBoost predictions for lactate from an exemplary culture obtained from the test dataset. The upper part of the image shows a near-ideal description of the fitted lactate process; in the lower part, the process can be represented with an R-value of 0.98. 2 To describe.

[0285] For validation, a study was first conducted to determine which model best described the interrelationships of features on the test dataset. To this end, the model was initially provided with only ten datasets for training. As the process progressed, the number of datasets per dataset was increased by ten. This resulted in twelve training cycles, during which the model received 10 to 120 datasets. After each training cycle, the target variable was estimated based on the test dataset. The respective RMSE was calculated. As mentioned above, the test dataset also consisted of 30 randomly selected datasets for validation. VCD was chosen as the target variable. This resulted in... Figure 13 The learning behaviors described in the text.

[0286] like Figure 13 As shown, compared to neural networks, both Random Forest and XGBoost achieve smaller errors when predicting the test dataset using a small dataset. However, this effect seems to diminish with increasing training dataset size, thus achieving comparable errors starting from approximately 80 datasets compared to the other two models. Random Forest achieves the lowest RMSE with a maximum of 120 datasets. However, the errors of all models remain within a very narrow range.

[0287] The estimates of 30 trained VCD prediction models on the test dataset were evaluated in detail. Although good results were achieved on the entire dataset (histogram, coefficient of determination, RMSE), some predictions still showed significantly larger biases. Figure 14 This shows a culture run in which the estimated VCD was significantly higher than the actual distribution.

[0288] Insufficient accuracy of estimations has been increasingly observed in the cultures of Project 1 and Project 3. In the cultures of these two projects, the cultured cells produced complex molecular forms.

[0289] The study found that VCDs based on IgG (which have the characteristic Y-shape of natural IgG antibodies or largely retain it) (items 2 and 4) were on average higher in cells with complex molecular forms as target products (items 1 and 3), and the calculated cell diameters were higher for items with complex molecular forms.

[0290] Figure 15 The mean cell diameter and standard deviation for each sample, grouped by Y-shaped IgG (IgG, items 2 and 4) and complex IgG (complex, items 1 and 3), are shown in a block diagram. The diagram shows the green blocks (complex protein format; left at each time point) above the blue blocks (Y-shaped IgG antibodies; right at each time point). At the start of culture, the two molecular forms are still relatively similar. As culture time progresses, cells with the complex molecular form as the target product only become larger. In contrast, cells with the standard antibody can be seen to consistently increase in size until day 7, but no further increase in cell diameter is observed thereafter.

[0291] The results showed that for the IgG form, the relationship between higher VCD and smaller cell diameter, and for the complex protein form, the relationship between smaller VCD and larger cells, led to inaccurate prediction of VCD.

[0292] It was also found that live cell volume (VCV) represents a more suitable target variable than VCD, applicable not only to cultures that produce complex antibody forms but also to cultures that produce Y-shaped IgG antibodies.

[0293] VCV is calculated using the following formula:

[0294]

[0295] Therefore, VCV is a better approximation of viable biomass in culture compared to VCD.

[0296] Since the calculated VCV value, like all other offline parameters, only contains the number of samples, the new target parameters are fitted using a third-order polynomial. The model is then trained and evaluated for the new target size, as described above for the other target parameters.

[0297] RMSE and coefficient of determination were used to evaluate individual models. In summary, the optimal model with 14 features achieved the following results:

[0298] Table 9: Comparison of RMSE and Coefficient of Determination for the Best Model for the Target Variable VCD

[0299] Model RMSE <![CDATA[R 2 ]]> MLPRegressor 0.23 0.86 Random Forest 0.20 0.89 XGBoost 0.22 0.85

[0300] For the target variable VCV, the computational error and coefficient of determination for a single model are summarized in Table 10 below:

[0301] Table 10: Comparison of RMSE and coefficient of determination of the best model for the target variable VCV

[0302] Model RMSE <![CDATA[R 2 ]]> MLPRegressor 0.14 0.94 Random Forest 0.11 0.97 XGBoost 0.12 0.96

[0303] By using VCV as the target variable instead of VCD, all models were able to achieve a coefficient of determination higher than 0.9. Model improvements can be made with lower RMSE and higher R-values. 2 As seen in the value.

[0304] To demonstrate the improved results in comparing live cell density with cell volume, scatter plots representing the estimates across the entire training and test datasets were obtained. Random forest estimation yielded the best results for VCD and VCV. The two scatter plots are shown below. Figure 16 As shown.

[0305] Comparing the two scatter plots, it's clear that VCV's predictions are closer to the ideal estimate, and the distribution of both the test and training datasets is significantly smaller compared to VCD's predictions. Considering only the training data (blue dots), the model learns better the relationship between features and cell volume, rather than the relationship with live cell density. Therefore, these features allow for a more accurate estimation of cell volume across the entire test dataset for all trained models.

[0306] The following study examines the extent to which dividing antibodies into different groups and using only a limited dataset affect the quality of method training.

[0307] If we consider all four items separately for the process of targeting the VCV, we obtain Figure 17 The block diagram shown illustrates this. It can be seen that the VCV in item 4 lies between one side of items 1 and 3 and the other side of item 2. This means that the datasets from items 1, 3, and 4 can also be classified as complex IgG antibody formats (Classification 2). Therefore, the calculation was repeated using this classification. Various combinations of training and test datasets were also tested. The results are shown in Table 11 and... Figure 18 and Figure 19 As shown in the image.

[0308] Table 11: RMSE for different combinations of training and test datasets.

[0309]

[0310] Different combinations show that the prediction using the random forest method achieves the best results, i.e., the lowest RMSE.

[0311] Compared to VCD, RMSE shows a significant improvement (reduction) in all combinations of training or test datasets when VCV is used as the target parameter.

[0312] Different combinations of training and test datasets show that choosing the dataset based on the molecular format affects the RMSE of the target parameters. This combination achieves the highest RMSE when training the model using a dataset in a standard format and estimating VCD or VCV in a complex format. Training with a dataset in a complex molecular format and predicting VCD or VCV results in a smaller RMSE. The minimum RMSE can be achieved by using a mixed dataset for both standard Y-IgG and complex molecular formats.

[0313] Furthermore, the model was evaluated based on estimates from both the training and test datasets to check for overfitting. The training model for the target variable VCV was estimated on both the training and test datasets. The estimates were evaluated based on RMSE, and the differences between the test and training datasets were then displayed as a bar chart. Figure 20 ).

[0314] Figure 20 The error rate of the MLPRegressor on the test dataset is lower than that on the training dataset. Therefore, the calculated difference is negative. Random Forest and XGBoost have larger errors on the test dataset, making the difference shown here positive. Therefore, both decision tree-based models tend to overfit.

[0315] Existing technology

[0316] Existing techniques use parameters such as glucose, lactate, ammonia, and VCD (all of which are offline parameters) as input variables for random forest regression analysis to explain the dynamic behavior of intracellular activities, but not to predict or model offline parameters.

[0317] Compared with the prior art, in this invention, the parameters used for the machine learning model are exclusive online parameters (which are used to control fermentation conditions).

[0318] Therefore, this invention utilizes typical online measurement parameters and statistical models generated throughout the culture process to estimate parameters such as VCV and glucose, without requiring additional sensors or sampling.

[0319] Summary and Outlook

[0320] By interpolating existing online and offline training datasets, a standardized and uniform dataset can be obtained, which is used for model generation to predict target parameters that are only available offline.

[0321] For offline data (considered the target variable in further processes), an interpolation must be found that can representatively describe the process of the corresponding target parameter. Since live cell density is related to the growth process of biological systems, conventional interpolations such as polynomial fitting or univariate spline fitting typically describe the target parameter with only insufficient accuracy. Incorrect extrapolation will lead to an inaccurate description of the target variable. Although the chosen difference produces a result consistent with R... 2 The results were comparable, but the difference chosen by M. Peleg

[27] best described the growth process of cell culture. The background of the interpolation strategy lies in the combination of continuous logic equations used to describe cell growth and mirror logic equations (Fermi equations) used to describe death behavior.

[0322] The results of correlation analysis are only partially unaffected by the choice of interpolation strategy.

[0323] The accuracy of VCD target variable estimation can be improved by adjusting the dataset to an adapted split ratio between the training and test datasets. To this end, a validation dataset was selected based on the distribution of the target variable so that their mean and standard deviation are as small as possible. The goal is not to artificially generate a better dataset for prediction. Instead, it is assumed that the previously generated test dataset is not sufficient to accurately describe the entire dataset. Cross-validation is referenced here as the appropriate method.

[0324] The calculation of cell volume and its correlation with cell size can describe a better approximation of biomass than VCD, thus leading to VCV as a new target parameter.

[0325] Calculated cell volume, as an approximation of biomass, provides a higher level of information about process characteristics than the viable cell density of cultures previously determined by analyzing samples. The average volume of a cell culture can be derived from the measured average cell diameter. This shows that cell size, especially for cells with complex target molecules as products, increases with culture time. However, viable cell density cannot map this relationship. Finally, viable cell volume describes the metabolic activity of cultured cells better than viable cell density.

[0326] To determine the target parameters in real time, estimations should be performed at defined time intervals (e.g., ten minutes). For CHO cells, this interval provides acceptable resolution, as their doubling time is approximately 24 hours.

[0327] The following embodiments and accompanying drawings are for illustrative purposes only. The scope of protection is defined by the pending claims. However, modifications can be made to the disclosed embodiments without departing from the principles of the invention.

[0328] References

[0329] [1]J.Glassey, et al., Biotechnol.J.6(2011)369-377.

[0330] [2] F. Garcia-Ochoa, et al., Biochem. Eng. J. 49 (2010) 289-307.

[0331] [3]E.Trummer,et al.,Biotechnol.Bioeng.94(2006)1033–1044.

[0332] [4] Y.-M. Huang, et al., Biotechnol. Prog. 26 (2010) 1400-1410.

[0333] [5]BC Mulukutla,et al.,Met.Eng.14(2012)138-149.

[0334] [6]R.Luttmann,et al.,Biotechnol.J.7(2012)1040-1048.

[0335] [7]T.Becker and D.Krause,Chem.Ing.Tech.82(2010)429-440.

[0336] [8]LZ Chen,et al.,Bioproc.Biosys.Eng.26(2004)191-195.

[0337] [9]P.Kroll,et al.,Biotechnol.Lett.39(2017)1667-1673。

[0338]

[10] S.Raschka and V.Mirjalili,Machine Learning with Python andScikit-Learn and Tensor-Flow:The comprehensive practice manual for datascience,deep learning and predictive analytics.Frechen:mitp,2.,updated andexpanded edition,2018.

[0339]

[11] Hsu,Chih-Wie,et al.,"A practical guide to support vectorclassification,"Taipei,pp.1-16,2003.

[0340]

[12] R.Kohavi et al.,Ijcai,14(1995)1137-1145.

[0341]

[13] WS McCulloch and W.Pitts,Bull.Math.Biophys.5(1943)115-133.

[0342]

[14] F.Rosenblatt,Psychol.Rev.65(1958)386-408.

[0343]

[15] Kriesel David,ed.,A brief overview of neural networks.

[0344]

[16] W.Lu,“Neural network models for distortional buckling behavior ofcold-formed steel compression members,”2000.

[0345]

[17] R.Rojas,Neural Networks:A Systematic Introduction.Berlin andHeidelberg:Springer,1996.

[0346]

[18] L.Breiman,“Random forests,”Machine Learn.45(2001)5-32。

[0347]

[19] RO Duda,et al.,Pattern Classification.sl:Wiley Interscience,2.Ed.,2012.

[0348]

[20] L.Breiman,Machine Learn.24(1996)123-140。

[0349]

[21] Y.Freund and RE Schapire,J.Comp.Syst.Sci.55:119-139(1997).

[0350]

[22] T.Chen and C.Guestrin,“Xgboost,”in the 22.ACM SIGKDDInternational Conference(B.Krishnapuram,M.Shah,A.Smola,C.Aggarwal,D.Shen,andR.Rastogi,eds.),pp.785-794.

[0351]

[23] L. Fahrmeir, et al., Statistics: The path to data analysis. Springer textbook, Berlin, Heidelberg and sl: Springer Berlin Heidelberg, fourth, improved edition ed., 2003.

[0352]

[24] Kozachenko, LF, et al., Prob. Peredachi Informat. 23 (1987) 9-16.

[0353]

[25] A. Kraskov, et al., Phys. Rev. E 69 (2004) Pt 2, p.066138.

[0354]

[26] BC Ross,PLoS one 9(2014)e87357.

[0355]

[27] M.Peleg,J.Sci.Food Agric.71 / 2(1996)225-230.

[0356] List of abbreviations

[0357] ambr Automated Microbioreactor

[0358] ATP adenosine triphosphate

[0359] Packaging Bootstrap Aggregation

[0360] CHO Chinese hamster ovary

[0361] CIP in-situ cleaning

[0362] FDA (Food and Drug Administration)

[0363] GMP (Good Manufacturing Practice) for Pharmaceuticals

[0364] ANN (Artificial Neural Network)

[0365] MLP Multilevel Perceptor

[0366] NADPH phosphate nicotinamide adenine dinucleotide phosphate

[0367] OUR oxygen uptake rate

[0368] Oxygen delivery rate (OTR)

[0369] PAT process analysis technology

[0370] Random Forest (RF)

[0371] SIP in-situ disinfection

[0372] VCD live cell density

[0373] VCV (Volume of Living Cells)

[0374] XGBoost Extreme Gradient Boosting

[0375] Symbol list

[0376]

[0377]

[0378]

[0379] Material

[0380] software:

[0381] Within the Spyder development environment, the entire project was conducted using the Python programming language. The implementation was done in object-oriented programming. Several classes were written, each implementing an individual task within the project.

[0382] Software used

[0383]

[0384] method

[0385] Data processing

[0386] The entire dataset comprises 155 culture runs. These are broken down into online and offline data. Data processing is implemented using Spyder in the Python programming language. The data is provided in CSV file format. The data is read in using the "csv" library. This allows for quick and easy reading and conversion of data into new data structures within the development environment. The "PIFileParser" class for online data and the "Off-lineDataParser" class for offline data have been implemented.

[0387] interpolation

[0388] Because the data had varying densities, appropriate interpolation was necessary. Linear interpolation and interpolation using moving averages were employed. Both functions were implemented using the "scipy" library: "linear interpolation interp1d" and "moving average convolution". This ensured that the interpolated values ​​always fell between the two original measurements. Therefore, the interpolations were always within the natural fluctuation range of the measured signals of the process variables. Since each process variable had a different timestamp in one file, a separate CSV file was created. "TimelineMapping" contained all start and end times for each culture and was created by another database query. Three different intervals were selected to parse the data:

[0389] • Timestamps of offline data sampling time

[0390] ·1 / 10 day

[0391] Five minutes

[0392] Due to the relatively low data density and non-linear data processes, linear interpolation was not applied to the offline data. Three different interpolation strategies were used for fitting:

[0393] Peleg fitting

[0394] Polynomial fitting

[0395] ·spline

[0396] According to M. Peleg's interpolation method, biological growth maps can be plotted by adding function terms, thus providing a good description of the growth process

[27] . Therefore, the raw data for live cell density conforms to all three interpolation methods. For glucose and lactate, polynomial and spline methods are used for interpolation because no biological behavior is assumed here. The online and offline datasets are merged at different time intervals and saved as CSV files for each culture. Correlation analysis is then performed based on these datasets.

[0397] Correlation analysis

[0398] Correlation analysis uses Perform. Use. Statistical analysis can be performed on datasets. Multivariate statistics are applied to online data (features) related to each target variable (lactic acid, glucose, VCD, VCV). The statistical significance and linearity of the data in describing the target variables are analyzed. According to Bravais-Pearson correlation analysis, the linear relationship between the independent and dependent variables is shown in the form of correlation coefficients.

[0399] Mutual Information

[0400] Another method for determining suitable features has been used in the form of mutual information. When determining features via mutual information, information content is defined, which is included in the independent variable X to describe the target variable Y. Dependencies are computed and implemented using sklearn via "mutual information regression". The information content for each culture is calculated separately based on the size of the dataset with a resolution of five minutes, and then the mean of the values ​​obtained from all cultures is formed.

[0401] Create feature matrix / result vector

[0402] The feature matrix is ​​created based on the results of correlation analysis and statistical evaluation of information content. This can be represented as a matrix, with each column containing a feature and a time point with the corresponding version of that feature. The feature matrix is ​​saved as a Pandas DataFrame. Therefore, a suitable file format can be used to train and test the model.

[0403] Modeling and Evaluation

[0404] With the help of correlation analysis results, a separate dataset was created for each target variable. Dividing the feature matrix into training and testing datasets is necessary for training the model. A complete validation dataset needs to be retained for later online prediction. The training dataset contains 80% of the entire dataset and therefore 123 training runs.

[0405] Since the prediction of all target variables is a constant target parameter, only the regressor is used as the model. Many hyperparameters, varying from model to model, can be used. Therefore, model training is used to tune the hyperparameters to map the target variables as accurately as possible.

[0406] For the training itself, the entire feature matrix is ​​normalized using the standard scaler from the Scikit-Learn library.

[0407] Hyperparameter optimization

[0408] Hyperparameters were optimized using RandomizedSearchCV and GridSearchCV from the Scikit-Learn library. All models were trained using RandomizedSearchCV from the Scikit-Learn library combined with 10x cross-validation on the training dataset. Each region of the hyperparameters was examined to obtain the minimum RMSE. The RandomizedSearch was performed 30 times. Therefore, a different, randomly selected set of hyperparameters was used in each iteration. The hyperparameters of the ten models with the minimum RMSE were output. Then, the hyperparameters of the GridSearch were further refined based on the hyperparameters from the RandomizedSearch. GridSearch was performed again with 10x cross-validation on the dataset. The model with the minimum error (minimum RMSE) was saved and then used to estimate the target variable from the test dataset.

[0409] Multilayer perceptron

[0410] The Scikit-Learn library is used to implement multilayer perceptrons (MLPs). The following list contains the hyperparameters used to train the model:

[0411] • Number of neurons in the input layer

[0412] • Number of neurons in the hidden layer

[0413] • Solver algorithms used to set weights (adam, lbfgs, sgd)

[0414] • Activation functions (identity, logistic, tanh, relu)

[0415] Learning rate

[0416] Maximum number of iterations

[0417] Random Forest

[0418] The Scikit-Learn library also implements random forests. The following candidates can be used as hyperparameters in this optimization:

[0419] Number of decision trees

[0420] • Number of features for each decision tree

[0421] • Maximum depth of decision tree

[0422] Minimum dataset size for creating a new node

[0423] • Method for selecting datasets (bootstrapping = true / false)

[0424] XGBoost

[0425] The XGBoost algorithm is integrated into the project structure through the XGBoost library. The following hyperparameter space corresponds to:

[0426] • Number of regression trees in the set

[0427] Maximum depth of regression tree

[0428] • Learning rate η

[0429] • Number of data sets per decision tree

[0430] Minimum weight of child nodes in a decision tree

[0431] • Gamma error evaluation

[0432] Used as a hyperparameter.

[0433] Model Evaluation

[0434] Model evaluation is primarily achieved by displaying an error histogram. This shows the error (residual) the model has when predicting the actual values ​​of the target parameters on the test dataset.

[0435] Calculate the RMSE of the target parameter estimation accuracy and compare it with the average value of the target parameter.

[0436] To check for overfitting of the model, the RMSE of the entire training dataset and the test dataset was calculated. The difference between the two errors was used as an indicator of model overfitting.

[0437] Overfitting = RMSE 测试 -RMSE 训练

[0438] To further characterize the model quality, the coefficients of determination for the entire test dataset and each culture considered individually were used.

[0439] Example 1

[0440] Ambr250 - Culture

[0441] 155 datasets were collected based on cultures cultured in the ambr250 system. The eukaryotic cells used were CHO cells expressing the target molecules extracellularly. Cultures were performed using a fed-batch process. The ambr system used allowed for up to twelve simultaneous cultures. The master culture duration was 13–14 days. A single-use bioreactor (250 mL) provided the necessary space for the reaction. Pre-cultures were performed in shake flasks and lasted for three weeks. Starting conditions regarding cell volume and quantity at seed were comparable across all bioreactors. The culture media used were chemically defined. Only one batch of culture medium was used per culture.

[0442] To provide optimal culture conditions within this system, numerous process variables can be used. The parameters to be controlled are the pH, temperature, and dissolved oxygen concentration in the culture medium. The table below contains a complete list of all process variables used in this work.

[0443] Table 12: Parameters Measured Online

[0444]

[0445]

[0446] All measurement variables are recorded throughout the culture period using a so-called PI system. The PI system contains only online measurement variables.

[0447] The parameters listed here can be used to monitor optimal culture conditions. BlueSens exhaust gas analysis can also be used for each reactor. This detects the O2 and CO2 content in the exhaust gas stream from the bioreactor, providing another important component for process control. These two measured variables in the exhaust gas stream can be used to determine OUR and OTR.

[0448] Samples were taken daily during the culture period. Then Cedex Bio was used. (Roche Diagnostics GmbH, Mannheim, Germany) Analyzes different concentrations of metabolites and product titers.

[0449] In addition, cell counting measurements are performed. These measurements provide information on viable cell density, total cell density, viability, aggregation rate, and cell diameter. These parameters can be used to infer the growth behavior of the culture. Offline dimensions are obtained via Cedex. (Roche Diagnostics GmbH, Mannheim, Germany) Cell counter measurements. The error of these cell counting and cell analysis systems is within 10%. All offline measurements used are shown in the table below.

[0450] Table 13: Variables measured offline.

[0451]

Claims

1. A method for determining one or more of the following during mammalian cell culture: viable cell density, viable cell volume, glucose concentration in the culture medium, or lactate concentration in the culture medium, comprising the steps of: i) At least determine the online process variable 'time' for cultivation; the temperature of the cooling element in °C, also abbreviated as 'CHT.PV'; the total CO2 in the exhaust gas in volume %, also abbreviated as 'ACOT.PV'; the cumulative feed 2 in ml, also abbreviated as 'FED2T.PV'; the fermenter weight in g, also abbreviated as 'GEW.PV'; the cumulative CO2 inflow in ml, also abbreviated as 'CO2T.PV'; the CO2 concentration in the exhaust gas in volume %, also abbreviated as 'ACO.PV'; the O2 concentration in the exhaust gas in volume %, also abbreviated as 'AO.PV'; the N2 inflow in ml / min, also abbreviated as 'N2.PV'; the cumulative basal addition in g, also abbreviated as 'LGE.PV'; the CO2 inflow in ml / min, also abbreviated as 'CO2.PV'; the feed 3 in ml... Accumulation, also abbreviated as 'FED3T.PV'; in mol / (l h) represents oxygen utilization rate, also abbreviated as 'OUR'; and pH, also expressed as 'PH.PV', is the current online measured value. as well as ii) Using the currently measured values ​​determined in i), determine the viable cell density, viable cell volume, glucose concentration in the culture medium, or lactate concentration in the culture medium using the data-driven model for mammalian cell culture. This model is generated using a feature matrix containing the process variables 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'. The method described therein is performed without sampling and uses only online measurements from the culture. The mammalian cells mentioned therein are CHO cells.

2. A method for adjusting glucose concentration to a target value during mammalian cell culture, comprising the following steps: a) At least the following process variables must be determined: 'Time' (in °C); 'CHT.PV' (temperature of cooling elements); 'ACOT.PV' (total CO2 in the exhaust gas, expressed as % by volume); 'FED2T.PV' (cumulative feed 2, expressed as ml); 'GEW.PV' (fermenter weight, expressed as g); 'CO2T.PV' (cumulative CO2 inflow, expressed as ml); 'ACO.PV' (CO2 concentration in the exhaust gas, expressed as % by volume); 'AO.PV' (O2 concentration in the exhaust gas, expressed as % by volume); 'N2.PV' (N2 inflow, expressed as ml / min); 'LGE.PV' (cumulative basal addition, expressed as g); 'CO2.PV' (CO2 inflow, expressed as ml / min); 'CO2.PV' (cumulative feed 3, expressed as ml). Accumulation, also abbreviated as 'FED3T.PV'; in mol / (l The oxygen utilization rate (h) is also abbreviated as 'OUR'; and pH is also expressed as 'PH.PV', the current online measured value. b) Using the values ​​determined in a), determine the current glucose concentration in the culture medium using the data-driven model for mammalian cell culture, which was generated using a feature matrix containing process variables 'time', 'CHT.PV', 'ACOT.PV', 'FED2T.PV', 'GEW.PV', 'CO2T.PV', 'ACO.PV', 'AO.PV', 'N2.PV', 'LGE.PV', 'CO2.PV', 'FED3T.PV', 'OUR', and 'PH.PV'. as well as c) If the current glucose concentration determined in b) is lower than the target value, add glucose until the target value is reached, thereby adjusting the glucose concentration to the target value. The method described therein is performed without sampling and uses only online measurements from the culture. The mammalian cells mentioned therein are CHO cells.

3. The method according to claim 1 or 2, characterized in that, Data-driven models are generated through machine learning.

4. The method according to claim 3, characterized in that, Data-driven models are generated using methods selected from a group that includes artificial neural networks and ensemble learning.

5. The method according to claim 4, characterized in that, The data-driven model is generated using the random forest method.

6. The method according to claim 4, characterized in that, The data-driven model is generated using the MLPRegressor method.

7. The method according to claim 4, characterized in that, The data-driven model is generated using the XGBoost method.

8. The method according to claim 1 or 2, characterized in that, Data-driven models are generated through supervised learning.

9. The method according to claim 8, characterized in that, Data-driven models are validated through cross-validation.

10. The method according to claim 9, characterized in that, Cross-validation is 10x cross-validation.

11. The method according to claim 1 or 2, characterized in that, The data-driven model is generated using a training dataset that contains at least 10 training runs.

12. The method according to claim 11, characterized in that, The training dataset contains at least 60 culture runs.

13. The method according to claim 3, characterized in that, Approximately 80% of the datasets generated for data-driven models can be used as training datasets, and the remainder can be used as test datasets.

14. The method according to claim 13, characterized in that, a) Randomly divide the datasets available for modeling into training and testing datasets at a ratio between 70:30 and 80:

20. b) Form a model, c) Determine the mean and standard deviation of the process variables used to determine the dataset from the training dataset, and determine the mean and standard deviation of the process variables used to determine the dataset from the test dataset. d) Repeat steps a) to c) until comparable mean and standard deviation are reached, i.e., within 10% of each other relative to the test and training datasets, where the split obtained under a) is different for / with each new run.

15. The method according to claim 14, characterized in that, The datasets used to generate data-driven models each contain the same number of data points.

16. The method according to claim 15, characterized in that, The data points used to generate the data-driven model in the dataset each correspond to the same cultivation time point.

17. The method according to claim 16, characterized in that, Missing data points in the dataset are obtained through interpolation.

18. The method according to claim 17, characterized in that, Missing data points for glucose concentration and / or live cell volume were obtained by fitting a third-order polynomial.

19. The method according to claim 17, characterized in that, Missing data points for lactic acid concentration were obtained through univariate spline fitting.

20. The method according to claim 17, characterized in that, Missing data points for live cell density were obtained by Peleg fitting.

21. The method according to claim 15, characterized in that, Each dataset contains at least one data point every 144 minutes.

22. The method according to claim 15, characterized in that, Each dataset contains at least one data point every 60 minutes.

23. The method according to claim 15, characterized in that, Each dataset contains approximately one data point every 5 to 10 minutes.

24. The method according to claim 1 or 2, characterized in that, Mammalian cells are CHO-K1 cells.

25. The method according to claim 1 or 2, characterized in that, Mammalian cells express and secrete therapeutic proteins.

26. The method according to claim 25, characterized in that, Mammalian cells express and secrete antibodies.

27. The method according to claim 26, characterized in that, The antibody is a monoclonal and / or therapeutic antibody.

28. The method according to claim 26, characterized in that, This antibody is not a standard IgG antibody, i.e., a wild-type full-length four-chain antibody or a complex antibody, i.e., an antibody containing additional antibody and / or non-antibody domains compared to standard antibodies.

29. The method according to claim 1 or 2, characterized in that, The data-driven model is generated using a training dataset that contains only complex IgG culture runs.

30. The method according to claim 1 or 2, characterized in that, The data-driven model is generated using a training dataset that also includes standard IgG culture runs.

31. The method according to claim 1 or 2, characterized in that, Mammalian cells express and secrete complex or standard IgG.

32. The method according to claim 1 or 2, characterized in that, The culture volume is 300 mL or less.

33. The method according to claim 32, characterized in that, The culture volume is 250 mL or less, 200 mL or less, 100 mL or less, 75 mL or less, between 200 and 250 mL, or between 50 and 100 mL.

34. The method according to claim 1 or 2, characterized in that, The culture was carried out using a supplemental feeding and batch culture method.

35. The method according to claim 1 or 2, characterized in that, The cultivation was carried out in a stirred tank reactor.

36. The method according to claim 1 or 2, characterized in that, Immersion aeration is involved in the cultivation process.

37. The method according to claim 1 or 2, characterized in that, The cultivation was carried out in a single-use bioreactor (SUB).

38. The method according to claim 1 or 2, characterized in that, Mammalian cells are cultured in suspension or are mammalian cells that grow in suspension.

39. The method according to claim 1 or 2, characterized in that, Data-driven models are generated through regression analysis.