Methods for predicting process outcomes in bioreactors and for modeling those processes.

The method addresses inefficiencies in bioreactor outcome prediction by using self-learning models to adapt and refine predictions based on historical and current data, enhancing process control and anomaly detection, thereby optimizing bioreactor operations.

JP7852898B2Active Publication Date: 2026-04-28CYTIVA SWEDEN AB
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CYTIVA SWEDEN AB
Filing Date
2018-06-18
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Current methods for predicting outcomes in bioreactors, such as disposable bioreactors, are inefficient and rely heavily on trial and error, lack self-learning frameworks, require multiple parameters, and fail to connect data sources for real-time process monitoring, leading to delayed detection of abnormalities and suboptimal process control.

Method used

A method and system that utilize a process model to predict outcomes by accessing historical and current data from bioreactors, incorporating data from online and offline sensors, and employing self-learning techniques like Kalman filters to adapt and refine predictions, enabling early detection of anomalies and optimizing process parameters.

Benefits of technology

Enables accurate prediction of bioreactor outcomes, reduces the need for manual model building, detects anomalies early, and optimizes processes, leading to time and cost savings in pharmaceutical workflows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852898000004
    Figure 0007852898000004
  • Figure 0007852898000005
    Figure 0007852898000005
  • Figure 0007852898000006
    Figure 0007852898000006
Patent Text Reader

Abstract

The present invention relates to a method for predicting the outcome of a process used to produce a sample in a bioreactor, the process belonging to a category. The method includes selecting (51) a process model based on the category, accessing (53) historically significant data related to past process runs to produce the sample, and accessing (54) current data obtained from a current process run of the process. The obtained current data based on the selected process model includes process strategy data, bioreactor equipment data, data from online sensors, and / or data from offline sensors. The method further includes predicting (62) the outcome of at least one selected parameter of the current process run to produce the sample based on the accessed historically significant data and the current data. The present invention also relates to a method for modeling a process and a control system (10) for controlling a process.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of predicting process outcomes in a bioreactor and modeling such processes, in particular to methods for predicting the outcomes of a process used to produce a sample intended for use in another system, and for modeling such a process. [Background technology]

[0002] Users of disposable bioreactors routinely engage in process development and optimization as part of their research and manufacturing activities, which can take months of work to achieve the best results. Furthermore, abnormalities in the cell culture process (cell therapy or bioprocess) during a run are not detected by instrument use, and without automated remote monitoring and diagnostic environments, human supervision is the only inspection necessary to maintain the batch.

[0003] Digital twins can be used to model processes, such as bioreactors, to predict outcomes from process runs ahead of schedule, provided the digital twin has access to a process model that accurately describes the process run. Users continue to expend considerable effort developing protocols to ensure the process run maximizes cell growth and to optimize the process. This currently relies on trial and error, conventional statistical techniques, and experience.

[0004] Thus, there is a need to develop procedures for creating process models that can serve as digital representations of processes used to prepare samples in biological systems such as bioreactors.

[0005] Cell proliferation is a highly nonlinear process that follows four phases: linear, exponential, stationary, and death. Furthermore, cell proliferation is highly variable and can be influenced in complex ways by numerous known and unknown environmental and genetic factors, resulting in two passaged cell culture batches exhibiting very different proliferation patterns.

[0006] In this context, the ability to accurately predict characteristics such as viable cell concentration, total cell concentration, products, or metabolites for any cell culture setup several days in advance, and furthermore, to identify patterns in cell proliferation, can lead to improved logistics and optimized pharmaceutical workflows.

[0007] Currently, methods used to solve similar problems include using standardized tools such as SIMCA. While these are broad-purpose tools, they have several weaknesses. 1) Tools like SIMCA use standard statistical methods, a. Not customized to solve a specific problem. b. It lacks a self-learning framework. 2) Additionally, these tools typically require multiple parameters.

[0008] The evaluation of data from bioprocesses is typically performed offline and post-run. Furthermore, the evaluation is not conducted in comparison to "expected outcomes." One reason for this is the lack of connectivity to data from different sources, such as log data, in-process controls, product quality data, and others. Another reason is the inability to adequately model the execution of the bioprocess. [Overview of the Initiative] [Problems that the invention aims to solve]

[0009] The purpose of this disclosure is to provide methods, devices, and computer programs configured to perform methods, which endeavor to mitigate, alleviate, or eliminate one or more of the defects and shortcomings identified above in the art, either individually or in any combination. [Means for solving the problem]

[0010] The objective is a method for predicting the outcome of a process used to produce a sample in a bioreactor, wherein the process belongs to a category, and this is achieved by the method. The method includes the steps of selecting a process model based on a category, accessing historically significant data related to past process runs for producing a sample, and accessing current data obtained from a current process run of the process. Based on the selected process model, the current data obtained includes process strategy data, bioreactor equipment data, data from online sensors, and / or data from offline sensors. The method further includes the step of predicting the outcome of at least one selected parameter of the current process run for producing a sample, based on the accessed historically significant data and current data.

[0011] The advantage is that undesirable characteristics of the selected parameters can be detected ahead of schedule, allowing for the initiation of actions that may affect the outcome.

[0012] The objective is further a method for modeling a process used to produce a sample in a bioreactor, wherein the process belongs to a category, and this is achieved by the method. The method includes the steps of selecting a process model based on a category, accessing history-important data related to past process runs for producing a sample, and accessing current data obtained from the current process run of the process. Based on the selected process model, the current data includes process strategy data, bioreactor equipment data, data from online sensors, and / or data from offline sensors. The method further includes the steps of predicting an outcome, at least one parameter of the current process run for producing a sample, and updating the process model based on the history-important data and at least one monitored parameter when the current process run is completed.

[0013] The advantage is that the process model used to model the process is automatically updated based on the results from previous process runs.

[0014] The object is furthermore a control system for controlling a process used for manufacturing a sample in a bioreactor, the process belonging to a category, which is achieved by the control system. The control system is configured to simulate the process, to select a process model based on the category, to access historical important data associated with past process runs for manufacturing the sample, and to access current data obtained from the current process run of the process. The data obtained, based on the selected process model, includes process strategy data, bioreactor equipment data, data from online sensors, and / or data from offline sensors. The control unit is further configured to predict the outcome of at least one selected parameter of the current process run for manufacturing the sample, and to control the process used for manufacturing the sample in the bioreactor based on the predicted outcome of at least one selected parameter of the current process run.

[0015] Further objects and advantages can be obtained by those skilled in the art from the detailed description.

Brief Description of the Drawings

[0016] [Figure 1] FIG. illustrating a control arrangement configuration for modeling a process used for manufacturing a sample in a bioreactor. [Figure 2] FIG. illustrating extracting current data from the current process run into a database. [Figure 3] FIG. showing an example of how different processes are categorized to assign a suitable process model to the process. [Figure 4] FIG. illustrating a system for generating a product based on a sample created in a cell culture process. [Figure 5]A diagram illustrating a flowchart for adaptive modeling of a process used to produce a sample in a bioreactor. [Figure 6] A diagram illustrating a flowchart for predicting the outcome of a process used to produce a sample in a bioreactor. [Figure 7] A diagram illustrating a flowchart for predicting features prior to producing a sample in a bioreactor. [Figure 8a] A diagram illustrating different steps in a process for predicting features prior to using the process described in relation to Figure 7. [Figure 8b] A diagram illustrating different steps in a process for predicting features prior to using the process described in relation to Figure 7. [Figure 8c] A diagram illustrating different steps in a process for predicting features prior to using the process described in relation to Figure 7. [Figure 9a] A diagram illustrating prior bioprocess control and data evaluation. [Figure 9b] A diagram illustrating model-based bioprocess control and data evaluation. **DETAILED DESCRIPTION OF THE INVENTION**

[0017] The term "process model" refers to a uniquely developed model of a cell culture process that can predict the outcome of interest and enable "what if" analysis for process optimization.

[0018] The term "feed" refers to a solution added to the culture to prevent nutrient depletion.

[0019] The term "medium" refers to a basal liquid or gel designed to support cell growth. A typical medium consists of amino acids, vitamins, inorganic salts, glucose, serum, and others.

[0020] The term "cell line" refers to a cell culture consisting of cells developed from a single cell and, therefore, possessing a uniform genetic structure.

[0021] The term "clone" refers to an organism or cell, or a group of organisms or cells, that are produced asexually from a single ancestor or lineage, and that organism or cell is genetically identical to its ancestor or lineage.

[0022] The term "outcome" refers to the measurable output / product of a cell culture. This can be cells, proteins, or by-products such as lactates, ammonium compounds, and others.

[0023] The term "strategy" refers to a protocol for process parameters such as the feed regime, instrument setpoints (e.g., pH, DO, CO2), and other parameters.

[0024] The term "supplement" refers to additional nutrients added separately from the feed and basal culture medium.

[0025] In the context of chromatography methods, the term "capture" refers to the initial chromatography step in which a large amount of the target compound is captured, or, in the case of a flow-through process, a large amount of impurities are captured.

[0026] Digital representations of processes used to produce samples in biological systems such as bioreactors are desirable because they allow for the evaluation and improvement of the processes before they are used to produce samples. A consequence of this premise is that the associated process models used as digital representations may require self-learning capabilities and novel analytical theories to faithfully represent the biological system.

[0027] The advantage of digital representation is that process outcomes, such as cell viability, cell count, product titer, product quality, etc., can be predicted. There is no direct causal relationship between instrument parameters (e.g., pH, rocking rate, rocking angle, temperature, oxygen / CO2 control, etc.), user control factors (e.g., feed, feed strategy, culture medium, clones, glucose storage solution, etc.), and these outcomes.

[0028] To obtain good predictions from digital representations (i.e., process models), outcomes must be modeled as a function of instrument parameters and parameters measured using offline / online sensors in a biological system, e.g., a cell culture process, during the current process run. Examples of parameters measured include pH, dissolved O2, CO2, glucose, glutamine, glutamate, lactate, ammonium, sodium ions, potassium ions, gravimetric osmolality, and others.

[0029] Recommended instrument parameters depend on the type of reactor used (e.g., shaking flask or stirring tank) and include pH, oscillation speed, oscillation angle, stirring speed, impeller speed, temperature, oxygen / CO2 control, aeration rate, and others. While studies have been conducted on which instrument parameters affect cell proliferation, to date, there is no comprehensive model to find optimal values ​​for instrument parameters and other related parameters such as the calculated amount of supplements provided, feed rate, feed strategy, and others. In conventional techniques, optimal settings are typically achieved after considerable process development and optimization effort. However, if analytical methods are implemented by inevitably including domain knowledge, process data, and historically significant data—that is, by integrating domain knowledge into a statistical understanding of process data—the effort required to arrive at optimal instrument parameters and perform "hypothetical" analysis will be significantly reduced.

[0030] Thus, the disclosed process models are inherently analytical because they self-learn from past process runs, or historically significant process runs, to predict outcomes from the current process run. The models can further refine their outcome predictions by incorporating information from soft sensors in the form of line sensors and from commercially available asset performance management solutions. These analytical models can automatically fine-tune their predictions of the current process for sample preparation, for example, by using a Kalman filter.

[0031] Furthermore, in conventional systems, anomalies, such as contaminants and metabolites outside boundary values, are detected only through human supervision. The disclosed process model can detect patterns from offline measurements performed, thereby enabling early detection of anomalies. This information can be used by the operator to take necessary corrective actions or a combination of actions to improve yield from the process run.

[0032] Figure 1 illustrates a control configuration for adaptive modeling of processes used to produce samples in biological processes, such as a bioreactor used in a cell culture process 11. The control configuration comprises a control system 10 configured to model a cell culture process and further configured to select a process model, access history-critical data, access current data from a current process run, and predict outcomes for at least one selected parameter of the current process run for producing a sample.

[0033] The process model is selected based on the type of process used to produce the sample in the cell culture process 11. A system for categorizing different processes is disclosed in relation to Figure 3, and each process is assigned to a category and stored in a database for the process model 12. Provenance-important data related to past process runs for producing the sample is accessed from the database with the provenance-important data 13, and current data is accessed from the current process run of the process, as described in relation to Figure 2. Provenance-important data includes data from completed process runs or experiments, and typical data is similar to the data obtained from the current process run.

[0034] As described above, the control unit is further configured to predict the outcome of at least one selected parameter, which may include cell viability, cell count, product titer, product quality, and others.

[0035] In some embodiments, database 14 is used to organize and consolidate historically significant data, current data, and data related to process models. In addition to data related to process models, all data related to current and past processes is organized, consolidated, and stored in one place for easy access.

[0036] The control unit 10 may also be configured to control the process used to produce the sample in the bioreactor based on the predicted outcome of at least one selected parameter of the current process run, as indicated by the dashed arrow 15.

[0037] Figure 2 illustrates the extraction of current data from the current process run into database 14. The current data is obtained from the cell culture process 11 and, based on a selected process model from the process model database 12, includes process strategy data 21, bioreactor data 22 (including instrument data and data from online sensors), and / or data 23 from offline measurements.

[0038] Process strategy data 21 includes strategic information regarding the process, and therefore includes, for example, culture medium, feed type, feed regime, supplements (type and concentration), and supplement regime, etc. Bioreactor data 22 includes process-related data from bioreactor equipment (e.g., stirring, aeration, etc.) and any available online sensors attached to the bioreactor (e.g., pH, dissolved O2, etc.). Offline measurement data 23 includes data related to process samples measured on offline sensors (e.g., data for cell count, product titer, metabolite concentration, gas partial pressure, etc.).

[0039] When extracting data from the cell culture process 11, this is done based on a selected process model to optimize the resources required to obtain the necessary data.

[0040] Figure 3 illustrates an example of how different processes are categorized to assign a suitable process model to each process. The system includes several levels representing the details of the categorization being performed, such as default level 30, first level 31, second level 32, third level 33, and others. This information is stored in a database for process models 12. It should be noted that the following description is merely an example of how categorization may be applied, and other features for category selection, such as scale-up / scale-down, may also be used.

[0041] In the example, default level 30, default process model "D" is the basic process model used and is assigned to any process that has not been previously categorized. The first level 31 is, in this example, cell lines C1, ..., C n This represents a process model based on the following, and different cell lines may require a process model that can be adapted to correctly predict outcomes. Each process model, in turn, is exemplified in Level 32, for example, reactor type R1, ..., R k Based on this, further adaptation is possible. In this example, cell line C1 and reactor type R k Some of the process models for this are illustrated in the third level 33, with media M1, ..., M j It has been further adapted based on that.

[0042] Depending on the process, the corresponding process model is selected by the control system 10 and used for the current process run. The process for categorizing processes into suitable process models is described in more detail in relation to Figures 5 and 6.

[0043] Figure 4 illustrates an example embodiment of a system 40 for generating products based on samples created in the cell culture process 11. The control system has access to an organized and integrated database 14, which is described in relation to Figure 1 and contains all data related to current and historically significant process runs.

[0044] In the context of a chromatography system, the term “sample” refers to a liquid containing two or more compounds to be separated. In this context, the terms “compound” or “product” are used in a broad sense to refer to any entity, such as a molecule, chemical compound, cell, or other. The terms “target compound” or “target product” as used herein mean any compound that is desired to be separated from a liquid containing one or more additional compounds. Thus, the “target compound” may be, for example, a compound desired as a drug, diagnostic agent, or vaccine, or alternatively, a contaminating or undesirable compound that should be removed from one or more desired compounds.

[0045] System 40 further includes a capture step, as exemplified in this example by a continuous chromatography system 41, into which a sample from the cell culture process 11 is fed. The sample contains the target product. The continuous chromatography system 41 captures the target product 42 to be delivered. The continuous chromatography system 41 measures several parameters to obtain an efficient, high-quality manufacturing process, and information 43 regarding impurities, product quality, and other matters may be transferred to the control system 10. This information may be used to further adapt the process model to increase the performance of the complete system, resulting in improved yields, and to control the cell culture process 11.

[0046] Figure 5 illustrates an adaptive model of the processes used to produce a sample in a bioreactor, illustrating a flowchart for adaptive modeling where the processes belong to categories. Each process used to produce the sample is assigned to a default category if no category is assigned to the current process.

[0047] The process begins at step 50, and in step 51, a process model is selected based on category, as illustrated in relation to Figure 3. The selected process model defines the relevant data that will be obtained from the cell culture process 11.

[0048] Step 52 is an optional step, as illustrated in relation to Figure 1, which involves organizing and integrating historically significant data, current data, and data related to the process model in the database.

[0049] In step 53, historically significant data is accessed either from a separate database or within a consolidated database. This historically significant data is related to past process runs used to prepare the sample and includes data from completed process runs or experiments.

[0050] In step 53, current data is accessed either directly from the cell culture process or from a database that is consolidated. Current data includes process strategy data, bioreactor instrument data, data from online sensors, and / or data from offline sensors, as illustrated in relation to Figure 2.

[0051] The data accessed is based on the selected process model, and in step 54, the current data is obtained from the current process run of the cell culture process. This step allows for the fitting of parameters with additional or less process data in comparison with the data required by the process model.

[0052] In some embodiments, input parameters are selected in step 54a. Input parameters are used that are either automatically based on available data and process models, or user-specified, and the system user selects parameters in addition to the required parameters as needed by the process model.

[0053] In some embodiments, the data of the selected parameters is read in step 54b so that it is available for further processing.

[0054] Step 55 is an optional step in which missing data in the current data obtained from the current process run is addressed, allowing the process model to perform data imputation when it encounters missing data.

[0055] In some embodiments, missing data values ​​are replaced in step 55a by imputed values ​​based on historical trends, interpolation, and predictions, based on other available data for the parameters.

[0056] According to some embodiments, data with missing data values ​​is removed in step 55b, i.e., data with missing values ​​is cleared. Preferably, data is removed only when it is determined that it does not substantially affect the prediction.

[0057] In step 56, at least one parameter of the current process run for producing the sample is monitored, and the process model is adapted in real time based on the historically significant data and / or the monitored at least one parameter as the current process run is completed. This step provides the process model with the ability to adapt it to different cell lines, clones, reactor types, media, and other factors. This reduces the need to manually build a process model for each variant. Self-learning helps to adjust the predictive errors of the process model for improved accuracy.

[0058] The purpose of this step is to train and update the process model by either updating the process model based on data from completed process runs or experiments, or by creating a new process model. Techniques used for self-learning include Kalman filters, fuzzy logic, and others.

[0059] According to some embodiments, the step of adapting the process model further includes step 56a, which involves updating the process model for a category using historically significant data for better prediction or forecasting.

[0060] According to some embodiments, the process used in the current process run is determined to belong to a new category, and the step of adapting the process model further includes a step of creating a new process model (step 56b) by assigning the process to the new category and storing the process model as a new process model.

[0061] The process ends in step 57.

[0062] The process described in relation to Figure 5 may be implemented in the form of a computer program for modeling the process used to produce a sample in a bioreactor. The computer program includes instructions that, when executed on at least one processor, cause at least one processor to perform the method described in Figure 5. The computer program may be stored on a computer-readable storage medium.

[0063] Figure 6 illustrates a flowchart for predicting the outcomes of processes used to produce a sample in a bioreactor, where each process belongs to a category. Each process used to produce the sample is assigned to a default category if no category is currently assigned to the process.

[0064] The process begins at 60 and continues through steps 51, 52, 53, 54, and 55, which are described in relation to Figure 5. These steps include:

[0065] Step 51 - Select a process model based on category.

[0066] Optional Step 52 - A step to organize and integrate historically significant data, current data, and data related to the process model in the database.

[0067] Step 53 - A step to access historically significant data related to past process runs for producing a sample in the step, and to access current data obtained from the current process run of the process.

[0068] Step 54 - A step to acquire current data based on the selected process model. Current data includes process strategy data, bioreactor equipment data, data from online sensors, and / or data from offline sensors. According to some embodiments, input parameters are selected in step 54a. Input parameters are used that are either automatically based on the available data and process model, or user-specified. According to some embodiments, the data for the selected parameters is read out in step 54b so that it is available for further processing.

[0069] Optional step 55 - A step to address missing data in the current data obtained from the current process run and to enable the process model to perform data imputation when it encounters missing data. According to some embodiments, missing data values ​​are replaced in step 55a by imputed values ​​based on historical trends, interpolation, and predictions based on other available data for the parameters. According to some embodiments, data with missing data values ​​is removed in step 55b, i.e., data with missing values ​​is cleared.

[0070] In an optional step 61, at least one parameter of the current process run for preparing the sample is monitored, and the process model is adapted based on the at least one parameter monitored during the current process run. This step provides the process model with the ability to adapt it to different cell lines, clones, reactor types, media, and other factors. This reduces the need to manually build a process model for each variant. Self-learning helps to adjust the predictive errors of the process model for improved accuracy.

[0071] The purpose of this step is to train and update the process model, either temporarily or permanently, based on data obtained from process runs or experiments. Techniques used for self-learning include Kalman filters, fuzzy logic, and others.

[0072] In some embodiments, the step of adapting the process model further includes the step of updating the process model for a category by using new data from the current process run and temporarily updating the process model as a “new” process model for the category, step 61a, i.e., applying the updated process model when predicting outcomes from the current process run. At the end of the process run, the original process model is restored for the category.

[0073] In some embodiments, the “new” process model is permanently stored for the category in step 61b, i.e., the updated process model is applied when predicting outcomes in future process runs that use processes belonging to this category.

[0074] According to some embodiments, the process used in the current process run is determined to belong to a new category, and the step of adapting the process run further includes the step of assigning the process to the new category and the step of storing the process model as a new process model.

[0075] In addition to refining the process model, actual predictions during process runs can be updated based on measured values ​​versus predicted values.

[0076] The process continues to the final step 62, where the outcome of at least one selected parameter of the current process run for producing the sample is predicted based on historically significant data and current data accessed.

[0077] According to some embodiments, the step of predicting an outcome for at least one selected parameter further includes at least one of the following. - Step 63: Predict the outcome of at least one parameter. The outcome of interest is predicted ahead of schedule until the end of the current process run. - Step 64: Predicting anomalies. Anomalies, such as metabolite concentrations outside the expected range over 24 hours if no adjustments are made to the cell culture process, are predicted in advance of the scheduled time. - Step 65: Determine and recommend actions to obtain conditions for improvement over the current process run. Define the conditions under which the process can be optimized.

[0078] The process described in relation to Figure 6 may be implemented in the form of a computer program for predicting the outcome, the current process run for producing a sample in a bioreactor. The computer program includes instructions that cause at least one processor to perform the method described in Figure 6 when executed on at least one processor.

[0079] Computer programs can be stored on computer-readable storage media.

[0080] As mentioned above, current solutions for predicting cell proliferation in bioreactors have several weaknesses. In contrast, our method learns over time to produce the best possible prediction.

[0081] Furthermore, the disclosed process is a lean approach. It can operate with a minimal set of parameters, but more parameters can be used if available.

[0082] In summary, a method and system are provided that have an autoevolutionary data-based approach to learning patterns of cell proliferation. a) Prediction of various outcome parameters / metabolites of the current process run, such as live cell concentration and product concentration. b) Identifying abnormal growth patterns, and c) To learn the relationship between cell proliferation outcome parameters and various experimental control methods. It can be used for this purpose.

[0083] A system for predicting characteristics such as viable cell concentration, total cell concentration, products, or metabolites will be disclosed a few days in advance. This is a stepwise approach in which coarser predictions are made in earlier steps and refined in later steps.

[0084] The system is required to have access to a database that contains historical data of output and input features from past process runs, which are stored in the database, indexed by experiment IDs, either in a raw format, or potentially as a knowledge tree based on some distance / proximity criterion such as Euclidean distance.

[0085] Once the current process run is complete, that process run will also be added to this database (in the appropriate place, if it is a tree).

[0086] Figure 7 illustrates a flowchart for an alternative process to predict characteristics in advance when preparing a sample in a bioreactor. The process is a curve-development method based on a sequence of steps to improve the likelihood of predicting outcomes or characteristics of the process for preparing a sample in a bioreactor during the process run.

[0087] The process is illustrated in conjunction with Figures 8a to 8c and has four steps in the form of a stepwise method, with rougher predictions made in earlier steps and refined in later steps.

[0088] Step 1) A step to fit the overall base model, as illustrated in Figure 8a, to the history-important data. This is a base model that takes a simple mean / median / polynomial fit of the history-important data in the database versus culture time. According to some embodiments, all process runs in the database are used. According to some embodiments, to give more relevant estimates, a subset of process runs only is used that satisfy certain conditions, i.e., prior knowledge (e.g., only experiments in the database that use the same medium as the current process run are selected).

[0089] Step 2) A step to improve the base model based on real-time curve development, as illustrated in Figure 8b. At any time period T from the start of the process run while the current process run is in progress, this step includes running a time-series pattern matching of the current process run up to the current time for all process runs in the database up to time T, and identifying the closest K matches that satisfy a fixed distance threshold. Now, all predictions from T to the end of the batch can be updated by repeating the Step 1 methodology, but only for the K closest matches. This ensures that the predictions issued by Step 1 are corrected to a more accurate number by learning from other historical process runs. This step is repeated for every T.

[0090] Step 3) The step of building the model from the feedback error, as illustrated in Figure 8c. At time T, the predictions made for time T+ after applying Step 2 are further refined by building the model based on the historical process runs in database f (or a subset selected from Step 2) to predict how the errors that will occur in all predictions up to the current time of the experiment can predict the errors that will occur in the future, and updating the current predictions made after Step 2 with these errors.

[0091] Step 4) Adding covariate information to further refine the prediction (optional step). Specifically, this step involves building a model of error g after applying Step 3 to metabolite information using measured metabolites such as glucose, lactate, and others, thereby further refining the prediction.

[0092] Predicting how cell proliferation / cellular product production will progress several days in advance is a difficult problem because cells are complex and influenced by numerous factors. The challenges involved include identifying abnormal proliferation patterns in advance, determining whether recovery from abnormal proliferation is possible by operationally altering the cellular environment, and if so, how.

[0093] The system described herein is a continuously learning predictive system that makes accurate predictions by learning from other historical process runs in the past, and even from previous data points of the current experiment in a feedback mode.

[0094] The step of learning from other process runs (Model 2) can also be used for other related functions, such as identifying bad or irrecoverable process runs and emerging anomalies by comparing them with other flagged anomaly process runs in the past, or deviations from normal process runs.

[0095] Furthermore, this system could be useful for experimenters for logistical purposes, as it could be used to predict interesting experimental outcomes such as the time when cell proliferation peaks, the time when cell viability reaches a certain threshold, and so on.

[0096] The systems described herein can be further extended to understand how cell proliferation is influenced by the parameters being measured and in the experimental design, potentially by considering multidimensional cluster patterns.

[0097] Step 71 is an optional step in which a base model for a process run and conditions for the BM are set. If this step is omitted, the base model is created based on historically significant data from all previous process runs. In the optional step 72, historically significant data that matches the conditions set in step 71 is retrieved from previous runs (usually stored in a system-accessible database).

[0098] As explained below, the generation of a base model requires the provision of historically significant data, and if the conditions are too strict as an example, it may be difficult to find relevant process runs in the historically significant data. This is checked in the optional step 73a. If negative, the process continues to step 73b, in which the conditions for the base model are updated before new historically significant data is acquired in step 72.

[0099] On the other hand, if a sufficient predetermined amount of historically significant data is obtained, the process continues to step 74 to create a model, referred to in this disclosure as the “base model,” based on the historically significant data. The historically significant data may be selected based on the conditions set in step 71.

[0100] In short, step 74 builds a rough model based on mean / median / arbitrary curve fitting, indicated by 80 in Figure 8a, of all data similar to the current experiment for which predictions must be made. Similarity can be similar conditions (similar reactor, medium, cell line) or a more elaborate measure of similarity.

[0101] For example, a prediction is made for the progress of a batch of a specific cell line (e.g., CHO cell line) grown in a 5-liter agitated bioreactor in a separate medium, starting on day "0". All previous process runs in the history database that have similar cell lines, are grown in similar media, and are grown in similar bioreactors are considered and selected. This sub-selected set of historically significant process runs is called E1. The mean / median / polynomial fit of all these process runs E1 is calculated to obtain an estimate of the prediction for the current process run, indicated by reference number 80 in Figures 8a and 8b.

[0102] It should be noted that this "similarity" can be defined by any means, and that its similarity threshold may be looser if a set of previous process runs from historically significant data does not match the conditions of the current process run, or it may have a stricter threshold if a run batch with the same conditions exists in the past.

[0103] This similarity threshold may also depend on the stage of the pharmaceutical workflow, with a stricter similarity threshold being desired in manufacturing and a looser threshold being desired in process development.

[0104] In step 75, as the current process run progresses over time, the set of similar process runs used to build the model continues to change. This helps to continuously refine the predictions.

[0105] Continuing the above example, a set of similar process runs E1 is selected to build the base model, while in step 75, all previous process runs in E1 that are closest in time to the current process run's time series up to that individual day are selected, and this model is called the Base Model with Learning (BMWL).

[0106] For example, when it is day "5", there are five days in the current process run. If a forecast for days 6 through 10 is desired, a subset of E1 is considered, where the time series (univariate or multivariate) of that subset from day 0 to day 5 is closest to the current process run. This subset is called E2, and then, in order to generate a forecast for the current process run from day 6 to day 10, the mean / median / arbitrary curve fit is calculated using only the process runs in E2, as indicated by reference number 81 in Figure 8b. This gives a finer forecast than the base model in step 74 because the set of process runs used to build the model is revised.

[0107] Please note that this list of process runs may change over time, as the current process runs will progress.

[0108] In step 76, the prediction issued by step 75 is reviewed to create a more detailed prediction. The error caused by the prediction, as explained above, is fed back to improve the future prediction at the next time instant. In other words, a model in the form of the following equation is built. ε T+LookAhead = f(ε T,T-1..0 ) However, ε T = actual value T - BMWL prediction T represents the error between the actual data and the prediction from the base model with learning, BMWL, at time T during the historically important process runs. This error is used to update the current prediction made using BMWL as follows. BMWL(+EC) prediction T+LookAhead = BMWL prediction T+LookAhead +ε T+LookAhead BMWL(+EC) represents the base model with learning and error correction.

[0109] In step 77, which is an optional step, residual error correction using metabolite information is applied. This step reviews the prediction from step 76 by regressing the remaining error between the step 76 prediction and the actual data for measured metabolite information such as glucose, lactate, ammonia, and others. The metabolite information is preferably measured beforehand and stored in a database accessible to the system. In other words, Model "step 77" prediction T+LookAhead = BMWL(+EC) prediction T+LookAhead +α T+LookAhead where α T+LookAhead = g(metabolite information T,T-1..0 ), α T = actual value T - BMWL(+EC) prediction T is.

[0110] In step 78, data from the current process run is stored as historically significant data for future process runs. Historically significant data is typically stored in a system-accessible database.

[0111] The technical advantage of the method described above is that it allows a system that continuously learns from new data to ensure that it can predict output features as accurately as possible. Since this is automated as a whole, it further eliminates the need to manually build and update custom solutions / models for each type of bioreactor or each experimental setup. As the system is continuously fed more data, it continues to become more powerful over time, as its learning repository grows larger.

[0112] The commercial advantage is that doing so makes the following possible: a) Time and cost savings in process development and manufacturing workflows in the pharmaceutical industry. b) Faster detection of anomalies that may help preserve the batch by adding necessary metabolites or making decisions regarding faster termination of the experiment, resulting in cost and labor savings. c) Better experimental design and design of the experiment d) This self-learning system could potentially be helpful even in emerging fields such as predicting cell therapy outcomes.

[0113] The process described in relation to Figure 7 may be implemented in the form of a computer program for predicting the outcome, the current process run for producing a sample in a bioreactor. The computer program includes instructions that, when executed on at least one processor, cause at least one processor to perform the method described in Figure 7. The computer program may be stored on a computer-readable storage medium.

[0114] (Examples) In the following sections, the average results across 100 test data with 75:25 CV splits for different treatments are presented.

[0115] In a retrospective analysis of 20 experiments for three different output features (live cell count, total cell count, and product titer), the inventors were able to achieve the accuracy listed in Tables 1 to 3 below for a look-ahead of 1 to 5 days. The metric used here for measurement is percentage points, and the error between the actual feature and the predicted feature is ≤20%.

[0116] Every step of the model improves accuracy in accordance with the learning method's strategy.

[0117] [Table 1]

[0118] Table 1 - Live cell count

[0119] [Table 2]

[0120] Table 2 - Total cell count

[0121] [Table 3]

[0122] Table 3 - Product titer

[0123] During a process run, in the process used to prepare a sample in a bioreactor, a method for predicting characteristics can be embodied as follows: A value related to the characteristics is continuously measured during the current process run, and the method is: - Step 74 involves creating a model for the current process run based on a selection of historically significant data, - Step 75 involves selecting the best fitting historically significant data related to the current process run after a predetermined time period, and updating the model based on the best fitting historically significant data. - Step 76: Review the predictions from the updated model based on the calculated error between the measured values ​​and the updated model. Includes.

[0124] A method, according to some embodiments, further comprising the step of performing residual error correction 77 in a revised forecast using metabolite information.

[0125] A method, according to some embodiments, further comprising step 71 of setting a set of conditions for a model, and step 72 of obtaining a predetermined amount of history-important data from a previous process run in order to form a selection of history-important data to be used to create the model.

[0126] According to some embodiments, the method further includes step 73a, which controls whether a predetermined amount of historically significant data obtained from a previous run falls within a predetermined interval; and step 73b, which, if the predetermined amount of historically significant data falls outside the predetermined interval, updates the conditions for the model and repeats step 72 to obtain the predetermined amount of historically significant data, or, if the predetermined amount of historically significant data falls within the predetermined interval, proceeds to step 74 to create a model.

[0127] According to some embodiments, a predetermined interval is selected such that it consists of historically significant data from at least 10 previous process runs.

[0128] According to some embodiments, a predetermined interval is selected such that it consists of historically significant data from 100 or fewer previous process runs.

[0129] According to some embodiments, the method further includes the step of storing data from the current process run as historically significant data for future runs.

[0130] As mentioned above, the evaluation of data from bioprocesses is typically performed offline and post-run, and not in comparison to expected outcomes. To provide an improved process, the following is implemented: - Connect all data sources in the bioprocess (as described in more detail below). - Establish process models that represent both the "cell" (metabolism, division and proliferation, product formation, etc.) and the physical environment (mass transport (e.g., oxygen), shearing, mixing, etc.). The model will be a build-up from a limited set of equations. The first process run will be used to determine the model constants, and the model will "self-learn / adapt" during subsequent process runs.

[0131] The model will be used with online data to perform online evaluation of the process and to build soft sensors, i.e., sensors implemented in the form of algorithms using measured data from the system, which will be used to improve process control. The overall goal is to leverage all available data and knowledge to detect deviations at an early stage and to tighten process control.

[0132] The concept is illustrated in relation to Figures 9a and 9b.

[0133] Figure 9a illustrates a conventional system setup 90 for bioprocess control and data evaluation. This setup can be used for a cell-related parameter, such as pH, generated in process 91. A set point 92, e.g., pH = 7.0, and a parameter value to be measured 93 (after process 91), e.g., pH = 7.6, is compared to the set point, and the difference, in this example pH = 0.6, is used as input to the controller 94. The controller 94 will affect the parameter to be measured by adding 95 (in this example, CO2 or base addition) to cause a change in the parameter measured in the process, thereby achieving the desired difference between the set point and the measured value. The output 96 from the process is a sample for further processing in downstream processes, and its characteristics, e.g., biomass concentration, are measured relative to the sample.

[0134] Figure 9b illustrates an improved system setup 100 for model-based control and data evaluation. This setup can be used to further improve the prediction of outcomes from a process. In addition to the features described in relation to Figure 9a, the improved system setup 100 further comprises a unit 97 for generating a process model. Additional inputs from 95 and model parameter estimation 98 are used to determine at least one selected parameter 99 (viable cell density (VCD), viability, titer (product concentration), metabolite concentration, partial pressure of carbon dioxide (pCO2), buffering capacity, PID parameters, oxygen mass transfer coefficient (k L a) Used to estimate cell-specific rate, lactate concentration, etc.

[0135] The value of at least one selected parameter is fed back to the controller 94 and used to control the process 91. Each parameter points to the benefit of the system's user and provides soft sensing.

[0136] When VCD or viability is estimated by the model, harvest time can be estimated. When titer, [metabolites], or pCO2 are estimated by the model, sampling frequency can be reduced. pCO2, buffering capacity, PID parameters, or K L When a is estimated, process control can be improved, and when the cell-specific rate is estimated, online data evaluation can be performed.

[0137] The advantage of the system setup described in relation to Figures 9a and 9b is that more rigorous and robust results are achieved from the process. This leads to more predictable product quality, which in turn means a reduction in the number of tests required to establish a good quality product.

[0138] Typically, multiple bioreactors are used in a chromatography system to generate samples for downstream capture processes. The amount of sample generated from the bioreactors must be adapted to match the capacity of the downstream capture process. This means that the downstream process may affect the input to the controller 94 and / or when building the model in unit 97. [Explanation of Symbols]

[0139] 10 Control systems, control units 11 Cell Culture Process 12 Process Models, Process Model Databases 13 Historically important data 14 Databases 15. Dashed arrow 21 Process Strategy Data 22 Bioreactor Data 23 Data from offline measurements, offline measurement data 30 Default level 31 Level 1 32 Second Level 33 Third Level 40 Systems 41 Continuous Chromatography Systems 42 Target product to be delivered 43 Information 90 Previous System Setup 91 Processes 92 Set Points 93 Parameter values 94 Controllers 95 Added 96 Output 97 units 98 Model Parameter Estimation 99 Selected Parameters 100 Improved System Setup

Claims

1. A method for predicting the outcome of a process used to produce a sample in a bioreactor, wherein the process belongs to a category, and the method is The process model is selected based on the categories that categorize the types of processes used to produce the sample in the cell culture process (51), Historical data related to past process runs for manufacturing the aforementioned sample, Step (53) to access current data (54) obtained from the current process run of the process, wherein the current data obtained, based on the selected process model, includes process strategy data, bioreactor equipment data, data from online sensors, and / or data from offline sensors. Step (61) of adapting the process model in real time based on the historical data and at least one monitored parameter of the current process run, Step (62) predicts the outcome of at least one selected parameter of the current process run for producing the sample, based on the accessed historical data and current data, wherein the at least one selected parameter is selected from cell viability, cell count, product titer, and product quality. A step of determining an updated model for predicting outcomes in future process runs, wherein the updated model is The aforementioned process model, and A decision step, which is determined at least partially based on additional data from the current process, Includes, The aforementioned process strategy data includes strategic information relating to a process that includes at least one of the following: culture medium, feed type, feed regime, supplement, or supplement regime. method.

2. The method according to claim 1, further comprising the step (52) of organizing and integrating historical data, current data, and data related to the process model in a database.

3. The steps to retrieve the current data are: Step (54a) of selecting parameters based on the base model to be used, The steps of reading the data for the selected parameters (54b) and The method according to claim 1 or 2, further comprising:

4. The method according to any one of claims 1 to 3, further comprising the step (55) of dealing with missing data in the current data obtained from the current process run.

5. The steps to handle missing data are: A step (55a) in which missing data values ​​are replaced with values ​​to be imputed, or Step (55b) to remove data with missing data values. The method according to claim 4, further comprising:

6. The method according to claim 1, wherein the step of adapting the process model further includes the step of updating the process model based on data from the current process run.

7. The method according to claim 6, further comprising the step (61a) of applying the updated process model when predicting the outcome from the current process run.

8. The method according to claim 6 or 7, further comprising the step (61b) of applying the updated process model when predicting the outcome in a future process run.

9. The process used in the current process run is determined to belong to a new category, and the step of adapting the process run is: The steps of assigning the process to the new category, A step of storing the aforementioned process model as a new process model. The method according to claim 1, further comprising:

10. The step of predicting the outcome of at least one selected parameter is: The step of predicting at least one of the parameters (63), and / or, Step (64) to predict an anomaly, and / or, Steps to determine (65) and recommend actions to obtain improved conditions for the current process run The method according to any one of claims 1 to 9, further comprising:

11. A computer program for predicting the outcome of a current process run for producing a sample in a bioreactor, comprising instructions for causing the at least one processor to perform the method according to any one of claims 1 to 10 when run on the at least one processor.

Citation Information

Patent Citations

  • Method for controlling culture of biological cell, control device for controlling culture apparatus and culture apparatus

    JP2003235544A

  • Cell culture control system and cell culture control method

    JP2015216886A

  • Method for on-line prediction of future performance of a fermentation unit

    US20090048816A1

  • Bio-process model predictions from optical loss measurements

    US20090104653A1

  • Model based controls for use with bioreactors

    US20120107921A1