Data-Driven Process Development and Biopharmaceutical Manufacturing

By integrating machine learning models with mechanistic models, the complexity of biopharmaceutical manufacturing is addressed, providing accurate predictions and reducing uncertainty in biopharmaceutical production processes.

JP2025535071APending Publication Date: 2025-10-22BIOCURIE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025519889
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-09-12
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing biopharmaceutical manufacturing processes, particularly for cell and gene therapy, face high uncertainty due to the complexity of developing scalable and reproducible processes, with machine learning techniques lacking sufficient training data and failing to effectively predict outcomes.

Method used

Integration of machine learning models with mechanistic models that leverage physical, chemical, and biological properties to create predictive models, reducing uncertainty by combining data-driven approaches with process characteristics.

Benefits of technology

The integrated models provide accurate predictions across various scales and applications, from 1 mL to 25,000 L, enhancing process development and manufacturing efficiency and reducing uncertainty in biopharmaceutical production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535071000001_ABST
    Figure 2025535071000001_ABST
Patent Text Reader

Abstract

Disclosed is a method implemented to output a model for developing or operating a CGT process, the method including receiving, storing, and accessing a data item, determining attributes of the data item, selecting one or more machine learning models based on the attributes, accessing one or more mechanism models, integrating the one or more machine learning models with the one or more mechanism models to obtain one or more integrated models, selecting one or more predictive models from the one or more machine learning models, the one or more mechanism models, and the one or more integrated models, applying the one or more predictive models to the data item, adjusting one or more values ​​of one or more parameters of the one or more predictive models to reduce uncertainty in the model predictions, and outputting the one or more predictive models having the one or more adjusted values.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Non-Provisional Patent Application No. 17 / 959,537, filed October 4, 2022, the entire contents of which are incorporated herein by reference. [Background technology]

[0002] background

[0002] A biopharmaceutical—also known as a biological drug, biologic, or biologic—is any pharmaceutical product manufactured in, extracted from, or semi-synthesized from a biological source. The production of these biopharmaceuticals involves complex process development and manufacturing, primarily due to uncertainties at nearly every stage of development and manufacturing. Here, process development refers to the development of a robust, scalable, and reproducible process aimed at cost-effectively producing safe and effective biopharmaceuticals, and manufacturing refers to the production of biopharmaceuticals for clinical trials or commercial supply.

[0003]

[0003] One area of ​​focus in modern biopharmaceuticals is cell or gene therapy (CGT), which further includes subareas such as cell therapy (CT), gene therapy (GT), nucleic acid (NA) therapy and vaccines, and regenerative medicine. In particular, CT generally refers to the ex vivo production and delivery of cells to a human subject or the in vivo production of cells in a human subject to achieve a therapeutic or prophylactic effect. CT may or may not involve genetic modification. GT generally refers to the in vivo delivery of a gene or genetic element to a human subject to achieve a therapeutic or prophylactic effect. Summary of the Invention [Means for solving the problem]

[0004] overview According to one aspect of the present disclosure, a method is provided for outputting one or more models for developing or operating a CGT process, the method comprising: receiving a plurality of data items; storing the plurality of data items in a hardware storage device; and accessing the plurality of data items using the data processing system. The method comprises determining, by the data processing system, one or more attributes of the plurality of data items and selecting one or more machine learning models based on the one or more attributes. The method comprises accessing one or more mechanism models. The method comprises integrating, by the data processing system, the one or more machine learning models with the one or more mechanism models to obtain one or more integrated models; and selecting one or more predictive models from the one or more machine learning models, the one or more mechanism models, and the one or more integrated models. The method comprises applying the one or more predictive models to the plurality of data items. The method comprises adjusting one or more values ​​of one or more parameters of the one or more predictive models to reduce uncertainty in the model predictions. The method comprises outputting, by the data processing system, one or more predictive models having one or more adjusted values ​​of the one or more parameters.

[0005] According to one aspect of the present disclosure, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium includes program instructions that, when executed, cause a data processing system to perform operations for developing or operating a process for CGT. The operations include receiving a plurality of data items, storing the plurality of data items in a hardware storage device, and accessing the plurality of data items. The operations include determining one or more attributes of the plurality of data items and selecting one or more machine learning models based on the one or more attributes. The operations include accessing one or more mechanism models. The operations include integrating the one or more machine learning models with the one or more mechanism models to obtain one or more integrated models, and selecting one or more predictive models from the one or more machine learning models, the one or more mechanism models, and the one or more integrated models. The operations include applying the one or more predictive models to the plurality of data items. The operations include adjusting one or more values ​​of one or more parameters of the one or more predictive models to reduce uncertainty in the model predictions. The operations include outputting one or more predictive models having one or more adjusted values ​​of the one or more parameters.

[0006] In some implementations of the method or non-transitory computer-readable medium, the one or more attributes include at least one of nonlinearity, non-normality, collinearity, or dynamics.

[0007]

[0007] In some implementations of the method or non-transitory computer-readable medium, one or more mechanistic models are accessed based on at least one of physical, chemical, or biological properties of the process.

[0008]

[0008] In some embodiments of the method or the non-transitory computer-readable medium, integrating one or more machine learning models with one or more mechanistic models includes arranging the one or more machine learning models and the one or more mechanistic models in a sequence including one or more first models and one or more second models, sending output of the first one or more models to the second one or more models, sending data to the second one or more models, and obtaining output of the second one or more models.

[0009]

[0009] In some embodiments of the method or the non-transitory computer-readable medium, integrating one or more machine learning models with one or more mechanistic models includes determining first one or more models and second one or more models from the one or more machine learning models and the one or more mechanistic models, sending input data to the first one or more models, constraining predictions of the first one or more models using the second one or more models, and obtaining outputs of the first one or more models.

[0010]

[0010] The method or the non-transitory computer readable medium embodiment can be applied to various CGT techniques. [Brief explanation of the drawings]

[0011] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1]

[0011] An exemplary flow for reducing uncertainty in CGT process development and manufacturing is provided, according to some embodiments. [Figure 2]

[0012] 1 provides an exemplary block diagram illustrating model selection and integration, according to some embodiments. [Figure 3A]

[0013] 1 provides an exemplary chart illustrating the selection of a machine learning model, according to some embodiments. [Figure 3B]

[0014] 1 illustrates an exemplary biomanufacturing process to which modeling is applied, according to some embodiments. [Figure 4A]

[0015] 1 illustrates an exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models, according to some embodiments. [Figure 4B]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4C]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4D]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4E]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4F]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4G]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4H]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4I]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 4J]

[0015] An exemplary mechanism for integrating one or more machine learning models with one or more mechanistic models is shown, according to some embodiments. [Figure 5]

[0016] 1 illustrates an exemplary CGT modality according to some embodiments. [Figure 6]

[0017] 1 illustrates exemplary techniques involved in CGT, according to some embodiments. [Figure 7]

[0018] 1 illustrates a flowchart of an exemplary method according to some embodiments. [Figure 8]

[0019] 1 illustrates a block diagram of an exemplary computer system according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0012]

[0020] The figures are not drawn to scale and like reference numbers refer to like elements.

[0013] Detailed Description

[0021] Biopharmaceutical process development and manufacturing has a significant impact on drug safety, efficacy, scalability, and cost due to the high complexity of the drug, especially in the case of CGT. The nature of uncertainty in process development and manufacturing requires mechanisms that leverage data, analytics, and modeling for prediction and optimization.

[0014]

[0022] Machine learning is often considered a potential solution for making predictions after training on similar types of sample data. Commonly adopted machine learning techniques include supervised and unsupervised learning, neural networks, natural language processing, symbolic inference, algebraic learning, support vector machines, ensemble methods, kernel methods, k-nearest neighbor methods, automated learning, reinforcement learning, and Bayesian optimization. While machine learning techniques have been widely adopted to describe systems in other industries, these techniques have so far been largely unadopted in the biopharmaceutical industry. Among the very few applications of machine learning in the biopharmaceutical industry, most have been limited to describing small molecule and recombinant protein-based therapeutics, such as antibodies. For the process development and manufacturing of complex biopharmaceuticals, such as CGT, single machine learning techniques face the challenge of lacking training samples with an adequate quantity and quality to make predictions.

[0015]

[0023] In view of the above problems, the present disclosure provides one or more mechanisms for integrating machine learning models with mechanistic models. In this disclosure, a model generally refers to a description of a system using mathematical concepts and language. A model describes a system by a set of variables and a set of equations that establish relationships between the variables. A model can be used to describe a system, study the effects of components of the system, and make predictions. A model can be implemented as software code on a computer.

[0016]

[0024] Unlike machine learning models that rely on input sample data to make predictions, mechanistic models are created based on the application of physical, chemical, and / or biological properties that describe the behavior of components of the modeled system. For example, a mechanistic model may receive input parameters for raw materials to a bioreactor and apply mathematical equations that describe the underlying biological processes to predict the production rate and quality of the resulting biopharmaceutical. Mechanistic models can describe phenomena that are intracellular and / or extracellular and / or involve multiple cell populations, and may obtain information from process and / or omic data, such as genomes, transcriptomes, proteomes, epigenomes, metabolomes, fluxomes, and glycemia. Mechanistic models can describe phenomena that occur in liquid or solid solutions, in single or multiple phases, or on surfaces, such as for nucleic acids (e.g., oligonucleotides) produced by solid-phase synthesis. By integrating machine learning models with mechanistic models, embodiments of the present disclosure enable data-driven predictions to reduce uncertainty in complex process development and manufacturing of biopharmaceuticals.

[0017]

[0025] 1 illustrates a flow 100 for reducing uncertainty during process development and manufacturing of CGTs, according to some embodiments. As shown in FIG. 1, the flow 100 receives input data 101 that is used to form one or more predictive models 103. The predictive models 103 are then applied to a particular process development or manufacturing to reduce uncertainty in the prediction of CGTs 105.

[0018]

[0026] In FIG. 1 , data 101 can be collected from a variety of sources, such as physical sensors (pressure, temperature), chemical sensors (e.g., pH, pIons, metabolites), smart sensors, spectral sensors (e.g., Raman, fluorescence, near-infrared), imaging equipment (e.g., microscopy, holography, hyperspectral), assays, automation and robotics, digital twinning, systems biology, data lakes, and the Internet of Things. In some embodiments, data 101 includes sample data for training machine learning models. In some embodiments, data 101 is updated periodically or autonomously as flow 100 reduces uncertainty in predicting CGT.

[0019]

[0027] In operation 102, flow 100 includes selecting and / or integrating models based on data 101. In operation 102, two types of models are involved: one or more machine learning models and one or more mechanistic models. After the machine learning models and mechanistic models are selected and accessed, these two types of models are integrated into one or more predictive models 103.

[0020]

[0028] In some embodiments, the predictive model 103 makes predictions to reduce uncertainty in the operations 104. The uncertainty can be mathematically represented and adjusted in several ways, such as parameter-adaptive extended Kalman filtering, parameter-adaptive Runberger observer, and Bayesian adaptive ensemble Kalman filtering. The parameters involved in the operations 104 can be extracted from the input data 101. The predictive model 103 can be used in a wide range of applications, including control strategy design, production maximization, production scale-up, production technology transfer (e.g., to a different site), process monitoring, root cause analysis (e.g., troubleshooting), and economic modeling.

[0021]

[0029] The predictive model 103 can be used in a variety of applications of the CGT 105, ranging in scale from 1 mL per production run to 25,000 L per production run. In addition, production can be automated or semi-automated, and can be open, semi-closed, or closed. In addition, production can be automated for one or more unit operations or for an end-to-end production system or facility.

[0022]

[0030] 2 provides an exemplary block diagram illustrating a model selection and integration operation 200, according to some embodiments. Operation 200 may correspond in whole or in part to operation 102 of FIG. 1. In some embodiments, operation 200 is performed by computer hardware including, for example, a data processing system 203 and storage 202. In this example, storage 202 includes a non-transitory hardware storage device that, in combination with the memory of data processing system 203, causes data processing system 203 to perform the functions described herein.

[0023]

[0031] According to Figure 2, operation 200 includes receiving input data 201 and storing it in storage 202. Similar to data 101 in Figure 1, data 201 in Figure 2 may include sample data for training a machine learning model. By accessing data 201 from storage 202, data processing system 203 may select machine learning model 204.

[0024]

[0032] In some implementations, the selection of the machine learning model 204 is based on one or more attributes of the data 201. Specifically, after accessing the data 201 from storage 202, the data processing system 203 determines one or more attributes about the data 201 and makes the selection based on the one or more attributes. The selection is described in more detail below with reference to FIG. 3A.

[0025]

[0033] In addition to selecting the machine learning model 204, the data processing system 203 accesses a mechanistic model 206. In some embodiments, the access is based on the physical, chemical, or biological properties 205 of the process (or manufacturing) being developed. For example, if a mechanistic model 206 is used to develop a nucleic acid therapeutic, the data processing system 203 may access the mechanistic model 206 based on the chemical properties of the nucleic acid molecule.

[0026]

[0034] Once machine learning model 204 and mechanistic model 206 have been selected / accessed, data processing system 203 integrates the two in operation 207 to create an integrated model, also referred to as hybrid model 208. Integration operation 207 is described in more detail below with reference to FIGS. 4A-4J. As described herein, a data structure stores data representing a model (e.g., model 204 and / or model 206). This data structure is stored in memory, and data processing system 203 applies values ​​of fields in the data structure to input data to, for example, generate output. That is, the data structure includes fields that store values ​​or other data representing the model itself. In some examples, data processing system 203 includes a parser to parse the input data, e.g., to identify the structure of the data, to identify fields, etc. From the parsed data, data processing system 203 uses techniques described herein to identify fields and values ​​to be input to and / or applied to the model—from the structure of the data.

[0027]

[0035] The integrated model 208 can be applied to the input data 201 to make predictions in the process development or manufacturing of CGT. Because the integrated model 208 has both a machine learning component and a mechanism component, in some embodiments, the data processing system 203 can select whether to make the predictions using (i) only the machine learning component, (ii) only the mechanism component, or (ii) an integration of the machine learning and mechanism components. The selection can be based, for example, on the nature and degree of uncertainty of the data 201.

[0028]

[0036] In some embodiments, one or more values ​​of one or more parameters of integrated model 208 are adjusted to reduce prediction uncertainty. For example, after applying integrated model 208 to data 201 for CGT development, data processing system 203 evaluates the predicted results based on the actual results of CGT development. Thus, data processing system 203 determines to change the values ​​of one or more parameters of integrated model 208 according to the actual results, resulting in adjusted integrated model 208'. The adjustments can be performed iteratively for a limited number of times or automatically as process development or manufacturing progresses.

[0029]

[0037] Exemplary parameters of the integrated model 208 include prefactors in biochemical rate expressions, prefactors in cellular uptake rates, time scales for transport of molecules between the nucleus and organelles, and molecular diffusion rates. Exemplary tuning methods include Kalman filtering, parameter-adaptive Runberger observer, and Bayesian adaptive ensemble Kalman filtering. The integrated model 208 and the tuned integrated model 208' used in making predictions may be collectively referred to as a prediction model, as described in FIG. 1.

[0030]

[0038] As described above with reference to Figures 1 and 2, embodiments of the present disclosure combine machine learning features with mechanistic models to reduce uncertainty in CGT development. The embodiments do not require complex training datasets and take process / manufacturing characteristics into account. As such, the embodiments can be used in a wide range of CGT development applications, from scales as low as 1 mL per production run to 25,000 L per production run.

[0031]

[0039] FIG. 3A provides an exemplary chart illustrating the selection 300 of one or more machine learning models according to some embodiments. Consistent with the above description with reference to FIG. 2, the selection can be made by a data processing system based on one or more attributes of the input data. FIG. 3A illustrates three exemplary attributes: nonlinearity, collinearity, and dynamics. FIG. 3A also illustrates several candidate machine learning models available for selection. These machine learning models include: Algebraic Learning via Elastic Nets (ALVEN), Canonical Variable Analysis (CVA), Dynamic ALVEN (DALVEN), Elastic Nets, Multivariate Output Error State Space (MOESP), Partial Least Squares (PLS), Random Forest (RF), Recurrent Neural Networks (RNN), Ridge Regression (RR), Spare PLS, State Space Autoregressive Extraneous (SSARX), and Kernel Support Vector Regression (kSVR).

[0032]

[0040] In selection 300, the data is characterized based on three attributes: nonlinearity, collinearity, and dynamics. Specifically, by accessing each data item, the data processing system can determine whether the data item sufficiently exhibits the attributes of nonlinearity, collinearity, or dynamics. The data may demonstrate one, two, or all three attributes. In the triangle chart of FIG. 3A, each vertex corresponds to data exhibiting one attribute, each edge corresponds to data exhibiting two attributes simultaneously, and the center of the triangle corresponds to data exhibiting all three attributes simultaneously. Next to the vertices / edges / centers are examples of machine learning models that can be used for the corresponding data. As an example, data exhibiting only collinearity corresponds to the lower left vertex of the triangle. Selection 300 can select a machine learning model based on one or more of PLS, sparse PLS, RR, or elastic nets to process this type of data. As another example, data exhibiting both collinearity and dynamics corresponds to the lower edge of the triangle. Selection 300 allows for the selection of a machine learning model based on one or more of CVA, MOSEP, or SSARX to process this type of data. A description of the use of these machine learning models for general data analysis can be found in Sun and Braatz, "Smart process analytics for predictive modeling," Computers & Chemical Engineering, vol. 144, 107134, January 4, 2021, which is incorporated by reference herein. In some embodiments, in addition to the three exemplary attributes, other attributes, such as non-normality, can also be used in selecting a machine learning model.

[0033]

[0041] FIG. 3B illustrates an exemplary biomanufacturing process 350 to which modeling is applied, according to some embodiments. In FIG. 3B, input parameters define the operation of the biomanufacturing process 350, in which bioreactions related to the production of cells and / or their products occur. These input parameters, such as dissolved oxygen (DO), pH value, and impeller speed (measured in revolutions per minute [RPM]), affect physical, chemical, and / or biological phenomena in the biomanufacturing process 350. By applying mathematical equations that are functions of the input parameters, a model (e.g., a mechanistic model) can predict internal cellular states and / or process performance, such as the production rate and quality of cells and / or products produced by a bioreactor. Examples of internal cellular states include the concentration of species within a cell and / or its intracellular structure.

[0034]

[0042] Using one or more machine learning models selected based on data attributes and one or more mechanism models accessed based on process or manufacturing characteristics, the two types of models are integrated to create an integrated model. Many mechanisms are available for integrating the two types of models, and some examples are described below with reference to Figures 4A-4J. In the following description, the term "model" means "one or more models." Similarly, the terms "variable" and "parameter" mean "one or more variables" and "one or more parameters," respectively.

[0035]

[0043] FIG. 4A shows an exemplary mechanism 400A for integrating one or more machine learning models 404 with one or more mechanistic models 406, according to some embodiments. In 400A, the machine learning model 404 and the mechanistic model 406 are arranged in series, forming a two-stage data path. Input 401, which can be data 101 in FIG. 1 or data 201 in FIG. 2, is the initial input to the mechanistic model 406. Based on physical, chemical, or biological properties involved in process development or manufacturing, the mechanistic model 406 makes predictions from the input 401 and outputs one or more intermediate parameters or variables 411. The intermediate parameters or variables 411 are then fed to the machine learning model 404, which outputs a final integrated prediction 421. In this example, because the intermediate parameters or variables 411 are the result of the mechanistic prediction according to the physical / chemical / biological properties, subsequent machine learning predictions based on the intermediate parameters or variables 411 can be more relevant to actual development or manufacturing than machine learning predictions based purely on the raw input data 401.

[0036]

[0044] FIG. 4B illustrates an exemplary mechanism 400B for integrating one or more machine learning models with one or more mechanism models, according to some embodiments. Similar to 400A, in 400B, the machine learning model 404 and the mechanism model 406 are arranged in series. There are two possible data paths for the input 401: a two-step path with both the mechanism model 406 and the machine learning model 404, and a one-step path with only the machine learning model 404. Thus, in addition to receiving intermediate parameters or variables 411 for making machine learning predictions, the machine learning model 404 also receives the input 401 directly as another input source. The machine learning model 404 can potentially filter out outliers or errors in the intermediate parameters or variables 411, using the two input sources, for example, as references for each other. This mechanism can therefore improve prediction accuracy.

[0037]

[0045] 4C illustrates an exemplary mechanism 400C for integrating one or more machine learning models with one or more mechanistic models, according to some embodiments. Similar to 400A, in 400C, machine learning model 404 and mechanistic model 406 are arranged in series, forming a two-stage data path. Unlike the arrangement in 400A, machine learning model 404 is arranged first in 400C, receiving input 401 and outputting intermediate parameters or variables 411, and mechanistic model 406 outputs a final prediction 421 of the integration based on the intermediate parameters or variables 411. In this example, the prediction of mechanistic model 406 can be considered a refinement of the machine learning prediction based on physical / chemical / biological properties involved in process development or manufacturing.

[0038]

[0046] 4D illustrates an exemplary mechanism 400D for integrating one or more machine learning models with one or more mechanism models, according to some embodiments. In this example, a subset selector 402 is introduced to divide the data items of input 401 into two subsets 1 and 2, which may or may not overlap. Subset 1 is fed into a two-stage data path similar to mechanism 400C of FIG. 4C. Subset 2 is fed directly into a one-stage data path as an input to mechanism model 406.

[0039]

[0047] 4D , the selection of a data path for supplying a data item can be based on various factors. As an example, if the subset selector 402 determines that the machine learning model 404 is sufficiently trained to process a particular type of data item, the subset selector 402 can send the particular data item to subset 1 to be processed first by the machine learning model 404 and then by the mechanism model 406. Otherwise, the subset selector 402 can send the particular data item to subset 2 to be processed directly by the mechanism model 406. As another example, if the subset selector 402 determines that the machine learning model 404 has limited computational power and will become a bottleneck in the flow, the subset selector 402 can, at its own discretion or according to some predetermined rules, send some data items to subset 2 to bypass the machine learning model 404, thereby reducing computational latency.

[0040]

[0048] FIG. 4E illustrates an exemplary mechanism 400E for integrating one or more machine learning models with one or more mechanism models, according to some implementations. Similar to 400D in FIG. 4D, 400E uses a subset selector 402 to divide input 401 into subsets 1 and 2, with each subset passing through a two-stage data path in parallel. Subsets 1 and 2 may or may not overlap. Subset 1 first passes through mechanism model 406-1 and then through machine learning model 404-1, while subset 2 first passes through machine learning model 404-2 and then through mechanism model 406-2. Predictions from both subsets are then input to machine learning model 404-3 as a third stage for integrated predictions. In 400E, the two instances of the mechanism model 406-1 and 406-2 may or may not be the same, and the three instances of the machine learning model 404-1 through 404-3 may or may not be the same. By branching input 401 into multiple paths, mechanism 400E can potentially improve computation speed. Additionally, by having three stages of prediction, mechanism 400E can potentially improve prediction accuracy compared to two-stage and one-stage mechanisms.

[0041]

[0049] 4F shows an exemplary mechanism 400F for integrating one or more machine learning models with one or more mechanistic models, according to some embodiments. In 400F, subsets 1 and 2 each go through a one-stage path with machine learning model 404 and mechanistic model 406, respectively. The outputs of the two parallel paths are then combined by combiner 403 as the integration output 421.

[0042]

[0050] 4G illustrates an exemplary mechanism 400G for integrating one or more machine learning models with one or more mechanistic models, according to some embodiments. Unlike 400F, 400G replaces combiner 403 with machine learning model 404-2 as a two-stage prediction. The predictions made by machine learning model 404-2 form output 421 of the integration.

[0043]

[0051] 4H illustrates an exemplary mechanism 400H for integrating one or more machine learning models with one or more mechanism models, according to some embodiments. In 400H, a machine learning model 404 is incorporated within a mechanism model 406 to assist the mechanism model 406 in making predictions. For example, the built-in machine learning model 404 can assist the mechanism model 406 in data acquisition and / or classification to accelerate the mechanism prediction process.

[0044]

[0052] FIG. 4I shows an exemplary mechanism 400I for integrating one or more machine learning models with one or more mechanistic models, according to some embodiments. In 400I, the mechanistic model 406 is embedded within the machine learning model 404. Specifically, the machine learning model 404 has built-in constraints imposed according to the physical / chemical / biological properties of the mechanistic model 406. Through this mechanism, the output 421, which is the machine learning prediction, leverages information from the mechanistic model 406. Thus, the output 421 more closely describes the actual process than a pure machine learning model.

[0045]

[0053] 4J illustrates an exemplary mechanism 400J for integrating one or more machine learning models with one or more mechanism models, according to some embodiments. In addition to incorporating machine learning model 404 into mechanism model 406-1, 400J introduces mechanism model 406-2 for second-stage prediction. The addition of mechanism model 406 can potentially improve prediction accuracy.

[0046]

[0054] In the above-described integration mechanism, the input data can be structured in a variety of ways. As one example, each data item can be structured as a one-dimensional or multi-dimensional vector, with each element corresponding to an aspect or characteristic of process development or manufacturing. As another example, each data item can be structured as a tree, with the root corresponding to a major aspect or characteristic and each branch below it corresponding to a sub-aspect or sub-characteristic. Many other exemplary data structures are available. After reading this disclosure, one of ordinary skill in the art will be able to implement the above-described integration mechanism using one or more data structures suitable for process development or manufacturing.

[0047]

[0055] Once integrated, the integrated model can be used in a wide variety of applications to make predictions and reduce uncertainty in CGT process development and / or manufacturing. Exemplary modalities and techniques for CGT are described below with reference to Figures 5 and 6.

[0048]

[0056] Figure 5 shows exemplary CGT modalities according to some embodiments. Generally, CGT includes CT and GT. CT includes ex vivo or in vivo production and may or may not include genetic modification. GT involves direct delivery of therapeutic genetic elements to a human subject. GT may be achieved using a delivery vehicle, such as a viral vector, a bacterial vector, a non-viral vector, and / or a physical method, or may be delivered without a delivery vehicle. The genetic elements can encode, for example, full-length proteins, protein fragments, polypeptides, peptides, regulatory elements, or transposition elements. Examples of encoded proteins and fragments include enzymes, structural proteins, regulatory proteins, antibodies, cytokines, antigens, and transcription factors. Examples of regulatory elements include promoters, enhancers, gene switches, and logic gates. Examples of transposition elements include transposons such as Sleeping Beauty, piggyBac, and Tol2. Examples of gene switches include on switches, off switches, on-off switches, and dimmer switches.

[0049]

[0057] Ex vivo implementation of CT includes steps 501-503. In 501, cells, such as stem cells or non-stem cells, are extracted from a human body or a non-human animal, which may or may not be a human recipient. In 502, the extracted cells may be manipulated, modified, and / or amplified. Additionally, the cells may be genetically modified with a payload delivered via a viral vector, a bacterial vector, a non-viral carrier, and / or by physical means. Examples of payloads include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), non-naturally occurring nucleic acids, peptides, and / or proteins. In step 503, the resulting cell therapy product, which may be single-cell or multicellular, is transferred into the body of a human subject for therapeutic or prophylactic purposes.

[0050]

[0058] In vivo implementations of CT include step 511, in which a payload is delivered directly to the body of a human subject. The payload may be delivered via a viral vector, a bacterial vector, a non-viral carrier, and / or by physical means. Examples of payloads include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), non-naturally occurring nucleic acids, peptides, and / or proteins. Step 511 may similarly be used in GT implementations, which are also in vivo, to deliver gene sequences to the body of a human subject.

[0051]

[0059] The integrated model, as well as many other techniques involved in CGT, can be applied to any of steps 501-503 and 511. Examples of these techniques are described below with reference to FIG.

[0052]

[0060] 6 shows exemplary techniques involved in CGT, according to some embodiments. These techniques are broadly divided into four categories: GT, NA, genetically modified CT, and non-genetically modified CT. Production scales for these techniques can range from 1 mL per production run to 25,000 L per production run, and production modes can be, for example, batch, fed-batch, perfusion, continuous, semi-continuous, or a hybrid of fed-batch and perfusion.

[0053]

[0061] In GTs, the payload can include one or more genes and / or regulatory sequences and can include one or more non-coding sequences. Payloads can be used, for example, for gene replacement, gene activation, gene inactivation, introduction of new or modified genes, cell reprogramming, transdifferentiation, and / or gene editing. The payload can be delivered via a viral vector, a non-viral vector such as a bacterium, a non-viral carrier, and / or a physical delivery means. The GT can include at least one targeting moiety, which can be combined with the payload or can be inherent to the payload. The targeting moiety can include an NA sequence, a protein (e.g., an antibody), a protein fragment (e.g., an antibody fragment), a peptide, a monosaccharide, a polysaccharide, an aptamer, a dendrimer, a small molecule, or a centrin. The moiety can be combined, fused, conjugated, or attached to the vector or payload.

[0054]

[0062] GT can involve performing transient transfection, stable transfection, or transduction of suspension or adherent cells to produce a therapeutic agent, such as a viral vector carrying a gene of interest. GT can involve performing stable producer transfection or packaging of therapeutic-producing cell lines grown in suspension or as adherent cells. Exemplary cell lines include HEK293 and its variants (e.g., HEK293T), Sf9, HeLa, A469, CAP, AGELHN, PER.C6, NS01, COS-7, BHK, CHO, VERO, MDCK, BRL3A, HepG2, primary human cells, peripheral blood mononuclear cells (PBMC), immune cells, T cells, human stem cells, induced pluripotent stem cells, or somatic cells. Additionally, GT can involve producing therapeutic agents in transfection-free systems, such as self-attenuated adenovirus-based systems for viral vector production (e.g., systems based on tetracycline-responsive self-silencing adenovirus [TESSA]) or oncolytic viruses that selectively replicate in and kill target cells. Examples of viral vectors include adeno-associated viruses (AAV), lentiviruses (LV), adenoviruses (Ad), baculoviruses, herpes simplex viruses (HSV), retroviruses, oncolytic viruses, parvoviruses, anelloviruses, and bacteriophages.

[0055]

[0063] All of the above GT techniques can use integrated models to reduce uncertainty. For example, integrated models can be used in generating viral vectors or non-viral carriers, generating and delivering payloads, performing transient transfections, etc.

[0056]

[0064] NA therapy and vaccines can involve producing and delivering an NA-based therapy or vaccine that encodes a therapeutic and / or protective moiety. Delivery can be in vivo or ex vivo. NA therapy and vaccines can be applied to a variety of cell types, with or without a specific target. Such cell types include immune cells, tumor cells, cardiac cells, ocular cells, retinal cells, lung cells, skin cells, muscle cells, liver cells, pancreatic cells, intestinal cells, brain cells, and neural cells, to name a few.

[0057]

[0065] NA therapy or vaccines can include DNA, plasmid DNA (pDNA) (including bacmids, nanoplasmids, linearized pDNA, etc.), RNA, messenger RNA (mRNA), small activating RNA (saRNA), small interfering RNA (also known as short interfering RNA, silencing RNA, or siRNA), microRNA (miRNA), circular RNA, antisense oligonucleotides (ASO), doggybone DNA (dbDNA), minicircle DNA (mcDNA), minimal immunologically defined gene expression (MIDGE), closed DNA (ceDNA), synthetic DNA, or non-natural NAs containing non-natural or potentially modified nucleotides or nucleosides, as well as peptides containing non-natural chemical and multidimensional structures.

[0058]

[0066] An NA therapy or vaccine can include, for example, one or more non-identical NA molecules, each encoding a different sequence. It can also include one or more non-NA elements. For example, an NA therapy can include a protein or protein fragment, such as a ribonucleoprotein (RNP) for gene editing. It can also include one or more targeting moieties, such as for enhancing delivery to a specific organ, tissue, cell type, or intracellular compartment. The one or more targeting moieties can include an NA sequence, a protein (e.g., an antibody), a protein fragment (e.g., an antibody fragment), a peptide, a monosaccharide, a polysaccharide, an aptamer, a dendrimer, a small molecule, or a centrin. The moiety can be a ligand that is combined, fused, conjugated, or bound to the NA payload. The moiety can also be encoded on or inherent in the payload itself.

[0059]

[0067] NA therapy and vaccines can further include producing NA, chemically or enzymatically modifying NA, and delivering NA either by combining NA with a non-viral carrier and / or via a physical delivery method. Examples of non-viral carriers include lipid nanoparticles (LNPs), solid lipid nanoparticles (SLNs), nanostructured lipid carriers (NLCs), liposomes, lipoplexes, polymeric nanoparticles, lipid-polymer hybrid nanoparticles, inorganic nanoparticles, exosomes, virus-like particles, extracellular vesicles, cell-penetrating peptides, cationic polymers (e.g., PEI, PLA, PLGA, chitosan), dendrimers, aptamers, and centrins. Examples of physical delivery methods include electroporation, cell squeezing, needles (including microneedles and nanoneedles), patches, iontophoresis, biolistic delivery (including gene guns and particle bombardment), sonoporation, ultrasound-mediated microbubbles, hydroporation, photoporation, and magnetofection.

[0060]

[0068] All of the above NA-based therapy and vaccine techniques can use integrated models to reduce uncertainty, for example, in the generation of NA production, modification, and delivery, the production of NA sequences (e.g., either produced together in the same reaction or produced in separate reactions and then mixed into a single product), the application of therapies and vaccines to various cell types, the generation of non-viral carriers, and the implementation of physical delivery methods.

[0061]

[0069] CT can be produced, for example, by transduction with a viral vector or transfection with NA. In CT, the resulting cells can be genetically modified or non-genetically modified, stem cell-based or non-stem cell-based, and unicellular or multicellular. In addition, the resulting cells can be autologous (patient-specific). Autologous CT involves obtaining cells from a source from a human subject (e.g., stem cells, human pluripotent stem cells including induced pluripotent stem cells and embryonic stem cells, non-stem cells, or cell lines derived from various sources such as peripheral blood, bone marrow, umbilical cord blood, placenta, skin, eye, and muscle), culturing and expanding the cells outside the body (ex vivo), and reintroducing the resulting CT product into the same subject. This process can include enrichment of one or more specific cell types or phenotypes. This process can include genetic modification. In addition, this process can include gene editing to produce one or more gene edits.

[0062]

[0070] The cells produced can be allogeneic (used to treat multiple patients). Allogeneic CT involves obtaining cells (e.g., stem cells, human pluripotent stem cells, including induced pluripotent stem cells and embryonic stem cells, non-stem cells, or cell lines) from various sources, such as human peripheral blood from healthy donors, umbilical cord blood, placenta, and skin, to create a master cell bank (MCB), which is used as a source to create cell populations that are processed according to the requirements of a specific therapy. The final cell population is then used to treat one or more patients. This process can include enrichment for one or more specific cell types or phenotypes. This process can include genetic modification. In addition, this process can include gene editing to produce one or more gene edits.

[0063]

[0071] Genetically modified CT can involve the modification of specific genes and / or regulatory sequences within cells. Genetically modified CT can be applied to a variety of cell types, including immune cells, tumor cells, cardiac cells, ocular cells, retinal cells, lung cells, skin cells, pancreatic cells, intestinal cells, muscle cells, liver cells, brain cells, and neural cells, to name a few.

[0064]

[0072] Genetically modified CT can be applied to tumor cells associated with hematological malignancies and solid tumors. For example, CT can include the production of genetically modified chimeric antigen receptor T cells (CAR T cells), gamma delta T cells, natural killer (NK) cells, engineered T cell receptors (TCRs), tumor-infiltrating lymphocytes (TILs), macrophages, dendritic cells, hematopoietic stem cells (HSCs), or mesenchymal stem / stromal cells (MSCs).

[0065]

[0073] Genetically modified CT can include one or more targeting moieties. One or more targeting moieties can include NA sequences, proteins (e.g., antibodies), protein fragments (e.g., antibody fragments), peptides, monosaccharides, polysaccharides, aptamers, dendrimers, small molecules, or centrins. The moieties can be combined with cells, fused to cells, conjugated to cells, or attached to cells. The moieties can also be encoded by the cells themselves, or can be native to the cells themselves, or expressed on the cell surface. In genetically modified CT, cells can be generated and modified ex vivo or in vivo.

[0066]

[0074] Non-genetically modified CT can include regenerative medicine, stem cell therapy, or tissue engineering. Non-genetically modified CT can be applied to a variety of cells, including immune cells, tumor cells, cardiac cells, ocular cells, retinal cells, lung cells, pancreatic cells, intestinal cells, muscle cells, skin cells, bone cells, liver cells, brain cells, and nerve cells, to name a few.

[0067]

[0075] All of the above genetically modified and non-genetically modified CT techniques can use integrated models to reduce uncertainty. For example, integrated models can be used in the generation of autologous or allogeneic cells, cell editing, and the production of viral or non-viral delivery vehicles.

[0068]

[0076] In addition to the techniques described in each category, the integrated model can be applied to CGT techniques for producing cell lines. For example, using input data obtained from a first type of cell line or cell population, the integrated model can be applied to produce a second type of cell line or cell population that is different from the first type. The two cell lines / populations can have heterogeneous or clonal cell populations, and the heterogeneous cell population can have intracellular or cell surface heterogeneity. Production can be automated or semi-automated, and can be a semi-closed or closed system.

[0069]

[0077] In some embodiments, the integrated model can be applied to stable cell line production or cell line packaging. In some embodiments, the integrated model can be applied to performing transfection (e.g., transient transfection) or transduction of one or more stable producer host cell lines or one or more packaging host cell lines that can be used, for example, in GT.

[0070]

[0078] The techniques described above are just a few of the many examples where integrated models can be applied to reduce uncertainty. Due to their high scalability, customizability, and predictive accuracy, embodiments of the present disclosure can be used in many applications in biopharmaceutical process development and manufacturing.

[0071]

[0079] 7 shows a flowchart of an exemplary method 700 according to some embodiments. Method 700 may be implemented as software code on a computer. One or more steps of method 700 may correspond to steps or acts described with reference to FIGS. 1 and 2.

[0072]

[0080] At 702, the method 700 includes receiving a plurality of data items, such as data 101 of FIG. 1 or data 201 of FIG.

[0073]

[0081] At 704, the method 700 includes storing the plurality of data items in a hardware storage device, such as storage 202 of FIG.

[0074]

[0082] At 706, the method 700 includes accessing a plurality of data items using a data processing system, such as data processing system 203 of FIG.

[0075]

[0083] At 708, the method 700 includes determining, by the data processing system, one or more attributes of the plurality of data items. These attributes may include those described with reference to FIG. 3A.

[0076]

[0084] At 710, the method 700 includes selecting one or more machine learning models based on the one or more attributes. The one or more machine learning models may include those described with reference to FIG. 3A.

[0077]

[0085] At 712, method 700 includes accessing one or more mechanistic models. Consistent with the above description, accessing the one or more mechanistic models can be based on one or more physical / chemical / biological properties involved in process development or manufacturing.

[0078]

[0086] At 714, the method 700 includes integrating, by the data processing system, the one or more machine learning models with the one or more mechanistic models to obtain one or more integrated models. The integration may correspond to step 207 of Figure 2 and may use one or more mechanisms described with reference to Figures 4A-4J.

[0079]

[0087] At 716, method 700 includes selecting one or more predictive models from the one or more machine learning models, the one or more mechanistic models, and the one or more integrated models. The selection can be based on the input data items, the nature of the uncertainty, and available computing resources. Depending on the selection, either one or more machine learning models alone, one or more mechanistic models alone, or an integration of the two can be selected as the one or more predictive models for reducing uncertainty.

[0080]

[0088] At 718, the method 700 includes applying one or more predictive models to the plurality of data items. This application can be used in the CGT techniques described with reference to FIGS.

[0081]

[0089] At 720, the method 700 includes adjusting one or more values ​​of one or more parameters of the one or more predictive models to reduce uncertainty in the model predictions. The one or more parameters and their adjustments can be similar to those described with reference to FIG. 2.

[0082]

[0090] At 722, method 700 includes outputting, by the data processing system, one or more predictive models having the one or more adjusted values ​​of the one or more parameters. The output one or more predictive models can be similar to adjusted integrated model 208' with the parameters adjusted.

[0083]

[0091] As discussed above, the integration of mechanistic and machine learning models using the features described herein advantageously improves predictive power and efficiency in biopharmaceutical process development and manufacturing, resulting in significantly increased scalability and reduced costs.

[0084]

[0092] 8 is a block diagram of an exemplary computer system 800 according to an embodiment of the present disclosure. Storage 202 and data processing system 203 may be implemented, for example, as components of computer system 800. System 800 includes a processor 810, memory 820, a storage device 830, and one or more input / output interface devices 840. Each of components 810, 820, 830, and 840 may be interconnected using, for example, a system bus 850.

[0085]

[0093] Processor 810 is capable of processing instructions for execution within system 800. As used herein, the term "execution" refers to the technique of program code causing a processor to execute one or more processor instructions. In some implementations, processor 810 is a single-threaded processor. In some implementations, processor 810 is a multi-threaded processor. Processor 810 is capable of processing instructions stored in memory 820 or storage device 830. Processor 810 may perform operations as described with reference to other figures described herein.

[0086]

[0094] Memory 820 stores information within system 800. In some implementations, memory 820 is a computer-readable medium. In some implementations, memory 820 is a volatile memory unit. In some implementations, memory 820 is a non-volatile memory unit.

[0087]

[0095] Storage device 830 can provide mass storage for system 800. In some implementations, storage device 830 is a non-transitory computer-readable medium. In various different implementations, storage device 830 can include, for example, a hard disk device, an optical disk device, a solid-state drive, a flash drive, a magnetic tape, or some other mass storage device. In some implementations, storage device 830 can be a cloud storage device, e.g., a logical storage device including one or more physical storage devices distributed over and accessed using a network. In some examples, the storage device can store long-term data. Input / output interface device 840 provides input / output operations to system 800. In some implementations, input / output interface device 840 can include one or more network interface devices, e.g., an Ethernet interface, a serial communication device, e.g., an RS-232 interface, and / or a wireless interface device, e.g., an 802.11 interface, a 3G wireless modem, a 4G wireless modem, a 5G wireless modem, etc. The network interface devices enable system 800 to communicate, e.g., to send and receive data. In some implementations, the input / output devices may include driver devices configured to receive input data and send output data to other input / output devices, such as keyboards, printers, and display devices 860. In some implementations, mobile computing devices, mobile communication devices, and other devices may be used.

[0088]

[0096] A server may be implemented in a distributed manner over a network, such as a server farm or a widely distributed set of servers, or in a single virtual device that includes multiple distributed devices operating in conjunction with each other. For example, one of the devices may control the other devices, or the devices may operate under a coordinated set of rules or protocols, or the devices may cooperate in another way. The cooperation of multiple distributed devices makes them appear to operate as a single device.

[0089]

[0097] In some examples, system 800 is contained within a single integrated circuit package. This type of system 800, in which both processor 810 and one or more other components are contained within a single integrated circuit package and / or fabricated as a single integrated circuit, may be referred to as a microcontroller. In some implementations, the integrated circuit package includes pins corresponding to input / output ports that can be used, for example, to communicate signals to or from one or more of input / output interface devices 840.

[0090]

[0098] While an exemplary processing system is depicted in Figure 8, embodiments of the subject matter and functional operations described herein can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. In some embodiments, computation (e.g., data processing) may occur at a central location and / or at distributed locations, e.g., distributed locations involving edge computing. Computing may also include quantum computing in some embodiments.

[0091]

[0099] Software implementations of the described subject matter can be implemented as one or more computer programs. Each computer program can include one or more modules of computer program instructions encoded on a tangible, non-transitory, computer-readable computer storage medium for execution by or to control the operation of a data processing apparatus. Alternatively or additionally, the program instructions can be encoded in / on an artificially generated propagated signal. In one example, the signal can be a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a receiver apparatus suitable for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of computer storage media.

[0092]

[0100] The terms “data processing device,” “computer,” and “computing device” (or equivalents as understood by those skilled in the art) refer to data processing hardware. For example, a data processing device can encompass any type of apparatus, device, and machine for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. An apparatus can also include special-purpose logic circuitry, including, for example, a central processing unit (CPU), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some implementations, a data processing device or special-purpose logic circuitry (or a combination of data processing devices or special-purpose logic circuitry) can be hardware-based or software-based (or a combination of both). An apparatus can optionally include code that creates an execution environment for a computer program, such as processor firmware, a protocol stack, a database management system, an operating system, or code that constitutes a combination of the execution environment. The present disclosure contemplates the use of a data processing device with or without a conventional operating system, such as LINUX, UNIX, Windows, MAC OS, ANDROID, or iOS.

[0093]

[0101] A computer program may also be referred to as or described as a program, software, software application, module, software module, script, or code, and may be written in any form of programming language. Programming languages ​​may include, for example, compiled, interpreted, declarative, or procedural languages. A program may be deployed in any form, including as a stand-alone program, module, component, subroutine, or unit for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data, such as one or more scripts stored in a markup language document, a single file dedicated to the program in question, or multiple associated files that store one or more modules, subprograms, or code portions. A computer program may be deployed to run on one computer or on multiple computers, for example, located at a single site or distributed across multiple sites interconnected by a communications network. While portions of the programs depicted in various figures may be depicted as individual modules that implement various features and functionality through various objects, methods, or processes, a program may instead include multiple sub-modules, third-party services, components, and libraries. Conversely, features and functionality of various components may be combined into a single component as desired. The thresholds used to make the computational determination may be determined statically, dynamically, or both statically and dynamically.

[0094]

[0102] The methods, processes, or logic flows described herein may be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The methods, processes, or logic flows may also be performed by, and an apparatus may be implemented as, special purpose logic circuitry, such as a CPU, FPGA, Arduino, or ASIC.

[0095]

[0103] A computer suitable for executing a computer program can be based on one or more of general-purpose and special-purpose microprocessors and other types of CPUs. Elements of a computer are a CPU for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a CPU can receive instructions and data from (and write data to) memory. A computer may also include, or be operatively coupled to, one or more mass storage devices for storing data. In some implementations, a computer can receive data from and transfer data to a mass storage device, including, for example, a magnetic disk, a magneto-optical disk, or an optical disk. Furthermore, a computer can be incorporated into another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GNSS sensor or receiver, or a portable storage device such as a universal serial bus (USB) flash drive.

[0096]

[0104] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media that can store, contain, or transport instructions and / or data. Computer-readable media may also include non-transitory media that can store data and do not involve carrier waves and / or transitory electronic signals propagated wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may store code and / or machine-executable instructions, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0097]

[0105] Computer-readable media suitable for storing computer program instructions and data (either transiently or non-transiently, as appropriate) can include all forms of permanent / non-permanent and volatile / non-volatile memory, media, and memory devices. Computer-readable media can include, for example, semiconductor memory devices such as random access memory (RAM), read-only memory (ROM), phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices. Computer-readable media can also include, for example, magnetic devices such as tapes, cartridges, cassettes, and internal / removable disks. Computer-readable media can also include magneto-optical disks and optical memory devices, as well as technologies including, for example, digital video disks (DVDs), CD-ROMs, DVD+ / -Rs, DVD-RAMs, DVD-ROMs, HD-DVDs, and Blu-ray discs. The memory may store a variety of objects or data, including caches, classes, frameworks, applications, modules, backup data, jobs, web pages, web page templates, data structures, database tables, repositories, and dynamic information. The types of objects and data stored in the memory may include parameters, variables, algorithms, instructions, rules, constraints, and references. Additionally, the memory may include logs, policies, security or access data, and report files. The processor and memory may be supplemented by or incorporated in special purpose logic circuitry.

[0098]

[0106] While this specification contains details of many specific embodiments, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be inherent in particular embodiments. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, separately, or in any suitable subcombination. Furthermore, while the features described above may be described as working in a particular combination, and may even be initially claimed as such, in some cases, one or more features from a claimed combination may be deleted from that combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0099]

[0107] Specific embodiments of the present subject matter have been described. Other embodiments, modifications, and permutations of the described embodiments, as will be apparent to those skilled in the art, are within the scope of the following claims. Although operations are shown in the figures or in the claims in a particular order, this should not be understood as requiring such operations to be performed in the particular order or sequential order shown, or that all illustrated operations be performed, to achieve desirable results (some operations may be considered optional). In particular situations, multitasking or parallel processing (or a combination of multitasking and parallel processing) may be advantageous and performed as deemed appropriate.

[0100]

[0108] Furthermore, the separation or integration of various system modules and components in the foregoing embodiments should not be understood to require such separation or integration in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.

[0101]

[0109] Accordingly, the foregoing exemplary embodiments do not define or limit this disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of this disclosure.

Claims

1. 1. A method implemented by a data processing system for outputting one or more models for developing or operating a process for cell or gene therapy (CGT), comprising: receiving a plurality of data items; storing the plurality of data items in a hardware storage device; accessing the plurality of data items using the data processing system; determining, by the data processing system, one or more attributes of the plurality of data items; selecting one or more machine learning models based on the one or more attributes; accessing one or more mechanical models; integrating, by the data processing system, the one or more machine learning models with the one or more mechanistic models to obtain one or more integrated models; selecting one or more predictive models from the one or more machine learning models, the one or more mechanistic models, and the one or more integrated models; applying the one or more predictive models to the plurality of data items; adjusting one or more values ​​of one or more parameters of the one or more predictive models to reduce uncertainty in the model predictions; outputting, by the data processing system, the one or more predictive models having the one or more adjusted values ​​of the one or more parameters; A method comprising:

2. The method of claim 1 , wherein the one or more attributes include at least one of nonlinearity, collinearity, non-normality, or dynamics.

3. The method of claim 1 , wherein the one or more mechanistic models are accessed based on at least one of physical, chemical, or biological properties of the process.

4. Integrating the one or more machine learning models with the one or more mechanistic models includes: arranging the one or more machine learning models and the one or more mechanistic models into a sequence including first one or more models and second one or more models; sending outputs of the first one or more models to the second one or more models; transmitting data to the second one or more models; obtaining outputs of the second one or more models; The method of claim 1 , comprising:

5. Integrating the one or more machine learning models with the one or more mechanistic models includes: determining first one or more models and second one or more models from the one or more machine learning models and the one or more mechanistic models; sending input data to the first one or more models; using the second one or more models to constrain predictions of the first one or more models; obtaining an output of the first one or more models; The method of claim 1 , comprising:

6. the plurality of data items are obtained from a first type of cell population; The method comprises: producing a second type of cell population using the one or more output models; The method of claim 1 , wherein the second type is different from the first type.

7. 7. The method of claim 6, wherein the cell population of the first type and the cell population of the second type each comprise at least one of a heterogeneous cell population or a clonal cell population.

8. The method of claim 7 , wherein the heterogeneous cell population has at least one of intracellular heterogeneity or cell surface heterogeneity.

9. 10. The method of claim 1, further comprising producing a stable cell line using one or more output models.

10. The stable cell line comprises:

10. The method of claim 9, comprising at least one of HEK293 cells, HEK293T cells, Sf9 cells, HeLa cells, A469 cells, CAP cells, AGELHN cells, Per. C6 cells, NS01 cells, COS-7 cells, BHK cells, CHO cells, VERO cells, MDCK cells, BRL3A cells, HepG2 cells, primary human cells, peripheral blood mononuclear cells (PBMCs), immune cells, T cells, human stem cells, induced pluripotent stem cells, or somatic cells.

11. 10. The method of claim 1, wherein the scale of the CGT ranges from 1 mL per production run to 25,000 L per production run.

12. The CGT is 10. The method of claim 1, wherein the one or more output models are used on cells grown in at least one of batch, fed-batch, perfusion, continuous, semi-continuous, or a hybrid of fed-batch and perfusion modes.

13. The method of claim 1 , wherein the CGT uses one or more output models to generate automated or semi-automated production.

14. 14. The method of claim 13, wherein the production is in a closed or semi-closed system.

15. The method of claim 1 , wherein the CGT comprises gene therapy.

16. The gene therapy comprises:

16. The method of claim 15, comprising using one or more payloads for at least one of gene replacement, gene activation, gene inactivation, introduction of new or modified genes, or gene editing.

17. 16. The method of claim 15, further comprising using the one or more output models to generate one or more viral vectors for said gene therapy.

18. The viral vector is 18. The method of claim 17, comprising at least one of an adeno-associated virus, a lentivirus, an adenovirus, a baculovirus, a herpes simplex virus, a retrovirus, an oncolytic virus, a parvovirus, anellovirus, or a bacteriophage.

19. 16. The method of claim 15, further comprising using one or more output models to perform transient transfection, stable transfection, or transduction for said gene therapy.

20. 20. The method of claim 18, further comprising performing transient transfection, stable transfection, or transduction of suspension or adherent cells.

21. The floating cells or adherent cells are 21. The method of claim 20, comprising at least one of HEK293 cells, HEK293T cells, Sf9 cells, HeLa cells, A469 cells, CAP cells, AGELHN cells, Per. C6 cells, NS01 cells, COS-7 cells, BHK cells, CHO cells, VERO cells, MDCK cells, BRL3A cells, HepG2 cells, primary human cells, peripheral blood mononuclear cells (PBMCs), immune cells, T cells, human stem cells, induced pluripotent stem cells, or somatic cells.

22. 16. The method of claim 15, further comprising using one or more output models to effect transfection or transduction of one or more stable producer host cell lines or one or more packaging host cell lines for said gene therapy.

23. 16. The method of claim 15, further comprising producing the viral vector in the system without transfection.

24. the plurality of data items are obtained from transient transfections; The method comprises:

16. The method of claim 15, further comprising using one or more output models to grow and / or produce stable producer or packaging cell lines for said gene therapy.

25. 16. The method of claim 15, wherein the gene therapy comprises one or more targeting moieties.

26. the one or more targeting moieties 26. The method of claim 25, comprising at least one of a nucleic acid sequence, a protein, a protein fragment, a peptide, a monosaccharide, a polysaccharide, a small molecule, an aptamer, a dendrimer, or a centilin.

27. 10. The method of claim 1, further comprising using one or more output models to produce a nucleic acid-based therapy or vaccine for said GCT.

28. 28. The method of claim 27, further comprising using the one or more output models to produce nucleic acids for the nucleic acid-based therapy or vaccine.

29. The nucleic acid therapy or vaccine comprises:

28. The method of claim 27, comprising at least one of DNA, plasmid DNA (pDNA), RNA, messenger RNA (mRNA), small activating RNA (saRNA), small interfering RNA (also known as short interfering RNA, silencing RNA, or siRNA), microRNA (miRNA), circular RNA, antisense oligonucleotide (ASO), doggybone DNA (dbDNA), closed-loop DNA (ceDNA), synthetic DNA, or non-naturally occurring nucleic acid.

30. 28. The method of claim 27, further comprising effecting chemical or enzymatic modification of the nucleic acid.

31. 28. The method of claim 27, wherein the nucleic acid is combined with a non-viral carrier.

32. The production 28. The method of claim 27, comprising at least one of a non-viral carrier or a physical delivery method.

33. 28. The method of claim 27, further comprising using a plurality of nucleic acid molecules to produce one or more sequences for the nucleic acid-based therapy or vaccine.

34. 28. The method of claim 27, wherein the nucleic acid-based therapy or vaccine comprises one or more nucleic acid molecules and one or more targeting moieties.

35. the one or more targeting moieties 35. The method of claim 34, comprising at least one of a nucleic acid sequence, a protein, a protein fragment, a peptide, a monosaccharide, a polysaccharide, a small molecule, an aptamer, a dendrimer, or a centilin.

36. 28. The method of claim 27, wherein the nucleic acid-based therapy or vaccine comprises one or more nucleic acid molecules and one or more non-nucleic acid molecules.

37. 37. The method of claim 36, wherein the one or more non-nucleic acid molecules comprise a protein, a protein fragment, or a peptide.

38. 28. The method of claim 27, wherein the nucleic acid-based therapy or vaccine is applied to at least one of immune cells, tumor cells, cardiac cells, eye cells, retinal cells, lung cells, muscle cells, skin cells, liver cells, pancreatic cells, intestinal cells, brain cells, or neural cells.

39. 33. The method of claim 32, wherein the non-viral carrier comprises at least one of a lipid nanoparticle, a solid lipid nanoparticle, a nanostructured lipid carrier, a liposome, a lipoplex, a polymeric nanoparticle, a lipid-polymer hybrid nanoparticle, an inorganic nanoparticle, an exosome, a virus-like particle, an extracellular vesicle, a cell-penetrating peptide, a cationic polymer, an aptamer, a dendrimer, or a centilin.

40. 33. The method of claim 32, wherein the physical delivery method comprises at least one of electroporation, cell squeezing, needles, patches, iontophoresis, biolistic delivery, sonoporation, ultrasound-mediated microbubbles, hydroporation, photoporation, and magnetofection.

41. The method of claim 1 , wherein the CGT comprises cell therapy.

42. generating one or more cells for the cell therapy based on the one or more output models; 42. The method of claim 41, wherein the one or more cells comprise at least one of autologous or allogeneic cells.

43. 42. The method of claim 41, further comprising a cell therapy produced by transduction with a viral vector or transfection with a nucleic acid.

44. 42. The method of claim 41, wherein the cell therapy is applied to at least one of immune cells, tumor cells, cardiac cells, eye cells, retinal cells, lung cells, pancreatic cells, intestinal cells, kidney cells, muscle cells, skin cells, liver cells, brain cells, or nerve cells.

45. 45. The method of claim 44, wherein the cell therapy is applied to the tumor cells associated with a hematological malignancy or a solid tumor.

46. 42. The method of claim 41 , wherein the cell therapy comprises production of at least one of modified chimeric antigen receptor T cells (CAR T cells), gamma delta T cells, natural killer (NK) cells, engineered T cell receptors (TCR), tumor infiltrating lymphocytes (TILs), macrophages, dendritic cells, hematopoietic stem cells (HSCs), or mesenchymal stem / stromal cells (MSCs).

47. the one or more cells are provided from a source comprising at least one of stem cells, pluripotent stem cells, non-stem cells, or cell lines; 42. The method of claim 41, wherein the autologous cells are derived from a source comprising at least one of peripheral blood, bone marrow, umbilical cord blood, placenta, skin, eye, muscle, or tumor.

48. 42. The method of claim 41, wherein the one or more cells are provided from a source comprising at least one of peripheral blood mononuclear cells (PBMCs), umbilical cord blood, stem cells, or skin cells.

49. 42. The method of claim 41, further comprising editing the one or more cells.

50. 42. The method of claim 41, further comprising editing one or more genes in said one or more cells.

51. 42. The method of claim 41, wherein the cell therapy comprises one or more targeting moieties.

52. the one or more targeting moieties 52. The method of claim 51, comprising at least one of a nucleic acid sequence, a protein, a protein fragment, a peptide, a monosaccharide, a polysaccharide, a small molecule, an aptamer, a dendrimer, or a centilin.

53. 42. The method of claim 41, wherein the cell therapy comprises ex vivo cell therapy.

54. 42. The method of claim 41, wherein the cell therapy comprises at least one of regenerative medicine, stem cell therapy, or tissue engineering.

55. 42. The method of claim 41, wherein the cell therapy comprises in vivo cell therapy.

56. 56. The method of claim 55, wherein the in vivo cell therapy comprises at least one of endogenous production of modified chimeric antigen receptor T cells (CAR T cells), natural killer (NK) cells, engineered T cell receptors (TCRs), tumor infiltrating lymphocytes (TILs), or macrophages.

57. 10. The method of claim 1, wherein the CGT comprises a non-genetically modified cell therapy.

58. 58. The method of claim 57, wherein the non-genetically modified cell therapy comprises at least one of regenerative medicine or tissue engineering.

59. 1. A non-transitory computer readable medium containing program instructions that, when executed, cause a data processing system to perform operations for developing or operating a process for cell or gene therapy (CGT), the operations comprising: receiving a plurality of data items; storing the plurality of data items in a hardware storage device; accessing the plurality of data items; determining one or more attributes of the plurality of data items; selecting one or more machine learning models based on the one or more attributes; accessing one or more mechanical models; Integrating the one or more machine learning models with the one or more mechanistic models to obtain one or more integrated models; selecting one or more predictive models from the one or more machine learning models, the one or more mechanistic models, and the one or more integrated models; applying the one or more predictive models to the plurality of data items; adjusting one or more values ​​of one or more parameters of the one or more predictive models to reduce uncertainty in the model predictions; outputting the one or more predictive models having the one or more adjusted values ​​of the one or more parameters; 1. A non-transitory computer-readable medium comprising: