Selection of chromatographic parameters for the production of therapeutic proteins
Machine learning models facilitate the efficient selection of chromatography parameters for therapeutic protein purification, addressing resource-intensive challenges and enhancing process efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- AMGEN INC
- Filing Date
- 2021-04-21
- Publication Date
- 2026-07-22
AI Technical Summary
The conventional selection of chromatography parameters for therapeutic protein purification is resource-intensive in terms of time, cost, labor, and equipment, requiring extensive trial and error, and there is a need for faster and more efficient design and execution of manufacturing processes.
The use of machine learning models to predict chromatography process performance metrics based on process parameters and molecular descriptors, allowing for the efficient selection of chromatography parameters and reducing the need for experimental trials.
This approach significantly reduces the time and resources required for designing and executing downstream manufacturing processes, providing insights into molecular design and improving process efficiency.
Smart Images

Figure 0007893749000006 
Figure 0007893749000007 
Figure 0007893749000008
Abstract
Description
Technical Field
[0001] This application generally relates to the manufacture of biopharmaceutical products, and more particularly, to techniques for modeling a chromatography process (such as a chromatography purification process) to facilitate the selection of chromatography parameters when manufacturing therapeutic proteins.
Background Art
[0002] In the biopharmaceutical industry, large and complex protein molecules known as biopharmaceuticals or therapeutic proteins are derived from biological systems. At a high level, the process of manufacturing a therapeutic protein includes the following steps: (1) a host cell selection stage where a host cell in the selection stage and a universal cell line containing the gene for the desired protein are generated (e.g., using Chinese hamster ovary (CHO) cells); (2) a cell culture stage where a defined culture medium is used to grow a large number of cells that produce the protein in a bioreactor; (3) a purification stage where the product is recovered and purified from the previous stage to isolate the protein; and (4) a formulation and fill-finish-packaging stage where the protein is prepared for use by a physician or patient.
[0003] Figure 2 shows a typical therapeutic protein manufacturing process 10. In the first stage 12, the “upstream” manufacturing process begins by receiving the cryopreserved cells, after optimal cells producing high concentrations of the desired protein have been engineered and cryopreserved in vials or cell bags. The cells are typically thawed in small T-flasks, shaking flasks, or spinner flasks and grown in increased numbers and increased flask sizes to achieve inoculation into a seed bioreactor in stage 14. Throughout the growth process, the cells are maintained under controlled conditions (temperature, pH, and / or nutrients) for sustained growth. After one or more stages of culture volume expansion (indicated as stage 16 in Figure 2), the cells are inoculated into a manufacturing bioreactor in stage 18. During stage 18, the therapeutic protein is expressed by the cells. After this step, the “downstream” process begins. In the downstream process, in stage 20, centrifugation or deep filtration is performed to separate the culture medium from the cells and / or separate the desired protein from other molecules in the bioreactor. In Stage 22, the chromatographic purification process further isolates the desired protein from host cells and impurities or other undesirable substances (e.g., degraded or aggregated proteins). Various filtration technologies may be used in Stage 22 to isolate and purify the protein based on its size, molecular weight, and charge. The resulting material undergoes viral filtration in Stage 24. The purified protein is typically prepared with excipients to produce a sterile solution that can be injected or infused. In Stage 26, the material is concentrated and placed in a target buffer to produce a formulation which is then placed in a container (e.g., a vial or syringe) for labeling, long-term storage, and transport. While this exemplary therapeutic agent manufacturing process 10 is provided for illustrative purposes, it will be understood that the selection of chromatographic parameters described herein is readily applicable to other therapeutic protein manufacturing processes that involve chromatography.
[0004] Generally, "chromatography" (as performed, for example, in Stage 22) refers to a separation process in which molecules are distributed between two phases: (1) a stationary phase, often a chromatography resin; and (2) a mobile phase, which in the case of protein separation is a solvent such as water or chloroform. Molecules that are more strongly attracted to the stationary phase are attracted more slowly through the system compared to molecules that are more strongly attracted to the mobile phase. For commercial manufacturing and purification, chromatography is typically performed as column chromatography, due to considerations of scale. In a typical chromatographic operation, a certain amount of sample is injected into a column. The eluate is then pumped through the column, and molecules are separated based on their relative affinity to the stationary phase resin and the eluate. Different molecules will elute from the column at different time points and after different amounts of eluate have passed through the column. Thus, therapeutic proteins can be separated from other substances elute from the column at different time points. This information is captured in a chromatogram, which is a plot of concentration versus time eluten from the column.
[0005] Hydrophobic interaction chromatography can be used to separate proteins based on their hydrophobicity differences, affinity chromatography can be used to separate molecules based on their affinity for target ligands attached to a chromatography resin, and ion exchange chromatography can be used to separate molecules based on their molecular charge differences. As a more specific example, cation exchange chromatography (CEX) is an ion exchange chromatography used when the molecule of interest is positively charged. Proteins have amino acids with acidic and basic side chains. Depending on the acidity level (pH) of the solution surrounding the biopharmaceutical, the molecule can be positively charged, negatively charged, or neutral. The isoelectric point (pI) is the pH at which the number of protonating and deprotonating groups is equal, and the protein has no net charge. If the pH is higher than pI, the protein will have a net negative charge, and if the pH is lower than pI, the protein will have a net positive charge. The pI of a protein can be determined by the primary amino acid sequence of the protein and thus calculated, and a buffer can be selected to guarantee the known net charge of the protein of interest. Proteins with different pI values will have varying degrees of charge at a given pH; therefore, different proteins will bind to resins with different strengths, facilitating their separation through the column. Other common types of chromatography include size exclusion chromatography (SEC) and protein A chromatography, where molecules in solution are separated by size and / or molecular weight. [Overview of the project] [Problems that the invention aims to solve]
[0006] Conventionally, selecting chromatography parameters (e.g., pH of elution buffer, conductivity of elution buffer, molar concentration of elution buffer, gradient slope, linear velocity, loading, and collection time) and determining how the purification stage is carried out for a particular product / molecule (e.g., using a specific solution, specific pH, etc.) can be highly resource-intensive in terms of time, cost, labor, and equipment use, and may require a great deal of trial and error by setting up and running many experiments to obtain empirical measurements. However, as the pace of biotechnology advances and the processing of additional molecules in the pipeline becomes increasingly emphasized, the need to design and execute manufacturing processes, including chromatographic purification processes, more quickly is increasing. [Means for solving the problem]
[0007] Embodiments described herein relate to systems and methods for creating and applying one or more models to predict the performance of a purification process in the production of therapeutic proteins. The therapeutic proteins may be any preferred type of protein, such as monoclonal antibodies ("mAb") or bispecific or other multispecific antibodies. More specifically, in these embodiments, machine learning models are used to predict performance metrics (e.g., product yield and / or quality measures) of a chromatographic purification process, such as CEX, SEC, Protein A, or any other preferred chromatographic process, based on various process parameters (e.g., buffer and / or elution buffer and / or load pH, molar concentration of elution buffer, conductivity of elution buffer, gradient slope, linear velocity, load conductivity, load factor, collection stop, column volume, substantial CEX load (if CEX is used), load flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end) and molecular descriptors (e.g., mathematical representation of the physical properties of a molecule). This process can lead to better selection of chromatography process parameters compared to conventional processes, resulting in a substantial reduction in the amount of time required to design / develop / execute downstream manufacturing processes and / or a substantial reduction in the use of other resources (e.g., labor, equipment, costs) (e.g., by eliminating or reducing the need to perform experiments). Furthermore, this process can elucidate how the physical characteristics of different molecules (related to various molecular descriptors) affect the performance of chromatography, thereby providing insights into molecular design.
[0008] Performance index values are typically intended to depend on process parameters. As described herein, non-null process parameters can be predicted based on one or more performance index values, which allows for efficient prediction of process parameters (or their precision range) based on one or more desired process parameters. Several embodiments herein relate to systems and methods for constructing and applying one or more models that facilitate the selection of chromatographic parameters for a purification process based on one or more desired performance index values. These methods can be used to facilitate the selection of chromatographic parameters for a purification process during the production of therapeutic proteins. In these embodiments, machine learning models are used to predict process parameters (e.g., buffer and / or elution buffer and / or load pH, molar concentration of elution buffer, conductivity of elution buffer, gradient slope, linear velocity, load conductivity, load factor, collection stop, column volume, substantial CEX load (if CEX is used), load flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end) based on various process parameters and molecular descriptors (e.g., mathematical representation of the physical properties of molecules), and based on one or more performance indicators (e.g., product yield and / or quality measure) of a suitable chromatography process, such as CEX, SEC, Protein A or any other suitable chromatography process.
[0009] Furthermore, interpretable machine learning algorithms can be used to identify the most important input features (e.g., molecular descriptors and process parameters or performance metrics) for making accurate predictions. This can be particularly beneficial given the large number of process parameters, performance metrics, and especially potential molecular descriptors (e.g., hundreds or even thousands of potential molecular descriptors). Thus, it may be possible to create sufficiently accurate descriptors for purification processes using a relatively small number of input features, eliminating the need to measure or compute many other parameters and / or descriptors. Knowledge of the correlation between input parameters / descriptors and predicted targets can provide scientific insights and give rise to hypotheses for further research that may lead to improvements in future bioprocesses.
[0010] Those skilled in the art will understand that the drawings described herein are included for illustrative purposes only and do not limit the disclosure. The drawings are not necessarily to scale and instead focus on illustrating the principles of the disclosure. In some cases, various aspects of the embodiments described may be shown in exaggeration or enlargement to facilitate understanding of the embodiments described. In the drawings, similar reference numerals throughout the drawings refer to components that are generally functionally and / or structurally similar. [Brief explanation of the drawing]
[0011] [Figure 1] This is a simplified block diagram of an exemplary system in which the technologies described herein may be implemented. [Figure 2] This shows the prior art process for manufacturing the active pharmaceutical ingredient. [Figure 3] This is a flowchart illustrating the exemplary process for generating a machine learning model for use in the system shown in Figure 1. [Figure 4A] This document presents an exemplary feature importance scale for predicting experimental yield or SE-HPLC HMW using the eXtreme gradient boost model. [Figure 4B]This document presents an exemplary feature importance scale for predicting experimental yield or SE-HPLC HMW using the eXtreme gradient boost model. [Figure 5A] This is a flowchart illustrating an exemplary method for facilitating the selection of chromatographic parameters for the synthesis process during the manufacturing of therapeutic proteins. [Figure 5B] This is a flowchart illustrating an exemplary method for facilitating the selection of chromatographic parameters for the synthesis process during the manufacturing of therapeutic proteins. [Modes for carrying out the invention]
[0012] The various concepts described above as an introduction and discussed in more detail below can be implemented in any of many ways, and the concepts described are not limited to any particular form of embodiment. Examples of embodiments are provided for illustrative purposes.
[0013] Figure 1 is a simplified block diagram of an exemplary system 100 in which the techniques described herein may be implemented. System 100 includes a computing system 102 that is communicably connected to a training server 104 via a network 106. Generally, the computing system 102 and / or the training server 104 are configured to train one or more machine learning (ML) models 108, and the trained models are used to predict the performance (e.g., yield and / or product quality measures) of a hypothetical chromatography process that can be used in the production of therapeutic proteins. It should be understood that the term “hypothetical” as used herein does not necessarily mean that a corresponding real-world process does not exist. For example, the predicted performance can be compared to the performance measured by running one of the ML models 108 concurrently with or after a corresponding real-world chromatography purification process. The chromatography purification process may include at least one of the CEX process, SEC process, protein A chromatography process, or any other suitable chromatography process.
[0014] The ML model 108 can predict performance based on process parameters (e.g., pH of elution buffer, salt concentration, column volume, etc.), molecular descriptors (e.g., parameters including or relating to molecular charge, hydrophobicity, isoelectric point, dipole moment, etc.), and / or other numerical and / or categorical parameters (e.g., in the form of monoclonal antibodies (mAbs), etc.) or bispecific antibodies, etc.). The computing system 102 is also configured to allow one or more users, which may be located locally or remotely, to utilize the predictive capabilities of the computing system 102 and to provide users with various interactive capabilities as discussed elsewhere in this specification.
[0015] Network 106 may be a single communication network or may include one or more types of communication networks (e.g., one or more wired and / or wireless local area networks (LANs) and / or one or more wired and / or wireless wide area networks (WANs) such as the Internet). In various embodiments, the training server 104 trains and / or uses the ML model 108 as a “cloud” service (e.g., Amazon Web Services), or the training server 104 may be a local server. However, in the illustrated embodiment, the ML model 108 is trained by the server 104 and, if necessary, transferred to the computing system 102 via Network 106. In other embodiments, one, some, or all of the ML models 108 may be trained on the computing system 102 and then uploaded to the server 104. In yet another embodiment, the computing system 102 trains and maintains / stores the model 108, in which case system 100 can exclude both Network 106 and the training server 104, or the server 104 may be part of the computing system 102.
[0016] The computing system 102 may include one or more general-purpose computers specifically programmed to perform the operations discussed herein, and / or one or more special-purpose computing devices. As is evident from Figure 1, the computing system 102 includes a processing unit 120, a network interface 122, a display 124, a user input device 126, and a memory unit 128. In embodiments in which the computing system 102 includes two or more computers (either located in the same place or in separate locations), the operations described herein relating to at least the processing unit 120, the network interface 122, and / or the memory unit 128 can be divided among each of the multiple processing units, multiple network interfaces, and / or multiple memory units. Furthermore, although the display 124 and the user input device 126 are referred to singly herein, each of them may include multiple displays and multiple user input devices. For example, the display 124 may include at least one display in each of many remote user-specific client devices, and the user input device 126 may include at least one user input device for each of those client devices.
[0017] The processing unit 120 includes one or more processors, each of which can be a programmable microprocessor that executes software instructions stored in the memory unit 128 to perform some or all of the functions of the computing system 102 described herein. The processing unit 120 can include, for example, one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs). Alternatively or in addition, some of the processors within the processing unit 120 can be other types of processors (e.g., application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.), and some of the functions of the computing system 102 described herein can be implemented in hardware instead.
[0018] The network interface 122 can include any suitable hardware (e.g., front-end transmitter and receiver hardware), firmware, and / or software configured to communicate with the training server 104 via the network 106 using one or more communication protocols. For example, the network interface 122 can be or include an Ethernet interface that enables the computing system 102 to communicate with the training server 104 over the Internet or an intranet, etc.
[0019] The display 124 may use any suitable display technology (e.g., LED, OLED, LCD, etc.) to present information to the user, and the user input device 126 may be a keyboard or other suitable input device. In some embodiments, the display 124 and the user input device 126 are integrated within a single device (e.g., a touchscreen display). Generally, the display 124 and the user input device 126 can be combined to enable the user to interact with a graphical user interface (GUI) provided by the computing system 102. However, in certain embodiments where the computing system 102 interacts with other computing devices or systems (e.g., third-party client devices) to enable user interaction with its device or system, the display 124 and / or the user input device 126 may be excluded.
[0020] The memory unit 128 may include one or more volatile and / or non-volatile memories. It may include one or more any suitable memory types, such as read-only memory (ROM), random access memory (RAM), flash memory, solid-state drives (SSDs), and hard disk drives (HDDs). Collectively, the memory unit 128 may store one or more software applications, data received / used by those applications, and data output / generated by those applications. These applications, when executed by the processing unit 120, include a chromatography modeling application 130 that predicts the performance (e.g., yield and / or quality measures) of a hypothetical chromatography process for purification during the production of therapeutic proteins. In some embodiments, the various “units” of application 130 considered herein may be distributed among various software applications, and / or the functionality of any one such device may be divided among two or more software applications.
[0021] In the exemplary system 100, the application 130 includes a data collection unit 132, a prediction unit 134, and a visualization unit 136. Generally, the data collection unit 132 receives (e.g., reads) parameters that are applied as inputs to the local machine learning (ML) model 138 for the prediction unit 134 to predict performance metric values (e.g., yield or product quality metrics) or process parameters (or their accuracy ranges). In the illustrated embodiment, the ML model 138 is a local copy of the model 108 trained by the training server 104 and can be stored, for example, in the RAM of the memory unit 128. However, as described above, in some embodiments, the server 104 can utilize / run the entire model 108, and in any case, there is no local copy that needs to be present in the memory unit 128, or not all of the model 108 is read from the training server 104 as needed, but rather can be present in the continuous memory of the memory unit 128. The data collection unit 132 can receive values from GUI (e.g., on the display 124) user input parameters / values generated or added by the visualization unit 136, and / or can receive values, for example, as one or more files or other data transfers (e.g., using a file path specified by the user via such a GUI).
[0022] The visualization unit 136 may also generate and / or add a GUI for viewing and / or interacting with, for example, the predicted results of a modeling process (e.g., performance index value output or process parameter output by the prediction unit 134 using model 138). Depending on the embodiment, the visualization unit 136 may also provide tools for the user to develop a useful model (e.g., to identify the most predictive features for a given performance index value or process parameter, to optimize hyperparameters for the model), and / or tools to utilize such a model when the user designs (e.g., optimizes) a chromatographic purification process (e.g., to achieve a process with high yield and excellent quality attributes such as high consistency / reproducibility).
[0023] In the exemplary system 100, the memory unit 128 also stores software instructions for a molecular manipulation environment (MOE) application 139 that provides homology modeling for therapeutic proteins of interest. Generally, the MOE application 139 is configured to generate descriptors for molecules (e.g., mathematical representations of molecular physical features such as charge, hydrophobicity, dipole moment, isoelectric point) based on input information about the molecule. For example, a user can use the MOE application 139 to input the amino acid sequence of a molecule (e.g., via the user input device 126) and select an appropriate molecular template. The MOE application 139 can then attempt to "fit" the amino acid sequence to the selected template. Alternatively, or in addition, the MOE application 139 can generate descriptors based on experimental / measurement results for the molecule. In alternative embodiments, the MOE application 139 is stored and executed by a computing device or system other than the computing system 102 (e.g., by a third-party computing device or system).
[0024] Herein, the operation of system 100 according to one embodiment is described in more detail. First, the training server 104 trains the ML model 108 using historical data stored in the training database 140. The training database 140 may include a single database stored in a single memory (e.g., HDD, SSD, etc.) or multiple databases stored in one or more memories. The ML model 108 may include many different types of machine learning models, such as eXtreme gradient boost (or "xgboost") models, regression (or "decision" or "ID") tree models, elastic net models, lasso models, ridge models, stochastic gradient descent (SGD) normalized loss linear models, linear support vector machine (SVM) models, partial least squares (PLS) regression models and / or one or more other suitable model types. Furthermore, various models of ML model 108 can be trained to predict different performance indicator values (e.g., yield, specific CEX readout, specific SEC readout, etc.) or different process parameters (e.g., buffer and / or elution buffer and / or load pH, elution buffer molar concentration, elution buffer conductivity, gradient slope, linear velocity, load conductivity, load factor, collection stop, column volume, actual CEX load (if CEX is used), load flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end). In some embodiments, for example, ML model 108 includes, in detail, a regression tree model for predicting one or more performance indicator values of a first set, an xgboost model for predicting one or more performance indicator values of a different second set, and an elastic net model for predicting one or more performance indicator values of a different third set. In some embodiments, for example, the ML model 108 includes, in detail, a regression tree model for predicting one or more process parameters of a first set, an xgboost model for predicting one or more process parameters of a different second set, and an elastic net model for predicting one or more process parameters of a different third set.Furthermore, in some embodiments, the ML model 108 may include two or more models of any given type (e.g., two or more models of the same type trained on different historical datasets using different feature sets and / or having different hyperparameters). As discussed in some embodiments and in more detail in conjunction with Figures 4A and 4B, each of the ML models 108 may be used to identify which feature (e.g., process parameters, molecular descriptors, etc.) is the best predictor of a particular performance metric value (where applicable) and / or may be trained or retrained using a feature set that includes only the features that are the best predictors of a particular performance metric value or process parameter.
[0025] For each different model within the ML model 108, the training database 140 may store corresponding sets of training data (e.g., input / feature data and corresponding labels), which may overlap between training datasets. To train a model to predict yield percentages, for example, the training database 140 may include many sets of input / feature data that may be produced by analytical instruments, each containing historical process parameters (e.g., pH level, load flow rate, salt concentration, etc.), as well as molecular descriptors calculated by software for the protein being manufactured (e.g., MOE application 139 or similar software) (e.g., descriptors relating to charge, hydrophobicity, isoelectric point, etc.) and possibly other information (e.g., the form of the protein being manufactured), along with labels for each set of feature data. In this embodiment, the labels for each set of feature data indicate the yield percentage if a particular protein were measured in a chromatographic process. In some embodiments, all features and labels are numerical, and non-numerical classifications or categories are mapped to numerical values (for example, the acceptable values for modality features / inputs [Monoclonal, Bispecific Format 1, Bispecific Format 2, Bispecific Format 1 or 2] are mapped to values [00, 10, 01, 11]).
[0026] In some embodiments, the training server 104 uses additional labeled datasets in the training database 140 to validate the trained ML models 108 (for example, to ensure that one given ML model 108 provides at least some minimum acceptable accuracy). In some embodiments, the training server 104 also continuously updates / improves one or more ML models 108. For example, after the ML models 108 have been initially trained to provide a sufficient level of accuracy, additional measurements of chromatographic performance metrics (and corresponding inputs / features) can be used to improve predictive accuracy.
[0027] Application 130 may retrieve a specific ML model 108 corresponding to a performance metric of interest from the training server 104 via network 106 and network interface 122. The performance metric may be one specified by the user via a GUI generated or added by the visualization unit 136, for example. Once the model is retrieved, the computing system 102 saves a local copy as the local ML model 138. In other embodiments, as described above, the model is not retrieved, and instead, the input / data is sent to the training server 104 (or another server) to use the appropriate model of model 108 as needed, or all models 108 may reside only in the computing system 102.
[0028] The performance metrics used to train a specific model 108 or model 138 for prediction may include metrics of any aspect of performance, such as yield or product quality (e.g., purity). Furthermore, the performance metrics may be general to different types of chromatography (e.g., yield) or specific to a particular type of chromatography (e.g., CEX, SEC, Protein A, etc.). For example, and without limitation, ML Model 108 or 138 may have the following index values: process yield, CEX acidity (%), CEX main (%), CEX basic (%), SEC high molecular weight (HMW) (%), SEC main (%), SEC low molecular weight (LMW) (%), reduced sample preparation (rCE-SDS) main (%), rCE-SDS LMW (%), rCE-SDS capillary electrophoresis with light chain (LC) + heavy chain (HC) (%), reduced sample preparation (nrCE-SDS) main (%), rCE-SDS pre-LC (%), rCE-SDS LC (%), rCE-SDS non-glycosylated heavy chain (NGHC) (%), rCE-SDS HC (%), rCE-SDS It is possible to predict any of the following: HMW (%), capillary electrophoresis sodium dodecyl sulfate (nrCE-SDS) without rCE-SDS pre-LC+LC+HC (%), pool conductivity (mS / cm), capillary isoelectric focusing (cIEF) acidity (%), cIEF basicity (%), cIEF main (%), nrCE-SDS pre-peak (%), host cell protein (HCP), and SE-HPLC HMW. In some embodiments, the process parameters to which a particular model 108 or model 138 is trained to predict may include index values for any aspect of the process or its accuracy range (e.g., confidence intervals such as 80%, 85%, 90%, or 95%).For example, and without limitation, ML Model 108 or 138 can predict any of the following process parameters: buffer and / or elution buffer and / or load pH, molar concentration of elution buffer, conductivity of elution buffer, gradient slope, linear velocity, load conductivity, load factor, collection stop, column volume, actual CEX load (if CEX is used), load flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end.
[0029] The data acquisition unit 132 collects the necessary data according to the feature set used by the model 138. For example, the data acquisition unit 132 can receive user-inputted process parameters (or performance metrics as needed) and molecular descriptor output from the MOE application 139 for the therapeutic protein of interest (for example, after the user inputs or otherwise provides the amino acid sequence of the protein to the MOE application 139). The process parameters and performance metrics may be as described herein. For example, process parameters may include, for example and without limitation, any parameters related to the conditions or characteristics of a hypothetical chromatography process, such as buffer pH, elution buffer conductivity (mS / cm), elution buffer molar concentration (mM), elution buffer pH, gradient slope (mM / CV), linear velocity (cm / hr), loading conductivity (mS / cm), loading factor (g / Lr), loading pH, collection stop (%), column volume, actual CEX loading, loading flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end.
[0030] The molecular descriptors generated by MOE application 139 include, for example and without limitation, the following descriptors that are publicly known to users of MOE software: pH, HI, pro_Fv_net_charge, U, asa_hyd, viscosity, hyd_idx, pro_helicity, apol, asa_hph, pro_net_charge, hyd_idx_cdr, pro_henry, b_1rotR, volume, amphipathicity, hyd_strength, pro_hyd_moment, b_rotR, mobility, ASPmax, hyd_strength_cdr, pro_mass, density, helicity, BSA, Packing Score, pro_mobility, ens_dipole, henry, BSA_HC, pro_affinity, pro_pI_3D, mass, net_charge, BSA_LC_HC, pro_app_charge, pro_pI_seq, pI_seq, app_charge, contact This may include any or all of the following preferred descriptor types: energy, pro_asa_hph, pro_r_gyr, pI_3D, dipole_moment, DRT, pro_asa_hyd, pro_r_solv, coeff_fric, hyd_moment, E bond, pro_asa_vdw, pro_sed_const, coeff_diff, zeta, E ele, pro_cdr_net_charge, pro_stability, r_gyr, zdipole, E sol, pro_coeff_diff, pro_volume, r_solv, zquadrupole, E vdw, pro_coeff_fric, pro_zdipole, sed_const, Eint_VL_VH, pro_dipole_moment, pro_zeta, eccen, GB / VI, pro_eccen, pro_zquadrupole, and / or asa_vdw. Preferably, in some embodiments, the molecular descriptor includes at least one descriptor that is a function (or otherwise dependent on) of the pH level of the molecular environment and, if applicable, more such descriptors.For example, various descriptors may depend on the surface charge of the protein molecule, and the surface charge may, in turn, depend on the pH of the molecular environment.
[0031] After the data acquisition unit 132 has collected process parameters (or performance metrics, if applicable) and molecular descriptors for a specific hypothetical chromatography process (including, if applicable, other data such as the protein type entered by the user), the prediction unit 134 triggers an ML model 138 to operate on its inputs / features to predict desired performance metrics (or process parameters) for the hypothetical chromatography process. In some embodiments and / or scenarios, the prediction unit 134 may obtain multiple different local ML models 138 from the training server 104 to predict different performance metrics (or process parameters, if applicable) for the same hypothetical chromatography process (e.g., in parallel or sequentially), where it is understood that the local ML models 138 operate on the same or different features to generate each predictor.
[0032] The visualization unit 136 causes the GUI displayed on the display 124 to present predicted performance metric values (or process parameters, if applicable) and / or other information derived from the predicted performance metric values (or process parameters, if applicable). For example, the visualization unit 136 may cause the GUI to display whether the predicted performance metric values meet one or more pass criteria (for example, after application 130 compares the performance metric values to one or more respective thresholds). For example, the visualization unit 136 may cause the GUI to display the accuracy range of the predicted process parameters.
[0033] The above prediction / visualization process can be repeated across many different hypothetical chromatography processes (for example, for different combinations of process parameters and fixed sets of molecular descriptors for therapeutic proteins), thereby enabling users to quickly test different process designs. Users can also quickly test specific aspects of the design, such as how small perturbations affect predicted performance indicators for a particular input (e.g., reflecting predicted ranges in elution buffer pH or load flow rate). The visualization unit 136 can generate or add one or more GUIs to help the viewing user understand and consider the prediction results for various hypothetical chromatography processes. In this way, the viewing user can be informed and select which chromatography process parameters to use in a real-world commercial manufacturing process (undergoing any necessary certification tests). The selection of chromatography process parameters should generally attempt to maximize yield while minimizing impurities (with emphasis on one of these goals rather than other goals that may depend on the use case / project goals), and can be determined by one or more users based on the displayed information, or can be fully automated according to some predefined selection criteria. In some cases, the techniques described herein can be used not only for selecting chromatography process parameters, but also, or alternatively, to provide relative insights into the interrelationship between novel molecules and the purification process (e.g., by fine-tuning molecular descriptors and observing their effects on various performance metrics). These insights may help guide future molecular designs by identifying key molecular properties that affect purification effectiveness.
[0034] To avoid the time and expense required to implement and collect extremely large amounts of labeled historical data, an interpretable machine learning model can be used as Model 108. For example, a training server 104 can train one of Model 108 on hundreds of features, and then the training server 104 (or a human reviewer) can analyze the trained model (e.g., the weights assigned to each feature) to determine the most predictive features (e.g., about 10 features or about 50 features). That particular Model 108, or a new version of that Model 108 trained using only the most predictive features, can then be used with a much smaller set of features. Identifying highly predictive features can also be useful for other purposes, such as providing new scientific insights that may give rise to new hypotheses (which could then lead to improvements in bioprocesses).
[0035] Various techniques for determining which model is best suited to a particular performance metric (and / or process parameter) and for identifying the most predictive features for a given model or use case are described below with reference to Figures 3 and 4.
[0036] Generally, models that perform well for specific performance metrics can be identified by training many different model types using real-world historical training data from previous chromatographic purification processes and comparing the results. Figure 3 illustrates an exemplary process 300 that may be used for this purpose for a particular performance metric of interest. In the first stage 302 of process 300, data related to the performance metric is selected (i.e., identified and obtained). However, historical data is often inconsistent with different types and / or forms of data captured for different drug products or projects, for example. Therefore, in stage 304, it may be necessary to attribute missing values and / or perform other steps (e.g., normalization, outlier removal) to ensure a robust set of training data.
[0037] In Stage 306, each candidate model is trained on at least a portion of the historical data using hyperparameters optimized for that specific model. Stage 306 may include performing k-fold verification for each model (e.g., using k=10 (where the model is trained and evaluated 10 times across various 90 / 10 sections of the dataset selected in Stage 302 and augmented in Stage 304) or k=5, etc.). Stage 306 may include tuning the hyperparameters of each model using Bayesian search techniques. Bayesian techniques perform a computationally more efficient Bayesian guided search than grid search or random search, but achieve a similar level of performance to random search. Stage 306 may include selecting model hyperparameters through several iterations of Bayesian search and k-fold verification.
[0038] In Stage 308, various candidate models (including their finely tuned hyperparameters) are evaluated, and the best model (for the performance metric or process parameter of interest) is selected. The "best" model can be selected using any appropriate criterion, for example, the coefficient of determination (R).2 Algorithmic performance metrics such as ) and / or root mean square error (RMSE) can be captured for each model, and their respective mean values are obtained based on a cross-checking process. 2 teeth,
number
number
number
number
[0039] RMSE indicates the accuracy / error of the model within an easily understandable unit of the predicted performance metric (or process parameter), so R 2 It could be a better measure than R 2 The metric can, in some cases, produce extremely negative values when using several cross-checking sets, which can distort model comparisons when averaged across the entire set. RMSE may be preferable to mean absolute error (MAE) because the former imposes a larger error between predictions and actual results.
[0040] Subsequently, in Stage 310, the final model for predicting performance metrics (or process parameters) is output / selected, for example, based on a comparison of RMSE (and / or one or more other metrics) for different models. The final model can be retrained on the entire dataset. The final generated model is then saved as the next trained model (e.g., one of ML Model 108) and ready to make predictions for novel / future chromatography processes.
[0041] In one embodiment, process 300 is performed by the training server 104 in Figure 1 (with human input at various stages, such as selecting performance metrics of interest and selecting candidate models, if applicable). Process 300 can be repeated for any suitable number of performance metrics (e.g., 5, 10, 20, etc.), including each performance metric of interest and any of the exemplary performance metrics discussed elsewhere in this specification. Once the final models for different performance metrics are output at each iteration of stage 310, the training server 104 can add these final models to the ML model 108. Subsequently and before making predictions about a particular hypothetical chromatographic purification process in the manner discussed herein (see, for example, Figure 1), the computing system 102 or the training server 104 can select an appropriate final model from the ML model 108. The selection can be made, for example, based on user input indicating a desired performance metric.
[0042] In general, following Process 300, many models have been identified as having superior performance in various different performance indicators for chromatographic purification processes during the production of therapeutic proteins, based on RMSE and as shown in Table 1 below.
[0043] [Table 1]
[0044] The results in Table 1 are all related to historical datasets for various monoclonal antibodies. For each performance metric, the model that produced the lowest RMSE was interpreted as the "best-running" model. As is clear from Table 1, no model ran best to predict all performance metrics. Rather, the regression tree model ran best for 12 performance metrics, the xgboost model ran best for 8 performance metrics, and the elastic net ran best for 2 performance metrics. Other models evaluated using process 300 (lasso, ridge, SGD, linear SVM, and PLS) did not run best for any of the performance metrics in Table 1. The regression tree and xgboost models ran particularly well with both a large and small number of observational findings (training datasets).
[0045] Other processes may have different performance characteristics due to a larger or smaller number of training datasets, such as evaluating the model based on different performance metrics (and / or combinations of different performance metrics), or evaluating the model using different measures (other than RMSE). For example, to predict yield and R 2 When evaluating machine learning models for predicting SEC-HMW percentage according to a scale (particularly for monoclonal antibodies), the best-performing model was found to be the xgboost model. For this latter evaluation, process parameter inputs to the xgboost model included column volume, loading pH, actual CEX loading, loading flow rate, elution flow rate, buffer concentration, elution pH, gradient slope, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end.
[0046] As described above, if the “best” model is identified / output in stage 310, it may be beneficial to learn which features are most important for a particular model so that only the features that best predict the desired performance metric are used. Figures 4A and 4B show example feature importance scales (relative feature importance and correlation with features, respectively) plots 400 and 420 for both predicting experimental yield and SE-HPLC HMW using the eXtreme gradient boost (xgboost) model. Plots 400 and 420 can be generated, for example, by the visualization unit 136 and presented on the GUI via the display 124.
[0047] Plots such as plots 400 and 420, for example, can allow users (e.g., scientists) to easily identify the most important factors for predicting specific performance indicators (in this case, experimental yield and SE-HPLC HMW). This can also provide greater insight into how molecular structure affects the purification process. For example, while it is common knowledge that hydrophobicity plays a role in increasing impurities in the CEX process, it is generally believed that molecular charge will have the greatest impact. However, feature importance plots similar to plots 400 and / or 420 (or correlational thermal maps, etc.) consistently and surprisingly rank hydrophobicity as a more important indicator of higher levels of impurities (specifically high HMW) according to the CEX process. Furthermore, not all forms of hydrophobicity have an equal impact on impurity levels. For example, plots similar to plot 420 (or correlational thermal maps, etc.) show that some forms of hydrophobicity are indicators of lower impurities (in particular, low HMW), at least compared to others.
[0048] Figure 5A is a flowchart of an exemplary method 500 for facilitating the selection of chromatographic parameters for a purification process during the production of therapeutic proteins. Method 500 may be executed by one or more processors of a computing system 102 or, for example, a server 104 (e.g., in the execution of a cloud service), at least in part, by executing software instructions of an application 130 stored in a memory unit 128.
[0049] Block 502 receives one or more process parameter values related to a hypothetical chromatography (e.g., CEX, SEC, or Protein A chromatography) process. Process parameter values can be received via the user interface and / or by importing, for example, a file or other data. As an example and without limitation, process parameter values may include one or more of the following: buffer pH, elution buffer conductivity (mS / cm), elution buffer molar concentration (mM), elution buffer pH, gradient slope (mM / CV), linear velocity (cm / hr), load conductivity (mS / cm), load factor (g / Lr), load pH, collection stop (%), column volume, actual CEX load, load flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end. While illustrative units are shown in the preceding list, it should be understood that these units are for illustrative purposes only, and these parameters may be communicated in any suitable units. Therefore, it will be understood that exemplary units may be excluded from the preceding list or any other list of process parameters herein.
[0050] In block 504, one or more molecular descriptors describing a therapeutic protein are received. Molecular descriptors may be received via a user interface and / or by importing, for example, a file or other data (e.g., from MOE application 139). In some embodiments, method 500 further includes determining one or more molecular descriptors based on sequence information associated with the therapeutic protein (e.g., amino acid sequence information entered into the MOE software) and / or experimental measurements of the physical properties of the therapeutic protein (e.g., measurement results entered into the MOE software). In some embodiments, at least one molecular descriptor is a function of the pH of the environment surrounding the molecule (e.g., a mathematical function that varies if the pH is known / specified).Saturation: pH, HI, pro_Fv_net_charge, U, asa_vis_hyd, and hycosity d_idx、pro_helicity、apol、asa_hph、pro_net_charge、hyd_idx_cdr、pro_henry 、b_1rotR、volume、amphipathicity、hyd_strength、pro_hyd_moment、b_rotR、mobility、ASPmax、hyd_strength_cdr、pro_mass、density、BSA helicity、Packing Score、pro_mobility、ens_dipole、henry、BSA_HC、pro_affinity、pro_pI_3D、mass、net_charge、BSA_LC_HC、pro_app_charge、pro_pI_seq、pI_seq_app_contact energy、pro_asa_hph、pro_r_gyr、pI_3D、dipole_moment、DRT、pro_asa_hyd、pro_r_solv、coeff_fric、hyd_moment、E bond、pro_asa_vdw、pro_sed_const、coeff_diff、zeta、E ele、pro_cdr_net_charge、pro_stability、r_gyr、zdipole、E sol、pro_coeff_diff、pro_volumev、zquaolev vdw、pro_coeff_fric、pro_zdipole、sed_const、Eint_VL_VH、pro_dipole_moment、p ro_zeta、eccen、GB / VI、pro_eccen、pro_zquadrupole.
[0051] In block 506, performance metrics for the hypothetical chromatography process are predicted by analyzing process parameters received in block 502 and molecular descriptors received in block 504 using a machine learning model. The machine learning model may be a regression tree model, an eXtreme gradient boost (xgboost) model, or an elastic network model. For example and without limitation, predicted performance indicators are as follows: process yield, CEX acidity (%), CEX main (%), CEX basic (%), SEC HMW (%), SEC main (%), SEC low molecular weight (LMW) (%), capillary electrophoresis of sodium dodecyl sulfate (rCE-SDS) with reductive sample preparation (rCE-SDS) main (%), rCE-SDS LMW (%), rCE-SDS light chain (LC) + heavy chain (HC) (%), capillary electrophoresis of sodium dodecyl sulfate (nrCE-SDS) without reductive sample preparation (nrCE-SDS) main (%), rCE-SDS pre-LC (%), rCE-SDS LC (%), rCE-SDS non-glycosylated heavy chain (NGHC) (%), rCE-SDS HC (%), rCE-SDS high molecular weight (HMW) (%), rCE-SDS Pre-LC+LC+HC (%), pool conductivity (mS / cm), capillary isoelectric focusing (cIEF) acidic (%), cIEF basic (%), cIEF main (%), nrCE-SDS pre-peak (%), HCP and / or SE-HPLC HMW may be one of these.
[0052] In some embodiments, block 506 includes using a regression tree model to predict nrCE-SDS LC+HC (%), rCE-SDS pre-LC (%), CEX basicity (%), SEC HMW (%), SEC main (%), SEC LMW (%), rCE-SDS HC (%), rCE-SDS HMW (%), rCE-SDS pre-LC+LC_HC (%), pooled conductivity (%), or nrCE-SDS pre-peak (%). In other embodiments, block 506 includes using an eXtreme gradient boost model to predict CEX acidity (%), CEX main (%), process yield, rCE-SDS main (%), rCE-SDS LMW (%), cIEF acidity (%), cIEF basicity (%), or cIEF main (%). Alternatively, block 506 may include using an eXtreme gradient boost model to predict SEC HMW (%) or yield. In yet another embodiment, block 506 includes using an elastic net model to predict rCE-SDS LC+HC(%) or to predict rCE-SDS LC(%).
[0053] In block 508, the performance index values predicted in block 506 and / or the predicted performance index values are presented to the user via a user interface (e.g., a GUI generated or added by the visualization unit 136 and presented on the display 124 in Figure 1) that indicate whether one or more acceptance criteria (e.g., exceeding or falling below several thresholds) are met in order to facilitate the selection of chromatographic parameters for real-world purification processes during the production of therapeutic proteins (e.g., manual selection by the user).
[0054] In some embodiments, Method 500 includes one or more additional blocks not shown in Figure 5A. For example, Method 500 may include two additional blocks, both performed before block 502: a first additional block in which data indicating performance metric values of interest is received from the user via a user interface (e.g., a GUI generated or added by the visualization unit 136 and presented on the display 124); and a second additional block in which a machine learning model (later used in block 506) is selected from among several machine learning models (e.g., ML model 108) trained to predict different performance metrics.
[0055] As another example, method 500 may include four additional blocks similar to blocks 506 and 508 (or 502-508) to generate a second performance metric of interest using a second machine learning model. For example, the first and second machine learning models may be xgboost models trained for different purposes, one predicting an experimental yield percentage and the other predicting an SEC HMW percentage.
[0056] As yet another example, method 500 may include two additional blocks, each occurring after block 506: a first additional block in which one or more process parameter values are selected for a (real-world) chromatographic process for a therapeutic protein based on performance index values and / or instructions presented in block 508, and a second additional block in which the chromatographic process is carried out on the therapeutic protein according to one or more selected process parameter values.
[0057] Figure 5B is a flowchart of another exemplary method 520 for facilitating the selection of chromatographic parameters for the purification process during the production of therapeutic proteins. Method 520 may be executed by one or more processors of a computing system 102 or, for example, a server 104 (e.g., in the execution of cloud services), at least in part, by executing software instructions of an application 130 stored in a memory unit 128.
[0058] Block 522 receives one or more performance metric values related to a hypothetical chromatography process (e.g., CEX, SEC, or Protein A chromatography). These performance metric values can be received via the user interface and / or by importing, for example, a file or other data. For example and without limitation, performance indicators are as follows: process yield, CEX acidity (%), CEX main (%), CEX basic (%), SEC HMW (%), SEC main (%), SEC low molecular weight (LMW) (%), capillary electrophoresis of sodium dodecyl sulfate (rCE-SDS) with reductive sample preparation (rCE-SDS) main (%), rCE-SDS LMW (%), rCE-SDS light chain (LC) + heavy chain (HC) (%), capillary electrophoresis of sodium dodecyl sulfate (nrCE-SDS) without reductive sample preparation (nrCE-SDS) main (%), rCE-SDS pre-LC (%), rCE-SDS LC (%), rCE-SDS non-glycosylated heavy chain (NGHC) (%), rCE-SDS HC (%), rCE-SDS high molecular weight (HMW) (%), rCE-SDS This may include one or more values from pre-LC+LC+HC (%), pool conductivity (mS / cm), capillary isoelectric focusing (cIEF) acidity (%), cIEF basicity (%), cIEF main (%), nrCE-SDS pre-peak (%), HCP and / or SE-HPLC HMW.
[0059] In block 524, one or more molecular descriptors describing the therapeutic protein are received. The molecular descriptors may be received via a user interface and / or by importing, for example, a file or other data (e.g., from MOE application 139). In some embodiments, method 520 further includes determining one or more molecular descriptors based on sequence information associated with the therapeutic protein (e.g., amino acid sequence information entered into the MOE software) and / or experimental measurements of the physical properties of the therapeutic protein (e.g., measurement results entered into the MOE software). In some embodiments, at least one molecular descriptor is a function of the pH of the environment surrounding the molecule (e.g., a mathematical function that varies if the pH is known / specified).Saturation: pH, HI, pro_Fv_net_charge, U, asa_vis_hyd, and hycosity d_idx、pro_helicity、apol、asa_hph、pro_net_charge、hyd_idx_cdr、pro_henry 、b_1rotR、volume、amphipathicity、hyd_strength、pro_hyd_moment、b_rotR、mobility、ASPmax、hyd_strength_cdr、pro_mass、density、BSA helicity、Packing Score、pro_mobility、ens_dipole、henry、BSA_HC、pro_affinity、pro_pI_3D、mass、net_charge、BSA_LC_HC、pro_app_charge、pro_pI_seq、pI_seq_app_contact energy、pro_asa_hph、pro_r_gyr、pI_3D、dipole_moment、DRT、pro_asa_hyd、pro_r_solv、coeff_fric、hyd_moment、E bond、pro_asa_vdw、pro_sed_const、coeff_diff、zeta、E ele、pro_cdr_net_charge、pro_stability、r_gyr、zdipole、E sol、pro_coeff_diff、pro_volumev、zquaolev vdw、pro_coeff_fric、pro_zdipole、sed_const、Eint_VL_VH、pro_dipole_moment、p ro_zeta、eccen、GB / VI、pro_eccen、pro_zquadrupole.
[0060] In block 526, the process parameter values of the hypothetical chromatography process are predicted by analyzing the performance index values received in block 522 and the molecular descriptors received in block 524 using a machine learning model. The machine learning model may be a regression tree model, an eXtreme gradient boost (xgboost) model, or an elastic net model. For example and without limitation, process parameters for which values are predicted may be: buffer pH, elution buffer conductivity (mS / cm), elution buffer molar concentration (mM), elution buffer pH, gradient slope (mM / CV), linear velocity (cm / hr), load conductivity (mS / cm), load factor (g / Lr), load pH, collection stop (%), column volume, actual CEX load, load flow rate, elution flow rate, buffer concentration, gradient length, gradient start point, gradient end point, pool volume, protein concentration, pool start and / or pool end. While exemplary units are shown in the preceding list, it should be understood that these units are for illustrative purposes only, and that these parameters can be communicated in any suitable unit. Therefore, it will be understood that exemplary units may be excluded from the preceding list or any other list of process parameters herein.
[0061] In some embodiments, block 526 includes using a regression tree model to predict process parameter values based on a molecular descriptor and one or more of the following: nrCE-SDS LC+HC (%), rCE-SDS pre-LC (%), CEX basicity (%), SEC HMW (%), SEC main (%), SEC LMW (%), rCE-SDS HC (%), rCE-SDS HMW (%), rCE-SDS pre-LC+LC_HC (%), pool conductivity, and / or nrCE-SDS (%) pre-peaks. In other embodiments, block 526 includes using an eXtreme gradient boost model to predict process parameters based on a molecular descriptor and one or more of the following: CEX acidity (%), CEX main (%), process yield, rCE-SDS main (%), rCE-SDS LMW (%), cIEF acidity (%), cIEF basicity (%), and / or cIEF main (%). Alternatively, block 506 may include using an eXtreme gradient boost model to predict process parameter values based on a molecular descriptor and either or both of SEC HMW (%) and yield. In yet another embodiment, block 506 may include using an elastic net model to predict process parameter values based on a molecular descriptor and rCE-SDS LC+HC (%) or a molecular descriptor and rCE-SDS LC (%).
[0062] In block 528, the process parameter values predicted in block 526 and / or the predicted precision range of the predicted process parameter values are presented to the user via a user interface (e.g., a GUI generated or added by the visualization unit 136 and presented on the display 124 in Figure 1) to facilitate the selection of chromatographic parameters for real-world purification processes during the production of therapeutic proteins (e.g., manual selection by the user).
[0063] In some embodiments, Method 520 includes one or more additional blocks not shown in Figure 5B. For example, Method 520 may include two additional blocks, both performed before block 522: a first additional block in which data indicating process parameters of interest is received from the user via a user interface (e.g., a GUI generated or added by the visualization unit 136 and presented on the display 124); and a second additional block in which a machine learning model (e.g., ML model 108) is selected from among several machine learning models (e.g., ML model 108) trained to predict values for different process parameters (which are later used in block 526).
[0064] As another example, method 520 may include additional blocks similar to blocks 526 and 528 (or 522-528) to generate a second process parameter of interest using a second machine learning model. For example, the first and second machine learning models may be xgboost models trained for different purposes, one predicting the pH of a buffer and the other predicting a loading factor.
[0065] As yet another example, method 520 may include two additional blocks, each occurring after block 526: a first additional block in which one or more process parameter values are selected for a (real-world) chromatographic process for a therapeutic protein based on the information presented in block 528, and a second additional block in which the chromatographic process is performed on the therapeutic protein according to one or more selected process parameter values.
[0066] Systems, methods, apparatus, and their components have been described in terms of exemplary embodiments, but they are not limited to these exemplary embodiments. Detailed descriptions are to be interpreted as examples only, and since it would be impractical, if not impossible, to describe all possible embodiments, not all possible embodiments of the present invention are described. Many alternative embodiments can be carried out using either the current art or art developed after the filing date of this patent, and these still fall within the scope of the claims defining the present invention.
[0067] Those skilled in the art will understand that various modifications, variations, and combinations of the above embodiments can be made without departing from the scope of the present invention, and that such modifications, variations, and combinations are to be interpreted as being within the scope of the concept of the present invention.
Claims
1. A method for facilitating the selection of chromatographic parameters for the purification process during the production of therapeutic proteins, A process in which one or more processors in a computing system receive one or more process parameter values associated with a hypothetical chromatography process; A step of receiving one or more molecular descriptors describing the therapeutic protein using one or more processors; A step of predicting performance index values of the hypothetical chromatography process as output of the machine learning model by applying at least one of the process parameter values and one or more molecular descriptors as input to the machine learning model using one or more processors, wherein the machine learning model is selected from the group consisting of (i) a regression tree model, (ii) an extreme gradient boost model, and (iii) an elastic net model, the regression tree model is used to predict a first set of one or more performance index values, the extreme gradient boost model is selected to predict a second set of one or more different performance index values, and the elastic net model is selected for a third set of one or more different performance index values. The first set of one or more performance index values is: nrCE-SDS LC+HC (%); rCE-SDS Pre-LC (%); CEX Basicity (%); SEC HMW (%); SEC Main (%); SEC LMW (%); rCE-SDS HC (%); rCE-SDS HMW (%); rCE-SDS Pre-LC + LC_HC (%); Pool conductivity; or nrCE-SDS Pre-peak (%) Includes, The second set of one or more performance indicator values is: CEX acidity (%); CEX Main (%); Process yield; rCE-SDS Main (%); rCE-SDS LMW (%); cIEF acidic (%); cIEFF basicity (%); or cIEFF Main (%) Includes, The third set of one or more performance index values is: rCE-SDS LC+HC (%); or rCE-SDS LC (%) Processes including; and The process of having one or more processors present to the user, via a user interface, (i) the predicted performance index value, and (ii) an indicator of whether the predicted performance index value meets one or more pass criteria. A method that includes this.
2. A method for facilitating the selection of chromatographic parameters for the purification process during the production of therapeutic proteins, A step of receiving one or more performance index values associated with a hypothetical chromatography process by one or more processors of a computing system, wherein the one or more performance index values can be selected from a first set of one or more performance index values, a second set of one or more performance index values, or a third set of one or more performance index values. The first set of one or more performance index values is: nrCE-SDS LC+HC (%); rCE-SDS Pre-LC (%); CEX Basicity (%); SEC HMW (%); SEC Main (%); SEC LMW (%); rCE-SDS HC (%); rCE-SDS HMW (%); rCE-SDS Pre-LC + LC_HC (%); Pool conductivity; or nrCE-SDS Pre-peak (%) Includes, The second set of one or more performance indicator values is: CEX acidity (%); CEX Main (%); Process yield; rCE-SDS Main (%); rCE-SDS LMW (%); cIEF acidic (%); cIEFF basicity (%); or cIEFF Main (%) Includes, The third set of one or more performance index values is: rCE-SDS LC+HC (%); or rCE-SDS LC (%) Processes including; A step of receiving one or more molecular descriptors describing the therapeutic protein using one or more processors; A step of predicting process parameter values of a hypothetical chromatography process as output of a machine learning model by applying at least one of the performance index values and one or more molecular descriptors as input to a machine learning model using one or more processors, wherein the machine learning model is selected from the group consisting of (i) a regression tree model, (ii) an extreme gradient boost model, and (iii) an elastic net model, the regression tree model is selected to predict the process parameters when the one or more performance index values are included in a first set of the one or more performance index values, the extreme gradient boost model is selected to predict the process parameters when the one or more performance index values are included in a second set of the one or more performance index values, and the elastic net model is selected to predict the process parameters when the one or more performance index values are included in a third set of the one or more performance index values; and The process of having one or more processors present to the user via a user interface either (i) the predicted process parameter values and (ii) the predicted accuracy range of the predicted process parameter values. A method that includes this.
3. The aforementioned hypothetical chromatography process is Hypothetical cation exchange chromatography (CEX) process; Hypothetical size exclusion chromatography (SEC) process; and Protein A chromatography process The method according to claim 1 or 2, wherein the process is selected from the group consisting of the following.
4. The method according to any one of claims 1 to 3, further comprising the step of determining at least one of the one or more molecular descriptors based on sequence information associated with the therapeutic protein using the one or more processors.
5. The method according to any one of claims 1 to 4, further comprising the step of determining at least one of the one or more molecular descriptors based on experimental measurements of the physical properties of the therapeutic protein using the one or more processors.
6. The method according to any one of claims 1 to 5, wherein at least one of the one or more molecular descriptors is a function of pH level.
7. The one or more process parameter values mentioned above are: pH of the buffer solution; pH of the elution buffer; Conductivity of the elution buffer solution; Molar concentration of elution buffer; Gradient; linear velocity; Load conductivity; Loading factor; Load pH; or Collection stopped The method according to any one of claims 1 to 6, comprising one or more of the above.
8. The performance index value includes SEC HMW (%), A step of predicting the yield of the hypothetical chromatography process by analyzing at least process parameters and molecular descriptors using one or more processors and an additional machine learning model, wherein the additional machine learning model is another Extreme gradient boost model; and The process of having one or more processors present to the user, via the user interface, (i) the predicted yield and (ii) an indicator of whether the predicted yield meets one or more additional acceptance criteria. The method according to any one of claims 1 to 7, further comprising:
9. A step of selecting one or more process parameter values for the chromatography process for the therapeutic protein based on the presented performance index values and / or the presented index; and A step of performing the chromatography process for the therapeutic protein according to the selected process parameter values. The method according to any one of claims 1 or 3 to 8, further comprising:
10. A step of selecting one or more process parameter values for the chromatography process for the therapeutic protein based on the presented predicted process parameter values and / or the predicted accuracy range; and A step of performing the chromatography process for the therapeutic protein according to the selected process parameter values. The method according to any one of claims 2 to 8, further comprising:
11. One or more non-temporary computer-readable media that, when executed by one or more processors of a computing system, store instructions causing the computing system to perform the method according to any one of claims 1 to 10.
12. A computing system, One or more processors; and One or more non-temporary computer-readable media that, when executed by the one or more processors, store instructions causing the computing system to carry out the method according to any one of claims 1 to 10. A computing system that includes this.