Computational-based methods for improving protein purification
Iteratively trained machine learning models streamline protein purification by predicting molecular binding properties, addressing the inefficiencies of current experimental methods and accelerating the development of therapeutic antibodies.
Patent Information
- Application Number
- JP2025507723
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-15
- Filing Date
- 2023-08-14
- Publication Date
- 2025-09-17
AI Technical Summary
Current protein purification techniques are laborious and time-consuming, leading to costly inefficiencies in developing therapeutic antibodies and other immunotherapies due to reliance on experimental methods that are difficult to optimize.
Utilizing iteratively trained machine learning models, particularly ensemble learning models, to predict molecular binding properties of proteins, reducing the need for extensive experimentation by identifying desirable proteins in silico for streamlined protein purification processes.
Accelerates the development and manufacturing of therapeutic antibodies by reducing experimental duration and inefficiencies, enabling early identification of promising candidates through computational model-based predictions.
Smart Images

Figure 2025530653000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 398,168, filed August 15, 2022, entitled "COMPUTATIONAL-BASED METHODS FOR IMPROVING PROTEIN PURIFICATION," the disclosure of which is incorporated herein by reference in its entirety.
[0002] This application relates generally to protein purification, and more particularly to computationally based methods for improving protein purification. [Background technology]
[0003] Cell cultures utilizing engineered mammalian or bacterial cell lines can be used to produce target proteins of interest, for example, by inserting a recombinant plasmid containing the gene for the target protein. Because the cell line itself is a living organism, the cell line may produce proteins other than the target protein and require a complex growth medium containing, for example, various sugars, amino acids, and growth factors. Obtaining a highly purified composition of the target protein is often desirable, if not necessary, particularly when the target protein is used as a therapeutic active agent, for example, when the target protein is a therapeutic antibody. Therefore, the produced target protein must be purified from these other components in the cell culture, which can involve a complex sequence of processes, each of which involves many variables, such as chromatographic stationary phases, mobile phases, salt concentrations, pH, and other operating conditions, such as temperature.
[0004] For example, a protein purification process sequence can include (a) obtaining a cell culture sample containing the target protein, (b) one or more capture steps, e.g., affinity capture using Protein A, (c) one or more conditioning steps, (d) one or more depth filtration steps, (e) one or more ion exchange chromatography steps, e.g., cation exchange chromatography or anion exchange chromatography, or a mixed mode thereof, optionally including hydrophobic interaction chromatography, (f) one or more hydrophobic interaction chromatography steps, or a mixed mode thereof, (g) a virus filtration step, and (h) one or more ultrafiltration steps. Each of these column chromatography techniques can include various conditions utilized to purify the target protein. Specifically, purification techniques involve many variables important for efficiently producing a highly purified composition of the target protein, including considerations related to the target protein itself, as well as other operating conditions, such as the chromatographic stationary phase, mobile phase, salt concentration, pH, and temperature.
[0005] Currently, purification techniques and their operating conditions are determined experimentally in laboratories. Therefore, determining how to purify a target protein can be very laborious and time-consuming. Furthermore, relying solely on the aforementioned experimental purification techniques to separate and isolate a target protein can result in useful feedback regarding candidate proteins that may be difficult to purify without laborious experimentation. This can therefore lead to costly inefficiencies in the development and manufacturing process of therapeutic antibodies or other similar immunotherapies. Therefore, it would be useful to provide a technology that optimizes and streamlines the protein purification process to identify target proteins (e.g., antibodies) and accelerate the selection process for therapeutic antibody candidates. Summary of the Invention
[0006] Embodiments of the present disclosure relate to one or more computing devices, methods, and non-transitory computer-readable media that can utilize iteratively trained machine learning models to identify target proteins (e.g., antibodies) and generate predictions of molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates. This streamlined process of identifying target proteins (e.g., antibodies) in silico can facilitate and accelerate the downstream development and manufacture of, for example, one or more therapeutic monoclonal antibodies (mAbs), bispecific antibodies (bsAbs), trispecific antibodies (tsAbs), or other similar immunotherapies that can be utilized to treat various patient diseases. In some embodiments, the machine learning model comprises an ensemble machine learning model that includes multiple models.
[0007] For example, once trained, a machine learning model (e.g., a "boosting" ensemble learning model) can be utilized to generate predictions of molecular binding properties of one or more proteins (e.g., predictions of percent protein binding at one or more particular pH values and particular salt concentrations and / or particular salt species and chromatography resins) by utilizing the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training of the machine learning model, as well as a selected k-best matrix of feature vectors of molecular descriptor matrices representing sets of amino acid sequences corresponding to one or more proteins of interest.
[0008] Specifically, according to presently disclosed embodiments, once trained, the machine learning model may utilize the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training to predict percent protein binding (e.g., the percentage of a set of proteins predicted to bind to a ligand in solution for a given pH value and salt concentration) for one or more target proteins based solely on, as input, a set of amino acid sequences corresponding to one or more proteins of interest and a selected k-best matrix of feature vectors of molecular descriptor matrices representing one or more sets of pH values and salt concentrations associated with the binding properties of the one or more proteins of interest.
[0009] In this manner, by providing a computational model-based prediction of percent protein binding for one or more proteins of interest, the molecular binding and elution characteristics of one or more proteins of interest may be determined without significant upstream experimentation. That is, desirable proteins of one or more proteins of interest may be identified and distinguished in silico from undesirable proteins of one or more proteins of interest, and those in silico identified desirable proteins may be further utilized to facilitate and accelerate the downstream development and production of one or more therapeutic mAbs, bsAbs, tsAbs, or other similar immunotherapies that may be utilized to treat various patient diseases (e.g., by reducing upstream experimental duration and experimental inefficiencies and providing feedback in silico that candidate proteins may be difficult to purify and thus ultimately manufacture).
[0010] For example, during training of a machine learning model (e.g., a "boosting" ensemble learning model), hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) may be iteratively refined and optimized. The iterations may include: 1) reducing a molecular descriptor matrix representing a set of amino acid sequences by clustering similar feature vectors of the molecular descriptor matrix based on a distance metric. For example, the distance metric may be calculated based on Pearson's correlation, mutual information, or maximum information coefficient (MIC), or other distance metric. Feature vectors other than those closest to the cluster centroid may be discarded to generate a reduced molecular descriptor matrix. The iterations may then include determining the k-best most predictive feature vectors of the reduced molecular descriptor matrix based on a k-best process and a maximum information coefficient (MIC) to determine correlations between the feature vectors of the reduced molecular descriptor matrix and experimentally determined protein binding percentages and / or first principal component (PC) values for one or more specific pH values and salt concentrations. The iteration may then include calculating n cross-validation losses based on the k-best most predictive feature vectors and the experimentally determined protein binding percentages and / or first PC values. Finally, the iteration may include updating hyperparameters (e.g., general parameters, booster parameters, learning task parameters) based on the n cross-validation losses.
[0011] In certain embodiments, the aforementioned feature dimensionality reduction and feature selection techniques may reduce a molecular descriptor matrix, which may include a large set of amino acid sequence-based descriptors, so that the regression model successfully converges to an accurately trained regression model, as opposed to suffering from overfitting due to unnecessary or noisy descriptors. Furthermore, it should be understood that in some embodiments, instead of utilizing an MIC as part of the determination of the correlation between the feature vectors of the reduced molecular descriptor matrix and the experimentally determined percent protein binding and / or first PC value, distance correlation, mutual information, or other similar non-linear correlation metrics, or linear correlation metrics (e.g., Pearson's correlation) may be utilized.
[0012] In certain embodiments, one or more computing devices, methods, and non-transitory computer-readable media may access a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins. In some embodiments, the molecular descriptor matrix may be generated by a first machine learning model (e.g., a matrix generation machine learning model) that is different from the machine learning model (e.g., an ensemble learning model). The first machine learning model is trained to generate the molecular descriptor matrix based on the set of amino acid sequences. For example, in some embodiments, the first machine learning model may include a neural network trained to generate an M×N descriptor matrix representing the set of amino acid sequences, where N comprises the number of sets of amino acid sequences and M comprises the number of nodes in the output layer of the neural network.
[0013] In certain embodiments, the one or more computing devices may then refine a set of hyperparameters associated with a machine learning model that is trained to generate predictions of molecular binding properties of the one or more proteins. In some embodiments, the machine learning model may include one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a categorical boosting (CatBoost) model.
[0014] In some embodiments, the predictions of molecular binding properties of one or more proteins can be generated by a computational model-based column process. In some embodiments, the computational model-based chromatography process can include one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process.
[0015] The methods provided herein can be used to predict the molecular binding properties of one or more proteins in any chromatographic setting, including any set of chromatographic techniques or combinations of specific operating conditions. Chromatographic techniques generally involve a stationary phase and a mobile phase. The stationary phase can include moieties designed to interact with the target protein (e.g., in bind-and-elute mode chromatography) or not interact with the target protein (e.g., in flow-through mode chromatography). One or more mobile phases used in chromatographic techniques (e.g., loading, wash, elution mobile phases) can have many variables, including one or more salt concentrations, pH, and solvent gradients. Furthermore, chromatographic techniques can be performed under a variety of conditions, such as elevated temperatures.
[0016] In some embodiments, the computational model-based chromatography process includes an affinity chromatography process. In some embodiments, the affinity chromatography process can include an affinity ligand, such as from any of Protein A chromatography, Protein G chromatography, Protein A / G chromatography, Protein L chromatography, and Kappa chromatography. In some embodiments, the affinity chromatography process can include an elution mobile phase, such as a mobile phase having a set pH.
[0017] In some embodiments, the computational model-based chromatography process may include an ion exchange chromatography process. Ion exchange chromatography allows for separation based on electrostatic interactions (anions and cations) between ligands of an ion exchange stationary phase and components of a sample, such as target or non-target proteins. In some aspects, the ion exchange chromatography utilizes a cation exchange (CEX) stationary phase. In some embodiments, the ion exchange chromatography may include a strong CEX stationary phase. In some embodiments, the ion exchange chromatography may include a weak CEX stationary phase. In some embodiments, the ion exchange chromatography resin may be functionalized with a ligand containing an anionic functional group, such as a carboxyl group or a sulfonate group.
[0018] In some embodiments, the ion exchange chromatography stationary phase can include an anion exchange (AEX) stationary phase. In some embodiments, the ion exchange chromatography can include a strong AEX stationary phase. In some embodiments, the ion exchange chromatography can include a weak AEX stationary phase. In some embodiments, the ion exchange chromatography resin can be functionalized with a ligand containing a cationic functional group, such as a quaternary amine. In some embodiments, the ion exchange chromatography can include a multimodal ion exchange (MMIEX) stationary phase. The MMIEX chromatography stationary phase can include both cation exchange and anion exchange components and / or features. In some embodiments, the MMIEX stationary phase can include a multimodal anion / cation exchange (MM-AEX / CEX) stationary phase.
[0019] In some embodiments, the ion exchange chromatography can include a ceramic hydroxyapatite chromatographic stationary phase. In some embodiments, the ion exchange chromatography stationary phase can be selected from the group consisting of sulfopropyl (SP) Sepharose® Fast Flow (SPSFF), quaternary ammonium (Q) Sepharose® Fast Flow (QSFF), SP Sepharose® XL (SPXL), Streamline™ SPXL, ABx™ (MM-AEX / CEX medium), Poros™ XS, Poros™ 50HS, diethylaminoethyl (DEAE), dimethylaminoethyl (DMAE), trimethylaminoethyl (TMAE), quaternary aminoethyl (QAE), mercaptoethylpyridine (MEP)-Hypercel™, HiPrep™ Q XL, Q Sepharose® XL, and HiPrep™ SPXL. In some embodiments, the ion exchange chromatography process may include an elution step mobile phase that includes an increased salt concentration, such as an increased salt concentration compared to a binding or wash mobile phase.
[0020] In some embodiments, the computational model-based chromatography process may include a mixed-mode chromatography process. The mixed-mode chromatography process may include a stationary phase that combines charge-based components (i.e., characteristics of ion-exchange chromatography) with hydrophobicity-based components. In some embodiments, the mixed-mode chromatography process may include a bind-and-elute mode of operation. In some embodiments, the mixed-mode chromatography process may include a flow-through mode of operation. In some embodiments, the mixed-mode chromatography process may include a stationary phase selected from the group consisting of Capto MMC and Capto Adhere.
[0021] In some embodiments, the computational model-based chromatography process may include a hydrophobic interaction chromatography (HIC) process. The hydrophobic interaction chromatography process may include a hydrophobic stationary phase. In some embodiments, the mixed-mode chromatography process may include a bind-and-elute mode of operation. In some embodiments, the hydrophobic interaction chromatography process may include a flow-through mode of operation. In some embodiments, the hydrophobic interaction chromatography process may include a stationary phase comprising an inert matrix, for example, a substrate such as cross-linked agarose, sepharose, or a resin matrix. In some embodiments, at least a portion of the substrate of the hydrophobic interaction chromatography stationary phase may include a surface modification comprising a hydrophobic ligand.
[0022] In some embodiments, the hydrophobic interaction chromatography ligand is a ligand containing about 1 to 18 carbons. In some embodiments, the hydrophobic interaction chromatography ligand can contain one or more carbons, such as two or more carbons, three or more carbons, four or more carbons, five or more carbons, six or more carbons, seven or more carbons, eight or more carbons, nine or more carbons, ten or more carbons, eleven or more carbons, twelve or more carbons, thirteen or more carbons, fourteen or more carbons, fifteen or more carbons, sixteen or more carbons, seventeen or more carbons, or eighteen or more carbons. In some embodiments, the hydrophobic interaction chromatography ligand can contain one carbon, two carbons, three carbons, four carbons, five carbons, six carbons, seven carbons, eight carbons, nine carbons, ten carbons, eleven carbons, twelve carbons, thirteen carbons, fourteen carbons, fifteen carbons, sixteen carbons, seventeen carbons, or eighteen carbons. In some embodiments, the hydrophobic ligand is selected from the group consisting of an ether group, a methyl group, an ethyl group, a propyl group, an isopropyl group, a butyl group, a t-butyl group, a hexyl group, an octyl group, a phenyl group, and a polypropylene glycol group.
[0023] In some embodiments, the HIC medium is a hydrophobic charge induction chromatography medium. In some embodiments, the hydrophobic interaction chromatography process can include a mobile phase containing high salt conditions. For example, high salt conditions can be used to reduce solvation of the target, thereby exposing hydrophobic regions, which can then interact with the hydrophobic interaction chromatography stationary phase. In some embodiments, the hydrophobic interaction chromatography process can include a mobile phase containing low salt conditions, for example, no salt or no added salt. In some embodiments, the hydrophobic interaction chromatography stationary phase is selected from the group consisting of Bakerbond WP HI-Pro™, Phenyl Sepharose® Fast Flow (Phenyl-SFF), Phenyl Sepharose® Fast Flow Hi-sub (Phenyl-SFF HS), Toyopearl® Hexyl-650, Poros™ Benzyl Ultra, and Sartobind® Phenyl. In some embodiments, Toyopearl® Hexyl-650 is Toyopearl® Hexyl-650M. In some embodiments, Toyopearl® Hexyl-650 is Toyopearl® Hexyl-650C. In some embodiments, Toyopearl® Hexyl-650 is Toyopearl® Hexyl-650S.
[0024] In certain embodiments, predicting the molecular binding properties of one or more proteins may include identifying target proteins of one or more proteins. In another embodiment, predicting the molecular binding properties of one or more proteins may use quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of one or more proteins. In another embodiment, predicting the molecular binding properties of one or more proteins may include predicting the molecular binding properties for each amino acid sequence in a set of amino acid sequences corresponding to one or more proteins. In another embodiment, predicting the molecular binding properties for each amino acid sequence may include isolating desirable amino acid molecules from undesirable amino acid molecules based on a computational model. In one embodiment, the machine learning model (e.g., an ensemble learning model) may be further trained to generate predictions of molecular elution properties of one or more proteins. In another embodiment, the machine learning model may be further trained to generate predictions of flow-through properties of one or more proteins.
[0025] In certain embodiments, one or more computing devices may iteratively refine the set of hyperparameters by executing the process until a desired accuracy is reached. For example, in certain embodiments, one or more computing devices may execute the process by first reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters. In one embodiment, each of the feature vector clusters includes similar feature vectors. For example, in some embodiments, reducing the molecular descriptor matrix may include performing clustering using a correlation distance metric calculated, for example, based on Pearson's correlation of the feature vectors of the molecular descriptor matrix to generate a plurality of feature vector clusters. In some cases, the clustering of the descriptor sets may be based on the correlation distance between descriptors, which may be calculated from Pearson's correlation (e.g., 1-abs(Pearson's correlation)). In one embodiment, the selected one representative feature vector in each of the plurality of feature vector clusters may include a centroid feature vector for each of the plurality of feature vector clusters that is utilized to represent two or more of the similar feature vectors.
[0026] In certain embodiments, the one or more computing devices may then perform the process by determining one or more most predictive feature vectors of the selected representative feature vectors for each of the plurality of feature vector clusters based on correlations between the selected representative feature vectors and predetermined batch binding data associated with one or more proteins. For example, in some embodiments, determining one or more representative feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters may include selecting a k-best matrix of feature vectors for the selected representative feature vectors in each of the plurality of feature vector clusters. In one embodiment, the k-best matrix of feature vectors for the selected representative feature vectors is determined based on a predetermined k-best process. In certain embodiments, the correlation between the selected representative feature vectors and the predetermined batch binding data is determined based on Pearson's correlation, mutual information, maximum information coefficient (MIC), or other metric between the selected representative feature vectors and the predetermined batch binding data. In other embodiments, instead of utilizing MIC as part of the correlation determination, distance correlation, mutual information, or other similar nonlinear and / or linear correlation metrics may be utilized.
[0027] In particular embodiments, the one or more computing devices may then perform the process by calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data. For example, calculating the one or more cross-validation losses may further include evaluating a cross-validation loss function based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and a set of learnable parameters associated with the machine learning model, and further minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the predetermined batch combined data, and the set of hyperparameters remain constant.
[0028] For example, in some embodiments, minimizing the cross-validation loss function may include optimizing a set of hyperparameters. For example, in one embodiment, the set of hyperparameters may include one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters. In some embodiments, minimizing the cross-validation loss function may further include minimizing the loss between a prediction of percent protein binding for the one or more proteins and an experimentally determined percent protein binding for the one or more proteins. In one embodiment, the given batch binding data may include experimentally determined percent protein binding for one or more pH values and salt concentrations associated with molecular binding properties of the one or more proteins. In one embodiment, the set of learnable parameters may include one or more weights or decision variables determined by the machine learning model based at least in part on the one or more most predictive feature vectors and the given batch binding data.
[0029] In some embodiments, calculating one or more cross-validation losses may include calculating n cross-validation losses, where n includes an integer between 1 and n. In some embodiments, calculating one or more cross-validation losses may include determining n individual train-test splits based on the one or more most predictive feature vectors and the given batch binding data, where n includes an integer between 1 and n. In some embodiments, calculating one or more cross-validation losses may include calculating n cross-validation losses and generating a prediction of a molecular binding property of the one or more proteins based on an average of the n cross-validation losses.
[0030] In certain embodiments, the one or more computing devices may then perform the process by updating a set of hyperparameters based on one or more cross-validation losses. For example, the updated set of hyperparameters may include one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters. In some embodiments, after refining the set of hyperparameters, the one or more computing devices may output, by the machine learning model, a prediction of molecular binding properties of the one or more proteins based at least in part on the updated set of hyperparameters.
[0031] In certain embodiments, following refining the set of hyperparameters, the one or more computing devices may further access a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins, reduce the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix, and determine one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data related to the one or more second proteins. The one or more second most predictive feature vectors may be input to a machine learning model trained to generate predictions of molecular binding properties of the one or more second proteins, and the machine learning model may output predictions of molecular binding properties of the one or more second proteins based at least in part on the updated set of hyperparameters. For example, the predictions of molecular binding properties of the one or more second proteins may include predictions of protein binding percentages for the one or more second proteins.
[0032] In certain embodiments, the one or more computing devices may further optimize the machine learning model based on a Bayesian model optimization process. In some embodiments, the one or more computing devices may then utilize group K-fold cross-validation to train and evaluate the optimized machine learning model based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and the set of learnable parameters. In some embodiments, the group K-fold cross-validation may be stratified so that the cross-validation training and evaluation splits include a diverse range of regression target values. In some embodiments, stratification may be achieved using labels generated by binning the regression target values into several quantiles.
[0033] In certain embodiments, one or more computing devices, methods, and non-transitory computer-readable media may access a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins, and obtain predictions of molecular binding properties of the one or more proteins based at least in part on the molecular descriptor matrix using a machine learning model, the machine learning model being trained by accessing a training molecular descriptor matrix representing a training set of amino acid sequences corresponding to one or more empirically evaluated proteins, and iteratively performing a process of refining a set of hyperparameters associated with the machine learning model until a desired accuracy is reached, the process including: reducing the training molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster including similar feature vectors; determining one or more most predictive feature vectors among the selected representative feature vectors for each feature vector cluster based on correlations between the selected representative feature vectors and predetermined batch binding data associated with the one or more empirically evaluated proteins; calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch binding data; and updating the set of hyperparameters based on the one or more cross-validation losses.
[0034] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Brief explanation of the drawings]
[0035] [Figure 1] FIG. 1 shows a diagram illustrating an experimental example for performing one or more protein purification processes compared to a computational model-based example for performing one or more protein purification processes, according to various embodiments.
[0036] [Figure 2] FIG. 1 illustrates a high-level workflow diagram for performing feature generation, feature dimensionality reduction, regression model optimization, and model output-based feature selection, according to various embodiments.
[0037] [Figure 3A] FIG. 1 shows a workflow diagram for optimizing hyperparameters and learnable parameters of machine learning models for implementing one or more computational model-based protein purification processes, according to various embodiments.
[0038] [Figure 3B] FIG. 1 shows a workflow diagram for optimizing a machine learning model for implementing one or more computational model-based protein purification processes, according to various embodiments.
[0039] [Figure 4] FIG. 1 shows a flow diagram of a method for generating predictions of the molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to identify target proteins, according to various embodiments.
[0040] [Figure 5] 1 illustrates an exemplary computing system, according to various embodiments.
[0041] [Figure 6] 6 illustrates a diagram of an exemplary artificial intelligence (AI) architecture included as part of the exemplary computing system of FIG. 5, according to various embodiments.
[0042] [Figure 7] FIG. 1 illustrates another high-level workflow diagram for performing feature generation, feature dimensionality reduction, regression model optimization, and model output-based feature selection, according to various embodiments.
[0043] [Figure 8]FIG. 1 shows another workflow diagram for optimizing hyperparameters and learnable parameters of machine learning models for implementing one or more computational model-based protein purification processes, according to various embodiments.
[0044] [Figure 9] 1 illustrates a process for training a machine learning model to predict molecular binding properties, according to various embodiments.
[0045] [Figure 10A] 1 shows exemplary plots illustrating how principal component analysis can be used to predict molecular binding properties, according to various embodiments. [Figure 10B] 1 shows exemplary plots illustrating how principal component analysis can be used to predict molecular binding properties, according to various embodiments. [Figure 10C] 1 shows exemplary plots illustrating how principal component analysis can be used to predict molecular binding properties, according to various embodiments. [Figure 10D] 1 shows exemplary plots illustrating how principal component analysis can be used to predict molecular binding properties, according to various embodiments.
[0046] [Figure 11A] 10A-10C show exemplary heat maps illustrating the relationship between experimental conditions and experimental Kp values and modeled Kp values, respectively, according to various embodiments. [Figure 11B] 10A-10C show exemplary heat maps illustrating the relationship between experimental conditions and experimental Kp values and modeled Kp values, respectively, according to various embodiments. [Figure 11C] 10A-10C show exemplary heat maps illustrating the relationship between experimental conditions and experimental Kp values and modeled Kp values, respectively, according to various embodiments. [Figure 11D] 10A-10C show exemplary heat maps illustrating the relationship between experimental conditions and experimental Kp values and modeled Kp values, respectively, according to various embodiments. [Figure 11E]10A-10C show exemplary heat maps illustrating the relationship between experimental conditions and experimental Kp values and modeled Kp values, respectively, according to various embodiments. [Figure 11F] 10A-10C show exemplary heat maps illustrating the relationship between experimental conditions and experimental Kp values and modeled Kp values, respectively, according to various embodiments.
[0047] [Figure 12] FIG. 1 shows a flow diagram of a method for generating predictions of molecular binding properties of one or more target proteins as part of an alternative streamlined process of protein purification to identify target proteins, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0048]
[0003] Embodiments of the present embodiments relate to one or more computing devices, methods, and non-transitory computer-readable media that can utilize iteratively trained machine learning models to identify target proteins (e.g., antibodies) and generate predictions of the molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates. This streamlined process of identifying target proteins (e.g., antibodies) in silico can facilitate and accelerate the downstream development and manufacturing of, for example, one or more therapeutic monoclonal antibodies (mAbs), bispecific antibodies (bsAbs), trispecific antibodies (tsAbs), or other similar immunotherapies that can be utilized to treat various patient diseases.
[0049] For example, once trained, a machine learning model (e.g., an ensemble learning model or a "boosting" ensemble learning model) can be utilized to generate predictions of molecular binding properties of one or more proteins (e.g., predictions of percent protein binding at one or more particular pH values and particular salt concentrations and / or particular salt species and chromatography resins) by utilizing the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training of the machine learning model, as well as a selected k-best matrix of feature vectors of molecular descriptor matrices representing sets of amino acid sequences corresponding to one or more proteins of interest.
[0050] Specifically, according to disclosed embodiments, once trained, the machine learning model may utilize optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training to (i) predict percent protein binding (e.g., percentage of a set of proteins predicted to bind to a ligand in solution) for a given pH value and salt concentration or multiple different combinations of pH values and salt concentrations, (ii) predict percent protein binding (e.g., percentage of a set of proteins predicted to bind to a ligand in solution) for a set of pH values and salt concentrations, and / or (iii) predict principal components (PCs) representing a set of pH values and salt concentrations for one or more target proteins based solely on a selected k-best matrix of feature vectors of a set of amino acid sequences corresponding to one or more proteins of interest and a molecular descriptor matrix representing a set of pH values and salt concentrations associated with the binding properties of the one or more proteins of interest as input.
[0051] In this manner, by providing computational model-based predictions of percent protein binding and / or PC values for one or more proteins of interest, the molecular binding and elution characteristics of one or more proteins of interest may be determined without significant upstream experimentation. That is, desirable proteins of one or more proteins of interest may be identified and distinguished in silico from undesirable proteins of one or more proteins of interest, and those desirable proteins identified in silico may be further utilized to facilitate and accelerate the downstream development of one or more therapeutic mAbs, bsAbs, tsAbs, or other similar immunotherapies that may be utilized to treat various patient diseases (e.g., by reducing upstream experimental duration and experimental inefficiencies and providing feedback in silico that candidate proteins may be difficult to purify and thus ultimately manufacture).
[0052] For example, during training of a machine learning model (e.g., a "boosting" ensemble learning model), hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) may be iteratively refined and optimized by: 1) reducing a molecular descriptor matrix representing a set of amino acid sequences by clustering similar feature vectors of the molecular descriptor matrix based on a correlation distance metric, which may be calculated based on Pearson's correlation, and discarding feature vectors other than those closest to the cluster centroid; 2) determining the k-best most predictive feature vectors of the reduced molecular descriptor matrix based on a k-best process and a correlation coefficient (e.g., maximum information coefficient (MIC)) to determine the correlation between the feature vectors of the reduced molecular descriptor matrix and experimentally determined percent protein binding for one or more particular pH values and salt concentrations; 3) calculating n cross-validation losses based on the k-best most predictive feature vectors and the experimentally determined percent protein binding; and 4) updating the hyperparameters (e.g., general parameters, booster parameters, learning task parameters) based on the n cross-validation losses. In one or more examples, the clustering of one or more sets of descriptors may be based on correlation distances between the descriptors calculated from Pearson's correlation (eg, 1-abs(Pearson's correlation)).
[0053] In certain embodiments, the aforementioned feature dimensionality reduction and feature selection techniques may reduce a molecular descriptor matrix, which may comprise a large set of amino acid sequence-based descriptors, so that the regression model successfully converges to an accurately trained regression model, as opposed to suffering from overfitting due to unnecessary or noisy descriptors. Furthermore, it should be understood that in some embodiments, instead of utilizing MIC as part of the determination of correlation between the feature vectors of the reduced molecular descriptor matrix and the experimentally determined percent protein binding, distance correlation, mutual information, or other similar non-linear correlation metrics, or linear correlation metrics (e.g., Pearson's correlation, f-statistic-based metrics) may be utilized.
[0054] As used herein, the terms "polypeptide" and "protein" may refer interchangeably to polymers of amino acid residues and are not limited to a minimum length. For example, such polymers of amino acid residues may contain natural or unnatural amino acid residues and may include, but are not limited to, peptides, oligopeptides, dimers, trimers, and multimers of amino acid residues. Both full-length proteins and fragments thereof, for example, are encompassed by the definition. The terms "polypeptide" and "protein" may also include post-translational modifications of the polypeptide, such as glycosylation, sialylation, acetylation, phosphorylation, and the like.
[0055] 1 shows a diagram 100 illustrating an example experiment 102 for performing one or more protein purification processes compared to a computational model-based example 104 for performing one or more protein purification processes, according to disclosed embodiments. As shown, on the one hand, the experiment duration in the example experiment 102 for performing one or more protein purification processes may extend over several weeks, while the execution time in the computational model-based example 104 for performing one or more protein purification processes may be only a few minutes.
[0056] For example, an example experiment 102 for performing one or more protein purification processes may include receiving an amino acid sequence at block 106, selecting a plasmid at block 108, engineering the protein by cell line and cell culture at blocks 110 and 112, respectively, performing one or more chromatography processes (e.g., an affinity chromatography process, an ion exchange chromatography process (IEX), a hydrophobic interaction chromatography process (HIC), or a mixed-mode chromatography process (MMC) process) at block 114, and performing high-throughput screening (HTS) at block 116 to determine the partition coefficient (K), all as part of a laborious and time-consuming protein purification process. p ) to quantify protein binding. Molecular evaluation of one or more target proteins may then be performed in block 118.
[0057] In certain embodiments, a computational model-based example 104 for performing one or more protein purification processes in accordance with the presently disclosed technology may include, in block 106, accessing amino acid sequences corresponding to one or more proteins of interest; in block 120, generating a molecular descriptor matrix based on the amino acid sequences and reducing the molecular descriptor matrix; and, in block 122, utilizing a machine learning model (e.g., an ensemble learning model) to identify target proteins (e.g., antibodies) and generate predictions of molecular binding properties of the one or more target proteins as part of an optimized and streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates in accordance with the presently disclosed embodiments.
[0058] Indeed, as described in further detail below with respect to Figures 2-4, the machine learning model may utilize the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training to predict protein binding percentages (e.g., the percentage of a set of proteins predicted to bind to a ligand in solution for a given pH value and salt concentration) for one or more target proteins based solely on, as input, the selected k-best matrix of feature vectors of the molecular descriptor matrix generated in block 120 and one or more sets of pH values and salt concentrations associated with the binding properties of one or more proteins of interest. Molecular evaluation of one or more target proteins may then be performed in block 118 without significant upstream experimentation (e.g., compared to example experiment 102 for performing one or more protein purification processes). That is, desirable proteins of one or more proteins of interest may be identified and distinguished in silico from undesirable proteins of one or more proteins of interest, and those desirable proteins identified in silico may be further utilized to facilitate and accelerate the downstream development of one or more therapeutic mAbs, bsAbs, tsAbs, or other similar immunotherapies that may be utilized to treat various patient diseases (e.g., by reducing upstream experimental duration and experimental inefficiencies and providing in silico feedback that candidate proteins may be difficult to purify and therefore ultimately manufacture). As an example, based at least in part on the molecular descriptor matrix, a machine learning model may be configured to obtain predictions of molecular binding properties of one or more proteins. From the molecular binding properties, desirable proteins may be identified.
[0059] Figure 2 shows a high-level workflow diagram 200 for performing feature generation 202, feature dimensionality reduction 204, model output-based feature selection 206, and regression model optimization 208, according to disclosed embodiments. Specifically, as described with respect to Figure 2, high-level examples for performing feature generation 202, feature dimensionality reduction 204, model output-based feature selection 206, and regression model optimization 208 may be described in more detail below with respect to Figures 3A and 3B, and it should be understood that they may be performed by machine learning (e.g., a matrix generation machine learning model) in conjunction with another machine learning model (e.g., an ensemble learning model) according to presently disclosed embodiments.
[0060] That is, as described below with respect to Figures 3A and 3B, feature generation 202 may be performed by machine learning model 301, feature dimensionality reduction 204 may be performed by feature dimensionality reduction models 307A and 307B of machine learning models 302A and 302B, model output-based feature selection 206 may be performed by feature selection models 309A and 309B of machine learning models 302A and 302B, and regression model optimization 208 may be performed by regression models 311A and 311B of machine learning models 302A and 302B.
[0061] For example, as described in more detail below, performing feature generation 202 may include generating, for example, 1024 molecular descriptors (e.g., descriptors based on amino acid sequences). In particular embodiments, performing feature dimensionality reduction 204 may include clustering and reducing the 1024 molecular descriptors (e.g., descriptors based on amino acid sequences), for example, to remove redundant features or other features determined to be highly similar. In particular embodiments, performing feature selection based on model output 206 may include generating a k-best feature matrix to reduce the molecular descriptors to only their k-best, most predictive features. As one example, the number of molecular descriptors may be 1024 based on the particular model used to generate the descriptors. As another example, the number of molecular descriptors may be more or less, such as 2048 descriptors, 320 descriptors, etc. In particular embodiments, performing regression model optimization 208 may include, for example, optimizing hyperparameters and learnable parameters associated with regression models 311A, 311B of machine learning models 302A, 302B. In particular embodiments, as described in more detail below, feature dimensionality reduction 204 and model output-based feature selection 206 may be provided to filter a large set of amino acid sequence-based descriptors that may, in some embodiments, be generated as part of feature generation 202. In this manner, reducing a large set of amino acid sequence-based descriptors by feature dimensionality reduction 204 and model output-based feature selection 206 may ensure that the regression model successfully converges to an accurately trained regression model, as opposed to suffering from overfitting due to unnecessary or noisy descriptors.
[0062] FIG. 3A shows a detailed workflow diagram 300A for optimizing hyperparameters and learnable parameters of a machine learning model 302A (e.g., an ensemble learning model) and utilizing the machine learning model 302A to identify target proteins (e.g., antibodies) and generate predictions of molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibodies or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates, according to disclosed embodiments. In certain embodiments, as shown in FIG. 3A , the workflow diagram 300A may be implemented using hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a system on a chip (SoC), a microcontroller, a field programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a vision processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), or a processor for processing genomics data, proteomics data, metabolomics data, metabolomics data, or metabolomics data). The machine learning model 301 (e.g., the matrix generation machine learning model) and the machine learning model 302A (e.g., shown by dashed lines) may be executed in conjunction with one or more processing devices (e.g., computing device 500 and artificial intelligence architecture 600, described below in connection with Figures 5 and 6), which may include software (e.g., instructions operating / executing on one or more processors), firmware (e.g., microcode), or any other processing device that may be suitable for processing tagenomics data, transcriptomic data, or other omics data and making one or more decisions based thereon.
[0063] In particular embodiments, machine learning models 302A may include any number of individual machine learning or other predictive models (e.g., feature dimensionality reduction model 307A, feature selection model 309A, and regression model 311A) that may be trained and run in tandem (e.g., serially, in parallel, or end-to-end) to perform one or more predictions in sequence, such that the output of one or more earlier models in a pipeline serves as input to one or more subsequent models in an ensemble until a final overall prediction is output (e.g., "boosting"). For example, in some embodiments, machine learning models 302A may include one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a categorical boosting (CatBoost) model.
[0064] 3A and described in more detail below, machine learning model 301 may perform one or more feature generation and data import tasks 303, while machine learning model 302A may include a feature dimensionality reduction model 307A, a feature selection model 309A, and a regression model 311A. One or more hyperparameter optimization tasks 314 may also be performed to refine a set of hyperparameters associated with machine learning model 302A.
[0065] Workflow diagram 300A may begin at functional block 304, in which machine learning model 301 imports amino acid sequences in a set of one or more P proteins. For example, in certain embodiments, machine learning model 301 may include one or more pre-trained artificial neural networks (ANNs), convolutional neural networks (CNNs), or other neural networks, which may be suitable for generating a large set of amino acid sequence-based descriptors, e.g., in a supervised, weakly supervised, semi-supervised, or unsupervised manner. According to embodiments disclosed herein, amino acid sequence-based descriptors may be utilized (e.g., as opposed to structure-based descriptors) because they may be more effective for training machine learning model 302A to generate predictions of molecular binding properties of one or more target proteins. Indeed, as described in more detail below, feature dimensionality reduction models 307A, 307B and feature selection models 309A, 309B may be provided in some embodiments to filter a large set of amino acid sequence-based descriptors that may be output by machine learning model 301. In this way, reducing a large set of amino acid sequence-based descriptors by the feature dimensionality reduction models 307A, 307B and feature selection models 309A, 309B may allow the regression models 311A, 311B to successfully converge to accurately trained regression models, as opposed to suffering from overfitting due to unnecessary or noisy descriptors.
[0066] In function block 305, predetermined batch binding data for a set of one or more P proteins may also be imported for use by machine learning model 302A. For example, in certain embodiments, the predetermined batch binding data may include one or more particular pH values and salt concentrations (e.g., sodium chloride (NaCl), phosphate (PO4) concentrations, etc.). 3-) concentration) and / or salt species (e.g., sodium acetate (CHCOONa) species, sodium phosphate (NaPO) species) and experimentally determined percent protein binding for the chromatography resin. Workflow diagram 300A may then proceed to function block 306, where machine learning model 301 generates a molecular descriptor matrix of size M×N. For example, in certain embodiments, from the amino acid sequence of each protein of the one or more P proteins, machine learning model 301 (e.g., a neural network, a convolutional neural network (CNN), a deep neural network (DNN)) may generate a molecular descriptor matrix of size M×N, where M is the number of descriptors (M=1024) and N is the number of amino acids in a given protein of the set of one or more P proteins.
[0067] In certain embodiments, workflow diagram 300A may then continue to function block 308 with generating a weighted average of the descriptors (M) in the molecular descriptor matrix across all amino acids (N). For example, in certain embodiments, a weighted average of the descriptors (M) in the molecular descriptor matrix across all amino acids (N) may be calculated, resulting in a descriptor vector of size M×1 for each protein in the set of one or more P proteins. For example, in some embodiments, machine learning model 301 may generate one or more M×1 vectors of descriptors for each protein in the set of one or more P proteins. In certain embodiments, workflow diagram 300A may then continue in function block 310 with representing the descriptor vectors for all proteins (P) as a protein descriptor matrix of size M×P.
[0068] In particular embodiments, function block 312 of workflow diagram 300A may represent an iteration of an already trained machine learning model 302A, where a set of hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and a set of learnable parameters (e.g., regression model weights, decision variables) were identified during training of machine learning model 302A. Specifically, as part of one or more hyperparameter optimization tasks 314 performed on machine learning model 302A, in function block 316, the hyperparameters that minimize the average score of the n-cycle regression-based model of machine learning model 302A are determined.
[0069] For example, in one embodiment, a baseline set of hyperparameters (e.g., general parameters, booster parameters, learning task parameters) may be selected and then iteratively updated to minimize the average score of a 10-cycle regression-based model of the machine learning model 302A. In particular embodiments, in function block 318, the machine learning model 302A may be iteratively trained until a desired accuracy is reached, thereby refining the set of hyperparameters by updating the selected hyperparameters (e.g., general parameters, booster parameters, learning task parameters) at each successive iteration. The selected hyperparameters may be updated based on one or more cross-validation losses. For example, in one embodiment, the desired accuracy is reached when a given set of selected hyperparameters minimizes one or more cross-validation losses (e.g., reaches the lowest possible value or error on a scale of 0.0 to 1.0). For example, as described further below, in some embodiments, minimizing one or more cross-validation losses may include minimizing the loss between the predicted percent protein binding and the experimentally determined percent protein binding. Thus, in some embodiments, the desired accuracy of the machine learning model 302A is reached when a given set of selected hyperparameters minimizes the loss between the predicted and experimentally determined percent protein binding.
[0070] In accordance with the presently disclosed technology, in certain embodiments, in function block 320, the hyperparameters may be optimized by evaluating a cross-validation loss function based on the most predictive k-best feature vectors of a given batch binding data, the given batch binding data (e.g., experimentally determined percent protein binding for one or more particular pH values and salt concentrations and / or salt species and chromatography resins), a baseline set of hyperparameters (e.g., general parameters, booster parameters, learning task parameters), and a set of learnable parameters associated with and determined by the machine learning model 302A (e.g., regression model weights, decision variables). The machine learning model 302A may then minimize the cross-validation loss function by varying the set of learnable parameters while the k-best most predictive feature vectors, the given batch binding data, and the set of hyperparameters remain constant.
[0071] In particular embodiments, as previously described, machine learning model 302A may include feature dimensionality reduction model 307A, feature selection model 309A, and regression model 311A. The feature dimensionality reduction task may reduce the molecular descriptor matrix by selecting one representative feature vector for each of multiple feature vector clusters. Starting with feature dimensionality reduction model 307A, in particular embodiments, workflow diagram 300A may proceed to function block 322, where machine learning model 302A evaluates the similarity of different descriptors by comparing sets of M feature vectors of size 1×P.
[0072] For example, in certain embodiments, the similarity of different descriptors may be assessed by comparing a set of M feature vectors of size 1×P with respect to the protein descriptor matrix (M×P). In certain embodiments, the workflow diagram 300A may then proceed to function block 324, where the machine learning model 302A calculates the correlations (of size 1×P) between the feature vectors. For example, in certain embodiments, the machine learning model 302A may calculate a correlation distance metric between each of the feature vectors (of size 1×P), which may be calculated using, for example, Pearson's correlation. In one or more examples, the clustering of the descriptors may be based on the correlation distance between the descriptors calculated from Pearson's correlation (e.g., 1-abs(Pearson's correlation)).
[0073] In particular embodiments, the workflow diagram 300A may then proceed to function block 326, where the machine learning model 302A clusters the feature vectors to group together redundant features that capture similar information. For example, in particular embodiments, utilizing an agglomerative clustering process and a calculated distance correlation metric, which may be calculated based on Pearson's correlation, the machine learning model 302A may cluster the feature vectors to group together any and all redundant features that contain similar information (similar feature vectors). In particular embodiments, the workflow diagram 300A may then proceed to function block 328, where the machine learning model 302A determines the centroid of each cluster as a representative of clusters that are useful for feature selection. Furthermore, selecting the centroid of each cluster may allow for the selection of an orthogonal set of features, which may reduce multicollinearity. In particular embodiments, the workflow diagram 300A may then proceed to function block 330, where the machine learning model 302A iteratively evaluates the number of clusters (C) to determine which will result in optimal performance of the machine learning model 302A.
[0074] In certain embodiments, the machine learning model 302A may also include a feature selection model 309A that may determine one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster. Workflow diagram 300A continues with function block 332 with the machine learning model 302A, which starts with a reduced descriptor matrix (size C×P) and calculates correlations between the feature vectors (1×P) in the reduced descriptor matrix (C×P) and given batch binding data in function block 334. For example, in certain embodiments, using a nonlinear correlation metric (e.g., maximum information coefficient (MIC), distance correlation, mutual information, or other similar nonlinear correlation metric) and / or a linear correlation metric (e.g., Pearson's correlation), the machine learning model 302A may calculate correlations between the selected representative feature vectors (1×P) in the reduced descriptor matrix (C×) and given batch binding data (associated with one or more proteins) to rank which features and / or descriptors capture information suitable for predicting an output.
[0075] In certain embodiments, the workflow diagram 300A may then proceed to function block 336, where the machine learning model 302A determines the top K feature vectors (1×P) that are most predictive of a given batch of combined data to generate a k-best feature matrix (K×P). For example, in certain embodiments, utilizing a k-best process, the machine learning model 302A may select the top K feature vectors (1×P) that are most predictive of a given batch of combined data (e.g., as scored by MIC, distance correlation, mutual information, or other similar non-linear correlation metric) to generate the k-best feature matrix (K×P). Specifically, in one embodiment, the k-best feature matrix (K×P) may retain the top K feature vectors (1×P), where K is an integer value indicating the number of feature vectors retained. In another embodiment, the k-best feature matrix (K×P) may retain the top K feature vectors, where K is a percentage value indicating the percentage of feature vectors retained. In particular embodiments, the workflow diagram 300A may then proceed to function block 338, where the machine learning model 302A iteratively evaluates the K feature vectors to determine which results in optimal performance of the machine learning model 302A.
[0076] In particular embodiments, machine learning model 302A may include regression model 311. For example, workflow diagram 300A may proceed to function block 340 with machine learning model 302A starting from baseline hyperparameters selected and updated as part of hyperparameter optimization task 314, and machine learning model 302A may perform cross-validation utilizing n unique train-test splits (e.g., group K-fold cross-validation, stratified K-fold cross-validation). Cross-validation may include calculating one or more cross-validation losses based at least in part on one or more most predictive feature vectors and a given batch of combined data.
[0077] For example, starting with baseline hyperparameters, the k-best feature matrices, and a given batch of combined data selected and updated as part of the hyperparameter optimization task 314, the machine learning model 302A may perform cross-validation using 10 unique train-test splits of the k-best feature matrices and the given batch of combined data (e.g., a training dataset). In one embodiment, the machine learning model 302A may perform cross-validation using two or more, five or more, ten or more, or other amounts of unique train-test splits of the k-best feature matrices and the given batch of combined data (e.g., a training dataset), e.g., to reduce the possibility of overfitting or miscalculation of the accuracy of the machine learning model 302A due to the train-test splits. In other embodiments, the machine learning model 302A may perform cross-validation using any n integer unique train-test splits, as long as the integer n is less than or equal to the number of data points corresponding to the training dataset, e.g.
[0078] In certain embodiments, workflow diagram 300A may then proceed to function block 342, in which machine learning model 302A adjusts the weights given to data for a given batch of binding data (e.g., percent protein binding at various pH values and salt concentrations and / or salt species and chromatography resins) to weight data within transition regions with higher importance. For example, machine learning model 302A may adjust the weights given to each point in a given batch of binding data to weight data within transition regions (e.g., partially bound proteins) to be more important than fully bound or fully unbound proteins. In certain embodiments, workflow diagram 344 involving machine learning model 302A predicts percent protein binding for a set of proteins P and optimizes machine learning model 302A by minimizing the loss between the predicted percent protein binding and the experimentally determined percent protein binding. In certain embodiments, workflow diagram 300A may then proceed to function block 346, in which machine learning model 302A repeats model optimization n times on a unique train-test split and reports the average score.
[0079] Specifically, the regression task of machine learning model 302A may include receiving given batch binding data and a k-best feature matrix, and predicting (at function block 346) percent protein binding for a set of proteins P based on the given batch binding data and the k-best feature matrix. Machine learning model 302A may then predict percent protein binding for one or more particular pH values and salt concentrations (e.g., sodium chloride (NaCl), phosphate (PO4 3-The pH value and salt concentration and / or salt species (e.g., sodium acetate (CHCOONa) species, sodium phosphate (NaPO) species) and chromatography resin may be optimized (in function block 346) by minimizing the loss (e.g., sum of squared errors (SSE)) between the predicted and experimentally determined percent bound protein for the pH value and salt concentration and / or salt species and chromatography resin. In some embodiments, the pH value and salt concentration and / or salt species and chromatography resin may be related to the molecular binding properties of one or more proteins.
[0080] Thus, as shown by workflow diagram 300A in Figure 3A, a machine learning model 302A can be iteratively trained to identify target proteins (e.g., antibodies) and generate predictions of molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates. For example, a streamlined process of identifying target proteins (e.g., antibodies) in silico can facilitate and accelerate the downstream development and manufacturing of one or more therapeutic mAbs, bsAbs, tsAbs, 2+1 Abs, or other similar immunotherapies that can be utilized to treat various diseases.
[0081] For example, once trained, the machine learning model 302A (e.g., a "boosting" ensemble learning model) may be utilized to generate predictions of molecular binding properties of one or more proteins (e.g., predictions of percent protein binding at one or more particular pH values and particular salt concentrations and / or particular salt species and chromatography resins) by utilizing the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training of the machine learning model 302A, as well as a selected k-best matrix of feature vectors of molecular descriptor matrices representing sets of amino acid sequences corresponding to one or more proteins of interest.
[0082] Specifically, according to disclosed embodiments, once trained, the machine learning model 302A utilizes the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training to calculate protein binding percentages (e.g., percentage of a set of proteins predicted to bind to a ligand in solution for a given pH value and salt concentration) and / or Log(K) for one or more target proteins based solely on a selected k-best matrix of feature vectors of a molecular descriptor matrix representing, as input, a set of amino acid sequences corresponding to one or more proteins of interest and one or more sets of pH values and salt concentrations and / or salt species and / or chromatography resins associated with the binding properties of the one or more proteins of interest. p ) values (a logit transformation of percent binding). In some embodiments, instead of predicting percent binding for a given pH and salt concentration, the first principal component (PC1) of Log(K p The first principal component (PC1) of ) values (logit transformation of percent binding) can be predicted from data across the design space (several data sets covering a range of pH / salt concentrations) for a given resin.
[0083] In this manner, by providing computational model-based predictions of percent protein binding for one or more proteins of interest, the molecular binding and elution characteristics of one or more proteins of interest can be determined without significant upstream experimentation. That is, desirable proteins of one or more proteins of interest can be identified and distinguished in silico from undesirable proteins of one or more proteins of interest, and these in silico-identified desirable proteins can be further utilized to facilitate and accelerate the downstream development of one or more therapeutic mAbs, bsAbs, tsAbs, or other similar immunotherapies that can be utilized to treat various patient diseases (e.g., by reducing upstream experimental duration and experimental inefficiencies and providing in silico feedback that can be difficult to purify and ultimately manufacture candidate proteins). As an example, based at least in part on the molecular descriptor matrix, a machine learning model can be configured to obtain predictions of molecular binding characteristics of one or more proteins. From the molecular binding characteristics, desirable proteins can be identified. Although the present embodiments are described herein primarily with respect to the machine learning model 302A generating predictions of molecular binding properties of one or more target proteins, it should be understood that the trained machine learning model 302A may also generate predictions of elution properties of one or more proteins, or may generate predictions of flow-through properties of one or more proteins, in accordance with the presently disclosed embodiments.
[0084] 3B shows a detailed workflow diagram 300B for optimizing a machine learning model 302A as described above with respect to FIG. 3A and utilizing the optimized machine learning model 302B to generate predictions of molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to identify target proteins (e.g., antibodies) and accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates, according to disclosed embodiments. Specifically, as will be further understood below, workflow diagram 300B may represent an improvement over workflow diagram 300A, as described above with respect to FIG. 3A. For example, as described below, workflow diagram 300B may include performing one or more Bayesian optimization processes (e.g., sequential model-based optimization (SMBO), expected improvement (EI)) to iteratively optimize and evaluate machine learning model 302B, e.g., by selectively determining which functional blocks of feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B to execute, and the order in which the determined functional blocks of feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B are executed.
[0085] In certain embodiments, as shown in FIG. 3B , workflow diagram 300B may be performed utilizing one or more processing devices (e.g., computing device 500 and artificial intelligence architecture 600, described below in connection with FIGS. 5 and 6 ), which may include hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), or any other processing device that may be suitable for processing genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data and making one or more decisions based thereon), software (e.g., instructions running / executing on one or more processors), firmware (e.g., microcode), or some combination thereof.
[0086] In certain embodiments, workflow diagram 300B may begin with function block 348 importing amino acid sequences in a set of one or more P proteins. For example, in some embodiments, function block 348 imports experimental amino acid sequences for the set of one or more P proteins and / or one or more partition coefficients (K p ) screening may be imported. Workflow diagram 300B may then continue in function block 350 with formatting the amino acid sequences in the set of one or more P proteins to generate a molecular descriptor matrix of size M×N.
[0087] In certain embodiments, M may be the number of descriptors (M=1024), and N may be the number of amino acids in a given protein of the set of one or more P proteins for both the light chain (LC) amino acid sequence and the heavy chain (HC) amino acid sequence. In function block 350, workflow diagram 300B may also include generating a weighted average of the descriptors (M) in the molecular descriptor matrix across all amino acids (N). For example, in certain embodiments, a weighted average of the descriptors (M) in the molecular descriptor matrix across all amino acids (N) may be calculated, resulting in a descriptor vector of size M×1 for each protein of the set of one or more P proteins. For example, in some embodiments, machine learning model 301 (as described above with respect to FIG. 3A ) may generate one or more M×1 vectors of descriptors for each protein of the set of one or more P proteins.
[0088] In certain embodiments, the workflow diagram 300B then, in function block 352, removes amino acid sequence data with high salt precipitation and weights the experimental data to identify binding transition regions (e.g., −2 <Log[K p ]<+2 or -0.5 <Log[K p ]<+2. ) may continue to preprocess the descriptor vectors by prioritizing the descriptor vectors. In certain embodiments, as previously described, workflow diagram 300B may be provided to optimize machine learning model 302A as described above with respect to FIG. 3A , and the optimized machine learning model 302B may then be utilized to generate predictions of molecular binding properties of one or more target proteins according to the presently disclosed embodiments. For example, workflow diagram 300B may continue, in function block 354, selectively determine which of feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B function blocks to execute, as well as the order in which to execute the feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B function blocks.
[0089] For example, in particular embodiments, as part of a process for optimizing machine learning model 302B (e.g., an ensemble learning model), workflow diagram 300B of function block 354 may perform one or more Bayesian optimization processes (e.g., sequential model-based optimization (SMBO), expected improvement (EI)) to optimize and evaluate machine learning model 302B. For example, in one embodiment, a Bayesian optimization process (e.g., SMBO, EI) may include one or more probability-based objective functions that may be constructed and utilized to, for example, select the most predictive or most promising of the function blocks of feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B and to select the order in which to execute these function blocks. These function blocks of feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B are described below.
[0090] In certain embodiments, based on the feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B functional blocks selected for execution, workflow diagram 300B at function block 354 may further estimate the accuracy of machine learning model 302B, e.g., utilizing nested cross-validation with group K-fold cross-validation. In this manner, workflow diagram 300B may optimize machine learning model 302B to more efficiently generate predictions of molecular binding properties of one or more target proteins (e.g., reducing the execution time of machine learning model 302B and the database capacity suitable for storing machine learning model 302B) compared to machine learning model 302A, e.g., as described above with respect to FIG. 3A .
[0091] In certain embodiments, workflow diagram 300B may then continue with training and evaluating optimized machine learning model 302B in function block 356. For example, in some embodiments, optimized machine learning model 302B (e.g., as optimized in function block 354) may be trained and evaluated based on descriptor vectors representing the amino acid sequences of a set of one or more P proteins (e.g., as calculated in function block 352), and the feature dimensionality reduction model 307B, feature selection model 309B, and regression model 311B function blocks selected for execution.
[0092] In certain embodiments, workflow diagram 300B at function block 356 may further include applying the optimized set of hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and the optimized set of learnable parameters (e.g., regression model weights, decision variables) (e.g., iteratively optimized, as described above with respect to workflow diagram 300A of FIG. 3A ) to optimized machine learning model 302B and utilizing optimized machine learning model 302B to generate predictions of molecular binding properties of one or more target proteins according to presently disclosed embodiments. In certain embodiments, workflow diagram 300B may then conclude at function block 358 by storing the optimized machine learning model 302B, the optimized set of hyperparameters (e.g., general parameters, booster parameters, learning task parameters), and the optimized set of learnable parameters (e.g., regression model weights, decision variables) utilized in subsequent predictions of molecular binding properties of one or more target proteins.
[0093] In certain embodiments, for example, during the inference phase (e.g., after the optimized machine learning model 302B has been trained and stored with the optimized set of learnable hyperparameters as described above), as further illustrated by FIG. 3B , the feature dimensionality reduction model 307B of the machine learning model 302B may receive or import a molecular descriptor matrix and scale and normalize one or more sets of descriptors in the descriptor matrix. For example, the molecular descriptor matrix may represent a set of amino acid sequences corresponding to a set of P proteins. In certain embodiments, the feature dimensionality reduction model 307B may perform clustering of one or more sets of descriptors by determining the correlation distance (e.g., 1-abs (Pearson's correlation)) between the descriptors and then store only the descriptors closest to the centroid. For example, in some embodiments, utilizing the calculated correlation distance metric, which may be calculated based on Pearson's correlation, the feature dimensionality reduction model 307B may cluster the feature vectors to group any and all redundant features containing similar information (similar feature vectors) and determine the centroid of each cluster as representing the cluster. In particular embodiments, the feature dimensionality reduction model 307B may then optimize the number of selected descriptors.
[0094] In certain embodiments, the feature selection model 309B may calculate the nonlinear correlation between descriptors and output a percent protein binding using an MIC metric, distance correlation, mutual information, or other similar nonlinear correlation metric, and / or other linear correlation metric (e.g., Pearson's correlation). In one or more other embodiments, the feature selection model 309B may calculate the nonlinear correlation between descriptors and output a percent protein binding using distance correlation, mutual information, or other similar nonlinear correlation metric. For example, the feature selection model 309B may determine the k-best most predictive feature vector of the reduced molecular descriptor matrix based on a k-best process and MIC to determine the correlation between the feature vector of the reduced molecular descriptor matrix and experimentally determined percent protein binding at one or more particular pH values and salt concentrations and / or salt species and chromatography resins. In other embodiments, instead of using MIC as part of the correlation determination, distance correlation, mutual information, or other similar nonlinear correlation metric may be used. The feature selection model 309B may then select highly correlated descriptors and optimize the selected descriptors. In particular embodiments, the feature selection model 309B may select a set of descriptors based on their impact on the overall performance (e.g., processing speed, storage capacity) of the machine learning model 302B. For example, in some embodiments, the feature selection model 309B may iteratively evaluate K descriptors to determine the number of results that yields the optimal performance of the machine learning model 302B. In some embodiments, the feature selection model 309B may be entirely selective in selecting a set of descriptors based on their impact on overall performance.
[0095] For example, in other embodiments, feature selection model 309B may implement one or more Volta feature selection algorithms, one or more SHapley Additive exPlanations (SHAP) feature selection algorithms, or other similar recursive feature elimination algorithms, e.g., to select K descriptors and optimize the percentage of the number of K descriptors selected. In particular embodiments, regression model 311B of machine learning model 302B may then receive as input the pH values, salt concentrations, and descriptor sequence-based descriptors, and then output a prediction of the percentage of proteins bound to the set of proteins P, optimizing machine learning model 302B by minimizing the loss between the predicted protein binding percentage and the experimentally determined protein binding percentage.
[0096] 3B, the machine learning model 302B is iteratively trained to identify target proteins and generate predictions of molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates. For example, the streamlined process of identifying target proteins (e.g., antibodies) in silico can facilitate and accelerate the downstream development and manufacturing of one or more therapeutic mAbs, bsAbs, tsAbs, or other similar immunotherapies that can be utilized to treat various patient diseases.
[0097] For example, once trained, the machine learning model 302A (e.g., a "boosting" ensemble learning model) may be utilized to generate predictions of molecular binding properties of one or more proteins (e.g., predictions of percent protein binding at one or more particular pH values and particular salt concentrations and / or particular salt species and chromatography resins) by utilizing the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training of the machine learning model 302A, as well as a selected k-best matrix of feature vectors of molecular descriptor matrices representing sets of amino acid sequences corresponding to one or more proteins of interest.
[0098] Specifically, according to presently disclosed embodiments, once trained, the machine learning model 302B may utilize the optimized hyperparameters (e.g., general parameters, booster parameters, learning task parameters) and learnable parameters (e.g., regression model weights, decision variables) learned during training to predict percent protein binding (e.g., percentage of a set of proteins predicted to bind to a ligand in solution for a given pH value and salt concentration) for one or more target proteins based solely on a selected k-best matrix of feature vectors of molecular descriptor matrices representing, as input, a set of amino acid sequences corresponding to one or more proteins of interest and one or more sets of pH values and salt concentrations and / or salt species and chromatography resins associated with the binding properties of the one or more proteins of interest. Furthermore, according to disclosed embodiments, the ensemble learning 302B may be further optimized using one or more Bayesian optimization processes to more efficiently generate predictions of molecular binding properties (e.g., predictions of percent protein binding at one or more particular pH values and particular salt concentrations and / or particular salt species and chromatography resins).
[0099] In this manner, by providing an optimized prediction based on a computational model of percent protein binding for one or more proteins of interest, the molecular binding and elution characteristics of one or more proteins of interest may be determined without significant upstream experimentation. That is, desirable proteins of one or more proteins of interest may be identified and distinguished in silico from undesirable proteins of one or more proteins of interest, and these in silico identified desirable proteins may be further utilized to facilitate and accelerate the downstream development of one or more therapeutic mAbs, bsAbs, tsAbs, or other similar immunotherapies that may be utilized to treat various patient diseases (e.g., by reducing upstream experimental duration and experimental inefficiencies and providing in silico feedback that may be difficult to purify and ultimately manufacture candidate proteins). As an example, based at least in part on a molecular descriptor matrix, a machine learning model may be configured to obtain predictions of molecular binding characteristics of one or more proteins. From the molecular binding characteristics, desirable proteins may be identified. Although the present embodiments are described herein primarily with respect to machine learning model 302B generating predictions of molecular binding properties of one or more target proteins, it should be understood that the trained machine learning model 302B may also generate predictions of elution properties of one or more proteins or generate predictions of flow-through properties of one or more proteins in accordance with embodiments disclosed herein.
[0100] FIG. 4 shows a flow diagram of a method 400 for identifying target proteins (e.g., antibodies) and generating predictions of the molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates, according to disclosed embodiments. Method 400 may be performed utilizing one or more processing devices (e.g., the computing devices and artificial intelligence architectures described below in connection with FIGS. 5 and 6) that may include hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a vision processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), or any other processing device that may be suitable for processing genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data to create one or more processors), firmware (e.g., microcode), or some combination thereof.
[0101] Method 400 may begin, at block 402, with one or more processing devices accessing a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins. Method 400 may then continue, at block 404, by one or more processing devices refining a set of hyperparameters associated with a machine learning model trained to generate predictions of molecular binding properties of the one or more proteins. As shown, method 400 may then proceed to an iterative subprocess of optimizing the set of hyperparameters by iteratively performing subprocesses (e.g., indicated by dashed lines around portions of method 400 in FIG. 4 ) until the machine learning model reaches a desired accuracy.
[0102] For example, method 400 may continue at block 406 with one or more processing devices reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each of the feature vector clusters containing similar feature vectors. Method 400 may then continue at block 408, where one or more processing devices may determine one or more most predictive feature vectors among the selected representative feature vectors for each of the plurality of feature vector clusters based on correlations between the selected representative feature vectors and predetermined batch binding data associated with one or more proteins. Method 400 may then continue at block 410, where one or more processing devices calculate one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch binding data. Method 400 may then end at block 412 with one or more processing devices updating a set of hyperparameters based on the one or more cross-validation losses.
[0103] 5 illustrates an example of one or more computing devices 500 that may be utilized to identify target proteins (e.g., antibodies) and generate predictions of molecular binding properties of one or more target proteins as part of a streamlined process of protein purification to accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates through early identification of promising therapeutic antibody candidates, according to disclosed embodiments. In certain embodiments, one or more computing devices 500 may perform one or more steps of one or more methods described or illustrated herein. In certain embodiments, one or more computing devices 500 provide functionality described or illustrated herein. In certain embodiments, software running on one or more computing devices 500 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Certain embodiments include one or more portions of one or more computing devices 500.
[0104] The present disclosure contemplates any suitable number of computing systems 500. The present disclosure contemplates one or more computing devices 500 taking any suitable physical form. By way of example and not limitation, the one or more computing devices 500 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the one or more computing devices 500 may be single or distributed, spanning multiple locations, spanning multiple machines, spanning multiple data centers, or in a cloud, which may include one or more cloud components of one or more networks.
[0105] Where appropriate, one or more computing devices 500 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computing devices 500 may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. One or more computing devices 500 may, where appropriate, perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.
[0106] In particular embodiments, one or more computing devices 500 include a processor 502, a memory 504, a database 506, an input / output (I / O) interface 508, a communication interface 510, and a bus 512. While this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In particular embodiments, processor 502 includes hardware for executing instructions, such as instructions making up a computer program. By way of example and not limitation, to execute instructions, processor 502 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 504, or database 506, decode and execute those instructions, and then write one or more results to an internal register, an internal cache, memory 504, or database 506. In particular embodiments, processor 502 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 502 including any suitable number of any suitable internal caches, where appropriate. By way of example and not limitation, processor 502 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 504 or database 506, and the instruction cache may speed up retrieval of those instructions by processor 502.
[0107] The data in the data cache may be a copy of data in memory 504 or database 506 upon which instructions executing in processor 502 operate, results of previous instructions executed in processor 502 for access by subsequent instructions executing in processor 502 or for writing to memory 504 or database 506, or other suitable data. The data cache may speed up read or write operations by processor 502. The TLB may speed up virtual address translation for processor 502. In particular embodiments, processor 502 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processors 502 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 502 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include more than one processor 502. While this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0108] In particular embodiments, memory 504 includes main memory for storing instructions for processor 502 to execute or data on which processor 502 operates. By way of example and not limitation, one or more computing devices 500 may load instructions into memory 504 from database 506 or another source (e.g., another one or more computing devices 500). Processor 502 may then load the instructions from memory 1204 into an internal register or cache. To execute the instructions, processor 502 may retrieve the instructions from the internal register or cache and decode them. During or after executing the instructions, processor 502 may write one or more results (which may be intermediate or final results) to an internal register or cache. Processor 502 may then write one or more of those results to memory 504.
[0109] In particular embodiments, processor 502 executes instructions only from one or more internal registers, internal caches, or memory 504 (as opposed to database 506 or elsewhere) and operates only on data in one or more internal registers, internal caches, or memory 504 (as opposed to database 506 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may connect processor 502 to memory 504. Bus 512 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 502 and memory 504 and facilitate accesses to memory 504 requested by processor 502. In particular embodiments, memory 504 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. This RAM may be dynamic RAM (DRAM) or static RAM (SRAM), where appropriate. Moreover, this RAM may be single-ported RAM or multi-ported RAM, where appropriate. The present disclosure contemplates any suitable RAM. Memory 504 may include, where appropriate, one or more memories 504. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.
[0110] In particular embodiments, database 506 includes mass storage for data or instructions. By way of example and not limitation, database 506 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more thereof. Database 506 may include removable or non-removable (or fixed) media, where appropriate. Database 506 may be internal or external to one or more computing devices 500, where appropriate. In particular embodiments, database 506 is non-volatile solid-state memory. In particular embodiments, database 506 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), flash memory, or a combination of two or more thereof. The present disclosure contemplates mass database 506 taking any suitable physical form. Database 506 may include, where appropriate, one or more storage control units that facilitate communication between processor 502 and database 506. Where appropriate, database 506 may include one or more databases 506. Although this disclosure describes and illustrates particular storage devices, this disclosure contemplates any suitable storage device.
[0111] In particular embodiments, I / O interface 508 includes hardware, software, or both that provide one or more interfaces for communication between one or more computing devices 500 and one or more I / O devices. One or more computing devices 500 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and one or more computing devices 500. By way of example and not limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interface 508 therefor. Where appropriate, I / O interface 508 may include one or more device or software drivers that enable processor 502 to drive one or more of these I / O devices. I / O interface 1208 may include, where appropriate, one or more of I / O interfaces 508. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.
[0112] In particular embodiments, communication interface 510 includes hardware, software, or both that provide one or more interfaces for communication (e.g., packet-based communication) between one or more computing devices 500 and one or more other computing devices 500 or one or more networks. By way of example and not limitation, communication interface 510 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 510 for that network.
[0113] By way of example, and not limitation, one or more computing devices 500 may communicate with an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, one or more computing devices 500 may communicate with a wireless PAN (WPAN) (e.g., a BLUETOOTH WPAN), a Wi-Fi network, a Wi-MAX network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), other suitable wireless networks, or a combination of two or more of these. One or more computing devices 500 may include any suitable communication interface 510 for any of these networks, where appropriate. The communication interface 510 may include one or more communication interfaces 510, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.
[0114] In particular embodiments, bus 512 includes hardware, software, or both that connects one or more components of computing device 500 to one another. By way of example, and not limitation, bus 512 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, another suitable bus, or a combination of two or more of these. Bus 512 may include one or more buses 512, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0115] As used herein, one or more computer-readable non-transitory storage media may comprise, where appropriate, one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Non-transitory computer-readable storage media may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.
[0116] FIG. 6 shows a diagram 600 of an exemplary artificial intelligence (AI) architecture 602 (which may be included as part of one or more computing devices 500, as described above in connection with FIG. 5) that may be utilized to generate predictions of molecular binding properties of one or more target proteins (e.g., antibodies) as part of a streamlined process of protein purification to identify target proteins and accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates, according to disclosed embodiments. In particular embodiments, the AI architecture 602 may be implemented utilizing one or more processing devices, which may include, for example, hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and / or any other processing device that may be suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructions running / executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.
[0117] 6, AI architecture 602 may include machine learning (ML) algorithms and functions 604, natural language processing (NLP) algorithms and functions 606, expert systems 608, computer-based vision algorithms and functions 610, speech recognition algorithms and functions 612, planning algorithms and functions 614, and robotics algorithms and functions 616. In particular embodiments, ML algorithms and functions 604 may include any statistical-based algorithms that may be suitable for finding patterns across large amounts of data (e.g., "big data" such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data). For example, in particular embodiments, ML algorithms and functions 604 may include deep learning algorithms 618, supervised learning algorithms 620, and unsupervised learning algorithms 622.
[0118] In particular embodiments, the deep learning algorithm 618 may include any artificial neural network (ANN) that can be utilized to learn deep-level representations and abstractions from large amounts of data. For example, the deep learning algorithm 618 may include ANNs such as perceptrons, multi-layer perceptrons (MLPs), autoencoders (AEs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memories (LSTMs), grated recurrent units (GRUs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks, neural autoregressive distribution estimation (NADEs), adversarial networks (ANs), attention models (AMs), spiking neural networks (SNNs), deep reinforcement learning, etc.
[0119] In particular embodiments, supervised learning algorithm 620 may include any algorithm that can be utilized to apply what has been learned in the past to new data, e.g., using labeled examples to predict future events. For example, starting from an analysis of a known training data set, supervised learning algorithm 620 may create an inferred function to make a prediction about output values. Supervised learning algorithm 500 may also compare its output with the correct intended output and find errors in order to correct supervised learning algorithm 620 accordingly. On the other hand, unsupervised learning algorithm 622 may include any algorithm that can be applied, e.g., when the data used to train unsupervised learning algorithm 622 is neither classified nor labeled. For example, unsupervised learning algorithm 622 may study and analyze how a system can infer functions to describe hidden structure from unlabeled data.
[0120] In particular embodiments, NLP algorithms and functions 606 may include any algorithms or functions that may be suitable for automatically manipulating natural language, such as speech and / or text. For example, in some embodiments, NLP algorithms and functions 606 may include content extraction algorithms or functions 624, classification algorithms or functions 626, machine translation algorithms or functions 628, question answering (QA) algorithms or functions 630, and text generation algorithms or functions 632. In particular embodiments, content extraction algorithms or functions 624 may include means for extracting text or images from electronic documents (e.g., web pages, text editor documents, etc.) for use in other applications, for example.
[0121] In particular embodiments, classification algorithm or function 626 may include any algorithm that may utilize a supervised learning model (e.g., logistic regression, naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machines (SVMs), etc.) to learn from and make new observations or classifications based on data input into the supervised learning model. Machine translation algorithm or function 628 may include any algorithm or function that may be suitable for automatically converting source text in one language into text in another language, for example. QA algorithm or function 630 may include any algorithm or function that may be suitable for automatically answering questions posed by humans in natural language, such as those performed by a voice-controlled personal assistant device, for example. Text generation algorithm or function 632 may include any algorithm or function that may be suitable for automatically generating natural language text.
[0122] In particular embodiments, expert system 608 may include any algorithms or functions that may be suitable for simulating the judgment and actions of a human or organization with expertise and experience in a particular field (e.g., stock trading, medicine, sports statistics, etc.). Computer-based vision algorithms and functions 610 may include any algorithms or functions that may be suitable for automatically extracting information from images (e.g., photographic images, video images). For example, computer-based vision algorithms and functions 610 may include image recognition algorithms 634 and machine vision algorithms 636. Image recognition algorithms 634 may include any algorithms that may be suitable, for example, for automatically identifying and / or classifying objects, places, people, etc. that may be included in one or more image frames or other display data. Machine vision algorithms 636 may include any algorithms that may be suitable for enabling a computer to "see" or that may be suitable, for example, for relying on image sensor cameras with specialized optics to acquire images in order to process, analyze, and / or measure various data characteristics for decision-making purposes.
[0123] In particular embodiments, speech recognition algorithms and functions 612 may include any algorithms or functions that may be suitable for recognizing and translating spoken language into text, such as through automatic speech recognition (ASR), computer speech recognition, speech-to-text (STT) 638, or text-to-speech (TTS) 640, for computing purposes to communicate with one or more users via voice. In particular embodiments, planning algorithms and functions 614 may include any algorithms or functions that may be suitable for generating a sequence of actions, where each action may include its own set of preconditions to be satisfied before performing the action. Examples of AI planning may include classical planning, reduction to other problems, temporal planning, probabilistic planning, preference-based planning, conditional planning, etc. Finally, robotics algorithms and functions 616 may include any algorithms, functions, or systems that may enable one or more devices to replicate human behavior, for example, through movements, gestures, performance tasks, decision-making, emotions, etc.
[0124]
[0003] What is described herein includes processes related to predicting molecular binding properties of one or more proteins, as described above. This may include importing an amino acid sequence of the protein and generating a molecular descriptor matrix based on the amino acid sequence. A protein molecule is formed from the amino acid sequence. The amino acid sequence may be represented by a character string (e.g., a string). In one or more examples, the amino acid sequence may be input into a machine learning model (e.g., a neural network) to generate the molecular descriptor matrix. In one or more examples, the machine learning model may be pre-trained using the amino acid sequence. For example, the machine learning model may include a protein language model. In another example, the machine learning model may be pre-trained in an unsupervised manner. In some embodiments, the machine learning model may be configured to generate structure-based descriptors representing the sequence used to generate the protein structure.
[0125] The generated molecular feature matrix can be used to predict molecular binding properties of the corresponding protein. In some embodiments, the molecular descriptor matrix can be a multidimensional matrix (i.e., a tensor) composed of multiple feature vectors representing descriptors for each amino acid in each protein's sequence. To determine predicted molecular binding properties, in one or more instances, the dimensionality of the multidimensional molecular descriptor matrix (e.g., descriptor tensor) can be reduced. In some embodiments, the multidimensional molecular descriptor matrix can be reduced to a two-dimensional molecular feature matrix (with a feature vector for each amino acid in each molecule) by averaging the feature vectors across all amino acids in each molecule. In some embodiments, feature dimensionality reduction techniques used to reduce the number of feature vectors in the molecular descriptor matrix can include, among other things, removing redundant feature vectors after averaging. For example, some feature vectors (and / or the features contained therein) can be highly correlated, and therefore a single representative feature vector can be identified to represent a set of highly correlated feature vectors. In some embodiments, clustering techniques (e.g., hierarchical / agglomerative clustering techniques) may be used to identify similar feature vectors (e.g., whose corresponding embeddings are less than a threshold distance apart from each other in the embedding space). From the identified feature vectors, one or more representative feature vectors may be selected from each cluster of similar feature vectors as being "representatives" of that cluster.
[0126] As described above, the representative feature vectors can be input into a machine learning model to obtain predictions of the molecular binding properties of proteins. These proteins can be proteins of interest for potential drug discovery assays. The machine learning model can receive one or more representative feature vectors describing one or more proteins as input, and can be trained to output predictions of the molecular binding properties of the proteins based on the representative feature vectors.
[0127] In some embodiments, the machine learning model may be trained by aligning molecular descriptors (from a training molecular descriptor matrix generated by machine learning model 301 of FIG. 3A based on the imported amino acid sequences of one or more empirically evaluated proteins) with predetermined batch binding data associated with the empirically evaluated proteins. Once aligned, supervised regression may be performed to train the machine learning model. In one or more examples, the regressor used may include a bagged decision tree, a bagged linear model, an unbagged linear model, a random forest, a linear forest, or another type of regressor, or a combination thereof.
[0128] In some embodiments, part of the training step includes optimizing a set of hyperparameters of the machine learning model. In one or more examples, the hyperparameters may include a regularization parameter, the number of estimators, a maximum tree depth, etc. In one or more examples, the pipeline (e.g., the feature dimensionality reduction model 307A, the feature selection models 309A, 309B, and the regression models 311A, 311B) may be optimized together. The feature dimensionality reduction model may be configured to reduce the number of feature vectors in the molecular descriptor matrix using correlation clustering, recursive feature elimination, and / or other techniques.
[0129] In some embodiments, the training step may also include a cross-validation step in which an optimized set of learnable parameters for the machine learning model is identified, and then cross-validation tests are performed iteratively until the optimized set of learnable parameters is determined. For example, if the machine learning model includes a decision tree structure (e.g., a random forest), the number of learnable parameters may include the number of trees and / or the tree depth. The optimized set of learnable parameters may be selected to optimize the performance of the machine learning model. The optimized set of learnable parameters may be used to train the machine learning model to generate predictions of molecular binding properties of new amino acid sequences that are not part of the training set.
[0130] In some embodiments, one or more additional steps may be performed to predict the molecular binding properties of one or more protein molecules based on their amino acid sequences. One of the goals of the disclosed technology includes predicting the properties of molecules to be evaluated. In particular, how well a protein molecule binds to a resin provides valuable clinical information and / or valuable manufacturing process development information that can be used in the development of new therapeutics. The above describes a set of additional / alternative steps to those described above that can be performed to predict molecular binding properties based on their amino acid sequences.
[0131] The techniques described herein offer many technical advantages. For example, as illustrated with respect to FIG. 1 , testing the binding properties of molecules, e.g., for drug discovery purposes, is a complex and time-consuming process. For example, the experimental duration of example experiment 102 of FIG. 1 for performing one or more protein purification processes can extend to several weeks. In contrast, the execution time of a computational model-based example 104 (e.g., a machine learning model described herein) for performing one or more protein purification processes can be only a few minutes. Thus, example experiment 102 describes a non-ideal process for testing all potential molecules. The machine learning models described herein can reduce the time spent on testing by increasing the number of molecules that can be screened in a given amount of time or by a given researcher. As another example, molecular descriptor matrices can be generated using various existing protein language models (e.g., molecular descriptors 120 of FIG. 1 ). Thus, existing technology can be utilized to generate inputs for the machine learning model, thereby reducing the amount of additional data that needs to be collected and the amount of additional model training required. As yet another example, the machine learning models described herein can be trained using less data while maintaining or improving model accuracy. For example, instead of inputting a molecular descriptor matrix into a machine learning model (which may contain a very large number of descriptors), the molecular descriptor matrix can be reduced to determine the most predictive feature vectors (and used as input to the machine learning model). This descriptor reduction process can further optimize the training process for the machine learning model. For example, each training molecular descriptor matrix can be reduced by determining the most predictive feature vector, and the model can be trained based on the most predictive feature vectors.
[0132] 7 shows another high-level workflow diagram 700 for performing feature generation 202, feature dimensionality reduction 204, feature filtering 206, recursive model-based feature elimination 207, and regression model optimization 208, according to various embodiments. The descriptions of feature generation 202, feature dimensionality reduction 204, feature filtering 206, and regression model optimization 208 may apply here as well.
[0133] However, unlike diagram 200, diagram 700 may further include recursive model-based feature elimination 207. Recursive model-based feature elimination 207 may include additional models to further reduce the number of features in the feature set. In particular, recursive model-based feature elimination 207 may help prevent or reduce the possibility of overfitting. As an example, referring to FIG. 8, recursive model-based feature elimination 207 may implement machine learning model 820 of FIG. 8. FIG. 8 includes similar components to FIG. 3A, and similar labels are used to refer to those components. For example, workflow 800 may include model 301 and machine learning model 820. Machine learning model 820 may include feature dimensionality reduction model 307A, feature filtering model 309A, recursive feature elimination model 801, and regression model 311A. Workflow 800 may follow a path similar to that of workflow 300A, except that the most predictive feature vectors may include those reduced via recursive feature elimination model 801. As an example, determining the one or more most predictive feature vectors may further include implementing a recursive feature elimination model 801 to further reduce the number of feature vectors. In particular, some embodiments include the number of feature vectors included in the further reduced number of feature vectors being less than or equal to the number of training items.
[0134] In some embodiments, in function block 802, the recursive feature removal model 801 may be configured to fit a model to the representative feature vectors. The model may be, for example, a regression model. In function block 804, a feature importance score may be calculated based on the fitted model. The feature importance score may indicate the importance of each representative feature vector. In function block 806, one or more feature vectors from the representative feature vectors may be removed based on their respective feature importance scores to obtain a subset of representative feature vectors. For example, the least important features or feature vectors may be removed from the representative feature vectors. In one or more examples, the most predictive feature vectors may include one or more feature vectors from the subset of representative feature vectors. In function block 808, the recursive feature removal model 801 may iteratively perform blocks 802-806 until the number of feature vectors included in the subset satisfies a feature criterion. For example, the satisfied feature criterion may include the number of feature vectors included in the subset of representative feature vectors being less than or equal to a threshold number of feature vectors. In some examples, the threshold number of feature vectors may include the same or similar number of features from the training data used to train the machine learning model 820. In some examples, the number of feature vectors included in the subset of representative feature vectors may comprise one of a set of hyperparameters. In some examples, the number of feature vector clusters included in the plurality of feature vector clusters comprises one of a set of hyperparameters.
[0135] FIG. 9 illustrates a process for training a machine learning model to predict molecular binding properties, according to various embodiments. As an example, compared to the process described above with respect to FIGS. 3A-4, process 900 of FIG. 9 may organize the data used to train a regression model (e.g., in step 930) differently, as described herein. As an example, in the above process, for each empirically evaluated protein, the data used to train the machine learning model (e.g., machine learning pipeline 908) includes a predetermined amount of experimental conditions. The experimental conditions may specify the molecular binding properties of the protein for a given set of experimental conditions. For example, the data may include the measured molecular binding level of the protein at a first salt concentration and a first pH level, the measured molecular binding level of the protein at a second salt concentration and a first pH level, the measured molecular binding level of the protein at a first salt concentration and a second pH level, etc. In one or more examples, the given amount of experimental conditions for a given batch binding may include 12 or more experimental conditions (e.g., 4 salt concentrations, 3 pH levels), 24 or more experimental conditions (e.g., 6 salt concentrations, 4 pH levels), etc. A trained machine learning model may use the experimental conditions (e.g., pH levels and salt concentrations) as inputs in addition to a molecular descriptor matrix to predict molecular binding properties of one or more proteins, as described above.
[0136] In some embodiments, experimental conditions need not be input into the machine learning model; instead, predicted molecular binding properties may be determined for a set of experimental conditions. However, to do so, the training data and training process may be adjusted, as shown in FIG.
[0137] FIG. 9 shows a workflow diagram of a process 900 for optimizing hyperparameters and learnable parameters of a machine learning model for implementing one or more computational model-based protein purification processes, according to various embodiments. Process 900 differs from that described above with respect to FIGS. 3A-4 in that a transformed representation of empirically evaluated molecular binding properties of a protein can be used to train the machine learning model. The trained machine learning model may output values corresponding to the transformed representation of the molecular binding properties, which can be used to predict all binding conditions for all experimental conditions for a given protein molecule. Thus, the amount of training data required to train the machine learning model can be reduced from N empirically derived binding indices for N different experimental conditions (e.g., salt concentration levels and pH levels) to a single transformed binding index that can be used to resolve the N empirically derived binding indices.
[0138] In process 900, sequence data 902 corresponding to one or more amino acid sequences of protein P may be provided to a matrix-generating machine learning (ML) model 904. In some embodiments, machine learning model 904 may be the same as or similar to machine learning model 301 of FIG. 3A , and the foregoing description may apply. In some embodiments, matrix-generating ML model 904 may be trained to generate a molecular descriptor matrix 906 from sequence data 902 representing the amino acid sequence of protein P. Matrix-generating ML model 904 may comprise a neural network that may generate features X structured as molecular descriptor matrix 906. Molecular descriptor matrix 906 may be the same as or similar to the molecular descriptor matrix generated in function block 306 of FIG. 3A . In one or more examples, molecular descriptor matrix 906 may include 100 or more features, 500 or more features, 1,000 or more features, 2,000 or more features, 10,000 or more features, or some other amount of features. The features in the molecular descriptor matrix 906 may then be analyzed to determine which, if any, correlate with the molecular binding properties of the corresponding protein molecule.
[0139] The molecular descriptor matrix 906 may have dimensions M, the number of molecules by N, the number of descriptors (e.g., features). For a given molecule, the amino acid sequence may be represented using a series of letters (e.g., alphabet) that form the protein being tested.
[0140] In some embodiments, sequence 902 may be experimentally analyzed. The experiment may generate empirically derived protein binding data 912. The empirically derived protein binding data 912 may include molecular binding property values for a set of experimental conditions 914. For example, the empirically derived protein binding data 912 may indicate that for a given sequence (e.g., sequence A) and first experimental conditions (e.g., a first salt concentration level and a first pH level), the molecular binding property is Y1. Similarly, the empirically derived protein binding data 912 may indicate that for a sequence (e.g., sequence A) and second experimental conditions (e.g., a second salt concentration level and a first pH level), the molecular binding property is Y2. Furthermore, the empirically derived protein binding data 912 may indicate that for a sequence (e.g., sequence A) and third experimental conditions (e.g., a first salt concentration level and a second pH level), the molecular binding property is Y3. In some embodiments, a given batch of binding data can be formulated similarly with molecules as rows and experimental conditions 914 as columns.
[0141] Process 900 can be configured to train a machine learning model (e.g., machine learning model 820) to predict the molecular binding properties of a protein for a set of experimental conditions. Testing the binding properties of molecules, such as for drug discovery purposes, is a complex and time-consuming process (e.g., it takes 2-6 weeks to grow, purify, and then test a molecule, so it can take several weeks to fully evaluate each molecule). Testing every potential molecule is not ideal. Therefore, the goal of this model is to increase the number of molecules that can be screened in a given amount of time or by a given researcher. Another goal is to increase the number of molecules that can be screened without incurring timeline delays or additional experimental burden.
[0142] Process 900 may be trained using a small number of training examples (e.g., a small number of molecules) and a large number of descriptors (e.g., 100 or more features, 500 or more features, 1,000 or more features, 2,000 or more features, 10,000 or more features, etc.). Process 900 may systematically sort the descriptors to train a machine learning model to predict molecular binding properties 910. Furthermore, process 900 may utilize descriptors that have a relationship to one or more physical attributes of the protein. Thereby, machine learning pipeline 908 may be configured to find the descriptors (e.g., features) that best predict the molecular binding properties of the protein based on the molecular descriptor matrix. The ML model may then try and determine which descriptors are most predictive.
[0143] In an illustrative example, the given batch binding data 912 includes empirically measured binding properties of each protein analyzed for a set of experimental conditions. In some embodiments, the process 900 may include performing a linearization transformation 916 to the empirically measured binding properties, e.g., using the computing system 500 of FIG. 5 . For example, the empirically measured binding properties may include a percent binding indicator (e.g., the protein is Y% bound to the resin). The process 900 may convert the empirically measured binding properties of percent binding stored in the given batch binding data 912 to a linearized or pseudo-linear representation of the empirically measured binding properties. For example, a logit transformation operation may be performed.
[0144] The logit transformation involves calculating the logarithm of the ratio of bound / unbound protein concentrations. This transformation results in binding going from 0.0 to 1.0 (i.e., 0% binding to 100% binding) to negative infinity to positive infinity (log(K)). p ) space). Once the data is linearized, a linear model such as a PCA model can be used, which will converge better.
[0145] In some embodiments, process 900 may include applying one or more dimensionality reduction techniques (e.g., principal component analysis (PCA) 918) to a linear representation of the empirically measured binding properties of each analyzed protein. PCA 918 may be configured to derive first, second, etc. principal components (PCs) of a linearizing transformation (e.g., a logit transform) of those empirically measured binding properties. Once performed, PCA 918 may represent the linear representation of the empirically measured protein binding properties in a more concise representation. In particular, the number of experimental conditions, C, defines the number of data points in a given batch binding data 912. PCA 918 may reduce the number of data points from C to C or less. For example, if 24 experimental conditions were used to obtain a given batch binding data 912, PCA may accept the given batch binding data 912 as less than (or equal to) 24 data points. PCA 918 may be configured to output a transformed representation 920 that represents a transformed version of the empirically measured molecular binding properties. The number of molecules tested can be 1 or more, 5 or more, 10 or more, 20 or more, 50 or more, or other values.
[0146] A PCA model can decompose data (e.g., a given batch of binding data 912) into a set of low-dimensional vectors. For example, for 24 experimental conditions (e.g., 24 experimental data points), a PCA model can identify the first eigenvector of the data, which can capture the variance of the dataset. PCA therefore allows for the use of low-dimensional projections to describe the behavior of binding data. In one or more examples, when predicting the average binding efficiency of a molecule, PCA provides a more representative and valuable result than any of the individual experimental conditions. Furthermore, PCA's ability to succinctly summarize (in a low-dimensional representation) trends in noisy, multidimensional data can be useful to scientists. As will be appreciated by those skilled in the art, any number of principal components can be identified by PCA 918, including, but not limited to, a first principal component and / or a second principal component.
[0147] In some embodiments, the predicted molecular binding properties 910 may be compared to transformed representations 920 of empirically measured molecular binding properties. In one or more examples, a cross-validation loss may be calculated to determine how well the machine learning model 908 predicted the empirically measured molecular binding properties of a given protein. In particular, the prediction indicates how well the machine learning pipeline 908 predicts the transformed representations of the empirically measured molecular binding properties.
[0148] In some embodiments, at 930, a cross-validation loss may be calculated. As mentioned above, one or more examples may use a k-fold cross-validation technique. Additionally or alternatively, at 930, stratified k-fold cross-validation may be calculated. Stratified k-fold cross-validation involves taking molecules of a training set and ranking them into bins based on their molecular binding properties. For example, the bins may include a first bin corresponding to weakly binding proteins, a second bin corresponding to moderately binding proteins, a third bin corresponding to tightly binding proteins, etc. The stratified k-fold cross-validation may then evaluate the performance of the machine learning pipeline 908 by selecting a representative subset, such as an even number of representatives of each of the binned proteins.
[0149] 10A-10D show exemplary plots illustrating how principal component analysis can be used to predict molecular binding properties, according to various embodiments. FIG. 10A, for example, shows a plot 1000 of the results of principal component analysis of a set of molecules. In plot 1000, the X-axis represents the first component of each molecule in the set. の The Y-axis corresponds to the principal component value, and the Y-axis corresponds to the second の corresponds to the principal components. The red and green ellipses represent the first and second standard deviations from the centroid of the cluster of data points. As can be seen in plot 1000, the molecules (e.g., data points) are fairly well distributed around the x-axis.
[0150] FIG. 10B shows a plot 1020 of isotherms for a given molecule for various values of the first principal component, according to various embodiments. In plot 1020, the x-axis represents the salt concentration level used during the corresponding experiment to determine the molecule's protein binding properties, and the y-axis represents the protein binding level. Isotherms 1022-1030 correspond to different principal component (PC) values. For example, isotherm 1022 represents how protein binding changes as the salt concentration level is changed for the first PC (e.g., PC=-6). Isotherm 1024 represents how protein binding changes as the salt concentration level is changed for the second PC (e.g., PC=-3). Isotherm 1026 represents how protein binding changes as the salt concentration level is changed for the third PC (e.g., PC=0). Isotherm 1028 represents how protein binding changes as the salt concentration level is changed for the fourth PC (e.g., PC=+3). Isotherm 1030 represents how protein binding changes as salt concentration levels are changed for the fifth PC (eg, PC=+6).
[0151] Isotherms 1022-1030 of plot 1020 can be calculated using fixed pH levels. As can be seen from plot 1020, the binding behavior changes as the first PC value (e.g., -6 for curve 1022) goes to a very large value (e.g., +6 for curve 1030). In the example of plot 1020, the percent binding is approximately 100% at low salt concentrations and approximately 0% at high salt concentration values.
[0152] When producing a protein, one or more protein purification steps may be performed to remove molecules that are not the protein of interest. Protein purification steps involve binding or otherwise facilitating the binding of the protein of interest to a resin (e.g., a chromatography column). Ideally, the resin binds all of the protein of interest. To remove the protein from the resin, a wash may then be applied to deposit the protein of interest into a solution. The wash may contain salt at a particular salt concentration level (and / or pH level). The salt concentration level may affect whether the protein binds to the resin. For example, at lower salt concentration levels, the protein may remain bound to the resin, while at higher salt concentration levels, the protein may detach from the resin. Once removed, assays or other studies may be performed on the solution / protein.
[0153] Generally, for all resins, molecules that are fully or nearly bound throughout the design space (the typical range of pH and salt over which the protein is stable) may not be suitable because they will not be able to elute the protein, resulting in low yield. For resins typically operated in bind-and-elute mode, molecules that are fully or nearly bound throughout the design space may not be suitable because they will not be able to bind to the protein. Therefore, it is desirable to have a protein that exhibits variability in its binding, which should transition from a bound state (e.g., greater than 90% bound) to an unbound state (e.g., less than 10% bound). The first principal component, as shown in Figures 10A-10B, uses a single value (e.g., principal component) instead of a set of experimental conditions (e.g., 24 salt / pH combinations) to represent the transition of the protein from a bound to an unbound state (e.g., as seen by isotherms 822-830).
[0154] In some embodiments, the first principal component can visually represent average binding as percent binding, as shown in plot 1020 of Figure 10B. For example, for a first principal component of -6, the protein may bind tightly to the resin, as seen by isotherm curve 1022. Isotherm curve 1022 can be flagged as problematic because, regardless of the salt concentration level, for a particular pH level and first principal component value, the protein under analysis is unlikely to debind from the resin.
[0155] Looking at plot 1040 in FIG. 10C, different sets of principal component values may be analyzed. The protein and pH levels in the analysis used in the example plot 1040 may be similar to those in plot 1020 in FIG. 10B, but this is not required. As seen in plot 1040, isotherms 1042-1050 show how binding percentages change as the salt concentration level of the wash is changed for different values of the first principal component. For example, isotherm 1042 represents how protein binding changes as the salt concentration level is changed for a first PC (e.g., PC=-10). Isotherm 1044 represents how protein binding changes as the salt concentration level is changed for a second PC (e.g., PC=-5). Isotherm 1046 represents how protein binding changes as the salt concentration level is changed for a third PC (e.g., PC=0). Isotherm 1048 represents how protein binding changes as salt concentration levels are changed for the fourth PC (e.g., PC=+5). Isotherm 1050 represents how protein binding changes as salt concentration levels are changed for the fifth PC (e.g., PC=+10).
[0156] Looking at isotherm curve 1042, the percent binding of the protein does not change significantly as the salt concentration level is changed. Isotherm curve 1042 may then be flagged as problematic because the protein is bound to the resin and cannot be removed. As another example, isotherm curve 1050 may have a substantially static percent binding, regardless of salt concentration level. However, unlike isotherm curve 1042, the protein in this example may not be able to bind to the resin. Therefore, isotherm curve 1050 may also be flagged as problematic because purification cannot be performed because all of the protein is washed away. Isotherm curves 1044-1048 represent a more desirable situation in which the percent binding transitions from bound to unbound as the salt concentration level is changed.
[0157] In some embodiments, by predicting the first principal component, percent binding can be determined for an infinite amount of salt concentrations (and / or pH). In contrast, without PCA, percent binding predictions for all experimental conditions (e.g., points along an isotherm) are required. Thus, using PCA to predict the first principal component significantly simplifies the process of predicting molecular binding properties of proteins without sacrificing accuracy.
[0158] In some embodiments, PCA may output more than the first principal component. For example, a second principal component may also be determined and used to guide the decision-making process. As an example, referring to FIG. 10D , plot 1060 shows isotherms 1062-1070 for a second principal component of a protein. Isotherms 1062-1070 show how percent binding of the protein changes as the salt concentration level changes for a set of second principal component values. For example, isotherm 1062 represents how protein binding changes as the salt concentration level changes for a first PC value (e.g., second PC value = -6). Isotherm 1064 represents how protein binding changes as the salt concentration level changes for a second PC value (e.g., second PC value = -3). Isotherm 1066 represents how protein binding changes as the salt concentration level changes for a third PC value (e.g., second PC value = 0). Isotherm curve 1068 represents how protein binding changes as the salt concentration level changes with respect to a fourth PC value (e.g., a second PC value of +1). Isotherm curve 1070 represents how protein binding changes with respect to a first PC value (e.g., a second PC value of +2). The first principal component can shift the transition from bound to unbound. In some instances, the transition may not occur. The second principal component does not shift the transition as much, but it does change the steepness of isotherm curves 1062-1070. For example, isotherm curve 1062 may include a second PC value of -6, which is very steep compared to isotherm curve 1070, which has a second PC value of +2, as shown, and is not as steep (reaching approximately 0% percent binding). In some embodiments, other principal components may be used. 10D, isotherm 1066 may represent an "ideal" curve. In this example, the first principal component is set to zero and the second principal component may be varied.
[0159] In some embodiments, the machine learning pipeline 908 may be trained to output the first principal component, the second principal component, other principal components, or a combination thereof. The machine learning pipeline 908 may output the principal components together or sequentially.
[0160] In some embodiments, process 900 can reduce the number of data points required to train a machine learning model. For example, the number of principal components can be limited by the number of empirically measured protein data points. In one or more examples, the number of principal components can be less than or equal to the number of experimental conditions. For example, the process described by Figures 3A-4 may require N data points for N experimental conditions, while process 900 of Figure 9 can reduce that number to one data point.
[0161] 11A-11F show experimental conditions and experimental K according to various embodiments. p Values, experimental conditions, and modeled K p 11 shows exemplary heat maps 1100-1150, each showing the relationship between the log(K) value and the salt concentration level. The heat maps 1100-1150 contain a color gradient that represents how tightly the protein is bound (in units of percent binding). The x-axis of the maps 1100-1150 represents salt concentration levels, and the y-axis represents pH levels. The "red" areas of the heat maps 1100-1150 represent higher log(K) values. p ) values (e.g., molecular binding properties), and "green" indicates lower log(K p ) values. Heat maps 1100-1150 can be generated based on one or more empirically evaluated proteins. For example, Figures 11A-11B show experimental K values for ion exchange resins. p Screen and model predictions p 11C-11D show experimental K values for the hydrophobic resin. p Screen and model predictions p 11E-11F show heat maps 1120-1130 showing the experimental K values for mixed-mode resins. p Screen and model predictions pHeat maps 1140-1150 showing the screen may be shown. As an example, at a pH level of pH=5.5, the protein of interest may bind until the salt concentration level used reaches approximately 250 mM in experimental data. Meanwhile, modeled data may show that at a pH level of pH=5.5, the protein of interest may bind until the salt concentration level reaches approximately 175 mM.
[0162] 12 shows a flow diagram of a method 1200 for generating predictions of molecular binding properties of one or more target proteins as part of an alternative streamlined process of protein purification to identify target proteins, according to various embodiments. Method 1200, according to disclosed embodiments, can accelerate the selection process for therapeutic antibody candidates or other immunotherapy candidates by early identification of the most promising therapeutic antibody candidates. Method 1200 may be performed utilizing one or more processing devices (e.g., the computing devices and artificial intelligence architectures described below in connection with FIGS. 5 and 6) that may include hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a vision processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), or any other processing device that may be suitable for processing genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data, to create one or more processors), firmware (e.g., microcode), or some combination thereof.
[0163] In some embodiments, method 1200 may begin at block 1210. Block 1210 may form part of the steps performed to train the machine learning pipeline 908. At block 1210, a training molecular descriptor matrix may be accessed that represents a training set of amino acid sequences corresponding to one or more empirically evaluated proteins. For example, the training molecular matrix may be generated for experimentally evaluated proteins under one or more experimental conditions (e.g., salt concentration levels, pH levels, etc.).
[0164] In block 1220, an iterative process may be performed to refine the set of hyperparameters associated with the ensemble learning model until a desired accuracy is reached. For example, the process may be repeated until the machine learning pipeline 908 predicts molecular binding properties with a threshold level of accuracy. Block 1220 may include steps performed during each iteration of block 1220. For example, in step 1222, the training molecular descriptor matrix may be reduced by selecting one representative feature vector for each of multiple feature vector clusters. Each feature vector cluster may contain similar feature vectors. For example, two feature vectors having a distance (e.g., in the embedding space) less than a threshold distance may be classified as "similar." The selected representative feature vector may represent all feature vectors contained within a given feature vector cluster.
[0165] In step 1224, one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster may be determined based on correlation between the selected representative feature vectors and a given batch of binding data associated with the empirically evaluated protein. The most predictive feature vectors may be determined based on principal component analysis that identifies a first principal component.
[0166] At step 1226, one or more cross-validation losses may be calculated based at least in part on the most predictive feature vector and the given batch combined data. A set of hyperparameters of the machine learning pipeline 908 may be updated based on the cross-validation losses. At step 1228, a set of hyperparameters may be updated based on the one or more cross-validation losses.
[0167] Blocks 1210-1220 (including steps 1222-1228) may comprise a "training" portion. The results of blocks 1210-1220 may include a trained machine learning model (e.g., machine learning model 908) that can be used during inference. For example, in block 1230, a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins may be accessed. In block 1240, predictions of molecular binding properties of the one or more proteins may be obtained by the ML model trained based at least in part on the molecular descriptor matrix.
[0168] In some embodiments, the protein may be a protein of interest. In one or more examples, a machine learning model (e.g., a protein language model implemented using a neural network) may be trained to receive data representing a set of amino acid sequences corresponding to the protein and generate a molecular descriptor matrix that describes the amino acid sequences. In some examples, the molecular descriptor matrix may include multiple descriptors (e.g., features). The descriptors may be configured as a feature vector.
[0169] In some embodiments, the machine learning pipeline 908 may be trained to analyze the molecular descriptor matrix and perform dimensionality reduction. Dimensionality reduction may reduce the molecular descriptor matrix by selecting representative feature vectors. The selected representative feature vectors may be selected from clusters of similar feature vectors in the molecular descriptor matrix. In one or more examples, each cluster may have a representative feature vector. A most predictive feature vector of the representative feature vectors may be determined. The most predictive feature vector may then be used to generate a predicted molecular binding profile. In one or more examples, the predicted molecular binding profile may represent a first principal component.
[0170] As used herein, "or" is inclusive and not exclusive, unless expressly stated otherwise or clear from the context. Thus, as used herein, "A or B" means "A, B, or both," unless expressly stated otherwise or indicated otherwise by the context. Moreover, "and" is both jointly and severally, unless expressly stated otherwise or indicated otherwise by the context. Thus, as used herein, "A and B" means "A and B jointly or severally," unless expressly stated otherwise or clear from the context.
[0171] As used herein, "automatically" and its derivatives mean "without human intervention" unless expressly indicated otherwise or indicated otherwise by context.
[0172] The embodiments disclosed herein are merely examples, and the scope of the disclosure is not limited thereto. Embodiments according to the present disclosure are disclosed in the appended claims, particularly those directed to methods, storage media, systems, and computer program products. Any feature recited in one claim category, e.g., a method, may also be claimed in another claim category, e.g., a system. Dependencies or references in the appended claims are chosen for formality reasons only. However, just as any combination of a claim and its features may be disclosed and claimed without regard to the dependencies recited in the appended claims, any subject matter resulting from an intentional reference to any preceding claim (e.g., multiple dependencies) may likewise be claimed. Subject matter that may be claimed includes not only combinations of features as recited in the appended claims, but also any other combinations of features within the scope of the claims, and each feature recited in a claim may be combined with any other feature or combination of features within the scope of the claim. Furthermore, any of the embodiments and features described or illustrated in this specification may be claimed in a separate claim and / or in any combination with any of the embodiments or features described or illustrated in this specification or with any of the features of the accompanying claims.
[0173] The scope of the present disclosure encompasses all changes, substitutions, variations, changes, and modifications to the exemplary embodiments described or illustrated herein that would be understood by one skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although the present disclosure describes and illustrates each embodiment herein as including particular components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by one skilled in the art. Furthermore, references in the appended claims to a device or system or a component of a device or system that is arranged, arranged, enabled, configured, enabled, operable, or operates to perform a particular function encompass that device, system, or component, so long as it is so arranged, arranged, enabled, configured, enabled, operable, or operates, regardless of whether it or that particular function is activated, turned on, or released. Furthermore, although this disclosure may describe or illustrate particular embodiments as providing certain advantages, the particular embodiments may provide none, some, or all of these advantages.
[0174] Illustrative Embodiments Embodiments disclosed herein may include the following. 1. A method for predicting molecular binding properties of one or more proteins, the method comprising: accessing, by one or more computing devices, a molecular descriptor matrix representing a set of amino acid sequences corresponding to the one or more proteins; refining a set of hyperparameters associated with a machine learning model trained to generate predictions of the molecular binding properties of the one or more proteins, wherein refining the set of hyperparameters includes iteratively performing the process until a desired accuracy is reached, the process including: reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, wherein each feature vector cluster includes similar feature vectors; determining one or more most predictive feature vectors among the selected representative feature vectors for each feature vector cluster based on correlations between the selected representative feature vectors and predetermined batch binding data related to the one or more proteins; calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch binding data; and updating the set of hyperparameters based on the one or more cross-validation losses; and outputting, by the machine learning model, predictions of the molecular binding properties of the one or more proteins based at least in part on the updated set of hyperparameters. 2. The method of embodiment 1, wherein calculating the one or more cross-validation losses further comprises: evaluating a cross-validation loss function based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and a set of learnable parameters associated with the machine learning model; and minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the predetermined batch combined data, and the set of hyperparameters remain constant. 3. The method of embodiment 2, wherein minimizing the cross-validation loss function includes optimizing a set of hyperparameters, the set of hyperparameters including one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters. 4. The method of embodiment 2 or 3, wherein minimizing the cross-validation loss function comprises minimizing the loss between predictions of percent protein binding for the one or more proteins and experimentally determined percent protein binding for the one or more proteins. 5. The method of any one of embodiments 2 to 4, wherein the predetermined batch binding data comprises experimentally determined percent protein binding for one or more pH values and salt concentrations related to molecular binding properties of the one or more proteins. 6. The method of any one of embodiments 2 to 5, wherein the set of learnable parameters includes one or more weights or decision variables determined by the machine learning model based at least in part on one or more most predictive feature vectors and a given batch of combined data. 7. The method of any one of embodiments 1 to 6, further comprising: after refining the set of hyperparameters, accessing a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins; reducing the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix; determining one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data related to the one or more second proteins; inputting the one or more second most predictive feature vectors into a machine learning model trained to generate predictions of molecular binding properties of the one or more second proteins; and outputting, by the machine learning model, predictions of molecular binding properties of the one or more second proteins based at least in part on the updated set of hyperparameters. 8. The method of embodiment 7, wherein predicting the molecular binding properties of the one or more second proteins comprises predicting percent protein binding for the one or more second proteins. 9. The method of any one of embodiments 1 to 8, wherein the updated set of hyperparameters includes one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters. 10. The method of any one of embodiments 1 to 9, wherein calculating one or more cross-validation losses includes calculating n cross-validation losses, where n comprises an integer between 1 and n. 11. The method of any one of embodiments 1 to 10, wherein calculating one or more cross-validation losses comprises determining n individual train-test splits based on one or more most predictive feature vectors and the given batch combined data, where n comprises an integer between 1 and n. 12. The method of any one of embodiments 1 to 11, wherein calculating one or more cross-validation losses comprises calculating n cross-validation losses, and the method further comprises generating a prediction of a molecular binding property of the one or more proteins based on an average of the n cross-validation losses. 13. Molecular descriptor matrix is 2 n 13. The method of any one of embodiments 1 to 12, wherein n comprises feature vectors, and n comprises the dimension of the molecular descriptor matrix. 14. The method of any one of embodiments 1 to 13, wherein the molecular descriptor matrix is generated by a first machine learning model that is different from the machine learning model. 15. The method of embodiment 14, wherein the first machine learning model is trained to generate a molecular descriptor matrix based on a set of amino acid sequences. 16. The method of embodiment 14 or 15, wherein the first machine learning model comprises a neural network trained to generate an M×N descriptor matrix representing a set of amino acid sequences, where N comprises the number of sets of amino acid sequences and M comprises the number of nodes in the output layer of the neural network. 17. The method of any one of embodiments 1 to 16, wherein the machine learning model comprises one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a categorical boosting (CatBoost) model. 18. The method of any one of embodiments 1 to 17, wherein the machine learning model is further trained to generate predictions of molecular elution properties of one or more proteins. 19. The method of any one of embodiments 1 to 18, wherein the machine learning model is further trained to generate predictions of flow-through properties of one or more proteins. 20. A method according to any one of embodiments 1 to 19, wherein reducing the molecular descriptor matrix comprises performing a Pearson correlation of the feature vectors of the molecular descriptor matrix to generate a plurality of feature vector clusters. 21. The method of embodiment 20, wherein the selected one representative feature vector in each of the plurality of feature vector clusters includes a centroid feature vector for each of the plurality of feature vector clusters that is utilized to represent two or more of the similar feature vectors. 22. A method according to any one of embodiments 1 to 21, wherein determining one or more representative feature vectors among the selected representative feature vectors in each of the plurality of feature vector clusters includes selecting the k-best matrix of feature vectors among the selected representative feature vectors in each of the plurality of feature vector clusters. 23. The method of embodiment 22, wherein the k-best matrix of feature vectors of the selected representative feature vectors is determined based on a predetermined k-best process. 24. A method according to any one of embodiments 1 to 23, wherein the correlation between the selected representative feature vector and the predetermined batch combined data is determined based on a maximum information coefficient (MIC) between the selected representative feature vector and the predetermined batch combined data. 25. The method of any one of embodiments 1 to 24, wherein the prediction of the molecular binding properties of one or more proteins comprises a chromatographic process based on a computational model. 26. The method of embodiment 25, wherein the computational model-based chromatography process comprises one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process. 27. The method of any one of embodiments 1 to 26, further comprising optimizing the machine learning model based on a Bayesian model optimization process. 28. The method of embodiment 27, further comprising: using group K-fold cross-validation to train and evaluate an optimized machine learning model based on one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and the set of learnable parameters. 29. The method of any one of embodiments 1 to 28, wherein predicting the molecular binding properties of the one or more proteins comprises identifying target proteins of the one or more proteins. 30. The method of any one of embodiments 1 to 29, wherein predicting the molecular binding properties of the one or more proteins comprises quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of the one or more proteins. 31. The method of any one of embodiments 1 to 30, wherein predicting the molecular binding properties of the one or more proteins comprises predicting the molecular binding properties for each amino acid sequence of a set of amino acid sequences corresponding to the one or more proteins. 32. A method for predicting molecular binding properties of one or more proteins, the method comprising: accessing, with one or more computing devices, a molecular descriptor matrix representing a set of amino acid sequences corresponding to the one or more proteins; obtaining, with a machine learning model, a prediction of the molecular binding properties of the one or more proteins based at least in part on the molecular descriptor matrix, the machine learning model being trained by accessing a training molecular descriptor matrix representing a training set of amino acid sequences corresponding to one or more empirically evaluated proteins; and iteratively performing a process of refining a set of hyperparameters associated with the machine learning model until a desired accuracy is reached, the process comprising: reducing the training molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster comprising similar feature vectors; determining one or more most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on correlations between the selected representative feature vectors and predetermined batch binding data associated with the one or more empirically evaluated proteins; calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch binding data; and updating the set of hyperparameters based on the one or more cross-validation losses. 33. The method of embodiment 32, wherein obtaining a prediction includes reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters of the molecular descriptor matrix, determining one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster based on correlation between the selected representative feature vectors and predetermined batch binding data related to one or more proteins, and inputting the one or more most predictive feature vectors into a machine learning model to obtain a prediction of molecular binding properties of the one or more proteins. 34. The method of any one of embodiments 32 to 33, wherein predicting the molecular binding properties of one or more proteins comprises predicting percent protein binding for one or more proteins. 35. The method of any one of embodiments 32 to 34, wherein calculating one or more cross-validation losses further comprises: evaluating a cross-validation loss function based on one or more most predictive feature vectors, a predetermined batch of combined data, a set of hyperparameters, and a set of learnable parameters associated with the machine learning model; and minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the predetermined batch of combined data, and the set of hyperparameters remain constant. 36. The method of embodiment 35, wherein minimizing the cross-validation loss function includes optimizing a set of hyperparameters, the set of hyperparameters including one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters. 37. The method of embodiment 35 or 36, wherein minimizing the cross-validation loss function comprises minimizing the loss between predictions of percent protein binding for the one or more proteins and experimentally determined percent protein binding for the one or more proteins. 38. The method of any one of embodiments 35 to 37, wherein the predetermined batch binding data comprises experimentally determined percent protein binding for one or more pH values and salt concentrations related to the molecular binding properties of the one or more proteins. 39. A method according to any one of embodiments 35 to 38, wherein the set of learnable parameters includes one or more weights or decision variables determined by a machine learning model based at least in part on one or more most predictive feature vectors and a given batch of combined data. 40. The method of any one of embodiments 32 to 39, wherein the molecular descriptor matrix comprises a first molecular descriptor matrix representing a first set of amino acid sequences corresponding to one or more first proteins, and wherein predicting the molecular binding properties comprises a first prediction of the molecular binding properties of the one or more first proteins, and the method further comprises accessing a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins, and obtaining, by a machine learning model, a second prediction of the molecular binding properties of the one or more second proteins based at least in part on the second molecular descriptor matrix. 41. The method of embodiment 40, wherein the machine learning model is trained to reduce the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix, determine one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data related to one or more second proteins, and input the one or more second most predictive feature vectors into the machine learning model trained to generate a second prediction. 42. The method of embodiment 40 or 41, wherein the second prediction of the molecular binding properties of the one or more second proteins comprises a prediction of percent protein binding for the one or more second proteins. 43. A method according to any one of embodiments 32 to 42, wherein the updated set of hyperparameters includes one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters. 44. The method of any one of embodiments 32 to 43, wherein the machine learning model used to generate a prediction of the molecular binding properties of one or more proteins comprises an updated set of hyperparameters. 45. The method of any one of embodiments 32 to 44, wherein calculating one or more cross-validation losses includes calculating n cross-validation losses, where n comprises an integer between 1 and n. 46. The method of any one of embodiments 32 to 45, wherein calculating one or more cross-validation losses includes determining n individual train-test splits based on one or more most predictive feature vectors and the given batch combined data, where n comprises an integer between 1 and n. 47. The method of any one of embodiments 32 to 46, wherein calculating one or more cross-validation losses comprises calculating n cross-validation losses, and the method further comprises generating a prediction of a molecular binding property of the one or more proteins based on an average of the n cross-validation losses. 48. The method of any one of embodiments 32 to 48, wherein the molecular descriptor matrix is generated by a first machine learning model different from the machine learning model. 49. The method of embodiment 48, wherein the first machine learning model is trained to generate a molecular descriptor matrix based on a set of amino acid sequences. 50. The method of embodiment 49, wherein the first machine learning model comprises a neural network trained to generate an M×N descriptor matrix representing the set of amino acid sequences. 51. The method of embodiment 49 or 50, wherein N comprises the number of sets of amino acid sequences and M comprises the number of nodes in the output layer of the neural network. 52. The method of any one of embodiments 32 to 51, wherein the machine learning model includes one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a categorical boosting (CatBoost) model. 53. The method of any one of embodiments 32 to 52, wherein the machine learning model is further trained to generate predictions of molecular elution properties of one or more proteins. 54. The method of any one of embodiments 32 to 53, wherein the machine learning model is further trained to generate predictions of flow-through properties of one or more proteins. 55. A method according to any one of embodiments 32 to 54, wherein reducing the molecular descriptor matrix comprises clustering similar feature vectors into multiple feature vector clusters based on correlation distance. 56. The method of embodiment 55, wherein the correlation distance is calculated using Pearson's correlation. 57. The method of embodiment 55 or 56, wherein the selected representative feature vector in each of the plurality of feature vector clusters includes a centroid feature vector for each of the plurality of feature vector clusters that is utilized to represent two or more of the similar feature vectors. 58. A method according to any one of embodiments 32 to 57, wherein determining one or more representative feature vectors among the selected representative feature vectors in each of the plurality of feature vector clusters includes selecting the k-best matrix of feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters. 59. The method of embodiment 58, wherein the k-best matrix of feature vectors of the selected representative feature vectors is determined based on a predetermined k-best process. 60. A method according to any one of embodiments 32 to 59, wherein the correlation between the selected representative feature vector and the predetermined batch combined data is determined based on a maximum information coefficient (MIC) between the selected representative feature vector and the predetermined batch combined data. 61. The method of any one of embodiments 32 to 60, wherein the prediction of the molecular binding properties of one or more proteins comprises a chromatographic process based on a computational model. 62. The method of embodiment 61, wherein the computational model-based chromatography process comprises one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process. 63. The method of any one of embodiments 32 to 62, further comprising optimizing the machine learning model based on a Bayesian model optimization process. 64. The method of embodiment 63, further comprising: using group K-fold cross-validation to train and evaluate an optimized machine learning model based on one or more most predictive feature vectors, a given batch of combined data, a set of hyperparameters, and a set of learnable parameters. 65. The method of embodiment 63 or 64, further comprising using stratified K-fold cross-validation to train and evaluate an optimized machine learning model based on one or more of the most predictive feature vectors, a predetermined batch of combined data, a set of hyperparameters, and a set of learnable parameters. 66. The method of any one of embodiments 32 to 65, wherein predicting the molecular binding properties of one or more proteins comprises identifying target proteins of the one or more proteins. 67. The method of any one of embodiments 32 to 66, wherein predicting the molecular binding properties of one or more proteins comprises quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of one or more proteins. 68. The method of any one of embodiments 32 to 67, wherein predicting the molecular binding properties of one or more proteins comprises predicting the molecular binding properties for each amino acid sequence of a set of amino acid sequences corresponding to one or more proteins. 69. The method of any one of embodiments 32 to 68, wherein for each of the one or more empirically evaluated proteins, a corresponding predetermined batch binding is measured for each of a set of experimental conditions. 70. The method of embodiment 69, wherein the set of experimental conditions comprises 24 experimental conditions. 71. The method of embodiment 70, wherein the set of experimental conditions comprises a first subset of salt concentrations and a second subset of pH values. 72. The method of embodiment 70 or 71, wherein the set of experimental conditions is input into the machine learning model along with a molecular descriptor matrix, and the prediction of molecular binding properties of the one or more proteins includes predicting molecular binding properties of the one or more proteins for each of the set of experimental conditions. 73. The method of any one of embodiments 32 to 72, further comprising converting the prediction of the molecular binding properties of one or more proteins into a linear representation. 74. The method of embodiment 73, wherein a logit transform is used to generate the linear representation. 75. The method of embodiment 73 or 74, further comprising performing principal component analysis (PCA) on the linear representation to obtain at least a first principal component. 76. The method of any one of embodiments 32 to 75, wherein the predetermined batch binding data associated with the one or more empirically evaluated proteins comprises experimentally determined binding values measured for each of a set of experimental conditions for each of the one or more empirically evaluated proteins. 77. The method of embodiment 76, wherein the correlation between the selected representative feature vector and the given batch binding data comprises: generating, for each of the one or more empirically evaluated proteins and for each of the set of experimental conditions, a linear representation of the experimentally determined binding values of the empirically evaluated proteins based on a logit transform applied to the experimentally determined binding values of the empirically evaluated proteins; and performing principal component analysis (PCA) on the linear representation of the experimentally determined binding values of the one or more empirically evaluated proteins to obtain at least a first principal component. 78. The method of embodiment 77, further comprising using the machine learning model to generate training predictions of molecular binding properties of one or more empirically evaluated proteins, and comparing the training predictions to the first principal component to calculate one or more cross-validation losses. 79. The method of embodiment 77 or 78, wherein the first principal component represents the average batch binding value. 80. The method of any one of embodiments 32 to 79, further comprising generating a set of functions representing the behavior of the one or more proteins with respect to the set of experimental conditions based on the predictions, and selecting at least one of the one or more proteins for one or more drug discovery assays based on the behavior of the one or more proteins with respect to the set of experimental conditions. 81. The method of any one of embodiments 32 to 80, wherein the correlation between the selected representative feature vectors and predetermined batch binding data associated with one or more empirically evaluated proteins comprises a correlation between the representative feature vectors and principal components calculated based on the predetermined batch binding data. 82. The method of embodiment 81, wherein one or more cross-validation losses are calculated based on predicted and empirical molecular binding properties. 83. The method of embodiment 82, wherein the predicted molecular binding properties include principal components calculated based on representative feature vectors, and the empirical molecular binding properties include principal components calculated based on predetermined batch binding data. 84. The method of any one of embodiments 32 to 33, wherein determining one or more most predictive feature vectors further comprises: (i) fitting a model to the representative feature vectors; (ii) calculating a feature importance score for each representative feature vector based on the model; and (iii) removing one or more feature vectors from the representative feature vectors based on the respective feature importance scores of the representative feature vectors to obtain a subset of representative feature vectors, wherein the one or more most predictive feature vectors include one or more feature vectors from the subset of representative feature vectors. 85. The method of embodiment 84, further comprising iteratively performing steps (i) to (iii) until the number of feature vectors included in the subset meets the feature criterion. 86. The method of embodiment 85, wherein the feature criterion being met includes the number of feature vectors included in the subset of representative feature vectors being less than or equal to a threshold number of feature vectors. 87. The method of embodiment 86, wherein the threshold number of feature vectors includes the same or similar number of features from the training data used to train the machine learning model. 88. The method of any one of embodiments 84 to 88, wherein the number of feature vectors included in the subset of representative feature vectors comprises one of the set of hyperparameters. 89. The method of any one of embodiments 32 to 88, wherein one of the sets of hyperparameters represents the number of feature vector clusters included in the plurality of feature vector clusters. 90. A system including one or more computing devices, further comprising: one or more non-transitory computer-readable storage media containing instructions; and one or more processors coupled to the one or more storage media, wherein the one or more processors are configured to execute the instructions to perform the method of any one of embodiments 1 to 89. 91. A non-transitory computer-readable medium containing instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to perform operations including a method described in any one of embodiments 1 to 89.
Claims
1. 1. A method for predicting molecular binding properties of one or more proteins, comprising, by one or more computing devices: accessing a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins; refining a set of hyperparameters associated with a machine learning model trained to generate predictions of molecular binding properties of the one or more proteins, wherein refining the set of hyperparameters is performed iteratively until a desired accuracy is reached. wherein the process comprises: reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster containing similar feature vectors; determining one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster based on correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more proteins; calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data; updating the set of hyperparameters based on the one or more cross-validation losses; elaborating, including outputting, by the machine learning model, the prediction of the molecular binding properties of the one or more proteins based at least in part on the updated set of hyperparameters; A method comprising:
2. Calculating the one or more cross-validation losses includes: Evaluating a cross-validation loss function based on the one or more most predictive feature vectors, the predetermined batch of combined data, the set of hyperparameters, and a set of learnable parameters associated with the machine learning model; minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the given batch combined data, and the set of hyperparameters remain constant; The method of claim 1 further comprising:
3. 3. The method of claim 2, wherein minimizing the cross-validation loss function comprises optimizing the set of hyperparameters, the set of hyperparameters comprising one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters.
4. 3. The method of claim 2, wherein minimizing the cross-validation loss function comprises minimizing the loss between a prediction of percent protein binding for the one or more proteins and an experimentally determined percent protein binding for the one or more proteins.
5. 3. The method of claim 2, wherein the predetermined batch binding data comprises experimentally determined percent protein binding for one or more pH values and salt concentrations associated with the molecular binding properties of the one or more proteins.
6. 3. The method of claim 2, wherein the set of learnable parameters includes one or more weights or decision variables determined by the machine learning model based at least in part on the one or more most predictive feature vectors and the predetermined batch of combined data.
7. After refining the set of hyperparameters, accessing a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins; reducing the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix; determining one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more second proteins; inputting the one or more second most predictive feature vectors into the machine learning model trained to generate predictions of molecular binding properties of the one or more second proteins; outputting, by the machine learning model, the prediction of the molecular binding property of the one or more second proteins based at least in part on the updated set of hyperparameters; The method of claim 1 further comprising:
8. 8. The method of claim 7, wherein the prediction of the molecular binding properties of the one or more second proteins comprises a prediction of percent protein binding for the one or more second proteins.
9. 2. The method of claim 1 , wherein the updated set of hyperparameters comprises one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters.
10. The method of claim 1 , wherein computing one or more cross-validation losses comprises computing n cross-validation losses, where n comprises an integer from 1 to n.
11. 2. The method of claim 1 , wherein calculating the one or more cross-validation losses comprises determining n individual train-test splits based on the one or more most predictive feature vectors and the given batch combined data, where n comprises an integer from 1 to n.
12. Calculating the one or more cross-validation losses includes calculating n cross-validation losses, and the method further comprises: generating the prediction of the molecular binding property of the one or more proteins based on an average of the n cross-validation losses; The method of claim 1 further comprising:
13. The molecular descriptor matrix is n 10. The method of claim 1, wherein n comprises feature vectors, and n comprises the dimension of the molecular descriptor matrix.
14. The method of claim 1 , wherein the molecular descriptor matrix was generated by a first machine learning model that is different from the machine learning model.
15. 15. The method of claim 14, wherein the first machine learning model is trained to generate the molecular descriptor matrix based on the set of amino acid sequences.
16. 15. The method of claim 14, wherein the first machine learning model comprises a neural network trained to generate an M×N descriptor matrix representing the set of amino acid sequences, where N comprises the number of sets of amino acid sequences and M comprises the number of nodes in an output layer of the neural network.
17. 2. The method of claim 1, wherein the machine learning model comprises one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a category boosting (CatBoost) model.
18. 10. The method of claim 1, wherein the machine learning model is further trained to generate predictions of molecular elution properties of the one or more proteins.
19. 10. The method of claim 1, wherein the machine learning model is further trained to generate a prediction of flow-through properties of the one or more proteins.
20. The method of claim 1 , wherein reducing the molecular descriptor matrix comprises performing a Pearson correlation of feature vectors of the molecular descriptor matrix to generate the plurality of feature vector clusters.
21. 21. The method of claim 20, wherein the selected representative feature vector in each of the plurality of feature vector clusters comprises a centroid feature vector for each of the plurality of feature vector clusters that is utilized to represent two or more of the similar feature vectors.
22. 2. The method of claim 1 , wherein determining the one or more representative feature vectors from the selected representative feature vectors in each of the plurality of feature vector clusters comprises selecting a k-best matrix of feature vectors from the selected representative feature vectors in each of the plurality of feature vector clusters.
23. The method of claim 22, wherein the k-best matrix of feature vectors of the selected representative feature vectors is determined based on a predetermined k-best process.
24. 2. The method of claim 1, wherein the correlation between the selected representative feature vector and the given batch combined data is determined based on a maximum information coefficient (MIC) between the selected representative feature vector and the given batch combined data.
25. 10. The method of claim 1, wherein the prediction of the molecular binding properties of the one or more proteins comprises a chromatographic process based on a computational model.
26. 26. The method of claim 25, wherein the computational model-based chromatography process comprises one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process.
27. The method of claim 1 , further comprising optimizing the machine learning model based on a Bayesian model optimization process.
28. 28. The method of claim 27, further comprising: training and evaluating an optimized machine learning model based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and the set of learnable parameters using group K-fold cross-validation.
29. 10. The method of claim 1, wherein the prediction of the molecular binding properties of the one or more proteins comprises identification of target proteins of the one or more proteins.
30. 10. The method of claim 1, wherein the prediction of the molecular binding properties of the one or more proteins comprises quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of the one or more proteins.
31. 2. The method of claim 1, wherein the prediction of the molecular binding properties of the one or more proteins comprises predicting a molecular binding property for each amino acid sequence of the set of amino acid sequences corresponding to the one or more proteins.
32. 1. A system including one or more computing devices, one or more non-transitory computer-readable storage media containing instructions; one or more processors coupled to the one or more storage media; and wherein the one or more processors further comprise: instructions to access a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins; and instructions for refining a set of hyperparameters associated with a machine learning model trained to generate predictions of molecular binding properties of the one or more proteins, wherein refining the set of hyperparameters comprises iteratively performing a process until a desired accuracy is reached, the process comprising: instructions for reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster containing similar feature vectors; instructions for determining one or more representative feature vectors of the selected representative feature vectors for each of a plurality of feature vector clusters based on a correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more proteins; instructions for calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data; and instructions for updating the set of hyperparameters based on the one or more cross-validation losses Refining instructions, including: instructions for outputting, by the machine learning model, the prediction of the molecular binding properties of the one or more proteins based at least in part on the updated set of hyperparameters. The system is configured to run
33. The instructions to calculate the one or more cross-validation losses include: instructions for evaluating a cross-validation loss function based on the one or more most predictive feature vectors, the predetermined batch of combined data, the set of hyperparameters, and a set of learnable parameters associated with the machine learning model; instructions for minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the given batch of combined data, and the set of hyperparameters remain constant; 33. The system of claim 32, further comprising:
34. 34. The system of claim 33, wherein the instructions to minimize the cross-validation loss function comprise instructions to optimize the set of hyperparameters, the set of hyperparameters comprising one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters.
35. 34. The system of claim 33, wherein the instructions for minimizing the cross-validation loss function comprise instructions for minimizing the loss between a prediction of percent protein binding for the one or more proteins and an experimentally determined percent protein binding for the one or more proteins.
36. 34. The system of claim 33, wherein the predetermined batch binding data comprises experimentally determined percent protein binding for one or more pH values and salt concentrations associated with the molecular binding properties of the one or more proteins.
37. 34. The system of claim 33, wherein the set of learnable parameters includes one or more weights or decision variables determined by the machine learning model based at least in part on the one or more most predictive feature vectors and the predetermined batch of combined data.
38. The instruction: After refining the set of hyperparameters, instructions to access a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins; instructions for reducing the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix; instructions for determining one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more second proteins; inputting the one or more second most predictive feature vectors into the machine learning model trained to generate a prediction of a molecular binding property of the one or more second proteins; and instructions for outputting, by the machine learning model, the prediction of the molecular binding property of the one or more second proteins based at least in part on the updated set of hyperparameters; 33. The system of claim 32, further comprising:
39. 39. The system of claim 38, wherein the prediction of the molecular binding properties of the one or more second proteins comprises a prediction of percent protein binding for the one or more second proteins.
40. 33. The system of claim 32, wherein the updated set of hyperparameters comprises one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters.
41. 33. The system of claim 32, wherein the instructions to compute one or more cross-validation losses further comprise instructions to compute n cross-validation losses, where n comprises an integer from 1 to n.
42. 33. The system of claim 32, wherein the instructions to calculate the one or more cross-validation losses further comprise instructions to determine n individual train-test splits based on the one or more most predictive feature vectors and the predetermined batch combined data, where n comprises an integer from 1 to n.
43. The instructions to calculate one or more cross-validation losses further include instructions to calculate n cross-validation losses, the instructions comprising: instructions for generating the prediction of the molecular binding property of the one or more proteins based on an average of the n cross-validation losses 33. The system of claim 32, further comprising:
44. The molecular descriptor matrix is n 33. The system of claim 32, wherein n comprises feature vectors, and n comprises the dimension of the molecular descriptor matrix.
45. 33. The system of claim 32, wherein the molecular descriptor matrix was generated by a first machine learning model that is different from the machine learning model.
46. 46. The system of claim 45, wherein the first machine learning model is trained to generate the molecular descriptor matrix based on the set of amino acid sequences.
47. 46. The system of claim 45, wherein the first machine learning model comprises a neural network trained to generate an M×N descriptor matrix representing the set of amino acid sequences, where N comprises the number of sets of amino acid sequences and M comprises the number of nodes in an output layer of the neural network.
48. 33. The system of claim 32, wherein the machine learning model comprises one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a category boosting (CatBoost) model.
49. 33. The system of claim 32, wherein the machine learning model is further trained to generate a prediction of the molecular elution properties of the one or more proteins.
50. 33. The system of claim 32, wherein the machine learning model is further trained to generate a prediction of flow-through properties of the one or more proteins.
51. 33. The system of claim 32, wherein the instructions for reducing the molecular descriptor matrix further comprise instructions for performing a Pearson correlation of feature vectors of the molecular descriptor matrix to generate the plurality of feature vector clusters.
52. 52. The system of claim 51 , wherein the selected one representative feature vector in each of the plurality of feature vector clusters comprises a centroid feature vector for each of the plurality of feature vector clusters that is utilized to represent two or more of the similar feature vectors.
53. 33. The system of claim 32, wherein the instructions for determining the one or more representative feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters further comprise instructions for selecting a k-best matrix of feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters.
54. 54. The system of claim 53, wherein the k-best matrix of feature vectors for the selected representative feature vectors is determined based on a predetermined k-best process.
55. 33. The system of claim 32, wherein the correlation between the selected representative feature vectors and the given batch combined data is determined based on a maximum information coefficient (MIC) between the selected representative feature vectors and the given batch combined data.
56. 33. The system of claim 32, wherein the prediction of the molecular binding properties of the one or more proteins comprises a chromatographic process based on a computational model.
57. 57. The system of claim 56, wherein the computational model-based chromatography process comprises one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process.
58. 33. The system of claim 32, wherein the instructions further comprise instructions for optimizing the machine learning model based on a Bayesian model optimization process.
59. 59. The system of claim 58, wherein the instructions further comprise instructions for training and evaluating an optimized machine learning model based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and the set of learnable parameters using group K-fold cross-validation.
60. 33. The system of claim 32, wherein the prediction of the molecular binding properties of the one or more proteins comprises identification of target proteins of the one or more proteins.
61. 33. The system of claim 32, wherein the prediction of the molecular binding properties of the one or more proteins comprises quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of the one or more proteins.
62. 33. The system of claim 32, wherein the prediction of the molecular binding properties of the one or more proteins comprises a prediction of a molecular binding property for each amino acid sequence of the set of amino acid sequences corresponding to the one or more proteins.
63. A non-transitory computer-readable medium containing instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: accessing a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins; refining a set of hyperparameters associated with a machine learning model trained to generate predictions of molecular binding properties of the one or more proteins, wherein refining the set of hyperparameters is performed iteratively until a desired accuracy is reached. wherein the process comprises: instructions for reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster containing similar feature vectors; instructions for determining one or more representative feature vectors of the selected representative feature vectors for each of a plurality of feature vector clusters based on a correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more proteins; instructions for calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data; and instructions for updating the set of hyperparameters based on the one or more cross-validation losses; elaborating, including outputting, by the machine learning model, the prediction of the molecular binding properties of the one or more proteins based at least in part on the updated set of hyperparameters; A non-transitory computer-readable medium for causing
64. The instructions to calculate the one or more cross-validation losses include: instructions for evaluating a cross-validation loss function based on the one or more most predictive feature vectors, the predetermined batch of combined data, the set of hyperparameters, and a set of learnable parameters associated with the machine learning model; instructions for minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the given batch of combined data, and the set of hyperparameters remain constant; 64. The non-transitory computer-readable medium of claim 63, further comprising:
65. 65. The non-transitory computer-readable medium of claim 64, wherein the instructions to minimize the cross-validation loss function comprise instructions to optimize the set of hyperparameters, wherein the set of hyperparameters comprises one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters.
66. 65. The non-transitory computer-readable medium of Claim 64, wherein the instructions for minimizing the cross-validation loss function comprise instructions for minimizing the loss between a prediction of percent protein binding for the one or more proteins and an experimentally determined percent protein binding for the one or more proteins.
67. 65. The non-transitory computer-readable medium of claim 64, wherein the predetermined batch binding data comprises experimentally determined percent protein binding for one or more pH values and salt concentrations associated with the molecular binding properties of the one or more proteins.
68. 65. The non-transitory computer-readable medium of claim 64, wherein the set of learnable parameters includes one or more weights or decision variables determined by the machine learning model based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data.
69. The instruction: After refining the set of hyperparameters, instructions to access a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins; instructions for reducing the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix; instructions for determining one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more second proteins; inputting the one or more second most predictive feature vectors into the machine learning model trained to generate a prediction of a molecular binding property of the one or more second proteins; and instructions for outputting, by the machine learning model, the prediction of the molecular binding property of the one or more second proteins based at least in part on the updated set of hyperparameters; 64. The non-transitory computer-readable medium of claim 63, further comprising:
70. 69. The non-transitory computer-readable medium of Claim 68, wherein the prediction of the molecular binding properties of the one or more second proteins comprises a prediction of a second protein binding percentage for the one or more second proteins.
71. 64. The non-transitory computer-readable medium of claim 63, wherein the updated set of hyperparameters comprises one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters.
72. 64. The non-transitory computer-readable medium of claim 63, wherein the instructions to compute one or more cross-validation losses further comprise instructions to compute n cross-validation losses, where n comprises an integer from 1 to n.
73. 64. The non-transitory computer-readable medium of claim 63, wherein the instructions to calculate the one or more cross-validation losses further comprise instructions to determine n individual train-test splits based on the one or more most predictive feature vectors and the predetermined batch combined data, where n comprises an integer from 1 to n.
74. The instructions to calculate one or more cross-validation losses further include instructions to calculate n cross-validation losses, the instructions comprising: instructions for generating the prediction of the molecular binding property of the one or more proteins based on an average of the n cross-validation losses 64. The non-transitory computer-readable medium of claim 63, further comprising:
75. The molecular descriptor matrix is n 64. The non-transitory computer-readable medium of claim 63, comprising n feature vectors, wherein n comprises a dimension of the molecular descriptor matrix.
76. 64. The non-transitory computer-readable medium of claim 63, wherein the molecular descriptor matrix was generated by a first machine learning model that is different from the machine learning model.
77. 77. The non-transitory computer-readable medium of Claim 76, wherein the first machine learning model is trained to generate the molecular descriptor matrix based on the set of amino acid sequences.
78. 77. The non-transitory computer-readable medium of claim 76, wherein the first machine learning model comprises a neural network trained to generate an M×N descriptor matrix representing the set of amino acid sequences, where N comprises the number of sets of amino acid sequences and M comprises the number of nodes in an output layer of the neural network.
79. 64. The non-transitory computer-readable medium of claim 63, wherein the machine learning model comprises one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a category boosting (CatBoost) model.
80. 64. The non-transitory computer-readable medium of claim 63, wherein the machine learning model is further trained to generate a prediction of molecular elution properties of the one or more proteins.
81. 64. The non-transitory computer-readable medium of claim 63, wherein the machine learning model is further trained to generate a prediction of flow-through properties of the one or more proteins.
82. 64. The non-transitory computer-readable medium of claim 63, wherein the instructions for reducing the molecular descriptor matrix further comprise instructions for performing a Pearson correlation of feature vectors of the molecular descriptor matrix to generate the plurality of feature vector clusters.
83. 83. The non-transitory computer-readable medium of claim 82, wherein the selected one representative feature vector in each of the plurality of feature vector clusters comprises a centroid feature vector for each of the plurality of feature vector clusters utilized to represent two or more of the similar feature vectors.
84. 64. The non-transitory computer-readable medium of claim 63, wherein the instructions for determining the one or more representative feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters further comprise instructions for selecting a k-best matrix of feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters.
85. 85. The non-transitory computer-readable medium of claim 84, wherein the k-best matrix of feature vectors for the selected representative feature vectors is determined based on a predetermined k-best process.
86. 64. The non-transitory computer-readable medium of claim 63, wherein the correlation between the selected representative feature vectors and the predetermined batch combined data is determined based on a maximum information coefficient (MIC) between the selected representative feature vectors and the predetermined batch combined data.
87. 64. The non-transitory computer-readable medium of claim 63, wherein the prediction of the molecular binding properties of the one or more proteins comprises a column chromatography process based on a computational model.
88. 88. The non-transitory computer-readable medium of claim 87, wherein the computational model-based chromatography process comprises one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process.
89. 64. The non-transitory computer-readable medium of claim 63, wherein the instructions further comprise instructions for optimizing the machine learning model based on a Bayesian model optimization process.
90. 90. The non-transitory computer-readable medium of claim 89, wherein the instructions further comprise instructions for training and evaluating an optimized machine learning model based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and a set of learnable parameters using group K-fold cross-validation.
91. 64. The non-transitory computer-readable medium of Claim 63, wherein the prediction of the molecular binding properties of the one or more proteins comprises identification of a target protein of the one or more proteins.
92. 64. The non-transitory computer-readable medium of claim 63, wherein the prediction of the molecular binding properties of the one or more proteins comprises quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of the one or more proteins.
93. 64. The non-transitory computer-readable medium of Claim 63, wherein the prediction of the molecular binding properties of the one or more proteins comprises predicting a molecular binding property for each amino acid sequence of the set of amino acid sequences corresponding to the one or more proteins.
94. 1. A method for predicting molecular binding properties of one or more proteins, comprising, by one or more computing devices: accessing a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins; and obtaining, with a machine learning model, a prediction of molecular binding properties of the one or more proteins based at least in part on the molecular descriptor matrix. wherein the machine learning model comprises: accessing a training molecular descriptor matrix representing a training set of amino acid sequences corresponding to one or more empirically evaluated proteins; and iteratively performing a process of refining a set of hyperparameters associated with the machine learning model until a desired accuracy is reached; The process is trained by reducing the training molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster containing similar feature vectors; determining one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster based on correlations between the selected representative feature vectors and predetermined batch binding data associated with the one or more empirically evaluated proteins; calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data; updating the set of hyperparameters based on the one or more cross-validation losses; A method comprising:
95. Obtaining the prediction comprises: reducing the molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters of the molecular descriptor matrix; determining one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster based on correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more proteins; inputting the one or more most predictive feature vectors into the machine learning model to obtain the prediction of the molecular binding property of the one or more proteins; 95. The method of claim 94, comprising:
96. 95. The method of claim 94, wherein the prediction of the molecular binding properties of the one or more proteins comprises a prediction of percent protein binding for the one or more proteins.
97. Calculating the one or more cross-validation losses includes: Evaluating a cross-validation loss function based on the one or more most predictive feature vectors, the predetermined batch of combined data, the set of hyperparameters, and a set of learnable parameters associated with the machine learning model; minimizing the cross-validation loss function by varying the set of learnable parameters while the one or more most predictive feature vectors, the given batch combined data, and the set of hyperparameters remain constant; 95. The method of claim 94, further comprising:
98. 98. The method of claim 97, wherein minimizing the cross-validation loss function comprises optimizing the set of hyperparameters, wherein the set of hyperparameters comprises one or more of a set of general parameters, a set of booster parameters, or a set of learning task parameters.
99. 98. The method of Claim 97, wherein minimizing the cross-validation loss function comprises minimizing the loss between a prediction of percent protein binding for the one or more proteins and an experimentally determined percent protein binding for the one or more proteins.
100. 98. The method of claim 97, wherein the predetermined batch binding data comprises experimentally determined percent protein binding for one or more pH values and salt concentrations associated with the molecular binding properties of the one or more proteins.
101. 98. The method of claim 97, wherein the set of learnable parameters includes one or more weights or decision variables determined by the machine learning model based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data.
102. the molecular descriptor matrix comprises a first molecular descriptor matrix representing a first set of amino acid sequences corresponding to one or more first proteins, and the prediction of the molecular binding properties comprises a first prediction of molecular binding properties of the one or more first proteins, and the method further comprises: accessing a second molecular descriptor matrix representing a second set of amino acid sequences corresponding to one or more second proteins; obtaining, with the machine learning model, a second prediction of a molecular binding property of the one or more second proteins based at least in part on the second molecular descriptor matrix; and 95. The method of claim 94, further comprising:
103. The machine learning model is reducing the second molecular descriptor matrix by selecting one representative feature vector for each of a second plurality of feature vector clusters of the second molecular descriptor matrix; determining one or more second most predictive feature vectors of the selected representative feature vectors for each feature vector cluster based on a second correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more second proteins; inputting the one or more second most predictive feature vectors into the machine learning model trained to generate the second prediction; 103. The method of claim 102, wherein the subject is trained to:
104. 103. The method of claim 102, wherein the second prediction of the molecular binding properties of the one or more second proteins comprises a prediction of percent protein binding for the one or more second proteins.
105. 95. The method of claim 94, wherein the updated set of hyperparameters comprises one or more of an updated set of general parameters, an updated set of booster parameters, or an updated set of learning task parameters.
106. 95. The method of Claim 94, wherein the machine learning model used to generate the prediction of the molecular binding property of the one or more proteins comprises the updated set of hyperparameters.
107. 95. The method of claim 94, wherein computing one or more cross-validation losses comprises computing n cross-validation losses, where n comprises an integer from 1 to n.
108. 95. The method of claim 94, wherein calculating the one or more cross-validation losses comprises determining n individual train-test splits based on the one or more most predictive feature vectors and the given batch combined data, where n comprises an integer from 1 to n.
109. Calculating the one or more cross-validation losses includes calculating n cross-validation losses, and the method further comprises: generating the prediction of the molecular binding property of the one or more proteins based on an average of the n cross-validation losses.
95. The method of claim 94, further comprising:
110. 95. The method of claim 94, wherein the molecular descriptor matrix was generated by a first machine learning model that is different from the machine learning model.
111. 111. The method of claim 110, wherein the first machine learning model is trained to generate the molecular descriptor matrix based on the set of amino acid sequences.
112. 112. The method of claim 111, wherein the first machine learning model comprises a neural network trained to generate an MxN descriptor matrix representing the set of amino acid sequences.
113. 113. The method of claim 112, wherein N comprises the number of sets of amino acid sequences and M comprises the number of nodes in the output layer of the neural network.
114. 95. The method of claim 94, wherein the machine learning model comprises one or more of a gradient boosting model, an adaptive boosting (AdaBoost) model, an eXtreme gradient boosting (XGBoost) model, a light gradient boosting machine (LightGBM) model, or a category boosting (CatBoost) model.
115. 95. The method of claim 94, wherein the machine learning model is further trained to generate a prediction of the molecular elution properties of the one or more proteins.
116. 95. The method of claim 94, wherein the machine learning model is further trained to generate a prediction of flow-through properties of the one or more proteins.
117. 95. The method of claim 94, wherein reducing the molecular descriptor matrix comprises clustering the similar feature vectors into the plurality of feature vector clusters based on correlation distance.
118. 118. The method of claim 117, wherein the correlation distance is calculated using Pearson's correlation.
119. 118. The method of claim 117, wherein the selected one representative feature vector in each of the plurality of feature vector clusters comprises a centroid feature vector for each of the plurality of feature vector clusters that is utilized to represent two or more of the similar feature vectors.
120. 95. The method of claim 94, wherein determining the one or more representative feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters comprises selecting a k-best matrix of feature vectors of the selected representative feature vectors in each of the plurality of feature vector clusters.
121. 121. The method of claim 120, wherein the k-best matrix of feature vectors for the selected representative feature vectors is determined based on a predetermined k-best process.
122. 95. The method of claim 94, wherein the correlation between the selected representative feature vectors and the given batch combined data is determined based on a maximum information coefficient (MIC) between the selected representative feature vectors and the given batch combined data.
123. 95. The method of claim 94, wherein the prediction of the molecular binding properties of the one or more proteins comprises a chromatographic process based on a computational model.
124. 124. The method of claim 123, wherein the computational model-based chromatography process comprises one or more of a computational model-based affinity chromatography process, an ion exchange chromatography (IEX) process, a hydrophobic interaction chromatography (HIC) process, or a mixed-mode chromatography (MMC) process.
125. 95. The method of claim 94, further comprising optimizing the machine learning model based on a Bayesian model optimization process.
126. 126. The method of claim 125, further comprising: training and evaluating an optimized machine learning model based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and the set of learnable parameters using group K-fold cross-validation.
127. 126. The method of claim 125, further comprising: training and evaluating an optimized machine learning model based on the one or more most predictive feature vectors, the predetermined batch combined data, the set of hyperparameters, and the set of learnable parameters using stratified K-fold cross-validation.
128. 95. The method of Claim 94, wherein said prediction of said molecular binding properties of said one or more proteins comprises identification of target proteins of said one or more proteins.
129. 95. The method of claim 94, wherein the prediction of the molecular binding properties of the one or more proteins comprises quantitative structure-property relationship (QSPR) or quantitative structure-activity relationship (QSAR) modeling of the one or more proteins.
130. 95. The method of Claim 94, wherein the prediction of the molecular binding properties of the one or more proteins comprises predicting a molecular binding property for each amino acid sequence of the set of amino acid sequences corresponding to the one or more proteins.
131. 95. The method of claim 94, wherein for each of the one or more empirically evaluated proteins, a corresponding predetermined batch binding is measured for each of a set of experimental conditions.
132. 132. The method of claim 131, wherein the set of experimental conditions comprises 24 experimental conditions.
133. 132. The method of claim 131, wherein the set of experimental conditions comprises a first subset of salt concentrations and a second subset of pH values.
134. 132. The method of Claim 131, wherein the set of experimental conditions along with the molecular descriptor matrix are input into the machine learning model, and wherein the prediction of the molecular binding properties of the one or more proteins comprises a prediction of the molecular binding properties of the one or more proteins for each of the set of experimental conditions.
135. 95. The method of claim 94, further comprising converting the prediction of the molecular binding properties of the one or more proteins into a linear representation.
136. 136. The method of claim 135, wherein a logit transform is used to generate the linear representation.
137. 136. The method of claim 135, further comprising performing principal component analysis (PCA) on the linear representation to obtain at least a first principal component.
138. 95. The method of Claim 94, wherein the predetermined batch binding data associated with the one or more empirically evaluated proteins comprises experimentally determined binding values measured for each of a set of experimental conditions for each of the one or more empirically evaluated proteins.
139. The correlation between the selected representative feature vectors and the given batch combined data is For each of the one or more empirically evaluated proteins, and for each of the set of experimental conditions, generating a linear representation of the experimentally determined binding values of the empirically evaluated protein based on a logit transform applied to the experimentally determined binding values of the empirically evaluated protein; performing a principal component analysis (PCA) on the linear representation of the experimentally determined binding values of the one or more empirically evaluated proteins to obtain at least a first principal component; 139. The method of claim 138, comprising:
140. using the machine learning model to generate training predictions of molecular binding properties of the one or more empirically evaluated proteins; comparing the training predictions to the first principal component to calculate the one or more cross-validation losses; 140. The method of claim 139, further comprising:
141. 140. The method of claim 139, wherein the first principal component represents an average batch binding value.
142. generating a set of functions representing the behavior of the one or more proteins with respect to a set of experimental conditions based on the predictions; selecting at least one of the one or more proteins for one or more drug discovery assays based on the behavior of the one or more proteins with respect to the set of experimental conditions; 95. The method of claim 94, further comprising:
143. Correlation between the selected representative feature vectors and the given batch binding data associated with the one or more empirically evaluated proteins is Correlations between the representative feature vectors and principal components calculated based on the given batch of combined data.
95. The method of claim 94, comprising:
144. 144. The method of claim 143, wherein the one or more cross-validation losses are calculated based on predicted and empirical molecular binding properties.
145. 145. The method of claim 144, wherein the predicted molecular binding properties comprise principal components calculated based on the representative feature vectors, and the empirical molecular binding properties comprise principal components calculated based on the predetermined batch binding data.
146. Determining the one or more most predictive feature vectors includes: (i) fitting a model to the representative feature vector; (ii) calculating a feature importance score for each representative feature vector based on the model; (iii) removing one or more feature vectors from the representative feature vectors based on the feature importance score of each of the representative feature vectors to obtain a subset of representative feature vectors, wherein the one or more most predictive feature vectors include one or more feature vectors from the subset of representative feature vectors; 95. The method of claim 94, further comprising:
147. 147. The method of claim 146, further comprising iteratively performing steps (i) through (iii) until the number of feature vectors included in the subset satisfies a feature criterion.
148. 148. The method of claim 147, wherein the feature criterion being met comprises the number of feature vectors in the subset of representative feature vectors being less than or equal to a threshold number of feature vectors.
149. 149. The method of claim 148, wherein the threshold number of feature vectors comprises the same or similar number of features from training data used to train the machine learning model.
150. 148. The method of claim 147, wherein the number of feature vectors included in the subset of representative feature vectors comprises one of the sets of hyperparameters.
151. 95. The method of claim 94, wherein one of the sets of hyperparameters represents a number of feature vector clusters included in the plurality of feature vector clusters.
152. 1. A system including one or more computing devices, one or more non-transitory computer-readable storage media containing instructions; one or more processors coupled to the one or more storage media; and wherein the one or more processors further comprise: instructions for accessing a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins; instructions for obtaining, with a machine learning model, a prediction of a molecular binding property of the one or more proteins based at least in part on the molecular descriptor matrix. and the machine learning model is configured to: accessing a training molecular descriptor matrix representing a training set of amino acid sequences corresponding to one or more empirically evaluated proteins; reducing the training molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster containing similar feature vectors; determining one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster based on correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more empirically evaluated proteins; and Iteratively performing a process of refining a set of hyperparameters associated with the machine learning model until a desired accuracy is reached. The process is trained by calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data; updating the set of hyperparameters based on the one or more cross-validation losses; Including, the system.
153. A non-transitory computer-readable medium containing instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: accessing a molecular descriptor matrix representing a set of amino acid sequences corresponding to one or more proteins; obtaining, with a machine learning model, a prediction of molecular binding properties of the one or more proteins based at least in part on the molecular descriptor matrix; and performing an operation including: accessing a training molecular descriptor matrix representing a training set of amino acid sequences corresponding to one or more empirically evaluated proteins; reducing the training molecular descriptor matrix by selecting one representative feature vector for each of a plurality of feature vector clusters, each feature vector cluster containing similar feature vectors; determining one or more most predictive feature vectors from the selected representative feature vectors for each feature vector cluster based on correlation between the selected representative feature vectors and predetermined batch binding data associated with the one or more empirically evaluated proteins; and Iteratively performing a process of refining a set of hyperparameters associated with the machine learning model until a desired accuracy is reached. The process is trained by calculating one or more cross-validation losses based at least in part on the one or more most predictive feature vectors and the predetermined batch combined data; updating the set of hyperparameters based on the one or more cross-validation losses; 1. A non-transitory computer-readable medium comprising: