Systems and methods for identifying a ligand binding to products of a gene combination
Patent Information
- Application Number
- PCT/US2026/020152
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-21
- Filing Date
- 2026-03-20
- Publication Date
- 2026-09-24
Smart Images

Figure IMGF000052_0001_TABLE 
Figure IMGF000053_0001_TABLE 
Figure IMGF000054_0001_TABLE
Abstract
Description
Attorney Docket: 27576-2000240SYSTEMS AND METHODS FOR IDENTIFYING A LIGAND BINDING TO PRODUCTS OF A GENE COMBINATIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This patent application claims priority to U.S. Provisional Application No.63 / 775,479, filed on March 21, 2025, the contents of which are hereby incorporated herein by reference in their entirety.FIELD
[0002] The present application relates to methods and systems for identifying and validating a Synergistic Multifunctional Therapeutic (SMTh) (e.g. output ligand) targeting a duolog or multilog (gene combination) that is associated with a trait. In certain aspects, the methods and systems provided herein are useful for identifying ligands that can be used to treat a condition, such as a disease, in an individual by being capable of binding to more than one protein.BACKGROUND
[0003] Conventional drug discovery approaches are built on a foundation of identifying a single compound that specifically modulates a single target (e.g. protein). Such monogenic drug discovery methods include identifying a gene and / or protein associated with a condition (e.g., a disease) and then screening candidate compounds based on modulation of said gene and / or protein to treat the condition. As our understanding of cellular mechanisms grows, it is evident that monogenic approaches to drug discovery are limited in their ability to: address more complex conditions caused by more than one gene (or derivatives thereof), and / or advance beyond treatments focused on a single target. Taking a polygenic approach to identify underlying genes associated with polygenic traits complicates drug discovery because under the single compound, single target compound, the number of compounds needed to discover and test also increases. Once identified, the single targets also need to be tested and validated in combinations. There remains a need for approaches to identifying ligands and / or drags that modulate more than one target gene to target polygenic traits, such as human diseases, with a single compound.BRIEF SUMMARY
[0004] Provided herein are methods and systems for identifying an output ligand capable of binding (e.g. binds to) to two or more proteins comprising training and using machine learning1MF-367788979Attorney Docket: 27576-2000240models trained to predict binding of a ligand to one or more proteins. The methods and systems can be used to identify and validate an output ligand as binding to two or more proteins associated with a trait, such as a human disease. The methods and systems can be used to identify and output ligands capable of binding to two or more proteins that are not expected to bind to the same ligand based on differences in sequence and / or differences in binding pocket sizes. The methods and systems can be used to identify and output ligands that bind to two or more proteins via different combinations of interacting functional groups on both the ligand and each of the proteins.
[0005] Also provided herein are methods and systems for identifying a gene combination with an output ligand capable of binding to two or more protein products of the gene combination. In some embodiments, the methods and systems are for assessing co-drugability of a gene combination based on identification of a ligand capable of binding to the protein products of the gene combination. The methods and systems can be used to increase efficiency in drug design when multiple synergistic gene combinations have been identified as potential targets (e.g. proteins) for treatment of a disease.
[0006] In some aspects, the methods and systems comprise two or more machine learning models each trained to predict binding of a ligand to a protein based on calculated binding of training ligands to the protein. The two or more machine learning models can be used to predict binding of test ligands to each of the two or more proteins and an output ligand can be identified based on the two or more predicted bindings.
[0007] In some aspects, the methods and systems comprise a machine learning model trained to predict binding of a ligand to two or more proteins. The machine learning model can be used to predict binding of test ligands to the two or more proteins and an output ligands can be identified based on the predicted binding.
[0008] Provided herein are methods of identifying an output ligand capable of binding to two or more proteins; the method comprising: by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determining binding of each training ligand in the training chemical structure data to each of two or more proteins to generate2MF-367788979Attorney Docket: 27576-2000240training data; training, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins; inputting at least a subset of the chemical structure data into the two or more machine learning models, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands; outputting from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; and identifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding of the test ligand to the two or more proteins.
[0009] In some aspects, identifying the output ligand comprises selecting the output ligand if the predicted binding of the test ligand is above a first predetermined cutoff a first protein of the two or more proteins and above a second predetermined cutoff for a second protein of the two or more proteins. In some aspects, the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the predicted binding of the test ligands to the first protein. In some aspects, the second predetermined cutoff is selected based on a top about 10 million of test ligands as ranked by the predicted binding of the test ligands to the second protein.
[0010] In some aspects, identifying the output ligand comprises selecting the output ligand if the predicted binding of the output ligand to each of the two or more proteins is below a predefined energy. In some aspects, the predefined energy is about -8.0 Kcal / mol.
[0011] Also provided herein are methods of identifying an output ligand capable of binding to two or more proteins; the method comprising: by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determining binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; training, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins; inputting at least a subset of the chemical structure data into the machine learning model, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands; outputting from the machine learning models, predicted binding of each test ligand in the subset3MF-367788979Attorney Docket: 27576-2000240of the chemical structure data to the two or more proteins; and identifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
[0012] In some aspects, the predicted binding comprises a first predicted binding for the ligand to a first protein of the two or more proteins and a second predicted binding for the ligand to a second protein of the two or more proteins. In some aspects, identifying the output ligand comprises selecting the output ligand if the first predicted binding is above a first predetermined cutoff and the second predicted binding is above a second predetermined cutoff. In some aspects, the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the first predicted bindings. In some aspects, the second predetermined cutoff is selected based on a about 10 million of test ligands as ranked by the second predicted bindings.
[0013] In some aspects, the predicted binding comprises a combined predicted binding for the ligand to a first protein of the two or more proteins and to a second proteins of the two or more proteins. In some aspects, identifying the output ligand comprises selecting the output ligand if the combined predicted binding is above a first predetermined cutoff. In some aspects, the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the combined predicted bindings.
[0014] In some aspects, the methods comprise determining binding to the two or more proteins for the output ligand. In some aspects, determining binding comprises computing an empirical docking score in silico using the chemical structure data and protein structure data and / or an affinity screening method to measuring binding of the training ligand to the protein in vitro.
[0015] In some aspects, the protein structure data comprises a 2D representation of a 3D structure of a protein. In some aspects, the 3D representation comprises Xray, NMR, cryo-EM, or deep learning structural models data. In some aspects, the empirical docking score comprises a Wscore, a SiteMap score, a Glide score, a DOCK3.7 score or a DOCK3.8 score. In some aspects, the affinity screening comprises array -based screening technology.
[0016] In some aspects, the methods comprise calculating an absolute binding free energy (ABFE) of the output ligand. In some aspects, identifying the output ligand is based on an ABFE of the test ligand to the two or more proteins. In some aspects, calculating ABFE comprises simulating a transition of the test ligand and a protein of the two or more proteins from an unbound state to a bound state and computing a free energy difference between the unbound and4MF-367788979Attorney Docket: 27576-2000240bound state. In some aspects, calculating ABFE comprises sampling free energy of one or more intermediate states between the unbound state and the bound state.
[0017] In some aspects, the plurality of ligands comprises between about 48 million ligands and 1 billion ligands. In some aspects, the plurality of training ligands comprises about 0.1% of the plurality of ligands. In some aspects, the plurality of training ligands comprises about 48,000 of the plurality of ligands. In some aspects, the plurality of training ligands comprises about 100,000 of the plurality of ligands.
[0018] In some aspects, selecting training chemical structure data comprises selecting chemical structure data corresponding to the training ligands. In some aspects, selecting the training ligands is based on chemical diversity of the plurality of ligands. In some aspects, the training ligands represent the about 0.1% most diverse structures in the plurality of ligands. In some aspects, the 0.1% most diverse structures in the plurality of ligands are selected based on variation in chemical structures of the ligands in the plurality of ligands or variation in chemical fingerprints in the chemical structures of the ligands in the plurality of ligands.
[0019] In some aspects, the two or more proteins have sequence similarity between about 10% and 30%. In some aspects, the two or more proteins have consensus binding site sequence similarity between about 10% and 50%. In some aspects, binding site volumes of the two or more proteins is not predictive of the predicted binding to the test ligands. In some aspects, the output ligand is capable of binding to each of the two or more proteins in different orientations.
[0020] In some aspects, the two or more proteins have been determined to have a synergistic effect on a disease. In some aspects, the two or more proteins have been identified as a gene combination with a synergistic effect on a trait.
[0021] In some aspects, the machine learning model comprises an AQ / DC, a graphical convolution neural network (GCNN), and or a deep neural network (DNN). In some aspects, training the machine learning model comprises performing a 5-fold cross validation scheme to select hyperparameters for the machine learning model. In some aspects, the hyperparameters comprise model algorithm, number of layers, number of training epochs, and a normalization scheme applied to a raw response variable.
[0022] In some aspects, the machine learning models comprises a multi dataset machine learning model.5MF-367788979Attorney Docket: 27576-2000240
[0023] In some aspects, the chemical structures comprise a 2 dimensional (2D) representation of the plurality of ligands. In some aspects, the 2D representation is a Simplified Molecular Input Line Entry System (SMILES) representation or a morgan fingerprint.
[0024] In some aspects, the plurality of ligands comprises chemical compounds.
[0025] Also provided herein are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; train, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins; input the at least a subset of the chemical structure data into the two or more machine learning models, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands; output from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; and identify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
[0026] Also provided herein are systems: comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; train, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins; input the at least a subset of the chemical structure data into the machine learning model, wherein the subset of chemical structure data comprises chemical structures for the6MF-367788979Attorney Docket: 27576-2000240plurality of test ligands; output from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; and identify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.BRIEF DESCRIPTION OF THE DRAWINGS
[0027] FIG. I describes an example flowchart describing a method for identifying an output ligand capable of binding to two or more targets in accordance with various embodiments.
[0028] FIG. 2 describes an example flowchart describing training of a machine learning model in accordance with various embodiments.
[0029] FIG. 3 describes an example flowchart describing active learning of a machine learning model in accordance with various embodiments.
[0030] FIG. 4 illustrates an example system for identifying an output ligand capable of binding to two or more targets in accordance with various embodiments.
[0031] FIG. 5A provides a structural representation of an example compound docked in the MAPK3 protein.
[0032] FIG. 5B provides a structural representation of an example compound docked in the FDFT1 protein.
[0033] FIG. 6 shows the predicted binding of ligands to FDFT1 in Kcal / mol compared to the predicted binding of the ligands to MAPK3 in Kcal / mol, colored by the belonging to the ligand to the 3B library ligands used for testing, 3B library ligands used for training, and 48M used for testing.
[0034] FIG. 7 provides the structure of ligands predicted to bind to both FDFT1 and MAPK3.
[0035] FIG 8 shows binding of one a ligand predicted to bind to both FDFT1 and MAPK3 docked in the proteins in different orientations.DETAILED DESCRIPTION
[0036] The present application provides, in certain aspects, methods and systems for identifying and validating Synergistic Multifunctional Therapeutics (SMThs) (described herein as output ligands capable of binding to two or more proteins). SMThs represent a novel approach to therapy for polygenic human diseases. Once gene combinations that have a synergistic effect on a trait (e.g. duologs and multilogs) are identified, the methods described herein can be used to 7MF-367788979Attorney Docket: 27576-2000240identify and validate ligands for use in treatments that can target the gene combinations to achieve synergistic effects without the need to develop two or more separate therapeutic compounds to account for each gene in the combination.
[0037] The inventions provided herein are based, at least in part, on the inventors’ unique perspectives and findings regarding synergistic relationships between genes from seemingly distinct biological networks in explaining human traits and disease. The methods and systems can be used as to prioritize synergistic gene combinations or gene combinations that are part of synergistic networks as part of a drug discovery pipelines. A gene combination with identified and validated SMThs, according to the methods described herein, is a gene combination with co¬ druggable protein products.
[0038] Identification of ligands that bind to one protein target is a time intensive and expensive process. Identification of a ligand that binds two or more protein multiplies the complexity of the problem. By leveraging artificial intelligence methods in novel ways, the inventions allow for quick sorting of large libraries of ligands and multiple binding configurations for each library. The methods allow for identification of SMThs that would otherwise be impractical with conventional methods.
[0039] It is conventionally believed that similar proteins and / or proteins with similar binding pockets will be able to bind to the same ligand. The inventors identified that sequence similarity between protein targets and / or similarly in binding pockets does not reliably predict the existence of ligands able to bind to two or more proteins. Other approaches to co-drug proteins with low sequence similarity or different binding pockets rely on combing modules that individually target each protein, for example bifunctional tethered molecules.
[0040] Previous methods designed to identify a ligand with overlapping binding to multiple targets rely on target proteins that are already known to be similar or capable of being bound by the same ligands. The finding demonstrates that the small and directed screens based on a priori knowledge of protein similarity, such as those previously described, are not sufficient to identify SMThs or to prioritize combinations of proteins based on co-draggability. Rather, exceptionally large libraries of ligands are needed to test if and how many known ligands are capable of binding to two or more target proteins.
[0041] As the size of ligand libraries increases and the number of protein combinations to which the methods are applied increases, the traditional computational chemistry or empirical methods8MF-367788979Attorney Docket: 27576-2000240used to determine binding between a ligand and a target protein become increasingly inefficient. First, computationally expensive and / or lengthy experimental methods would need to be applied to each target protein individually and second, the results would need to be overlapped and compared. The methods and systems provided herein rely on active learning and use of machine learning models to predict binding without the need to determine bindings for a full library of ligands. Active learning as described herein is intended to encompass active training methods. In some aspects, the methods described herein use multi data set machine learning models that can be trained and fine-tuned using multiple types of data to improve predictions without needing to increase the size of training datasets or determine (e.g. calculate) empirical binding for a full training library and multiple targets. By training models with multiple independent datasets, the similarities and differences between the datasets can be used in generating practical and accurate binding predictions. The multi data set machine learning models may enable identification of compounds capable of binding from large libraries. The approaches exponentially decrease the time and computational resources needed to identify ligands able to bind to tvoo or more diverse protein targets and / or to identify a co-druggable combination of proteins from a group of synergistic network. The methods and systems are able to identify ligands that can bind to two or more proteins even when the ligand has a different binding conformation or orientation in each protein and can be applied multiple times efficiently to identify co-druggable protein combinations from a larger list of synergistic protein combinations.
[0042] The approach described herein provides an efficient solution to identify a ligand capable of binding to any combination of two or more target proteins and / or ranking a combination of two or more proteins based on co-druggability without relying on a priori knowledge of similarity between the two or more target proteins. Unlike previous approaches, the methods and systems are highly flexible and can be used to identify ligands capable of binding to two or more target proteins with highly distinct sequences, binding pocket structures and that bind in distinct conformations.I. Definitions
[0043] Unless otherwise defined, all the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.9MF-367788979Attorney Docket: 27576-2000240
[0044] The terms “polypeptide” and “protein,” as used herein, may be used interchangeably to refer to a polymer comprising amino acid residues, and are not limited to a minimum length. Such polymers may contain natural or non-natural amino acid residues, or combinations thereof, and include, but are not limited to, peptides, polypeptides, oligopeptides, dimers, trimers, and multimers of amino acid residues. Full-length polypeptides or proteins, and fragments thereof, are encompassed by this definition. The terms also include modified species thereof, e.g., post-translational modifications of one or more residues, for example, methylation, phosphorylation glycosylation, sialylation, or acetylation.
[0045] The term “treating” or “treatment,” as used herein, is an approach for obtaining beneficial or desired results including clinical results. For purposes of this application, beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms resulting from the disease, diminishing the extent of the disease, stabilizing the disease (e.g., preventing or delaying the worsening of the disease), preventing or delaying the spread (e.g., metastasis) of the disease, preventing or delaying the recurrence of the disease, delay or slowing the progression of the disease, ameliorating the disease state, providing a remission (e.g., partial or total) of the disease, decreasing the dose of one or more other medications required to treat the disease, delaying the progression of the disease, increasing the quality of life, and / or prolonging survival.
[0046] The term “individual” refers to a mammal and includes, but is not limited to, human, bovine, horse, feline, canine, mouse, rodent, or primate.
[0047] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0048] As used herein, the terms “comprising” (and any form or variant of comprising, such as “comprise” and “comprises”), “having” (and any form or variant of having, such as “have” and “has”), “including” (and any form or variant of including, such as “includes” and “include”), or “containing” (and any form or variant of containing, such as “contains” and “contain”), are inclusive or open-ended and do not exclude additional, un-recited additives, components, integers, elements, or method steps.
[0049] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for10MF-367788979Attorney Docket: 27576-2000240convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Therefore, the description of a range should be interpreted as having explicitly included all potential sub-ranges and specific numerical values within that range. For example, when a range of values is specified, it is assumed that each value, up to one-tenth of the unit of the lower limit, unless the context indicates otherwise, between the upper and lower limits of that range, along with any other mentioned or intermediate values within that range, is included in the disclosure, except for any limits explicitly excluded within the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure. In some embodiments, two opposing and open-ended ranges are provided for a feature, and in such description it is envisioned that combinations of those two ranges are provided herein. For example, in some embodiments, it is described that a feature is greater than about 10 units, and it is described (such as in another sentence) that the feature is less than about 20 units, and thus, the range of about 10 units to about 20 units is described herein.
[0050] As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value. Reference to “about” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se.
[0051] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.II. Methods of identifying an output ligand capable of binding to two or more proteins
[0052] Provided herein are methods that can be used to identify an output ligand capable of binding to two or more proteins. In some embodiments, the methods can be used to identify an output ligands that binds to two or more proteins. The methods may be computer implemented methods wherein the methods are performed by one or more computing device comprising one or more processors and memory. The methods can be used to identify one or more output ligands that can be further tested and validated to select a ligand that can be used for treating a disease, or be used as a starting point to further derive a ligand that can be used for treating a disease. The methods rely on chemical structure data for a plurality of ligands, for example 48 million or 2 billion ligands and binding calculation methods that would be unable to determine the binding of 11MF-367788979Attorney Docket: 27576-2000240each of the ligands in the plurality of ligands effectively and / or efficiently. Thus, the methods comprise training and using machine learning models trained to predict binding to a protein from a chemical structure.
[0053] The methods comprise, by or more processors and memory; receiving chemical structure data for a plurality of ligands; selecting training chemical structure data; determining binding of chemical structures for the training ligands to two or more proteins to generate training data. The training data may be used to train one or more machine learning models either by training a machine learning model to predict binding of a ligand to each protein of the two or more proteins or by training a machine learning model to predict binding of a ligand to all proteins of the two or more proteins (e.g. combined model). At least a subset of the chemical structure data is input into the two or more machine learning models and predicted binding of test ligands to the two or more proteins is outputted. The predicted binding of the test ligands to the two or more proteins is used to identify an output ligand.
[0054] In some embodiments, the two or more proteins may be protein products of genes from a gene combination. The proteins combination and or the gene combination may have been determined to have a synergistic effect on a trait or disease. For example, methods such as those described in U. S. Application 19 / 025,625 and the PCT / US2025 / 011900, hereby incorporated by reference in their entirety may have been used to identify the gene combination with a synergistic effect. The protein products of a gene combination with a synergistic effect on a disease may be targets for a therapy treating the disease by modulating all the proteins in the combination in conjunction. The treatment effect may be synergistic, meaning treatment targeting the combination is more effective than an additive effect of targeting the protein products independently and in some cases targeting one protein may have no effect on its own.
[0055] In some embodiments, the methods comprise selecting the two or more proteins based on a determined synergistic effect. In some embodiments, the determined synergistic effect is a synergistic effect on a trait when the two or more proteins are modulated. In some embodiments, the methods comprise selecting two or more proteins from a combination of genes. In some embodiments, the two or more proteins are protein products of the combination of genes. In some embodiments, the combination of genes has a synergistic effect on a trait. In some embodiments, the two or more proteins have been selected based on a determined synergistic effect.12MF-367788979Attorney Docket: 27576-2000240
[0056] The output ligand identified using the methods described herein is capable of binding to two or more proteins. According to the methods described herein binding refers to an interaction between the output ligand and the protein. In some embodiments, a ligand is considered as capable of binding and / or is bound to a protein if the reaction has a KD of about less than 10 micromolar. An output ligand may be capable of binding the two proteins with the same binding affinity or with different binding affinities. In some embodiments, output ligand may bind to the two or more proteins in a different configuration. In some embodiments, different aspects of the output ligand may drive binding to the two or more proteins. As used herein, an output ligand capable of binding to two or more proteins comprises an output ligand wherein the binding domains for binding of the output ligand to the two or more protein overlap at least partially. For example, at least some of the chemical moieties or functional groups involved in binding of the output ligand to the first protein are involved in binding of the output ligand to a second protein. The overlap of the binding domains does not need to be complete. In some embodiments, the output ligand is not a tethered bifunctional compound.
[0057] In some embodiments, the methods described herein can be used to identify an output ligand capable of binding to two or more proteins that one of skill in the art would not have expected would bind to a single ligand. For example, ligands capable of binding to the two or more proteins are not known in the art. While others have successfully drugged combinations of kinases, (See Luciano et. al., The multi-tyrosine kinase inhibitor ponatinib for chronic myeloid leukemia: Real-world data, 105(1), Fur. J. Haematol. (2020)), the methods described herein aim at identifying output ligands with low sequence similarity (e.g. below 30%) and those skilled in the art do not routinely know how to drug without resorting to tethered bifunctional molecules. The tethered bifunctional molecules may lead to difficulties with other drug properties. The methods described herein are designed to enable these more difficult combinations in a way that yields good drug properties by solving the partially overlapping binding.
[0058] In some embodiments, the methods described herein can be used to identify an output ligand capable of binding to two or more proteins with low sequence similarity. In some embodiments, the methods comprise selecting the two or more proteins based on sequence similarity between the two or more proteins. In some embodiments, the two or more proteins have been selected based on sequence similarity between the two or more proteins. In some embodiments, the methods comprise measuring the sequence similarity between the two or more13MF-367788979Attorney Docket: 27576-2000240proteins. In some embodiments, the methods comprise selecting two or more proteins with sequence similarity between about 10% and about 30%. In some embodiments, the methods comprise selecting two or more proteins with sequence similarity between about 15% and about 30%, about 20% and about 30%, about 25% and about 30%, about 15% and about 25%, or about 15% and about 20%. In some embodiments, the sequence similarity is about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, or about 30%.
[0059] In some embodiments, the two or more proteins have sequence similarity between about 10% and about 30%. In some embodiments, the two or more proteins have sequence similarity between about 15% and about 30%, about 20% and about 30%, about 25% and about 30%, about 15% and about 25%, or about 15% and about 20%. In some embodiments, the sequence similarity is about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, or about 30%. In some embodiments, the two or more proteins comprise at least three proteins and the sequence similarity between two or more of the at least three proteins have sequence similarity between 10% and 30%.
[0060] In some embodiments, the methods described herein can be used to identify an output ligand capable of binding to two or more proteins with low consensus binding site sequence similarity. In some embodiments, the methods comprise measuring the consensus binding site sequence similarity between the two or more proteins. In some embodiments, the methods comprise selecting the two or more proteins based on consensus binding site sequence similarity between the two or more proteins. In some embodiments, the two or more proteins have been selected based on consensus binding site sequence similarity between the two or more proteins. In some embodiments, the methods comprise selecting two or more proteins with consensus binding site sequence similarity between about 10% and about 30%. In some embodiments, the methods comprise selecting two or more proteins with consensus binding site sequence similarity between about 15% and about 30%, about 20% and about 30%, about 25% and about 30%, about 15% and about 25%, or about 15% and about 20%, In some embodiments, the consensus binding site sequence similarity is about 10%, about 11%, about 12%, about 13%, about 14%, about14MF-367788979Attorney Docket: 27576-200024015%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, or about 30%.
[0061] In some embodiments, the two or more proteins have consensus binding site sequence similarity between about 10% and about 30%, In some embodiments, the two or more proteins have consensus binding site sequence similarity between about 15% and about 30%, about 20% and about 30%, about 25% and about 30%, about 15% and about 25%, or about 15% and about 20%. In some embodiments, the consensus binding site sequence similarity is about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, or about 30%. In some embodiments, the two or more proteins comprise at least three proteins and the consensus binding site sequence similarity between two or more of the at least three proteins have consensus binding site sequence similarity between 10% and 30%.
[0062] FIG. 1 provides an exemplary flowchart for methods of identifying an output ligand capable of binding to two or more proteins according to several embodiments as described herein.
[0063] At block 102, chemical structure data are received. The chemical structure data comprise chemical structures for a plurality of ligands, wherein the plurality of ligands comprise training ligands and test ligands. In some embodiments, the chemical structures comprise a 2 dimensional (2D) representation of the plurality of ligands. In some embodiments, the 2D representation is a simplified Molecular Input Line Entry System (SMILES) representation or a morgan fingerprint. In some embodiments, the SMILES representation may comprise one or more SMILES strings, deep simplified molecular-input line entry system (DeepSMILES) stings, or self-referencing embedding strings (SELFIES). In some embodiments, the 2D representations represent an original chemical structure of a ligand in the plurality of ligands. In some embodiments, the 2D representation is a predicted chemical structure of a ligand in the plurality of ligands.
[0064] In some embodiments, the chemical structure data received at block 102 comprise chemical structure data for a plurality of ligands, wherein the plurality of ligands comprises about 20 million, about 25 million, about 30 million, about 35 million, about 40 million, about 50 million, about 55 million, about 60 million, about 65 million, about 70 million, about 75 million, about 80 million, about 85 million, about 90 million, or about 100 million ligands. In15MF-367788979Attorney Docket: 27576-2000240some embodiments, the plurality of ligands comprises 48 million ligands. In some embodiments, the plurality of ligands comprises between about 20 million and 100 million ligands, between about 25 million and 100 million ligands, between about 30 million and 100 million ligands, between about 35 million and 100 million ligands, between about 40 million and 100 million ligands, between about 45 million and 100 million ligands, between about 50 million and 100 million ligands, between about 55 million and 100 million ligands, between about 60 million and 100 million ligands, between about 65 million and 100 million ligands, between about 70 million and 100 million ligands, between about 75 million and 100 million ligands, between about 80 million and 100 million ligands, between about 85 million and 100 million ligands, between about 90 million and 100 million ligands, or between about 95 million and 100 million ligands.
[0065] In some embodiments, the plurality of ligands comprises between about 20 million and 95 million ligands, between about 20 million and 90 million ligands, between about 20 million and 80 million ligands, between about 20 million and 75 million ligands, between about 20 million and 70 million ligands, between about 25 million and 65 million ligands, between about 20 million and 60 million ligands, between about 20 million and 55 million ligands, between about 20 million and 50 million ligands, between about 20 million and 45 million ligands, between about 20 million and 40 million ligands, between about 20 million and 35 million ligands, between about 20 million and 30 million ligands, or between about 20 million and 25 million ligands.
[0066] In some embodiments, the chemical structure data received at block 102 comprise chemical structure data for a plurality of ligands, wherein the plurality of ligands comprises about 100 million, about 200 million, about 300 million, about 400 million, about 500 million, about 600 million, about 700 million, about 800 million, about 900 million, about 1 billion, about 1.5 billion, about 2 billion, about 2.5 billion, about 3 billion, about 3.5 billion, about 4 billion, about 4.5 billion, or about 5 billion ligands. In some embodiments, the plurality of ligands comprises 3 billion ligands. In some embodiments, the plurality of ligands comprises more than 5 billion ligands.
[0067] In some embodiments, the plurality of ligands comprises between about 100 million and about 5 billion, between about 200 million and about 5 billion, between about 300 million and about 5 billion, between about 400 million and about 5 billion, between about 500 million and about 5 billion, between about 600 million and about 5 billion, between about 700 million and16MF-367788979Attorney Docket: 27576-2000240about 5 billion, between about 800 million and about 5 billion, between about 900 million and about 5 billion, between about 1 billon and about 5 billion, between about 1.5 billon and about 5 billion, between about 2 billon and about 5 billion, between about 2.5 billon and about 5 billion, between about 3 billon and about 5 billion, between about 3.5 billon and about 5 billion, between about 4 billon and about 5 billion, between about 4.5 billon and about 5 billion ligands.
[0068] In some embodiments, the plurality of ligands comprises between about 100 million and about 4.5 billion, between about 100 million and about 4 billion, between about 100 million and about 3.5 billion, between about 100 million and about 3 billion, between about 100 million and about 2.5 billion, between about 100 million and about 2 billion, between about 100 million and about 1.5 billion, between about 100 million and about 1 billion, between about 100 million and about 900 million, between about 100 million and about 800 million, between about 100 million and about 700 million, between about 100 million and about 600 million, between about 100 million and about 500 million, between about 100 million and about 400 million, between about 100 million and about 300 million, or between about 100 million and about 200 million ligands.
[0069] In some embodiments, the plurality of ligands comprises chemical compounds. In some embodiments, the ligands comprise small molecules. In some embodiments, small molecules weigh about 2500 Da or less. In some embodiments, a ligand in the plurality of ligands is a part of chemical compound. The part of the chemical compound may comprise a part of a chemical compound capable of binding to a protein. In some embodiments, the ligands may be therapeutic compounds. In some embodiments, the ligands are identified from a database of therapeutic small molecules, including but not limited to Enamine, Wuxi Galaxi, or Zinc. In some embodiments, the therapeutic compounds comprise regulatory approved compounds, such as FDA approved compounds.A. Generating training data
[0070] At block 104, the chemical structure data from 102 are used for selecting training chemical structure data. The training chemical structure data comprises chemical structures for the plurality of training ligands. In some embodiments, selecting the training chemical structure data comprises selecting chemical structure data corresponding to the training ligands.
[0071] In some embodiments, the plurality of training ligands comprises about 0.1% of the plurality of ligands. In some embodiments, the plurality of training ligands comprises about 0.05%, about.1%, about 1.5%, about 2%, about 2.5%, about 3%, about 3.5%, about 4%, aboutMF-367788979Attorney Docket: 27576-20002404.5% or about 5% of the plurality of ligands. In some embodiments, the plurality of training ligands comprises about 48,000 ligands. In some embodiments, the plurality of training ligands comprises about 100,000 ligands.
[0072] In some embodiments, selecting the training ligands is based on chemical diversity of the training ligands. Chemical diversity may be based on the class or structure of the ligand, such as based on a chemical fingerprint (e.g. Morgan fingerprints). A high diversity of training ligands may improve training of the machine learning models described herein. As the diversity of the training ligands increases, the number of training ligands needed to train the machine learning model to predict binding to a protein with high accuracy may or may not decrease. Because the rate limiting step for the method may be calculating the binding of the training ligands to the two or more proteins, the efficiency of the method may be increased by decreasing the number of training ligands needed to train the machine learning models.
[0073] In some embodiments, the most diverse structures in the plurality of ligands are selected based on variation in chemical structures of the ligands in the variation in chemical fingerprints in the chemical structures of the ligands. In some embodiments, the method comprises applying a dimensionality reduction method to the chemical structure data and selecting ligands that are separated in the reduced space. In some embodiments, the training ligands represent the about 0.1% most diverse structures in the plurality of ligands. In some embodiments, the training ligands represent the about 0.05%, about.1%, about 1.5%, about 2%, about 2.5%, about 3%, about 3.5%, about 4%, about 4.5% or about 5% most diverse structures in the plurality of ligands.
[0074] As described herein, the training ligands may be selected as part of active learning of the machine learning models. In some embodiments, the training ligands may be selected based on uncertainty of predicted binding from a previous round of training and / or predicted binding of the ligand to the protein from the previous round of training. In some embodiments, active learning methods may decrease the number of training ligands that are needed to train the machine learning models to predict binding to a protein with high accuracy.
[0075] At block 106, the training chemical structure data from block 104 are used to determine binding of the training ligands to each of the two or more proteins to generate training data. The training data may comprise chemical structures for each training ligand and an associated determined binding to one or more proteins. In some embodiments, the determined binding is18MF-367788979Attorney Docket: 27576-2000240represented as an energy (Kcal / mol), wherein stronger binding of a ligand to a protein is represented by a lower Kcal / mol. Determining the binding may comprise calculating the binding of the test ligands as well as the training ligands which would be inefficient and expensive using methods known in the art for determining binding of a ligand to a protein.
[0076] In some embodiments, determining binding comprises computing an empirical docking score. The score may be computed in silico using chemical structure data for the ligand and protein structure data for the protein. In some embodiments, the protein structure data comprises a 2D representation of a 3D structure of a protein. In some embodiments, the 3D representation comprises X-ray, NMR, cryo- EM. or deep learning structural data (e.g. AlphaFold data). In some embodiments, the empirical docking score comprises a WScore, SiteMap score, Glide score, or a DOCK3.7 or DOCK3.8 score, but not limited to these. The empirical docking score may be calculated using tools available in the Schrodinger suite of tools according to provided instructions. In some embodiments, the empirical docking score can be a WScore, as described in Murphey et al., WScore: A Flexible and Accurate Treatment of Explicit Water Molecules in Ligand-Receptor Docking, 2016 May 12; J. Med. Chem., 59(9):4364-84, or as implemented in the Schrodinger suite. WScore uses water energetics of a binding site to score binding of a ligand to a protein. In some embodiments, the empirical docking score may be calculated with Glide (e.g., Glide v. 5.6) within the Schrodinger software suite. (Schrodinger, LLC) (Mohamadi, F. et al. J. Comput. Chem. 1990; 11(4):440-467). In some embodiments, the empirical docking score is a SiteMap score calculated by Schrodinger's SiteMap program (SiteMap, version 3.0, Schrodinger, LLC, New York, N. Y., 2014). SiteMap searches the protein structure for likely binding sites and highlights regions within the binding site suitable for occupancy by hydrophobic groups, hydrogen- bond donors, acceptors, or metal-binding functionality of the ligand.
[0077] Empirical docking score methods are intrinsically slow and do not scale effectively for determining docking at the scale of the ligand libraries described herein. The lack of scalability may come from the increased computing power that would be needed to combine intermediate results computed on multiple CPUs or GPUs outweighs the computational improvement realized by performing the task on more CPUs or GPUs. Algorithms such as those used to determine empirical docking scores in the Schrodinger suite are computationally expensive and slow, thus a19MF-367788979Attorney Docket: 27576-2000240brute force method to calculate an empirical docking score for all of the ligands in the plurality of ligands would not be practical.
[0078] In some embodiments, determining binding comprises an affinity screening method to measure binding of the training ligand to the protein in vitro. The affinity screening methods may comprise measuring binding of the training ligand to the protein in vitro. In some embodiments, determining binding comprises an affinity screening method. In some embodiments, the affinity screening method comprises DNA encoded library (DEL) screening or similar type array based screening methods. Array-based screening technology refers to methods used to evaluate multiple samples or compounds simultaneously in a high-throughput manner. DEL or other array based screening may be used to screen one or more proteins for binding to a compound, (see Shi et al, DNA-encoded libraries (DELs): a review ofon-DNA chemistries and their output, 4, RSC Advances (2021). Other array based methods may utilize nanoparticle beads or optical screening, such as a tArray. Methods for determining binding in vitro may comprise standard assays known in the art, including but not limited to competition binding probe assays. These methods are inefficient because the ligands need to be synthesized, and the protein needs to be isolated in the lab.
[0079] In some embodiments, the machine learning models described herein are multi dataset machine learning models. Multi dataset machine learning models according to the embodiments described herein may be trained with computational, experimental methods such as affinity screening methods determinations as described herein and / or additional datasets related to the prediction task (e.g. to ligand binding, SMTh detection). In some embodiments, the additional datasets used for training may comprise pharmaceutical properties of the ligands. The additional datasets may relate to binding of the first target to the training ligands or a subset of the training ligands. The additional datasets may relate to binding of the second target to the training ligands or a subset of the training ligands. The additional datasets may relate to binding of the two or more targets to the training ligands or a subset of the training ligands. In some embodiments, the methods comprise obtaining the additional training datasets,
[0080] In some embodiments, the additional training datasets comprise structure based virtual screening (SB VS) datasets. SBVS datasets comprise diverse sets of compounds that were virtually screened against a target based on structural information. The structural information may include docking scores or binding poses. In some embodiments, the SBVS datasets20MF-367788979Attorney Docket: 27576-2000240comprise predicted binding affinities that have been generated based on the machine learning models described herein. In some embodiments, SBVS datasets comprise empirical docking scores generated according to methods described herein.
[0081] In some embodiments, the additional training datasets comprise NMR fragment screening data sets. NMR fragment datasets comprise fragments of ligands that have been screened with NMR methods to identify binding interactions to a target (e.g, first target and / or second target). In some embodiments, the NMR fragment data sets may comprise empirical binding affinity measurements or qualitative information on binding interactions to generate fingerprints. The fingerprints may be parts of a ligand that are predicted to bind to a target.
[0082] In some embodiments, the additional training datasets comprise data from affinity¬ screening methods as described herein. In some embodiments, the additional training datasets comprise data from DEL screens.B. Training the machine learning models
[0083] At block 108, the training data generated at block 106 are used to train one or more machine learning models to predict binding of a test ligand to a protein. In some embodiments, at block 108, a machine learning model is trained per protein of the two or more proteins. The two or more machine learning models are each trained to predict binding of a test ligand to a protein of the two or more proteins. In some embodiments, each machine learning model is trained based on the chemical structure data for the training ligands and the determined binding of the training ligands to a protein of the two or more proteins. The machine learning model may be trained to input chemical structure data (e.g. a SMILES representation) for a test ligand and output a predicted binding of the ligand to the protein. The one or more machine learning models may be trained concurrently or sequentially.
[0084] In some embodiments, the machine learning model comprises a regression model, a random forest model or a neural network. In some embodiments, the machine learning model comprises an AQ / DC model. The AQ / DC model may be trained to predict binding of a ligand to a protein using the methods described in Yang et. al., Efficient Exploration of Chemical Space with Docking and Deep Learning, 17, J. of Chem. Theory and Computation (2021). The AQ / DC model may be an automated neural network implementing a graph-convolution neural network (GCNN). In some embodiments, the machine learning model comprises a GCNN, or a deep neural network (DNN).21MF-367788979Attorney Docket: 27576-2000240
[0085] In some embodiments, any artificial neural networks (ANNs) that may be utilized to learn deep levels of representations and abstractions from large amounts of data may be used. For example, the machine learning models may include ANNs, such as a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth.
[0086] In some embodiments, the machine learning model may comprise a classification algorithm or function may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon.
[0087] In some embodiments, training the machine learning model comprises performing a 5-fold cross-validation scheme to select hyperparameters for the machine learning model. In some embodiments, the hyperparameters comprise model algorithm, number of layer, number of training epochs, and a normalization scheme applied to a raw response variable. The 5-fold cross validation scheme may comprise training 5 models each to a different fold of data for a given hyperparameter set. See Yang et. al., Efficient Exploration of Chemical Space with Docking and Deep Learning, 17, J. of Chem Theory and Computation (2021).
[0088] In some embodiments, at block 108, a machine learning model is trained to predict binding of a test ligand to the two or more proteins. In some embodiments, training the machine learning model to predict binding of a test ligand to two or more proteins at the same time improves performance of the methods disclosed herein because the machine learning model parameters can be tuned to account for variation in binding between similar and different ligands to more than one protein. Training a machine learning model to predict binding of a ligand to two or more proteins may improve speed and efficiency of the model as the number of proteins increases. Rather than training and using a machine learning model to predict binding of a ligand to each protein independently, the machine learning model can be trained once for the22MF-367788979Attorney Docket: 27576-2000240combination of proteins. Training the machine learning model to predict binding of a ligand to two or more proteins may allow for more efficient or more accurate optimization of binding affinity to the two or more proteins The models may be able to balance the strength of affinity to the two or more proteins in generating predictions and a different search path may be utilized compared to models trained to individually predict binding of a ligand to a single protein.
[0089] In some embodiments, the machine learning model is trained based on the chemical structure data for the training ligands and the determined binding of the training ligands to each protein of the two or more proteins. The machine learning model may be trained to input chemical structure data (e.g. a SMILES representation) for a test ligand and output a predicted binding of the ligand to each of the proteins in the two or more proteins. In some embodiments, the machine learning model is trained to classify a ligand as likely to bind to the two or more proteins (e.g. combined machine learning model).
[0090] The machine learning model may be a model capable of predicting more than one target variable from a single input. In some embodiments, the machine learning model is a neural network, such as a GCNN or a DNN. Training the machine learning model may comprise performing a 5-fold cross-validation scheme to select hyperparameters for the machine learning model. In some embodiments, the hyperparameters comprise model algorithm, number of layer, number of training epochs, and a normalization scheme applied to two or more raw response variables.
[0091] In some embodiments, training the machine learning models (e.g. machine learning model to predict binding to one protein or machine learning model to predict binding to two or more proteins) comprises performing DEL screens as shown in FIG. 2. At block 202 and 204 DEL screens are performed (or the data are obtained) for binding of the DEL library compounds to target 1 and target 2 respectively. The data are used to generate a hit list comprising ligands capable of binding to protein 1 and protein 2 at blocks 206 and 208 respectively. The DEL library and the hit lists at 206 and 208 may be used to train the machine learning model at block 210. The trained machine learning model, 210, may be used to predict binding for a ligand library at block 212, and the DEL library may be rescreened based on the output of the machine learning model. The rescreening and predicted bindings from the ligand library may be used to identify an output ligand at block 214.23MF-367788979Attorney Docket: 27576-2000240
[0092] In some embodiments, training the machine learning model (e.g. machine learning model to predict binding to one protein or machine learning model to predict binding to two or more proteins) comprises active learning. An exemplary flowchart for active training is shown in FIG.3. At block 302, the plurality of ligands such as the plurality of ligands in the chemical structure data at block 102 is obtained. From the plurality of ligands at block 302, a subset of training ligands is obtained, using methods such as those used to select training chemical structure data at block 104. At block 306, binding of the training ligands to the one or more proteins is determined using methods such as those described at block 106. At block 308, the machine learning model is trained to estimate binding for test ligands. The trained machine learning model is used to predict binding of the full set of plurality of ligands. In some embodiments the predicted binding comprises a binding score and a measurement of uncertainty in the prediction. The predicted binding, including the predicted binding and performance scores related to the uncertainty in the predicted binding are used to select a new subset of training ligands for a second round of training and / or fine tuning of the machine learning model. In some embodiments, the uncertainty score may be a standard deviation in the binding prediction for the ligands from the cross validation in training (e.g. 5 fold cross validation). The new subset of training ligands may be selected according to the methods described above and may comprise fewer training ligands than the training ligands in the first round of training. For example, the ligands with the highest predicted binding and / or the highest uncertainly in binding to may be selected for the new subset of training ligands. In some embodiments, the active learning cycle may be performed, 1, 2, 3, 4, or 5 times. In some embodiments, the active learning cycle may be performed until accuracy of the model does not improve from additional rounds of training. The benefits of active learning and active training methods for training a machine learning model to predict binding of a test ligand to a single protein can be found in Yang et al., Efficient Exploration of Chemical Space with Docking and Deep Learning, 17, J. of Chem Theory and Computation (2021).
[0093] In some embodiments, the machine learning models described herein are multi dataset machine learning models. The multi dataset machine learning models may be machine learning models capable of predicting binding of a ligand to one or more targets. In some embodiments, the multi dataset machine learning models use multitask learning, transfer learning, ensemble methods, and / or domain adaptation. In multitask learning, a single model is trained on multiple related tasks that involve different training datasets in order to learn shared representations and24MF-367788979Attorney Docket: 27576-2000240improve performance across tasks. The multiple related tasks may be predicting binding to two or more targets. As a non-limiting example, these methods may be used when it is possible to determine binding using the methods described herein for only some of the two or more proteins and NMR data for proteins with similar fragments to the other proteins is all that is available.
[0094] In some embodiments, transfer learning comprises pretraining a model with one of the data sets and then fine tuning a model based on a second data set. This method may be effective for training machine learning models using datasets that are related but differ in size or distribution, such as but not limited to datasets with features associated with binding of training ligands to a protein measured using different methods. In some embodiments, transfer learning may be used to incorporate data obtained from experimental assays or computationally expensive data (e.g. ABFE screening) to refine the machine learning models described herein.
[0095] In some embodiments, the multi data set machine learning models may comprise ensemble models. Accordingly, separate models can be trained with features from individual datasets and then combined to improve overall performance of the methods. In some embodiments, combining the predictions may comprise averaging the outputs or use of a meta learner technique. As a non-limiting example, these methods may be used to combine results from two different types of screens. In some cases, a structure driven virtual screen (computationally determining binding between training ligand and the first protein) as described herein may be performed for one target but the second protein may not have structural information available, but wherein there are known ligands that bind to the second protein. A shape similarity screen may be performed and combined with the virtual screen performed for the first protein in order to train a machine learning model to predict binding to both proteins. Experimental screens performed in vitro or in vivo may be incorporated using ensemble methods.
[0096] In some embodiments, domain adaptation may comprise applying techniques to adapt models trained on one data set to perform well on another data set. As a non-limiting example, domain adaptation methods may be used to train a model with SBVS data for application using NMR fragment screening data.
[0097] In some embodiments, a machine learning model as described herein (a regression model, random forest model, neural network) using structural features associated with binding of a ligand to a protein. The structural features may include, but are not limited to, the determined25MF-367788979Attorney Docket: 27576-2000240binding of a training ligand to target structural properties of the training login and / or molecular descriptors related to the binding of the training ligand to the protein. The similarities and differences between structural features and features from an NMR dataset may be analyzed in order to apply transfer learning methods. The determined binding scores may be compared to empirical affinities, the scale of the binding affinities calculated and estimated from different methods (e.g., continuous values in SBVS vs. binary or categorical in NMR) and the diversity and size of the training ligands may differ between training datasets used in multi dataset methods.
[0098] In some embodiments, feature transformation techniques may be used to make datasets compatible. In some embodiments, feature transformation may comprise domain adversarial training. In some embodiments, the multi dataset machine learning model may be trained with a domain classifier that encourages the model to learn features that are invariant to domain differences. In some embodiments, feature transformation may comprise feature normalization wherein features are adjusted to minimize discrepancies across datasets for example using scaling or dimensionality reduction.
[0099] In some embodiments, methods described herein comprise fine tuning the machine learning model. In some embodiments, the machine learning model may be fine-tuned using NMR fragments data. Fine tuning may comprise transfer learning techniques described herein wherein additional training is performed adjusting the learning rates to prevent overfitting. In some embodiments, fine tuning may comprise focusing on features that are common between two or more dataset and / or are relevant to predicting binding of a training ligand to the one or more proteins.
[0100] In some embodiments, training the machine learning models comprises performing cross¬ domain validation. In some embodiments, validation may comprise validating the fine-tuned model using NMR fragment screening datasets and performance metrics such as RMSE, R2, or area under the curve (AUC) to ensure the model effectively predicts binding for the fragments.
[0101] In some embodiments, training the machine learning models comprises iterative refinement and continual training using additional training datasets, such as NMR fragment datasets, and evaluating performance against the original data set and new datasets. In some embodiments, additional insights from binding interactions identifying in NMR fragment datasets may be used to iteratively improve model prediction.26MF-367788979Attorney Docket: 27576-2000240C Predicting binding
[0102] The methods described herein comprise the use of machine learning models that can be used to predict binding of the test ligands to two or more proteins. Use of the machine learning models alleviates the need to determine (e.g. calculate) binding for the test ligands using the computationally or empirically more difficult methods. This is particularly important as the number of test ligands increases, which may be necessary to identify a ligand capable of binding to two or more proteins with low sequence similarity or binding site similarity as described herein.
[0103] At block 110, the machine learning models (1 machine learning model per protein or combined machine learning model) from block 108 is used to predict binding of the test ligands for the plurality of ligands to the two or more proteins. The methods comprise, inputting the at least a subset of the chemical structure data into the machine learning models, wherein the at least a subset of chemical structure data comprise structures for the plurality of test ligands. In some embodiments, the chemical structure data are input into two or more machine learning models. In some embodiments, the chemical structure data are input into one machine learning model. As described above, the machine learning models are trained to predict binding to one or more proteins using chemical structure data.
[0104] At block 112, predicted binding of each test ligand in the subset of clinical structure data to the two or more proteins is outputted from block 110. In some embodiments, a predicted binding of each test ligand to each protein of the two or more proteins is outputted. In some embodiments, a predicted binding of each test ligand to two or more proteins is outputted. In some embodiments, predicted binding comprises a combined predicted binding for each ligand to the two or more proteins (e.g. predicted binding to a first protein of the two or more proteins and a predicted binding to a second protein of the two or more proteins).
[0105] In some embodiments, the input chemical structure data from block 110 can be fed into the one or more machine learning models at block 112 sequentially. In some embodiments, the input chemical structure data may be input into a machine learning model trained to output a predicted binding for one of the two or more proteins. The predicted bindings of a ligand to the first of the two or more proteins can be used to reduce the size of the input library for input into a subsequent machine learning model. Ligands predicted to bind to the first protein of the two or more proteins may have predicted binding above any of the predetermined thresholds asMF-367788979Attorney Docket: 27576-2000240described herein. In some embodiments, the input chemical structure data for the test ligands that are predicted to bind to the first protein are input into the second machine learning model trained to predict binding of a test ligand to the second of the two or more proteins. This method can be expanded for any number of proteins. This method may increase efficiency because ligands that are not predicted to bind to any one of the two or more proteins will not be an output ligand. Each round may reduce the computing power needed to identify an output ligand.
[0106] In some embodiments, the predicted binding of the test ligands to the two or more proteins is not predictive of the binding site volume of the two or more proteins. In some embodiments, the combined predicted binding is not predictive of the binding site volume of the two or more proteins. In some embodiments, the predicted binding of the test ligands to the two or more proteins is not predictive of the number of test ligands identified as an output ligand. In some embodiments, the combined predicted binding of the test ligand to the two or more proteins is not predictive in the number or test ligands identified as an output ligand.D. Identifying an output ligand
[0107] At block 114, an output ligand is identified from the plurality of test ligands based on the predicted binding from block 112. The methods comprise identifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding of the test ligand to the two or more proteins. In some embodiments, the methods comprise identifying a plurality of output ligands wherein each output ligand is capable of binding to the two or more proteins.
[0108] In some embodiments, identifying the output ligand comprises selecting the output ligand if the predicted binding of the test ligand is above a first predetermined cutoff a first protein of the two or more proteins and above a second predetermined cutoff for a second protein of the two or more proteins. In some embodiments, identifying the output ligand comprises selecting the output ligand if the predicted binding of the test ligand is above a first predetermined cutoff a first protein of the two or more proteins, above a second predetermined cutoff for a second protein of the two or more proteins, above a third, fourth, fifth predetermined cutoff for a third protein, fourth protein, fifth protein.
[0109] In some embodiments, a predefined cutoff may be determined for each of the two or more proteins (e.g. first predefined cutoff, second predefined cutoff, third predefined cutoff). In some embodiments, the first predefined cutoff and the second predefined cutoff are the same. In28MF-367788979Attorney Docket: 27576-2000240some embodiments, the first predefined cutoff and the second predefined cutoff are different. Determined predefined cutoffs for each protein may provide flexibility in selecting an output ligand, such that the ligand is expected to bind to two of the two or more proteins or all proteins of the two or more proteins. Different predefined cutoffs allow for consideration of known properties of the protein as they relate to binding such as properties related to the structure and / or composition of the proviant.
[0110] In some embodiments, the first predefined cutoff is selected based on a ranking of the predicted bindings of the test ligand to the first protein. In some embodiments, the ranking is based on the uncertainty in predicted binding of the test ligand to the first protein. In some embodiments, a first predefined cutoff is selected based on the top about 10 million test ligands. In some embodiments, a first predefined cutoff is selected based on the top about 1 million, about 5 million, about 10 million, about 15 million, about 25 million, or about 30 million test ligands. In some embodiments, a first predefined cutoff is selected based on the top about 1 million to about 30 million, about 1 million to about 25 million, about 1 million to about 20 million, about 1 million to about 15 million, about 1 million to about 10 million, or about 1 million to about 5 million. In some embodiments, a first predefined cutoff is selected based on the top about 5 million to about 30 million, about 10 million to about 30 million, about 15 million to about 30 million, about 20 million to about 30 million, or about 25 million to about 30 million.
[0111] In some embodiments, the second predefined cutoff is selected based on a ranking of the predicted bindings of the test ligand to the second protein. In some embodiments, the ranking is based on the uncertainty in predicted binding of the test ligand to the second protein. In some embodiments, a second predefined cutoff is selected based on the top about 10 million test ligands. In some embodiments, a second predefined cutoff is selected based on the top about 1 million, about 5 million, about 10 million, about 15 million, about 25 million, or about 30 million test ligands. In some embodiments, a second predefined cutoff is selected based on the top about 1 million to about 30 million, about 1 million to about 25 million, about 1 million to about 20 million, about 1 million to about 15 million, about 1 million to about 10 million, or about 1 million to about 5 million. In some embodiments, a second predefined cutoff is selected based on the top about 5 million to about 30 million, about 10 million to about 30 million, about29MF-367788979Attorney Docket: 27576-200024015 million to about 30 million, about 20 million to about 30 million, or about 25 million to about 30 million.
[0112] In some embodiments, a third, fourth, or fifth predefined cutoff may be selected based on a similar method as those described for the first and second predefined cutoffs.
[0113] In some embodiments, identifying an output ligand comprises selecting the output ligand if the predicted binding of the output ligand to each of the two or more proteins is below a predefined energy. In some embodiments, the predefined energy may be selected based on a determined binding for a positive control ligand and protein pair. Like the predefined cutoffs, the predefined energy may be the same or different for each protein and may be determined based on chemical properties of the protein that effect binding of the protein to a ligand. In some embodiments, the predefined energy is about -8.0 Kcal / mol. In some embodiments, the predefined energy is about -8.0 Kcal / mol, about -8.5 Kcal / mol, about -9.0 Kcal / mol, about -9.5 Kcal / mol, about -10.0 Kcal / mol, about -10.5 Kcal / mol, or about -11.0 Kcal / mol. In some embodiments, the predefined energy is between about -8.0 Kcal / mol and about -11.0 Kcal / mol, about -8.0 Kcal / mol and about -10.5 Kcal / mol, about -8.0 Kcal / mol and about -10.0 Kcal / mol, about -8.0 Kcal / mol and about -9.5 Kcal / mol, about -8.0 Kcal / mol and about -9.0 Kcal / mol, about -8.0 Kcal / mol and about -8.5 Kcal / mol, or about -8.0 Kcal / mol and about -11.0 Kcal / mol. In some embodiments, the predefined energy is between about -8.5 Kcal / mol and about - 11.0 Kcal / mol, about -9.0 Kcal / mol and about -11.0 Kcal / mol, about -9.5 Kcal / mol and about - 11.0 Kcal / mol, about -10.0 Kcal / mol and about -11.0 Kcal / mol, or about -10.5 Kcal / mol and about -11.0 Kcal / mol.
[0114] In some embodiments, identifying the output ligand comprises selecting an output ligand based on a combined predicted binding. In some embodiments, the combined predicted binding is the output of the machine learning model at block 112. In some embodiments, the combined predicted binding is determined from the output of the machine learning models at block 112. In some embodiments, the predicting binding may relate to a mean predicted binding of the test ligand to the two or more proteins, a sum predicted binding of the test ligand to the two or more proteins, or a median predicted binding of the test ligand to the two or more proteins.
[0115] In some embodiments, identifying the output ligand comprises selecting the output ligand if the combined predicted binding is above a first predetermined cutoff. In some embodiments, the first predetermined cutoff is any of the predetermined cutoffs described here. In some30MF-367788979Attorney Docket: 27576-2000240embodiments, the first predefined cutoff is selected based on a ranking of the combined predicted bindings of the test ligands. In some embodiments, the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the combined predicted bindings. In some embodiments, a first predefined cutoff is selected based on the top about 1 million, about 5 million, about 10 million, about 15 million, about 25 million, or about 30 million test ligands as ranked by the combined predicted bindings. In some embodiments, a first predefined cutoff is selected based on the top about 1 million to about 30 million, about 1 million to about 25 million, about 1 million to about 20 million, about 1 million to about 15 million, about 1 million to about 10 million, or about 1 million to about 5 million as ranked by the combined predicted bindings. In some embodiments, a first predefined cutoff is selected based on the top about 5 million to about 30 million, about 10 million to about 30 million, about 15 million to about 30 million, about 20 million to about 30 million, or about 25 million to about 30 million as ranked by the combined predicted bindings.
[0116] In some embodiments, the output ligand-binds to each of the two or more proteins in different orientations. In certain aspects, -the output ligand-sits in the binding site of-a first protein differently and / or has a different binding mode-as compared to how the ligand sits in the binding site of a second protein. In some embodiments, different aspects of the output ligand, such as chemical functional groups, moieties, and / or atoms, -are involved in binding to the first and second proteins, wherein there is some degree of overlap of aspects of the output ligand involved in binding to the first and second proteins. There may be a different sets of interactions between atoms of the ligand and the first protein as compared to interaction with the second protein The output ligand may bind to each of the two or more proteins with at least partially overlapping aspects of the ligand involved with binding domainsE. Validation methods
[0117] In some embodiments, the methods further comprise determining binding (e.g. calculating binding) to the two or more proteins for the output ligand. In some embodiments, the methods may comprise determining binding to the two or more proteins for plurality of output ligands in order to reduce the number of output ligands that proceed to additional testing methods in a drug discovery process.
[0118] In some embodiments, determining binding to the two or more proteins for the output ligand may be performed using any of the methods for determining binding as described herein.31MF-367788979Attorney Docket: 27576-2000240In some embodiments, two or more orthogonal computational methods can be used to determine binding. For example, binding of the output ligand to the two or more proteins may be calculated using a Wscore, SiteMap, and / or Glide. Determining binding between the output ligand and the two or more proteins using two or more orthogonal computational methods provides additional evidence the ligand will bind to the two or more proteins in vitro and in vivo. In some embodiments, the methods may comprise synthesizing the output ligand and determining binding of the output ligand to the two or more proteins experimentally as described herein.
[0119] As the number of output ligands will likely be much lower than the number of ligands in the plurality of test ligands, confirming and / or validating the predicted bindings with the same or an orthogonal method will likely be practical wherein testing all the test ligand would not be. Accordingly, the methods described herein increase efficiency by decreasing the need to determine binding for all the test ligands in the plurality of ligands. In some embodiments, the determined binding for the output ligands may be used to validate the output ligands as capable of binding to the two or more proteins. In some embodiments, the determined bindings for the output ligands may be used for an additional round of active learning for the machine learning models described here.
[0120] In some embodiments, the determined binding of an output ligand to the two or more proteins may be compared to a positive control ligand and protein binding combination. The positive control may be selected based on the two or more proteins. For example, determined binding of a ligand to the two or more proteins may be compared to binding between a known ligand capable of binding to the first protein of the two or more proteins and a second protein of the two or more proteins. An output ligand may be validated if the determined binding of the output ligand to the two or more proteins is as strong or stronger than the positive controls.
[0121] In some embodiments, the methods comprise calculating an absolute binding free energy (ABFE) of the output ligand. In some embodiments, the output ligand may be identified based on an ABFE of the test ligand to the two or more proteins. ABFE is a method known in the art to predict binding affinities and potential for use of ligands in drug discovery pipelines, Aldega et a, Accurate Prediction of Ligand Binding Free Energies in a Large Dataset of Protein-Ligand Complexes with Absolute Binding Free Energy Calculations, Journal of Chemical Information and Modeling (2016), Mobley and Gilson., Absolution Alchemical Free Energy Calculations for Ligand Binding: A Beginners Guide, Journal of Computational Chemistry (2017).32MF-367788979Attorney Docket: 27576-2000240
[0122] In some embodiments, calculating ABFE comprises simulating a transition of the test ligand and a protein of the two or more proteins from an unbound state to a bound state and computing a free energy difference between the unbound and bound state. In some embodiments, calculating ABFE comprises sampling free energy.
[0123] Like determining binding with other methods described herein, calculating ABFE is computationally expensive and slow making it impractical to calculate ABFE for all ligands in the plurality of ligands and / or all of the test ligands. By predicting binding to identify output ligands using the methods described herein, the methods can help prioritize ligands for ABFE calculation.
[0124] In some embodiments, the method comprises experimental validation of the output ligand. In some embodiments, the methods comprise experimentally validating the output ligand is capable of binding to two or more proteins. In some embodiments, the methods comprise experimentally validating the output ligand can modulate biochemical activity of each protein. In some embodiments, the methods comprise experimentally validating that the output ligand is capable of binding to two or more proteins.
[0125] In some embodiments, the experimental validation comprises a cell-based assay, as described herein. In some embodiments, experimental validation comprises pharmacological validation. Pharmacological validation can be used to show the output ligand can achieve the functional synergistic effect in a cell-based assay.
[0126] In some embodiments, the experimental validation comprises a cell-based assay. In some embodiments, the cell-based assay is selected from the group consisting of a cell viability assay, a cell growth assay, a cell proliferation assay, a growth inhibition assay, an ELISA assay, and a metabolic assay. In some embodiments, a cell-based assay comprises treatment of the cells in the cell-based assay to model one or more disease or trait conditions. In some embodiments, the cellbased assay comprises measuring a response in cells that have been subjected to a condition modeling one or more disease or trait conditions. In some embodiments, the cell-based assay comprises condition testing one or more conditions associated with a trait. For example, condition testing for type 2 diabetes may comprise testing of cellular response to high glucose environments. In some embodiments, the cell-based assay can be performed in cells, in tissue, or in one or more model organism.33MF-367788979Attorney Docket: 27576-2000240
[0127] The experimental validation techniques encompass many read-outs such as to obtain a measurable or observable metric (e.g., a metric associated with a characteristic). In some embodiments, the read-out is cell growth inhibition. In some embodiments, the read-out may be cell viability, apoptosis, cell death, or cytotoxicity. In some embodiments, the read-out may be cell growth inhibition, pro- or anti-apoptotic activity, inhibition or stimulation of cellular stress responses, such as by glucose or reactive oxygen species, modulation of glucose metabolism, modulation of insulin-dependent metabolism, production or inhibition of production of disease- related polypeptides, or cytokine production, immune cell co-culture toxicity assay changes, or mitochondrial activity changes.
[0128] In some embodiments, the read-out may be a change in a phenotype consistent with treating or preventing a human condition. In some embodiments, the phenotype may be cell viability, apoptosis, cell death, or cytotoxicity. In some embodiments, the read-out may be cell growth inhibition, pro- or anti-apoptotic activity, inhibition or stimulation of cellular stress responses, such as by glucose or reactive oxygen species, modulation of glucose metabolism, modulation of insulin-dependent metabolism, production or inhibition of production of disease-related polypeptides, or cytokine production, immune cell co-culture toxicity assay changes, or mitochondrial activity changes.III. Methods of training a machine learning model to predict binding
[0129] Also provided herein are methods of training a machine learning model to predict binding of a test ligand to two or more proteins (e.g. combined machine learning model). The machine learning models may be any of the machine learning models described herein. The methods comprise generating training data comprising estimated binding of training ligands to two or more proteins, using the methods described here, and training a machine learning model to predict binding of a test ligand to the two or more proteins. The methods may comprise training the machine learning model to output a predicted binding of a test ligand to two or more proteins. In some embodiments, the machine learning models may be multi dataset machine learning models as described herein.
[0130] In some embodiments, any artificial neural networks (ANNs) that may be utilized to learn deep levels of representations and abstractions from large amounts of data may be used for training. For example, the machine learning models may include ANNs, such as a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural34MF-367788979Attorney Docket: 27576-2000240network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth.
[0131] In some embodiments, the machine learning model may comprise a classification algorithm or function may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon.
[0132] In some embodiments, training multi dataset machine learning models may comprise multitask learning, transfer learning, ensemble methods, and domain adaptation as described herein. In some embodiments, training multi dataset machine learning models may comprise feature transformation or identification of cross-domain features as described herein,
[0133] In some embodiments, the method may comprise active machine learning model training as described herein.IV. Additional usesA. Accessing co-druggability
[0134] In some embodiments, the methods described herein can be used to assess druggability of one or more gene combinations. In some embodiments, the gene combinations are between combinations from a combination of gene networks determined to have a synergistic effect on a trait. In some embodiments, the two or more proteins may be protein products from two or more gene combinations. In some embodiments, the methods comprise ranking gene combinations based on the identification of an output ligand for the protein products of the gene in the gene combination according to the methods described herein.
[0135] In some embodiments, the methods can be used to identify an output ligand capable of binding to two or more proteins from any combination of two or more protein products of two or more gene combinations that are part of a synergistic network pair. For example, two or more gene combinations may have been determined to have a synergistic effect on a trait using the methods described in U. S. Application 19 / 025,625 and the PCT / US2025 / 011900. Protein 35MF-367788979Attorney Docket: 27576-2000240products from each of the identified genes combinations can be used as the two or more proteins in the methods described herein. The gene combination may be ranked for co-druggabilitiy based on the identification of an output ligand capable of binding to the protein products of the gene combination.
[0136] In some embodiments, the methods comprise identifying a first output ligand capable of binding to a first set of two or more proteins and identifying a second output ligand capable of binding to a second set of two or more proteins using any of the methods described herein. In some embodiments, the first set of two or more proteins and the second set of two or more proteins have been selected because the first set of two or more proteins and the second set of two or more proteins are sets of protein products from two gene combinations identified as having a synergistic effect on a trait. In some embodiments, the methods comprise comparing the quantity and / or quality of the first output ligand and second output ligand. In some embodiments, the methods comprise identifying a first, second, third, fourth and / or fifth output ligands for a first, second, third, fourth and / or fifth set of two or more proteins and comparing the output ligands to assess co-druggability of the sets of two or more proteins.
[0137] In some embodiments, the methods comprise identifying a first output ligand capable of binding to a first set of two or more proteins; the method comprising: by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determining binding of each training ligand in the training chemical structure data to each of two or more proteins in the first set of two or more proteins to generate training data; training, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the first set of two or more proteins; inputting at least a subset of the chemical structure data into the two or more machine learning models, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands; outputting from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins in the first set of two or more proteins; and identifying a first output ligand of the plurality of test ligands capable of binding to the two or more proteins in the first set of two or36MF-367788979Attorney Docket: 27576-2000240more proteins based on the predicted binding of the test ligand to the two or more proteins in the first set of two or more proteins.
[0138] In some embodiments, the methods comprise identifying a second output ligand capable of binding to a second set of two or more proteins; the method comprising: by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determining binding of each training ligand in the training chemical structure data to each of two or more proteins in the second set of two or more proteins to generate training data; training, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the second set of two or more proteins; inputting at least a subset of the chemical structure data into the two or more machine learning models, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligand; outputting from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins in the second set of two or more proteins; and identifying a second output ligand of the plurality of test ligands capable of binding to the two or more proteins in the first set of two or more proteins based on the predicted binding of the test ligand to the two or more proteins in the second set of two or more proteins.B. Designing a therapeutic
[0139] The output ligand according the methods described herein may be used as a therapeutic compound or a therapeutic compound may be derived therefrom using medicinal chemistry methods known in the art.
[0140] In some embodiments, the methods described herein can be used to design a therapeutic comprising an output ligand capable of binding to two or more proteins. In some embodiments, the output ligand has been identified using any of the methods described herein. In some embodiments, the methods for designing a therapeutic comprise performing lead optimization methods known in the art on the output ligand. In some embodiments, the methods comprise optimizing the output ligand to increase stability and / or therapeutic efficacy of the output ligand.37MF-367788979Attorney Docket: 27576-2000240Also provided herein is therapeutic designed by identifying an output ligand capable of binding to two or more proteins using the methods described herein.V. Computing Devices and Systems
[0141] In certain aspects, provided herein are computer implemented methods for identifying an output ligand capable of binding to two or more proteins. For example, the methods described in FIG. 1, FIG. 2 and FIG. 3 and accompanying embodiments, can be implemented using one or more suitable computing devices capable of displaying a user interface to a user and recording and / or transmitting user inputs to a user interface. The computer implemented methods may be executed using the systems described herein.
[0142] In certain aspects, provided herein are systems configured for performing the methods, such as the computer implemented methods provided herein, or aspects thereof. In some embodiments, the systems may comprise one or more components, connected directly (such as by hardware) and / or indirectly (such as via wireless connection and / or data transfer). In some embodiments, the components of the system may communicate with one another using a network, such as the internet.
[0143] In some embodiments, the system comprises: one or more processors; and memory storing one or more programs, the one or more programs configured to be executed by one or more processors, the one or more programs including instructions for performing a method disclosed herein, or an aspect thereof. For example, in some embodiments, the instructions comprise information for performing a step of identifying an output ligand capable of binding to two or more proteins according to the disclosure provided herein.
[0144] In some embodiments, the system comprises one or more processors; and memory storing one or more programs, the one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the system to: receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; train, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a 38MF-367788979Attorney Docket: 27576-2000240test ligand to a protein of the two or more proteins; input the at least a subset of the chemical structure data into the two or more machine learning models, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands; output from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; and identify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
[0145] In some embodiments, the system comprises one or more processors; and memory storing one or more programs, the one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the system to: receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; train, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins; input the at least a subset of the chemical structure data into the machine learning model, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands; output from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; and identify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
[0146] In some embodiments, the system comprises one or more machine learning models. In some embodiments, the instructions comprise running a machine learning model, such as a machine learning model trained to predict binding of a test ligand to one or more proteins based on chemical structure data for ligand.
[0147] In some embodiments, the system is the exemplary system in FIG. 4, comprising computer system 400. As example and not by way of limitation, computing system 400 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer39MF-367788979Attorney Docket: 27576-2000240systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, computing system 400 may include one or more computing systems 400; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. In some embodiments, the computing system 400 includes a processor 402, memory 404, storage 406, an input / output (I / O) interface 408, a communication interface 410.
[0148] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In certain embodiments, processor 402 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions, processor 402 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 404, or storage 406; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 404, or storage 406. In certain embodiments, processor 402 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 402 including any suitable number of any suitable internal caches, where appropriate. As an example, and not by way of limitation, processor 402 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs).Instructions in the instruction caches may be copies of instructions in memory 404 or storage 406, and the instruction caches may speed up retrieval of those instructions by processor 402.
[0149] Data in the data caches may be copies of data in memory 404 or storage 406 for instructions executing at processor 402 to operate on; the results of previous instructions executed at processor 402 for access by subsequent instructions executing at processor 402 or for writing to memory 404 or storage 406; or other suitable data. The data caches may speed up read or write operations by processor 402. The TLBs may speed up virtual-address translation for processor 402. In some embodiments, processor 402 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 402 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 402 may include one or more arithmetic logic units (ALUs); be a multi-core processor;40MF-367788979Attorney Docket: 27576-2000240or include one or more processors 402. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0150] In some embodiments, memory 404 includes main memory for storing instructions for processor 402 to execute or data for processor 402 to operate on. As an example, and not by way of limitation, computing system 400 may load instructions from storage 406 or another source (such as, for example, another computing system 400) to memory 404. Processor 402 may then load the instructions from memory 404 to an internal register or internal cache. To execute the instructions, processor 402 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 402 may write one or more results (which may be intermediate or final results) to the internal register or internal cache. Processor 402 may then write one or more of those results to memory 404. In certain embodiments, processor 402 executes only instructions in one or more internal registers or internal caches or in memory 404 (as opposed to storage 406 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 404 (as opposed to storage 406 or elsewhere).
[0151] In some embodiments, the one or more processors, e.g. 402, may comprise hardware processor such as a central processing unit (CPU), a graphical processing unit (GPU), a general-purpose processing unit, or other computing platform such as a cloud-based platform. The processor may be comprised of any of a variety of suitable integrated circuits, microprocessors, logic devices, field programmable gate arrays (FGPAs) and the like. In some embodiments, the reference to a processor may be to other types of integrated circuits and logic devices. The processor may have any suitable data operation capability.
[0152] In some embodiments, the storage, e.g. 406, can be any suitable device that provides storage, such as an electric, magnetic, or optical memory device including RAM, cache, hard drive, or removable disk. In some embodiments, storage can be in the form of an external computing cloud. In certain embodiments, storage 406 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 406 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 406 may include removable or non-removable (or fixed) media, where appropriate. Storage 406 may be internal or external to computing system 400, where appropriate. In certain41MF-367788979Attorney Docket: 27576-2000240embodiments, storage 406 is non-volatile, solid-state memory. In certain embodiments, storage 306 includes read-only memory (ROM ). Where appropriate, this ROM may be mask- programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 406 taking any suitable physical form. Storage 406 may include one or more storage control units facilitating communication between processor 402 and storage 406, where appropriate. Where appropriate, storage 406 may include one or more storages 406. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.
[0153] In some embodiments, the systems may comprise a user device (e.g. VO Interface 408 and or Communication interface 410). The user device may be a computing device configured to interface with various components of the system to control one or more tasks, cause one or more actions to be performed or effectuate other operations. In some embodiments, a user device may be a desktop computer, server, mobile computer, smart device, wearable device, cloud computing platform. In some embodiments, the user device may include one or more processors, memory, communications components, display components, audio capture / output devices, image capture components, or other components, or combinations thereof. The user device may include any type of wearable device, mobile terminal, fixed terminal, or other device.
[0154] In some embodiments, VO interface 408 includes hardware, software, or both, providing one or more interfaces for communication between computing system 400 and one or more I / O devices. Computing system 400 may include one or more of these VO devices, where appropriate. One or more of these I / O devices may enable communication between a person and the computing system 400. As an example, and not by way of limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device or a combination of two or more of these. An VO device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces 406 for them. Where appropriate, VO interface 408 may include one or more device or software drivers enabling processor 402 to drive one or more of these VO devices. I / O interface 408 may include one or more I / O interfaces 406, where appropriate. Although this disclosure describes and illustrates a particular VO interface, this disclosure contemplates any suitable I / O interface.42MF-367788979Attorney Docket: 27576-2000240
[0155] In some embodiments, communication interface 410 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between the computing system 400 and one or more other computer systems 400 or one or more networks. As an example, and not by way of limitation, communication interface 410 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 410 for it.
[0156] In some embodiments, bus 412 includes hardware, software, or both coupling components of computing system 400 to each other. As an example, and not by way of limitation, bus 412 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 412 may include one or more buses 412, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0157] Also described herein are computer-readable non -transitory storage medium comprising instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; train, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins; input the at least a subset of the chemical structure data into the two or more machine43MF-367788979Attorney Docket: 27576-2000240learning models, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands; output from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; and identify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.[0158j Also described herein are computer-readable non-transitory storage media comprising instructions that, when executed by one or more processors of one or more computing devices, cause the one or more processors to: receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data; determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data; train, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins; input the at least a subset of the chemical structure data into the machine learning model, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands; output from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; and identify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
[0159] In some embodiments, the computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.44MF-367788979Attorney Docket: 27576-2000240EXEMPLARY EMBODIMENTS
[0160] The following exemplary embodiments are representative of some aspects of the invention:
[0161] Embodiment 1: A method of identifying an output ligand capable of binding to two or more proteins; the method comprising:by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data:determining binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;training, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins;inputting at least a subset of the chemical structure data into the two or more machine learning models, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands;outputting from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; andidentifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding of the test ligand to the two or more proteins.
[0162] Embodiment 2. The method of embodiment 1, wherein identifying the output ligand comprises selecting the output ligand if the predicted binding of the test ligand is above a first predetermined cutoff a first protein of the two or more proteins and above a second predetermined cutoff for a second protein of the two or more proteins.
[0163] Embodiment 3. The method of embodiment 2, wherein the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the predicted binding of the test ligands to the first protein.45MF-367788979Attorney Docket: 27576-2000240
[0164] Embodiment 4. The method of embodiment 2 or 3, wherein the second predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the predicted binding of the test ligands to the second protein.
[0165] Embodiment 5. The method of any of embodiments 1-4, wherein identifying the output ligand comprises selecting the output ligand if the predicted binding of the output ligand to each of the two or more proteins is below a predefined energy.
[0166] Embodiment 6. The method of embodiment 5, wherein the predefined energy is about -8.0 Kcal / mol.
[0167] Embodiment 7. A method of identifying an output ligand capable of binding to two or more proteins; the method comprising:by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data;determining binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;training, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins; inputting at least a subset of the chemical structure data into the machine learning model, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands;outputting from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; and identifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
[0168] Embodiment 8. The method of embodiment 7, wherein the predicted binding comprises a first predicted binding for the ligand to a first protein of the two or more proteins and a second predicted binding for the ligand to a second protein of the two or more proteins.46MF-367788979Attorney Docket: 27576-2000240
[0169] Embodiment 9. The method of embodiment 8, wherein identifying the output ligand comprises selecting the output ligand if the first predicted binding is above a first predetermined cutoff and the second predicted binding is above a second predetermined cutoff.
[0170] Embodiment 10, The method of embodiment 8, wherein the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the first predicted bindings.
[0171] Embodiment 11. The method of embodiment 8 or 9, wherein the second predetermined cutoff is selected based on about 10 million test ligands as ranked by the second predicted bindings.
[0172] Embodiment 12. The method of embodiment 7, wherein the predicted binding comprises a combined predicted binding for the ligand to a first protein of the two or more proteins and to a second proteins of the two or more proteins.
[0173] Embodiment 13, The method of embodiment 12, wherein identifying the output ligand comprises selecting the output ligand if the combined predicted binding is above a first predetermined cutoff.
[0174] Embodiment 14. The method of embodiment 13, wherein the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the combined predicted bindings.
[0175] Embodiment 15. The method of any of embodiments 1-14, comprising determining binding to the two or more proteins for the output ligand.
[0176] Embodiment 16. The method of any of embodiments 1-15, wherein determining binding comprises computing an empirical docking score in silico using the chemical structure data and protein structure data and / or an affinity screening method to measuring binding of the training ligand to the protein in vitro.
[0177] Embodiment 17. The method of any of embodiments 1-16, wherein the protein structure data comprises a 2D representation of a 3D structure of a protein.
[0178] Embodiment 18, The method of claim 17, wherein the 3D representation comprises Xray, NMR, cryo-EM, or deep learning structural models data.
[0179] Embodiment 19. The method of any of embodiments 16-18, wherein the empirical docking score comprises a Wscore, a SiteMap score, a Glide score, a DOCK3.7 score or a DOCK3.8 score.47MF-367788979Attorney Docket: 27576-2000240
[0180] Embodiment 20. The method of any of embodiments 16-19, wherein the affinity screening comprises array-based screening technology.
[0181] Embodiment 21. The method of any of embodiments 1-20, comprising calculating an absolute binding free energy (ABFE) of the output ligand.
[0182] Embodiment 22. The method of any of embodiments 1-21, wherein identifying the output ligand is based on an ABFE of the test ligand to the two or more proteins.
[0183] Embodiment 23. The method of embodiment 22, wherein calculating ABFE comprises simulating a transition of the test ligand and a protein of the two or more proteins from an unbound state to a bound state and computing a free energy difference between the unbound and bound state.
[0184] Embodiment 24. The method of embodiment 22, wherein calculating ABFE comprises sampling free energy of one or more intermediate states between the unbound state and the bound state.
[0185] Embodiment 25. The method of any of embodiments 1-24, wherein the plurality of ligands comprises between about 48 million ligands and 1 billion ligands.
[0186] Embodiment 26. The method of any of embodiments 1-25, wherein the plurality of training ligands comprises about 0.1% of the plurality of ligands.
[0187] Embodiment 27. The method of any one of embodiments 1-26, wherein the plurality of training ligands comprises about 48,000 of the plurality of ligands.
[0188] Embodiment 28. The method of any one of embodiments 1-27, wherein the plurality of training ligands comprises about 100,000 of the plurality of ligands.
[0189] Embodiment 29. The method of any of embodiments 1-28, wherein selecting training chemical structure data comprises selecting chemical structure data corresponding to the training ligands.
[0190] Embodiment 30. The method of any of embodiments 1-29, wherein selecting the training ligands is based on chemical diversity of the plurality of ligands.
[0191] Embodiment 31. The method of any of embodiments 1-30, wherein the training ligands represent the about 0.1% most diverse structures in the plurality of ligands.
[0192] Embodiment 32. The method of embodiment 31, wherein the 0.1% most diverse structures in the plurality of ligands are selected based on variation in chemical structures of the48MF-367788979Attorney Docket: 27576-2000240ligands in the plurality of ligands or variation in chemical fingerprints in the chemical structures of the ligands in the plurality of ligands.
[0193] Embodiment 33. The method of any of embodiments 1-32, wherein the two or more proteins have sequence similarity between about 10% and 30%,
[0194] Embodiment 34. The method of any of embodiments 1-33, wherein the two or more proteins have consensus binding site sequence similarity between about 10% and 50%.
[0195] Embodiment 35. The method of any of embodiments 1-34, wherein binding site volumes of the two or more proteins is not predictive of the predicted binding to the test ligands.
[0196] Embodiment 36. The method of any of embodiments 1-35, wherein the output ligand is capable of binding to each of the two or more proteins in different orientations.
[0197] Embodiment 37. The method of any of embodiments 1-36, wherein the two or more proteins have been determined to have a synergistic effect on a disease.
[0198] Embodiment 38. The method of any of embodiments 1-37, wherein the two or more proteins have been identified as a gene combination with a synergistic effect on a trait.
[0199] Embodiment 39. The method or any of embodiments 1-38, wherein the machine learning model comprises an AQ / DC, a graphical convolution neural network (GCNN), and or a deep neural network (DNN).
[0200] Embodiment 40. The method of any of embodiments 1-39, wherein training the machine learning model comprises performing a 5-fold cross validation scheme to select hyperparameters for the machine learning model.
[0201] Embodiment 41. The method of embodiment 40, wherein the hyperparameters comprise model algorithm, number of layers, number of training epochs, and a normalization scheme applied to a raw response variable.
[0202] Embodiment 42. The method of any of embodiments 1-41, wherein the machine learning models comprises a multi dataset machine learning model.
[0203] Embodiment 43. The method of any of embodiments 1-42, wherein the chemical structures comprise a 2 dimensional (2D) representation of the plurality of ligands.
[0204] Embodiment 44. The method of embodiment 43, wherein the 2D representation is a Simplified Molecular Input Line Entry System (SMILES) representation or a morgan fingerprint.
[0205] Embodiment 45. The method of any of embodiments 1-44, wherein the plurality of ligands comprises chemical compounds.49MF-367788979Attorney Docket: 27576-2000240
[0206] Embodiment 46. A system comprising:one or more processors; anda memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands: select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data:determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;train, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins;input the at least a subset of the chemical structure data into the two or more machine learning models, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands;output from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; andidentify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins,
[0207] Embodiment 47, A system comprising:one or more processors; anda memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data;50MF-367788979Attorney Docket: 27576-2000240determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;train, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins: input the at least a subset of the chemical structure data into the machine learning model, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands;output from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; andidentify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.EXAMPLESExample 1[0208j This example demonstrates a method for identifying an output ligand capable of biding to FDFT1 and MAPK3 according to embodiments described herein.
[0209] FDFT1 and MAPK3 have less than 30% sequence similarity and the binding sites are unrelated to each other. FIG. 5A and FIG. 5B shows the structure of MAPK3 and FDFT1 with an exemplary compound (ligand) docked pose highlighted as spheres. The consensus binding site similarity was calculated as 11%. The binding site volume of FDFT1 was measured as 1135.53 A3and the binding site volume of MAPK3 was measured as 1051.64 A3. A binding site D-score was calculated for FDFT1 and MAPK3 as 1.105 and 1.076 respectively. The D-score was calculated using tools in the Schrödinger tool suite.
[0210] A first machine learning model was trained to predict binding of a ligand to FDFT1. A second machine learning model was trained to predict binding of a ligand to MAPK3. Both machine learning models were trained with a subset of training ligands from a first library comprised 48 million compounds (48M library) and estimated binding scores for the each of the training ligands to either FDFT1 or MAPK3 respectively. The binding scores were calculated using GlideSP available as part of the Schrödinger tools suite.51MF-367788979Attorney Docket: 27576-2000240
[0211] A third machine learning model was trained to predict binding of a ligand to FDFT1. A fourth machine learning model was trained to predict binding of a ligand to MAPK3. Both machine learning models were train with a subset of training ligands from a second library comprised 3 billion compounds (3B library) and calculated binding scores for the each of the training ligands to either FDFT1 or MAPK3 respectively. The binding scores were calculated using GlideSP available as part of the Schrödinger tools suite.
[0212] The predicted binding from the trained machine learning models of each ligand to both FDFT1 and MAPK3 is shown in FIG. 6. The ligands represented in green were in the 48M (also included in the 3B library). The ligands represented in orange were from the 3B library and were used for training the third and fourth machine learning models. The ligands represented in blue were only in the set of ligands in the 3B library test set.
[0213] FIG. 7 shows 4 ligands identified as capable of binding to FDFT1 and MAPK3 with a predicted binding score of less than -9.5 Kcal / mol. As shown in FIG. 8, it was found that the same ligand exhibited two different binding modes in binding to FDFT1 and MAPK3 respectively.
[0214] Table 1 shows the number of ligands meeting different binding score cutoffs for the predicted binding between FDFT1, MAPK3 and each of the ligands in the 48M library and 3B library. It was found that using the 3B library, 7285 ligands were found with predicted binding below -9 Kcal / mol for BOTH proteins. Of those only 7 were used in training the machine learning models.
[0215] Together these results demonstrated that to identify compounds that bind to both molecules with a score of less than -9 Kcal / mol and -10 Kcal / mol, it was necessary to predict binding using the 3B library.Table 1:Test TrainingLigand LigandLibrary 3B 48M 3B 48MCutoff < -10 < -9 < -10 < -9 < -10 < -9 < -10 < -9 (Kcal / mole)FDFT1 2164 38987 218 4470 4 32 0 0 MAPK3 4086 71982 89 2079 2 43 0 0Both 39 7285 0 105 0 7 0 052MF-367788979Attorney Docket: 27576-2000240
[0216] Additional approaches were used to further characterize the binding of the ligands to FDFT1 and MAPK3. A computational method was used to estimate the strength of the interaction between the ligand and each of the proteins (data not shown). The method quantified the absolute binding free energy (ABFE) change associated with the binding process, providing a rigorous thermodynamic measure of binding affinity. The ABFE calculation involved simulating the transition of the small molecule from an unbound state in solution to its bound state within the protein’s binding pocket. Molecular dynamics (MD) simulations, often enhanced with alchemical transformations, were employed to compute the free energy differences between these states. Key techniques included free energy perturbation (FEP) and thermodynamic integration (TI), which rely on sampling intermediate states (lambda states) to estimate the free energy change accurately. The methods can be found in Aldeghi et al. Accurate Prediction of Ligand Binding Free Energies in a Large Dataset of Protein-Ligand Complexes with Absolute Binding Free Energy Calculations. J. Chemical Information and Modeling (2016) and Mobley and Gilson, Absolute Alchemical Free Energy Calculations for Ligand Binding: A Beginner’s Guide, J. of Computational Chemistry (2017).Example 2
[0217] The methods described in example 1 and the 48 million compound library were used to estimate binding to ligands for 6 additional protein pairs. As shown in table 2, like FDFT1 and MAPK3 the sequence similarity, binding site similarity and binding site volume differed greatly between the proteins in each pair. The number of overlapping binding hits as predicted by the machine learning models is shown in Table 2 as Overlapping hit count.Table 2Protein 1 Protein 2 Consensus Full Protein Protein Binding Overlappin binding site sequence 1 2 site D- g hit count similarity similarity binding binding score (score <9) (%) site site (proteinl,volume volume protein2)(A3) (A3)CDC42 TNFR1 4 11 185.22 622.89 1.036, 01.210FDFT1 DUSP3 8 14 1134.53 664.73 1.105, 01.00653MF-367788979Attorney Docket: 27576-2000240FDFT1 MAPK3 11 20 1134.53 1051.64 1.105, 1071.076ADAM 17 AOC3 15 28 545.37 535.42 1.037, 91.078PAK1 PDK2 17 24 509.01 621.17 1.012, 43560.960ADAM 17 ALDH3 18 24 545.37 339.57 1.037, 17391Al 1.094 MAPK11 PRKAA 19 41 781.01 195.5 1.085, 233381.139
[0218] Below a threshold of binding site sequence similarity of 10%, there were no ligands identified that could simultaneously bind both proteins. In addition, there was no strict correlation between binding site similarity and the number of ligands that could simultaneously bind both proteins (hits). For example, the ADAM 17-AOC3 pair with sequence identity 15% produces 4356 hits while PAK1-PDK2 with sequence identity of 17% produces only 9 hits. In addition, the volume and size of the binding site was not correlated with the numbers of hits produced. For example, FDFT1 / MAPK3 with average pocket size of -1000 A3produced only 107 hits while ADAM 17 and AOC3 with average pocket size of -500 A3produced over 4000 hits.54MF-367788979
Claims
Attorney Docket: 27576-2000240CLAIMS1. A method of identifying an output iigand capable of binding to two or more proteins; the method comprising:by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data;determining binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;training, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins;inputting at least a subset of the chemical structure data into the two or more machine learning models, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands;outputting from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; andidentifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding of the test ligand to the two or more proteins.
2. The method of claim 1, wherein identifying the output ligand comprises selecting the output ligand if the predicted binding of the test ligand is above a first predetermined cutoff a first protein of the two or more proteins and above a second predetermined cutoff for a second protein of the two or more proteins.
3. The method of claim 2, wherein the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the predicted binding of the test ligands to the first protein.55MF-367788979Attorney Docket: 27576-20002404. The method of claim 2 or 3, wherein the second predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the predicted binding of the test ligands to the second protein.
5. The method of any one of claims 1-4, wherein identifying the output li gand comprises selecting the output ligand if the predicted binding of the output ligand to each of the two or more proteins is below a predefined energy.
6. The method of claim 5, wherein the predefined energy is about -8.0 Kcal / mol.
7. A method of identifying an output ligand capable of binding to two or more proteins; the method comprising:by one or more computing devices comprising one or more processors and memory: receiving chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; selecting training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data;determining binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;training, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins; inputting at least a subset of the chemical structure data into the machine learning model, wherein the at least a subset of chemical structure data comprises chemical structures for the plurality of test ligands;outputting from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; and identifying an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.56MF-367788979Attorney Docket: 27576-20002408. The method of claim 7, wherein the predicted binding comprises a first predicted binding for the ligand to a first protein of the two or more proteins and a second predicted binding for the ligand to a second protein of the two or more proteins.
9. The method of claim 8, wherein identifying the output ligand comprises selecting the output ligand if the first predicted binding is above a first predetermined cutoff and the second predicted binding is above a second predetermined cutoff.
10. The method of claim 8, wherein the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the first predicted bindings.
11. The method of claim 8 or 9, wherein the second predetermined cutoff is selected based on about 10 million test ligands as ranked by the second predicted bindings.
12. The method of claim 7, wherein the predicted binding comprises a combined predicted binding for the ligand to a first protein of the two or more proteins and to a second proteins of the two or more proteins.
13. The method of claim 12, wherein identifying the output ligand comprises selecting the output ligand if the combined predicted binding is above a first predetermined cutoff.
14. The method of claim 13, wherein the first predetermined cutoff is selected based on a top about 10 million test ligands as ranked by the combined predicted bindings.
15. The method of any one of claims 1-14, comprising determining binding to the two or more proteins for the output ligand.
16. The method of any one of claims 1-15, wherein determining binding comprises computing an empirical docking score in silico using the chemical structure data and protein structure data and / or an affinity screening method to measuring binding of the training ligand to the protein in vitro.57MF-367788979Attorney Docket: 27576-200024017. The method of any one of claims 1-16, wherein the protein structure data comprises a 2D representation of a 3D structure of a protein.
18. The method of claim 17, wherein the 3D representation comprises Xray, NMR, cryo-EM, or deep learning structural model data.
19. The method of any one of claims 16-18, wherein the empirical docking score comprises a Wscore, a SiteMap score, a Glide score, a DOCK3.7 score or a DOCK3.8 score.
20. The method of any one of claims 16-19, wherein the affinity screening comprises arraybased screening technology.
21. The method of any one of claims 1-20, comprising calculating an absolute binding free energy (ABFE) of the output ligand.
22. The method of any one of claims 1-21, wherein identifying the output ligand is based on an ABFE of the test ligand to the two or more proteins.
23. The method of claim 22, wherein calculating ABFE comprises simulating a transition of the test ligand and a protein of the two or more proteins from an unbound state to a bound state and computing a free energy difference between the unbound and bound state.
24. The method of claim 22, wherein calculating ABFE comprises sampling free energy of one or more intermediate states between the unbound state and the bound state.
25. The method of any one of claims 1-24, wherein the plurality of ligands comprises between about 48 million ligands and 1 billion ligands.
26. The method of any one of claims 1-25, wherein the plurality of training ligands comprises about 0.1% of the plurality of ligands.58MF-367788979Attorney Docket: 27576-200024027. The method of any one of claims 1-26. wherein the plurality of training ligands comprises about 48,000 of the plurality of ligands.
28. The method of any one of claims 1-27, wherein the plurality of training ligands comprises about 100,000 of the plurality of ligands.
29. The method of any one of claims 1-28, wherein selecting training chemical structure data comprises selecting chemical structure data corresponding to the training ligands.
30. The method of any one of claims 1-29, wherein selecting the training ligands is based on chemical diversity of the plurality of ligands.
31. The method of any one of claims 1-30, wherein the training ligands represent the about 0.1% most diverse structures in the plurality of ligands.
32. The method of claim 31, wherein the 0.1 % most diverse structures in the plurality of ligands are selected based on variation in chemical structures of the ligands in the plurality of ligands or variation in chemical fingerprints in the chemical structures of the ligands in the plurality of ligands.
33. The method of any one of claims 1-32, wherein the two or more proteins have sequence similarity between about 10% and 30%.
34. The method of any one of claims 1-33, wherein the two or more proteins have consensus binding site sequence similarity between about 10% and 50%.
35. The method of any one of claims 1-34, wherein binding site volumes of the two or more proteins is not predictive of the predicted binding to the test ligands.
36. The method of any one of claims 1-35, wherein the output ligand is capable of binding to each of the two or more proteins in different orientations.59MF-367788979Attorney Docket: 27576-200024037. The method of any one of claims 1-36, wherein the two or more proteins have been determined to have a synergistic effect on a disease.
38. The method of any one of claims 1-37, wherein the two or more proteins have been identified as a gene combination with a synergistic effect on a trait,39. The method or any one of claims 1-38, wherein the machine learning model comprises an AQ / DC, a graphical convolution neural network (GCNN), and or a deep neural network (DNN).
40. The method of any one of claims 1-39, wherein training the machine learning model comprises performing a 5-fold cross validation scheme to select hyperparameters for the machine learning model.
41. The method of claim 40, wherein the hyperparameters comprise model algorithm, number of layers, number of training epochs, and a normalization scheme applied to a raw¬ response variable.
42. The method of any one of claims 1-41, wherein the machine learning models comprises a multi dataset machine learning model.
43. The method of any one of claims 1-42, wherein the chemical structures comprise a 2 dimensional (2D) representation of the plurality of ligands.
44. The method of claim 43, wherein the 2D representation is a Simplified Molecular Input Line Entry System (SMILES) representation or a morgan fingerprint.
45. The method of any one of claims 1-44, wherein the plurality of ligands comprises chemical compounds.
46. A system comprising:60MF-367788979Attorney Docket: 27576-2000240one or more processors; anda memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data:determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;train, using the training data, two or more machine learning models, wherein each machine learning model is trained to predict binding of a test ligand to a protein of the two or more proteins;input the at least a subset of the chemical structure data into the two or more machine learning models, w'herein the subset of chemical structure data comprises chemical structures for the plurality of test ligands;output from the two or more machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to each of the two or more proteins; andidentify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.
47. A system comprising:one or more processors; anda memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:receive chemical structure data comprising chemical structures for a plurality of ligands, wherein the plurality of ligands comprises training ligands and test ligands; select training chemical structure data, comprising chemical structures for the plurality of training ligands from the chemical structure data;61MF-367788979Attorney Docket: 27576-2000240determine binding of each training ligand in the training chemical structure data to each of two or more proteins to generate training data;train, using the training data, a machine learning model, wherein the machine learning model is trained to predict binding of a test ligand to two or more proteins: input the at least a subset of the chemical structure data into the machine learning model, wherein the subset of chemical structure data comprises chemical structures for the plurality of test ligands;output from the machine learning models, predicted binding of each test ligand in the subset of the chemical structure data to the two or more proteins; andidentify an output ligand of the plurality of test ligands capable of binding to the two or more proteins based on the predicted binding the test ligands to the two or more proteins.62MF-367788979