Ranking candidate ligands by competitive co-folding

EP4744050A1Pending Publication Date: 2026-05-20ISOMORPHIC LABS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
ISOMORPHIC LABS LTD
Filing Date
2024-10-28
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Current methods for predicting binding affinities of ligands for proteins are costly and time-consuming, often requiring abundant labeled training data which is scarce due to the expense and duration of experiments.

Method used

A system utilizing a co-folding neural network to generate co-folding data for pairs of candidate ligands, characterizing their relative binding affinity for a protein by processing joint 3D structures of the protein and ligands, without requiring explicit binding affinity data.

Benefits of technology

This approach enables efficient ranking of candidate ligands based on predicted binding affinities, reducing the need for extensive experimental data and improving computational efficiency, while achieving higher accuracy in ligand ranking compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024080476_08052025_PF_FP_ABST
    Figure EP2024080476_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a ranking of candidate ligands from a set of candidate ligands that is indicative of respective predicted binding affinities of each of the candidate ligands for a protein. In one aspect, a method comprises: generating respective co-folding data for each of a plurality of pairs of candidate ligands, comprising, for each pair of candidate ligands: processing data defining the pair of candidate ligands and the protein using a co-folding neural network to generate data defining a joint three-dimensional (3D) structure of the pair of candidate ligands and the protein; and generating the co-folding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein; and processing the co-folding data to generate the ranking.
Need to check novelty before this filing date? Find Prior Art

Description

RANKING CANDIDATE LIGANDS BY COMPETITIVE CO-FOLDINGCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Application No. 63 / 594,320, filed on October 30, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND

[0002] This specification relates to ranking candidate ligands based on a respective binding affinity of each candidate ligand for a protein using or more machine learning models.

[0003] Predictions can be made using machine learning models. Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY

[0004] This specification describes a system implemented as computer programs on one or more computers in one or more locations that can generate a ranking of candidate ligands in a set of candidate ligands based on a respective predicted binding affinity of each of the candidate ligands for a protein.

[0005] A “protein” can be understood to refer to any biological molecule that is specified by one or more sequences (or “chains”) of amino acids. For example, the term protein can refer to a protein domain, e.g., a portion of an amino acid chain of a protein that can undergo protein folding nearly independently of the rest of the protein. As another example, the term protein can refer to a protein complex, i.e., that includes multiple amino acid chains that jointly fold into a protein structure.

[0006] A “ligand” can refer to a molecule or compound that binds to a target molecule, e.g., a protein. Ligands can include, e.g., inorganic molecules, organic molecules, proteins, biomolecules, and so forth.

[0007] A “multiple sequence alignment” (MSA) for an amino acid chain in a protein specifies a sequence alignment of the amino acid chain with multiple additional amino acid chains, e.g., from other proteins, e.g., homologous proteins. More specifically, the MSA can define a correspondence between the positions in the amino acid chain and corresponding positions in multiple additional amino acid chains. A MSA for an amino acid chain can be generated, e.g., by processing a database of amino acid chains using any appropriate computational sequence alignment technique, e.g., progressive alignment construction. The amino acid chains in the MSA can be understood as having an evolutionary relationship, e.g., where each amino acid chain in the MSA may share a common ancestor. The correlations between the amino acids in the amino acid chains in a MSA for an amino acid chain can encode information that is relevant to predicting the structure of the amino acid chain.

[0008] A “binding pocket” on a protein can refer to a specific three-dimensional cavity or crevice within the structure of the protein where a ligand can bind to the protein. The binding pocket can, in some cases, be understood as a "lock" that fits the shape and chemical properties of ligands that act as "keys" for the lock. In other cases, the ligand may initially not fit perfectly into the binding pocket, e.g., due to structural differences or slight mismatches in shape or chemical groups, but conformational changes during binding can cause the interaction between the ligand and the binding pocket to become more complementary and specific, e.g., as in induced-fit binding. Examples of binding pockets include, e.g., orthosteric binding pockets, allosteric binding pockets, and cryptic binding pockets.

[0009] A “binding affinity” of a ligand for a protein refers to the strength or degree of attraction between the ligand and the protein when they interact to form a complex.

[0010] According to one aspect there is provided a method performed by one or more computers, the method comprising: obtaining data identifying: (i) a protein, and (ii) a set of candidate ligands; generating respective co-folding data for each of a plurality of pairs of candidate ligands, wherein: each pair of candidate ligands comprises a respective first candidate ligand from the set of candidate ligands and a respective second candidate ligand from the set of candidate ligands; the respective co-folding data for each pair of candidate ligands characterizes a relative binding affinity for the protein of the respective first candidate ligand in the pair in comparison to the respective second candidate ligand in the pair; and generating the co-folding data for the pair of candidate ligands comprises: processing data defining the pair of candidate ligands and the protein using a co-folding neural network, in accordance with values of a set of co-folding neural network parameters, to generate data defining a joint three-dimensional (3D) structure of the pair of candidate ligands and theprotein; and generating the co-folding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein; and processing the respective co-folding data for each of the plurality of pairs of candidate ligands to generate a ranking of the candidate ligands from the set of candidate ligands that is indicative of respective predicted binding affinities of each of the candidate ligands for the protein.

[0011] In some implementations, for each pair of candidate ligands, generating the co-folding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein comprises: determining the co-folding data based on a difference between: (i) a distance of the first candidate ligand of the pair from a binding pocket of the protein in the joint 3D structure of the pair of candidate ligands and the protein, and (ii) a distance of the second candidate ligand of the pair from the binding pocket of the protein in the joint 3D structure of the pair of candidate ligands and the protein.

[0012] In some implementations, for each pair of candidate ligands, generating the co-folding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein comprises: determining a joint 3D structure of the first candidate ligand of the pair and the protein; determining a joint 3D structure of the second candidate ligand of the pair and the protein; and determining the co-folding data for the pair of candidate ligands based on the respective joint 3D structures of: (i) the pair of candidate ligands and the protein, (ii) the first candidate ligand of the pair and the protein, and (ii) the second candidate ligand of the pair and the protein.

[0013] In some implementations, determining the joint 3D structure of the first candidate ligand of the pair and the protein comprises: processing data defining the first candidate ligand of the pair and the protein using the co-folding neural network.

[0014] In some implementations, determining the joint 3D structure of the second candidate ligand of the pair and the protein comprises: processing data defining the second candidate ligand of the pair and the protein using the co-folding neural network.

[0015] In some implementations, determining the co-folding data for the pair of candidate ligands based on the respective joint 3D structures of: (i) the pair of candidate ligands and the protein, (ii) the first candidate ligand of the pair and the protein, and (ii) the second candidate ligand of the pair and the protein comprises: determining a measure of displacement of the first candidate ligand between the respective joint 3D structures of: (i) the first candidate ligand and the protein, and (ii) the pair of candidate ligands and the protein; determining a measure of displacement of the second candidate ligand between the respective joint 3D structures of: (i) the second candidate ligand and the protein, and (ii) the pair of candidateligands and the protein; and determining the co-folding data based on a difference between the measure of displacement of the first candidate ligand and the measure of displacement of the second candidate ligand.

[0016] In some implementations, processing the respective co-folding data for each of the plurality of pairs of candidate ligands to generate the ranking of the candidate ligands comprises: generating the ranking of the candidate ligands as a solution of a game characterized by a payoff matrix defined by the co-folding data for the plurality of pairs of candidate ligands.

[0017] In some implementations, the game is a two-player game where each player selects a respective candidate ligand, and wherein the payoff received by each player depends on a relative binding affinity for the protein of the candidate ligand selected by the player in comparison the candidate ligand selected by the other player.

[0018] In some implementations, processing the respective co-folding data for each of the plurality of pairs of candidate ligands to generate the ranking of the candidate ligands comprises: processing the co-folding data for each of the plurality of pairs of candidate ligands using a ranking machine learning model, in accordance with trained values of a set of ranking machine learning model parameters, to generate data defining the ranking of the candidate ligands.

[0019] In some implementations, the ranking machine learning model comprises a neural network.

[0020] In some implementations, the co-folding neural network has been trained on a set of training data that comprises a plurality of training examples that each include: (i) a training input that defines a protein and one or more ligands, and (ii) a target output that defines a joint 3D structure of the protein and the one or more ligands.

[0021] In some implementations, the data identifying the protein comprises data defining each of one or more amino acid sequences of the protein.

[0022] In some implementations, the data identifying each ligand comprises a text string defining a chemical structure of the candidate ligand.

[0023] In some implementations, one or more of the candidate ligands are small molecules.

[0024] In some implementations, the protein comprises an enzyme, receptor, or signaling protein that has been identified as being involved in a disease process.

[0025] In some implementations, the method further comprises: selecting one or more candidate ligands from the set of candidate ligands based on the ranking; and physically synthesizing the selected candidate ligands.

[0026] In some implementations, the method further comprises, for each of the selected candidate ligands, performing experiments using physically synthesized instances of the candidate ligand to determine one or more of: an absorption of the candidate ligand, a distribution of the candidate ligand, a metabolism of the candidate ligand, or an excretion of the candidate ligand.

[0027] In some implementations, the ranking of the candidate ligands ranks the candidate ligands from highest predicted binding affinity for the protein to lowest predicted binding affinity for the protein.

[0028] According to another aspect there is provided a system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the methods described herein.

[0029] According to another aspect there are provided one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the methods described herein.

[0030] According to another aspect there is provided a ligand that has been synthesized by performing the methods described herein.

[0031] According to another aspect there are provided one or more non-transitory computer storage media storing ligand data defining a ligand, wherein the ligand was selected from a set of candidate ligands by performing operations comprising: generating a ranking of candidate ligands in the set of candidate ligands that is indicative of respective predicted binding affinities of each of the candidate ligands for a protein by performing the methods described herein; and selecting the ligand from the set of candidate ligands based on the ranking.

[0032] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0033] Drug discovery can involve identifying specific molecules within the body that are involved in a disease process. These molecules are often proteins, such as enzymes, receptors, or signaling proteins, that play a key role in the disease's development or progression. A ligand, often a small molecule (e.g. molecular weight equal to or less than 1000 daltons), peptide, or antibody, can be selected to bind specifically to an identified target protein and modify its biological activity. When a drug that includes the ligand is administered to a patient, the ligand can bind to the target protein with high affinity and in doing so contribute to achieving a therapeutic effect in the patient. For instance, if the target protein is an enzyme involved in adisease process, the ligand can inhibit its activity, thus disrupting the disease pathway. More generally, the interaction between the ligand and the target protein can activate, inhibit, or alter the function of the target protein to achieve a therapeutic effect.

[0034] Therefore, identifying ligands with high (or low) binding affinity for a protein can be a crucial step in the process of drug discovery. (Identifying ligands with low binding affinities for a protein can be desirable, e.g., when that protein is not the intended target protein and the binding of the ligand to such an “off-target” protein may cause undesirable side effects). However, determining binding affinities of ligands for proteins, e.g., through computational simulations or physical experiments, can be expensive and time consuming.

[0035] The system described in this specification can generate a ranking of candidate ligands in a set of candidate ligands based on their binding affinity for a protein (e.g., for a particular binding pocket on a protein), without requiring that binding affinities be explicitly determined for the candidate ligands. Thus, the system can determine the ranking of the candidate ligands based on their binding affinities for the protein without explicitly determining a respective binding affinity for each protein. The system can use the ranking to efficiently select candidate ligands that have a high (or low) binding affinity for the protein. The system can then determine binding affinities (e.g., using computational methods or physical experiments) for only the selected candidate ligands, rather than for every candidate ligand in the set of candidate ligands, thus enabling more efficient and targeted use of resources.

[0036] The system can generate a ranking of a set of candidate ligands using a co-folding neural network that has been trained to predict the three-dimensional (3D) structure of a complex that includes a protein and at least one ligand. In particular, the system can use the cofolding neural network to generate respective “co-folding data” for each of multiple pairs of candidate ligands. Co-folding data for a pair of candidate ligands characterizes a relative binding affinity for the protein of a first candidate ligand in the pair in comparison to a second candidate ligand in the pair, more particularly which of the two candidate ligands of the pair is closer to a binding pocket in a joint 3D structure of the protein and the pair of candidate ligands.

[0037] To generate co-folding data for a pair of candidate ligands, the system processes data defining the protein and the pair of candidate ligands using the co-folding neural network (any co-folding neural network will do) to predict the joint 3D structure of the protein and the pair of candidate ligands. The pair of candidate ligands can be understood as “competing” to occupy the binding pocket on the protein, and the candidate ligand having the higher binding affinity may tend to occupy the binding pocket with a higher likelihood than the other candidate ligand. Based on this rationale, the system can process data defining joint 3D structures of the proteinand pairs of candidate ligands to generate co-folding data. The system can then derive the ranking of the candidate ligands from the co-folding data.

[0038] The co-folding neural network can be trained to generate 3D structures of molecular complexes and then applied to generate co-folding data, as described above, in a “zero-shot” manner, i.e., without requiring any training related to directly predicting binding affinities. The system thus enables reduced consumption of computational resources, e.g., memory and computing power, by obviating the need to fine-tune a neural network on the task of predicting binding affinity.

[0039] Moreover, the system can achieve a higher accuracy in ranking candidate ligands than can be achieved by conventional systems, e.g., that individually determine binding affinities for each candidate ligand and then rank the candidate ligands based on their individually- determined binding affinities. First, the system described in this specification generates the ranking based on a richer set of intermediate data, in particular, the co-folding data, that exploits and characterizes interactions between pairs of candidate ligands rather than considering each candidate ligand separately. Second, the co-folding neural network can be trained on more plentiful “unlabeled” data defining molecular complexes, without requiring training data that specifically labels proteins and ligands with binding affinities.

[0040] The system provides a technical solution to the technical problem of data scarcity that arises when using computational methods to predict binding affinity. More specifically, training a machine learning model to predict binding affinities of ligands for proteins using conventional approaches would require abundant “labeled” training examples, i.e., where particular ligands and proteins are labeled with an associated binding affinity value. However, experiments to assess binding affinity are time consuming and expensive, and there is a scarcity of labeled training examples for training machine learning models to predict binding affinity. Training a machine learning model in the absence of sufficient training data may be infeasible, e.g., because the training may fail to converge, and because the trained machine learning model may overfit the training data and fail to generalize to previously unseen proteins and ligands. The system described in this specification is trained on more plentiful unlabeled training data defining 3D structures of molecule complexes, without requiring training data that specifically labels proteins and ligands with binding affinities, and thereby provides a solution to the problem of data scarcity in relation to binding affinity data.

[0041] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, andadvantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG. 1 shows an example ligand ranking system.

[0043] FIG. 2 is a flow diagram of an example process for generating a ranking for a set of candidate ligands based on a respective predicted binding affinity of each candidate ligand for a protein.

[0044] FIG. 3 is a flow diagram of an example process for generating a co-folding value for a pair of candidate ligands that does not rely on data defining the location of the binding pocket in the protein.

[0045] FIG. 4 illustrates an example of a joint 3D structure of a pair of candidate ligands and a protein.

[0046] FIG. 5 illustrates co-folding data for pairs of candidate ligands being processed by a ranking engine to generate a ranking of candidate ligands based on a respective predicted binding affinity of each candidate ligand for a protein.

[0047] FIG. 6 illustrates an example of upper bounding and lower bounding binding affinities of non-anchor ligands based on: (i) a ranking of a set of candidate ligands, and (ii) the known binding affinities of anchor ligands.

[0048] FIG. 7 illustrates an example of determining a co-folding value for a pair of candidate ligands using a process that does not require data defining the location of the binding pocket on the protein.

[0049] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0050] FIG. 1 shows an example ligand ranking system 100. The ligand ranking system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0051] The ligand ranking system 100 (“the system”) is configured to receive an input that includes data identifying: (i) a protein, and (ii) a set candidate ligands. The system 100 processes the input to generate data defining a ranking 118 of the set of candidate ligands based on a respective predicted (relative) binding affinity of each of the candidate ligands for theprotein. More specifically this data can characterize or define the protein and each of the candidate ligands.

[0052] The data identifying the protein 102 can include data defining one or more amino acid sequences of the protein 102. In some cases, the protein 102 can be a single-chain protein that includes a single amino acid sequence. In some cases, the protein 102 can be a multi-chain protein that includes multiple amino acid sequences, e.g., the protein 102 can be a protein complex, e.g., a group of two or more protein molecules that come together and interact with each other to perform a biological function. In this case, the data defining the protein complex can include data defining one or more amino acid sequences of each protein molecule in the protein complex.

[0053] The system 100 can further receive other data characterizing the protein 102, e.g., as an alternative to or in combination with the data defining the one or more amino acid sequences of the protein. For instance, the system 100 can receive data characterizing a respective multiple sequence alignment (MSA) for each amino acid chain of the protein. As another example, the system 100 can receive data characterizing a predicted structure of the protein, e.g., a set of structure parameters defining a predicted structure of the protein, e.g., a set of structure parameters defining a respective 3D spatial position of some or all of the atoms in a 3D structure of the protein.

[0054] The set of candidate ligands can include any appropriate number of candidate ligands, e.g., 2 candidate ligands, or 10 candidate ligands, or 100 candidate ligands, or 1000 candidate ligands, or 100,000 candidate ligands. The system 100 can receive any appropriate data identifying each candidate ligand 104. For instance, for one or more candidate ligands, the system 100 can receive a textual representation of the chemical structure of the ligand, e.g., that characterizes one or more of the individual atoms in the ligand, the bonds between the atoms in the ligand, any cyclic structures in the ligand, charge or radicals present in the ligand, and so forth. A chemical structure of a candidate ligand can represented textually, e.g., by way of a Simplified Molecular Input Line Entry System (SMILES) string identifying the ligand. As another example, for one or more of the candidate ligands, the system 100 can receive a representation of the ligand by way of graph data representing a graph, e.g., where the nodes in the graph represent atoms in the ligand and the edges in the graph represent bonds between atoms in the ligand. As another example, one or more of the candidate ligands can be proteins, and for each protein ligand, the system 100 can receive data characterizing one or more of one or more amino acid sequences of the protein, a structure of the protein, an MSA for the amino acid chains of the protein, and so forth.

[0055] The ranking 118 of the candidate ligands can define an ordering of the candidate ligands based on a respective predicted binding affinity of each candidate ligand for the protein 102. For instance, the ranking 118 can order the candidate ligands 104 from a first candidate ligand, having a lowest predicted binding affinity for the protein 102, to a last candidate ligand, having a highest predicted binding affinity for the protein 102. As another example, the ranking 118 can order the candidate ligands from a first candidate ligand, having a highest predicted binding affinity for the protein 102, to a last candidate ligand, having a lowest predicted binding affinity for the protein 102.

[0056] The system 100 can generate the ranking 118 without explicitly generating a respective (relative or absolute) predicted binding affinity (e.g., measured in units of molarity) of each candidate ligand for the protein 102. In particular, the system 100 can infer the ranking 118, even without explicitly generating predicting binding affinities, based on “co-folding data” corresponding to pairs of candidate ligands, as will be described in more detail below.

[0057] Optionally, the system can generate data characterizing a respective predicted binding affinity of each of one or more of the candidate ligands 104 for the protein 102. A few examples of processes involving generating predicted binding affinities for one or more of the candidate ligands 104 are described next.

[0058] In one example, the system 100 can generate data characterizing a predicted binding affinity of each of one or more of the candidate ligands 104 based on: (i) the ranking 118, and (ii) one or more “anchor” ligands included in the set of candidate ligands. An anchor ligand refers to a ligand having a known binding affinity for the protein, e.g., where the binding affinity has been determined through computational simulations or through physical experiments. The system 100 can use the binding affinities of the anchor ligands to establish an upper bound, or a lower bound, or both, for the respective binding affinity of each “nonanchor” ligand (i.e., for which the binding affinity of the ligand for the protein is not known). Example techniques for upper bounding and lower bounding binding affinities of non-anchor ligands based on: (i) the ranking 118, and (ii) the known binding affinities of anchor ligands, are described in more detail with reference to FIG. 6.

[0059] In another example, the system 100 can select a proper subset of the set of candidate ligands 104 based on the ranking 118, and then generate a predicted, e.g., absolute rather than relative, binding affinity of each of the selected ligands. For instance, the system 100 can select a predefined number (or fraction) of the candidate ligands that are predicted to have the highest (or lowest) binding affinity for the protein based on the ranking 118. In some cases, the system 100 can select less than 50%, or less than 25%, or less than 10%, or less than 5%, or less than1% of the ligands in the set of candidate ligands. The system 100 can then determine a respective predicted binding affinity of each of the selected candidate ligands for the protein 102. The system 100 can determine a predicted binding affinity of a ligand for a protein in any of a variety of possible ways, e.g., using structure-based docking methods, quantitative structure-activity relationship (QSAR) models, a machine learning model (e.g., a neural network) configured to predict binding affinities, or using free energy perturbation (FEP) or molecular dynamics (MD) simulations.

[0060] In another example, the system 100 can select a proper subset of the set of candidate ligands 104 based on the ranking 118, and a respective physical (laboratory) experiment can be performed for each selected ligand to determine a binding affinity of the selected ligand for the protein 102. As described above, the system can select a predefined number (or fraction) of the candidate ligands that are predicted to have the highest (or lowest) binding affinity for the protein based on the ranking 118. In some cases, the system 100 can select less than 50%, or less than 25%, or less than 10%, or less than 5%, or less than 1% of the ligands in the set of candidate ligands as subjects for physical experiments.

[0061] Optionally, after determining binding affinities for selected ligands in the set of candidate ligands (e.g., through computational methods or physical experiments, as described above), the system 100 can provide the selected ligands as anchor ligands. The system 100 can then use the anchor ligands to upper bound, or lower bound, or both, the binding affinities of the remaining candidate ligands using the ranking 118, as described in more detail with reference to FIG. 6.

[0062] Determining binding affinities of candidate ligands for the protein 102, e.g., through computational methods or physical experiments, can be expensive and time consuming. The ranking 118 generated by the system 100 can be used to select a proper subset of the total set of candidate ligands that are expected to have desirable properties, e.g., high or low binding affinity for the protein. Binding affinities can then be determined for only the selected candidate ligands, rather than for every candidate ligand, thus enabling more efficient and targeted use of resources.

[0063] The ranking 118 can be used to select one or more candidate ligands (e.g., those having the highest or lowest predicted binding affinity from among the set of candidate ligands, according to the ranking 118) for experimental testing and validation, e.g., for use in a drug that achieves a therapeutic effects in patients. Optionally, based on the ranking 118, one or more of the candidate ligands can be selected for physical synthesis and then tested for a variety of properties, e.g., absorption, distribution, metabolism, and / or excretion, by a living organismor cell culture or tissue model. One or more of the candidate ligands can be selected for inclusion in a drug, e.g., based at least in part on results of the testing. A drug that includes one or more of the candidate ligands can be synthesized using any appropriate drug synthesis technique.

[0064] That is, the system 100 can be used for identifying a drug, i.e. a ligand with a therapeutic effect (e.g. an agonist or antagonist of a receptor or enzyme). This can involve using the ranking of the candidate ligands to select one or more ligands, e.g. one or more with a ranking that predicts a highest (or lowest) binding affinity for the protein, as a putative drug. Optionally the selected ligand(s) can then be further screened in silico, or after physical synthesis, in vitro (e.g. in a cell culture or tissue model), or in vivo. For example the selected ligand(s) can be screened for further useful properties, e.g. according to a degree to which binding is accompanied by a biological (therapeutic) effect such as facilitating a biological mechanism or directly or indirectly inhibiting a biological disease mechanism (e.g. inhibiting a bacteria or virus from entering a cell); toxicity; clearance time; and so forth.

[0065] The system 100 uses a co-folding neural network 110 and a ranking engine 116 to process the input identifying the protein 102 and the set of candidate ligands 104 and to generate the ranking 118 of the candidate ligands. The co-folding neural network 110 and the ranking engine 116 are described in more detail next. The described “competitive binding” approach does not rely on any particular co-folding neural network architecture; any protein folding neural network that can fold a protein and a pair of ligands can be used. Some general characteristics of the co-folding neural network 110 are described below. One particular example of a neural network that can be used as the co-folding neural network 110 is AlphaFold 3 (Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3”, Nature 630, 493-500, 2024; also AlphaFoldServer.com).

[0066] The co-folding neural network 110 is configured to process data defining: (i) the protein 102, and (ii) one or more candidate ligands, to generate data defining a joint three-dimensional (3D) structure 112 of the candidate ligands and the protein 102. That is, the co-folding neural network 110 generates data that jointly defines the 3D conformation of an atomic system that includes the protein 102 and the one or more candidate ligands.

[0067] In more detail, the co-folding neural network can be configured to process any appropriate data defining the protein and the one or more candidate ligands. For instance, the data defining the protein can include data defining one or more amino acid sequences of the protein, an MSA for the protein, a structure of the protein, or a combination thereof. As anotherexample, the data defining each ligand can include a textual representation of the chemical structure of the ligand, e.g., as a SMILES string.

[0068] The co-folding neural network can be configured to generate any appropriate data that represents the joint 3D structure (conformation) of the protein and the one or more candidate ligands. For instance, the co-folding neural network can generate data that defines a respective 3D spatial position of each of the atoms included within each of the one or more candidate ligands and the protein. As another example, the co-folding neural network can generate data defines a respective 3D spatial position of only a proper subset of the atoms included within the one or more candidate ligands and the protein. For instance, the co-folding neural network can generate data defining respective 3D spatial positions of each of the atoms included within the one or more candidate ligands but of only the backbone atoms in the amino acids of the protein. The modelled atoms may comprise only so-called “heavy” atoms (e.g. C, N, O, S), i.e. not including hydrogen.

[0069] The co-folding neural network can have any appropriate neural network architecture that enables the co-folding neural network to perform its described functions, e.g., processing data defining the protein and one or more candidate ligands to generate a joint 3D structure of the one or more candidate ligands and the protein. For instance, the co-folding neural network can include any appropriate types of neural network layers (e.g., fully-connected layers, convolutional layers, attention layers, pooling layers, etc.) in any appropriate number (e.g., e.g., 5 layers, 10 layers, or 100 layers) and connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).

[0070] The system can train the co-folding neural network on a set of training examples, where some or all of the training examples include: (i) a training input to the co-folding neural network that defines a protein and one or more ligands that bind to the protein, and (ii) a target 3D structure that defines a joint 3D structure of the protein and the one or more ligands, e.g., when the one or more ligands are bound to the protein. Training the co-folding neural network on the set of training examples can include, for each training example, training the co-folding neural network to process the training input of the training example to generate a predicted 3D structure that matches the target 3D structure of the training example. The target 3D structures of the training data may have been determined, e.g., through physical experiments. There are also many public databases that can be used including those listed in the Supplementary Methods section of Abramson et al, ibid, such as the Protein Data Bank (wwpdb.org), and many others.

[0071] More specifically, the system can train the set of neural network parameters of the cofolding neural network to optimize an objective function that, for each training example, measures an error between: (i) the predicted 3D structure generated by the co-folding neural network for the training example, and (ii) the target 3D structure specified by the training example. The objective function can measure the error between a predicted 3D structure generated by the co-folding neural network and a target 3D structure specified by a training example in any of a variety of possible ways. For instance, the objective function can measure a root mean square deviation (RMSD) or a global distance test - total score (GDT-TS, Zemla, "LGA: A method for finding 3D similarities in protein structures", Nucleic Acids Research, 31 (13), 3370-3374, 2003) between 3D spatial positions of respective atoms in the predicted 3D structure and the target 3D structure.

[0072] The system can train the co-folding neural network on the set of training examples using an appropriate machine learning training technique, e.g., stochastic gradient descent. In particular, the system can train the co-folding neural network over a sequence of training iterations. At each training iteration, the system can sample (e.g., randomly sample) a batch of training examples from the set of training examples. For each training example in the batch, the system can process the training input of the training example using the co-folding neural network, in accordance with current values of the set of neural network parameters of the cofolding neural network, to generate a predicted 3D structure. The system can then determine gradients of the objective function (that measures an error between the predicted 3D structure and the target 3D structure of the training example), e.g., using backpropagation, and uses the gradients to update the current values of the set of neural network parameters of the co-folding neural network, e.g., using the update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.

[0073] For each of multiple pairs of candidate ligands from the set of candidate ligands 104, the system 100 can process the data defining the pair of candidate ligands (e.g., including a first candidate ligand 106 and a second candidate ligand 108) and the protein using the cofolding neural network 110 to generate data defining a joint 3D structure 112 of the pair of candidate ligands and the protein.

[0074] For each pair of candidate ligands, the system 100 can process the joint 3D structure 112 of the pair of candidate ligands and the protein 102 (i.e., as generated by the co-folding neural network 110) to generate co-folding data 114 that characterizes a respective binding affinity for the protein 102 of the first candidate ligand 106 in the pair in comparison to the second candidate ligand 108 in the pair. That is, the co-folding data 114 can characterize whichof the two candidate ligands 106, 108 has a higher binding affinity for the protein 102. In particular, the pair of candidate ligands can be understood as “competing” to occupy the binding pocket on the protein, and the candidate ligand having the higher binding affinity may tend to occupy the binding pocket with a higher likelihood than the other candidate ligand. Based on this rationale, the system can process data defining joint 3D structures of the protein and pairs of candidate ligands to generate co-folding data. Example techniques for processing a joint 3D structure 112 of the protein 102 and a pair of candidate ligands 106, 108 to generate co-folding data 114 for the pair of candidate ligands 106, 108 are described in more detail with reference to FIG. 2. For example, the co-folding data for a pair of candidate ligands, based on the joint 3D structure of the pair of candidate ligands and the protein, may comprise data that defines which candidate ligand is displaced most from a binding pocket of the protein, more specifically that defines a difference in the distance of the candidate ligands from the binding pocket in the joint 3D structure.

[0075] The ranking engine 116 is configured to process the respective co-folding data 114 for the pairs of candidate ligands to generate the ranking 118 of the set of candidate ligands 104 based on a respective predicted binding affinity of each of the candidate ligands for the protein. The ranking engine 116 can generate the ranking 118 from the co-folding data 114 in a variety of possible ways. For example there are many known techniques for converting pairwise comparisons into rankings. As some particular examples, useful in system 100, the ranking engine 116 can generate the ranking as the solution of a game characterized by a payoff matrix defined by the co-folding data. As another example, the ranking engine 116 can process the cofolding data using a ranking machine learning model, trained in accordance with a machine learning training technique, to generate the ranking 118. Example techniques for generating the ranking 118 from the co-folding data 114 are described in more detail with reference to FIG. 2.

[0076] The system 100 can receive the data identifying the protein 102 and the set of candidate ligands 104 from any appropriate source, e.g., from a user or from another system, by way of an appropriate interface, e.g., an application programming interface (API) or a user interface (e.g., a graphical user interface). After generating the ranking 118 of the candidate ligands, the system 100 can, e.g., store data defining the ranking 118 in a memory, or transmit the data defining the ranking 118 over a data communication network, or provide data defining the ranking directly to a system that performs downstream processing based on the ranking 118.

[0077] In some cases, the ligand ranking system can be used to identify ligands that can act as a molecular glue that binds together two different proteins, e.g., by simultaneously binding toa respective binding site on each protein and thus gluing together the two proteins. For instance, the ligand ranking system can receive data identifying multiple proteins, e.g., rather than a single protein, and can be used to rank a set of candidate ligands based on their simultaneous binding affinities for multiple binding sites on the multiple proteins.

[0078] FIG. 2 is a flow diagram of an example process 200 for generating a ranking for a set of candidate ligands based on a respective predicted binding affinity of each candidate ligand for a protein. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a ligand ranking system, e.g., the ligand ranking system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200.

[0079] The system receives data identifying (e.g. characterizing or defining): (i) a protein, and (ii) a set of candidate ligands (202). The protein can be, e.g., an enzyme, receptor, or signaling protein that is involved in a human or animal disease process. The set of candidate ligands can be obtained from any appropriate source. For instance, the candidate ligands can be obtained from an existing library of ligands, such as so-called compound libraries (e.g. available commercially), libraries generated by combinatorial techniques, and other sources (e.g. the previously mentioned examples of databases of training data). As another example, the candidate ligands can be obtained by randomly sampling a space of possible ligands. As another example, the candidate ligands can be generated by applying random modifications (e.g., adding or removing bonds or atoms) from a known ligand for the protein.

[0080] The system identifies a set of pairs of candidate ligands from the set of candidate ligands (204). Each pair of candidate ligands includes a respective first candidate ligand and a respective second (different) candidate ligand from the set of candidate ligands. The system can identify any appropriate number of pairs of candidate ligands. For instance, the system can identify every possible pair of candidate ligands (where each possible pair of candidate ligands comprise two different ligands from the set of candidate ligands). As another example, the system can identify fewer than every possible pair of candidate ligands, e.g., the system can randomly sample a collection of pairs of candidate ligands from the set of candidate ligands.

[0081] For each pair of candidate ligands, the system generates data defining one or more joint 3D structures of the pair of candidate ligands and the protein (206). Each joint 3D structure of the pair of candidate ligands and the protein defines a 3D conformation of an atomic system that includes the protein and the pair of candidate ligands, as described above with reference to FIG. 1.

[0082] The system can generate a joint 3D structure of the pair of candidate ligands and the protein using the co-folding neural network. In particular, the system can process data defining the pair of candidate ligands and the protein using the co-folding neural network, in accordance with (trained) values of a set of co-folding neural network parameters, to generate the joint 3D structure.

[0083] Optionally, the system can generate multiple (different) joint 3D structures for the pair of candidate ligands and the protein. Each joint 3D structure can represent a possible joint conformation of the pair of candidate ligands and the protein. A few example techniques for generating multiple joint 3D structures for the pair of candidate ligands and the protein are described next.

[0084] In one example, the system can maintain an ensemble of multiple co-folding neural networks. Each co-folding neural network in the ensemble can have a respective set of cofolding neural network parameters having values that are specific to the co-folding neural network (and in particular, are different from each other co-folding neural network in the ensemble). Each co-folding neural network may be trained on a different set of training examples, or may have the values of the co-folding neural network parameters initialized randomly (and separately from each other co-folding neural network) during training. The system can use the ensemble of co-folding neural networks to generate multiple joint 3D structures for the pair of candidate ligands and the protein. In particular, the system can process data identifying the pair of candidate ligand and the protein using each co-folding neural network in the ensemble to generate a respective joint 3D structure as the output of each cofolding neural network in the ensemble.

[0085] In another example, one or more neural network layers in a co-folding neural network can be configured to perform operations that involve random sampling. For instance, a penultimate neural network layer of the co-folding neural network can be configured to generate, for each atom included in the protein and the pair of candidate ligands, a respective distribution over a space of possible spatial positions of the atom. The output layer of the neural network can be configured to generate a joint 3D structure of the protein and the pair of candidate ligands by sampling a respective 3D spatial position of each atom included in the protein and the pair of candidate ligands from the corresponding distribution. In implementations where the co-folding neural network includes neural network layers that perform operations involving random sampling, the system can generate multiple joint 3D structures of the protein and the pair of candidate ligands using the co-folding neural network. For instance, the system can perform multiple independent runs of the random samplingoperations of the neural network layers of the co-folding neural network and then process the sampled data using downstream neural network layers of the co-folding neural network to generate multiple (different) joint 3D structures.

[0086] For each pair of candidate ligands, the system generates respective co-folding data for the pair of candidate ligands that characterizes a relative binding affinity for the protein of the first candidate ligand in the pair in comparison to the second candidate ligand in the pair (208). In a particular example, the co-folding data for the pair of candidate ligands includes a set of one or more co-folding values (where each co-folding value can be, e.g., a scalar value). The system can generate each co-folding value by processing a respective joint 3D structure of the pair of candidate ligands and the protein. (The number of co-folding values may be equal to the number of joint 3D structures of the pair of candidate ligands and the protein).

[0087] The system can process a joint 3D structure of a pair of candidate ligands and the protein to generate a corresponding co-folding value in any of a variety of possible ways. A few example techniques for generating co-folding values are described next.

[0088] In some implementations, the location of the binding pocket in the protein is known, and the system can generate a co-folding value from a joint 3D structure of a pair of candidate ligands and the protein based on the known location of the binding pocket. The location of the binding pocket can be defined, e.g., by data identifying a set of amino acids in the protein that are included in the binding pocket, or by data identifying a set of atoms in the protein that are included in the binding pocket. The location of the binding pocket may have been determined, e.g., through physical experiments or through computational (e.g., simulation-based) methods.

[0089] To generate the co-folding value for a pair of candidate ligands when the location of the binding pocket on the protein is known, the system can determine a respective distance measure for each candidate ligand that characterizes a distance of the candidate ligand from the binding pocket of the protein in the joint 3D structure. The system can then determining the co-folding value for the pair of candidate ligands based on the respective distance measure of each candidate ligand from the binding pocket. For instance, the system can determine the co-folding value as a difference between: (i) the distance measure of the first candidate ligand in the pair from the binding pocket, and (ii) the distance measure of the second candidate ligand in the pair from the binding pocket. The co-folding value can thus reflect the intuition that, as the pair of candidate ligands “compete” the occupy the binding pocket, the candidate ligand with a higher binding affinity for the binding pocket is more likely to occupy the binding pocket, and thus have a lower distance measure from the binding pocket, than the other candidate ligand in the joint 3D structure.

[0090] The system can determine a distance measure that characterizes a distance of a candidate ligand from the binding pocket of the protein in any of a variety of possible ways. For instance, the system can determine the distance measure for a candidate ligand based on fraction of the amino acids of the binding pocket that have at least one atom within a threshold distance (e.g., 8 Angstroms) of at least one atom of the candidate ligand (i.e., according to the joint 3D structure). In another example, the system can determine the distance measure for a candidate ligand based on a measure of distance (e.g. an RMSD or a Wasserstein distance) between the atoms of the candidate ligand and the atoms of the binding pocket (i.e., according to the joint 3D structure).

[0091] In some implementations, to determine a co-folding value from a joint 3D structure of a pair of candidate ligands and the protein, the system additionally generates: (i) a joint 3D structure of a first candidate ligand from the pair and the protein, and (ii) a joint 3D structure of a second candidate ligand from the pair and the protein. The system can then determine the co-folding value by measuring relative displacements of the positions of the candidate ligands between these joint 3D structures, e.g., to determine a difference between the relative displacements. Generating the co-folding value in this manner does not require information about the location of the binding pocket on the protein, and may thus be particularly appropriate when the location of the binding pocket is unknown. An example process for generating a cofolding value based on displacements of the positions of the candidate ligands between respective joint 3D structures that include the first candidate ligand, or the second candidate ligand, or both, is described in more detail with reference to FIG. 3.

[0092] The system processes the respective co-folding data for the pairs of candidate ligands to generate the ranking of the candidate ligands in the set of candidate ligands based on a respective predicted binding affinity of each of the candidate ligands for the protein (210). A few example techniques for processing co-folding data for pairs of candidate ligands to generate the ranking of the candidate ligands are described next.

[0093] In some implementations, the system generates the ranking of the candidate ligands as the solution of a game characterized by a “payoff matrix” defined by the co-folding data for the pairs of candidate ligands. More specifically, the co-folding data can define a so-called payoff matrix that characterizes outcomes in a two-player game where each player selects a respective candidate ligand (a player choice or “strategy”), and the “payoff’ received by each player depends on the relative binding affinity of the candidate ligand selected by the player compared to the binding affinity of the candidate ligand selected by the other player, e.g. comprises a co-folding value as described above. This can result in a zero sum game (as theeffect of players exchanging selected ligands is the same joint 3D structure). The ranking engine 116 can leverage one or more game theory algorithms to determine a so-called “solution” to the two-player game characterized by the payoff matrix, where the solution to the game defines the ranking of the candidate ligands. In one example, the ranking engine 116 can determine the solution to the game characterized by the payoff matrix using the Elo rating system, e.g., as described in Elo, Arpad E. (August 1967). “The Proposed USCF Rating System, Its Development, Theory, and Applications”. Chess Life. XXII (8): 242-247. In another example, the ranking engine 116 can determine the solution to the game characterized by the payoff matrix using the Alpha rank system, e.g., as described in: Omidshafiei, Shayegan, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M. Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos, “a-rank: Multi-agent evaluation by evolution.” Scientific reports 9, no. 1 (2019): 9937. In another example, the ranking engine 116 can determine the solution to the game characterized by the payoff matrix using the Disc rating system, e.g., as described in: Quentin Bertrand, Wojciech Marian Czarnecki, Gauthier Gidel: On the Limitations of the Elo, Real -World Games are Transitive, not Additive. AISTATS 2023: 2905-2921.

[0094] In some implementations, the system generates the ranking of the candidate ligands using a ranking machine learning model. More specifically, the system processes co-folding data for the pairs of candidate ligands using the ranking machine learning model, in accordance with trained values of a set of ranking machine learning model parameters, to generate data defining a ranking of the candidate ligands.

[0095] The output of the ranking machine learning model can define the ranking of the candidate ligands in any of a variety of possible ways. For instance, the ranking machine learning model can be implemented as a ranking neural network with an output layer that generates a respective numerical “activation” value for each of the candidate ligands characterized by the co-folding data provided as an input to the ranking neural network. The activation values for the candidate ligands can define the ranking of the candidate ligands, e.g., a ranking from highest activation value to lowest activation value, or a ranking from lowest activation value to highest activation value.

[0096] The ranking machine learning model can have any appropriate machine learning model architecture that enables the ranking machine learning model to perform its described functions, e.g., processing co-folding data for pairs of candidate ligands to generate data defining a ranking of the candidate ligands. For instance, the ranking machine learning model can be implemented as a ranking neural network that includes any appropriate types of neuralnetwork layers (e.g., attention layers, fully-connected layers, convolutional layers, etc.) in any appropriate number (e.g., 5 layers, 10 layers, or 50 layers) and connected in any appropriate configuration (e.g., as a linear sequence of layers or as a directed graph of layers).

[0097] The system can train the ranking machine learning model on a set of training examples that each include: (i) training co-folding data for a set of training ligands, and (ii) a target ranking of the set of training ligands. In particular, the training co-folding data can include respective co-folding data for multiple pairs of training ligands from the set of training ligands. The co-folding data for a pair of training ligands characterizes a relative binding affinity for a training protein of a first training ligand in the pair in comparison to a second training ligand in the pair. The target ranking can define a ranking of the training ligands based on a respective binding affinity of each training ligand for the training protein. The training co-folding data can be generated using a co-folding neural network, e.g., as described throughout this document. The target ranking can be determined, e.g., through computational simulations or physical experiments.

[0098] Training the ranking machine learning model on the set of training examples can include, for each training example, training the ranking machine learning model to process the training co-folding data of the training example to generate data defining a predicted ranking that conforms with the target ranking specified by the training example. The training objective function can be to minimize a difference between the predicted ranking and the target ranking.

[0099] FIG. 3 is a flow diagram of an example process 300 for generating a co-folding value for a pair of candidate ligands that does not rely on data defining the location of the binding pocket in the protein. The co-folding value for the pair of candidate ligands characterizes a relative binding affinity for the protein of a first candidate ligand in the pair in comparison to a second candidate ligand in the pair. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a ligand ranking system, e.g., the ligand ranking system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 300.

[0100] The system receives data defining a joint 3D structure of the pair of candidate ligands and the protein (302). The joint 3D structure of the pair of candidate ligands and the protein defines a 3D conformation of an atomic system that includes the protein and the pair of candidate ligands, and can be generated using a co-folding neural network, as described above with reference to FIG. 2.

[0101] The system generates data defining one or more joint 3D structures of first candidate ligand (from the pair of candidate ligands) and the protein (304). The joint 3D structure of thefirst candidate ligand and the protein defines a 3D conformation of an atomic system that includes the first candidate ligand and the protein, but does not include the second candidate ligand. The system can generate a joint 3D structure of the first candidate ligand and the protein by processing data defining the first candidate ligand and the protein using the co-folding neural network (described above with reference to FIG. 1). The system can generate multiple joint 3D structures of the first candidate ligand and the protein, e.g., using an ensemble of cofolding neural networks, or by using a co-folding network that includes one or more neural network layers that perform operations involving random sampling, as described above with reference to FIG. 2.

[0102] The system determines a measure of displacement of the first candidate ligand between: (i) the one or more joint 3D structures of the first candidate ligand and the protein, and (ii) the joint 3D structure of the pair of candidate ligands and the protein (306). That is, the system determines a measure of displacement that characterizes a change in the 3D spatial position of the first candidate ligand resulting from including the second candidate ligand in an atomic system that includes the first candidate ligand and the protein. Where multiple joint 3D structures are used, the measure of displacement can, e.g., be averaged.

[0103] As previously described a measure of displacement can be based, e.g., on atomic distances. Optionally as part of determining the measure of displacement of the first candidate ligand, the system can align: (i) the one or more joint 3D structures of the first candidate ligand and the protein, and (ii) the joint 3D structure of the pair of candidate ligands and the protein, e.g. by aligning the two 3D protein structures. For instance, the system can apply a respective affine transformation (e.g., a transformation that includes translation and rotation operations) to each of the 3D structures to minimize a distance (e.g. an RMSD or Wasserstein distance) between the positions of the atoms of the protein among the 3D structures.

[0104] After aligning the 3D structures, the system can determine the measure of displacement of the first candidate ligand in any of a variety of possible ways. For instance, for each of the one or more joint 3D structures of the first candidate ligand and the protein, the system can determine a respective distance (e.g., an RMSD or Wasserstein distance) between the positions of the atoms in the first candidate ligand between: (i) the joint 3D structure of the first candidate ligand and the protein, and (ii) the joint 3D structure of the pair of candidate ligands and the protein. The system can then determine the measure of displacement of the first candidate ligand as a measure of central tendency, e.g., a mean or median, of the determined distances.

[0105] The system generates data defining one or more joint 3D structures of second candidate ligand (from the pair of candidate ligands) and the protein (308). The joint 3D structure of thesecond candidate ligand and the protein defines a 3D conformation of an atomic system that includes the second candidate ligand and the protein, but does not include the first candidate ligand. Example techniques for generating a joint 3D structure of a candidate ligand and a protein are described above with reference to step 304.

[0106] The system determines a measure of displacement of the second candidate ligand between: (i) the one or more joint 3D structures of the second candidate ligand and the protein, and (ii) the joint 3D structure of the pair of candidate ligands and the protein (310). Example techniques for determining the measure of displacement of a candidate ligand are described above with reference to step 306.

[0107] The system determines the co-folding value as a difference between: (i) the measure of displacement of the first candidate ligand, and (ii) the measure of displacement of the second candidate ligand (312). The co-folding value can thus reflect the intuition that the candidate ligand with a higher binding affinity for the protein will be displaced less as a result of the introduction of the other candidate ligand into the atomic system.

[0108] FIG. 4 illustrates an example of a joint 3D structure of a pair of candidate ligands and a protein. In this example, a first ligand is bound to a binding pocket of the protein, while a second ligand is outside any binding pockets of the ligand, e.g., indicating that the second ligand may not bind to any binding pockets on the protein, or that the second ligand has a lower binding affinity for the protein than the first ligand.

[0109] FIG. 5 illustrates co-folding data for pairs of candidate ligands being processed by a ranking engine (e.g., of the ligand ranking system, as described with reference to FIG. 1) to generate a ranking of candidate ligands based on a respective predicted binding affinity of each candidate ligand for a protein.

[0110] In the illustration of FIG. 5, the set of candidate ligands includes “Ligand A,” “Ligand B,” “Ligand C,” and so forth. The co-folding data for pairs of candidate ligands is illustrated as an array, where each entry in the array corresponds to co-folding data for a pair of candidate ligands. For instance, the co-folding data for candidate ligands B and C is illustrated as a shaded box in the array. It will be appreciated that the representation of the co-folding data as a square array is for illustrative purposes and is not generally how the co-folding data would be stored or represented by the ligand ranking system. For instance, the diagonal entries of the array are not required (since the ligand ranking system does not generate co-folding data for pairs of identical ligands) and the array as a whole would be symmetric (and thus storing the entire square array would be unnecessary).

[0111] The ranking engine processes the co-folding data, e.g., using a game theory algorithm or using a ranking machine learning model (as described above with reference to FIG. 2) to generate the ranking of the candidate ligands.

[0112] FIG. 6 illustrates an example of upper bounding and lower bounding binding affinities of non-anchor ligands based on: (i) a ranking of a set of candidate ligands, and (ii) the known binding affinities of anchor ligands. An anchor ligand refers to a ligand with a known binding affinity for a protein, whereas a non-anchor ligand refers to a ligand with an unknown binding affinity for the protein, as described above with reference to FIG. 1. The ranking of the set of candidate ligands can be generated by a ligand ranking system, e.g., as described with reference to FIG. 1, and can rank the candidate ligands in order of their binding affinity for the protein.

[0113] The binding affinity of a non-anchor ligand for the protein can be: (i) bounded on one extreme by the binding affinity of any anchor ligand ranked above the non-anchor ligand in the ranking, and (ii) bounded on the other extreme by the binding affinity of any anchor ligand ranked lower than the non-anchor ligand in the ranking. For instance, in FIG. 6, the binding affinities of candidate ligands A and B can be upper bounded by the binding affinity of anchor ligand #1 and lower bounded by the binding affinity of anchor ligand #2.

[0114] FIG. 7 illustrates an example of determining a co-folding value for a pair of candidate ligands using a process that does not require data defining the location of the binding pocket on the protein. The process illustrated with reference to FIG. 7 is described in detail with reference to FIG. 3. To determine the co-folding value for ligand #1 and ligand #2, the ligand ranking system can determine a displacement of ligand #1 and a displacement of ligand #2. The displacement of ligand #1 measures a displacement in the position of ligand #1 between: (i) a joint 3D structure of ligand #1 and the protein, and (ii) a joint 3D structure of ligand #1, ligand #2, and the protein. The displacement of ligand #2 measures a displacement in the position of ligand #2 between: (i) a joint 3D structure of ligand #2 and the protein, and (ii) a joint 3D structure of ligand #1, ligand #2, and the protein. In the example illustrated in FIG. 7, ligand #1 is displaced by less than ligand #2, which suggests that ligand #1 has a higher binding affinity for the protein than ligand #2.

[0115] Aspects of the disclosure of this specification are further described in Appendix A.

[0116] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particularoperations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.

[0117] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0118] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0119] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules,sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0120] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0121] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0122] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0123] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0124] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT(cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0125] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.

[0126] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.

[0127] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0128] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, whichacts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0129] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0130] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0131] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

CLAIMS1. A method performed by one or more computers, the method comprising: obtaining data identifying: (i) a protein, and (ii) a set of candidate ligands; generating respective co-folding data for each of a plurality of pairs of candidate ligands, wherein: each pair of candidate ligands comprises a respective first candidate ligand from the set of candidate ligands and a respective second candidate ligand from the set of candidate ligands; the respective co-folding data for each pair of candidate ligands characterizes a relative binding affinity for the protein of the respective first candidate ligand in the pair in comparison to the respective second candidate ligand in the pair; and generating the co-folding data for the pair of candidate ligands comprises: processing data defining the pair of candidate ligands and the protein using a co-folding neural network, in accordance with values of a set of co-folding neural network parameters, to generate data defining a joint three-dimensional (3D) structure of the pair of candidate ligands and the protein; and generating the co-folding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein; and processing the respective co-folding data for each of the plurality of pairs of candidate ligands to generate a ranking of the candidate ligands from the set of candidate ligands that is indicative of respective predicted binding affinities of each of the candidate ligands for the protein.

2. The method of claim 1, wherein for each pair of candidate ligands, generating the cofolding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein comprises: determining the co-folding data based on a difference between:(i) a distance of the first candidate ligand of the pair from a binding pocket of the protein in the joint 3D structure of the pair of candidate ligands and the protein, and(ii) a distance of the second candidate ligand of the pair from the binding pocket of the protein in the joint 3D structure of the pair of candidate ligands and the protein.

3. The method of claim 1, wherein for each pair of candidate ligands, generating the cofolding data for the pair of candidate ligands based on the joint 3D structure of the pair of candidate ligands and the protein comprises: determining a joint 3D structure of the first candidate ligand of the pair and the protein; determining a joint 3D structure of the second candidate ligand of the pair and the protein; and determining the co-folding data for the pair of candidate ligands based on the respective joint 3D structures of: (i) the pair of candidate ligands and the protein, (ii) the first candidate ligand of the pair and the protein, and (ii) the second candidate ligand of the pair and the protein.

4. The method of claim 3, wherein determining the joint 3D structure of the first candidate ligand of the pair and the protein comprises: processing data defining the first candidate ligand of the pair and the protein using the co-folding neural network.

5. The method of any one of claims 3-4, wherein determining the joint 3D structure of the second candidate ligand of the pair and the protein comprises: processing data defining the second candidate ligand of the pair and the protein using the co-folding neural network.

6. The method of any one of claims 3-5, wherein determining the co-folding data for the pair of candidate ligands based on the respective joint 3D structures of: (i) the pair of candidate ligands and the protein, (ii) the first candidate ligand of the pair and the protein, and (ii) the second candidate ligand of the pair and the protein comprises: determining a measure of displacement of the first candidate ligand between the respective joint 3D structures of: (i) the first candidate ligand and the protein, and (ii) the pair of candidate ligands and the protein; determining a measure of displacement of the second candidate ligand between the respective joint 3D structures of: (i) the second candidate ligand and the protein, and (ii) the pair of candidate ligands and the protein; and determining the co-folding data based on a difference between the measure of displacement of the first candidate ligand and the measure of displacement of the secondcandidate ligand.

7. The method of any preceding claim, wherein processing the respective co-folding data for each of the plurality of pairs of candidate ligands to generate the ranking of the candidate ligands comprises: generating the ranking of the candidate ligands as a solution of a game characterized by a payoff matrix defined by the co-folding data for the plurality of pairs of candidate ligands.

8. The method of claim 7, wherein the game is a two-player game where each player selects a respective candidate ligand, and wherein the payoff received by each player depends on a relative binding affinity for the protein of the candidate ligand selected by the player in comparison the candidate ligand selected by the other player.

9. The method of any one of claims 1-6, wherein processing the respective co-folding data for each of the plurality of pairs of candidate ligands to generate the ranking of the candidate ligands comprises: processing the co-folding data for each of the plurality of pairs of candidate ligands using a ranking machine learning model, in accordance with trained values of a set of ranking machine learning model parameters, to generate data defining the ranking of the candidate ligands.

10. The method of claim 9, wherein the ranking machine learning model comprises a neural network.

11. The method of any preceding claim, wherein the co-folding neural network has been trained on a set of training data that comprises a plurality of training examples that each include: (i) a training input that defines a protein and one or more ligands, and (ii) a target output that defines a joint 3D structure of the protein and the one or more ligands.

12. The method of any preceding claim, wherein the data identifying the protein comprises data defining each of one or more amino acid sequences of the protein.

13. The method of any preceding claim, wherein the data identifying each ligand comprises a text string defining a chemical structure of the candidate ligand.

14. The method of any preceding claim, wherein one or more of the candidate ligands are small molecules.

15. The method of any preceding claim, wherein the protein comprises an enzyme, receptor, or signaling protein that has been identified as being involved in a disease process.

16. The method of any preceding claim, further comprising: selecting one or more candidate ligands from the set of candidate ligands based on the ranking; and physically synthesizing the selected candidate ligands.

17. The method of claim 16, further comprising, for each of the selected candidate ligands, performing experiments using physically synthesized instances of the candidate ligand to determine one or more of: an absorption of the candidate ligand, a distribution of the candidate ligand, a metabolism of the candidate ligand, or an excretion of the candidate ligand.

18. The method of any preceding claim, wherein the ranking of the candidate ligands ranks the candidate ligands from highest predicted binding affinity for the protein to lowest predicted binding affinity for the protein.

19. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-18.

20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operationsof the respective method of any one of claims 1-18.

21. A ligand that has been synthesized by performing the method of claim 16.

22. One or more non-transitory computer storage media storing ligand data defining a ligand, wherein the ligand was selected from a set of candidate ligands by performing operations comprising: generating a ranking of candidate ligands in the set of candidate ligands that is indicative of respective predicted binding affinities of each of the candidate ligands for a protein by performing the method of any one of claims 1-15; and selecting the ligand from the set of candidate ligands based on the ranking.