Designing ligands using generative models
The generative model system efficiently generates ligands that meet design criteria and bind to target molecules by directly mapping from target molecule data, reducing computational costs and accounting for conformational changes, addressing the limitations of traditional ligand identification methods.
Patent Information
- Application Number
- PCT/EP2025/058453
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional methods for identifying ligands that bind to target molecules are computationally expensive, require advance knowledge of 3D structures, and fail to account for conformational changes, while traditional screening approaches are limited to known molecule databases, missing a vast array of possible ligands.
A generative model system that generates ligands by directly mapping from target molecule data, incorporating ligand design criteria, and using a generative diffusion model to denoise atom state data, ensuring the generated ligands meet desired properties and binding affinity.
Reduces computational resources and generates a diverse range of ligands that meet specific design criteria, including binding affinity, absorption, distribution, and toxicity, without requiring prior knowledge of 3D structures, thus overcoming limitations of traditional methods.
Smart Images

Figure EP2025058453_02102025_PF_FP_ABST
Abstract
Description
DESIGNING LIGANDS USING GENERATIVE MODELSBACKGROUND
[0001] This specification relates to using generative models to computationally design ligands that are predicted to bind to target molecules, e.g., proteins.
[0002] Predictions can be made using machine learning models. Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY
[0003] This specification describes a system implemented as computer programs on one or more computers in one or more locations that can use a generative model to computationally design a ligand for binding to a target molecule, such as a protein.
[0004] A “protein” can refer to any biological molecule that is specified by one or more sequences (or “chains”) of amino acids. For example, the term protein can refer to a protein domain, e.g., a portion of an amino acid chain of a protein that can undergo protein folding nearly independently of the rest of the protein. As another example, the term protein can refer to a protein complex, i.e., that includes multiple amino acid chains that jointly fold into a protein structure. The term “protein” also encompasses proteins that have undergone post- translational modifications, such as phosphorylation, glycosylation, ubiquitination, methylation, acetylation, or proteolytic cleavage.
[0005] A “ligand” can refer to a molecule or compound that binds to another, “target” molecule in a molecule complex, e.g., a protein. Ligands can include, e.g., small organic molecules, complex organic molecules, proteins, nucleic acids (e.g., deoxyribonucleic acid (DNA), ribonucleic acid (RNA)), lipids, carbohydrates, and so forth.
[0006] A “component” of a molecule can refer to any building block or unit of the molecule, e.g., a component can refer to an atom, or to an amino acid, or to a nucleotide, or to a functional group, and so forth. For example, if the target molecule is a protein, a component of the target molecule can refer to an amino acid. In some examples, a component of a molecule can referto a subunit of the molecule that comprises two or more atoms of the molecule, e.g., two or more, but not all the atoms of the molecule.
[0007] A “molecule complex” (or “complex”) can refer to an assembly of two or more molecules, e.g., that are held together by interactions such as hydrogen bonding, ionic interactions, van der Waals forces, or hydrophobic effects.
[0008] A “multiple sequence alignment” (MSA) for an amino acid chain in a protein specifies a sequence alignment of the amino acid chain with multiple additional amino acid chains, e.g., from other proteins, e.g., homologous proteins. More specifically, the MSA can define a correspondence between the positions in the amino acid chain and corresponding positions in multiple additional amino acid chains. An MSA for an amino acid chain can be generated, e.g., by processing a database of amino acid chains using any appropriate computational sequence alignment technique, e.g., progressive alignment construction. The amino acid chains in the MSA can be understood as having an evolutionary relationship, e.g., where each amino acid chain in the MSA may share a common ancestor. The correlations between the amino acids in the amino acid chains in an MSA for an amino acid chain can encode information that is relevant to predicting the structure of the amino acid chain.
[0009] A “binding pocket” or “binding site” on a target molecule can refer to a specific three- dimensional cavity or crevice within the structure of the target molecule where a ligand can bind to the target molecule. The binding pocket can, in some cases, be understood as a "lock" that fits the shape and chemical properties of ligands that act as "keys" for the lock. In other cases, the ligand may initially not fit perfectly into the binding pocket, e.g., due to structural differences or slight mismatches in shape or chemical groups, but conformational changes (“relaxation”) during binding can cause the interaction between the ligand and the binding pocket to become more complementary and specific, e.g., as in induced-fit binding. Examples of binding pockets include, e.g., orthosteric binding pockets, allosteric binding pockets, and cryptic binding pockets.
[0010] A “template” protein for a given protein can refer to a protein that is “similar” to the given protein, e.g., such that the value of a similarity measure between the template protein and the given protein satisfies (e.g., exceeds) a threshold (e.g., 0.8, or 0.9, or 0.99, or any other appropriate threshold). Similarity between a first protein and a second protein can be measured using any appropriate similarity measure, e.g., a sequence identity or percent identity similarity measure between the respective amino acid sequence(s) of the first protein and the second protein.
[0011] A first neural network can be referred to as a “subnetwork” of a second neural network if the first neural network is included in the second neural network.
[0012] A “block” (e.g., a “self-attention block”) in a neural network can refer to a group of one or more neural network layers in the neural network.
[0013] An “embedding” of an entity (e.g., an atom, or a ligand, or a protein) can refer to a representation of the entity as an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values.
[0014] “Conditioning” a model (e.g., a generative model) or aneural network (e.g., a denoising neural network) or an operation (e.g., a self-attention operation) on conditioning data (e.g., an embedding representing a protein or a set of ligand design criteria) can refer to providing the conditioning data as an input (e.g., a side input) to the model, neural network, or operation, such that outputs generated by the model, neural network, or operation are influenced by (depend on) the conditioning data.
[0015] A “binding affinity” of a ligand for a target molecule refers to the strength or degree of attraction between the ligand and the target molecule when they interact to form a complex.
[0016] A 3D spatial position of an atom can be represented by a set of coordinates in an appropriate coordinate system, e.g., a 3D Cartesian coordinate system or a spherical coordinate system.
[0017] “Absorption” properties of a molecule (e.g., a ligand) characterize how well the molecule is absorbed in a subject following administration, and can include one or more of: a solubility of the molecule, a permeability of the molecule, a chemical stability of the molecule, and so forth.
[0018] “Distribution” properties of a molecule (e.g., a ligand) characterize how the molecule spreads through the body of the subject after being administered to the subject, and can include one or more of: a lipophilicity of the molecule, a strength of plasma protein binding of the molecule, a volume of distribution of the molecule, and so forth.
[0019] “Metabolism” properties of a molecule (e.g., a ligand) characterize how the molecule is chemically transformed in the body, and can include one or more of: properties characterizing which enzymatic pathways are responsible for metabolizing the molecule, metabolic rate properties, properties characterizing the activity of metabolites that are generated by metabolism of the molecule, and so forth.
[0020] “Excretion” properties of a molecule (e.g., a ligand) characterize the mechanisms and pathways through which the molecule and its metabolites are removed from the body, and cancharacterize whether the molecule is excreted through renal excretion, or hepatic excretion, or pulmonary excretion, and so forth.
[0021] “Toxicity” properties of a molecule (e.g., a ligand) characterize the potential of the molecule to cause harmful effects to a subject, and can include one or more of: acute toxicity properties, chronic toxicity properties, carcinogenicity properties, organ toxicity properties, reproductive toxicity properties, developmental toxicity properties, and so forth.
[0022] According to one aspect, there is provided a method performed by one or more computers for computationally designing a ligand for binding to a target molecule, the method comprising: obtaining target molecule data characterizing at least a portion of a target molecule; processing a network input comprising the target molecule data using an embedding neural network to generate latent conditioning data representing the target molecule; generating, using a generative model and while the generative model is conditioned on the latent conditioning data representing the target molecule, predicted ligand data defining a predicted ligand that is predicted to bind to the target molecule; and providing the predicted ligand data defining the predicted ligand.
[0023] In some implementations, the method further comprises: obtaining ligand design criteria that define desired characteristics of the ligand; wherein the network input to the embedding neural network further comprises the ligand design criteria; and wherein the latent conditioning data jointly represents the target molecule and the ligand design criteria.
[0024] In some implementations, the embedding neural network and the generative model are jointly trained to optimize an objective function that encourages the embedding neural network and the generative model to attempt to generate ligand data defining a ligand that binds to the target molecule and that satisfies the ligand design criteria.
[0025] In some implementations, the ligand design criteria define a respective target value for each of one or more global properties of the ligand that characterize the ligand as a whole.
[0026] In some implementations, the ligand design criteria define a respective target value for a set of global properties of the ligand characterizing one or more of: a binding affinity of the ligand for the target molecule, absorption properties of the ligand, distribution properties of the ligand, metabolism properties of the ligand, excretion properties of the ligand, toxicity properties of the ligand, a number of rings in the ligand, a molecular weight of the ligand, a lipophilicity of the ligand, an ability of the ligand to donate or accept hydrogen bonds, a total polar surface area of the ligand, a number of rotatable bonds in the ligand, a number of chiral centers in the ligand, a number of electrophilic centers in the ligand, a number of nucleophilic centers in the ligand.
[0027] In some implementations, the ligand design criteria define a respective target value for each of one or more atom-specific properties of the ligand that each relate to a specific atom in the ligand.
[0028] In some implementations, the ligand design criteria define a respective target value for a set of atom-specific properties of the ligand that, for each of one or more atoms in the ligand, characterize one or more of: an elemental type of the atom, a hybridization state of the atom, or a partial charge of the atom.
[0029] In some implementations, the ligand design criteria comprises target molecule scaffolding data, or ligand scaffolding data, or both.
[0030] In some implementations, the target molecule scaffolding data defines a respective target three-dimensional (3D) spatial position of each of one or more atoms in the target molecule when in a complex with the ligand.
[0031] In some implementations, the ligand scaffolding data defines a respective target 3D spatial position of each of one or more atoms in the ligand when in a complex with the target molecule.
[0032] In some implementations, the embedding neural network comprises a target molecule embedding neural network and a design embedding neural network, and wherein processing the network input comprising the target molecule data and the ligand design criteria using the embedding neural network to generate the latent conditioning data comprises: processing the target molecule data using the target molecule embedding neural network to generate a target molecule embedding of the target molecule; processing the ligand design criteria using the design embedding neural network to generate a design embedding of the ligand design criteria; and processing the target molecule embedding and the design embedding to generate the latent conditioning data.
[0033] In some implementations, processing the ligand design criteria using the design embedding neural network to generate the design embedding of the ligand design criteria comprises: generating a collection of initial atom embeddings based on the ligand design criteria, wherein the collection of initial atom embeddings includes a respective atom embedding for each ligand atom in a set of possible ligand atoms that are eligible for inclusion in the ligand; processing the collection of initial atom embeddings, by a plurality of neural network layers of the design embedding neural network, to generate a collection of final atom embeddings that includes a respective final atom embedding for each ligand atom in the set of possible ligand atoms; wherein the collection of final atom embeddings defines the design embedding of the ligand design criteria.
[0034] In some implementations, the ligand design criteria specify a respective target value for each of one or more global properties of the ligand; and generating the collection of initial atom embeddings based on the ligand design criteria comprises: including the respective target value for each of the one or more global properties of the ligand in each initial atom embedding in the collection of initial atom embeddings.
[0035] In some implementations, the ligand design criteria specify a respective target value for each of one or more atom-specific properties of the ligand; and generating the collection of initial atom embeddings based on the ligand design criteria comprises, for each atom-specific property of the ligand: including the target value of the atom-specific property only in the initial atom embedding representing the ligand atom that is characterized by the atom-specific property.
[0036] In some implementations, a number of ligand atom embeddings in the set of ligand atom embeddings defines a maximum number of atoms that can be selected for inclusion in the ligand.
[0037] In some implementations, the plurality of neural network layers of the design embedding neural network include one or more self-attention neural network layers.
[0038] In some implementations, the ligand design criteria leave undefined at least part of a chemical structure of the ligand.
[0039] In some implementations, the target molecule embedding comprises a respective component embedding of each component in the target molecule; the design embedding comprises a respective atom embedding of each ligand atom in a set of ligand atoms that are eligible for inclusion in the ligand, wherein a number of ligand atom embeddings in the set of ligand atom embeddings defines a maximum number of atoms that can be selected for inclusion in the ligand; and processing the target molecule embedding and the design embedding to generate the latent conditioning data comprises: generating data defining a ID sequence of component embeddings and atom embeddings by concatenating: (i) the component embeddings of the target molecule embedding, and (ii) the ligand atom embeddings of the design embedding; and the latent conditioning data is derived from the ID sequence of component embeddings and atom embeddings.
[0040] In some implementations, processing the target molecule embedding and the design embedding to generate the latent conditioning data further comprises: transforming the ID sequence of component embeddings and atom embeddings into a two-dimensional (2D) array of embeddings; wherein the latent conditioning data is derived from the 2D array of embeddings.
[0041] In some implementations, the 2D array of embeddings comprises a plurality of atom - atom embeddings that are each derived from a respective pair of atom embeddings of the design embedding.
[0042] In some implementations, the 2D array of embeddings comprises a plurality of component - component embeddings that are each derived from a respective pair of component embeddings of the target molecule embedding.
[0043] In some implementations, the 2D array of embeddings comprises a plurality of component - atom embeddings that are each derived from: (i) a respective atom embedding of the design embedding, and (ii) a respective component embedding of the target molecule embedding.
[0044] In some implementations, transforming the ID sequence of component embeddings and atom embeddings into the 2D array of embeddings comprises: applying an outer product operation to the ID sequence of component embeddings and atom embeddings; or applying a 2D concatenation operation to the ID sequence of component embeddings and atom embeddings.
[0045] In some implementations, the embedding neural network further comprises a fusion neural network; and processing the target molecule embedding and the design embedding to generate the latent conditioning data further comprises: processing the 2D array of embeddings using the fusion neural network to generate an updated 2D array of embeddings; wherein the updated 2D array of embeddings defines the latent conditioning data.
[0046] In some implementations, the fusion neural network comprises a sequence of selfattention blocks, wherein each self-attention block is configured to perform operations comprising: apply one or more self-attention operations to an input 2D array of embeddings to update the input 2D array of embeddings.
[0047] In some implementations, for one or more of the self-attention blocks, the self-attention operations comprise one or more row-wise self-attention operations.
[0048] In some implementations, for one or more of the self-attention blocks, the self-attention operations comprise one or more column-wise self-attention operations.
[0049] In some implementations, for one or more of the self-attention blocks, the self-attention operations comprise one or more triangle self-attention operations.
[0050] In some implementations, the target molecule is a protein and the target molecule data comprises data defining one or more of: an amino acid sequence of the target molecule; a multiple sequence alignment (MSA) for the target molecule; or a respective structure of each of one or more template target molecules.
[0051] In some implementations, the generative model is a generative diffusion model that comprises a denoising neural network.
[0052] In some implementations, generating, using the generative model and while the generative model is conditioned on the latent conditioning data, the predicted ligand data defining the predicted ligand comprises: generating respective atom state data for each atom in the target molecule and for each ligand atom in a set of ligand atoms that are eligible for inclusion in the ligand; denoising the atom state data over a sequence of time steps using the denoising neural network and while the denoising neural network is conditioned on the latent conditioning data; and generating the predicted ligand data based on the atom state data after a final time step in the sequence of time steps.
[0053] In some implementations, for each ligand atom in the set of ligand atoms, the atom state data for the ligand comprises features characterizing: (i) a 3D spatial position of the ligand atom, and (ii) one or more of: a partial charge of the ligand atom, a hybridization state of the ligand atom, an elemental type of the ligand atom, or a type of an amino acid that includes the ligand atom, or a type of a nucleotide that includes the ligand atom.
[0054] In some implementations, generating respective atom state data for each atom in the target molecule and for each ligand atom in the set of ligand atoms comprises, for one or more atoms: stochastically sampling the atom state data for the atom.
[0055] In some implementations, for one or more atoms, stochastically sampling the atom state data for the atom comprises, for each of one or more continuous features of the atom: stochastically sampling a value of the continuous feature of the atom from a probability distribution; and including the stochastically sampled value of the continuous feature of the atom in the atom state data for the atom.
[0056] In some implementations, the one or more continuous features of the atom comprise respective features defining one or more of: a 3D spatial position of the atom or a partial charge of the atom.
[0057] In some implementations, for one or more atoms, stochastically sampling the atom state data for the atom comprises, for each of one or more categorical features of the atom: stochastically sampling a distribution over possible values of the categorical feature from a probability distribution; and including the stochastically sampled distribution over possible values of the categorical feature of the atom in the atom state data for the atom.
[0058] In some implementations, the one or more categorical features of the atom comprise respective features defining one or more of: a hybridization state of the atom or an elemental type of the atom.
[0059] In some implementations, generating the predicted ligand data based on the atom state data after a final time step in the sequence of time steps comprises: selecting a subset of the ligand atoms in the set of ligand atoms for inclusion in the ligand, wherein fewer than all of the ligand atoms in the set of ligand atoms are selected for inclusion in the ligand; and filtering the set of ligand atoms to remove any ligand atom that is not selected for inclusion in the ligand.
[0060] In some implementations, selecting a subset of the ligand atoms in the set of ligand atoms for inclusion in the ligand comprises, for each ligand atom: selecting the ligand atom for inclusion in the ligand only if a 3D spatial position of the ligand atom, as defined by the atom state data for the ligand atom, is at least a threshold distance from a predefined throw-away position; wherein the generative model has been trained to move respective 3D spatial positions of ligand atoms that are not included in the ligand to the throw-away position.
[0061] In some implementations, generating the predicted ligand data based on the atom state data after the final time step in the sequence of time steps comprises, for each ligand atom in the set of ligand atoms: determining a respective value of each of one or more continuous features of the ligand atom from the atom state data for the ligand atom, comprising, for each continuous feature: extracting a value of the continuous feature from one or more corresponding dimensions of the atom state data for the ligand atom.
[0062] In some implementations, generating the predicted ligand data based on the atom state data after the final time step in the sequence of time steps comprises, for each ligand atom in the set of ligand atoms: determining a respective value of each of one or more categorical features of the ligand atom from the atom state data for the ligand atom, comprising, for each categorical feature: extracting a distribution over possible values of the categorical feature from a plurality of corresponding dimensions of the atom state data for the ligand atom; and determining the value of the categorical feature based on the distribution over possible values of the categorical feature.
[0063] In some implementations, for one or more categorical features of the ligand atom, determining the value of the categorical feature based on the distribution over possible values of the categorical feature comprises: stochastically sampling the value of the categorical feature from the distribution over possible values of the categorical feature; or setting the value of the categorical feature equal to a possible value of the categorical feature having a highest score under the distribution over possible values of the categorical feature.
[0064] In some implementations, denoising the atom state data over the sequence of time steps using the denoising neural network and while the denoising neural network is conditioned on the latent condition data comprises, at each of one or more time steps in the sequence of timesteps: receiving current atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at the time step; generating a denoising output using the denoising neural network and while the denoising neural network is conditioned on the latent conditioning data; and generating atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at a next time step using the denoising output.
[0065] In some implementations, the denoising output comprises a respective predicted error in the current atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at the time step.
[0066] In some implementations, generating the denoising output using the denoising neural network and while the denoising neural network is conditioned on the latent conditioning data comprises: generating a set of atom embeddings using an encoder block of the denoising neural network, wherein each atom embedding represents one or more atoms in the target molecule or in the set of ligand atoms and is based at least in part on the respective current atom state data of the one or more atoms at the time step; processing the set of atom embeddings using an update block of the denoising neural network to generate a set of updated atom embeddings; and processing the set of updated atom embeddings to generate the denoising output.
[0067] In some implementations, the set of atom embeddings includes a respective atom embedding representing each ligand atom in the set of ligand atoms.
[0068] In some implementations, the set of atom embeddings includes a respective atom embedding representing each atom included in each component of the target molecule.
[0069] In some implementations, for each component in the target molecule, the set of atom embeddings includes a respective atom embedding that j ointly represents all the atoms included in the component.
[0070] In some implementations, each atom embedding in the set of atom embeddings is based on, for each of the one or more atoms represented by the atom embedding: (i) the current atom state data of the atom at the time step, and (ii) a respective conditioning embedding for the atom that is selected from a collection of embeddings included in the latent conditioning data.
[0071] In some implementations, for each atom embedding that represents a ligand atom, the conditioning embedding for the atom comprises an atom - atom embedding corresponding to the ligand atom in the latent conditioning data.
[0072] In some implementations, for each atom embedding that represents an atom that is included in a component of the target molecule, the conditioning embedding for the atom comprises a component - component embedding corresponding to the component in the latent conditioning data.
[0073] In some implementations, the update block of the denoising neural network comprises a sequence of self-attention blocks; wherein each of the self-attention blocks are configured to apply one or more self-attention operations to a set of current atom embeddings to update the set of current atom embeddings; wherein each of the one or more self-attention operations are conditioned on the latent conditioning data.
[0074] In some implementations, applying a self-attention operation to the set of current atom embeddings to update the set of current atom embeddings comprises: generating, based on the current set of atom embeddings, a respective intermediate attention score for each pair of current atom embeddings from the set of current atom embeddings; generating, based on the latent conditioning data, a respective attention score bias for each pair of current atom embeddings from the set of current atom embeddings; generating a respective final attention score for each pair of current atom embeddings from the set of current atom embeddings based on the intermediate attention scores and the attention score biases; and updating the set of current atom embeddings using the final attention scores.
[0075] In some implementations, for each pair of current atom embeddings from the set of current atom embeddings, generating the attention score bias for the pair of current atom embeddings comprises: summing the intermediate attention score for the pair of current atom embeddings and the attention score bias for the pair of current atom embeddings.
[0076] In some implementations, for each pair of current atom embeddings from the set of current atom embeddings, generating the attention score bias for the pair of current atom embeddings comprises: processing a respective conditioning embedding selected from a collection of embeddings included in the latent conditioning data using a projection neural network to generate the attention score bias.
[0077] In some implementations, for each pair of current atom embeddings that includes: (i) a first atom embedding representing a first ligand atom, and (ii) a second atom embedding representing a second ligand atom, the selected conditioning embedding comprises an atom - atom embedding corresponding to the first ligand atom and the second ligand atom in the latent conditioning data.
[0078] In some implementations, for each pair of current atom embeddings that includes: (i) a first atom embedding representing a first atom included in an component, and (ii) a second atom embedding representing a ligand atom, the selected conditioning embedding comprises an component - atom embedding corresponding to: (i) the component that includes the first atom, and (ii) the ligand atom, in the latent conditioning data.
[0079] In some implementations, for each pair of current atom embeddings that includes: (i) a first atom embedding representing a first atom included in a first component, and (ii) a second atom embedding representing a second atom included in a second component, the selected conditioning embedding comprises an component - component embedding corresponding to: (i) the first component, and (ii) the second component, in the latent conditioning data.
[0080] In some implementations, for each pair of current atom embedding that includes: (i) a first atom embedding that jointly represents all atoms in a first component, and (ii) a second atom embedding that represents all atoms in a second component, the selected conditioning embedding comprises a component - component embedding corresponding to: (i) the first component, and (ii) the second component, in the latent conditioning data.
[0081] In some implementations, for each pair of current atom embeddings that includes: (i) a first atom embedding that jointly represents all atoms in a component, and (ii) a second atom embedding that represents a ligand atom, the selected conditioning embedding comprises an component - atom embedding corresponding to: (i) the component, and (ii) the ligand atom, in the latent conditioning data.
[0082] In some implementations, denoising the atom state data over the sequence of time steps using the denoising neural network further comprises, at each of one or more time steps in the sequence of time steps: processing at least some of the current atom state data using a property prediction neural network to generate a predicted value of a ligand property of a ligand characterized by the current atom state data; and determining gradients of a conditioning objective function with respect to at least some of the current atom state data, wherein the conditioning objective function measures a discrepancy between: (i) the predicted value of the ligand property, and (ii) a target value of the ligand property; wherein generating atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at the next time step using the denoising output comprises: combining the gradients of the conditioning objective function with the denoising output.
[0083] In some implementations, the property prediction neural network processes the current atom state data for the atoms in the target molecule and for the ligand atoms in the set of ligand atoms.
[0084] In some implementations, the property prediction neural network generates a predicted value of a binding affinity of the ligand for the target molecule.
[0085] In some implementations, the method further comprises: receiving a respective atom embedding for each atom in the ligand; and processing a model input based on the atom embeddings of the atoms in the ligand using a bond prediction machine learning model togenerate bond data that defines, for each pair of atoms in the ligand, whether the pair of atoms are bonded.
[0086] In some implementations, for each atom in the ligand, the atom embedding of the atom is generated as an intermediate output of the generative model.
[0087] In some implementations, the generative model is a generative diffusion model comprising a denoising neural network; and wherein for each atom in the ligand, the atom embedding of the atom is generated as an intermediate output of the denoising neural network at a final time step in a sequence of denoising time steps.
[0088] In some implementations, processing a model input based on the atom embeddings of the atoms in the ligand using the bond prediction machine learning model to generate the bond data comprises: processing a one-dimensional (ID) sequence of atom embeddings of the atoms in the ligand to generate a two-dimensional (2D) array of pair embeddings; processing the 2D array of pair embeddings using the bond prediction machine learning model to generate the bond data.
[0089] In some implementations, the model input to the bond prediction machine learning model further comprises data defining a respective 3D spatial position of each atom in the ligand.
[0090] In some implementations, the method further comprises physically synthesizing the predicted ligand.
[0091] In some implementations, the embedding neural network and the generative model have been trained by performing operations comprising: obtaining data defining a set of target molecule - ligand complexes; determining, for each target molecule - ligand complex in the set of target molecule - ligand complexes, whether the target molecule - ligand complex satisfies a set of filtering criteria; and during at least one training stage, training the embedding neural network and the generative model based only on training examples derived from target molecule - ligand complexes that satisfy the set of filtering criteria.
[0092] In some implementations, for each target molecule - ligand complex in the set of target molecule - ligand complexes, the target molecule - ligand complex satisfies the set of filtering criteria only if the ligand of the target molecule - ligand complex has one or more specified ligand properties.
[0093] In some implementations, the one or more specified ligand properties comprise functional properties, comprising one or more of: full agonism, partial agonism, antagonism, or inverse agonism.
[0094] In some implementations, for each target molecule - ligand complex in the set of target molecule - ligand complexes, the target molecule - ligand complex satisfies the set of filtering criteria if either: (i) the ligand of the target molecule - ligand complex has one or more specified ligand properties and the target molecule of the target molecule - ligand complex is included in a specified target molecule class, or (ii) the target molecule of the target molecule - ligand complex is not included in the specified target molecule class.
[0095] In some implementations, the one or more specified ligand properties comprise functional properties, comprising one or more of: full agonism, partial agonism, antagonism, or inverse agonism; and wherein the target molecule class is a G-protein-coupled receptor (GPCR) protein class.
[0096] In some implementations, during at least one training stage, training the embedding neural network and the generative model based only on training examples derived from target molecule - ligand complexes that satisfy the set of filtering criteria comprises: at a first training stage, training the embedding neural network and the generative model based on training examples derived from all target molecule - ligand complexes in the set of target molecule - ligand complexes; and at a second training stage, training the embedding neural network and the generative model based only on training examples derived from target molecule - ligand complexes that satisfy the set of filtering criteria.
[0097] In some implementations, the target molecule is a protein.
[0098] In some implementations, the ligand is a protein.
[0099] According to another aspect, there is provided a method comprising: generating a collection of ligands for a target molecule using the methods described herein; determining, for each ligand in the collection of ligands, one or more respective properties of the ligand; and selecting one or more ligands in the collection of ligands for physical synthesis based at least in part on the properties of the ligands.
[0100] In some implementations, the method further comprises physically synthesizing the one or more selected ligands.
[0101] According to another aspect there is provided a method of obtaining a ligand, wherein the ligand is a drug or a ligand of an industrial enzyme, the method comprising: performing the methods described herein to determine a plurality of candidate ligands for a target molecule; evaluating a respective interaction of each candidate ligand with the target molecule; and selecting one or more of the candidate ligands dependent on a result of the evaluating.
[0102] In some implementations, the target molecule comprises a receptor or enzyme, and wherein each candidate ligand is an agonist or antagonist of the receptor or enzyme.
[0103] In some implementations, the target molecule comprises an antibody or aptamer target, in particular a virus or cancer cell protein, and wherein the ligand binds to the antibody or aptamer target to provide a therapeutic effect.
[0104] According to another aspect, there is provided a method of obtaining a ligand, wherein the ligand is for modifying one or both of plant growth and stress resistance of an agricultural plant, the method comprising: performing the methods described herein to determine a plurality of candidate ligands for a target molecule identified as being involved in one or both of plant growth and stress resistance of the agricultural plant; evaluating a respective interaction of each candidate ligand with the target molecule; and selecting one or more of the candidate ligands as the ligand dependent on a result of the evaluating.
[0105] According to another aspect, there is provided a system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the methods described herein.
[0106] According to another aspect, there are provided one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the methods described herein.
[0107] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0108] Drug discovery can involve identifying specific molecules within the body that are involved in a disease process. These molecules are often proteins, such as enzymes, receptors, or signaling proteins, that play a key role in the disease's development or progression. A ligand, often a small molecule, peptide, or antibody, can be selected to bind specifically to an identified target molecule and modify its biological activity. When a drug that includes the ligand is administered to a patient, the ligand can bind to the target molecule with high affinity and in doing so contribute to achieving a therapeutic effect in the patient. For instance, if the target molecule is an enzyme involved in a disease process, the ligand can inhibit its activity, thus disrupting the disease pathway. More generally, the interaction between the ligand and the target molecule can activate, inhibit, or alter the function of the target molecule to achieve a therapeutic effect. Therefore, identifying ligands with high binding affinity for target molecules can be a crucial step in the process of drug discovery, and also in other areas such as pest or pathogen control in agriculture, modifying plant species to improve growth and stress resistance, and so forth, as will be described in more detail later in this specification.
[0109] Traditionally, identifying a ligand that will bind to a target molecule involves screening large databases of known molecules using molecular docking. Molecular docking of a target molecule and a ligand involves obtaining data defining respective 3D structures of the target molecule and the ligand, and performing a search through a space of possible poses of the target molecule and the ligand to optimize a scoring function. The scoring function can measure, e.g., the energy of each joint conformation of the target molecule and the ligand. Conventional molecular docking can be computationally expensive, e.g., because optimizing the scoring function requires searching a large space of possible poses of the target molecule and the ligand. Moreover, conventional molecular docking requires advance knowledge of the individual 3D structures of the target molecule and the ligand, and of the binding site(s) on the target molecule. Further, even if the 3D structure of the target molecule is known, e.g., from crystallography, the 3D structure of the target molecule may deform through a process of conformational change as the ligand interacts with the target molecule, e.g., to bind to a binding site on the target molecule. However, the process of conventional molecular docking does not account for potential conformational changes of the target molecule as a result of interaction with the ligand which can lead to inaccurate results. Further, traditional approaches for identifying ligands that bind to a target molecule by screening databases of known molecules can only identify ligands that are included in molecule databases which are screened. However, such molecule databases, however large, cover only a small fraction of the space of possible ligands.
[0110] The design system described in this specification addresses these issues by directly mapping from data defining a target molecule, e.g., the amino acid sequence of a protein and an MSA for the protein, to data defining a ligand that is predicted to bind to the target molecule. The design system does not require advance knowledge of the 3D structure of the target molecule or of the binding site(s) on the target molecule. To generate a ligand that will bind to a target molecule, the design system can generate latent conditioning data that represents the target molecule (and, optionally, also represents a set of ligand design criteria, as will be described in more detail below). The system can then use the latent conditioning data to condition a generative model that can generate one or more ligands that will bind to the target molecule by sampling from a distribution over a space of possible ligands.
[0111] The design system thus overcomes many of the disadvantages of traditional ligand identification approaches described above. For instance, the design system may consume significantly fewer computational resources (e.g., memory and computing power) than a conventional molecular docking approach, as the design system can generate a ligand by asingle forward pass through a neural network system. In contrast, a conventional molecular docking approach requires iteratively searching through a large space of possible poses of the target molecule and a candidate ligand to evaluate whether the candidate ligand binds to the protein, and then this must be repeated for each candidate ligand in a molecule database. As another example, the generative model of the design system generates a ligand by sampling from a distribution over a space of possible ligands, and can thus generate a vastly larger variety of ligands than a conventional screening approach that can only identify ligands from a predefined molecule library.
[0112] In many practical applications, merely identifying a ligand that will bind to a target molecule may be insufficient. In particular, to be practically useful, a ligand may be required to satisfy additional design criteria, e.g., relating to absorption, distribution, metabolism, excretion, and toxicity characteristics of the molecule. The design system can address this issue in a variety of ways, a few of which are described next.
[0113] For instance, the design system can enable a user (or another upstream system) to provide an input that specifies both: (i) a target molecule, and (ii) a set of ligand design criteria specifying one or more target (desired) properties of the ligand. The target properties of the ligand can include both global properties of the ligand, i.e., that characterize the ligand as a whole, such as absorption, distribution, metabolism, and so forth, and also atom-specific properties of the ligand, e.g., such as the elemental types of particular atoms in the ligand. The embedding neural network of the design system can jointly embed data representing the target molecule and the ligand design criteria in a set of latent conditioning data, and then condition the generative model on the latent conditioning data. The generative model, when conditioned on the latent conditioning data, is configured through training to generate ligands that both bind to the target molecule and also satisfy the ligand design criteria.
[0114] As another example, the generative model of the design system can be implemented as a generative diffusion model, and the design system can actively steer the iterative denoising process implemented by the generative diffusion model to increase the likelihood that the generated ligand satisfies particular ligand design criteria. More specifically, the generative diffusion model can generate a ligand by progressively denoising atom state data for the atoms in the ligand over a sequence of denoising iterations (e.g., time steps). At each denoising iteration, the design system can process the current atom state data for the ligand using a property prediction neural network to generate a predicted value of a property of the ligand. The design system can determine gradients, with respect to the current atom state data for the ligand, of a conditioning objective function that measures a discrepancy between: (i) thepredicted value of the property of the ligand, and (ii) a target (desired) value of the property of the ligand. The system can then use these gradients as part of adjusting the current atom state data in order to increase the likelihood that the ligand will satisfy ligand design criteria related to the property.
[0115] As another example, the design system can train the embedding neural network and the generative model on training examples derived from a set of target molecule - ligand complexes that have been filtered in a manner that encourages the generation of ligands satisfying particular design criteria, e.g., having specified functional properties. More specifically, the system can filter a set of target molecule - ligand complexes using a set of filtering criteria, and then generate training examples for training the embedding neural network and the generative model based on only the target molecule - ligand complexes that remain after the filtering. The system can select the filtering criteria to increase the likelihood that the embedding neural network and the generative model, when trained on training examples derived from the filtered set of target molecule - ligand complexes, will generate ligands having desired properties.
[0116] By implementing conditioning mechanisms that facilitate the generation of ligands satisfying particular design criteria, the design system can enable reduced consumption of computational resources, e.g., memory and computing power. More specifically, in the absence of such conditioning mechanisms, identifying a ligand that satisfies particular design criteria may require generating very large numbers of candidate ligands using the design system, and then individually screening the candidate ligands (e.g., using computational methods) to identify those satisfying the design criteria. Providing conditioning mechanisms that encourage the generation of ligands satisfying design criteria can thus reduce computational resource consumption by reducing the number of candidate ligands that must be generated by the design system, and also reduce the number of screening operations that are performed to screen candidate ligands generated by the design system.
[0117] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0118] FIG. 1 shows an example design system.
[0119] FIG. 2 is a flow diagram of an example process for generating and screening ligands for binding to a target molecule.
[0120] FIG. 3A shows an example embedding neural network.
[0121] FIG. 3B is a flow diagram of an example process for processing ligand design criteria using a design embedding neural network.
[0122] FIG. 3C provides an illustration of a collection of initial atom embeddings generated by an embedding block of a design embedding neural network.
[0123] FIG. 4A is a flow diagram of an example process for processing a target molecule embedding of a target molecule and a design embedding of ligand design criteria using a fusion neural network to generate latent conditioning data that jointly represents the target molecule and the ligand design criteria.
[0124] FIG. 4B illustrates operations performed by the design system to generate the latent conditioning data.
[0125] FIG. 5 is a flow diagram of an example process for generating a data defining a ligand that is predicted to bind to a target molecule using a generative diffusion model that includes a denoising neural network.
[0126] FIG. 6 is a flow diagram of an example process for generating a denoising output using a denoising neural network conditioned on latent conditioning data.
[0127] FIG. 7 is a flow diagram of an example process for updating a set of current atom embeddings using a self-attention operation that is implemented by a self-attention block of the denoising neural network and that is conditioned on the latent conditioning data.
[0128] FIG. 8 is a flow diagram of an example process generating gradients of a conditioning objective function with respect to current atom state data for the atoms in a target moleculeligand complex at a current denoising time step in a sequence of denoising time steps.
[0129] FIG. 9 is a flow diagram of an example process for generating data defining a ligand based on the denoised atom state data for each ligand atom.
[0130] FIG. 10 is a flow diagram of an example process for jointly training the embedding neural network and the generative model of the design system.
[0131] FIG. 11 is a flow diagram of an example process for jointly training an embedding neural network and a generative diffusion model on a training example.
[0132] FIG. 12 is a flow diagram of an example process for generating bond data for a ligand generated by the design system.
[0133] FIG. 13 is a flow diagram of an example process for in-painting and / or out-painting a protein-ligand complex.
[0134] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0135] FIG. 1 shows an example design system 100. The design system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0136] The design system 100 is configured to process an input including target molecule data 102 characterizing at least a portion of one or more target molecules to generate an output including data defining one or more ligands 116 that are predicted to form a molecule complex with the one or more target molecules.
[0137] In some cases, the target molecule data 102 characterizes (at least a portion of) a single target molecule. In other cases, the target molecule data 102 characterizes (at least portions of) multiple target molecules.
[0138] In some cases, the design system 100 generates data characterizing a single ligand 116 that is predicted to form a complex with the one or more target molecules. In other cases, the design system generates data characterizing multiple ligands 116 that jointly form a complex with the one or more target molecules. The number of ligands that are defined by the output of the design system 100 can be a predefined default value (e.g., one), or can be specified in a set of ligand design criteria 104 that can be provided as an additional input to the design system 100 (as will be described in more detail below).
[0139] The target molecule data 102 can include any appropriate data characterizing the one or more target molecules. A few examples of target molecule data 102 are described in more detail next.
[0140] In one example, the target molecule data 102 can include data characterizing a protein, such as, e.g., data defining one or more amino acid sequences of the protein, or data defining an MSA for the protein, or data characterizing a respective structure of each of one or more template proteins, or data characterizing a structure of the protein itself, or a combination thereof. In some cases, the target molecule data 102 can characterize a full protein or protein complex. In other cases, the target molecule data 102 can include data characterizing only a portion of a protein, e.g., only a binding pocket of a protein.
[0141] In another example, the target molecule data 102 can include data characterizing a nucleic acid molecule, e.g., a deoxyribonucleic acid (DNA) molecule or a ribonucleic acid (RNA) molecule. For instance, the target molecule data 102 can include data defining one or more nucleotide sequences of the nucleic acid molecule, data characterizing a structure of thenucleic acid molecule, data indicating the presence and type of any modified nucleotides (e.g., methylation, pseudouridine), data characterizing a respective structure of each of one or more template nucleic acid molecules, or a combination thereof.
[0142] In another example, the target molecule data 102 can include data characterizing a lipid molecule, e.g., a fatty acid, or a triglyceride, or a phospholipid, or a steroid (e.g., cholesterol, hormones), and so forth. For instance, the target molecule data 102 can include data defining a chemical structure of the lipid molecule (e.g., the length of one or more hydrocarbon chains included in the lipid molecule; the number and position of double bonds in the lipid molecule; the presence of phosphate, hydroxyl, carboxyl, or other polar head groups; and / or the presence of linear vs. branched chains), data characterizing a 3D structure of the lipid, and so forth.
[0143] In another example, the target molecule data 102 can include data characterizing a carbohydrate molecule, e.g., a monosaccharide (e.g., glucose, fructose), or a disaccharide (e.g., sucrose, lactose), or a polysaccharide (e.g., starch, glycogen, cellulose). For instance, the target molecule data 102 can include data defining one or more sequences of monosaccharide units that are included in the carbohydrate molecule, data identifying an anomeric configuration of the monosaccharides (e.g., whether the monosaccharides are in the alpha or beta configuration at the anomeric carbon), data characterizing a 3D structure of the carbohydrate molecule, and so forth.
[0144] In some implementations, the target molecule data for a target molecule can include data characterizing one or more known binders for the target molecule, e.g., ligands that are already known to bind to the target molecule. The known binders can be any appropriate type of molecule (e.g., small molecule, protein, nucleic acid, and so forth), and the target molecule data can characterize the chemical structure, 3D structure, or both, of the known binders.
[0145] Generally, the target molecule data 102 can characterize any appropriate number of target molecules (e.g., 1, or 3, or 5 target molecules) of any appropriate types (e.g., a combination of one or more of: protein molecules, or nucleic acid molecules, or lipid molecules, or carbohydrate molecules, and so forth).
[0146] For convenience, in the following description, the target molecule data 102 may be referred to as characterizing “a” target molecule, but it is understood that more generally the target molecule data 102 can characterize any appropriate number of target molecules. Similarly, in the following, the output of the design system 100 may be referred to as identifying “a” ligand 116, but more generally, the output of the design system 100 can characterize one or more ligands 116.
[0147] Optionally, the input to the design system 100 can further include ligand design criteria 104 that specify: (i) a respective target (desired) value for each of one or more properties of the ligand (i.e., the target ligand properties 106), or (ii) scaffolding data 108, or (iii) both. Target ligand properties 106 and scaffolding data 108 are each described in more detail next (and throughout this specification). Optionally, the ligand design criteria 104 can further specify a number of ligands to be generated by the design system 100. The number of ligands can be, e.g., one ligand, or more than one ligand, e.g., 2 ligands, or 3 ligands, or 4 ligands. In cases where the ligand design criteria specify that multiple ligands are to be generated by the design system 100, the ligand design criteria 104 can include respective target ligand properties 106, scaffolding data 108, or both, for each ligand.
[0148] The target ligand properties 106 can specify a target (desired) value for any appropriate “global” or “atom-specific” properties of a ligand.
[0149] Global properties of a ligand can refer to properties that characterize the ligand as a whole rather than being specific to a single atom, and can include properties characterizing one or more of: a binding affinity of the ligand for a target molecule, functional properties of the ligand, absorption properties of the ligand, distribution properties of the ligand, metabolism properties of the ligand, excretion properties of the ligand, toxicity properties of the ligand, a number of rings (e.g., aromatic rings) in the ligand, a molecular weight of the ligand, a lipophilicity (e.g., a logP value specifying the value of the logarithm of an octanol / water partition coefficient for the ligand) of the ligand, an ability of the ligand to donate or accept hydrogen bonds, a total polar surface area of the ligand, a number of rotatable bonds in the ligand, a number of chiral centers (stereocenters) in the ligand, a number of electrophilic (electron-accepting) centers in the ligand, a number of nucleophilic (electron-donating) centers in the ligand, and so forth. Functional properties of the ligand characterize the biological or physiological effect that the ligand induces when it binds to a target molecule, e.g., in terms of efficacy (e.g., full agonism, partial agonism, antagonism, or inverse agonism), potency (e.g., the concentration or amount of the ligand needed to produce a certain level of response in a subject), selectivity (e.g., how selectively the ligand binds to target molecules of a particular type over others), or duration of action (e.g., the length of time the ligand exerts its functional effects once bound to a target molecule). Global properties can also therefore include pharmacokinetic and / or pharmacodynamic properties of the ligand.
[0150] In some cases, global properties of a ligand can identify a binding site on a target molecule where the ligand is to bind in the complex comprising the ligand and the targetmolecule. For instance, the ligand design criteria 104 can identify a binding site on a target molecule by identifying a subset of the atoms that are included in the binding site.
[0151] In some cases, global ligand properties of the ligand can characterize one or more “off- target” molecules. In contrast to the target molecule(s), to which the ligand is intended to bind, the design system designs the ligand to avoid binding to the “off-target” molecules. The off- target molecules can be any appropriate types of molecules (e.g., proteins or nucleic acids), and the global ligand properties can specify any appropriate number of off-target molecules, e.g., 1, or 3, or 10, or 100 off-target molecules. The global ligand properties can include data characterizing the chemical structure, or the 3D spatial structure, or both, of the off-target molecules.
[0152] Atom-specific properties of a ligand can refer to properties that relate to specific atoms in the ligand rather than the entire ligand, and can include properties characterizing one or more of: an elemental type of an atom (e.g., carbon, oxygen, nitrogen, and so forth), a hybridization state of an atom (e.g., sp hybridization, or sp2hybridization, or sp3hybridization, or sp3d hybridization, or sp2d2hybridization, and so forth), a partial charge of an atom, and so forth.
[0153] In some cases, atom-specific properties of a ligand can further identify pairwise covalent or non-covalent interactions between: (i) an atom or functional group in a ligand, and (ii) an atom or functional group in a target molecule, in the complex comprising the ligand and the target molecule. The non-covalent interactions can include, e.g., hydrogen bonds, ionic bonds, Van der Waals forces, hydrophobic interactions, metal coordination, and so forth. As another example, an atom-specific property can specify a stereochemistry of an atom in the ligand, e.g., an R or S designation for the atom, e.g., as defined by application of the Cahn- Ingold-Prelog rules.
[0154] The scaffolding data 108 can specify target molecule scaffolding data, or ligand scaffolding data, or both. The target molecule scaffolding data can specify (at least) a portion of a 3D structure of a target molecule, in particular, by specifying a respective 3D spatial position of each of one or more atoms in the target molecule. The ligand scaffolding data can specify (at least) a portion of the 3D structure of the ligand, in particular, by specifying a respective 3D spatial position of each of one or more atoms in the ligand.
[0155] Optionally, the ligand design criteria can include data specifying a number of ligands to be generated by the design system 100, data identifying a respective “type” of each ligand (e.g., whether the ligand is a small molecule, or a protein, or a carbohydrate, or a nucleic acid, and so forth), or both.
[0156] In cases where the input to the design system 100 includes ligand design criteria 104, the design system 100 attempts to generate a ligand 116 (or, in some cases, multiple ligands 116) that satisfy the ligand design criteria 104. For instance, if the ligand design criteria 104 include target ligand properties 106, then the design system 100 attempts to generate a ligand 116 that has the target ligand properties 106. As another example, if the ligand design criteria 104 include target molecule scaffolding data, then the design system 100 attempts to generate a ligand 116 that binds to the target molecule when the target molecule has the conformation defined by the target molecule scaffolding data. As another example, if the ligand design criteria 104 include ligand scaffolding data, then the design system 100 attempts to generate a ligand 116 that has the structure defined by the ligand scaffolding data.
[0157] In particular, the design system 100 has a set of design system parameters that are configured through training (e.g., by a machine learning training technique) to encourage the generation of a ligand 116 that satisfies the ligand design criteria 104. The set of design system parameters can include respective sets of parameters of an embedding neural network 300 and of a generative model 112, as will be described in more detail below.
[0158] In some cases, a ligand 116 generated by the design system 100 may satisfy all the ligand design criteria 104 specified in the input to the design system 100. However, in other cases, a ligand 116 generated by the design system 100 may satisfy certain ligand design criteria 104 only approximately, or may entirely fail to satisfy certain ligand design criteria 104. This may occur, for instance, if the ligand design criteria 104 are mutually incompatible (e.g., if there does not exist a ligand that binds to a target molecule and that simultaneously satisfies all the ligand design criteria 104), or if the design system parameters have not been trained on a sufficient amount or type of training data to enable the precise generation of a ligand that binds to a target molecule and that satisfies all the ligand design criteria 104. In particular, providing ligand design criteria 104 to the design system 100 increases the likelihood, but does not guarantee, that the design system 100 will generate a ligand 116 that satisfies the ligand design criteria 104.
[0159] Generally, the ligand design criteria 104 do not specify the full chemical structure and 3D spatial structure of the ligand 116. In particular, the ligand design criteria 104 leave undefined at least parts of the chemical structure or 3D spatial structure of the ligand.
[0160] The data defining the ligand 116 can include, for each atom in the ligand 116, respective atom state data for the atom that defines at least a respective 3D spatial position of the atom (i.e., in a complex that includes the target molecule(s)). The atom state data for an atom canfurther include any other appropriate atom-specific properties of the atom, such as an elemental type of the atom, a hybridization state of the atom, a partial charge of the atom, and so forth.
[0161] Optionally, the data defining the ligand 116 can further include bond data that defines, for each pair of atoms in the ligand, whether the pair of atoms are connected by a bond. The bond data can further define one or more respective properties of each bond in the ligand, e.g., the type of the bond, e.g., single, double, or triple covalent bond, or ionic bond, or coordinate covalent bond, and so forth.
[0162] Optionally, when the ligand 116 is a protein, the data defining the ligand 116 can include data defining one or more amino acid sequences of the ligand.
[0163] Optionally, when the ligand 116 is a nucleic acid molecule (e.g., a DNA or RNA molecule), the data defining the ligand 116 can include data specifying one or more nucleic acid sequences of the ligand.
[0164] The design system 100 includes an embedding neural network 300, a generative model 112, and (optionally) a bond prediction machine learning model 114, which are each described in more detail next (and throughout this specification).
[0165] The embedding neural network 300 is configured to process a network input representing the target molecule data 102 and (optionally) the ligand design criteria 104, in accordance with values of a set of embedding neural network parameters, to generate latent conditioning data 110. The latent conditioning data 110 represents the target molecule(s) 102, and when the input to the embedding neural network includes ligand design criteria 104, jointly represents the target molecule(s) 102 and the ligand design criteria 104. The latent conditioning data is “latent,” e.g., in the sense that is represented in a latent embedding space (e.g., a Euclidean space having an appropriate dimensionality), and is “conditioning data” in the sense that is used for conditioning the generative model 112, as will be described in more detail below.
[0166] The embedding neural network 300 can have any appropriate neural network architecture that enables the embedding neural network 300 to perform its described functions. In particular, the embedding neural network 300 can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g., 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers). A particular example of a possible architecture of the embedding neural network 300 is described in more detail with reference to FIG. 3.
[0167] The generative model 112, when conditioned on the latent conditioning data 110, is configured to generate data defining a ligand 116 (in some cases, multiple ligands 116), in particular, to generate respective atom state data for each atom in the ligand that defines at least a respective 3D spatial position of the atom (as described above). The generative model 112 can be any appropriate conditional generative model. More specifically, the generative model 112 can be any appropriate model that, when conditioned on the latent conditioning data 110, can generate samples from a distribution over a space of possible ligands. For instance, the generative model 112 can be implemented as a generative diffusion model, or a generative adversarial neural network (GAN) model, or a flow-based neural network model (normalizing flow model), and so forth. An example process for generating ligands using a generative diffusion model is described detail with reference to FIG. 5.
[0168] The bond prediction machine learning model 114 is configured to process a model input characterizing the ligand 116 to generate bond data that identifies bonds present in the ligand, and optionally, one or more properties of the bonds present in the ligand (as described above). An example of a bond prediction machine learning model 114 is described in more detail below with reference to FIG. 12.
[0169] In some implementations, rather than generating a single ligand 116, the design system 100 can generate multiple distinct ligands 116. In particular, the design system 100 can use the generative model 112 to generate multiple samples from the distribution over the space of possible ligands, each of which represents a respective different ligand that is predicted to form a complex with the target molecule(s) and to satisfy any ligand design criteria 104. Each of the ligands 116 generated by the design system 100 can have different chemical structures and properties and can satisfy any ligand design criteria 104 to different degrees.
[0170] The design system 100 jointly trains the embedding neural network 300 and the generative model 112 to encourage that the generative model, when conditioned on latent conditioning data 110 that is generated by the embedding neural network and that represents a target molecule(s) and (optionally) ligand design criteria, generates ligands that form a complex with the target molecule(s) and that satisfy the ligand design criteria. In particular, the design system 100 can jointly train the embedding neural network 300 and the generative model 112 (and, optionally, the bond prediction machine learning model 114), on a set of training data using an appropriate machine learning training technique. The training data can include a set of training examples, where each training example corresponds to a molecule complex of a target molecule(s) and a ligand(s), e.g., where each ligand is bound to a respective binding site on a target molecule. Training examples are available from publicly or commercially availabledatabases, such as: ChEMBL(https: / / www.ebi. ac.uk / chembl / ), which is a manually curated database of bioactive molecules with drug-like properties; BioLip, which is a ligand-protein binding database; DUD-E (https: / / dude.docking.org / ); SitesBase; etc. The embedding neural network 300 and the generative model 112 can be jointly trained to optimize an objective function that encourages the embedding neural network and the generative model to attempt to generate ligand data defining a ligand that binds to the target molecule and that satisfies the ligand design criteria. That is, by jointly training the embedding neural network 300 and the generative model 112 to optimize the objective function, the embedding neural network 300 and the generative model 112 become configured to generate (or have at least a greater tendency to generate) ligand data defining a ligand that binds to the target molecule and that satisfies the ligand design criteria. An example process for jointly training the embedding neural network 300 and the generative model 112 (and, optionally, the bond prediction machine learning model 114) is described in more detail with reference to FIG. 10.
[0171] The design system 100 can be used to design ligands 116 for any of a variety of applications. A few example applications of ligands 116 generated by the design system 100 are described next.
[0172] In one example, the design system can be used for drug design, in particular, to generate one or more ligands that are predicted to bind to a target molecule (e.g., protein) that is identified as being involved in a disease process, e.g., associated with cancer, Alzheimer’s disease, heart disease, infectious diseases (e.g., bacterial diseases, viral diseases, parasitic diseases, fungal diseases, prion diseases, etc.), and so forth. By binding to the target molecule, a ligand can modulate (e.g., inhibit or activate) the activity of the target molecule thereby disrupting the disease process and contributing to treating the disease.
[0173] As another example, the design system can be used for pest or pathogen control in agriculture, in particular, to generate one or more ligands that are predicted to bind to target molecules (e.g., proteins) in agricultural pests or pathogens (e.g., insects, fungi, or bacteria). Such ligands can be used as part of targeted pesticides that bind to target molecules that are essential to the survival of pests or pathogens and that are not found in non-target species.
[0174] As another example, the design system can be used for modifying plant species cultivated for agricultural purposes, in particular, to generate one or more ligands that are predicted to bind to target molecules identified as being involved in plant growth, or stress resistance, or both. (Stress resistance in a plant species can characterize an ability of the plant species to survive and adapt to adverse conditions such as drought, high soil salinity, extremetemperatures, and so forth). Such ligands can be used for modulating the behavior of target molecules in the plant species to increase crop yields and stress resistance.
[0175] The design system 100 can receive requests to perform “in-painting” or “out-painting” of a molecule complex. A request to in-paint the complex identifies, for each atom property of each atom in the complex, whether the atom property is a “static” property or a “variable” property. In-painting the complex refers to generating new ligands that have the static atom properties of the input ligand and that bind to target molecules that have the static atom properties of the input target molecules. Out-painting the complex refers to generating new ligands that expand on the original ligand, e.g., by inclusion of one or more new atoms. An example process for in-painting and out-painting a complex is described with reference to FIG. 13.
[0176] In a particular example, the design system can be used to jointly design an enzyme and a cofactor associated with a particular substrate. An enzyme is a biological catalyst, typically a protein, that accelerates the rate of a specific chemical reaction without being consumed in the process. A cofactor is a non-protein chemical compound or metallic ion that is required to enable the biological activity of an enzyme. A substrate is a specific reactant molecule upon which an enzyme acts. In this example, the target molecule data can characterize the substrate, while the ligand design criteria can specify the generation of two “ligands” to form a complex with the substrate - one being the enzyme, the other being the cofactor.
[0177] In another particular example, the design system can be used to design a bi-specific ligand that binds concurrently to two different target molecules. For instance, the bi-specific ligand can be designed to concurrently bind to both a cancer cell and an immune cell to promote a targeted immune response. In this example, the target molecule data can characterize both the target molecules, and the ligand design criteria can specify the generation of one ligand to form a complex with both target molecules.
[0178] In another particular example, the design system can be used to redesign a ligand for binding to a target molecule. More specifically, in this example, the system can receive an input ligand that is a protein or nucleic acid that is known to have a 3D structure that enables effective binding to a target molecule. The system can generate one or more redesigned versions of the input ligand that are predicted have the same (or a similar) 3D structure as the input ligand, but that are characterized by a different underlying component sequence, e.g., a different amino acid sequence for a protein ligand, or a different nucleotide sequence for a nucleic acid ligand. In particular, the system can generate redesigned versions of the input ligand by generating and processing ligand design criteria 104 that include scaffolding data defining some or all the 3Dstructure of the input ligand. The system can then process the ligand design criteria (along with target molecule data for the target molecule) to generate output ligands with the same (or a similar) 3D structure as the input ligand but with different underlying component sequences. The redesigned versions of the input ligand may have properties (e.g., absorption, distribution, metabolism, excretion, toxicity, solubility, etc.) that are more desirable than the original input ligand.
[0179] In a particular example, the target molecule may be an antigen and the input ligand may be an antibody that is known to bind to the antigen. In this example, the system can be used to redesign the amino acid sequence of the antibody to generate new antibodies that bind to the antigen but have more desirable properties, e.g., that are less likely to trigger an anti-drug response in animal or human subjects.
[0180] FIG. 2 is a flow diagram of an example process 200 for generating and screening ligands for binding to a target molecule. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200.
[0181] The system receives target molecule data characterizing at least a portion of a target molecule, and optionally, ligand design criteria (202). The ligand design criteria specify criteria to be satisfied by a ligand to be generated by the system. The ligand design criteria can include target ligand properties, or scaffolding data, or both. The target ligand properties can specify a respective target value for each of one or more properties of the ligand. The scaffolding data can include target molecule scaffolding data (specifying at least a portion of the 3D structure of the target molecule) or ligand scaffolding data (specifying at least a portion of the 3D structure of the ligand) or both.
[0182] The system can receive the target molecule data and the ligand design criteria from any appropriate source, e.g., from a user or from another system, by way of an appropriate interface, e.g., an application programming interface (API) or a user interface (e.g., a graphical user interface).
[0183] The system processes the target molecule data and any ligand design criteria using the embedding neural network to generate latent conditioning data representing the target molecule(s) and any ligand design criteria (204). An example architecture of the embedding neural network is described in more detail below with reference to FIG. 3.
[0184] The system generates, using a generative model and while the generative model is conditioned on the latent conditioning data representing the target molecule(s) and any liganddesign criteria, data defining a set of ligands that are each predicted to form a complex with the target molecule(s) (206). The system generates each ligand in a manner that encourages the ligand to satisfy any ligand design criteria. The set of ligands can include any appropriate number of ligands, e.g., 1, 10, 100, or 1000 ligands. An example process for generating data defining a ligand using a generative diffusion model is described in more detail with reference to FIG. 5.
[0185] The system generates, for each ligand in the set of ligands and using a bond prediction machine learning model, bond data for the ligand that defines, for each pair of atoms in the ligand, whether the pair of atoms are connected by a bond (208). The bond data for a ligand can further define one or more respective properties of each bond in the ligand, e.g., the type of the bond, e.g., single, double, or triple covalent bond, or ionic bond, or coordinate covalent bond, and so forth.
[0186] An example process for generating bond data for a ligand using a bond prediction machine learning model that processes atom embeddings generated for the atoms in the ligand by a generative model implemented as a generative diffusion model is described with reference to FIG. 12.
[0187] The system filters the set of ligands to remove any ligand that fails to satisfy each acceptance criterion in a set of one or more acceptance criteria (210). The set of acceptance criteria can include any appropriate acceptance criteria. A few examples of possible acceptance criteria are described next.
[0188] In one example, an acceptance criterion for a ligand can be that the ligand does not contain any structures that are designated as being physically impossible or highly unstable, e.g., a square planar carbon structure.
[0189] In another example, an acceptance criterion for a ligand can be that each atom in the ligand is bonded to at least one other atom in the ligand.
[0190] In another example, an acceptance criterion for a ligand can be that the value of a molecular property of the ligand satisfies one or more property-specific thresholds. The molecular property of the ligand can be, e.g., any of the global ligand properties or atomspecific ligand properties described earlier. The property-specific threshold for a molecular property can include, e.g., a lower bound on the value of the property, or an upper bound on the value of the property, or both.
[0191] The system can set the one or more property-specific thresholds for a ligand property in any appropriate way. For instance, the system can set a property-specific threshold for a molecular property based on a user input received from a user of the system by way of a userinterface or an API made available by the system. As another example, the system can set a property-specific threshold for a molecular property based on a target value specified for the molecular property in the ligand design criteria processed by the system. For instance, if the ligand design criteria specify a target value for a molecular property, the system can set property-specific thresholds for the molecular property that include a lower bound and an upper bound on the value of the molecular property, where the lower and upper bound jointly define a range of values centered on the target value for the molecular property. Thus, the system can include acceptance criteria requiring that a ligand, in order to avoid being filtered from the set of possible ligands, must at least approximately satisfy some or all of the ligand design criteria.
[0192] As another example, an acceptance criterion for a ligand can be that, when the set of ligands are ranked based on the value of a particular molecular property, the ligand is included in a predefined number (e.g., 10, or 100, or 1000) or predefined percentage (e.g., 1%, or 5%, or 10%) of highest (or lowest) ranked ligands. For instance, the molecular property can be binding affinity for the protein, and the acceptance criterion can require that a ligand, in order to avoid being filtered from the set of possible ligands, must be among a predefined number or percentage of top-ranked ligands in the set of ligand when ranked based on binding affinity for a target molecule.
[0193] In order to evaluate whether a ligand satisfies an acceptance criterion (as described above), the system can computationally generate a value of a molecular property for the ligand. Certain molecular properties, such as molecular weight, number of rotatable bonds, number of rings, and so forth, may be directly and unambiguously derivable from the chemical structure of the ligand. However, in order to obtain the values of other molecular properties such as binding affinity, toxicity, absorption, and so forth, the system can process data characterizing the ligand (and, optionally, a target molecule) using a property prediction model that is configured to generate a predicted value of the molecular property.
[0194] The system can implement a property prediction model for a molecular property, e.g., as a machine learning model, e.g., a neural network, or a random forest, or a support vector machine, and so forth, having any appropriate machine learning model architecture. The system can train the property prediction model on a set of training examples and using a machine learning training technique.
[0195] Each training example can correspond to a molecule and can include: (i) a training input characterizing the molecule, and (ii) an actual value of the molecular property. For each training example, the system can train the property prediction model to reduce a discrepancy between: (i) the actual value of the molecular property specified by the training example, and (ii) apredicted value of the molecular property generated by processing the training input specified by the training example using the property prediction model. The discrepancy between an actual value and a predicted value of a molecular property can be measured, e.g., as an absolute error, or a squared error, or in any other appropriate way. The machine learning training technique can be any appropriate technique appropriate for training the type of machine learning model used to implement the property prediction model. For instance, for a property prediction model implemented as a neural network, the design system can train the property prediction model using a stochastic gradient descent training technique.
[0196] Filtering the set of ligands to remove any ligands that do not satisfy the acceptance criteria can have the effect of reducing the number of ligands in the set of ligands, e.g., by 50%, or 90%, or 99%.
[0197] After screening the set of ligands, the system provides the set of ligands as an output (212), e.g., by storing data defining the set of ligands in a memory, or by transmitting data defining the set of ligands over a data communication network, or by providing data defining the set of ligands directly to a system that performs downstream processing based on the ligands.
[0198] Optionally, one or more of the remaining ligands, i.e., that satisfy the acceptance criteria and remain in the set of ligands after the filtering, can be selected for further computational or physical validation.
[0199] Computational validation of a ligand can include, e.g., performing computational simulations such as quantum mechanics simulations (e.g., electronic structure calculations or molecular orbitals analysis), or molecular mechanics simulations (e.g., molecular dynamics simulations or Monte Carlo simulations), or docking simulations (e.g., protein-ligand docking simulations), and so forth. Computational simulations of a ligand can generate additional data characterizing the behavior, interactions, and properties of the ligand. In some cases, performing a computational simulation of a ligand can be computationally intensive, in particular, can require significant computational resources such as memory and computing power. Filtering the set of ligands and performing computational validation of only the ligands remaining after the filtering can thereby significantly reduce consumption of computational resources, e.g. as compared to performing computational validation of the entire set of ligands generated by the system.
[0200] Physical validation of a ligand can include, e.g., physically synthesizing the ligand and (in some cases) experimentally measuring one or more characteristics of the ligand, e.g., by measuring the binding affinity of the ligand for a target molecule, or by administering a drugincluding the ligand to a subject (e.g., a cell, or a collection of cells, or an animal, or a person) to assess the absorption, or distribution, or metabolism, or excretion, or toxicity of the drug that includes the ligand. Performing physical validation of a ligand can require significant resources, e.g., laboratory resources, chemical resources, personnel resources, and so forth. Filtering the set of ligands and performing physical validation of only the ligands remaining after the filtering can thereby significantly reduce consumption of resources, e.g., as compared to performing physical validation of the entire set of ligands generated by the system.
[0201] FIG. 3A shows an example embedding neural network 300, e.g., that is included in the design system described with reference to FIG. 1. The embedding neural network 300 is configured to process: (i) target molecule data 102 characterizing a target molecule(s), and optionally, (ii) ligand design criteria 104 specifying target ligand properties and / or scaffolding data, to generate latent conditioning data 110 representing the target molecule(s) 102 and (optionally) the ligand design criteria 104.
[0202] The embedding neural network 300 includes a target molecule embedding neural network 302, a design embedding neural network 304, and a fusion neural network 310, which are each described in more detail next (and throughout this specification).
[0203] The target molecule embedding neural network 302 is configured to process the target molecule data 102 characterizing the target molecule(s) to generate a target molecule embedding 306 of the target molecule(s). The target molecule embedding 306 can include, e.g., a respective component embedding for each “component” in each target molecule. A “component” can be, e.g., an atom, or an amino acid, or a nucleic acid, etc. (dependent upon the type of the target molecule). For instance, for a target molecule that is a protein, the target molecule embedding 306 can include a respective amino acid embedding for each amino acid in the protein. As another example, for a target molecule that is a nucleic acid, the target molecule embedding 306 can include a respective nucleotide embedding for each nucleotide in the nucleic acid.
[0204] In cases where the ligand design criteria 104 include target molecule scaffolding data 108, the target molecule embedding neural network 302 can process both the target molecule data 102 and the target molecule scaffolding data.
[0205] The target molecule embedding neural network 302 can have any appropriate neural network architecture that enables the target molecule embedding neural network 302 to perform its described functions. In particular, the target molecule embedding neural network 302 can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g., 5 layers, 10 layers,or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0206] A particular example of a possible architecture of the target molecule embedding neural network 302 is the “Evoformer” neural network described in Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature, Vol 596, 26 August 2021. The Evoformer can process a network input derived from: (i) the amino acid sequence of a protein, (ii) an MSA for the protein, and (iii) the 3D structures of one or more template amino acid sequences, to generate an output that includes a “single representation” that defines a respective embedding of each position in each amino acid sequence of the protein. Optionally, the target molecule embedding neural network can process any protein scaffolding data in a similar manner as the 3D structures of the template amino acid sequences.
[0207] In another example, the target molecule embedding neural network can include an embedding layer that is configured to generate a respective initial embedding for each component in each target molecule. The embedding of each component can be a predefined embedding that is manually designed (e.g., a one-hot embedding) or learned (e.g., along with the parameters of the embedding neural network by an end-to-end machine learning training procedure). The target molecule embedding neural network can include a sequence of selfattention neural network layers (e.g., 3, or 5, or 10 self-attention layers) that operate on the set of component embeddings generated by the embedding layer to generate an output set of component embeddings that collectively define the target molecule embedding.
[0208] In some cases, when the target molecule data characterizes multiple target molecules, the target molecule embedding neural network can generate a separate molecule embedding for each target molecule. The embedding neural network 300 can then combine (e.g., concatenate) the respective molecule embeddings of each target molecule to generate the overall target molecule embedding 306.
[0209] The design embedding neural network 304 is configured to process the target ligand properties and / or ligand scaffolding data specified by the ligand design criteria 104 to generate a design embedding 308. If the input to the embedding neural network 300 does not include any target ligand properties or ligand scaffolding data, then the design system 100 can bypass the operations of the design embedding neural network 304 and initialize the design embedding 308 as a default embedding.
[0210] The design embedding 308 can include a respective atom embedding representing each atom in a set of possible atoms that are eligible for inclusion in the ligand. The number of atoms that are included in the ligand to be generated by the design system may be unknown.Therefore, the number of atom embeddings in the design embedding 308 can be defined by the ligand design criteria 104 or, if no ligand design criteria 104 are provided to the design system, can be set to a default value, e.g., 100, 500, 1000, or 2000 atom embeddings. Each atom embedding in the design embedding thus represents an atom that may (or may not) be selected for inclusion in the ligand generated by the generative model when conditioned on the latent conditioning data 110, as will be described in more detail below. The number of atom embeddings in the design embedding 308 can thus define a maximum number of atoms that can be included in a ligand generated by the design system. The atoms in the set of possible atoms that are eligible for inclusion in the ligand may be referred to for convenience as “ligand atoms”.
[0211] In some cases, the ligand design criteria 104 can identify that the design system is intended to generate more than one ligand. Optionally, the embedding neural network can partition the set of atom embeddings of the ligand atoms into a plurality of subsets, where each subset of atom embeddings is assigned to represent a respective ligand and includes data identifying that ligand. (For instance, each ligand embedding that is associated with a particular ligand can be “tagged,” e.g., concatenated, with one or more numerical values that uniquely identify that ligand).
[0212] If the input to the embedding neural network 300 does not include any target ligand properties or ligand scaffolding data, then the design system can initialize the design embedding 308 as a default embedding. In particular, the default design embedding 308 can include a respective atom embedding representing each ligand atom, where each atom embedding is defined as a default embedding, e.g., an embedding where each value in the embedding is set to a predefined value, e.g., a value of zero or one.
[0213] The design embedding neural network 304 can have any appropriate neural network architecture that enables the design embedding neural network 304 to perform its described functions. In particular, the design embedding neural network 304 can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g., 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0214] An example of processing ligand design criteria 104 using a design embedding neural network 304 to generate a design embedding 308 is described in more detail with reference to FIG. 3B.
[0215] The fusion neural network 310 is configured to process the target molecule embedding 306 and the design embedding 308 to generate the latent conditioning data 110 that jointlyrepresents the target molecule(s) 102 and any ligand design criteria 104. The fusion neural network 310 can have any appropriate neural network architecture that enables the fusion neural network 310 to perform its described functions. In particular, the fusion neural network 310 can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g., 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0216] An example of processing a target molecule embedding 306 and a design embedding 308 using a fusion neural network 310 to generate latent conditioning data 110 is described in more detail with reference to FIG. 4A.
[0217] FIG. 3B is a flow diagram of an example process 312 for processing ligand design criteria using a design embedding neural network (e.g., that is included in the embedding neural network described with reference to FIG. 3A) to generate a design embedding. For convenience, the process 312 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 312.
[0218] The system receives, by an input layer of the design embedding neural network, ligand design criteria specifying one or more target (desired) properties of the ligand and / or ligand scaffolding data (314). The ligand design criteria can specify global properties of the ligand (i.e., that characterize the ligand as a whole rather than being specific to a single atom) or atomspecific properties of the ligand (i.e., that relate to specific atoms in the ligand rather than the entire ligand) or both. The ligand scaffolding data can specify a respective 3D spatial position for each of one or more ligand atoms.
[0219] The system processes the ligand design criteria, by an embedding block of the design embedding neural network, to generate a collection of initial atom embeddings that includes a respective initial atom embedding for each ligand atom (316). If the ligand design criteria specify a maximum number of atoms that can be selected for inclusion in the ligand, then the system can generate a number of initial atom embeddings equal to the maximum number of atoms that can be selected for inclusion in the ligand. If the ligand design criteria do not specify a maximum number of atoms that can be selected for inclusion in the ligand, then the system can generate a default (e.g., predefined) number of initial atom embeddings.
[0220] The system can include data representing the global properties of the ligand in all of the initial atom embeddings. For each ligand atom, the system can include: (i) any atom-specific properties of the ligand atom, and (ii) any ligand scaffolding data specifying a 3D spatial position of the atom, in the initial atom embedding of the atom. Thus, for each ligand atom, the system can generate the initial atom embedding for the atom based on “atom property data” that includes: (i) any global properties of the ligand, (ii) any atom-specific properties of the atom, and (iii) any ligand scaffolding data for the atom.
[0221] To generate the initial atom embedding for an atom, the system can generate an array (e.g., vector) of numerical values having a predefined dimensionality, where each component of the array is assigned to represent a respective type of atom property data, e.g., global ligand property data, atom-specific property data, and ligand scaffolding data. For instance, the array can include one or more components assigned to represent a target binding affinity of the ligand for the protein, and one or more components assigned to represent an elemental type of the atom, and one or more components assigned to represent ligand scaffold data for the atom, and so forth. The system can populate the array with atom property data for the atom that includes: (i) any global properties of the ligand, (ii) any atom-specific properties of the atom, and (iii) any ligand scaffolding data for the atom. Any components of the array that are not populated using the atom property data are masked, e.g., are set to a default value (e.g., negative one).
[0222] In some implementations, the array of atom property data for an atom directly defines the initial atom embedding of the atom. In other implementations, the system processes the array of atom property data for an atom using one or more neural network layers (e.g., fully connected layers) of the embedding block to generate the initial atom embedding for the atom.
[0223] For one or more of the ligand atoms, the array of atom property data for the atom may be partially or fully masked, e.g., if the ligand design criteria do not specify any global properties of the ligand, any atom-specific properties of the atom, or any ligand scaffolding data for the atom.
[0224] The system processes the collection of initial atom embeddings, by an update block of the design embedding neural network, to generate a collection of final atom embeddings that includes a respective final atom embedding for each ligand atom (318). In some implementations, the update block includes a sequence of one or more self-attention neural network layers that are each configured to receive a collection of atom embeddings, apply one or self-attention operations (e.g., query-key-value (QKV) self-attention operations) to the collection of input atom embeddings to update each of the atom embeddings, and then output the collection of updated atom embeddings. The first self-attention layer in the sequence of self-attention layers can receive as input the collection of initial atom embeddings generated by the embedding block, and each subsequent self-attention layer can receive as input thecollection of atom embeddings output by the preceding self-attention layer. The final selfattention layer in the sequence of self-attention layers can output collection of final atom embeddings.
[0225] The system provides the collection of final atom embeddings, by an output layer of the design embedding neural network, as the design embedding representing the ligand design criteria (320).
[0226] FIG. 3C provides an illustration of a collection of initial atom embeddings generated by an embedding block of a design embedding neural network, as described at step 316 of FIG. 3. Each initial atom embedding represents a respective ligand atom. Any global ligand properties specified by the ligand design criteria are broadcast across all of the initial atom embeddings. Any atom-specific properties or ligand scaffolding data that are specific to a particular atom are separately represented in the initial atom embedding for that atom.
[0227] FIG. 4A is a flow diagram of an example process 400 for processing a target molecule embedding of a target molecule and a design embedding of ligand design criteria using a fusion neural network to generate latent conditioning data that jointly represents the target molecule(s) and the ligand design criteria. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
[0228] The system receives the target molecule embedding of the target molecule(s) and the design embedding of the ligand design criteria (402). The target molecule embedding can be generated by a target molecule embedding neural network, as described with reference to FIG. 3 A, and can include a respective component embedding (e.g., amino acid embedding) for each of a plurality of components (e.g., amino acids) included in the target molecule(s). The design embedding can be generated by a design embedding neural network, as described with reference to FIG. 3B, and can include a respective atom embedding for each ligand atom in a set of ligand atoms.
[0229] The system concatenates the target molecule embedding and the design embedding to generate a one-dimensional (ID) sequence of embeddings (404). The ID sequence of embedding includes the component embeddings of the target molecule embedding and the atom embeddings of the design embedding. The embeddings included in the ID sequence of embeddings can be ordered in any appropriate way, e.g., the ID sequence of embeddings can be ordered to have the component embeddings followed by the atom embeddings, or to have the atom embeddings followed by the component embeddings. The length of the ID sequenceof embeddings can be a sum of: (i) the number of component embeddings in the target molecule embedding, and (ii) the number of atom embeddings in the design embedding. The ID sequence of embeddings can be represented by data having dimensionality NumT okens x d, where NumT okens is given by the sum of: (i) the number of component embeddings in the target molecule embedding, and (ii) the number of atom embeddings in the design embedding, and d is a positive integer value defining the number of channel dimensions in each component embedding and atom embedding.
[0230] The system processes the ID sequence of component embeddings and atom embeddings to generate data defining a two-dimensional (2D) array of embeddings (406). The 2D array of embeddings can be represented by data having dimensionality NumT okens x NumT okens x d' where (as above) NumT okens is given by the sum of: (i) the number of component embeddings in the target molecule embedding, and (ii) the number of atom embeddings in the design embedding, and d’ is a positive integer value defining the number of channel dimensions in each embedding (d' can be equal to d, i.e., the number of channel dimensions in each component embedding and atom embedding). The system can generate the 2D array of embeddings from the ID sequence of embeddings in any of a variety of ways, e.g., using a tensor product or other bilinear map. For instance, the system can generate the 2D array of embeddings as a result of an element-wise outer product of the ID sequence of embeddings with itself. As another example, the system can generate the 2D array by an appropriate 2D concatenation operation, e.g., where the embedding at each position (i,y) in the 2D array of embeddings is generated by concatenating: (i) the embedding at position i. and (ii) the embedding at position y, in the ID sequence of embeddings (where indices i. j G {1, ... A}, where N is the length of the ID sequence of embeddings).
[0231] Each embedding in the 2D array of embeddings can be, e.g.: (i) an atom - atom embedding, or (ii) a component - component embedding, or (iii) a component - atom embedding. Each atom - atom embedding is derived from a pair of atom embeddings representing ligand atoms. Each component - component embedding is derived from a pair of component embeddings representing components in the target molecule(s). Each component - atom embedding is derived from a pair of embeddings that includes a component embedding representing a component in the target molecule(s) and an atom embedding representing a ligand atom. In some implementations where the target molecule is a protein, the component - component embeddings can comprise amino acid - amino acid embeddings.
[0232] The system processes the 2D array of embeddings by a set of neural network layers of the fusion neural network to generate an updated 2D array of embeddings defining the latent conditioning data (408). The updated 2D array of embeddings can have the same dimensionality as the original 2D array of embeddings, e.g., NumT okens x NumT okens x d' (where NumTokens and d’ are defined as above). To generate the updated 2D array of embeddings, the fusion neural network can process the 2D array of embeddings using a sequence of one or more self-attention blocks. Each self-attention block can be configured to receive the current 2D array of embeddings as an input, to update the current 2D array of embeddings by one or more self-attention operations (e.g., single-head or multi-head querykey-value (QKV) self-attention operations), and to provide the updated 2D array of embeddings to a subsequent neural network layer (e.g., to another self-attention block, or to an output layer of the fusion neural network).
[0233] The self-attention blocks of the fusion neural network can implement any appropriate self-attention operations. A few examples of self-attention operations that can be implemented by self-attention blocks of the fusion neural network are described next.
[0234] In some implementations, one or more of the self-attention blocks of the fusion neural network implement “row-wise” or “column-wise” self-attention over the current 2D array of embeddings (i.e., that is provided as an input to the self-attention block). In a row-wise selfattention operation, a self-attention layer updates each given embedding in the 2D array of embeddings using a self-attention operation over only embeddings located in the same row as the given embedding in the 2D array of embeddings. In a column-wise self-attention operation, a self-attention block updates each given embedding in the 2D array of embeddings using a self-attention operation over only embeddings located in the same column as the given embedding in the 2D array of embeddings.
[0235] In some implementations, one or more of the self-attention blocks of the fusion neural network implement triangle self-attention operations. An example implementation of triangle self-attention operations are described in Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature, Vol 596, 26 August 2021.
[0236] In some implementations, one or more of the self-attention blocks of the fusion neural network implement a full self-attention operation over the current 2D array of embeddings, e.g., by updating each embedding in the 2D array of embeddings using attention over the entire 2D array of embeddings.
[0237] The fusion neural network can include other neural network layers, i.e., in addition to the sequence of self-attention blocks, e.g., other neural network layers (such as fully connectedlayers or normalization layers) that are interleaved among the self-attention blocks. The fusion neural network can also include features such as skip connections, e.g., to implement residual blocks in the fusion neural network.
[0238] The system outputs the 2D array of embeddings generated by the fusion neural network as the latent conditioning data (410). The system can provide the latent conditioning data generated by the fusion neural network, e.g., for conditioning the generative model, as will be described in more detail below with reference to FIG. 5.
[0239] FIG. 4B illustrates operations performed by the design system to generate the latent conditioning data. The system generates a target molecule embedding that includes a sequence of component embeddings 412 (one for component in each target molecule) and a design embedding that includes a sequence of atom embeddings 414 (one for each ligand atom in a set of ligand atoms). The design system concatenates the sequence of component embeddings and the sequence of atom embeddings into a ID sequence of embeddings, and then transforms the ID sequence of embeddings (e.g., by an outer product operation) into a 2D array of embeddings. The 2D array of embeddings can include: (i) atom - atom embeddings 426, (ii) component - component embeddings 420, and (iii) component - atom embeddings 422, 424. Each atom - atom embedding is derived from a pair of atom embeddings representing ligand atoms. Each component - component embedding is derived from a pair of component embeddings representing components in the target molecule(s). Each component - atom embedding is derived from a pair of embeddings that includes a component embedding representing a component in the target molecule(s) and an atom embedding representing a ligand atom. The design system can process the 2D array of embeddings 418 using a sequence of one or more self-attention blocks (e.g., that implement row-wise attention, or column-wise attention, or triangle self-attention, or full self-attention) to generate the latent conditioning data.
[0240] FIG. 5 is a flow diagram of an example process 500 for generating a data defining a ligand that is predicted to form a complex with one or more target molecules using a generative diffusion model that includes a denoising neural network. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.
[0241] The system generates respective initial atom state data for each atom in the target molecule(s) and for each ligand atom in each of one or more ligand(s) (502).
[0242] The atom state data for an atom can include features that can be categorized as “continuous features” or “categorical features.”
[0243] A continuous feature refers to a feature that can take on values from a continuous range of possible values, e.g., in a continuous interval. The atom state data for an atom can include continuous features such as a feature defining a 3D spatial position of the atom (e.g., which can assume values from in the continuous space M3). a feature defining a partial charge of the atom (e.g., which can assume values in the continuous interval [-1,+1 ]), and so forth.
[0244] A categorical feature can refer to a feature that can assume a value in a finite set of possible values. The atom state data for an atom can include categorical features such as a feature defining an elemental type of the atom (e.g., which can assume values from a discrete set of possible elemental types, e.g., carbon, oxygen, nitrogen, etc.), a feature defining a hybridization state of the atom (e.g., which can assume values from a discrete set of possible hybridization states, e.g., sp hybridization, or sp2hybridization, or sp3hybridization, or sp3d hybridization, or sp2d2hybridization, etc.), a feature defining a type of an amino acid in which the atom is included (in particular, when the ligand is a protein), a feature defining a type of a nucleotide in which the atom is included (in particular, when the ligand is a nucleic acid molecule), and so forth. As another example, an atom-specific property can specify a stereochemistry of an atom in the ligand, e.g., an R or S designation for the atom.
[0245] Generally, for each atom, the atom state data for the atom specifies at least a 3D spatial position of the atom. Optionally, the atom state data for each atom can further specify one or more additional atom-specific properties of the atom, e.g., including one or more of: an elemental type of the atom, a hybridization state of the atom, a partial charge of the atom, a type of an amino acid in which the atom is included (in particular, when the ligand is a protein), a type of a nucleotide in which the atom is included (in particular, when the ligand is a nucleic acid molecule), and so forth.
[0246] The system can stochastically sample the initial atom state data for some or all of the atoms in the complex, i.e., for some or all of the atoms in the target molecule(s) and for some or all of the ligand atoms in the ligand(s).
[0247] In particular, for some or all of the continuous features in the atom state data for an atom, the system can sample a value of the feature from a probability distribution over the space of possible values of the feature. The system can then include the sampled value of the feature in the initial atom state data for the atom.
[0248] For instance, the system can sample a 3D spatial position of an atom from a probability distribution over 3D space, and then include the sampled 3D spatial position in the initial atomstate data for the atom. The probability distribution over 3D space can be, e.g., a standard Normal distribution over 3D space. The system can represent the sampled 3D spatial position in the initial atom state data for the atom, e.g., in Cartesian coordinates or in spherical coordinates or in any other appropriate coordinate system over 3D space.
[0249] As another example, the system can sample a value for a partial charge of an atom from a probability distribution (e.g., a uniform distribution) over a range of possible partial charges (e.g., the interval [- 1 ,+ 1]). The system can then include the sampled value of the partial charge in the initial atom state data for the atom.
[0250] Further, for some or all of the categorical features in the atom state data for an atom, the system can sample a distribution over a set of possible values of the feature. The system can then include the distribution over possible values of the feature in the initial atom state data for the atom. The distribution over the set of possible values of the feature can define a respective score for each value in the set of possible values of the feature. The system can sample the score for each possible value of the feature from a probability distribution over a range of possible scores. Optionally, the system can normalize the distribution over the set of possible values of the feature prior to including the distribution in the initial atom state data, e.g., by processing the distribution using a soft-max function.
[0251] For instance, the system can generate a distribution over a set of possible atomic element types of an atom, e.g., by sampling a respective score for each possible atomic element type from a probability distribution over a range of possible scores, e.g., a uniform distribution over the range [0,1], The system can then include the distribution over the set of possible atomic element types of the atom in the initial atom state data for the atom.
[0252] As another example, the system can generate a distribution over a set of possible atomic hybridization states of an atom, e.g., by sampling a respective score for each possible atomic hybridization state from a probability distribution over a range of possible scores, e.g., a uniform distribution over the range [0,1], The system can then include the distribution over the set of possible atomic hybridization states of the atom in the initial atom state data for the atom.
[0253] As another example, when the ligand is a protein, the system can generate a distribution over a set of possible amino acid types in which the atom may be included, e.g., by sampling a respective score for each possible amino acid type from a probability distribution over a range of possible scores, e.g., a uniform distribution over the range [0,1], The system can then include the distribution over the set of possible amino acid types of the atom in the initial atom state data for the atom.
[0254] As another example, when the ligand is a nucleic acid, the system can generate a distribution over a set of possible nucleotide types in which the atom may be included, e.g., by sampling a respective score for each possible nucleotide type from a probability distribution over a range of possible scores, e.g., a uniform distribution over the range [0,1], The system can then include the distribution over the set of possible nucleotide types of the atom in the initial atom state data for the atom.
[0255] In some cases, the system receives scaffolding data (e.g., target molecule scaffolding data, ligand scaffolding data, or both) that specifies a respective target 3D spatial position of each of one or more atoms in the complex comprising the target molecule(s) and the ligand(s). If the system receives scaffolding data specifying a target 3D spatial position of an atom in the complex, then the system can generate initial atom state data for that atom that includes the target 3D spatial position of the atom rather than stochastically sampled 3D spatial position data for the atom (as described above).
[0256] In some cases, the system receives ligand design criteria that specify one or more atomspecific target properties of ligand atoms in the ligand(s) (e.g., target hybridization state data, or target partial charge data). If the system receives ligand design criteria specifying an atomspecific target property of a ligand atom, then the system can generate initial atom state data for that atom that includes the atom-specific target property of the ligand atom rather than stochastically sampled atom property data (as described above).
[0257] In particular, if the system receives ligand design criteria specifying a particular value of a continuous feature of an atom (e.g., a 3D spatial position or a partial charge of the atom), then the system can include that feature value in the initial atom state data for the atom (e.g., instead of a stochastically sampled value for that feature).
[0258] If the system receives ligand design criteria specifying a particular value of a categorical feature of an atom (e.g., an element type of the atom or a hybridization state of the atom), then the system can generate a distribution over a set of possible values of the feature that assigns a first value (e.g., one) to the actual value of the feature and a second value (e.g., zero) to each other value of the feature. For instance, the system can generate a one-hot distribution over the set of possible values of the feature that uniquely identifies the actual value of the feature. The system can then include the generated distribution over possible values of the feature in the initial atom state data for the atom (e.g., instead of a stochastically sampled distribution over possible values of the feature).
[0259] Optionally, the system can generate initial bond data for each of multiple pairs of atoms in the ligand (in some cases, for every pair of atoms in the ligand). The bond data for a pair ofligand atoms can be a categorical feature that assumes values from a set bond types that includes a “no bond” type (i.e. , indicating that the pair of atoms are not bonded), and one or more other bond types, e.g., single bond, double bond, triple bond, etc. The system can generate the initial bond data by, for each of multiple pairs of atoms in the ligand, sampling a distribution over the set of bond types, as described above for other categorical feature types. The system can then include the respective distribution over possible values of the set of bond types in the initial bond data for each pair of atoms.
[0260] For each atom in the complex, the initial atom state data for the atom thus defines a “noisy” representation of the state of the atom. The system performs steps 504 - 512, which are described next, over a sequence of iterations that may be referred to as “denoising time steps,” to progressively denoise the initial atom state data for the atoms as part of designing the ligand(s) that form a complex with the target molecule(s). More specifically, the system progressively denoises the initial atom state data for each atom in the complex to cause the noisy representation of the atom state to converge on a denoised representation that accurately represents the atom state.
[0261] In implementations where the system generates initial bond data (as described above) that defines a “noisy” representation of the bonding between ligand atoms, by iteratively performing steps 504-512, the system can progressively denoise the initial bond data to converge on a denoised representation that accurately represents bonds between pairs of atoms in the ligand.
[0262] The description of steps 504 - 510 which follows will reference a “current” denoising time step for convenience; the current denoising time step can be any denoising time step in the sequence of denoising time steps. The system can perform the steps 504 - 510 over any appropriate number of denoising time steps, e.g., 3 denoising time steps, 10 denoising time steps, or 100 denoising time steps. The number of denoising time steps can be a predetermined number of time steps.
[0263] The system generates a denoising output by processing respective current atom state data for the atoms in the complex (and, optionally, current bond data for the ligand) using a denoising neural network that is conditioned on the latent conditioning data (504). If the current denoising time step is the first denoising time step, then the current atom state data for the atoms in the complex may be the initial atom state data, e.g., as generated at step 502. Further, the current bond data for the ligand may be the initial bond data, e.g., as generated at step 502. If the current denoising time step is after the first denoising time step, then the current atom state data for the atoms in the complex may be generated at the preceding denoising time step,e.g., as in step 512 of the process 500, which will be described in more detail below. Further, the current bond data for the ligand may be generated at the preceding denoising time step, e.g., as in step 512 of the process 500, which will be described in more detail below.
[0264] The denoising output can be any appropriate data that enables estimation of the denoised (ground truth) atom state data of each atom in the complex (and, optionally, estimation of denoised bond data for pairs of atoms in the ligand). For instance, the denoising output can define, for each atom in the complex, a predicted error in the atom state data of the atom at the current denoising time step. Further, the denoising output can define, for each pair of atoms in the ligand, a predicted error in the bond data for the pair of atoms at the current denoising time step. As another example, the denoising output can directly define, for each atom in the complex, predicted atom state data for the atom. Further, the denoising output can directly define, for each pair of atoms in the ligand, predicted bond data for the pair of atoms. As another example, the denoising output can define, for each atom in the complex, both: (i) a predicted error in the atom state data for the atom at the current denoising time step, and (ii) predicted atom state data for the atom. Further, the denoising output can define, for each pair of atoms in the ligand, both: (i) a predicted error in the bond data for the pair of atoms at the current denoising time step, and (ii) predicted bond data for the pair of atoms. As another example, the denoising output can define, for each atom in the complex, a prediction for an array of values that is a linear combination of: (i) actual atom state data for the atom, and (ii) an error between the atom state data for the atom at the current denoising time step and the actual atom state data for the atom, e.g., as implemented by the v-parametrization described in: Tim Salimans, Jonathan Ho, “Progressive distillation for fast sampling of diffusion models,” ICLR 2022, arXiv: 2202.00512v2. As another example, the denoising output can define, for each pair of atoms in the ligand, a prediction for an array of values that is a linear combination of: (i) actual bond data for the pair of atoms, and (ii) an error between the bond data for the pair of atoms at the current denoising time step and the actual bond data for the pair of atoms. An example process for generating a denoising output using the denoising neural network is described in more detail with reference to FIG. 6.
[0265] Optionally, in addition to generating the denoising output, the system can generate gradients, with respect to the current atom state data for some or all of the atoms in the complex (and, optionally, the current bond data for the ligand), of each of one or more conditioning objective functions (506). Each conditioning objective function is associated with a respective ligand property of the ligand and measures a discrepancy between: (i) a predicted value of the ligand property for the ligand as defined by the current atom state data (and, optionally, thecurrent bond data), and (ii) a target (desired) value of the ligand property. The system can generate the predicted value of the ligand property, e.g., by processing the current atom state data for some or all of the atoms in the complex (and, optionally, the current bond data) using a respective property prediction neural network, as will be described in more detail with reference to FIG. 8. The system can receive the target (desired) value of the ligand property, e.g., as an input from a user of the system or from an upstream system, e.g., by way of a user interface or an API.
[0266] The gradients of a conditioning objective function associated with a ligand property can be used to update the current atom state data (and, optionally, the current bond data) to reduce a discrepancy between the predicted value of the ligand property for the ligand defined by the current atom state data and the target value of the ligand property. Thus, the system can use the gradients (in addition to the denoising output of the denoising neural network) to steer (influence) the process of denoising the current atom state data (and, optionally, the bond data) to increase the likelihood that the resulting ligand will assume the target value for the ligand property. An example of using gradients of a conditioning objective function to update the current atom state data (and, optionally, the current bond data) is described in more detail below with reference to step 508 of the process 500.
[0267] An example process for determining gradients of a conditioning obj ective function with respect to the current atom state data (and, optionally, the current bond data) for some or all of the atoms in the complex is described in detail with reference to FIG. 8.
[0268] The system generates a current estimate of the denoised atom state data for each atom in the complex (and, optionally, of the denoised bond data for the ligand) using the denoising output generated by the denoising neural network, and optionally, respective gradients of each conditioning objective function (508).
[0269] In implementations where the system generates gradients of a conditioning objective function, the system can combine the gradients of the conditioning objective function with the denoising output of the denoising neural network prior to using the denoising output to generate the current estimate of the denoised atom state data (and, optionally, of the denoised bond data). The system can combine the gradients of the conditioning objective function with the denoising output of the denoising neural network, e.g., by scaling the gradients of the conditioning objective function by a scaling constant (that can depend on the current time step) and then adding the gradients of the conditioning objective function to the denoising output of the denoising neural network.
[0270] More specifically, the denoising output of the denoising neural network can include a respective value, referred to for convenience as a denoising value, for each component of the atom state data for each atom in the complex (and, optionally, for each component of the bond data for each pair of atoms in the ligand). The gradients of the conditioning objective function include a respective gradient value for some or all of the components of the atom state data for some or all of the atoms in the complex (and, optionally, for some or all of the components of the bond data). The system can scale the gradients of the conditioning objective function by the scaling constant (as described above), and then combine (e.g., by addition) each gradient value with the corresponding denoising value that is associated with the same component of the atom state data (or bond data) as the gradient value.
[0271] After (optionally) combining the gradients of each conditioning objective function with the denoising output, the system can generate the current estimate of the denoised atom state data for each atom in the complex (and, optionally, the current estimate of the denoised bond data for the ligand) in any appropriate way, depending on the form of the denoising output. A few example techniques for generating the current estimate of the denoised atom state data for each atom in the complex (and, optionally, the current estimate of the denoised bond data) using the denoising output are described next.
[0272] In one example, the denoising output defines, for each atom, a respective prediction for the atom state data of the atom. In this example, the respective predicted atom state data for each atom defines the current estimate of the denoised atom state data for the atom. Further, the denoising output can define, for each pair of atoms in the ligand, a respective prediction for the bond data of the pair of atoms. The respective predicted bond data for each pair of atoms defines the current estimate of the denoised bond data for the pair of atoms.
[0273] In another example, the denoising output defines, for each atom, a predicted error in the atom state data for the atom at the current time step. In this example, the system can generate the current estimate for the denoised atom state data for each atom as a linear combination of: (i) the current atom state data for the atom, and (ii) the predicted error in the atom state data for the atom. Each term in the linear combination can be scaled by a respective constant value that is dependent on the time step. For instance, the system can generate the current estimate for the denoised atom state data xt-rfor an atom in the complex as:where t indexes the current time step, at, at, and crtare constants specific to time step t, and egxt, t) is the predicted error in the atom state data of the atom (e.g., as generated by thedenoising neural network at the time step). (In the notation of equation (1), the time steps decrement, such that time step t — 1 is the “next” time step after time step t). The constants in equation (1) (atand at) can be selected in accordance with a predefined noise schedule. Further, the denoising output can define, for each pair of ligand atoms, a predicted error in the bond data for the pair of ligand atoms at the current time step. The system can generate the current estimate for the denoised bond data for the pair of ligand atoms as a linear combination of: (i) the current bond data for the pair of ligand atoms, and (ii) the predicted error in the bond data for the pair of ligand atoms.
[0274] In another example, the denoising output defines, for each atom, both: (i) predicted atom state data for the atom, and (ii) a predicted error in the atom state data for the atom at the current time step. In this example, the system can generate the current estimate for the denoised atom state data for the atom as a combination (e.g., an average) of: (i) the predicted atom state data for the atom as specified by the denoising output, and (ii) predicted atom state data for the atom that is derived from the predicted error in the atom state data for the atom at the current time step, e.g., using equation (1). Further, the denoising output can define, for each pair of ligand atoms, both: (i) predicted bond data for the pair of ligand atoms, and (ii) a predicted error in the bond data for the pair of ligand atoms at the current time step. In this example, the system can generate the current estimate for the denoised bond data for the pair of ligand atoms as a combination (e.g., an average) of: (i) the predicted bond data for the pair of ligand atoms as specified by the denoising output, and (ii) predicted bond data for the pair of ligand atoms that is derived from the predicted error in the bond data for the pair of ligand atoms at the current time step, e.g., using equation (1).
[0275] In another example, the denoising output is expressed using a v-parametrization, and the system generates a respective current estimate for the denoised atom state data for each atom (and, optionally, for the denoised bond data for the ligand) using the techniques described in Tim Salimans, Jonathan Ho, “Progressive distillation for fast sampling of diffusion models,” ICLR 2022, arXiv:2202.00512v2.
[0276] Optionally, the system can generate a respective confidence measure for the current estimate of the respective denoised atom state data for each atom in the complex. For instance, as part of generating the denoising output, the denoising neural network can generate a respective atom embedding for each atom in the complex, e.g., as the output of the update block of the denoising neural network, as described with reference to step 606 of FIG. 6. The system can process each atom embedding using one or more neural network layers (e.g., a combination of one or more of: fully connected layers, or attention layers, or pooling layers) to generate arespective confidence estimate for the current estimate of the denoised atom state data of the atom(s) represented by the atom embedding. The confidence measure for the current estimate of the denoised atom state data for an atom can characterize a predicted error in the current estimate of the atom state data for the atom.
[0277] Optionally, the system can generate a respective confidence measure for the current estimates of the 3D spatial positions of pairs of atoms in the complex. The current estimate for the 3D spatial position of an atom in the complex refers to the 3D spatial position of the atom that is defined by the current estimate of the denoised atom state data for the atom. For instance, for a first atom in the complex and a second atom in the complex, the system can process an atom embedding representing the first atom and an atom embedding representing the second atom (e.g., as generated by the update block of the denoising neural network) using one or more neural network layers (e.g., a combination of one or more of: fully connected layers, or attention layers, or pooling layers) to generate a confidence estimate for the current estimates of the 3D spatial positions of the first atom and the second atom. The confidence measure can characterize, e.g., a predicted error in the relative 3D displacement of the first atom and the second atom.
[0278] Optionally, the system can generate a confidence measure for the structure of the target molecule(s) as defined by the current estimates of the 3D spatial positions of the atoms in the target molecule(s), e.g., by combining (e.g., summing or averaging) the confidence measures for the individual atoms included in the target molecule(s).
[0279] Optionally, the system can generate a confidence measure for the ligand(s) as defined by the current estimates of the denoised atom state data for the atoms in the ligand(s), e.g., by combining (e.g., summing or averaging) the confidence measures for the individual atoms included in the ligand(s).
[0280] Optionally, the system can generate a confidence measure for the structure of an interface between a target molecule and a ligand, e.g., by combining (e.g., summing or averaging) the confidence measures of pairs of atoms included in the interface. A pair of atoms can be referred to as being included in the interface, e.g., if the pair includes: (i) a first atom included in the ligand, and (ii) a second atom included in the target molecule, where the relative displacement between the 3D spatial positions of the atoms is less than a threshold, e.g., 2 Angstroms, or 3 Angstroms, or 8 Angstroms. An example of generating a confidence measure for a pair of atoms is described above.
[0281] If the current denoising time step is not the final denoising time step, the system generates respective atom state data for each atom in the complex (and, optionally, respectivebond data for each pair of ligand atoms) for the next denoising time step based on the current estimates of the denoised atom state data for the atoms (and optionally based on the current estimates of the denoised bond data) (as generated at step 508) using an appropriate diffusion sampling technique (512). A few examples of possible diffusion sampling techniques are described next.
[0282] In one example, the system can generate the atom state data for each atom in the complex at the next denoising time step by combining random noise with the current estimate of the denoised atom state data for the atom. For instance, for each atom, the system can add respective random noise to the current estimate of the denoised atom state data for the atom. The random noise can be sampled from an appropriate probability distribution. The probability distribution can vary based on the time step, e.g., such that the variance of the noise combined with the updated atom state data for the atoms decreases over the sequence of time steps.
[0283] Similarly, the system can generate the bond data for each pair of atoms in the ligand at the next denoising time step by combining random noise with the current estimate of the denoised bond data for the ligand. For instance, for each pair of ligand atoms, the system can add respective random noise to the current estimate of the denoised bond data for the pair of ligand atoms. The random noise can be sampled from an appropriate probability distribution. The probability distribution can vary based on the time step, e.g., such that the variance of the noise combined with the updated bond data for the pairs of ligand atoms decreases over the sequence of time steps
[0284] As another example, the system can generate the atom state data for each atom in the complex at the next denoising time step using a deterministic diffusion sampling technique, i.e., that does not rely on random noise. An example of a deterministic diffusion sampling technique is the denoising diffusion implicit model (DDIM), e.g., as described in: Jiaming Song, Chenlin Meng, Stefano Ermon, “Denoising diffusion implicit models,” ICLR 2021, arXiv:2010.02502v4. Similarly, the system can generate bond data for each pair of ligand atoms in the ligand at the next denoising time step using the deterministic diffusion sampling technique.
[0285] Optionally, the system can refrain from applying the diffusion sampling technique to any component of the atom state data for an atom (or bond data) that was initialized at step 502 to represent target values specified by the ligand design criteria. More specifically, for any component of the atom state data for an atom (or bond data) that was initialized at step 502 to represent a target value specified by the ligand design criteria, the system can set the value of the component in the atom state data for the atom (or bond data) for the next denoising timestep to be the same as the value of the component in the current estimate of the denoised atom state data (or denoised bond data).
[0286] If the current denoising time step is the final denoising time step in the sequence of denoising time steps, the system generates data defining the ligand(s) based on the denoised atom state data for the ligand atoms (and, optionally, based on the bond data for the ligand) (514). The denoised atom state data for an atom refers to the current estimate of the denoised atom state data generated for the atom at the final denoising time step in the sequence of denoising time steps. Similarly, the denoised bond data for a pair of ligand atoms refers to the current estimate of the denoised bond data for the pair of ligand atoms at the final denoising time step in the sequence of denoising time steps. An example process for generating data defining the ligand(s) based on the denoised atom state data (and, optionally, the denoised bond data) is described with reference to FIG. 9.
[0287] Optionally, the system can additionally output data specifying a 3D structure of the target molecule(s), i.e., as defined by the 3D spatial position data included in the denoised atom state data for each atom in the protein.
[0288] Optionally, the system can additionally output data specifying a 3D structure of the entire target molecule-ligand complex, i.e., as defined by the 3D spatial position data included in the denoised atom state data for each atom in the target molecule(s) and for each atom in the ligand(s).
[0289] Optionally, the system can additionally output data identifying a binding pocket where a ligand is bound to a target molecule. In particular, the system can identify any component in a target molecule which includes an atom that is within a threshold distance (e.g., 2 Angstroms) of at least one atom in a ligand as being included in the binding pocket. The data identifying the binding pocket can include data identifying each of the components in the target molecule that are included in the binding pocket.
[0290] Optionally, the system can perform the process 500 multiple times to generate multiple ligands that can each form a complex with the target molecule(s). Each execution of the steps of the process 500 can result in the generation of a different ligand(s), e.g., as a result of stochasticity in the random sampling performed to generate the initial atom state data for the atoms (at step 502), and in some cases, as a result of stochasticity in the diffusion sampler (at step 512).
[0291] FIG. 6 is a flow diagram of an example process 600 for generating a denoising output using a denoising neural network conditioned on latent conditioning data. For convenience, the process 600 will be described as being performed by a system of one or more computers locatedin one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 600.
[0292] The system receives: (i) data defining respective current atom state data for each atom in the complex, and (ii) the latent conditioning data (602). Optionally, the system also receives data defining current bond data for each pair of ligand atoms in the ligand. The latent conditioning data can include: (i) atom - atom embeddings, (ii) component - component embeddings, and (iii) component - atom embeddings, as described in more detail above with reference to FIG. 4A. The system can optionally receive additional inputs, e.g., an input defining a current denoising time step in a diffusion process being implemented by a generative diffusion model, as described above with reference to FIG. 5.
[0293] The system generates a respective atom embedding for each atom in the complex, using an encoder block of the generative neural network, based at least in part on the current atom state data for the atom (604). Optionally, the system can generate the atom embeddings for the atoms based at least in part on the latent conditioning data (and, optionally, the current time step in the diffusion process), e.g., in addition to the current atom state data for the atoms. For instance, for each atom, the system can generate the atom embedding of the atom based on both: (i) the current atom state data for the atom, and (ii) a respective conditioning embedding selected from the collection of embeddings included in the latent conditioning data. For an atom included in a component of a target molecule, the conditioning embedding for the atom can be the component - component embedding (i.e., from the latent conditioning data) of the component that includes the atom (i.e., the component - component embedding corresponding to a pair of components that includes two copies of the component). For a ligand atom, the conditioning embedding for the atom can be the atom - atom embedding (i.e., from the latent conditioning data) of the atom (i.e., the atom - atom embedding corresponding to a pair of atoms that includes two copies of the atom).
[0294] The system can generate the atom embedding for an atom based on the current atom state data for the atom (and, optionally, a conditioning embedding for the atom) in any of a variety of possible ways. For instance, the system can generate the atom embedding for an atom by processing the atom state data for the atom using an encoder block of the denoising neural network. As another example, the system can generate the atom embedding for an atom by processing both: (i) the atom state data for the atom, and (ii) the conditioning embedding for the atom, using an encoder subnetwork of the denoising neural network. As another example, the system can generate the atom embedding for an atom by concatenating: (i) anoutput generated by an encoder subnetwork of the denoising neural network by processing the atom state data for the atom, and (ii) the conditioning embedding for the atom.
[0295] Optionally, for each atom, the system can further generate the respective atom embedding of the atom based on data, e.g., a set of binary flags, indicating which (if any) components of the initial atom state data are initially set to target values specified by the ligand design criteria, e.g., as opposed to being stochastically sampled, as described above with reference to step 502 of FIG. 5.
[0296] Optionally, the system can generate a respective bond embedding for each pair of ligand atoms in the ligand, using the encoder block of the generative neural network, based at least in part on the current bond data for the pair of ligand atoms. For instance, the system can generate the bond embedding for a pair of ligand atoms by processing the bond data for the pair of ligand atoms using the encoder block of the denoising neural network.
[0297] The encoder block of the denoising neural network can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, and so forth) in any appropriate number (e.g., 1 layer, 3, layers or 5 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers). In a particular example, the encoder block includes a sequence of fully connected neural network layers and is configured to, for each atom, process data defining the atom state data for the atom and the conditioning embedding for the atom to generate the atom embedding of the atom.
[0298] In some implementations, for each component (e.g., amino acid) in each target molecule, the system can generate one atom embedding that jointly represents all the atoms in the component. Thus, in these implementations, the number of atom embeddings may be equal to a sum of: (i) the number of ligand atoms, and (ii) the number of components (e.g., amino acids) in each target molecule (e.g., protein). The system can generate an atom embedding that jointly represents all the atoms in a component in any appropriate way. For instance, the system can generate a respective atom embedding for each atom in the component (as described above), and then combine the atom embeddings for the atoms in the component, e.g., using a pooling operation (e.g., a max pooling or a summation pooling operation), or by processing the atom embeddings for the atoms in the component using one or more neural network layers (e.g., fully connected layers or self-attention layers) to generate the embedding that jointly represents all the atoms in the component. Generating atom embeddings that jointly represent all the atoms in a component can significantly reduce the overall number of atom embeddings and thus reduce consumption of computational resources, e.g., memory and computing power,resulting from operations performed by an update block of the denoising neural network, as will be described next.
[0299] The system processes the atom embeddings for the atoms in the complex (and, optionally, the bond embeddings for pairs of ligand atoms), using an update block of the denoising neural network, to generate a respective updated atom embedding for each atom in the complex (and, optionally, an updated bond embedding for each pair of ligand atoms) (606). (For convenience, the description of FIG. 6 which follows will refer primarily to atom embeddings, but the described operations can be naturally extended to bond embeddings). The update block of the denoising neural network can include a sequence of self-attention blocks. Each self-attention block can be configured to receive a respective current atom embedding for each atom in the complex, to apply a self-attention operation to the current atom embeddings of the atoms in the complex, and to provide the updated atom embeddings, e.g., for processing by a subsequent neural network layer.
[0300] Each self-attention block included in the update block can apply any appropriate selfattention operation to the current atom embeddings of the atoms in the complex, e.g., a singlehead or multi-head query-key-value (QKV) self-attention operation. Optionally, the system can condition the self-attention operations of one or more of the self-attention blocks on the latent conditioning data. An example process for implementing a self-attention operation conditioned on the latent conditioning data is described in more detail with reference to FIG. 7.
[0301] The update block of the denoising neural network can include any appropriate number of self-attention blocks (e.g., 1 self-attention block, or 10 self-attention blocks, or 50 selfattention blocks) and can optionally include additional neural network layers of any appropriate type (e.g., fully connected layers, convolutional layers, and so forth) in any appropriate number (e.g., 1 layer, 3, layers or 5 layers) and connected in any appropriate configuration (e.g., interleaved with the self-attention blocks).
[0302] The system processes the updated atom embeddings (i.e., as generated by the update block of the denoising neural network) using a decoder block of the denoising neural network to generate the denoising output (608). Examples of denoising outputs are described above with reference to FIG. 5. The decoder block of the denoising neural network can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, and so forth) in any appropriate number (e.g., 1 layer, 3, layers or 5 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0303] In a particular example, the decoder block can include a sequence of fully connected neural network layers that are configured to operate separately on each updated atomembedding to generate a predicted error in the atom state data of the corresponding atom at the current time step, or to generate predicted atom state data of the corresponding atom, or both.
[0304] In implementations where one atom embedding jointly represents all the atoms in component (as described above with reference to step 604), the decoder block can process that (updated) atom embedding to generate respective denoising outputs for all the atoms in the component. For instance, the decoder block can process an updated atom embedding that jointly represents all the atoms in an amino acid to generate a respective predicted error in the atom state data of each atom in the component, or to generate respective atom state data of each atom in the component, or both.
[0305] FIG. 7 is a flow diagram of an example process 700 for updating a set of current atom embeddings (and, optionally, a set of current bond embeddings) using a self-attention operation that is implemented by a self-attention block of the denoising neural network and that is conditioned on the latent conditioning data. For convenience, the process 700 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 700. For convenience, the description of FIG. 7 which follows will refer primarily to the set of current atom embeddings, but the described operations can be naturally extended to processing both: (i) the set of current atom embeddings, and (ii) the set of current bond embeddings.
[0306] The system receives: (i) a set of current atom embeddings, and (ii) latent conditioning data (702). The set of current atom embeddings includes a respective atom embedding for each atom in the complex. (In some cases, for each component in a target molecule, all the atoms in the component are jointly represented by one atom embeddings, as described above with reference to FIG. 6). The set of current atom embeddings can be generated, e.g., by the encoder block of the denoising neural network or by a previous self-attention layer in the update block of the denoising neural network, as described above with reference to FIG. 6. The latent conditioning data can be generated by an embedding neural network, e.g., as described above with reference to FIG. 3A. The latent conditioning data can include: (i) atom - atom embeddings, (ii) component - component embeddings, and (iii) component - atom embeddings.
[0307] The system generates a set of intermediate attention scores based on the set of current atom embeddings (704). The set of intermediate attention scores includes a respective attention score for each pair of current atom embeddings from the set of current atom embeddings.
[0308] The system can generate the intermediate attention scores in any of a variety of possible ways. For instance, the system can generate a respective query embedding for each current atom embedding by processing the current atom embedding using a query neural network, e.g., as:Q = WQ- E (2) where Q is a matrix where each column (or row) defines a respective query embedding, WQis a matrix of parameter values (defining the query neural network in this example), and E is a matrix where each column (or row) defines a respective current atom embedding. Further, the system can generate a respective key embedding for each current atom embedding by processing the current atom embedding using a key neural network, e.g., as:K = WK• E (3) where K is a matrix where each column (or row) defines a respective key embedding, WKis a matrix of parameter values (defining the key neural network in this example), and E is a matrix where each column (or row) defines a respective current atom embedding. The system can generate the intermediate attention scores based on the query embeddings and the key embeddings, e.g., as:A = Q - KV(4) where A is a matrix of intermediate attention scores, Q is the matrix of query embeddings, and K is the matrix of key embeddings.
[0309] The system generates a set of attention score biases based on the latent conditioning data (706). The set of attention score biases includes a respective attention score bias for each pair of current atom embeddings from the set of current atom embeddings.
[0310] The system can generate the set of attention score biases in any of a variety of possible ways. For instance, for each pair of current atom embeddings, the system can generate the attention score bias for the pair of current atoms embeddings by processing a corresponding conditioning embedding selected from the collection of embeddings included in the latent conditioning data using a projection neural network.
[0311] For a pair of current atom embeddings that includes: (i) a first current atom embedding representing a first ligand atom, and (ii) a second current atom embedding representing a second ligand atom, the conditioning embedding can be the atom - atom embedding for the first ligand atom and the second ligand atom in the latent conditioning data.
[0312] For a pair of current atom embeddings that includes: (i) a first current atom embedding representing an atom included in a component of a target molecule, and (ii) a second currentatom embedding representing a ligand atom, the conditioning embedding can be the component - atom embedding corresponding to: (i) the component in the target molecule that includes the atom, and (ii) the ligand atom, in the latent conditioning data.
[0313] For a pair of current atom embeddings that includes: (i) a first current atom embedding representing an atom included in a first component in a first target molecule, and (ii) a second current atom embedding representing an atom included in a second component in a second target molecule, the conditioning embedding can be the component - component embedding corresponding to: (i) the first component, and (ii) the second component, in the latent conditioning data.
[0314] For a pair of current atom embeddings that includes: (i) a first current atom embedding that jointly represents the atoms in a first component in a target molecule, and (ii) a second current atom embedding that jointly represents the atoms in a second component in the target molecule, the conditioning embedding can be the component - component embedding corresponding to: (i) the first component, and (ii) the second component, in the latent conditioning data.
[0315] For a pair of current atom embeddings that includes: (i) a first current atom embedding that jointly represents the atoms in a component in a target molecule, and (ii) a second current atom embedding that represents a ligand atom, the conditioning embedding can be the component - atom embedding corresponding to: (i) the component in the target molecule, and (ii) the ligand atom, in the latent conditioning data.
[0316] The projection neural network can have any appropriate neural network architecture that enables the projection neural network to perform its described functions, e.g., processing a conditioning vector to generate an attention score bias. In particular, the projection neural network can include any appropriate number of neural network layers (e.g., 1 layer, or 5 layers, or 10 layers) in any appropriate number (e.g., 1 layer, 3, layers or 5 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0317] The system generates a set of final attention scores by combining: (i) the intermediate attention scores, and (ii) the attention score biases (708). The set of final attention scores includes a respective final attention score for each pair of current atom embeddings in the set of current atom embeddings. The system can generate the final attention score for a pair of current atom embeddings by combining (e.g., summing): (i) the intermediate attention score for the pair of current atom embeddings, and (ii) the attention score bias for the pair of current atom embeddings. Optionally, the system can apply further processing operations to the set offinal atention scores, e.g., by applying a soft-max operation to some or all of the final atention scores.
[0318] The system generates a set of updated atom embeddings using: (i) the set of current atom embeddings, and (ii) the set of final atention scores (710). For instance, to generate the set of updated atom embeddings, the system can generate a respective value embedding for each current atom embedding in the set of current atom embeddings by processing the current atom embedding using a value neural network, e.g., as:V = Wv- E (5) where V is a matrix where each column (or row) defines a respective value embedding, Wvis a matrix of parameter values (defining the value neural network in this example), and E is a matrix where each column (or row) defines a respective current atom embedding. The system can then generate the set of updated atom embeddings, e.g., as:E' = V • A (6) where each column (or row) of E' defines a respective updated atom embedding, each column (or row) of V defines a respective value embedding, and A denotes the set of final atention scores arranged into a matrix.
[0319] In implementations where the self-atention block implements a multi -head atention operation, each head of the atention operation can individually perform the steps of the process 700, and the updated atom embeddings generated by each atention head can be combined (e.g., concatenated) to define the overall output of the multi-head atention operation. Each atention head can have a respective set of neural network parameters, having values that are specific to each atention head, that are used for generating the intermediate atention scores and the attention score biases.
[0320] FIG. 8 is a flow diagram of an example process 800 generating gradients of a conditioning objective function with respect to current atom state data for the atoms in a target molecule-ligand complex at a current denoising time step in a sequence of denoising time steps. For convenience, the process 800 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 800. Optionally, the process 800 can also include generating gradients of the conditioning objective function with respect to current bond data for pairs of atoms in the ligand at a current denoising time step in a sequence of denoising time steps. For convenience, the description of FIG. 8 which follows will mainly refer to determining gradients with respectto current atom state data, but the same operations can be extended to determining gradients with respect to current bond data.
[0321] The system receives current atom state data for some or all of the atoms in the complex at the current denoising time step in the sequence of denoising time steps over which the generative diffusion model denoises the current atom state data (802). The system can receive current atom state data for some or all of the atoms in the ligand(s), and optionally, for some or all of the atoms in the target molecule(s). For each atom for which the system receives current atom state data, the system can receive the full set of current atom state data for the atom, or a subset of the full set of current atom state data for the atom.
[0322] The system processes the current atom state data using a property prediction neural network to generate a predicted value of a ligand property of the ligand characterized by the current atom state data (804). The ligand property of the ligand can characterize, e.g., one or more of: a binding affinity of the ligand for a target molecule, absorption properties of the ligand, distribution properties of the ligand, metabolism properties of the ligand, excretion properties of the ligand, toxicity properties of the ligand, and so forth.
[0323] In some implementations, the property prediction neural network is configured to process atom state data for only the atoms in the ligand. In other implementations, the property prediction neural network is configured to process atom state data for the atoms in the ligand and for some or all of the atoms in the target molecule(s). For instance, predicting the binding affinity of the ligand for a target molecule may require the property prediction neural network to process current atom state data for some or all of the atoms in the target molecule as well as for the atoms in the ligand.
[0324] In some implementations, the property prediction neural network is configured to process a network input that includes both: (i) the current atom state data, and (ii) data identifying the current denoising time step in the sequence of denoising time steps over which the generative diffusion model denoises the atom state data for the atoms in the complex.
[0325] The property prediction neural network can have any appropriate neural network architecture that enables the property prediction neural network to perform its described functions, e.g., including processing atom state data for some or all of the atoms in the complex to generate a predicted value of a ligand property of the ligand. In particular, the property prediction neural network can include any appropriate types of neural network layers (e.g., fully connected layers, attention layers, pooling layers, convolutional layers, and so forth) in any appropriate number (e.g., 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0326] The system can train the property prediction neural network on a set of training examples using a machine learning training technique. Each training example in the set of training examples corresponds to a respective ligand and includes: (i) a training network input to the property prediction neural network, and (ii) an actual value of the property of the ligand.
[0327] For each training example, the training network input to the property prediction neural network can include: (i) respective atom state data for some or all of the atoms in the corresponding ligand, and optionally, for some or all of the atoms in a target molecule to which the ligand can bind, and optionally, (ii) data identifying a denoising time step in a sequence of denoising time steps. The system can generate the atom state data in the training network input by obtaining target (ground truth) atom state data for the atoms in the ligand (and, optionally, for the atoms in the target molecule), and then combining random noise with the target atom state data. The system can scale the random noise combined with the target atom state data by a constant that depends on the denoising time step, e.g., where the constant is defined by the same noise schedule implemented by the diffusion sampler of the generative diffusion model, as described with reference to step 512 of FIG. 5.
[0328] The system can train the property prediction neural network to optimize a loss function (objective function) that, for each training example, measures a discrepancy between: (i) the actual value of the ligand property specified by the training example, and (ii) a predicted value of the ligand property that is generated by the property prediction neural network by processing the training network input of the training example. The loss function can measure the discrepancy in any appropriate way, e.g., using an absolute error metric or a squared error metric.
[0329] The system can train the property prediction neural network on the set of training examples using any appropriate machine learning training technique, e.g., using a stochastic gradient descent training technique.
[0330] The system determines gradients of the conditioning objective function with respect to some or all of the current atom state data that was processed by the property prediction neural network to generate the predicted value of the ligand property (as described at step 804) (806). The conditioning objective function measures a discrepancy between: (i) the predicted value of the ligand property, and (ii) a target (desired) value of the ligand property. The conditioning objective function can measure the discrepancy using any appropriate error metric, e.g., using an absolute error metric or a squared error metric. The system can determine the gradients of the conditioning objective function using any appropriate technique, e.g., backpropagation.
[0331] The system provides the gradients of the conditioning objective function (806), e.g., for generating a current estimate of the denoised atom state data for the atoms in the complex at the current denoising time step, as described above with reference to step 508 of FIG. 5.
[0332] FIG. 9 is a flow diagram of an example process 900 for generating data defining a ligand based on the denoised atom state data for each ligand atom. For convenience, the process 900 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 900.
[0333] The system receives respective denoised atom state data for each ligand atom (902). The denoised atom state data for a ligand atom refers to a current estimate of the denoised atom state data generated for the ligand atom at the final denoising time step in the sequence of denoising time steps, as described above with reference to FIG. 5. Optionally, the system also receives respective denoised bond data for each pair of ligand atoms. The denoised bond data for a pair of ligand atoms refers to a current estimate of the denoised bond data generated for the pair of ligand atoms at the final denoising time step in the sequence of denoising time steps, as described above with reference to FIG. 5.
[0334] The system selects a subset of the set of ligand atoms for inclusion in the ligand (904). In particular, as described above, each ligand atom in the set of ligand atoms represents an atom that is eligible for inclusion in the ligand. The system is not aware, prior to generating the denoised atom state data for the ligand atoms, which ligand atoms will be selected for inclusion in the ligand and which ligand atoms will be discarded. The number of ligand atoms in the set of ligand atoms thus represents the maximum number of atoms that can be included in the ligand.
[0335] For each ligand atom, the system can determine whether to select the ligand atom for inclusion in the ligand based on the 3D spatial position of the ligand atom as defined by the denoised atom state data for the ligand atom. For instance, the system can determine that a ligand atom should be selected for inclusion in the ligand only if the 3D spatial position of the ligand atom is at least a threshold distance from a predefined 3D spatial position referred to as the “throw-away” position. In this example, the system may have trained the generative diffusion model to move ligand atoms that are not included in the ligand to the throw-away position. Training the generative diffusion model to move unneeded ligand atoms to the throwaway position is described in more detail below with reference to FIG. 11.
[0336] The system filters the set of ligand atoms to remove any ligand atoms that are not selected for inclusion in the ligand (906). Filtering the set of ligand atoms to remove ligandatoms that are not selected for inclusion in the ligand may reduce the number of ligand atoms in the set of the ligand atoms by any appropriate amount, e.g., by 10%, or 50%, or 90%. After the filtering, a ligand atom is included in the set of ligand atoms if and only if it represents an atom that is selected for inclusion in the ligand. In some cases, the system may select all the ligand atoms in the original set of ligand atoms for inclusion in the ligand and thus no ligand atoms are filtered from the set of ligand atoms.
[0337] For each ligand atom in the set of ligand atoms, the system selects a respective value for each categorical feature that is represented in the denoised atom state data for the ligand atom (908). A categorical feature can refer to a feature having a value that is selected from a finite set of possible values. Examples of categorical features can include, e.g., atomic hybridization state or atomic element type. The denoised atom state data for an atom can represent a categorical feature by a distribution over the set of possible values of the categorical feature, where the distribution assigns a respective score to each possible value of the categorical feature.
[0338] The system can select a value for a categorical feature of a ligand atom based on the distribution over possible values of the feature that is included in the denoised atom state data for the atom in any of a variety of possible ways. For instance, the system can select the value that is assigned the highest score by the distribution over the set of possible values of the feature as the value of the categorical feature. As another example, the system can sample the value of the categorical feature from a probability distribution defined by the distribution over the set of possible values of the feature.
[0339] Optionally, for each pair of ligand atoms in the set of ligand atoms, the system selects a respective value for the bond feature that is represented in the denoised bond data for the pair of ligand atoms. As described above, the system can select a value for the bond feature for a pair of ligand atoms based on the distribution over possible bond types that is included in the denoised bond data for the pair of ligand atoms.
[0340] For each ligand atom in the set of ligand atoms, the system selects a respective value for each continuous feature that is represented in the denoised atom state data for the ligand atom (910). A continuous feature can refer to a feature having a value that is selected from a continuous set of possible values. Examples of continuous features can include, e.g., 3D spatial position, atomic partial charge, and so forth. The denoised atom state data for an atom can directly represent a continuous feature in one or more assigned dimensions of the denoised atom state data, and the system can select the value for the continuous feature by extracting thecorresponding dimensions representing the value of the continuous feature from the denoised atom state data for the atom.
[0341] The system provides data defining the ligand (912). The data defining the ligand includes data identifying each atom that is included in the ligand, and a set of features associated with each atom that is included in the ligand. The set of features associated with an atom that is included in the ligand can include continuous features, e.g., the 3D spatial position of the atom and the partial charge of the atom, and categorical features, e.g., the hybridization state of the atom and the element type of the atom. When the ligand is a protein and when the categorical features for the ligand atoms include an amino acid type feature, the data defining the ligand can additionally include, for each atom in the ligand, data identifying a type of an amino acid in which the atom is included. When the ligand is a nucleic acid and when the categorical features for the ligand atoms include a nucleotide type feature, the data defining the ligand can additionally include, for each atom in the ligand, data identifying a type of a nucleotide in which the atom is included.
[0342] Optionally, the data defining the ligand can also define, for each pair of atoms in the ligand, whether the pair of atoms is connected by a bond, and if so, the type of the bond.
[0343] In some implementations, the ligand design criteria specify that multiple ligands are generated by the design system. In these implementations, the set of ligand atoms can be associated with a partition into multiple subsets that are each associated with a respective ligand, as described above with reference to FIG. 3. The system can carry out the steps of the process 900 for each subset of ligand atoms that are associated with a respective ligand to generate the output data defining the ligand.
[0344] In some implementations, the ligand is a protein, but the categorical features associated with the ligand atoms do not include an amino acid type feature. In these implementations, the system can process the data defining the 3D structure of the ligand using a neural network, referred to as a “sequence design neural network,” to generate one or more predicted amino acid sequences for the ligand. More specifically, each predicted amino acid sequence is predicted to: (i) achieve the 3D structure that was generated as an output of the generative machine learning model, and (ii) bind to the target molecule. An example of a sequence design neural network is ProteinMPNN, as described with reference to: Dauparas, Justas, et al. "Robust deep learning-based protein sequence design using ProteinMPNN." Science 378.6615 (2022): 49-56.
[0345] FIG. 10 is a flow diagram of an example process 1000 for jointly training the embedding neural network and the generative model of the design system. In the exampleprocess 1000, the generative model is a model that implements a differentiable generative process. For instance, the generative model can be a generative diffusion model implemented using a denoising neural network, as described above with reference to FIG. 5. For convenience, the process 1000 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1000.
[0346] The system receives data characterizing a set of target molecule - ligand complexes (1002). More specifically, for each target molecule - ligand complex, the system receives: (i) a respective set of atom features of each atom in the target molecule(s) and each atom in the ligand(s), and optionally, (ii) a set of ligand properties of the ligand. For each atom in the complex, the set of atom features of the atom defines at least a 3D spatial position of the atom in a complex where the ligand is bound to a respective binding site on a target molecule. For each ligand atom in the ligand, the set of atom features of the atom can include additional features such as the hybridization state of the atom, the partial charge of the atom, and the element type of the atom. The set of ligand properties of the ligand can include global ligand properties (that characterize the ligand as a whole rather than being specific to a single atom in the ligand) or atom-specific ligand properties (that characterize specific atoms in the ligand rather than the entire ligand) or both. Examples of global ligand properties and atom-specific ligand properties are described with reference to FIG. 1. The set of ligand properties of the ligand can also define, for each pair of atoms in the ligand, whether the pair of atoms is connected by a bond, and if so, the type of the bond.
[0347] Generally, for a given molecule complex that is used for training the embedding neural network and the generative model, the complex can be partitioned into one or more target molecules and one or more ligands in any appropriate manner. That is, any molecule in the complex can be designated as a target molecule, and any molecule in the complex can be designated as a ligand molecule. During training, any molecule designated as a ligand molecule will be generated at least partially de novo by the design system, and any molecule designated as a target molecule will not be generated de novo by the design system.
[0348] The system can receive data characterizing molecule complexes that include a variety of type of molecules. For instance, the system can receive data characterizing: protein-protein complexes, small molecule - protein complexes, nucleic acid-protein complexes, lipid-protein complexes, protein-carbohydrate complexes, nucleic acid-nucleic acid complexes, small molecule-nucleic acid complexes, lipid-lipid complexes, carbohydrate-carbohydratecomplexes, small molecule-lipid complexes, ion-protein complexes, metal ion-nucleic acid complexes, multi-subunit enzyme complexes, or any combination thereof.
[0349] The data characterizing the target molecule - ligand complexes can be obtained, e.g., through physical experiments, or through computational methods (e.g., physics-based simulations or predictions generated by machine learning systems), or through a combination of both. Suitable data can be obtained from public databases, such as the protein data bank (PDB).
[0350] Optionally, the system can filter the set of target molecule - ligand complexes using one one or more filtering criteria (1003). More specifically, the system can determine, for each target molecule - ligand complex, whether the target molecule - ligand complex satisfies the filtering criteria. In response to determining that a target molecule - ligand complex satisfies the filtering criteria, the system can determine that one or more training examples should be generated based on the target molecule - ligand complex, e.g., as described at step 1004. Conversely, in response to determining that a target molecule - ligand complex does not satisfy the filtering criteria, the system can remove the target molecule - ligand complex from the set of target molecule - ligand complexes, e.g., so that no training examples are generated from the target molecule - ligand complex.
[0351] The set of filtering criteria can be based on features of the ligand(s) and / or the target molecule(s). The set of filtering criteria can be selected (e.g., by a user) to encourage the embedding neural network and the generative model, when trained on the training examples derived from the filtered set of target molecule - ligand complexes, to generate ligands having desired properties. A few examples of possible filtering criteria are described next.
[0352] In one example, the set of filtering criteria can be defined so that a target molecule - ligand complex satisfies the filtering criteria only if the ligand of the target molecule - ligand complex has one or more specified properties, e.g., global ligand properties or atom-specific ligand properties. In a particular example, the target molecule can be a protein, and the specified properties may be functional properties of the ligand, so that a protein - ligand complex satisfies the filtering criteria only if the ligand is a full agonist for the protein, or a partial agonist for the protein, or an antagonist for the protein, or an inverse agonist for the protein, etc. Training the embedding neural network and the generative model on training examples derived from a set of protein - ligand complexes that have been filtered in this manner encourages the embedding neural network and the generative model to generate ligands having the specified properties.
[0353] In another example, the set of filtering criteria can be defined so that a target molecule - ligand complex satisfies the filtering criteria if either: (i) the ligand has one or more specified ligand properties and the target molecule is included in a specified target molecule class, or (ii) the target molecule is not included in the specified target molecule class. The specified ligand properties can be any appropriate global ligand properties or atom-specific ligand properties. The target molecule class can be any group of target molecules defined by their structure, function, or role in biological processes. Examples of protein classes can include, e.g., G- protein-coupled receptors (GPCRs), enzymes, structural proteins, transport proteins, signaling proteins, receptor proteins, motor proteins, storage proteins, defense proteins, regulatory proteins, or chaperone proteins. In a particular example, the specified ligand property can be a functional property, e.g., agonism, or partial agonism, or antagonism, or inverse agonism, and so forth; and the specified protein class can be GPCRs. Training the embedding neural network and the generative model on training examples derived from a set of target molecule - ligand complexes that have been filtered in this manner encourages the embedding neural network and the generative model to generate ligands having the specified ligand properties when the target molecule is included in the specified target molecule class. Defining that a target molecule - ligand complex satisfies the filtering criteria for any target molecule that is not included in the specified target molecule class reduces the number of target molecule - ligand complexes that are filtered increases the diversity of the filtered set of target molecule - ligand complexes.
[0354] The system generates a set of training examples (1004). Each training example corresponds to a respective target molecule - ligand complex and includes data defining: (i) a training input to the design system, and (ii) a target output of the design system.
[0355] For each training example, the training input includes target molecule data and, optionally, a set of ligand design criteria.
[0356] The target molecule data can include any appropriate data characterizing the target molecule(s), e.g., for a target molecule that is a protein, the target molecule data can define one or more amino acid sequences of the protein, or data defining an MSA for the protein, or data characterizing a respective structure of each of one or more template proteins, or a combination thereof.
[0357] Optionally, the system can generate ligand design criteria for inclusion in the training input. The ligand design criteria can specify a set of target ligand properties of the ligand, or scaffolding data, or both. The set of target ligand properties can include global ligand properties of the ligand, atom-specific properties of the ligand (e.g., that are derived from the atomfeatures of the atoms in the ligand), or both. The scaffolding data can include target molecule scaffolding data (that specifies at least a portion of the 3D structure of the target molecule(s)), ligand scaffolding data (that specifies a portion of the 3D structure of the ligand), or both. The system can derive the scaffolding data from the 3D spatial positions of the atoms in the complex as specified by the atom feature data.
[0358] The target output of the design system can be based on the atom feature data of the atoms in the ligand, and optionally, of the atoms in the target molecule(s). For instance, the target output of the design system can include the atom feature data of each atom in the ligand and of each atom in the target molecule(s). A particular example of a target output of a generative diffusion model is described in more detail below with reference to FIG. 11.
[0359] The set of training examples can include training examples with varying amounts of ligand design criteria. For instance, the set of training examples can include some training examples that include no ligand design criteria, some training examples that include ligand design criteria specifying target ligand properties but not scaffolding data, some training examples that include ligand design criteria specifying scaffolding data but not target ligand properties, and some training examples that include ligand design criteria specifying target ligand properties and scaffolding data.
[0360] The system can generate multiple training examples corresponding to a single target molecule -ligand complex, where each training example includes a different set of ligand design criteria based on the known properties of the ligand the known structure of the target molecule-ligand complex. For instance, for a given target molecule-ligand complex, the system can generate a first training example that does not include any ligand design criteria, a second training example that includes target ligand properties but not scaffolding data, and a third training example that includes scaffolding data but not target ligand properties.
[0361] The system jointly trains the embedding neural network and the generative model on the set of training examples by a machine learning training technique (1006). More specifically, for each training example, the system can process the training input of the training example using the embedding neural network and the generative model to generate a predicted output of the generative model. The system can evaluate an objective function that measures an error (e.g., a root mean square deviation (RMSD), or a mean absolute error (MAE), or a mean squared error (MSE)) between: (i) the predicted output of the generative model, and (ii) the target output of the generative model. The system can determine gradients of the objective function with respect to the parameters of the embedding neural network and the generative model, e.g., using backpropagation. (The parameters of the generative model can include, e.g.,a set of neural network parameters of a neural network implemented by the generative model). The system can then update the current values of the parameters of the embedding neural network and the generative model using the gradients, e.g., by the update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0362] In some implementations, the system trains the embedding neural network and the generative model on training examples derived from a filtered set of target molecule - ligand complexes, e.g., to encourage the generation of ligands having desired properties. Filtering a set of target molecule - ligand complexes is described above at step 1003.
[0363] In some implementations, the system first trains the embedding neural network and the generative model on training examples derived from the full set of target molecule - ligand complexes (i.e. , without performing any filtering), and then trains (fine-tunes) the embedding neural network and the generative model on training examples derived from a filtered set of target molecule - ligand complexes. Filtering a set of target molecule - ligand complexes is described above at step 1003.
[0364] Specific aspects of the training (e.g., the operations of the generative model during the training and the objective function) may depend on the implementation of the generative model. An example process for training a generative diffusion model that includes a denoising neural network on a training example is described in more detail next with reference to FIG. 11. In the example process of FIG. 11, the operations of the generative diffusion model are modified during training (e.g., as compared to the operations of the generative diffusion model during inference, e.g., as described with reference to FIG. 5), as will be described in more detail below.
[0365] In some implementations, the generative model includes one or more “confidence estimation” neural network layers that process atom embeddings generated by the generative model to generate confidence measures for the predicted atom state data of atoms or pairs of atoms in a complex (as described above with reference to FIG. 5). The system can jointly train the confidence estimation neural network layers along with the embedding neural network and the generative model. For instance, for each atom in the complex, the system can generate a confidence measure for the atom that defines a predicted error in the atom state data of the atom, and the objective function can include a term that measures a difference between: (i) the predicted error in the atom state data of the atom, and (ii) the actual error in the atom state data of the atom. As another example, for each pair of atoms in the complex, the system can generate a confidence measure that defines a predicted error in the relative 3D displacement of the pair of atoms, and the objective function can include a term that measures a difference between: (i)the predicted relative 3D displacement of the pair of atoms, and (ii) the actual relative 3D displacement of the pair of atoms. The system can backpropagate gradients through the confidence estimation neural network layers, and optionally, into the generative model and / or the embedding neural network.
[0366] FIG. 11 is a flow diagram of an example process 1100 for jointly training an embedding neural network and a generative diffusion model on a training example. For convenience, the process 1100 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1100.
[0367] The system generates latent conditioning data by processing: (i) the target molecule data, and (ii) any ligand design criteria, included in the training input of the training example using the embedding neural network (1102).
[0368] As part of generating the latent conditioning data, the embedding neural network can generate a design embedding that represents any ligand design criteria provided in the input to the embedding neural network, as described with reference to FIG. 3A. The design embedding can include a respective atom embedding representing each ligand atom in a set of possible ligand atoms that are eligible for inclusion in the ligand. The number of ligand atoms represented in the design embedding define the maximum number of atoms that the system can select for inclusion in the ligand. As part of generating the latent conditioning data, the system can cause the embedding neural network to generate a design embedding that includes a number of atom embeddings that is at least as great as the number of atoms included in the ligand of the training example.
[0369] For each atom in the complex, i.e., for each atom in the target molecule and for each ligand atom in the set of possible ligand atoms, the system generates target atom state data for the atom based on the set of atom features of the atom (1104). The set of atom features of the atom can include continuous features, e.g., that define the 3D spatial position of the atom and the partial charge of the atom, and categorical features, e.g., that define the hybridization state and element type of the atom. For each continuous feature in the atom state data of the atom, the system can populate one or more corresponding dimensions of the target atom state data for the atom with numerical values defining the continuous feature. For each categorical feature in the atom state data, the system can populate a set of corresponding dimensions of the target atoms state data for the atom with a distribution over a set of possible values of the categorical10feature, e.g., a one-hot distribution over the set of possible values of the categorical feature that uniquely identifies the actual value of the categorical feature.
[0370] If the number of ligand atoms in the set of possible ligand atoms is greater than the actual number of atoms included in the ligand, then the set of possible ligand atoms includes one or more “extra” ligand atoms that will not be selected for inclusion in the ligand. For each extra ligand atom, the system can generate target atom state data that defines the 3D spatial position of the atom as being a predefined “throw-away” position. The system can set the other dimensions of the target atom state data for the extra ligand atom (i.e., the dimensions that do not define the 3D spatial position of the extra ligand atom) to masked (default) values.
[0371] The system samples a time step from the sequence of denoising time steps (1106). More specifically, during inference, the generative diffusion model can be configured to perform a sequence of denoising time steps, e.g., as described with reference to steps 504-512 of FIG. 5. During training, the system can randomly sample a single denoising time step from the sequence of denoising time steps, e.g., in accordance with a uniform distribution over the sequence of denoising time steps.
[0372] The system generates respective noisy atom state data for each atom in the complex by combining random noise with the target atom state data of the atom (1108). For instance, for each atom in the complex, the system can generate the noisy atom state data for the atom by adding random noise to the target atom state data of the atom. The system can scale the random noise combined with the target atom state data of the atoms by a constant that depends on the sampled time step, e.g., where the values of the constants corresponding to the denoising time steps are defined by a noise schedule. Optionally, the system can refrain from combining random noise with any dimensions of the target atom state data for an atom that are designated as being scaffolding data.
[0373] The system generates a denoising output using the denoising neural network while the denoising neural network is conditioned on the latent conditioning data (1110). An example process for generating a denoising output is described in detail with reference to FIG. 6. At step 602 of FIG. 6, the current atom state data for each atom in the complex can be defined as the noisy atom state data for each atom in the complex.
[0374] The system determines gradients of an objective function that depends on the denoising output, and uses the gradients to update the parameter values of the denoising neural network and the embedding neural network (1112). The objective function can measure an error between: (i) the denoising output of the denoising neural network, and (ii) a target output of the denoising neural network. The target output of the denoising neural network can define anoutput of the denoising neural network that, if used to generate a current estimate of the atom state data of the atoms in the complex (as described in step 508 of FIG. 5), would cause the current estimate of the atom state data of the atoms to match the target atom state data of the atoms in the target molecule - ligand complex of the training example.
[0375] In some cases, the dimensions of the target atom state data for the extra ligand atoms (i.e., that are not included in the ligand) other than those that define the 3D spatial position of the extra ligand (i.e., as the throw-away position) are set to masked values. The system can exclude the masked dimensions of the target atom state data for the extra ligand from the computation of the objective function, i.e., such that the objective function is independent of the values of the masked dimensions of the target atom state data.
[0376] FIG. 12 is a flow diagram of an example process 1200 for generating bond data for a ligand generated by the design system. For convenience, the process 1200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1200.
[0377] The system receives a respective atom embedding for each atom in the ligand (1202).
[0378] In some implementations, the atom embeddings of the atoms of the ligand are generated, e.g., as an intermediate output of the generative model that generates the ligand. For instance, for a generative diffusion model implemented using a denoising neural network, as described with reference to FIG. 5, the system can receive a respective atom embedding for each atom in the ligand that is generated by the denoising neural network during a final denoising time step in a sequence of denoising time steps performed by generative diffusion model. In particular, for each atom in the ligand, the system can receive a respective atom embedding for the atom that is generated as an output of an update block of the denoising neural network at the final denoising time step in the sequence of denoising time steps, as described at step 606 of FIG. 6.
[0379] In some implementations, for each atom in the ligand, the atom embedding of the atom comprises a collection of atom features characterizing the atom, e.g., including one or more of: the 3D spatial position of the atom (e.g., in a complex with atarget molecule), the hybridization state of the atom, the partial charge of the atom, the element type of the atom, and so forth. The atom features characterizing the atoms in the ligand can generated as part of the output of the generative model that generates the ligand, e.g., as described in FIG. 5.
[0380] The system processes the ID sequence of atom embeddings of the atoms in the ligand to generate a 2D array of “pair” embeddings (1204). The 2D array of embeddings can berepresented by data having dimensionality NumAtoms x NumAtoms x d where NumAtoms is the number of atoms in the ligand and d is a positive integer value defining the number of channel dimensions in each embedding. The system can generate the 2D array of pair embeddings from the ID sequence of atom embeddings in any of a variety of ways. For instance, the system can generate the 2D array of pair embeddings as a result of an element- wise outer product of the ID sequence of atom embeddings with itself. As another example, the system can generate the 2D array by an appropriate 2D concatenation operation, e.g., where the embedding at each position (i,y) in the 2D array of pair embeddings is generated by concatenating: (i) the atom embedding at position i, and (ii) the atom embedding at position y, in the ID sequence of atom embeddings (where indices i, j G {1, ... A}, where N is the number of atoms in the ligand).
[0381] Optionally, the system receives data identifying a respective 3D spatial position of each atom in the ligand (1206). The 3D spatial positions of the atoms in the ligand can be generated the generative model, as described throughout this specification.
[0382] Optionally, the system processes the 3D spatial positions of the atoms in the ligand to generate a 2D spatial distance array that, at position (i,j), represents a spatial distance between atom i and atom j in the ligand (1208).
[0383] Optionally, the system updates the 2D array of pair embeddings based on the 2D spatial distance array (1210). For instance, the system can channel-wise concatenate the 2D spatial distance array to the 2D array of pair embeddings, i.e., such that the 2D array of pair embeddings has dimensionality NumAtoms x NumAtoms x (d + 1), where (as above) NumAtoms is the number of atoms in the ligand, and d is a positive integer value defining the number of channel dimensions in each pair embedding prior to the channel-wise concatenation.
[0384] The system processes the 2D array of pair embeddings using a bond prediction machine learning model to generate bond data that bond data that defines, for each pair of atoms in the ligand, whether the pair of atoms are connected by a bond (1212). The bond data can further define one or more respective properties of each bond in the ligand, e.g., the type of the bond, e.g., single, double, or triple covalent bond, or ionic bond, or coordinate covalent bond, and so forth.
[0385] For instance, the bond prediction machine learning model can generate a model output that defines, for each pair of atoms in the ligand, a distribution over a set of possible bond categories. The set of possible bond categories includes a “no bond” category, indicating that the pair of the atoms are not connected by a bond, and a respective category for each of multipletypes of bonds that can exist between the pair of atoms, e.g., single, double, or triple covalent bond, or ionic bond, or coordinate covalent bond, and so forth.
[0386] For each pair of atoms in the ligand, the bond prediction machine learning model can select a bond category for the pair of atoms based on the corresponding distribution over the set of possible bond categories. For instance, the system can select the bond category for the pair of atoms as the bond category having the highest score under the distribution over the set of bond categories. As another example, the system can sample the bond category for the pair of bonds from distribution over the set of bond categories.
[0387] The bond prediction machine learning model can be any appropriate type of machine learning model. For instance, the bond prediction machine learning model can be a neural network, or a random forest, or a support vector machine, and so forth. In a particular example, the bond prediction machine learning model can be implemented as a neural network having any appropriate neural network architecture, in particular, an architecture that includes any appropriate types of neural network layers (e.g., fully connected layers, attention layers, convolutional layers, and so forth) in any appropriate number (e.g., 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0388] The system can train the bond prediction machine learning model on a set of training examples using a machine learning training technique. Each training example corresponds to a ligand and can include: (i) a training input to the bond prediction machine learning model, and (ii) target bond data for the ligand. The training input to the bond prediction machine learning model can include respective atom embedding for each atom in the ligand, and optionally, 3D spatial position data for each atom in the ligand. For each training example, the system trains the bond prediction machine learning model to reduce a discrepancy between: (i) target bond data specified by the training example, and (ii) predicted bond data generated by the bond prediction machine learning model by processing the training input of the training example.
[0389] In some implementations, the bond prediction machine learning model is implemented as a neural network, the generative model is implemented as a generative diffusion model parameterized by a denoising neural network, and the bond prediction machine learning model receives atom embeddings that are generated as an intermediate output of the denoising neural network (as described above). In these implementations, the system can train the bond prediction machine learning model jointly with the embedding neural network and the denoising neural network.
[0390] More specifically, for each training example, the training input includes atom embeddings generated as an intermediate output of the denoising neural network. As part oftraining the bond prediction machine learning model on a training example, the system can determine gradients (e.g., by backpropagation) of an objective function (e.g., a cross-entropy objective function) that measures a discrepancy between: (i) the target bond data specified by the training example, and (ii) predicted bond data generated by the bond prediction machine learning model by processing the training input of the training example. In particular, the system can determine gradients of the objective function with respect to not only the set of parameters of the bond prediction machine learning model, but also with respect to parameters of the denoising neural network and the embedding neural network. The system can then adjust the parameters of the bond prediction machine learning model, the denoising neural network, and the embedding neural network using the gradients, e.g., in accordance with the update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam. That is, the system can backpropagate gradients of the objective function through the bond prediction neural network and into the denoising neural network and the embedding neural network.
[0391] FIG. 13 is a flow diagram of an example process for in-painting and / or out-painting a target molecule-ligand complex. For convenience, the process 1300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a design system, e.g., the design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1300.
[0392] The system receives respective atom properties for each atom in a complex comprising a input target molecule and an input ligand (1302). Atom properties for an atom can define one or more of: a 3D spatial position of the atom (e.g., when the protein and the ligand are bound in a complex), a hybridization state of the atom, a partial charge of the atom, or an element type of the atom.
[0393] The system receives a request to perform in-painting of the complex, or out-painting of the complex, or both (1304).
[0394] A request to in-paint the complex identifies, for each atom property of each atom in the complex, whether the atom property is a “static” property or a “variable” property. In-painting the complex refers to generating new ligands that have the static atom properties of the input ligand and that bind to target molecules that have the static atom properties of the input target molecule.
[0395] In implementations where the system receives a request to perform in-painting, the system receives input data that, for each atom property of each atom in the complex, identifies the atom property as a static property or as a variable property. The input data identifies at least one atom property of at least one atom in the ligand as being a variable property. The systemcan receive the input that identifies the static and variable atom properties, e.g., from a user or from an upstream system, e.g., by way of a user interface (e.g., a graphical user interface) made available by the system, or by an API made available by the system. For instance, a graphical user interface can allow a user to dynamically select parts of the complex as variable or static, e.g., by interacting with a 3D representation of the complex presented to the user by way of the user interface.
[0396] A request to out-paint the complex defines a request to generate new ligands that expand on the input ligand, e.g., by including one or more new (additional) atoms relative to the input ligand.
[0397] In implementations where the system receives a request to perform in-painting, the system generates ligand design criteria that include the static atom properties for the atoms in the complex (1306). The ligand design criteria can include, e.g., one or more of: ligand scaffolding data, atom-specific properties of atoms in the ligand, protein scaffolding data, or atom-specific properties of atoms in the protein.
[0398] The system processes an input including any ligand design criteria, e.g., using the embedding neural network and the generative model described throughout this specification, to generate data defining one or more ligands (1308). In implementations where the system performs in-painting, the ligand design criteria are selected (as described above) such that the system attempts to generate ligands that have the static atom properties of the ligand atoms and that bind to target molecules that have the static atom properties of the target molecule atoms. That is, the system attempts to generate ligands that in-paint the variable atom properties in a variety of possible ways. In some cases, the system generates ligands that each differ from the ligand received as an input at step 1302, e.g., as a result of stochasticity in the operations of the generative model.
[0399] In implementations where the system performs out-painting of the ligand, the system can generate a design embedding of the ligand design criteria that includes a respective atom embedding for each atom in a set of atoms that includes: (i) all the atoms included in the input ligand, and (ii) one or more “out-painted” ligand atoms. The out-painted ligand atoms represent atoms that the generative model can select for inclusion in a new ligand as part of out-painting the input ligand.
[0400] The system can design ligands for use in various applications. For example, the system can be used to obtain a ligand that is a drug or a ligand of an industrial enzyme, or a ligand for a target molecule identified as being involved in one or both of plant growth and stressresistance of an agricultural plant. In some implementations, the ligand can be used for pest or pathogen control in agriculture.
[0401] By way of example, the system can be used to determine a plurality of candidate ligands for a target molecule. A respective interaction of each candidate ligand with the target molecule can then be determined and one or more of the candidate ligands selected dependent on a result of the evaluating. Optionally, the selected ligands can be synthesized, e.g., for testing for biological activity of the ligand in vitro or in vivo.
[0402] The evaluation of the interaction of a candidate ligand with the target molecule may be performed using a computer-aided approach in which graphical models of the candidate ligand and target molecule structure are displayed for user-manipulation, and / or the evaluation may be performed partially or completely automatically, for example using standard molecular (e.g., protein-ligand) docking software. In some implementations the evaluation may include determining an interaction score for the candidate ligand, where the interaction score includes a measure of an interaction between the candidate ligand and the target molecule. The interaction score may be dependent upon a strength and / or specificity of the interaction, e.g., a score dependent on binding free energy. A candidate ligand may be selected dependent upon its score.
[0403] In some implementations the target molecule includes a receptor or enzyme and the ligand is an agonist or antagonist of the receptor or enzyme. In some implementations, the method may be used to identify the structure of a cell surface marker. This may then be used to identify a ligand, e.g., an antibody or aptamer or a label such as a fluorescent label, which binds to the cell surface marker. This may be used to identify and / or treat cancerous cells.
[0404] In some implementations, the ligand is a drug and the interaction of each of a plurality of target molecules with each of the candidate ligands is evaluated. Then one or more of the candidate ligands may be selected either to obtain a ligand that (functionally) interacts with each of the target molecules, or to obtain a ligand that (functionally) interacts with only one of the target molecules. For example, in some implementations it may be desirable to obtain a drug that is effective against multiple drug targets. Also or instead, it may be desirable to screen a drug for off-target effects. For example, in agriculture it can be useful to determine that a drug designed for use with one plant species does not interact with another, different plant species and / or an animal species.
[0405] The system and method described herein can also be used to obtain a diagnostic antibody or aptamer marker of a disease. There is also provided a method that comprises selecting a target molecule that is to be recognized by the antibody or aptamer marker, anddetermining a plurality of candidate antibodies or aptamers for the target molecule. The method may also involve, evaluating an interaction between each of the one or more candidate antibodies or aptamers and the target molecule, and selecting one of the one or more of the candidate antibodies or aptamers as the diagnostic antibody or aptamer marker dependent on a result of the evaluating, e.g., selecting one or more candidate antibodies or aptamers that have the highest affinity to the target protein. The method may include making the diagnostic antibody or aptamer marker. The diagnostic antibody or aptamer marker may be used to diagnose a disease by detecting whether it binds to the target molecule (e.g., protein) in a sample obtained from a patient, e.g., a sample of bodily fluid. As described above, a corresponding technique can be used to obtain a therapeutic antibody or aptamer (e.g., polypeptide or polynucleotide ligand).
[0406] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0407] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0408] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including byway of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0409] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0410] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0411] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0412] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a centralprocessing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0413] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0414] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0415] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
[0416] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0417] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0418] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0419] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0420] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules andcomponents in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0421] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers for computationally designing a ligand for binding to a target molecule, the method comprising: obtaining target molecule data characterizing at least a portion of a target molecule; processing a network input comprising the target molecule data using an embedding neural network to generate latent conditioning data representing the target molecule; generating, using a generative model and while the generative model is conditioned on the latent conditioning data representing the target molecule, predicted ligand data defining a predicted ligand that is predicted to bind to the target molecule; and providing the predicted ligand data defining the predicted ligand.
2. The method of claim 1, further comprising obtaining ligand design criteria that define desired characteristics of the ligand; wherein the network input to the embedding neural network further comprises the ligand design criteria; and wherein the latent conditioning data jointly represents the target molecule and the ligand design criteria.
3. The method of claim 2, wherein the embedding neural network and the generative model are jointly trained to optimize an objective function that encourages the embedding neural network and the generative model to attempt to generate ligand data defining a ligand that binds to the target molecule and that satisfies the ligand design criteria.
4. The method of any one of claims 2-3, wherein the ligand design criteria define a respective target value for each of one or more global properties of the ligand that characterize the ligand as a whole.
5. The method of claim 4, wherein the ligand design criteria define a respective target value for a set of global properties of the ligand characterizing one or more of: a binding affinity of the ligand for the target molecule, absorption properties of the ligand, distribution properties of the ligand, metabolism properties of the ligand, excretion properties of the ligand, toxicity properties of the ligand, a number of rings in the ligand, a molecular weight of the ligand, a lipophilicity of the ligand, an ability of the ligand to donate or accept hydrogen bonds, a totalpolar surface area of the ligand, a number of rotatable bonds in the ligand, a number of chiral centers in the ligand, a number of electrophilic centers in the ligand, a number of nucleophilic centers in the ligand.
6. The method of any one of claims 2-5, wherein the ligand design criteria define a respective target value for each of one or more atom-specific properties of the ligand that each relate to a specific atom in the ligand.
7. The method of claim 6, wherein the ligand design criteria define a respective target value for a set of atom-specific properties of the ligand that, for each of one or more atoms in the ligand, characterize one or more of: an elemental type of the atom, a hybridization state of the atom, or a partial charge of the atom.
8. The method of any one of claims 2-7, wherein the ligand design criteria comprises target molecule scaffolding data, or ligand scaffolding data, or both.
9. The method of claim 8, wherein the target molecule scaffolding data defines a respective target three-dimensional (3D) spatial position of each of one or more atoms in the target molecule when in a complex with the ligand.
10. The method of any one of claims 8-9, wherein the ligand scaffolding data defines a respective target 3D spatial position of each of one or more atoms in the ligand when in a complex with the target molecule.
11. The method of any one of claims 2-10, wherein the embedding neural network comprises a target molecule embedding neural network and a design embedding neural network, and wherein processing the network input comprising the target molecule data and the ligand design criteria using the embedding neural network to generate the latent conditioning data comprises: processing the target molecule data using the target molecule embedding neural network to generate a target molecule embedding of the target molecule; processing the ligand design criteria using the design embedding neural network to generate a design embedding of the ligand design criteria; and processing the target molecule embedding and the design embedding to generate thelatent conditioning data.
12. The method of claim 11, wherein processing the ligand design criteria using the design embedding neural network to generate the design embedding of the ligand design criteria comprises: generating a collection of initial atom embeddings based on the ligand design criteria, wherein the collection of initial atom embeddings includes a respective atom embedding for each ligand atom in a set of possible ligand atoms that are eligible for inclusion in the ligand; processing the collection of initial atom embeddings, by a plurality of neural network layers of the design embedding neural network, to generate a collection of final atom embeddings that includes a respective final atom embedding for each ligand atom in the set of possible ligand atoms; wherein the collection of final atom embeddings defines the design embedding of the ligand design criteria.
13. The method of claim 12, wherein the ligand design criteria specify a respective target value for each of one or more global properties of the ligand; and wherein generating the collection of initial atom embeddings based on the ligand design criteria comprises: including the respective target value for each of the one or more global properties of the ligand in each initial atom embedding in the collection of initial atom embeddings.
14. The method of any one of claims 12-13, wherein the ligand design criteria specify a respective target value for each of one or more atom-specific properties of the ligand; and wherein generating the collection of initial atom embeddings based on the ligand design criteria comprises, for each atom-specific property of the ligand: including the target value of the atom-specific property only in the initial atom embedding representing the ligand atom that is characterized by the atom-specific property.
15. The method of any one of claims 12-14, wherein a number of ligand atom embeddings in the set of ligand atom embeddings defines a maximum number of atoms that can be selected for inclusion in the ligand.
16. The method of any one of claims 11-15, wherein the plurality of neural network layers of the design embedding neural network include one or more self-attention neural network layers.
17. The method of any one of claims 2-16, wherein the ligand design criteria leave undefined at least part of a chemical structure of the ligand.
18. The method of any one of claims 11-17, wherein the target molecule embedding comprises a respective component embedding of each component in the target molecule; wherein the design embedding comprises a respective atom embedding of each ligand atom in a set of ligand atoms that are eligible for inclusion in the ligand, wherein a number of ligand atom embeddings in the set of ligand atom embeddings defines a maximum number of atoms that can be selected for inclusion in the ligand; and wherein processing the target molecule embedding and the design embedding to generate the latent conditioning data comprises: generating data defining a ID sequence of component embeddings and atom embeddings by concatenating: (i) the component embeddings of the target molecule embedding, and (ii) the ligand atom embeddings of the design embedding; wherein the latent conditioning data is derived from the ID sequence of component embeddings and atom embeddings.
19. The method of claim 18, wherein the processing the target molecule embedding and the design embedding to generate the latent conditioning data further comprises: transforming the ID sequence of component embeddings and atom embeddings into a two-dimensional (2D) array of embeddings; wherein the latent conditioning data is derived from the 2D array of embeddings.
20. The method of claim 19, wherein the 2D array of embeddings comprises a plurality of atom - atom embeddings that are each derived from a respective pair of atom embeddings of the design embedding.
21. The method of any one of claims 19-20, wherein the 2D array of embeddings comprises a plurality of component - component embeddings that are each derived from a respective pairof component embeddings of the target molecule embedding.
22. The method of any one of claims 19-21 , wherein the 2D array of embeddings comprises a plurality of component - atom embeddings that are each derived from: (i) a respective atom embedding of the design embedding, and (ii) a respective component embedding of the target molecule embedding.
23. The method of any one of claims 19-22, wherein transforming the ID sequence of component embeddings and atom embeddings into the 2D array of embeddings comprises: applying an outer product operation to the ID sequence of component embeddings and atom embeddings; or applying a 2D concatenation operation to the ID sequence of component embeddings and atom embeddings.
24. The method of any one of claims 19-23, wherein the embedding neural network further comprises a fusion neural network; and wherein processing the target molecule embedding and the design embedding to generate the latent conditioning data further comprises: processing the 2D array of embeddings using the fusion neural network to generate an updated 2D array of embeddings; wherein the updated 2D array of embeddings defines the latent conditioning data.
25. The method of claim 24, wherein the fusion neural network comprises a sequence of self-attention blocks, wherein each self-attention block is configured to perform operations comprising: apply one or more self-attention operations to an input 2D array of embeddings to update the input 2D array of embeddings.
26. The method of claim 25, wherein for one or more of the self-attention blocks, the selfattention operations comprise one or more row-wise self-attention operations.
27. The method of any one of claims 25-26, wherein for one or more of the self-attention blocks, the self-attention operations comprise one or more column-wise self-attentionoperations.
28. The method of any one of claims 25-27, wherein for one or more of the self-attention blocks, the self-attention operations comprise one or more triangle self-attention operations.
29. The method of any preceding claim, wherein the target molecule is a protein and the target molecule data comprises data defining one or more of: an amino acid sequence of the target molecule; a multiple sequence alignment (MSA) for the target molecule; or a respective structure of each of one or more template target molecules.
30. The method of any preceding claim, wherein the generative model is a generative diffusion model that comprises a denoising neural network.
31. The method of claim 30, wherein generating, using the generative model and while the generative model is conditioned on the latent conditioning data, the predicted ligand data defining the predicted ligand comprises: generating respective atom state data for each atom in the target molecule and for each ligand atom in a set of ligand atoms that are eligible for inclusion in the ligand; denoising the atom state data over a sequence of time steps using the denoising neural network and while the denoising neural network is conditioned on the latent conditioning data; and generating the predicted ligand data based on the atom state data after a final time step in the sequence of time steps.
32. The method of claim 31, wherein for each ligand atom in the set of ligand atoms, the atom state data for the ligand comprises features characterizing: (i) a 3D spatial position of the ligand atom, and (ii) one or more of: a partial charge of the ligand atom, a hybridization state of the ligand atom, an elemental type of the ligand atom, or a type of an amino acid that includes the ligand atom, or a type of a nucleotide that includes the ligand atom.
33. The method of any one of claims 31-32, wherein generating respective atom state data for each atom in the target molecule and for each ligand atom in the set of ligand atoms comprises, for one or more atoms:stochastically sampling the atom state data for the atom.
34. The method of claim 33, wherein for one or more atoms, stochastically sampling the atom state data for the atom comprises, for each of one or more continuous features of the atom: stochastically sampling a value of the continuous feature of the atom from a probability distribution; and including the stochastically sampled value of the continuous feature of the atom in the atom state data for the atom.
35. The method of claim 34, wherein the one or more continuous features of the atom comprise respective features defining one or more of: a 3D spatial position of the atom or a partial charge of the atom.
36. The method of any one of claims 33-35, wherein for one or more atoms, stochastically sampling the atom state data for the atom comprises, for each of one or more categorical features of the atom: stochastically sampling a distribution over possible values of the categorical feature from a probability distribution; and including the stochastically sampled distribution over possible values of the categorical feature of the atom in the atom state data for the atom.
37. The method of claim 36, wherein the one or more categorical features of the atom comprise respective features defining one or more of: a hybridization state of the atom or an elemental type of the atom.
38. The method of any one of claims 31-37, wherein generating the predicted ligand data based on the atom state data after a final time step in the sequence of time steps comprises: selecting a subset of the ligand atoms in the set of ligand atoms for inclusion in the ligand, wherein fewer than all of the ligand atoms in the set of ligand atoms are selected for inclusion in the ligand; and filtering the set of ligand atoms to remove any ligand atom that is not selected for inclusion in the ligand.
39. The method of claim 38, wherein selecting a subset of the ligand atoms in the set of ligand atoms for inclusion in the ligand comprises, for each ligand atom: selecting the ligand atom for inclusion in the ligand only if a 3D spatial position of the ligand atom, as defined by the atom state data for the ligand atom, is at least a threshold distance from a predefined throw-away position; wherein the generative model has been trained to move respective 3D spatial positions of ligand atoms that are not included in the ligand to the throw-away position.
40. The method of claim 39, wherein generating the predicted ligand data based on the atom state data after the final time step in the sequence of time steps comprises, for each ligand atom in the set of ligand atoms: determining a respective value of each of one or more continuous features of the ligand atom from the atom state data for the ligand atom, comprising, for each continuous feature: extracting a value of the continuous feature from one or more corresponding dimensions of the atom state data for the ligand atom.
41. The method of any one of claims 39-40, wherein generating the predicted ligand data based on the atom state data after the final time step in the sequence of time steps comprises, for each ligand atom in the set of ligand atoms: determining a respective value of each of one or more categorical features of the ligand atom from the atom state data for the ligand atom, comprising, for each categorical feature: extracting a distribution over possible values of the categorical feature from a plurality of corresponding dimensions of the atom state data for the ligand atom; and determining the value of the categorical feature based on the distribution over possible values of the categorical feature.
42. The method of claim 41, wherein for one or more categorical features of the ligand atom, determining the value of the categorical feature based on the distribution over possible values of the categorical feature comprises: stochastically sampling the value of the categorical feature from the distribution over possible values of the categorical feature; or setting the value of the categorical feature equal to a possible value of the categorical feature having a highest score under the distribution over possible values of the categoricalfeature.
43. The method of any one of claims 31-42, wherein denoising the atom state data over the sequence of time steps using the denoising neural network and while the denoising neural network is conditioned on the latent condition data comprises, at each of one or more time steps in the sequence of time steps: receiving current atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at the time step; generating a denoising output using the denoising neural network and while the denoising neural network is conditioned on the latent conditioning data; and generating atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at a next time step using the denoising output.
44. The method of claim 43, wherein the denoising output comprises a respective predicted error in the current atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at the time step.
45. The method of any one of claims 43-44, wherein generating the denoising output using the denoising neural network and while the denoising neural network is conditioned on the latent conditioning data comprises: generating a set of atom embeddings using an encoder block of the denoising neural network, wherein each atom embedding represents one or more atoms in the target molecule or in the set of ligand atoms and is based at least in part on the respective current atom state data of the one or more atoms at the time step; processing the set of atom embeddings using an update block of the denoising neural network to generate a set of updated atom embeddings; and processing the set of updated atom embeddings to generate the denoising output.
46. The method of claim 45, wherein the set of atom embeddings includes a respective atom embedding representing each ligand atom in the set of ligand atoms.
47. The method of any one of claims 45-46, wherein the set of atom embeddings includes a respective atom embedding representing each atom included in each component of the targetmolecule.
48. The method of any one of claims 45-46, wherein for each component in the target molecule, the set of atom embeddings includes a respective atom embedding that jointly represents all the atoms included in the component.
49. The method of any one of claims 45-48, wherein each atom embedding in the set of atom embeddings is based on, for each of the one or more atoms represented by the atom embedding: (i) the current atom state data of the atom at the time step, and (ii) a respective conditioning embedding for the atom that is selected from a collection of embeddings included in the latent conditioning data.
50. The method of claim 49, wherein for each atom embedding that represents a ligand atom, the conditioning embedding for the atom comprises an atom - atom embedding corresponding to the ligand atom in the latent conditioning data.
51. The method of any one of claims 49-50, wherein for each atom embedding that represents an atom that is included in a component of the target molecule, the conditioning embedding for the atom comprises a component - component embedding corresponding to the component in the latent conditioning data.
52. The method of any one of claims 45-51, wherein the update block of the denoising neural network comprises a sequence of self-attention blocks; wherein each of the self-attention blocks are configured to apply one or more selfattention operations to a set of current atom embeddings to update the set of current atom embeddings; wherein each of the one or more self-attention operations are conditioned on the latent conditioning data.
53. The method of claim 52, wherein applying a self-attention operation to the set of current atom embeddings to update the set of current atom embeddings comprises: generating, based on the current set of atom embeddings, a respective intermediate attention score for each pair of current atom embeddings from the set of current atom embeddings;generating, based on the latent conditioning data, a respective attention score bias for each pair of current atom embeddings from the set of current atom embeddings; generating a respective final attention score for each pair of current atom embeddings from the set of current atom embeddings based on the intermediate attention scores and the attention score biases; and updating the set of current atom embeddings using the final attention scores.
54. The method of claim 53, wherein for each pair of current atom embeddings from the set of current atom embeddings, generating the attention score bias for the pair of current atom embeddings comprises: summing the intermediate attention score for the pair of current atom embeddings and the attention score bias for the pair of current atom embeddings.
55. The method of any one of claims 53-54, wherein for each pair of current atom embeddings from the set of current atom embeddings, generating the attention score bias for the pair of current atom embeddings comprises: processing a respective conditioning embedding selected from a collection of embeddings included in the latent conditioning data using a projection neural network to generate the attention score bias.
56. The method of claim 55, wherein for each pair of current atom embeddings that includes: (i) a first atom embedding representing a first ligand atom, and (ii) a second atom embedding representing a second ligand atom, the selected conditioning embedding comprises an atom - atom embedding corresponding to the first ligand atom and the second ligand atom in the latent conditioning data.
57. The method of any one of claims 55-56, wherein for each pair of current atom embeddings that includes: (i) a first atom embedding representing a first atom included in an component, and (ii) a second atom embedding representing a ligand atom, the selected conditioning embedding comprises an component - atom embedding corresponding to: (i) the component that includes the first atom, and (ii) the ligand atom, in the latent conditioning data.
58. The method of any one of claims 55-57, wherein for each pair of current atom embeddings that includes: (i) a first atom embedding representing a first atom included in afirst component, and (ii) a second atom embedding representing a second atom included in a second component, the selected conditioning embedding comprises an component - component embedding corresponding to: (i) the first component, and (ii) the second component, in the latent conditioning data.
59. The method of any one of claims 55-58, wherein for each pair of current atom embedding that includes: (i) a first atom embedding that jointly represents all atoms in a first component, and (ii) a second atom embedding that represents all atoms in a second component, the selected conditioning embedding comprises an component - component embedding corresponding to: (i) the first component, and (ii) the second component, in the latent conditioning data.
60. The method of any one of claims 55-59, wherein for each pair of current atom embeddings that includes: (i) a first atom embedding that jointly represents all atoms in an component, and (ii) a second atom embedding that represents a ligand atom, the selected conditioning embedding comprises an component - atom embedding corresponding to: (i) the component, and (ii) the ligand atom, in the latent conditioning data.
61. The method of any one of claims 31-60, wherein denoising the atom state data over the sequence of time steps using the denoising neural network further comprises, at each of one or more time steps in the sequence of time steps: processing at least some of the current atom state data using a property prediction neural network to generate a predicted value of a ligand property of a ligand characterized by the current atom state data; and determining gradients of a conditioning objective function with respect to at least some of the current atom state data, wherein the conditioning objective function measures a discrepancy between: (i) the predicted value of the ligand property, and (ii) a target value of the ligand property; wherein generating atom state data of each atom in the target molecule and of each atom in the set of ligand atoms at the next time step using the denoising output comprises: combining the gradients of the conditioning objective function with the denoising output.
62. The method of claim 61, wherein the property prediction neural network processes the current atom state data for the atoms in the target molecule and for the ligand atoms in the set of ligand atoms.
63. The method of any one of claims 61-62, wherein the property prediction neural network generates a predicted value of a binding affinity of the ligand for the target molecule.
64. The method of any preceding claim, further comprising: receiving a respective atom embedding for each atom in the ligand; and processing a model input based on the atom embeddings of the atoms in the ligand using a bond prediction machine learning model to generate bond data that defines, for each pair of atoms in the ligand, whether the pair of atoms are bonded.
65. The method of claim 64, wherein for each atom in the ligand, the atom embedding of the atom is generated as an intermediate output of the generative model.
66. The method of claim 65, wherein the generative model is a generative diffusion model comprising a denoising neural network; and wherein for each atom in the ligand, the atom embedding of the atom is generated as an intermediate output of the denoising neural network at a final time step in a sequence of denoising time steps.
67. The method of any one of claims 64-66, wherein processing a model input based on the atom embeddings of the atoms in the ligand using the bond prediction machine learning model to generate the bond data comprises: processing a one-dimensional (ID) sequence of atom embeddings of the atoms in the ligand to generate a two-dimensional (2D) array of pair embeddings; processing the 2D array of pair embeddings using the bond prediction machine learning model to generate the bond data.
68. The method of any one of claims 64-67, wherein the model input to the bond prediction machine learning model further comprises data defining a respective 3D spatial position of each atom in the ligand.
69. The method of any preceding claim, further comprising physically synthesizing the predicted ligand.
70. The method of any preceding claim, wherein the embedding neural network and the generative model have been trained by performing operations comprising: obtaining data defining a set of target molecule - ligand complexes; determining, for each target molecule - ligand complex in the set of target molecule - ligand complexes, whether the target molecule - ligand complex satisfies a set of filtering criteria; and during at least one training stage, training the embedding neural network and the generative model based only on training examples derived from target molecule - ligand complexes that satisfy the set of filtering criteria.
71. The method of claim 70, wherein for each target molecule - ligand complex in the set of target molecule - ligand complexes, the target molecule - ligand complex satisfies the set of filtering criteria only if the ligand of the target molecule - ligand complex has one or more specified ligand properties.
72. The method of claim 71, wherein the one or more specified ligand properties comprise functional properties, comprising one or more of: full agonism, partial agonism, antagonism, or inverse agonism.
73. The method of claim 70, wherein for each target molecule - ligand complex in the set of target molecule - ligand complexes, the target molecule - ligand complex satisfies the set of filtering criteria if either: (i) the ligand of the target molecule - ligand complex has one or more specified ligand properties and the target molecule of the target molecule - ligand complex is included in a specified target molecule class, or (ii) the target molecule of the target molecule - ligand complex is not included in the specified target molecule class.
74. The method of claim 73, wherein the one or more specified ligand properties comprise functional properties, comprising one or more of: full agonism, partial agonism, antagonism, or inverse agonism; and wherein the target molecule class is a G-protein-coupled receptor (GPCR) protein class.
75. The method of any one of claims 70-74, wherein during at least one training stage, training the embedding neural network and the generative model based only on training examples derived from target molecule - ligand complexes that satisfy the set of filtering criteria comprises: at a first training stage, training the embedding neural network and the generative model based on training examples derived from all target molecule - ligand complexes in the set of target molecule - ligand complexes; and at a second training stage, training the embedding neural network and the generative model based only on training examples derived from target molecule - ligand complexes that satisfy the set of filtering criteria.
76. The method of any preceding claim, wherein the target molecule is a protein.
77. The method of any preceding claim, wherein the ligand is a protein.
78. A method comprising: generating a collection of ligands for a target molecule using the method of any one of claims 1-77; determining, for each ligand in the collection of ligands, one or more respective properties of the ligand; and selecting one or more ligands in the collection of ligands for physical synthesis based at least in part on the properties of the ligands.
79. The method of claim 78, further comprising physically synthesizing the one or more selected ligands.
80. A method of obtaining a ligand, wherein the ligand is a drug or a ligand of an industrial enzyme, the method comprising: performing the method of any one of claims 1-77 to determine a plurality of candidate ligands for a target molecule; evaluating a respective interaction of each candidate ligand with the target molecule; and selecting one or more of the candidate ligands dependent on a result of the evaluating.
81. The method of claim 80, wherein the target molecule comprises a receptor or enzyme, and wherein each candidate ligand is an agonist or antagonist of the receptor or enzyme.
82. The method of claim 80, wherein the target molecule comprises an antibody or aptamer target, in particular a virus or cancer cell protein, and wherein the ligand binds to the antibody or aptamer target to provide a therapeutic effect.
83. A method of obtaining a ligand, wherein the ligand is for modifying one or both of plant growth and stress resistance of an agricultural plant, the method comprising: performing the method of any one of claims 1-77 to determine a plurality of candidate ligands for a target molecule identified as being involved in one or both of plant growth and stress resistance of the agricultural plant; evaluating a respective interaction of each candidate ligand with the target molecule; and selecting one or more of the candidate ligands as the ligand dependent on a result of the evaluating.
84. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-77.
85. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-77.
Citation Information
Cited By
Unified structure model for molecule property and structure prediction
WO2026052718A1