Predicting properties of molecule complexes using single molecule embedding neural networks and a molecule complex embedding neural network
A modular neural network architecture efficiently generates representations of molecule complexes, addressing computational intensity and data scarcity issues by using single molecule embeddings and 3D structure encoding, enabling accurate predictions.
Patent Information
- Application Number
- PCT/US2025/030836
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-05-23
- Publication Date
- 2025-12-26
AI Technical Summary
Existing methods for generating representations of molecule complexes, such as antibody-antigen complexes, are computationally intensive and require scarce training data, making them unsuitable for screening large numbers of complexes.
A modular neural network architecture comprising single molecule embedding neural networks and a molecule complex embedding neural network is used to generate representations of molecule complexes, leveraging plentiful single molecule data and reducing computational resources by encoding 3D structure information.
The system efficiently generates dense and information-rich representations of molecule complexes, consuming fewer computational resources and requiring less training data, enabling accurate predictions of properties like binding affinity and stability.
Smart Images

Figure US2025030836_26122025_PF_FP_ABST
Abstract
Description
PREDICTING PROPERTIES OF MOLECULE COMPLEXES USING SINGLE MOLECULE EMBEDDING NEURAL NETWORKS AND A MOLECULE COMPLEX EMBEDDING NEURAL NETWORKTECHNICAL FIELD
[0001] This specification relates to generating predictions characterizing molecule complexes using neural networks.BACKGROUND
[0002] Predictions can be made using machine learning models. Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
[0003] A molecule complex refers to an association of two or more molecules which can be held together by non-covalent interactions such as hydrogen bonding, ionic interactions, Van der Waals forces, or hydrophobic effects, and that forms a distinct, functional entity with specific properties or activities.SUMMARY
[0004] This specification describes a system implemented as computer programs on one or more computers in one or more locations that can generate predictions characterizing molecule complexes using neural networks.
[0005] Throughout this specification, an “embedding” of an entity can refer to a representation of the entity as an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values.
[0006] Throughout this specification, a “latent space” can refer to a space of embeddings, e.g., a Euclidean spacewhere N can be any positive integer value.
[0007] Throughout this specification, an “intermediate output” of a neural network can refer to an output generated by one or more hidden layers of the neural network, where each hidden layer is a layer in the neural network other than an input layer or an output layer.
[0008] Throughout this specification, a “small molecule’' can refer to a low molecular weight compound (e.g., organic compound), e.g., typically less than 900 Daltons. Small molecules can regulate biological processes by interacting with specific biomolecules and are often used in drug development due to their ability (in some cases) to easily diffuse across cell membranes.
[0009] Throughout this specification, an "antigen" can refer to a molecule (e.g., a protein or polysaccharide) that can be found, e.g., on the surface of pathogens such as bacteria, viruses, or fungi, or on the surface of infected cells, or as free molecules within the body, or on the surface of immune cells. Generally, an antigen can be any molecule that an antibody can recognize and bind to.
[0010] Throughout this specification, an “antibody"’ (e.g., a monoclonal antibody) can refer to a Y-shaped protein molecule that specifically binds to antigens, e.g., to neutralize pathogens or to mark pathogens for destruction by other immune cells. A “multi-specific” antibody can refer to an antibody that can recognize and bind to multiple antigens. An antibody can include two types of amino acid chains: heavy chains and light chains. More specifically, an antibody includes two heavy chains and two light chains. The heavy chains can determine the class and effector functions of the antibody and are typically larger than the light chains. The light chains contribute to antigen binding. Each heavy / light chain includes include a variable region and a constant region. The variable region is located at the N-terminus and can vary' greatly among different antibodies. The constant region is located at the C-terminus and is more conserved among different antibodies.
[0011] Throughout this specification, the “binding affinity” of a molecule complex can refer to the strength of the interaction between the molecules included in the molecule complex.
[0012] Throughout this specification, a “paratope” of an antibody refers to a part of the antibody that recognizes and binds to an antigen.
[0013] Throughout this specification, an “epitope” of an antigen refers to a part of the antigen that is recognized and bound by an antibody.
[0014] Throughout this specification, the “stability” of a molecule complex can refer to how resistant the molecule complex is to dissociation into its individual components or unfolding from its native structure. (Note that the stability of a molecule complex may be related to the binding affinity of the molecules in the molecule complex).
[0015] Throughout this specification, “physically synthesizing” a molecule or molecule complex can refer to creating the molecule or molecule complex in the real-world, e.g., by an appropriate biological or chemical production method.
[0016] Throughout this specification, a "multiple sequence alignment’' (MSA) refers to an alignment of three or more biological sequences (such as DNA, RNA. or protein sequences). An MSA encodes information about the conservation, covariation, and evolutionary relationships among the aligned sequences, e.g., by enabling identification of conserved regions, which may indicate functionally or structurally important areas, as well as variable regions that can provide insights into evolutionary changes and adaptations.
[0017] Throughout this specification, the terms ‘"molecule complex” and “molecular complex” are interchangeable.
[0018] Throughout this specification, the terms “single molecule” and “individual molecule” are interchangeable.
[0019] According to a first aspect there is provided a method performed by one or more computers, comprising: receiving data characterizing a plurality of molecules, generating, for each of the plurality of molecules, a respective representation of the molecule as a sequence of embeddings in a latent space using a respective single-molecule embedding neural network, jointly processing the respective representation of each of the plurality of molecules using a molecule complex embedding neural network, in accordance with values of a set of molecule complex embedding neural network parameters, to generate a representation of a molecule complex comprising the plurality of molecules; and processing the representation of the molecule complex to generate one or more predictions characterizing the molecule complex.|0020| In some implementations, the plurality of molecules comprise an antibody molecule and an antigen molecule.
[0021] In some implementations, the antibody molecule is a multi-specific antibody molecule.
[0022] In some implementations, generating, for each of the plurality of molecules, a respective representation of the molecule as a sequence of embeddings comprises, for the antibody molecule: processing data characterizing a variable region of a heavy chain of the antibody molecule using a variable heavy embedding neural network to generate a representation of the variable region of the heavy chain of the antibody molecule as respective a sequence of embeddings, processing data characterizing a variable region of a light chain of the antibody molecule using a variable light embedding neural network to generate a representation of the variable region of the light chain of the antibody molecule as a respective sequence of embeddings; and generating the representation of the antibody molecule based on the sequence of embeddings representing the variable region of the heavy chain and the sequence of embeddings representing the variable region of the light chain of the antibodymolecule. In some implementations, the variable heavy embedding neural network has been trained to perform a protein language modeling task.
[0023] In some implementations, the variable light embedding neural network has been trained to perform a protein language modeling task.
[0024] In some implementations, for one or more of the plurality of molecules, generating the representation of the molecule as a sequence of embeddings comprises: generating an initial sequence of embeddings representing the molecule using a respective single-molecule embedding neural network, generating, for each embedding in the initial sequence of embeddings, a corresponding positional embedding based at least in part on data characterizing a three-dimensional (3D) structure of the molecule, and combining each embedding in the initial sequence of embeddings representing the molecule with the corresponding positional embedding.
[0025] In some implementations, the molecule is a protein molecule and the initial sequence of embeddings representing the molecule comprises a respective embedding representing each of a plurality of amino acid residues in the molecule: and generating, for each embedding in the initial sequence of embeddings, a corresponding positional embedding based at least in part on data characterizing the 3D structure of the molecule comprises: obtaining one or more contact maps that each characterize respective 3D spatial distances between pairs of amino acid residues in the molecule; performing, for each of the one or more contact maps, an eigen- decomposition of a matrix representation of the contact map to generate a respective set of eigenvectors: and determining the positional embeddings based on the respective sets of eigenvectors.
[0026] In some implementations, wherein the molecule complex embedding neural network has been trained to perform a language modeling task.
[0027] In some implementations, training the molecule complex embedding neural network to perform a language modeling task comprises: obtaining a set of training examples, wherein each training example comprises a respective sequence of embeddings representing: (i) a portion of a molecule, or (ii) a molecule, or (iii) a molecule complex comprising a plurality’ of molecules, and training the molecule complex embedding neural network on the set of training examples to perform the language modeling task.
[0028] In some implementations, training the molecule complex embedding neural network on the set of training examples to perform the language modeling task comprises performing the training over a first training stage and a second training stage; wherein the first training stage comprises training the molecule complex embedding neural network only on training examplesthat each represent: (i) a portion of a molecule, or (ii) a molecule, and wherein the second training stage comprises training the molecule complex embedding neural network on training examples that each represent molecule complexes.
[0029] In some implementations, the set of training examples comprises a plurality of training examples that each represent a variable region of a light chain of an antibody molecule.
[0030] In some implementations, the set of training examples comprises a plurality of training examples that each represent a variable region of a heavy chain of an antibody molecule.
[0031] In some implementations, the set of training examples comprises a plurality of training examples that each jointly represent both: (i) a variable region of a light chain, and (ii) a variable region of a heavy chain, of an antibody molecule.
[0032] In some implementations, the set of training examples comprises a plurality of training examples that each represent a respective antigen molecule.
[0033] In some implementations, the set of training examples comprises a plurality of training examples that each represent a respective complex comprising an antibody molecule and an antigen molecule.
[0034] In some implementations, training the molecule complex embedding neural network on the set of training examples comprises, for each training example: generating a masked sequence of embeddings by masking one or more embeddings in the sequence of embeddings included in the training example, processing the masked sequence of embeddings using the molecule complex embedding neural network to generate a network output;, processing the network output of the molecule complex embedding neural network using a decoder neural network to generate, for each masked embedding in the masked sequence of embeddings, a prediction for an unmasked version of the embedding, and backpropagating gradients of an objective function that measures an error in the predictions for the unmasked version of the embeddings through the decoder neural network and the molecule complex embedding neural network.
[0035] In some implementations, the molecule complex embedding neural network has further been trained to perform an auxiliary task of 3D structure prediction, comprising: obtaining a set of training examples, wherein each training example comprises: (i) a sequence of embeddings representing a molecular entity, wherein the molecular entity is a portion of a molecule, or a molecule, or a molecule complex, and (ii) target 3D structure data characterizing a 3D structure of the molecular entity; training the molecule complex embedding neural network on the set of training examples, comprising, for each training example: processing the sequence of embeddings included in the training example using the molecule complexembedding neural network to generate a network output, processing the network output of the molecule complex embedding neural network using a structure prediction neural network to generate predicted 3D structure data characterizing a predicted 3D structure of the molecular entity corresponding to the training example, and backpropagating gradients of an objective function that measures an error between: (i) the target 3D structure data, and (ii) the predicted 3D structure data, through the structure prediction neural network and into the molecule complex embedding neural network.
[0036] In some implementations, for each training example, the target 3D structure data comprises a contact map.
[0037] In some implementations, the plurality of molecules comprise a protein molecule and a small molecule ligand.
[0038] In some implementations, the plurality of molecules comprise a ribonucleic acid (RNA) molecule and a protein molecule.
[0039] In some implementations, the plurality of molecules comprise an RNA molecule and a small molecule ligand.
[0040] In some implementations, the one or more predictions characterizing the molecule complex comprises a predicted binding affinity of the plurality’ of molecules in the molecule complex.
[0041] In some implementations, the one or more predictions characterizing the molecule complex comprises a predicted stability of the molecule complex.
[0042] In some implementations, the one or more predictions characterizing the molecule complex comprises a three-dimensional (3D) structure of the molecule complex.
[0043] In some implementations, the one or more predictions characterizing the molecule complex comprises, for each of one or more molecules of the plurality of molecules, data identifying a portion of the molecule that is predicted to be included in a binding interface.
[0044] In some implementations, the method further comprises performing docking of the plurality of molecules based at least in part on the data identifying the portion of the molecule that is predicted to be included in the binding interface.
[0045] In some implementations, the method further comprises selecting the plurality of molecules for physical synthesis based on the one or more predictions characterizing the molecule complex.
[0046] In some implementations, the method further comprises physically synthesizing the plurality of molecules.
[0047] In some implementations, the method further comprises, after physically synthesizing the plurality of molecules: physically synthesizing molecule complexes comprising the plurality of molecules, and performing physical experiments to test one or more properties of the molecule complexes comprising the plurality of molecules.
[0048] In another aspect, there is provided a method comprising: obtaining data identifying: (i) an antigen, and (ii) a plurality of candidate antibodies; performing, for a plurality of molecule complexes that each comprise the antigen and a respective one of the plurality of candidate antibodies, the methods described herein to generate a predicted binding affinity of the molecule complex; and determining a ranking of the plurality of candidate antibodies based on the binding affinities.
[0049] In some implementations, the method further comprises selecting one or more candidate antibodies for physical synthesis based on the ranking.
[0050] In some implementations, the method further comprises, physically synthesizing the one or more candidate antibodies selected for physical synthesis.
[0051] In some implementations, the method further comprises selecting one or more candidate antibodies for inclusion in a therapeutic based on a ranking.
[0052] In some implementations, the method further comprises physically synthesizing the therapeutic comprising the one or more selected candidate antibodies.
[0053] In some implementations, the therapeutic comprises a vaccine.|0054| In another aspect, there is provided a therapeutic that includes one or more antibodies selected according to the methods described herein.
[0055] In some implementations, the therapeutic comprises a vaccine.
[0056] In some implementations, the molecule complex embedding neural network has been jointly trained along with an epitope prediction neural network to perform an auxiliary task of antigen epitope prediction, comprising: obtaining a set of training examples, wherein each training example comprises: (i) a sequence of embeddings that jointly represents an antibody and an antigen, and (ii) target epitope data that identifies, for each amino acid in the antigen, whether the amino acid is included in an epitope that is bound by the antibody; training the molecule complex embedding neural network and the epitope prediction neural network on the set of training examples, comprising, for each training example: processing the sequence of embeddings included in the training example using the molecule complex embedding neural network to generate a network output; processing the network output of the molecule complex embedding neural network using the epitope prediction neural network to generate predicted epitope data that identifies, for each amino acid in the antigen, whether the aminoacid is predicted to be included in the epitope that is bound by the antibody; and backpropagating gradients of an objective function that measures an error between: (i) the target epitope data, and (ii) the predicted epitope data, through the epitope prediction neural network and into the molecule complex embedding neural network.
[0057] In some implementations, the network output of the molecule complex embedding neural network is an intermediate output produced by one or more hidden layers of the molecule complex embedding neural network.
[0058] In some implementations, the intermediate output produced by the one or more hidden layers of the molecule complex embedding neural network comprises embeddings representing amino acids in the antigen which are produced by a self-attention layer of the molecule complex embedding neural network.
[0059] In some implementations, the epitope prediction neural network has a graph neural network architecture.
[0060] In some implementations, processing the network output of the molecule complex embedding neural network using the epitope prediction neural network to generate the predicted epitope data comprises: instantiating a representation of the antigen as a graph of nodes and edges, wherein each node represents a respective amino acid in the antigen and wherein each edge connects a respective pair of nodes; augmenting the graph representation of the antigen to associate each node in the graph with an embedding representing a corresponding amino acid in the antigen, wherein the embedding is derived from the network output generated by the molecule complex embedding neural network; and processing the graph representation of the antigen using the epitope prediction neural network, including by a plurality of message passing layers of the epitope prediction neural network, to generate the predicted epitope data.
[0061] In some implementations, the one or more predictions characterizing the molecule complex comprises a prediction for which amino acids in the molecule complex are surface accessible.
[0062] According to another aspect, there is provided a method comprising: obtaining data identifying: (i) an antibody, and (ii) a plurality of candidate antigens; performing, for a plurality of molecule complexes that each comprise the antibody and a respective one of the plurality of candidate antigens, the methods described herein to generate a predicted binding affinity' of the molecule complex; and determining a ranking of the plurality of candidate antigens based on the binding affinities.
[0063] In some implementations, the method further comprises selecting one or more candidate antigens for physical synthesis based on the ranking.
[0064] In some implementations, the method further comprises physically synthesizing the one or more candidate antigens selected for physical synthesis.
[0065] In some implementations, the method further comprises selecting one or more candidate antigens for inclusion in a therapeutic based on the ranking.
[0066] In some implementations, the therapeutic comprises a vaccine or an antibody-based therapeutic.
[0067] In some implementations, the method further comprises: obtaining data identifying an antibody and an antigen; determining a plurality of mutated versions of the antibody, comprising: selecting a set of mutation positions in the antibody; determining, using the molecule complex embedding neural network and for each mutation position in the antibody, a score distribution over a set of possible amino acids; selecting, for each mutation position in the antibody, a set of alternative amino acids for the position using the score distribution over the set of possible amino acids; and generating each mutated version of the antibody by, for each of one or more mutation positions in the antibody, replacing an amino acid at the mutation position with an alternative amino acid selected from the set of alternative amino acids for the position; and determining a ranking of the plurality of mutated versions of the antibody.
[0068] In some implementations, determining the ranking of the plurality of mutated version of the antibody comprises: determining, for each mutated version of the antibody, a qualitymetric for the mutated version of the antibody, comprising: generating data characterizing a predicted 3D structure of a complex comprising the mutated version of the antibody and the antigen; and generating a qualify metric based on a spatial proximity of the mutated version of the antibody to the antigen in the 3D structure of the complex; and determining the ranking of the plurality of mutated versions of the antibody based on the quality metrics.
[0069] In some implementations, the method further comprises, for each mutated version of the antibody, determining an additional quality metric based on a joint probability of the mutations in the mutated version of the antibody.
[0070] In some implementations, the method further comprises selecting one or more mutated versions of the antibody for physical synthesis based at least in part on the ranking of the plurality of mutated versions of the antibody.
[0071] In some implementations, the method further comprises: obtaining data identifying an antigen and a library of antibodies; determining, for each antibody in the library- of antibodies, a corresponding predicted epitope on the antigen using the molecule complex embedding neural network.
[0072] In some implementations, the method further comprises identifying one or more of the predicted epitopes on the antigen as a novel epitope.
[0073] In another aspect, there is provided a system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the methods described herein.
[0074] In another aspect, there are provided one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the methods described herein.
[0075] Particular embodiments of the subject matter described in this specification can be implemented to realize one or more of the following advantages.
[0076] The system described in this specification can generate a representation of a molecule complex (e.g., an antibody-antigen complex), and then process that representation to generate one or more predictions characterizing the molecule complex, e.g.. predictions characterizing the three-dimensional (3D) structure of the molecule complex, the stability of the molecule complex, the binding affinity of the molecule complex, and so forth. The system can be used in any of a variety of applications, e.g., for screening a collection of possible antibodies to identify one or more antibodies appropriate for binding to an antigen, e.g., in the context of designing a vaccine, as will be described in more detail throughout this specification.
[0077] One approach to generating a representation of amolecule complex relies on leveraging multiple sequence alignments (MSAs) for the molecules included in the complex. However, generating an MSA for a molecule requires performing the steps of a computational multiple sequence alignment algorithm. Multiple sequence alignment algorithms are computationally intensive, e.g., because many such algorithms rely on dynamic programming techniques to find the optimal alignment, which involves comparing all possible pairs of sequences at each position, and because calculating optimal alignments involves complex scoring mechanisms that consider gap penalties and substitution matrices. Thus, an approach to generating a representation of a molecule complex that relies on multiple sequence alignments may be slow and computationally intensive, and thus unsuitable for applications that involve screening large numbers of possible molecule complexes.
[0078] Another approach to generating a representation of a molecule complex relies on a neural network that has been trained to perform a language modeling task. Training a neural network to perform a language modeling task involves providing the neural network with amasked representation of a biological sequence (e.g., a DNA, RNA, or protein sequence), and training the neural network to generate a representation of the biological sequence that enables reconstruction of the masked portions of the sequence. Training a neural network to perform a language modeling task can require a large volume of training examples that each include a respective biological sequence. However, training examples with biological sequences of molecule complexes are relatively scarce. A neural network that has been trained to perform a language modeling task on only the small number of available training examples with biological sequences of molecule complexes may generate representations of molecule complexes that fail to capture relevant information about the complexes and that are unsuitable for use in downstream prediction tasks.
[0079] The system described in this specification addresses these issues. In particular, to generate a representation of a molecule complex, the system generates a separate representation of each molecule in the complex using a respective neural network, referred to for convenience as a “single molecule embedding” neural network. The system can then jointly process the representations of the individual molecules in the complex using another neural network, referred to for convenience as a “molecule complex embedding neural network,” to generate a representation of the molecule complex. The modularity of the architecture, i.e., with separate single molecule embedding neural networks and a molecule complex embedding neural network, greatly increases the amount of training data available for training.|0080| More specifically, each single molecule embedding neural network can be trained, e.g., to perform a language modeling task, on biological sequence data from single molecules or portions of single molecules, and such training data is relatively plentiful (as compared to, e.g., training data for entire molecule complexes). Further, the molecule complex embedding neural network can be trained to perform a language modeling task on biological sequence data for portions of molecules, and for single molecules, and for entire molecule complexes. For instance, the molecule complex embedding neural network can be trained on biological sequence data for any combination of one or more of: the variable region of the heavy chain of antibody molecules, the variable region of the light chain of antibody molecules, and antigen molecules.
[0081] The modular architecture of the described system thus greatly increases the amount of training data available for training and enables the generation of dense and information rich molecule complex representations as compared to an approach that relies on a language modeling neural network that is only trained on biological sequences for entire molecule complexes. Further, the described system may consume significantly fewer computationalresources, e.g.. memory and computing power, as compared to an approach that relies on MSAs. In particular, after training, the system can generate a representation of a molecule complex with a single forward pass through a collection of neural networks, which may consume significantly fewer computational resources than an approach that relies on dynamically generating MSAs for some or all the molecules in the complex.
[0082] Further, the modularity of the architecture enables the pre-trained single molecule embedding neural networks to be included as a component of the neural network system without requiring additional fine-tuning, which can enable reduced consumption of computational resources. For instance, various single molecule embedding neural networks can be swapped into and out of the architecture, and the system can adapt the entire architecture to account for the variations in the single molecule embedding neural networks by parameterefficient adaptation of the molecule complex embedding neural network, w hich again enables reduced consumption of computational resources.
[0083] To generate a representation of a molecule complex, the system can generate a respective representation of each molecule in the molecule complex as a respective sequence of embeddings using single molecule embedding neural networks, and then jointly process the sequences of embeddings representing each of the individual molecules using a molecule complex embedding neural netw ork, as described above. After generating an initial sequence of embeddings representing an individual molecule, the system can generate positional embeddings based on the 3D structure of the molecule, and then combine these positional embeddings with the initial sequence of embeddings representing the molecule. The system can thus encode information that characterizes the 3D structure of the molecule into the sequence of embeddings representing the molecule. Enriching the information content of the embeddings provided to the molecule complex embedding neural network with 3D molecule structure information in this manner can reduce the amount of training data required for training the molecule complex embedding neural network and thus reduce consumption of computational resources such as memory and computing power during training. The system can obtain data characterizing the 3D structures of the individual molecules in the complex, e.g., from experimentally determined structures, or by prediction based on intermediate outputs of the single molecule embedding neural networks, or in any other appropriate way.
[0084] In addition to training the molecule complex embedding neural network to perform a main task, such as a language modeling task, the system can also train the molecule complex embedding neural network to perform an auxiliary task of 3D molecule structure prediction. For instance, the system can jointly train the molecule complex embedding neural networkalong with a structure prediction neural network. The structure prediction neural network can process a network output (e.g., an intermediate output) generated by the molecule complex embedding neural network to generate data characterizing a predicted 3D structure of the molecule complex (e.g., in the form of a contact map). Jointly training the molecule complex embedding neural network and the structure prediction neural network, e.g., by backpropagating gradients through the structure prediction neural network and into the molecule complex embedding neural network, can encourage the molecule complex embedding neural network to generate richer intermediate feature representations of molecule complexes. In particular, training the molecule complex embedding neural network to perform the auxiliary 3D molecule structure prediction task provides an additional training signal that can reduce the amount of training data required for training the molecule complex embedding neural network and can thus reduce consumption of computational resources such as memory and computing power during training.
[0085] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS|0086| FIG. 1 shows an example prediction system.
[0087] FIG. 2 is a flow diagram of an example process for generating one or more predictions that characterize a molecule complex.
[0088] FIG. 3 is a flow diagram of an example process for generating a representation of an antibody molecule as a sequence of embeddings using a single embedding neural network.
[0089] FIG. 4 is a flow diagram of an example process for generating positional embeddings based on data characterizing a predicted 3D structure of a molecule.
[0090] FIG. 5 is a flow diagram of an example process for training a molecule complex embedding neural network.
[0091] FIG. 6 is a flow diagram of an example process for training an embedding neural network to perform a language modeling task.
[0092] FIG. 7 is a flow7diagram of an example process for training a molecule embedding neural netw ork to perform a 3D molecule structure prediction task.
[0093] FIG. 8 shows an example implementation of a part of the prediction system that generates respective sequences of embeddings representing antibodies and antigens.
[0094] FIG. 9 shows an example implementation of a part of the prediction system that generates embeddings of antibody -antigen complexes.
[0095] FIG. 10 shows a flow diagram of an example process for extracting and parsing data defining molecules and / or molecule complexes (e.g., antibody, antigen, and antibody-antigen complexes) from a Cry stallographic Information File (CIF).
[0096] FIG. 11 shows a flow diagram of an example process for finding and extracting all possible single chain antigen complexes from a data file.
[0097] FIG. 12 shows a flow diagram of an example process for filtering data defining antibody-antigen complexes to remove contacts due to crystal packing issues.
[0098] FIG. 13 shows an example process deduplicating a set of data defining antigens, variable light chains, variable heavy chains, variable heavy - variable light chain pairs, variable heavy chain - antigen pairs, and variable heavy chain - variable light chain - antigen complexes.
[0099] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0100] FIG. 1 shows an example prediction system 100. The prediction system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0101] The prediction system includes one or more single molecule embedding neural networks 106a-n, a molecule complex embedding neural network 110, and a prediction generator 114. The prediction system 100 processes input data 102 that characterizes multiple (e.g., 2. 5, 8, etc.) molecules to generate one or more predictions 116 characterizing a complex that includes the molecules defined in the input data.
[0102] The molecules in the input data can include one or more of: protein molecules (e.g., antigens, antibodies, NANOBODIES® etc.), ribonucleic acid (RNA) molecules, deoxyribonucleic acid (DNA) molecules, small molecules (e.g., small molecule ligands), or combinations thereof. For instance, in one example, the input data can include data identifying an antibody molecule and an antigen molecule. As another example, the input data can include data identifying a protein molecule and a small molecule ligand. As another example, the input data can include data identifying an RNA molecule and a protein molecule. As another example, the input data can include data identifying an RNA molecule and a small molecule ligand. As another example, the input data can include data identifying a DNA molecule and aprotein molecule. As another example, the input data can include data identifying a DNA molecule and a small molecule ligand. As another example, the input data can include data identifying a first protein molecule and a second protein molecule. As another example, the input data can include data identify ing N protein molecules, where N is a positive integer value that is greater than two.
[0103] The input data 102 includes respective molecule data 104a-n for each molecule in a set of molecules. The molecule data for a molecule can include any appropriate data characterizing (identifying) the molecule. A few examples of molecule data characterizing molecules are described next.
[0104] In one example, the input data 102 can include molecule data 104 characterizing a protein molecule, including data identifying one or more amino acid sequences of the protein molecule. In particular, for each position in each amino acid sequence of the protein molecule, the molecule data 104 can identify' the amino acid that occupies the position.
[0105] In another example, the input data 102 can include molecule data 104 characterizing an RNA molecule, including data identifying one or more nucleotide sequences of the RNA molecule. In particular, for each position in each nucleotide sequence of the RNA molecule, the molecule data 104 can identify’ the nucleotide that occupies the position.
[0106] In another example, the input data 102 can include molecule data 104 characterizing a DNA molecule, including data identifying one or more nucleotide sequences of the DNA molecule. In particular, for each position in each nucleotide sequence of the DNA molecule, the molecule data 104 can identify the nucleotide that occupies the position.
[0107] In another example, the input data 102 can include molecule data 104 characterizing a small molecule, including data identify ing a chemical structure of the small molecule, e.g., the arrangement of atoms and bonds in the ligand, the atom types in the small molecule, the chirality of the bonds in the small molecule, any functional groups (e.g., hydroxyl groups, amino groups, carboxyl groups, and so forth) included in the small molecule, and so forth. For instance, the molecule data 104 characterizing the small molecule can include a textual representation the small molecule, e.g., as a simplified molecular-input line-entry system (SMILES) string.
[0108] Each single molecule embedding neural network 106 is associated with a respective type of molecule and is configured to process a network input that includes molecule data characterizing a molecule of the associated type to generate a molecule embedding 108 of the molecule. A molecule embedding 108 of a molecule is a representation of the molecule as a sequence of embeddings in a latent space.
[0109] The number of embeddings in the sequence of embeddings representing the molecule can depend on the type of molecule. For instance, for a protein molecule, the number of embeddings in the sequence of embeddings representing the protein molecule can be equal to the length of the protein molecule, e.g., the number of amino acids in the protein molecule. As another example, for an RNA molecule, the number of embeddings in the sequence of embeddings representing the protein molecule can be equal to the length of the RNA molecule, e.g., the number of nucleotides in the RNA molecule. As another example, for a DNA molecule, the number of embeddings in the sequence of embeddings representing the DNA molecule can be equal to a length of the DNA molecule, e.g., the number of nucleotides in the DNA molecule. As another example, for a small molecule, the number of embeddings in the sequence of embeddings representing the small molecule can be equal to a number of atoms in the small molecule, or equal to the number of heavy atoms in the small molecule.
[0110] The system 100 can include any appropriate number of single molecule embedding neural networks. For instance, the system 100 can include a single molecule embedding neural network that is configured to process data characterizing protein molecules, or a single molecule embedding neural network that is configured to process data characterizing RNA molecules, or a single molecule embedding neural network that is configured to process data characterizing DNA molecules, or a single molecule embedding neural network that is configured to process data characterizing small molecules, or a single molecule embedding neural network that is configured to process data characterizing antigen molecules, or a single molecule embedding neural network that is configured to process data characterizing antibody molecules, or any combination thereof.
[0111] Each single molecule embedding neural network 106 can have any appropriate neural network architecture that enables the single molecule embedding neural network 106 to perform its described functions, e.g., processing molecule data characterizing a molecule to generate a representation of the molecule as a sequence of embeddings in a latent space. In particular, each single molecule embedding neural network 106 can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, message passing layers, attention layers, etc.) in any appropriate number (e.g.. 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0112] The system can train each single molecule embedding neural network in any appropriate way. For instance, the system can train one or more of the single molecule embedding neural networks to perform a language modeling task. An example process fortraining a neural network to perform a language modeling task is described with reference to FIG. 6.
[0113] In a particular example, the system 100 can include a single molecule embedding neural network that is configured to process molecule data identifying a protein molecule, and the single molecule embedding neural network can be based on the ESM-2 neural network described in Lin et al., “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science. 16 March 2023, Vol 379, Issue 6637, pp. 1123-1130.
[0114] In another particular example, the system 100 can include a single molecule embedding neural network that is configured to process molecule data identify ing an antibody, and the single molecule embedding neural network can be based on the AntiBERTy neural network described in Ruffolo et al., “Deciphering antibody affinity maturation with language models and weakly supervised learning,” arXiv:2112.07782, or on the AbLang neural network described in Olsen et al., “AbLang: an antibody language model for completing antibody sequences,” Bioinformatics Advances, Volume 2, Issue 1, 17 June 2022.
[0115] Each single molecule embedding neural network included in the system 100 can be different from each other single molecule embedding neural network included in the system 100, e.g., as a result of having a different neural network architecture, or being trained on a different set of training data, or both.
[0116] The molecule complex embedding neural network 110 jointly processes the respective molecule embedding 108 of each molecule in the input data 102 to generate a molecule complex embedding 1 12 that is a representation of a complex that includes each molecule in the input data 102. More specifically, the network input to the molecule complex embedding neural network 110 can include a sequence of embeddings generated (at least in part) by concatenating the respective sequence of embeddings from each molecule embedding 108. The molecule complex embedding 112 generated by the molecule complex embedding neural network 110 can include a sequence of embeddings, e.g., an updated version of the sequence of embeddings provided as an input to the molecule complex embedding neural network 110.
[0117] The molecule complex embedding neural netw ork 110 can have any appropriate neural network architecture that enables the molecule complex embedding neural network 1 12 to perform its described functions, e.g., processing multiple molecule embeddings 108 to generate a molecule complex embedding 112. In particular, the molecule complex embedding neural network 112 can include any appropriate types of neural netw ork layers (e g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g., 5 layers, 10layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0118] In a particular example, the molecule complex embedding neural network 110 can include a sequence of self-attention neural network layers. Each self-attention neural network layer is configured to receive a set of embeddings, to update the set on embeddings at least in part by applying one or more self-attention operations to the set of embeddings, and then to output the updated set of embeddings. The first self-attention layer can be configured to receive a layer input that includes the network input to the molecule complex embedding neural network 110, and each subsequent self-attention layer (i.e., after the first self-attention layer) can be configured to receive a set of embeddings generated by a preceding self-attention layer in the sequence of self-attention layers.
[0119] The system 100 can train the molecule complex embedding neural network 1 10 in any appropriate way. An example process for training the molecule complex embedding neural network to perform a language modeling task is described with reference to FIG. 6, and an example process for training the molecule complex embedding neural network to perform an auxiliary task of 3D structure prediction is described with reference to FIG. 7.
[0120] In some implementations, the system 100 trains the molecule complex embedding neural network 110 separately from each of the single molecule embedding neural networks 106. That is, the system 100 can initially train each single molecule embedding neural network, e.g., to perform a language modeling task (as described above). The system 100 can then freeze the parameter values of the single molecule embedding neural networks 106, i.e., by fixing the trained values of the neural network parameters of the single molecule embedding neural networks 106 as static values, and then train the molecule complex embedding neural network 110 with reference to the molecule embeddings generated by the trained single molecule embedding neural networks 106.
[0121] The prediction generator 114 processes the molecule complex embedding 112 to generate one or more predictions 116 characterizing the molecule complex. The prediction generator 114 can be implemented as a machine learning model having a set of machine learning model parameters. For instance, the prediction generator 114 can be implemented as a neural network, or as a random forest, or as a support vector machine, or linear regression model, and so forth.
[0122] The prediction generator 114 can, in some cases, process only a proper subset (i.e., less than all) of the molecule complex embedding 112 in order to generate a prediction 116 characterizing the molecule complex. For instance, the molecule complex embedding neuralnetwork can be configured to generate a molecule complex embedding 112 that include a property-specific embedding that is associated with a particular molecule complex property. In this example, the prediction generator 114 can be configured to process only the propertyspecific embedding (e.g., rather than the entire molecule complex embedding 112) in order to generate the corresponding predicted property of the molecule complex.
[0123] The system 100 can train the prediction generator 114 on a set of training examples by a machine learning training technique. Each training example corresponds to a respective training molecule complex and includes: (i) a training input that includes a training molecule complex embedding 112 generated by the molecule complex embedding neural network 110, and (ii) a target prediction characterizing the training molecule complex. For each training example, the system 100 trains the prediction generator 114 to optimize an objective function that measures discrepancy between: (i) the target prediction specified by the training example, and (ii) a prediction generated by the prediction generator 114 by processing the training input specified by the training example. The objective function can be, e.g., a mean squared error objective function, or a cross-entropy objective function, or any other appropriate objective function.
[0124] A few examples of possible predictions 116 that can be generated by the prediction generator 114 are described next.
[0125] In one example, the prediction generator 114 can generate a prediction 116 defining a predicted binding affinity of the molecules included in the molecule complex.
[0126] In another example, the prediction generator 1 14 can generate a prediction 1 1 defining a predicted stability of the molecule complex.
[0127] In another example, the prediction generator 114 can generate a prediction 116 for which amino acids in the molecule complex are surface accessible. For instance, for a molecule complex that includes an antibody and an antigen, the prediction 116 can characterize which amino acids in the antigen are surface accessible.
[0128] In another example, the prediction generator 114 can generate a prediction 116 characterizing a 3D structure of the molecule complex, e.g., a prediction 116 defining a respective 3D spatial position of each atom in the molecule complex. For instance, the prediction generator 114 can generate a predicted 3D structure of the molecule complex by processing the molecule complex embedding 112 using the “folding trunk” and “structure module” described in Lin et al., “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science. 16 March 2023, Vol 379, Issue 6637, pp. 1123- 1130.
[0129] In another example, the prediction generator can generate a prediction 116 that, for each of one or more molecules included in the molecule complex, identifies a portion of the molecule that is predicted to be included in a binding interface. The binding interface prediction can subsequently be used to perform docking of the plurality of molecules included in the molecule complex, e.g., by initializing the docking such that the predicted binding interfaces are already in proximity.
[0130] For instance, for a molecule complex that includes a protein, the prediction generator can generate a prediction that includes, for each amino acid in the protein, a prediction for whether any atom in the amino acid is within a threshold distance (e.g., 8 Angstroms) of an atom in a different molecule in the complex.
[0131] For instance, for a molecule complex that includes an RNA molecule, the prediction generator can generate a prediction that includes, for each nucleotide in the RNA molecule, a prediction for whether any atom in the nucleotide is within a threshold distance (e.g., 8 Angstroms) of an atom in a different molecule in the complex.
[0132] For instance, for a molecule complex that includes a DNA molecule, the prediction generator can generate a prediction that includes, for each nucleotide in the DNA molecule, a prediction for whether the atom is within a threshold distance (e.g., 8 Angstroms) of an atom in a different molecule in the complex.
[0133] For instance, for a molecule complex that includes a small molecule, the prediction generator can generate a prediction that includes, for each atom in the small molecule, a prediction for whether the atom is within a threshold distance (e.g., 8 Angstroms) of an atom in a different molecule in the complex.
[0134] For instance, for a molecule complex that includes an antibody and an antigen, the prediction generator can generate a prediction that includes: (i) for each atom in the antibody, a prediction for whether the atom is included in the paratope of the antibody, and (ii) for each atom in the antigen, a prediction for whether the atom is included in the epitope of the antigen.
[0135] The system 100 can train the prediction generator 114 on a set of training examples by a machine learning training technique, as described above. In some cases, the system 100 trains the prediction generator 114 jointly with the molecule complex embedding neural network 110. For instance, the prediction generator 114 can be implemented as a property prediction neural network that is configured to process at least a portion of the molecule complex embedding 112 to generate a predicted property of the molecule complex. The predicted property can be, e.g., a predicted binding affinity, or a predicted stability, and so forth. As part of training the property prediction neural network to optimize an objective function that measures an accuracyof property predictions generated by the property prediction neural network, the system 100 can backpropagate gradients of the objective function through the property prediction neural network and into the molecule complex embedding neural network 1 10. In other cases, the system 100 trains the prediction generator 114 by a training procedure that is separate from and independent of the training of the molecule complex embedding neural network 110, e.g., in implementations where the prediction generator 114 is implemented as a non-differentiable model.
[0136] Predictions 116 generated by the system 100 can be used in any of a variety of downstream applications, a few examples of which are described next.
[0137] In one example, the set of molecules included in the molecule complex can be selected for physical synthesis based at least in part on predictions 116 generated by the system 100 for the molecule complex. For instance, the set of molecules included in the molecule complex can be selected for physical synthesis based at least in part on a predicted binding affinity' or a predicted stability of the complex satisfying (e.g., exceeding) a threshold value. The set of molecules can be physically synthesized using an appropriate synthesis technique, and physical experiments can be performed to test the properties of the molecule complex that includes the physically synthesized molecules. For instance, physical experiments can be performed to assess the structure of the complex (e.g., x-ray cry stallography experiments), or the binding affinity of the complex, or the thermal stability of the complex, and so forth.|0138| In another example, for each candidate antibody in a set of candidate antibodies, the system 100 can generate a predicted binding affinity of a molecule complex that includes: (i) the candidate antibody, and (ii) an antigen. The candidate antibodies can then be ranked based on their predicted binding affinity for the antigen. One or more of the candidate antibodies can then be selected for physical synthesis or for inclusion in an antibody-based therapy based at least in part on the ranking, e.g., each candidate antibody having a binding affinity for the antigen that exceeds a threshold can be selected, or a predefined number of candidate antibodies having the highest binding affinity for the antigen can be selected. The selected candidate antibodies can then be physically synthesized, or the antibody-based therapy that includes the selected candidate antibodies can be physically synthesized.
[0139] In another example, for each candidate antigen in a set of candidate antigens, the system 100 can generate a predicted binding affinity7of a molecule complex that includes: (i) the candidate antigen, and (ii) an antibody. The candidate antigens can then be ranked based on their predicted binding affinity for the antibody. One or more of the candidate antigens can then be selected for physical synthesis or for inclusion in a vaccine based at least in part on theranking, e.g., each candidate antigen having a binding affinity for the antibody that exceeds a threshold can be selected, or a predefined number of candidate antigens having the highest binding affinity for the antibody can be selected. The selected candidate antigens can then be physically synthesized, or a vaccine that includes the selected candidate antigens can be physically synthesized. The system can thus perform antibody-guided vaccine design by reverse engineering a known potent antibody to determine an antigen that will train the body to produce an effective immune response.
[0140] The antibody-based therapy or vaccine can target, e.g., a viral pathogen (e.g., influenza virus, or human papillomavirus, or hepatitis A virus, or hepatitis B virus, measles virus, or mumps virus, or rubella virus, or rabies virus, or SARS-CoV-2 virus), or bacterial pathogens (e.g., streptococcus pneumoniae, or Neisseria meningitidis, or mycobacterium tuberculosis), or a parasite (e.g., plasmodium falciparum), or (in the case of antibody-based therapies) tumor antigens for cancer treatment.
[0141] FIG. 2 is a flow diagram of an example process 200 for generating one or more predictions that characterize a molecule complex. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200.
[0142] The system receives data characterizing a set of multiple molecules (202). The set of molecules can include, e.g., one or more of: protein molecules, RNA molecules, DNA molecules, small molecules, and so forth. The system can receive the data identifying the set of multiple molecules, e.g., from a user or from an upstream system, e.g., by way of a graphical user interface (GUI) or by way of an application programming interface (API) made available by the system.
[0143] The system generates, for each molecule in the set of molecules, a respective representation of the molecule as a sequence of embeddings in a latent space using a respective single molecule embedding neural network that is configured to process that type of molecule (204). An example process for generating a representation of an antibody molecule as a sequence of embeddings is described with reference to FIG. 3.
[0144] In some implementations, to generate a representation of a molecule as a sequence of embeddings, the system processes data characterizing the molecule using a single molecule embedding neural network to generate an initial sequence of embeddings representing the molecule. The system also generates, for each embedding in the initial sequence of embeddings, a corresponding positional embedding based at least in part on data characterizinga predicted 3D structure of the molecule. The system then combines each embedding in the initial sequence of embeddings representing the molecule with the corresponding positional embedding, e.g., by element-wise summing, element-wise averaging, element-wise multiplying, or concatenating the embedding in the initial sequence of embeddings and the corresponding positional embedding. The system can thus incorporate data characterizing the 3D structure of the molecule in the individual embeddings representing the molecule. An example process for generating positional embeddings based on data characterizing a predicted 3D structure of a molecule is described in detail with reference to FIG. 4.
[0145] Optionally, the system projects the respective sequence of embeddings representing each molecule in the set of molecules into a common latent space (206). More specifically, each molecule is represented by a respective sequence of embeddings where each embedding in the sequence of embeddings has a respective dimensionality (e.g., defined by the number of entries in a tensor defining the embedding). In some cases, the embeddings in a sequence of embeddings representing a first molecule can have a different dimensionality7than the embeddings in a sequence of embeddings representing a second molecule. A variation in dimensionality among the embeddings can be problematic, e.g., when the sequences of embeddings representing the respective molecules are concatenated to form the network input to the molecule complex embedding neural network (e.g., at step 208). To address this issue, for each molecule, the system can process the sequence of embeddings representing the molecule using a respective projection neural network that operates individually on each embedding to project the embedding into the common latent space. Projecting the embeddings representing each molecule in the set of molecules into the common latent space causes the embeddings representing each of the molecules to have a same dimensionality.
[0146] Optionally, for each molecule in the set of molecules, the system can combine (e.g., sum or concatenate) a respective molecule identifier embedding with each embedding in the sequence of embeddings representing the molecule. In some cases, the molecule identifier embedding can define a type of the molecule, e.g., whether the molecule is a protein molecule, or a DNA molecule, or an RNA molecule, or a small molecule. In some cases, particularly when at least two of the molecules are of the same type (e.g.. the set of molecules includes at least two protein molecules), the molecule identifier embedding can also serve to distinguish between different instances of molecules of the same t pe. In some cases, the system can generate multiple molecule identifier embeddings for a single molecule and then combine each molecule identifier with a respective subset of embeddings in the sequence of embeddings representing the molecule. For instance, for a sequence of embeddings representing anantibody, the system can combine a first molecule identifier embedding with the embeddings representing the variable region of the light chain, and a second molecule identifier embedding with the embeddings representing the variable region of the heavy chain.
[0147] The system jointly processes the respective representation of each molecule in the set of molecules using a molecule complex embedding neural network, in accordance with values of a set of molecule complex embedding neural network parameters, to generate a representation of a molecule complex that includes the set of molecules (208).
[0148] The system processes the representation of the molecule complex to generate one or more predictions characterizing the molecule complex (210). For instance, the system can generate one or more of: a predicted binding affinity of the molecules in the molecule complex, or a predicted stability of the molecule complex, or a predicted 3D structure of the molecule complex, or data identifying predicted binding interfaces among the molecules in the molecule complex.
[0149] The system can provide the predictions characterizing the molecule complex, e.g., for presentation on a display of a user device, or for storage in a memory, or for transmission over a data communications network.
[0150] FIG. 3 is a flow diagram of an example process 300 for generating a representation of an antibody molecule as a sequence of embeddings using a single embedding neural network. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g.. the prediction system 100 of FIG. 1 , appropriately programmed in accordance with this specification, can perform the process 300.
[0151] The system receives data identifying: (i) an amino acid sequence of a variable region of a heavy chain of the antibody molecule, and (ii) an amino acid sequence of a variable region of a light chain of the antibody molecule (302).
[0152] The system processes the data identifying the amino acid sequence of the variable region of the heavy chain of the antibody molecule using a variable heavy embedding neural network that is included in the single embedding neural network to generate a representation of the variable region of the heavy chain as a sequence of embeddings (304). More specifically, the network input to the variable heavy embedding neural network can include a sequence of embeddings with a respective embedding (e.g., one-hot embedding) representing the respective amino acid at each position in the variable region of the heavy chain. The netw ork output generated by the variable heavy embedding neural network can similarly include a respectiveembedding of the amino acid at each position in the amino acid sequence of the variable region of the heavy chain.
[0153] The system can train the variable heavy embedding neural network to perform a language modeling task. An example process for training a neural network to perform a language modeling task is described with reference to FIG. 6. In some cases, the system trains the variable heavy embedding neural network substantially or exclusively on training data that defines amino acid sequences of variable regions of heavy chains of antibody molecules, e.g.. such that the variable heavy embedding neural network is specialized for processing variable heavy chains. In other cases, the system trains the variable heavy embedding neural network on training data that includes at least some amino acid sequences other than those from variable regions of heavy chains.
[0154] The variable heavy embedding neural network can have any appropriate neural network architecture that enables the variable heavy embedding neural network to perform its described functions. In particular, the variable heavy embedding neural network can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g.. 5 layers. 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0155] In a particular example, the variable heavy embedding neural network can be based on the AntiBERTy neural network described in Ruffolo et al., "Deciphering antibody affinity maturation with language models and weakly supervised learning / ’ arXiv:2112.07782, or on the AbLang neural network described in Olsen et al., “AbLang: an antibody language model for completing antibody sequences,” Bioinformatics Advances, Volume 2, Issue 1, 17 June 2022.
[0156] The system processes the data identifying the amino acid sequence of the variable region of the light chain of the antibody molecule using a variable light embedding neural network that is included in the single embedding neural network to generate a representation of the variable region of the light chain as a sequence of embeddings (306). More specifically, the network input to the variable light embedding neural network can include a sequence of embeddings with a respective embedding (e.g., one-hot embedding) representing the respective amino acid at each position in the variable region of the light chain. The netw ork output generated by the variable light embedding neural network can similarly include a respective embedding of the amino acid at each position in the amino acid sequence of the variable region of the light chain.
[0157] The system can train the variable light embedding neural network to perform a language modeling task. An example process for training a neural network to perform a language modeling task is described with reference to FIG. 6. In some cases, the system trains the variable light embedding neural network substantially or exclusively on training data that defines amino acid sequences of variable regions of light chains of antibody molecules, e.g., such that the variable light embedding neural network is specialized for processing variable light chains. In other cases, the system trains the variable light embedding neural network on training data that includes at least some amino acid sequences other than those from variable regions of light chains.
[0158] The variable light embedding neural network can have any appropriate neural network architecture that enables the variable light embedding neural network to perform its described functions. In particular, the variable light embedding neural network can include any appropriate types of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) in any appropriate number (e.g., 5 layers, 10 layers, or 20 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0159] In a particular example, the variable light embedding neural network can be based on the AntiBERTy neural network described in Ruffolo et al., “Deciphering antibody affinity maturation with language models and weakly supervised learning,” arXiv:2112.07782, or on the AbLang neural network described in Olsen et al.. “AbLang: an antibody language model for completing antibody sequences.” Bioinformati.es Advances, Volume 2, Issue 1, 17 June 2022.
[0160] The system generates the representation of the antibody molecule based on the sequence of embeddings representing the variable region of the heavy chain and the sequence of embeddings representing the variable region of the light chain of the antibody molecule (308). The system can, for example, concatenate the sequence of embeddings representing the variable region of the heavy chain and the sequence of embeddings representing the variable region of the light chain of the antibody molecule.
[0161] FIG. 4 is a flow diagram of an example process 400 for generating positional embeddings based on data characterizing a predicted 3D structure of a molecule. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
[0162] The description of the process 400 will reference a sequence of “elements’' of the molecule. For a protein, the sequence of elements can refer to a sequence of amino acids. For an RNA molecule or a DNA molecule, the sequence of elements can refer to a sequence of nucleotides.
[0163] The system obtains one or more contact maps for the molecule (402). Each contact map characterizes respective 3D spatial distances between pairs of elements in the sequence of elements of the molecule. More specifically, each contact map can define, for each of multiple pairs of elements from the sequence of elements of the molecule, whether the pair of elements are separated by less than a threshold distance in the 3D structure of the molecule. The threshold distance can be any appropriate threshold distance, e.g., 4 Angstroms, or 8 Angstroms, or 12 Angstroms.
[0164] A pair of elements (e.g., amino acids or nucleotides) in a molecule can be referred to as being separated by less than a threshold distance, e.g., if at least one atom in the first element is separated by less than the threshold distance from at least one atom in the second element, or if a designated atom in the first element is separated by less than the threshold distance from a designated atom in the second element. For instance, for an element that is an amino acid, the “designated atom” in the element can be, e.g., an appropriate backbone atom in the amino acid, e.g., the alpha carbon atom.
[0165] In some cases, the system can obtain a single contact map for the entire molecule, i.e., a contact map that, for each pair of elements in the sequence of elements of the molecule, defines whether the pair of elements are separated by less than the threshold distance. In other cases, the system can obtain multiple contact maps for a molecule, where each of the contact maps is associated with a respective subset of the sequence of elements of the molecule. For instance, for an antibody molecule, the system can obtain a first contact map for the variable region of the heavy chain of the antibody, and a second contact map for the variable region of the light chain of the antibody.
[0166] The system can obtain the one or more contact maps for the molecule in any appropriate way. For instance, the system can generate each contact map by providing an intermediate output generated by a single molecule embedding neural network by processing data characterizing the molecule as an input to a structure prediction neural network. The structure prediction neural network can process the intermediate output generated by the single molecule embedding neural network to generate the contact map for the molecule.
[0167] For instance, in one example, the single molecule embedding neural network can include a sequence of self-attention neural network layers. As part of processing datacharacterizing the molecule, each self-attention neural network layer can generate an array of attention weights as part of applying a self-attention operation to a sequence of embeddings processed by the self-attention neural network layer. In this example, the system can extract the respective array of attention weights generated by one or more of the self-attention neural network layers, and provide these for processing by the structure prediction neural network in order to generate the contact map. An example implementation of a single molecule prediction neural network (and techniques for predicting contact maps from arrays of attention weights) is described with reference to: Roshan Rao et al., “Transformer protein language models are unsupervised structure learners,” International Conference on Learning Representations, October 2020.
[0168] The system can generate multiple contact maps for respective portions of a molecule by processing respective intermediate outputs generated by the single embedding neural network. For instance, FIG. 3 describes an example architecture of a single embedding neural network for generating a molecule embedding of an antibody molecule, where the single embedding neural network includes: (i) a variable heavy embedding neural network, and (ii) a variable light embedding neural network. In this example, the system can generate a contact map for the variable region of the heavy chain of the antibody by processing an intermediate output generated by the variable heavy embedding neural network. Similarly, the system can generate a contact map for the variable region of the light chain of the antibody by processing an intermediate output generated by the variable light embedding neural network.
[0169] In some cases, the system can obtain contact maps that have been determined by physical experiments, e.g., as an alternative to or in combination with generating contact maps using computational prediction approaches (as described above).
[0170] The system generates, for each of the contact maps, an eigen-decomposition of a matrix representation of the contact map to generate a set of eigenvectors of the matrix representation of the contact map (404). (The phrase “contact map” as used here may be understood to refer to a representation of the contact map as an adjacency matrix, e.g., where the molecule is represented as an undirected graph without self-loops. Further, the system can convert the adjacency matrix to a normalized Laplacian matrix). More specifically, each contact map can be represented as a two-dimensional (2D) array of binary values, where each entry (i,y) of the matrix holds a binary value indicating whether element i and element j in the sequence of elements of the molecule are separated by less than the threshold distance. The system can perform the eigen-decomposition of a contact map matrix using any appropriate numerical technique, e g., the QR algorithm, the Jacobi method, or singular value decomposition (SVD).Each eigenvector of each contact map can correspond to a respective element in the sequence of elements of the molecule.
[0171] The system determines the positional embeddings based on the eigenvectors of the one or more contact maps (406). For instance, the system can identify each eigenvector as the positional embedding for a respective element in the sequence of elements of the molecule. Positional embeddings generated in this manner may be referred to as ‘"spectral embeddings.”
[0172] Optionally, as part of performing the steps of the process 400, the system can apply various additional transformation or projection operations to the matrix representations of the contact maps (adjacency matrices) and / or the eigenvectors of the contact maps (adjacency matrices). Example techniques for using spectral embeddings as positional encodings are described with reference to: Dwivedi et al., ‘"A generalization of transformer networks to graphs,” arXiv:2012.09699v2, 24 January 2021.
[0173] The system can use the positional embeddings for the elements (e.g., amino acids) in a molecule to augment an initial sequence of embeddings representing the molecule, e.g., as described above with reference to FIG. 2. In this manner, the system can “infuse” structural information characterizing a molecule into the sequence of embeddings representing the molecule. Infusing structural information into embeddings representing molecules facilitates using the embeddings in downstream processes that generate predictions characterizing molecules. The 3D structure of a molecule influences and is integrally related to many properties of the molecule, and therefore incorporating structural information into embeddings representing molecules can facilitate using those embeddings to generate downstream predictions. In particular, infusing structural information into molecule embeddings can reduce an amount of training data and the number of training iterations required for training a machine learning model - e.g., the molecule complex embedding neural network - to process those embeddings to perform prediction tasks. Further, infusing structural information into the molecule embeddings can enable a machine learning model operating on those embeddings - e.g., the molecule complex embedding neural network - to be implemented using a less complex architecture than would otherwise be required. Thus infusing structural information into molecule embeddings can reduce consumption of computational resources, including both compute and memory resources, both at training and during inference.
[0174] A technical issue that arises when incorporating structural information into molecule embeddings is that the 3D structure of a molecule, expressed in terms of 3D atomic coordinates, is heavily dependent upon arbitrary translation and rotation parameters. More specifically, the 3D structure of a molecule is the same regardless of the particular location and rotation of themolecule in 3D space, but most parametrizations of the 3D structure have the undesirable property of being dependent upon arbitrary’ choices of translation and rotation. An effect of this dependence upon translation and rotation is that infusing 3D structural information into molecule embeddings may cause a machine learning model trained on these embeddings to overfit the training data and fail to generalize to previously unseen molecules.
[0175] The system described in this specification addresses this issue. In particular, the system generates positional embeddings representing the 3D structural information for a molecule that are invariant to rotations and translations of the 3D structure of the molecule. More specifically, the system leverages contact maps which characterize pairwise distance between elements (e.g., amino acids) in the molecule, and 3D structural information expressed in this manner is rotation and translation invariant. Thus, the manner in which the system infuses 3D structural information into molecule embeddings addresses technical issues such as overfitting that arise during machine learning training.
[0176] FIG. 5 is a flow diagram of an example process 500 for training a molecule complex embedding neural network. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.
[0177] The system obtains a set of training examples for training the molecule complex embedding neural network (502). Each training example includes a respective sequence of embeddings representing: (i) a portion of a molecule, or (ii) a molecule, or (iii) a molecule complex that includes multiple molecules.
[0178] In some implementations, the set of training examples includes one or more training examples representing a portion of an antibody molecule, e.g.: (i) a variable region of a light chain of an antibody molecule, or (ii) a variable region of a heavy chain of an antibody molecule, or (iii) both a variable region of a light chain and a variable region of a heavy chain of an antibody molecule.
[0179] In some implementations, the set of training examples includes one or more training examples representing respective antigen molecules.
[0180] In some implementations, the set of training examples includes one or more training examples representing respective complexes that each include an antigen molecule and an antibody molecule.
[0181] In some implementations, the set of training examples includes one or more training examples representing respective complexes that each include an antigen molecule and a NANOBODY® molecule.
[0182] In some implementations, the set of training examples includes one or more training examples that each represent a respective protein molecule.
[0183] In some implementations, the set of training examples includes one or more training examples that each represent a respective DNA molecule.
[0184] In some implementations, the set of training examples includes one or more training examples that each represent a respective RNA molecules.
[0185] In some implementations, the set of training examples includes one or more training examples that each represent a respective small molecule.
[0186] In some implementations, the set of training examples includes one or more training examples that each represent a complex that includes a protein and a DNA molecule.
[0187] In some implementations, the set of training examples includes one or more training examples that each represent a complex that includes a protein and an RNA molecule.
[0188] In some implementations, the set of training examples includes one or more training examples that each represent a complex that includes a protein and a small molecule.
[0189] For each training example that represents a portion of a molecule, the system generates the sequence of embeddings included in the training example by processing data representing the portion of the molecule using a single embedding neural network to generate the sequence of embeddings representing the portion of the molecule. For instance, the system can generate a sequence of embeddings representing a variable region of a heavy chain of an antibody molecule by processing data defining the variable region of the heavy chain of the antibody molecule using a variable heavy embedding neural network, as described with reference to FIG. 3. As another example, the system can generate a sequence of embeddings representing a variable region of a light chain of an antibody molecule by processing data defining the variable region of the light chain of the antibody molecule using a variable light embedding neural network, as described with reference to FIG. 3.
[0190] For each training example that represents a molecule, the system generates the sequence of embeddings included in the training example by processing data representing the molecule using a respective single embedding neural network.
[0191] For each training example that represents a molecule complex, the system generates the sequence of embeddings included in the training example by, for each molecule in the molecule complex, processing data representing the molecule using a respective single embedding neuralnetwork to generate a sequence of embeddings representing the molecule. The system then concatenates the respective sequence of embeddings representing each of the molecules in the complex to generate the sequence of embeddings that is included in the training example.
[0192] After obtaining the set of training examples, the system trains the molecule complex embedding neural network on the set of training examples. For instance, the system can train the molecule complex embedding neural network to perform a language modeling task, and optionally, an auxiliary task of 3D molecule structure prediction. An example process for training a neural network to perform a language modeling task is described in detail with reference to FIG. 6. An example process for training the molecule complex embedding neural network to perform a 3D molecule structure prediction task is described with reference to FIG. 7.
[0193] For each training example, the system can train the molecule complex embedding neural network on the training example to perform the language modeling task, or the 3D molecule structure prediction task, or both. In some cases, for multiple training examples, the system trains the molecule complex embedding neural network on the training example to perform the language modeling task but not the 3D molecule structure prediction task (e.g., in cases where the 3D structure of the molecule or molecule complex is unknown). In some cases, for multiple training examples, the system trains the molecule complex embedding neural network on the training example to perform the 3D molecule structure prediction task but not the language modeling task. In some cases, for multiple training examples, the system trains the molecule the molecule complex embedding neural network on the training example to perform both the language modeling task and the 3D molecule structure prediction task.
[0194] Optionally, for one or more training examples, the system can train the molecule complex embedding neural network on the training example to perform one or more additional auxiliary tasks, e.g., binding affinity prediction, heavy and light chain mis-pairing prediction, and so forth.
[0195] The system can train the molecule complex embedding neural network on the set of training examples by a multi-stage training procedure, including a first training stage (step 504, described below) and a second training stage (step 506. described below).
[0196] At the first training stage, the system trains the molecule complex embedding neural network only on training examples that each represent: (i) a portion of a molecule, or (ii) a molecule (504). In particular, at the first training stage, the system can (optionally) refrain from training the molecule complex embedding neural network on training examples that represent molecule complexes.
[0197] At the second training stage, the system trains the molecule complex embedding neural network on training examples that each represent a molecule complex (506). Optionally, at the second training stage, the system can train the molecule complex embedding neural network on both: (i) training examples that represent molecule complexes, and (ii) training examples that represent molecules or portions of molecules.
[0198] FIG. 6 is a flow diagram of an example process 600 for training an embedding neural network to perform a language modeling task. For convenience, the process 600 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 600.
[0199] The embedding neural network that is trained by the process 600 can be. e.g., a single embedding neural network or a molecule complex embedding neural network.
[0200] The system obtains a training example that includes a sequence of embeddings (602). The sequence of embeddings can represent, e.g., a portion of a molecule, or a molecule, or a molecule complex.
[0201] In some cases, e.g., when the embedding neural network is a molecule complex embedding neural network, the sequence of embedding in the training example is generated by one or more single embedding neural networks, e.g., as described above with reference to FIG. 5.|0202| In some cases, e.g., when the embedding neural network is a single embedding neural network, the sequence of embeddings can be, e.g., a one-hot sequence of embedding. For instance, if the sequence of embeddings represents a protein, then the sequence of embeddings can include, for each position in each amino acid sequence of the protein, a respective one-hot embedding that identifies the amino acid at the position. As another example, if the sequence of embeddings represents a portion of a protein (e.g., a variable region of a light chain or a heavy chain of an antibody), then the sequence of embeddings can include, for each of one or more positions in one or more amino acid sequences of the protein, a respective one-hot embedding that identifies the amino acid at the position. As another example, if the sequence of embeddings represents a DNA molecule or an RNA molecule, then the sequence of embeddings can include, for each position in each nucleotide sequence of the DNA molecule, a respective one-hot embedding that identifies the nucleotide at the position. (It will be appreciated that a one-hot embedding is a particular choice of predefined embedding representing an entity such as an amino acid or a nucleotide and that, more generally, the system can use any predefined embedding scheme for amino acids or nucleotides).
[0203] The system generates a masked sequence of embeddings by masking one or more embeddings in the sequence of embeddings included in the training example (604). Masking an embedding can refer to replacing the embedding by a default (predefined) embedding, e.g., an embedding with a zero (or another default number) in every entry, or with an embedding sampled from a probability distribution (e.g., a Normal distribution or a uniform distribution) over a latent space. Masking an embedding has the effect of removing the information content of the embedding.
[0204] The system can select embeddings from the sequence of embeddings for masking in any appropriate way. For instance, the system can randomly select a predefined percentage (e.g., 10%, or 30%, or 50%) of the embeddings in the sequence of embeddings for masking. As another example, for a sequence of embeddings that represents an antibody and that includes one or more embeddings representing the variable region of the heavy chain of the antibody and one or more embeddings representing the variable region of the light chain of the antibody, the system can mask only the embeddings representing the variable region of the heavy7chain (while leaving the embeddings representing the variable region of the light chain unmasked). As another example, for a sequence of embeddings that represents an antibody (as in the previous example), the system can mask only the embedding representing the variable region of the light chain (while leaving the embeddings representing the variable region of the heavy chain unmasked). As another example, for a sequence of embeddings that represents a molecule complex of multiple molecules, the system can mask the embeddings representing one or more of the molecules while leaving the embeddings representing the remaining molecules unmasked.
[0205] The system processes the masked sequence of embeddings using the embedding neural network to generate a network output (606). The network output can be a sequence of embeddings, e.g., having the same number of embeddings as are included in the network input to the embedding neural network.
[0206] The system processes the network output of the embedding neural network using a decoder neural network to generate, for each masked embedding in the masked sequence of embeddings, a prediction for an unmasked version of the embedding (608). For instance, for a masked embedding that is a masked version of an embedding representing an amino acid, the decoder neural network can generate a score distribution over a set of possible amino acids (where the score for each amino acid defines a likelihood that the original embedding represents the amino acid). As another example, for a masked embedding that is a masked version of an embedding representing a nucleotide, the decoder neural network can generate ascore distribution over a set of possible nucleotides (where the score for each nucleotide defines a likelihood that the original embedding represents the nucleotide). As another example, for one or more of the masked embeddings in the sequence of embeddings, the decoder neural network can generate a predicted embedding that is a prediction for the unmasked version of the embedding.
[0207] The decoder neural network can have any appropriate neural network architecture that enables the decoder neural network to perform its described functions. In particular, the decoder neural network can include any appropriate types of neural network layers (e.g., attention layers, fully connected layers, convolutional layers, and so forth) in any appropriate number (e.g., 5 layers, 10 layers, or 50 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0208] The system backpropagates gradients of an objective function that measures an error in the predictions for the unmasked version of the embeddings through the decoder neural network and the embedding neural network (610). More specifically, the system determines gradients of the objective function with respect to the neural network parameters of the decoder neural network and the neural network parameters of the embedding neural network, e.g.. by backpropagation. The system then updates the values of the neural network parameters of the decoder neural network and the embedding neural network using the gradients, e.g., by an update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam. |0209| The objective function can be any appropriate objective that can measure an error in a prediction for an unmasked version of an embedding. For instance, the objective function can be a cross-entropy objective function, or an objective function that measures a norm (e.g., an LI or L2 norm) of a difference between: (i) an embedding, and (ii) a prediction for the embedding.
[0210] FIG. 7 is a flow diagram of an example process 700 for training a molecule embedding neural network to perform a 3D molecule structure prediction task. For convenience, the process 700 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 700.
[0211] For convenience, the process 700 describes training the molecule embedding neural network to perform the 3D molecule structure prediction task with reference to a particular training example. It will be appreciated that the process 700 can be applied to train the moleculeembedding neural network on a set of training examples that can include any appropriate number of training examples.
[0212] The system obtains a training example that includes: (i) a sequence of embeddings representing a molecular entity, where the molecular entity is a portion of a molecule, or a molecule, or a molecule complex, and (ii) target 3D structure data characterizing a 3D structure of the molecular entity (702).
[0213] The sequence of embeddings representing the molecular entity can be generated by one or more single embedding neural networks, as described above with reference to FIG. 5.
[0214] The target 3D structure data can be any appropriate data that characterizes the 3D structure of the molecular entity. For instance, the target 3D structure data can define a contact map that, for each pair of elements (e.g., amino acids, nucleotides, or atoms) of the molecular entity, defines whether the pair of elements are separated by less a threshold distance in the 3D structure of the molecular entity. The threshold distance can be, e.g., 4 Angstroms, or 8 Angstroms, or 12 Angstroms. As another example, the target 3D structure data can define a respective 3D spatial position of each element in the molecular entity, where each element can be. e.g., an amino acid, or a nucleotide, or an atom.
[0215] The target 3D structure data can be obtained from any appropriate source. For instance, the target 3D structure data can be obtained from physical experiments to assess the structure of the molecular entity, e.g., x-ray crystallography. As another example, the target 3D structure data can be generated as an output of a separate computational system, e.g., a separate neural network system dedicated to predicting molecule structures.
[0216] The system processes the sequence of embeddings included in the training example using the molecule complex embedding neural network to generate a network output (704). The network output can include an output generated by an output layer of the molecule complex embedding neural network, or an intermediate output generated by the molecule complex embedding neural network, or both.
[0217] For instance, in one example, the molecule complex embedding neural network can include a sequence of self-attention neural network layers. As part of processing the sequence of embeddings, each self-attention neural network layer can generate an array of attention weights as part of applying a self-attention operation to a layer input processed by the selfattention neural network layer. In this example, the system can extract the respective array of attention weights generated by one or more of the self-attention neural network layers, and provide these as an intermediate output generated by the molecule complex embedding neural network.
[0218] The system processes the network output of the molecule complex embedding neural network using a structure prediction neural network to generate predicted 3D structure data characterizing a predicted 3D structure of the molecular entity corresponding to the training example (706). The predicted 3D structure data can be, e.g., a predicted contact map that, for each pair of elements (e.g., amino acids, nucleotides, or atoms) of the molecular entity', defines a prediction for yvhether the pair of elements are separated by less a threshold distance in the 3D structure of the molecular entity’. The threshold distance can be. e.g., 4 Angstroms, or 8 Angstroms, or 12 Angstroms. As another example, the predicted 3D structure data can define a respective predicted 3D spatial position of each element in the molecular entity', where each element can be, e.g., an amino acid, or a nucleotide, or an atom.
[0219] The structure prediction neural network can have any appropriate neural network architecture that enables the structure prediction neural network to perform its described functions. In particular, the structure prediction neural netyvork can include any appropriate ty pes of neural network layers (e.g., attention layers, fully connected layers, convolutional layers, and so forth) in any appropriate number (e.g., 5 layers, 10 layers, or 50 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0220] The system jointly trains the structure prediction neural netyvork and the molecule complex embedding neural netyvork to optimize an objective function that depends on the predicted 3D structure generated by the structure prediction neural network (708). More specifically, the system determines gradients of the objective function with respect to the neural network parameters of the structure prediction neural network and the neural network parameters of the molecule complex embedding neural netyvork, e.g., by backpropagation. The system then updates the values of the neural network parameters of the structure prediction neural netyvork and the molecule complex embedding neural network using the gradients, e.g., by an update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0221] The objective function can be any appropriate objective that can measure an error between: (i) the target 3D structure data, and (ii) the predicted 3D structure data. For instance, the objective function can be a mean squared error (MSE) objective function, or a root-meansquare deviation (RMSD) objective function, or a global distance test (GDT) objective function.
[0222] FIG. 8 shows an example implementation of a part of the prediction system described in this specification which can generate embeddings of antibodies and antigens. In the example shown in FIG. 8, the system includes: (i) a first single molecule embedding neural network thatincludes a variable heavy embedding neural network 808 and a variable light embedding neural network 810 that collectively generate a sequence of embeddings 820, 822 representing an antibody, and (ii) a second single molecule embedding neural network 812 that generates a sequence of embeddings representing an antigen. The variable heavy embedding neural network 808, the variable light embedding neural network 810, and the single molecule embedding neural network 812 that generates the sequence of embeddings representing the antigen can be trained by the process described with reference to FIG. 5. The training procedure can involve training one or more of the described neural networks on training examples that include whole antibodies and / or training examples that include only heavy chains or light chains individually, as described in more detail with reference to FIG. 5.
[0223] The variable heavy embedding neural network 808 is configured to process data defining the amino acid sequence 802 of the variable region of the heavy chain of an antibody to generate a sequence of embeddings representing the variable region of the heavy chain. The variable heavy embedding neural network 808 can be implemented, e.g., as a pre-trained antibody-specific protein language model such as AntiBERTy, AbLang, HuBERT. etc. The variable heavy embedding neural network 808 can. in some cases, be trained (e.g.. to perform a language modeling task) substantially or entirely on training examples representing amino acid sequences of the variable region of the heavy chains of antibodies.
[0224] The system can process an intermediate output of the variable heavy embedding neural network 808 using a structure prediction neural network to generate a predicted contact map 814 characterizing the 3D structure of the variable region of the heavy chain of the antibody. The system can process the contact map 814 to generate positional embeddings characterizing the 3D structure of the variable region of the heavy chain, and then combine these positional embeddings with the sequence of embeddings representing the variable region of the heavy chain.
[0225] The system can process the amino acid sequence of the variable region of the heavy chain using the ANARCI (Antigen Receptor Numbering and Receptor Classification by IMGT (International Immunogenetic Information System)) technique 832 to assign each amino acid in the variable region of the heavy chain to a respective category (e.g., FWR1, CDR1, ... . CDR3, FWR4) 834, and then combine category embeddings representing the categories with the sequence of embeddings generated by the variable heavy embedding neural network. (Instead of using ANARCI, the system can instead use IMGT numberings obtained from another source).
[0226] The system can provide a molecule embedding 826 of the variable region of the heavy chain of the antibody (derived from the sequence of embeddings generated by the variable heavy embedding neural network, the positional embeddings, and the category embeddings) for processing by the molecule complex embedding neural network, as will be described with reference to FIG. 9.
[0227] The variable light embedding neural network 810 is configured to process data defining the amino acid sequence 804 of the variable region of the light chain of an antibody to generate a sequence of embeddings 822 representing the variable region of the light chain. The variable light embedding neural network can be implemented, e.g., as a pre-trained antibody-specific protein language model such as AntiBERTy, AbLang. HuBERT, etc. The variable light embedding neural network 810 can, in some cases, be trained (e.g., to perform a language modeling task) substantially or entirely on training examples representing amino acid sequences of the variable region of the light chains of antibodies.
[0228] The system can process an intermediate output of the variable light embedding neural network 810 using a structure prediction neural network to generate a predicted contact map 816 characterizing the 3D structure of the vanable region of the light chain of the antibody. The system can process the contact map 816 to generate positional embeddings characterizing the 3D structure of the variable region of the light chain, and then combine these positional embeddings with the sequence of embeddings representing the variable region of the light chain.
[0229] The system can process the amino acid sequence of the variable region of the light chain using the ANARCI (Antigen Receptor Numbering and Receptor Classification by IMGT (International Immunogenetic Information System)) 832 technique to assign each amino acid in the variable region of the light chain to a respective category’ (e.g.. FWR1, CDR1. ... , CDR3, FWR4) 834, and then combine category embeddings 834 representing the categories with the sequence of embeddings generated by the variable light embedding neural network. (Instead of using ANARCI, the system can instead use IMGT numberings obtained from another source).
[0230] The system can provide a molecule embedding 828 of the variable region of the light chain of the antibody (derived from the sequence of embeddings generated by the variable light embedding neural network, the positional embeddings, and the category’ embeddings) for processing by the molecule complex embedding neural network, as will be described with reference to FIG. 9.
[0231] The second single molecule embedding neural network 812 is configured to process data defining the amino acid sequence 806 of an antigen to generate a sequence of embeddings 824 representing the antigen. The second single molecule embedding neural network 812 can be implemented, e.g., as a pre-trained general protein language model such as ESM-2.
[0232] The system can process an intermediate output of the second single molecule embedding neural network 812 using a structure prediction neural network to generate a predicted contact map 818 characterizing the 3D structure of the antigen. The system can process the contact map 818 to generate positional embeddings characterizing the 3D structure of the antigen, and then combine these positional embeddings with the sequence of embeddings representing the antigen.
[0233] The system can provide an antigen embedding 830 of the antigen (derived from the sequence of embeddings generated by the second single molecule embedding neural network and the positional embeddings) for processing by the molecule complex embedding neural network, as will be described with reference to FIG. 9.
[0234] FIG. 9 shows an example implementation of a part of the prediction system described in this specification which can generate embeddings of antibody-antigen complexes.
[0235] The system processes a sequence of embeddings 826 representing the variable region of the heavy chain of the antibody using a projection layer 902 to project the sequence of embeddings 826 representing the variable region of the heavy chain into a common latent space. The system also generates a molecule identifier embedding 950 for the heavy chain, and combines the molecule identifier embedding with the sequence of embeddings 908 representing the heavy chain.
[0236] The system processes a sequence of embeddings 828 representing the variable region of the light chain of the antibody using a projection layer 904 to project the sequence of embeddings 828 representing the variable rejection of the light chain into the common latent space. The system also generates a molecule identifier embedding 952 for the light chain, and combines the molecule identifier embedding 952 with the sequence of embeddings 910 representing the light chain.
[0237] The system processes a sequence of embeddings 830 representing the antigen using a projection layer 906 to project the sequence of embeddings 830 representing the antigen into the common latent space. The system also generates a molecule identifier embedding 954 for the antigen, and combines the molecule identifier embedding with the sequence of embeddings 912 representing the antigen.
[0238] The system processes a concatenation of the respective sequences of embeddings representing the variable region of the heavy chain of the antibody, the variable region of the light chain of the antibody, and the antigen using a molecule complex embedding neural network 914 to generate an embedding 916 of the molecule complex. The molecule complex embedding neural network 914 is trained to perform a language modeling task, and optionally, a 3D molecule structure prediction task. In the 3D molecule structure prediction task, a structure prediction neural network processes an intermediate output generated by the molecule complex embedding neural network to generate a predicted contact map 920 for the antibodyantigen complex.
[0239] Optionally, the network input to the molecule complex embedding neural network 914 can additionally include a property-specific embedding 922 that is associated with a particular molecule complex property (e.g., binding affinity, stability, and so forth). The molecule complex embedding neural network 914 can process the property-specific embedding 922 jointly with the sequences of embeddings representing the variable region of the heavy chain, the variable region of the light chain, and the antigen to generate a network output that includes an updated version 918 of the property-specific embedding. The system can process the updated property-specific embedding 918 generated by the molecule complex embedding neural network 914 using a property prediction neural network to generate a predicted value of a property of the antibody-antigen complex, e.g., binding affinity, stability, etc.(0240] FIG. 10 - FIG. 13, which are described next, illustrate various parts of a data processing pipeline for extracting data defining parts of molecules, molecules, and molecule complexes from data files and preparing the extracted molecule data, e.g., for use in training or validating the system described in this specification.
[0241] FIG. 10 shows a flow diagram of an example process for extracting and parsing data defining molecules and / or molecule complexes (e.g., antibody, antigen, and antibody-antigen complexes) from a Crystallographic Information File (CIF).
[0242] FIG. 11 shows a flow diagram of an example process for finding and extracting all possible single chain antigen complexes from a data file.FIG. 12 shows a flow diagram of an example process for filtering data defining antibodyantigen complexes to remove contacts due to crystal packing issues. A 12 Angstrom threshold is used to find a contact.
[0243] FIG. 13 shows an example process deduplicating a set of data defining antigens, variable light chains, variable heavy chains, variable heavy - variable light chain pairs, variableheavy chain - antigen pairs, and variable heavy chain - variable light chain - antigen complexes.
[0244] FIG. 14 is a flow diagram of an example process 1400 for training a molecule embedding neural network and an epitope prediction neural network to perform an epitope prediction task. For convenience, the process 1400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1400.
[0245] For convenience, the process 1400 describes training the molecule embedding neural network and the epitope prediction neural network to perform the epitope prediction task with reference to a particular training example. It will be appreciated that the process 1400 can be applied to train the molecule embedding neural network and the epitope prediction neural network on a set of training examples that can include any appropriate number of training examples.
[0246] The system obtains a training example that includes: (i) a sequence of embeddings that jointly represents an antibody and an antigen, and (ii) target epitope data that identifies, for each amino acid in the antigen, whether the amino acid is included in an epitope that is bound by the antibody (1402). For instance, the target epitope data can define that an amino acid in the antigen is included in the epitope if the amino acid is located within a threshold 3D spatial distance of the antibody in a 3D structure of the antibody-antigen complex. The target epitope data may be determined based on the 3D structures of antibody-antigen complexes that are determined, e.g., by physical experiments (e.g., x-ray crystallography) or by a separate computational system, e.g., a separate neural network system dedicated to predicting antibodyantigen structures.
[0247] The sequence of embeddings representing the antibody and the antigen can be generated by one or more single embedding neural networks, as described above with reference to FIG. 5.
[0248] The system processes the sequence of embeddings included in the training example using the molecule complex embedding neural network to generate a network output (1404). The network output can include an output generated by an output layer of the molecule complex embedding neural network, or an intermediate output generated by the molecule complex embedding neural network, or both.
[0249] For instance, in one example, the molecule complex embedding neural network can include a sequence of self-attention neural network layers. As part of processing the sequenceof embeddings, each self-attention neural network layer can intermediate embeddings representing the antibody and the antigen. In this example, the system can extract an array of embeddings generated by one or more of the self-attention neural network layers and which represents the antigen, and provide these as an intermediate output generated by the molecule complex embedding neural network. In some cases, the intermediate output generated by the molecule complex embedding neural network includes an array of inter-chain attention vectors.
[0250] The system processes the network output of the molecule complex embedding neural network using the epitope prediction neural network to generate predicted epitope data that identifies, for each amino acid in the antigen, whether the amino acid is predicted to be included in the epitope that is bound by the antibody (1406). The predicted epitope data can define, for each amino acid in the antigen, a probability that the amino acid is included in the epitope.
[0251] The epitope prediction neural network can have any appropriate neural network architecture that enables the epitope prediction neural network to perform its described functions. In particular, the epitope prediction neural network can include any appropriate types of neural network layers (e.g., attention layers, fully connected layers, convolutional layers, and so forth) in any appropriate number (e.g., 5 layers, 10 layers, or 50 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0252] In a particular example, the epitope prediction neural network can be implemented as a graph neural network that includes a sequence of message passing neural network layers. In this example, the epitope prediction neural network can be configured to process a network input that includes a graph representation of the antigen. The graph representation of the antigen can include a set of nodes and a set of edges, where each node represents a respective amino acid in the antigen and where each edge connects a respective pair of nodes. For instance, a pair of nodes may be connected by an edge if the corresponding pair of amino acids in the antigen are separated by less than a threshold 3D spatial distance (e.g., 8 Angstroms) in the 3D structure of the antigen. The system can augment the graph representation of the antigen with node embeddings derived from the network output produced by the molecule embedding neural network. The system can then process the graph representation of the antigen using the epitope prediction neural network, including by the sequence of message passing layers included within the epitope prediction neural network, to generate the predicted epitope data.
[0253] The system jointly trains the epitope prediction neural network and the molecule complex embedding neural network to optimize an objective function that depends on the predicted epitope data generated by the epitope prediction neural network (1408). More specifically, the system determines gradients of the objective function with respect to the neuralnetwork parameters of the epitope prediction neural network and the neural network parameters of the molecule complex embedding neural network, e.g., by backpropagation. The system then updates the values of the neural network parameters of the epitope prediction neural network and the molecule complex embedding neural network using the gradients, e.g., by an update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0254] In some cases, the system trains the epitope prediction neural network while freezing the parameter values of the molecule complex embedding neural network, e.g., to increase the efficiency of training the epitope prediction neural network.
[0255] The objective function can be any appropriate objective that can measure an error between: (i) the target epitope data, and (ii) the predicted epitope data. For instance, the objective function can be a cross-entropy objective function.
[0256] FIG. 15 is a flow diagram of an example process 1500 determining mutated antibodies with higher binding affinities for an antigen. For convenience, the process 1500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1500.
[0257] The system obtains data identify an antibody that is known to bind to an antigen (1502). The data identifying the antibody can include data that defines the respective amino acid sequences of the heavy chain and the light chain of the antibody. The system also receives data identifying the amino acid sequence of the antigen. The system can receive the data characterizing the antibody and the antigen from any appropriate source, e.g., from a user by way of a user interface, or from an upstream computational or experimental system. The system can also receive data identifying the region of the antibody that binds to the antigen, i.e., the paratope of the antibody.
[0258] The system generates a data identifying a set of mutated versions of the antibody (1504). In particular, each mutated version of the antibody has one or more mutations in the amino acid sequence of the heavy chain, or the light chain, or both, as compared to the original antibody.
[0259] To generate the mutated versions of the antibody, the system selects a set of mutation positions in the antibody. For instance, the system can select some or all of the amino acids in the paratope of the antibody, or some or all of the amino acids that are located within a threshold distance (e.g., spatial distance or sequence distance) of the paratope of the antibody.
[0260] For each mutation position in the antibody, the system generates a score distribution over a set of possible amino acids for the mutation position. The set of possible amino acidscan refer to, e.g., the set of 20 standard amino acids. To generate the amino acid score distribution for the mutation position, the system generates a “masked” version of the amino acid sequence of the antibody, e.g., by replacing a token representing the identity of the amino acid at the mutation position with a default “masking” token. The system then processes: (i) the masked version of the antibody, and (ii) the antigen, using the single molecule embedding neural networks and the molecule complex embedding neural network described with reference to FIG. 1 to generate a sequence of embeddings representing the antibody-antigen complex. The system then processes the embedding that corresponds to the mutation position (and that was previously masked) using a decoder neural network to generate a score distribution over the set of possible amino acids. (An example of a decoder neural network is described with reference to FIG. 6). The system can then select each amino acid that is assigned a score having at least a threshold value under the score distribution as an “alternative” amino acid for the position.
[0261] To generate a mutated version of the antibody, the system can select a number of mutation positions (e.g., by randomly sampling a predefined number of mutation positions from the set of mutation positions). The system then selects, for each selected mutation position, an alternative amino acid to occupy the mutation position, e.g., from the set of alternative amino acids for the mutation position, as described above.
[0262] The system generates one or more quality metrics for each mutated version of the antibody (1506). A few examples of quality metrics are described next.
[0263] In one example, to generate a quality metric for a mutated version of the antibody, the system generates data characterizing a predicted 3D structure of a complex that includes the mutated version of the antibody and the antigen. The system then generates the quality metric based on how closely the paratope of the mutated version of the antibody is located to the epitope of the antigen in the 3D structure of the complex. For instance, the system can generate a contact map for the complex comprising the mutated version of the antibody and the antigen, e.g., by feeding the respective amino acid sequences of the mutated version of the antibody and the antigen through the single molecule embedding neural networks, the molecule complex embedding neural network, and the prediction generator described with reference to FIG. 1. The contact map can define, for each pair of amino acids in the complex, a distance between the pair of amino acids or whether the amino acids are separated by less than a threshold distance. The system can then generate the quality metric for the mutated version of the antibody by combining (e.g., summing or averaging) the entries in the contact map thatcorresponds to pairs of amino acids that include: (i) an amino acid in the epitope of the antigen, and (ii) an amino acid in the paratope of the mutated antibody.
[0264] A quality metric that characterizes the spatial proximity of the mutated version of the antibody to the antigen in a complex can be shown to be correlated with the binding affinity of the mutated version of the antibody for the antigen. More specifically, if the mutated version of the antibody is spatially closer to the antigen when bound, then the mutated version of the antibody is likely to have a higher binding affinity for the antigen.
[0265] Directly predicting the binding affinity of an antibody for an antigen may be challenging. For instance, there is only a very limited availability of “labeled” training data, i.e., that associates antibody-antigen complexes with experimentally measured binding affinity values. Therefore, training a machine learning model to predict binding affinity values for antibody-antigen complexes is challenging, i.e., due to the scarcity of labeled training data. However, 3D structures of antibody-antigen complexes can be predicted with a relatively high degree of accuracy, e.g., by the system described with reference to FIG. 1, and these predictions can be used to generate a quality metric (as described above) that is correlated with binding affinity. The system thus provides a technical solution to the technical problem of scarcity of labeled binding affinity training data for antibody-antigen complexes.
[0266] As another example, the system can generate a quality metric for a mutated version of the antibody based on a probability of the mutated version of the antibody. The system can compute the probability of the mutated version of the antibody as a combination (e.g., product) of, for each mutated position, the score for the alternative amino acid occupying the mutated position under the score distribution over the set of possible amino acids for the mutation position (as described above). For instance, the system can generate the quality metric as a logarithm of the joint mutational probability for the mutated version of the antibody.
[0267] As another example, the system can generate a quality metric for a mutated version of the antibody by generating, for each amino acid in the antigen, a probability that the amino acid is included in an epitope that is bound by the mutated version of the antibody. The system can generate the quality metric, e.g., based on a sum of the epitope inclusion probabilities for the amino acids in the antigen. The system can generate epitope inclusion probabilities for amino acids in the antigen using the single molecule embedding neural networks, the molecule complex embedding neural network, and the prediction generator described with reference to FIG. 1.
[0268] The system generates a ranking of the mutated versions of the antibody based on the quality metrics (1508). For instance, the system can generate an overall score for each mutatedversion of the antibody by combining (e.g., by a linear combination or a product operation) the quality metrics for the mutated version of the antibody. The system can then rank the mutated versions of the antibody based on the overall scores. Generally, mutated versions of the antibody that have a higher binding affinity for the antigen and that are more likely to be stable may tend to be ranked higher.
[0269] Optionally, certain mutated versions of the antibody can be selected for physical synthesis based at least in part on the ranking. For instance, one or more highest ranked mutated versions of the antibody may be selected for physical synthesis. The selected mutated versions of the antibody may then be physically synthesized and tested, e.g., for binding affinity to the antigen. Certain mutated versions of the antibody, e.g., those with higher binding affinity than the original antibody, may be selected for inclusion in antibody-based therapies.
[0270] FIG. 16 is a flow diagram of an example process 1600 for identifying epitopes on an antigen. For convenience, the process 1600 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 1600.
[0271] The system obtains data identifying an antigen and a library of antibodies (1602). The library of antibodies can include any appropriate number of antibodies, e.g., 1000, or 10,000, or 100,000 antibodies.
[0272] The system determines, for each antibody, a predicted epitope on the antigen, i.e., to which the antibody has a highest likelihood of binding to from among possible locations on the antigen (1604). For instance, the system can process data identifying the antibody and the antigen using the single molecule embedding neural networks, molecule complex embedding neural network, and prediction generator described with reference to FIG. 1 (and throughout this specification) to generate a data identifying a predicted epitope (i.e., binding region) on the antigen.
[0273] Optionally, the system can apply a deduplication operation to the set of predicted epitopes on the antigen to remove duplicates. The system can also apply a clustering operation (e.g., an expectation maximization or k-means clustering operation) to the predicted epitopes to produce a reduced number of “clustered” epitopes.
[0274] By computationally predicting the sites of epitopes on an antigen with reference to a library of known antibodies, the system can identify novel (previously unknown) epitopes on the antigen that may be targeted by antibodies to achieve therapeutic effects.
[0275] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0276] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0277] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0278] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as astand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0279] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0280] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.(0281) Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0282] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way ofexample semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0283] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensoiy feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0284] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
[0285] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0286] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0287] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. Therelationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0288] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0289] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0290] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0291] What is claimed is:
Claims
CLAIMS1. A method performed by one or more computers, the method comprising: receiving data characterizing a plurality of molecules; generating, for each of the plurality of molecules, a respective representation of the molecule as a sequence of embeddings in a latent space using a respective single molecule embedding neural network; jointly processing the respective representation of each of the plurality of molecules using a molecule complex embedding neural network, in accordance with values of a set of molecule complex embedding neural network parameters, to generate a representation of a molecule complex comprising the plurality of molecules; and processing the representation of the molecule complex to generate one or more predictions characterizing the molecule complex.
2. The method of claim 1, wherein the plurality of molecules comprise an antibody molecule and an antigen molecule.
3. The method of claim 2, wherein the antibody molecule is a multi -specific antibody molecule.
4. The method of any one of claims 2-3, wherein generating, for each of the plurality of molecules, a respective representation of the molecule as a sequence of embeddings comprises, for the antibody molecule: processing data characterizing a variable region of a heavy chain of the antibody molecule using a variable heavy embedding neural network to generate a representation of the variable region of the heavy chain of the antibody molecule as respective a sequence of embeddings; processing data characterizing a variable region of a light chain of the antibody molecule using a variable light embedding neural network to generate a representation of the variable region of the light chain of the antibody molecule as a respective sequence of embeddings; and generating the representation of the antibody molecule based on the sequence of embeddings representing the variable region of the heavy chain and the sequence ofembeddings representing the variable region of the light chain of the antibody molecule.
5. The method of claim 4, wherein the variable heavy embedding neural network has been trained to perform a protein language modeling task.
6. The method of any one of claims 4-5, wherein the variable light embedding neural network has been trained to perform a protein language modeling task.
7. The method of any preceding claim, wherein for one or more of the plurality of molecules, generating the representation of the molecule as a sequence of embeddings comprises: generating an initial sequence of embeddings representing the molecule using a respective single molecule embedding neural network; generating, for each embedding in the initial sequence of embeddings, a corresponding positional embedding based at least in part on data characterizing a three- dimensional (3D) structure of the molecule; and combining each embedding in the initial sequence of embeddings representing the molecule with the corresponding positional embedding.
8. The method of claim 7, wherein for one or more of the plurality of molecules: the molecule is a protein molecule and the initial sequence of embeddings representing the molecule comprises a respective embedding representing each of a plurality of amino acid residues in the molecule; and generating, for each embedding in the initial sequence of embeddings, a corresponding positional embedding based at least in part on data characterizing the 3D structure of the molecule comprises: obtaining one or more contact maps that each characterize respective 3D spatial distances between pairs of amino acid residues in the molecule; performing, for each of the one or more contact maps, an eigen-decomposition of a matrix representation of the contact map to generate a respective set of eigenvectors; and determining the positional embeddings based on the respective sets of eigenvectors of the one or more contact maps.
9. The method of any preceding claim, wherein the molecule complex embedding neural network has been trained to perform a language modeling task.
10. The method of claim 9, wherein training the molecule complex embedding neural network to perform a language modeling task comprises: obtaining a set of training examples, wherein each training example comprises a respective sequence of embeddings representing: (i) a portion of a molecule, or (ii) a molecule, or (iii) a molecule complex comprising a plurality of molecules; and training the molecule complex embedding neural network on the set of training examples to perform the language modeling task.
11. The method of claim 10, wherein training the molecule complex embedding neural network on the set of training examples to perform the language modeling task comprises performing the training over a first training stage and a second training stage; wherein the first training stage comprises training the molecule complex embedding neural network only on training examples that each represent: (i) a portion of a molecule, or (ii) a molecule; and wherein the second training stage comprises training the molecule complex embedding neural network on training examples that each represent molecule complexes.
12. The method of any one of claims 10-1 1 , wherein the set of training examples comprises a plurality' of training examples that each represent a variable region of a light chain of an antibody molecule.
13. The method of any one of claims 10-12, wherein the set of training examples comprises a plurality of training examples that each represent a variable region of a heavy chain of an antibody molecule.
14. The method of any one of claims 10-13. wherein the set of training examples comprises a plurality of training examples that each jointly represent both: (i) a variable region of a light chain, and (ii) a variable region of a heavy chain, of an antibody molecule.
15. The method of any one of claims 10-14, wherein the set of training examples comprises a plurality of training examples that each represent a respective antigen molecule.
16. The method of any one of claims 10-15, the set of training examples comprises a plurality of training examples that each represent a respective complex comprising an antibody molecule and an antigen molecule.
17. The method of any one of claims 10-16, wherein training the molecule complex embedding neural network on the set of training examples comprises, for each training example: generating a masked sequence of embeddings by masking one or more embeddings in the sequence of embeddings included in the training example; processing the masked sequence of embeddings using the molecule complex embedding neural network to generate a network output; processing the network output of the molecule complex embedding neural netw ork using a decoder neural network to generate, for each masked embedding in the masked sequence of embeddings, a prediction for an unmasked version of the embedding; backpropagating gradients of an objective function that measures an error in the predictions for the unmasked version of the embeddings through the decoder neural netw ork and the molecule complex embedding neural network.
18. The method of any one of claims 9-17, wherein the molecule complex embedding neural netw ork has further been trained to perform an auxiliary task of 3D structure prediction, comprising: obtaining a set of training examples, w herein each training example comprises: (i) a sequence of embeddings representing a molecular entity7, wherein the molecular entity is a portion of a molecule, or a molecule, or a molecule complex, and (ii) target 3D structure data characterizing a 3D structure of the molecular entity7; training the molecule complex embedding neural network on the set of training examples, comprising, for each training example: processing the sequence of embeddings included in the training example using the molecule complex embedding neural network to generate a network output; processing the network output of the molecule complex embedding neural network using a structure prediction neural network to generate predicted 3D structure datacharacterizing a predicted 3D structure of the molecular entity corresponding to the training example; and backpropagating gradients of an objective function that measures an error between: (i) the target 3D structure data, and (ii) the predicted 3D structure data, through the structure prediction neural network and into the molecule complex embedding neural network.
19. The method of claim 18, wherein for each training example, the target 3D structure data comprises a contact map.
20. The method of any one of claims 9-19, wherein the molecule complex embedding neural network has been jointly trained along with an epitope prediction neural network to perform an auxiliary task of antigen epitope prediction, comprising: obtaining a set of training examples, wherein each training example comprises: (i) a sequence of embeddings that jointly represents an antibody and an antigen, and (ii) target epitope data that identifies, for each amino acid in the antigen, whether the amino acid is included in an epitope that is bound by the antibody; training the molecule complex embedding neural network and the epitope prediction neural network on the set of training examples, comprising, for each training example: processing the sequence of embeddings included in the training example using the molecule complex embedding neural netw ork to generate a network output; processing the network output of the molecule complex embedding neural network using the epitope prediction neural network to generate predicted epitope data that identifies, for each amino acid in the antigen, whether the amino acid is predicted to be included in the epitope that is bound by the antibody; and backpropagating gradients of an objective function that measures an error between: (i) the target epitope data, and (ii) the predicted epitope data, through the epitope prediction neural network and into the molecule complex embedding neural network.
21. The method of claim 21, wherein the netw ork output of the molecule complex embedding neural network is an intermediate output produced by one or more hidden layers of the molecule complex embedding neural network.
22. The method of claim 22, wherein the intermediate output produced by the one or more hidden layers of the molecule complex embedding neural network comprises embeddings representing amino acids in the antigen which are produced by a self-attention layer of the molecule complex embedding neural network.
23. The method of any one of claims 19-22, wherein the epitope prediction neural network has a graph neural network architecture.
24. The method of any one of claims 19-23, wherein processing the network output of the molecule complex embedding neural network using the epitope prediction neural network to generate the predicted epitope data comprises: instantiating a representation of the antigen as a graph of nodes and edges, wherein each node represents a respective amino acid in the antigen and wherein each edge connects a respective pair of nodes; augmenting the graph representation of the antigen to associate each node in the graph with an embedding representing a corresponding ammo acid in the antigen, wherein the embedding is derived from the netw ork output generated by the molecule complex embedding neural netw ork; and processing the graph representation of the antigen using the epitope prediction neural network, including by a plurality of message passing layers of the epitope prediction neural network, to generate the predicted epitope data.
25. The method of any preceding claim, wherein the plurality of molecules comprise a protein molecule and a small molecule ligand.
26. The method of any preceding claim, wherein the plurality of molecules comprise a ribonucleic acid (RNA) molecule and a protein molecule.
27. The method of any preceding claim, wherein the plurality of molecules comprise an RNA molecule and a small molecule ligand.
28. The method of any preceding claim, wherein the one or more predictions characterizing the molecule complex comprises a predicted binding affinity of the plurality ofmolecules in the molecule complex.
29. The method of any preceding claim, wherein the one or more predictions characterizing the molecule complex comprises a predicted stability of the molecule complex.
30. The method of any preceding claim, wherein the one or more predictions characterizing the molecule complex comprises a prediction for which amino acids in the molecule complex are surface accessible.
31. The method of any preceding claim, wherein the one or more predictions characterizing the molecule complex comprises a three-dimensional (3D) structure of the molecule complex.
32. The method of any preceding claim, wherein the one or more predictions characterizing the molecule complex comprises, for each of one or more molecules of the plurality of molecules, data identifying a portion of the molecule that is predicted to be included in a binding interface.
33. The method of claim 32, further comprising performing docking of the plurality of molecules based at least in part on the data identifying the portion of the molecule that is predicted to be included in the binding interface.
34. The method any preceding claim, further comprising selecting the plurality of molecules for physical synthesis based on the one or more predictions characterizing the molecule complex.
35. The method of claim 34, further comprising physically synthesizing the plurality of molecules.
36. The method of claim 35, further comprising, after physically synthesizing the plurality of molecules: physically synthesizing molecule complexes comprising the plurality of molecules: and performing physical experiments to test one or more properties of the moleculecomplexes comprising the plurality7of molecules.
31. A method comprising: obtaining data identifying: (i) an antibody, and (ii) a plurality of candidate antigens; performing, for a plurality7of molecule complexes that each comprise the antibody and a respective one of the plurality of candidate antigens, the method of claim 28 to generate a predicted binding affinity of the molecule complex; and determining a ranking of the plurality of candidate antigens based on the binding affinities.
38. The method of claim 37, further comprising selecting one or more candidate antigens for physical synthesis based on the ranking.
39. The method of claim 38, further comprising physically synthesizing the one or more candidate antigens selected for physical synthesis.
40. The method of any one of claims 37-39, further comprising selecting one or more candidate antigens for inclusion in a therapeutic based on the ranking.
41. The method of claim 40. wherein the therapeutic comprises a vaccine.
42. A method comprising: obtaining data identifying: (i) an antigen, and (ii) a plurality of candidate antibodies; performing, for a plurality of molecule complexes that each comprise the antigen and a respective one of the plurality of candidate antibodies, the method of claim 28 to generate a predicted binding affinity of the molecule complex; and determining a ranking of the plurality7of candidate antibodies based on the binding affinities.
43. The method of claim 42, further comprising selecting one or more candidate antibodies for physical synthesis based on the ranking.
44. The method of claim 43, further comprising physically synthesizing the one or more candidate antibodies selected for physical synthesis.
45. The method of any one of claims 42-44, further comprising selecting one or more candidate antibodies for inclusion in a therapeutic based on the ranking.
46. The method of claim 46. further comprising physically synthesizing the therapeutic comprising the one or more selected candidate antibodies.
47. The method of any one of claims 45-46, wherein the therapeutic is an antibody -based therapy.
48. The method of any preceding claim, further comprising: obtaining data identifying an antibody and an antigen; determining a plurality of mutated versions of the antibody, comprising: selecting a set of mutation positions in the antibody; determining, using the molecule complex embedding neural network and for each mutation position in the antibody, a score distribution over a set of possible amino acids; selecting, for each mutation position in the antibody, a set of alternative amino acids for the position using the score distribution over the set of possible amino acids; and generating each mutated version of the antibody by, for each of one or more mutation positions in the antibody, replacing an amino acid at the mutation position with an alternative amino acid selected from the set of alternative amino acids for the position; and determining a ranking of the plurality of mutated versions of the antibody.
49. The method of claim 48, wherein determining the ranking of the plurality of mutated version of the antibody comprises: determining, for each mutated version of the antibody, a quality metric for the mutated version of the antibody, comprising: generating data characterizing a predicted 3D structure of a complex comprising the mutated version of the antibody and the antigen; and generating a quality metric based on a spatial proximity of the mutated version of the antibody to the antigen in the 3D structure of the complex; and determining the ranking of the plurality of mutated versions of the antibody based onthe quality metrics.
50. The method of claim 49, further comprising, for each mutated version of the antibody, determining an additional quality metric based on ajoint probability of the mutations in the mutated version of the antibody.
51. The method of any one of claims 48-50. further comprising selecting one or more mutated versions of the antibody for physical synthesis based at least in part on the ranking of the plurality of mutated versions of the antibody.
52. The method of any preceding claim, further comprising: obtaining data identifying an antigen and a library of antibodies; determining, for each antibody in the library of antibodies, a corresponding predicted epitope on the antigen using the molecule complex embedding neural network.
53. The method of claim 52. further comprising identifying one or more of the predicted epitopes on the antigen as a novel epitope.
54. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective method of any one of claims 1-53.
55. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective method of any one of claims 1-53.
56. A therapeutic that includes one or more antigens selected according to the method of claim 40 or one or more antibodies selected according to the method of claim 45.
57. The method of claim 56, wherein the therapeutic comprises a vaccine or an antibodybased therapeutic.