Protein structure prediction from amino acid sequences using self-attention neural networks
By using a multi-layer self-attention neural network to process the multi-sequence alignment results of proteins, the prediction structure of the protein is generated by pairs embedded and determined, which solves the problems of large computing resources and insufficient accuracy in the prior art, and achieves efficient and accurate protein structure prediction.
Patent Information
- Application Number
- CN202080069683.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-02
- Filing Date
- 2020-12-02
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-12-02
AI Technical Summary
In the prior art, when predicting protein structure, computing resources are consumed and insufficient accuracy is not sufficient, making it difficult to effectively encode the relationship between amino acid pairs.
Multi-layer self-attention neural network is used to process the multi-sequence alignment results of proteins, and pair embeddings are generated, and the predicted structure of proteins is determined through folding neural networks.
It reduces the consumption of computing resources, improves the accuracy of protein structure prediction, and can generate a three-dimensional configuration of proteins with higher efficiency and accuracy.
Smart Images

Figure CN114503203B_ABST
Abstract
Description
Background Art
[0001] This specification relates to predicting protein structure.
[0002] Proteins are designated (specified) by a sequence of amino acids. Amino acids are organic compounds that include amino and carboxyl functional groups and side chains (i.e., groups of atoms) that are specific to the amino acids. Protein folding refers to the physical process by which a sequence of amino acids folds into a three-dimensional configuration. The structure of a protein is defined by the three-dimensional configuration of the atoms in the amino acid sequence of the protein after protein folding. When in a sequence linked by peptide bonds, the amino acids may be referred to as amino acid residues.
[0003] Predictions can be made using machine learning models. A machine learning model receives an input and generates an output (e.g., a predicted output) based on the received input. Some machine learning models are parametric models and generate an output based on the received input and the values of the model's parameters. The structure of a protein can be predicted based on the amino acid sequence of a given protein.
[0004] Some machine learning models are deep models that employ multi-layer models to generate outputs for received inputs. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a nonlinear transformation to a received input to generate an output. Summary of the invention
[0005] This specification describes a system implemented as a computer program on one or more computers at one or more locations that performs protein structure prediction.
[0006] According to a first aspect, a method for determining a predicted structure of a protein specified by an amino acid sequence, performed by one or more data processing devices, is provided. The method includes obtaining a multiple sequence alignment (MSA) for the protein. The method may further include determining a corresponding initial embedding of the amino acid pair from the multiple sequence alignment and for each amino acid pair in the amino acid sequence of the protein. The method may further include processing the initial embedding of the amino acid pairs using a pair embedding neural network including a plurality of self-attention neural network layers to generate a final embedding for each amino acid pair. The method may then include determining the predicted structure of the protein based on the final embedding for each amino acid pair.
[0007] Some advantages of this method are described later. For example, implementations of the method generate "pairwise" embeddings that encode the relationships between pairs of amino acids in a protein. (For example, pairwise embeddings for pairs of amino acids in a protein may encode the relationships between corresponding specified atoms (e.g., carbon alpha atoms) in the amino acid pair) These can then be processed by subsequent neural network layers to determine additional information, particularly the predicted structure of the protein. In an implementation, the folded neural network processes the pairwise embeddings to determine the predicted structure, for example, in terms of structural parameters such as atomic coordinates or backbone torsion angles for the carbon atoms of the protein. An example implementation of such a folded neural network is described later, but others may also be used. In some implementations, the final pairwise embeddings are used to determine the initial embeddings for each (single) amino acid in the amino acid sequence, and the single embeddings may be used alone or in conjunction with the pairwise embeddings to determine the predicted structure.
[0008] In implementation, each attention neural network layer receives the current embedding of each amino acid pair and uses attention on the current embedding of the amino acid pair or on a proper subset (appropriate subset) of these embeddings to update the current embedding. For example, the self-attention neural network layer can determine a set of attention weights to apply to the current embedding (or subset) to update the current embedding, such as a weighted sum of the attention based on the current embedding.
[0009] In some implementations, the current embeddings of the amino acid pairs are arranged (arranged) into a two-dimensional array (array). Then, the self-attention neural network layer may include row-wise and / or column-wise self-attention neural network layers, for example, in alternating sequences. The row (or column)-wise self-attention neural network layer may update the current embedding of the amino acid pair using attention only on the current embedding of the amino acid pair that is in the same row (or column) as the current embedding of the amino acid pair.
[0010] It has been found that using self-attention significantly improves the accuracy of the predicted structure, for example, by generating embeddings that are easier to process to determine the predicted structure parameters. Row and column processing helps to achieve this with reduced computational resources.
[0011] The initial embedding of each amino acid pair can be determined by dividing the MSA into a set of clustered amino acid sequences and a set of (larger) additional amino acid sequences. Then, the embedding of each set can be generated, and the embedding of the additional amino acid sequence can be used to update the embedding of the clustered amino acid sequence. Then, the updated embedding of the clustered amino acid sequence can be used to determine the initial embedding of the amino acid pair. In this way, for computational efficiency, a large set of additional amino acid sequences can be used to enrich the information of the embedding of a smaller set of amino acid sequences (clustered amino acid sequences). The processing can be performed by a cross-attention neural network. The cross-attention neural network may include a neural network that uses attention (i.e., attention weights applied to the embedding of the additional amino acid sequence) on the embedding of the additional amino acid sequence to update the embedding of the clustered amino acid sequence. Optionally, the method may further involve using attention on the embedding of the clustered amino acid sequence to update the embedding of the additional amino acid sequence. The updating of the embedding of the clustered amino acid sequence and, optionally, the additional amino acid sequence can be performed repeatedly, for example sequentially in time or using sequential neural network layers.
[0012] In a second aspect, a system is provided, comprising: one or more computers; and one or more storage devices communicatively connected (coupled) to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations (execute) including the operations of the method of the first aspect.
[0013] In a third aspect, one or more non-transitory computer storage media storing instructions are provided that, when executed by one or more computers, cause the one or more computers to perform operations including the operations of the method of the first aspect.
[0014] Methods and systems described herein can be used to obtain ligands such as ligands of drugs or industrial enzymes. For example, the method for obtaining a ligand may include obtaining a target (target) amino acid sequence, particularly the amino acid sequence of a target protein, and using the target amino acid sequence as the amino acid sequence to carry out a computer-implemented method as described above or herein, to determine the (tertiary) structure of the target protein, i.e., the predicted protein structure. Then, the method may include assessing the interaction of one or more candidate ligands with the structure of the target protein. The method may further include selecting one or more of the candidate ligands as a ligand depending on the result of the assessment of the interaction.
[0015] In some implementations, evaluating the interaction may include evaluating the binding of a candidate ligand to a structure of a target protein. For example, evaluating the interaction may include identifying a ligand that binds with sufficient affinity for a biological effect. In some other implementations, evaluating the interaction may include evaluating the association of a candidate ligand with a structure of a target protein that has an effect on the function of the target protein, such as an enzyme. The evaluation may include evaluating the affinity between the structure of the candidate ligand and the target protein or evaluating the selectivity of the interaction.
[0016] (Multiple) candidate ligands can be obtained from a database of candidate ligands, and / or can be obtained by changing (modifying) ligands in a database of candidate ligands (for example, by changing the structure or amino acid sequence of the candidate ligands), and / or can be obtained by stepwise or iterative assembly / optimization of candidate ligands.
[0017] The evaluation of the interaction of the candidate ligand with the target protein structure can be performed using a computer-assisted approach in which a graphical model of the structure of the candidate ligand and the target protein is displayed to the user, and / or the evaluation can be performed partially or completely automatically, for example using standard molecular (protein-ligand) docking software. In some implementations, the evaluation may include determining an interaction score for the candidate ligand, wherein the interaction score includes a measure of the interaction between the candidate ligand and the target protein. The interaction score may depend on the strength and / or specificity of the interaction, for example, the score depends on the binding free energy. The candidate ligand may be selected depending on its score.
[0018] In some implementations, the target protein includes a receptor or an enzyme, and the ligand is an agonist or antagonist of the receptor or enzyme. In some implementations, the method can be used to identify the structure of a cell surface marker. This can then be used to identify a ligand, such as an antibody or a label such as a fluorescent label, that is bound to a cell surface marker. This can be used to identify and / or process (treat) cancer cells.
[0019] In some implementations, the candidate ligand(s) may include small molecule ligands, such as organic compounds with a molecular mass < 900 Daltons. In some other implementations, the candidate ligand(s) may include polypeptide ligands, i.e., defined by an amino acid sequence.
[0020] Some implementations of this method can be used to determine the structure of a candidate polypeptide ligand, such as a ligand for a drug or industrial enzyme. Its interaction with a target protein structure can then be assessed; the target protein structure can be determined using computer-implemented methods as described herein or using conventional physical research techniques such as x-ray crystallography and / or magnetic resonance techniques.
[0021] Therefore, in another aspect, a method for obtaining a polypeptide ligand (e.g., a molecule or sequence thereof) is provided. The method may include obtaining the amino acid sequence of one or more candidate polypeptide ligands. The method may further include performing a computer-implemented method as described above or herein, using the amino acid sequence of the candidate polypeptide ligand as the amino acid sequence to determine the (tertiary) structure of the candidate polypeptide ligand. The method may further include obtaining the target protein structure of the target protein via in silico and / or by physical studies, and evaluating the interaction between the respective structures of one or more candidate polypeptide ligands and the target protein structure. The method may further include selecting one or more of the candidate polypeptide ligands as the polypeptide ligand depending on the results of the evaluation.
[0022] As previously described, evaluating the interaction may include evaluating the binding of a candidate polypeptide ligand to a structure of a target protein, e.g., identifying a ligand that binds with sufficient affinity for a biological effect, and / or evaluating the association of a candidate polypeptide ligand with a structure of a target protein that has an effect on the function of a target protein, such as an enzyme, and / or evaluating the affinity between a candidate polypeptide ligand and a structure of a target protein, or evaluating the selectivity of the interaction. In some implementations, the polypeptide ligand may be an aptamer.
[0023] Implementation of the method may further include synthesizing, ie, making, a small molecule or polypeptide ligand. The ligand may be synthesized by any conventional chemical technique and / or may be already available, for example, from a compound library or may be synthesized using combinatorial chemistry.
[0024] The method may further include testing the biological activity of the ligand in vitro and / or in vivo. For example, the ADME (absorption, distribution, metabolism, excretion) and / or toxicological properties of the ligand may be tested to screen out unsuitable ligands. Testing may include, for example, contacting a candidate small molecule or polypeptide ligand with a target protein and measuring changes in the expression or activity of the protein.
[0025] In some implementations, candidate (polypeptide) ligands may include: isolated antibodies, fragments of isolated antibodies, single variable domain antibodies, bi- or multi-specific antibodies, multivalent antibodies, bi-variable domain antibodies, immunoconjugates, fibronectin molecules, adnectins, DARPins, avimers, affibodies, anticalins, affilins, protein epitope mimetics, or combinations thereof. Candidate (polypeptide) ligands may include antibodies with mutated or chemically modified amino acid Fc regions, for example, which prevent or reduce ADCC (antibody-dependent cellular toxicity) activity and / or increase half-life when compared to wild-type Fc regions.
[0026] Misfolded proteins are associated with many diseases. Therefore, on the other hand, a method for identifying the presence of a protein misfolding disease is provided. The method may include: obtaining the amino acid sequence of the protein, and using the amino acid sequence of the protein to perform a computer-implemented method as described above or herein to determine the structure of the protein. The method may further include obtaining the structure of a version of the protein obtained from a human or animal body, for example by conventional (physical) methods. Then, the method may further include: comparing the structure of the protein with the structure of the version obtained from the body, and identifying the presence of a protein misfolding disease depending on the result of the comparison. That is, by comparison with the structure determined via computer simulation, the misfolding of the version of the protein from the body can be determined.
[0027] In some other aspects, the computer-implemented methods as described above or herein can be used to identify active / binding / blocking sites on a target protein from its amino acid sequence.
[0028] Particular implementations of the subject matter described in this specification can be implemented to realize one or more of the following advantages.
[0029] The protein structure prediction system described in this specification can predict the structure of the protein by: a single forward propagation through a collection of jointly trained neural networks, which can take less than one second. In contrast, some conventional systems predict the structure of the protein by: an extended search process in the space of possible protein structures to optimize a scalar score function, for example, using simulated annealing or gradient descent techniques. Such a search process may require millions of search iterations and consume hundreds of central processing unit (CPU) hours. Compared with a system that predicts protein structure through an iterative search process, predicting protein structure through a single forward propagation through a collection of neural networks can enable the structure prediction system described in this specification to consume less computing resources (e.g., storage and computing power). In addition, the structure prediction system described in this specification can (in some cases) predict protein structure with an accuracy comparable to or higher than that of a more computationally intensive structure prediction system.
[0030] To generate predicted protein structures, the structure prediction system described in this specification generates a set of embeddings representing proteins by processing a multiple sequence alignment (MSA) corresponding to the protein using a learning operation implemented by a neural network layer. In contrast, some conventional systems rely on generating MSA features through artificial feature engineering techniques. Using a learned neural network operation rather than an artificial feature engineering technique to process the MSA enables the structure prediction system described in this specification to predict protein structures more accurately than some conventional systems.
[0031] The structure prediction system described in this specification processes MSA to generate a set of "paired" embeddings, where each pair of embeddings corresponds to a corresponding pair of amino acids in a protein. The system then enriches the pair of embeddings by processing the pairwise embeddings through a series of self-attention neural network layers, and uses the enriched pairwise embeddings to predict protein structure. Generating and enriching pairwise embeddings enables the structure prediction system to develop an effective representation of the relationship between amino acid pairs in proteins, thereby contributing to accurate protein structure prediction. In contrast, some conventional systems rely on "single" embeddings (i.e., each of which corresponds to a corresponding amino acid, rather than an amino acid pair), which may be less effective for developing an effective representation for protein structure prediction. Using pairwise embeddings enables the structure prediction system described in this specification to predict protein structure more accurately than other methods.
[0032] The structure of a protein determines the biological function of the protein. Therefore, determining the protein structure can help understand life processes (e.g., including the mechanism of many diseases) and design proteins (e.g., as drugs, or as enzymes for industrial processes). For example, which molecules (e.g., drugs) will bind to proteins (and where the binding will occur) depends on the structure of the protein. Since the effectiveness of drugs can be affected by the degree to which they are bound to proteins (e.g., in the blood), determining the structure of different proteins can be an important aspect of drug development. However, using physical experiments (e.g., by x-ray crystallography) to determine protein structure can be time-consuming and very expensive. Therefore, the protein prediction system described in this specification can help in the field of biochemical research and engineering (e.g., drug development) involving proteins.
[0033] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Displays an example protein structure prediction system.
[0035] Figure 2 An example illustrating the "grid transformer" neural network architecture.
[0036] Figure 3 Shows an example multiple sequence alignment (MSA) embedded system.
[0037] Figure 4 is a graphic representation of an unfolded protein and a folded protein.
[0038] Figure 5 is a flow chart of an example process for determining a predicted protein structure.
[0039] Like reference numbers and designations in the various drawings represent like elements. DETAILED DESCRIPTION
[0040] This specification describes a structure prediction system for predicting the structure of a protein, that is, for predicting the three-dimensional configuration of an amino acid sequence in a protein after the protein undergoes protein folding.
[0041] To predict the structure of a protein, the structure prediction system processes a multiple sequence alignment (MSA) corresponding to the protein to generate a representation of the protein as a collection of "pairwise" embeddings. Each pairwise embedding corresponds to a pair of amino acids in the protein, and the structure prediction system enriches the pairwise embeddings by repeatedly updating the pairwise embeddings using a series of self-attention neural network layers that share information between the pairwise embeddings.
[0042] The structure prediction system processes the pairwise embeddings to generate corresponding "single" embeddings for each amino acid in the protein, and processes the single embeddings using a folded neural network to generate a predicted structure for the protein. The predicted structure of a protein can be defined by a set of structural parameters that define the spatial position and, optionally, rotation of each amino acid in the protein.
[0043] These and other features are described in more detail below.
[0044] Figure 1 An example protein structure prediction system 100 is shown. Protein structure prediction system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described below are implemented.
[0045] The structure prediction system 100 is configured to process data defining an amino acid sequence 102 of a protein 104 to generate a predicted structure 106 of the protein 104. Each amino acid in the amino acid sequence 102 is an organic compound that includes an amino functional group and a carboxyl functional group and a side chain (i.e., a group of atoms) specific to the amino acid. The predicted structure 106 defines an estimate of the three-dimensional (3-D) configuration of the atoms in the amino acid sequence 102 of the protein 104 after the protein 104 undergoes protein folding.
[0046] The term "protein" as used throughout this specification may be understood to refer to any biological molecule specified by one or more amino acid sequences. For example, the term protein may be understood to refer to a protein domain (i.e., a portion of an amino acid sequence that can undergo protein folding almost independently of the rest of the amino acid sequence) or a protein complex (i.e., which is specified by multiple related amino acid sequences).
[0047] The amino acid sequence 102 can be represented in any suitable digital format. For example, the amino acid sequence 102 can be represented as a series of one-hot vectors. In this example, each one-hot vector represents a corresponding amino acid in the amino acid sequence 102. For each different amino acid (e.g., a predetermined number of amino acids, such as 21), the one-hot vector has different components. The one-hot vector representing a specific amino acid has a value of 1 (or some other predetermined value) in the component corresponding to the specific amino acid and has a value of 0 (or some other predetermined value) in the other components.
[0048] The predicted structure 106 of the protein 104 is defined by the values of a set of structural parameters. For each amino acid in the protein 104, the set of structural parameters may include: (i) a position parameter, and (ii) a rotation parameter.
[0049] The position parameters for amino acids can specify the predicted 3-D spatial position of a specified atom in the amino acid in the protein structure. The specified atom can be the alpha carbon atom in the amino acid, i.e., the carbon atom to which the amino functional group, the carboxyl functional group, and the side chain of the amino acid are bonded. The position parameters for amino acids can be expressed in any suitable coordinate system, such as a three-dimensional [x, y, z] Cartesian coordinate system.
[0050] The rotation parameters for an amino acid can specify the predicted "orientation" of the amino acid in the protein structure. More specifically, the rotation parameters can specify a 3-D spatial rotation operation that, if applied to the coordinate system of the position parameters, causes the three "backbone" atoms in the amino acid to assume fixed positions relative to the rotated coordinate system. The three backbone atoms in an amino acid refer to the connected series of nitrogen, alpha carbon, and carbonyl carbon atoms in the amino acid (or, for example, alpha carbon, nitrogen, and oxygen atoms in the amino acid). The rotation parameters for an amino acid can be represented, for example, as a standard orthogonal 3×3 matrix with a determinant (determinant) equal to 1.
[0051] Typically, the position and rotation parameters for an amino acid define an egocentric reference frame for the amino acid. In this reference frame, the side chain for each amino acid can start from the origin, and the first bond along the side chain (i.e., the α carbon-β carbon bond) can be along a defined direction.
[0052] To generate the predicted structure 106, the structure prediction system 100 obtains a multiple sequence alignment (MSA) 108 corresponding to the amino acid sequence 102 of the protein 104. The MSA 108 specifies a sequence alignment of the amino acid sequence 102 with a plurality of additional amino acid sequences (e.g., from other, e.g., homologous proteins). The MSA 108 can be generated, for example, by processing a database of amino acid sequences using any suitable computational sequence alignment technique (e.g., progressive alignment construction). The amino acid sequences in the MSA 108 can be understood to have an evolutionary relationship, e.g., where each amino acid sequence in the MSA 108 can share a common ancestor. The correlation between the amino acid sequences in the MSA 108 can encode information related to the structure of the predicted protein 104. The MSA 108 can be used as a one-hot encoding identical to the amino acid sequence, with additional categories for insertions (or deletions) of amino acid residues.
[0053] The structure prediction system 100 generates a predicted structure 106 from an amino acid sequence 102 and a MSA 108 using: (1) a multiple sequence alignment embedding system 110 (i.e., "MSA embedding system"), (2) a pairwise embedding neural network 112, and (3) a folded neural network 114, each of which is described in more detail below. As used herein, an "embedding" can be an ordered collection of values, such as a vector or matrix of values.
[0054] The MSA embedding system 110 is configured to process the MSA 108 to generate an MSA embedding 116. The MSA embedding 116 is an alternative representation of the MSA 108 that includes a respective embedding for each amino acid in each amino acid sequence corresponding to the MSA 108. The MSA embedding may be represented as a matrix having dimensions M×N×E, where M is the number of sequences in the MSA 108, N is the number of amino acids in each amino acid sequence in the MSA 108 (which may be the same as the length of the amino acid sequence), and E is the dimension of the embedding corresponding to each amino acid of the MSA 108.
[0055] In one example, the MSA embedding system 110 may generate the MSA embedding M as follows:
[0056]
[0057] in is the representation of the amino acid sequence 102 as a 1-D array of one-hot amino acid embedding vectors, h θ (·) is applied to 1-D arrays The learned linear projection operation of each embedding of is a representation of the MSA 108 as a 2-D array of one-hot amino acid embedding vectors (where each row of the 2-D array corresponds to a corresponding amino acid sequence of the MSA), h φ(·) is applied to 2-D arrays The learned linear projection operation of each embedding of The operation represents Add to Each row.
[0058] refer to Figure 3 Another example of MSA embedded system 110 is described in more detail. Figure 3 The described MSA embedding system 110 may generate embeddings corresponding only to a proper subset of the sequences in the MSA 108, for example, to reduce computational resource consumption. Figure 3 The described MSA embedding system 110 can use self-attention neural network layers to “share” information within and between amino acid sequences in the MSA 108 , thereby generating a more informative MSA embedding 116 that can facilitate more efficient prediction of the protein structure 106 .
[0059] After generating the MSA embedding 116, the structure prediction system 100 converts the MSA embedding 116 into an alternative representation as a set of pairwise embeddings 118. Each pairwise embedding 118 corresponds to an amino acid pair in the amino acid sequence 102 of the protein. An amino acid pair in the amino acid sequence 102 refers to a first amino acid and a second amino acid in the amino acid sequence 102, and the set of possible amino acid pairs is given by:
[0060] {(A i , A j ): 1≤i, j≤N} (2)
[0061] Where N is the number of amino acids in the amino acid sequence 102, i.e., the length of the amino acid sequence 102, and A i is the amino acid at position i in amino acid sequence 102, A j is the amino acid at position j in the sequence, and N is the number of amino acids in amino acid sequence 102 (i.e., the length of amino acid sequence 102). The set of pairwise embeddings 118 can be represented as a matrix having dimensions N×N×F, where N is the number of amino acids in amino acid sequence 102, and F is the dimension of each pairwise embedding 118.
[0062] The structure prediction system 100 can generate a pairwise embedding 118 from the MSA embedding 116, for example, by computing the outer product of the MSA embedding 116 with itself, marginalizing (e.g., averaging) the MSA dimensions of the outer product, and combining the embedding dimensions of the outer product. For example, the pairwise embedding 118 can be generated as:
[0063]
[0064]
[0065]
[0066] where M represents the MSA embedding 116 (e.g., having a first dimension indexing the sequences in the MSA, a second dimension indexing the amino acids in each sequence in the MSA, and a third dimension indexing the channels of the embedding for each amino acid), represents the weight matrix that is multiplied by M to generate the “left” encoded embedding L, represents the weight matrix that is multiplied by M to generate the “right” encoding embedding R, Denotes the weight matrix that combines the left-encoded and right-encoded embeddings to generate the pairwise embedding P. Equations (3) to (5) use Einstein summation notation, where only the repeated indices that appear on the right side of the equation are implicitly summed over.
[0067] In other words, the structure prediction system 100 can generate the pairwise embeddings 118 from the MSA embeddings 116 by computing the “outer product mean” of the MSA embeddings 116, where the MSA embeddings are viewed as an M×N array of embeddings (i.e., where each embedding of the array has dimension E). The outer product mean defines a series of operations that, when applied to the MSA embeddings 116, generates an N×N array of embeddings (i.e., where N is the number of amino acids in the protein) that define the pairwise embeddings 118.
[0068] To compute the outer product mean of the MSA embedding 116, the system 100 may apply the outer product mean operation to the MSA embedding 116, and identify (confirm) the pairwise embedding 118 as a result of the outer product mean operation. To compute the outer product mean, the system generates a tensor A(·), for example, given by:
[0069]
[0070] Where res1, res2 ∈ {1, ..., N}, where N is the number of amino acids in the protein, ch1ch2 ∈ {1, ..., E}, where E is the number of channels in each embedding of the M×N array of embeddings representing the MSA embeddings, M is the number of rows in the M×N array of embeddings representing the MSA embeddings, LeftAct(m, res1, ch1) is a linear operation (e.g., defined by matrix multiplication) applied to the channel ch1 of the embedding located at the row indexed by “m” and the column indexed by “res1” in the M×N array of embeddings representing the MSA embeddings, and RightAct(m, res2, ch2) is a linear operation (e.g., defined by matrix multiplication) applied to the channel ch2 of the embedding located at the row indexed by “m” and the column indexed by “res2” in the M×N array of embeddings representing the MSA embeddings. The result of the outer product mean is generated by flattening and linearly projecting the (ch1, ch2) dimension of the tensor A. Optionally, the system may perform one or more layer normalization operations as part of computing the outer product mean (e.g., as described in Jimmy Lei Ba et al., “Layer Normalizaiton,” arXiv:1607.06450).
[0071] Typically, the MSA embedding 116 may be expressed with explicit reference to the multiple amino acid sequences of the MSA 108, e.g., as a 2-D array of the embedding, where each row of the 2-D array corresponds to a corresponding amino acid sequence of the MSA 108. Thus, the format of the MSA embedding 116 may not be suitable for predicting the structure 106 of a protein 104 that does not have an explicit dependency on the individual amino acid sequences of the MSA 108. In contrast, the pairwise embedding 118 characterizes the relationship between corresponding pairs of amino acids in the protein 104 and is expressed without explicit reference to the multiple amino acid sequences from the MSA 108, and is therefore a more convenient data representation for predicting the protein structure 106.
[0072] The pairwise embedding neural network 112 is configured to process a set of pairwise embeddings to update the values of the pairwise embeddings, i.e., to generate updated pairwise embeddings 120. The pairwise embedding neural network 112 updates the pairwise embeddings by processing them using one or more "self-attention" neural network layers. As used throughout this document, a self-attention layer generally refers to a neural network layer that updates a set of embeddings, i.e., receives a set of embeddings and outputs updated embeddings. To update a given embedding, the self-attention layer determines a corresponding "attention weight" between the given embedding and each of one or more selected embeddings, and then updates the given embedding using (i) the attention weights and (ii) the selected embedding. For convenience, the self-attention layer can be said to update the given embedding using attention "on" the selected embedding.
[0073] In one example, the self-attention layer can receive a set of input embeddings And to update the embedding x i , the self-attention layer can determine the attention weight in and a j Represents x i and x j The attention weights between are as follows:
[0074]
[0075]
[0076] Where W q and W k is the learning parameter matrix, softmax(·) represents the soft-max normalization operation, and c is a constant. Using the attention weights, the self-attention layer can update the embedding x as follows i :
[0077]
[0078] Where W v is the learning parameter matrix.
[0079] In addition to the self-attention neural network layer, the pairwise embedding neural network may include other neural network layers, such as linear neural network layers, which may be interleaved with the self-attention layer. In one example, the pairwise embedding neural network 112 has a neural network architecture referred to herein as a "grid transformer neural network architecture" which will be referred to as Figure 2 Describe in more detail.
[0080] The structure prediction system 100 uses the pairwise embeddings 118 to generate a corresponding "single" embedding 122 corresponding to each amino acid in the protein 104. In one example, the structure prediction system may generate a single embedding S corresponding to amino acid i. i ,as follows:
[0081]
[0082] Where P i,j is the amino acid pair (A i , A j ), A i is the amino acid at position i in amino acid sequence 102, and A jis the amino acid at position j in the amino acid sequence 102. In another example, the structure prediction system may further multiply the right hand side of equation (9) by a factor of 1 / N, i.e., perform mean pooling instead of sum pooling. In another example, the structure prediction system may identify (confirm) the single embedding as the diagonal of the pairwise embedding matrix, i.e., such that the single embedding S corresponding to amino acid i is i By P i,i (i.e., corresponding to the amino acid pair (A i , A j ). In another example, before processing the pairwise embeddings 118 using the pairwise embedding neural network 112, the structure prediction system may append the embeddings of the additional rows to the pairwise embeddings (where the embeddings of the additional rows may be initialized with random values or default values). In this example, the set of pairwise embeddings may be represented as a matrix having dimensions (N+1)×N×F, where N is the number of amino acids in the amino acid sequence and F is the dimension of each pairwise embedding. In this example, after processing the pairwise embeddings using the pairwise embedding neural network 112, the structure prediction system may extract the embeddings of the appended rows from the updated pairwise embeddings 120, and identify (confirm) the extracted embeddings as a single embedding 122.
[0083] The folded neural network 114 is configured to generate an output specifying a value of a structural parameter that defines a predicted structure 106 of the protein by processing an input including a single embedding 122, a pairwise embedding 120, or both. In one example, the folded neural network 114 may process the single embedding 122, the pairwise embedding 120, or both using one or more neural network layers (e.g., convolutional or fully connected neural network layers) to generate an output specifying a value of a structural parameter of the protein 104.
[0084] The training engine can train the structure prediction system 100 from end to end to optimize the structure loss 126, and optionally one or more auxiliary losses such as the reconstruction loss 128, the distance prediction loss 130, or both, each of which will be described in more detail below. The training engine can train the structure prediction system 100 on a set of training data including a plurality of training examples. Each training example can specify: (i) a training input including an amino acid sequence and a corresponding MSA, and (ii) a target protein structure that should be generated by the structure prediction system 100 by processing the training input. The target protein structure used to train the structure prediction system 100 can be determined using experimental techniques (e.g., x-ray crystallography).
[0085] The structural loss 126 can characterize the similarity between: (i) a predicted protein structure generated by the structure prediction system 100, and (ii) a target protein structure that should have been generated by the structure prediction system. It can be given by:
[0086]
[0087]
[0088]
[0089] Where N is the number of amino acids in the protein, t i represents the predicted position parameter for amino acid f, R i represents the 3×3 rotation matrix specified by the predicted rotation parameters for amino acid i, is the target position parameter for amino acid i, represents the 3×3 rotation matrix specified by the target rotation parameters for amino acid i, A is a constant, is the predicted rotation parameter R i The inverse (reciprocal) of the specified 3×3 rotation matrix, The target rotation parameter The inverse of the specified 3×3 rotation matrix, and (·) + Represents a rectified linear unit (ReLU) operation.
[0090] The structural loss defined by equations (10)-(12) can be understood as the loss of each amino acid pair in the protein Find the average. ij defines the predicted spatial position of amino acid j in the predicted reference frame of amino acid i, and Defines the actual spatial position of amino acid j in the actual reference frame of amino acid i. These terms are sensitive to the predicted and actual rotations of amino acids i and j, and therefore carry richer information than loss terms that are only sensitive to the predicted and actual distances between amino acids. Optimizing the structural losses forces the structure prediction system 100 to generate predicted protein structures that accurately mimic (approximate) real protein structures.
[0091] The (unsupervised) reconstruction loss 128 measures the accuracy of the structure prediction system 100 in predicting the identities of amino acids from the MSA 108 that was "corrupted" prior to the generation of the MSA embedding 116. More specifically, the structure prediction system 100 may corrupt the MSA 108 prior to the generation of the MSA embedding 116 by randomly selecting a token representing the identity of an amino acid at a particular position in the MSA 108 to modify or mask. The structure prediction system 100 may modify the token representing the identity of an amino acid at a position in the MSA 108 by replacing it with a randomly selected token that identifies a different amino acid. The structure prediction system 100 may consider the frequency with which different amino acids appear in the MSA 108 as part of modifying the tokens that identify amino acids in the MSA 108. For example, the structure prediction system 100 may generate a probability distribution over a set of possible amino acids, where the probability associated with an amino acid is based on the frequency with which the amino acid appears in the MSA 108. The structure prediction system 100 can use the amino acid probability distribution in modifying the MSA 108 by randomly sampling disrupted amino acid identities according to the probability distribution. The structure prediction system 100 can mask the glyphs representing the amino acid identities at positions in the MSA by replacing the glyphs with specially designated empty glyphs.
[0092] To evaluate the reconstruction loss 128, the structure prediction system 100 processes the MSA embedding 116 (e.g., using a linear neural network layer) to predict the true identity of the amino acid at each disrupted position in the MSA 108. The reconstruction loss 128 can be, for example, a cross-entropy loss between the predicted and true identities of the amino acids at the disrupted positions in the MSA 108. During the calculation of the reconstruction loss 128, the identities of the amino acids at the undisrupted positions of the MSA 108 can be disregarded.
[0093] The distance prediction loss 130 measures the accuracy of the structure prediction system 100 in predicting the physical distances between pairs of amino acids in the protein structure. To evaluate the distance prediction loss 130, the structure prediction system 100 processes the pairwise embeddings 118 generated by the pairwise embedding neural network 112 (e.g., using a linear neural network layer) to generate a distance prediction output. The distance prediction output characterizes the predicted distance between each pair of amino acids in the protein structure 106. In one example, the distance prediction output may specify a probability distribution over a set of possible distance ranges between the amino acid pairs for each pair of amino acids in the protein. The set of possible distance ranges may be given by, for example, {(0, 5], (5, 10], (10, 20], (20, ∞)}, where the distance may be measured in angstroms (or any other suitable unit of measurement). The distance prediction loss 130 may be, for example, a cross entropy loss between predicted and actual distances (ranges) between pairs of amino acids in a protein, such as a cross entropy loss between a 1-hot range bin and a ground truth structure. (The distance between a first amino acid and a second amino acid may refer to the distance between a specified atom in the first amino acid and a corresponding specified atom in the second amino acid. The specified atom may be, for example, a carbon alpha atom). In some implementations, the distance prediction output may be understood to define or be part of a structural parameter that defines the predicted protein structure 108.
[0094] In general, training the structure prediction system 100 to optimize the auxiliary loss can cause the structure prediction system to generate intermediate representations (e.g., MSA embeddings and pairwise embeddings) that encode information relevant to the ultimate goal of protein structure prediction. Therefore, optimizing the auxiliary loss can enable the structure prediction system 100 to more accurately predict protein structures.
[0095] The training engine may train the structure prediction system 100 on the training data for multiple training iterations, for example using a stochastic gradient descent training technique.
[0096] Figure 2 An example of a "grid transformer" neural network architecture 200 is illustrated. The grid transformer neural network is configured to receive a set of embeddings having initial values, process the set of embeddings through one or more residual blocks 202, and output updated values for the set of embeddings. For convenience, the following description will refer to a set of embeddings that are logically organized into a two-dimensional array of embeddings. For example, each embedding may be a pairwise embedding corresponding to a pair of amino acids in a protein, as described in reference to Figure 1In this example, the pairwise embeddings may be understood as being logically arranged in a two-dimensional array, wherein the embedding at position (i, j) in the array corresponds to the pairwise embedding of the amino acid pair at positions i and j in the amino acid sequence of the protein. In another example, each embedding may be an embedding of an amino acid from an MSA embedding. In this example, the embeddings may be understood as being logically arranged in a two-dimensional array, wherein the embedding at position (i, j) in the array corresponds to the embedding of the jth amino acid in the i-th amino acid sequence in the MSA.
[0097] Each residual block 202 of the grid transformer architecture includes a self-attention layer 204 implementing "row-wise self-attention", a self-attention layer 206 implementing "column-wise self-attention", and a transition block 208. To reduce computational resource consumption, the row-wise and column-wise self-attention layers update each input embedding using attention only on a proper subset of the input embeddings, as will be described in more detail below.
[0098] The row-wise self-attention layer 204 updates each embedding using attention only on those embeddings in the same row as the embeddings in the two-dimensional array of embeddings. For example, the row-wise self-attention layer updates the embedding 210 (i.e., indicated by the blackened square) using attention on the embeddings in the same row (i.e., indicated by the hatched area 212). The residual block 202 may also add the input to the row-wise self-attention layer 204 to the output of the row-wise self-attention layer 204.
[0099] The column-wise self-attention layer 206 updates each embedding using attention only on those embeddings in the same column as the embeddings in the two-dimensional array of embeddings. For example, the column-wise self-attention layer updates the embedding 214 (i.e., indicated by the blackened square) using self-attention on the embeddings in the same column (i.e., indicated by the hatched area 216). The residual block 202 may also add the input to the column-wise self-attention layer to the output of the column-wise self-attention layer.
[0100] The transition block 208 may update the embedding by processing the embedding using one or more linear neural network layers. The residual block 202 may also add the input to the transition block to the output of the transition block.
[0101] The grid converter architecture 200 may include a plurality of residual blocks. For example, in different implementations, the grid converter architecture 200 may include 5 residual blocks, 15 residual blocks, or 45 residual blocks.
[0102] Figure 3 An example multiple sequence alignment (MSA) embedded system 110 is shown. The MSA embedded system 110 is an example of a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described below are implemented.
[0103] The MSA embedding system 110 is configured to process the multiple sequence alignment (MSA) 108 to generate an MSA embedding 116. Generating an MSA embedding 116 including a corresponding embedding for each amino acid of each amino acid sequence in the MSA 108 may be computationally intensive, for example, in cases where the MSA includes a large number of long amino acid sequences. Therefore, in order to reduce computational resource consumption, the MSA embedding system 110 may generate an MSA embedding 116 for only a portion of the amino acid sequences in the entire MSA (referred to as a "clustered sequence" 302).
[0104] The MSA embedding system 110 may select clustered sequences 302 from the MSA 108, for example, by randomly selecting a predetermined number of amino acid sequences, such as 128 or 256 amino acid sequences, from the MSA 108. The remaining amino acid sequences in the MSA 108 that are not selected as clustered sequences 302 will be referred to herein as “additional sequences” 304. Typically, the number of clustered sequences 302 is selected to be significantly smaller than the number of additional sequences 304, for example, the number of clustered sequences 302 may be an order of magnitude (one order of magnitude) smaller than the number of additional sequences 304.
[0105] The MSA embedding system 110 may generate clustered embeddings 306 for the clustered sequence 302 and additional embeddings 308 for the additional sequences 304, for example, using the techniques described above with reference to equation (1).
[0106] The MSA embedding system 110 can then process the cluster embedding 306 and the additional embedding 308 using a crisscross attention neural network 310 to generate an updated cluster embedding 312. The crisscross attention neural network 310 includes one or more crisscross attention neural network layers. The crisscross attention neural network layer receives two independent sets of embeddings, one of which can be designated as "mutable" and the other can be designated as "static", and uses attention on the other (static) set of embeddings to update one (mutable) set of embeddings. That is, the crisscross attention layer can be understood as updating the mutable set of embeddings with information extracted from the static set of embeddings.
[0107] In one example, the crisscross attention neural network 310 may include a series of residual blocks, each of which includes two crisscross attention layers and (one) transition block. The first crisscross attention layer may receive the cluster embeddings 306 as a set of variable embeddings and the additional embeddings 308 as a set of static embeddings, and then update the cluster embeddings 306 with information from the additional embeddings 308. The second crisscross attention layer may receive the additional embeddings 308 as a set of variable embeddings and the cluster embeddings 306 as a set of static embeddings, and then update the additional embeddings 308 with information from the cluster embeddings 306. Each residual block may further include a transition block after the second crisscross attention layer, which updates the cluster embeddings using one or more linear neural network layers. For both the crisscross attention layers and the transition blocks, the residual blocks may add the input to the layer / block to the output of the layer / block.
[0108] In some cases, to improve computational efficiency, the criss-cross attention layer can implement a form of "global" criss-cross attention by determining a corresponding attention weight for each amino acid sequence (e.g., in a cluster or additional sequence), i.e., rather than determining a corresponding attention weight for each amino acid. In these cases, each amino acid in an amino acid sequence shares the same attention weight as the amino acid sequence.
[0109] The crisscross attention neural network 310 can be understood as enriching the cluster embedding 306 (and vice versa) with information from the additional embedding 308. After the operation performed by the crisscross attention neural network 310 is completed, the additional embedding 308 is discarded and only the cluster embedding 312 is used to generate the MSA embedding 116.
[0110] The MSA embedding system 110 uses a self-attention neural network 314 (e.g., with reference Figure 2 The cluster embedding 312 is processed by the grid transformer architecture described in , to generate the MSA embedding 116. Processing the cluster embedding 312 using a self-attention neural network 314 can enhance the cluster embedding by "sharing" information in the embedding within and between cluster sequences.
[0111] The MSA embedding system 110 may provide the output of the self-attention neural network 314 (ie, corresponding to the updated cluster embedding 306 ) as the MSA embedding 116 .
[0112] Figure 4 is a diagram of an unfolded protein and a folded protein. An unfolded protein is a random coil of amino acids. An unfolded protein undergoes protein folding and folds into a 3D configuration. Protein structures often include stable local folding patterns, such as alpha helices (e.g., as shown in 402) and beta sheets.
[0113] Figure 5 5 is a flow chart of an example process 500 for determining a predicted protein structure. For convenience, process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a protein structure prediction system (e.g., Figure 1 The protein structure prediction system 100) can perform process 500.
[0114] The system obtains a multiple sequence alignment of amino acid sequences for a specified protein (502).
[0115] The system processes the multiple sequence alignment to determine corresponding initial embeddings for each pair of amino acids in the protein amino acid sequence (504).
[0116] The system processes the initial embeddings of the amino acid pairs using a pairwise embedding neural network including a plurality of self-attention neural network layers to generate a final embedding for each amino acid pair (506). The respective attention layers of the pairwise embedding neural network can be configured to: (i) receive a current embedding for each amino acid pair, and (ii) use attention on the current embedding of the amino acid pair to update the current embedding for each amino acid pair. In some cases, the self-attention layers of the pairwise embedding neural network can alternate between row-wise self-attention layers and column-wise self-attention layers.
[0117] The system determines a predicted structure of the protein based on the final embeddings for each amino acid pair (508). For example, the system can process the final embeddings for the amino acid pairs to generate a corresponding "single" embedding for each amino acid in the protein, and process the single embeddings using a folded neural network to determine the predicted structure of the protein.
[0118] This specification uses the term "configuration" with respect to systems and computer program components. A system of one or more computers configured to perform a particular operation or action means that the system has installed thereon software, firmware, hardware, or a combination thereof that, in operation, causes the system to perform those operations or actions. One or more computer programs configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operations or actions.
[0119] The subject matter and implementation of the functional operations described in this specification may be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more thereof. The implementation of the subject matter described in this specification may be implemented as follows: one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-volatile storage medium, for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more thereof. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device.
[0120] The term "data processing apparatus" refers to data processing hardware and includes all kinds of apparatus, devices and machines for processing data, including, for example, a programmable processor, a computer or a multiprocessor or a computer. The apparatus may also be or further include a dedicated logic circuit, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to the hardware, the apparatus may optionally include code that creates an execution environment for a computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0121] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may correspond to a file in a file system, but need not correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple collaborative files, such as files storing portions of one or more modules, subroutines, or code. A computer program may be deployed to execute on one computer or on multiple computers located in one location or distributed in multiple locations and interconnected by a data communications network.
[0122] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0123] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by a special purpose logic circuit (e.g., FPGA or ASIC), or by a combination of a special purpose logic circuit and one or more programmed computers.
[0124] Computers suitable for executing computer programs can be based on general or special microprocessors or both, or any other kind of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more storage devices for storing instructions and data. The central processing unit and the memory can be supplemented or incorporated therein by a dedicated logic circuit. Typically, the computer will also include or be operably connected to one or more large-capacity storage devices (e.g., disks, magneto-optical disks or optical disks) for storing data to receive data from it or to transmit data to it or both. However, the computer does not need to have such a device. In addition, the computer can be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name a few.
[0125] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CDROM and DVD-ROM disks.
[0126] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser in response to a request received from a web browser on the user's device. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user as feedback.
[0127] A data processing device for implementing a machine learning model may also include, for example, dedicated hardware accelerator units for processing common and computationally intensive portions of machine learning training or production (i.e., inference, workloads).
[0128] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.
[0129] Implementations of the subject matter described in this specification may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0130] A computing system may include a client and a server. The client and the server are usually remote from each other and typically interact through a communication network. The relationship between the client and the server is generated due to computer programs running on the corresponding computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to the user device, for example, in order to display data to a user interacting with the device acting as a client and receive user input from the user. Data generated at the user device (e.g., the result of the user interaction) can be received from the device at the server.
[0131] Although this specification contains many specific implementation details, these should not be interpreted as limitations on the scope of any invention or the scope that can be claimed, but rather as descriptions of features that may be unique to a specific embodiment of a specific invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. In addition, although features may be described above as working in certain combinations, and even initially claimed as such, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may involve a sub-combination or a variant of a sub-combination.
[0132] Similarly, although operations are described in the claims and depicted in the drawings in a specific order, this should not be understood as requiring that such operations be performed in the specific order or in a sequential order, or that all operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0133] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the order of the sequence to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for determining the predicted structure of a protein specified by an amino acid sequence, performed by one or more data processing devices, the method comprising: include: Obtaining a multiple sequence alignment for proteins; determining a multiple sequence alignment embedding from the multiple sequence alignment using a multiple sequence alignment embedding system; Converting the multiple sequence alignment embedding into an array of N×N pairwise embeddings, where each pairwise embedding corresponds to a pair of amino acids among the N amino acids in the amino acid sequence of the protein; processing each pairwise embedding in the array of pairwise embeddings using a pairwise embedding neural network comprising a plurality of self-attention neural network layers including one or more row-wise self-attention neural network layers to generate an updated pairwise embedding, the processing comprising: for each row-wise self-attention neural network layer, updating the pairwise embedding using attention on pairwise embeddings of amino acid pairs located in the same row of the array; and The updated pairwise embeddings are processed using a folded neural network comprising one or more neural network layers to determine a predicted structure of the protein based on the updated pairwise embeddings, wherein the multiple sequence alignment embedding system, the pairwise embedding neural network, and the folded neural network are part of a structure prediction system that has been trained end-to-end to optimize a structure loss that characterizes the similarity between a predicted protein structure generated by the structure prediction system and a corresponding target protein structure that should be generated by the structure prediction system.
2. The method according to claim 1, in: One or more of the self-attention neural network layers are column-based self-attention neural network layers; as well as For each column-wise self-attention neural network layer, attention is used on the pairwise embeddings of amino acid pairs that are in the same column of the array to update the pairwise embeddings.
3. The method of claim 2, wherein the plurality of self-attention neural network layers embedded in pairs in a neural network comprises an alternating sequence of row-wise self-attention neural network layers and column-wise self-attention neural network layers.
4. The method according to any one of claims 1 to 3, wherein the predicted structure of the protein is determined based on the updated pairwise embeddings include: determining a corresponding initial embedding for each amino acid in the amino acid sequence of the protein based on the updated pairwise embeddings; The predicted structure of the protein is determined based on the initial embedding of each amino acid in the amino acid sequence.
5. The method according to any one of claims 1 to 3, wherein a multiple sequence alignment embedding system is used to determine the multiple sequence alignment embedding from the multiple sequence alignment. include: dividing the multiple sequence alignment into: (i) a set of clustered amino acid sequences and (ii) a set of additional amino acid sequences; generating: (i) an embedding of a set of clustered amino acid sequences and (ii) an embedding of a set of additional amino acid sequences; updating the embedding of the clustered amino acid sequence by processing a network input including: (i) the embedding of the clustered amino acid sequence and (ii) the embedding of the additional amino acid sequence using a criss-cross attention neural network; as well as The pairwise embeddings are determined based on the updated embeddings of the clustered amino acid sequences.
6. The method according to claim 5, wherein a crisscross attention neural network is used to process include: Updating the embeddings of the clustered amino acid sequences by taking into account (i) the embeddings of the clustered amino acid sequences and (ii) the embeddings of the additional amino acid sequences comprises repeatedly performing the following operations: Use attention on the embeddings of additional amino acid sequences to update the embeddings of clustered amino acid sequences; as well as Attention is used on the embeddings of clustered amino acid sequences to update the embeddings of additional amino acid sequences.
7. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1-6.
8. A system comprising one or more computers and one or more storage devices storing instructions, wherein when the instructions are executed by the one or more computers, the one or more computers are caused to perform corresponding operations of the method according to any one of claims 1 to 6.
9. A method for obtaining a ligand, wherein the ligand is a ligand of a drug or an industrial enzyme, the method include: obtaining a target amino acid sequence, wherein the target amino acid sequence is an amino acid sequence of a target protein; Using the target amino acid sequence as an amino acid sequence to perform the method according to any one of claims 1 to 6 to determine the predicted structure of the target protein; evaluating the interaction of one or more candidate ligands with the predicted structure of the target protein; and One or more of the candidate ligands are selected as the ligand depending on the results of the evaluation.
10. The method of claim 9, wherein the target protein comprises a receptor or an enzyme, and wherein the ligand is an agonist or antagonist of the receptor or enzyme.
11. A method for obtaining a polypeptide ligand, wherein the ligand is a ligand of a drug or an industrial enzyme, the method comprising obtaining amino acid sequences of one or more candidate polypeptide ligands; For each of the one or more candidate polypeptide ligands, using the amino acid sequence of the candidate polypeptide ligand as the amino acid sequence, performing the method according to any one of claims 1 to 6 to determine the predicted structure of the candidate polypeptide ligand; obtaining a target protein structure of a target protein; evaluating the interaction between the predicted structure of each of the one or more candidate polypeptide ligands and the target protein structure; as well as One of the one or more candidate polypeptide ligands is selected as the polypeptide ligand depending on the result of the evaluation.
12. The method of any one of claims 9-11, wherein evaluating the interaction of one of the candidate ligands comprises determining an interaction score for the candidate ligand, wherein the interaction score comprises a measure of the interaction between the candidate ligand and the target protein.
13. The method according to any one of claims 9-11, further comprising synthesizing the ligand.
14. The method of claim 13, further comprising testing the biological activity of the ligand in vitro.
15. A method for identifying the presence of a protein misfolding disease, include: Obtaining the amino acid sequence of the protein; performing the method according to any one of claims 1 to 6 using the amino acid sequence of the protein as a sequence of amino acids or amino acid residues to determine a predicted structure of the protein; Obtaining the structure of a version of a protein obtained from the human or animal body; comparing the predicted structure of the protein to the structure of a version of the protein obtained from humans or animals; as well as The presence of a protein misfolding disease is identified depending on the results of the comparison.