Prediction of pathogenicity of protein mutations using amino acid fraction profiles
By developing a neural network system, using multi-sequence alignment and embedded subnetwork to process protein data and calculate pathogenicity scores, the problem of pathogenicity prediction of mutant proteins is solved, and efficient and accurate pathogenicity assessment is achieved.
Patent Information
- Application Number
- CN202380081905.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2023-10-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively predict the pathogenicity of mutant proteins, affecting the accuracy of disease diagnosis and biological control.
A neural network system was developed to predict the pathogenicity of mutant proteins by generating pathogenicity scores. The system uses multi-sequence alignment (MSA) representation and embedded subnetwork to process network inputs to generate fractional distributions on the amino acid set and calculates pathogenic fractions based on the fractional difference of the original and alternative amino acids.
Accurate prediction of the pathogenicity of mutant proteins is achieved, the accuracy of disease diagnosis and biological control is improved, and the consumption of computing resources is reduced.
Smart Images

Figure CN120226084A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 415,117, filed on October 11, 2022, and U.S. Provisional Application No. 63 / 479,653, filed on January 12, 2023. The disclosures of the prior applications are considered part of the disclosure of this application and are incorporated herein by reference. Background Art
[0003] This specification relates to predicting the pathogenicity of mutant proteins.
[0004] A protein is specified by one or more amino acid sequences ("chains"). An amino acid is an organic compound that includes an amino functional group and a carboxyl functional group as well as a side chain (i.e., a group of atoms) that is unique to the amino acid. Protein folding refers to the physical process used to fold one or more amino acid sequences into a three - dimensional (3 - D) configuration. The structure of a protein defines the 3 - D configuration of the atoms in the protein's amino acid sequence after the protein has undergone protein folding. When amino acids are linked in a sequence by peptide bonds, they can be referred to as amino acid residues.
[0005] Predictions can be made using machine - learning models. A machine - learning model receives an input and generates an output based on the received input, such as a predicted output. Some machine - learning models are parametric models and generate an output based on the received input and the values of the model parameters. Some machine - learning models are deep models, and deep models employ multiple - layer models to generate an output for the received input. For example, a deep neural network is a deep machine - learning model that includes an output layer and one or more hidden layers, and each of the one or more hidden layers applies a non - linear transformation to the received input to generate an output. Summary of the Invention
[0006] This specification describes a neural - network system implemented as a computer program on one or more computers at one or more locations for predicting the pathogenicity of mutant proteins.
[0007] As used throughout this specification, the term "protein" can be understood to refer to any biomolecule specified by one or more amino acid sequences (or "chains"). For example, the term "protein" can refer to a protein domain, e.g., a part of an amino acid chain of a protein that can fold almost independently of the rest of the protein. As another example, the term "protein" can refer to a protein complex, i.e., a collection of multiple amino acid chains that co - fold into a protein structure.
[0008] A "mutant" protein can refer to a protein encoded by a mutated version of a gene (e.g., a mutated version of a deoxyribonucleic acid (DNA) nucleotide sequence). The mutant protein can have an amino acid sequence different from that of the "original" protein (i.e., the protein encoded by the non-mutated version of the gene). The original protein can be any reference protein, i.e., a protein selected as a reference; for example, it can but need not be defined by the wild-type version of the protein.
[0009] A gene can mutate in any of a variety of possible ways. For example, a gene can include a "missense mutation". A missense mutation can refer to a point mutation where the change of a single DNA nucleotide results in a codon encoding a different amino acid.
[0010] Since a mutant protein has an amino acid sequence different from that of the original (non-mutated) protein, the mutant protein can have properties different from those of the original protein. For example, the mutant protein can have different stability, different structure, or different binding interfaces from the original protein.
[0011] In some cases, one or more biological functions of a protein in a living organism (e.g., a human) depend on the properties of the protein, such as the stability, structure, or binding interface of the protein. Mutation of the protein can affect the properties of the protein and, in turn, can affect the ability of the protein to perform its biological function in the organism. Specifically, certain protein mutations can be "pathogenic", e.g., leading to a higher probability of disease, such as genetic diseases, such as sickle cell anemia, cystic fibrosis, cancer, etc.
[0012] Unless otherwise specified, an "amino acid set" can refer to the standard twenty (20) amino acids, e.g.: alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, or any suitable subset or superset of the standard twenty (20) amino acids, such as proteinogenic amino acids.
[0013] A "fractional distribution" or "amino acid fractional distribution" over an amino acid set can refer to data defining the respective fractions of each amino acid in the amino acid set.
[0014] A "multiple sequence alignment" (MSA) of a particular amino acid chain in a protein specifies the alignment of that amino acid chain with multiple additional amino acid chains (e.g., from other proteins, such as homologous proteins). More specifically, the MSA can define the correspondence between positions in the amino acid chain and corresponding positions in the multiple additional amino acid chains. An MSA of an amino acid chain can be generated, for example, by processing a database of amino acid chains using any suitable computational sequence alignment technique (e.g., progressive alignment construction). The amino acid chains in the MSA can be understood to have an evolutionary relationship. For example, each amino acid chain in the MSA can have a common ancestor. In an MSA of an amino acid chain, the correlation between amino acids in the amino acid chain can encode information related to predicting the structure of the amino acid chain.
[0015] An "embedding" of an entity (such as an amino acid pair) can refer to representing the entity as an ordered set of numerical values, such as a vector or matrix of numerical values.
[0016] The structure of a protein can be defined by a set of structural parameters. The set of structural parameters defining the protein structure can be represented as an ordered set of numerical values. Several examples of possible structural parameters for defining the protein structure will be described in more detail below.
[0017] In one example, for each amino acid in a protein, the structural parameters defining the protein structure include: (i) a position parameter, and (ii) a rotation parameter.
[0018] The position parameter of an amino acid can specify the predicted 3-D spatial position of a specified atom in the amino acid in the protein structure. The specified atom can be the alpha carbon atom in the amino acid, i.e., the carbon atom in the amino acid that is bonded to the amino functional group, the carboxyl functional group, and the side chain. The position parameter of an amino acid can be represented in any suitable coordinate system (e.g., a three-dimensional Cartesian coordinate system).
[0019] The rotation parameter of an amino acid can specify the predicted "orientation" of the amino acid in the protein structure. More specifically, the rotation parameter can specify a 3-D spatial rotation operation that, if applied to the coordinate system of the position parameter, results in three "backbone" atoms in the amino acid being in fixed positions relative to the rotated coordinate system. These three backbone atoms in the amino acid can refer to the sequence of the nitrogen atom, alpha carbon atom, and carbonyl carbon atom linked in the amino acid. The rotation parameter of an amino acid can be represented as, for example, an orthogonal matrix with a determinant equal to 1.
[0020] Generally, the position parameters and rotation parameters of an amino acid define an ego - centric reference frame of the amino acid. In this reference frame, the side chain of each amino acid can start from the origin, and along the first bond of the side chain (i.e., the alpha - carbon - beta - carbon bond), it can be in the defined direction.
[0021] In another example, the structural parameters defining a protein structure can include a "distance map" that characterizes the corresponding estimated distances (e.g., measured in angstroms) between each pair of amino acids in the protein. The distance map can characterize the estimated distance between amino - acid pairs, for example, by a probability distribution over the set of possible distances between amino - acid pairs.
[0022] In another example, the structural parameters defining a protein structure can define the three - dimensional (3D) spatial positions of each atom in each amino acid in the protein structure.
[0023] According to a first aspect, there is provided a method performed by one or more computers, the method comprising: generating a pathogenicity score that characterizes the likelihood that a protein mutation is a pathogenic mutation, where the mutation modifies the amino - acid sequence of the protein by replacing an original amino acid with a substitute amino acid at a mutation position in the amino - acid sequence of the protein, and where generating the pathogenicity score includes: generating a network input for a pathogenicity prediction neural network, where the network input includes an MSA representation representing a multiple - sequence alignment (MSA) of the protein; processing the network input using the pathogenicity prediction neural network to generate a score distribution over an amino - acid set; and generating the pathogenicity score based on the difference between: (i) the score of the original amino acid under the score distribution, and (ii) the score of the substitute amino acid under the score distribution.
[0024] In some implementations, the MSA representation includes corresponding embeddings for each position in the amino - acid sequence of each protein in the MSA; and generating the MSA representation includes masking the embedding in the MSA representation corresponding to the mutation position in the amino - acid sequence of the protein.
[0025] In some implementations, processing the network input using the pathogenicity prediction neural network to generate the score distribution over the amino - acid set includes: processing the MSA representation using an embedding sub - network of the pathogenicity prediction neural network to generate an updated MSA representation, where the updated MSA representation includes corresponding updated embeddings for each position in the amino - acid sequence of each protein in the MSA; and processing the updated embedding corresponding to the mutation position in the amino - acid sequence of the protein using a projection sub - network of the pathogenicity prediction neural network to generate the score distribution over the amino - acid set.
[0026] In some implementations, the embedding sub-network of the pathogenicity prediction neural network includes one or more self-attention neural network layers.
[0027] In some implementations, the embedding sub-network of the pathogenicity prediction neural network includes one or more row-wise or column-wise self-attention neural network layers.
[0028] In some implementations, the pathogenicity prediction neural network has been trained by operations including: training the pathogenicity prediction neural network to perform a protein unmasking task; and training the pathogenicity prediction neural network to perform a pathogenicity prediction task.
[0029] In some implementations, training the pathogenicity prediction neural network to perform the protein unmasking task includes, for each of a plurality of training proteins: generating a network input for the pathogenicity prediction neural network based on the training protein, where: the network input includes an MSA representation representing the MSA of the training protein; and the MSA representation includes respective masked embeddings corresponding to each of one or more masked positions in the amino acid sequence of the training protein; processing the network input of the pathogenicity prediction neural network to generate a network output that defines a respective prediction for the amino acid identity at each masked position in the amino acid sequence of the training protein; and backpropagating the gradient of a masking loss through the pathogenicity prediction neural network, where the masking loss measures the accuracy of the respective predictions generated by the pathogenicity prediction neural network for each masked position.
[0030] In some implementations, the method further includes training the pathogenicity prediction neural network to perform a protein structure prediction task.
[0031] In some implementations, training the pathogenicity prediction neural network to perform the protein structure prediction task includes, for each of a plurality of training proteins: generating a network input for the pathogenicity prediction neural network based on the training protein, where the network input includes an MSA representation representing the MSA of the training protein; using the pathogenicity prediction neural network to process the network input to generate an intermediate output of the pathogenicity prediction neural network; using a folding neural network to process the intermediate output of the pathogenicity prediction neural network to generate a set of structural parameters defining the predicted structure of the training protein; and backpropagating the gradient of a structure loss through the folding neural network into the pathogenicity prediction neural network, where the structure loss measures the error of the predicted structure of the training protein.
[0032] In some implementations, the pathogenicity prediction neural network is pre-trained to perform the protein de-masking task and the protein structure prediction task before being trained to perform the pathogenicity prediction task.
[0033] In some implementations, the pathogenicity prediction neural network has been trained on at least two training phases including a first training phase and a second training phase; wherein the first training phase includes: training the pathogenicity prediction neural network on a set of pathogenicity training examples, where each pathogenicity training example defines: (i) the amino acid sequence of a training protein, (ii) one or more mutations of the amino acid sequence of the training protein, and (iii) a corresponding target pathogenicity score for each mutation of the one or more mutations; determining a corresponding filtering score for each training example; and filtering the set of pathogenicity training examples using these filtering scores.
[0034] In some implementations, for each training example, determining the filtering score of the training example includes: for each mutation included in the training example, using the pathogenicity prediction neural network to determine the predicted pathogenicity score of the mutation; and for each mutation included in the training example, determining an error score for the mutation based on the error between: (i) the predicted pathogenicity score of the mutation, and (ii) the target pathogenicity score of the mutation; and determining the filtering score of the training example at least in part based on these error scores of the mutations included in the training example.
[0035] In some implementations, filtering the set of pathogenicity training examples using these filtering scores includes: determining a probability distribution on the set of pathogenicity training examples based on these filtering scores; sampling a subset of these pathogenicity training examples from the set of pathogenicity training examples according to the probability distribution; and removing the sampled subset of pathogenicity training examples from the set of pathogenicity training examples.
[0036] In some implementations, the second training phase includes: retraining the pathogenicity prediction neural network on the filtered set of pathogenicity training examples.
[0037] In some implementations, the method further includes: determining whether the pathogenicity score of the mutation meets a threshold; and in response to determining that the pathogenicity score of the mutation meets the threshold: classifying the mutation as a pathogenic mutation.
[0038] In some implementations, determining that the pathogenicity score of the mutation meets the threshold includes: determining that the pathogenicity score of the mutation exceeds the threshold.
[0039] In some implementations, the method further includes using the pathogenicity score to identify the cause of a disease in a subject.
[0040] In some implementations, the methods described herein are used to obtain a mutated protein for pest control, for example, including: determining the pathogenicity score of a plurality of mutations of a reference protein, each mutation defining a corresponding mutated protein; and selecting, based on the determined pathogenicity score, one of the mutations to define the mutated protein for pest control.
[0041] In some implementations, the methods described herein are used to screen for the presence of a protein with a pathogenic mutation in one or more living organisms, for example, including: obtaining the amino acid sequence of the protein for each of the organisms; determining the pathogenicity score of the mutation for each amino acid sequence that includes a mutation in the obtained amino acid sequence; and processing each determined pathogenicity score to determine whether the mutation is pathogenic.
[0042] In some implementations, the methods described herein are used to determine the pathogenicity level of a bacterium or virus to a living organism, for example, including: obtaining the amino acid sequence of an artificial protein, where the artificial protein is made inside the living organism due to the living organism being infected by the bacterium or virus, and where the artificial protein is a mutation of a naturally occurring protein in the organism; determining the pathogenicity score of the mutation of the naturally occurring protein; and determining the pathogenicity level of the bacterium or virus based on the pathogenicity score.
[0043] In some implementations, the method further includes providing the pathogenicity score for display on a user interface.
[0044] According to another aspect, a system is provided that includes: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, where the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the methods described herein.
[0045] According to another aspect, one or more non-transitory computer storage media are provided, the computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the methods described herein.
[0046] According to another aspect, a method is provided that includes: maintaining a mutation-pathogenicity database that includes data defining a respective pathogenicity score for each of one or more mutations of each of a plurality of proteins, where the pathogenicity score has been generated using the methods described herein; obtaining genetic material from a subject; identifying, based on the genetic material from the subject, one or more proteins encoded in the genetic material of the subject that each have one or more mutations; and determining a diagnosis for the subject based on: (i) the mutation-pathogenicity database, and (ii) the protein mutations identified from the genetic material of the subject.
[0047] Obtaining genetic material from a subject can include, for example, obtaining a biological sample from the subject, such as, for example, a blood sample, or a saliva sample, or an oral swab, or a tissue biopsy. The genetic material, such as deoxyribonucleic acid (DNA), can then be extracted from the biological sample. A genetic test, such as, for example, DNA sequencing, DNA microarray analysis, polymerase chain reaction (PCR) analysis, or next-generation sequencing (NGS), can be performed to identify one or more proteins encoded in the genetic material of the subject that each have one or more mutations compared to a reference genome.
[0048] In some cases, the mutation-pathogenicity database includes data characterizing a plurality of proteins in the human proteome.
[0049] In some cases, the mutation-pathogenicity database includes data characterizing all proteins in the human proteome.
[0050] In some cases, the subject is a human subject.
[0051] In some cases, determining the diagnosis for the subject includes: determining that a particular protein mutation identified from the genetic material of the subject is associated in the mutation-pathogenicity database with a pathogenicity score that meets a threshold; and in response, diagnosing the subject with a medical condition associated with the particular protein mutation. The threshold can be any suitable threshold, such as 0.5, 0.8, 0.9, or 0.99.
[0052] Particular embodiments of the subject matter described in this specification can be implemented so as to achieve one or more of the following advantages.
[0053] This specification describes a pathogenicity prediction neural network that can generate pathogenicity scores that predict whether a protein mutation is pathogenic, e.g., whether it is likely to cause a disease. To determine whether a mutation at a position in the amino acid sequence of a protein is pathogenic, the pathogenicity prediction neural network can process the MSA of the protein to generate a score distribution over the set of amino acids. A pathogenicity score for the mutation can then be generated based on the difference between (i) the score of the original amino acid at that position under the score distribution, and (ii) the score of the substituted amino acid at that position under the score distribution. Processing the MSA enables the pathogenicity prediction neural network to leverage the protein evolutionary history that encodes information relevant to the pathogenicity of the mutation (as will be described in more detail later in this specification). Additionally, calculating the pathogenicity score as the difference between the amino acid scores assigned to the original and substituted amino acids encodes an inductive bias, namely, that assessing pathogenicity involves comparing the original amino acid to the substituted amino acid.
[0054] The pathogenicity prediction neural network can be pre-trained (or co-trained) to perform an auxiliary task of protein structure prediction. The structure of a protein plays an important role in determining whether a protein mutation will be pathogenic, and training the pathogenicity prediction neural network to perform the auxiliary protein structure prediction task can enable the pathogenicity prediction neural network to implicitly reason about the protein structure. Thus, training the pathogenicity prediction neural network to perform the auxiliary protein structure prediction task can allow the pathogenicity prediction neural network to be trained on less training data, or in fewer training iterations, or both, thereby reducing the consumption of computing resources.
[0055] The pathogenicity prediction neural network can be pre-trained (or co-trained) to perform an auxiliary task of protein demasking. The demasking task requires the pathogenicity prediction neural network to decode the identity of the amino acid at a masked position in the protein amino acid sequence based on context information in the remaining unmasked portion of the amino acid sequence. Learning to effectively perform the demasking task can enable the pathogenicity prediction neural network to implicitly reason about protein biochemistry, and thereby can allow the pathogenicity prediction neural network to be trained on less training data, or in fewer training iterations, or both, thereby reducing the consumption of computing resources.
[0056] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 An example pathogenicity prediction system is shown.
[0058] Figure 2 Shows an example architecture of a pathogenicity prediction neural network.
[0059] Figure 3 Shows an example training system for training a pathogenicity prediction neural network.
[0060] Figure 4 Is a flowchart of an example process for generating a pathogenicity score that characterizes the likelihood that a protein mutation is a pathogenic mutation.
[0061] Figure 5 Is a flowchart of an example process of a multi-stage process for training a pathogenicity prediction neural network.
[0062] Figure 6A Shows an example of an MSA representation of an MSA of a protein.
[0063] Figure 6B Shows an example of a score distribution on an amino acid set generated by a pathogenicity prediction neural network.
[0064] Figure 7 Shows an example protein structure prediction system.
[0065] Figure 8 Shows an example architecture of an embedding neural network.
[0066] Figure 9 Shows an example architecture of an update block of an embedding neural network.
[0067] Figure 10 Shows an example architecture of an MSA update block.
[0068] Figure 11 Shows an example architecture of a pairwise update block.
[0069] Figure 12 Shows an example architecture of a folding neural network.
[0070] Figure 13 Shows the torsion angles between bonds in an amino acid.
[0071] Figure 14 Is a diagram of an unfolded protein and a folded protein.
[0072] Figure 15 Is a flowchart of an example process for predicting a protein structure.
[0073] Figure 16 Shows an example process for generating an MSA representation of an amino acid chain in a protein.
[0074] Figure 17Shows an example process for generating corresponding paired embeddings for each amino acid pair in a protein.
[0075] Figures 18A to 18B Shows the performance (e.g., classification accuracy) of the pathogenicity prediction system.
[0076] In the various figures, the same reference numerals and names indicate the same elements. Detailed Description
[0077] Figure 1 Shows an example pathogenicity prediction system 120. The pathogenicity prediction system 120 is an example of a system implemented as a computer program on one or more computers at one or more locations, where the systems, components, and techniques described below are implemented.
[0078] The pathogenicity prediction system 120 is configured to process data defining the following to generate a pathogenicity score 134: (i) the amino acid sequence of a protein 122, and (ii) a mutation 132 in the amino acid sequence of the protein 122. The pathogenicity score 134 characterizes the likelihood that the mutation 132 is a pathogenic mutation (e.g., a mutation that will result in a higher likelihood of the organism having a disease). The amino acid sequence of the protein can be any reference sequence, such as a wild-type sequence. The organism can be an animal (e.g., a human) or a plant or a microorganism (such as a bacterium).
[0079] The mutation 132 can specify: (i) the position of the mutation in the amino acid sequence of the protein 122, and (ii) the alternative amino acid at the position of the mutation in the amino acid sequence (i.e., under the mutation, it will replace the original amino acid at the position of the mutation in the amino acid sequence). The mutation 132 can be, for example, the result of a missense mutation in the gene encoding the protein 122.
[0080] The pathogenicity score 134 can be represented as, for example, a scalar numerical value, e.g., where a higher value of the pathogenicity score can indicate a higher likelihood that the protein mutation 122 is pathogenic.
[0081] The pathogenicity score 134 generated by the pathogenicity prediction system 120 for a protein mutation can be used in any of a variety of possible applications. For example, the pathogenicity score 134 can be used to identify specific genes that cause genetic diseases. More specifically, a subject suspected of having a genetic disease can be genotyped, and the subject's genetic data can be processed to identify a set of mutated protein-coding genes in the subject's genome, for example, by comparing the gene sequence of the protein to a reference gene sequence. The pathogenicity prediction system 120 described in this specification can be used to generate a corresponding pathogenicity score for each protein mutation caused by each mutated protein-coding gene in the set of mutated protein-coding genes. Mutated protein-coding genes associated with a high pathogenicity score (e.g., a score above a threshold or relatively higher than the scores of other proteins) can then be designated as candidate causes of the subject's genetic disease. Some further applications of the system will be described later.
[0082] The pathogenicity prediction system 120 includes a pathogenicity prediction neural network 126 and an evaluation engine 130, which will be described in detail separately below.
[0083] To generate the pathogenicity score 134 for the mutation 132, the pathogenicity prediction system 120 obtains data defining a multiple sequence alignment (MSA) of the protein 122 and generates an MSA representation 124 that represents the MSA of the protein 122. (Throughout this document, the MSA of a protein should be understood to include the protein itself; that is, the MSA of a protein includes: (i) the protein (the reference version of the protein), and (ii) a set of other proteins that are homologous to the (reference) protein).)
[0084] The MSA representation 124 includes a set of embeddings that includes corresponding embeddings for each position in the amino acid sequence of each protein included in the MSA. Example techniques for generating the MSA representation 124 of the protein 122 are described in more detail below with reference to Figure 7 More detailed description of example techniques for generating the MSA representation 124 of the protein 122.
[0085] Optionally, the pathogenicity prediction system 120 can "mask" the embeddings in the MSA representation 124 that correspond to the mutation positions (i.e., as specified by the mutation 132) in the amino acid sequence of the (reference) protein. A masked embedding refers to replacing the embedding with a default (masked) embedding (e.g., a predefined embedding (e.g., having predefined values in each entry, such as zeros)) or a random embedding (e.g., the values in each entry of the embedding are sampled from a probability distribution (e.g., a Gaussian distribution)). Generally speaking, masking an embedding refers to changing the embedding.
[0086] The MSA encoding of a protein encodes information related to predicting the pathogenicity of protein mutations. For example, for each position in the amino acid sequence of a protein, the MSA encoding indicates information about the frequency at which the amino acid at that position has changed to a different amino acid during the evolutionary history of the protein. If an amino acid at a position in a protein amino acid sequence has frequently mutated during the evolutionary history of the protein, then a mutation at that position may be less likely to be pathogenic. Conversely, if an amino acid at a position in a protein amino acid sequence has rarely mutated during the evolutionary history of the protein, then a mutation at that position may be more likely to be pathogenic. Intuitively, if a mutation at a position in a protein amino acid sequence causes pathogenicity, then evolutionary pressure may cause the frequency of the amino acid sequence at that position where the mutation occurs to decrease in the evolutionary history of the protein.
[0087] The pathogenicity prediction system 120 uses the pathogenicity prediction neural network 126 to process the network input including the MSA representation 124 to generate a score distribution 128 over the set of amino acids. An example architecture of the pathogenicity prediction neural network 126 will be described in more detail below with reference to Figure 2 More detailed description of the example architecture of the pathogenicity prediction neural network 126.
[0088] The evaluation engine 130 generates a pathogenicity score 134 for the mutation 132 at the mutation position based on the amino acid score distribution 128 generated by the pathogenicity prediction neural network 126. For example, the evaluation engine 130 can generate the pathogenicity score 134 for the mutation 132 based on the difference (e.g., by determining a measure of the difference) between: (i) the score of the original amino acid at the mutation position under the score distribution 128 (i.e., as given by it), and (ii) the score of the alternative amino acid at the mutation position under the score distribution 128. The original amino acid at the mutation position refers to the amino acid at the mutation position in the original (i.e., unmutated) amino acid sequence of the protein 122. That is, the evaluation engine 130 can process the MSA representation with a masked embedding at the mutation position (as described above) to determine the score distribution 128, and then can generate the pathogenicity score 134 for the mutation 132 based on the scores of the original amino acid and the alternative amino acid in the score distribution (e.g., based on the difference between these scores).
[0089] In some implementations, the evaluation engine 130 can use the pathogenicity score 134 to classify the mutation 132 as "pathogenic" or "benign". For example, if the pathogenicity score 134 of the mutation 132 meets (e.g., exceeds) a threshold, the evaluation engine 130 can classify the mutation 132 as pathogenic. If the pathogenicity score 134 of the mutation does not meet (e.g., does not exceed) the threshold, the evaluation engine 130 can classify the mutation 132 as benign. Generally, the system can be configured to indicate pathogenic mutations with relatively large scores or relatively small scores.
[0090] The pathogenicity prediction system 120 can provide a pathogenicity score 134, a mutation classification 132, or both, for example, for storage in a memory, or for presentation to a user via a user interface, or for use by a downstream system, or for any other suitable purpose.
[0091] Figure 2 An example architecture of the pathogenicity prediction neural network 126 is shown. The pathogenicity prediction neural network 126 can be included in a pathogenicity prediction system, such as the pathogenicity prediction system 120 described with reference to Figure 1 the pathogenicity prediction system 120 described above.
[0092] The pathogenicity prediction neural network 126 is configured to receive a network input including an MSA representation 124 of an MSA including a protein. The MSA representation 124 includes a set of embeddings that includes a respective embedding corresponding to each position in the amino acid sequence of each protein included in the MSA. In an implementation, the embedding corresponding to a "mutation" position in the protein amino acid sequence is a masked embedding.
[0093] The pathogenicity prediction neural network 126 processes the network input to generate an amino acid score distribution 128. The pathogenicity prediction system 120 can use the amino acid score distribution 128 to generate a pathogenicity score that characterizes the pathogenic likelihood of a mutation at a mutation position in the protein amino acid sequence, as described above with reference to Figure 1 described above.
[0094] The pathogenicity prediction neural network 126 includes an embedding neural network 136 (also referred to as an embedding subnetwork) and a projection neural network 140 (also referred to as a projection subnetwork), which will be described in detail separately below.
[0095] The embedding neural network 136 is configured to process the network input of the pathogenicity prediction neural network 126 (including the MSA representation 124) to generate an updated MSA representation 138. The updated MSA representation 138 includes an updated version of each embedding included in the original MSA representation 124, including an updated version of the embedding corresponding to the mutation position in the protein amino acid sequence.
[0096] The embedding neural network 136 can have any suitable neural network architecture such that the embedding neural network 136 can perform the functions described thereof, for example, processing the MSA representation 124 to generate the updated MSA representation 138. Specifically, the embedding neural network 136 can include any suitable number (e.g., 5 layers, 10 layers, or 15 layers) and any suitable type of neural network layers (e.g., self-attention layers, cross-attention layers, fully connected layers, convolutional layers, etc.) connected in any suitable configuration (e.g., as a linear sequence of layers).
[0097] The following will refer to Figure 8 describe an example implementation of an embedded neural network. Referring to Figure 8 the example embedded neural network described is configured to receive an input that additionally includes a set of "pairwise embeddings" (i.e., in addition to the MSA representation 124), and is configured to generate an output that additionally includes an updated set of pairwise embeddings (i.e., in addition to the updated MSA representation 138), as will be described in more detail below with reference to Figure 8 Generally, pairwise embeddings can be embeddings obtained from pairs of amino acids in an amino acid sequence.
[0098] The projection neural network 140 is configured to process the updated embeddings corresponding to the mutation positions in the amino acid sequence of the protein from the updated MSA representation 138 to generate a score distribution over the set of amino acids.
[0099] The projection neural network 140 can have any suitable neural network architecture such that the projection neural network 140 can perform the functions described for it, e.g., processing embeddings from the updated MSA representation to generate an amino acid score distribution. Specifically, the projection neural network 140 can include any suitable number (e.g., 1 layer, 5 layers, 10 layers, or 15 layers) and any suitable type of neural network layers (e.g., linear layers, self-attention layers, cross-attention layers, fully connected layers, convolutional layers, etc.) connected in any suitable configuration (e.g., as a linear sequence of layers).
[0100] Figure 3 An example training system 142 is shown. The training system 142 is an example of a system implemented as a computer program on one or more computers in one or more locations, where the systems, components, and techniques described below are implemented.
[0101] The training system 142 is configured to train the pathogenicity prediction neural network 126 on a training example set 144. Generally, the pathogenicity prediction neural network is trained to perform a pathogenicity prediction task; it can also be trained to perform a protein demasking task (removing the influence of masked embeddings) and / or a protein structure prediction task (predicting the structure of the training protein).
[0102] Each training example can be, for example: a pathogenicity training example, a masked training example, or a structure training example. Including masked training examples and structure training examples can be useful but is not necessary.
[0103] Pathogenicity training example definition: (i) the amino acid sequence of a training protein, (ii) one or more mutations of the training protein, and (iii) the corresponding target pathogenicity score for each mutation among the one or more mutations. Each mutation among the mutations can specify: (i) the mutation position in the amino acid sequence of the training protein, and (ii) the alternative amino acid at the mutation position in the amino acid sequence of the training protein. Some ways of generating the target pathogenicity score will be described later. The training protein can include a protein of the organism for which the pathogenicity score is generated, but is not limited to such a protein.
[0104] Masked training example definition: (i) the amino acid sequence of a training protein, and (ii) one or more positions in the amino acid sequence of the training protein that are designated as "masked" positions.
[0105] Structure training example definition: (i) the amino acid sequence of a training protein, and (ii) the structure of the training protein (e.g., represented by a set of structure parameters).
[0106] The training system 142 can train the pathogenicity prediction neural network 126 on the following: (i) pathogenicity training examples to optimize the pathogenicity loss 154, (ii) masked training examples to optimize the masking loss 152, and (iii) structure training examples to optimize the structure loss 150, as will be described in more detail next.
[0107] To train the pathogenicity prediction neural network 126 on pathogenicity training examples, the training system 142 can provide the pathogenicity prediction neural network 126 with a network input that characterizes the corresponding training protein. The network input can include an MSA representation of the MSA of the training protein, where the MSA representation includes corresponding embeddings for each position in the amino acid sequence of each protein in the MSA. For each mutation specified by the training example, the training system 142 masks the embedding corresponding to the mutation position of the mutation in the amino acid sequence of the training protein.
[0108] The training system 142 processes the network input of the training protein using the pathogenicity prediction neural network 126 to generate a corresponding amino acid score distribution for the mutation position of each mutation specified by the training example. For each mutation specified by the training example, the training system then generates a corresponding predicted pathogenicity score 134 based on the amino acid score distribution generated by the pathogenicity prediction neural network 126 for the mutation position. (Example techniques for generating the pathogenicity score of a mutation from an amino acid score distribution were described above Figure 1 ).
[0109] The training system 142 evaluates the pathogenicity loss 154 for each mutation specified by a pathogenic training example and then determines the overall pathogenicity loss of the pathogenic training example as a combination (e.g., average or sum) of the pathogenicity losses of each mutation. The training system 142 can evaluate the pathogenicity loss of a mutation based on the error (e.g., hinge loss, cross-entropy loss, or any other suitable loss) between: (i) the predicted pathogenicity score of the mutation, and (ii) the target pathogenicity score of the mutation. The target pathogenicity score of a mutation can be obtained as described later with reference to Figure 5 As an example, the target pathogenicity score can have a value indicative of the prevalence of the mutation in a population or group of organisms. Then, the organism or group of organisms can include such organisms: for which the pathogenicity prediction system 120 (at inference time) is used to generate a pathogenicity score - but since there can be substantial similarity between different organisms, the pathogenicity prediction system 120 does not have to be trained on data from the same organism that it is used on at inference time.
[0110] The training system 142 can, for example, use backpropagation to determine the gradient of the pathogenicity loss 154 with respect to the parameters of the pathogenicity prediction neural network 126. The training system 142 can then use the gradient (e.g., using a suitable gradient descent optimization technique) to adjust the current values of the parameters of the pathogenicity prediction neural network 126.
[0111] To train the pathogenicity prediction neural network 126 on a masked training example, the training system can provide the pathogenicity prediction neural network 126 with network inputs that characterize the corresponding training protein. The network inputs can include an MSA representation representing the MSA of the training protein, where the MSA representation includes corresponding embeddings corresponding to each position in the amino acid sequence of each protein in the MSA. The training system 142 can mask each embedding in the MSA representation 124 that corresponds to a position in the amino acid sequence of the training protein that has been specified as a masked position by the masked training example.
[0112] The training system 142 then processes the network inputs of the training protein using the pathogenicity prediction neural network 126 to generate a corresponding amino acid score distribution corresponding to each masked position in the amino acid sequence of the training protein. The training system 142 evaluates the masking loss for each masked position and then determines the overall masking loss of the training example as a combination (e.g., average or sum) of the masking losses of each masked position. The training system 142 can evaluate the masking loss of a masked position based on the score of the amino acid at the masked position in the training protein under the score distribution at the masked position (e.g., based on the change in score due to masking). The masking loss can measure the accuracy of the corresponding prediction generated by the pathogenicity prediction neural network for each masked position.
[0113] The training system 142 can determine, for example, using backpropagation, the gradient of the masking loss 152 with respect to the parameters of the pathogenicity prediction neural network 126. The training system 142 can then adjust the current values of the parameters of the pathogenicity prediction neural network 126 using the gradient (e.g., using a suitable gradient descent optimization technique).
[0114] By training the pathogenicity prediction neural network 126 against the masking loss 152, the training system 142 trains the pathogenicity prediction neural network 126 to perform a demasking task. The demasking task requires the pathogenicity prediction neural network to decode the identity of the amino acid at a masked position in a protein amino acid sequence based on context information in the remaining unmasked portion of the amino acid sequence. Learning to perform the demasking task effectively can implicitly encode an understanding of protein biochemistry in the parameter values of the pathogenicity prediction neural network and can thus improve the performance of the pathogenicity prediction neural network 126 on the pathogenicity prediction task.
[0115] To train the pathogenicity prediction neural network 126 on a structural training example, the training system 142 can provide the pathogenicity prediction neural network 126 with a network input that characterizes the corresponding training protein. The network input can include an MSA representation that represents the MSA of the training protein. The training system 142 processes the network input of the training protein using the pathogenicity prediction neural network 126 and provides an intermediate output of the pathogenicity prediction neural network to the folding neural network 146. (An intermediate output of a neural network refers to an output produced by one or more hidden layers of the neural network; any layer can be used). The folding neural network 146 is configured to process the intermediate output of the pathogenicity prediction neural network 126 to generate data that defines the predicted structure 148 of the training protein.
[0116] The folding neural network 146 can be configured to receive any suitable intermediate output of the pathogenicity prediction neural network 126. For example, the pathogenicity prediction neural network 126 can include an embedding neural network, as described in reference Figure 8 and the folding neural network 146 can be configured to receive some or all of the outputs of the embedding neural network. For an embedding neural network having the architecture described in reference Figure 8 below, the folding neural network can receive the updated MSA representation and / or the updated pairwise embeddings generated by the embedding neural network.
[0117] The folding neural network can have any suitable neural network architecture such that the folding neural network is capable of performing the functions described thereof, particularly processing the intermediate output of the pathogenicity prediction neural network to generate a predicted protein structure. Specifically, the folding neural network can include any suitable number (e.g., 5 layers, 10 layers, or 15 layers) and any suitable type of neural network layers (e.g., self-attention layers, cross-attention layers, fully connected layers, convolutional layers, etc.) connected in any suitable configuration (e.g., as a linear sequence of layers). An example implementation of the folding neural network is described below with reference to Figure 12 describe an example implementation of the folding neural network.
[0118] The structural loss 150 of the structural training example measures the error between: (i) the predicted protein structure 148 generated by the folding neural network 146, and (ii) the target protein structure specified by the structural training example. An example of the structural loss is described below, for example, with reference to equations (2)-(4). The training system 142 determines the gradient of the structural loss 150 with respect to the parameters of the pathogenicity prediction neural network 126 and the folding neural network 146, for example, using backpropagation. The training system 142 can then adjust the current values of the parameters of the pathogenicity prediction neural network 126 and the folding neural network 146 using the gradient (e.g., by appropriate gradient descent optimization techniques). That is, the training system 142 backpropagates the gradient of the structural loss 150 through the folding neural network 146 into the pathogenicity prediction neural network 126.
[0119] After training the pathogenicity prediction neural network 126, the folding neural network 146 is not required. Thus, the folding neural network 146 can be discarded after training the pathogenicity prediction neural network 126.
[0120] By training the pathogenicity prediction neural network 126 and the folding neural network 146 with respect to the structural loss 150, the training system 142 trains the pathogenicity prediction neural network 126 and the folding neural network 146 to jointly perform the task of protein structure prediction. Learning to effectively perform the protein structure prediction task can implicitly encode the understanding of protein biochemistry in the parameter values of the pathogenicity prediction neural network, and thereby improve the performance of the pathogenicity prediction neural network on the pathogenicity prediction task.
[0121] More specifically, the structure of a protein encodes information about the possible pathogenicity of protein mutations. For example, a protein can fold to nest hydrophobic amino acids inside the protein while exposing hydrophilic amino acids outside the protein. A protein mutation that places a hydrophobic amino acid outside the protein may significantly alter the properties of the protein, such as causing the protein to become unstable or changing the shape of the protein, and thus potentially being pathogenic. Training the pathogenicity prediction neural network 126 against the structure loss 150 enables the pathogenicity prediction neural network 126 to reason about the protein structure and thus perform pathogenicity prediction more accurately.
[0122] In some implementations, the training system 142 first trains the pathogenicity prediction neural network 126 to optimize the masking loss 152 and the structure loss 150. Then, the training system 142 disables the masking loss 152 and the structure loss 150 and trains the pathogenicity prediction neural network 126 only against the pathogenicity loss 154. That is, the training system 142 can pre-train the pathogenicity prediction neural network 126 against the masking loss 152 and the structure loss 150 and then fine-tune the pathogenicity prediction neural network 126 against the pathogenicity loss. In some cases, during fine-tuning, the training system 142 can train the pathogenicity prediction neural network against both the pathogenicity loss and the structure loss 150 (i.e., rather than disabling the structure loss 150).
[0123] Pre-training the pathogenicity prediction neural network against the masking loss 152 and the structure loss 150 before fine-tuning against the pathogenicity loss 154 can improve the performance of the pathogenicity prediction neural network on the pathogenicity prediction task. Specifically, pre-training can generate an effective initialization for the parameter values of the pathogenicity prediction neural network and can enable the pathogenicity prediction neural network to achieve acceptable performance on the pathogenicity prediction task using little training data and fewer training iterations.
[0124] In some cases, the pathogenicity training examples used to train the pathogenicity prediction neural network may be "noisy", e.g., may be generated by a process that produces a non-negligible number of pathogenicity training examples with inaccurate target pathogenicity scores. To address this issue, the training system 142 can filter (i.e., remove) the noisy pathogenicity training examples from the training data during the training process and thus regularize and improve the training of the pathogenicity prediction neural network. An example of a multi-stage training process for training the pathogenicity prediction neural network 126 while filtering noisy pathogenicity training examples from the training data is described in more detail below. Figure 5 More specifically, an example of a multi-stage training process for training the pathogenicity prediction neural network 126 while filtering noisy pathogenicity training examples from the training data is described.
[0125] Figure 4is a flowchart of an example process 158 for generating a pathogenicity score that characterizes the likelihood that a protein mutation is a pathogenic mutation. A mutation modifies the amino acid sequence of a protein by replacing the original amino acid with a substituted amino acid at the mutation position in the amino acid sequence of the protein. In an implementation, the process involves obtaining data representing the mutation. For convenience, process 158 will be described as being performed by a system of one or more computers located in one or more locations. For example, a pathogenicity prediction system (e.g., Figure 1 's pathogenicity prediction system 120) can perform process 158 after being appropriately programmed according to this specification.
[0126] The system generates a network input (160) for a pathogenicity prediction neural network. The network input includes an MSA representation representing a multiple sequence alignment (MSA) of the protein. In an implementation, the embedding corresponding to the mutation position in the protein amino acid sequence in the masked MSA representation is masked.
[0127] The system processes the network input using a (trained) pathogenicity prediction neural network to generate a score distribution (162) over an amino acid set. In an implementation, the pathogenicity prediction neural network has been trained to optimize (minimize) a pathogenicity loss that characterizes the difference between the pathogenicity of a protein mutation and the prediction of the pathogenicity of the protein mutation by the pathogenicity prediction system (i.e., derived from the score distribution over the amino acid set, more specifically from the pathogenicity score of the protein mutation).
[0128] The system generates a pathogenicity score based on the difference between: (i) the score of the original amino acid under the score distribution, and (ii) the score of the substituted amino acid under the score distribution (164).
[0129] Figure 5 is a flowchart of an example process 166 for implementing a multi-stage process for training a pathogenicity prediction neural network. For convenience, process 166 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system (e.g., Figure 3 's training system 142) can perform process 166 after being appropriately programmed according to this specification.
[0130] The system generates a set of pathogenicity training examples (168). Each pathogenicity training example defines: (i) the amino acid sequence of a training protein, (ii) one or more mutations of the training protein, and (iii) the corresponding target pathogenicity score for each of the one or more mutations.
[0131] In some implementations, the system can generate or obtain a set of pathogenicity training examples based on a dataset representing the frequency of occurrence of protein mutations in a population of organisms (e.g., a human population, or a population of primates, or a population of mammals or other animals, or a plant population, or a microbial population). Additionally or alternatively, the system can use known biological data (e.g., from a suitable database) to obtain a measure of the pathogenicity of the mutations.
[0132] Specifically, the system can identify each protein mutation that occurs in at least a threshold percentage of the organisms in the population as a "benign" mutation, i.e., a non-pathogenic mutation. Intuitively, protein mutations with severe pathogenicity are less likely to occur frequently in the population, e.g., because organisms with severe pathogenic mutations are more likely to be removed from the population and less likely to reproduce. Thus, protein mutations that occur in the population at least at a threshold frequency are unlikely to have severe pathogenicity and can be identified as benign mutations for use in generating training examples for training a pathogenicity prediction neural network.
[0133] If a protein mutation has not occurred in the population of organisms, the system can designate the protein mutation as a pathogenic mutation. Thus, to identify pathogenic mutations, the system can randomly generate a set of protein mutations, remove any mutations that have already occurred in the population of organisms, and then designate the remaining mutations as pathogenic mutations. Intuitively, a protein mutation may not occur in the population because: (i) the population (effective population size) is not large enough for the mutation to be reflected in at least one organism in the population, or (ii) the protein mutation has severe pathogenicity such that organisms with the protein mutation are extremely unlikely to reproduce or survive.
[0134] After identifying the set of benign mutations and the set of pathogenic mutations as described above, the system can proceed to generate or obtain pathogenicity training examples. Specifically, to generate pathogenicity training examples, the system can identify training proteins and one or more protein mutations associated with the training proteins. The system can associate a first pathogenicity score (e.g., zero (0)) with each training protein mutation that has been designated as benign, and a second pathogenicity score (e.g., one (1)) with each training protein mutation that has been designated as pathogenic.
[0135] It should be understood that the above process of identifying benign and pathogenic protein mutations may result in the generation of "noisy" pathogenic training examples. For example, pathogenic protein mutations occur in a population, and thus the fact that a protein mutation occurs at least a threshold number of times in the population does not necessarily mean that the protein mutation is benign. Conversely, a protein mutation may not occur in the population because the population (effective population size) is not large enough, rather than because the protein mutation is pathogenic. To address this issue, the system can filter the set of pathogenic training examples to remove potentially inaccurate pathogenic training examples, as described below with reference to step (172).
[0136] The system trains a pathogenicity prediction neural network on the training example set (170). The training example set includes the pathogenic training examples generated at step (168), as well as masked training examples and structural training examples. Reference Figure 3 describes in more detail example techniques for training a pathogenicity prediction neural network on a training example set that includes pathogenic training examples, masked training examples, and structural training examples.
[0137] In some implementations, the system filters the set of pathogenic training examples (172). Specifically, for each protein mutation, the system can use the pathogenicity prediction neural network to generate a predicted pathogenicity score for the protein mutation. The system then generates an error score for the protein mutation based on the error between (i) the target pathogenicity score of the protein mutation and (ii) the predicted pathogenicity score. Here and elsewhere, the target pathogenicity score of a protein mutation can be, for example, a binary 0 / 1 score indicating whether the protein mutation is designated as benign or pathogenic, such as at step (168). Intuitively, a larger error score for a protein mutation indicates that the trained pathogenicity prediction neural network generates a (confident) prediction opposite to the initial designation of whether the protein mutation is benign or pathogenic. A larger error score for a protein mutation can indicate that the protein mutation is mislabeled as benign or pathogenic.
[0138] The system can generate a filtering score for each pathogenic training example based at least in part on the error scores of the mutations included in the amino acid sequence of the corresponding training protein. For example, the system can generate a filtering score for a pathogenic training example based at least in part on a combination (e.g., product) of the error scores of the mutations included in the amino acid sequence of the corresponding training protein. The system can use the filtering score in a variety of ways to filter the pathogenic training examples. As an example, the system can use the filtering score (e.g., via softmax) to generate a probability distribution over the set of pathogenic training examples, sample multiple pathogenic training examples according to the probability distribution (e.g., such that examples with higher filtering scores are more likely to be selected), and filter (remove) the sampled pathogenic training examples from the set of pathogenic training examples.
[0139] The filtering score of a pathogenicity training example can additionally be based on other aspects of the corresponding training protein, i.e., in addition to the error score of the protein mutation. For example, if the number of (pathogenic or benign) mutations included in the corresponding training protein is significantly deviated (e.g., deviated by at least a threshold amount) from the expected (pathogenic or benign) mutation number typically observed in proteins, then the filtering score of the pathogenicity training example can have a higher value. The pathogenicity prediction neural network can easily distinguish a training protein with a significantly higher or lower number of (pathogenic or benign) mutations than the expected (pathogenic or benign) mutation number from a native protein, which may affect the effectiveness of training.
[0140] Filtering the set of pathogenicity training examples based on the filtering score can improve the quality of the set of pathogenicity training examples, e.g., by removing training examples with mislabeled protein mutations and by removing training examples with unnatural artifacts.
[0141] The system can retrain the pathogenicity prediction neural network (174) on a set of training examples including the filtered set of pathogenicity training examples. That is, during the retraining of the pathogenicity prediction neural network, the filtered (removed) pathogenicity training examples are not used. Thus, for example, compared with the original pathogenicity prediction neural network trained in step (170), the retrained pathogenicity prediction neural network can be trained to achieve higher prediction accuracy in fewer training iterations.
[0142] Figure 6A An example of an MSA representation showing the MSA of a protein is shown. Each row of the MSA representation represents the corresponding MSA protein in the MSA. Each row includes a corresponding set of embeddings, and each position in the amino acid sequence of the MSA protein represented by that row includes an embedding. The second row of the MSA representation represents the protein, and the embedding 176 represents the mutation position in the amino acid sequence of the protein and is masked in the example MSA representation. The pathogenicity prediction neural network can process the MSA representation to generate a pathogenicity score that characterizes the pathogenic likelihood of the mutation at the mutation position in the amino acid sequence of the protein.
[0143] Figure 6B An example of a score distribution on an amino acid set generated by the pathogenicity prediction neural network is shown. In Figure 6BIn this case, amino acids are represented using their single-letter abbreviations. Amino acid L is the original amino acid at the mutation position in the protein amino acid sequence. The pathogenicity prediction system can evaluate the pathogenic likelihood of the mutation from amino acid L to amino acid N at the mutation position based on the difference between the score of amino acid L (0.55) and the score of amino acid N (0.44) under the score distribution. The pathogenicity prediction system can evaluate the likelihood of mutating from amino acid L to amino acid R at the mutation position based on the difference between the score of amino acid L (0.55) and the score of amino acid R (0.01) under the score distribution. The score distribution indicates that the mutation of amino acid R is more likely to be pathogenic than the mutation of amino acid N ( The difference is less than the difference).
[0144] Figure 7 An example protein structure prediction system 100 is shown. The protein structure prediction system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations, where the systems, components, and techniques described below are implemented.
[0145] System 100 is configured to process data defining one or more amino acid chains 102 of a protein 104 to generate a set of structural parameters 106 defining a predicted protein structure 108 (i.e., a prediction of the structure of protein 104). That is, the predicted structure 108 of protein 104 can be defined by a set of structural parameters 106 that jointly define the predicted three-dimensional structure of the protein after the protein undergoes protein folding.
[0146] The structural parameters 106 defining the predicted protein structure 108 can be as described above. For example, they can include, for example, position parameters and rotation parameters for each amino acid in protein 104, a distance map characterizing the estimated distance between each pair of amino acids in the protein, the corresponding spatial positions of each atom or backbone atom in each amino acid in the protein structure, or a combination thereof, as described above.
[0147] To generate the structural parameters 106 defining the predicted protein structure 108, system 100 generates: (i) a multiple sequence alignment (MSA) representation 110 of the protein, and (ii) a set of "pairwise" embeddings 112 of the protein, as will be described in more detail below.
[0148] The MSA representation 110 of the protein includes the corresponding representations of the MSA of each amino acid chain in the protein. The MSA representation of the amino acid chain in the protein can be represented as an embedding array (i.e., a 2-D embedding array having M rows and N columns), where N is the number of amino acids in the amino acid chain. Each row of the MSA representation can correspond to the corresponding MSA sequence of the amino acid chain in the protein. Refer toFigure 16 Describes an example process for generating a multiple sequence alignment (MSA) representation of an amino acid chain in a protein.
[0149] System 100 generates an MSA representation 110 of protein 104 based on the MSA representation of the amino acid chains in the protein.
[0150] If the protein includes only a single amino acid chain, system 100 can identify the MSA representation 110 of protein 104 as the MSA representation of the single amino acid chain in the protein.
[0151] If the protein includes multiple amino acid chains, system 100 can generate the MSA representation 110 of the protein by assembling the MSA representations of these amino acid chains in the protein into a block diagonal 2-D embedding array, i.e., the MSA representations of the amino acid chains in the protein form blocks on the diagonal. System 100 can initialize the embedding at each position outside the blocks on the diagonal in the 2-D array to a default embedding, such as a zero vector. The amino acid chains in the protein can be assigned any order, and the MSA representations of the amino acid chains in the protein can be sorted accordingly in the block diagonal matrix. For example, the MSA representation of the first amino acid chain (i.e., according to the sorting) can be the first block on the diagonal, the MSA representation of the second amino acid chain can be the second block on the diagonal, and so on.
[0152] Generally, the MSA representation 110 of a protein can be represented as a 2-D embedding array. Throughout this specification, a "row" of the MSA representation of a protein refers to a row of the 2-D embedding array that defines the MSA representation of the protein. Similarly, a "column" of the MSA representation of a protein refers to a column of the 2-D embedding array that defines the MSA representation of the protein.
[0153] The set of pairwise embeddings 112 includes corresponding pairwise embeddings corresponding to each pair of amino acids in protein 104. Generally, a pairwise embedding represents information encoding the relationship between pairs of amino acids in a protein. A pair of amino acids refers to an ordered tuple including a first amino acid and a second amino acid in the protein, i.e., such that the set of possible pairs of amino acids in the protein is given as follows:
[0154]
[0155] where N is the number of amino acids in the protein, indexes the amino acids in the protein, is the amino acid in the protein indexed by i, and is the amino acid in the protein indexed by If the protein includes multiple amino acid chains, the amino acids in the protein can be in the order of the amino acid chains in the protein from Sequential indexing. That is, first index the amino acids from the first amino acid chain in order, then the amino acids from the second amino acid chain, then the amino acids from the third amino acid chain, and so on. The set of pairwise embeddings 112 can be represented as a 2-D pairwise embedding array, for example, where the rows of the 2-D array are indexed by and the columns of the 2-D array are indexed by and the position in the 2-D array is occupied by the pairwise embedding of this amino acid pair .
[0156] Reference Figure 17 describes an example process for generating (initializing) the corresponding pairwise embeddings corresponding to each amino acid pair in a protein.
[0157] System 100 uses both the MSA representation 110 and the pairwise embeddings 112 to generate the structural parameters 106 that define the predicted protein structure 108, because both have complementary properties. The structure of the MSA representation 110 can explicitly depend on the number of amino acid chains in the MSA corresponding to each amino acid chain in the protein. Thus, the MSA representation 110 may not be suitable for directly predicting the protein structure because the protein structure 108 has no explicit dependence on the number of amino acid chains in the MSA. In contrast, the pairwise embeddings 112 characterize the relationships between the corresponding amino acid pairs in the protein 104 and do not explicitly refer to the MSA when expressed, and are thus a convenient and effective data representation for predicting the protein structure 108.
[0158] System 100 uses the embedding neural network 200 to process the MSA representation 110 and the pairwise embeddings 112 according to the values of the parameter set of the embedding neural network 200 to update the MSA representation 110 and the pairwise embeddings 112. That is, the embedding neural network 200 processes the MSA representation 110 and the pairwise embeddings 112 to generate the updated MSA representation 114 and the updated pairwise embeddings 116.
[0159] The embedding neural network 200 updates the MSA representation 110 and the pairwise embeddings 112 by sharing information between the MSA representation 110 and the pairwise embeddings 112. More specifically, the embedding neural network 200 alternates between updating the current MSA representation 110 based on the current pairwise embeddings 112 and updating the current pairwise embeddings 112 based on the current MSA representation 110.
[0160] Reference Figure 8 describes the example architecture of the embedding neural network 200 in more detail.
[0161] System 100 generates a network input for the folded neural network 600 from the updated pair embeddings 116, the updated MSA representation 114, or both, and processes the network input using the folded neural network 600 to generate the structural parameters 106 that define the predicted protein structure.
[0162] In some implementations, the folded neural network 600 processes the updated pair embeddings 116 to generate a distance map that includes, for each amino acid pair in the protein, a probability distribution over a set of possible distances between the amino acid pair in the protein structure. For example, to generate a probability distribution over a set of possible distances between amino acid pairs in a protein structure, the folded neural network can apply one or more fully-connected neural network layers to the updated pair embeddings 116 corresponding to the amino acid pair.
[0163] In some implementations, the folded neural network 600 generates the structural parameters 106 by processing a network input obtained from both the updated MSA representation 114 and the updated pair embeddings 116 using geometric attention operations that explicitly reason about the 3-D geometry of the amino acids in the protein structure. Refer to Figure 12 An example architecture of the folded neural network 600 that implements a geometric attention mechanism is described.
[0164] The training engine can train the protein structure prediction system 100 end-to-end to optimize an objective function herein referred to as the structure loss. The training engine can train the system 100 on a training dataset that includes a plurality of training examples. Each training example can specify: (i) a training input that includes an initial MSA representation and initial pair embeddings of a protein, and (ii) a target protein structure that the system 100 should generate by processing the training input. Experimental techniques (e.g., X-ray crystallography) can be used to determine the target protein structure for training the system 100.
[0165] The structure loss can characterize the similarity between: (i) the predicted protein structure generated by the system 100, and (ii) the target protein structure that the system should have generated.
[0166] For example, if the predicted structural parameters define the predicted position parameters and predicted rotation parameters for each amino acid in the protein, the structure loss can be given by:
[0167]
[0168] where N is the number of amino acids in the protein, t i represents the predicted position parameter of amino acid i, and R i represents the rotation matrix specified by the predicted rotation parameter of amino acid i is the target position parameter of amino acid i, representing the rotation matrix specified by the target rotation parameter of amino acid i, A is a constant, refers to the i specified by the predicted rotation parameter R inverse of the rotation matrix, refers to the specified by the target rotation parameter inverse of the rotation matrix, and represents a rectified linear unit (ReLU) operation.
[0169] The structure loss defined by reference equations (2)-(4) can be understood as the loss for each amino acid pair in the protein averaged. The term defines the predicted spatial position of the amino acid in the predicted reference frame of amino acid i, and defines the actual spatial position of the amino acid in the actual reference frame of amino acid i. These terms are sensitive to the predicted rotation and actual rotation of amino acid i and thus carry more information compared to the loss terms that are only sensitive to the predicted distance and actual distance between amino acids.
[0170] As another example, if the predicted structural parameters define the predicted spatial position of each atom in each amino acid of the protein, then the structure loss can be the average error (e.g., squared error) between the following two terms: (i) the predicted spatial position of the atom, and (ii) the target (e.g., ground truth) spatial position of the atom.
[0171] Optimizing the structure loss prompts system 100 to generate a predicted protein structure that accurately approximates the true protein structure.
[0172] In addition to optimizing the structure loss, the training engine can also train system 100 to optimize one or more auxiliary losses. The auxiliary losses can penalize predicted structures with properties that are unlikely to occur in nature, such as based on the bond angles and / or bond lengths of the bonds between atoms in the amino acids of the predicted structure, or based on the proximity of atoms in different amino acids in the predicted structure.
[0173] The training engine can train the structure prediction system 100 on the training data through multiple training iterations, e.g., using a stochastic gradient descent training technique.
[0174] Figure 8Shows an example architecture of an embedding neural network 200 configured to process an MSA representation 110 and a pairwise embedding 112 to generate an updated MSA representation 114 and an updated pairwise embedding 116.
[0175] The embedding neural network 200 includes a sequence of update blocks 202-A to 202-N. Throughout this specification, a "block" refers to a part of a neural network, e.g., a neural network subnetwork including one or more neural network layers.
[0176] Each update block in the embedding neural network is configured to receive a block input including an MSA representation and a pairwise embedding and process the block input to generate a block output including an updated MSA representation and an updated pairwise embedding.
[0177] The embedding neural network 200 provides the MSA representation 110 and the pairwise embedding 112 included in the network input of the embedding neural network 200 to the first update block (i.e., in the sequence of update blocks). The first update block processes the MSA representation 110 and the pairwise embedding 112 to generate an updated MSA representation and an updated pairwise embedding.
[0178] For each update block after the first update block, the embedding neural network 200 provides the MSA representation and the pairwise embedding generated by the previous update block to the update block and provides the updated MSA representation and the updated pairwise embedding generated by this update block to the next update block.
[0179] The embedding neural network 200 gradually enriches the information content of the MSA representation 110 and the pairwise embedding 112 by repeatedly updating them using the sequence of update blocks 202-A to 202-N.
[0180] The embedding neural network 200 may provide the updated MSA representation 114 and the updated pairwise embedding 116 generated by the final update block (i.e., in the sequence of update blocks) as the network output.
[0181] Figure 9 Shows an example architecture of an update block 300 of the embedding neural network 200 (i.e., as described in reference Figure 8 ).
[0182] The update block 300 receives a block input including a current MSA representation 302 and a current pairwise embedding 304 and processes the block input to generate an updated MSA representation 306 and an updated pairwise embedding 308.
[0183] The update block 300 includes an MSA update block 400 and a pairwise update block 500.
[0184] The MSA update block 400 updates the current MSA representation 302 using the current pairwise embedding 304, and the pairwise update block 500 updates the current pairwise embedding 304 using the updated MSA representation 306 (i.e., generated by the MSA update block 400).
[0185] Generally, the MSA representation and the pairwise embedding can encode complementary information. For example, the MSA representation can encode information about the correlations between the amino acid identities at different positions in an evolutionarily related set of amino acid chains, while the pairwise embedding can encode information about the relationships between the amino acids in a protein. The MSA update block 400 uses the complementary information encoded in the pairwise embedding to enrich the information content of the MSA representation, and the pairwise update block 500 uses the complementary information encoded in the MSA representation to enrich the information content of the pairwise embedding. Due to this enrichment, the updated MSA representation and the updated pairwise embedding encode information that is more relevant to predicting the protein structure.
[0186] This document describes the update block 300 as first using the current pairwise embedding 304 to update the current MSA representation 302, and then using the updated MSA representation 306 to update the current pairwise embedding 304. This description should not be construed as limiting the update block to performing operations in this order. For example, the update block can first use the current MSA representation to update the current pairwise embedding, and then use the updated pairwise embedding to update the current MSA representation.
[0187] The update block 300 is described in this document as including an MSA update block 400 (i.e., updating the current MSA representation) and a pairwise update block 500 (i.e., updating the current pairwise embedding). This description should not be construed as limiting the update block 300 to including only one MSA update block or only one pairwise update block. For example, the update block 300 can include multiple MSA update blocks that update the MSA representation multiple times before providing the MSA representation to the pairwise update block for updating the current pairwise embedding. As another example, the update block 300 can include multiple pairwise update blocks that use the MSA representation to update the pairwise embedding multiple times.
[0188] The MSA update block 400 and the pairwise update block 500 can have any suitable architecture to enable them to perform the functions described for them.
[0189] In some implementations, the MSA update block 400, the pairwise update block 500, or both include one or more "self-attention" blocks. As used throughout this document, a self-attention block generally refers to a neural network block that updates a set of embeddings (i.e., receives a set of embeddings and outputs updated embeddings). To update a given embedding, the self-attention block can determine the respective "attention weights" between the given embedding and each of one or more selected embeddings, and then update the given embedding using: (i) the attention weights, and (ii) the selected embeddings. For convenience, the self-attention block can be said to update the given embedding using the attention "applied" to the selected embeddings.
[0190] For example, the self-attention block can receive a set of input embeddings , where N is the number of amino acids in the protein, and to update the embedding , the self-attention block can determine the attention weights , where represents the attention weight between and , as follows:
[0191]
[0192] where and are learned parameter matrices, represents the soft-max normalization operation, and c is a constant. Using the attention weights, the self-attention layer can update the embedding , as follows:
[0193]
[0194] where is a learned parameter matrix. ( can be referred to as the "query embedding" of the input embedding , can be referred to as the "key embedding" of the input embedding , and can be referred to as the "value embedding" of the input embedding ).
[0195] The parameter matrices ("query embedding matrix"), ("key embedding matrix"), and ("value embedding matrix") are trainable parameters of the self-attention block. The parameters of any self-attention block included in the MSA update block 400 and the pairwise update block 500 can be understood as parameters of the update block 300, which can be used as a reference Figure 1Part of the end-to-end training of the described protein structure prediction system 100 is trained. Generally, the (trained) parameters of the query embedding matrix, the key embedding matrix, and the value embedding matrix are different for different self-attention blocks. For example, the self-attention blocks included in the MSA update block 400 can have different query embedding matrices, key embedding matrices, and value embedding matrices with parameters different from those of the self-attention blocks included in the pairwise update block 500.
[0196] In some implementations, the MSA update block 400, the pairwise update block 500, or both include one or more self-attention blocks conditioned on pairwise embeddings, i.e., perform self-attention operations conditioned on pairwise embeddings. To condition the self-attention operation on pairwise embeddings, the self-attention block can process the pairwise embeddings to generate corresponding "attention biases" corresponding to each attention weight. For example, in addition to determining the attention weights according to equations (5)-(6) the self-attention block can also generate a corresponding set of attention biases where denotes the attention bias between and The self-attention block can generate the attention bias by applying a learned parameter matrix to the pairwise embeddings (i.e., the pairs of amino acids indexed by in the protein).
[0197] The self-attention block can determine a set of "biased attention weights" by, for example, summing (or otherwise combining) the attention weights and the attention biases, where denotes the biased attention weight between and For example, the self-attention block can determine the biased attention weight between the embedding and
[0198]
[0199] where is the attention weight between and is the attention bias between and The self-attention block can use the biased attention weights to update each input embedding
[0200]
[0201] wherein is the learned parameter matrix.
[0202] Generally, pairwise embeddings encode information representing the protein structure and the relationships between pairs of amino acids in the protein structure. Applying a self-attention operation conditioned on the pairwise embeddings to the input embedding set allows the input embeddings to be updated in a way that is influenced by the protein structure information encoded in the pairwise embeddings. The update block of the embedding neural network can use a self-attention block conditioned on the pairwise embeddings to update and enrich the MSA representation and the pairwise embeddings themselves.
[0203] Optionally, the self-attention block can have multiple "heads", each head generating a corresponding updated embedding corresponding to each input embedding, i.e., such that each input embedding is associated with multiple updated embeddings. For example, each head can generate an updated embedding according to different values of the parameter matrices , and . A self-attention block with multiple heads can implement a "gating" operation to combine the updated embeddings generated by these heads for the input embeddings, i.e., to generate a single updated embedding corresponding to each input embedding. For example, the self-attention block can process the input embeddings using one or more neural network layers (e.g., fully connected neural network layers) to generate corresponding gating values for each head. The self-attention block can then combine the updated embeddings corresponding to the input embeddings according to the gating values. For example, the self-attention block can generate an updated embedding for the input embedding as follows:
[0204]
[0205] wherein indexes the head, is the gating value of the head , and is the updated embedding generated by the head for the input embedding . .
[0206] Reference Figure 10 describes an example architecture of the MSA update block 400 using a self-attention block conditioned on pairwise embeddings. The example MSA update block described in reference Figure 10 updates the current MSA representation by processing the rows of the current MSA representation using a self-attention block conditioned on the current pairwise embeddings, based on the current pairwise embeddings.
[0207] Reference Figure 11Describes an example architecture of a pairwise update block 500 that uses a self-attention block conditioned on pairwise embeddings. Refer to Figure 11 The described example pairwise update block updates the current pairwise embedding based on the updated MSA representation by calculating the mean of the outer products of the updated MSA representation, adding the result of the outer product mean to the current pairwise embedding, and processing the current pairwise embedding using a self-attention block conditioned on the current pairwise embedding.
[0208] Figure 10 Illustrates an example architecture of an MSA update block 400. The MSA update block 400 is configured to receive a current MSA representation 302 and update the current MSA representation 306 based at least in part on the current pairwise embedding.
[0209] To update the current MSA representation 302, the MSA update block 400 uses a self-attention operation conditioned on the current pairwise embedding (i.e., a "row-wise" self-attention operation that operates only on the embeddings in a specific row) to update the embeddings in each row of the current MSA representation. More specifically, the MSA update block 400 provides the embeddings in each row of the current MSA representation 302 to a "row-wise" self-attention block 402 conditioned on the current pairwise embedding (e.g., as described in reference Figure 9 ), to generate updated embeddings for each row of the current MSA representation 302. Optionally, the MSA update block can add the input of the row-wise self-attention block 402 to the output of the row-wise self-attention block 402. Making the row-wise self-attention block 402 conditioned on the current pairwise embedding enables the MSA update block 400 to use information from the current pairwise embedding to enrich the current MSA representation 302.
[0210] Then, the MSA update block uses a self-attention operation not conditioned on the current pairwise embedding (i.e., a "column-wise" self-attention operation that operates only on the embeddings in a specific column) to update the embeddings in each column of the current MSA representation. More specifically, the MSA update block 400 provides the embeddings in each column of the current MSA representation 302 to a "column-wise" self-attention block 404 not conditioned on the current pairwise embedding to generate updated embeddings for each column of the current MSA representation 302. Since it is not conditioned on the current pairwise embedding, the column-wise self-attention block 404 uses attention weights (e.g., as described in reference to equations (5)-(6)) instead of biased attention weights (e.g., as described in reference to equation (8)) to generate updated embeddings for each column of the current MSA representation. Optionally, the MSA update block can add the input of the column-wise self-attention block 404 to the output of the column-wise self-attention block 404.
[0211] Then, the MSA update block processes the current MSA representation 302 using a transition block that, for example, applies one or more fully-connected neural network layers to the current MSA representation 302. Optionally, the MSA update block 400 can add the input of the transition block 406 to the output of the transition block 406.
[0212] The MSA update block can output an updated MSA representation 306 resulting from the operations performed by the row-wise self-attention block 402, the column-wise self-attention block 404, and the transition block 406.
[0213] Figure 11 An example architecture of the pairwise update block 500 is shown. The pairwise update block 500 is configured to receive the current pairwise embedding 304 and update the current pairwise embedding 304 (at least in part) based on the updated MSA representation 306.
[0214] To update the current pairwise embedding 304, the pairwise update block 500 applies an outer product mean operation 502 to the updated MSA representation 306 and adds the result of the outer product mean operation 502 to the current pairwise embedding 304.
[0215] The outer product mean operation defines a sequence of operations that, when applied to an MSA representation expressed as an embedding array, generates an embedding array, i.e., where N is the number of amino acids in the protein. The current pairwise embedding 304 can also be expressed as an embedding array, and adding the result of the outer product mean 502 to the current pairwise embedding 304 means adding these two embedding arrays.
[0216] To compute the outer product mean, the pairwise update block generates a tensor , for example, given by:
[0217]
[0218] where , , , , where C is the number of channels in each embedding of the MSA representation, is the number of rows in the MSA representation, is a linear operation (e.g., defined by matrix multiplication) applied to the channels of the embedding at the row indexed by " " and the column indexed by " " in the MSA representation, and is a linear operation (e.g., defined by matrix multiplication) applied to the channels of the embedding at the row indexed by " " and the column indexed by " ”The rows indexed by “ ” and the embedded channels at the columns indexed by “ ” are linearly operated (e.g., defined by matrix multiplication). The result of the outer product mean is generated by flattening and linearly projecting the dimensions of tensor A. Optionally, as part of calculating the outer product mean, the pairwise update block may perform one or more layer normalization operations (e.g., as described in the following reference: “Layer Normalization” by Jimmy Lei Ba et al., arXiv:1607.06450).
[0219] Generally, the updated MSA representation 306 encodes information about the correlations between the amino acid identities at different positions in an evolutionarily related amino acid chain set. The information encoded in the updated MSA representation 306 is related to predicting the structure of a protein, and by incorporating the information encoded in the updated MSA representation into the current pairwise embedding (i.e., via the outer product mean 502), the pairwise update block 500 can enhance the information content of the current pairwise embedding.
[0220] After using the updated MSA representation (i.e., via the outer product mean 502) to update the current pairwise embedding 304, the pairwise update block 500 updates each row of the permutation of the current pairwise embedding to an array using a self-attention operation conditioned on the current pairwise embedding (i.e., a “row-wise” self-attention operation). More specifically, the pairwise update block 500 provides each row of the current pairwise embedding to a “row-wise” self-attention block 504 (e.g., as described in the reference Figure 9 ), which is also conditioned on the current pairwise embedding, to generate an updated pairwise embedding for each row. Optionally, the pairwise update block may add the input of the row-wise self-attention block 504 to the output of the row-wise self-attention block 504.
[0221] Then, the pairwise update block 500 uses a self-attention operation (i.e., a “column-wise” self-attention operation) conditioned on the current pairwise embedding to update the current pairwise embedding in each column of the array. More specifically, the pairwise update block 500 provides each column of the current pairwise embedding to a “column-wise” self-attention block 506, which is also conditioned on the current pairwise embedding, to generate an updated pairwise embedding for each row. Optionally, the pairwise update block may add the input of the column-wise self-attention block 506 to the output of the column-wise self-attention block 506.
[0222] Then, the paired update block 500 processes the current paired embedding using a transformation block that applies, for example, one or more fully connected neural network layers to the current paired embedding. Optionally, the paired update block 500 can add the input of the transformation block 508 to the output of the transformation block 508.
[0223] The paired update block can output the updated paired embedding 308 resulting from the operations performed by the row self-attention block 504, the column self-attention block 506, and the transformation block 508.
[0224] Figure 12 An example architecture of a folded neural network 600 is shown that generates a set of structural parameters 106 that define a predicted protein structure 108. The folded neural network 600 can be included in the protein structure prediction system 100 described in the reference Figure 1 described.
[0225] The folded neural network 600 generates structural parameters for each amino acid in the protein, which can include: (i) position parameters and (ii) rotation parameters. As previously described, the position parameter of an amino acid can specify the predicted 3-D spatial position of a specified atom in the amino acid in the protein structure. The rotation parameter of an amino acid can specify the predicted "orientation" of the amino acid in the protein structure. More specifically, the rotation parameter can specify a 3-D spatial rotation operation that, if applied to the coordinate system of the position parameter, causes three "backbone" atoms in the amino acid to be in fixed positions relative to the rotated coordinate system.
[0226] In an implementation, the folded neural network 600 receives inputs from the final MSA representation, the final paired embedding, or both, and generates final values of the structural parameters 106 that define the predicted structure of the protein. For example, the folded neural network 600 can receive the following inputs: (i) the corresponding paired embedding 116 for each amino acid pair in the protein, (ii) the initial value of the "single" embedding 602 for each amino acid in the protein, and (iii) the initial value of the structural parameter 604 for each amino acid in the protein. The folded neural network 600 processes the inputs to generate final values of the structural parameters 106 that jointly characterize the predicted structure 108 of the protein.
[0227] The protein structure prediction system 100 can provide the paired embedding generated as the output of the embedding neural network to the folded neural network 600 as described in the reference Figure 1 described.
[0228] The protein structure prediction system 100 can generate the initial single embedding 602 of the amino acid according to the MSA representation 114, i.e., as described in the reference Figure 1The embedding is generated as the output of the embedded neural network. For example, as described above, the MSA representation 114 can be represented as a 2-D embedding array having the number of amino acids in the protein, where each column is associated with the corresponding amino acid in the protein. The protein structure prediction system 100 can generate an initial single embedding for each amino acid in the protein by summing (or otherwise combining) the embeddings from the column associated with that amino acid in the MSA representation 114. As another example, the protein structure prediction system 100 can generate an initial single embedding for an amino acid in the protein by extracting the embedding from the row in the MSA representation 114 corresponding to the amino acid sequence of the protein whose structure is being estimated.
[0229] The protein structure prediction system 100 can generate initial structure parameters 604 with default values. For example, where the position parameter for each amino acid is initialized to the origin (e.g., in a Cartesian coordinate system ), and the rotation parameter for each amino acid is initialized to the identity matrix.
[0230] The folding neural network 600 can generate the final structure parameters 106 by repeatedly updating the current values of the single embedding 606 and the structure parameters 608 (i.e., starting from their initial values). More specifically, the folding neural network 600 includes a sequence of update neural network blocks 610, where each update block 610 is configured to update the current single embedding 606 (i.e., generate an updated single embedding 612) and update the current structure parameters 608 (i.e., generate updated structure parameters 614). In addition to the update blocks, the folding neural network 600 can also include other neural network layers or blocks, such as can be interleaved with the update blocks.
[0231] Each update block 610 can include: (i) a geometric attention block 616, and (ii) a folding block 618, each of which will be described in more detail below.
[0232] The geometric attention block 616 updates the current single embedding using a "geometric" self-attention operation that explicitly reasons about the 3-D geometry of the amino acids in the protein structure (i.e., the geometry defined by the structure parameters). More specifically, to update a given single embedding, the geometric attention block 616 determines the corresponding attention weights between the given single embedding and each of one or more selected single embeddings, where the attention weights depend on the current single embedding, the current structure parameters, and the pairwise embeddings. Then, the geometric attention block 616 updates the given single embedding using: (i) the attention weights, (ii) the selected single embeddings, and (iii) the current structure parameters.
[0233] To determine the attention weights, the geometric attention block 616 processes each current single embedding to generate corresponding "symbol query" embedding, "symbol key" embedding, and "symbol value" embedding. For example, for the single embedding corresponding to the i-th amino acid , the geometric attention block 616 can generate the symbol query embedding , the symbol key embedding , and the symbol value embedding , as follows:
[0234]
[0235] where refers to a linear layer with independently learned parameter values.
[0236] The geometric attention block 616 additionally processes each current single embedding to generate corresponding "geometric query" embedding, "geometric key" embedding, and "geometric value" embedding. The geometric query embedding, geometric key embedding, and geometric value embedding of each single embedding are initially generated in the local reference frame of the corresponding amino acid and then rotated and translated to 3-D points in the global reference frame using the structural parameters of that amino acid. For example, for the single embedding h i corresponding to the i-th amino acid, the geometric attention block 616 can generate the geometric query embedding , the geometric key embedding , and the geometric value embedding , as follows:
[0237]
[0238] where refers to linear layers with independently learned parameter values that project h i onto 3-D points (the superscript p indicates that the quantity is a 3-D point), R i represents the rotation matrix specified by the rotation parameters of the i-th amino acid, and t i represents the position parameter of the i-th amino acid.
[0239] To update the single embedding h i corresponding to amino acid i, the geometric attention block 616 can generate the attention weights , where N is the total number of amino acids in the protein, and is the attention weight between amino acid i and amino acid j, as follows:
[0240]
[0241] where represents the symbol query embedding of amino acid i, represents amino acid Symbol key embedding, denotes and the dimension of, denotes the learned parameter, denotes the geometric query embedding of amino acid i, denotes amino acid the geometric key embedding of, is the norm, and is the pairwise embedding 116 corresponding to the amino acid pair including amino acid i and amino acid , and w is the learned weight vector (or some other learned projection operation).
[0242] Generally, the pairwise embedding of an amino acid pair implicitly encodes information about the relationship between the amino acids in the pair (e.g., the distance between the amino acids in the pair). By determining the attention weights between amino acid i and partially based on the pairwise embedding of, the folding neural network 600 enriches the attention weights with information from the pairwise embedding and thus improves the accuracy of the predicted folded structure. between amino acid i and amino acid
[0243] In some implementations, the geometric attention block 616 generates multiple sets of geometric query embeddings, geometric key embeddings, and geometric value embeddings, and uses each generated set of geometric embeddings to determine the attention weights.
[0244] After generating the attention weights for the single embedding h i corresponding to amino acid i, the geometric attention block 616 uses the attention weights to update the single embedding h i . Specifically, the geometric attention block 616 uses the attention weights to generate a "symbol return" embedding and a "geometric return" embedding, and then uses the symbol return embedding and the geometric return embedding to update the single embedding. The geometric attention block 124 can generate the symbol return embedding o i of amino acid i, for example, as follows:
[0245]
[0246] where denotes the attention weights (e.g., as defined in reference equation (16)), and each denotes the symbol value embedding of amino acid j. The geometric attention block 616 can generate the geometric return embedding of amino acid i, for example, as follows:
[0247]
[0248] where the geometric return embedding is a 3-D point, represents the attention weights (e.g., as defined in reference equation (16)), is the inverse of the rotation matrix specified by the rotation parameters of amino acid i, and t i is the position parameter of amino acid i. It should be understood that the geometric return embedding is initially generated in the global reference frame and then rotated and translated into the local reference frame of the corresponding amino acid.
[0249] The geometric attention block 616 can use the corresponding symbolic return embedding o i (e.g., generated according to equation (17)) and the geometric return embedding (e.g., generated according to equation (18)) to update the single embedding h i of amino acid i, for example, as follows:
[0250]
[0251] where is the updated single embedding of amino acid i, is the norm, such as the L2 norm, and represents the layer normalization operation, for example, as described in reference to the following: "Layer Normalization" by J.L. Ba, J.R. Kiros, G.E. Hinton, arXiv:1607.06450 (2016).
[0252] Using the specific 3-D geometric embedding to update the single embedding of amino acids 606 (e.g., as described in reference to equations (13)-(15)) enables the geometric attention block 616 to reason about 3-D geometry when updating the single embedding. In addition, each update block updates the single embedding and the structure parameters in a way that is invariant to the rotation and translation of the overall protein structure. For example, applying the same global rotation and translation operations to the initial structure parameters provided to the folding neural network 600 will cause the folding neural network 600 to generate a predicted structure that is globally rotated and translated in the same way but otherwise remains unchanged. Therefore, applying global rotation and translation operations to the initial structure parameters does not affect the accuracy of the predicted protein structure generated by the folding neural network 600 starting from the initial structure parameters. The rotation and translation invariance of the representations generated by the folding neural network 600 facilitates training, for example, because the folding neural network 600 automatically learns to generalize over all rotations and translations of the protein structure.
[0253] The updated single embedding of the amino acid can be further transformed by one or more additional neural network layers (e.g., linear neural network layers) in the geometric attention block 616 and then provided to the folding block 618.
[0254] After the geometric attention block 616 updates the current single embedding 606 of the amino acid, the folding block 618 uses the updated single embedding 612 to update the current structural parameters 608. For example, the folding block 618 can update the current position parameter t of amino acid i i , as follows:
[0255]
[0256] where is the updated position parameter, represents a linear neural network layer, and represents the updated single embedding of amino acid i. In another example, the rotation parameter R of amino acid i i can specify a rotation matrix, and the folding block 618 can update the current rotation parameter R i , as follows:
[0257]
[0258] where w i is a three-dimensional vector, is a linear neural network layer, is the updated single embedding of amino acid i, represents a quaternion with a real part of 1 and an imaginary part of w i , and represents an operation that transforms the quaternion into an equivalent rotation matrix. Using equations (21)-(22) to update the rotation parameter ensures that the updated rotation parameter defines a valid rotation matrix, such as an orthogonal matrix with a determinant of one.
[0259] The folding neural network 600 can provide the updated structural parameters generated by the final update block 610 as the final structural parameters 106 that define the predicted protein structure 108. The folding neural network 600 can include any suitable number of update blocks, such as 5 update blocks, 25 update blocks, or 125 update blocks. Optionally, each update block in the folding neural network can share a set of parameter values that are jointly updated during the training of the folding neural network. Sharing parameter values between the update blocks 610 reduces the number of trainable parameters of the folding neural network and can therefore facilitate the effective training of the folding neural network, for example, by stabilizing the training and reducing the likelihood of overfitting.
[0260] During training, the training engine can train parameters of the structure prediction system, including parameters of the folded neural network 600, as described above, based on a structural loss that evaluates the accuracy of the final structure parameters 106. In some implementations, the training engine can further evaluate the auxiliary structural loss of one or more update blocks in the update block 610 before the final update block (i.e., generating the final structure parameters). The auxiliary structural loss of the update block evaluates the accuracy of the updated structure parameters generated by the update block.
[0261] Optionally, during training, the training engine can apply "stop gradient" operations to prevent gradients from backpropagating through certain neural network parameters of each update block (e.g., neural network parameters used to calculate updated rotation parameters (as described in Equations (21)-(22))). Applying these stop gradient operations can improve the numerical stability of the gradients calculated during training.
[0262] In general, the similarity between the predicted protein structure 108 generated by the folded neural network 600 and the corresponding reference true protein structure can be measured, for example, by a similarity metric that assigns a corresponding accuracy score to each of a plurality of atoms in the predicted protein structure. For example, the similarity metric can assign a corresponding accuracy score to each alpha carbon atom in the predicted protein structure. The accuracy score of an atom in the predicted protein structure can characterize the degree of agreement between the position of the atom in the predicted protein structure and the actual position of the atom in the reference true protein structure. An example of a similarity metric that can compare a predicted protein structure with a reference true protein structure to generate an accuracy score for the atoms in the predicted protein structure is the lDDT similarity metric described in reference to the following document: V. Mariani et al. "lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests," Bioinformatics, Nov. 1, 2013; 29(21) 2722-2728.
[0263] The folded neural network 600 can be configured to generate corresponding confidence estimates 650 for each of one or more atoms in the predicted protein structure 108. The confidence estimate 650 of an atom in the predicted protein structure characterizes the predicted accuracy score of the atom in the predicted protein structure (e.g., the lDDT accuracy score), i.e., the predicted accuracy score will be generated by a similarity metric that compares the predicted protein structure with the (possibly unknown) ground truth protein structure. In one example, the confidence estimate 650 of an atom in the predicted protein structure can define a discrete probability distribution over a set of intervals that partition the range of possible values of the atom accuracy score. The discrete probability distribution can associate a corresponding probability with each of the intervals in the set, which defines the likelihood that the actual accuracy score is included in that interval. For example, the range of possible values of the accuracy score can be , and the confidence estimate 650 can define a probability distribution over the set of intervals . In another example, the confidence estimate 650 of an atom in the predicted protein structure can be a numerical value, i.e., directly predict the accuracy score of the atom.
[0264] In some implementations, the folded neural network 600 generates corresponding confidence estimates 650 for a specified atom (e.g., the alpha carbon atom) in each amino acid of the protein. The folded neural network 600 can generate the confidence estimate 650 of the specified atom in the amino acid of the protein by processing the updated single embedding of the amino acid generated by the last update block in the folded neural network using, for example, one or more neural network layers (e.g., fully connected layers).
[0265] The structure prediction system can generate corresponding confidence scores for each amino acid in the protein based on the confidence estimates 650 of the atoms in the predicted protein structure. For example, the structure prediction system can generate the confidence score of an amino acid as the expected value of the probability distribution over the possible values of the accuracy score of the alpha carbon atom in the amino acid.
[0266] The structure prediction system can generate a confidence score for the entire predicted structure, e.g., as the average of the confidence scores of the amino acids in the protein.
[0267] During the training of the structure prediction system, the training engine can adjust the parameter values of the structure prediction system by backpropagating the gradient of an auxiliary loss that measures the error between: (i) the confidence estimates generated by the folded neural network 600, and (ii) the accuracy scores generated by comparing the predicted protein structure with the ground truth protein structure. The error can be, for example, the cross-entropy error.
[0268] Confidence estimates generated by a structure prediction system can be used in a variety of ways. For example, the confidence estimate of an atom in a predicted protein structure can indicate which parts of the structure have been reliably estimated and are thus suitable for further downstream processing or analysis. As another example, the confidence score for each protein can be used to rank a set of predictions of protein structures (e.g., predictions generated by the same structure prediction system by processing different inputs characterizing the same protein, or predictions generated by different structure prediction systems).
[0269] The position parameter and the rotation parameter specified by the structure parameter 106 can define the spatial position of the backbone atoms in a protein amino acid (e.g., in Cartesian coordinate representation). However, the structure parameter 106 does not necessarily define the spatial position of the remaining atoms in a protein amino acid (e.g., the atoms in the side chain of the amino acid). Specifically, the spatial position of the remaining atoms in an amino acid depends on the values of the torsion angles between the bonds in the amino acid, such as, for example, the ω-angle (omega-angle), the φ-angle (phi-angle), the ψ-angle (psi-angle), the χ1-angle (chi1-angle), the χ2-angle (chi2-angle), the χ3-angle (chi3-angle), and the χ4-angle (chi4 angle), as shown in the reference Figure 13 indicated.
[0270] Optionally, one or more of the update blocks 610 in the update block 600 of the folding neural network 600 can generate an output defining the corresponding predicted spatial position of each atom in each amino acid of the protein. To generate the predicted spatial position of an atom in an amino acid, the update block can process the updated single embedding of the amino acid using one or more neural network layers to generate a predicted value of the torsion angle of the bond between the atoms in the amino acid. The neural network layer can be, for example, a fully connected neural network layer embedded with residual connections. Each torsion angle can be represented as, for example, a 2-D vector.
[0271] The update block can determine the spatial positions of atoms in an amino acid based on: (i) the values of the torsion angles of the amino acid, and (ii) the updated structural parameters of the amino acid (e.g., position parameters and rotation parameters). For example, the update block can process the torsion angles according to a predefined function to generate the spatial positions of the atoms in the amino acid in the local reference frame of the amino acid. The update block can generate the spatial positions of the atoms in the amino acid in the global reference frame (i.e., common to all amino acids in the protein) by rotating and translating the spatial positions of the atoms according to the updated structural parameters of the amino acid. For example, the update block can determine the spatial positions of the atoms in the global reference frame by applying a rotation operation defined by the updated rotation parameters to the spatial positions of the atoms in the local reference frame to generate a rotated spatial position, and then applying a translation operation defined by the updated position parameters to the rotated spatial position.
[0272] In some implementations, instead of or in combination with outputting the final structural parameters, the folding neural network 600 outputs the predicted spatial positions of the atoms in the protein amino acids generated by the final update block.
[0273] Reference Figure 12 The folding neural network 600 described herein is characterized as receiving as input an MSA representation 114 and a pairwise embedding 116 based on (e.g., as described in reference Figure 8 ). However, in general, the inputs to the folding neural network (e.g., the single embedding 602 and the pairwise embedding 116) can be generated using any suitable technique. Additionally, aspects of the operations performed by the folding neural network (e.g., predicting the spatial positions of the atoms in each amino acid of the protein) can be performed by other folding neural networks (e.g., having different architectures that receive different inputs).
[0274] Figure 13 Illustrates the torsion angles between the bonds in an amino acid, such as the ω angle, φ angle, ψ angle, χ1 angle, χ2 angle, χ3 angle, χ4 angle, and χ5 angle.
[0275] Figure 14 Is an illustration of an unfolded protein and a folded protein. The unfolded protein is a random coil structure of amino acids. The unfolded protein undergoes protein folding and folds into a 3D configuration. Protein structures typically include stable local folding patterns, such as α helices (e.g., shown at 802) and β sheets.
[0276] Figure 15is a flowchart of an example process 900 for predicting the structure of a protein comprising one or more chains, where each chain specifies an amino acid sequence. For convenience, process 900 will be described as being performed by a system of one or more computers located in one or more locations. For example, a protein structure prediction system (e.g., Figure 1 protein structure prediction system 100) can perform process 900 after being appropriately programmed according to this specification.
[0277] The system obtains an initial MSA representation (902) representing a corresponding multiple sequence alignment (MSA) for each chain in the protein.
[0278] The system obtains a corresponding initial pairwise embedding for each amino acid pair in the protein (904).
[0279] The system processes the input comprising the initial MSA representation and the initial pairwise embeddings using an embedding neural network to generate an output comprising a final MSA representation and a corresponding final pairwise embedding for each amino acid pair in the protein.
[0280] The embedding neural network comprises a sequence of update blocks. Each update block has a corresponding set of update block parameters and is configured to receive a current MSA representation and a corresponding current pairwise embedding for each amino acid pair in the protein. Each update block: (i) updates the current MSA representation based on the current pairwise embeddings, and (ii) updates the current pairwise embeddings based on the updated MSA representation (906).
[0281] The system determines the predicted structure of the protein based on using the final MSA representation, the final pairwise embeddings, or both (908).
[0282] Figure 16 illustrates an example process 1000 for generating an MSA representation 1010 of an amino acid chain in a protein. Refer to Figure 1 The protein structure prediction system 100 described can implement the operations of process 1000.
[0283] To generate the MSA representation 100 of an amino acid chain in a protein, system 100 obtains the MSA 1002 of the protein, which can include, for example, thousands of MSA sequences.
[0284] System 100 divides the MSA sequence set into a "core" MSA sequence set 1004 and an "extra" MSA sequence set 1006. The core MSA sequence set can be smaller than the extra MSA sequence set 1006 (e.g., by an order of magnitude). System 100 can divide the MSA sequence set into the core MSA sequences 1004 and the extra MSA sequences 1006, for example, by randomly selecting a predetermined number of MSA sequences as the core MSA sequences and identifying the remaining MSA sequences as the extra MSA sequences 1006.
[0285] For each extra MSA sequence 1006, System 100 can determine the corresponding similarity metric (e.g., based on the Hamming distance) between the extra MSA sequence and each core MSA sequence 1004. Then, System 100 can associate each extra MSA sequence 1006 with the corresponding core MSA sequence 1004 that is most similar to that extra MSA sequence 1006 (i.e., according to the similarity metric). The set of extra MSA sequences 1006 associated with the core MSA sequences 1004 can be referred to as the "MSA sequence cluster" 1008. That is, System 100 determines the corresponding MSA sequence cluster 1008 for each core MSA sequence 1004, where the MSA sequence cluster 1008 corresponding to the core MSA sequence 1004 includes the set of extra MSA sequences 1006 that are most similar to the core MSA sequence 1004.
[0286] System 100 can generate an MSA representation of the amino acid chain in the protein based on the core MSA sequences and the MSA sequence clusters 1008. The MSA representation 1010 can be represented by an embedding array, where M is the number of core MSA sequences (i.e., such that each core MSA sequence is associated with a corresponding row of the MSA representation), and N is the number of amino acids in the amino acid chain. The embedding in the MSA representation can be indexed by to.
[0287] To generate the embedding at position in the MSA representation 1010, System 100 can obtain the embedding (e.g., one-hot embedding) that defines the identity of the amino acid at position j in the core MSA sequence i. System 100 can also determine the probability distribution over the set of possible amino acids based on the relative frequency of occurrence of each possible amino acid at position j in the extra MSA sequences 1006 in the MSA sequence cluster 1008 corresponding to the core MSA sequence i. Then, System 100 can determine the position in the MSA representation by combining (e.g., concatenating) the following Embedding at (i) an embedding defining the identity of the amino acid at position j in the core MSA sequence i, and (ii) a probability distribution over the possible amino acids corresponding to position j in the core MSA sequence i.
[0288] In some cases, the (ground truth) protein structure of one or more of the core MSA sequences may be known. Specifically, for one or more of the core MSA sequences, the torsion angle values (e.g., ω angle, φ angle, ψ angle, etc.) between the bonds in the amino acids in the core MSA sequence may be known. If the torsion angle value of the amino acid in core MSA sequence i is known, then system 100 can generate the embedding at position at least in part based on the torsion angle value of amino acid j in core MSA sequence i. For example, the system can use one or more neural network layers to generate an embedding of the torsion angle value and concatenate the embedding of the torsion angle value to the embedding at position in the MSA representation.
[0289] Figure 17 Illustrates an example process 1100 for generating (initializing) corresponding pairwise embeddings 112 for each amino acid pair in a protein. Refer to Figure 1 The protein structure prediction system 100 described can implement the operations of process 1100.
[0290] System 100 can use the MSA representation 1102 of the protein to generate the pairwise embeddings 112. Refer to Figure 7 and Figure 16 describe in more detail the generation of the MSA representation of a protein. System 100 can generate the MSA representation 1102 of the protein based on the corresponding MSA representations of each amino acid chain in the protein. To generate the MSA representation of an amino acid chain in a protein, system 100 can use more MSA sequences (e.g., an order of magnitude more) than the MSA sequences used to generate the MSA representation 110 described in Figure 1 . Thus, the MSA representation 1102 used by system 100 to generate the pairwise embeddings 112 can have more rows (e.g., an order of magnitude more) than the MSA representation 110 described in Figure 1 . In some implementations, system 100 can use the additional MSA sequences 1006 described in Figure 16 to generate the MSA representation 1102.
[0291] After generating the MSA representation 1102, system 100 processes the MSA representation 1102 to generate pairwise embeddings 1104 according to the MSA representation 1102, e.g., by applying an outer product mean operation to the MSA representation 1102 and identifying the pairwise embeddings 1104 as the result of the outer product mean operation.
[0292] System 100 uses an embedded neural network 1106 to process the MSA representation 1102 and the pairwise embeddings 1104. The embedded neural network 1106 can update the MSA representation 1102 and the pairwise embeddings 1104 by sharing information between the MSA representation 1102 and the pairwise embeddings 1104. More specifically, the embedded neural network 1106 can alternate between updating the MSA representation 1102 based on the pairwise embeddings 1104 and updating the pairwise embeddings 1104 based on the MSA representation 1102.
[0293] The embedded neural network 1106 can have an architecture based on the Figures 2 to 5 described architecture of the embedded neural network, i.e., using row-wise self-attention blocks and column-wise self-attention blocks to update the pairwise embeddings 1104 and the MSA representation 1102.
[0294] In some implementations, the embedded neural network 1106 can use a column-wise "global" self-attention operation to update the embeddings in each column of the MSA representation 1102. More specifically, the embedded neural network 1106 can provide the embeddings in each column of the MSA representation to a column-wise global self-attention block to generate updated embeddings for each column of the current MSA representation. To implement global column-wise self-attention, the self-attention block can generate corresponding query embeddings for each embedding in the column and then average the query embeddings to generate a single "global" query embedding for the column. The column-wise self-attention block then uses the single global query embedding to perform a self-attention operation, which can reduce the complexity of the self-attention operation from quadratic (i.e., the number of embeddings per column) to linear. Using the global self-attention operation can reduce the computational complexity of the column-wise self-attention operation, enabling the column-wise self-attention operation to be performed on columns of the MSA representation 1102 that include a large number (e.g., thousands) of embeddings.
[0295] After updating the pairwise embeddings 1104 and the MSA representation 1102 using the embedded neural network 1106, the system 100 can identify the pairwise embeddings 112 as the updated pairwise embeddings generated by the embedded neural network 1106. The system 100 can discard the updated MSA representation generated by the embedded neural network 1106 or use it in any appropriate way.
[0296] As part of generating the paired embeddings 112, the system 100 can include relative position encoding information in the corresponding paired embeddings corresponding to each pair of amino acids in the protein. The system can include the relative position encoding information of pairs of amino acids included in the same amino acid chain in the corresponding paired embeddings by: calculating the signed difference representing the number of amino acids between the pair of amino acids in the amino acid chain, truncating the result to a predefined interval, using a one-hot encoded vector to represent the truncated value, applying a linear transformation to the one-hot encoded vector, and adding the result of the linear transformation to the corresponding paired embedding. The system can include the relative position encoding information of pairs of amino acids not included in the same amino acid chain in the corresponding paired embeddings by: adding a default encoding vector to the corresponding paired embedding to indicate that the pair of amino acids is not included in the same amino acid chain.
[0297] The system 100 can also generate the paired embeddings 112 based at least in part on a set of one or more template sequences 1110. Each template sequence 1110 is an MSA sequence of an amino acid chain in a protein, where the folded structure of the template sequence 1110 is known, for example, through physical experiments.
[0298] The system 100 can generate a corresponding template representation 1112 for each template sequence 1110. The template representation 1112 of the template sequence 1110 includes corresponding embeddings corresponding to each pair of amino acids in the template sequence. For example, the template representation 1112 of a template sequence of length n (i.e., having n amino acids) can be represented as an embedding array. The system 100 generates the embedding at position in the template representation 1112 of the template sequence 1110 based on, for example: (i) corresponding embeddings (e.g., one-hot embeddings) representing the identities of the amino acids at positions i and j in the template sequence, (ii) a unit vector defined by the spatial position difference of the corresponding alpha carbon atoms in the amino acids at positions i and j in the template sequence (i.e., in the folded structure of the template sequence), where the unit vector is calculated in the reference frame of amino acid i or amino acid j, and (iii) a discretized / binned representation of the distance between the spatial positions of the corresponding alpha carbon atoms in the amino acids at positions i and j in the template sequence.
[0299] System 100 can process each template representation using a sequence of one or more template update blocks 1114 to generate corresponding updated template representations 1116 corresponding to each template sequence 1110. The template update block can include, for example, a row-wise self-attention block (e.g., which updates the embeddings in each row of the template representation), a column-wise self-attention block (e.g., which updates the embeddings in each column of the template representation), and a transformation block (e.g., which applies one or more neural network layers to each embedding in the embeddings of the template representation).
[0300] After generating the updated template representation 1116, system 100 uses the updated template representation 1116 to update the pairwise embeddings 112. For example, system 100 can use "cross-attention" on the embeddings at the corresponding positions of the updated template representation 1116 to update the corresponding pairwise embeddings 112 at each position . In the cross-attention operation for updating the pairwise embeddings 112 at a position , the query embedding is generated from the pairwise embedding at the position , and the key embedding and value embedding are generated from the embeddings at the corresponding positions of the updated template representation 1116.
[0301] Updating the pairwise embeddings 112 using the template sequence 1110 enables system 100 to enrich the pairwise embeddings with information characterizing the protein structures of evolutionarily related template sequences 1110, thereby enhancing the information content of the pairwise embeddings 112 and improving the accuracy of the protein structures predicted using the pairwise embeddings.
[0302] As previously described, the pathogenicity prediction system 120 described in this specification has a variety of possible applications, including using pathogenicity scores to identify the causes of diseases in animal or plant subjects.
[0303] In another example application, the system is used to obtain mutated proteins for pest control. This can include determining the pathogenicity scores of multiple protein mutations (e.g., relative to a reference amino acid sequence), each mutation defining a corresponding mutated protein. Then, one of these mutations can be selected based on the pathogenicity scores of these mutations to define the mutated protein for pest control. For example, the pathogenicity scores can be ranked, and the most pathogenic mutation can be selected according to the scores, e.g., the mutation with the highest score. As another example, a mutation with a medium pathogenicity score can be selected, e.g., in order to facilitate the spread of the mutation in the pest population. To use the mutated protein for pest control, the pests can be bred to produce the selected mutated protein.
[0304] In another example application, the system is used to screen one or more living organisms, such as animals (including humans) or plants, to determine whether there are proteins with pathogenic mutations. This can include obtaining the amino acid sequences of the proteins in each of these organisms, for example, by the gene sequencing described previously. Generally, some of these sequences will correspond to mutated versions of the protein. For each amino acid sequence obtained that includes a mutation (e.g., compared to the reference amino acid sequence of a reference protein), the system determines the pathogenicity score of the mutation. Then, each determined pathogenicity score can be processed to determine whether the mutation is pathogenic, for example, by classifying the pathogenicity score as benign or pathogenic. The pathogenicity score or classification can, for example, provide input for clinical decisions on whether or how to provide treatment.
[0305] In another example application, the system is used to determine the pathogenicity level of bacteria or viruses to living organisms. This can include obtaining the amino acid sequences of artificial proteins, where the artificial proteins are produced within a living organism, such as by a living organism or bacteria, as a result of the living organism being infected by bacteria or viruses. The artificial proteins can be toxins produced by bacteria or proteins produced by a living organism, such as viral proteins, and can be mutations of proteins that naturally occur in the organism. The pathogenicity score of the mutation of the naturally occurring protein can be determined and used to determine the pathogenicity level of the bacteria or viruses. For example, the pathogenicity score can be classified as benign or pathogenic, or the relative pathogenicity of different strains of bacteria or viruses can be compared. This can be used, for example, to determine whether the proportion of pathogenic strains in a population is increasing or decreasing, for example, to trigger an alarm; or to select bacteria or viruses for vaccine production based on the determined pathogenicity level, for example, to identify bacterial or viral strains with relatively low pathogenicity for vaccine production.
[0306] Figure 18A and Figure 18BShows the performance (e.g., classification accuracy) of the systems described in this specification on clinically curated classification benchmarks, as evaluated by the area under the receiver operator curve (auROC). Error bars show the 95% confidence intervals of 1000 bootstrap resamplings. Chart 1802 compares the performance of the current system (“AlphaMissense”) with other predictors on 82,868 ClinVar missense variants held out. Chart 1804 compares the performance of the current system (“AlphaMissense”) and other predictors on 868 cancer mutations in hotspots from 202 driver genes versus 1734 randomly selected negative variants. Chart 1806 compares the performance of the current system (“AlphaMissense”) and other predictors in distinguishing de novo variants from patients in the Deciphering Developmental Disorders (DDD) cohort and healthy controls; 353 patient variants and 57 control variants in 215 DDD-related genes were considered. Chart 1808 compares the performance of the current system (“AlphaMissense”) and other predictors in classifying ClinVar variants (3430 pathogenic variants and 1185 benign variants) in regions of higher evolutionary constraint.
[0307] This specification uses the term “configured” when referring to systems and computer program components. For a system of one or more computers that is configured to perform a particular operation or action, it means that the system has installed thereon software, firmware, hardware, or a combination thereof that in operation causes the system to perform the operation or action. For one or more computer programs that are configured to perform a particular operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0308] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0309] The term "data processing apparatus" refers to data processing hardware and includes all kinds of devices, apparatus, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The apparatus may also be or further include special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to the hardware, the apparatus may optionally include code that creates an execution environment for a computer program, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0310] A computer program, which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language (including compiled or interpreted languages or declarative or procedural languages); and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may or may not correspond to a file in a file system. A program may be stored in a part of a file that holds other programs or data (such as one or more scripts in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (such as files that hold one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0311] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same one or more computers.
[0312] The processes and logical flows described in this specification may be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows may also be performed by special purpose logic circuitry, such as an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0313] A computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data therefrom or to transfer data thereto or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0314] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0315] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page in response to a request received from a web browser on the user's device. Further, a computer can interact with a user by sending text messages or other forms of messages to a personal device, such as a smart phone running a messaging application, and receiving a response message in return from the user.
[0316] The data processing device for implementing a machine learning model may also include, for example, a special purpose hardware accelerator unit for processing the general and computationally intensive parts of machine learning training or production (i.e., inference, workload).
[0317] Machine learning frameworks (e.g., TensorFlow framework, Microsoft Cognitive Toolkit framework, Apache Singa framework, or Apache MXNet framework) can be used to implement and deploy machine learning models.
[0318] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.
[0319] The computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and causing a client-server relationship between them. In some embodiments, the server sends data (e.g., an HTML page) to a user device, e.g., for the purpose of displaying data to and receiving user input from a user interacting with the device acting as a client. Data generated at the user device, e.g., the result of a user interaction, can be received at the server from the device.
[0320] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination may cover a sub-combination or a variant of a sub-combination.
[0321] Similarly, although the operations are depicted in the drawings and recited in the claims in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0322] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.
[0323] Certain novel aspects of the subject matter of this specification are set forth in the appended claims and further described in Appendix A.
Claims
1. A method performed by one or more computers, the method comprising: Generating a pathogenicity score that characterizes the likelihood that a protein mutation is a pathogenic mutation, wherein the mutation modifies the amino acid sequence of the protein by replacing an original amino acid with a replacement amino acid at a mutation position in the amino acid sequence of the protein, wherein generating the pathogenicity score comprises: Generating a network input for a pathogenicity prediction neural network, wherein the network input comprises an MSA representation representing a multiple sequence alignment (MSA) of the protein; Processing the network input using the pathogenicity prediction neural network to generate a score distribution over an amino acid set; and Generating the pathogenicity score based on a difference between (i) the score of the original amino acid under the score distribution and (ii) the score of the replacement amino acid under the score distribution.
2. The method according to claim 1, wherein the MSA representation comprises corresponding embeddings corresponding to each position in the amino acid sequence of each protein in the MSA; And wherein generating the MSA representation comprises: Masking an embedding in the MSA representation corresponding to the mutation position in the amino acid sequence of the protein. The method according to claim 2, wherein processing the network input using the pathogenicity prediction neural network to generate the score distribution over the amino acid set comprises: Processing the MSA representation using an embedding subnetwork of the pathogenicity prediction neural network to generate an updated MSA representation, wherein the updated MSA representation comprises corresponding updated embeddings corresponding to each position in the amino acid sequence of each protein in the MSA; and Processing the updated embedding corresponding to the mutation position in the amino acid sequence of the protein using a projection subnetwork of the pathogenicity prediction neural network to generate the score distribution over the amino acid set.
3. The method according to claim 2, wherein the embedding subnetwork of the pathogenicity prediction neural network comprises one or more self-attention neural network layers.
4. The method according to claim 3, wherein the embedding subnetwork of the pathogenicity prediction neural network comprises one or more row-wise or column-wise self-attention neural network layers.
5. The method according to any one of the preceding claims, wherein the pathogenicity prediction neural network has been trained by operations comprising: Training the pathogenicity prediction neural network to perform a protein unmasking task; and Training the pathogenicity prediction neural network to perform a pathogenicity prediction task.
6. The method according to claim 5, wherein training the pathogenicity prediction neural network to perform the protein unmasking task comprises, for each of a plurality of training proteins: Generating a network input for the pathogenicity prediction neural network based on the training protein, wherein: The network input comprises an MSA representation representing an MSA of the training protein; And The MSA representation comprises corresponding masked embeddings corresponding to each of one or more masked positions in the amino acid sequence of the training protein. Process the network input of the pathogenicity prediction neural network to generate a network output that defines a corresponding prediction for the amino acid identity at each masked position in the amino acid sequence of the training protein; and Backpropagate the gradient of the masking loss through the pathogenicity prediction neural network, where the masking loss measures the accuracy of the corresponding predictions generated by the pathogenicity prediction neural network for each masked position.
7. The method according to any one of claims 5 to 6, further comprising: Training the pathogenicity prediction neural network to perform a protein structure prediction task.
8. The method according to claim 7, wherein training the pathogenicity prediction neural network to perform the protein structure prediction task comprises, for each training protein in a plurality of training proteins: Generating a network input for the pathogenicity prediction neural network based on the training protein, where the network input includes an MSA representation representing the MSA of the training protein; Processing the network input using the pathogenicity prediction neural network to generate an intermediate output of the pathogenicity prediction neural network; Processing the intermediate output of the pathogenicity prediction neural network using a folding neural network to generate a set of structure parameters that define the predicted structure of the training protein; and Backpropagating the gradient of the structure loss through the folding neural network into the pathogenicity prediction neural network, where the structure loss measures the error of the predicted structure of the training protein.
9. The method according to any one of claims 5 to 8, wherein the pathogenicity prediction neural network is pre-trained to perform the protein unmasking task and the protein structure prediction task before being trained to perform the pathogenicity prediction task.
10. The method according to any one of the preceding claims, wherein the pathogenicity prediction neural network has been trained over at least two training phases including a first training phase and a second training phase; wherein the first training phase comprises: Training the pathogenicity prediction neural network on a set of pathogenicity training examples, where each pathogenicity training example defines: (i) the amino acid sequence of a training protein, (ii) one or more mutations in the amino acid sequence of the training protein, and (iii) a corresponding target pathogenicity score for each of the one or more mutations; Determining a corresponding filtering score for each training example; and Filtering the set of pathogenicity training examples using the filtering score.
11. The method according to claim 10, wherein determining the filtering score for each training example comprises: For each mutation included in the training example, using the pathogenicity prediction neural network to determine the predicted pathogenicity score of the mutation; and For each mutation included in the training example, determine an error score for the mutation based on the error between: (i) the predicted pathogenicity score of the mutation, and (ii) the target pathogenicity score of the mutation; and Determine the filtering score of the training example based at least in part on the error scores of the mutations included in the training example.
12. The method according to any one of claims 10 or 11, wherein filtering the set of pathogenicity training examples using the filtering score comprises: Determine a probability distribution on the set of pathogenicity training examples based on the filtering score; Sample a subset of the pathogenicity training examples from the set of pathogenicity training examples according to the probability distribution; And Remove the sampled subset of the pathogenicity training examples from the set of pathogenicity training examples.
13. The method according to any one of claims 10 to 12, wherein the second training phase comprises: Retrain the pathogenicity prediction neural network on the filtered set of pathogenicity training examples.
14. The method according to any one of the preceding claims, further comprising: Determine whether the pathogenicity score of the mutation satisfies a threshold; And In response to determining that the pathogenicity score of the mutation satisfies the threshold: Classify the mutation as a pathogenic mutation.
15. The method according to claim 14, wherein determining that the pathogenicity score of the mutation satisfies the threshold comprises: Determine that the pathogenicity score of the mutation exceeds the threshold.
16. The method according to any one of the preceding claims, further comprising using the pathogenicity score to identify the cause of a disease in a subject.
17. The method according to any one of claims 1 to 16, wherein the method is used to obtain a mutated protein for pest control, and the method further comprises: Determine the pathogenicity scores of a plurality of mutations of a reference protein, each mutation defining a corresponding mutated protein; And Based on the determined pathogenicity scores, select one of the mutations to define the mutated protein for pest control.
18. The method according to any one of claims 1 to 16, wherein the method is used to screen for the presence of a protein with a pathogenic mutation in one or more living organisms, and the method further comprises: Obtain the amino acid sequence of the protein for each organism in the organisms; For each amino acid sequence including a mutation in the obtained amino acid sequences, determine the pathogenicity score of the mutation; And Process each determined pathogenicity score to determine whether the mutation is pathogenic.
19. The method according to any one of claims 1 to 16, wherein the method is used to determine the degree of pathogenicity of a bacterium or virus to a living organism, and the method further comprises: Obtain the amino acid sequence of an artificial protein, wherein the artificial protein is made inside the living organism due to the living organism being infected by the bacterium or virus, and wherein the artificial protein is a mutation of a naturally occurring protein in the organism; Determine the pathogenicity score of the mutation of the naturally occurring protein; Determine the pathogenicity level of the bacterium or virus based on the pathogenicity score.
20. The method according to any one of the preceding claims, further comprising providing the pathogenicity score for display on a user interface.
21. A system, comprising: One or more computers; And One or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 20.
22. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1 to 20.
23. A method, comprising: Maintain a mutation-pathogenicity database that includes data defining a corresponding pathogenicity score for each mutation of one or more mutations of each protein among a plurality of proteins, wherein the pathogenicity score has been generated using the method according to any one of claims 1 to 20; Obtain genetic material from a subject; Based on the genetic material from the subject, identify one or more proteins encoded in the genetic material of the subject each having one or more mutations; Determine a diagnosis for the subject based on: (i) the mutation-pathogenicity database, and (ii) the protein mutations identified from the genetic material of the subject.
24. The method according to claim 23, wherein the mutation-pathogenicity database includes data characterizing a plurality of proteins in the human proteome.
25. The method according to claim 24, wherein the mutation-pathogenicity database includes data characterizing all proteins in the human proteome.
26. The method according to any one of claims 23 to 25, wherein the subject is a human subject.
27. The method according to any one of claims 23 to 26, wherein determining the diagnosis for the subject comprises: Determine that a specific protein mutation identified from the genetic material of the subject is associated with a pathogenicity score that meets a threshold in the mutation-pathogenicity database; And In response, diagnose the subject with a medical condition associated with the specific protein mutation.