Predicting protein structure using auxiliary folding networks

By predicting protein structure through a single forward pass of a neural network system, the problem of high computational resource consumption in existing technologies has been solved, enabling rapid, efficient, and high-quality protein structure prediction, which promotes protein biology research and drug development.

CN116325002BActive Publication Date: 2026-03-27GDM HOLDINGS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing protein structure prediction methods are typically time-consuming and computationally intensive, making it difficult to efficiently generate high-quality protein structure predictions.

Method used

A neural network system, including a main folded neural network and an auxiliary folded neural network, is used to predict protein structures through a single forward pass. An embedding neural network is used to process multiple sequence alignments and embedding information, and geometric attention operations are combined to generate high-quality structure predictions.

Benefits of technology

This enables rapid, low-resource-consumption, high-quality protein structure prediction, improving the accuracy and efficiency of prediction and promoting the understanding of protein biological functions and drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116325002B_ABST
    Figure CN116325002B_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a structure prediction neural network, including an embedding neural network and a main fold neural network. According to one aspect, a method includes obtaining a training network input characterizing a training protein; processing the training network input using the embedding neural network and the main fold neural network to generate a main structure prediction; for each auxiliary fold neural network in a set of one or more auxiliary fold neural networks, processing at least a corresponding intermediate output of the embedding neural network to generate an auxiliary structure prediction; determining a gradient of an objective function, the objective function including a respective auxiliary structure loss term for each of the auxiliary fold neural networks; and updating current values of embedding network parameters and main fold parameters based on the gradient.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 118,921, filed November 28, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This manual relates to the prediction of protein structure. Background Technology

[0004] Proteins are defined by a sequence of one or more amino acids (“chains”). Amino acids are organic compounds that include amino and carboxyl functional groups, as well as amino acid-specific side chains (i.e., atomic groups). Protein folding refers to the physical process by which one or more amino acid sequences fold into a three-dimensional (3-D) conformation. The structure of a protein defines the 3-D conformation of the atoms in the protein’s amino acid sequence after protein folding. When in a sequence linked by peptide bonds, an amino acid can be referred to as an amino acid residue.

[0005] Machine learning models can be used for prediction. A machine learning model takes input and generates an output, such as a predicted output, based on that input. Some machine learning models are parametric models and generate outputs based on the received input and the model's parameter values. Some machine learning models are deep models, which employ multiple layers to generate outputs for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, each of which applies a non-linear transformation to the received input to generate an output. Summary of the Invention

[0006] This specification describes a neural network system implemented as a computer program on one or more computers at one or more locations for predicting protein structures.

[0007] As used throughout this specification, the term "protein" can be understood to mean any biomolecule specified by one or more amino acid sequences (or "chains"). For example, the term protein can refer to a protein domain, such as a portion of an amino acid chain of a protein, which is capable of folding almost independently of the rest of the protein. As another example, the term protein can refer to a protein complex, that is, a protein complex comprising multiple amino acid chains that fold together to form a protein structure.

[0008] Multiple sequence alignment (MSA) of an amino acid chain in a protein specifies the alignment of that amino acid chain with multiple other amino acid chains—e.g., from other proteins, such as homologous proteins. More specifically, an MSA can define the correspondence between positions in that amino acid chain and corresponding positions in multiple other amino acid chains. An MSA of an amino acid chain can be generated, for example, by processing a database of amino acid chains using any suitable computational sequence alignment technique (e.g., progressive alignment construction). The amino acid chains in an MSA can be understood as having evolutionary relationships, for example, where each amino acid chain in the MSA may share a common ancestor. The correlations between amino acids in the amino acid chains in an MSA of an amino acid chain can encode information related to predicting the structure of the amino acid chain.

[0009] The “embedding” of an entity (e.g., an amino acid pair) can refer to a representation of an entity as an ordered set of numerical values, such as a vector or matrix of numerical values.

[0010] The structure of a protein can be defined by a set of structural parameters. This set of structural parameters can be represented as an ordered set of values. Below are some examples of possible structural parameters used to define the structure of a protein, described in more detail.

[0011] In one example, for each amino acid in a protein, the structural parameters that define the structure of the protein include: (i) positional parameters and (ii) rotational parameters.

[0012] The position parameter of an amino acid specifies the predicted 3D spatial position of a designated atom within the amino acid in the structure of a protein. The designated atom can be an alpha carbon atom in the amino acid, i.e., a carbon atom bonded to the amino group, carboxyl group, and side chain. As another example, the designated atom can be a beta carbon atom in the amino acid. The position parameter of an amino acid can be represented in any suitable coordinate system, such as a three-dimensional [x,y,z] Cartesian coordinate system.

[0013] The rotation parameters of an amino acid can specify its predicted "orientation" within the protein structure. More specifically, the rotation parameters can specify a 3-D spatial rotation operation, which, when applied to a coordinate system of position parameters, causes the three "backbone" atoms in the amino acid to take fixed positions relative to the rotation coordinate system. The three backbone atoms in an amino acid can refer to a series of nitrogen, alpha carbon, and carbonyl carbon atoms linked together in the amino acid. The rotation parameters of an amino acid can be represented, for example, as an orthogonal 3×3 matrix with a determinant of 1.

[0014] Typically, the position and rotation parameters of an amino acid define its egocentric reference frame. In this frame, each amino acid's side chain can start from the origin and proceed along the first bond of the side chain (i.e., the alpha-beta bond) in a defined direction.

[0015] In another example, the structural parameters defining the structure of a protein can include a “distance map” that characterizes the estimated distances (e.g., measured in angstroms) between each pair of amino acids in the protein. The distance map can characterize the estimated distances between amino acid pairs, for example, through a probability distribution of a set of possible distances between amino acid pairs.

[0016] In another example, structural parameters that define the structure of a protein can define the three-dimensional (3D) spatial location of each atom in each amino acid within the protein's structure.

[0017] The protein structure prediction system described herein can be used to obtain ligands, such as ligands for drugs or industrial enzymes. For example, methods for obtaining ligands may include obtaining the target amino acid sequence, particularly the amino acid sequence of a target protein (e.g., a drug target), and using the protein structure prediction system to process the input based on the target amino acid sequence to determine the (tertiary) structure of the target protein, i.e., predicting the protein structure. The method may then include evaluating the interaction between one or more candidate ligands and the target protein structure. The method may also include selecting one or more candidate ligands as ligands based on the evaluation results of the interaction.

[0018] In some embodiments, evaluating the interaction may include assessing the binding of a candidate ligand to a target protein structure. For example, evaluating the interaction may include identifying ligands that bind with sufficient affinity to produce a biological effect. In some other embodiments, evaluating the interaction may include assessing the association of a candidate ligand to a target protein structure that has an impact on the function of the target protein (e.g., an enzyme). The assessment may include assessing the affinity between the candidate ligand and the target protein structure, or assessing the selectivity of the interaction. Candidate ligands may be selected based on the candidate ligand with the highest affinity.

[0019] Candidate ligands (or multiple ligands) can be derived from a database of candidate ligands, and / or can be derived by modifying ligands in the database, for example by modifying the structure or amino acid sequence of the candidate ligands, and / or by stepwise or iterative assembly / optimization of the candidate ligands.

[0020] The evaluation of the interaction between a candidate ligand and the target protein structure can be performed using computer-aided methods, wherein a graphical model displaying the candidate ligand and the target protein structure is used for user manipulation, and / or the evaluation can be performed partially or fully automatically, for example using standard molecular (protein-ligand) docking software. In some embodiments, the evaluation may include determining the interaction score of the candidate ligand, wherein the interaction score comprises a measurement of the interaction between the candidate ligand and the target protein. The interaction score may depend on the strength and / or specificity of the interaction, for example, on the fraction of the binding free energy. The candidate ligand may be selected based on its score, for example, when the candidate ligand has the best interaction score.

[0021] In some embodiments, the target protein includes a receptor or enzyme, and the ligand is an agonist or antagonist of the receptor or enzyme. In some embodiments, the method can be used to identify the structure of a cell surface marker. This can then be used to identify ligands, such as antibodies or markers (e.g., fluorescent markers), that bind to the cell surface marker. This can be used to identify and / or treat cancer cells.

[0022] In some embodiments, the ligand is a drug, and a predicted structure for each of a plurality of target proteins is determined, and the interaction of one or more candidate ligands with the predicted structure of each of the target proteins is evaluated. One or more candidate ligands can then be selected to obtain a ligand that interacts with each (functional) target protein, or to obtain a ligand that interacts with only one (functional) target protein. For example, in some embodiments, it is desirable to obtain a drug that is effective against multiple drug targets. Alternatively or concurrently, it is desirable to screen for off-target effects of the drug. For example, in agriculture, it can be useful to determine that a drug designed for one plant species does not interact with another different plant species and / or animal species.

[0023] In some embodiments, candidate ligands (or multiple ligands) may include small molecule ligands, such as organic compounds with a molecular weight <900 Daltons. In other embodiments, candidate ligands (or multiple ligands) may include peptide ligands, i.e., peptide ligands defined by an amino acid sequence.

[0024] In some cases, protein structure prediction systems can be used to determine the structure of candidate polypeptide ligands (e.g., ligands for drugs or industrial enzymes). Their interaction with the target protein structure can then be evaluated; the target protein structure can already be determined using structure prediction neural networks or conventional physical probing techniques—such as X-ray crystallography and / or magnetic resonance imaging and / or cryo-electron microscopy.

[0025] On the other hand, a method is provided for obtaining peptide ligands (e.g., molecules or their sequences) using a protein structure prediction system. This method may include obtaining the amino acid sequences of one or more candidate peptide ligands. The method may also include using a protein structure prediction system to determine the (tertiary) structure of the candidate peptide ligands. The method may further include obtaining the target protein structure of the target protein through bioinformatics and / or physical probing, and evaluating the interaction between the structure of each of the one or more candidate peptide ligands and the target protein structure. The method may also include selecting one or more candidate peptide ligands as peptide ligands based on the evaluation results.

[0026] As previously described, evaluating interactions can include assessing the binding of candidate peptide ligands to target protein structures, such as identifying ligands that bind with sufficient affinity to achieve biological effects, and / or assessing the association of candidate peptide ligands with target protein structures that influence the function of the target protein (e.g., enzymes), and / or assessing the affinity between candidate peptide ligands and target protein structures, or assessing the selectivity of the interaction. In some embodiments, the peptide ligand may be an aptamer. Similarly, peptide candidate ligands (or multiple ligands) may be selected based on the peptide candidate ligand with the highest affinity.

[0027] As previously described, the selected peptide ligand may include a receptor or an enzyme, and the ligand may be an agonist or antagonist of the receptor or enzyme. In some embodiments, the peptide ligand may include an antibody, and the target protein may include an antibody target, such as a virus, particularly a viral capsid protein, or a protein expressed on cancer cells. In these embodiments, the antibody binds to the antibody target to provide a therapeutic effect. For example, the antibody may bind to the target and act as an agonist of a specific receptor; alternatively, the antibody may prevent another ligand from binding to the target and thus prevent activation of the associated biological pathway.

[0028] The implementation of this method may also include synthesis, i.e., the preparation of small molecule or peptide ligands. The ligands can be synthesized by any conventional chemical technique and / or can be already available, for example, derived from a compound library or synthesized using combinatorial chemistry.

[0029] This method may also include testing the bioactivity of the ligand in vitro and / or in vivo. For example, the ADME (absorption, distribution, metabolism, excretion) and / or toxicological properties of the ligand may be tested to screen out unsuitable ligands. Testing may include, for example, contacting the candidate small molecule or peptide ligand with the target protein and measuring changes in protein expression or activity.

[0030] In some embodiments, candidate (peptide) ligands may include: separating antibodies, fragments of separating antibodies, monovariable domain antibodies, bispecific or multispecific antibodies, multivalent antibodies, bivariable domain antibodies, immunoconjugates, fibronectin molecules, adnectin, DARPin, aviser, affinity molecules, anticarrier proteins, affilin, protein epitope mimics, or combinations thereof. Candidate (peptide) ligands may include antibodies with mutated or chemically modified amino acid Fc regions, for example, which, compared to wild-type Fc regions, inhibit or reduce ADCC (antibody-dependent cytotoxicity) activity and / or increase half-life. Candidate (peptide) ligands may include antibodies with different CDRs (complementarity-determining regions).

[0031] The protein structure prediction system described herein can also be used to obtain diagnostic antibody biomarkers for diseases. A method is also provided in which, for each of one or more candidate antibodies, such as as described above, the method uses the protein structure prediction system to determine the predicted structure of the candidate antibody. The method may also involve obtaining the target protein structure of the target protein, evaluating the interaction between the predicted structure of each of one or more candidate antibodies and the target protein structure, and selecting one of one or more candidate antibodies as a diagnostic antibody biomarker based on the evaluation results, for example, selecting one or more candidate antibodies with the highest affinity for the target protein structure. The method may include the preparation of the diagnostic antibody biomarker. The diagnostic antibody biomarker can be used to diagnose diseases by detecting whether it binds to a target protein in a sample obtained from a patient (e.g., a body fluid sample). As described above, corresponding techniques can be used to obtain therapeutic antibodies (peptide ligands).

[0032] Misfolded proteins are associated with many diseases. Therefore, in another aspect, a method is provided to identify the presence of protein misfolding diseases using a protein structure prediction system. This method may include obtaining the amino acid sequence of a protein and using a protein structure prediction system to determine the protein's structure. The method may also include obtaining the structure of a protein version obtained from a human or animal body, for example, through conventional (physical) methods. The method then includes comparing the protein's structure with the structure of the version obtained from the body and identifying the presence of a protein misfolding disease based on the comparison result. That is, misfolding of the protein version from the body can be determined by comparing it with the structure determined through bioinformatics.

[0033] Typically, identifying the presence of a protein misfolding disorder may involve obtaining the amino acid sequence of a protein, using the amino acid sequence to determine the protein's structure, as described herein, and comparing the protein's structure to the structure of a baseline version of the protein. The presence of the protein misfolding disorder is identified based on the results of the comparison. For example, the structures being compared may be those of a mutant and a wild-type protein. In this implementation, a wild-type protein may be used as the baseline version, but in principle, either can be used as the baseline version.

[0034] In some other respects, computer-implemented methods, such as those described above or herein, can be used to identify active / binding / blocking sites on target proteins from their amino acid sequences.

[0035] According to one aspect, a method for training a structure prediction neural network is provided, wherein the structure prediction neural network includes: (i) an embedding neural network having a plurality of embedding parameters and configured to receive network inputs characterizing a protein and process the network inputs according to the embedding parameters to generate an embedding output about the network inputs; and (ii) a master folding neural network having a plurality of master folding parameters and configured to receive embedding outputs and process master folding inputs according to the master folding parameters to generate a master structure prediction defining a predicted structure of a protein, the method comprising: obtaining training network inputs characterizing a training protein and data on a target protein structure specifying the training protein; processing the training network inputs using the embedding neural network and according to current values ​​of the embedding parameters to generate a training embedding output about the training network inputs; processing the training embedding outputs using the master folding neural network and according to current values ​​of the master folding parameters to generate a master structure prediction defining a master predicted structure of the training protein; for each... Each of one or more auxiliary folding neural networks in a set of corresponding auxiliary folding parameters is used to process at least the corresponding intermediate output of the embedded neural network based on the current value of the corresponding auxiliary folding parameter of the auxiliary folding neural network to generate an auxiliary structure prediction that defines the auxiliary prediction structure of the training protein; the gradient of an objective function is determined, the objective function comprising: a main structure loss term, representing the similarity between (i) the main prediction structure defined by the main structure prediction and (ii) the target protein structure of the training protein; and a corresponding auxiliary structure loss term for each auxiliary folding neural network, representing the similarity between (i) the auxiliary prediction structure defined by the auxiliary structure prediction generated by the auxiliary structure prediction neural network and (ii) the target protein structure of the training protein; and the current values ​​of the embedded network parameters, the main folding parameter, and the corresponding auxiliary folding parameters of the one or more auxiliary folding neural networks are updated based on the gradient.

[0036] In some implementations, each auxiliary folded neural network has the same neural network architecture as the main folded neural network.

[0037] In some implementations, an ensemble of one or more auxiliary folded neural networks includes multiple auxiliary folded neural networks, and an objective function constrains the auxiliary folded neural networks to share parameter values.

[0038] In some implementations, the objective function constrains the auxiliary folded neural network and the main folded neural network to share parameter values.

[0039] In some implementations, the network input includes a corresponding initial pair embedding for each amino acid pair in the protein, the embedding neural network includes an update block sequence, and each update block performs the following operations: receiving a block input including a corresponding current pair embedding for each amino acid pair in the protein; and updating the corresponding current pair embedding for each amino acid pair in the protein to generate a corresponding updated pair embedding for each amino acid pair in the protein, and the embedding output includes at least the updated pair embedding generated by the last update block in the sequence.

[0040] In some implementations, each auxiliary folding neural network corresponds to a different update block in the sequence that is not the last update block in the sequence, and each auxiliary folding neural network is configured to receive at least the pair embeddings of updates generated by the corresponding update block as input.

[0041] In some implementations, the network input also includes an initial multiple sequence alignment (MSA) embedding, which represents the corresponding multiple sequence alignment for each strand in the protein, and the block input for each update block also includes the current MSA embedding; and the operation performed by each update block also includes updating the current MSA embedding to generate an updated MSA embedding.

[0042] In some implementations, the input to each auxiliary folded neural network also includes an updated MSA embedding generated by the corresponding update block.

[0043] In some implementations, each auxiliary folding neural network is configured to: generate a transformed structural prediction from the auxiliary structural prediction generated by the auxiliary folding neural network, the transformed structural prediction having the same dimension as the updated pair embedding; combine the transformed structural prediction with the updated pair embedding to generate a further updated pair embedding; and provide the further updated pair embedding as input to an update block in the sequence following the current update block.

[0044] In some implementations, the assisted folding prediction includes structural parameters that specify the predicted 3-D spatial location of a specified atom in the amino acid within the protein structure for each amino acid; and generating the transformed structural prediction includes: generating a distance map from the predicted 3-D spatial locations of the amino acids specified by the structural parameters, wherein for each pair of amino acids in the protein, the distance map characterizes the corresponding estimated distance between the pairs of amino acids in the protein structure; and generating a transformed distance map from the distance map having the same dimensions as the updated pair embeddings.

[0045] In some implementations, the auxiliary folding prediction includes specifying structural parameters for a distance map, whereby, for each amino acid pair in the protein, the distance map characterizes the estimated distance between the corresponding amino acid pairs in the protein's structure; and

[0046] Generating the transformed structure prediction involves generating a transformed distance map with the same dimensions as the initial pair of embeddings from the distance map specified by the structure parameters.

[0047] In some implementations, the method further includes: obtaining new network inputs characterizing the new protein after training; and processing the new network inputs using a trained structure prediction neural network to generate new master structure predictions that define the predicted structure of the new protein.

[0048] Specific embodiments of the subject matter described in this specification may be implemented in order to achieve one or more of the following advantages.

[0049] The system described in this specification uses a neural network to predict the structure of proteins. At least during training, this neural network includes one or more auxiliary folded neural networks in addition to a main folded neural network that generates the actual predicted output of the protein structure and an embedded neural network that generates the input to the main folded neural network. Each auxiliary folded neural network corresponds to one of a plurality of update blocks within the embedded neural network and receives input from one of the plurality of update blocks within the embedded neural network. Including auxiliary folded neural networks as part of the neural network during training allows the embedded neural network to receive richer training signals and thus learn to generate higher-quality embeddings. For example, compared to learning based solely on errors in predictions made by the main folded neural network and optionally on errors in the outputs of one or more auxiliary tasks, once the neural network has been trained, this results in higher-quality predictions being generated by the main folded neural network.

[0050] In some cases, auxiliary neural networks are also included in the main neural network after training, and each auxiliary neural network uses the auxiliary structure predictions generated by the main neural network to modify the output of the corresponding update block. This can further improve the quality of the generated embeddings provided as input to the main folding neural network, and further improve the quality of the predictions generated by the main folding neural network after training.

[0051] The system described in this specification can predict protein structures using an ensemble of jointly trained neural networks in a single forward pass, potentially in less than a second. In contrast, some conventional systems predict protein structures by optimizing a scalar score function through an expanded search process across the space of possible protein structures, using techniques such as simulated annealing or gradient descent. Such a search process can require millions of search iterations and consume hundreds of central processing unit (CPU) hours. Predicting protein structures via a single forward pass of an ensemble of neural networks allows the system described in this specification to consume far fewer computational resources (e.g., memory and computing power) than systems that predict protein structures through an iterative search process.

[0052] The structure of a protein determines its biological function. Therefore, determining protein structure can facilitate the understanding of life processes (e.g., the mechanisms of many diseases) and the design of proteins (e.g., as drugs, or as enzymes for industrial processes). For example, which molecules (e.g., drugs) will bind to a protein (and where this binding will occur) depends on the protein's structure. Since the effectiveness of drugs can be influenced by the extent to which they bind to proteins (e.g., in the blood), determining the structure of different proteins is an important aspect of drug development. However, determining protein structure using physical experiments (e.g., via X-ray crystallography) can be time-consuming and very expensive. Therefore, the protein prediction system described in this specification can facilitate biochemical research and engineering fields involving proteins (e.g., drug development).

[0053] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims. Attached Figure Description

[0054] Figure 1 An example protein structure prediction system is shown.

[0055] Figure 2 An example architecture for embedding a neural network is shown.

[0056] Figure 3 An example architecture is shown that embeds an update block of a neural network.

[0057] Figure 4This shows an example schema for the MSA update block.

[0058] Figure 5 An example schema for updating blocks is shown.

[0059] Figure 6 An example architecture of a folded neural network is shown.

[0060] Figure 7 The torsion angle between bonds in an amino acid is shown.

[0061] Figure 8 These are illustrations of unfolded and folded proteins.

[0062] Figure 9 This is a flowchart of an example process for training a structure prediction neural network.

[0063] Figure 10 An example process is shown for generating an MSA representation of the amino acid chain in a protein.

[0064] Figure 11 An example process is shown for generating corresponding pair embeddings for each amino acid pair in a protein.

[0065] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation

[0066] Figure 1 An example protein structure prediction system 100 is shown. The protein structure prediction system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components, and techniques described below are implemented.

[0067] System 100 is configured to process data of one or more amino acid chains 102 defining protein 104 to generate a master structure prediction 106, which defines a set of structural parameters that define a predicted protein structure 108—that is, a prediction of the structure of protein 104. In other words, the predicted structure 108 of protein 104 can be defined by a set of structural parameters (e.g., included in the master structure prediction 106) that collectively define the predicted three-dimensional structure of the protein after protein folding.

[0068] The structural parameters for predicting the protein structure 108 can be defined as described above. For example, they may include, for instance, positional and rotational parameters of each amino acid in protein 104, a distance map characterizing the estimated distances between each pair of amino acids in the protein, the corresponding spatial position of each atom or backbone atom in each amino acid in the protein's structure, or combinations thereof, as described above.

[0069] In order to generate a master structure prediction 106 that defines the predicted protein structure 108, the system 100 may generate: (i) a multiple sequence alignment (MSA) representation of the protein 110 (in some embodiments), and (ii) a set of “pair” embeddings of the protein 112, as will be described in more detail below.

[0070] The MSA representation of a protein includes a corresponding representation of the MSA for each amino acid chain in the protein. The MSA representation of an amino acid chain in a protein can be represented as an embedded M×N array (i.e., an embedded 2D array with M rows and N columns), where N is the number of amino acids in the amino acid chain. Each row of the MSA representation can correspond to the corresponding MSA sequence of the amino acid chain in the protein. Reference Figure 10 An exemplary process is described for generating an MSA representation of the amino acid chain in a protein.

[0071] System 100 generates the MSA representation of protein 104 from the MSA representation of the amino acid chain in the protein 110.

[0072] If a protein consists of only a single amino acid chain, then system 100 can recognize the MSA representation 110 of protein 104 as the MSA representation of a single amino acid chain in the protein.

[0073] If a protein comprises multiple amino acid chains, system 100 can generate an MSA representation 110 of the protein by assembling the MSA representations of the amino acid chains in the protein into an embedded block-diagonal 2-D array, i.e., where the MSA representations of the amino acid chains in the protein form blocks on the diagonal. System 100 can initialize the embeddings at each position in the 2-D array outside the blocks on the diagonal to default embeddings, such as zero vectors. The amino acid chains in the protein can be assigned in any order, and the MSA representations of the amino acid chains in the protein can be ordered accordingly in the block-diagonal matrix. For example, the MSA representation of the first amino acid chain (i.e., according to the order) could be the first block on the diagonal, the MSA representation of the second amino acid chain could be the second block on the diagonal, and so on.

[0074] Typically, a protein's MSA representation 110 can be represented as an embedded 2D array. Throughout this specification, a "row" of a protein's MSA representation refers to the row that defines the embedded 2D array of the protein's MSA representation. Similarly, a "column" of a protein's MSA representation refers to the column that defines the embedded 2D array of the protein's MSA representation.

[0075] A set of embeddings 112 includes corresponding embeddings for each amino acid pair in protein 104. Typically, an embedding represents, i.e., encodes, information about the relationship between a pair of amino acids in a protein. A pair of amino acids refers to an ordered tuple comprising the first and second amino acids in the protein, i.e., such that the set of possible amino acid pairs in the protein is given by:

[0076] {(A i A j ):1≤i,j≤N} (1)

[0077] Where N is the number of amino acids in the protein, i,j∈{1,…,N} indexes the amino acids in the protein, and A i It is the amino acid in the protein indexed by i, and A j The amino acids in the protein are indexed by j. If the protein consists of multiple amino acid chains, the amino acids in the protein can be indexed sequentially from {1,…,N} according to the order of the amino acid chains in the protein. That is, amino acids from the first amino acid chain are indexed first, then amino acids from the second amino acid chain, then amino acids from the third amino acid chain, and so on. This set of embeddings 112 can be represented as a 2-DN×N array of embeddings, for example, where the rows of the 2-D array are indexed by i∈{1,…,N}, the columns of the 2-D array are indexed by j∈{1,…,N}, and the position (i,j) in the 2-D array is indexed by the amino acid pair (A i A j ) to the embedded occupancy.

[0078] refer to Figure 11 An example process for generating (initializing) the corresponding pair embeddings that correspond to each amino acid pair in a protein is described.

[0079] System 100 uses both MSA representation 110 and pair embeddings 112 to generate structural parameters defining the predicted protein structure 108 because they are complementary. The structure of MSA representation 110 can explicitly depend on the number of amino acid chains in the MSA corresponding to each amino acid chain in the protein. Therefore, MSA representation 110 may not be suitable for directly predicting protein structure because protein structure 108 does not have an explicit dependence on the number of amino acid chains in the MSA. Conversely, pair embeddings 112 characterize the relationship between corresponding amino acid pairs in protein 104 and are expressed without explicit reference to the MSA, thus providing a convenient and efficient data representation for predicting protein structure 108.

[0080] System 100 processes the MSA representation 110 and the pair of embeddings 112 using the values ​​of a set of parameters of the embedding neural network 200 to update the MSA representation 110 and the pair of embeddings 112. That is, the embedding neural network 200 processes the MSA representation 110 and the pair of embeddings 112 to generate an updated MSA representation 114 and an updated pair of embeddings 116.

[0081] The embedding neural network 200 updates the MSA representation 110 and the pair of embeddings 112 by sharing information between the MSA representation 110 and the pair of embeddings 112. More specifically, the embedding neural network 200 alternates between updating the current MSA representation 110 based on the current pair of embeddings 112 and updating the current pair of embeddings 112 based on the current MSA representation 110.

[0082] refer to Figure 2 A more detailed description of the example architecture of the embedded neural network 200.

[0083] System 100 generates network input for master fold neural network 600 from updated pair embedding 116, updated MSA representation 114, or both, and uses master fold neural network 600 to process the network input to generate master structure prediction 106, that is, to generate structural parameters that define the predicted protein structure.

[0084] In some implementations, the master folded neural network 600 processes the updated pair embeddings 116 to generate a distance map, which, for each amino acid pair in the protein, includes a probability distribution of a set of possible distances between amino acid pairs in the protein structure. For example, to generate a probability distribution over a set of possible distances between amino acid pairs in the protein structure, the folded neural network can apply one or more fully connected neural network layers to the updated pair embeddings 116 corresponding to that amino acid pair.

[0085] In some implementations, the master-fold neural network 600 generates structural parameters by processing the network inputs derived from both the updated MSA representation 114 and the updated pair embeddings 116 using a geometric attention operation that explicitly derives the 3D geometry of the amino acids in the protein structure. (Reference) Figure 6 Describe an example architecture of a main folded neural network 600 that implements the geometric attention mechanism.

[0086] The training engine trains the protein structure prediction system 100 end-to-end to optimize a loss function, which includes a term referred to herein as structure loss. The training engine can train the system 100 on a training dataset that includes multiple training examples. Each training example can specify: (i) training input, including the initial MSA representation of the protein and initial pair embeddings; and (ii) the target protein structure that should be generated by the system 100 by processing the training input. The target protein structure used to train the system 100 can be determined using experimental techniques such as X-ray crystallography or cryo-electron microscopy.

[0087] The structural loss can characterize the similarity between (i) the predicted protein structure generated by the master folding neural network 600 and (ii) the target protein structure that should be generated by the master folding neural network 600.

[0088] For example, if the predicted structural parameters define the predicted position and rotation parameters for each amino acid in the protein, then the structural loss can be given by the following equation:

[0089]

[0090]

[0091]

[0092] Where N is the number of amino acids in the protein, t i R represents the predicted position parameter of amino acid i. i This represents a 3×3 rotation matrix specified by the predicted rotation parameters of amino acid i. It is the target position parameter of amino acid i. This represents a 3×3 rotation matrix specified by the target rotation parameters of amino acid i, where A is a constant. Refers to the predicted rotation parameter R i The inverse of a specified 3×3 rotation matrix, Refers to the target rotation parameters The inverse of the specified 3×3 rotation matrix, and (·) + This represents the ReLU operation.

[0093] The structural loss defined by equations (2)-(4) can be understood as the loss of each amino acid pair in a protein. Take the average. t ij The term defines the predicted spatial position of amino acid j in the prediction reference frame of amino acid i, and Define the actual spatial position of amino acid j in the actual reference frame of amino acid i. These terms are sensitive to both the prediction and actual rotation of amino acids i and j, and therefore carry richer information than the loss term, which is only sensitive to the prediction and actual distance between amino acids.

[0094] As another example, if the predicted structural parameters define the predicted spatial location of each atom in each amino acid of a protein, the structural loss can be the average error (e.g., squared error) between (i) the predicted spatial location of the atom and (ii) the target (e.g., ground truth) spatial location of the atom.

[0095] Optimizing structural loss encourages the system to generate predicted protein structures that accurately approximate the real protein structure.

[0096] In addition to optimizing the structural loss, the training engine can also train System100 to optimize one or more auxiliary losses; that is, the loss function can include additional terms corresponding to each of the auxiliary losses. The auxiliary losses can, for example, penalize predicted structures with features unlikely to occur in the natural world based on the bond angles and / or bond lengths between atoms in amino acids in the predicted structure, or based on the proximity of atoms in different amino acids in the predicted structure.

[0097] The training engine can, for example, use stochastic gradient descent training techniques to train the structure prediction system 100 on the training data through multiple training iterations.

[0098] Additionally, at least during training, system 100 includes one or more auxiliary folded neural networks 620. Each auxiliary folded neural network 620 typically has the same architecture as the main folded neural network 600 and generates a corresponding auxiliary structure prediction 622, which is the same type of output as the main structure prediction 106, but based on different information.

[0099] In particular, as will be described in more detail below, each auxiliary folded neural network 620 receives input including different intermediate outputs of the embedded neural network 200, i.e., instead of the updated pair of embeddings 116 (and optionally the updated MSA representation 114), each auxiliary folded neural network receives different intermediate sets of the embeddings (and optionally corresponding intermediate MSA representations) as input.

[0100] Each auxiliary folding neural network 620 processes the input generated from the intermediate output received from the network 620 to generate an auxiliary structure prediction 622, which, like the main structure prediction 106, predicts the structure of the input protein.

[0101] To account for the inclusion of auxiliary folding neural networks (or more) 620, the training engine augments the loss function to include a corresponding additional structural loss term for each auxiliary folding neural network 620. The structural loss term for each neural network 620 characterizes (i) the similarity between the predicted protein structure generated by the auxiliary folding neural network 620 and (ii) the target protein structure that should be generated by the system. Each auxiliary structural loss term typically has the same form as the structural loss used to evaluate the main structure prediction 106 generated by the main folding neural network 600 (i.e., one of the aforementioned structural losses).

[0102] During training, the training engine backpropagates the gradient of each auxiliary structural loss term into the embedded neural network 200, such that, in addition to the main structural loss term, the parameters of those components of the embedded neural network 200 involved in generating intermediate outputs are updated based on the gradients of the auxiliary structural loss terms. In other words, the auxiliary structural predictions generated by any given auxiliary folded neural network 620 are used to update the parameter values ​​of the given auxiliary folded neural network 620 and the parameter values ​​of the components of the embedded neural network 200, and the generation of intermediate outputs by the given folded neural network 620 to generate the auxiliary structural predictions is conditioned on the parameter values ​​of the given auxiliary folded neural network 620 and the parameter values ​​of the components of the embedded neural network 200.

[0103] Optionally, each of the auxiliary structure loss terms can be assigned a lower weight in the overall loss function than the main structure loss, i.e., such that the gradient signal generated as a result of the main structure prediction is given more weight than the gradient signal generated as a result of any given auxiliary structure prediction.

[0104] In some implementations, the training engine imposes constraints on the values ​​of parameters of the auxiliary neural networks (or more) 620 during training. As an example, the training engine may constrain the values ​​of parameters of the auxiliary neural networks (or more) 620 such that the networks 620 share parameter values; that is, constraining the value of any given parameter is the same for all auxiliary neural networks 620. As another example, the training engine may constrain the values ​​of parameters of the auxiliary neural networks (or more) 620 and the main fold network 600 such that the networks 620 and 600 share parameter values; that is, constraining the value of any given parameter is the same for all auxiliary neural networks 620 and the main fold network 600.

[0105] Optionally, the training engine can also utilize additional auxiliary loss when training the neural network. Specifically, below... Figure 6 The description describes an additional auxiliary loss that can be calculated based on the intermediate structural parameters generated by the main folding network 600. The same auxiliary loss can also be applied to the intermediate structural parameters generated by each auxiliary folding network 620.

[0106] Optionally, an auxiliary neural network 620 may also be included as part of system 100 after training. In these cases, each auxiliary folding neural network 620 uses the prediction output 622 generated by the neural network 620 to update the intermediate output of the embedding neural network 200, which is received as input by the auxiliary folding neural network 620. The auxiliary neural network then provides the updated intermediate output to the embedding neural network for further processing, i.e., replacing the original intermediate output. Thus, in these cases, the predictions made by the auxiliary folding neural networks (or multiple) 620 can be used to improve the quality of the updated embedding generated by the embedding neural network 200, which is provided as input to the main folding neural network 600.

[0107] In some implementations, instead of feeding back the prediction 622 made by any given auxiliary folded neural network 620, or in addition to feeding back the prediction 622 made by any given auxiliary folded neural network 620, the neural network 620 may provide features generated by the network 620 as part of the generated prediction 622 to the embedded neural network. Specifically, the neural network 620 may select any suitable intermediate output generated by the network 620, transform that output into features, and then feed those features back into the embedded neural network.

[0108] Figure 2 An example architecture of an embedded neural network 200 is shown, which is configured to process MSA representation 110 and pair embedding 112 to generate updated MSA representation 114 and updated pair embedding 116.

[0109] The embedded neural network 200 includes updating the block sequence 202-AN. Throughout the specification, "block" refers to a portion of a neural network, such as a subnetwork of a neural network that includes one or more neural network layers.

[0110] Each update block in the embedded neural network is configured to receive a block input including an MSA representation and an embedding, and to process the block input to generate a block output including an updated MSA representation and an updated embedding.

[0111] In particular, although Figure 2The description of the following figures illustrates the architecture of the embedding neural network 200, in which each block updates both the embedding and the MSA representation. However, the techniques described for incorporating auxiliary folded neural networks into a structure prediction system can be used with any suitable architecture of the embedding neural network comprising multiple blocks, where each block updates the embedding. As a specific example of an alternative architecture to the architecture described below, the embedding neural network 200 can adopt an architecture in which the embeddings 112 are initialized using the MSA representation 110 as described above, and where each block receives only the current embeddings and updates those embeddings by applying self-attention as described below, without relying on the MSA representation.

[0112] The embedded neural network 200 provides the MSA representation 110 and the pair embeddings 112 included in the network input of the embedded neural network 200 to a first update block (i.e., in the update block sequence). The first update block processes the MSA representation 110 and the pair embeddings 112 to generate an updated MSA representation and an updated pair embedding.

[0113] For each update block following the first update block, the embedding neural network 200 provides the update block with the MSA representation and pair embedding generated by the previous update block, and for each update block except the last update block, provides the next update block with the updated MSA representation and updated pair embedding generated by that update block.

[0114] The embedded neural network 200 gradually enriches the information content of the MSA representation 110 and the pair of embeddings 112 by repeatedly updating the MSA representation 110 and the pair of embeddings 112 using the sequence of update blocks 202-AN.

[0115] The embedded neural network 200 can provide an updated MSA representation 114 and an updated pair embedding 116 generated by the final update block (i.e., in the update block sequence) as the network output.

[0116] like Figure 2 As shown, the system also includes a single auxiliary folded neural network 620 corresponding to update block 202-M (where block 202-M is any update block in the sequence of update blocks; in some embodiments, it may be the first update block 202-A, but it is generally not the last update block 202-N) and generating auxiliary structure prediction 622. The auxiliary folded neural network 620 receives at least the updated pair embeddings generated by the corresponding update block—i.e., update block 202-M—as input, and uses those updated pair embeddings to generate the auxiliary structure prediction 622. Optionally, as described above, the auxiliary folded neural network 620 may also receive the updated MSA representation generated by update block 202-M. The auxiliary folded network 620 may generate prediction 622 in any of the ways described above and below with reference to the main folded network 600.

[0117] During training, the training engine uses the structural loss computed from the auxiliary structure prediction 622 to update the parameter values ​​of the auxiliary folded neural network 620 via backpropagation, as well as the parameter values ​​of update blocks 202M and the update blocks preceding update block 202M in the update block sequence—that is, update blocks 202A-L. Including this signal as part of the gradient signal backpropagated through the system provides richer feedback to the update blocks and results in higher quality final output generated by the embedded neural network 200.

[0118] Although Figure 2 The diagram shows only a single auxiliary fold network 620, but the system may contain multiple auxiliary neural networks 620, each corresponding to a different update block within the update block. As some examples, each update block may have a corresponding auxiliary neural network 620, every other update block in the sequence may have a corresponding auxiliary neural network 620, or only update blocks following a specific position in the sequence may have a corresponding auxiliary neural network 620.

[0119] Therefore, more generally, each of the auxiliary neural networks 620 corresponds to a different one in the update blocks and receives at least the updated pair embeddings generated by the corresponding update block as input. Each auxiliary fold network 620 uses the received input to generate an auxiliary structure prediction, and the training engine uses the loss computed from the auxiliary structure prediction to update the parameter values ​​of the auxiliary neural network, the corresponding update block, and any update blocks preceding the corresponding update block in the update block sequence via backpropagation.

[0120] In some implementations, auxiliary neural networks (or multiple) 620 are not used after the system has been trained.

[0121] In other embodiments, auxiliary neural networks (or multiple networks) 620 remain a part of the system after training. In these embodiments, during and after training, i.e., at inference, each auxiliary folding network 620 uses the predictions 622 generated by network 620 to further update the pair embeddings received by network 620 from the corresponding update block as input. The auxiliary folding network 620 then provides these further updated pair embeddings as input to the update block immediately following the corresponding update block in the sequence, i.e., replacing the updated embeddings initially generated by the corresponding update block.

[0122] exist Figure 2 In the example, the auxiliary folding network 620 updates the pair embeddings received from update block 202-M and provides further updated pair embeddings as input to the next update block—that is, update block 202-N—instead of the pair embeddings generated by update block 202-M.

[0123] The auxiliary folding network 620 is able to use the auxiliary structure prediction 622 to update the pair embeddings received from the update block 202-M in any of a variety of ways.

[0124] However, typically, network 620 transforms the structure prediction 622 into an embedding with the same dimensions as the updated pair embedding, and then combines the transformed structure prediction with the updated pair embedding to generate a further updated pair embedding.

[0125] In some implementations, network 620 directly combines the transformed structure prediction and updated pair embeddings, for example, network 620 adds, averages, or cascades the transformed structure prediction and updated pair embeddings.

[0126] In some other implementations, network 620 can combine updated pair embeddings and transformed structural predictions by applying a learned combination—that is, a combination of joint learning with the training of the embedding neural network and the main folding neural network. As a specific example, network 620 can apply a learned cross-attention mechanism to incorporate information from the transformed structural predictions. Unlike the self-attention mechanism described below, in the cross-attention mechanism, the query is generated from the updated pair embeddings, while the key and value are generated from the transformed structural predictions.

[0127] Specifically, as described above, in some embodiments, the auxiliary structure prediction 622 is a distance map characterizing the corresponding estimated distances between each pair of amino acids in the protein. In these cases, the network 620 can project the distance map to the same dimension as the updated pair embeddings to generate a transformed structure prediction.

[0128] For example, the system can discretize each distance in the distance graph by assigning each distance to a warehouse in a plurality of warehouses, each warehouse corresponding to a different range of distance values. The system can then embed each discretized distance, for example, by multiplying the one-hot encoded representation of the discretized distance by an embedding matrix. The one-hot encoded representation represents the discretized distance as a vector that includes "1" at the position corresponding to the discretized distance and "0" at all other positions corresponding to all other discretized distances.

[0129] As another example, the system can apply a kernel function to each distance in the distance graph to generate a representation of the distance, and then embed the representation. For example, the kernel function could be equal to the function exp(-distance) or more generally proportional to the function exp(-distance).

[0130] As another example, as described above, in some other embodiments, the auxiliary structure prediction 622 includes structure parameters that define the positions of specified atoms in each amino acid of the protein. The network 620 can construct a distance map based on the positions defined by the structure parameters by calculating the distances between specified pairs of atoms in the amino acids of the protein. In these cases, the network 620 can project the distance map to the same dimension as the updated pair embeddings to generate the transformed structure prediction as described above.

[0131] Figure 3 An example architecture of the update block 300 embedded in the neural network 200 is shown, i.e., as referenced Figure 2 As described.

[0132] Update block 300 receives block input including current MSA representation 302 and current pair embedding 304, and processes the block input to generate updated MSA representation 306 and updated pair embedding 308.

[0133] Update block 300 includes MSA update block 400 and update block 500.

[0134] MSA update block 400 uses the current pair embedding 304 to update the current MSA representation 302, and update block 500 uses the updated MSA representation 306 (i.e., generated by MSA update block 400) to update the current pair embedding 304.

[0135] Typically, MSA representations and embeddings encode complementary information. For example, an MSA representation encodes information about the correlation between the identities of amino acids at different positions in a set of evolutionarily related amino acid chains, and an embedding encodes information about the relationships between amino acids in a protein. MSA update block 400 uses the complementary information encoded in the embeddings to enrich the information content of the MSA representation, and embedding update block 500 uses the complementary information encoded in the MSA representation to enrich the information content of the embeddings. As a result of this enrichment, the updated MSA representation and the updated embeddings encode information that is more relevant to predicting protein structure.

[0136] Update block 300 is described herein as first updating the current MSA representation 302 using the current pair embedding 304, and then updating the current pair embedding 304 using the updated MSA representation 306. This description should not be construed as limiting the update block to performing operations in this order; for example, the update block could first update the current pair embedding using the current MSA representation, and then subsequently update the current MSA representation using the updated pair embedding.

[0137] Update block 300 is described herein as including MSA update block 400 (i.e., which updates the current MSA representation) and pair update block 500 (i.e., which updates the current pair embedding). This description should not be construed as limiting update block 300 to including only one MSA update block or only a pair of update blocks. For example, update block 300 can include multiple MSA update blocks that update the MSA representation multiple times before the MSA representation is provided to the pair update block for updating the current pair embedding. As another example, update block 300 can include multiple pair update blocks that update the pair embedding multiple times using the MSA representation.

[0138] MSA update block 400 and update block 500 can have any suitable architecture that enables them to perform the functions they describe.

[0139] In some implementations, the MSA update block 400, the update block 500, or both include one or more “self-attention” blocks. As used throughout this document, a self-attention block generally refers to a neural network block that updates a set of embeddings—that is, receives a set of embeddings and outputs updated embeddings. To update a given embedding, a self-attention block is able to determine a corresponding “attention weight” between the given embedding and each of one or more selected embeddings, and then update the given embedding using (i) the attention weights and (ii) the selected embeddings. For convenience, one could say that a self-attention block updates a given embedding using attention “on” the selected embeddings.

[0140] For example, a self-attention block can receive a set of input embeddings. Where N is the number of amino acids in the protein, and in order to update the embedding x i Self-attention blocks can determine attention weights. Where a i,j x represents i With x j The attention weights between them are:

[0141]

[0142]

[0143] Among them, W q and W k This is the parameter matrix to be learned, softmax(·) denotes the soft-max normalization operation, and c is a constant. Using attention weights, the self-attention layer can embed x... i Updated to:

[0144]

[0145] Among them, W v It is the parameter matrix learned. (W) q xi It can be called an input embedding x i "Query embedding", W k x j It can be called an input embedding x i "Keyword embedding", and W v x j It can be called an input embedding x i (value embedding).

[0146] Parameter matrix W q ("Query Embedding Matrix"), W k ("Keyword Embedding Matrix") and (W) v The "value embedding matrix" is the trainable parameter of the self-attention block. The parameters of any self-attention blocks included in MSA update block 400 and update block 500 can be understood as parameters of update block 300, which can be used as a reference. Figure 1 The protein structure prediction system 100 described is trained as part of the end-to-end training. Typically, the (trained) parameters of the query, keyword, and value embedding matrices are different for different self-attention blocks, for example, such that the self-attention blocks included in MSA update block 400 can have different query, keyword, and value embedding matrices, which have different parameters than those for the self-attention blocks included in update block 500.

[0147] In some implementations, the MSA update block 400, the update block 500, or both include one or more self-attention blocks conditioned on embeddings, i.e., one or more self-attention operations that implement self-attention operations conditioned on embeddings. To make the self-attention operation conditioned on embeddings, the self-attention block is able to process the embeddings to generate a corresponding “attention bias” for each attention weight. For example, in addition to determining the attention weights according to equations (5)-(6)... In addition, self-attention blocks can also generate a corresponding set of attention biases. Among them, b i,j x represents i With x j Attention bias between them. Self-attention blocks can apply the learned parameter matrix to the embedding pair h. i,j That is, for amino acid pairs in proteins indexed by (i,j), to generate attention bias b i,j .

[0148] Self-attention blocks can determine a set of "biased attention weights" by, for example, summing attention weights and attention biases (or otherwise combining them). Where c i,j x represents i With x jAttention weights with bias between them. For example, a self-attention block can embed x... i With x j The attention weights for the bias between them are determined as follows:

[0149] c i,j =a i,j +b i,j

[0150] Among them, a i,j It is x i With x j The attention weights between them, and b i,j It is x i With x j Attention bias between inputs. Self-attention blocks can use the biased attention weights to update each input embedding x. i ,For example:

[0151]

[0152] Among them W v It is the parameter matrix for learning.

[0153] Typically, embeddings encode information characterizing the structure of a protein and the relationships between amino acid pairs within that structure. Applying a self-attention operation conditioned on embeddings to a set of input embeddings allows the input embeddings to be updated in a manner notified by the protein structural information encoded in the embeddings. Update blocks in embedding neural networks can use self-attention blocks conditioned on embeddings to update and enrich both the MSA representation and the embeddings themselves.

[0154] Optionally, the self-attention block can have multiple "heads," each head generating a corresponding updated embedding for each input embedding, i.e., such that each input embedding is associated with multiple updated embeddings. For example, each head can be a parameter matrix W described by reference equations (5)-(7). q W k and W v Different values ​​of the input embedding are used to generate updated embeddings. Self-attention blocks with multiple heads can implement a "gating" operation to combine updated embeddings generated by the heads used for the input embedding; that is, to generate a single updated embedding corresponding to each input embedding. For example, a self-attention block can use one or more neural network layers (e.g., fully connected neural network layers) to process the input embedding to generate a corresponding gating value for each head. The self-attention block can then combine the updated embeddings corresponding to the input embeddings based on the gating values. For example, a self-attention block can combine the input embedding x... i The updated embedding is generated as follows:

[0155]

[0156] Among them, k indexes the head, α k It is the gating value of the head k. It is the header k embedded in the input x i The generated updated embedding.

[0157] refer to Figure 4 This describes an example architecture for updating block 400 using MSA with self-attention blocks conditioned on embedding. (Reference) Figure 4 The example MSA update block described updates the current MSA representation based on the current pair embedding by processing the rows of the current MSA representation using a self-attention block conditioned on the current pair embedding.

[0158] refer to Figure 5 This describes an example architecture for updating block 500 using self-attention blocks conditioned on embedding. (Reference) Figure 5 The example described describes how the update block computes the outer product mean of the updated MSA representation, adds the result of the outer product mean to the current pair embedding, and processes the current pair embedding using a self-attention block conditioned on the current pair embedding, updating the current pair embedding based on the updated MSA representation.

[0159] like Figure 3 As shown, update block 300 has a corresponding auxiliary folding network 620. Auxiliary folding network 620 receives the pair embeddings 308 of the update generated from update block 500 (and optionally the MSA representation 306 of the update generated from MSA update 400) as input, and generates auxiliary structure prediction 622 from the received input as described above. The training engine then uses the auxiliary structure prediction 622 to update the parameter values ​​of the auxiliary folding network 600, the components of update block 300, and the components of any block in the update block sequence preceding update block 300.

[0160] In an implementation where the auxiliary folding network 620 is also used during training, the auxiliary folding network 620 also uses the auxiliary structure prediction 620 to further update the updated pair embedding 308 before providing the pair embedding to the next block, i.e., such that the pair embedding provided to the next block in the sequence is a further updated pair embedding, which is a combination of the updated pair embedding 308 as described above and the transformed structure prediction.

[0161] Figure 4 An example architecture of MSA update block 400 is shown. MSA update block 400 is configured to receive the current MSA representation 302 to update the current MSA representation 306 (at least in part) based on the current pair embedding.

[0162] To update the current MSA representation 302, the MSA update block 400 uses a self-attention operation conditioned on the current pair of embeddings (i.e., a "line-by-line" self-attention operation) to update the embeddings in each line of the current MSA representation. More specifically, the MSA update block 400 provides the embeddings in each line of the current MSA representation 302 to a "line-by-line" self-attention block 402 conditioned on the current pair of embeddings, for example, as referenced. Figure 3 The method described above generates updated embeddings for each row of the current MSA representation 302. Optionally, the MSA update block can add the input of the line-by-line self-attention block 402 to the output of the line-by-line self-attention block 402. By conditioned the line-by-line self-attention block 402 on the current pair of embeddings, the MSA update block 400 can use information from the current pair of embeddings to enrich the current MSA representation 302.

[0163] Then, the MSA update block updates the embeddings in each column of the current MSA representation using a self-attention operation that is not conditioned on the current pair of embeddings (i.e., a “column-by-column” self-attention operation). More specifically, MSA update block 400 provides the embeddings in each column of the current MSA representation 302 to a “column-by-column” self-attention block 404 that is not conditioned on the current pair of embeddings to generate updated embeddings for each column of the current MSA representation 302. As a result of not being conditioned on the current pair of embeddings, column-by-column self-attention block 404 generates updated embeddings for each column of the current MSA representation using attention weights (e.g., as described in reference equations (5)-(6)) instead of biased attention weights (e.g., as described in reference equation (8)). Optionally, the MSA update block is able to add the input of column-by-column self-attention block 404 to the output of column-by-column self-attention block 404.

[0164] The MSA update block then processes the current MSA representation 302 using, for example, a transformation block that applies one or more fully connected neural network layers to the current MSA representation 302. Optionally, the MSA update block 400 can add the input of the transformation block 406 to the output of the transformation block 406.

[0165] The MSA update block can output an updated MSA representation 306 resulting from the operations performed by the row-by-row self-attention block 402, the column-by-column self-attention block 404, and the transformation block 406.

[0166] Figure 5 An example architecture for update block 500 is shown. Update block 500 is configured to receive the current pair embedding 304 and (at least in part) update the current pair embedding 304 based on the updated MSA representation 306.

[0167] To update the current pair embedding 304, the update block 500 applies the outer product mean operation 502 to the updated MSA representation 306 and adds the result of the outer product mean operation 502 to the current pair embedding 304.

[0168] The outer product mean operation defines a series of operations that, when applied to an MSA representation of an embedded M×N array, produce an embedded N×N array, i.e., where N is the number of amino acids in the protein. The current embedding 304 can also be represented as an embedded N×N array, and adding the result of the outer product mean 502 to the current embedding 304 means summing the two embedded N×N arrays.

[0169] To compute the mean of the outer product, a tensor A(·) is generated for the update block, for example, given by the following equation:

[0170]

[0171] Where res1, res2 ∈ {1,…,N}, ch1, ch2 ∈ {1,…,C}, where C is the number of channels in each embedding of the MSA representation, |rows| is the number of rows in the MSA representation, LeftAct(row,res1,ch1) is a linear operation (e.g., defined by matrix multiplication) applied to the channel ch1 of the embedding of the MSA representation located at the row indexed by “row” and the column indexed by “res1”, and RightAct(row,res2,ch2) is a linear operation (e.g., defined by matrix multiplication) applied to the channel ch2 of the embedding of the MSA representation located at the row indexed by “row” and the column indexed by “res2”. The result of the outer product mean is generated by flattening and linearly projecting the (ch1,ch2) dimension of tensor A. Optionally, one or more layer normalization operations (e.g., as described in “Layer Normalization” by Jimmy Lei Ba et al., arXiv: 1607.06450) can be performed on the updated block as part of the calculation of the outer product mean.

[0172] Typically, the updated MSA representation 306 encodes information about the correlation between the identities of amino acids at different positions in a set of evolutionarily related amino acid chains. The information encoded in the updated MSA representation 306 is related to the predicted structure of the protein, and the updated block 500 can enhance the information content of the current pair embedding by incorporating the information encoded in the updated MSA representation into the current pair embedding (i.e., by means of the outer product mean 502).

[0173] After updating the current pair embeddings 304 using the updated MSA representation (i.e., via the outer product mean 502), the update block 500 is updated with an N×N array of current pair embeddings in each row of the current pair embeddings arrangement using a self-attention operation conditioned on the current pair embeddings (i.e., a "row-by-row" self-attention operation). More specifically, the update block 500 provides each row of the current pair embeddings to a "row-by-row" self-attention block 504, also conditioned on the current pair embeddings, for example, as referenced. Figure 3 The method is described above, to generate a pair of embeddings for updating each row. Optionally, the pair of update blocks can add the input of the line-by-line self-attention block 504 to the output of the line-by-line self-attention block 504.

[0174] Then, the update block 500 is updated with a self-attention operation (i.e., a "column-by-column" self-attention operation) that is also conditioned on the current pair embeddings, to update the current pair embeddings in each column of the N×N array of current pair embeddings. More specifically, the update block 500 provides each column of the current pair embeddings to a "column-by-column" self-attention block 506, which is also conditioned on the current pair embeddings, to generate updated pair embeddings for each column. Optionally, the update block can add the input of the column-by-column self-attention block 506 to the output of the column-by-column self-attention block 506.

[0175] Then, the current pair of embeddings is processed by applying, for example, one or more fully connected neural network layers to the transformation block of the current pair of embeddings. Optionally, the update block 500 can add the input of the transformation block 508 to the output of the transformation block 508.

[0176] The update block can output the updated pair embedding 308 resulting from the operations performed by the row-by-row self-attention block 504, the column-by-column self-attention block 506, and the transformation block 508.

[0177] Figure 6 An example architecture of a master folded neural network 600 is shown, which generates a set of structural parameters 106 that define a predicted protein structure 108. For example, in one implementation, the folded neural network 600 determines the predicted protein structure based on the pair embeddings by processing inputs including corresponding pair embeddings 116 for each amino acid pair in the protein to generate values ​​for the structural parameters 106. The folded neural network 600 may be included in a reference... Figure 1 The protein structure prediction system 100 is described. (Although reference is made to the main folded neural network 600 description.) Figure 6 The architecture shown is different, but the auxiliary folded neural networks (or multiple ones) 620 may each have the same architecture, the difference being that each auxiliary folded neural network 620 will receive input from the corresponding update block that is not the last update block in the sequence of update blocks embedded in the neural network.

[0178] In this implementation, the folded neural network 600 generates structural parameters for each amino acid in the protein, which may include: (i) position parameters and (ii) rotation parameters. As previously described, the position parameters of an amino acid can specify the predicted 3-D spatial position of a specified atom within the amino acid in the protein structure. The rotation parameters of an amino acid can specify the predicted “orientation” of the amino acid in the protein structure. More specifically, the rotation parameters can specify a 3-D spatial rotation operation that, when applied to the coordinate system of the position parameters, causes the three “main chain” atoms in the amino acid to take fixed positions relative to the rotation coordinate system.

[0179] In an implementation, the folded neural network 600 generates structural parameters for each amino acid in the protein. These structural parameters can include: (i) position parameters and (ii) rotation parameters. As previously described, the position parameters of an amino acid can specify the predicted 3-D spatial position of a specified atom within the amino acid in the protein's structure. The rotation parameters of an amino acid can specify the predicted "orientation" of the amino acid within the protein's structure. More specifically, the rotation parameters can specify a 3-D spatial rotation operation that, when applied to the coordinate system of the position parameters, causes the three "main chain" atoms in the amino acid to take fixed positions relative to the rotation coordinate system.

[0180] In one implementation, the folded neural network 600 receives input derived from the final MSA representation, the final pair embeddings, or both, and generates final values ​​for structure parameters 106 that define the predicted structure of the protein. For example, the folded neural network 600 may receive input including: (i) the corresponding pair embeddings 116 for each amino acid pair in the protein, (ii) initial values ​​for "single" embeddings 602 for each amino acid in the protein, and (iii) initial values ​​for structure parameters 604 for each amino acid in the protein. The folded neural network 600 processes the input to generate the final values ​​for structure parameters 106 that collectively characterize the predicted structure 108 of the protein.

[0181] The protein structure prediction system 100 can provide the folded neural network 600 with embedded pairs generated as the output of the embedded neural network, as shown in the reference. Figure 1 As stated above.

[0182] The protein structure prediction system 100 is capable of generating initial single embeddings 602 of amino acids from the MSA representation 114, i.e., they are generated as the output of an embedding neural network, as shown in the reference. Figure 1As described above, the MSA representation 114 can be represented as a 2-D array of embeddings, the number of columns of which equals the number of amino acids in the protein, where each column is associated with a corresponding amino acid in the protein. The protein structure prediction system 100 is able to generate an initial single embedding for each amino acid in the protein by summing (or otherwise combining) the embeddings from the columns of the MSA representation 114 associated with the amino acids. As another example, the protein structure prediction system 100 is able to generate an initial single embedding for an amino acid in the protein by extracting embeddings from the rows of the MSA representation 114, the rows corresponding to the amino acid sequence of the protein whose structure is being estimated.

[0183] The protein structure prediction system 100 can generate initial structure parameters 604 with default values, for example, where the position parameter of each amino acid is initialized to the origin (e.g., [0,0,0] in a Cartesian coordinate system), and the rotation parameter of each amino acid is initialized to a 3×3 identity matrix.

[0184] The folded neural network 600 generates the final structure parameter 106 by repeatedly updating the current values ​​of a single embedding 606 and the structure parameter 608 (i.e., starting from their initial values). More specifically, the folded neural network 600 includes a sequence of updating neural network blocks 610, wherein each updating block 610 is configured to update the current single embedding 606 (i.e., generate an updated single embedding 612) and update the current structure parameter 608 (i.e., generate an updated structure parameter 614). In addition to the updating blocks, the folded neural network 600 may also include other neural network layers or blocks, for example, other neural network layers or blocks that may be interleaved with the updating blocks.

[0185] Each update block 610 can include: (i) a geometric attention block 616, and (ii) a folding block 618, each of which will be described in more detail below.

[0186] Geometric attention block 616 uses a “geometric” self-attention operation to update the current single embedding, which explicitly infers the 3-D geometry of the amino acids in the protein structure (i.e., as defined by the structure parameters). More specifically, to update a given single embedding, geometric attention block 616 determines a corresponding attention weight between the given single embedding and each of one or more selected single embeddings, where the attention weight depends on the current single embedding, the current structure parameters, and the pair of embeddings. Geometric attention block 616 then updates the given single embedding using (i) the attention weight, (ii) the selected single embedding, and (iii) the current structure parameters.

[0187] To determine attention weights, geometric attention block 616 processes each current individual embedding to generate corresponding "symbol query" embeddings, "symbol keyword" embeddings, and "symbol value" embeddings. For example, geometric attention block 616 can target a single embedding h corresponding to the i-th amino acid. i Generate symbol query embedding q i Symbolic keyword embedding k i and symbolic value embedding v i ,as follows:

[0188] q i =Linear(h i (10)

[0189] k i =Linear(h i (11)

[0190] v i =Linear(h i (12)

[0191] Here, Linear(·) refers to a linear layer with independent learning parameter values.

[0192] The geometric attention block 616 additionally processes each current individual embedding to generate a corresponding "geometric query" embedding, "geometric keyword" embedding, and "geometric value" embedding. The geometric query embedding, geometric keyword embedding, and geometric value embedding for each individual embedding are each initially generated as a 3-D point in the local reference frame of the corresponding amino acid, then rotated and translated to the global reference frame using the amino acid's structural parameters. For example, the geometric attention block 616 can target a single embedding h corresponding to the i-th amino acid... i Generate geometric query embeddings Geometric Keyword Embedding and geometric value embedding as follows:

[0193]

[0194]

[0195]

[0196] Among them, Linear p (·) refers to a linear layer with independent learning parameter values, which will h i Projected onto 3-D points (superscript p indicates the quantity is 3-D points), R i Let represent the rotation matrix specified by the rotation parameter of the i-th amino acid, and t i This represents the position parameter of the i-th amino acid.

[0197] In order to update the single embedding h corresponding to amino acid i i Geometric attention block 616 can generate attention weights. Where N is the total number of amino acids in the protein, and a j The attention weights between amino acid i and amino acid j are as follows:

[0198]

[0199] Where, q i The symbolic query embedding of amino acid i, k j The symbolic keyword embedding represents amino acid j, and m represents q. i and k j The dimension, α represents the learning parameters. The geometric query embedding of amino acid i. Let |·|2 represent the geometric key embedding of amino acid j, where |·|2 is the L2 norm, and b i,j It is the pair embedding 116 corresponding to the amino acid pair including amino acid i and amino acid j, and w is the learned weight vector (or some other learned projection operation).

[0200] Typically, the pair embeddings of amino acid pairs implicitly encode information related to the relationship between the amino acids in the pair, such as the distance between the amino acids in the pair. By determining the attention weights between amino acids i and j in part based on the pair embeddings of amino acids i and j, the folded neural network 600 utilizes information from the pair embeddings to enrich the attention weights, thereby improving the accuracy of predicting folded structures.

[0201] In some implementations, the geometric attention block 616 generates multiple sets of geometric query embeddings, geometric keyword embeddings, and geometric value embeddings, and uses each set of generated geometric embeddings when determining attention weights.

[0202] In generating a single embedding h corresponding to amino acid i i After the attention weights, the geometric attention block 616 uses the attention weights to update the individual embedding h. i Specifically, geometric attention block 616 uses attention weights to generate "symbolic return" embeddings and "geometric return" embeddings, and then uses the symbolic return embeddings and geometric return embeddings to update individual embeddings. Geometric attention block 124 can use the symbolic return embedding o of amino acid i. i Generate as, for example:

[0203]

[0204] in, This indicates attention weights (e.g., as defined in equation (16)), and each vj This represents the symbolic value embedding of amino acid j. Geometric attention block 616 can return the geometry of amino acid i as an embedding. Generate as, for example:

[0205]

[0206] Among them, geometric return embedding It is a 3D point. This indicates attention weight (e.g., as defined in equation (16)). It is the inverse of the rotation matrix specified by the rotation parameter of amino acid i, and t i This is the position parameter of amino acid i. It's understandable that the geometry-returned embedding is initially generated in the global reference frame, then rotated and translated to the local reference frame of the corresponding amino acid.

[0207] Geometric attention block 616 can use the corresponding symbol to return the embedded o i (e.g., generated according to equation (17)) and geometry return embedding (For example, according to equation (18)) the single embedding of amino acids into h i Updated to, for example:

[0208]

[0209] in, It is the single embedding of the updated amino acid i, |·| is the norm, for example, the L2 norm, and LayerNorm(·) denotes the layer normalization operation, for example, as described with reference to: JLBa, JRKiros, GEHinton, “Layer Normalization” arXiv:1607.06450 (2016).

[0210] Individual embeddings 606 of amino acids are updated using specific 3-D geometric embeddings, for example, as described in reference equations (13)-(15), such that the geometric attention block 616 is able to infer the 3-D geometry when updating individual embeddings. Furthermore, each update block updates individual embeddings and structural parameters in a manner invariant to rotation and translation over the entire protein structure. For example, applying the same global rotation and translation operations to the initial structural parameters provided to the folded neural network 600 will cause the folded neural network 600 to generate a predicted structure that is globally rotated and translated in the same way, but otherwise identical. Therefore, the global rotation and translation operations applied to the initial structural parameters do not affect the accuracy of the predicted protein structure generated by the folded neural network 600 from the initial structural parameters. The rotation and translation invariance of the representation generated by the folded neural network 600 aids training, for example, because the folded neural network 600 automatically learns to generalize over all rotations and translations of the protein structure.

[0211] The updated single embedding of an amino acid can be further transformed by one or more additional neural network layers (e.g., linear neural network layers) in the geometric attention block 616 before being provided to the folded block 618.

[0212] After geometric attention block 616 updates the current single embedding 606 of amino acid, fold block 618 uses the updated single embedding 612 to update the current structural parameter 608. For example, fold block 618 can update the current position parameter t of amino acid i. i Updated to:

[0213]

[0214] in, These are the updated position parameters; Linear(·) represents a linear neural network layer, and... This represents a single embedding of the update for amino acid i. In another example, the rotation parameter R of amino acid i... i A rotation matrix can be specified, and fold block 618 can display the current rotation parameter R. i Updated to:

[0215]

[0216]

[0217] Among them, w i It is a three-dimensional vector, and Linear(·) is a linear neural network layer. It is a single embedding of the updated amino acid i, 1+w i It indicates that it has a real part 1 and an imaginary part w. iThe quaternion is defined by the operation of transforming the quaternion into an equivalent 3×3 rotation matrix. The rotation parameters are updated using equations (21)-(22) to ensure that the updated rotation parameters define a valid rotation matrix, such as an orthogonal matrix with a determinant of 1.

[0218] The folded neural network 600 can provide updated structural parameters generated by the final update block 610 as final structural parameters 106 defining the predicted protein structure 108. The folded neural network 600 can include any suitable number of update blocks, such as 5, 25, or 125 update blocks. Optionally, each update block of the folded neural network can share a single set of parameter values ​​that are jointly updated during the training of the folded neural network. Sharing parameter values ​​among update blocks 610 reduces the number of trainable parameters of the folded neural network and can therefore facilitate efficient training of the folded neural network, for example, by stabilizing training and reducing the likelihood of overfitting.

[0219] During training, the training engine can train the parameters of the structure prediction system, including the parameters of the folded neural network 600, based on the structural loss that evaluates the accuracy of the final structural parameters 106, as described above. In some embodiments, the training engine can further evaluate one or more auxiliary structural losses in update blocks 610 prior to the final update block (i.e., generating the final structural parameters). The auxiliary structural loss of the update block evaluates the accuracy of the updated structural parameters generated by the update block.

[0220] Optionally, during training, the training engine can apply "stop gradient" operations to prevent gradient backpropagation through certain neural network parameters of each update block, such as the neural network parameters used to compute the updated rotation parameters (as described in equations (21)-(22)). Applying these stop gradient operations can improve the numerical stability of gradients computed during training.

[0221] Typically, the similarity between the predicted protein structure 108 generated by the folded neural network 600 and the corresponding benchmark protein structure can be measured, for example, by a similarity measure that assigns a corresponding accuracy score to each of the multiple atoms in the predicted protein structure. For example, the similarity measure can assign a corresponding accuracy score to each carbon alpha atom in the predicted protein structure. The accuracy score of the atoms in the predicted protein structure can characterize the degree of agreement between the positions of the atoms in the predicted protein structure and the actual positions of the atoms in the benchmark protein structure. An example of a similarity measure that can compare the predicted protein structure with the benchmark protein structure to generate an accuracy score for the atoms in the predicted protein structure is the lDDT similarity measure described in the following literature: V. Mariani et al., “lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests”, Bioinformatics, 1 Nov 2013; 29(21)2722-2728.

[0222] The folded neural network 600 can be configured to generate a corresponding confidence estimate 650 for each of one or more atoms in the predicted protein structure 108. The confidence estimate 650 for the atoms in the predicted protein structure characterizes the prediction accuracy score (e.g., lDDT accuracy score) of the atoms in the predicted protein structure; that is, it will be generated by a similarity measure comparing the predicted protein structure to a (potentially unknown) benchmark ground truth protein structure. In one example, the confidence estimate 650 for the atoms in the predicted protein structure can be defined as a discrete probability distribution over a set of intervals that form partitions of the possible range of accuracy scores for the atoms. The discrete probability distribution can associate a corresponding probability with each of the intervals, defining the likelihood that the actual accuracy score is included in the interval. For example, the possible range of accuracy scores could be [0, 100], and the confidence estimate 650 could define a probability distribution over a set of intervals {[0, 2), [2, 4), ..., [98, 100]}. In another example, the confidence estimate of 650 for predicting atoms in a protein structure can be a numerical value, i.e., the accuracy score for directly predicting atoms.

[0223] In some implementations, the folded neural network 600 generates a corresponding confidence estimate 650 for a specified atom (e.g., an alpha carbon atom) in each amino acid of the protein. The folded neural network 600 is capable of generating confidence estimates 650 for specified atoms in amino acids of the protein, for example, by processing a single embedding of an updated amino acid generated by the last update block in the folded neural network using one or more neural network layers (e.g., fully connected layers).

[0224] The structure prediction system can generate a corresponding confidence score for each amino acid in a protein based on a confidence estimate of 650 for the atoms in the predicted protein structure. For example, the structure prediction system can generate a confidence score for an amino acid as an expected value of the probability distribution over the possible values ​​of the accuracy score for the alpha carbon atom in the amino acid.

[0225] The structure prediction system can generate a confidence score for the entire predicted structure, for example, as the average confidence score of amino acids in a protein.

[0226] During training of the structure prediction system, the training engine can adjust the parameter values ​​of the structure prediction system by backpropagating the gradient of an auxiliary loss that measures the error between (i) the confidence estimate generated by the folded neural network 600, and (ii) the accuracy score generated by comparing the predicted protein structure with the benchmark true protein structure. The error can be, for example, cross-entropy error.

[0227] Confidence estimates generated by structure prediction systems can be used in various ways. For example, confidence estimates of atoms in a protein structure can indicate which parts of the structure have been reliably estimated and are therefore suitable for further downstream processing or analysis. As another example, per-protein confidence scores can be used to rank a set of predictions of a protein's structure, such as predictions generated by the same structure prediction system by processing different inputs characterizing the same protein, or predictions generated by different structure prediction systems.

[0228] The position and rotation parameters specified by structural parameter 106 define the spatial positions (e.g., in [x,y,z] Cartesian coordinates) of the main chain atoms in the amino acids of a protein. However, structural parameter 106 does not necessarily define the spatial positions of the remaining atoms in the amino acids of a protein (e.g., atoms in the side chains of amino acids). In particular, the spatial positions of the remaining atoms in an amino acid depend on the value of the twist angle between bonds in the amino acid, e.g., ω angle. Angles, psi angles, chi1 angles, chi2 angles, chi3 angles, and chi4 angles, as shown in the reference. Figure 7 As shown.

[0229] Optionally, one or more of the update blocks 610 of the folded neural network 600 can generate outputs that define the corresponding predicted spatial location of each atom in each amino acid of the protein. To generate the predicted spatial location of the atom in the amino acid, the update block can use one or more neural network layers to process a single embedding of the updated amino acid to generate predicted values ​​of the twist angles of the bonds between the atoms in the amino acid. The neural network layers can be, for example, fully connected neural network layers with residual connections embedded. Each twist angle can be represented, for example, as a 2-D vector.

[0230] The update block can determine the spatial positions of atoms in an amino acid based on (i) the value of the amino acid's twist angle and (ii) the updated structural parameters of the amino acid (e.g., position and rotation parameters). For example, the update block can process the twist angle according to a predefined function to generate the spatial positions of atoms in the amino acid within the amino acid's local reference frame. The update block can generate the spatial positions of atoms in the amino acid within the global reference frame (i.e., the reference frame common to all amino acids in a protein) by rotating and translating the spatial positions of atoms according to the updated structural parameters of the amino acid. For example, the update block can determine the spatial positions of atoms in the global reference frame by applying a rotation operation defined by the updated rotation parameters to the spatial positions of atoms in the local reference frame to generate rotated spatial positions, and then applying a translation operation defined by the updated position parameters to the rotated spatial positions.

[0231] In some implementations, instead of outputting the final structural parameters or in combination with them, the folded neural network 600 outputs the predicted spatial locations of atoms in the amino acids of the protein generated by the final update block.

[0232] refer to Figure 6 The folded neural network 600 described herein is characterized as receiving input based on MSA representation 114 and embeddings 116 generated by an embedding neural network, for example, as referenced Figure 2 However, generally, any suitable technique can be used to generate the inputs to a folded neural network (e.g., a single embedding 602 and a pair of embeddings 116). Furthermore, aspects of the operations performed by a folded neural network (e.g., predicting the spatial location of atoms in each amino acid of a protein) can be performed by other folded neural networks, for example, with different architectures that receive different inputs.

[0233] Figure 7 Showing the twist angle between bonds in amino acids, such as the ω angle, Angle, psi angle, chi1 angle, chi2 angle, chi3 angle, chi4 angle and chi5 angle.

[0234] Figure 8This is an illustration of unfolded and folded proteins. An unfolded protein is a random coil of amino acids. An unfolded protein undergoes protein folding and folds into a 3D conformation. Protein structures typically include stable local folding patterns, such as alpha helices (e.g., as shown in 802) and beta folds.

[0235] Figure 9 This is a flowchart of an example process 900 for training a structure prediction neural network including embedded neural networks and master fold neural networks. For convenience, process 900 will be described as being performed by a system of one or more computers located in one or more locations. For example, a protein structure prediction system appropriately programmed according to this specification (e.g., Figure 1 The protein structure prediction system 100 can execute process 900.

[0236] The system receives training input including data characterizing a given protein (step 902). For example, the data may include: (i) initial multiple sequence alignment (MSA) representations, which represent the corresponding MSA for each strand in the protein; and (ii) the corresponding initial pair embeddings for each amino acid pair in the protein.

[0237] The system obtains data on the target protein structure that should be generated by the system 100 by processing the training input (step 904).

[0238] The system uses embedded neural networks and master fold neural networks to process the training input, for example, as described above, to generate master structure predictions (step 906).

[0239] For each of one or more auxiliary folded neural networks, the system processes the input to the embedded network, including an update corresponding to one of the generated updates from the "hidden" update blocks in the embedded neural network, to generate an auxiliary structure prediction (step 908). The hidden update block is an update block that is not the last update block in the sequence. In embodiments where the auxiliary structure prediction is fed back into the embedded neural network, the system performs step 908 concurrently with step 906, i.e., because the output of the embedded neural network depends on the auxiliary structure predictions (or more) generated by the auxiliary folded neural networks. In embodiments where the auxiliary structure prediction is not fed back into the embedded neural network, the system may perform step 908 concurrently with step 906, or after step 906, i.e., after generating the main structure prediction.

[0240] The system calculates the gradient of each parameter of the neural network with respect to the loss function (step 910), the loss function comprising at least: (a) a main structure loss characterizing the similarity between (i) the predicted protein structure defined by the main structure prediction generated by the main folded neural network and (ii) the target protein structure that should be generated by the system; and (b) the corresponding auxiliary structure loss for each auxiliary neural network (or more). The auxiliary structure loss of a given auxiliary neural network characterizes the similarity between (i) the predicted protein structure defined by the auxiliary structure prediction generated by the given auxiliary folded neural network and (ii) the target protein structure that should be generated by the system.

[0241] Once a threshold condition is met, for example, when the gradients of the entire batch of training inputs have been computed, the system can compute updates to the parameter values ​​of the embedded neural network, the main folded neural network, and the auxiliary folded neural network from the gradients according to the update rules of the optimizer used for training (e.g., Adam, rmsProp, or SGD). The system can then apply the updates to the current values ​​of the parameters, i.e., subtract the updates from the current values ​​or add the updates to the current values.

[0242] After training, the system can perform process 900 without executing steps 904 and 910, and in an implementation in which the output from the auxiliary folding network is not fed back into the embedded neural network, step 908, then provides the main structure prediction or other information specifying the structure defined by the main structure prediction as the output of the system characterizing the protein by the input.

[0243] Figure 10 An example procedure 1000 is shown for generating the MSA representation 1010 of the amino acid chain in a protein. Reference Figure 1 The protein structure prediction system 100 described is capable of performing the operation of process 1000.

[0244] In order to generate an MSA representation 1010 of the amino acid chain in a protein, system 100 obtains an MSA 1002 of the protein, which may include, for example, thousands of MSA sequences.

[0245] System 100 divides the set of MSA sequences into a "core" MSA sequence group 1004 and an "extra" MSA sequence group 1006. The core MSA sequence group can be smaller than the extra MSA sequence group 1006 (e.g., an order of magnitude smaller). System 100 can divide the set of MSA sequences into core MSA sequences 1004 and extra MSA sequences 1006, for example, by randomly selecting a predetermined number of MSA sequences as core MSA sequences and identifying the remaining MSA sequences as extra MSA sequences 1006.

[0246] For each additional MSA sequence 1006, system 100 is able to determine a corresponding similarity measure (e.g., based on Hamming distance) between the additional MSA sequence and each core MSA sequence 1004. System 100 then associates each additional MSA sequence 1006 with its corresponding core MSA sequence 1004 that is most similar to the additional MSA sequence 1006 (i.e., according to the similarity measure). The group of additional MSA sequences 1006 associated with the core MSA sequence 1004 can be referred to as an "MSA sequence cluster" 1008. That is, system 100 determines a corresponding MSA sequence cluster 1008 for each core MSA sequence 1004, wherein the MSA sequence cluster 1008 corresponding to the core MSA sequence 1004 includes a group of additional MSA sequences 1006 that are most similar to the core MSA sequence 1004.

[0247] System 100 is capable of generating an MSA representation of the amino acid chain in a protein based on a core MSA sequence and an MSA sequence cluster 1008. The MSA representation 1010 can be represented by an embedded M×N array, where M is the number of core MSA sequences (i.e., such that each core MSA sequence is associated with a corresponding row of the MSA representation), and N is the number of amino acids in the amino acid chain. The embeddings in the MSA representation can be indexed by (i,j)∈{(i,j): i=1,…,M,j=1,…,N}.

[0248] To generate the embedding at position (i,j) in the MSA representation 1010, system 100 is able to obtain the embedding (e.g., a one-hot embedding) of the identity of the amino acid at position j in the core MSA sequence i. System 100 is also able to determine a probability distribution of a set of possible amino acids based on the relative frequency of each possible amino acid at position j in the additional MSA sequence 1006 in the MSA sequence cluster 1008 corresponding to the core MSA sequence i. System 100 is then able to determine the embedding at position (i,j) in the MSA representation by combining (e.g., cascading) the following: (i) the embedding of the identity of the amino acid at position j in the core MSA sequence i, and (ii) the probability distribution of the possible amino acid corresponding to position j in the core MSA sequence.

[0249] In some cases, for one or more core MSA sequences, the (benchmark true) protein structure can be known. Specifically, for one or more core MSA sequences, the torsion angles (e.g., ω angles) between bonds in the amino acids of the core MSA sequence can be known. The values ​​of the torsion angle (e.g., psi angle, etc.) can be known. If the value of the torsion angle of the amino acid in the core MSA sequence i is known, then the system 100 can generate an embedding at position (i,j) in the MSA representation based at least in part on the value of the torsion angle of the amino acid j in the core MSA sequence i. For example, the system can use one or more neural network layers to generate embeddings of the torsion angle values ​​and concatenate the embeddings of the torsion angle values ​​to the embeddings at position (i,j) in the MSA representation.

[0250] Figure 11 An example process 1100 is shown for generating (initializing) the corresponding pair embedding 112 for each amino acid pair in a protein. (See reference) Figure 1 The protein structure prediction system 100 described is capable of performing the operation of process 1100.

[0251] System 100 is capable of generating pair embeddings using the MSA representation of proteins 1102. (Reference) Figure 1 and Figure 10 The MSA representation of the generated protein is described in more detail. System 100 is capable of generating an MSA representation 1102 of the protein based on the corresponding MSA representation of each amino acid chain in the protein. To generate an MSA representation of the amino acid chains in the protein, system 100 can use a reference [database name missing] to generate the MSA representation. Figure 1 The described MSA representation 110 contains more (e.g., an order of magnitude more) MSA sequences. Therefore, the MSA representation 1102 used by system 100 to generate the embedding 112 can have a greater number of MSA sequences than the reference representation. Figure 1 The described MSA represents 110 or more rows (e.g., an order of magnitude more). In some implementations, system 100 is able to use references Figure 10 The additional MSA sequence 1006 described is used to generate the MSA representation 1102.

[0252] After generating the MSA representation 1102, the system 100 processes the MSA representation 1102 to generate a pair of embeddings 1104 from the MSA representation 1102, for example, by applying an outer product mean operation to the MSA representation 1102 and recognizing the pair of embeddings 1104 as the result of the outer product mean operation.

[0253] System 100 uses an embedded neural network 1106 to process the MSA representation 1102 and the pair embedding 1104. The embedded neural network 1106 is able to update the MSA representation 1102 and the pair embedding 1104 by sharing information between the MSA representation 1102 and the pair embedding 1104. More specifically, the embedded neural network 1106 is able to alternate between updating the MSA representation 1102 based on the pair embedding 1104 and updating the pair embedding 1104 based on the MSA representation 1102.

[0254] Embedded neural network 1106 can have reference-based Figure 2-5 The architecture described is an embedded neural network architecture that uses row-wise and column-wise self-attention blocks to update the embedding 1104 and the MSA representation 1102.

[0255] In some implementations, the embedding neural network 1106 is capable of updating the embeddings in each column of the MSA representation 1102 using a column-wise "global" self-attention operation. More specifically, the embedding neural network 1106 is capable of providing the embeddings in each column of the MSA representation to a column-wise global self-attention block to generate updated embeddings for each column of the current MSA representation. To implement global column-wise self-attention, the self-attention block is capable of generating a corresponding query embedding for each embedding in the column, and then averaging the query embeddings to generate a single "global" query embedding for the column. The column-wise self-attention block then performs the self-attention operation using the single global query embedding, which reduces the complexity of the self-attention operation from quadratic (i.e., the number of embeddings per column) to linear. Using a global self-attention operation reduces the computational complexity of the column-wise self-attention operation, making it possible to perform column-wise self-attention operations on columns of the MSA representation 1102 that include a large number (e.g., thousands) of embeddings.

[0256] After updating the pair embedding 1104 and MSA representation 1102 using the embedding neural network 1106, the system 100 is able to recognize the pair embedding 112 as the updated pair embedding generated by the embedding neural network 1106. The system 100 is able to discard the updated MSA representation generated by the embedding neural network 1106, or use it in any appropriate manner.

[0257] As part of generating pair embeddings 112, system 100 is capable of including relative position encoding information in the corresponding pair embedding for each amino acid pair in a protein. The system can include the relative position encoding information of amino acid pairs included in the same amino acid chain within the corresponding pair embedding by: calculating a signed difference representing the number of amino acids in the separated amino acid pairs; clipping the result to a predefined interval; representing the clipping value using a one-hot encoding vector; applying a linear transformation to the one-hot encoding vector; and adding the result of the linear transformation to the corresponding pair embedding. The system can also include the relative position encoding information of amino acid pairs in the same amino acid chain not included in the corresponding pair embedding by adding a default encoding vector to the corresponding pair embedding, the default encoding vector indicating that the amino acid pair is not included in the same amino acid chain.

[0258] System 100 is also capable of generating pairs of embeddings 112 based at least in part on a set of one or more template sequences 1110. Each template sequence 1110 is an MSA sequence of an amino acid chain in a protein, wherein the folding structure of the template sequence 1110 is known, for example, from physical experiments.

[0259] System 100 is capable of generating a corresponding template representation 1112 for each template sequence 1110. The template representation 1112 of template sequence 1110 includes a corresponding embedding corresponding to each amino acid pair in the template sequence, for example, such that the template representation 1112 of a template sequence of length n (i.e., having n amino acids) can be represented as an n×n array of embeddings. System 100 generates the embedding at position (i,j) in the template representation 1112 of template sequence 1110 based on, for example, the following: (i) a corresponding embedding (e.g., a uniquely heated embedding) representing the identity of the amino acids at positions i and j in the template sequence, (ii) a unit vector defined by the difference in the spatial positions of the corresponding carbon alpha atoms in the amino acids at positions i and j in the template sequence (i.e., in the folded structure of the template sequence), wherein the unit vector is calculated in the reference frame of amino acid i or amino acid j, and (iii) a discrete / packed representation of the distance between the spatial positions of the corresponding carbon alpha atoms in the amino acids at positions i and j in the template sequence.

[0260] System 100 can process each template representation using a sequence of one or more template update blocks 1114 to generate a corresponding updated template representation 1116 for each template sequence 1110. The template update blocks can include, for example, row-by-row self-attention blocks (e.g., which update the embeddings in each row of the template representation), column-by-column self-attention blocks (e.g., which update the embeddings in each column of the template representation), and transformation blocks (e.g., which apply one or more neural network layers to each embedding in the template representation).

[0261] After generating the updated template representation 1116, system 100 uses the updated template representation 1116 to update the pair embeddings 112. For example, system 100 can use "cross-attention" of the embeddings at the corresponding (i,j) positions in the updated template representation 1116 to update the corresponding pair embeddings 112 at each position (i,j). In the cross-attention operation for updating the pair embeddings 112 at position (i,j), a query embedding is generated from the pair embedding at position (i,j), and key and value embeddings are generated from the embeddings at the corresponding (i,j) positions in the updated template representation 1116.

[0262] Updating the embedding 112 using the template sequence 1110 enables the system 100 to enrich the embedding with information characterizing the protein structure of the evolution-related template sequence 1110, thereby enhancing the information content of the embedding 112 and improving the accuracy of the protein structure predicted using the embedding.

[0263] This specification uses the term "configuration" in conjunction with system and computer program components. For a system of one or more computers configured to perform a specific operation or action, this means that software, firmware, hardware, or a combination thereof have been installed on the system, which, in operation, causes the system to perform the operation or action. For one or more computer programs configured to perform a specific operation or action, this means that one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.

[0264] Embodiments of the subject matter and functional operation described in this specification can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their equivalents), or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium, for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof. Alternatively or additionally, program instructions can be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0265] The term "data processing apparatus" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The apparatus may also be or include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.

[0266] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not necessarily, correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data, for example, as one or more scripts stored in a markup language document, as a single file dedicated to the program in question, or as multiple harmonized files, for example, as a file storing one or more modules, subroutines, or code portions. A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a data communication network.

[0267] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in others, multiple engines can be installed and run on the same one or more computers.

[0268] The processes and logic flows described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGA or ASIC), or by a combination of special-purpose logic circuitry and one or more programmable computers.

[0269] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be supplemented by or incorporated into special-purpose logic circuitry. Typically, a computer will also include one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to, or both. However, a computer does not necessarily need to have such devices. Furthermore, a computer can be embedded in another device, to name just a few, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.

[0270] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0271] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser in response to a request received from a web browser on the user's device. Furthermore, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user as a response.

[0272] The data processing apparatus used to implement machine learning models may also include, for example, dedicated hardware accelerator units for handling the common and computationally intensive parts of machine learning training or production, namely inference and workloads.

[0273] It is possible to implement and deploy machine learning models using machine learning frameworks (such as TensorFlow, Microsoft's Cognitive Toolkit, Apache's Singa, or Apache's MXNet).

[0274] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface, web browser, or app that a user can interact with through an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0275] A computing system can include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship is established by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., HTML pages) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from it. It is possible to receive data generated at the user device, such as the result of user interaction, from the device at the server.

[0276] While this specification contains numerous details of specific implementation, these should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, one or more features from a claimed combination can be removed from the combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0277] Similarly, although operations are depicted in the accompanying drawings and described in a specific order in the claims, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all of the shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0278] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in different orders and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method for training a structure prediction neural network, wherein, The structural prediction neural network includes: (i) an embedding neural network having multiple embedding parameters and configured to receive network input characterizing a protein and process the network input according to the embedding parameters to generate an embedding output about the network input; and (ii) a master fold neural network having multiple master fold parameters and configured to receive the embedding output and process the master fold input according to the master fold parameters to generate a master structure prediction defining the predicted structure of the protein, the method comprising: Obtain data on the training network input characterizing the training protein and the target protein structure for the specified training protein; The training network input is processed using an embedded neural network and the current values ​​of the embedding parameters to generate a training embedding output about the training network input. The master fold neural network is used and the training embedding output is processed according to the current value of the master fold parameters to generate master structure predictions that define the master prediction structure of the trained protein. For each auxiliary folding neural network in a set of one or more auxiliary folding neural networks with corresponding multiple auxiliary folding parameters, the corresponding intermediate output of the embedded neural network is processed using the auxiliary folding neural network and based on the current value of the corresponding auxiliary folding parameter of the auxiliary folding neural network to generate an auxiliary structure prediction that defines the auxiliary prediction structure of the trained protein. Determine the gradient of the objective function, which includes: The principal structure loss term characterizes the similarity between: (i) the principal prediction structure defined by the principal structure prediction and (ii) the target protein structure of the training protein; and The corresponding auxiliary structure loss term for each auxiliary folding neural network characterizes the similarity between: (i) the auxiliary prediction structure defined by the auxiliary structure prediction generated by the auxiliary structure prediction neural network and (ii) the target protein structure of the training protein; and The current values ​​of the embedded network parameters, the main fold parameters, and the corresponding auxiliary fold parameters of one or more auxiliary fold neural networks are updated based on gradients.

2. The method according to claim 1, wherein, Each auxiliary folded neural network has the same neural network architecture as the main folded neural network.

3. The method according to claim 2, wherein, An ensemble of one or more auxiliary folded neural networks includes multiple auxiliary folded neural networks, wherein an objective function constrains the auxiliary folded neural networks to share parameter values.

4. The method according to claim 2, wherein, The objective function constrains the auxiliary folded neural network and the main folded neural network to share parameter values.

5. The method according to any one of the preceding claims, wherein: The network input includes the corresponding initial pair embeddings for each amino acid pair in the protein. The embedded neural network includes a sequence of update blocks, wherein each update block performs an operation, the operation including: Receive block input, which includes the corresponding current pair embedding for each amino acid pair in the protein; and Update the corresponding current pair embedding for each amino acid pair in the protein to generate the corresponding updated pair embedding for each amino acid pair in the protein, and The embedded output includes at least the pair of embeddings of the updates generated by the last updated block in the sequence.

6. The method according to claim 5, wherein, Each auxiliary folding neural network corresponds to a different update block in the sequence that is not the last update block in the sequence, and wherein each auxiliary folding neural network is configured to receive at least the pair embeddings of updates generated by the corresponding update block as input.

7. The method according to claim 6, wherein: The network input also includes initial multiple sequence alignment (MSA) embeddings, which represent the corresponding multiple sequence alignments to each strand in the protein. The block input to each of the update blocks further includes the current MSA embedding, and The operations performed by each update block further include: Update the current MSA embedding to generate an updated MSA embedding.

8. The method according to claim 7, wherein, The input to each auxiliary folded neural network also includes the updated MSA embedding generated from the corresponding update block.

9. The method according to any one of claims 7 or 8, wherein, Each auxiliary folded neural network is configured as follows: A transformed structure prediction is generated from the auxiliary structure prediction generated by the auxiliary folding neural network, and the transformed structure prediction has the same dimension as the updated pair embedding; The transformed structural prediction is combined with the updated pair embeddings to generate further updated pair embeddings; and Provides further updated pairs of embeddings as input to the next update block in the sequence following the current update block.

10. The method according to claim 9, wherein: The auxiliary folding prediction includes structural parameters, which specify the predicted 3-D spatial position of a specific atom in the protein structure for each amino acid. as well as Generating transformed structure predictions includes: Distance maps are generated from the predicted 3-D spatial locations of amino acids specified by structural parameters. For each amino acid pair in the protein, the distance map characterizes the corresponding estimated distance between that amino acid pair in the protein structure; and Generate a transformed distance map from the distance map, having the same dimensions as the updated pair embeddings.

11. The method according to claim 9, wherein: Auxiliary folding prediction includes structural parameters for a distance map that, for each amino acid pair in a protein, characterizes the estimated distance between that amino acid pair in the protein structure. and Generating transformed structure predictions includes: Generate a transformed distance map with the same dimensions as the initial pair of embeddings from the distance map specified by the structural parameters.

12. The method according to claim 1, further comprising: After training, new network inputs representing new proteins are obtained; as well as New network inputs are processed using a trained structure prediction neural network to generate new master structure predictions that define the predicted structure of new proteins.

13. A system comprising: One or more computers; as well as One or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operation of the corresponding method according to any one of claims 1-12.

14. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause one or more computers to perform the operations of the corresponding method according to any one of claims 1-12.

15. A method for obtaining a ligand, wherein, The ligand is a ligand for a drug or an industrial enzyme, and the method includes: Perform the method according to any one of claims 1-12 to determine the predicted structure of the target protein; Evaluate the interaction between one or more candidate ligands and the predicted structure of the target protein; and One or more candidate ligands are selected as ligands based on the evaluation results.

16. The method according to claim 15, wherein, The target proteins include receptors or enzymes, and the ligands are agonists or antagonists of the receptors or enzymes.

17. The method according to claim 15 or 16, wherein, The ligand is a drug, and the method includes: Perform the method according to any one of claims 1-12 to determine the predicted structure of each of the plurality of target proteins; Evaluate the interaction between one or more candidate ligands and the predicted structure of each target protein; and Select one or more candidate ligands as ligands to i) obtain ligands that interact with each of the target proteins, or ii) obtain ligands that interact with only one of the target proteins.

18. A method for obtaining polypeptide ligands, wherein, The ligand is a ligand for a drug or an industrial enzyme, and the method includes... For each of one or more candidate polypeptide ligands, the method according to any one of claims 1-12 is performed to determine the predicted structure of the candidate polypeptide ligand; Obtain the target protein structure of the target protein; Evaluate the interaction between the predicted structure of each of one or more candidate polypeptide ligands and the target protein structure; as well as Depending on the evaluation results, one of one or more candidate peptide ligands is selected as the peptide ligand.

19. The method according to claim 18, wherein, The target protein includes a receptor or enzyme, and wherein the ligand is an agonist or antagonist of the receptor or enzyme; or wherein the polypeptide ligand contains an antibody, and the target protein contains an antibody target, particularly a viral or cancer cell protein, and wherein the antibody binds to the antibody target to provide a therapeutic effect.

20. A method for obtaining a diagnostic antibody biomarker for a disease, the method comprising: For each of one or more candidate antibodies, the method according to any one of claims 1-12 is performed to determine the predicted structure of the candidate antibody; Obtain the target protein structure of the target protein; Evaluate the interaction between the predicted structure of each of one or more candidate antibodies and the target protein structure; as well as Depending on the evaluation results, one of one or more candidate antibodies may be selected as a diagnostic antibody biomarker.

21. The method of claim 20, wherein, Evaluating the interaction of one of the candidate ligands involves determining the interaction score of the candidate ligand, where the interaction score is a measure of the interaction between the candidate ligand and the target protein.

22. The method of claim 15, further comprising: Synthetic ligands.

23. The method of claim 22, further comprising: The bioactivity of the ligands was tested in vitro and in vivo.

24. A method for identifying the presence of a protein misfolding disease, comprising: Perform the method according to any one of claims 1-12 to determine the predicted structure of the protein; To obtain the structure of a protein version obtained from a human or animal body; Compare the predicted structure of a protein with the structure of a protein version obtained from a human or animal. as well as The presence of protein misfolding diseases can be identified based on the comparison results.

25. A system comprising: One or more computers; as well as One or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform: The structure prediction neural network includes: (i) an embedding neural network having multiple embedding parameters and configured to receive network inputs characterizing proteins and process the network inputs according to the embedding parameters to generate embedding outputs about the network inputs; and (ii) a master fold neural network having multiple master fold parameters and configured to receive embedding outputs and process master fold inputs according to the master fold parameters to generate master structure predictions that define the predicted structure of the protein. A collection of one or more auxiliary folding neural networks, wherein each auxiliary folding neural network has a corresponding plurality of auxiliary folding parameters and is configured to process at least the intermediate output of the embedded neural network based on the current value of the corresponding auxiliary folding parameter of the auxiliary folding neural network to generate an auxiliary structure prediction that defines the auxiliary predicted structure of a protein. The training subsystem is configured to perform operations, including: Obtain data on the training network input characterizing the training protein and the target protein structure for the specified training protein; The training network input is processed using an embedded neural network and the current values ​​of the embedding parameters to generate a training embedding output about the training network input. The master fold neural network is used and the training embedding output is processed according to the current value of the master fold parameters to generate master structure predictions that define the master prediction structure of the trained protein. For each of the one or more auxiliary folded neural networks in the set, the corresponding intermediate output of the embedded neural network is processed using the auxiliary folded neural network to generate an auxiliary structure prediction that defines the auxiliary prediction structure of the trained protein. Determine the gradient of the objective function, which includes: The principal structure loss term characterizes the similarity between: (i) the principal prediction structure defined by the principal structure prediction and (ii) the target protein structure of the training protein; and The corresponding auxiliary structure loss term for each auxiliary folding neural network characterizes the similarity between: (i) the auxiliary predicted structure defined by the auxiliary structure prediction generated by the auxiliary structure prediction neural network and (ii) the target protein structure of the training protein; and The current values ​​of the embedded network parameters, the main fold parameters, and the corresponding auxiliary fold parameters of one or more auxiliary fold neural networks are updated based on gradients.