Methods, systems, and storage media for predicting symmetric protein structures
Patent Information
- Application Number
- CN202610948902.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-28
- Filing Date
- 2021-11-23
- Publication Date
- 2026-09-11
Smart Images

Figure CN122738579A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on November 23, 2021, with application number 202180068764.2 and invention title "Using Symmetry Extended Transformation to Predict Symmetric Protein Structure".
[0002] Cross-reference to related applications
[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 118,914, filed on November 28, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0004] This manual relates to the prediction of protein structure. Background Technology
[0005] Proteins are defined by one or more amino acid sequences (“chains”). Amino acids are organic compounds that include amino and carboxyl functional groups, as well as amino acid-specific side chains (i.e., atomic groups). Protein folding refers to the physical process by which one or more amino acid sequences fold into a three-dimensional (3-D) conformation. The structure of a protein defines the 3-D conformation of the atoms in the protein’s amino acid sequence after protein folding. When in a sequence linked by peptide bonds, an amino acid can be referred to as an amino acid residue.
[0006] Machine learning models can be used for prediction. A machine learning model takes input and generates an output, such as a predicted output, based on that input. Some machine learning models are parametric models and generate outputs based on the received input and the model's parameter values. Some machine learning models are deep models, which employ multiple layers to generate outputs for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, each of which applies a non-linear transformation to the received input to generate an output. Summary of the Invention
[0007] This specification describes a protein structure prediction system, implemented as a computer program on one or more computers in one or more locations, capable of predicting the structure of symmetrical proteins.
[0008] The term "protein" can be understood as any biomolecule specified by a sequence (or "chain") of one or more amino acids. For example, the term protein can refer to a protein domain, such as a portion of the amino acid chain of a protein, which is capable of undergoing protein folding almost independently of the rest of the protein. As another example, the term protein can refer to a protein complex, which comprises multiple amino acid chains that fold together to form the protein structure.
[0009] Multiple sequence alignment (MSA) of an amino acid sequence in a protein specifies the alignment of an amino acid sequence with multiple other amino acid sequences (called "MSA sequences," such as those from other proteins, like homologous proteins). More specifically, an MSA defines the correspondence between positions in an amino acid chain and corresponding positions in multiple MSA sequences. An MSA of an amino acid sequence can be generated, for example, by processing a database of amino acid sequences using any suitable computational sequence alignment technique, such as progressive alignment construction. MSA sequences can be interpreted as having evolutionary relationships, for example, where each MSA sequence may share a common ancestor. The correlations between amino acids in an amino acid chain within an MSA of an amino acid chain can encode information related to predicting the structure of the amino acid chain.
[0010] The “embedding” of an entity (e.g., an amino acid pair) can refer to a representation of an entity as an ordered set of numerical values, such as a vector or matrix of numerical values.
[0011] The structure of a protein can be defined by a set of structural parameters. This set of structural parameters can be represented as an ordered set of values. Below are some examples of possible structural parameters used to define the structure of a protein, described in more detail.
[0012] In one example, for each amino acid in a protein, the structural parameters that define the structure of the protein include: (i) positional parameters and (ii) rotational parameters.
[0013] Amino acid position parameters specify the predicted 3D spatial position of a designated atom within the amino acid group of a protein. The designated atom can be the alpha carbon atom in the amino acid, i.e., the carbon atom bonded to the amino group, carboxyl group, and side chain of the amino acid. Amino acid position parameters can be represented in any suitable coordinate system, such as three-dimensional space. Cartesian coordinate system.
[0014] The rotation parameters of an amino acid can specify its predicted "orientation" within the protein structure. More specifically, the rotation parameters can specify a 3-D spatial rotation operation, which, when applied to a coordinate system of position parameters, causes the three "backbone" atoms in the amino acid to take fixed positions relative to the rotation coordinate system. The three backbone atoms in an amino acid can refer to a series of nitrogen, alpha carbon, and carbonyl carbon atoms linked together in the amino acid. The rotation parameters of an amino acid can be represented, for example, as an orthogonal 3×3 matrix with a determinant of 1.
[0015] Typically, the position and rotation parameters of an amino acid define its egocentric reference frame. In this frame, each amino acid's side chain can start from the origin and proceed along the defined direction along the first bond of the side chain (i.e., the alpha-beta bond).
[0016] In another example, the structural parameters defining the structure of a protein can include a “distance map” that characterizes the estimated distances (e.g., measured in angstroms) between each pair of amino acids in the protein. The distance map can characterize the estimated distances between amino acid pairs, for example, through a probability distribution of a set of possible distances between amino acid pairs.
[0017] In another example, structural parameters that define the structure of a protein can define the three-dimensional (3D) spatial location of each atom in each amino acid of the protein.
[0018] The protein structure prediction system described herein can be used to obtain ligands, such as ligands for drugs or industrial enzymes. For example, a method for obtaining a ligand may include obtaining the target amino acid sequence, particularly the amino acid sequence of a target protein (e.g., a drug target), and using the protein structure prediction system to process the input based on the target amino acid sequence to determine the (tertiary) structure of the target protein, i.e., predicting the protein structure. The method may then include evaluating the interaction between one or more candidate ligands and the target protein structure. The method may also include selecting one or more candidate ligands as ligands based on the evaluation results of the interaction.
[0019] In some embodiments, evaluating the interaction may include assessing the binding of a candidate ligand to a target protein structure. For example, evaluating the interaction may include identifying ligands that bind with sufficient affinity to achieve a biological effect. In some other embodiments, evaluating the interaction may include assessing the association of a candidate ligand to a target protein structure that has an impact on the function of the target protein (e.g., an enzyme). The assessment may include assessing the affinity between the candidate ligand and the target protein structure, or assessing the selectivity of the interaction. Candidate ligands may be selected based on the candidate ligand with the highest affinity.
[0020] Candidate ligands (or multiple ligands) can be derived from a database of candidate ligands, and / or can be derived by modifying ligands in the database, for example by modifying the structure or amino acid sequence of the candidate ligands, and / or by stepwise or iterative assembly / optimization of the candidate ligands.
[0021] The evaluation of the interaction between the candidate ligand and the target protein structure can be performed using computer-aided methods, wherein a graphical model displaying the candidate ligand and target protein structure is used for user manipulation, and / or the evaluation can be performed partially or fully automatically, for example using standard molecule (protein-ligand) docking software. In some embodiments, the evaluation may include determining the interaction score of the candidate ligand, wherein the interaction score includes a measurement of the interaction between the candidate ligand and the target protein. The interaction score may depend on the strength and / or specificity of the interaction, for example, on the fraction of the binding free energy. The candidate ligand may be selected based on its score.
[0022] In some embodiments, the target protein includes a receptor or enzyme, and the ligand is an agonist or antagonist of the receptor or enzyme. In some embodiments, the method can be used to identify the structure of a cell surface marker. This can then be used to identify ligands, such as antibodies or markers (e.g., fluorescent markers), that bind to the cell surface marker. This can be used to identify and / or treat cancer cells.
[0023] In some embodiments, the ligand is a drug, and a predicted structure for each of a plurality of target proteins is determined, and the interaction of one or more candidate ligands with the predicted structure of each of the target proteins is evaluated. One or more candidate ligands can then be selected to obtain a ligand that interacts with each target protein (functionally), or to obtain a ligand that interacts with only one target protein (functionally). For example, in some embodiments, it is desirable to obtain a drug that is effective against multiple drug targets. Alternatively or concurrently, it is desirable to screen for off-target effects of the drug. For example, in agriculture, it can be useful to determine that a drug designed for one plant species does not interact with another different plant species and / or animal species.
[0024] In some implementations, the ligand is a drug, and a predicted structure of a target protein is determined, which is a protein complex, such as a dimer or multimer. Evaluating the interaction of one or more candidate ligands with the predicted structure of the target protein can then include identifying candidate ligands that interact with the protein complex and can therefore be expected to influence the formation or stability of the complex. This can then be confirmed through experimental screening. Therefore, this process can be used to identify drugs that can disrupt protein complexes or inhibit their formation. Some diseases, such as neurodegenerative diseases like dementia, are caused by protein aggregation. Therefore, this method can be used to identify ligands that could serve as drugs for treating such diseases.
[0025] In some embodiments, candidate ligands (or multiple ligands) may include small molecule ligands, such as organic compounds with a molecular weight <900 Daltons. In other embodiments, candidate ligands (or multiple ligands) may include peptide ligands, i.e., peptide ligands defined by an amino acid sequence.
[0026] In some cases, protein structure prediction systems can be used to determine the structure of candidate polypeptide ligands (e.g., ligands for drugs or industrial enzymes). Their interaction with the target protein structure can then be evaluated; target protein structures can already be determined using structure prediction neural networks or conventional physical probing techniques such as X-ray crystallography and / or magnetic resonance imaging or cryo-electron microscopy.
[0027] On the other hand, a method is provided for obtaining peptide ligands (e.g., molecules or their sequences) using a protein structure prediction system. This method may include obtaining the amino acid sequences of one or more candidate peptide ligands. The method may also include using a protein structure prediction system to determine the (tertiary) structure of the candidate peptide ligands. The method may further include obtaining the target protein structure of a target protein through bioinformatics and / or physical probing, and evaluating the interaction between the structure of each of the one or more candidate peptide ligands and the target protein structure. The method may also include selecting one or more candidate peptide ligands as peptide ligands based on the evaluation results.
[0028] As previously described, evaluating interactions can include assessing the binding of candidate peptide ligands to target protein structures, such as identifying ligands that bind with sufficient affinity to achieve biological effects, and / or assessing the association of candidate peptide ligands with target protein structures that influence the function of the target protein (e.g., enzymes), and / or assessing the affinity between candidate peptide ligands and target protein structures, or assessing the selectivity of the interaction. In some embodiments, the peptide ligand may be an aptamer. Similarly, peptide candidate ligands (or multiple ligands) may be selected based on the peptide candidate ligand with the highest affinity.
[0029] As previously described, the selected polypeptide ligand may include a receptor or an enzyme, and the ligand may be an agonist or antagonist of the receptor or enzyme. In some embodiments, the polypeptide ligand may include an antibody, and the target protein may include an antibody target, such as a virus, particularly a viral capsid protein, or a protein expressed on cancer cells. In these embodiments, the antibody binds to the antibody target to provide a therapeutic effect. For example, the antibody may bind to the target and act as an agonist of a specific receptor; alternatively, the antibody may prevent another ligand from binding to the target and thus prevent activation of the associated biological pathway.
[0030] The implementation of this method may also include synthesis, i.e., the preparation of small molecule or peptide ligands. The ligands can be synthesized by any conventional chemical technique and / or can be already available, for example, derived from a compound library or synthesized using combinatorial chemistry.
[0031] This method may also include testing the bioactivity of the ligand in vitro and / or in vivo. For example, the ADME (absorption, distribution, metabolism, excretion) and / or toxicological properties of the ligand may be tested to screen out unsuitable ligands. Testing may include, for example, contacting the candidate small molecule or peptide ligand with the target protein and measuring changes in protein expression or activity.
[0032] In some embodiments, candidate (peptide) ligands may include: separating antibodies, fragments of separating antibodies, monovariable domain antibodies, bispecific or multispecific antibodies, multivalent antibodies, bivariable domain antibodies, immunoconjugates, fibronectin molecules, adnectin, DARPin, avimer, affinity molecules, anticarrier proteins, affilin, protein epitope mimics, or combinations thereof. Candidate (peptide) ligands may include antibodies with mutated or chemically modified amino acid Fc regions, for example, which, compared to wild-type Fc regions, inhibit or reduce ADCC (antibody-dependent cytotoxicity) activity and / or increase half-life. Candidate (peptide) ligands may include antibodies with different CDRs (complementarity-determining regions).
[0033] The protein structure prediction system described herein can also be used to obtain diagnostic antibody biomarkers for diseases. A method is also provided in which, for each of one or more candidate antibodies, such as as described above, the method uses the protein structure prediction system to determine the predicted structure of the candidate antibody. The method may also involve obtaining the target protein structure, evaluating the interaction between the predicted structure of each of one or more candidate antibodies and the target protein structure, and selecting one of one or more candidate antibodies as a diagnostic antibody biomarker based on the evaluation results, for example, selecting one or more candidate antibodies with the highest affinity for the target protein structure. The method may include the preparation of diagnostic antibody biomarkers. Diagnostic antibody biomarkers can be used to diagnose diseases by detecting whether they bind to a target protein in a sample obtained from a patient (e.g., a body fluid sample). As described above, corresponding techniques can be used to obtain therapeutic antibodies (peptide ligands).
[0034] Misfolded proteins are associated with many diseases. Therefore, in another respect, a method is provided to identify the presence of protein misfolding diseases using a protein structure prediction system. This method may include obtaining the amino acid sequence of a protein and using a protein structure prediction system to determine the protein's structure. The method may also include obtaining the structure of a protein version obtained from a human or animal body, for example, through conventional (physical) methods. The method then includes comparing the protein's structure with the structure of the version obtained from the body and identifying the presence of a protein misfolding disease based on the comparison result. That is, misfolding of the protein version from the body can be determined by comparing it with the structure determined through bioinformatics.
[0035] Typically, identifying the presence of a protein misfolding disorder may involve obtaining the amino acid sequence of a protein, using the amino acid sequence to determine the protein's structure, as described herein, and comparing the protein's structure to the structure of a baseline version of the protein. The presence of the protein misfolding disorder is identified based on the results of the comparison. For example, the structures being compared may be those of a mutant and a wild-type protein. In this implementation, a wild-type protein may be used as the baseline version, but in principle, either can be used as the baseline version.
[0036] In some other respects, computer-implemented methods, such as those described above or herein, can be used to identify active / binding / blocking sites on target proteins from their amino acid sequences.
[0037] According to one aspect, a method for predicting the structure of a protein comprising multiple amino acid chains is provided, executed by one or more data processing devices. The method includes: obtaining initial structural parameters of a first amino acid chain in the protein, wherein the structural parameters of the first amino acid chain in the protein define predicted three-dimensional (3D) spatial positions of amino acids in the first amino acid chain in the structure of the protein; obtaining data identifying symmetry groups, wherein the predicted protein folds into a structure symmetrical with respect to the symmetry groups; processing the input including the initial structural parameters of the first amino acid chain and the data identifying symmetry groups using a folding neural network to generate an output defining a final predicted structure of the protein symmetrical with respect to the symmetry groups, wherein the folding neural network includes an update block sequence, wherein each update block in the update block sequence has multiple update block parameters, and performs operations including: receiving current structural parameters of the first amino acid chain and the data identifying symmetry groups; applying a symmetry extension transformation to the current structural parameters of the first amino acid chain to generate corresponding current structural parameters of each other amino acid chain in the protein to define the current predicted structure of the protein symmetrical with respect to the symmetry groups; and processing the current structural parameters of the amino acid chains in the protein according to the values of the update block parameters of the update blocks to update the current structural parameters of the first amino acid chain.
[0038] In some embodiments, the structural parameters of the first amino acid chain in the protein include the corresponding amino acid structural parameters of each amino acid in the first amino acid chain, and the amino acid structural parameters of each amino acid define the 3D spatial position and orientation of the amino acid in the reference frame of the first amino acid chain.
[0039] In some implementations, the structural parameters of the first amino acid chain in the protein include global structural parameters that define the 3D spatial position and orientation of the first amino acid chain in the protein's reference frame.
[0040] In some implementations, applying a symmetry extension transformation to the current structural parameters of the first amino acid chain to generate corresponding current structural parameters for each other amino acid chain in the protein includes, for each other amino acid chain in the protein: generating global structural parameters for the other amino acid chains by applying a predefined transformation to the global structural parameters of the first amino acid chain, wherein the predefined transformation depends on: (i) the number of amino acid chains in the protein, and (ii) symmetry groups; and determining the amino acid structural parameters of the amino acids in the other amino acid chains that match the amino acid structural parameters of the amino acids in the first amino acid chain.
[0041] In some embodiments, the symmetric group is a cyclic symmetric group, a dihedral symmetric group, or a cubic symmetric group.
[0042] In some implementations, the input processed by the folded neural network further includes: (i) the corresponding initial amino acid embedding for each amino acid in the first amino acid chain, and (ii) the initial global embedding for the first amino acid chain.
[0043] In some implementations, the operations performed by each update block further include receiving the corresponding current amino acid embedding and the current global embedding of each amino acid in the first amino acid chain; and processing the current structural parameters of the amino acid chain in the protein to update the current structural parameters of the first amino acid chain includes: updating the current amino acid embedding and the current global embedding of the first amino acid chain based on the current structural parameters of the amino acid chain in the protein; and updating the current structural parameters of the first amino acid chain based on the updated amino acid embedding and the updated global embedding of the first amino acid chain.
[0044] In some implementations, updating the current amino acid embedding of the first amino acid chain based on the current structural parameters of the amino acid chains in the protein includes: for each other amino acid chain in the protein, determining the corresponding current amino acid embedding of each amino acid in the other amino acid chains based on the current amino acid embedding of the corresponding amino acid in the first amino acid chain; and updating the current amino acid embedding of the first amino acid chain and the current global embedding using attention to the current amino acid embedding of the amino acid chain, wherein the attention to the current amino acid embedding of the amino acid chain is conditioned on the current structural parameters of the amino acid chain.
[0045] In some implementations, updating the current global embedding of the first amino acid chain using attention to the current amino acid embedding of the amino acid chain includes:
[0046] For each amino acid in each amino acid chain, the corresponding attention weight between the current global embedding of the first amino acid chain and the current amino acid embedding of the amino acid is determined at least in part based on the following: (i) the global structural parameters of the first amino acid chain, and (ii) the amino acid structural parameters of the amino acid and the global structural parameters of the amino acid chain; and the current global embedding of the first amino acid chain is updated based on (i) the attention weight and (ii) the current amino acid embedding of the amino acid chain.
[0047] In some implementations, for each amino acid in each amino acid chain, determining the attention weight between the current global embedding of the first amino acid chain and the current amino acid embedding of that amino acid includes: generating a geometric query embedding corresponding to the current global embedding of the first amino acid chain, including: processing the current global embedding of the first amino acid chain using one or more neural network layers to generate a 3D embedding; rotating and translating the 3D embedding into a protein reference frame using global structural parameters of the first amino acid chain; generating a geometric keyword embedding corresponding to that amino acid, including: processing the current amino acid embedding of that amino acid using one or more neural network layers to generate a 3D embedding; rotating and translating the 3D embedding into a protein reference frame using amino acid structural parameters of the amino acid and global structural parameters of the amino acid chain; and determining the attention weight based on the spatial distance between (i) the geometric query embedding corresponding to the current global embedding of the first amino acid chain and (ii) the geometric keyword embedding corresponding to that amino acid.
[0048] In some implementations, updating the current structural parameters of the first amino acid chain based on the updated amino acid embedding and the updated global embedding of the first amino acid chain includes: for each amino acid in the first amino acid chain, updating the amino acid structural parameters of that amino acid based on the updated amino acid embedding; and updating the global structural parameters of the first amino acid chain based on the updated global embedding of the first amino acid chain.
[0049] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages.
[0050] This specification describes a system for predicting the structure of proteins comprising multiple identical amino acid chains and expected to fold into a structure symmetrical with respect to symmetry groups (e.g., cyclic, dihedral, or cubic symmetry groups). Specifically, the system uses a neural network to predict the protein structure, which iteratively refines the currently predicted structure of the protein while explicitly enforcing, at each iteration, that the currently predicted structure of the iteratively refined protein is symmetrical with respect to the symmetry groups. Explicitly enforcing the known symmetry of the protein structure throughout the iterative refinement of the predicted protein structure can significantly improve the accuracy of structure predictions made by the neural network.
[0051] To explicitly enforce protein structural symmetry, neural networks can internally represent the currently predicted protein structure as a function of the predicted structures of individual amino acid chains within the protein. The neural network directly updates the predicted structure of only a single amino acid chain, while the rest of the protein structure is indirectly updated because it is defined as a function of the predicted structures of individual amino acid chains. This allows neural networks to perform fewer operations and thus consume fewer computational resources (e.g., memory and computing power), for example, compared to neural networks that perform operations to independently update the predicted structure of every single amino acid chain in the protein.
[0052] The structure of a protein determines its biological function. Therefore, determining protein structure can facilitate the understanding of life processes (e.g., the mechanisms of many diseases) and the design of proteins (e.g., as drugs, or as enzymes for industrial processes). For example, which molecules (e.g., drugs) will bind to a protein (and where this binding will occur) depends on the protein's structure. Since the effectiveness of drugs can be influenced by the extent to which they bind to proteins (e.g., in the blood), determining the structure of different proteins is an important aspect of drug development. However, determining protein structure using physical experiments (e.g., via X-ray crystallography) can be time-consuming and very expensive. Therefore, the protein prediction system described in this specification can facilitate biochemical research and engineering fields involving proteins (e.g., drug development).
[0053] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims. Attached Figure Description
[0054] Figure 1The diagram illustrates the chain structure parameters of the first amino acid chain in a protein processed by a symmetric extension block to generate the structure parameters of the other amino acid chains in the protein. These structure parameters collectively define the protein structure symmetrical with respect to the C4 symmetric group.
[0055] Figure 2 An example architecture of a folded neural network is shown.
[0056] Figure 3 An example protein structure prediction system is shown.
[0057] Figure 4 An example embedded system is shown.
[0058] Figure 5 An example architecture for embedding a neural network is shown.
[0059] Figure 6 An example architecture is shown that embeds an update block of a neural network.
[0060] Figure 7 This shows an example schema for the MSA update block.
[0061] Figure 8 An example schema for updating blocks is shown.
[0062] Figure 9 An example embedded extension system is shown.
[0063] Figure 10 This is a flowchart of an example process for predicting the structure of proteins that consist of multiple amino acid chains.
[0064] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation
[0065] This specification describes a protein structure prediction system (“System”) configured to predict the structure of proteins composed of multiple amino acid chains. The amino acid chains can be identical or nearly identical. Such proteins include, but are not limited to, protein complexes. If the amino acid chains consist of the same amino acid sequence, they are said to be “identical”.
[0066] In many cases, proteins composed of multiple, for example, identical amino acid chains fold into symmetrical protein structures. A protein structure is said to be symmetrical with respect to a symmetry group (i.e., a transformation operation) if the protein structure remains invariant under transformation operations according to the symmetry group—that is, if applying any transformation operation according to the symmetry group to the protein structure results in the orientation of the protein structure being effectively preserved. Transformation operations in the symmetry group can include, for example, rotations and reflections around a plane.
[0067] Examples of symmetrical groups include cyclic symmetrical groups (e.g., C2, C3, C4, etc.), dihedral symmetrical groups (e.g., D2, D3, D4, etc.), and cubic symmetrical groups. For example, a C4 symmetrical group refers to a transformation category defined by multiples of 90 degrees of rotation around an axis (e.g., 90 degrees, 180 degrees, 270 degrees, 360 degrees, etc.).
[0068] As part of predicting the structure of symmetrical proteins composed of multiple, for example, identical amino acid chains, the protein structure prediction system uses a symmetry extension block 100, i.e., as a reference... Figure 2 A part of a more detailed description of folded neural networks.
[0069] The symmetry expansion block 100 is configured to receive block input characterizing an amino acid chain in a protein, which, for convenience, is referred to herein as the “first” amino acid chain of the protein. (The term “first amino acid chain” as used in this document is used only as a convenient way to distinguish one of the amino acid chains of a protein from the other amino acid chains. Any amino acid chain of a protein can be the “first amino acid chain”). The block input of the symmetry expansion block includes: (i) structural parameters of the first amino acid chain (i.e., “chain structure parameters” 104), and (ii) data identifying symmetry groups of the protein complex (i.e., such that the protein structure is symmetrical with respect to the symmetry groups).
[0070] The chain structure parameter 104 of the first amino acid chain of the protein complex can include: (i) the corresponding "amino acid" structure parameter for each amino acid in the first amino acid chain, and (ii) the "global" structure parameter of the first amino acid chain.
[0071] The amino acid structural parameters for each amino acid in an amino acid chain can include, for example, position parameters defining the 3-D spatial position of a specified atom within the amino acid, and rotation parameters specifying the orientation of the amino acid. The position parameters of an amino acid can be, for example, determined by... Cartesian coordinates are used, and the rotation parameters of the amino acids can be represented, for example, by an orthogonal 3×3 matrix with a determinant of 1, as described above.
[0072] The global structural parameters of an amino acid chain can include both "global" positional parameters and "global" rotational parameters. The global positional parameters of an amino acid chain can be, for example, derived from... Cartesian coordinates are used, and the global rotation parameters of the amino acid chain can be represented, for example, by an orthogonal 3×3 matrix with a determinant of 1.
[0073] The amino acid structural parameters of an amino acid chain and global structural parameters together define the (predicted) positions and orientations of amino acids within the amino acid chain in the structure of a protein complex. For example, the spatial position of each amino acid can be obtained by summing its positional parameters and the global positional parameters of the amino acid chain. The orientation of each amino acid can be obtained by combining (i.e., matrix multiplication) the rotation matrix representing the amino acid rotational parameters and the global rotational parameters of the amino acid chain.
[0074] Generally, the amino acid structural parameters of an amino acid chain can be understood as defining the spatial position and orientation of the amino acids within the chain in a local reference frame. The global structural parameters of an amino acid chain can be understood as defining translation and rotation operations that move the entire amino acid chain to its position in the global reference frame of the protein structure, i.e., relative to other amino acid chains.
[0075] Data identifying the symmetric group 106 of a protein can be represented, for example, by a uniquely heated vector.
[0076] The symmetry expansion block 100 processes data defining the symmetry group 106 of protein 110 and the chain structure parameters 104 of the first amino acid chain in the protein to generate chain structure parameters (i.e., "protein structure parameters" 108) for each other amino acid chain in the protein. Specifically, the symmetry expansion block 100 generates chain structure parameters for the other amino acid chains in the protein, which collectively define the structure of the protein complex symmetrical with respect to the symmetry group 106.
[0077] Typically, the amino acid structural parameters are the same for each amino acid chain in a protein, because the structure of each amino acid chain is identical within its local reference frame. However, the global structural parameters are different for each amino acid chain in a protein.
[0078] The symmetry expansion block 100 generates corresponding global structural parameters for each of the other amino acid chains by applying a "symmetry expansion transformation" to the global structural parameters of the first amino acid chain. The symmetry expansion transformation is a predefined function that, when applied to the global structural parameters of the first amino acid chain, generates corresponding global structural parameters for each of the other amino acid chains in the protein complex, such that the resulting protein complex structure is symmetrical with respect to the symmetry group 106. Typically, the symmetry expansion block 100 uses different predefined symmetry expansion transformations depending on (i) the symmetry group 106 of the protein structure and (ii) the number of amino acid chains in the protein.
[0079] In one example, if symmetry group 106 is a C2 symmetry group and the protein complex has two amino acid chains, the symmetry extension transformation can generate the global structural parameters of the other amino acid chains by applying a 180-degree rotation operation to the global rotation parameters of the first amino acid chain. In this example, the global position parameters can be the same for both amino acid chains in the protein complex.
[0080] As another example, if symmetry group 106 is a C4 symmetry group and the protein complex has four amino acid chains, the symmetry extension transformation can generate the global structural parameters of the other three amino acid chains by applying rotation operations of 90 degrees, 180 degrees, and 270 degrees to the global rotation parameter of the first amino acid chain. In this example, the global position parameter can be the same for all amino acid chains in the protein complex.
[0081] Figure 1 A symmetry extension block 100 is shown, which processes the chain structure parameter 104 of the first amino acid chain 102 to generate the structure parameters of other amino acid chains to collectively define a symmetric protein complex 110 relative to the C4 symmetry group.
[0082] Protein structure prediction systems can use folded neural networks to predict protein structures. These networks iteratively refine the chain structure parameters of the first amino acid chain in the protein. Folded neural networks can utilize symmetric extension blocks, such as reference... Figure 1 The entire structure of a symmetrical protein is defined as a function of the chain structure parameters of the first amino acid chain in the protein. Using symmetry extension blocks allows folded neural networks to explicitly enforce known symmetries in protein structure prediction, thereby improving the accuracy of protein structure prediction.
[0083] Figure 2 An example architecture of a folded neural network 200 is shown, which generates protein structure parameters 226 that define the predicted structure 228 of a symmetrical protein with multiple identical amino acid chains. In other words, the predicted structure 228 of a protein can be defined by a set of structure parameters 226, which collectively define the predicted three-dimensional structure of the protein after it has undergone protein folding.
[0084] The protein structure prediction system provides input to the folded neural network 200, which includes: (i) the corresponding amino acid embedding 202 for each amino acid in the first amino acid chain of the protein, (ii) the chain structure parameters 204 of the first amino acid chain, (iii) data defining the symmetry groups 218 of the protein, and (iv) a set of interaction embeddings 210.
[0085] The protein structure prediction system can use an embedding system to generate amino acid embeddings of the first amino acid chain, which will refer to... Figure 4 To describe in more detail.
[0086] The structural parameters 204 of the first amino acid chain include: (i) the corresponding amino acid structural parameters of each amino acid in the first amino acid chain, and (ii) the global structural parameters of the first amino acid chain, as referenced above. Figure 1 The protein structure prediction system can initialize the amino acid structure parameters of the first amino acid chain to default (predefined) values. For example, the position parameter of each amino acid can be initialized to the origin (e.g., [0,0,0]), and the rotation parameter of each amino acid can be initialized to a 3×3 identity matrix. Similarly, the protein structure prediction system can initialize the global structure parameters of the first amino acid chain to default values. Optionally, the protein structure prediction system can generate initial values for the global structure parameters of the first amino acid chain by processing interaction embeddings using one or more neural network layers (e.g., fully connected neural network layers).
[0087] Data defining the symmetry group 218 of a protein structure can be represented, for example, by an one-hot vector. The protein structure prediction system can internally predict the symmetry group 218 of a protein structure, as referenced... Figure 4 A more detailed description.
[0088] A set of interaction embeddings 210 can be represented as having Each row (i.e., where) (This refers to the number of amino acids in the first amino acid chain) and Columns (i.e., where) A 2D array of inter-amino acid chains (the number of amino acid chains in a protein) is a 2D array of inter-amino acid chains. In other words, the number of columns in the 2D array of inter-amino acid chains can be equal to the total number of amino acids in the protein.
[0089] Typically, each interaction embedding 210 corresponds to a specific amino acid pair in a protein and characterizes the relationship between amino acid pairs in the protein, for example by encoding information characterizing the spatial distance between amino acid pairs in the protein's structure. For example, amino acids in the first amino acid chain can be based on... Indexing allows amino acids in the entire protein complex (i.e., including all amino acid chains) to be indexed from... Indexing, and the position in the array Interactions at the site can characterize the amino acids in the first amino acid chain. Amino acids in protein complexes The relationship between them. Protein structure prediction systems can use embedding systems to generate interaction embeddings, such as referencing... Figure 4 A more detailed description.
[0090] In addition to obtaining the corresponding amino acid embeddings 202 of the amino acids in the first amino acid chain, the folded neural network 200 obtains the “global” embedding of the first amino acid chain. For example, as a result of the interaction embeddings 210 being processed by one or more neural network layers (e.g., fully connected neural network layers), the folded neural network 200 is able to generate (i.e., initialize) the global embedding of the first amino acid chain.
[0091] To generate protein structure parameters 226 that define the predicted protein structure 228, the folded neural network 200 is capable of repeatedly updating the current values of amino acid embeddings 206, the current values of structure parameters 208 of the first amino acid chain (i.e., starting from their initial values), and the current global embedding of the first amino acid chain. More specifically, the folded neural network 200 includes a sequence of update blocks 220, wherein each update block 220 is configured to update the current amino acid embedding 206 of the first amino acid chain (i.e., generate an updated amino acid embedding 222 of the first amino acid chain), update the current structure parameters 208 of the first amino acid chain (i.e., generate an updated structure parameter 224 of the first amino acid chain), and update the current global embedding of the first amino acid chain. In addition to update blocks, the folded neural network 200 may also include other neural network layers or blocks, for example, other neural network layers or blocks that may be interleaved with the update blocks.
[0092] Each update block 220 may include: (i) a symmetric extension block 100, (ii) a geometric attention block 214 and (ii) a folding block 216, each of which will be described in more detail below.
[0093] The symmetry expansion block 100 processes the current structural parameters 208 of the first amino acid chain and the data of the symmetry group 218 that identifies the protein to generate the current structural parameters of each of the other amino acid chains in the protein. Specifically, the symmetry expansion block 100 generates the current structural parameters of the other amino acid chains in the protein such that the current structural parameters of the amino acid chains collectively define the protein structure symmetrical with respect to the symmetry group 218, as referenced. Figure 1 More detailed description.
[0094] In addition to generating the current chain structure parameters for other amino acid chains in the protein, the update block also associates each other amino acid chain with its corresponding amino acid embedding. More specifically, the update block associates each amino acid in each other amino acid chain with its corresponding amino acid embedding. Specifically, for each other amino acid chain, update block 220 will assign the position of each amino acid in the other amino acid chain to its corresponding embedding. The amino acid at that position and its position in the first amino acid chain The current amino acid embedding at a given location is associated with that of the amino acid. In other words, the updated block "tiles" the amino acid embedding of the amino acid in the first amino acid chain across every other amino acid chain.
[0095] Geometric attention block 214 and folding block 216 use geometric attention operations to jointly update the amino acid structure parameters of the first amino acid chain and the global structure parameters, as will be described in more detail below.
[0096] To implement geometric attention operations, geometric attention block 214 determines a corresponding "symbol query" embedding, "symbol keyword" embedding, and "symbol value" embedding for each amino acid in the first amino acid chain. For example, geometric attention block 214 can process the corresponding amino acid embeddings. To generate amino acids Symbol query embedding Symbol keyword embedding and symbolic value embedding :
[0097] (1)
[0098] (2)
[0099] (3)
[0100] in, It refers to a linear layer with independent learning parameter values.
[0101] Geometric attention block 214 associates each amino acid in each of the other amino acid chains with the symbolic keyword embedding and symbolic value embedding of the corresponding amino acid in the first amino acid chain. Specifically, for each other amino acid chain, geometric attention block 214 associates the position of each amino acid in the other amino acid chain with the symbolic keyword embedding and symbolic value embedding. The amino acid at that position and its position in the first amino acid chain The symbolic keyword embedding and symbolic value embedding of amino acids at a given location are associated. That is, the symbolic keyword embedding and symbolic value embedding of amino acids in the first amino acid chain are associated with the amino acid tile that spans each other amino acid chain in the update block.
[0102] Geometric attention block 214 also generates "geometric query" embeddings, "geometric keyword" embeddings, and "geometric value" embeddings for each amino acid in each amino acid chain of the protein. Each amino acid's geometric query embedding, geometric keyword embedding, and geometric value embedding are 3D points, initially generated in the amino acid's local reference frame, and then rotated and translated to the protein's global reference frame using the structural parameters corresponding to that amino acid and the global structural parameters of the amino acid chain. For example, geometric attention block 214 can process the corresponding amino acid embeddings... To generate amino acids Geometric query embedding Geometric keyword embedding and geometric value embedding :
[0103]
[0104] in, This refers to a linear layer with independent learning parameter values, which will Projected onto 3-D points (superscript) This indicates that the quantity is a 3-D point. Indicates that it is composed of amino acids The rotation parameters specify the rotation matrix. Indicates that it is composed of amino acids The rotation matrix is specified by the global rotation parameters of the amino acid chain, where × denotes matrix multiplication. Indicates amino acids The position parameters, and Indicates amino acids The global positional parameters of the amino acid chain.
[0105] In order to update the amino acids in the first amino acid chain Amino acid embedding, geometric attention block 214 can generate attention weights ,in It is the total number of amino acids in a protein, and It is an amino acid With amino acids The attention weights between them are as follows:
[0106]
[0107] in, Indicates amino acids Symbol query embedding, Indicates amino acids Symbolic keyword embedding, express and Dimensions The parameters representing the learning process, Indicates amino acids Geometric query embedding, Indicates amino acids Geometric keyword embedding, yes Norm, Position in an interactively embedded 2D array The interaction embedding at 210, and It is the learned weight vector (or some other learning projection operation).
[0108] Typically, the interaction embeddings of amino acid pairs implicitly encode information related to the relationship between the amino acids in the pair, such as the distance between the amino acids in the pair. This is achieved by partially based on amino acid interactions. and amino acids Interactions and intercalation to determine amino acids and amino acids By enriching the attention weights between interactions, the folded neural network 200 utilizes information from the interaction embeddings to improve the accuracy of predicting folded structures.
[0109] Targeting the amino acids in the first amino acid chain Amino acid intercalation After generating attention weights, geometric attention block 214 uses the attention weights to update the amino acid embeddings. Specifically, geometric attention block 214 uses attention weights to generate "symbolic return" embeddings and "geometric return" embeddings, and then uses the symbolic return embeddings and geometric return embeddings to update the amino acid embeddings. Geometric attention block 214 is capable of generating amino acid embeddings. The symbol returns the embedded For example, as follows:
[0110]
[0111] in, This indicates attention weight (e.g., refer to the definition in equation (7)). Index all amino acids in the protein, and each Indicates amino acids The symbol value embedding. Geometric attention block 214 can generate amino acids. Geometry return embedding For example, as follows:
[0112]
[0113] Among them, geometric return embedding It is a 3D point. This indicates attention weight (e.g., refer to the definition in equation (7)). Indexing all amino acids in a protein, It is composed of amino acids in the first amino acid chain The rotation parameters specify the rotation matrix. It is a rotation matrix specified by the global rotation parameters of the first amino acid chain. It is an amino acid in the first amino acid chain. The position parameters, and These are the global rotation parameters for the first amino acid chain. It's understandable that the geometrically returned embedding is initially generated in the protein's global reference frame, then rotated and translated to the amino acid level. The local reference frame.
[0114] Geometric attention block 214 can return the embedding using the corresponding symbol. (e.g., generated according to equation (8)) and geometry return embedding (For example, according to equation (9)) to update the amino acids in the first amino acid chain. amino acid intercalation For example, as follows:
[0115]
[0116] in, It is an amino acid The updated amino acid intercalation, It is a norm, for example, norm, and The layer normalization operation is described, for example, as described below: J.1. Ba, J.R. Kiros, GE. Hinton, “Layer Normalization,” arXiv:1607.06450 (2016).
[0117] Amino acid embeddings 206 are updated using specific 3-D geometric embeddings, for example, as described in reference equations (4)-(6), such that geometric attention blocks 214 can derive 3-D geometry when updating amino acid embeddings. Furthermore, each update block updates the amino acid embeddings and structural parameters in a rotation- and translation-invariant manner over the entire protein structure. For example, applying the same global rotation and translation operations to the initial structural parameters provided to the folded neural network 200 will cause the folded neural network 200 to generate a predicted structure that is globally rotated and translated in the same way, but otherwise identical. Therefore, the global rotation and translation operations applied to the initial structural parameters do not affect the accuracy of the predicted protein structure generated by the folded neural network 200 from the initial structural parameters. The rotation- and translation invariance of the representation generated by the folded neural network 200 aids training, for example, because the folded neural network 200 automatically learns to generalize over all rotations and translations of the protein structure.
[0118] The updated amino acid embedding of the first amino acid chain can be further transformed by one or more additional neural network layers (e.g., linear neural network layers) in the geometric attention block 214 before being provided to the folded block 216.
[0119] In addition to updating the amino acid embeddings of the amino acids in the first amino acid chain, geometric attention block 214 also updates the global embeddings of the first amino acid chain.
[0120] To update the global embedding of the first amino acid chain, geometric attention block 214 processes the global embedding. To generate the corresponding symbol query embedding For example, as follows:
[0121]
[0122] in, This refers to a linear neural network layer. Geometric attention block 214 also handles global embeddings. To generate the corresponding geometric query embedding For example, as follows:
[0123]
[0124] in, Indicate Linear neural network layers projected onto 3-D points (superscript) This indicates that the quantity is a 3-D point. This represents the rotation matrix specified by the global rotation parameters of the first amino acid chain, and This represents the global position parameter of the first amino acid chain.
[0125] Geometric attention block 214 uses symbolic query embedding and geometric query embedding for global embedding of the first amino acid chain to generate attention weights. ,in It is the total number of amino acids in a protein, and It is an amino acid The attention weights between global embeddings and amino acid embeddings are as follows:
[0126]
[0127] in, Symbol query embedding representing the global embedding of the first amino acid chain. Indicates amino acids Symbolic keyword embedding, express and Dimensions The parameters representing the learning process, Geometric query embedding representing the global embedding of the first amino acid chain. Indicates amino acids The geometric keyword embedding, and yes Norm.
[0128] After generating the attention weights for the global embedding of the first amino acid chain, geometric attention block 214 uses the attention weights to update the global embedding. Specifically, geometric attention block 214 uses the attention weights to generate "symbol-return" embeddings and "geometric-return" embeddings, and then uses the symbol-return embeddings and geometric-return embeddings to update the global embedding. Geometric attention block 214 is capable of generating the symbol-return embedding of the global embedding. For example, as follows:
[0129]
[0130] in, This represents the attention weight between global embedding and amino acid embedding. Index all amino acids in the protein, and each Indicates amino acids The symbolic value embedding. Geometric attention block 214 can generate a globally embedded geometric return embedding. For example, as follows:
[0131]
[0132] Among them, geometric return embedding It is a 3D point. This represents the attention weight between global embedding and amino acid embedding. Indexing all amino acids in a protein, It is a rotation matrix specified by the global rotation parameters of the first amino acid chain, and It is the global rotation parameter of the first amino acid chain.
[0133] Geometric attention block 214 can return the embedding using the corresponding symbol. and geometry return embedding To update the global embedding of the first amino acid chain, for example, as follows:
[0134]
[0135] in, It is a global embedding of the updated first amino acid chain. It is a norm, for example, norm, and Presentation layer normalization operation.
[0136] After the geometric attention block 214 updates the amino acid embedding 206 of the amino acids in the first amino acid chain and the global embedding of the first amino acid chain, the fold block 216 uses the updated amino acid embedding 222 and the updated global embedding to update the current structural parameters 208 of the first amino acid chain.
[0137] For example, fold block 216 can fold amino acids in the first amino acid chain Current position parameters Updated to:
[0138]
[0139] in, It is the updated position parameter. Represents a linear neural network layer, and Indicates amino acids The updated amino acid embedding.
[0140] In another example, fold block 216 can store the current global position parameter of the first amino acid chain. Updated to:
[0141]
[0142] in, It is an updated global position parameter. Represents a linear neural network layer, and This represents the global embedding of the update of the first amino acid chain.
[0143] In another example, the amino acids in the first amino acid chain rotation parameters A rotation matrix can be specified, and fold block 216 can display the current rotation parameters. Updated to:
[0144]
[0145]
[0146] in, It is a three-dimensional vector. It is a linear neural network layer. It is an amino acid The updated amino acid intercalation, It indicates that it has a real part of 1 and an imaginary part. quaternions, and This represents the operation of transforming a quaternion into an equivalent 3×3 rotation matrix. Updating the rotation parameters using equations (12)-(13) ensures that the updated rotation parameters define a valid rotation matrix, such as an orthogonal matrix with a determinant of 1.
[0147] In another example, the global rotation parameter of the first amino acid chain A rotation matrix can be specified, and fold block 216 can display the current global rotation parameters. Updated to:
[0148]
[0149]
[0150] in, It is a three-dimensional vector. It is a linear neural network layer. It is an updated global embedding. It indicates that it has a real part of 1 and an imaginary part. quaternions, and This represents the operation of transforming a quaternion into an equivalent 3×3 rotation matrix.
[0151] As described above, the operations of geometric attention block 214 and folding block 216 jointly update the following two: (i) the structural parameters of the amino acids in the first amino acid chain, and (ii) the global structural parameters of the first amino acid chain. Updating the structural parameters of the amino acids in the first amino acid chain has the effect of updating the local structure of each amino acid chain in the protein complex. Updating the global structural parameters of the first amino acid chain has the effect of rotating and translating the positions of the amino acid chains in the protein structure relative to each other.
[0152] The final update block in the update block sequence of the folded neural network can generate the final structural parameters of the first amino acid chain, that is, the final structural parameters of each amino acid in the first amino acid chain and the final global structural parameters of the first amino acid chain.
[0153] The folded neural network 200 is able to use symmetry extension blocks to process (i) the final structural parameters of the first amino acid chain and (ii) the data that identifies symmetry groups 218 to generate protein structural parameters 226 that define the predicted symmetry structure 228 of the protein.
[0154] The folded neural network 200 may include any suitable number of update blocks, such as 5, 25, or 125 update blocks. Optionally, each update block of the folded neural network may share a single set of parameter values that are jointly updated during the training of the folded neural network. Sharing parameter values among update blocks 220 reduces the number of trainable parameters of the folded neural network and thus can facilitate efficient training of the folded neural network, for example, by stabilizing training and reducing the possibility of overfitting.
[0155] During training, the training engine can train the parameters of the protein structure prediction system, including the parameters of the folded neural network 200, based on the structural loss that evaluates the accuracy of protein structure parameter 226, as will be described in more detail below. In some embodiments, the training engine can further evaluate one or more auxiliary structural losses in update blocks 220 preceding the final update block. The auxiliary structural loss evaluation of the update block is the accuracy of the protein structure parameters defined by processing the updated structural parameters generated by the update block for the first amino acid chain using a symmetric expansion block.
[0156] Optionally, during training, the training engine can apply "stop gradient" operations to prevent gradient backpropagation through certain neural network parameters of each update block, such as the neural network parameters used to compute the updated rotation parameters (as described in equations (12)-(13)). Applying these stop gradient operations can improve the numerical stability of gradients computed during training.
[0157] Figure 3 Showing references Figure 2 The example protein structure prediction system 300 is described as a folded neural network 200. The protein structure prediction system 300 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components, and techniques described below are implemented.
[0158] System 300 is configured to generate a set of protein structure parameters 226 that define a predicted symmetrical protein structure 228 of a protein including multiple identical amino acid chains 304.
[0159] In order to generate structural parameters 226 that define the predicted protein structure 228, system 300 generates: (i) a multiple sequence alignment (MSA) representation 308 of the first amino acid chain in the protein, and (ii) a set of "pair" embeddings 306 of the first amino acid chain in the protein, as will be described in more detail below.
[0160] MSA 308 represents the MSA corresponding to the first amino acid chain in a protein. The MSA of an amino acid chain can be represented as an intercalation. Array (i.e., having) lines and (2-D array of embedded columns), where It is the number of MSA sequences in the MSA, and This refers to the number of amino acids in the first amino acid chain. Each line of the MSA representation can correspond to a corresponding MSA sequence. The system can initialize the MSA representation in any suitable manner. For example, system 300 can initialize each position in the MSA representation 308. The embedding at position i is initialized to the position defined in the MSA sequence. The unique heat vector for the identity of the amino acid at that location. Throughout the specification, "row" in MSA refers to the row that defines the 2D array of embedded MSA representations. Similarly, "column" in MSA refers to the column that defines the 2D array of embedded MSA representations.
[0161] This set of intercalations 306 includes corresponding intercalations for each amino acid pair in the first amino acid chain of the protein. An amino acid pair refers to an ordered tuple comprising the first and second amino acids in the first amino acid chain, i.e., such that a set of possible amino acid pairs in the first amino acid chain is given by the following formula:
[0162]
[0163] in, It refers to the number of amino acids in the first amino acid chain. Indexing the amino acids in the first amino acid chain, It is by The amino acids in the first amino acid chain of the index, and It is by The amino acids in the first amino acid chain of the index. This group of intercalations 306 can be represented as intercalations of 2-D. Arrays, for example, where the rows of a 2-D array are ∈ Indexing, the columns of a 2-D array are... Indexed, and the position in the 2-D array From amino acid pairs The embedded occupancy.
[0164] The system can initialize embeddings, for example, by applying the outer product mean operation to MSA representation 308, and recognize embedding 306 as the result of the outer product mean operation. The outer product mean operation defines a series of operations that, when applied to an embedding... When representing the MSA of the array, the embedded information is generated. Array, that is, where It is the number of amino acids in the first amino acid chain.
[0165] To compute the outer product average, the system generates a tensor. For example, given by the following formula:
[0166]
[0167] in, , ,in MSA represents the number of channels in each embedding. This is the number of rows in the MSA representation. It is applied to the location of " "The indexed rows and those by " "The MSA at the indexed column indicates the embedded channel" Linear operations (e.g., defined by matrix multiplication), and It is applied to the location of " "The indexed rows and those by " "The MSA at the indexed column indicates the embedded channel" Linear operations (e.g., defined by matrix multiplication). The result of the outer product mean is obtained by applying a tensor... of The dimensions are generated by flattening and linear projection. Optionally, the system can perform one or more layer normalization operations (e.g., as described in “Layer Normalization” by Jimmy Lei Ba et al., arXiv: 1607.06450) as part of the calculation of the outer product mean.
[0168] System 300 uses both MSA representation 308 and the pairwise embedding 306 to generate structural parameters 226 that define the predicted protein structure 228 because they have complementary properties. The structure of MSA representation 308 can be explicitly determined by the number of amino acid chains in the MSA. Therefore, MSA representation 308 may not be suitable for directly predicting protein structure because protein structure 228 does not have an explicit dependence on the number of amino acid chains in the MSA. Conversely, the pairwise embedding 306 characterizes the relationship between corresponding amino acid pairs in protein 302 and is expressed without explicit reference to the MSA, thus providing a convenient and efficient data representation for predicting protein structure 228.
[0169] System 300 uses embedding system 400 to process MSA representation 308 and pair embedding 306 to generate inputs for folded neural network 200, namely, interaction embedding 210, amino acid embedding 202 of the first amino acid chain, and symmetry group 218 of the protein.
[0170] refer to Figure 4 The example embedded system 400 is described in more detail.
[0171] System 300 generates network inputs for folded neural network 200 from interaction embeddings 210, amino acid embeddings 202, and symmetry groups 218, and uses folded neural network 200 to process the network inputs to generate structural parameters 226 that define the predicted protein structure.
[0172] The training engine can train the protein structure prediction system 300 end-to-end to optimize the objective function referred to in this paper as structure loss. The training engine can train the system 300 on a training dataset comprising multiple training samples. Each training sample can specify: (i) training input, which includes the MSA representation of the protein and its embeddings, and (ii) the target protein structure that should be generated by the system 300 by processing the training input. The target protein structure used to train the system 300 can be determined using experimental techniques such as X-ray crystallography or cryo-EM.
[0173] Structural loss can characterize the similarity between (i) the predicted protein structure generated by system 300 and (ii) the target protein structure that should be generated by the system.
[0174] For example, if the predicted structural parameters define the predicted position and rotation parameters for each amino acid in a protein, then the structural loss... It can be given by the following formula:
[0175]
[0176]
[0177]
[0178] in, It refers to the number of amino acids in a protein. Indicates amino acids Predicted location parameters, Indicates that it is composed of amino acids The predicted rotation parameters specify a 3×3 rotation matrix. It is an amino acid Target position parameters, Indicates that it is composed of amino acids The target rotation parameters specify a 3×3 rotation matrix. It is a constant. Refers to the predicted rotation parameters The inverse of a specified 3×3 rotation matrix, Refers to the target rotation parameters The inverse of the specified 3×3 rotation matrix, and This represents the ReLU operation.
[0179] Structural loss, as defined by equations (15)-(17), can be understood as the loss of each amino acid pair in a protein. Take an average. Item definition amino acid In amino acids The predicted spatial location in the predictive reference frame, and Define amino acids In amino acids The actual spatial location in the actual reference frame. These terms relate to amino acids. and The predictions and actual rotations are sensitive, thus carrying richer information than loss terms that are only sensitive to the predicted and actual distances between amino acids.
[0180] As part of evaluating structural loss, training determines which amino acid chain in the predicted protein structure corresponds to which amino acid chain in the benchmark protein structure. In some implementations, the training engine calculates the structural loss for every possible mapping from the amino acid chain in the predicted protein structure to the amino acid chain in the benchmark protein structure, and trains the system using the minimum of these structural losses.
[0181] Optimizing structural loss encourages System 300 to generate predicted protein structures that accurately approximate the real protein structure.
[0182] In addition to optimizing the structural loss, the training engine can also train System300 to optimize one or more auxiliary losses. Auxiliary losses can, for example, penalize predicted structures with features unlikely to occur in the natural world based on bond angles and / or bond lengths between atoms in amino acids within the predicted structure, or based on the proximity of atoms in different amino acids within the predicted structure.
[0183] The training engine can, for example, use stochastic gradient descent training techniques to train the structure prediction system 300 on training data through multiple training iterations.
[0184] Figure 4 An example embedded system 400 is shown. Embedded system 400 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components and technologies described below are implemented.
[0185] System 400 processes the MSA representation 308 and the pair of embeddings 306 using the values of a set of parameters of the embedding neural network 500 to update the MSA representation 308 and the pair of embeddings 306. That is, the embedding neural network 500 processes the MSA representation 308 and the pair of embeddings 306 to generate an updated MSA representation 402 and an updated pair of embeddings 404.
[0186] The embedding neural network 500 updates the MSA representation 308 and the pair embedding 306 by sharing information between the MSA representation 308 and the pair embedding 306. More specifically, the embedding neural network 500 alternates between updating the current MSA representation 308 based on the current pair embedding 306 and updating the current pair embedding 306 based on the current MSA representation 308.
[0187] refer to Figure 5 A more detailed description of an example architecture for embedding a neural network.
[0188] The embedding system 400 is capable of generating amino acid embeddings 202 for amino acids in the first amino acid chain from an updated MSA representation 402. For example, the updated MSA representation 402 can be represented as a 2-D array of embeddings, the number of columns being equal to the number of amino acids in the first amino acid chain, where each column is associated with a corresponding amino acid in the first amino acid chain. The embedding system 400 is capable of generating an initial amino acid embedding for that amino acid in the first amino acid chain by summing (or otherwise combining) the embeddings from the columns of the MSA representation 308 associated with each amino acid. As another example, the embedding system 400 is capable of generating an initial amino acid embedding for an amino acid in the first amino acid chain by extracting embeddings from the rows of the MSA representation 308 corresponding to the amino acid sequence of the first amino acid chain.
[0189] To generate the symmetric group 218, the embedding system 400 uses one or more neural network layers to process: (i) updated pair embeddings, and (ii) data identifying the number of amino acid chains in the protein (e.g., as a one-hot vector), to generate a probability distribution of a set of possible symmetric groups. The embedding system 400 is able to use the probability distribution to select the symmetric group 218, for example, by sampling possible symmetric groups according to the probability distribution, or by selecting the symmetric group with the highest probability under the probability distribution.
[0190] The embedding system uses the embedding extension system 400 to process updated pairs of embeddings 404 to generate interacting embeddings 210. (See reference) Figure 9 The example embedded extension system 900 is described in more detail.
[0191] Figure 5 An example architecture of an embedded neural network 500 is shown, which is configured to process MSA representation 308 and pair embedding 306 to generate updated MSA representation 402 and updated pair embedding 404.
[0192] The embedded neural network 500 includes an update block sequence 502-AN. Throughout the specification, "block" refers to a portion of a neural network, such as a subnetwork of a neural network that includes one or more neural network layers.
[0193] Each update block in the embedded neural network is configured to receive a block input including an MSA representation and an embedding, and to process the block input to generate a block output including an updated MSA representation and an updated embedding.
[0194] The embedding neural network 500 provides the MSA representation 308 and the pair embeddings 306 included in its network input to a first update block (i.e., in the update block sequence). The first update block processes the MSA representation 308 and the pair embeddings 306 to generate updated MSA representations and updated pair embeddings.
[0195] For each update block following the first update block, the embedding neural network 500 provides the update block with the MSA representation and pair embedding generated by the previous update block, and provides the next update block with the updated MSA representation and updated pair embedding generated by the update block.
[0196] The embedded neural network 500 gradually enriches the information content of the MSA representation 308 and the embedding 306 by repeatedly updating the MSA representation 308 and the embedding 306 by using the sequence of update block 502-AN.
[0197] The embedded neural network 500 can provide an updated MSA representation 402 and an updated pair embedding 404 generated by the final update block (i.e., in the update block sequence) as network outputs.
[0198] Figure 6 An example architecture of an update block 600 embedded in a neural network 500 is shown, i.e., as referenced Figure 5 As described.
[0199] Update block 600 receives block input including current MSA representation 602 and current pair embedding 604, and processes the block input to generate updated MSA representation 606 and updated pair embedding 608.
[0200] Update block 600 includes MSA update block 700 and update block 800.
[0201] MSA update block 700 uses the current pair embedding 604 to update the current MSA representation 602, and update block 800 uses the updated MSA representation 606 (i.e., generated by MSA update block 700) to update the current pair embedding 604.
[0202] Typically, MSA representations and pair embeddings encode complementary information. For example, an MSA representation encodes information about the correlation between the identities of amino acids at different positions in a set of evolutionarily related amino acid chains, and a pair embedding encodes information about the relationships between amino acids in a protein. MSA update block 700 uses the complementary information encoded in the pair embeddings to enrich the information content of the MSA representation, and pair update block 800 uses the complementary information encoded in the MSA representation to enrich the information content of the pair embeddings. As a result of this enrichment, the updated MSA representation and the updated pair embeddings encode information that is more relevant to predicting protein structure.
[0203] Update block 600 is described herein as first updating the current MSA representation 602 using the current pair embedding 604, and then updating the current pair embedding 604 using the updated MSA representation 606. This description should not be construed as limiting the update block to performing operations in this order; for example, the update block could first update the current pair embedding using the current MSA representation, and then subsequently update the current MSA representation using the updated pair embedding.
[0204] Update block 600 is described herein as including MSA update block 700 (i.e., which updates the current MSA representation) and pair update block 800 (i.e., which updates the current pair embedding). This description should not be construed as limiting update block 600 to including only one MSA update block or only one pair update block. For example, update block 600 can include multiple MSA update blocks that update the MSA representation multiple times before the MSA representation is provided to the pair update blocks for updating the current pair embedding. As another example, update block 600 can include multiple pair update blocks that update the pair embedding multiple times using the MSA representation.
[0205] MSA update block 700 and update block 800 can have any suitable architecture that enables them to perform the functions they describe.
[0206] In some implementations, the MSA update block 700, the update block 800, or both include one or more “self-attention” blocks. As used throughout this document, a self-attention block generally refers to a neural network block that updates a set of embeddings—that is, receives a set of embeddings and outputs updated embeddings. To update a given embedding, a self-attention block is able to determine a corresponding “attention weight” between the given embedding and each of one or more selected embeddings, and then uses (i) the attention weights and (ii) the selected embeddings to update the given embedding. For convenience, one could say that a self-attention block uses attention “on” the selected embeddings to update the given embedding.
[0207] For example, a self-attention block can receive a set of input embeddings. ,in It refers to the number of amino acids in the first amino acid chain, and is used for updating the intercalation. Self-attention blocks can determine attention weights. ,in express and The attention weights between them are:
[0208]
[0209]
[0210] in, and It is the parameter matrix learned. This indicates the soft-max normalization operation, and It is a constant. Using attention weights, the self-attention layer can embed... Updated to:
[0211]
[0212] in, This is the parameter matrix learned. Can be called about input embedding "Query embedding" Can be called about input embedding "Keyword embedding", and Can be called about input embedding (value embedding).
[0213] Parameter matrix ("Query Embedding Matrix") ("Keyword Embedding Matrix") and ( The "value embedding matrix" is a trainable parameter of the self-attention block. The parameters of any self-attention blocks included in MSA update block 700 and update block 800 can be understood as the parameters of update block 600, which can be used as a reference. Figure 3 The protein structure prediction system 300 described is trained as part of the end-to-end training. Typically, the (trained) parameters of the query embedding matrix, keyword embedding matrix, and value embedding matrix are different for different self-attention blocks, for example, such that the self-attention block included in the MSA update block 700 can have different query embedding matrices, keyword embedding matrices, and value embedding matrices, which have different parameters than the self-attention block included in the update block 800.
[0214] In some implementations, the MSA update block 700, the update block 800, or both include one or more self-attention blocks conditioned on embeddings, i.e., one or more self-attention blocks that implement self-attention operations conditioned on embeddings. To make the self-attention operation conditioned on embeddings, the self-attention blocks are able to process embeddings to generate a corresponding “attention bias” for each attention weight. For example, in addition to determining the attention weights according to equations (18)-(19)... In addition, self-attention blocks can also generate a corresponding set of attention biases. ,in, express and Attention bias between them. Self-attention blocks can apply the learned parameter matrix to the embedding. That is, for those caused by Indexed amino acid pairs in proteins to generate attentional biases .
[0215] Self-attention blocks can determine a set of "biased attention weights" by, for example, summing (or otherwise combining) attention weights and attention biases. ,in express and Attention weights with bias between them. For example, self-attention blocks can embed... and Attention weights of bias between Determined as:
[0216]
[0217] in, yes and The attention weights between them, and yes and Attentional bias between inputs. Self-attention blocks can use the biased attentional weights to update each input embedding. ,For example:
[0218]
[0219] in, It is the parameter matrix for learning.
[0220] Typically, embeddings encode information characterizing the structure of a protein and the relationships between amino acid pairs within that structure. Applying a self-attention operation conditioned on embeddings to a set of input embeddings allows the input embeddings to be updated in a manner notified by the protein structural information encoded in the embeddings. Update blocks in embedding neural networks can use self-attention blocks conditioned on embeddings to update and enrich both the MSA representation and the embeddings themselves.
[0221] Optionally, the self-attention block can have multiple "heads," each generating a corresponding updated embedding for each input embedding, i.e., such that each input embedding is associated with multiple updated embeddings. For example, each head can be a parameter matrix described by reference equations (18)-(21). , Different values are used to generate updated embeddings. Self-attention blocks with multiple heads can implement a "gating" operation to combine updated embeddings generated from the heads used for the input embeddings; that is, to generate a single updated embedding corresponding to each input embedding. For example, a self-attention block can use one or more neural network layers (e.g., fully connected neural network layers) to process the input embeddings to generate corresponding gating values for each head. The self-attention block can then combine the updated embeddings corresponding to the input embeddings based on the gating values. For example, a self-attention block can combine the input embeddings... The updated embedding is generated as follows:
[0222]
[0223] in, Index the header. It is the head The gating value, and It is from the head For input embedding The generated updated embedding.
[0224] refer to Figure 7 This describes an example architecture for updating block 700 using MSA with self-attention blocks conditioned on embedding. (Reference) Figure 7 The example MSA update block described updates the current MSA representation based on the current pair embedding by processing the rows of the current MSA representation using a self-attention block conditioned on the current pair embedding.
[0225] refer to Figure 8 This describes an example architecture for updating block 800 using self-attention blocks conditioned on embedding. (Reference) Figure 8 The example described describes how the update block computes the outer product mean of the updated MSA representation, adds the result of the outer product mean to the current pair embedding, and processes the current pair embedding using a self-attention block conditioned on the current pair embedding, updating the current pair embedding based on the updated MSA representation.
[0226] Figure 7 An example architecture of MSA update block 700 is shown. MSA update block 700 is configured to receive the current MSA representation 302 and update the current MSA representation 602 (at least in part) based on the current pair embedding.
[0227] To update the current MSA representation 602, the MSA update block 700 uses a self-attention operation conditioned on the current pair of embeddings (i.e., a "line-by-line" self-attention operation) to update the embeddings in each line of the current MSA representation. More specifically, the MSA update block 700 provides the embeddings in each line of the current MSA representation 602 to the "line-by-line" self-attention block 702 conditioned on the current pair of embeddings, for example, as referenced... Figure 6The method described above generates updated embeddings for each row of the current MSA representation 602. Optionally, the MSA update block can add the input of the line-by-line self-attention block 702 to the output of the line-by-line self-attention block 702. By conditioned the line-by-line self-attention block 702 on the current pair of embeddings, the MSA update block 600 can use information from the current pair of embeddings to enrich the current MSA representation 302.
[0228] Then, the MSA update block uses a self-attention operation that is not conditioned on the current pair of embeddings (i.e., a “column-by-column” self-attention operation) to update the embeddings in each column of the current MSA representation. More specifically, MSA update block 700 provides the embeddings in each column of the current MSA representation 602 to a “column-by-column” self-attention block 704 that is not conditioned on the current pair of embeddings to generate updated embeddings for each column of the current MSA representation 602. As a result of not being conditioned on the current pair of embeddings, column-by-column self-attention block 704 generates updated embeddings for each column of the current MSA representation using attention weights (e.g., as described in reference equations (18)-(19)) instead of biased attention weights (e.g., as described in reference equation (21)). Optionally, the MSA update block is able to add the input of column-by-column self-attention block 704 to the output of column-by-column self-attention block 704.
[0229] Then, the MSA update block processes the current MSA representation 602 using, for example, a transformation block that applies one or more fully connected neural network layers to the current MSA representation 602. Optionally, the MSA update block 700 can add the input of the transformation block 706 to the output of the transformation block 706.
[0230] The MSA update block can output an updated MSA representation 606 resulting from the operations performed by the row-by-row self-attention block 702, the column-by-column self-attention block 704, and the transformation block 706.
[0231] Figure 8 An example architecture for update block 800 is shown. Update block 800 is configured to receive the current pair embedding 604 and (at least in part) update the current pair embedding 604 based on the updated MSA representation 602.
[0232] To update the current pair embedding 604, the update block 800 applies the outer product mean operation 802 to the updated MSA representation 602 and adds the result of the outer product mean operation 802 to the current pair embedding 404.
[0233] Typically, the updated MSA representation 602 encodes information about the correlation between the identities of amino acids at different positions in a set of evolutionarily related amino acid chains. The information encoded in the updated MSA representation 602 is related to the predicted structure of the protein, and the updated block 800 can enhance the information content of the current pair embedding by incorporating the information encoded in the updated MSA representation into the current pair embedding (i.e., by means of the outer product mean 802).
[0234] After updating the current pair embedding 604 using the updated MSA representation (i.e., via the outer product mean 802), a self-attention operation conditioned on the current pair embedding (i.e., a "row-by-row" self-attention operation) is applied to update block 800 to update the current pair embedding in each row of the current pair embedding arrangement. Array. More specifically, update block 800 provides each row of the current pair of embeddings to a "line-by-line" self-attention block 804, which is also conditioned on the current pair of embeddings, for example, as referenced. Figure 6 The method described above generates a pair of embeddings for updating each row. Optionally, the pair update block can add the input of the line-by-line self-attention block 804 to the output of the line-by-line self-attention block 804.
[0235] Then, update block 800 using a self-attention operation also conditioned on the current pair of embeddings (i.e., a "column-by-column" self-attention operation). The current pair embeddings in each column of the array. More specifically, the update block 800 provides each column of the current pair embeddings to the column-by-column self-attention block 506, which also conditions the current pair embeddings to generate updated pair embeddings for each column. Optionally, the update block can add the input of the column-by-column self-attention block 806 to the output of the column-by-column self-attention block 806.
[0236] Then, the update block 800 is processed using, for example, applying one or more fully connected neural network layers to the transformation block of the current pair of embeddings. Optionally, the update block 800 can add the input of the transformation block 808 to the output of the transformation block 808.
[0237] The update block can output the updated pair embedding 608 resulting from the operations performed by the row-by-row self-attention block 804, the column-by-column self-attention block 806, and the transformation block 808.
[0238] Figure 9 An example embedded extension system 900 is shown, for example, which can be included in a reference. Figure 3 The protein structure prediction system 300 is described. The embedded extension system 900 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components, and techniques described below are implemented.
[0239] System 900 is configured to process a set of pairs of intercalations 404 of the first amino acid chain in a protein to generate a set of interacting intercalations 210.
[0240] To generate the interaction intercalation 210, system 900 generates a corresponding copy of the intercalation pair 404 for each amino acid chain in the protein. For example, system 900 can generate a copy 904-AN of the intercalation pair, where each pair of intercalations 904-AN corresponds to a corresponding amino acid chain in the protein.
[0241] System 900 generates corresponding "symmetry group relative position encoding" data 902 for each amino acid chain in the protein. The symmetry group relative position encoding data 902 defines the corresponding encoding vector for each amino acid chain of the protein, that is, distinguishes each amino acid chain from each other amino acid chain. For example, for the C4 symmetry group, the encoding vector for the first amino acid chain can be [1,0,0,0], the encoding vector for the second amino acid chain can be [0,1,0,0], the encoding vector for the third amino acid chain can be [0,0,1,0], and the encoding vector for the fourth amino acid chain can be [0,0,0,1].
[0242] For each amino acid chain in a protein, system 900 uses an encoding neural network to process the encoding vector defined by the relative position encoding 902 of the symmetric groups of the amino acid chain to generate a relative position embedding with the same number of channels as the pair embeddings. System 900 is then able to add (or otherwise combine) the relative position embeddings generated for the amino acid chain to each pair embedding of the amino acid chain.
[0243] After combining the symmetry group relative position encoding data 902 with the amino acid chain pairings 904-AN, system 900 updates each pairing 904-AN by processing each pairing 904-AN using one or more pairing update blocks 906-AN. Each pairing update block can update the pairings using row-wise self-attention blocks and column-wise self-attention blocks, for example, as referenced. Figure 8 As described.
[0244] After updating the pair embedding 904-AN using the update block 906-AN, system 900 can concatenate each pair embedding 904-AN row by row to obtain the interacting embedding 210. More specifically, each pair embedding 904-AN can be represented as a pair of embeddings... Array, in which It refers to the number of amino acids in each amino acid chain, and the system 900 can cascade these intercalation arrays row by row to obtain interacting intercalations. Array, in which It refers to the number of amino acid chains in a protein.
[0245] Optionally, system 900 can update the array of interaction embeddings 210 by processing the array of interaction embeddings using one or more row-by-row self-attention blocks and one or more column-by-column self-attention blocks.
[0246] Figure 10 This is a flowchart of an exemplary process 1000 for predicting the structure of a protein comprising multiple amino acid chains. For convenience, process 1000 will be described as being performed by a system of one or more computers located in one or more locations. For example, a protein structure prediction system appropriately programmed according to this specification (e.g., Figure 3 The protein structure prediction system 300 can execute process 1000.
[0247] The system obtains the initial structural parameters (1002) of the first amino acid chain in a protein. The structural parameters of the amino acid chain in a protein define the predicted three-dimensional spatial positions of the amino acids in the amino acid chain within the protein's structure.
[0248] The system obtains data that identifies symmetric groups, in which proteins are predicted to fold into structures symmetrical with respect to the symmetric groups (1004).
[0249] The system uses a folded neural network to process input data including the initial structural parameters of the first amino acid chain and the identification of symmetric groups.
[0250] The folded neural network consists of a sequence of update blocks. Steps 1006 to 1010, described below, are performed by each update block in the folded neural network.
[0251] The update block receives the current structural parameters of the first amino acid chain and data on the identification of symmetric groups (1006).
[0252] The update block applies the symmetry extension transformation to the current structural parameters of the first amino acid chain to generate the corresponding current structural parameters for each of the other amino acid chains in the protein, thus defining the current predicted structure of the protein symmetrical with respect to the symmetry group (1008).
[0253] The update block processes the current structural parameters of the amino acid chains in the protein based on the values of the update block parameters to update the current structural parameters (1010) of the first amino acid chain.
[0254] The system uses the structural parameters of the first amino acid chain generated from the final update block to generate the final predicted structure of the protein that is symmetrical with respect to the symmetric groups (1012).
[0255] This specification uses the term "configuration" in conjunction with system and computer program components. For a system of one or more computers configured to perform a specific operation or action, this means that software, firmware, hardware, or a combination thereof have been installed on the system, which, in operation, causes the system to perform the operation or action. For one or more computer programs configured to perform a specific operation or action, this means that one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0256] Embodiments of the subject matter and functional operation described in this specification can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their equivalents), or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium, for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof. Alternatively or additionally, program instructions can be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0257] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The device may also be or include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0258] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not necessarily, correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data, for example, as one or more scripts stored in a markup language document, as a single file dedicated to the program in question, or as multiple harmonized files, for example, as a file storing one or more modules, subroutines, or code portions. A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a data communication network.
[0259] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in others, multiple engines can be installed and run on the same one or more computers.
[0260] The processes and logic flows described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGA or ASIC), or by a combination of special-purpose logic circuitry and one or more programmable computers.
[0261] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be supplemented by or incorporated into special-purpose logic circuitry. Typically, a computer will also include one or more mass storage devices (such as disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to, or both. However, a computer does not need to have such devices. Furthermore, a computer can be embedded in another device, to name just a few, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.
[0262] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0263] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser in response to a request received from a web browser on the user's device. Furthermore, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user as a response.
[0264] The data processing apparatus used to implement machine learning models may also include, for example, dedicated hardware accelerator units for handling the common and computationally intensive parts of machine learning training or production, namely inference and workloads.
[0265] It is possible to implement and deploy machine learning models using machine learning frameworks such as TensorFlow, Microsoft's Cognitive Toolkit, Apache's Singa, or Apache's MXNet.
[0266] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface, web browser, or app that a user can interact with through an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0267] A computing system can include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship is established by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., HTML pages) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from it. It is possible to receive data generated at the user device, such as the result of user interaction, from the device at the server.
[0268] While this specification contains numerous details of specific implementation, these should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, one or more features from a claimed combination can be removed from the combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0269] Similarly, although operations are depicted in the accompanying drawings and described in a specific order in the claims, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all of the shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0270] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in different orders and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for predicting the structure of a protein comprising multiple amino acid chains, performed by one or more data processing devices, the method comprising: Obtain the initial structural parameters of the first amino acid chain in a protein, wherein the structural parameters of the first amino acid chain in a protein define the predicted three-dimensional (3D) spatial location of the amino acids in the first amino acid chain in the structure of the protein. Data is obtained to identify symmetry groups, wherein proteins are predicted to fold into structures symmetrical with respect to the symmetry groups; and The folded neural network is used to process inputs including initial structural parameters of the first amino acid chain and data identifying symmetry groups to generate an output that defines the final predicted structure of the protein. The folded neural network is configured to generate the final predicted structure of the protein by forcing a constraint that the final predicted structure of the protein needs to be symmetric with respect to the symmetry groups.
2. The method of claim 1, wherein, Folded neural networks involve updating block sequences. Each update block in the update block sequence has multiple update block parameters and performs the following operations: Receive the current structural parameters of the first amino acid chain and the data for recognizing symmetric groups; The symmetry extension transformation is applied to the current structural parameters of the first amino acid chain to generate the corresponding current structural parameters for each other amino acid chain in the protein, defining the current predicted structure of the protein symmetrical with respect to the symmetry group; and The current structural parameters of the amino acid chains in the protein are processed based on the values of the update block parameters to update the current structural parameters of the first amino acid chain.
3. The method of claim 2, wherein, The structural parameters of the first amino acid chain in a protein include global structural parameters, which define the 3D spatial position and orientation of the first amino acid chain in the protein's reference frame.
4. The method of claim 3, wherein, Applying the symmetric expansion transformation to the current structural parameters of the first amino acid chain to generate the corresponding current structural parameters for each of the other amino acid chains in the protein includes, for each of the other amino acid chains in the protein: Global structural parameters for other amino acid chains are generated by applying predefined transformations to the global structural parameters of the first amino acid chain, wherein the predefined transformations depend on: (i) the number of amino acid chains in the protein, and (ii) symmetric groups; and Determine the amino acid structure parameters of other amino acid chains that match the amino acid structure parameters of the first amino acid chain.
5. The method of claim 1, wherein, The structural parameters of the first amino acid chain in a protein include the corresponding amino acid structural parameters of each amino acid in the first amino acid chain, wherein the amino acid structural parameters of each amino acid define the 3D spatial position and orientation of that amino acid in the reference frame of the first amino acid chain.
6. The method of claim 1, wherein, Symmetrical groups are cyclic symmetrical groups, dihedral symmetrical groups, or cubic symmetrical groups.
7. The method of claim 2, wherein, The inputs processed by the folded neural network also include: (i) the corresponding initial amino acid embedding for each amino acid in the first amino acid chain, and (ii) the initial global embedding for the first amino acid chain.
8. The method of claim 7, wherein, The operations performed by each update block also include receiving the corresponding current amino acid embedding for each amino acid in the first amino acid chain and the current global embedding for the first amino acid chain; and The process of processing the current structural parameters of the amino acid chains in a protein to update the current structural parameters of the first amino acid chain includes: Update the current amino acid embedding and current global embedding of the first amino acid chain based on the current structural parameters of the amino acid chains in the protein; and The current structural parameters of the first amino acid chain are updated based on the updated amino acid embedding and the updated global embedding.
9. The method of claim 8, wherein, Updating the current amino acid embedding of the first amino acid chain based on the current structural parameters of the amino acid chains in the protein includes: For each other amino acid chain in the protein, based on the current amino acid embedding of the corresponding amino acid in the first amino acid chain, determine the corresponding current amino acid embedding of each amino acid in the other amino acid chains; and The attention of the current amino acid embedding of the amino acid chain is used to update the current amino acid embedding of the first amino acid chain and the current global embedding, wherein the attention of the current amino acid embedding of the amino acid chain is conditioned on the current structural parameters of the amino acid chain.
10. The method of claim 9, wherein, Updating the current global embedding of the first amino acid chain using attention to the current amino acid embedding of the amino acid chain includes: For each amino acid in each amino acid chain, the corresponding attention weight between the current global embedding of the first amino acid chain and the current amino acid embedding of that amino acid is determined at least in part based on: (i) the global structural parameters of the first amino acid chain, and (ii) the amino acid structural parameters of that amino acid and the global structural parameters of the amino acid chain; and The current global embedding of the first amino acid chain is updated based on (i) attention weights and (ii) the current amino acid embedding of the amino acid chain.
11. The method of claim 10, wherein, For each amino acid in each amino acid chain, determining the attention weight between the current global embedding of the first amino acid chain and the current amino acid embedding of that amino acid includes: Generate a geometric query embedding corresponding to the current global embedding of the first amino acid chain, including: The current global embedding of the first amino acid chain is processed using one or more neural network layers to generate a 3D embedding; The 3D embedding is rotated and translated into the protein's reference frame using the global structural parameters of the first amino acid chain. Generate the geometric keyword embedding corresponding to the amino acid, including: The current amino acid embedding is processed using one or more neural network layers to generate a 3D embedding; and Using the amino acid structure parameters and the global structure parameters of the amino acid chain, 3D embeddings are rotated and translated into a protein reference frame; and Attention weights are determined based on the spatial distance between (i) the geometric query embedding corresponding to the current global embedding of the first amino acid chain and (ii) the geometric keyword embedding corresponding to that amino acid.
12. The method of any one of claims 8-11, wherein, Updating the current structural parameters of the first amino acid chain based on the updated amino acid embedding and the updated global embedding includes: For each amino acid in the first amino acid chain, the amino acid structure parameter of that amino acid is updated based on the updated amino acid embedding; and The global structural parameters of the first amino acid chain are updated by updating the global embedding based on the first amino acid chain.
13. A system comprising: One or more computers; and One or more storage devices are communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the corresponding methods according to any one of claims 1-12.
14. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause one or more computers to perform the operation of the corresponding method according to any one of claims 1 to 12.
15. A method of obtaining a ligand, wherein, The ligand is a ligand for a drug or an industrial enzyme, and the method includes: Perform the method according to any one of claims 1-12 to determine the predicted structure of the target protein; Evaluate the interaction between one or more candidate ligands and the predicted structure of the target protein; and One or more candidate ligands are selected as ligands based on the evaluation results.
16. The method of claim 15, wherein, Target proteins include receptors or enzymes, and the ligands therein are agonists or antagonists of the receptors or enzymes.
17. The method of claim 15, wherein, The ligand is a drug, and the method includes: Perform the method according to any one of claims 1-12 to determine the predicted structure of each of the plurality of target proteins; Evaluate the interaction between one or more candidate ligands and the predicted structure of each target protein; and Select one or more candidate ligands as ligands to i) obtain ligands that interact with each target protein, or ii) obtain ligands that interact with only one target protein.
18. A method of obtaining a polypeptide ligand, wherein, The ligand is a ligand for a drug or an industrial enzyme, and the method includes: For each of one or more candidate polypeptide ligands, perform the method of any one of claims 1-12 to determine the predicted structure of the candidate polypeptide ligand; Obtain the target protein structure; Evaluate the interaction between the predicted structure of each of one or more candidate polypeptide ligands and the target protein structure; and Depending on the evaluation results, one of one or more candidate peptide ligands is selected as the peptide ligand.
19. The method of claim 18, wherein, The target protein includes a receptor or enzyme, and the ligand is an agonist or antagonist of the receptor or enzyme; or the polypeptide ligand includes an antibody, and the target protein includes an antibody target, particularly a viral or cancer cell protein, and the antibody binds to the antibody target to provide a therapeutic effect.
20. A method for obtaining a diagnostic antibody biomarker for a disease, the method comprising: For each of one or more candidate antibodies, perform the method of any one of claims 1-12 to determine the predicted structure of the candidate antibody; Obtain the target protein structure; Evaluate the interaction between the predicted structure of each of one or more candidate antibodies and the target protein structure; as well as Depending on the evaluation results, one of one or more candidate antibodies may be selected as a diagnostic antibody biomarker.
21. The method of claim 15, wherein, Evaluating the interaction of one of the candidate ligands involves determining the interaction score of the candidate ligand, where the interaction score comprises a measurement of the interaction between the candidate ligand and the target protein.
22. The method of claim 21, further comprising synthesizing a ligand.
23. The method of claim 15 further includes testing the bioactivity of the ligand in vitro and in vivo.
24. A method for identifying the presence of a protein misfolding disease, comprising: Perform the method according to any one of claims 1-12 to determine the predicted structure of the protein; To obtain the structure of a protein version obtained from a human or animal body; Compare the predicted structure of a protein with the structure of a protein version obtained from a human or animal. as well as The presence of protein misfolding diseases can be identified based on the results of the comparison.