Methods, systems, and media for training a protein structure prediction neural network using simplified multiple sequence alignment
Patent Information
- Application Number
- CN202180067160.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-29
- Filing Date
- 2021-08-12
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2041-08-12
AI Technical Summary
[0026] Identifying the underlying protein structures can be expensive and time-consuming, and the underlying structures of many proteins may be unknown. The training system described in this specification enables a structure prediction neural network to be trained to effectively predict the structures of a wide variety of proteins even when the underlying structures of many proteins are not available.
Smart Images

Figure CN116635940B_ABST
Abstract
Description
Technical Field
[0001] This manual relates to training neural networks to predict protein structures. Background Technology
[0002] A protein is defined by one or more amino acid sequences. Amino acids are organic compounds that include amino and carboxyl functional groups, as well as side chains (i.e., atomic groups) specific to each amino acid. Protein folding refers to the physical process by which an amino acid sequence folds into a three-dimensional (3-D) conformation. The structure of a protein defines the 3-D conformation of the atoms in its amino acid sequence after folding. When in a sequence linked by peptide bonds, an amino acid can be referred to as an amino acid residue.
[0003] Machine learning models can be used for prediction. A machine learning model takes input and generates an output based on that input, such as a predicted output. Some machine learning models are parametric models and generate an output based on the received input and the model's parameter values. Some machine learning models are deep models that employ multiple layers to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, each of which applies a nonlinear transformation to the received input to generate an output. Summary of the Invention
[0004] This specification describes a training system implemented as a computer program on one or more computers in one or more locations for training a structure prediction neural network capable of predicting protein structures.
[0005] As used throughout this specification, the term "protein" may be understood to refer to any biomolecule specified by one or more amino acid sequences. For example, the term protein may be understood to refer to a protein domain (e.g., a portion of an amino acid sequence that can undergo protein folding almost independently of the rest of the amino acid sequence) or a protein complex (e.g., specified by multiple associated amino acid sequences).
[0006] The methods and systems described herein can be used to train a structure prediction neural network to obtain ligands, such as drugs or ligands for industrial enzymes. For example, a method for obtaining a ligand may include obtaining a target amino acid sequence, specifically the amino acid sequence of a target protein, and using a structure prediction neural network to process the input based on the target amino acid sequence to determine the (tertiary) structure of the target protein, i.e., predict the protein structure. The method may then include evaluating the interaction between one or more candidate ligands and the structure of the target protein. The method may also include selecting one or more candidate ligands as the ligand based on the evaluation results of the interaction.
[0007] In some embodiments, evaluating the interaction may include assessing the binding of a candidate ligand to the structure of a target protein. For example, evaluating the interaction may include identifying ligands that bind with sufficient affinity for the biological effect. In some other embodiments, evaluating the interaction may include assessing the association between a candidate ligand and the structure of a target protein, which has an impact on the function of the target protein (e.g., an enzyme). The assessment may include assessing the affinity between the candidate ligand and the structure of the target protein, or assessing the selectivity of the interaction.
[0008] Candidate ligands can be derived from a candidate ligand database, and / or can be derived by modifying ligands in the candidate ligand database, for example by modifying the structure or amino acid sequence of the candidate ligand, and / or can be derived by stepwise or iterative assembly / optimization of the candidate ligand.
[0009] Assessing the interaction between a candidate ligand and the target protein structure can be performed using a computer-aided method. This method displays a graphical model of the candidate ligand and target protein structures for user manipulation, and / or the assessment can be performed partially or fully automatically, for example, using standard molecular (protein-ligand) docking software. In some embodiments, the assessment may include determining the interaction score of the candidate ligand, where the interaction score is a measure of the interaction between the candidate ligand and the target protein. The interaction score may depend on the strength and / or specificity of the interaction; for example, the score depends on the binding free energy. Candidate ligands can be selected based on their scores.
[0010] In some embodiments, the target protein includes a receptor or enzyme, and the ligand is an agonist or antagonist of the receptor or enzyme. In some embodiments, the method can be used to identify the structure of a cell surface marker. This can then be used to identify ligands that bind to the cell surface marker, such as antibodies or markers such as fluorescent labels. This can be used to identify and / or treat cancer cells.
[0011] In some embodiments, candidate ligands may include small molecule ligands, such as organic compounds with a molecular weight <900 Daltons. In other embodiments, candidate ligands may include peptide ligands, i.e., peptide ligands defined by an amino acid sequence.
[0012] In some cases, a structure prediction neural network trained using the techniques described herein can be used to determine the structure of candidate peptide ligands (e.g., drugs or ligands for industrial enzymes). The interaction of that structure with the target protein structure can then be evaluated; the target protein structure may have been determined using a structure prediction neural network or conventional physical investigation techniques such as X-ray crystallography and / or magnetic resonance imaging.
[0013] Therefore, in another aspect, a method is provided using a structure prediction neural network trained with the techniques described herein to obtain peptide ligands (e.g., molecules or their sequences). This method may include obtaining the amino acid sequences of one or more candidate peptide ligands. The method may also include using the structure prediction neural network to determine the (tertiary) structure of the candidate peptide ligands. The method may further include obtaining the target protein structure via computer simulation (in silico) and / or by physical investigation, and evaluating the interaction between the structure of each of the one or more candidate peptide ligands and the target protein structure. The method may also include selecting one or more of the candidate peptide ligands as the peptide ligand based on the evaluation results.
[0014] As previously described, evaluating interactions may include assessing the binding of candidate peptide ligands to the structure of a target protein, such as identifying ligands that bind with sufficient affinity for biological effects, and / or assessing the association between candidate peptide ligands and the structure of a target protein that influences the function of the target protein (e.g., an enzyme), and / or assessing the affinity between candidate peptide ligands and the structure of a target protein, or assessing the selectivity of the interaction. In some embodiments, the peptide ligand may be an aptamer.
[0015] Implementations of this method may also include the synthesis (i.e., fabrication) of small molecule or peptide ligands. The ligands can be synthesized using any conventional chemical technique and / or may already be available, for example, from a compound library or possibly synthesized using combinatorial chemistry. Synthesis can be manual, semi-automatic, or fully automated. The synthesized small molecule or peptide ligands may be pharmaceuticals.
[0016] This method may also include testing the bioactivity of the ligand in vitro and / or in vivo. For example, the ADME (absorption, distribution, metabolism, excretion) and / or toxicological properties of the ligand can be tested to screen out unsuitable ligands. Testing may include, for example, contacting the candidate small molecule or peptide ligand with the target protein and measuring changes in protein expression or activity.
[0017] In some embodiments, candidate (peptide) ligands may include: isolated antibodies, fragments of isolated antibodies, monovariable domain antibodies, bispecific or multispecific antibodies, multivalent antibodies, bivariate domain antibodies, immunoconjugates, fibronectin molecules, adhesion proteins, DARPin, antibodies, affinities, antitransporters, affinity proteins, protein epitope mimics, or combinations thereof. Candidate (peptide) ligands may include antibodies with a mutated or chemically modified amino acid Fc region, for example, which, compared to a wild-type Fc region, prevents or reduces ADCC (antibody-dependent cytotoxicity) activity and / or increases half-life. Therefore, in some embodiments, this method is used to obtain peptide ligands comprising antibodies.
[0018] Misfolded proteins are associated with a variety of diseases. Therefore, in another aspect, a method is provided using a structure prediction neural network trained with the techniques described herein to identify the presence of protein misfolding diseases. This method may include obtaining the amino acid sequence of a protein and using the structure prediction neural network to determine the protein's structure. The method may also include obtaining the structure of a version of the protein obtained from a human or animal body, for example, through conventional (physical) methods such as X-ray crystallography, NMR spectroscopy, or electron microscopy. The method may then include comparing the protein's structure with the structure of the version obtained from the body and identifying the presence of a protein misfolding disease based on the comparison result. That is, misfolding of a protein version from the body can be determined by comparing it with a structure determined via computer simulation.
[0019] In some other respects, computer-implemented methods, as described above or herein, can be used to identify active / binding / blocking sites on a target protein from its amino acid sequence.
[0020] According to another aspect, a system is provided comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to operate to implement the techniques described herein. The system may include a subsystem for producing proteins obtained using the techniques, such as a robotic protein synthesis subsystem.
[0021] Specific embodiments of the subject matter described in this specification may be implemented in order to achieve one or more of the following advantages.
[0022] This specification describes a training system that can train a structure prediction neural network using both "paired" and "unpaired" training examples. Each paired training example includes a multiple sequence alignment (MSA) of a protein and a ground truth (e.g., an actual) protein structure, and the training system trains the structure prediction neural network to process the MSA to generate a predicted protein structure that matches the ground truth protein structure. Each unpaired training example includes the MSA of a protein, but the ground truth structure of the protein may be unknown. To train the structure prediction neural network based on the unpaired training examples, the training system generates a target protein structure for each unpaired training example by processing the MSA from the unpaired training examples using the structure prediction neural network. The training system then trains the structure prediction neural network to process a "simplified" MSA (i.e., where some data in the MSA has been removed or masked) for each unpaired training example to generate a predicted protein structure that matches the corresponding target protein structure.
[0023] By training a structure prediction neural network using unpaired training examples, the training system can improve the performance (e.g., prediction accuracy) of the network by reducing the likelihood of it overfitting to paired training examples. The network can, for example, "overfit" paired training examples by learning to predict the underlying real-world protein structure specified by the paired examples based on irrelevant variations in the MSA (Material Scale Analysis), rather than on implicit reasoning based on inferred biochemical principles. Furthermore, the number of available unpaired training examples can be significantly greater than the number of available paired training examples; therefore, training the network on unpaired examples allows it to learn to effectively predict the structures of a wider variety of proteins.
[0024] This specification describes a training system for training a "student" structure prediction neural network that can predict the structure of a protein by processing input including a representation of the protein's amino acid sequence but excluding the protein's MSA (Massively Sequence Aspect). To increase the amount of training data available, rather than just paired training examples (i.e., where the underlying real-world protein structure is known), the training system trains a "teacher" structure prediction neural network that can accurately predict the structure of a protein by processing input including the protein's MSA. The training system uses the teacher structure prediction neural network to generate target protein structures by processing input including MSAs from unpaired training examples, thereby generating a predicted target for each unpaired training example. The training system then trains the student structure prediction neural network to generate a predicted protein structure that matches the target protein structure of the training example for each unpaired training example, without processing the protein MSA.
[0025] By using a teacher structure prediction neural network to generate prediction targets, this training system can significantly increase the amount of training data available for training a student structure prediction neural network, thereby enabling the student structure prediction neural network to be trained to achieve higher prediction accuracy. After training, the student structure prediction neural network can be used to predict the structure of any protein, regardless of whether the protein's MSA is available, thus making it widely applicable to any task requiring protein structure prediction.
[0026] Identifying the underlying protein structures can be expensive and time-consuming, and the underlying structures of many proteins may be unknown. The training system described in this specification enables a structure prediction neural network to be trained to effectively predict the structures of a wide variety of proteins even when the underlying structures of many proteins are not available.
[0027] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the specification, drawings, and claims. Attached Figure Description
[0028] Figure 1A -B describes a training system for training a structure prediction neural network that can predict the structure of a protein by processing inputs including the protein's MSA.
[0029] Figure 2 A training system for training a structure prediction neural network is described, which can predict the structure of a protein without processing the protein's MSA.
[0030] Figure 3 This is a diagram of unfolded and folded proteins.
[0031] Figure 4 This is a flowchart of an example process for training a structure prediction neural network, which is configured to generate structural parameters characterizing the structure of a protein by processing network inputs, including multiple sequence alignment representations of the protein.
[0032] Figure 5 This is a flowchart of an example process for training a structure prediction neural network that can generate structural parameters characterizing the structure of a protein without handling multiple sequence alignments of the protein.
[0033] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation
[0034] This specification describes a training system that can, for example, train a protein structure prediction neural network with a set of model parameters by repeatedly adjusting the current values of the model parameters to determine the training values of the model parameters based on the initial values of the model parameters.
[0035] Throughout this specification, a protein structure prediction neural network (or "structure prediction neural network") refers to a neural network that processes inputs characterizing a protein to generate an output set of structure parameters that characterize the protein's predicted structure. The structure of a protein refers to the three-dimensional (3-D) conformation of the atoms within the protein after it has undergone folding. Figure 3 Illustrations of unfolded and folded proteins are provided.
[0036] For simplicity, this specification will primarily concern training neural networks for protein structure prediction. However, the techniques described herein are broadly applicable to training any machine learning model (i.e., one with a trainable set of model parameters) for protein structure prediction. Other examples of machine learning models may include, for instance, random forest models and support vector machine models.
[0037] To generate structural parameters characterizing the predicted structure of a protein, a structure prediction neural network can process inputs including a representation of the protein's amino acid sequence, and in some cases, a representation of the protein's multiple sequence alignment (MSA). An MSA can specify the alignment of a protein's amino acid sequence with, for example, multiple additional amino acid sequences from other proteins (e.g., homologous proteins). More specifically, an MSA can define the correspondence between positions in the protein's amino acid sequence and corresponding positions in the amino acid sequences of multiple additional proteins. An MSA can be generated, for example, by processing a database of amino acid sequences using any suitable computational sequence alignment technique (e.g., progressive alignment construction). The amino acid sequences in an MSA can be understood as having evolutionary relationships, for example, where each amino acid sequence in the MSA can share a common ancestor. The correlations between amino acid sequences in an MSA can encode information relevant to the predicted protein's structure. An MSA can be obtained using any known technique, such as the one summarized in https: / / en.wikipedia.org / wiki / Multiple_sequence_alignment.
[0038] The representation of a protein's amino acid sequence can be an ordered set of embeddings, comprising corresponding embeddings (i.e., ordered numerical sets, such as vectors or matrices of numbers) for each position in the amino acid sequence. The corresponding embedding for each position in the amino acid sequence can be, for example, a one-hot vector, which defines the identity of the amino acid at that position in the sequence. The one-hot vector has distinct components corresponding to each possible amino acid (e.g., each of a predetermined number of possible amino acids). The one-hot vector representing a specific amino acid has a value of one (or some other predetermined value) in the component corresponding to that specific amino acid and a value of zero (or some other predetermined value) in the other components.
[0039] The representation of a protein's MSA can be an ordered set of embeddings, comprising the corresponding embeddings for each position in the sequence of each amino acid in the MSA. The corresponding embedding for each position in the sequence of each amino acid can be, for example, a one-hot vector that defines the identity of amino acids at that position in the amino acid sequence. In some cases, the representation of a protein's MSA can be a set of features derived from the MSA, such as second-order statistical features, like those described in S. Seemayer, M. Gruber, and J. Soding, "CCMpred: fast and precise prediction of protein residue-residue contacts from correlated mutations," Bioinformatics, 2014.
[0040] In some embodiments, the structural parameters generated by the protein structure prediction neural network may include a sequence of three-dimensional (3D) numerical coordinates, where each coordinate represents the spatial position of a corresponding atom in an amino acid of the protein (in some given reference frame). In a particular example, the structural parameters may include a sequence of 3D numerical coordinates representing the corresponding spatial position of an α-carbon atom in an amino acid of the protein. In this specification, an α-carbon atom, which may be referred to as a main-chain atom, refers to the carbon atom to which the amino functional group, carboxyl functional group, and side chain are attached in the amino acid. Alternatively or additionally, the structural parameters may include a sequence of torsion (i.e., dihedral) angles between specific atoms in an amino acid of the protein. For example, the structural parameters may be a sequence of phi (φ), psi (ψ), and omega (ω) dihedral angles between main-chain atoms in an amino acid of the protein.
[0041] In some implementations, the structural parameters generated by the protein structure prediction neural network may include a “distance map” characterizing the estimated distances (e.g., measured in angstroms) between each pair of amino acids in the protein. In some examples, the distance map can characterize the estimated distances between amino acid pairs as a probability distribution over a set of possible distances between amino acid pairs. The distance map can be represented as an ordered set of numerical classes, such as a vector or matrix of numbers.
[0042] Generally, the structure prediction neural networks described in this specification can have any suitable neural network architecture that enables them to perform the functions they describe. For example, a structure prediction neural network can have a corresponding architecture that includes any suitable type of neural network layers (e.g., fully connected layers, convolutional layers, pooling layers, self-attention layers, etc.) arranged in any suitable configuration (e.g., as a linear sequence of layers).
[0043] The following will describe it in more detail. Figure 1A -B describes a training system for training a structure prediction neural network that can predict the structure of a protein by processing inputs including the protein's MSA.
[0044] The following will describe it in more detail. Figure 2 A training system for training a neural network that can predict the structure of a protein without processing the protein's MSA is described.
[0045] Figure 1A An example training system 100 is shown. The training system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, wherein the systems, components and techniques described below are implemented.
[0046] The training system 100 is configured to train a structure prediction neural network 102, which can generate structural parameters characterizing the structure of a protein by processing inputs, including corresponding representations of: (i) the amino acid sequence of the protein; and (ii) the MSA of the protein.
[0047] Training system 100 uses supervised training system 104 and self-supervised training system 106 to train structure prediction neural network 102.
[0048] The supervised training system 104 trains the structure prediction neural network 102 based on a set of “paired” training examples 108. Each paired training example 108 corresponds to a specific protein and includes data defining the following: (i) the training input to the structure prediction neural network, including the protein's amino acid sequence and MSA; and (ii) the protein's basic real-world structure. The protein's basic real-world structure refers to the known structure of the protein, which may have been experimentally determined using physical laboratory techniques (e.g., X-ray crystallography) or some other technique using a physical (i.e., real-world) instance of the protein. The protein's basic real-world structure may be in the form of corresponding values of multiple basic real-world structure parameters. These basic real-world structure parameters may be structural parameters generated by the structure prediction neural network.
[0049] The supervised training system 104 can train the structure prediction neural network 102 to generate structure parameters that match the base real-world structure parameters specified by the paired training example 108. More specifically, the supervised training system 104 can train the structure prediction neural network 102 to optimize an objective function that measures the error between (i) the structure parameters generated by the structure prediction neural network and (ii) the base real-world structure parameters specified by the paired training example. The objective function can measure the error between the corresponding sets of structure parameters, for example, as a squared error, or in any other suitable manner. The supervised training system 104 can train the structure prediction neural network 102 using any suitable training technique, such as stochastic gradient descent.
[0050] Optionally, the supervised training system 104 can train the structure prediction neural network 102 to generate one or more auxiliary outputs. Training the structure prediction neural network to generate auxiliary outputs allows for faster training of the structure prediction neural network and, for example, enables the structure prediction neural network to generate more effective internal representations of proteins, thus achieving higher prediction accuracy. Several examples of auxiliary outputs are described below.
[0051] In one example, the supervised training system 104 can train a structure prediction neural network to process inputs representing a protein to generate an auxiliary output that estimates the confidence level of the accuracy of the structural parameters generated by the structure prediction neural network for the protein. More specifically, the auxiliary output can estimate (i) the error (e.g., squared error) between (i) the structural parameters generated by the protein's structure prediction neural network and (ii) the protein's underlying real structural parameters.
[0052] In another example, the supervised training system 104 can mask the identity of corresponding amino acids at one or more positions in one or more amino acid sequences within an MSA provided as input to a structure prediction neural network 102. In this example, the supervised training system 104 can train the structure prediction neural network 102 to generate auxiliary outputs that predict the identity of each masked amino acid in the input MSA. "Masking" the identity of an amino acid at a position in the MSA can refer to replacing the data identifying the amino acid at that position with a predefined masking identifier (token). The supervised training system can randomly select the positions of the amino acids to be masked in the MSA.
[0053] The self-supervised training system 106 trains the structure prediction neural network based on a set of "unpaired" training examples 110. Each unpaired training example 110 corresponds to a specific protein and includes data defining the training input to the structure prediction neural network, which includes the protein's amino acid sequence and MSA. Unlike paired training examples 108, the underlying real-world protein structure may be unknown for some or all of the unpaired training examples.
[0054] To train the structure prediction neural network 102, the self-supervised training system 106 can process the MSA included in each unpaired training example to generate a “simplified” MSA, for example, by randomly removing or masking data from the full MSA (i.e., the entire MSA in training example 110, which typically includes corresponding data for substantially every amino acid in the corresponding protein). The self-supervised training system 106 can generate data defining the corresponding “target” structure parameters for each unpaired training example 110 based on the set of structure parameters generated by the structure prediction neural network 102 through processing the input, which includes the full (i.e., unsimplified) MSA from the unpaired training examples. The self-supervised training system 106 can then train the structure prediction neural network to process the simplified MSA for each unpaired training example to generate structure parameters that match the target structure parameters of the training examples. (Reference) Figure 1B An example of the self-supervised training system 106 is described in more detail.
[0055] Training system 100 trains structure prediction neural network 102 using both supervised training system 104 and self-supervised training system 106. For example, training system 100 may first use supervised training system 104 and then use self-supervised training system 106 to train structure prediction neural network 102. In some embodiments, training system 100 may repeatedly alternate between training structure prediction neural network 102 using supervised training system 104 and self-supervised training system 106.
[0056] Training the structure prediction neural network 102 using the self-supervised training system 106 can improve the performance (e.g., prediction accuracy) of the structure prediction neural network 102 by reducing the likelihood of the structure prediction neural network overfitting the paired training examples 108. The structure prediction neural network 102 can “overfit” the paired training examples, for example, by learning to predict the underlying real-world protein structures specified by the paired training examples based on irrelevant variations in the training inputs, rather than on implicit reasoning based on inferred biochemical principles. Furthermore, the number of available unpaired training examples can be much greater than the number of available paired training examples, thus the self-supervised training system 106 can enable the structure prediction neural network 102 to learn to effectively predict the structures of a wider variety of proteins.
[0057] Figure 1B An example self-supervised training system 106 is shown. The self-supervised training system 106 is an example of a system implemented as a computer program on one or more computers in one or more locations, wherein the systems, components, and techniques described below are implemented.
[0058] The self-supervised training system 106 trains the structure prediction neural network 102 based on a set of unpaired training examples 110. Each unpaired training example 110 corresponds to a specific protein and includes data defined as training input to the structure prediction neural network 102, which includes (i) the amino acid sequence of the protein; and (ii) the protein's "complete" (i.e., unsimplified) MSA. Typically, the underlying real-world protein structure may be unknown for some or all of the unpaired training examples.
[0059] As part of training the structure prediction neural network 102, the self-supervised training system 106 generates a corresponding set of target structure parameters 112 for each unpaired training example 110. The target structure parameters 112 of the unpaired training examples characterize the predicted structure of the protein corresponding to the unpaired training example. When processing a simplified MSA of a protein instead of the full MSA, the target structure parameters 112 provide a prediction target for the structure prediction neural network 102, as will be described in more detail below.
[0060] To generate target structural parameters 112 for the unpaired training example 110, the structure prediction neural network 102 processes inputs including representations of the complete MSA 114 and the amino acid (AA) sequence 116 of the protein to generate output structural parameters 118. The self-supervised training system 106 then determines the target structural parameters 112 based on the structural parameters 118 generated by the structure prediction neural network 102 through processing the complete MSA 114 and AA sequence 116. In some embodiments, the self-supervised training system 106 may determine the target structural parameters 112 as equal to the structural parameters 118 generated by the structure prediction neural network. In some embodiments, the self-supervised training system 106 may determine the target structural parameters 112 by adding random noise values to the structural parameters 118 generated by the structure prediction neural network 102. Adding random noise values to the structural parameters 118 generated by the structure prediction neural network 102 as part of generating the target structural parameters 112 may reduce the likelihood of overfitting and thereby regularize the training of the structure prediction neural network 102.
[0061] In addition to generating the target structure parameters 112 for each unpaired training example 110, the self-supervised training system 106 also uses a simplification engine 120 to process the full MSA 114 from each unpaired training example 110 to generate a corresponding “simplified” MSA 122. The simplification engine 120 can process the full MSA 114 to generate the simplified MSA 122, for example, by randomly removing or masking data from the full MSA 114. Several examples of operations that can be performed by the simplification engine 120 to generate the simplified MSA 122 from the full MSA 114 are described in more detail below.
[0062] In some implementations, simplification engine 120 may randomly remove one or more amino acid sequences from the full MSA 114 as part of generating a simplified MSA 122. Simplification engine 120 may use a random process to determine how many amino acid sequences to remove from the full MSA 114 and which specific amino acid sequences to remove. For example, simplification engine 120 may sample reduction parameter values based on a probability distribution over a space of possible reduction parameter values, which define the number of amino acid sequences to be removed from the full MSA 114. The space of possible reduction parameter values may be, for example, an interval of (0,1), and the sampled reduction parameter values may define a portion of the amino acid sequences to be removed from the full MSA 114. For example, sampling a reduction parameter value of 0.15 may define that 15% of the amino acid sequences in the full MSA 114 should be removed. After sampling the reduction parameter values, simplification engine 120 may randomly remove a specified number of amino acid sequences from the full MSA 114.
[0063] In some implementations, the simplification engine 120 can randomly mask the identity of corresponding amino acids at one or more positions in one or more amino acid sequences within the complete MSA 114. “Mask” the identity of amino acids at a position in the complete MSA 114 can refer to replacing the data identifying the amino acid at that position with a predefined masking identifier (token). In one example, the simplification engine 120 can sample masking parameter values based on a probability distribution over a space of possible masking parameter values (e.g., intervals (0, 0.05)). The masking parameter values can define the probability that the identity of a corresponding amino acid at any position in any amino acid sequence of the complete MSA should be masked. After sampling the masking parameter values, the simplification engine 120 can mask the identity of each amino acid in each amino acid sequence in the MSA with the probability defined by the masking parameter values.
[0064] The self-supervised training system 106 trains the structure prediction neural network 102 to process (i) the AA sequence 116 and (ii) the simplified MSA 122 representation for each unpaired training example to generate structure parameters 126 that match the target structure parameters 112 of the unpaired training example. More specifically, the self-supervised training system 106 uses a training engine 124 to train the structure prediction neural network 102 to optimize an objective function. For each unpaired training example, the objective function can measure the error between (i) the structure parameters 126 generated by the structure prediction neural network from the simplified MSA 122 and (ii) the target structure parameters 112 generated from the full MSA 114. The objective function can measure the error between the corresponding sets of structure parameters, for example as a squared error, or measure the error in any other suitable manner.
[0065] The self-supervised training system 106 can use the training engine 124 to train the structure prediction neural network 102 using any suitable training technique (e.g., stochastic gradient descent via a training iteration sequence). More specifically, in each training iteration, the training engine 124 can sample a batch of unpaired training examples. For each unpaired training example in the batch, the structure prediction neural network 102 can process the corresponding simplified MSA 122 and AA sequence 116 based on the current values of the model parameters 128 of the structure prediction neural network 102 to generate the corresponding structure parameters 126. The training engine 124 can then evaluate an objective function that measures the error between (i) the target structure parameters 112 and (ii) the structure parameters 126 generated by the structure prediction neural network 102 for the unpaired training examples in the current batch. The training engine 124 can determine the gradient of the objective function relative to the model parameters of the structure prediction neural network, for example, through backpropagation, and use the gradient to update the current values of the model parameters using any suitable gradient descent optimization technique (e.g., RMSprop or Adam).
[0066] In some implementations, the self-supervised training system 106 can train the structure prediction neural network 102 to generate one or more auxiliary outputs, such as auxiliary outputs that predict the identity of each masked amino acid in the MSA 122.
[0067] The structural parameters 118 generated by the structural prediction neural network based on the full MSA114 may be imprecise for one or more of the unpaired training examples. Therefore, the target structural parameters 112 of these training examples may be imprecise, and using these target structural parameters 112 during training can, for example, reduce the performance of the structural prediction neural network 102 by amplifying the error caused by the structural prediction neural network 102.
[0068] To reduce the likelihood that inaccurate target structural parameters 112 negatively impact the training of the structure prediction neural network 102, the self-supervised training system 106 can estimate the corresponding confidence level of the target structural parameters 112 for each unpaired training example. In some embodiments, the self-supervised training system 106 can suppress training the structure prediction neural network based on any target structural parameters 112 whose confidence estimates do not meet a threshold. In some embodiments, the self-supervised training system 106 can adjust the objective function based on the confidence estimate in the target structural parameters 112 for each training example, for example, to reduce the impact of low-confidence target structural parameters 112 on the objective function. For example, the objective function It can be given by the following formula:
[0069]
[0070] Where i is the index of N training examples, c i T represents the scaling factor for the confidence of the target structure parameters based on training example i. i P represents the target structure parameters of training example i. i Let represent the structure parameters generated by the structure prediction neural network based on the simplified MSA of the training example i, and Err(·,·) represent the error measurement, such as the squared error.
[0071] The self-supervised training system 106 can determine the confidence estimate of the target structure parameter 112 in various ways. Several examples of methods for determining the confidence estimate of the target structure parameter 112 for training examples are described in more detail below.
[0072] In one example, the self-supervised training system 106 can obtain a confidence estimate of the target structure parameters 112 of the training examples as an auxiliary output generated by the structure prediction neural network 102 through processing the full MSA 114 of the training examples. (See reference) Figure 1A The generation of confidence estimates as an auxiliary output of the structure prediction neural network 102 is described in more detail.
[0073] In another example, the self-supervised training system 106 can obtain a confidence estimate of the target structural parameter 112 of the training example based on an estimated distance map of the protein corresponding to the training example. The distance map can be defined for each pair of amino acids in the protein as a probability distribution over the possible physical distances between amino acid pairs in the protein structure. The self-supervised training system 106 can obtain the distance map as an auxiliary or main output generated by the structure prediction neural network by processing the full MSA 114 of the training example. For each pair of amino acids, the self-supervised training system 106 can determine the confidence estimate based on (i) the difference between the probability distribution defined by the distance map over the possible distances between amino acid pairs and (ii) the "background" probability distribution.
[0074] The background probability distribution can be a predefined probability distribution within a possible range of distances, reflecting the statistical distribution of distances between amino acid pairs in a known protein structure. Differences between corresponding probability distributions can be identified, for example, as Kullback-Leibler divergence. Typically, a large difference between the probability distribution defined by the distance map and the background probability distribution can indicate a higher confidence level in the target structure parameters generated by the structure prediction neural network for the training examples.
[0075] After training the structure prediction neural network 102 to process simplified MSA 122 to generate the corresponding target structure parameters 112 for each training example, the self-supervised training system 106 can provide the training model parameters 128 of the structure prediction neural network 102 as output.
[0076] Optionally, the self-supervised training system 106 can generate new target structure parameters 112 for training example 110 based on the training values of the model parameters 128 of the structure prediction neural network 102, and repeat the above process to continue training the structure prediction neural network 102. In some cases, the self-supervised training system 106 can continue to iteratively repeat the self-supervised training process used to train the structure prediction neural network 102 until the termination criterion is met.
[0077] In some implementations, the self-supervised training system 106 increases the expected amount of data removed or masked from the full MSA 114 by the simplification engine 120 in each iteration of the training process. For example, the self-supervised training system 106 may increase the mean of a probability distribution over possible reduction parameter values in each iteration of the training process, based on which the simplification engine samples reduction parameter values for portions of amino acid sequences that define to be removed from the full MSA. Increasing the expected amount of data removed or masked from the full MSA in each iteration of the training process can improve the performance of the structure prediction neural network 102 in predicting protein structures by processing MSAs that include several amino acid sequences.
[0078] Figure 2 An example teacher-student training system 200 is shown. The teacher-student training system 200 is an example of a system implemented as a computer program on one or more computers in one or more locations, wherein the systems, components and techniques described below are implemented.
[0079] The training system 200 uses the “teacher” structure prediction neural network 202 to train the “student” structure prediction neural network 204, specifically by using the teacher structure prediction neural network to generate target structure parameters that will be used as prediction targets by the student structure prediction neural network.
[0080] The teacher structure prediction neural network 202 is configured to process inputs including both: (i) a representation of the protein's amino acid (AA) sequence 206; and (ii) a representation of the protein's MSA 208. The teacher structure prediction neural network 202 processes the inputs to generate structural parameters 210 characterizing the predicted structure of the protein.
[0081] The student structure prediction neural network 204 is configured to process inputs that include a representation of the protein's AA sequence but not a representation of the protein's MSA 208. In some embodiments, the student structure prediction neural network 204 processes inputs that include only a representation of the protein's AA sequence. The student structure prediction neural network 204 processes the inputs to generate structure parameters 212 characterizing the predicted structure of the protein (specifically, in contrast to the teacher structure prediction neural network 202, which does not process the representation of the protein's MSA).
[0082] The teacher structure prediction neural network 202 can be trained using any suitable machine learning training technique. For example, a reference can be used. Figure 1A The supervised training system described in 104 or as referenced Figure 1A -B describes a combination of supervised training system 104 and self-supervised training system 106 to train teacher structure prediction neural network 202.
[0083] The training system 200 trains a student structure prediction neural network based on a set of unpaired training examples 214. Each unpaired training example 214 corresponds to a specific protein and includes data defining the following: (i) the amino acid sequence of the protein; and (ii) multiple sequence alignments of the protein.
[0084] Typically, the underlying real-world protein structure may be unknown for some or all of the unpaired training examples 214. Therefore, the training system 200 uses the teacher structure prediction neural network 202 to generate a set of target structure parameters 216 representing the predicted structure of the corresponding protein for each unpaired training example 214.
[0085] To generate the target structure parameters 216 for training example 214, training system 200 uses teacher structure prediction neural network 202 to generate a set of structure parameters 210 for each training example 214. Teacher structure prediction neural network 202 generates the structure parameters 210 for each training example by processing the input, which includes the corresponding representations of: (i) the AA sequence 206 from the training example; and (ii) the MSA 208 from the training example.
[0086] The training system 200 determines the target structural parameter 216 for each training example based on the structural parameters 210 generated by the teacher structural prediction neural network 202 for the training examples. In some embodiments, the training system 200 may determine the target structural parameter 216 as equal to the structural parameters 210 generated by the teacher structural prediction neural network 202. In some embodiments, the training system 200 may determine the target structural parameter 216 by adding random noise values to the structural parameters 210 generated by the teacher structural prediction neural network 202. Adding random noise values to the structural parameters 210 generated by the teacher structural prediction neural network 202 as part of generating the target structural parameter 216 may reduce the likelihood of overfitting and thereby regularize the training of the student structural prediction neural network 204.
[0087] The training system 200 can train a student structure prediction neural network 204 to process the representation of the AA sequence 206 of the training example for each training example to generate structure parameters 212 that match the target structure parameters 216 of the training example. More specifically, the training system 200 uses a training engine 218 to train the student structure prediction neural network 204 to optimize an objective function. For each training example, the objective function can measure (i) the error between the structure parameters 212 generated by the student structure prediction neural network by processing the AA sequence 206 of the training example and (ii) the target structure parameters 216 of the training example. The objective function can measure the error between the corresponding sets of structure parameters, for example as a squared error, or measure the error in any other suitable manner.
[0088] The training system 200 can train the student structure prediction neural network 204 using any suitable training technique (e.g., stochastic gradient descent via a training iteration sequence). More specifically, in each training iteration, the training engine 218 can sample a batch of training examples. For each training example in the batch, the student structure prediction neural network 204 processes the representation of the corresponding AA sequence 206 based on the current values of the model parameters 220 of the student structure prediction neural network 204 to generate structure parameters 212. The training engine 218 then evaluates an objective function that measures the error between (i) the target structure parameters 216 and (ii) the structure parameters 212 generated by the student structure prediction neural network 204 for the training examples in the current batch. The training engine 218 determines the gradient of the objective function relative to the model parameters of the student structure prediction neural network and uses the gradient to update the current values of the model parameters of the student structure prediction neural network using any suitable gradient descent optimization technique. The training engine can use the determined gradient to improve the model parameters, for example, through backpropagation, and the gradient descent optimization technique can be, for example, RMSprop or Adam.
[0089] In some cases, the structural parameters 210 generated by the teacher structural prediction neural network may be imprecise for one or more of the training examples. Therefore, the target structural parameters 216 of these training examples may be imprecise, and using these target structural parameters 216 during training may degrade the performance of the student structural prediction neural network 204.
[0090] To reduce the likelihood that inaccurate target structure parameters 216 negatively impact the training of the student structure prediction neural network 204, the training system 200 can estimate the corresponding confidence level of the target structure parameters 216 for each training example. In some implementations, the training system 200 can suppress any training example from being used to train the student structure prediction neural network if the confidence level of the target structure parameters 216 based on the training example does not meet a threshold. In some implementations, the training system 200 can adjust the objective function based on the confidence level of the target structure parameters 216 for each training example, for example, to reduce the impact of low-confidence target structure parameters 216 on the objective function. For example, the objective function It can be given by the following formula:
[0091]
[0092] Where i is the index of N training examples, c i T represents the scaling factor for the confidence of the target structure parameter 216 based on training example i. i P represents the target structure parameters of training example i. iLet represent the structural parameters generated by the student structure prediction neural network for training example i, and Err(·,·) represent the error measurement, such as the squared error.
[0093] The training system 200 can determine the confidence estimate of the target structure parameter 216 in various ways. Several examples of methods for determining the confidence estimate of the target structure parameter 216 for training examples are described in more detail below.
[0094] In one example, training system 200 can obtain confidence estimates of the target structure parameters 216 of the training examples as auxiliary outputs of the teacher structure prediction neural network 202 used for training examples. (See reference) Figure 1A The generation of confidence estimates as an auxiliary output of a structure prediction neural network is described in more detail.
[0095] In another example, training system 200 can obtain a confidence estimate of the target structure parameter 216 of the training example based on an estimated distance map corresponding to the protein of the training example. The distance map can be defined for each pair of amino acids in the protein as a probability distribution within the range of possible physical distances between amino acid pairs in the protein structure. Training system 200 can obtain the distance map as an auxiliary or main output of the teacher structure prediction neural network 202. (See above reference...) Figure 1B The confidence estimation of the structural parameter set based on the estimated distance map is described in more detail.
[0096] After training the student structure prediction neural network 204, the training system 200 can provide the training model parameters 220 of the student structure prediction neural network 204 as output.
[0097] The student structure prediction neural network 204 can predict the structure of any protein based on its amino acid sequence, without requiring the protein's MSA (Material Sequence Aspect). Therefore, the student structure prediction neural network can be more widely applied than the teacher structure, for example, because the MSA can be unavailable for many proteins.
[0098] Figure 3 This is an illustration of unfolded and folded proteins. An unfolded protein is a random coil of amino acids. An unfolded protein undergoes protein folding and folds into a 3D conformation. Protein structures typically include stable local folding patterns, such as α-helices (e.g., as depicted in 302) and β-sheets.
[0099] Figure 4This is a flowchart of an example process 400 for training a structure prediction neural network configured to generate structural parameters characterizing the structure of a protein by processing network inputs, including representations of multiple sequence alignments of the protein. For convenience, process 400 will be described as being executed by a system of one or more computers located in one or more locations. For example, a training system appropriately programmed according to this specification (e.g., Figure 1A The training system 100 can execute process 400.
[0100] The system obtains complete multiple sequence alignments for each of the multiple proteins (402).
[0101] The system generates target structural parameters (404) for each protein, representing the structure of the protein from its complete multiple sequence alignment. More specifically, the system uses a structure prediction neural network to process the representation of the complete multiple sequence alignment for each protein to generate output structural parameters, and determines the target structural parameters of the protein based on the output structural parameters of the protein.
[0102] The system determines the simplified multiple sequence alignment of a protein for each protein, for example, by removing or masking data from the protein’s full multiple sequence alignment (406).
[0103] The system trains the structure prediction neural network to process simplified multiple sequence alignment representations of one or more proteins to generate structural parameters that match the target structural parameters of the protein (408).
[0104] Figure 5 This is a flowchart of an example process 500 for training a structure prediction neural network that can generate structural parameters characterizing the structure of a protein without processing multiple sequence alignments of the protein. For convenience, process 500 will be described as being executed by a system of one or more computers located in one or more locations. For example, a teacher-student training system appropriately programmed according to this specification (e.g., Figure 2 The teacher-student training system 200 can execute process 500.
[0105] The system trains a teacher structure prediction neural network, which is configured to generate structural parameters characterizing the structure of a protein by processing inputs, including corresponding representations of: (i) the amino acid sequence of the protein; and (ii) multiple sequence alignments of the protein (502).
[0106] The system uses a teacher structure prediction neural network to generate corresponding target structure parameters for each of a variety of proteins (504).
[0107] The system trains a student structure prediction neural network configured to generate structural parameters (506) characterizing the structure of a protein by processing inputs, the inputs of which (i) include a representation of the protein's amino acid sequence; and (ii) do not include a representation of the protein's multiple sequence alignments. The system trains the student structure prediction neural network to generate structural parameters characterizing the structure of each protein, which are matched with the protein's target structural parameters.
[0108] This specification uses the term "configuration" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform specific operations or actions, this means that the system has software, firmware, hardware, or a combination thereof installed thereon that causes the system to perform those operations or actions in operation. For one or more computer programs to be configured to perform specific operations or actions, this means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform those operations or actions.
[0109] Embodiments of the subject matter and functional operation described in this specification may be implemented using digital electronic circuit systems, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed herein and their equivalents), or combinations thereof. Embodiments of the subject matter described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination thereof. Alternatively or additionally, program instructions may be encoded in artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals generated for encoding information to be transmitted to a suitable receiver device for execution by the data processing apparatus).
[0110] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus may also be or include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0111] Computer programs can be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can also be referred to or described as programs, software, software applications, applications, modules, software modules, scripts, or code. They can be deployed in any form, including as standalone programs or modules, components, subroutines, or other units suitable for a computing environment. A program may, but is not required to, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program under development, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a data communication network.
[0112] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed on and run on the same one or more computers.
[0113] The processes and logic flows described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform functions by manipulating input data and generating output. These processes and logic flows can also be executed by a dedicated logic circuit system (such as an FPGA or ASIC), or by a combination of a dedicated logic circuit system and one or more programmable computers.
[0114] A computer suitable for executing computer programs can be based on a general-purpose microprocessor or a special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The essential components of a computer are the central processing unit for making or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or incorporated into a special-purpose logic circuit system. Typically, a computer will also include one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, or the computer may be operatively coupled to receive data from or send data to or both of such mass storage devices. However, a computer does not necessarily need to have such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.
[0115] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0116] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor; and a keyboard and pointing device, such as a mouse or trackball, through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound input, voice input, or tactile input. Additionally, the computer can interact with the user by sending documents to the device used by the user and receiving documents from that device (e.g., by a web browser on the user's device sending a webpage in response to a request received from a web browser). Furthermore, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user in return.
[0117] The data processing apparatus used to implement machine learning models may also include, for example, dedicated hardware accelerator units for handling the general and computationally intensive portions of machine learning training or production (i.e., inference) workloads.
[0118] Machine learning models can be implemented and deployed using machine learning frameworks such as TensorFlow, Microsoft Cognitive Toolkit, Apache Singa, or Apache MXNet.
[0119] Embodiments of the subject matter described in this specification can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., client computers with a graphical user interface, web browser, or application through which users can interact with embodiments of the subject matter described in this specification), or computing systems that include one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0120] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. A client-server relationship is created by computer programs running on respective computers that have a client-server relationship with each other. In some embodiments, the server sends data (e.g., HTML pages) to a user device, for example, to display data to a user interacting with a device acting as a client and to receive user input from that user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.
[0121] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of the invention. Certain features described in this specification within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described within the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although these features may be described above as functioning in certain combinations and initially claimed even equally, in some cases one or more features from a claimed combination may be removed from that combination, and the claimed combination may involve sub-combinations or variations thereof.
[0122] Similarly, although operations are depicted and described in a specific order in the accompanying drawings and claims, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or as requiring all illustrated operations to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0123] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions described in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.
Claims
1. A method for training a structure prediction neural network by one or more data processing devices, the structure prediction neural network being configured to generate structural parameters characterizing the structure of a protein by processing network inputs, the network inputs including a multiple sequence alignment representation of the protein, the method comprising: For each of the multiple proteins, obtain complete multiple sequence alignment of the protein; Generate target structural parameters for each of the multiple proteins, characterizing the structure of the protein from the complete multiple sequence alignment of the protein, including: The structure prediction neural network is used to process the complete multiple sequence alignment representation of the protein to generate output structure parameters characterizing the structure of the protein; and The target structural parameters of the protein are determined based on the output structural parameters of the protein. For each of the multiple proteins, a simplified multiple sequence alignment of the protein is determined, including removing or masking data from the full multiple sequence alignment of the protein; and The structure prediction neural network is trained to process the simplified multiple sequence alignment representation of one or more of the proteins to generate structural parameters that match the target structural parameters of the protein.
2. The method of claim 1, wherein removing data from the complete multiple sequence alignment of each of the plurality of proteins comprises: Remove one or more amino acid sequences from the multiple sequence alignment of the protein.
3. The method of claim 2, wherein removing one or more amino acid sequences from the multiple sequence alignment of the protein comprises: Sampling of reduction parameter values is performed from the set of possible reduction parameter values based on a probability distribution over the set of possible reduction parameter values, wherein the reduction parameter value specifies the number of amino acid sequences to be removed from the complete multiple sequence alignment of the protein; as well as Remove a specified number of amino acid sequences from the complete multiple sequence alignment of the protein.
4. The method of claim 3, wherein removing the specified number of amino acid sequences from the complete multiple sequence alignment of the protein comprises: The amino acid sequence to be removed from the complete multiple sequence alignment of the protein is randomly selected.
5. The method according to any preceding claim, wherein masking data from the complete multiple sequence alignment of the proteins for each of the plurality of proteins comprises: Masking the identity of corresponding amino acids at one or more positions in one or more amino acid sequences in the complete multiple sequence alignment of the protein.
6. The method of claim 5, wherein masking the identity of corresponding amino acids at one or more positions in one or more amino acid sequences during the complete multiple sequence alignment of the protein comprises: The positions in the amino acid sequence to be masked in the complete multiple sequence alignment of the protein are randomly sampled.
7. The method according to claim 6, further comprising: The structure prediction neural network is trained to process the simplified multiple sequence alignment representation of each of the multiple proteins to generate an auxiliary output that predicts the identity of each masked amino acid in the simplified multiple sequence alignment of the protein.
8. The method according to any one of claims 1 to 4, wherein determining the target structural parameters of the protein based on the output structural parameters of the protein comprises: Random noise values are added to the output structural parameters of the protein.
9. The method according to any one of claims 1 to 4, wherein the structure prediction neural network is configured to process network inputs comprising both of the following: (i) a representation of multiple sequence alignments of a protein; and (ii) a representation of the amino acid sequence of the protein.
10. The method according to any one of claims 1 to 4, further comprising: For each of the multiple proteins, a confidence estimate of the target structural parameter of that protein is determined.
11. The method according to claim 10, further comprising: Identify one or more proteins whose confidence estimate of the target structural parameters of the protein does not meet a threshold; as well as Suppress the training of the structure prediction neural network based on the identified proteins.
12. The method of claim 11, wherein training the structure prediction neural network comprises: Determine the gradient of an objective function that measures for one or more of the plurality of proteins (i) the error between the structural parameters generated by the structure prediction neural network by processing the simplified multiple sequence alignment representation of the protein and (ii) the target structural parameters of the protein, wherein the error is scaled by a function of the confidence estimate of the target structural parameters of the protein.
13. The method of claim 12, wherein for each of the plurality of proteins: The confidence estimate of the target structure parameters of the protein is generated by processing the representation of the complete multiple sequence alignment of the protein as an auxiliary output of the structure prediction neural network; The confidence estimate of the target structural parameters of the protein is defined as (i) an estimate of the error between the output structural parameters generated by the structure prediction neural network by processing the complete multiple sequence alignment of the protein and (ii) the underlying real structural parameters characterizing the underlying real structure of the protein.
14. The method according to any one of claims 1 to 4, the method further comprising: The structure prediction neural network is trained to process multiple sequence alignment representations of one or more other proteins to generate structure parameters that match the underlying real structure parameters of the other proteins.
15. The method of claim 14, wherein the basic real-world structural parameters of the other proteins are determined by physical experiments.
16. A method performed by one or more data processing devices, the method comprising: A teacher structure prediction neural network is trained, the teacher structure prediction neural network being configured to generate structural parameters characterizing the structure of a protein by processing inputs, the inputs including (i) a representation of the amino acid sequence of the protein; and (ii) a representation of multiple sequence alignments of the protein; as well as A student structure prediction neural network is trained, the student structure prediction neural network being configured to generate structural parameters characterizing the structure of a protein by processing an input, the input (i) including a representation of the amino acid sequence of the protein, the representation of the amino acid sequence including a corresponding embedding for identifying an amino acid at each position in the amino acid sequence; (ii) excludes the representation of multiple sequence alignments of the proteins, wherein the training for each of the multiple proteins includes: The teacher structure is used to predict neural network to generate target structure parameters characterizing the structure of the protein; as well as The student structure prediction neural network is trained to generate structural parameters characterizing the structure of the protein, the structural parameters being matched with the target structural parameters of the protein.
17. The method of claim 16, wherein for each of the plurality of proteins, generating target structure parameters characterizing the structure of the protein using the teacher structure prediction neural network comprises: Processing input, wherein input (i) includes a representation of the amino acid sequence of the protein; (ii) includes a representation of multiple sequence alignments of the protein, and uses the teacher structure prediction neural network to generate output structural parameters characterizing the structure of the protein; as well as The target structural parameters of the protein are determined based on the output structural parameters of the protein.
18. The method of claim 17, wherein determining the target structural parameters of the protein based on the output structural parameters of the protein comprises: Random noise values are added to the output structural parameters of the protein.
19. The method of any one of claims 16 to 18, wherein the representation of the multiple sequence alignment of the protein processed by the teacher structure predictive neural network includes features derived from the multiple sequence alignment of the protein.
20. The method according to any one of claims 16 to 18, the method further comprising: For each of the multiple proteins, a confidence estimate of the target structural parameter of that protein is determined.
21. The method according to claim 20, further comprising: Identify one or more proteins whose confidence estimate of the target structural parameters of the protein does not meet a threshold; as well as Suppress the training of the student structure prediction neural network based on the identified proteins.
22. The method of claim 21, wherein training the student structure prediction neural network comprises: Determine the gradient of an objective function that measures for one or more of the plurality of proteins the error between (i) the structural parameters generated by the student structure prediction neural network for the protein and (ii) the target structural parameters of the protein, wherein the error is scaled by a function of the confidence estimate of the target structural parameters of the protein.
23. The method of claim 22, wherein for each of the plurality of proteins: The confidence estimate of the target structure parameters of the protein is generated as an auxiliary output of the teacher structure prediction neural network; The confidence estimate of the target structural parameters of the protein is defined as (i) an estimate of the error between the structural parameters generated by the teacher structural prediction neural network for the protein and (ii) the underlying real structural parameters characterizing the underlying real structure of the protein.
24. The method of claim 23, wherein the basic real-world structure parameters characterizing the basic real-world structure are determined by physical experiments.
25. The method according to any one of claims 16 to 18, wherein the structural parameters include one or both of a plurality of torsion angles and a plurality of atomic coordinates.
26. The method according to any one of claims 16 to 18, the method further comprising: The amino acid sequence of the protein is obtained and the structure of the protein is determined using a trained structure prediction neural network.
27. The method according to claim 26, further comprising: The protein is extracted from a human or animal body and the amino acid sequence is obtained from the extracted protein.
28. The method of claim 27, further comprising: Obtain the structure of the protein pattern from a human or animal body, compare the structure of the protein with the structure of the pattern obtained from the human or animal body, and identify the presence of protein misfolding diseases based on the comparison results.
29. A method for selecting a drug or ligand for use as an industrial enzyme, the method comprising: A structure prediction neural network is trained by the method according to any of the preceding claims; The structure prediction neural network described above is used to determine the structure of the target protein; Evaluate the interaction between one or more candidate proteins and the structure of the target protein; as well as One of the candidate proteins, which is to be used as an industrial enzyme, is selected based on the evaluated interactions of each candidate drug protein.
30. A method for selecting a drug or ligand for use as an industrial enzyme, the method comprising: A structure prediction neural network is trained by the method according to any of the preceding claims; The structure prediction neural network described above is used to determine the structure of one or more candidate proteins; Based on the determined structures of the one or more candidate proteins, evaluate the interactions between the one or more candidate proteins and the target protein; as well as One of the candidate proteins, which is to be used as the drug or ligand for the industrial enzyme, is selected based on the evaluated interactions of each candidate drug protein.
31. The method of claim 30, wherein the assessment of the interaction between the one or more candidate proteins and the target protein is based on the structure of the target protein determined by physical experiments.
32. The method according to any one of claims 29 to 31, wherein the target protein comprises a receptor or an enzyme, and wherein the protein is an agonist or antagonist of the receptor or enzyme.
33. The method according to any one of claims 29 to 31, further comprising: Prepare the selected protein.
34. A computer-implemented system, comprising: One or more computers; as well as One or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to operate the method according to any one of claims 1 to 33.
35. The system of claim 34, further comprising a subsystem for producing the selected protein.
36. A non-transitory computer storage medium for storing one or more instructions, which, when executed by one or more computers, cause the one or more computers to operate the method according to any one of claims 1 to 33.
Citation Information
Patent Citations
Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
US20190287515A1