Deep learning system and prediction method based on integrated sequence and structural features of protein engineering

By integrating a deep learning system with sequence and structural characteristics in protein engineering, the problem of large demand for experimental data and insufficient prediction accuracy of multi-point mutations in the prior art is solved, and high-precision protein mutation prediction is achieved, especially in the case of insufficient experimental data.

CN115954050BActive Publication Date: 2025-05-30SHANGHAI MATWINGS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211209310.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-05-30
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

Existing deep learning requires a lot of experimental data in the prediction of proteins from sequence to function, and the prediction accuracy of multi-point mutations lacking experimental data is insufficient.

Method used

A deep learning system based on protein engineering is adopted that integrates sequence and structural characteristics, including local encoder, global encoder, structural encoder, attention layer and output layer. Through technical means such as multiple sequence comparison methods, protein language model and unsupervised structural model, protein sequence and structural information are integrated for prediction.

Benefits of technology

Reliance on experimental sample size is reduced and the prediction accuracy of higher-order mutation effects is improved. Especially in the absence of experimental data, it can effectively predict the effect of protein mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115954050B_ABST
    Figure CN115954050B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep learning system and a prediction method based on protein engineering that integrate sequence and structural features. First, a deep learning model for predicting the effects of protein mutations by integrating sequence and structural information is established. Then, combined with specific data augmentation strategies, the dependence of the deep learning model on the experimental sample size is reduced. Specifically, a large number of low-quality prediction results from unsupervised models are first used to pre-train the deep learning model, and then a limited number of high-quality experimental results with experimental results are used to fine-tune the model. Experiments show that when the amount of experimental data for subsequent fine-tuning is less than 40 or there is no experimental data, the deep learning model obtained only through pre-training can achieve very high accuracy in the task of predicting the effects of higher-order mutations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of protein mutation prediction, and particularly relates to a deep learning system and prediction method integrating sequence and structural features based on protein engineering. Background Art

[0002] Proteins play an important role in maintaining life activities. They have various functions such as catalysis, binding, and transportation, and carry out most of the metabolic activities in cells. Nature provides a large number of proteins with potential practical applications. However, these proteins often do not have the optimal functions that can meet the needs of bioengineering. Directed evolution is a technology for optimizing protein functions through local search. In this process, mutations are selected and accumulated through an iterative process, that is, hundreds of mutations are tested in each generation to obtain mutants with mutations at multiple amino acid positions. The exploration degree of this technology for the protein sequence space is very limited. Therefore, directed evolution technology requires efficient high-throughput screening or a large number of experimental tests to obtain mutants with more powerful required functions (especially deep mutants with multiple amino acid mutations), which poses a huge challenge to experimental technology and experimental costs.

[0003] Since experimental screening is the bottleneck of directed evolution, it becomes very important to detect the effects of mutations (especially high-order mutations) on a computer.

[0004] Deep learning has been widely used in protein design engineering. However, the current high-precision algorithms require a large amount of protein experimental data, and the prediction accuracy of the effects of multi-site mutations is still insufficient in the absence of experimental data. Summary of the Invention

[0005] The present invention aims to solve the technical problems that existing deep learning has a large demand for experimental data in the prediction of proteins from sequence to function and the prediction accuracy of multi-site mutations lacking experimental data is insufficient, and thus provides a deep learning system and prediction method integrating sequence and structural features based on protein engineering.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0007] A deep learning system integrating sequence and structural features based on protein engineering, comprising a local encoder, a global encoder, a structural encoder, an attention layer, and an output layer;

[0008] The input of the local encoder is a mutant sequence, and the local encoder encodes the mutant sequence using a multiple sequence alignment method to output a tensor I encoding the evolutionary information of homologous proteins, and the size of the tensor I is L×256.

[0009] In the local encoder, the mutant sequence is converted into a tensor I to be trained through the Bi-LSTM layer 1 , the tensor I 1 has a size of L×128; this tensor I 1 formally satisfies the constraints including the self-constraints of amino acids and the coupling constraints between amino acids;

[0010] Using the multiple sequence alignment method to obtain the tensor I of the homologous constraint relationship of the wild-type sequence 2 , the tensor I 2 has a size of L×128; the homologous sequence constraint relationship includes the self-constraints of amino acids and the coupling constraints between amino acids;

[0011] After splicing the tensor I 1 and the tensor I 2 , the tensor I with the evolutionary information of homologous proteins is obtained, and the tensor I is the MSA coding sequence.

[0012] The input of the global encoder is the mutant sequence. The global encoder uses a protein language model to encode the mutant sequence and outputs a tensor II encoding the common biochemical characteristics and evolutionary information of proteins; the size of the tensor II is L×256;

[0013] The protein language model used is a protein language model that has completed large-scale training. After being trained with a large amount of data (the sequence data volume exceeds tens of millions), the protein language model will have excellent feature extraction ability; moreover, the protein language model used is an unsupervised learning-based protein language model; the protein language model trained with a large amount of data can fully represent the biological characteristics and evolutionary diversity of proteins. These information more represent the commonalities of all protein sequences in the entire nature and are encoded into the tensor output by this module. Because this protein language model is unsupervised, it can be directly used in the present invention. After inputting the mutant sequence, the encoded tensor II can be obtained.

[0014] The protein language model will first encode the mutant sequence into a vector, and the encoding process will incorporate the connections between amino acids. This protein language model is pre-trained on the UniRef50 database with 250 million sequences. During training, the protein language model will adjust the parameters according to the differences between the encoded sequences with some amino acids masked and the encoded complete sequences.

[0015] The input of the structural encoder is the mutant sequence and the wild-type structure. The structural encoder uses an open-source unsupervised model to evaluate the probability of the mutant sequence folding into the wild-type structure, and outputs a tensor III containing protein structure information. Tensor III is a one-dimensional output vector of length L, and this tensor III is the score of the unsupervised model. This score aims to evaluate the ability of the sequence to fold into the wild-type structure of a given protein. A higher score means that the mutations contained in the mutant sequence are more favorable than others.

[0016] In the structural encoder, for the wild-type structure, the open-source esm-if1 model is used to obtain the saturated single-point mutation scoring matrix, and the size of the saturated single-point mutation scoring matrix is L×20;

[0017] The mutant sequence is one-hot encoded to obtain an encoding matrix, and the size of the encoding matrix is L×20;

[0018] Calculate the cross-entropy of the saturated single-point mutation scoring matrix and the encoding matrix at each amino acid position. After softmax operation on the calculation results, tensor III is obtained. The elements in tensor III are the probabilities representing whether the amino acid at the corresponding position in the mutant sequence is the optimal amino acid recognized by the esm-if1 model.

[0019] The input of the attention layer is tensor IV representing protein sequence information. Tensor IV is formed by concatenating tensor I and tensor II after layer normalization. Tensor I and tensor II have the same size, so they can be directly aligned and concatenated and transformed into a single vector through the attention mechanism. The size of tensor IV is L×512;

[0020] In the attention layer, tensor IV will obtain sequence attention weights under the attention mechanism;

[0021] The acquisition of sequence attention weights is the attention mechanism in deep learning. Each amino acid in tensor IV encoding sequence information will obtain a weight evaluating its degree of relevance to other amino acids in the sequence. This weight is calculated from tensor IV encoding protein sequence information and a learnable parameter matrix;

[0022] The calculation formula is:

[0023]

[0024] where, r i is the vector representing each amino acid in tensor IV, and W a is a learnable parameter matrix of size h×1;

[0025] Tensor III will obtain structure attention weights under the attention mechanism;

[0026] The structural attention weight is a tensor III encoding protein structure information and a parameter matrix to be learned together to form the structural attention weight, which is used to evaluate the association degree between each amino acid and the remaining amino acids in the output of the structure encoder;

[0027] The general formula for calculating the structural attention weight is:

[0028]

[0029] where p’ i represents the real number output at the position of each amino acid in tensor III;

[0030] Take the average of the sequence attention weight and the structural attention weight as the joint attention weight;

[0031] The attention layer performs weighted summation according to the joint attention weight and tensor IV to output an aggregated vector, and the size of the aggregated vector is 1×512;

[0032] The input of the output layer is the aggregated vector and the score of the unsupervised model, and the score of the unsupervised model is tensor III;

[0033] In the output layer, first process the aggregated vector with the ReLU function to obtain a hidden vector;

[0034] Calculate the dynamic weight using the Sigmoid function according to the hidden vector and the score of the unsupervised model, and this dynamic weight represents the degree of trust in the score of the unsupervised model;

[0035] Use a linear layer to calculate the mutation effect score of the hidden vector;

[0036] Finally, take the dynamic weight × mutation effect score + (1 - dynamic weight) × score of the unsupervised model as the output of the output layer.

[0037] The present invention also provides a prediction method for a deep learning system integrating sequence and structural features based on protein engineering, and the steps are as follows:

[0038] Train the deep learning model:

[0039] Obtain training data:

[0040] The training data includes the wild-type sequence, wild-type structure, mutation content of the protein, and the numerical score for evaluating the protein traits after mutation;

[0041] The mutation content refers to which amino acid at which position in the wild-type sequence mutates into which amino acid, and a piece of mutation data is not;

[0042] Divide the training set and the test set:

[0043] Divide the mutated content in the training data into a training set and a validation set at a ratio of 8:2;

[0044] Train the model:

[0045] Generate a set of mutated sequences from the wild-type sequence according to each mutated content in the training set;

[0046] Input each mutated sequence in the set of mutated sequences into the deep learning model for iterative training in turn;

[0047] At each iteration: the mutated sequence is encoded by the local encoder into a tensor I with the evolutionary information of homologous proteins;

[0048] The mutated sequence is encoded by the global encoder into a tensor II containing the common biochemical characteristics and evolutionary information of proteins;

[0049] The mutated sequence is encoded by the structure encoder into a tensor III containing the protein structure information, and the tensor III is the score of the unsupervised model;

[0050] After layer normalization of tensor I and tensor II respectively, they are concatenated into a tensor IV representing protein sequence information;

[0051] In the attention layer, tensor IV obtains the sequence attention weights under the attention mechanism;

[0052] Tensor III obtains the structure attention weights under the attention mechanism;

[0053] The average value of the sequence attention weights and the structure attention weights is used as the joint attention weight;

[0054] The aggregated vector obtained by the weighted sum of the joint attention weight and tensor IV is the output of the attention layer;

[0055] The aggregated vector is input into the output layer. In the output layer, the aggregated vector is first processed by the ReLU function to obtain a hidden vector;

[0056] The hidden vector and the score of the unsupervised model are used to calculate the dynamic weight using the Sigmoid function;

[0057] The hidden vector obtains the mutation effect score in the linear layer;

[0058] The mutation effect score and the score of the unsupervised model obtain the mutation prediction score of the mutated sequence under the distribution of the dynamic weight;

[0059] Calculate the loss function for the mutation prediction score and the numerical score for evaluating the protein characteristics after mutation. The form of the loss function used is the root mean square error (MSE), and update the parameters of the deep learning model to obtain the updated deep learning model;

[0060] Input the next mutant sequence in the mutant sequence set into the updated deep learning model for retraining until the loop ends, and obtain the trained deep learning model;

[0061] Use the validation set to validate the trained deep learning model to obtain the validated deep learning model;

[0062] Target mutation prediction:

[0063] Given the target mutation content, input the target mutant sequence into the validated deep learning model to obtain the mutation prediction score.

[0064] As a preferred embodiment of the present invention, the acquisition method of the training data is related to proteins. For proteins with sufficient experimental data, directly use the experimental data as the training data;

[0065] For proteins with insufficient or even missing experimental data, use the data augmentation strategy to obtain the training data.

[0066] As a preferred embodiment of the present invention, the data augmentation strategy is to select an unsupervised model;

[0067] Obtain a large amount of low-site mutant data of the protein by given mutation content;

[0068] Score each low-site mutant data using the unsupervised model to obtain the numerical score for evaluating the protein characteristics after mutation;

[0069] The low-site mutant data and the corresponding numerical scores for evaluating the protein characteristics after mutation are used as the training data,

[0070] The trained deep learning model predicts high-dimensional mutations.

[0071] As a preferred embodiment of the present invention, for proteins with missing experimental data, the unsupervised model is selected empirically according to the protein and the protein characteristics; for proteins with insufficient experimental data, first use each unsupervised model to predict all mutations in the existing experimental data respectively, and select the unsupervised model with the highest ranking correlation between the prediction results and the experimental results.

[0072] As a preferred embodiment of the present invention, for proteins with missing experimental data, directly use the low-site mutant data and the corresponding numerical scores for evaluating the protein characteristics after mutation as the training data to train the deep learning model, and perform high-dimensional mutation prediction after training;

[0073] For proteins with insufficient experimental data, use the low-site mutant data and the corresponding numerical scores for evaluating the protein characteristics after mutation as the training data to pre-train the deep learning model;

[0074] Then, the pre-trained deep learning model is retrained using the available experimental data.

[0075] After training is completed, high-dimensional mutations are predicted.

[0076] In the present invention, a deep learning model for predicting protein mutation effects by integrating sequence and structural information is first established. Then, combined with specific data augmentation strategies, the dependence of the deep learning model on the experimental sample size is reduced. Specifically, a large number of low-quality prediction results from unsupervised models are first used to pre-train the deep learning model, and then for those with experimental results, a limited number of high-quality experimental results are used to fine-tune the model. Experiments show that when the amount of experimental data for subsequent fine-tuning is less than 40 or there is no experimental data, the deep learning model obtained only through pre-training can achieve very high accuracy in the task of predicting the effects of higher-order mutations (especially deep mutations with more than 4 mutation sites). BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0078] Figure 1 It is the architecture diagram of the deep learning model of the present invention.

[0079] Figure 2 It is the architecture diagram of the local encoder of the present invention.

[0080] Figure 3 It is the architecture diagram of the structure encoder of the present invention.

[0081] Figure 4 It is the prediction flow chart of the present invention.

[0082] Figure 5 It is the comparison chart of the ranking correlation between the prediction results and the experimental results of the model trained with GFP protein under different amounts of experimental data of the present invention; among them, Figures A - D respectively show the ranking correlation between the prediction results and the experimental results of the model fine-tuned using 40, 100, 400, and 1084 single-site experimental data for the mutation effects at 2 - 8 sites.

[0083] Figure 6This is the model result of pre-training using the prediction results of the unsupervised model when there is no experimental data in the present invention. Among them, Figure A is the rank correlation between the prediction results and experimental results of the pre-trained model and the ESM-IF1 unsupervised model for the mutation effects at positions 2-8; Figure B is the rank correlation between the prediction results and experimental results of the pre-trained model and the ProGen2 unsupervised model for the mutation effects at positions 2-8. Detailed implementation mode

[0084] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0085] Embodiment:

[0086] A deep learning system integrating sequence and structural features based on protein engineering, as Figure 1 shown, includes a local encoder, a global encoder, a structure encoder, an attention layer, and an output layer;

[0087] The input of the local encoder is the mutant sequence. The local encoder uses the multiple sequence alignment method to encode the mutant sequence and outputs a tensor I encoding the evolutionary information of homologous proteins. The size of the tensor I is L×256.

[0088] In the local encoder,

[0089] The mutant sequence is converted into a tensor I to be trained through the Bi-LSTM layer 1 , the tensor I 1 has a size of L×128; this tensor I 1 formally satisfies the inclusion of amino acid self-constraints and coupling constraints between amino acids;

[0090] Use the multiple sequence alignment method to obtain the tensor I of the homologous constraint relationship of the wild-type sequence 2 , the tensor I 2 has a size of L×128; the homologous sequence constraint relationship includes amino acid self-constraints and coupling constraints between amino acids;

[0091] After splicing the tensor I 1 with the tensor I 2 , the tensor I with the evolutionary information of homologous proteins is obtained. The tensor I is the MSA-encoded sequence.

[0092] Since many existing models currently use multiple sequence alignment methods to extract the constraint relationships between residues during the evolution process and encode such information as important features for guiding mutation prediction. The schematic diagram of the principle specifically used in this embodiment is as Figure 2 shown;

[0093] This method will first search for homologous sequences of the given protein and then use HHsuite to construct sequence alignment information. After that, this method uses a statistical model to identify evolutionary couplings, and this model learns the homologous sequence alignment generation model by using a Markov random field. In this model, the probability of each sequence depends on an energy function.

[0094] This energy function is defined as the sum of the single-site constraint e i and all pairwise coupling constraints e ij :

[0095] E(x) = ∑ i e i (x i ) + ∑ i≠j e ij (x i , x j );

[0096] where i and j are the position numbers of residues on the sequence.

[0097] Therefore, the i-th amino acid x i in the sequence will be encoded into a vector, and the elements in the vector are set as the single-site energy term e i (x i ) and the pairwise coupling term e ij (x i , x j ), and j (j = 1,..., n) is the traversal of all residues (a total of n) in the sequence.

[0098] And these coupling parameters e i and e ij can be estimated by the regularized maximum pseudo-natural algorithm (completed by the open-source software CCMPred).

[0099] Finally, each amino acid in the homologous sequence will be represented as a vector with a length of L + 1 (L is the sequence length), so the entire input homologous sequence can be represented as an L × (L + 1) matrix by concatenating the vectors of each amino acid.

[0100] Since the representation length of the local evolutionary information of each amino acid is similar to the sequence length, its vector with a length of (L + 1) will be transformed into a vector with a fixed length of 128 through a fully connected layer to avoid overfitting problems, and this vector is the tensor I 2 .

[0101] The mutant sequence passes through the Bi-LSTM layer (Bidirectional Long Short-Term Memory network layer) and is transformed into a matrix of size L×128 for initial randomization. This matrix is the tensor I. 1 The parameters of this Bi-LSTM layer are randomly generated and will gradually converge during the training process.

[0102] The tensor I (size L×256) obtained by concatenating the above two matrices is the output result of this module.

[0103] The input of the global encoder is the mutant sequence. The global encoder uses a protein language model to encode the mutant sequence and outputs a tensor II that encodes the common biochemical features and evolutionary information of proteins. The size of tensor II is L×256.

[0104] The protein language model used is a protein language model that has completed large-scale training. After being trained with a large amount of data (the sequence data volume exceeds tens of millions), the protein language model will have excellent feature extraction ability. Moreover, the protein language model used is an unsupervised protein language model. The protein language model trained with a large amount of data can fully represent the biological characteristics and evolutionary diversity of proteins. These information more represent the commonalities of all protein sequences in the whole nature and are encoded into the tensor output by this module. Since this protein language model is unsupervised, it can be directly used in the present invention. After inputting the mutant sequence, the encoded tensor II can be obtained.

[0105] The protein language model will first encode the mutant sequence into a vector, and the encoding process will incorporate the connections between amino acids. This protein language model is pre-trained on the UniRef50 database with 250 million sequences. During training, the protein language model will adjust the parameters according to the differences between the encoded sequences with some amino acids masked and the encoded complete sequences.

[0106] The input of the structure encoder is the mutant sequence and the wild-type structure. The structure encoder uses an open-source unsupervised model to evaluate the probability of the mutant sequence folding into the wild-type structure, and outputs a tensor III containing protein structure information. Tensor III is a one-dimensional output vector of length L. This tensor III is the score of the unsupervised model, and this score aims to evaluate the ability of this sequence to fold into the wild-type structure of a given protein. A higher score means that the mutations contained in this mutant sequence are more favorable than others.

[0107] In the structure encoder, for the wild-type structure, an open-source esm-if1 model is used to obtain a saturation single-point mutation scoring matrix, and the size of the saturation single-point mutation scoring matrix is L×20.

[0108] Perform one-hot encoding on the mutant sequence to obtain an encoding matrix, where the size of the encoding matrix is L×20;

[0109] Calculate the cross-entropy of the saturation single-point mutation scoring matrix and the encoding matrix at each amino acid position. After performing softmax on the calculation results, obtain Tensor III. The elements in Tensor III are the probabilities representing whether the amino acid at the corresponding position in the mutant sequence is the optimal amino acid recognized by the esm-if1 model.

[0110] Specifically, when scoring, the present invention uses the esm-if1 model because the esm-if1 model will evaluate the probability of a given sequence folding into a given structure for a given structure. The ratio of the evaluation results of the mutant sequence and the wild-type sequence here is the score we use.

[0111] Esm-if1 will give a higher evaluation to sequences that are more likely to fold into the original structure. Then, if the mutant sequence has a higher score in the esm-if1 model, it means that this mutation is more beneficial for the sequence to fold into the original structure. This is a beneficial mutation, and the score used in our model will be higher.

[0112] Specifically, as Figure 3 shown, all possible single-point mutations at each position in the sequence will have corresponding scores (this score is the score of the corresponding mutant sequence), and their distribution will form a matrix of size L×20 (L is the sequence length), and this distribution is the saturation single-point mutation scoring matrix.

[0113] Then, perform one-hot encoding on the mutant sequence to obtain an encoding matrix, where the size of the encoding matrix is L××20. The element representing the amino acid at this position in each row is 1, and the rest are 0;

[0114] Calculate the cross-entropy of the saturation single-point mutation scoring matrix and the encoding matrix at each amino acid position. After performing softmax on the calculation results, obtain Tensor III. Tensor III is a one-dimensional output vector of length L, and the elements in Tensor III are the probabilities representing whether the amino acid at the corresponding position in the mutant sequence is the optimal amino acid recognized by the esm-if1 model.

[0115] The input of the attention layer is Tensor IV representing protein sequence information. Tensor IV is formed by splicing the normalized Tensor I and Tensor II. Tensor I and Tensor II have the same size, so they can be directly aligned and spliced and transformed into a single vector through the attention mechanism. The size of Tensor IV is L×512;

[0116] In the attention layer, Tensor IV will obtain sequence attention weights under the attention mechanism;

[0117] The acquisition of sequence attention weights is an attention mechanism in deep learning. Each amino acid in the tensor IV encoding sequence information will obtain a weight that evaluates its degree of relevance to other amino acids in the sequence. This weight is calculated from the tensor IV encoding protein sequence information and the parameter matrix to be learned;

[0118] The general calculation formula is:

[0119]

[0120] where r i is the vector representing each amino acid in the tensor IV, and W a is the parameter matrix to be learned with a size of h×1;

[0121] The tensor III will obtain structural attention weights under the attention mechanism;

[0122] For the structural attention weights, they are jointly composed of the tensor III encoding protein structure information and the parameter matrix to form the structural attention weights, which are used to evaluate the degree of association between each amino acid and the remaining amino acids in the output of the structure encoder;

[0123] The general calculation formula for the structural attention weights is:

[0124]

[0125] where p’ i represents the real number output at the position of each amino acid in the tensor III;

[0126] Take the average of the sequence attention weights and the structural attention weights as the joint attention weights; the general calculation formula for the joint attention weights is:

[0127] w = <w 1 , w 2 ,..., w L >;

[0128]

[0129] The attention layer performs weighted summation output of the aggregated vector according to the joint attention weights and the tensor IV, and the size of the aggregated vector is 1×512;

[0130] The representation of the aggregated vector is:

[0131]

[0132] The input of the output layer is the aggregated vector and the score of the unsupervised model, and the score of the unsupervised model is the tensor III;

[0133] In the output layer, first apply the ReLU function to the aggregated vector to obtain the hidden vector;

[0134] Calculate the dynamic weight using the Sigmoid function based on the hidden vector and the score of the unsupervised model. This dynamic weight indicates the degree of trust in the score of the unsupervised model. Here, the score of the unsupervised model is the ability of the mutant sequence mentioned in the structure encoder to fold into the given protein wild-type structure, that is, Tensor III;

[0135] Use a linear layer to calculate the mutant effect score of the hidden vector;

[0136] Finally, take the dynamic weight × mutant effect score + (1 - dynamic weight) × score of the unsupervised model as the output of the output layer.

[0137] Use the prediction method of the deep learning system that integrates sequence and structure features based on protein engineering. Since the acquisition method of training data is related to proteins, for proteins with sufficient experimental data, directly use the experimental data as training data;

[0138] For proteins with insufficient or even missing experimental data, use data augmentation strategies to obtain training data.

[0139] As Figure 4 shown, so below it will be divided into the prediction method of directly using experimental data and the prediction method of using data augmentation strategies.

[0140] For proteins with sufficient experimental samples, the steps are:

[0141] Train the deep learning model:

[0142] Obtain training data:

[0143] The training data is directly the existing experimental data, including the wild-type sequence of the protein, the wild-type structure, the mutation content, and the numerical score for evaluating the protein characteristics after mutation;

[0144] The mutation content refers to which amino acid at which position in the wild-type sequence mutates into which amino acid. One piece of experimental data is not necessarily a single-point mutation, but can also be an integration of multiple-point mutations. After the mutation occurs, the new protein will show differences in thermal stability or a certain function. The experiment will use specific numerical values to evaluate certain properties of these mutants, and the higher or lower the score indicates the strength of this property.

[0145] Divide the training set and the test set:

[0146] Divide the mutation content in the training data into a training set and a validation set according to a ratio of 8:2;

[0147] Train the model:

[0148] The wild-type sequence generates a set of mutant sequences according to each mutation content in the training set;

[0149] Each mutant sequence in the set of mutant sequences is sequentially input into the deep learning model for iterative training;

[0150] At each iteration: the mutant sequence is encoded by the local encoder into a tensor I with the evolutionary information of homologous proteins;

[0151] The mutant sequence is encoded by the global encoder into a tensor II containing the common biochemical characteristics and evolutionary information of proteins;

[0152] The mutant sequence is encoded by the structure encoder into a tensor III containing the protein structure information, and the tensor III is the score of the unsupervised model;

[0153] After layer normalization, tensor I and tensor II are concatenated into a tensor IV representing the protein sequence information;

[0154] In the attention layer, tensor IV obtains the sequence attention weights under the attention mechanism;

[0155] Tensor III obtains the structure attention weights under the attention mechanism;

[0156] The average value of the sequence attention weights and the structure attention weights is used as the joint attention weight;

[0157] The aggregated vector obtained by the weighted summation of the joint attention weight and tensor IV is the output of the attention layer;

[0158] The aggregated vector is input into the output layer. In the output layer, the aggregated vector is first processed by the ReLU function to obtain a hidden vector;

[0159] The hidden vector and the score of the unsupervised model are used to calculate the dynamic weights using the Sigmoid function;

[0160] The hidden vector obtains the mutant effect score in the linear layer;

[0161] The mutant effect score and the score of the unsupervised model obtain the mutant prediction score of the mutant sequence under the distribution of the dynamic weights;

[0162] The mutant prediction score and the numerical score of the mutant-post-evaluation protein trait are used to calculate the loss function. The form of the loss function is the root mean square error (MSE), and the parameters of the deep learning model are updated to obtain the updated deep learning model;

[0163] The next mutant sequence in the set of mutant sequences is input into the updated deep learning model for retraining until the loop ends, and the trained deep learning model is obtained;

[0164] Use the validation set to validate the trained deep learning model to obtain a validated deep learning model;

[0165] Target mutation prediction:

[0166] Given the target mutation content, the target mutation sequence is input into the verified deep learning model to obtain the mutation prediction score.

[0167] Since directed evolution only requires deep learning to guide the selection of mutation sites, the absolute value of the model prediction results is not very meaningful. The prediction results of all target mutations are sorted according to the scores, and the top mutations are considered by the model to have a higher probability of causing beneficial changes in the protein. The specific direction of the change depends largely on the protein properties measured by the mutation scores in the experimental data used for training.

[0168] For the prediction method using data augmentation strategy, the steps are:

[0169] For proteins with insufficient or even missing experimental data, data augmentation strategies are used to obtain training data.

[0170] The data enhancement strategy is suitable for two situations: a small amount of experimental data and missing experimental data.

[0171] For situations where there is a small amount of experimental data;

[0172] For the selection of unsupervised models, each unsupervised model is first used to predict all mutations in the existing experimental data, and the unsupervised model with the highest ranking correlation between the prediction results and the experimental results is selected;

[0173] Then, a large amount of low-level mutation data of the protein is obtained through the selected unsupervised model by given mutation content;

[0174] The mutation data of each low-level position are scored to obtain a numerical score for evaluating the protein characteristics after mutation;

[0175] Then, the low-level mutation data and the corresponding numerical scores for evaluating protein properties after mutation are used as training data to pre-train the deep learning model;

[0176] Next, the pre-trained deep learning model is trained again using the available experimental data;

[0177] After training, high-dimensional mutations are predicted;

[0178] Given high-dimensional mutation content, the high-dimensional mutation sequence is input into the deep learning model after secondary training to obtain the prediction result.

[0179] For the case of proteins with missing experimental data, the choice of unsupervised model is empirically selected based on the protein and protein characteristics; a large number of low-site mutation data of the protein are obtained by given mutation content;

[0180] Score each low-site mutation data using an unsupervised model to obtain a numerical score for evaluating protein characteristics after mutation;

[0181] The low-site mutation data and the corresponding numerical scores for evaluating protein characteristics after mutation are used as training data; the deep learning model obtained after training predicts high-dimensional mutations;

[0182] Given high-dimensional mutation content, and input the high-dimensional mutation sequence into the deep learning model after secondary training to obtain the prediction result.

[0183] Taking the GFP protein as an example, the present invention conducts training on the model with different amounts of experimental data, and uses the trained deep learning model to predict the mutation effects at 2-8 sites. The results are as follows Figure 5 shown. Among them, Figures A-D show the rank correlations between the prediction results and experimental results of the models fine-tuned with 40, 100, 400, and 1084 single-site experimental data on the GFP protein for the mutation effects at 2-8 sites. In each of the two rectangular bar groups in each figure, the right part represents the rank correlation between the prediction result pre-trained with the data generated by the ESM-IF1 unsupervised model in advance and the experimental result, and the left part represents the rank correlation between the prediction result of the model trained only with experimental data without using the data generated by the unsupervised model and the experimental result.

[0184] From Figure 5 it can be found that pre-training the model with the data generated by the unsupervised model will greatly improve the accuracy of the final prediction result, and this improvement is more obvious when the experimental data is less (the data volume is about 40).

[0185] Similarly, the present invention also takes the GFP protein as an example to show the model results of pre-training with the prediction results of the unsupervised model when there is no experimental data, as follows Figure 6 shown. Figures A and B show the rank correlations between the prediction results and experimental results of the pre-trained model and the corresponding unsupervised model used by it for the mutation effects at 2-8 sites. For the left part in the two rectangular bar groups, it is the result of the pre-trained model, and the right part is the result directly predicted by the unsupervised model. And the unsupervised model used in Figure A is ESM-IF1, and the model used in Figure B is ProGen2.

[0186] From Figure 6 it can be seen that when the experimental data is completely missing, the model trained only with the results generated by the unsupervised model can also achieve a relatively high prediction accuracy. AndFigure 6 The differences between Figure A and Figure B in the Chinese version prove that choosing different unsupervised models will produce different results. Therefore, using a small amount of experimental data to screen unsupervised models is of great help in improving the prediction accuracy, and the amount of experimental data required for this screening work is far less than that required by other supervised models, and it can be borne by any ordinary biochemical laboratory.

[0187] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0188] As mentioned above, only the specific preferred embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A deep learning system based on protein engineering integrating sequence and structural features, characterized in that: it includes a local encoder, a global encoder, a structure encoder, an attention layer, and an output layer; the input of the local encoder is the mutant sequence. The local encoder uses the multiple sequence alignment method to encode the mutant sequence and outputs a tensor I encoding the evolutionary information of homologous proteins. The size of tensor I is L×256; the input of the global encoder is the mutant sequence. The global encoder uses the protein language model to encode the mutant sequence and outputs a tensor II encoding the common biochemical features and evolutionary information of proteins; the size of tensor II is L×256; the input of the structure encoder is the mutant sequence and the wild-type structure. The structure encoder uses an open-source unsupervised model to evaluate the probability of folding the mutant sequence into the wild-type structure and outputs a tensor III containing protein structure information. Tensor III is a one-dimensional output vector of length L; the input of the attention layer is a tensor IV representing protein sequence information. Tensor IV is formed by concatenating tensor I and tensor II after layer normalization. The size of tensor IV is L×512; in the attention layer, tensor IV will obtain sequence attention weights under the attention mechanism; tensor III will obtain structure attention weights under the attention mechanism; the average value of the sequence attention weights and the structure attention weights is used as the joint attention weight; the attention layer outputs an aggregated vector according to the joint attention weight and tensor IV. The size of the aggregated vector is 1×512; the input of the output layer is the aggregated vector and the score of the unsupervised model; in the output layer, first process the aggregated vector with the ReLU function to obtain a hidden vector; calculate the dynamic weight according to the hidden vector and the score of the unsupervised model using the Sigmoid function. This dynamic weight indicates the degree of trust in the score of the unsupervised model; use a linear layer to calculate the mutant effect score of the hidden vector; finally, use dynamic weight × mutant effect score + (1 - dynamic weight) × score of the unsupervised model as the output of the output layer.

2. The deep learning system based on protein engineering integrating sequence and structural features according to claim 1, characterized in that: in the local encoder, The mutant sequence is converted into a tensor I to be trained through the Bi-LSTM layer 1 , the tensor I 1 has a size of L×128; this tensor I 1 formally satisfies the constraints including the self-constraints of amino acids and the coupling constraints between amino acids; Obtain the tensor I of the homologous constraint relationship of the wild-type sequence using the multiple sequence alignment method 2 , the tensor I 2 has a size of L×128; the homologous sequence constraint relationship includes the self-constraint of amino acids and the coupling constraint between amino acids; Tensor I 1 and Tensor I 2 After splicing, Tensor I with the evolutionary information of homologous proteins is obtained, and Tensor I is the MSA coding sequence.

3. The deep learning system based on protein engineering integrating sequence and structural features according to claim 1, characterized in that: in the structure encoder, for the wild-type structure, use the open-source esm-if1 model to obtain a saturated single-point mutation scoring matrix. The size of the saturated single-point mutation scoring matrix is L×20; perform one-hot encoding on the mutant sequence to obtain an encoding matrix. The size of the encoding matrix is L×20; calculate the cross-entropy at each amino acid position between the saturated single-point mutation scoring matrix and the encoding matrix. After softmax of the calculation results, obtain tensor III. The elements in tensor III represent whether the amino acid at the corresponding position in the mutant sequence is the optimal amino acid identified by the esm-if1 model.

4. A prediction method for the deep learning system based on protein engineering integrating sequence and structural features according to any one of claims 1-3, characterized in that: Training a deep learning model: Obtaining training data: The training data includes the wild-type sequence of the protein, the wild-type structure, the mutation content, and the numerical score for evaluating the protein characteristics after mutation; The mutation content refers to which amino acid at which position in the wild-type sequence is mutated into which amino acid; Dividing the training set and the test set: Dividing the mutation content in the training data into a training set and a validation set; Training the model: The wild-type sequence generates a set of mutant sequences according to each mutation content in the training set; Each mutant sequence in the set of mutant sequences is sequentially input into the deep learning model for iterative training; At each iteration: The mutant sequence is encoded by the local encoder into a tensor I with the evolutionary information of homologous proteins; The mutant sequence is encoded by the global encoder into a tensor II containing the common biochemical characteristics and evolutionary information of the protein; The mutant sequence is encoded by the structure encoder into a tensor III containing the protein structure information, and tensor III is the score of the unsupervised model; Tensor I and tensor II are respectively layer-normalized and then concatenated into a tensor IV representing the protein sequence information; In the attention layer, tensor IV obtains the sequence attention weights under the attention mechanism; Tensor III obtains the structure attention weights under the attention mechanism; The average value of the sequence attention weights and the structure attention weights is used as the joint attention weight; The aggregated vector obtained by the weighted sum of the joint attention weight and tensor IV is the output of the attention layer; The aggregated vector is input into the output layer. In the output layer, the aggregated vector is first processed by the ReLU function to obtain a hidden vector; The hidden vector and the score of the unsupervised model are used to calculate the dynamic weight using the Sigmoid function; The hidden vector obtains the mutation effect score in the linear layer; The mutation effect score and the score of the unsupervised model obtain the mutation prediction score of the mutant sequence under the distribution of the dynamic weight; The mutation prediction score and the corresponding numerical score for evaluating the protein characteristics after mutation calculate the loss function and update the parameters of the deep learning model to obtain the updated deep learning model; Input the next mutant sequence in the set of mutant sequences into the updated deep learning model for retraining until the loop ends to obtain the trained deep learning model; Use the validation set to validate the trained deep learning model to obtain the validated deep learning model; Target mutation prediction: Given the target mutation content, input the target mutant sequence into the validated deep learning model to obtain the mutation prediction score.

5. The prediction method of the deep learning system based on the integrated sequence and structure features of protein engineering according to claim 4, characterized in that: The acquisition method of the training data is related to the protein. For proteins with sufficient experimental data, the experimental data is directly used as the training data; For proteins with insufficient or even missing experimental data, data augmentation strategies are used to obtain the training data.

6. The prediction method of the deep learning system based on the integrated sequence and structure features of protein engineering according to claim 5, characterized in that: The data augmentation strategy is to select an unsupervised model; A large amount of low-site mutation data of the protein is obtained by given mutation content; Score each low-site mutation data using an unsupervised model to obtain a numerical score for evaluating protein characteristics after mutation; The low-site mutation data and the corresponding numerical scores for evaluating protein characteristics after mutation are used as training data.

7. The prediction method of the deep learning system based on integrated sequence and structural features of protein engineering according to claim 6, characterized in that: For proteins with missing experimental data, the unsupervised model is selected empirically based on the protein and its characteristics; for proteins with insufficient experimental data, each unsupervised model is first used to predict all mutations in the existing experimental data, and the unsupervised model with the highest ranking correlation between the prediction results and the experimental results is selected.

8. The prediction method of the deep learning system based on integrated sequence and structural features of protein engineering according to claim 6, characterized in that: For proteins with missing experimental data, directly use the low-site mutation data and the corresponding numerical scores for evaluating protein characteristics after mutation as training data to train the deep learning model, and predict high-dimensional mutations after training; For proteins with insufficient experimental data, use the low-site mutation data and the corresponding numerical scores for evaluating protein characteristics after mutation as training data to pre-train the deep learning model; Then use the available experimental data to perform secondary training on the pre-trained deep learning model; Predict high-dimensional mutations after training is completed.

Citation Information

Patent Citations

  • Attention-based neural network to predict peptide binding, presentation, and immunogenicity

    US20220122690A1

  • Protein database search using learned representations

    US20220165356A1