A Targeted Polypeptide Design Method Based on Multi-Task Pre-Training and Transfer Learning
By adopting multi-task autoregressive pre-training and interactively perceived transfer learning methods in targeted polypeptide design, a multi-pulmonary generative pre-training model is constructed and the protein language model is used to extract features, which solves the problem of low accuracy and efficiency of peptide design in the existing technology, and achieves a more efficient targeted polypeptide design.
Patent Information
- Application Number
- CN202411141336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-20
AI Technical Summary
Existing deep learning algorithms lack the ability to understand the protein-polypeptide interaction relationship and multi-task learning to optimize feature extraction in targeted polypeptide design, resulting in low design accuracy and efficiency.
A targeted polypeptide design method based on multi-task autoregressive pre-training and interactive sensing transfer learning is adopted. By constructing a polypeptide generative pre-training model, deep features are extracted using the protein language model, and the transfer generation of polypeptide sequences is guided through interactive sensing attention to form a targeted polypeptide design model.
The interaction relationship between protein sequence-polypeptide sequence is effectively utilized, which improves the accuracy and efficiency of targeted polypeptide design, and enhances the robustness and generalization ability of the model.
Smart Images

Figure CN118969088B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for designing targeted polypeptides based on multi-task autoregressive pre-training and interaction-aware transfer learning, belonging to the field of polypeptide design. Background Art
[0002] In recent years, polypeptide drugs have become a hot spot in drug research and development due to their advantages such as high selectivity, high efficiency, low toxicity, and good biocompatibility. Polypeptide drugs are biological macromolecules composed of several to dozens of amino acids connected by peptide bonds, and can achieve therapeutic effects through specific interactions with proteins. However, designing efficient targeted polypeptide drugs is a complex and challenging problem, which requires considering various factors such as the sequence, structure, and function of polypeptides.
[0003] Recently, deep learning algorithms have achieved great success in the field of bioinformatics, and a large number of studies have been devoted to tasks such as protein structure prediction, protein function prediction, and protein-protein interaction prediction. However, deep learning algorithms for targeted polypeptide design are urgently needed to be developed. At the same time, due to the complex non-linear relationship between the sequence, structure, and function of polypeptides, existing deep learning algorithms lack an understanding of the protein-polypeptide interaction relationship and the ability to optimize and extract features from multi-task learning, thus limiting the accuracy and efficiency of targeted polypeptide design. Therefore, it is a challenging problem to propose a method for designing targeted polypeptides based on multi-task autoregressive pre-training and interaction-aware transfer learning. Summary of the Invention
[0004] To solve the above problems, the present invention provides a method for designing targeted polypeptides based on multi-task autoregressive pre-training and interaction-aware transfer learning, and the technical solution is as follows:
[0005] A method for designing targeted polypeptides of the present invention constructs a polypeptide generative pre-training model based on multi-task autoregressive pre-training, and then performs transfer learning on the polypeptide generative pre-training model to obtain a targeted polypeptide design model. The method includes:
[0006] Step 1: Obtain a large-scale polypeptide sequence data and protein-polypeptide pairing data, encode the protein sequence and polypeptide sequence to obtain vector representations of the protein sequence and polypeptide sequence, and divide them into a training set and a test set;
[0007] Step 2: Construct the polypeptide generative pre-training model based on multi-task autoregressive pre-training, and perform self-supervised multi-task training on the encoded large-scale polypeptide sequence data;
[0008] Step 3: Extract deep features from the vector representation of the protein sequence using a protein language model as the initial hidden state of the polypeptide generative pre-training model, and guide the migratory generation of the polypeptide sequence through interactive perception attention to obtain a trained targeted polypeptide design model;
[0009] Step 4: Use the trained targeted polypeptide design model to generate polypeptide sequences for the protein sequence to complete the automatic design of targeted polypeptides;
[0010] The targeted polypeptide design model includes: a sequence encoding module, a polypeptide transformation module, a multi-task prediction module, an interactive perception module, and a transfer learning module;
[0011] The sequence encoding module is used to encode the polypeptide sequence and the sequence; the polypeptide transformation module includes: masked multi-head attention, a normalization layer, a feed-forward layer, and a residual connection; the multi-task prediction module includes: a sequence prediction head, a secondary structure prediction head, and a polypeptide function prediction head; the interactive perception module is used to calculate the interaction relationship between the protein sequence and the polypeptide sequence during the transfer learning process to obtain an interactive perception representation; the transfer learning module is used to splice the protein sequence and the polypeptide sequence, input it into the pre-trained polypeptide generative pre-training model, and learn to generate the polypeptide sequence from the protein sequence based on the interactive perception module to complete the mapping from the protein sequence to the polypeptide sequence and obtain a trained targeted polypeptide design model.
[0012] Optionally, step 1 includes:
[0013] Step 11: Use Python pyfastx to extract polypeptide sequence information from the PeptideAtlas polypeptide map database to obtain a large-scale polypeptide sequence sample, and encode the large-scale polypeptide sequence sample using the Byte Pair Encoding method to obtain a vector representation of the large-scale polypeptide sequence;
[0014] Step 12: Use Python requests to batch download known protein complex structure data from the RCSB PDB database, analyze the protein complex structure data using Biopython, extract the complexes with 2 chains where one sequence length is less than 50 and the other sequence length is greater than 50, consider them as protein-polypeptide complexes, and extract the paired protein sequence and polypeptide sequence;
[0015] Step 13: Perform character encoding on the protein sequence, regard each amino acid letter as a number, and then vectorize based on the number to obtain a vector representation of the protein sequence.
[0016] Optionally, the processing process of the sequence encoding module includes:
[0017] Perform BPE encoding on large-scale polypeptide sequences, that is, regard the collected large-scale polypeptide data as a corpus, count the frequencies of different amino acid combinations in the corpus, store them as a vocabulary through the formation of frequency-amino acid combinations, and set the maximum length of the vocabulary, which can take any value;
[0018] Perform character encoding on protein sequences, that is, regard each amino acid letter as a number and then vectorize based on this number.
[0019] Optionally, the calculation process of the polypeptide transformation module includes:
[0020] First, define the vector of the polypeptide as h p , and obtain the corresponding query matrix Q based on this vector p , key matrix K p , and value matrix V p . Then, the attention score matrix is obtained by calculating the dot product of Q p and K p and then dividing by the scaling factor:
[0021]
[0022] After obtaining the attention scores, apply a masking operation to set the scores at positions that should not be attended to as negative infinity, and then apply the Softmax function to obtain the attention weight matrix:
[0023] Attentionweights = Softmax(Mask(Score))
[0024] where Mask represents the masking operation;
[0025] Finally, perform weighted summation on the value matrix with the attention weights to obtain the output matrix:
[0026]
[0027] Then, pass the output matrix through a normalization layer, a feed-forward layer, and a residual connection in sequence, which is specifically expressed as:
[0028]
[0029] where LN and FF represent layer normalization and feed-forward layer respectively.
[0030] Optionally, the multi-task prediction module includes: a sequence prediction head, a secondary structure prediction head, and a polypeptide function prediction head;
[0031] The sequence prediction head is used to predict the next amino acid corresponding to the current amino acid; the secondary structure prediction head is used to predict the secondary structure of the current amino acid; and the polypeptide function prediction head is used to predict the biological activity and function of the polypeptide.
[0032] Optionally, the multi-task pre-training loss function includes: a sequence prediction loss function L seq , a secondary structure prediction loss function L ss and a polypeptide function prediction loss function L fun , then the multi-task pre-training loss function is expressed as:
[0033] L multi-pretraining = λ seq L seq + λ ss L ss + λ func L func
[0034] where λ seq , λ ss , λ func respectively represent the weights of each task.
[0035] Optionally, the interaction-aware module calculates the interaction-aware attention of the feature vector h t of the target protein sequence and the feature vector h p of the polypeptide sequence, learns the interaction relationship between the protein sequence and the polypeptide sequence, and obtains the interaction-aware representation:
[0036]
[0037] where h i represents the interaction-aware representation.
[0038] A target polypeptide design system of the present invention is used to implement the target polypeptide design method as described above, and the system includes:
[0039] A dataset construction module, configured to obtain large-scale polypeptide sequence data and protein-polypeptide pairing data, encode the protein and polypeptide sequences to obtain vector representations of the protein sequence and the polypeptide sequence, and divide them into a training set and a test set;
[0040] A pre-training module, configured to construct a polypeptide generative pre-training model based on multi-task autoregressive pre-training, and perform self-supervised multi-task training on the encoded large-scale polypeptide sequence data;
[0041] The target polypeptide design model training module is configured to extract deep features from the vector representation of the protein sequence using the protein language model ESM-2 as the initial hidden state of the polypeptide generative pre-training model, and guide the migratory generation of the polypeptide sequence through interactive perception attention to obtain a trained target polypeptide design model;
[0042] Among them, ESM-2 is a Transformer-based language model developed by the Fundamental AI Research (FAIR) protein team of Facebook, focusing on the representation learning of protein sequences. It uses the attention mechanism to learn the interaction patterns between amino acid pairs in the input sequence and can predict its deep features from a single protein sequence. ESM-2 can be replaced by other protein language models, such as ProtTrans. ESM-2 is only a preferred solution in the examples of the present invention.
[0043] The target polypeptide generation module is configured to use the trained target polypeptide design model to generate a polypeptide sequence for the protein sequence to complete the automatic design of the target polypeptide.
[0044] An electronic device, mobile phone, and air conditioner according to the present invention include a memory and a processor;
[0045] The memory is used to store a computer program;
[0046] The processor is configured to implement the target polypeptide design method as described in any one of the above when executing the computer program.
[0047] A computer-readable storage medium according to the present invention has a computer program stored thereon, and when the computer program is executed by a processor, the target polypeptide design method as described in any one of the above is implemented.
[0048] The beneficial effects of the present invention are:
[0049] The present invention can pre-train large-scale polypeptide sequence data, learn potential sequence patterns and rules for polypeptide sequence generation; the present invention improves the existing generative neural network using multi-task autoregressive pre-training and interactive perception transfer learning, effectively utilizes the protein sequence-polypeptide sequence interaction relationship, and solves the problem of low accuracy in target polypeptide design; the present invention uses a multi-task prediction module to fully learn the cooperative interaction relationship between multi-tasks, and further improves the robustness and generalization ability of the target polypeptide design model. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0051] Figure 1 It is a schematic diagram of the basic process of a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0052] Figure 2 It is a schematic diagram of the polypeptide coding word representation of a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0053] Figure 3 It is a schematic diagram of the pre-training process of a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0054] Figure 4 It is a schematic diagram of the transfer learning of a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0055] Figure 5 It is a schematic diagram of the structure of the interaction-aware module of a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0056] Figure 6 It is a structural effect diagram of the polypeptide designed for the MDM4 protein by a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0057] Figure 7 It is a structural effect diagram of the polypeptide designed for the KOR protein by a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention.
[0058] Figure 8 It is an energy distribution diagram of the polypeptide designed for the GCGR protein by a targeted polypeptide design method based on multi-task autoregressive pre-training and interaction-aware transfer learning provided by an embodiment of the present invention. Detailed implementation manners
[0059] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0060] Example 1:
[0061] This embodiment provides a method for designing targeted polypeptides. A polypeptide generative pre-training model is constructed based on multi-task autoregressive pre-training, and then transfer learning is performed on the polypeptide generative pre-training model to obtain a targeted polypeptide design model. The method includes:
[0062] Step 1: Obtain a large-scale polypeptide sequence data and protein-polypeptide pairing data, encode the protein sequence and polypeptide sequence to obtain vector representations of the protein sequence and polypeptide sequence, and divide them into a training set and a test set;
[0063] Step 2: Construct a polypeptide generative pre-training model based on multi-task autoregressive pre-training, and perform self-supervised multi-task training on the encoded large-scale polypeptide sequence data;
[0064] Step 3: Extract deep features of the protein sequence vector representation using a protein language model as the initial hidden state of the polypeptide generative pre-training model, and guide the migratory generation of the polypeptide sequence through interactive perception attention to obtain a trained targeted polypeptide design model;
[0065] Step 4: Use the trained targeted polypeptide design model to generate polypeptide sequences for the target protein to complete the automatic design of targeted polypeptides;
[0066] The targeted polypeptide design model includes: a sequence encoding module, a polypeptide transformation module, a multi-task prediction module, an interactive perception module, and a transfer learning module;
[0067] The sequence encoding module is used to encode the polypeptide sequence and the sequence; the polypeptide transformation module includes: masked multi-head attention, a normalization layer, a feed-forward layer, and a residual connection; the multi-task prediction module includes: a sequence prediction head, a secondary structure prediction head, and a polypeptide function prediction head; the interactive perception module is used to calculate the interaction relationship between the protein sequence and the polypeptide sequence during the transfer learning process to obtain an interactive perception representation; the transfer learning module is used to splice the protein sequence and the polypeptide sequence, input them into the pre-trained polypeptide generative pre-training model, and learn to generate polypeptide sequences from the protein sequence based on the interactive perception module to complete the mapping from the protein sequence to the polypeptide sequence, and obtain a trained targeted polypeptide design model.
[0068] Embodiment 2:
[0069] This embodiment provides a method for designing targeted polypeptides based on multi-task autoregressive pre-training and interactive perception transfer learning. See Figures 1 to 5 , the method includes:
[0070] Step 1: Obtain large-scale polypeptide sequence data and protein-polypeptide pairing data through bioinformatics databases, encode the protein sequences and polypeptide sequences to obtain vector representations of the protein sequences and polypeptide sequences, and divide them into training sets and test sets.
[0071] Step 11: Use Python pyfastx to extract polypeptide sequence information from the PeptideAtlas database to obtain large-scale polypeptide sequence samples, and encode the large-scale polypeptide sequence samples using Byte Pair Encoding (BPE) to obtain vector representations of the large-scale polypeptide sequences.
[0072] Step 12: Use Python requests to batch download known protein complex structure data from the RCSB PDB database, analyze the protein complex structure data using Biopython, extract complexes with 2 chains where one sequence length is less than 50 and the other sequence length is greater than 50, consider them as protein-polypeptide complexes, extract the paired protein sequences and polypeptide sequences, perform character encoding on the protein sequences, that is, regard each amino acid letter as a number, and then vectorize based on this number to obtain vector representations of the protein sequences.
[0073] Step 2: Build a polypeptide generative pre-training model based on multi-task autoregressive pre-training, perform self-supervised multi-task training on large-scale polypeptide sequences, and obtain a trained polypeptide generative pre-training model.
[0074] It should be noted that the finally constructed targeted polypeptide design model in this embodiment includes: a sequence encoding module, a polypeptide transformation module, a multi-task prediction module, an interaction perception module, and a transfer learning module.
[0075] Specifically, refer to Figure 3 , the sequence encoding module performs BPE encoding on large-scale polypeptide sequences. Specifically: regard the collected large-scale polypeptide data as a corpus, count the frequencies of different amino acid combinations in the corpus, form a vocabulary by frequency-amino acid combination storage, set the maximum length of the vocabulary, and this length can take any value, such as 100, 1000, and then vectorize the polypeptide sequences based on this table.
[0076] The sequence encoding module also needs to perform character encoding on protein sequences, that is, regard each amino acid letter as a number, and then vectorize based on this number.
[0077] Furthermore, refer to Figure 3 , the polypeptide transformation module includes: masked multi-head attention, normalization layer, feed-forward layer, and residual connection. The calculation process of the polypeptide transformation module is as follows:
[0078] First, define the vector of the polypeptide as h p , and based on this vector, its corresponding query matrix Q p , key matrix K p , and value matrix V p can be obtained. Then, the attention score matrix can be obtained by calculating the dot product of Q p and K p and then dividing by the scaling factor:
[0079]
[0080] After obtaining the attention scores, apply a masking operation to set the scores at positions that should not be attended to as negative infinity, and then apply the Softmax function to obtain the attention weight matrix:
[0081] Attentonweights = Softmax(Mask(Score))
[0082] where Mask represents the masking operation.
[0083] Finally, perform a weighted sum of the value matrix with the attention weights to obtain the output matrix:
[0084]
[0085] Then, pass the output matrix through a normalization layer, a feed-forward layer, and a residual connection in sequence, which can be specifically expressed as:
[0086]
[0087] where LN and FF represent layer normalization and the feed-forward layer respectively.
[0088] Further, the multi-task prediction module includes: a sequence prediction head, a secondary structure prediction head, and a polypeptide function prediction head. The sequence prediction head is used to predict what the next amino acid corresponding to the current amino acid is, the secondary structure prediction head is used to predict what the secondary structure of the current amino acid is (such as α-helix, β-sheet), and the polypeptide function prediction head is used to predict the biological activity and function of the polypeptide (such as antibacterial, antiviral, antitumor, etc.).
[0089] Further, the multi-task pre-training loss function includes: a sequence prediction loss function L seq , a secondary structure prediction loss function L ss , and a polypeptide function prediction loss function L fun . Then, the multi-task pre-training loss function can be expressed as:
[0090] L multi-pretraining = λ seq L seq+λ ss L ss +λ func L func
[0091] where λ seq ,λ ss ,λ func respectively represent the weights of each task. The losses of different tasks can be calculated using a variety of loss functions, such as the cross-entropy loss function.
[0092] Step 3: Extract deep features from the vector representation of the protein sequence using the protein language model ESM-2 as the initial hidden state of the pre-trained polypeptide generative pre-trained model, and perform transfer learning on the pre-trained network model by guiding the transfer generation of the polypeptide sequence through interaction-aware attention.
[0093] The ESM-2 used in this embodiment is a Transformer-based language model developed by the Fundamental AI Research (FAIR) protein team of Facebook, which focuses on the representation learning of protein sequences. It uses the attention mechanism to learn the interaction patterns between amino acid pairs in the input sequence and can predict its deep features from a single protein sequence. ESM-2 can be replaced by other protein language models, such as ProtTrans. ESM-2 is just a preferred solution in the examples of the present invention.
[0094] It should be noted that Figure 4 ,the transfer learning process involves an interaction-aware module and a transfer learning module.
[0095] The calculation process of the interaction-aware module can be referred to Figure 5 ,this module calculates the interaction-aware attention between the deep features h t extracted by the protein language model ESM-2 and the feature vector h p of the polypeptide sequence, learns the interaction relationship between the target protein and the polypeptide, and obtains the interaction-aware representation:
[0096]
[0097] Furthermore, the calculation process of the transfer learning module includes: concatenating the protein sequence and the polypeptide sequence, inputting them into the pre-trained polypeptide generative pre-trained model, and generating the polypeptide sequence from the protein sequence based on the learning of the interaction-aware module to complete the mapping from the protein sequence to the polypeptide sequence, and obtaining the trained target polypeptide design model.
[0098] S4: Use the trained target polypeptide design model to generate polypeptide sequences for the protein sequences in the test set to complete the automatic and accurate design of target polypeptides.
[0099] Example 3:
[0100] In this example, polypeptides were designed using MDM4 protein, KOR protein, and GCGR protein and compared with known natural binding polypeptides through contrast tests. The experimental results were compared by scientific demonstration means to verify the actual effects of this method.
[0101] Among them, MDM4 protein is a key regulatory factor in the p53 signaling pathway. p53 is a tumor suppressor protein that plays an important role in maintaining genomic stability and preventing cancer occurrence. MDM4 can inhibit the transcriptional activity of p53, thereby regulating processes such as the cell cycle, apoptosis, and DNA repair. Therefore, MDM4 is an anti-cancer drug target with potential therapeutic value. Designing polypeptides that bind to MDM4 can interfere with the interaction between MDM4 and p53, restore the normal function of p53, and thus inhibit the growth and spread of cancer cells. KOR protein regulates physiological processes such as pain, depression, anxiety, addiction, and immune response. Designing polypeptides that bind to KOR can provide new ideas for the development of novel analgesics, antidepressants, and anti-anxiety drugs. GCGR protein binds to glucagon and regulates physiological processes such as blood glucose level, lipolysis, and insulin release. Designing polypeptides that bind to GCGR can block the interaction between glucagon and GCGR, reduce blood glucose levels, and thus provide new strategies for diabetes treatment.
[0102] Extract the sequence information of the corresponding proteins from the Uniport database and input it into the trained targeted polypeptide design model to obtain 2 generated polypeptide sequences for each target. At the same time, collect the known binding polypeptides of the target proteins for comparison.
[0103] First, use AlphaFold2 to predict the complex structures of MDM4 protein, KOR protein, and their corresponding polypeptide sequences, and obtain their binding confidence scores. See Figures 6 - 7 , the protein is represented by a cartoon structure, the polypeptide is shown by a ball-and-stick structure and characterized by a rectangular frame. It can be seen that the polypeptides designed by the present invention can bind to the same position as the natural polypeptides, indicating that their binding sites are the same and have the potential to perform the same functions; at the same time, the polypeptides designed by the present invention can obtain higher binding confidence scores, indicating that they may have higher binding affinities with the target proteins.
[0104] To further verify the effectiveness and generalization of the method, in this example, molecular dynamics simulations were performed on GCGR protein and its corresponding polypeptide to obtain its docking energy score. See Figure 8, the polypeptides designed in the present invention can achieve lower docking energy scores, indicating that they can bind more tightly to the target proteins and have stronger binding affinities. Therefore, the method for designing the targeting polypeptides of the present invention shows good performance in practical applications.
[0105] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software programs can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0106] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for designing a targeting polypeptide, characterized in that: The method constructs a polypeptide generative pre-training model based on multi-task autoregressive pre-training, and then performs transfer learning on the polypeptide generative pre-training model to obtain a targeted polypeptide design model. The method includes: Step 1: Obtain large-scale peptide sequence data and protein-peptide pairing data, encode protein sequences and peptide sequences, obtain vector representations of protein sequences and peptide sequences, and divide them into training sets and test sets; Step 2: constructing the peptide generative pre-training model based on multi-task autoregressive pre-training, and performing self-supervised multi-task training on the encoded large-scale peptide sequence data; Step 3: Extract deep features from the vector representation of the protein sequence using the protein language model ESM-2 as the initial hidden state of the peptide generation pre-training model, and guide the migration generation of the peptide sequence through interactive perception attention to obtain a trained targeted peptide design model; Step 4: Generate peptide sequences for protein sequences using the trained targeted peptide design model to complete the automatic design of targeted peptides; The targeted peptide design model includes: a sequence encoding module, a peptide transformation module, a multi-task prediction module, an interactive perception module, and a transfer learning module; The sequence encoding module is used to encode the polypeptide sequence and sequence; the polypeptide transformation module includes: masked multi-head attention, normalization layer, feedforward layer and residual connection; the multi-task prediction module includes: sequence prediction head, secondary structure prediction head and polypeptide function prediction head; the interaction perception module is used to calculate the interaction relationship between protein sequence and polypeptide sequence in the transfer learning process to obtain the interaction perception representation; the transfer learning module is used to splice the protein sequence and polypeptide sequence, input them into the pre-trained polypeptide generation pre-training model, generate polypeptide sequence from protein sequence based on the interaction perception module learning, complete the mapping from protein sequence to polypeptide sequence, and obtain the trained targeted polypeptide design model; The calculation process of the polypeptide conversion module includes: First, define the characteristic vector of the polypeptide sequence as h p , based on this vector, obtain its corresponding query matrix Q p , key matrix K p , value matrix V p , then the attention score matrix is calculated by Q p and K p The dot product is then divided by the scaling factor to obtain: After obtaining the attention score, a mask operation is applied to it, the scores of the positions that should not be noticed are set to negative infinity, and then the Softmax function is applied to obtain the attention weight matrix: in, Indicates mask operation; Finally, the value matrix is weighted and summed using the attention weights to obtain the output matrix: Then, the output matrix is passed through the normalization layer, the feedforward layer, and the residual connection in sequence, which can be specifically expressed as follows: in, and Represent the normalization layer and the feed-forward layer respectively; The calculation method of the interactive perception representation is: in, The feature vector representing the target protein sequence, A feature vector representing a peptide sequence.
2. The method for designing a targeting polypeptide according to claim 1, characterized in that: The step 1 comprises: Step 11: Use Python pyfastx to extract polypeptide sequence information from the polypeptide atlas database PeptideAtlas to obtain a large-scale polypeptide sequence sample, and use Byte Pair Encoding to encode the large-scale polypeptide sequence sample to obtain a vector representation of the large-scale polypeptide sequence; Step 12: Use Python requests to batch download known protein complex structure data from the RCSB PDB database, use Biopython to analyze the protein complex structure data, extract the complexes with two chains and one sequence length less than 50 and the other sequence length more than 50, which are considered to be protein-peptide complexes, and extract the paired protein sequences and polypeptide sequences therein; Step 13: Character encoding is performed on the protein sequence, each amino acid letter is regarded as a number, and then vectorization is performed based on the number to obtain a vector representation of the protein sequence.
3. The method for designing a targeting polypeptide according to claim 1, characterized in that: The processing process of the sequence encoding module includes: BPE encoding is performed on large-scale peptide sequences, that is, the collected large-scale peptide data is regarded as a corpus, the frequencies of different amino acid combinations in the corpus are counted, and the frequency-amino acid combination is formed and stored as a vocabulary, and the maximum length of the vocabulary is set, which can take any value; The protein sequence is character-encoded, that is, each amino acid letter is regarded as a number, and then vectorized based on the number.
4. The method for designing a targeting polypeptide according to claim 1, characterized in that: The multi-task prediction module includes: a sequence prediction head, a secondary structure prediction head and a polypeptide function prediction head; The sequence prediction head is used to predict the next amino acid corresponding to the current amino acid; the secondary structure prediction head is used to predict the secondary structure of the current amino acid; and the polypeptide function prediction head is used to predict the biological activity and function of the polypeptide.
5. The method for designing a targeting polypeptide according to claim 1, characterized in that: Multi-task pre-training loss functions include: sequence prediction loss function , secondary structure prediction loss function And the peptide function prediction loss function , then the multi-task pre-training loss function is expressed as: in , , Represent the weight of each task.
6. A targeting polypeptide design system, characterized in that: The system is used to implement the targeting polypeptide design method according to any one of claims 1 to 5, and the system comprises: A data set construction module is configured to obtain large-scale peptide sequence data and protein-peptide pairing data, encode protein sequences and peptide sequences, obtain vector representations of protein sequences and peptide sequences, and divide them into training sets and test sets; A pre-training module is configured to construct a peptide generative pre-training model based on multi-task autoregressive pre-training, and perform self-supervised multi-task training on the encoded large-scale peptide sequence data; The targeted peptide design model training module is configured to extract deep features from the vector representation of the protein sequence using a protein language model as the initial hidden state of the peptide generation pre-training model, and guide the migration generation of the peptide sequence through interactive perception attention to obtain a trained targeted peptide design model; The targeted peptide generation module is configured to generate peptide sequences for protein sequences using the trained targeted peptide design model to complete the automatic design of targeted peptides.
7. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the targeting polypeptide design method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the targeting polypeptide design method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method for rapidly analyzing protein and strong-polarity long amino acid sequence glycopeptides in biological sample
CN110441428A
Method and system for predicting protein-polypeptide binding site
CN113593631A