Polypeptide design method based on graph attention network
Through the method based on the graph attention network, the problem of cumbersome and low success rate of existing targeted protein peptide design methods is solved, and more efficient and accurate peptide design is achieved, which improves the design efficiency and success rate.
Patent Information
- Application Number
- CN202510353094.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing targeted protein peptide design methods have problems such as cumbersome process and low design success rate, which are difficult to meet the needs of practical applications.
Using a graph attention network-based method, a targeted protein polypeptide is designed by constructing a graph structure model of a protein, combining attention mechanisms and graph neural networks. This method includes steps such as dataset construction, protein information modeling, graph attention network construction, model training and peptide design.
It improves the accuracy and efficiency of peptide design, enhances the design success rate, and provides a more comprehensive tool chain to support the entire process from protein structure prediction to peptide design.
Smart Images

Figure CN120220792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and computer applications, and particularly to a polypeptide design method based on a graph attention network. Background Art
[0002] Polypeptides targeting proteins have broad application prospects, such as polypeptide drugs, enzyme inhibitors, peptide agonists, etc. However, for proteins of typical length, there can be as many as 20 200 different sequences. Predicting the interactions of such a large number of molecules requires a very large amount of computation. Coupled with the great complexity of the protein-polypeptide interaction interface, which involves various properties such as surface shape, charge, hydrophilicity, etc., the prediction of protein-polypeptide interactions remains a challenging task to date.
[0003] The design of protein-targeting polypeptides is essentially a generation task, whose goal is to generate polypeptide sequences that can interact with a specific input protein. The core of this task lies in creating new activities, behaviors, or structures by rationally designing protein molecules, thereby promoting the understanding of protein functions. Traditional protein engineering techniques, such as directed evolution and phage display screening, although can successfully design functional polypeptides in some cases, these methods usually rely on experimental screening, with a cumbersome process and low efficiency. In addition, traditional methods are difficult to systematically explore the sequence space, especially when designing polypeptides with new structures and new functions, they often face great limitations.
[0004] In recent years, methods such as protein structure prediction based on machine learning have made remarkable progress. Especially the development of large models has provided powerful tools for protein design. These large models are usually based on the transformer architecture, and the core lies in the introduction of the attention mechanism. The attention mechanism selectively screens out a small amount of important information from a large amount of information by calculating attention coefficients and focuses on this important information, thereby ignoring most of the unimportant information. This mechanism performs well in processing complex protein sequence and structure data and can effectively capture the key features in protein-polypeptide interactions. In addition, the data structure of "graph" has a natural similarity to the three-dimensional structure of proteins, so using the graph structure for data modeling of proteins has great advantages. Graph neural networks are designed based on the graph structure, and using such networks can effectively capture the topological relationships and local features in protein structures, providing new ideas for protein design.
[0005] However, although machine learning and deep learning technologies have made significant progress in the field of protein structure prediction and design, in the field of polypeptide design, there is no mature machine learning method that combines the attention mechanism with graph neural networks. Existing methods often rely on a single machine learning model and lack a complete tool chain to support the entire process from protein structure prediction to polypeptide design. In addition, methods based on other machine learning models have obvious deficiencies in design accuracy and success rate and are difficult to meet the needs of practical applications.
[0006] In summary, the existing targeted protein polypeptide design methods have the problems of cumbersome processes and low design success rates, and urgent improvement is needed. By combining the advantages of the attention mechanism and graph neural networks, developing a new polypeptide design method is expected to significantly improve the design efficiency while increasing the design success rate and accuracy, providing a powerful tool for the research and application of protein-polypeptide interactions. Summary of the Invention
[0007] In view of the above, the purpose of the present invention is to provide a targeted protein polypeptide design method based on a graph attention network to solve the current defects in polypeptide design.
[0008] To achieve the above invention purpose, the targeted protein polypeptide design method based on a graph attention network provided by the present invention includes the following steps:
[0009] 1) Construct a data set:
[0010] Download a data set containing a pair of protein interactions from the RCSB online protein database, obtaining a total of 53,717 pairs of proteins, and dividing them into a training set of 42,973 and a test set of 10,744 according to the ratio of 80% and 20%. Screen polypeptides with an alpha-helix stable structure as the seed bank for subsequent design, with a total of 88,204 polypeptides;
[0011] 2) Model the protein information:
[0012] For each protein in the seed bank, calculate the charge, hydrophilicity-hydrophobicity, and Poisson-Boltzmann electrostatic potential information on its surface, and model the surface atoms as vertices in the graph structure. The backbone structure forms the format of G=(V, E), and each vertex Vi contains the chemical information of the protein. The vertices V={v1, v2, …… vn} and the edges E={eij} jointly record the geometric information of the protein, and use pymesh to store the above information;
[0013] 3) Construct a graph attention network:
[0014] Construct a graph attention network, whose basic structure consists of a plurality of neural network composite layers. Each composite layer includes an attention layer and a graph convolutional layer, and a residual connection operation is used between the two layers;
[0015] 4) Set hyperparameters:
[0016] Set the total number of iterations, the number of neural network composite layers, the temperature factor, and the activation function;
[0017] 5) Train model parameters:
[0018] Train two models respectively. One uses the deviation between the real hot spot region and the predicted hot spot region as the loss function, and can calculate the hot spot region on the protein surface. The other uses the deviation between the predicted value of the interaction between two protein surface regions and whether the interaction can actually occur as the loss function, and can calculate the possibility of interaction between two protein regions.
[0019] 6) Make predictions based on the model:
[0020] For a specific target protein, designing a polypeptide targeting this protein needs to be carried out according to the following steps:
[0021] 6.1) Use MSMS to calculate the surface information of the target protein, calculate its hydrophilicity and hydrophobicity according to the residue type, use APBS to calculate its Poisson-Boltzmann electrostatic potential, and store the data in the form of mesh at each vertex, and integrate to obtain the characteristic information of the target protein;
[0022] 6.2) Input the target protein mesh into the hot spot prediction neural network, sort according to the hot spot prediction score from high to low, and retain the top-k hot spot regions;
[0023] 6.3) Traverse each hot spot region, use the interaction judgment model to compare with the existing seeds one by one, and obtain the seeds with a confidence value exceeding the threshold t as the target polypeptide;
[0024] 7) Optimize the prediction results:
[0025] Use the Rosetta FastRelax module to optimize the polypeptide, and obtain the optimized polypeptide based on energy minimization;
[0026] 8) Visualize the results:
[0027] For the hot spot prediction model, use pymesh to record the hot spot scores of each region on the protein surface, and stain the result in the form of a three-dimensional heat map. Then use pymol to obtain the visualized three-dimensional structure image of the target protein and the polypeptide.
[0028] Compared with the prior art, the gain effects of the present invention at least include:
[0029] 1) Apply the graph attention network to the design of targeted protein polypeptides to improve the accuracy and effectiveness of polypeptide generation. The introduction of the attention mechanism also provides a certain degree of interpretability.
[0030] 2) Use two models to generate polypeptides step by step, making it possible to display the intermediate result of the binding region of the target protein. The two models can be used separately, which also increases the flexibility of the entire method.
[0031] 3) Can improve the efficiency of targeted protein polypeptide synthesis to a certain extent, make up for the problem of spatial limitation in the field of traditional targeted protein polypeptide design, and explore a broader new structure. Brief Description of the Drawings
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required in the description of the embodiments or the prior art.
[0033] Figure 1 It is the overall flowchart of targeted protein polypeptide design;
[0034] Figure 2 It is the flowchart block diagram of targeted protein polypeptide design;
[0035] Figure 3 It is the schematic diagram of the composite layer structure provided by the embodiment;
[0036] Figure 4 It is the schematic diagram of the graph attention network provided by the embodiment; Detailed Embodiments
[0037] The following further describes the present invention with reference to the drawings.
[0038] Refer to Figures 1 to 4 , A method for designing targeted protein polypeptides based on a graph neural network, comprising the following steps:
[0039] 1) Construct a data set:
[0040] Obtain the complex structure from the RCSB PDB database, download a data set containing a pair of protein interactions, and obtain 53,717 pairs of valid data after redundancy processing. Use stratified random sampling to divide the training set (42,973) and the test set (10,744) to ensure consistent class distribution. In addition, screen polypeptides with an alpha helix stable conformation from the RCSB database as the seed bank, a total of 88,204;
[0041] 2) Model the protein information:
[0042] Adopt a multi-scale feature fusion strategy: ① Electrostatic potential calculation: Use APBS to solve the non-linear Poisson-Boltzmann equation. ② Hydrophilicity and hydrophobicity: Calculate the hydrophilicity value through the Kyte Doolittle scale. ③ Geometric features: Use PyMesh to extract surface curvature and concavity. When constructing the dynamic graph network, the vertex feature vector includes features such as atomic type (One-hot), electrostatic potential, SASA value, hydrophilicity and hydrophobicity, etc.; the edge features include geometric descriptors such as Euclidean distance and dihedral angle. The modeling results are stored using a custom mesh object;
[0043] 3) Construct a graph attention network:
[0044] Design a composite module with a depth of L = 2. Each module contains: ① Multi-head graph attention layer. ② Dynamic edge convolution layer. ③ Layer normalization. ④ Residual connection.
[0045] Build a graph attention network;
[0046] 4) Set hyperparameters:
[0047] Set the total number of iterations epoch = 100, temperature factor τ = 0.5, initial learning rate α = 1e-4, batch size batch = 100, adopt the Gaussian activation function, and AdamW optimizer;
[0048] 5) Train model parameters:
[0049] Train two models separately. The task of one model is hotspot prediction, and the cross-entropy is calculated using the deviation between the true hotspot region and the predicted hotspot region as the loss function. The task of the other model is interaction prediction, and the deviation between the predicted value of the interaction between two protein surface regions and whether the interaction can actually occur is used as the loss function. The two tasks share the underlying feature extractor, and mixed precision is used for acceleration during training.
[0050] 6) Make predictions based on the model:
[0051] For a specific target protein, designing a polypeptide targeting this protein needs to be carried out according to the following steps:
[0052] 6.1) Surface feature calculation: Use MSMS to generate the molecular surface for feature calculation. Calculate its hydrophilicity and hydrophobicity according to the residue type, and use APBS to calculate its Poisson-Boltzmann electrostatic potential based on the multi-grid method. The data is stored in each vertex in mesh format, which is the feature information of the target protein;
[0053] 6.2) Hotspot region prediction: Input the target protein mesh into the hotspot prediction neural network, sort according to the hotspot prediction score from high to low, and retain the top-k (k = 10) hotspot regions;
[0054] 6.3) Peptide matching: Traverse each hotspot region, compare with existing seeds one by one using the interaction judgment model, and obtain the seeds with a confidence value exceeding the threshold t = 0.7 as the target peptides;
[0055] 7) Peptide structure optimization:
[0056] Optimize the peptide based on energy minimization using the FastRelax protocol of Rosetta 3.13;
[0057] 8) Result visualization:
[0058] Based on the hotspot prediction model, use pymesh to record the hotspot scores of each region on the protein surface and render the results in the form of a three-dimensional heat map. Based on the interaction prediction model, use pymol to render the three-dimensional structure images of the target protein and peptide.
Claims
1. A method for designing targeted protein peptides based on graph attention network, characterized in that: The following steps are involved: (1) Constructing a protein-peptide interaction dataset: Obtaining protein complex structure data from the RCSB PDB database, screening protein interaction pairs and performing redundancy removal, and dividing them into training and test sets; analyzing and constructing a peptide seed library containing stable α-helical conformations; (2) Protein surface feature modeling: Calculate the surface charge distribution, hydrophilicity and Poisson-Boltzmann electrostatic potential of the target protein, and model the surface atomic topology as a graph G = (V, E), where the vertex V contains the atomic-level chemical features; (3) Construct a hierarchical graph attention network: It contains L composite modules, each of which consists of a multi-head graph attention layer, a dynamic edge convolution layer, and a residual connection; (4) Dual-task model: peptide design is achieved through the hotspot prediction model and the interaction prediction model; (5) Peptide design prediction: extract multi-scale features from the target protein surface, and match candidate peptides from the seed library through hot spot area screening and approximate search; (6) Peptide conformation optimization: The Rosetta energy minimization protocol was used to relax the candidate peptide structure; (7) Three-dimensional visualization: Generate heat maps of hot spots on the protein surface and visualization models of the protein-peptide complex interaction network.
2. The method for designing targeted protein peptides based on graph attention network according to claim 1, characterized in that: In step (2), the feature calculation process is as follows: Computational geometric features: Use MSMS and pymesh to extract protein surface curvature and concavity, and use the result as a feature of the protein; Calculate electrostatic potential: Use APBS to solve the nonlinear Poisson-Boltzmann equation and use the result as a characteristic of the protein; Calculation of hydropathicity: The Kyte Doolittle scale is used to calculate the hydropathicity of residues and the result is used as a characteristic of the protein.
3. The method for designing targeted protein peptides based on graph attention network according to claim 1, characterized in that: In the step (3), the multi-head attention layer of the composite module adopts GATv2, and the dynamic edge convolution layer uses a Gaussian activation function.
4. The method for designing targeted protein peptides based on graph attention network according to claim 1, characterized in that: In step (4), both task models are gradient updated using the adam optimizer and the cross entropy loss function.
5. The method for designing targeted protein peptides based on graph attention network according to claim 1, characterized in that: In the step (6), the structural optimization introduces trRosetta distance constraints and executes multiple rounds of FastRelax cycles.