Protein interaction site prediction method based on local-global feature fusion

By combining local and global features, building protein maps and using deep learning models for feature extraction and fusion, the problem of difficult to effectively combine global information and local information in the existing technology to predict protein interaction sites is solved, and higher prediction accuracy and model expression ability are achieved.

CN119964639AActive Publication Date: 2025-05-09QINGDAO UNIV

Patent Information

Application Number
CN202510437638.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively combine global information with local information to predict protein interaction sites, resulting in insufficient prediction accuracy.

Method used

A local-global feature fusion method is used to extract the node characteristics and edge characteristics of proteins, and local graphs and global graphs are constructed, and feature extraction and fusion are used for graph attention network (GAT) and Transformer models, and finally classified prediction is performed through multi-layer perception machines (MLP).

Benefits of technology

It significantly improves the comprehensiveness and robustness of feature extraction, improves the accuracy of protein interaction site prediction and model expression ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964639A_ABST
    Figure CN119964639A_ABST
Patent Text Reader

Abstract

The invention provides a protein interaction site prediction method based on local-global feature fusion, and relates to the field of protein site prediction, and the method specifically comprises the following steps: extracting protein features including protein node features and protein edge features; constructing a protein map structure according to the extracted protein features; inputting the protein graph structure into a graph attention network GAT for further feature extraction to obtain multi-scale features and fusion features; and comparing the similarity of the loss capture multi-scale features and the fusion features, inputting the multi-scale features and the fusion features into a Transform-based attention feature fusion module, and finally inputting the multi-scale features and the fusion features into a multi-layer perceptron (MLP) for classification prediction. According to the technical scheme, the problem that in the prior art, global information and local information cannot be effectively combined for protein interaction site prediction is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of protein site prediction, and in particular to a protein interaction site prediction method based on local-global feature fusion. Background Art

[0002] Proteins are important components of biological cells and play a key role in coordinating various physiological activities. However, proteins often need to interact with other biomolecules to achieve their biological functions, among which protein-protein interactions (PPIs) are crucial in many biological processes. In these interactions, specific residue regions constitute protein interaction sites (PPIS), which are interface regions on the protein surface that directly participate in binding. Accurate identification of PPIS is of great significance for in-depth understanding of biological processes, exploration of disease mechanisms, and guiding the development of new drugs.

[0003] Information based on protein structure is one of the important strategies for predicting PPIS at present, and it can usually provide higher accuracy. In this method, proteins can be modeled as a graph structure, in which amino acids are regarded as nodes in the graph, and whether there is an edge connection is determined by the spatial distance between amino acids. The characteristic information of proteins is crucial to the performance of the model, but it is often difficult to obtain the best prediction results using only one feature. When fusing multiple features, although the prediction accuracy is improved, higher computational complexity is also introduced. In addition, features such as position-specific score matrix (PSSM) rely on large reference databases, which makes the feature extraction process time-consuming. In recent years, with the development of natural language processing (NLP) technology, protein representation learning methods based on deep learning have gradually emerged, providing new ideas for protein feature extraction.

[0004] Many current methods only consider proteins as a single global graph when modeling, while ignoring the hierarchical relationship between global and local structures. In real biological systems, the local structure of proteins (such as pockets, active sites) often directly determines the binding characteristics, while the global topological structure affects the overall spatial conformation. Therefore, constructing a protein graph only from a global perspective may lose important local information, while focusing only on local features may make it difficult to capture the overall structural relationship. Therefore, how to effectively combine global information and local information in the PPIS prediction task and construct a more refined feature expression remains a key challenge in current research.

[0005] Therefore, there is a need for a protein interaction site prediction method that can effectively combine global and local information. Summary of the invention

[0006] The main purpose of the present invention is to provide a protein interaction site prediction method based on local-global feature fusion to solve the problem that the prior art cannot effectively combine global information with local information to predict protein interaction sites.

[0007] To achieve the above object, the present invention provides a protein interaction site prediction method based on local-global feature fusion, which specifically comprises the following steps: S1, extract protein features, including: protein node features and protein edge features.

[0008] S2, construct the protein graph structure based on the extracted protein features.

[0009] S3, the protein graph structure is input into the graph attention network GAT for further feature extraction to obtain multi-scale features and fusion features.

[0010] S4 captures the similarity of multi-scale features and fused features through contrast loss, inputs the multi-scale features and fused features into the Transformer-based attention feature fusion module, and finally inputs them into the multi-layer perceptron MLP for classification prediction.

[0011] Furthermore, protein node characteristics include: evolutionary characteristics, secondary structure characteristics and physicochemical characteristics of amino acids.

[0012] Furthermore, the evolutionary features include: a position-specific score matrix PSSM and a hidden Markov model matrix HMM; and the values ​​in the position-specific score matrix PSSM and the hidden Markov model matrix HMM are normalized: ; in, Represents the original value, and are the minimum and maximum values ​​of a feature type in the training set, respectively. is the normalized value.

[0013] Furthermore, the secondary structure features are calculated and generated by the DSSP algorithm, and the size of the secondary structure feature matrix is , Represents the length of the amino acid sequence, and 14 is the dimension; among them, the 9-dimensional features are nine secondary structure states, represented by one-hot encoding; the 4-dimensional features are obtained by sine and cosine transforming the torsion angles PHI and PSI of the peptide chain main chain; the last 1-dimensional feature is converted from the solvent accessible surface area SASA to the relative solvent accessibility RSA.

[0014] Furthermore, the physicochemical characteristics of amino acids include: isoelectric point, polarity, pH, number of hydrogen bond acceptors, number of hydrogen bond donors, octanol-water partition coefficient logP and topological polar surface area TPSA.

[0015] Furthermore, the protein edge features extracted in step S1 are specifically: The protein edge features are The feature matrix is ​​represented by , where 3 represents the dimension; the first dimension is 0 or 1, if two nodes are directly connected by an edge, it is 1, otherwise it is 0; the second dimension is the node and Location and The Euclidean distance , as shown in formula (1); the third dimension is and Angle between The cosine value of , as shown in formula (2): (1); (2); in, for and The distance between is the initial coordinate position.

[0016] Furthermore, step S2 specifically includes the following steps: S2.1, the protein graph structure includes: local graph and global graph. The adjacency matrix of local graph and global graph is constructed as shown in equations (3) and (4) respectively: (3); (4); in, , represents the spatial distance between residues, and Node and Main chain Carbon atom coordinates, is the adjacency matrix of the local graph, is the adjacency matrix of the global graph, and Represents the threshold value.

[0017] S2.2, fusion of local graph and global graph information: (5); in, is the adjacency matrix of the fused graph.

[0018] Furthermore, step S3 specifically includes the following steps: S3.1, input the protein features extracted in step S1 and the adjacency matrix of the local graph into the first graph attention network GAT to extract the local features of the protein; then use the extracted local features as the initial features of the global graph and input them into the second graph attention network to obtain multi-scale features after convolution.

[0019] S3.2, the initial protein features and the adjacency matrix of the fusion graph are input into the third graph attention network, and the fusion features of the protein are extracted after convolution. The implementation of the first, second, and third graph attention networks is shown in equations (6), (7), and (8): (6); (7); (8); in, and Respectively represent nodes and its neighbor nodes The input feature vector is Representation Node and The edge feature vector between Represents a splicing operation, , , and represents the learnable parameter matrix of the linear layer at different positions in the graph attention network, represents the activation function, Representation Node The weight of is between 0 and 1. For Node and The attention score between Representation Node Updated embed, Represents the ReLU activation function.

[0020] Furthermore, step S4 specifically includes the following steps: S4.1, multi-scale features and fusion features To splice, , and then the concatenated features The input is processed based on the Transformer attention feature fusion module. The calculation process based on the Transformer encoder is as follows: (9); in, are query, key, and value matrices respectively; is the attention mechanism, is the learnable weight matrix, is the dimension of the attention head, is the activation function.

[0021] S4.2, Output of Transformer-based Attention Feature Fusion Module for: (10); in, It is A head of attention, is the output projection matrix, For connection operation.

[0022] S4.3, the final prediction result is: (11); in, is the normalization layer, is a multi-layer perceptron, is the final prediction result.

[0023] The present invention has the following beneficial effects: This paper innovatively introduces the concepts of global graph and local graph, and extracts the global topological features and local detail features of proteins through these two graph structures. The contrast loss function is used to capture the similarities between features at different levels, and the two are deeply integrated, which significantly improves the comprehensiveness and robustness of feature extraction.

[0024] The present invention comprehensively utilizes the sequence features, structural features and edge features between residues of proteins to construct a multi-dimensional feature representation. This multi-feature fusion strategy not only enriches feature information, but also significantly improves the accuracy of protein interaction site prediction.

[0025] This paper proposes a Transformer-based feature fusion mechanism that can adaptively capture the long-range dependencies between protein sequences and structures, and achieve efficient feature fusion through a self-attention mechanism. This mechanism further enhances the model's ability to model complex protein features and provides a more powerful tool for protein function prediction.

[0026] The present invention deeply explores the correlation between protein sequence, structure and topological information. At the same time, it studies the multimodal feature fusion strategy to improve the robustness of the model in the absence of structural information, so as to more effectively utilize existing data and improve prediction accuracy and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Obviously, the drawings described below are some implementations of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 A flow chart of a protein interaction site prediction method based on local-global feature fusion according to the present invention is shown.

[0028] Figure 2 A schematic diagram of constructing a local map and a global map of the present invention is shown. DETAILED DESCRIPTION

[0029] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] like Figure 1 A protein interaction site prediction method based on local-global feature fusion is shown, which specifically includes the following steps: S1, extract protein features, including: protein node features and protein edge features.

[0031] Proteins are composed of amino acid sequences, of which there are 20 kinds of amino acids, namely alanine (Ala, A), arginine (Arg, R), aspartic acid (Asp, D), asparagine (Asn, N), etc. In the present invention, proteins are regarded as a graph structure, in which each amino acid residue is modeled as a node in the graph, and the edges between nodes are constructed based on the spatial position relationship or covalent bond relationship of amino acids. Common graph construction methods include defining the connection mode of edges based on the adjacent relationship of the amino acid backbone (Backbone), based on the Euclidean distance threshold between residues, or based on the contact map (Contact Map).

[0032] Therefore, before predicting protein interaction sites, it is necessary to first prepare complete protein sequence information and three-dimensional structure information, and build a reasonable graph structure based on this. On this basis, the present invention further extracts node features and edge features to provide rich information for the training of deep learning models.

[0033] S2, constructing the protein graph structure based on the extracted protein features. The present invention uses AlphaFold3 to generate the protein structure file and Euclidean distance between carbon atoms Constructing the graph structure of proteins. Specifically, the construction of protein graphs depends on the spatial distance between residues. , as shown below:

[0034] in, and Amino acid nodes and Main chain Carbon atom coordinates. In order to fully extract the amino acid features of proteins, the present invention constructs three different graph structures, including global graph, local graph and fusion graph, to depict short-range and long-range topological information, and combines the contrastive learning method to further enhance the feature expression ability of proteins.

[0035] S3, input the protein graph structure into the graph attention network GAT for specific feature extraction to obtain multi-scale features and fusion features.

[0036] S4 captures the similarity of multi-scale features and fused features through contrast loss, inputs the multi-scale features and fused features into the Transformer-based attention feature fusion module, and finally inputs them into the multi-layer perceptron MLP for classification prediction.

[0037] Specifically, protein node features include: evolutionary features, secondary structure features, and physicochemical features of amino acids.

[0038] Specifically, the evolutionary features include: a position-specific scoring matrix PSSM and a hidden Markov model matrix HMM; to ensure the numerical stability of the features, the values ​​in the position-specific scoring matrix PSSM and the hidden Markov model matrix HMM are normalized: ; in, Represents the original value, and are the minimum and maximum values ​​of a feature type in the training set, respectively. is the normalized value.

[0039] The PSSM was generated by the alignment tool PSI-BLAST v2.10.1, with the number of iterations set to 3 and the E-value set to 0.001, while the HMM matrix was generated by the HHblits v3.0.3 algorithm with default parameters. The shapes of the PSSM and HMM matrices are ,in Represents the length of the amino acid sequence.

[0040] Specifically, the secondary structure features are calculated and generated by the DSSP algorithm, and the size of the secondary structure feature matrix is , Represents the number of amino acids, 14 is the dimension; among them, 9-dimensional features are nine secondary structure states, represented by one-hot encoding; 4-dimensional features are obtained by sine and cosine transforming the torsion angles PHI and PSI of the peptide chain main chain; the last 1-dimensional feature is converted from the solvent accessible surface area SASA to the relative solvent accessibility RSA.

[0041] Specifically, the physicochemical characteristics of amino acids include: isoelectric point, polarity, pH, number of hydrogen bond acceptors, number of hydrogen bond donors, octanol-water partition coefficient logP and topological polar surface area TPSA.

[0042] The secondary structure characteristics are calculated by the DSSP algorithm; the physicochemical characteristics of amino acids include: isoelectric point, polarity, pH, number of hydrogen bond acceptors, number of hydrogen bond donors, octanol-water partition coefficient logP and topological polar surface area TPSA. These features comprehensively consider the evolutionary information, structural characteristics and chemical properties of proteins, providing richer information support for subsequent model learning.

[0043] Specifically, the protein edge features extracted in step S1 are:

[0044] The protein edge features are The feature matrix is ​​represented by , where 3 represents the dimension; the first dimension is 0 or 1, if two nodes are directly connected by an edge, it is 1, otherwise it is 0; the second dimension is the node and Location and The Euclidean distance , as shown in formula (1); the third dimension is and Angle between The cosine value of , as shown in formula (2): (1); (2); in, for and The distance between is the initial coordinate position.

[0045] Specifically, step S2 includes the following steps: S2.1, the protein graph structure includes: local graph and global graph. The adjacency matrix of local graph and global graph is constructed as shown in equations (3) and (4) respectively: (3); (4); in, , represents the spatial distance between residues, and Node and Main chain Carbon atom coordinates, is the adjacency matrix of the local graph, is the adjacency matrix of the global graph, and Represents the threshold value.

[0046] The present invention sets two thresholds and To construct local and global graphs of proteins to accurately describe the topological relationships between amino acids, such as Figure 2 Specifically, when satisfy When and There are local topological relationships, and local adjacency matrices are constructed based on them. , the matrix can capture the short-range interaction characteristics of proteins; when When and There is a global topological relationship, which is used to construct the global adjacency matrix , the matrix can effectively model the association information between long-distance amino acids, thereby supplementing the lack of local information. Through this two-layer topological modeling method, the local and global structural features of the protein can be accurately extracted, so that the neural network can focus on both microscopic amino acid interactions and capture macroscopic global topological information, thereby improving the expressive power of the model. In addition, this method can enhance the integrity of protein structural information, improve the accuracy of protein function prediction, and provide stronger support for tasks such as protein interaction prediction and functional classification.

[0047] S2.2, fusion of local graph and global graph information: (5); in, is the adjacency matrix of the fused graph.

[0048] The fusion graph combines the information of the local graph and the global graph to provide a more comprehensive expression of topological relationships and lay the foundation for calculating the contrast loss for subsequent tasks. satisfy When the present invention constructs a fusion graph adjacency matrix for proteins , as shown in formula (5). The fusion graph not only integrates the short-range and long-range interaction information, but also alleviates the problem of incomplete information in the local graph and the global graph to a certain extent, so that the model can further enhance the topological representation ability of proteins based on sequence and structural features, improve the effect of comparative learning, and thus optimize the performance of protein interaction prediction or other downstream tasks.

[0049] Specifically, after the feature and protein graphs are constructed, they are input into the graph attention network (GAT) for further feature extraction. Step S3 specifically includes the following steps: S3.1, input the protein features extracted in step S1 and the adjacency matrix of the local graph into the first graph attention network GAT to extract the local features of the protein; then use the extracted local features as the initial features of the global graph and input them into the second graph attention network to obtain multi-scale features after convolution.

[0050] S3.2, the initial protein features and the adjacency matrix of the fusion graph are input into the third graph attention network, and the fusion features of the protein are extracted after convolution. The implementation of the first, second, and third graph attention networks is shown in equations (6), (7), and (8): (6); (7); (8); in, and Respectively represent nodes and its neighbor nodes The input feature vector is Representation Node and The edge feature vector between Represents a splicing operation, , , and represents the learnable parameter matrix of the linear layer at different positions in the graph attention network, represents the activation function, Representation Node The weight ranges from 0 to 1. The function takes its attention score Convert to get, For Node and The attention score between Representation Node Updated embed, Represents the ReLU activation function.

[0051] Specifically, after obtaining multi-scale features and fused features in the upstream task, the contrast loss is used to capture the similarity of the two features and make them as close as possible, thereby optimizing the parameters of the upstream model. Subsequently, the two features are input into the Transformer-based attention feature fusion module to further enhance the protein representation ability through feature fusion, and finally input into the multi-layer perceptron (MLP) for classification prediction.

[0052] Step S4 specifically includes the following steps: S4.1, multi-scale features and fusion features To splice, , and then the concatenated features The input is processed based on the Transformer attention feature fusion module. The calculation process based on the Transformer encoder is as follows: (9); in, are query, key, and value matrices respectively; is the attention mechanism, is the learnable weight matrix, is the dimension of the attention head, is the activation function.

[0053] S4.2, Output of Transformer-based Attention Feature Fusion Module for: (10); in, It is Attention head, is the output projection matrix, For connection operation.

[0054] S4.3, the final prediction result is: (11); in, is the normalization layer, is a multi-layer perceptron, is the final prediction result.

[0055] Finally, the parameters are fine-tuned. In the parameter fine-tuning and prediction stage, the present invention uses 80% of the protein data for training and 20% of the data for testing, and adopts five-fold cross validation to evaluate the stability and generalization ability of the model. The model is fine-tuned by adjusting the learning rate, batch size, optimizer (such as AdamW) and regularization parameters (such as Dropout and weight decay) to optimize its performance. During the training process, the early stopping strategy is used to prevent overfitting, and the accuracy, F1 score and other indicators of the model are evaluated on the test set. Finally, the reliability of the model is verified by the results of five-fold cross validation.

[0056] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A protein interaction site prediction method based on local-global feature fusion, characterized in that: The specific steps include: S1, extract protein features, including protein node features and protein edge features; S2, constructing the protein graph structure based on the extracted protein features; S3, input the protein graph structure into the graph attention network GAT for further feature extraction to obtain multi-scale features and fusion features; S4 captures the similarity of multi-scale features and fused features through contrast loss, inputs the multi-scale features and fused features into the Transformer-based attention feature fusion module, and finally inputs them into the multi-layer perceptron MLP for classification prediction.

2. A protein interaction site prediction method based on local-global feature fusion according to claim 1, characterized in that: Protein node characteristics include: evolutionary characteristics, secondary structure characteristics and physicochemical characteristics of amino acids.

3. A protein interaction site prediction method based on local-global feature fusion according to claim 2, characterized in that: The evolutionary features include: a position-specific score matrix PSSM and a hidden Markov model matrix HMM; the values ​​in the position-specific score matrix PSSM and the hidden Markov model matrix HMM are normalized: ; in, Represents the original value, and are the minimum and maximum values ​​of a feature type in the training set, respectively. is the normalized value.

4. A protein interaction site prediction method based on local-global feature fusion according to claim 2, characterized in that: The secondary structure features are calculated by the DSSP algorithm, and the size of the secondary structure feature matrix is , Represents the length of the amino acid sequence, and 14 is the dimension; among them, the 9-dimensional features are nine secondary structure states, represented by one-hot encoding; the 4-dimensional features are obtained by sine and cosine transforming the torsion angles PHI and PSI of the peptide chain main chain; the last 1-dimensional feature is converted from the solvent accessible surface area SASA to the relative solvent accessibility RSA.

5. The method for predicting protein interaction sites based on local-global feature fusion according to claim 2, characterized in that: The physicochemical characteristics of amino acids include isoelectric point, polarity, pH, number of hydrogen bond acceptors, number of hydrogen bond donors, octanol-water partition coefficient logP and topological polar surface area TPSA.

6. A protein interaction site prediction method based on local-global feature fusion according to claim 1, characterized in that: The specific steps of extracting protein edge features in step S1 are: The protein edge features are The feature matrix is ​​represented by , where 3 represents the dimension; the first dimension is 0 or 1, if two nodes are directly connected by an edge, it is 1, otherwise it is 0; the second dimension is the node and Location and The Euclidean distance , as shown in formula (1); the third dimension is and Angle between The cosine value of , as shown in formula (2): (1); (2); in, for and The distance between is the initial coordinate position.

7. The method for predicting protein interaction sites based on local-global feature fusion according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2.1, the protein graph structure includes: local graph and global graph. The adjacency matrix of local graph and global graph is constructed as shown in equations (3) and (4) respectively: (3); (4); in, , represents the spatial distance between residues, and Node and Main chain Carbon atom coordinates, is the adjacency matrix of the local graph, is the adjacency matrix of the global graph, and represents the threshold value; S2.2, fusion of local graph and global graph information: (5); in, is the adjacency matrix of the fused graph.

8. The method for predicting protein interaction sites based on local-global feature fusion according to claim 1, characterized in that: Step S3 specifically includes the following steps: S3.1, input the protein features extracted in step S1 and the adjacency matrix of the local graph into the first graph attention network GAT1 to extract the local features of the protein; then use the extracted local features as the initial features of the global graph and input them into the second graph attention network GAT2 to obtain multi-scale features after convolution; S3.2, the initial protein features and the adjacency matrix of the fusion graph are input into the third graph attention network GAT3, and the fusion features of the protein are extracted after convolution. The implementation of the first, second, and third graph attention networks is shown in equations (6), (7), and (8): (6); (7); (8); in, and Respectively represent nodes and its neighbor nodes The input feature vector is Representation Node and The edge feature vector between Represents a splicing operation, , , and represents the learnable parameter matrix of the linear layer at different positions in the graph attention network, represents the activation function, Representation Node The weight of is between 0 and 1. For Node and The attention score between Representation Node Updated embed, Represents the ReLU activation function.

9. The method for predicting protein interaction sites based on local-global feature fusion according to claim 1, characterized in that: Step S4 specifically includes the following steps: S4.1, multi-scale features and fusion features To splice, , and then the concatenated features The input is processed based on the Transformer attention feature fusion module. The calculation process based on the Transformer encoder is as follows: (9); in, are query, key, and value matrices respectively; is the attention mechanism, is the learnable weight matrix, is the dimension of the attention head, is the activation function; S4.2, Output of Transformer-based Attention Feature Fusion Module for: (10); in, It is Attention head, is the output projection matrix, For connection operation; S4.3, the final prediction result is: (11); in, is the normalization layer, is a multi-layer perceptron, is the final prediction result.

Citation Information

Patent Citations

  • Protein secondary structure prediction method based on multi-scale convolution attention neural network

    CN112767997A

  • Protein feature extraction method based on key point selection space graph convolution model

    CN118136098A

  • Protein melting temperature prediction method based on mixed deep learning strategy

    CN119446287A

  • Multi-modal fusion-based drug-protein interaction prediction model construction method

    CN119580818A

  • Protein interaction prediction method based on cross-modal enhanced representation learning

    CN119626312A

Cited By

  • Quantum and region sensing fused protein methylation site prediction method

    CN120727109A

  • Protein interaction site prediction method and system

    CN120877854A

  • A method and system for predicting protein interaction sites

    CN120877854B

  • RNA binding site prediction method based on Mangbar and graph neural network

    CN121256334A