A protein interaction prediction method based on dual synergy mechanism

By introducing task collaboration mechanisms of interactive attention protein collaboration and isomorphic multi-task learning in protein interaction prediction models, the shortcomings of existing models in knowledge sharing and cross-domain knowledge complementarity are solved, and higher prediction accuracy and generalization capabilities are achieved.

CN118609643BActive Publication Date: 2025-05-16OCEAN UNIV OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410260920.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-05-16
Estimated Expiration
2044-03-07

AI Technical Summary

Technical Problem

The existing protein interaction prediction models have insufficient knowledge sharing and cross-domain knowledge complementarity, resulting in insufficient prediction accuracy.

Method used

The introduction of protein collaboration mechanisms based on interactive attention and task collaboration mechanisms based on isomorphic multitask learning enables the protein interaction prediction model to achieve knowledge sharing between two proteins and achieve knowledge complementarity in different tasks.

Benefits of technology

Through protein collaboration and task collaboration, the prediction accuracy of protein interaction prediction models is significantly improved and the generalization ability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118609643B_ABST
    Figure CN118609643B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of protein prediction technology, and in particular to a protein interaction prediction method based on a dual synergy mechanism. The present invention introduces protein synergy into a twin architecture by utilizing the interactive attention between protein structure pairs, thereby realizing the sharing of interaction knowledge between proteins. The present invention proposes a protein synergy mechanism based on interactive attention, so that the determination of key residues in a protein depends not only on its own characteristics, but also on its cooperative proteins, thereby realizing knowledge sharing between two proteins and further improving the interaction prediction accuracy of the model. The present invention introduces protein function prediction tasks and subcellular location prediction tasks into the training process of the protein interaction prediction model, so that the knowledge shared between a pair of proteins can be complemented in different tasks, further improving the interaction prediction accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of protein prediction, and in particular to a protein interaction prediction method based on a dual synergistic mechanism. Background Art

[0002] Exploring protein interactions is of great significance for understanding biological processes such as pathway signaling and molecular function expression. Although experimental techniques such as mass spectrometry have been applied to the determination of protein interactions, they are still unable to bridge the gap between the explosive growth of protein data and limited interaction insights due to their expensive and time-consuming disadvantages. Therefore, this apparent dilemma has prompted researchers to develop computational methods to accelerate protein interaction prediction.

[0003] Early computational methods for protein interactions were based on molecular dynamics simulations, which determined the interactions between proteins based on the binding posture of a pair of protein complexes presented in the simulation environment. However, in addition to the large amount of time and computing resources required for simulation, the high-precision protein real structure required by this simulation-based method is not always available, and this demand for the real protein structure may sometimes not be met.

[0004] With the widespread application of data-driven deep learning technology in natural language processing, computer vision and other fields, the life science field represented by protein interaction prediction tasks has also been inspired by it. The surge in protein data caused by the development of sequencing technology in the 21st century has also laid the foundation for the application of data-driven deep learning technology in protein interaction prediction.

[0005] At present, deep learning-based protein interaction prediction methods can be divided into two categories according to their protein representation, namely, methods based on protein one-dimensional sequence and methods based on protein two-dimensional structure.

[0006] The interaction prediction model based on protein one-dimensional sequence represents a pair of proteins to be predicted to interact as two protein sequences, and applies two sets of hidden layers consisting of multiple one-dimensional convolutional layers or fully connected layers to the two encoded protein sequences, thereby automatically capturing high-order features from the protein-protein pair. The high-order features of the two proteins obtained from the last hidden layer are concatenated as the input of the interaction classifier, which then outputs a specific 0 / 1 value to indicate whether the two proteins interact.

[0007] The protein interaction prediction model based on two-dimensional structure is based on the one-dimensional sequence method. It represents two proteins as two residue contact graphs (residues are nodes and the contact relationship between residues is edges), thereby retaining the spatial structural information of the proteins. On this basis, two sets of hidden layers consisting of multiple graph neural network layers are respectively applied to the feature aggregation between residues in the two proteins to further extract the structural features of the proteins, thereby obtaining a higher protein interaction prediction accuracy than the one-dimensional sequence method.

[0008] The existing protein interaction prediction models based on one-dimensional sequences and two-dimensional structures both regard the two proteins to be predicted to interact as two separate entities, and extract features for the two proteins separately through two sets of neural network layers, thereby ignoring the knowledge sharing between the two proteins (such as the knowledge sharing of the importance of residues from the perspective of cooperative proteins), and the interaction prediction accuracy needs to be further improved; in addition, the existing interaction prediction models are only guided by interaction tasks during training, thereby ignoring the knowledge complementarity brought by cross-domain knowledge for interaction prediction. Therefore, the present invention introduces a protein coordination mechanism based on interactive attention and a task coordination mechanism based on isomorphic multi-task learning, so that the interaction prediction model can not only realize the knowledge sharing between the two proteins, but also realize the complementarity of this shared knowledge in cross-domain tasks, thereby further improving the prediction accuracy of the protein interaction prediction model. Summary of the invention

[0009] Exploring protein interactions is of great significance for understanding biological processes such as pathway signaling and molecular function expression. Although experimental techniques such as mass spectrometry have been applied to the determination of protein interactions, they are still unable to bridge the gap between the explosive growth of protein data and the limited insights into protein interactions due to their expensive and time-consuming disadvantages. Therefore, this apparent dilemma has prompted researchers to develop deep learning-based computational methods to accelerate protein interaction prediction.

[0010] Protein interaction prediction refers to inputting two proteins into a deep learning model so that the model can quickly determine whether the two proteins interact with each other, significantly reducing the economic and time costs of traditional biochemical experiments. In view of the shortcomings of previous protein interaction prediction models, the technical problems solved by the present invention are as follows:

[0011] (1) The present invention constructs a residue contact map from the protein structure predicted by AlphaFold2, and calculates the residue features through ESM-2, so as to fully construct a protein descriptor, so that the constructed protein descriptor contains biological rule knowledge such as protein spatial structure, co-evolution information, and functional sites.

[0012] (2) The present invention proposes a protein synergy mechanism based on interactive attention, so that the determination of key residues in a protein depends not only on its own characteristics, but also on its cooperating proteins, thereby realizing knowledge sharing between two proteins and further improving the interaction prediction accuracy of the model.

[0013] By collecting protein function and subcellular location information, the present invention proposes a task coordination mechanism based on isomorphic multi-task learning, and introduces protein function prediction tasks and subcellular location prediction tasks into the training process of the protein interaction prediction model, so that the knowledge shared between a pair of proteins can be complemented in different tasks, further improving the interaction prediction accuracy of the model.

[0014] In order to make up for the deficiencies of the prior art, the present invention provides a protein interaction prediction method based on protein synergy and task synergy.

[0015] The present invention is achieved through the following technical solution: a protein interaction prediction method based on protein synergy and task synergy, comprising the following steps:

[0016] Step 1: Construction of residue contact map

[0017] First, the Euclidean distance between any pair of residues is calculated based on the three-dimensional coordinates of Alpha-C in the protein structure predicted by AlphaFold2; then, if the Euclidean distance d between residue i and residue j is i,j If it is less than the preset threshold t, it is considered as contact, and so on, so as to obtain a complete contact map, which is defined as:

[0018]

[0019] Among them, A∈R N×N represents the residue contact map, R represents an arbitrary real number, N is the number of residues, and the threshold t is experimentally determined and set to Here represents angstrom, i.e. 10 -12 M, residue contact map A∈R N×N It is regarded as an adjacency matrix and is used to represent the connectivity between residues in the two-dimensional structure of a protein;

[0020] Step 2: Build pre-trained residual embeddings

[0021] The protein sequence of length N is sent to the pre-trained ESM-2, and any residue in the protein is converted into a continuous vector with a size of 1×1280 by ESM-2. Due to the implicit learning ability of ESM-2 for protein co-evolution information, functional sites and other knowledge, the obtained vector contains rich prior knowledge related to protein interactions and can be directly used as residue embedding. On this basis, the feature matrix X∈R composed of stacked residue embeddings N×1280 transfer to downstream protein tasks;

[0022] Step 3: Extract protein structural features using twin graph attention layer

[0023] After steps 1 and 2, a pair of proteins can be represented as two protein graphs and in V (1) and V (2) They are and The set of residue nodes, A (1) and A (2) They are and The adjacency matrix, X (1) and X (2) They are and feature matrix; on this basis, the twin graph attention layer is applied to extract the structural features of the protein pair;

[0024] Step 4: Protein collaboration based on interactive attention

[0025] After the above steps, the updated feature representation of the protein pair is obtained, namely X (1)′ and X (2)′ ,However, the update of each representation only considers its own intrinsic structural or topological features, but ignores the information enhancement of its ,cooperating proteins., Therefore, interactive attention is introduced to achieve collaboration among ,proteins and thus promote their knowledge sharing;

[0026] The key to interactive attention is the calculation of the attention score, which fully combines the characteristics of the two proteins in the protein pair. and To this end, three different strategies are designed for the calculation of attention scores:

[0027]

[0028] in, and denote the embedding of the i-th residue in the first protein and the j-th residue in the second protein, respectively, d denotes the feature dimension, (·) T and ⊙ denote the transpose and Hadamard product respectively, U∈R d×d , V∈R d×d and w∈R 1×d There are three sets of learnable model parameters for and Perform a learnable linear mapping so that the attention score q of the jth residue in the second protein to the ith residue in the first protein is ij can be obtained; after the SoftMax operation, the final interaction attention score in the form of a probability vector of a pair of proteins is obtained:

[0029]

[0030]

[0031] Here, exp is an exponential function with the natural constant e as the base. Thus, the attention score of the second protein on the i-th residue in the first protein is and the attention score of the first protein to the jth residue in the second protein The knowledge of protein interactions is fully shared between protein pairs, so that the residues in each protein that contribute more to the interaction can be identified based on their co-proteins; finally, the two attention scores are combined with the original protein representation and Multiply them together to adjust the residue importance and thus achieve synergy between protein pairs:

[0032]

[0033]

[0034] Step 5: Task collaboration based on homogeneous multi-task learning

[0035] The protein pair obtained in step 4 is represented by s (1) and (2) It is directly input into the fully connected classifier for binary interaction prediction. This step introduces homogeneous multi-task learning with two auxiliary tasks, protein function prediction and subcellular location prediction, to jointly fine-tune the learnable parameters of the twin attention layer, thereby achieving task collaboration, promoting knowledge complementarity in multiple biological fields, and improving the generalization ability of the interaction prediction model.

[0036] Furthermore, in step 3, the twin graph attention layer is a graph attention network with two sets of shared parameters, based on protein For example, its residues are embedded in is input into the first set of attention layers, N1 is protein The number of residues, the corresponding structural features output by the graph attention layer is represented as:

[0037]

[0038] Among them, W k ∈R d×1280 is the learnable parameter for the linear mapping, L i is the first-order neighbor residue of residue i, which is obtained from the adjacency matrix A (1) Get,|| means splicing, represents the k-th normalized attention coefficient calculated by K-head attention:

[0039]

[0040] a∈R 2d is the learnable weight vector, T is the transpose, W∈R d×1280 To calculate the learnable parameter of the attention coefficient, L i is the first-order neighbor residue of residue i; in this way, the protein The feature matrix of is updated by the graph attention layer, namely So as to achieve further feature extraction;

[0041] In addition, protein The above operation has also been applied to proteins The second set of graph attention layers Structural feature extraction is performed to is obtained, where the first and second groups of graph attention layer parameters are shared and together constitute the twin graph attention layer.

[0042] Furthermore, in step 4, q ij The calculation process is summarized as:

[0043] φ(X (1)′ , X (2)′ ) → Q

[0044] Where φ is: represents the strategy function, which is used to map a pair of feature maps of size N1×d and N2×d into an attention score matrix of size N1×N2

[0045] Afterwards, the row and column average operations are applied to Thus, we can get the attention score of the whole protein to a single residue in another protein. and It is defined as:

[0046]

[0047] Among them, N1 and N2 represent the first protein and the second protein The number of residues in .

[0048] Furthermore, the steps of isomorphic multi-task learning with auxiliary tasks of protein function prediction and subcellular location prediction in step 5 are as follows:

[0049] 1) Protein function

[0050] Gene ontology terms describing protein functions were collected from the UniProt database to accurately define the biological functions of each protein. The top 100 GO terms with the highest frequency were then recorded to form a fixed set of GO terms. On this basis, bag-of-words encoding was used to map the functions of proteins to this fixed set, thus avoiding the dimensionality explosion problem caused by the huge functional space. In this way, each protein can obtain a 100-dimensional protein function label y consisting of 0 and 1. func , where 0 / 1 indicates whether a protein has a certain function.

[0051] 2) Subcellular location After collecting the subcellular location information of proteins, the subcellular location of each protein is preprocessed in the same way as the protein functional label. Therefore, a 100-dimensional location label y can also be obtained. loc ; In addition, subcellular location prediction and function prediction are considered as multi-label classification problems, because a protein may appear in multiple subcellular locations or have multiple functions;

[0052] Afterwards, the three groups of stacked fully connected layers are regarded as classifiers for their respective tasks, namely protein function prediction, subcellular location prediction, and interaction prediction, defined as:

[0053]

[0054]

[0055]

[0056]

[0057]

[0058] in, Represents the functional label and subcellular location label of the first protein predicted by the model, represents the functional label and subcellular location label of the second protein predicted by the model, represents the predicted interaction label between two proteins, 0 represents no interaction between the two proteins, 1 represents interaction between the two proteins, [·] represents the vector concatenation operation, W and b are learnable parameters, shared by the two proteins;

[0059] In this way, protein functional labels, subcellular location labels, and predicted interaction labels were obtained;

[0060] On this basis, according to the binary cross entropy as the loss function, the losses of the three tasks, namely and can also be calculated; however, unlike the weighted summation or direct summation of predefined weights, the final loss is calculated through an uncertain loss mechanism, thereby dynamically balancing the loss scales between different tasks during the early training process to improve the generalization of the model, which is defined as:

[0061]

[0062] Among them, σ is used to quantify the variance and offset of a single task. It is not set manually, but is added to the set of learnable parameters so that it can be trained together with other model parameters. log(1+σ 2 ) represents the regularization term to avoid trivial solutions. Through joint training among the three tasks, the learnable parameters of the twin graph attention layer are jointly optimized by the three tasks, thereby promoting knowledge complementarity across biological fields and improving the prediction performance of the model in interaction prediction tasks.

[0063] Compared with the prior art, the present invention has the following advantages:

[0064] The present invention introduces protein synergy into the twin architecture by utilizing the interactive attention between protein structure pairs, thus realizing the knowledge sharing of interaction between proteins. In addition, task synergy based on isomorphic multi-task learning is introduced by the present invention, and protein synergy and task synergy are combined to realize the knowledge complementarity between tasks closely related to protein interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The present invention will be further described below in conjunction with the accompanying drawings.

[0066] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0067] The present invention is described in detail below with reference to the accompanying drawings.

[0068] Example 1

[0069] like Figure 1 As shown, a protein interaction prediction method based on protein synergy and task synergy proposed in the present invention comprises the following steps:

[0070] 1. Construction of residue contact map

[0071] First, the Euclidean distance between any pair of residues is calculated based on the three-dimensional coordinates of Alpha-C in the protein structure predicted by AlphaFold2; then, if the Euclidean distance d between residue i and residue j is i,j If it is less than the preset threshold t, it is considered as contact, and so on, so as to obtain a complete contact map, which can be defined as:

[0072]

[0073] Among them, A∈R N×N represents the residue contact map, R represents an arbitrary real number, N is the number of residues, and the threshold t is experimentally determined and is usually set to ( represents angstrom, i.e. 10 -12 M), residue contact map A∈R N×N It can be viewed as an adjacency matrix, which is used to represent the connection relationship between residues in the two-dimensional structure of a protein.

[0074] 2. Constructing Pre-trained Residual Embeddings

[0075] ESM-2 (Evolutionary Scale Modeling 2nd generation) is a protein language model pre-trained on the large sequence database UniRef released by the American research team Meta. It is used here to construct residue embeddings. Specifically, a protein sequence of length N is sent to the pre-trained ESM-2. ESM-2 can convert any residue in the protein into a continuous vector with a size of 1×1280. Due to ESM-2's implicit learning ability for protein co-evolution information, functional sites and other knowledge, the obtained vector contains rich prior knowledge related to protein interactions and can be directly used as residue embedding. On this basis, the feature matrix X∈R composed of stacked residue embeddings N×1280 Can be transferred to downstream protein tasks such as protein interaction prediction.

[0076] 3. Extracting protein structural features using twin graph attention layer

[0077] After the above two steps, a pair of proteins can be represented as two protein graphs and in V (1) and V (2) They are and The set of residue nodes, A (1) and A (2) They are and The adjacency matrix, X (1) and X (2) They are and On this basis, the twin graph attention layer is applied to extract the structural features of the protein pair. The biggest feature of the twin graph attention layer is that it is a graph attention network with two sets of parameters shared, with proteins For example, its residues are embedded in is input into the first set of attention layers, N1 is protein The number of residues, the corresponding structural features output by the graph attention layer Can be expressed as:.

[0078]

[0079] Among them, W k ∈R d×1280 is the learnable parameter for the linear mapping, L i is the first-order neighbor residue of residue i (from the adjacency matrix A (1) Get), || indicates splicing, represents the k-th normalized attention coefficient calculated by K-head attention:

[0080]

[0081] Here, a∈R 2d is the learnable weight vector, T is the transpose, W∈R d×1280 To calculate the learnable parameter of the attention coefficient, L i are the first-order neighbors of residue i. In this way, protein The feature matrix of is updated by the graph attention layer, namely This enables further feature extraction.

[0082] In addition, protein The above operation has also been applied to proteins The second set of graph attention layers Structural feature extraction is performed to can be obtained, where the first and second groups of graph attention layer parameters are shared and together constitute the twin graph attention layer.

[0083] 4. Protein collaboration based on interactive attention

[0084] After the above steps, the updated feature representation of the protein pair can be obtained, namely X (1)′ and X (2)′ However, the update of each representation only considers its own intrinsic structural or topological features, but ignores the information enhancement of its collaborative proteins. Therefore, interactive attention is introduced to achieve collaboration between proteins and promote their knowledge sharing.

[0085] The key to interactive attention is the calculation of the attention score, which fully combines the characteristics of the two proteins in the protein pair. and To this end, we designed three different strategies for calculating attention scores:

[0086]

[0087] in, and denote the embedding of the i-th residue in the first protein and the j-th residue in the second protein, respectively, d denotes the feature dimension, (·) T and ⊙ denote the transpose and Hadamard product respectively, U∈R d×d , V∈R d×d and w∈R 1×d There are three sets of learnable model parameters for and Perform a learnable linear mapping so that the attention score q of the jth residue in the second protein to the ith residue in the first protein is ij can be obtained. The calculation process can also be summarized as:

[0088] φ(X (1)′ , X (2)′ ) → Q

[0089] Where φ is: represents the strategy function, which is used to map a pair of feature maps of size N1×d and N2×d into an attention score matrix of size N1×N2 Afterwards, the row and column average operations are applied to This allows us to obtain the attention score of the entire protein to a single residue in another protein. and It can be defined as:

[0090]

[0091] Among them, N1 and N2 represent the first protein and the second protein On this basis, after the SoftMax operation, the final interaction attention score in the form of a probability vector of a pair of proteins can be obtained:

[0092]

[0093] Here, exp is an exponential function with the natural constant e as the base. Thus, the attention score of the second protein on the i-th residue in the first protein is and the attention score of the first protein to the jth residue in the second protein The knowledge of protein interactions is fully shared between protein pairs, so that the residues in each protein that contribute more to the interaction can be identified based on their co-proteins. Finally, the two attention scores are combined with the original protein representation and Multiply them together to adjust the residue importance and thus achieve synergy between protein pairs:

[0094]

[0095] 5. Task Collaboration Based on Isomorphic Multi-Task Learning

[0096] Typically, the resulting protein pair representation s (1) and (2) It can be directly input into a fully connected classifier for binary interaction prediction. However, inspired by the fact that a pair of proteins with the same subcellular location or the same function are more likely to interact, the present invention introduces homogeneous multi-task learning with two auxiliary tasks, protein function prediction and subcellular location prediction, to jointly fine-tune the learnable parameters of the twin attention layer, thereby achieving task collaboration, promoting knowledge complementarity in multiple biological fields, and improving the generalization ability of the interaction prediction model.

[0097] Example 2

[0098] Homogeneous multi-task learning with two auxiliary tasks of protein function prediction and subcellular location prediction is as follows:

[0099] 1) Protein function

[0100] A pair of interacting proteins usually perform a certain function together, so proteins with the same function are likely to interact. Here, the present invention collects Gene Ontology (GO) terms that describe protein functions from the UniProt database to accurately define the biological functions of each protein. Then the top 100 GO terms with the highest frequency of occurrence are recorded to form a fixed set of GO terms. On this basis, the functions of the protein are mapped to the fixed set using bag-of-words encoding, thereby avoiding the dimensionality explosion problem caused by the huge functional space. In this way, each protein can obtain a 100-dimensional protein function label y consisting of 0 and 1 func , where 0 / 1 indicates whether a protein has a certain function.

[0101] 2) Subcellular location

[0102] Subcellular location reveals the specific location of proteins in cells, such as plasma membrane, nucleus, nucleoplasm, etc., and a pair of proteins that are close to each other in structural space are more likely to interact, which can be inferred that proteins in the same subcellular location are more likely to interact. After collecting the subcellular location information of proteins, the present invention performs the same preprocessing operation on the subcellular location of each protein as the protein function label. Therefore, a 100-dimensional location label y can also be obtained. loc In addition, subcellular location prediction and function prediction can be viewed as multi-label classification problems, since a protein may appear in multiple subcellular locations or have multiple functions.

[0103] Afterwards, the three groups of stacked fully connected layers are regarded as classifiers for their respective tasks, which are used for different tasks, namely protein function prediction, subcellular location prediction, and interaction prediction, which can be defined as:

[0104]

[0105]

[0106]

[0107] in, Represents the functional label and subcellular location label of the first protein predicted by the model, represents the functional label and subcellular location label of the second protein predicted by the model, represents the predicted interaction label between two proteins, 0 represents no interaction between the two proteins, 1 represents interaction between the two proteins, [·] represents the vector concatenation operation, W and b are learnable parameters, shared by the two proteins.

[0108] In this way, protein function labels, subcellular location labels, and interaction prediction labels can be obtained. Based on this, according to the binary cross entropy as the loss function, the losses of the three tasks, namely and However, unlike the weighted summation or direct summation of predefined weights, the present invention calculates the final loss through an uncertain loss mechanism, thereby dynamically balancing the loss scales between different tasks during the early training process to improve the generalization of the model, which can be defined as:

[0109]

[0110] Among them, σ is used to quantify the variance and offset of a single task. It is not set manually, but is added to the set of learnable parameters so that it can be trained together with other model parameters. log(1+σ 2 ) represents the regularization term to avoid trivial solutions. Through joint training among the three tasks, the learnable parameters of the twin graph attention layer are jointly optimized by the three tasks, thereby promoting knowledge complementarity across biological fields and improving the prediction performance of the model in interaction prediction tasks.

Claims

1. A protein interaction prediction method based on protein synergy and task synergy, characterized in that: The following steps are involved: Step 1: Construction of residue contact map First, the Euclidean distance between any pair of residues is calculated based on the three-dimensional coordinates of Alpha-C in the protein structure predicted by AlphaFold2; then, if the Euclidean distance d between residue i and residue j is i,j If it is less than the preset threshold t, it is considered as contact, and so on, to obtain a complete contact map, which is defined as: Among them, A∈R N×N represents the residue contact map, N is the number of residues, and the threshold t is experimentally determined and set to Residue contact map A∈R N×N It is defined as an adjacency matrix, which is used to represent the connectivity between residues in the two-dimensional structure of a protein; Step 2: Build pre-trained residual embeddings The protein sequence of length N is sent to the pre-trained ESM-2, and any residue in the protein is converted into a continuous vector with a size of 1×1280 by ESM-2; the continuous vector is directly embedded as a residue; the feature matrix X∈R composed of stacked residue embeddings N×1280 transfer to downstream protein tasks; Step 3: Extract protein structural features using twin graph attention layer After steps 1 and 2, a pair of proteins is represented as two protein graphs. and in V (1) and V (2) They are and The set of residue nodes, A (1) and A (2) They are and The adjacency matrix, X (1) and X (2) They are and feature matrix; use the twin graph attention layer to extract the structural features of the protein pair; Step 4: Protein collaboration based on interactive attention After the above steps, the updated feature representation of the protein pair is obtained, namely X (1)′ and X (2)′ , interactive attention is introduced to achieve collaboration between proteins; Characteristics of each of the two proteins in a binding protein pair and Calculate the attention score of the interactive attention; for this, three different strategies are designed for the calculation of the attention score: in, and denote the embedding of the i-th residue in the first protein and the j-th residue in the second protein, respectively, d denotes the feature dimension, (·) T and ⊙ denote the transpose and Hadamard product respectively, U∈R d×d , V∈R d×d and w∈R 1×d There are three sets of learnable model parameters for and Perform a learnable linear mapping so that the attention score q of the jth residue in the second protein to the ith residue in the first protein is ij is obtained; after the SoftMax operation, the final interaction attention score in the form of a probability vector of a pair of proteins is obtained: Where exp is an exponential function with the natural constant e as the base; the attention score of the second protein on the i-th residue in the first protein is and the attention score of the first protein to the jth residue in the second protein The knowledge of protein interactions is fully shared between protein pairs, so that the residues in each protein that contribute more to the interaction are identified according to their co-proteins; finally, the two attention scores are combined with the original protein representation and Multiply them together to adjust the residue importance and thus achieve synergy between protein pairs: also, and The calculation process is summarized as: φ(X (1)′ ,X (2)′ )→Q in, represents the strategy function, which is used to map a pair of feature maps of size N1×d and N2×d into an attention score matrix of size N1×N2 Afterwards, the row and column average operations are applied to Thus, we can get the attention score of the whole protein to a single residue in another protein. and It is defined as: Among them, N1 and N2 represent the first protein and the second protein The number of residues in ; Step 5: Task collaboration based on homogeneous multi-task learning The protein pair obtained in step 4 is represented by s (1) and (2) It is directly input into the fully connected classifier for binary interaction prediction. This step introduces homogeneous multi-task learning with two auxiliary tasks of protein function prediction and subcellular location prediction to jointly fine-tune the learnable parameters of the twin attention layer.

2. A protein interaction prediction method based on protein synergy and task synergy according to claim 1, characterized in that: In step 3, the twin graph attention layer is a graph attention network with two sets of shared parameters. Its residues are embedded is input into the first set of attention layers, N1 is protein The number of residues, the corresponding structural features output by the graph attention layer is represented as: Among them, W k ∈R d×1280 is the learnable parameter for the linear mapping, L i is the first-order neighbor residue of residue i, which is obtained from the adjacency matrix A (1) Get, ∥ means splicing, represents the k-th normalized attention coefficient calculated by K-head attention: a∈R 2d is the learnable weight vector, T is the transpose, W∈R d×1280 To calculate the learnable parameter of the attention coefficient, L i is the first-order neighbor residue of residue i; in this way, the protein The feature matrix of is updated by the graph attention layer, namely So as to achieve further feature extraction; In addition, protein The above operation has also been applied to proteins The second set of graph attention layers Structural feature extraction is performed to is obtained, where the first and second groups of graph attention layer parameters are shared and together constitute the twin graph attention layer.

3. A protein interaction prediction method based on protein synergy and task synergy according to claim 1, characterized in that: The steps of isomorphic multi-task learning with auxiliary tasks of protein function prediction and subcellular location prediction in step 5 are as follows: 1) Protein function Gene ontology terms describing protein functions were collected from the UniProt database to accurately define the biological functions of each protein; the top 100 GO terms with the highest frequency were then recorded to form a fixed GO term set; On this basis, the bag-of-words encoding is used to map the functions of proteins to the fixed GO term set, thereby avoiding the dimensionality explosion problem caused by the huge functional space; in this way, each protein obtains a 100-dimensional protein function label y consisting of 0 and 1 func , where 0 / 1 indicates whether a protein has a certain function; 2) Subcellular location After collecting the subcellular location information of proteins, the subcellular location of each protein is subjected to the same preprocessing operation as the protein function label, and a 100-dimensional location label y is obtained. loc ; Afterwards, 3 sets of stacked fully connected layers are regarded as classifiers for their respective tasks, i.e. protein function prediction, subcellular location prediction and interaction prediction, defined as: in, Represents the functional label and subcellular location label of the first protein predicted by the model, represents the functional label and subcellular location label of the second protein predicted by the model, represents the predicted interaction label between two proteins, 0 represents no interaction between the two proteins, 1 represents interaction between the two proteins, [·] represents the vector concatenation operation, W and b are learnable parameters, shared by the two proteins; In this way, protein functional labels, subcellular location labels, and predicted interaction labels were obtained; On this basis, according to the binary cross entropy as the loss function, the losses of the three tasks, namely and is also calculated; however, unlike the weighted summation or direct summation of predefined weights, the final loss is calculated through an uncertain loss mechanism, thereby dynamically balancing the loss scales between different tasks during the early training process to improve the generalization of the model, which is defined as: Among them, σ is used to quantify the variance and offset of a single task. It is not set manually, but is added to the set of learnable parameters so that it can be trained together with other model parameters. log(1+σ 2 ) represents the regularization term to avoid trivial solutions. Through joint training among the three tasks, the learnable parameters of the twin graph attention layer are jointly optimized by the three tasks.