A Protein-Protein Interaction Prediction Method Based on Cross-Graph Representation Learning
By constructing a protein map, using GCN and self-attention module to learn the spatial structure and context information of proteins, combined with the interactive map module to learn residue information, the problem of performance degradation in existing methods is solved and more efficient protein interaction prediction is achieved.
Patent Information
- Application Number
- CN202411939454.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing methods for protein interaction prediction show performance degradation problems in different fields, especially ignoring the spatial information and fine-grained modeling of proteins.
Using a method based on cross-graph representation learning, a protein map is constructed, a spatial structure information is learned using GCN, and a self-attention module is used to learn information between proteins, and a residue information is learned through the dual interaction map module, which is finally generated a classifier for protein interaction.
It improves the accuracy and consistency of protein interaction prediction, can better capture the relationship between proteins, and improves prediction performance.
Smart Images

Figure CN119832991B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of protein - protein interactions, and in particular, to a method for predicting protein - protein interactions based on cross - graph representation learning. Background Art
[0002] Protein - protein interaction prediction (PPI) is to determine whether two proteins can interact through various technical means. Proteins are the basic substances of all life, and the interactions between proteins are crucial for life. Protein - protein interaction prediction is involved in almost all biological processes, including metabolism, genetic pathways, and signal cascades, etc., where they are used in DNA replication and transcription, RNA translation, energy production, signal transduction, immunity, etc. The interactions between proteins determine molecular and cellular mechanisms and are involved in most biological processes in living organisms. It has important applications in interpreting pathogenic variants, developing drugs targeting PPIs, and designing protein binders to regulate protein functions. There are various methods for studying protein - protein interaction prediction:
[0003] 1) Methods based on experimental analysis: Various experimental analyses have been widely used in the identification of protein - protein interaction prediction (PPI). PPI detection methods include widely used low - throughput methods such as co - immunoprecipitation (Co - IP) and bioluminescence resonance energy transfer (BRET), and a series of high - throughput methods such as yeast two - hybrid (Y2H) screening. X - ray crystallography and nuclear magnetic resonance (NMR) are common methods for determining the structure of protein complexes. However, the complexity of PPI makes experimental analysis time - consuming, labor - intensive, and very expensive.
[0004] 2) Methods based on machine learning: Machine learning algorithms such as support vector machines and random forests are used to extract features of protein sequences and structures and train to establish a protein - protein interaction prediction model. Such methods can be relatively simple to calculate, can automatically learn the relationship between proteins, and have a certain adaptability to different individuals and environments. The disadvantage is that the quality and selection of feature extraction have a greater impact on the results, and a large amount of labeled data is required for training.
[0005] 3) Deep learning-based methods have powerful non-linear transformation capabilities and have received extensive attention in various fields. Deep learning-based algorithms are also increasingly applied in the prediction of protein-protein interactions. Compared with machine learning, deep learning methods can automatically extract effective features from large-scale raw datasets without prior knowledge. Although great progress has been made in PPI prediction using deep learning methods, there are still some problems. Firstly, many methods mainly rely on recurrent neural networks or convolutional neural networks to capture the features of primary sequences in one-dimensional space, while ignoring the spatial information of protein molecules. Secondly, the regions of protein-protein interaction account for a relatively small part of the protein, so fine-grained modeling is quite necessary for PPI prediction, which has not been fully studied in previous research. Summary of the Invention
[0006] The object of the present invention is to provide a method for predicting protein-protein interactions based on cross-graph representation learning, which solves the problem of performance degradation of existing methods for predicting protein-protein interactions in different fields.
[0007] To achieve the above object, the present invention provides a method for predicting protein-protein interactions based on cross-graph representation learning, comprising the following steps:
[0008] S1, collect a prediction dataset of protein-protein interactions, and then perform feature processing on the proteins to construct a protein graph;
[0009] S2, use a graph encoder based on GCN to learn the spatial structure information of the protein graph;
[0010] S3, use an encoder based on a self-attention module to learn the information between receptor proteins and ligand proteins;
[0011] S4, use a dual interaction graph module to learn the information of residues between receptor proteins and ligand proteins;
[0012] S5, generate a classifier for predicting protein-protein interactions.
[0013] Preferably, in S1, it is necessary to clean the prediction dataset to remove too short sequences and too long sequences in the protein sequences.
[0014] Preferably, in S1, the specific steps for constructing the protein graph are as follows:
[0015] S11, input the cleaned protein sequences into the ESM2 pre-trained model;
[0016] S12, through the ESM2 pre-trained model, a series of protein sequences obtain feature embeddings to obtain a series of residue embedding representations;
[0017] S13 measures the similarity between residues, determines the edge relationship between residues, and is defined as follows:
[0018]
[0019] Among them, Z l is the feature embedding of the ligand protein, and M l (Z l ) is the relationship matrix between residue sequences;
[0020] S14. For the input protein pair, we use residues as nodes, and its edge relationship is defined as in S13 to construct a protein graph:
[0021]
[0022] Among them, G l represents the ligand protein graph, Z l represents the feature embedding of the ligand protein, and M l represents the edge relationship between nodes;
[0023] Finally, repeat the steps from S11 to S14 to obtain the receptor protein graph: (G r (Z r , M r ))
[0024] Preferably, the specific process of S2 is
[0025] Use GCN to learn the representation of nodes based on their neighbor nodes, and assign weights according to the distance between neighbor nodes and the central node. The graph encoding process is as follows:
[0026] First, normalize the adjacency matrix. The normalization process is to add self-loops (each node is connected to itself) to the adjacency matrix and normalize it using the degree matrix:
[0027]
[0028] where I is the identity matrix, representing the self-loop where each node is connected to itself, is the semi-inverse matrix of the degree matrix;
[0029] Secondly, perform feature aggregation. For each layer of GCN, aggregate the features of nodes through the normalized adjacency matrix The feature of each node in the graph is represented as X (k) at the k-th layer, and the node feature representation X (k+1) at the k + 1-th layer is then calculated by the following formula:
[0030]
[0031] where X(k) represents the node feature matrix of the k-th layer, W (k) represents the weight matrix of the k-th layer, represents the normalized adjacency matrix, σ represents the activation function, and the activation function used is RELU;
[0032] Then, a graph encoder is used to learn the ligand protein graph and the receptor protein graph, and the ligand protein internal feature vector and the receptor protein internal feature vector Taking the ligand protein internal feature vector as an example, the formula is as follows:
[0033]
[0034] where softmax(Z l ) represents the feature of the ligand protein generated by the graph encoder.
[0035] Preferably, the specific process of S3 is, taking the ligand protein internal feature vector as an example, inputting the ligand protein sequence into the self-attention model to obtain the context representation of each residue:
[0036]
[0037]
[0038] where Q, K, and V represent the query, key, and value vectors of the ligand protein, respectively represent the corresponding weights initialized by the neural network, d K represents the dimension of the feature vector K;
[0039] The residue context representation of the receptor protein is obtained in the same way.
[0040] Preferably, the dual interaction graph module includes a normalization layer, a multi-head cross-attention module, a residual connection, and a feed-forward neural network layer. The multi-head cross-attention module is as follows:
[0041]
[0042]
[0043] Q r is the query vector of the receptor protein, K l , V l are the key and value vectors of the ligand protein, and m is the number of cross-attention modules;
[0044] Then, a feed-forward neural network is used to learn the sequence features to obtain the receptor protein feature vector and the ligand protein feature vector wherein the ligand protein feature vector is obtained as follows:
[0045]
[0046] wherein represents the features of the ligand protein generated by the protein interaction module, represents the feature information learned by the inter-protein representation learning module;
[0047] Similarly, the receptor protein internal feature vector is input into a feed-forward neural network for learning to obtain the receptor protein feature vector
[0048] Preferably, in S5, a learnable parameter λ is set to fuse the receptor protein information and the ligand protein information.
[0049]
[0050]
[0051] where, H l represents the fused ligand protein information, H r represents the fused receptor protein information;
[0052] Then H l and H r are concatenated together, and a fully connected layer is used as the classifier for protein interaction, with binary cross-entropy loss as the objective:
[0053]
[0054]
[0055] wherein, is the number of training samples;
[0056] The probability of contact between each residue between proteins is predicted through the fused receptor protein information and ligand protein information.
[0057] Therefore, the protein interaction prediction method based on cross-graph representation learning with the above structure of the present invention has the following advantages:
[0058] 1) Considering that proteins have a certain spatial structure, a graph model is constructed with residues as nodes and edge information constructed based on the similarity between nodes, increasing the spatial information of proteins.
[0059] 2) To learn the context information of protein sequences, a self-attention encoder is adopted to learn the long-range dependencies between proteins.
[0060] 3) A dual interaction graph module is proposed to learn the fine-grained representations between proteins.
[0061] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0062] Figure 1 It is a schematic diagram of the technical network of a protein interaction prediction method based on cross-graph representation learning according to the present invention;
[0063] Figure 2 It is a schematic flowchart of a protein interaction prediction method based on cross-graph representation learning according to the present invention;
[0064] Figure 3 It is to compare the effectiveness of the number of graph convolutional layers of the protein inner representation learning module in a protein interaction prediction method based on cross-graph representation learning according to the present invention;
[0065] Figure 4 It is to compare the influence of the number of layers of the interaction graph module on protein representation learning in a protein interaction prediction method based on cross-graph representation learning according to the present invention. Detailed Embodiments
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0067] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0068] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0069] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0070] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "install", "connect" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0071] The following will describe in detail some embodiments of the present invention with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0072] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper", "lower", "left", "right", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0073] The specific model specifications need to be selected and determined according to the actual specifications of the device, etc. The specific selection calculation method adopts the existing technology in the art, so it will not be elaborated in detail.
[0074] Embodiment
[0075] As Figure 1 , Figure 2 shown, it includes the following steps:
[0076] S1. Collect the prediction dataset of protein-protein interactions, then perform feature processing on the proteins to construct a protein graph;
[0077] In S1, it is necessary to clean the prediction dataset to remove the too-short sequences and too-long sequences in the protein sequences.
[0078] The specific steps for constructing the protein graph in S1 are as follows:
[0079] S11. Input the cleaned protein sequences into the ESM2 pre-trained model;
[0080] As a pre-trained model, ESM2 is an unsupervised Transformer model trained on protein sequences in the UniRef database. ESM2 only has a simple masked language modeling objective during training, but it has effectively succeeded in various tasks by leveraging the diversity of protein sequences. Using the ESM2 model as a pre-trained model provides rich protein embedding information (i.e., representation information in a high-dimensional space) for each protein sequence. These embeddings implicitly confer structural information about each protein, generating 2560-dimensional embeddings for each protein sequence. It is used for the construction of ligand and receptor protein graph models.
[0081] S12, obtaining feature embeddings for a series of protein sequences through the ESM2 pre-trained model to get a series of residue embedding representations;
[0082] S13, measuring the similarity between residues, judging the relationship of the edges between residues, defined as follows:
[0083]
[0084] where, Z l is the feature embedding of the ligand protein, and M l (Z l ) is the relationship matrix between residue sequences;
[0085] S14, for the input protein pair, we construct a protein graph with residues as nodes and their edge relationships defined as in S13:
[0086]
[0087] where, G l represents the ligand protein graph, Z l represents the feature embedding of the ligand protein, and M l represents the edge relationship between nodes;
[0088] Finally, repeat steps S11 to S14 for the receptor protein graph: (G r (Z r , M r ))).
[0089] S2, using a GCN-based graph encoder to learn the spatial structure information of the protein graph;
[0090] The specific process of S2 is
[0091] Using GCN to learn the representation of nodes based on their neighbor nodes, and assigning weights according to the distance between neighbor nodes and the central node. The graph encoding process is as follows:
[0092] First, normalize the adjacency matrix to avoid the influence of degree differences in the graph on the convolution operation. The normalization method is to add self-loops to the adjacency matrix (each node is connected to itself), and use the degree matrix for normalization:
[0093]
[0094] where \(I\) is the identity matrix, representing the self-loop where each node is connected to itself, and \(\D^{-1 / 2}\) is the semi-inverse matrix of the degree matrix.
[0095] Secondly, perform feature aggregation. For each layer of GCN, aggregate the features of nodes through the normalized adjacency matrix . Represent the features of each node in the graph at the \(k\)-th layer as \(X^{(k)}\) (k) . Then the node feature representation \(X^{(k + 1)}\) at the \((k + 1)\)-th layer (k+1) can be calculated by the following formula:
[0096]
[0097] where \(X^{(k)}\) (k) represents the node feature matrix at the \(k\)-th layer, \(W^{(k)}\) (k) represents the weight matrix at the \(k\)-th layer, represents the normalized adjacency matrix, and \(\sigma\) represents the activation function. Here, the RELU (Rectified Linear Unit) activation function is usually used.
[0098] In the GCN-based graph encoding module, aggregate neighbor information through the adjacency structure of the graph, and at the same time eliminate the influence of node degree differences through the normalized adjacency matrix. Finally, stack multiple convolutional layers to extract more complex graph structure features, and update the network parameters through backpropagation. Then use the graph encoder to learn the ligand protein graph and the receptor protein graph, and obtain the ligand protein inner feature vector and the receptor protein inner feature vector Taking the ligand protein inner feature vector as an example, the formula is as follows:
[0099]
[0100] where \(\text{softmax}(Z^{(L)})\) l represents the features of the ligand protein generated by the graph encoder.
[0101] S3. Use the encoder based on the self-attention module to learn the information between the receptor protein and the ligand protein;
[0102] The specific process of S3 is as follows. Taking the ligand protein inner feature vector For example, the ligand protein sequence is input into the self-attention model to obtain the context representation of each residue:
[0103]
[0104]
[0105] where Q, K, and V represent the query, key, and value vectors of the ligand protein, representing the corresponding weights initialized by the neural network respectively, and d K represents the dimension of the feature vector K;
[0106] The residue context representation of the receptor protein is obtained in the same way.
[0107] For the feature vectors of the protein pair obtained from the previous module and First, a protein feature encoder is introduced to learn the information between the contexts of individual protein residues. Then, a protein interaction module is designed to learn the information between the two proteins in a fine-grained manner.
[0108] S4, using a dual interaction graph module to learn the information of the residues between the receptor protein and the ligand protein;
[0109] The dual interaction graph module includes a normalization layer, a multi-head cross-attention module, a residual connection, and a feed-forward neural network layer. The multi-head cross-attention module is as follows:
[0110]
[0111]
[0112] Q r is the query vector of the receptor protein, K l , V l are the key and value vectors of the ligand protein, and m is the number of cross-attention modules;
[0113] Then, a feed-forward neural network is used to learn the sequence features to obtain the receptor protein feature vector and the ligand protein feature vector where the ligand protein feature vector is obtained as follows:
[0114]
[0115] where represents the feature of the ligand protein generated by the protein interaction module, represents the feature information learned through the inter-protein representation learning module;
[0116] Similarly, the eigenvector within the receptor protein is input into the feedforward neural network for learning to obtain the receptor protein eigenvector
[0117] When the receptor protein is the query vector, the ligand protein serves as the key and value vectors, and the residues of the receptor protein learn fine-grained representations from the ligand protein. This module consists of a normalization layer, a multi-head cross-attention module, residual connections, and a feedforward neural network layer. Residual connections are employed around the multi-head cross-attention module and the feedforward layer. Residual connections can effectively learn the residuals, thereby preventing the vanishing gradient problem. The feedforward neural network layer is executed to represent the attention passed through the multi-head cross-attention module. In the case of the multi-head cross-attention module, it can be extended from the m-th single-head cross-attention module. Then, a feedforward neural network is used to learn the sequence features. The feedforward neural network consists of two fully connected layers and an activation function for performing non-linear transformations and mapping the representations at each position.
[0118] S5. Generate a classifier for protein interaction prediction.
[0119] In S5, set the learnable parameter λ to fuse the receptor protein information and the ligand protein information.
[0120]
[0121]
[0122] where, H l represents the fused ligand protein information, and H r represents the fused receptor protein information;
[0123] Then, concatenate H l and H r together and use a fully connected layer as the classifier for protein interaction, with binary cross-entropy loss as the objective:
[0124]
[0125]
[0126] where, is the number of training samples;
[0127] Predict the probability of contact between each residue of the protein-protein interaction through the fused receptor protein information and ligand protein information.
[0128] The present invention was verified on the commonly used datasets of yeast and H. sapiens for protein - interaction prediction, compared with other benchmark methods, and finally the effectiveness of the module was analyzed.
[0129] Experimental settings
[0130] Prediction datasets: The prediction datasets include yeast and H. sapiens datasets. The intra - species PPI dataset of yeast is a benchmark dataset widely used in state - of - the - art methods. It contains 2,497 proteins, among which there are 5,594 positive samples and 5,594 negative samples. The high - quality yeast dataset is extracted from the interacting protein database and only contains the most reliable physical interactions. Randomly pairing proteins without interaction evidence generates negative interactions. The H. sapiens dataset contains 15,816 proteins, among which there are 38,345 positive samples and 383,450 negative samples.
[0131] Model experiment details: Specific details: For the input of the sequence information of protein pairs, a pre - trained model (ESM2) is adopted to generate 2,560 - dimensional features. For the representation learning module within proteins, a graph encoder is adopted, which consists of three - layer GCN modules to generate 512 - dimensional feature representations for in - protein learning. For the representation learning module between proteins, a self - attention module is used as the encoder to learn the context information of the sequence. Then, a multi - head cross - attention module is used to learn the information between sequences. Finally, the residue representations of the protein pair are concatenated and fed into a fully - connected layer for binary prediction. The rectified linear unit (ReLU) is used as the activation function, and the dropout is set to 0.4 to avoid overfitting. The model is trained for 100 epochs using ADAM on an NVIDIA RTX 4090 GPU card with 24GB of memory, where the learning rate is 0.0001. The cross - entropy loss is used to train the model in an end - to - end manner.
[0132] Experimental results
[0133] The comparison of the performance of the model with benchmark methods on the yeast dataset is as follows.
[0134]
[0135] In the table, the model of the present invention was compared with several state-of-the-art methods on the yeast dataset, namely MCD-SVM, RF-LPQ, KNN-CTD, EELM-PCA, DeepPPI, SAE, DPPI, DNN-PPI, PIPR, and TAGPPI. This clearly shows that the method of the present invention achieves new state-of-the-art performance and obtains significant improvements on all yeast datasets. Specifically, the performance of the method of the present invention on ACC is 0.9798, on Pre is 0.9812, on Sen is 0.9785, on Spe is 0.9812, on F1 is 0.9798, on MCC is 0.9597, and on AUC is 0.9798. Among all the metrics, six metrics reach the highest and one metric ranks second. All of the above results prove the effectiveness of the framework proposed by the present invention. During the learning process, the gain is attributed to the information learning within the proposed proteins and the information propagation of the inter-protein learning, which better captures the encoding of the relationships between residues.
[0136] Analysis of Experimental Results
[0137] The present invention has achieved superior performance compared with the existing state-of-the-art methods. Therefore, in addition to performance evaluation, additional experiments were conducted to clarify how each module promotes the classification task and the sensitivity of the hyperparameters used in the method.
[0138] As Figure 3 shown, for a comprehensive comparison, experiments were conducted using multiple graph convolutional layers in the graph encoder for in-graph learning. It can be observed that the Accuracy of stacking three convolutional layers is higher than other cases. In addition, when only one convolutional layer is used, the aggregated information of neighbors is insufficient. Moreover, due to over-smoothing, more graph convolutional layers will lead to a decrease in performance.
[0139] As Figure 4 shown, to study the influence of the interaction graph module in our framework, multiple layers were stacked and their performances were compared. Among them, the performance varies with the number of layers of the interaction graph module. When the number of layers is set to 0, it means that the information of ligand and receptor proteins does not propagate to each other, and the Accuracy decreases significantly. On the contrary, when two layers are combined, the best Accuracy can be obtained. This shows that the information from another protein effectively improves the residue representation and promotes the recognition of residue pair relationships.
[0140] Therefore, the present invention adopts a protein interaction prediction method based on cross-graph representation learning with the above structure, which solves the problem that the existing protein interaction prediction methods often show performance degradation in different fields.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for predicting protein-protein interactions based on cross-graph representation learning, characterized in that: It includes the following steps: S1. Collect a prediction dataset of protein-protein interactions, then perform feature processing on the proteins to construct a protein graph; S2. Use a graph encoder based on GCN to learn the spatial structure information of the protein graph; S3. Use an encoder based on a self-attention module to learn the information between receptor proteins and ligand proteins; S4. Use a dual interaction graph module to learn the information of residues between receptor proteins and ligand proteins; S5. Generate a classifier for predicting protein-protein interactions; The specific process of S3 is as follows. Taking the eigenvector in the ligand protein as an example, the ligand protein sequence is input into the self-attention model to obtain the context representation of each residue: ; ; Among them 、 and represent the query, key, and value vectors of the ligand protein, , , respectively represent the corresponding weights for the initialization of the neural network, represents the feature vector dimension; The residue context representation of the receptor protein is obtained in the same way; The dual interaction graph module includes a normalization layer, a multi-head cross-attention module, a residual connection, and a feed-forward neural network layer. The multi-head cross-attention module is as follows: ; ; is the query vector of the receptor protein, , is the key and value vectors of the ligand protein, is the number of cross-attention modules; Then, a feed-forward neural network is used to learn sequence features to obtain the receptor protein feature vector and the ligand protein feature vector , where the ligand protein feature vector is obtained as follows: ; Among them represents the characteristics of the ligand protein generated by the general protein interaction module, represents the feature information learned through the protein-interaction-based representation learning module; Similarly, the eigenvector inside the receptor protein is input into the feedforward neural network for learning to obtain the eigenvector of the receptor protein.
2. The protein-protein interaction prediction method based on cross-graph representation learning according to claim 1, wherein: In S1, it is necessary to clean the prediction dataset to remove too short sequences and too long sequences in the protein sequences.
3. The method for predicting protein-protein interactions based on cross-graph representation learning according to claim 2, wherein: In S1, the specific steps for constructing the protein graph are as follows: S11. Input the cleaned protein sequences into the ESM2 pre-trained model; S12. Through the ESM2 pre-trained model, a series of protein sequences obtain feature embeddings to get a series of residue embedding representations; S13. Measure the similarity between residues, judge the edge relationship between residues, and define as follows: ; Among them, is the characteristic embedding of the ligand protein, is the relationship matrix between residue sequences; S14. For the input protein pair, use residues as nodes, and its edge relationship is defined as in S13 to construct a protein graph: ; Among them, represents the ligand protein map, represents the feature embedding of the ligand protein, represents the relationship of the edges between nodes; Finally, repeat steps S11 to S14 to obtain the receptor protein map: .
4. A method for predicting protein-protein interactions based on cross-graph representation learning according to claim 3, characterized in that: The specific process of S2 is Use GCN to learn the representation of nodes based on their neighbor nodes, and assign weights according to the distance between neighbor nodes and the central node. The graph encoding process is as follows: First, perform normalization processing on the adjacency matrix. The normalization process is to add a self-loop to the adjacency matrix and use the degree matrix for normalization: ; Among them is the identity matrix, representing the self-loop where each node is connected to itself, is the semi-inverse matrix of the degree matrix; Next, perform feature aggregation. For each layer of GCN, aggregate the features of nodes through the normalized adjacency matrix and represent the features of each node in the graph at the k-th layer as . The node feature representation at the (k + 1)-th layer is then calculated using the following formula: ; Among them represents the node feature matrix of the k-th layer, represents the weight matrix of the k-th layer, represents the normalized adjacency matrix, represents the activation function, and the activation function used is RELU; Then, the graph encoder is used to learn the ligand protein graph and the receptor protein graph, respectively obtaining the ligand protein internal feature vector and the receptor protein internal feature vector . Taking the ligand protein internal feature vector as an example, the formula is expressed as follows: ; Among them represents the features of the ligand protein generated by the graph encoder.
5. The protein interaction prediction method based on cross-graph representation learning according to claim 4, wherein: In S5, set a learnable parameter λ to fuse the receptor protein information and ligand protein information, ; ; Among them, represents the ligand protein information after fusion, represents the receptor protein information after fusion; Then and are concatenated together, and a fully connected layer is used as the classifier for protein-protein interactions, with binary cross-entropy loss as the objective: ; ; wherein, is the number of training samples; Through the fused receptor protein information and ligand protein information, predict the probability of contact between each residue between proteins.
Citation Information
Patent Citations
Construction method of anti-programmed death protein-1 monoclonal antibody treatment response prediction model
CN119152925A
Protein-protein interaction map inference using interacting domain profile pairs
WO2002074901A2