A protein interaction prediction method, device and storage medium

By combining the Doc2vec model and graph isomorphic convolutional network, the problem of low accuracy in protein interaction prediction in the prior art is solved, and more efficient and accurate protein interaction prediction is achieved, especially when processing new proteins.

CN116259358BActive Publication Date: 2025-05-13SHENZHEN XIANGGAN SCIENCE & TECHNOLOGY ACHIEVEMENTS TRANSFORMATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211678119.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-05-13
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

In the prior art, the prediction accuracy of protein interactions is low, and the sequence information and PPI network structure information are not fully utilized, resulting in unsatisfactory prediction effect of new proteins.

Method used

The Doc2vec model is used to embed amino acid sequences into low-dimensional vector space, and combined with the graph isomorphic convolution network, aggregate the adjacent information of proteins, optimize the node encoding characterization, and finally classify prediction through a multi-layer perceptron.

Benefits of technology

It improves the accuracy of protein interaction prediction, can process protein sequences of any length, makes full use of PPI network structure information, and improves the prediction performance of multiple types of PPIs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116259358B_ABST
    Figure CN116259358B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of biological technology, and in particular to a protein interaction prediction method, device, equipment and storage medium. The protein interaction prediction method of the present invention can process protein sequences of any length by adjusting the Doc2vec unsupervised paragraph vector learning model, embedding the feature information of protein sequences of indefinite length into a low-dimensional vector space, and solves the problem of initial feature selection of proteins. It utilizes the advantages of graph isomorphic convolutional networks, fully combines the information of the protein interaction PPI network structure, aggregates the information of adjacent proteins of each protein, optimizes the encoding representation of protein nodes, finds the protein interaction edge according to the PPI network structure information, combines the information of two protein nodes, and continuously learns more efficient and accurate classification predictions therefrom, thereby improving the accuracy of model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological technology, and in particular to a protein interaction prediction method, device, equipment and storage medium. Background Art

[0002] In existing technologies, protein-protein interaction (PPI) refers to the correlation between protein molecules, and this correlation is studied from the perspectives of biochemistry, signal transduction, and genetic networks. In recent years, with the development of high-throughput screening technology, the number of protein interactions detected by experimental methods has increased significantly, forming more and more protein interaction networks. The analysis of protein interaction networks can enhance the understanding of biological processes, and the functional prediction of PPIs in the network is of great significance in cell biology, such as new drug design and target therapy. Some existing PPI prediction methods have the following problems:

[0003] "Path2PPI: an R package to predict protein-protein interaction networks for a set of proteins" (Journal source: Bioinformatics, 2016, 32: 1427-1429) integrates homology methods and maps a pair of sequences to known interacting proteins by calculating the BLAST value of the protein, thereby inferring new PPIs. "Using support vector machine combined with auto covariance to predict protein-protein interactions from protein sequences" (Journal source: Nucleic acids research, 2008, 36 (9): 3025-3030) integrates neighbor effects and proposes a new feature representation that combines autocovariance (AC) and support vector machine (SVM). “Detection of protein-protein interactions from amino acid sequences using a rotation forest model with a novel pr-lpqdescriptor” (Journal source: In International Conference on Intelligent Computing, 2015, 10(8): 713-720) uses the physicochemical property response matrix (PR) to transform the sequence into a matrix, uses the texture descriptor of local phase quantization (LPQ) to extract the local phrase information matrix, and combines the random forest (RF) model with the new feature representation to detect PPI. “A method for predicting protein-protein interaction types” (Journal source: PLoS One, 2014, 9(3): e90904) integrates experimental techniques for detecting interactions and uses logistic regression (LR) to predict interaction types. “Predicting protein–protein interactions based only on sequences information” (Journal source: Proceedings of the National Academy of Sciences, 2007, 104(11): 4337-4341) integrates support vector machine (SVM) and combines the joint triplet features describing amino acids with sequence information to predict PPI.These methods rely heavily on the ability to extract and select better features, so their performance is limited by the PPI feature representation and model expression capabilities. "Deep neural network based predictions of protein interactions using primary sequences" (Journal source: Molecules, 2018, 23 (8): 1923-1935), "Predicting protein–protein interactions through sequence-based deep learning" (Journal source: Bioinformatics, 2018, 34 (17): i802-i810), "Multifaceted protein–protein interaction prediction based on siamese residual rcnn" (Journal source: Bioinformatics, 2019, 35 (14): i305-i314) use convolutional neural networks (CNN), recurrent neural networks (RNN) and regional convolutional neural networks (R-CNN) to extract high-dimensional information features from sequences, thereby improving the model prediction performance in PPI-related tasks. Neural Network for Inter-novel-protein Interaction Prediction" (Journal Source: https: / / arxiv.org / abs / 2105.06709, 2021.) for the first time extended graph neural networks to multi-label PPI classification. Compared with early machine learning methods, the above models have a certain depth, the nonlinear modeling ability has been enhanced, and the performance of complex tasks such as multi-type PPI prediction has been continuously improved, but the structural information of the PPI network has been ignored, there are certain limitations, and the accuracy needs to be improved.

[0004] The above methods do not make full use of sequence information, and need to process proteins of different lengths into fixed-length sequences; although some methods use PPI network structure information, due to the limitation of protein sequence representation ability, the prediction results of the effects on new proteins that have never appeared are not ideal. Therefore, improving the accuracy of protein interaction prediction results is a problem to be solved at present. Summary of the invention

[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem of low accuracy in protein interaction prediction in the prior art.

[0006] To solve the above technical problems, the present invention provides a method for predicting protein-protein interactions, including:

[0007] Obtain the amino acid sequence of the i-th protein in the protein-protein interaction network, where i = 1...n, and n is the total number of proteins in the protein-protein interaction network;

[0008] Starting from the j-th amino acid of the amino acid sequence, successively intercept multiple k-mers backward until the remaining number of amino acids < k, to obtain the j-th amino acid subsequence, where j = 1...k;

[0009] Use the Doc2vec model to predict the central k-mer in each sliding window of the j-th amino acid subsequence and perform an aggregation operation to obtain the j-th amino acid subsequence feature;

[0010] Perform a mean operation on the k amino acid subsequence features to obtain the i-th protein feature;

[0011] Aggregate the features of multiple neighbor proteins of the i-th protein to itself to obtain the updated i-th protein feature, and repeat this step multiple times to obtain the i-th protein aggregation feature;

[0012] Input the dot product of the i-th protein aggregation feature and the protein aggregation feature that interacts with it into a classifier to obtain the prediction result of the i-th protein-protein interaction.

[0013] Preferably, the step of using the Doc2vec model to predict the central k-mer in each sliding window of the j-th amino acid subsequence and perform an aggregation operation to obtain the j-th amino acid subsequence feature includes:

[0014] Calculate the average value of the embedded encoding of the j-th amino acid subsequence and its context k-mer in the t-th sliding window, and perform classification prediction according to the average value to obtain the central k-mer in the t-th sliding window of the j-th amino acid subsequence.

[0015] Preferably, after performing the mean operation on the k amino acid subsequence features to obtain the i-th protein feature, it further includes:

[0016] Input the i-th protein feature into a stacked one-dimensional convolutional network for processing, where the processing process of each layer of the one-dimensional convolutional network is:

[0017] Take the output of the previous layer of the one-dimensional convolutional network as the input of the current layer of the one-dimensional convolutional network. After performing multiple convolution operations on the input, process it through the Relu activation function and the max pooling layer to obtain the output of the current one-dimensional convolutional network;

[0018] The i-th protein feature after being processed by the stacked one-dimensional convolutional network is processed by a fully connected layer.

[0019] Preferably, the formula for aggregating the features of multiple neighbor proteins of the i-th protein into itself to obtain the updated i-th protein feature and repeating this step multiple times to obtain the aggregated feature of the i-th protein is:

[0020]

[0021] Where, is the i-th protein feature after the previous update, is the i-th protein feature after the current update, is the u-th protein feature after the previous update, N(i) represents the set of corresponding serial numbers of the neighbor proteins of the i-th protein, f() represents feature aggregation, and φ() represents a mapping function.

[0022] Preferably, the feature aggregation is performed in an accumulative manner, and the mapping function uses a multi-layer perceptron:

[0023]

[0024] Where ε is a hyperparameter, and MLP() represents a multi-layer perceptron.

[0025] Preferably, the classifier is a multi-layer perceptron.

[0026] Preferably, the formula for obtaining the prediction result of the interaction of the i-th protein by taking the dot product of the aggregated feature of the i-th protein and the aggregated feature of the protein that interacts with it and then inputting it into the classifier is:

[0027]

[0028] Where, h i is the aggregated feature of the i-th protein, and h i+1 is the aggregated feature of the protein that interacts with the aggregated feature of the i-th protein.

[0029] The present invention also provides a protein interaction prediction device, including:

[0030] An amino acid sequence acquisition module for acquiring the amino acid sequence of the i-th protein in the protein interaction network, where i = 1...n, and n is the total number of proteins in the protein interaction network;

[0031] An amino acid subsequence acquisition module for starting from the j-th amino acid of the amino acid sequence and sequentially intercepting multiple k-mers backward until the remaining number of amino acids < k to obtain the j-th amino acid subsequence, where j = 1...k;

[0032] The amino acid subsequence feature acquisition module is used to predict the central k-mer in each sliding window in the j-th amino acid subsequence using the Doc2vec model and perform aggregation operations to obtain the j-th amino acid subsequence feature;

[0033] The protein feature acquisition module is used to perform mean operation on the k amino acid subsequence features to obtain the i-th protein feature;

[0034] The protein aggregation feature acquisition module is used to aggregate the features of multiple neighboring proteins of the ith protein into itself to obtain the updated ith protein feature, and repeat this step multiple times to obtain the ith protein aggregation feature;

[0035] The protein interaction prediction module is used to perform dot product between the aggregation feature of the ith protein and the aggregation feature of the protein that interacts with it and then input the result into the classifier to obtain the prediction result of the ith protein interaction.

[0036] The present invention also provides a protein interaction prediction device, comprising:

[0037] The memory is used to store the computer program; the processor is used to implement the steps of the above-mentioned protein interaction prediction method when executing the computer program.

[0038] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned protein interaction prediction method are implemented.

[0039] The above technical solution of the present invention has the following advantages compared with the prior art:

[0040] The protein interaction prediction method described in the present invention embeds variable-length protein sequence feature information into a low-dimensional vector space by adjusting the Doc2vec unsupervised paragraph vector learning model, can process protein sequences of arbitrary length, solves the problem of preliminary protein feature selection, utilizes the advantages of graph isomorphic convolutional networks, fully combines the information of the protein interaction PPI network structure, aggregates the information of adjacent proteins of each protein, optimizes the encoding representation of protein nodes, finds protein interaction edges based on the PPI network structure information, combines the information of two protein nodes, and continuously learns more efficient and accurate classification predictions therefrom, thereby improving the accuracy of model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0042] Figure 1 is the implementation flowchart of the protein interaction prediction method of the present invention;

[0043] Figure 2 is the schematic flowchart of the central k-mer prediction;

[0044] Figure 3 is the framework diagram of a protein interaction prediction method that fuses Doc2vec and graph convolution;

[0045] Figure 4 is the micro-F1 score of each method on the public dataset SHS27k;

[0046] Figure 5 is the micro-F1 score of each method on the public dataset SHS148k;

[0047] Figure 6 is the structural block diagram of a protein interaction prediction device provided by an embodiment of the present invention. Detailed implementation manners

[0048] The core of the present invention is to provide a protein interaction prediction method, device, equipment and computer storage medium, which improves the accuracy of protein interaction prediction.

[0049] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] Please refer to Figure 1 , Figure 1 is the implementation flowchart of the protein interaction prediction method provided by the present invention; the specific operation steps are as follows:

[0051] S101: Obtain the amino acid sequence s of the i-th protein p in the protein interaction network, where i = 1...n, and n is the total number of proteins in the protein interaction network;

[0052] S102: Starting from the j-th amino acid of the amino acid sequence, successively intercept multiple k-mers backward until the remaining number of amino acids < k, to obtain the j-th amino acid subsequence s j , where j = 1...k;

[0053] S103: using the Doc2vec model to predict the central k-mer in each sliding window in the j-th amino acid subsequence, and performing aggregation operations to obtain the j-th amino acid subsequence feature;

[0054] like Figure 2 As shown, the embedded codes of the j-th amino acid subsequence and its context k-mer in the t-th sliding window are averaged, and classification prediction is performed based on the average value to obtain the central k-mer in the t-th sliding window in the j-th amino acid subsequence.

[0055] The amino acid sequence of the protein is regarded as a document, and the Doc2vec model is used to embed the amino acid sequence. Only the amino acid sequence is needed to encode the protein and sequence information of any length can be obtained. Each protein is represented as a low-dimensional vector and used as the preliminary feature of the multi-label classification task.

[0056] S104: performing a mean operation on the k amino acid subsequence features to obtain the i-th protein feature;

[0057] S105: Aggregate the features of multiple neighboring proteins of the ith protein to itself to obtain an updated feature of the ith protein, repeat this step multiple times to obtain an aggregated feature of the ith protein;

[0058] The information-passing graph isomorphic convolutional network (GINConv) is used. GINConv formalizes the convolution process into two functions: information passing and node information updating. Each node aggregates the information of its neighbors to its own node. The node information updating is to combine the node representation of the previous layer of the node with the aggregated neighbor information.

[0059] The formula is:

[0060]

[0061] in, is the i-th protein feature after the last update, is the i-th protein feature after the current update, is the feature of the u-th protein after the last update, N(i) represents the set of corresponding serial numbers of the neighbor proteins of the i-th protein, f() represents feature aggregation, and φ() represents the mapping function.

[0062] The feature aggregation adopts the accumulation method, and the mapping function adopts the multi-layer perceptron:

[0063]

[0064] Among them, ε is a hyperparameter or a learnable parameter, and MLP() represents a multi-layer perceptron.

[0065] S106: Perform a dot product between the ith protein aggregation feature and the protein aggregation feature that interacts with it and input the result into the classifier to obtain a prediction result of the ith protein interaction.

[0066] The classifier is a multi-layer perceptron, and the formula is expressed as:

[0067]

[0068] Among them, h i is the i-th protein aggregation feature, h i+1 is the protein aggregation feature that interacts with the i-th protein aggregation feature.

[0069] The protein interaction prediction method described in the present invention embeds variable-length protein sequence feature information into a low-dimensional vector space by adjusting the Doc2vec unsupervised paragraph vector learning model, can process protein sequences of arbitrary length, solves the problem of preliminary protein feature selection, utilizes the advantages of graph isomorphic convolutional networks, fully combines the information of the protein interaction PPI network structure, aggregates the information of adjacent proteins of each protein, optimizes the encoding representation of protein nodes, finds protein interaction edges based on the PPI network structure information, combines the information of two protein nodes, and continuously learns more efficient and accurate classification predictions therefrom, thereby improving the accuracy of model prediction.

[0070] like Figure 3 Based on the above embodiments, this embodiment further includes after step S104:

[0071] The i-th protein feature is input into the stacked one-dimensional convolutional network for processing, where the processing process of each layer of the one-dimensional convolutional network is:

[0072] The output of the previous one-dimensional convolutional network is used as the input of the current one-dimensional convolutional network. After multiple convolution operations on the input, the Relu activation function is used to prevent the gradient from disappearing and the maximum value pooling is used to extract the main features to obtain the output of the current one-dimensional convolutional network.

[0073] The i-th protein feature processed by the stacked one-dimensional convolutional network is processed by a fully connected layer.

[0074] The above operations can comprehensively observe protein sequence information and extract effective features for PPI multi-type prediction tasks, thereby improving the classification efficiency of the model.

[0075] The present invention provides a protein interaction prediction method that integrates Doc2vec and graph convolution, which is used to efficiently predict PPIs whose functions have not yet been discovered, so as to reduce the cost of biological experiments. The present invention plays a guiding role in the fields of cell biology and medicine, such as target therapy and new drug design. Its characteristics are: (1) using an unsupervised model of bag-of-words prediction tasks in the field of natural language processing to train the amino acid sequence of proteins, and using the output of the model as the preliminary features of protein sequence information; (2) using a one-dimensional convolutional neural network to extract effective features for the task; (3) using a graph neural network as a downstream model to characterize a single protein while aggregating the information of its neighboring proteins; (4) using a multi-layer perceptron to predict the final features. This method only uses amino acid sequence information and PPI network information, and while effectively processing sequence information of any length, it also simplifies the model depth, thereby efficiently and accurately predicting interactions between proteins, especially for multi-type predictions of PPIs between "new proteins" that have never appeared.

[0076] Based on the above embodiments, in order to evaluate the effectiveness of the performance of the present method, the present embodiment uses random search, breadth-first search (Bfs) and depth-first search (Dfs) strategies to partition the data set. As shown in the following figure, when the data set is partitioned using the three strategies respectively, under the condition of selecting the same number of PPIs, the test set protein nodes under the Bfs and Dfs partitioning strategies are often far less than the Random partitioning strategy, that is, when the data set is partitioned using Bfs and Dfs, a large number of "new proteins" that have not appeared in the training set can appear, and these new proteins can better test the prediction efficiency of the model.

[0077] Therefore, this method GDP adopts three division methods, Random, Bfs and Dfs, and conducts comparative experiments with the current six protein classification methods. Among them, RF and LR use random forest and logistic regression methods respectively, DPPI, DNN-PPI and PIPR use convolutional network methods, and GNN-PPI uses graph neural network method.

[0078] This method was validated on two real public datasets of different sizes. It divided PPIs into seven types, namely reaction, binding, post-translational modifications, activation, inhibition, catalysis and expression. Any pair of PPIs contains at least one of them.

[0079] like Figure 4 and Figure 5As shown, in terms of the micro-F1 score metric and three dataset partitioning modes, the performance of this method on the SHS27k dataset and the SHS148k dataset is better than that of other methods. In the SHS27k dataset, the micro-F1 score metric of the GDP method under the three partitioning methods of Random, Bfs, and Dfs has increased by 0.81, 8.87, and 2.6 respectively compared to the currently best-performing GNN-PPI method; in the SHS148k dataset, the micro-F1 metrics of the GDP method under the three partitioning methods of Random, Bfs, and Dfs have increased by 0.3, 10.84, and 1.5 respectively. It can be seen that the accuracy of the multi-type PPI prediction results of the GDP method has been greatly improved, with higher biological significance.

[0080] Please refer to Figure 6 , Figure 6 is the structural block diagram of a protein-protein interaction prediction device provided by an embodiment of the present invention; the specific device may include:

[0081] An amino acid sequence acquisition module 100, configured to acquire the amino acid sequence of the i-th protein in the protein-protein interaction network, where i = 1...n, and n is the total number of proteins in the protein-protein interaction network;

[0082] An amino acid subsequence acquisition module 200, configured to start from the j-th amino acid of the amino acid sequence, and sequentially intercept multiple k-mers backward until the remaining number of amino acids < k, to obtain the j-th amino acid subsequence, where j = 1...k;

[0083] An amino acid subsequence feature acquisition module 300, configured to use the Doc2vec model to predict the central k-mer in each sliding window of the j-th amino acid subsequence, and perform an aggregation operation to obtain the j-th amino acid subsequence feature;

[0084] A protein feature acquisition module 400, configured to perform a mean operation on the k amino acid subsequence features to obtain the i-th protein feature;

[0085] A protein aggregation feature acquisition module 500, configured to aggregate the features of multiple neighbor proteins of the i-th protein to itself to obtain the updated i-th protein feature, and repeat this step multiple times to obtain the i-th protein aggregation feature;

[0086] A protein-protein interaction prediction module 600, configured to input the dot product of the i-th protein aggregation feature and the protein aggregation feature that interacts with it into a classifier to obtain the prediction result of the i-th protein-protein interaction.

[0087] The protein interaction prediction device of the present embodiment is used to implement the aforementioned protein interaction prediction method, so the specific implementation methods of the protein interaction prediction device can be seen in the embodiment part of the aforementioned protein interaction prediction method, for example, the amino acid sequence acquisition module 100, the amino acid subsequence acquisition module 200, the amino acid subsequence feature acquisition module 300, the protein feature acquisition module 400, the protein aggregation feature acquisition module 500, the protein interaction prediction module 600, are respectively used to implement steps S101, S102, S103, S104, S105, S106 in the aforementioned protein interaction prediction method, so its specific implementation methods can refer to the description of the corresponding embodiments of each part, which will not be repeated here.

[0088] A specific embodiment of the present invention further provides a protein interaction prediction device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned protein interaction prediction method when executing the computer program.

[0089] A specific embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned protein interaction prediction method are implemented.

[0090] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0091] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0092] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0094] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.

Claims

1. A protein interaction prediction method, characterized in that: Including: Obtain the amino acid sequence of the \(i\)-th protein in the protein interaction network, where \(i = 1,\cdots,n\) and \(n\) is the total number of proteins in the protein interaction network; Starting from the \(j\)-th amino acid of the amino acid sequence, successively intercept multiple \(k\)-mers backward until the remaining number of amino acids is less than \(k\), to obtain the \(j\)-th amino acid subsequence, where \(j = 1,\cdots,k\); Use the Doc2vec model to predict the central \(k\)-mer in each sliding window of the \(j\)-th amino acid subsequence and perform an aggregation operation to obtain the \(j\)-th amino acid subsequence feature, including: Calculate the average value of the embedding encodings of the \(j\)-th amino acid subsequence and its context \(k\)-mer in the \(t\)-th sliding window, and perform classification prediction based on the average value to obtain the central \(k\)-mer in the \(t\)-th sliding window of the \(j\)-th amino acid subsequence; Perform a mean operation on the \(k\) amino acid subsequence features to obtain the \(i\)-th protein feature, and then input the \(i\)-th protein feature into a stacked one-dimensional convolutional network for processing. The processing process of each layer of the one-dimensional convolutional network is: use the output of the previous layer of the one-dimensional convolutional network as the input of the current layer of the one-dimensional convolutional network, perform multiple convolutional operations on the input, and then process it through the Relu activation function and the max pooling layer to obtain the output of the current one-dimensional convolutional network; input the \(i\)-th protein feature processed by the stacked one-dimensional convolutional network through a fully connected layer; Use the graph isomorphism convolutional network to aggregate the features of multiple neighbor proteins of the \(i\)-th protein to itself to obtain the updated \(i\)-th protein feature, and repeat this step multiple times to obtain the \(i\)-th protein aggregation feature; Perform a dot product on the \(i\)-th protein aggregation feature and the protein aggregation feature that interacts with it, and then input it into a classifier to obtain the prediction result of the interaction of the \(i\)-th protein.

2. The protein interaction prediction method according to claim 1, characterized in that: The formula for aggregating the features of multiple neighbor proteins of the \(i\)-th protein to itself to obtain the updated \(i\)-th protein feature and repeating this step multiple times to obtain the \(i\)-th protein aggregation feature is: ; in, is the i-th protein feature after the last update, is the i-th protein feature after the current update, is the feature of the u-th protein after the last update, N(i) represents the set of corresponding serial numbers of the neighbor proteins of the i-th protein, f() represents feature aggregation, Represents a mapping function.

3. The protein interaction prediction method according to claim 2, characterized in that: The feature aggregation is performed in an accumulative manner, and the mapping function uses a multi-layer perceptron: Among them, ε is a hyperparameter, and MLP() represents a multi-layer perceptron.

4. The protein interaction prediction method according to claim 1, characterized in that: The classifier is a multi-layer perceptron.

5. The protein interaction prediction method according to claim 4, characterized in that: The dot product of the i-th protein aggregation feature and the protein aggregation feature that interacts with it is input into the classifier, and the formula for obtaining the prediction result of the i-th protein interaction is expressed as: ; in, is the i-th protein aggregation feature, is the protein aggregation feature that interacts with the i-th protein aggregation feature.

6. A protein interaction prediction device, characterized in that: Including: An amino acid sequence acquisition module for obtaining the amino acid sequence of the \(i\)-th protein in the protein interaction network, where \(i = 1,\cdots,n\) and \(n\) is the total number of proteins in the protein interaction network; An amino acid subsequence acquisition module for starting from the \(j\)-th amino acid of the amino acid sequence and successively intercepting multiple \(k\)-mers backward until the remaining number of amino acids is less than \(k\) to obtain the \(j\)-th amino acid subsequence, where \(j = 1,\cdots,k\); An amino acid subsequence feature acquisition module for using the Doc2vec model to predict the central \(k\)-mer in each sliding window of the \(j\)-th amino acid subsequence and performing an aggregation operation to obtain the \(j\)-th amino acid subsequence feature, including calculating the average value of the embedding encodings of the \(j\)-th amino acid subsequence and its context \(k\)-mer in the \(t\)-th sliding window and performing classification prediction based on the average value to obtain the central \(k\)-mer in the \(t\)-th sliding window of the \(j\)-th amino acid subsequence; The protein feature acquisition module is used to perform mean operation on the features of k amino acid subsequences to obtain the i-th protein feature, and then input the i-th protein feature into the stacked one-dimensional convolutional network for processing. The processing process of each layer of the one-dimensional convolutional network is as follows: the output of the previous layer of the one-dimensional convolutional network is used as the input of the current layer of the one-dimensional convolutional network, and the input is subjected to multiple convolution operations, and then processed by the Relu activation function and the maximum pooling layer to obtain the output of the current one-dimensional convolutional network; the i-th protein feature processed by the stacked one-dimensional convolutional network is processed by the fully connected layer; The protein aggregation feature acquisition module is used to use a graph isomorphic convolutional network to aggregate the features of multiple neighboring proteins of the ith protein into itself to obtain the updated ith protein feature, and repeat this step multiple times to obtain the ith protein aggregation feature; The protein interaction prediction module is used to perform dot product between the aggregation feature of the ith protein and the aggregation feature of the protein that interacts with it and then input the result into the classifier to obtain the prediction result of the ith protein interaction.

7. A protein interaction prediction device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of a protein interaction prediction method as claimed in any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the protein interaction prediction method according to any one of claims 1 to 5 are implemented.