A method and device for predicting protein-protein interaction sites

Through the combination of multi-feature fusion and capsule network, the problem of insufficient prediction accuracy caused by the neglect of feature correlation in the prior art is solved, and higher prediction accuracy and generalization ability of protein-protein interaction sites are achieved.

CN119943144BActive Publication Date: 2025-07-04SUZHOU CITY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413894.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The prior art ignores the correlation between features when predicting protein-protein interaction sites, resulting in insufficient prediction accuracy, especially in the prediction tasks of large-scale unknown proteins.

Method used

Using a multi-feature fusion method, biometric and semantic features are extracted through deep learning models, feature stitching and dynamic routing are combined with capsule networks, and the potential correlation between features is captured using Transformer, convolutional neural network and bidirectional long and short-term memory network, and model parameters are optimized through Focal Loss function.

Benefits of technology

It improves the prediction accuracy of protein-protein interaction sites, enhances the generalization ability of the model, can better deal with the problems of data sparseness and uneven label distribution, and improves the recognition accuracy of different categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943144B_ABST
    Figure CN119943144B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of proteins, and discloses a method and device for predicting protein-protein interaction sites, including: horizontally splicing the output feature matrices of a biological feature matrix of a protein sequence to be predicted through a variety of different deep learning models to obtain an integrated feature matrix; obtaining an attention feature matrix based on the semantic feature matrix of the protein sequence to be predicted, splicing it with the integrated feature matrix and passing through a linear layer to obtain a spliced matrix; inputting the spliced matrix into a capsule network layer to obtain a classification vector of amino acid residues at each site in the protein sequence to be predicted, and passing through a linear layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to an interaction site. The present invention fully utilizes the feature capture capabilities of different models through multi-feature fusion, and feature splicing better captures the potential correlation between features, and combines with a capsule network to improve the accuracy of predicting protein-protein interaction sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of proteins, and in particular to a method and device for predicting protein-protein interaction sites. Background Art

[0002] Protein interaction sites refer to the specific regions on protein molecules that physically or chemically interact with other molecules. According to different objects, they can be divided into protein-protein interaction sites (Protein-Protein Interaction sites, PPIS), protein-peptide interaction sites (Protein-Peptide Interaction Sites, PPepIS), protein-small molecule ligand interaction sites (Protein-Ligand Interaction Sites, PLIS), etc. The detailed study of protein interaction sites is of crucial significance for deepening the understanding of protein functions, revealing the mechanisms of biological processes, drug research and development, and analyzing disease mechanisms.

[0003] Protein-protein interaction plays an important role in cell physiological processes, and the identification of its interaction sites is crucial for understanding protein function mechanisms and drug research and development. Protein-protein interaction usually occurs on the protein surface, that is, in the region with a relatively large accessible surface area. The experimental method for determining the interaction sites of protein complexes has the characteristic of high precision and can accurately identify the interaction sites in the complex. However, the experimental method has a long cycle and high cost, and researchers have also started to use computational methods to predict the interaction sites in protein complexes in a high-throughput manner. The computational method for predicting protein interaction sites can be described as solving the following problem: given a protein amino acid sequence P with a length of L, find the best mapping function F(P) to map it to A. The length of A is also L, and each item is 0 or 1, where 0 represents that the site is not an interaction site, and 1 represents that the site is an interaction site.

[0004] Currently, the commonly used computational methods for protein-protein interaction sites are mainly divided into two categories: structure-based methods and sequence-based methods.

[0005] The structure-based methods mainly include: MaSIF, dMaSIF, MPNP, GraphPPIS and EGRET. Among them, MaSIF predicts interaction sites by capturing the key features of biomolecular surface interactions, dMaSIF takes the initial three-dimensional coordinates and chemical types of atoms as input to identify interaction sites in an end-to-end manner, and MPNP uses the relationship structure within the model to predict interaction sites. Yuan et al. proposed the GraphPPIS model, which uses evolutionary features and residue structure features to represent nodes in the graph, and uses the distance between residues to represent edges in the graph. Mahbub et al. constructed the EGRET model based on the graph self-attention network, which uses the protein sequence abstract features of the pre-trained model as the features of the nodes in the graph, and the distance and relative direction between two residues as the features of the edges in the graph. Sequence-based methods have received more attention.

[0006] Most of these existing technologies tend to introduce protein structure information into the prediction of protein interaction sites. Although this can improve the prediction accuracy of the model, its generalization is limited and it is not suitable for the task of predicting interaction sites of large-scale unknown proteins. Therefore, when facing unknown proteins, sequence-based prediction methods are needed.

[0007] The methods based entirely on sequences mainly include: SCRIBER, DLPred and DELPHI. Among them, SCRIBER uses a data set covering various types of binding residues to predict interacting residues through a two-layer architecture. The DLPred model proposed by Zhang et al. uses a simplified bidirectional recurrent neural network. Without using protein structure information, it combines the protein relative solvent accessibility prediction task and adopts a multi-task joint learning strategy to predict the interaction sites of proteins, achieving better results. The DELPHI model proposed by Li uses the idea of ​​ensemble learning to obtain different aspects of feature information in the same protein sequence by integrating bidirectional recurrent neural networks and convolutional neural networks, achieving good results.

[0008] In practical applications, some features often do not exist in isolation, but are connected and influence each other in complex ways; however, most of the above studies use recurrent neural networks or convolutional neural networks to mine local features or contextual features of protein sequences, ignoring the importance of potential correlations between features and their state changes, lacking effective modeling of the correlations between features, and each feature is independently applied to the prediction model without fully considering the potential interactions and dependencies between features. This leads to limitations in the model in capturing the complex interactions and differences between features, which may restrict the improvement of prediction performance.

[0009] In summary, the existing methods for prediction based on protein structure information have poor generalization of the prediction model and are not applicable to the prediction of large-scale location proteins; the existing sequence-based prediction methods ignore the correlation between various features in the protein sequence and do not fully consider the interaction and dependence relationships between features, which affects the prediction accuracy. Summary of the Invention

[0010] To this end, the technical problem to be solved by the present invention is to overcome the problem that the existing technology ignores the correlation between features and leads to inaccurate prediction.

[0011] To solve the above technical problem, the present invention provides a method for predicting protein-protein interaction sites, including:

[0012] Obtain the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted;

[0013] After the biological feature matrix is respectively input into a variety of different deep learning models, the output feature matrices of each deep learning model are horizontally spliced to obtain an integrated feature matrix; the semantic feature matrix is input into a bidirectional long short-term memory network to output an attention feature matrix; after the integrated feature matrix and the attention feature matrix are horizontally spliced and passed through a linear layer, the splicing matrix of the protein sequence to be predicted is obtained;

[0014] Input the splicing matrix of the protein sequence to be predicted into a capsule network layer to obtain the classification vector of amino acid residues at each site in the protein sequence to be predicted, including:

[0015] Input the splicing matrix into multiple primary capsule units for convolution operations, vertically stack the capsule matrices output by each primary capsule unit to obtain the capsule tensor of the protein sequence to be predicted;

[0016] Input the capsule tensor into multiple classification capsules respectively; in each classification capsule, generate a real value for each amino acid residue in the protein sequence to be predicted;

[0017] Splice the real values selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue;

[0018] Input the classification vector of each amino acid residue in the protein sequence to be predicted into a linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to an interaction site.

[0019] Preferably, a variety of different deep learning models include: Transformer, convolutional neural network and bidirectional long short-term memory network.

[0020] Preferably, after the biometric matrix is input into multiple different deep learning models respectively, the output feature matrices of each deep learning model are horizontally concatenated to obtain an integrated feature matrix, including:

[0021] Input the biometric matrix into a Transformer, and through the multi-head attention mechanism, residual connection, normalization layer, feed-forward network, residual connection and normalization layer, capture the long-range dependence relationship between biometric features, and output a multi-head attention feature matrix;

[0022] Input the biometric matrix into a convolutional neural network, and after splicing through convolutional units with multiple different convolutional kernel sizes, output a local feature matrix;

[0023] Input the biometric matrix into a bidirectional long short-term memory network to obtain a time-dependent feature matrix;

[0024] Horizontally concatenate the multi-head attention feature matrix, the local feature matrix and the time-dependent feature matrix to obtain an integrated feature matrix.

[0025] Preferably, input the concatenated matrix into multiple primary capsule units for convolution operations to obtain the capsule matrix output by each primary capsule unit, including:

[0026] The primary capsule unit is a two-dimensional convolutional neural network, the hyperparameters of multiple two-dimensional convolutional neural networks are the same, the convolutional kernel size is 1×9, the stride is 1, the number of input channels and the number of output channels are both 1, and the padding operation is in the same mode;

[0027] In each primary capsule unit, after the concatenated matrix undergoes convolution operations and padding operations, a capsule matrix with the same size as the concatenated matrix is obtained.

[0028] Preferably, in each classification capsule, use the dynamic routing algorithm to generate a real value for each amino acid residue in the protein sequence to be predicted, including:

[0029] Obtain the initial weight coefficient of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted;

[0030] Normalize the initial weight coefficient using the Softmax function to obtain the normalized weight coefficient of each capsule matrix;

[0031] Use the normalized weight coefficient to perform weighted summation on the capsule matrices corresponding to all amino acid residues to obtain a weighted capsule tensor;

[0032] Activate the weighted capsule tensor using the Squashing function to obtain an activated capsule tensor;

[0033] After multiple iterations, update the weight coefficients and obtain the target capsule tensor corresponding to the convergence of the weight coefficients;

[0034] Obtain the target capsule vectors in the capsule matrix corresponding to each amino acid residue in the target capsule tensor;

[0035] Calculate the norm of the target capsule vector of each amino acid residue as the real value of each amino acid residue.

[0036] Preferably, the acquisition of the biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted includes:

[0037] The acquisition of the biological features of the amino acid residues at each site in the protein sequence to be predicted includes:

[0038] Based on the AAindex database, obtain the physical properties and physicochemical characteristics of the amino acid residues at each site in the protein sequence to be predicted;

[0039] Use Psi-Blast to calculate the position-specific matrix of the protein sequence to be predicted;

[0040] Based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of each type of amino acid, calculate the solvent accessible surface area of each amino acid residue in the protein sequence to be predicted;

[0041] Based on the length of the protein sequence to be predicted, obtain the relative position information of each amino acid residue in the protein sequence to be predicted;

[0042] Obtain the number of amino acid types and the number of amino acid categories within a window of a preset size at each site in the protein sequence to be predicted;

[0043] The acquisition of the semantic features of the amino acid residues at each site in the protein sequence to be predicted includes: calculating the semantic features of the amino acid residues at each site in the protein sequence to be predicted using ProtBERT.

[0044] Preferably, based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of each type of amino acid, calculate the solvent accessible surface area of each amino acid residue in the protein sequence to be predicted, expressed as:

[0045] ;

[0046] Wherein, represents the solvent accessible surface area when the th amino acid residue in the protein sequence to be predicted is the jth category of amino acid; , represents the length of the protein sequence to be predicted; , represents the th type among the amino acid types, with a maximum of 20; represents the relative solvent accessibility of the th amino acid residue in the protein sequence to be predicted, represents the maximum accessible surface area of the

[0047] Preferably, use Focal Loss to train multiple different deep learning models, bidirectional long short-term memory networks, and capsule network layer models, including:

[0048] Obtain multiple protein sequence samples and the true labels of whether the amino acid residues at each site are protein-protein interaction sites;

[0049] Input the protein sequence samples, pass them through multiple different deep learning models and bidirectional long short-term memory networks, and obtain the corresponding concatenated matrix;

[0050] Input the concatenated matrix into the capsule network layer, obtain the classification vector of each amino acid residue in the protein sequence, pass it through the output layer, obtain the probability that each amino acid residue belongs to an interaction site, and generate prediction labels;

[0051] Based on the prediction labels and true labels of the amino acid residues at each site in each protein sequence sample, calculate the probability of being predicted as the positive class;

[0052] Based on the probability of being predicted as the positive class, use the Focal Loss function to calculate the model prediction loss, and update the model parameters of multiple different deep learning models, bidirectional long short-term memory networks, and capsule network layers until the Focal Loss function converges, and obtain the trained multiple different deep learning models, bidirectional long short-term memory networks, and capsule network layers.

[0053] Preferably, the Focal Loss function is expressed as:

[0054] ;

[0055] where represents the Focal Loss function, represents the weight factor with a value of ; represents the probability of being predicted as the positive class, represents the focusing parameter, represents the sample adjustment factor.

[0056] This embodiment provides a protein-protein interaction site prediction device, including:

[0057] A feature extraction module, configured to obtain the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted;

[0058] An ensemble learning module, configured to respectively input the biological feature matrix into multiple different deep learning models, then horizontally splice the output feature matrices of each deep learning model to obtain an ensemble feature matrix; input the semantic feature matrix into a bidirectional long short-term memory network to output an attention feature matrix; horizontally splice the ensemble feature matrix and the attention feature matrix and then pass through a linear layer to obtain the splicing matrix of the protein sequence to be predicted;

[0059] A capsule network module, configured to input the splicing matrix of the protein sequence to be predicted into a capsule network layer to obtain the classification vectors of amino acid residues at each site in the protein sequence to be predicted, including: respectively inputting the splicing matrix into multiple primary capsule units for convolution operations, obtaining the capsule matrices output by each primary capsule unit and stacking them vertically to obtain the capsule tensor of the protein sequence to be predicted; respectively inputting the capsule tensor into multiple classification capsules; in each classification capsule, generating a real value for each amino acid residue in the protein sequence to be predicted; splicing the real values selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue;

[0060] A prediction module, configured to send the classification vectors of each amino acid residue in the protein sequence to be predicted into a linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to an interaction site.

[0061] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0062] The protein-protein interaction site prediction method described in the present invention analyzes and obtains the biological feature matrix and semantic feature matrix of the protein sequence to be predicted, and uses different deep learning models for feature extraction and splicing respectively. This multi-feature fusion method can make full use of the ability of different models to capture different features in the protein sequence, thereby improving the prediction accuracy. At the same time, by further splicing the integrated features and attention features, the potential correlation between features can be better captured, avoiding the problem of isolated feature processing. At the same time, combined with the capsule network, multiple parallel primary capsule units are used to further extract and integrate the feature correlation of the splicing matrix from different perspectives, capturing more diverse features, avoiding feature omission caused by a single perspective, and making the extracted features more complete and representative. The classification capsule uses its internal dynamic routing mechanism to more accurately match and distinguish the capsule tensor with different categories, adaptively adjust the routing weights according to the correlation and similarity of the features in the capsule tensor, and accurately allocate the feature information to the corresponding category capsules, thereby improving the recognition accuracy of different categories and enhancing the accuracy of protein-protein interaction site prediction.

[0063] At the same time, through the combination of ensemble learning and capsule network in the present invention, the problems of data sparsity and unbalanced label distribution can be better handled; the dynamic routing mechanism of the capsule network can also adaptively adjust the feature weights, thereby alleviating the impact caused by the imbalance between positive and negative samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention in conjunction with the drawings, where:

[0065] Figure 1 is the step flowchart of the protein-protein interaction site prediction method provided by the present invention;

[0066] Figure 2 is the model network structure diagram of the protein-protein interaction site prediction method provided by the present invention;

[0067] Figure 3 is the schematic diagram of the working process of the dynamic routing algorithm in the classification capsule;

[0068] Figure 4 is the schematic diagram of the functional modules of the protein-protein interaction site prediction method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The following further illustrates the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.

[0070] Referring to Figure 1 as shown, the flowchart of the steps of the method for predicting protein-protein interaction sites provided by the present invention, the specific steps include:

[0071] S101: Obtain the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted;

[0072] S102: After inputting the biological feature matrix into a variety of different deep learning models respectively, horizontally splice the output feature matrices of each deep learning model to obtain an integrated feature matrix; input the semantic feature matrix into a bidirectional long short-term memory network to output an attention feature matrix; horizontally splice the integrated feature matrix and the attention feature matrix and then pass through a linear layer to obtain the splicing matrix of the protein sequence to be predicted;

[0073] S103: Input the splicing matrix of the protein sequence to be predicted into a capsule network layer to obtain the classification vector of amino acid residues at each site in the protein sequence to be predicted, including:

[0074] S103-1: Input the splicing matrix into multiple primary capsule units for convolution operations respectively, vertically stack the capsule matrices output by each primary capsule unit to obtain the capsule tensor of the protein sequence to be predicted;

[0075] S103-2: Input the capsule tensor into multiple classification capsules respectively; in each classification capsule, generate a real value for each amino acid residue in the protein sequence to be predicted;

[0076] S103-3: Splice the real values selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue;

[0077] S104: Input the classification vector of each amino acid residue in the protein sequence to be predicted into a linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to an interaction site.

[0078] This application directly performs prediction based on protein sequence information, avoiding the dependence on protein structure information, making the model more suitable for the task of predicting interaction sites of large-scale unknown proteins, and improving the generalization of the model.

[0079] Referring to Figure 2 as shown, it is the model network structure diagram of the method for predicting protein-protein interaction sites provided by the present invention; the model mainly includes an ensemble learning network layer and a capsule network layer.

[0080] Specifically, in the integrated learning network layer, there are multiple different deep learning models, including: Transformer, Convolutional Neural Network (CNN), and Bidirectional Long Short-Term Memory Network (BiLSTM). After the biometric matrix is input into multiple different deep learning models respectively, the output features of each deep learning model are horizontally concatenated to obtain integrated features, including:

[0081] Input the biometric matrix into the Transformer. Through the multi-head attention mechanism, residual connection, normalization layer, feed-forward network, residual connection, and normalization layer, capture the long-range dependence relationship between biometric features, and output the multi-head attention feature matrix;

[0082] Input the biometric matrix into the Convolutional Neural Network. After splicing through convolutional units with different convolutional kernel sizes in multiple paths, output the local feature matrix;

[0083] Input the biometric matrix into the Bidirectional Long Short-Term Memory Network to obtain the time-dependent feature matrix;

[0084] Horizontally concatenate the multi-head attention feature matrix, local feature matrix, and time-dependent feature matrix to obtain the integrated feature matrix.

[0085] The integrated learning network layer performs feature extraction and splicing using different deep learning models respectively; this multi-feature fusion method can make full use of the feature capture capabilities of different models for different features in the protein sequence, thereby improving the prediction accuracy. At the same time, by further concatenating the integrated feature matrix and the attention feature matrix, the potential correlation between features can be better captured, avoiding the problem of isolated feature processing.

[0086] Specifically, in the capsule network layer, it includes primary capsule units and classification capsules.

[0087] The primary capsule unit is a two-dimensional convolutional neural network. The hyperparameters of multiple two-dimensional convolutional neural networks are the same. The convolutional kernel size is 1×9, the stride is 1, the number of input channels and output channels are both 1, and the padding operation is in the same mode; in each primary capsule unit, after the splicing matrix undergoes convolutional operation and padding operation, a capsule matrix with the same size as the splicing matrix is obtained.

[0088] In each classification capsule, a real value is generated for each amino acid residue in the protein sequence to be predicted using a dynamic routing algorithm, including: obtaining the initial weight coefficients of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted; normalizing the initial weight coefficients using the Softmax function to obtain the normalized weight coefficients of each capsule matrix; performing weighted summation on the capsule matrices corresponding to all amino acid residues using the normalized weight coefficients to obtain a weighted capsule tensor; activating the weighted capsule tensor using the non-linear transformation Squashing function to obtain an activated capsule tensor; after multiple iterations, updating the weight coefficients to obtain the target capsule tensor corresponding to the convergence of the weight coefficients; obtaining the target capsule vectors in the capsule matrices corresponding to each amino acid residue in the target capsule tensor; calculating the norm of the target capsule vector of each amino acid residue as the real value of each amino acid residue.

[0089] The capsule network layer uses multiple parallel primary capsule units to further extract and integrate the feature correlations of the concatenated matrix from different perspectives, capturing more diverse features, avoiding feature omissions that may be caused by a single perspective, and making the extracted features more complete and representative; the classification capsule uses its internal dynamic routing mechanism to more accurately match and distinguish the capsule tensor with different categories, adaptively adjusting the routing weights according to the correlations and similarities of the features in the capsule tensor, and accurately distributing the feature information to the corresponding category capsules, thereby improving the recognition accuracy of different categories and enhancing the accuracy of protein-protein interaction site prediction.

[0090] The existing technologies do not handle data sparsity and the high imbalance of label distribution (the proportion of interaction sites is extremely small) sufficiently; the current technologies only alleviate the problem of positive and negative sample imbalance from a single aspect (data or model), which will indeed have a certain effect but cannot better alleviate the impact brought by label imbalance. The present invention can better handle the problems of data sparsity and label distribution imbalance through the combination of ensemble learning and capsule network; the dynamic routing mechanism of the capsule network can also adaptively adjust the feature weights, thereby alleviating the impact brought by positive and negative sample imbalance.

[0091] In the embodiments of the present invention, the acquisition of the biological and semantic features of the amino acid residues at each site in the protein sequence to be predicted includes:

[0092] The acquisition of the biological features of the amino acid residues at each site in the protein sequence to be predicted includes:

[0093] Based on the AAindex database, obtain the physical properties and physicochemical characteristics of amino acid residues at each site in the protein sequence to be predicted; the physical properties include hydrogen bond ability, polarity, solubility, and molecular weight; the physicochemical characteristics include physical properties, chemical properties, and kinetic properties;

[0094] Use Psi-Blast to calculate the position-specific matrix of the protein sequence to be predicted; the size of the position-specific matrix is N×20, where N represents the length of the protein sequence to be predicted, and a vector of length 20 is generated for each amino acid residue at each site in the protein sequence;

[0095] Based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of each type of amino acid, calculate the solvent accessible surface area of each amino acid residue in the protein sequence to be predicted, expressed as: ; where represents the solvent accessible surface area when the th amino acid residue in the protein sequence to be predicted is the jth type of amino acid; , represents the length of the protein sequence to be predicted; , represents the th type in the amino acid types, with a maximum of 20; represents the relative solvent accessibility of the th amino acid residue in the protein sequence to be predicted, represents the th type of amino acid's maximum accessible surface area.

[0096] Based on the length of the protein sequence to be predicted, obtain the relative position information of each amino acid residue in the protein sequence to be predicted;

[0097] Obtain the number of amino acid types and the number of amino acid types within a window of a preset size at each site in the protein sequence to be predicted; the amino acid types include polar, non-polar, positively charged, negatively charged, and uncharged;

[0098] Obtaining the semantic features of amino acid residues at each site in the protein sequence to be predicted includes:

[0099] Use ProtBERT to calculate the semantic features of amino acid residues at each site in the protein sequence to be predicted.

[0100] In the embodiments of the present invention, use Focal Loss to train a variety of different deep learning models, bidirectional long short-term memory networks, and capsule network layer models, including:

[0101] Obtain multiple protein sequence samples and the true labels indicating whether the amino acid residues at each site are protein-protein interaction sites;

[0102] Input the protein sequence samples, and through a variety of different deep learning models and bidirectional long short-term memory networks, obtain the corresponding splicing matrix;

[0103] Input the splicing matrix into the capsule network layer, obtain the classification vector of each amino acid residue in the protein sequence, and through the output layer, obtain the probability that each amino acid residue belongs to the interaction site, and generate prediction labels;

[0104] Based on the prediction labels and true labels of the amino acid residues at each site in each protein sequence sample, calculate the probability of being predicted as the positive class;

[0105] Based on the probability of being predicted as the positive class, use the Focal Loss function to calculate the model prediction loss, and update the model parameters of a variety of different deep learning models, bidirectional long short-term memory networks, and capsule network layers until the Focal Loss function converges, and obtain the trained a variety of different deep learning models, bidirectional long short-term memory networks, and capsule network layers.

[0106] The Focal Loss function is expressed as:

[0107] ;

[0108] Where, represents the Focal Loss function, represents the weight factor with a value of , which balances the importance of positive and negative samples by giving high weights to rare classes; represents the probability of being predicted as the positive class, represents the focusing parameter, represents the sample adjustment factor, which makes the model more focused on difficult-to-classify samples by reducing the loss sharing of simple samples.

[0109] Specifically, in the embodiments of the present invention, is set to 0.25, is set to 4.

[0110] Referring to Figure 2 shown, it is the model network structure diagram of the protein-protein interaction site prediction method provided by the present invention; based on this network structure, after inputting the protein sequence to be predicted, it includes:

[0111] ① Obtain the biological features and semantic features of each amino acid residue in the protein sequence to be predicted, and form an N×58-dimensional biological feature matrix and an N×1024-dimensional semantic feature matrix of the protein sequence to be predicted;

[0112] ② Respectively pass the biometric matrix through Transformer, CNN, and BiLSTM:

[0113] BiLSTM: The output vector is N×128 dimensions. The forward and backward vectors are added together. Therefore, the size of the output through BiLSTM is 128 dimensions;

[0114] CNN: Convolutional neural network. Convolution operations are performed using three convolutional kernels of different sizes. The sizes of the convolutional kernels are (3, 3), (5, 5), and (7, 7) respectively. The output vector sizes of these three convolutional operations are all N×29 dimensions, and they are concatenated horizontally. Therefore, the output vector size through the CNN module is N×87 dimensions;

[0115] Transformer: The Encoder part of Transformer is used. This part does not change the size of the vector. Therefore, the output vector size of this part is N×58 dimensions;

[0116] Finally, horizontally concatenate these three parts of vectors. Therefore, for a biometric matrix of N×58 dimensions, the size of the integrated feature matrix obtained after passing through this module is N×273 dimensions.

[0117] ③ Input the N×1024-dimensional semantic feature matrix into BiLSTM. The hidden layer of BiLSTM is 128, that is, the sizes of the forward output and the backward output are 128 dimensions. For a protein of length N, after calculating the semantic features through ProtBERT, a matrix of shape N×1024 is obtained. This matrix is used as the input and fed into BiLSTM, and the shape size of the output attention feature matrix is N×128;

[0118] ④ Concatenate the N×273-dimensional integrated feature matrix obtained from the ensemble learning network layer with the N×128-dimensional attention feature matrix. The shape size of the obtained matrix is N×401; Input the concatenated matrix into a linear layer. The input size of this linear layer is 401 dimensions, and the output size is 256 dimensions. Therefore, the shape size of the obtained matrix is N×256.

[0119] ⑤N × 256 is the input for each primary capsule. Each primary capsule is a two-dimensional convolutional neural network with a convolutional kernel size of (1, 9), and the input and output sizes of the convolutional neural network remain unchanged. Therefore, the size passing through the convolutional neural network is still N × 256. At this time, each amino acid is represented by a 256-dimensional vector, which can be understood as each amino acid consisting of 256 features. Combining eight primary capsules results in a matrix of shape N × 256 × 8, which can be understood as each feature being converted from a previous single value to an 8-dimensional vector. This 8-dimensional vector is generally called a capsule vector. Therefore, an amino acid residue can be represented by 256 capsule vectors; the output of the primary capsule layer is the capsule tensor of N × 256 × 8;

[0120] Each primary capsule in the primary capsule layer is a two-dimensional convolutional neural network. The convolutional neural network hyperparameters of the eight primary capsules are the same. The number of input channels (in_channels) and output channels (out_channels) are both 1, the size of the convolutional kernel is (1, 9), the stride is 1, and the padding mode is "same", which ensures that the matrix shapes of the input and output of each primary capsule remain consistent.

[0121] Each capsule tensor is the combined result of eight primary capsules, and these eight primary capsules can be regarded as eight independent convolutional neural networks. Each convolutional neural network processes the input N × 256 matrix, and through the padding operation of the convolutional neural network, the input and output shapes remain unchanged. Therefore, after being processed by eight primary capsules, eight matrices of different shapes with a size of N × 256 are obtained. Stacking them vertically gives a matrix of N × 256 × 8, where N represents the length of the protein sequence, that is, the number of amino acids. Each amino acid is also represented as a 256 × 8 matrix, which can be understood as converting the 256-dimensional feature vector of each amino acid in the original sequence to 256 × 8. That is to say, each primary capsule processes each value in the feature vector to obtain a new value. The processing results of the eight primary capsules are concatenated to obtain a 1 × 8 vector, converting the real values of the 256-dimensional feature vector to an 8-dimensional vector. Therefore, each capsule tensor is the combined result of 8 primary capsules.

[0122] ⑥The input of the classification capsule layer is N × 256 × 8. Each classification capsule executes the dynamic routing algorithm, which will select the most suitable capsule vector to represent the amino acid residue. Therefore, after the dynamic routing algorithm, the matrix shape size will become N × 1 × 8; then the internal sum of the obtained capsule vectors is calculated, and the matrix shape will change to N × 1; in this embodiment, 16 classification capsules are used, so the output of the classification capsule layer is N × 16;

[0123] The classification capsule selects a capsule vector for each amino acid residue in the sequence. The size of the capsule vector is (1, 8), and calculating the vector length gives a real value. Therefore, a classification capsule generates a matrix of (N, 1) for a protein sequence; the output results of the 16 classification capsules in this embodiment are horizontally concatenated to output a matrix of (N, 16).

[0124] Refer to Figure 3 shown in the figure, which is a schematic diagram of the working process of the dynamic routing algorithm in the classification capsule; among them, represents the number of input capsule vectors, represents the initial weight coefficients of the th round of update the normalized weight coefficients of the

[0125] ⑦ The last output layer is a linear layer. The input data dimension is 16, and the output data dimension is 1. Through the output layer, the matrix shape size is transformed into N×1, and the value corresponding to each amino acid residue represents the probability that this amino acid residue is an interaction site.

[0126] Refer to Figure 4 shown in the figure, which is a schematic diagram of the functional modules of the protein-protein interaction site prediction method provided by the present invention; the embodiment of the present invention mainly includes four modules, namely: a protein information input module, a protein sequence feature calculation module, a protein interaction site prediction module, and a prediction result output module.

[0127] ① The protein information input module obtains protein sequence data according to the protein sequence information input by the user;

[0128] Transmit the protein sequence data to the protein sequence feature calculation module, and at the same time store all the input data and the user information of the submitted data by the system.

[0129] ② The protein sequence feature calculation module receives protein sequence data and extracts protein sequence attribute features, including: calculating the physicochemical properties and physical features of amino acids at each site, calculating the position-specific matrix (PSSM) of the protein through the PAI-Blast software (PSSM reflects the frequency of amino acid substitution at each position by other amino acids), calling a prediction tool to calculate the accessible surface area ASA of amino acids at each site, calculating the length and relative position information of the amino acid residues in the sequence, counting the number of 20 amino acids and five types of amino acids (polar, non-polar, positively charged, negatively charged, uncharged) within a window of size 21 at each site, and calculating the potential semantic features of each amino acid on the protein sequence using ProtBERT;

[0130] Transmit the extracted protein sequence attribute features to the protein-protein interaction site prediction module;

[0131] ③ The protein-protein interaction site prediction module takes the protein sequence attribute features as input, calculates the probability that the amino acid residues at each site are protein-protein interaction sites, and obtains the prediction result;

[0132] ④ The prediction result output module stores the prediction result, displays it on the browser page, and also generates a pdf file to plot the prediction result as a line chart, and sends an email to the corresponding user who submitted the data according to the task.

[0133] The present invention only relies on the sequence information of proteins to extract features for model training, enabling the present invention to adapt to more biological application scenarios. For the problems of data sparsity and highly imbalanced label distribution, three strategies are adopted in this embodiment to alleviate the problems of data sparsity and label imbalance; First, an oversampling technique oriented to the training set is adopted, and by reusing and augmenting high-quality data instances, the recognition ability of the model for minority class samples (i.e., interaction sites) is effectively improved; Second, drawing on the idea of ensemble learning, multiple base models are used to mine the features concerned by minority class samples from multiple dimensions; Finally, a loss function mechanism Focal loss based on different sample weights is introduced. By assigning higher weight values to minority class samples and focusing on difficult-to-classify samples, the learning intensity of the model for minority class samples during the training iteration process is enhanced, thereby guiding the model to pay more attention to and accurately capture minority class samples. For exploring the correlation between features, a capsule network architecture is introduced at the model architecture level. Through the primary capsule layer, multiple capsule vectors are generated. The capsule vectors contain the correlation between features, and then the low-level capsule features are integrated through the classification capsule layer to generate higher-level feature representations for each amino acid residue.

[0134] Based on the above embodiments, an apparatus for predicting protein-protein interaction sites provided by an embodiment of the present invention; the specific apparatus may include:

[0135] A feature extraction module, configured to obtain the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted;

[0136] An ensemble learning module, configured to respectively input the biological feature matrix into multiple different deep learning models, then horizontally splice the output features of each deep learning model to obtain ensemble features; input the semantic features into a bidirectional long short-term memory network to output attention features; horizontally splice the ensemble features and the attention features and then pass through a linear layer to obtain the splicing matrix of the protein sequence to be predicted;

[0137] A capsule network module, configured to input the splicing matrix of the protein sequence to be predicted into a capsule network layer to obtain the classification vectors of amino acid residues at each site in the protein sequence to be predicted, including: respectively inputting the splicing matrix into multiple primary capsule units for convolution operations, vertically stacking the capsule matrices output by each primary capsule unit to obtain the capsule tensor of the protein sequence to be predicted; respectively inputting the capsule tensor into multiple classification capsules; in each classification capsule, generating a real value for each amino acid residue in the protein sequence to be predicted; splicing the real values selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue;

[0138] A prediction module, configured to send the classification vectors of each amino acid residue in the protein sequence to be predicted into a linear output layer to obtain the probabilities that each amino acid residue in the protein sequence to be predicted belongs to the interaction sites.

[0139] The protein-protein interaction site prediction device of this embodiment is used to implement the foregoing protein-protein interaction site prediction method. Therefore, the specific implementation manners in the protein-protein interaction site prediction device can be seen in the embodiment part of the protein-protein interaction site prediction method in the foregoing text. For example, the feature extraction module, the ensemble learning module, the capsule network module, and the prediction module are respectively used to implement steps S101, S102, S103, and S104 in the foregoing protein-protein interaction site prediction method. Therefore, the specific implementation manners can refer to the descriptions of the corresponding various part embodiments and will not be elaborated here.

[0140] Due to the inherent differences of different prediction tools on the training datasets, capPPISp needs to conduct horizontal comparative analysis with a variety of different prediction methods on a series of publicly available and recognized benchmark datasets. In the embodiments of the present invention, benchmark datasets Dset_448, Dset_355, Dset_164, Dset_186, and Dset_72 are respectively selected to compare the performance of the protein-protein interaction site prediction method capPPISp provided by the present invention with that of other technologies. The comparison results are shown in Tables 1, 2, 3, 4, and 5 respectively;

[0141] Table 1 Performance comparison of capPPISp and other prediction tools on Dset_448

[0142] Prediction tool Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.202 0.870 0.194 0.781 0.198 0.071 0.517 0.159 SPRINT 0.183 0.873 0.183 0.781 0.183 0.057 0.570 0.167 PSIVER 0.191 0.874 0.191 0.783 0.191 0.066 0.581 0.170 SPRINGS 0.229 0.882 0.228 0.796 0.229 0.111 0.625 0.201 LORIS 0.264 0.887 0.263 0.805 0.263 0.151 0.656 0.228 CRFPPI 0.268 0.887 0.264 0.805 0.266 0.154 0.681 0.238 SSWRF 0.288 0.891 0.286 0.811 0.287 0.178 0.687 0.256 SCRIBER 0.334 <![CDATA 0.896 > <![CDATA 0.332 > <![CDATA 0.821 > 0.333 0.230 0.715 0.287 DELPHI <![CDATA 0.371 > 0.901 0.371 0.829 0.371 0.272 0.737 0.337 capPPISp (ours) 0.675 0.659 0.237 0.661 <![CDATA 0.351 > <![CDATA 0.235 > <![CDATA 0.728 > <![CDATA 0.310 >

[0143] Based on Table 1, it can be seen that capPPISp is close to the optimal prediction tool in terms of the important balanced classification performance metrics AUPRC, MCC, and F1, ranking second. capPPISp is higher than other prediction tools in the Recall metric, 30.4% higher than the existing optimal tool, proving that capPPISp has better positive class sample recognition ability.

[0144] Table 2 Performance comparison of capPPISp and other prediction tools on Dset_355

[0145] Prediction tool Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.180 0.889 0.180 0.804 0.180 0.068 0.515 0.138 SPRINT 0.168 0.886 0.167 0.801 0.168 0.054 0.571 0.150 PSIVER 0.178 0.888 0.177 0.803 0.177 0.065 0.583 0.155 SPRINGS 0.211 0.892 0.210 0.811 0.211 0.103 0.608 0.178 LORIS 0.242 0.896 0.240 0.818 0.241 0.137 0.637 0.203 CRFPPI 0.247 0.897 0.245 0.819 0.246 0.143 0.662 0.214 SSWRF 0.268 0.901 0.268 0.825 0.268 0.168 0.667 0.228 DLPred 0.308 0.906 0.308 0.835 0.308 0.214 0.724 0.272 SCRIBER 0.322 <![CDATA 0.908 > <![CDATA 0.322 > <![CDATA 0.838 > 0.322 0.230 0.719 0.275 DELPHI <![CDATA 0.364 > 0.914 0.364 0.848 0.364 0.278 0.746 0.326 capPPISp (ours) 0.656 0.676 0.215 0.673 <![CDATA 0.324 > <![CDATA 0.224 > <![CDATA 0.729 > <![CDATA 0.282 >

[0146] Based on Table 2, it can be seen that on the Dset_355 benchmark dataset, capPPISp performs excellently in the key metrics of balanced classification performance - AUPRC, MCC, and F1. capPPISp is higher than other prediction tools in the recall rate (Recall) metric, proving that capPPISp has better positive class sample recognition ability and further verifying the superiority of capPPISp in identifying protein interaction sites.

[0147] Table 3 Performance comparison of capPPISp and other prediction tools on Dset_164

[0148] Prediction tool Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.264 0.828 0.253 0.726 0.258 0.090 0.528 0.220 PSIVER 0.217 0.826 0.216 0.716 0.216 0.043 0.554 0.205 CRFPPI 0.280 0.841 0.280 0.739 0.280 0.121 0.608 0.267 SSWRF 0.266 0.838 0.266 0.734 0.266 0.103 0.606 0.243 DLPred 0.338 <![CDATA 0.854 > <![CDATA 0.338 > 0.760 0.338 <![CDATA 0.192 > <![CDATA 0.672 > 0.330 SCRIBER 0.327 0.851 0.327 0.856 0.327 0.179 0.657 0.301 DELPHI <![CDATA 0.352 > 0.857 0.352 <![CDATA 0.765 > <![CDATA 0.352 > 0.209 0.685 <![CDATA 0.332 > capPPISp (ours) 0.515 0.812 0.322 0.639 0.354 0.130 0.621 0.363

[0149] Based on Table 3, it can be seen that on the Dset_164 benchmark dataset, capPPISp is superior to other existing prediction tools in the AUPRC and F1 balanced classification performance metrics, 3.1% and 0.2% higher respectively. In the Recall metric for measuring positive class sample recognition ability, capPPISp is also 16.3% higher than the existing optimal model.

[0150] Table 4 Performance Comparison of capPPISp and Other Prediction Tools on Dset_186

[0151] Prediction tool Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.194 0.848 0.186 0.748 0.190 0.041 0.499 0.165 DLPred 0.320 <![CDATA 0.878 > <![CDATA 0.320 > <![CDATA 0.783 > 0.320 <![CDATA 0.198 > <![CDATA 0.694 > 0.290 SCRIBER 0.279 0.870 0.279 0.780 0.279 0.150 0.647 0.246 DELPHI <![CDATA 0.351 > 0.884 0.351 0.803 0.351 0.235 0.710 0.319 capPPISp (ours) 0.583 0.859 0.256 0.603 <![CDATA 0.331 > 0.131 0.637 0.323

[0152] As can be seen from Table 4, on the Dset_186 benchmark test set, capPPISp is superior to other existing prediction tools in terms of AUPRC, with a 0.4% higher than the existing best prediction tool, and Recall is also 23.2% higher than the existing best model.

[0153] Table 5 Performance Comparison of capPPISp and Other Prediction Tools on Dset_72

[0154] Prediction tool Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.188 0.898 0.179 0.823 0.183 0.084 0.522 0.134 PSIVER 0.152 0.899 0.152 0.820 0.152 0.052 0.604 0.141 CRFPPI 0.248 <![CDATA 0.911 > <![CDATA 0.248 > <![CDATA 0.840 > 0.248 <![CDATA 0.158 > 0.669 0.200 SSWRF 0.246 <![CDATA 0.911 > 0.246 0.840 0.246 0.157 0.678 0.198 DLPred 0.246 0.901 0.246 0.826 0.246 0.148 <![CDATA 0.688 > 0.215 SCRIBER 0.232 0.909 0.232 0.837 0.232 0.141 0.680 0.198 DELPHI <![CDATA 0.274 > 0.914 0.274 0.847 <![CDATA 0.274 > 0.189 0.711 <![CDATA 0.237 > capPPISp (ours) 0.481 0.885 0.216 0.655 0.275 0.115 0.625 0.282

[0155] As can be seen from Table 5, on the Dset_72 benchmark test set, capPPISp is superior to other existing prediction tools in terms of AUPRC and F1 balanced classification performance metrics, with 4.5% and 0.1% higher respectively. In terms of Recall, capPPISp is also 20.7% higher than the existing best tool.

[0156] The method for predicting protein-protein interaction sites described in the present invention analyzes and obtains the biological feature matrix and semantic feature matrix of the protein sequence to be predicted, and performs feature extraction using different deep learning models and then splices them; this multi-feature fusion method can make full use of the capture capabilities of different models for different features in the protein sequence, thereby improving the prediction accuracy. At the same time, by further splicing the integrated feature matrix and the attention feature matrix, the potential correlation between features can be better captured, avoiding the problem of isolated feature processing. At the same time, combined with the capsule network, multiple parallel primary capsule units are used to further extract and integrate the feature correlation of the spliced matrix from different angles, capturing more diverse features, avoiding feature omission caused by a single perspective, and making the extracted features more complete and representative; using classification capsules, through its internal dynamic routing mechanism, the capsule tensor can be more accurately matched and distinguished with different categories, and according to the correlation and similarity of the features in the capsule tensor, the routing weights are adaptively adjusted, and the feature information is accurately assigned to the corresponding category capsules, thereby improving the recognition accuracy of different categories and enhancing the accuracy of predicting protein-protein interaction sites. At the same time, through the combination of ensemble learning and the capsule network, the present invention can better handle the problems of data sparsity and unbalanced label distribution; the dynamic routing mechanism of the capsule network can also adaptively adjust the feature weights, thereby alleviating the impact caused by the imbalance between positive and negative samples.

[0157] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0158] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0161] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for predicting protein-protein interaction sites, characterized in that, Including: Obtain the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted; Input the biological feature matrix into multiple different deep learning models respectively, including: input the biological feature matrix into Transformer, through the multi-head attention mechanism, residual connection, normalization layer, feed-forward network, residual connection and normalization layer, capture the long-range dependence relationship between biological features, and output the multi-head attention feature matrix; input the biological feature matrix into the convolutional neural network, and splice after passing through convolutional units with multiple different convolutional kernel sizes to output the local feature matrix; input the biological feature matrix into the bidirectional long short-term memory network to obtain the time-dependent feature matrix; Horizontally splice the multi-head attention feature matrix, the local feature matrix and the time-dependent feature matrix to obtain the integrated feature matrix; Input the semantic feature matrix into the bidirectional long short-term memory network to output the attention feature matrix; horizontally splice the integrated feature matrix and the attention feature matrix and then pass through the linear layer to obtain the splicing matrix of the protein sequence to be predicted; Input the splicing matrix of the protein sequence to be predicted into the capsule network layer to obtain the classification vector of amino acid residues at each site in the protein sequence to be predicted, including: Input the splicing matrix into multiple primary capsule units respectively for convolution operation, stack the capsule matrices output by each primary capsule unit vertically to obtain the capsule tensor of the protein sequence to be predicted; Input the capsule tensor into multiple classification capsules respectively; in each classification capsule, generate a real value for each amino acid residue in the protein sequence to be predicted, including: obtain the initial weight coefficient of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted; use the Softmax function to normalize the initial weight coefficient to obtain the normalized weight coefficient of each capsule matrix; use the normalized weight coefficient to perform weighted summation on the capsule matrices corresponding to all amino acid residues to obtain the weighted capsule tensor; use the Squashing function to activate the weighted capsule tensor to obtain the activated capsule tensor; after multiple iterations, update the weight coefficient to obtain the target capsule tensor corresponding to the convergence of the weight coefficient; obtain the target capsule vector in the capsule matrix corresponding to each amino acid residue in the target capsule tensor; calculate the norm of the target capsule vector of each amino acid residue as the real value of each amino acid residue; Splice the real values selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue; Send the classification vector of each amino acid residue in the protein sequence to be predicted into the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

2. The method for predicting protein-protein interaction sites according to claim 1, wherein Input the splicing matrix into multiple primary capsule units respectively for convolution operation to obtain the capsule matrix output by each primary capsule unit, including: The main capsule unit is a two-dimensional convolutional neural network. The hyperparameters of multiple two-dimensional convolutional neural networks are the same. The convolutional kernel size is 1×9, the stride is 1, the number of input channels and the number of output channels are both 1, and the padding operation is in the same mode; In each main capsule unit, after the splicing matrix undergoes convolutional operation and padding operation, a capsule matrix with the same size as the splicing matrix is obtained.

3. The method for predicting protein-protein interaction sites according to claim 1, characterized in that, The acquisition of the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted includes: The acquisition of the biological features of amino acid residues at each site in the protein sequence to be predicted includes: Based on the AAindex database, the physical properties and physicochemical characteristics of amino acid residues at each site in the protein sequence to be predicted are obtained; Use Psi-Blast to calculate the position-specific matrix of the protein sequence to be predicted; Based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of each type of amino acid, calculate the solvent accessible surface area of each amino acid residue in the protein sequence to be predicted; Based on the length of the protein sequence to be predicted, obtain the relative position information of each amino acid residue in the protein sequence to be predicted; Obtain the number of amino acid types and the number of amino acid categories within a preset-size window at each site in the protein sequence to be predicted; The acquisition of the semantic features of amino acid residues at each site in the protein sequence to be predicted includes: Use ProtBERT to calculate the semantic features of amino acid residues at each site in the protein sequence to be predicted.

4. The method for predicting protein-protein interaction sites according to claim 3, wherein Based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of each type of amino acid, calculate the solvent accessible surface area of each amino acid residue in the protein sequence to be predicted, expressed as: ; Among them, represents the solvent accessible surface area when the th amino acid residue in the protein sequence to be predicted is the amino acid of the j-th category; , represents the length of the protein sequence to be predicted; , represents the th type among the amino acid types, with a maximum of 20; represents the relative solvent accessibility of the th amino acid residue in the protein sequence to be predicted, represents the maximum accessible surface area of the th type of amino acid.

5. The method for predicting protein-protein interaction sites according to claim 1, wherein Use the Focal Loss to train multiple different deep learning models, bidirectional long short-term memory networks, and capsule network layer models, including: Obtain multiple protein sequence samples and the true labels of whether the amino acid residues at each site are protein-protein interaction sites; Input the protein sequence samples, and through multiple different deep learning models and bidirectional long short-term memory networks, obtain the corresponding splicing matrix; Input the splicing matrix into the capsule network layer, obtain the classification vector of each amino acid residue in the protein sequence, and through the output layer, obtain the probability that each amino acid residue belongs to an interaction site, and generate prediction labels; Based on the prediction labels and true labels of amino acid residues at each site in each protein sequence sample, calculate the probability of being predicted as the positive class; Based on the probability of being predicted as the positive class, use the Focal Loss function to calculate the model prediction loss, and update the model parameters of multiple different deep learning models, bidirectional long short-term memory networks, and capsule network layers until the Focal Loss function converges, and obtain the trained multiple different deep learning models, bidirectional long short-term memory networks, and capsule network layers.

6. The method for predicting protein-protein interaction sites according to claim 5, wherein The FocalLoss function is expressed as: ; Among them, represents the Focal Loss function, represents the weight factor with a value of ; represents the probability of predicting the positive class, represents the focusing parameter, represents the sample adjustment factor.

7. A protein-protein interaction site prediction device, characterized in that, Including: A feature extraction module, which is used to obtain the biological features and semantic features of amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted; An ensemble learning module, which is used to input the biological feature matrix into multiple different deep learning models respectively, including: inputting the biological feature matrix into a Transformer, through a multi-head attention mechanism, residual connection, normalization layer, feed-forward network, residual connection and normalization layer, to capture the long-range dependence relationship between biological features and output a multi-head attention feature matrix; inputting the biological feature matrix into a convolutional neural network, and splicing after passing through convolutional units with multiple different convolutional kernel sizes to output a local feature matrix; inputting the biological feature matrix into a bidirectional long short-term memory network to obtain a time-dependent feature matrix; horizontally splicing the multi-head attention feature matrix, the local feature matrix and the time-dependent feature matrix to obtain an integrated feature matrix; inputting the semantic feature matrix into a bidirectional long short-term memory network to output an attention feature matrix; horizontally splicing the integrated feature matrix and the attention feature matrix and then passing through a linear layer to obtain a splicing matrix of the protein sequence to be predicted; A capsule network module, which is used to input the splicing matrix of the protein sequence to be predicted into a capsule network layer to obtain the classification vector of amino acid residues at each site in the protein sequence to be predicted, including: inputting the splicing matrix into multiple primary capsule units for convolutional operations respectively, obtaining the capsule matrices output by each primary capsule unit and stacking them vertically to obtain the capsule tensor of the protein sequence to be predicted; inputting the capsule tensor into multiple classification capsules respectively; in each classification capsule, generating a real value for each amino acid residue in the protein sequence to be predicted, including: obtaining the initial weight coefficient of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted; normalizing the initial weight coefficient by using the Softmax function to obtain the normalized weight coefficient of each capsule matrix; using the normalized weight coefficient to perform weighted summation on the capsule matrices corresponding to all amino acid residues to obtain a weighted capsule tensor; activating the weighted capsule tensor by using the Squashing function to obtain an activated capsule tensor; through multiple iterations, updating the weight coefficient to obtain the target capsule tensor corresponding to when the weight coefficient converges; obtaining the target capsule vector in the capsule matrix corresponding to each amino acid residue in the target capsule tensor; calculating the norm of the target capsule vector of each amino acid residue as the real value of each amino acid residue; splicing the real values selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue; A prediction module, which is used to send the classification vector of each amino acid residue in the protein sequence to be predicted into a linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to an interaction site.

Citation Information

Patent Citations

  • Protein interaction site prediction method based on deep learning

    CN113643756A

  • Protein relative solvent accessibility prediction method and device based on sequence

    CN118629515A