Protein-protein interaction site prediction method and device

Through the combination of multi-feature fusion and capsule network, the prediction inaccuracy problem caused by the neglect of feature correlation in the prior art is solved, and the accuracy and applicability of the prediction of protein-protein interaction sites are improved.

CN119943144AActive Publication Date: 2025-05-06SUZHOU CITY UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510413894.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The prior art ignores the correlation between features in the prediction of protein-protein interaction sites, resulting in inaccurate predictions.

Method used

By obtaining the biometric matrix and semantic feature matrix of the protein sequence to be predicted, and using multiple deep learning models for feature extraction and splicing, combining with the capsule network layer, the potential correlation and complex interactions between features are captured.

Benefits of technology

It improves the accuracy of protein-protein interaction site prediction, enhances the model's ability to capture complex relationships between features, and is suitable for prediction tasks of large-scale unknown proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943144A_ABST
    Figure CN119943144A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of protein, and discloses a protein-protein interaction site prediction method and device, and the method comprises the steps: enabling a biological feature matrix of a to-be-predicted protein sequence to be subjected to the transverse splicing of output feature matrixes of a plurality of different deep learning models, and obtaining an integrated feature matrix; obtaining an attention feature matrix based on a semantic feature matrix of a to-be-predicted protein sequence, splicing the attention feature matrix with the integrated feature matrix, and obtaining a spliced matrix through a linear layer; and inputting the splicing matrix into a capsule network layer, obtaining a classification vector of the amino acid residue at each site in the to-be-predicted protein sequence, and obtaining the probability that each amino acid residue in the to-be-predicted protein sequence belongs to the interaction site through a linear layer. According to the method, the capturing capability of different models on the features is fully utilized in a multi-feature fusion mode, the potential relevance between the features is better captured through feature splicing, and the accuracy of protein-protein interaction site prediction is improved in combination with the capsule network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of protein technology, and in particular to a method and device for predicting protein-protein interaction sites. Background Art

[0002] Protein interaction sites refer to specific areas on protein molecules that physically or chemically interact with other molecules. According to different objects, they can be divided into protein-protein interaction sites (PPIS), protein-peptide interaction sites (PPepIS) and protein-ligand interaction sites (PLIS). Detailed research on protein interaction sites is of vital importance for deepening the understanding of protein function, revealing the mechanism of biological processes, drug development and analysis of disease mechanisms.

[0003] Protein-protein interactions play an important role in cellular physiological processes, and the identification of their interaction sites is crucial for understanding protein functional mechanisms and drug development. Protein-protein interactions usually occur on the surface of proteins, that is, areas with large accessible surface areas. Experimental methods for determining the interaction sites of protein complexes have the characteristics of high precision and can accurately identify the interaction sites in the complex. However, experimental methods have long cycles and high costs, so researchers have also begun to use computational methods to predict interaction sites in protein complexes with high throughput. Computational methods for predicting protein interaction sites can be described as solving the following problem: given a protein amino acid sequence P with a length of L, find the best mapping function F(P) to map it to A. The length of A is also L, and each item is 0 or 1, where 0 represents that the site is not an interaction site and 1 represents that the site is an interaction site.

[0004] At present, the commonly used protein-protein interaction site calculation methods are mainly divided into two categories: structure-based methods and sequence-based methods.

[0005] The structure-based methods mainly include: MaSIF, dMaSIF, MPNP, GraphPPIS and EGRET. Among them, MaSIF predicts interaction sites by capturing the key features of biomolecular surface interactions, dMaSIF takes the initial three-dimensional coordinates and chemical types of atoms as input to identify interaction sites in an end-to-end manner, and MPNP uses the relationship structure within the model to predict interaction sites. Yuan et al. proposed the GraphPPIS model, which uses evolutionary features and residue structure features to represent nodes in the graph, and uses the distance between residues to represent edges in the graph. Mahbub et al. constructed the EGRET model based on the graph self-attention network, which uses the protein sequence abstract features of the pre-trained model as the features of the nodes in the graph, and the distance and relative direction between two residues as the features of the edges in the graph. Sequence-based methods have received more attention.

[0006] Most of these existing technologies tend to introduce protein structure information into the prediction of protein interaction sites. Although this can improve the prediction accuracy of the model, its generalization is limited and it is not suitable for the task of predicting interaction sites of large-scale unknown proteins. Therefore, when facing unknown proteins, sequence-based prediction methods are needed.

[0007] The methods based entirely on sequences mainly include: SCRIBER, DLPred and DELPHI. Among them, SCRIBER uses a data set covering various types of binding residues to predict interacting residues through a two-layer architecture. The DLPred model proposed by Zhang et al. uses a simplified bidirectional recurrent neural network. Without using protein structure information, it combines the protein relative solvent accessibility prediction task and adopts a multi-task joint learning strategy to predict the interaction sites of proteins, achieving better results. The DELPHI model proposed by Li uses the idea of ​​ensemble learning to obtain different aspects of feature information in the same protein sequence by integrating bidirectional recurrent neural networks and convolutional neural networks, achieving good results.

[0008] In practical applications, some features often do not exist in isolation, but are connected and influence each other in complex ways; however, most of the above studies use recurrent neural networks or convolutional neural networks to mine local features or contextual features of protein sequences, ignoring the importance of potential correlations between features and their state changes, lacking effective modeling of the correlations between features, and each feature is independently applied to the prediction model without fully considering the potential interactions and dependencies between features. This leads to limitations in the model in capturing the complex interactions and differences between features, which may restrict the improvement of prediction performance.

[0009] In summary, the existing prediction methods based on protein structure information have poor generalization of prediction models and are not suitable for the prediction of large-scale positional proteins; the existing sequence-based prediction methods ignore the correlation between multiple features in protein sequences and do not fully consider the interactions and dependencies between features, which affects the prediction accuracy. Summary of the invention

[0010] Therefore, the technical problem to be solved by the present invention is to overcome the problem that the prior art ignores the correlation between features and leads to inaccurate prediction.

[0011] In order to solve the above technical problems, the present invention provides a method for predicting protein-protein interaction sites, comprising: Obtaining biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted, and obtaining a biological feature matrix and a semantic feature matrix of the protein sequence to be predicted; After the biological feature matrix is ​​input into a variety of different deep learning models, the output feature matrix of each deep learning model is horizontally spliced ​​to obtain an integrated feature matrix; the semantic feature matrix is ​​input into a bidirectional long short-term memory network to output an attention feature matrix; the integrated feature matrix and the attention feature matrix are horizontally spliced ​​and passed through a linear layer to obtain a spliced ​​matrix of the protein sequence to be predicted; The concatenation matrix of the protein sequence to be predicted is input into the capsule network layer to obtain the classification vector of the amino acid residues at each site in the protein sequence to be predicted, including: The concatenated matrix is ​​input into multiple main capsule units for convolution operation, and the capsule matrix output by each main capsule unit is obtained for vertical stacking to obtain the capsule tensor of the protein sequence to be predicted; Input the capsule tensor into multiple classification capsules respectively; in each classification capsule, generate a real value for each amino acid residue in the protein sequence to be predicted; Concatenate the real values ​​selected by each classification capsule for each amino acid residue to obtain a classification vector for each amino acid residue; The classification vector of each amino acid residue in the protein sequence to be predicted is sent to the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

[0012] Preferably, a variety of different deep learning models include: Transformer, convolutional neural network and bidirectional long short-term memory network.

[0013] Preferably, after the biometric feature matrices are input into a plurality of different deep learning models respectively, the output feature matrices of each deep learning model are horizontally spliced ​​to obtain an integrated feature matrix, including: The biometric matrix is ​​input into the Transformer, and after passing through the multi-head attention mechanism, residual connection, normalization layer, feedforward network, residual connection and normalization layer, the long-distance dependency between biometric features is captured, and the multi-head attention feature matrix is ​​output; The biometric matrix is ​​input into the convolutional neural network, and then concatenated after passing through multiple convolution units with different convolution kernel sizes to output a local feature matrix. Input the biological feature matrix into the bidirectional long short-term memory network to obtain the time-dependent feature matrix; The multi-head attention feature matrix, local feature matrix and time-dependent feature matrix are horizontally spliced ​​to obtain an integrated feature matrix.

[0014] Preferably, the concatenated matrix is ​​input into a plurality of main capsule units for convolution operation respectively, and the capsule matrix output by each main capsule unit is obtained, including: The main capsule unit is a two-dimensional convolutional neural network, and the hyperparameters of multiple two-dimensional convolutional neural networks are the same, the convolution kernel size is 1×9, the step size is 1, the number of input channels and the number of output channels are both 1, and the padding operation is the same mode; In each main capsule unit, the splicing matrix undergoes convolution and padding operations to obtain a capsule matrix of the same size as the splicing matrix.

[0015] Preferably, in each classification capsule, a real value is generated for each amino acid residue in the protein sequence to be predicted using a dynamic routing algorithm, including: Obtaining the initial weight coefficient of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted; Use the Softmax function to normalize the initial weight coefficients to obtain the normalized weight coefficients of each capsule matrix; The capsule matrices corresponding to all amino acid residues are weighted and summed using the normalized weight coefficient to obtain a weighted capsule tensor; Use the Squashing function to activate the weighted capsule tensor and obtain the activated capsule tensor; After multiple iterations, the weight coefficient is updated, and the corresponding target capsule tensor is obtained when the weight coefficient converges; Get the target capsule vector in the capsule matrix corresponding to each amino acid residue in the target capsule tensor; The magnitude of the target capsule vector for each amino acid residue is calculated as a real value for each amino acid residue.

[0016] Preferably, obtaining the biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted includes: Acquisition of biological features of amino acid residues at each site in the protein sequence to be predicted, including: Based on the AAindex database, the physical properties and physicochemical characteristics of the amino acid residues at each site in the protein sequence to be predicted are obtained; Psi-Blast was used to calculate the position-specific matrix of the protein sequence to be predicted; Calculate the solvent accessible surface area of ​​each amino acid residue in the protein sequence to be predicted based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of ​​each type of amino acid; Based on the length of the protein sequence to be predicted, obtaining the relative position information of each amino acid residue in the protein sequence to be predicted; Obtaining the number of amino acid species and the number of amino acid types within a window of a preset size at each site in the protein sequence to be predicted; The acquisition of semantic features of the amino acid residues at each site in the protein sequence to be predicted includes: using ProtBERT to calculate the semantic features of the amino acid residues at each site in the protein sequence to be predicted.

[0017] Preferably, the solvent accessible surface area of ​​each amino acid residue in the protein sequence to be predicted is calculated based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of ​​each type of amino acid, expressed as: ; in, Indicates the first The solvent accessible surface area when the amino acid residue is an amino acid of the jth category; , Indicates the length of the protein sequence to be predicted; , Indicates the amino acid type Types, maximum 20; Indicates the first The relative solvent accessibility of the amino acid residues, Indicates The maximum accessible surface area of ​​an amino acid.

[0018] Preferably, Focal Loss is used to train a variety of different deep learning models, bidirectional long short-term memory networks, and capsule network layer models, including: Obtain multiple protein sequence samples and the true labels of whether the amino acid residues at each site are protein-protein interaction sites; The protein sequence samples are input and passed through a variety of deep learning models and bidirectional long short-term memory networks to obtain the corresponding splicing matrix; The splicing matrix is ​​input into the capsule network layer to obtain the classification vector of each amino acid residue in the protein sequence. After passing through the output layer, the probability of each amino acid residue belonging to the interaction site is obtained to generate a predicted label. Based on the predicted labels and true labels of the amino acid residues at each site in each protein sequence sample, the probability of predicting a positive class is calculated; Based on the probability of predicting the positive category, the Focal Loss function is used to calculate the model prediction loss, and the model parameters of various deep learning models, bidirectional long short-term memory networks, and capsule network layers are updated until the Focal Loss function converges to obtain the trained various deep learning models, bidirectional long short-term memory networks, and capsule network layers.

[0019] Preferably, the Focal Loss function is expressed as: ; in, represents the Focal Loss function, Indicates that the value is The weight factor of represents the probability of predicting the positive category, represents the focusing parameter, Represents the sample adjustment factor.

[0020] This embodiment provides a protein-protein interaction site prediction device, comprising: A feature extraction module is used to obtain the biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted; The integrated learning module is used to input the biological feature matrix into a variety of different deep learning models, and then horizontally splice the output feature matrix of each deep learning model to obtain an integrated feature matrix; input the semantic feature matrix into a bidirectional long short-term memory network to output an attention feature matrix; horizontally splice the integrated feature matrix and the attention feature matrix and pass them through a linear layer to obtain a splicing matrix of the protein sequence to be predicted; The capsule network module is used to input the concatenation matrix of the protein sequence to be predicted into the capsule network layer to obtain the classification vector of the amino acid residue at each site in the protein sequence to be predicted, including: inputting the concatenation matrix into multiple main capsule units for convolution operation, obtaining the capsule matrix output by each main capsule unit for vertical stacking, and obtaining the capsule tensor of the protein sequence to be predicted; inputting the capsule tensor into multiple classification capsules; generating a real value for each amino acid residue in the protein sequence to be predicted in each classification capsule; and concatenating the real values ​​selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue; The prediction module is used to send the classification vector of each amino acid residue in the protein sequence to be predicted to the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

[0021] The above technical solution of the present invention has the following beneficial effects compared with the prior art: The protein-protein interaction site prediction method of the present invention analyzes and obtains the biological feature matrix and the semantic feature matrix of the protein sequence to be predicted, and uses different deep learning models to extract and splice features respectively; this multi-feature fusion method can make full use of the different models' ability to capture different features in the protein sequence, thereby improving the accuracy of the prediction. At the same time, by further splicing the integrated features with the attention features, the potential correlation between the features can be better captured, avoiding the problem of isolated feature processing. At the same time, combined with the capsule network, multiple parallel main capsule units are used to further extract and integrate the feature correlation of the splicing matrix from different angles, capture more diverse features, avoid feature omissions that may be caused by a single perspective, and make the extracted features more complete and representative; the classification capsule is used to more accurately match and distinguish capsule tensors with different categories through its internal dynamic routing mechanism, and the routing weight is adaptively adjusted according to the correlation and similarity of the features in the capsule tensor, and the feature information is accurately allocated to the corresponding category capsule, thereby improving the recognition accuracy of different categories and improving the accuracy of protein-protein interaction site prediction.

[0022] At the same time, the present invention can better deal with the problems of data sparsity and uneven label distribution through the combination of ensemble learning and capsule network; the dynamic routing mechanism of capsule network can also adaptively adjust feature weights, thereby alleviating the impact of imbalance between positive and negative samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein: Figure 1 It is a flowchart of the steps of the method for predicting protein-protein interaction sites provided by the present invention; Figure 2 It is a model network structure diagram of the protein-protein interaction site prediction method provided by the present invention; Figure 3 It is a schematic diagram of the workflow of the dynamic routing algorithm in the classification capsule; Figure 4 It is a schematic diagram of the functional modules of the protein-protein interaction site prediction method provided by the present invention. DETAILED DESCRIPTION

[0024] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0025] Reference Figure 1 As shown, the flowchart of the method for predicting protein-protein interaction sites provided by the present invention includes the following specific steps: S101: Obtaining biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted, and obtaining a biological feature matrix and a semantic feature matrix of the protein sequence to be predicted; S102: After the biological feature matrix is ​​input into a plurality of different deep learning models respectively, the output feature matrix of each deep learning model is horizontally spliced ​​to obtain an integrated feature matrix; the semantic feature matrix is ​​input into a bidirectional long short-term memory network to output an attention feature matrix; the integrated feature matrix and the attention feature matrix are horizontally spliced ​​and passed through a linear layer to obtain a spliced ​​matrix of the protein sequence to be predicted; S103: Input the concatenation matrix of the protein sequence to be predicted into the capsule network layer to obtain the classification vector of the amino acid residue at each site in the protein sequence to be predicted, including: S103-1: Input the concatenated matrix into multiple main capsule units for convolution operation, obtain the capsule matrix output by each main capsule unit, stack them vertically, and obtain the capsule tensor of the protein sequence to be predicted; S103-2: inputting the capsule tensor into a plurality of classification capsules respectively; in each classification capsule, generating a real value for each amino acid residue in the protein sequence to be predicted; S103-3: concatenating the real values ​​selected by each classification capsule for each amino acid residue to obtain a classification vector for each amino acid residue; S104: Send the classification vector of each amino acid residue in the protein sequence to be predicted to the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

[0026] This application makes predictions directly based on protein sequence information, avoiding dependence on protein structure information, making the model more suitable for large-scale unknown protein interaction site prediction tasks and improving the generalization of the model.

[0027] Reference Figure 2 As shown, it is a model network structure diagram of the protein-protein interaction site prediction method provided by the present invention; the model mainly includes an integrated learning network layer and a capsule network layer.

[0028] Specifically, the integrated learning network layer contains a variety of different deep learning models, including: Transformer, Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory Network (BiLSTM). After the biological feature matrix is ​​input into a variety of different deep learning models, the output features of each deep learning model are horizontally spliced ​​to obtain integrated features, including: The biometric matrix is ​​input into the Transformer, and after passing through the multi-head attention mechanism, residual connection, normalization layer, feedforward network, residual connection and normalization layer, the long-distance dependency between biometric features is captured, and the multi-head attention feature matrix is ​​output; The biometric matrix is ​​input into the convolutional neural network, and then concatenated after passing through multiple convolution units with different convolution kernel sizes to output a local feature matrix. Input the biological feature matrix into the bidirectional long short-term memory network to obtain the time-dependent feature matrix; The multi-head attention feature matrix, local feature matrix and time-dependent feature matrix are horizontally spliced ​​to obtain an integrated feature matrix.

[0029] The ensemble learning network layer uses different deep learning models to extract features and then splice them together; this multi-feature fusion method can make full use of the ability of different models to capture different features in protein sequences, thereby improving the accuracy of prediction. At the same time, by further splicing the ensemble feature matrix with the attention feature matrix, the potential correlation between features can be better captured, avoiding the problem of isolated feature processing.

[0030] Specifically, the capsule network layer includes a main capsule unit and a classification capsule.

[0031] The main capsule unit is a two-dimensional convolutional neural network, and the hyperparameters of multiple two-dimensional convolutional neural networks are the same, the convolution kernel size is 1×9, the step size is 1, the number of input channels and the number of output channels are both 1, and the padding operation is the same mode; in each main capsule unit, after the convolution operation and the padding operation of the splicing matrix, a capsule matrix of the same size as the splicing matrix is ​​obtained.

[0032] In each classification capsule, a dynamic routing algorithm is used to generate a real value for each amino acid residue in the protein sequence to be predicted, including: obtaining the initial weight coefficient of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted; normalizing the initial weight coefficient using the Softmax function to obtain the normalized weight coefficient of each capsule matrix; performing weighted summation on the capsule matrices corresponding to all amino acid residues using the normalized weight coefficient to obtain a weighted capsule tensor; activating the weighted capsule tensor using the nonlinear transformation Squashing function to obtain an activated capsule tensor; after multiple iterations, updating the weight coefficient to obtain the target capsule tensor corresponding to the weight coefficient convergence; obtaining the target capsule vector in the capsule matrix corresponding to each amino acid residue in the target capsule tensor; and calculating the modulus of the target capsule vector for each amino acid residue as the real value of each amino acid residue.

[0033] The capsule network layer uses multiple parallel main capsule units to further extract and integrate the feature correlation of the splicing matrix from different angles, capture more diverse features, avoid feature omissions that may be caused by a single perspective, and make the extracted features more complete and representative; the classification capsule is used to more accurately match and distinguish capsule tensors with different categories through its internal dynamic routing mechanism, and adaptively adjust the routing weights according to the correlation and similarity of the features in the capsule tensor, so as to accurately assign the feature information to the corresponding category capsules, thereby improving the recognition accuracy of different categories and improving the accuracy of protein-protein interaction site prediction.

[0034] The existing technology does not adequately handle the sparsity of data and the high imbalance of label distribution (interaction sites account for a very small proportion); the current technology only alleviates the problem of positive and negative sample imbalance from a single aspect (data or model), which does have a certain effect but cannot effectively alleviate the impact of label imbalance. The present invention can better handle the problems of data sparsity and uneven label distribution through the combination of ensemble learning and capsule network; the dynamic routing mechanism of capsule network can also adaptively adjust feature weights, thereby alleviating the impact of positive and negative sample imbalance.

[0035] In an embodiment of the present invention, the biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted are obtained, including: Acquisition of biological features of amino acid residues at each site in the protein sequence to be predicted, including: Based on the AAindex database, the physical properties and physicochemical characteristics of the amino acid residues at each site in the protein sequence to be predicted are obtained; the physical properties include hydrogen bonding ability, polarity, solubility and molecular weight; the physicochemical characteristics include physical properties, chemical properties and kinetic properties; Psi-Blast was used to calculate the position-specific matrix of the protein sequence to be predicted; the size of the position-specific matrix was N×20, where N represented the length of the protein sequence to be predicted, and a vector of length 20 was generated for the amino acid residues at each site of the protein sequence; Based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of ​​each type of amino acid, the solvent accessible surface area of ​​each amino acid residue in the protein sequence to be predicted is calculated and expressed as: ;in, Indicates the first The solvent accessible surface area when the amino acid residue is an amino acid of the jth category; , Indicates the length of the protein sequence to be predicted; , Indicates the amino acid type Types, maximum 20; Indicates the first The relative solvent accessibility of the amino acid residues, Indicates The maximum accessible surface area of ​​an amino acid.

[0036] Based on the length of the protein sequence to be predicted, obtaining the relative position information of each amino acid residue in the protein sequence to be predicted; Obtain the number of amino acid species and the number of amino acid types within a window of a preset size at each site in the protein sequence to be predicted; the amino acid types include polar, non-polar, positively charged, negatively charged, and uncharged; The semantic features of the amino acid residues at each site in the protein sequence to be predicted are obtained, including: ProtBERT is used to calculate the semantic features of the amino acid residues at each site in the protein sequence to be predicted.

[0037] In an embodiment of the present invention, Focal Loss is used to train a variety of different deep learning models, bidirectional long short-term memory networks, and capsule network layer models, including: Obtain multiple protein sequence samples and the true labels of whether the amino acid residues at each site are protein-protein interaction sites; The protein sequence samples are input and passed through a variety of deep learning models and bidirectional long short-term memory networks to obtain the corresponding splicing matrix; The splicing matrix is ​​input into the capsule network layer to obtain the classification vector of each amino acid residue in the protein sequence. After passing through the output layer, the probability of each amino acid residue belonging to the interaction site is obtained to generate a predicted label. Based on the predicted labels and true labels of the amino acid residues at each site in each protein sequence sample, the probability of predicting a positive class is calculated; Based on the probability of predicting the positive category, the Focal Loss function is used to calculate the model prediction loss, and the model parameters of various deep learning models, bidirectional long short-term memory networks, and capsule network layers are updated until the Focal Loss function converges to obtain the trained various deep learning models, bidirectional long short-term memory networks, and capsule network layers.

[0038] Focal Loss function, expressed as: ; in, represents the Focal Loss function, Indicates that the value is The weight factor balances the importance of positive and negative samples by giving high weights to rare categories; represents the probability of predicting the positive category, represents the focusing parameter, Represents the sample adjustment factor, which reduces the loss sharing of simple samples and allows the model to focus more on samples with difficult classification.

[0039] Specifically, the embodiments of the present invention will Set to 0.25, Set it to 4.

[0040] Reference Figure 2 As shown in FIG. 1 , it is a model network structure diagram of the protein-protein interaction site prediction method provided by the present invention; based on the network structure, after the protein sequence to be predicted is input, it includes: ① Obtain the biological features and semantic features of each amino acid residue in the protein sequence to be predicted, and form an N×58-dimensional biological feature matrix and an N×1024-dimensional semantic feature matrix of the protein sequence to be predicted; ②The biometric matrix is ​​passed through Transformer, CNN and BiLSTM respectively: BiLSTM: The output vector is N×128 dimensional, and the forward and reverse addition is performed, so the size of the output through BiLSTM is 128 dimensions; CNN: Convolutional neural network, which is composed of three convolution kernels of different sizes, namely (3, 3), (5, 5) and (7, 7). The output vectors of these three convolution operations are all N×29 in size and are horizontally spliced; therefore, the output vector size of the CNN module is N×87 in size; Transformer: The Encoder part of the Transformer is used. This part does not change the size of the vector, so the output vector size of this part is N×58; Finally, these three vectors are horizontally spliced. Therefore, for the N×58-dimensional biological feature matrix, the size of the integrated feature matrix obtained after passing through this module is N×273 dimensions.

[0041] ③ Input the N×1024-dimensional semantic feature matrix into BiLSTM. The hidden layer of BiLSTM is 128, that is, the size of the forward output and reverse output is 128 dimensions. For a protein of length N, after calculating the semantic features through ProtBERT, a matrix of shape N×1024 is obtained. This matrix is ​​passed into BiLSTM as input, and the shape of the output attention feature matrix is ​​N×128; ④ Concatenate the N×273-dimensional integrated feature matrix obtained by the integrated learning network layer with the N×128-dimensional attention feature matrix, and the resulting matrix shape size is N×401; input the concatenated matrix into a linear layer with an input size of 401 dimensions and an output size of 256 dimensions, so the resulting matrix shape size is N×256.

[0042] ⑤N×256 is the input of each main capsule. Each main capsule is a two-dimensional convolutional neural network with a convolution kernel size of (1, 9), and the input and output sizes of the convolutional neural network are kept unchanged. Therefore, the size of the convolutional neural network is still N×256. At this time, each amino acid is represented by a 256-dimensional vector, which can be understood as each amino acid is composed of 256 features. The eight main capsules are merged to obtain a matrix shape of N×256×8, which can be understood as each feature is converted from a single value to an 8-dimensional vector. This 8-dimensional vector is generally called a capsule vector, so an amino acid residue can be represented by 256 capsule vectors; the output of the main capsule layer is an N×256×8 capsule tensor; Each main capsule in the main capsule layer is a two-dimensional convolutional neural network. The convolutional neural network hyperparameters of the eight main capsules are consistent. The number of input channels (in_channels) and the number of output channels (out_channels) are both 1, the size of the convolution kernel is (1,9), the stride is 1, and the padding mode is "same", which ensures that the matrix shape of each main capsule input and output remains consistent.

[0043] Each capsule tensor is the result of the calculation of eight main capsules. These eight main capsules can be regarded as eight independent convolutional neural networks. Each convolutional neural network will process the input N×256 matrix, and keep the input and output shapes unchanged through the padding operation of the convolutional neural network. Therefore, through the processing of eight main capsules, eight matrices of different shapes and sizes of N×256 will be obtained. We stack them vertically to get a matrix of N×256×8, where N represents the length of the protein sequence, that is, the number of amino acids. Each amino acid is also represented as a 256×8 matrix, which can be understood as converting the 256-dimensional feature vector of each amino acid in the original sequence into 256×8. That is to say, each main capsule will process each value in the feature vector to obtain a new value, and concatenate the processing results of the eight main capsules to obtain a 1×8 vector, and convert the real value of the 256-dimensional feature vector into an 8-dimensional vector. Therefore, each capsule tensor is the common result of the eight main capsules.

[0044] ⑥ The input of the classification capsule layer is N×256×8. Each classification capsule will execute the dynamic routing algorithm. The dynamic routing algorithm will select the most appropriate capsule vector to represent the amino acid residue. Therefore, after the dynamic routing algorithm, the matrix shape will become N×1×8. Then the capsule vectors are summed internally, and the matrix shape will be transformed into N×1. This embodiment uses 16 classification capsules, so the output of the classification capsule layer is N×16. The classification capsule selects a capsule vector for each amino acid residue in the sequence, and the size of the capsule vector is (1, 8). The vector length is calculated to obtain a real value, so a classification capsule generates a (N, 1) matrix for a protein sequence. The output results of the 16 classification capsules in this embodiment are horizontally spliced ​​to output a (N, 16) matrix.

[0045] Reference Figure 3 As shown in Figure 1, it is a schematic diagram of the workflow of the dynamic routing algorithm in the classification capsule; Represents the number of input capsule vectors, express The initial weight coefficient of the capsule, Indicates After the round update The normalized weight coefficient of the capsule, Indicates input capsules vector.

[0046] ⑦The final output layer is a linear layer with an input data dimension of 16 and an output data dimension of 1. The output layer matrix is ​​transformed into N×1, and the value corresponding to each amino acid residue indicates the probability that this amino acid residue is an interaction site.

[0047] Reference Figure 4 As shown, it is a schematic diagram of the functional modules of the protein-protein interaction site prediction method provided by the present invention; the embodiment of the present invention mainly includes four modules, namely: a protein information input module, a protein sequence feature calculation module, a protein interaction site prediction module and a prediction result output module.

[0048] ① Protein information input module, which obtains protein sequence data based on the protein sequence information input by the user; The protein sequence data is transmitted to the protein sequence feature calculation module, and all input data and user information of submitted data are stored by the system.

[0049] ② Protein sequence feature calculation module, which receives protein sequence data and extracts protein sequence attribute features, including: calculating the physicochemical properties and physical characteristics of the amino acids at each site, calculating the protein position specificity matrix (PSSM reflects the frequency of amino acids at each position being replaced by other amino acids) through PAI-Blast software, calling the prediction tool to calculate the accessible surface area ASA of the amino acids at each site, calculating the length and relative position information of the sequence where the amino acid residues are located, counting the number of 20 amino acids and five types of amino acids (polar, non-polar, positively charged, negatively charged, and uncharged) at each site within a window of size 21, and using ProtBERT to calculate the potential semantic features of each amino acid in the protein sequence; The extracted protein sequence attribute features are transmitted to the protein interaction site prediction module; ③ Protein interaction site prediction module, which takes protein sequence attribute features as input, calculates the probability that the amino acid residues at each site are protein interaction sites, and obtains the prediction results; ④ The prediction result output module stores the prediction results and displays them on the browser page. It also generates a PDF file to draw the prediction results into a line graph and sends emails to the corresponding users who submitted the data according to the task.

[0050] The present invention only relies on the sequence information of proteins to extract features for model training, so that the present invention can adapt to more biological application scenarios. In view of the problems of data sparsity and highly unbalanced label distribution, this embodiment adopts three strategies to alleviate the problems of data sparsity and label imbalance; first, an oversampling technology for training sets is adopted to effectively improve the model's recognition ability of minority samples (i.e., interaction sites) by reusing and augmenting high-quality data instances; second, by drawing on the idea of ​​ensemble learning, multiple base models are used to mine the features of minority samples from multiple dimensions; finally, a loss function mechanism based on different sample weights, Focal loss, is introduced to enhance the model's learning strength for minority samples during the training iteration process by giving minority samples higher weight values ​​and paying attention to difficult-to-classify samples, thereby guiding the model to pay more attention to and accurately capture minority samples. For exploring the correlation between features, this embodiment introduces a capsule network architecture at the model architecture level. Multiple capsule vectors are generated through the main capsule layer. The capsule vector contains the correlation between features, and then the low-level capsule features are integrated through the classification capsule layer to generate a higher level of feature representation for each amino acid residue.

[0051] Based on the above embodiments, an embodiment of the present invention provides a protein-protein interaction site prediction device; the specific device may include: A feature extraction module is used to obtain the biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted; The integrated learning module is used to input the biological feature matrix into a variety of different deep learning models, and then horizontally splice the output features of each deep learning model to obtain integrated features; input the semantic features into the bidirectional long short-term memory network and output the attention features; horizontally splice the integrated features and the attention features and pass them through the linear layer to obtain the splicing matrix of the protein sequence to be predicted; The capsule network module is used to input the concatenation matrix of the protein sequence to be predicted into the capsule network layer to obtain the classification vector of the amino acid residue at each site in the protein sequence to be predicted, including: inputting the concatenation matrix into multiple main capsule units for convolution operation, obtaining the capsule matrix output by each main capsule unit for vertical stacking, and obtaining the capsule tensor of the protein sequence to be predicted; inputting the capsule tensor into multiple classification capsules; generating a real value for each amino acid residue in the protein sequence to be predicted in each classification capsule; and concatenating the real values ​​selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue; The prediction module is used to send the classification vector of each amino acid residue in the protein sequence to be predicted to the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

[0052] The protein-protein interaction site prediction device of this embodiment is used to implement the aforementioned protein-protein interaction site prediction method. Therefore, the specific implementation method of the protein-protein interaction site prediction device can be seen in the embodiment part of the protein-protein interaction site prediction method in the previous text. For example, the feature extraction module, the integrated learning module, the capsule network module, and the prediction module are respectively used to implement steps S101, S102, S103 and S104 in the above-mentioned protein-protein interaction site prediction method. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part, which will not be repeated here.

[0053] Due to the inherent differences of different prediction tools in training data sets, capPPISp needs to be compared and analyzed with a variety of different prediction methods on a series of public and recognized benchmark test sets. In the embodiments of the present invention, benchmark data sets Dset_448, Dset_355, Dset_164, Dset_186, and Dset_72 are selected respectively to compare the performance of the protein-protein interaction site prediction method capPPISp provided by the present invention with other technologies. The comparison results are shown in Table 1, Table 2, Table 3, Table 4, and Table 5, respectively. Table 1 Performance comparison of capPPISp and other prediction tools on Dset_448 Prediction Tools Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.202 0.870 0.194 0.781 0.198 0.071 0.517 0.159 SPRINT 0.183 0.873 0.183 0.781 0.183 0.057 0.570 0.167 PSIVER 0.191 0.874 0.191 0.783 0.191 0.066 0.581 0.170 SPRINGS 0.229 0.882 0.228 0.796 0.229 0.111 0.625 0.201 LORIS 0.264 0.887 0.263 0.805 0.263 0.151 0.656 0.228 CRFPPI 0.268 0.887 0.264 0.805 0.266 0.154 0.681 0.238 SSWRF 0.288 0.891 0.286 0.811 0.287 0.178 0.687 0.256 SCRIBER 0.334 <![CDATA[ 0.896 ]]> <![CDATA[ 0.332 ]]> <![CDATA[ 0.821 ]]> 0.333 0.230 0.715 0.287 DELPHI <![CDATA[ 0.371 ]]> 0.901 0.371 0.829 0.371 0.272 0.737 0.337 capPPISp (ours) 0.675 0.659 0.237 0.661 <![CDATA[ 0.351 ]]> <![CDATA[ 0.235 ]]> <![CDATA[ 0.728 ]]> <![CDATA[ 0.310 ]]> Based on Table 1, we can see that capPPISp is close to the best prediction tool in terms of important balanced classification performance indicators AUPRC, MCC and F1, ranking second. capPPISp is higher than other prediction tools in terms of Recall indicator, 30.4% higher than the existing best tool, proving that capPPISp has better positive sample recognition ability.

[0054] Table 2 Performance comparison of capPPISp and other prediction tools on Dset_355 Prediction Tools Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.180 0.889 0.180 0.804 0.180 0.068 0.515 0.138 SPRINT 0.168 0.886 0.167 0.801 0.168 0.054 0.571 0.150 PSIVER 0.178 0.888 0.177 0.803 0.177 0.065 0.583 0.155 SPRINGS 0.211 0.892 0.210 0.811 0.211 0.103 0.608 0.178 LORIS 0.242 0.896 0.240 0.818 0.241 0.137 0.637 0.203 CRFPPI 0.247 0.897 0.245 0.819 0.246 0.143 0.662 0.214 SSWRF 0.268 0.901 0.268 0.825 0.268 0.168 0.667 0.228 DLPred 0.308 0.906 0.308 0.835 0.308 0.214 0.724 0.272 SCRIBER 0.322 <![CDATA[ 0.908 ]]> <![CDATA[ 0.322 ]]> <![CDATA[ 0.838 ]]> 0.322 0.230 0.719 0.275 DELPHI <![CDATA[ 0.364 ]]> 0.914 0.364 0.848 0.364 0.278 0.746 0.326 capPPISp (ours) 0.656 0.676 0.215 0.673 <![CDATA[ 0.324 ]]> <![CDATA[ 0.224 ]]> <![CDATA[ 0.729 ]]> <![CDATA[ 0.282 ]]> Based on Table 2, on the Dset_355 benchmark test set, capPPISp performs well in key indicators of balanced classification performance, namely AUPRC, MCC, and F1. capPPISp is higher than other prediction tools in the recall metric, proving that capPPISp has better positive sample recognition capabilities, further verifying the superiority of capPPISp in identifying protein interaction sites.

[0055] Table 3 Performance comparison of capPPISp and other prediction tools on Dset_164 Prediction Tools Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.264 0.828 0.253 0.726 0.258 0.090 0.528 0.220 PSIVER 0.217 0.826 0.216 0.716 0.216 0.043 0.554 0.205 CRFPPI 0.280 0.841 0.280 0.739 0.280 0.121 0.608 0.267 SSWRF 0.266 0.838 0.266 0.734 0.266 0.103 0.606 0.243 DLPred 0.338 <![CDATA[ 0.854 ]]> <![CDATA[ 0.338 ]]> 0.760 0.338 <![CDATA[ 0.192 ]]> <![CDATA[ 0.672 ]]> 0.330 SCRIBER 0.327 0.851 0.327 0.856 0.327 0.179 0.657 0.301 DELPHI <![CDATA[ 0.352 ]]> 0.857 0.352 <![CDATA[ 0.765 ]]> <![CDATA[ 0.352 ]]> 0.209 0.685 <![CDATA[ 0.332 ]]> capPPISp (ours) 0.515 0.812 0.322 0.639 0.354 0.130 0.621 0.363 Based on Table 3, we can see that on the Dset_164 benchmark test set, capPPISp outperforms other existing prediction tools in terms of AUPRC and F1 balanced classification performance indicators, which are 3.1% and 0.2% higher respectively. In terms of Recall, which measures the ability to recognize positive samples, capPPISp is also 16.3% higher than the existing optimal model.

[0056] Table 4 Performance comparison of capPPISp and other prediction tools on Dset_186 Prediction Tools Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.194 0.848 0.186 0.748 0.190 0.041 0.499 0.165 DLPred 0.320 <![CDATA[ 0.878 ]]> <![CDATA[ 0.320 ]]> <![CDATA[ 0.783 ]]> 0.320 <![CDATA[ 0.198 ]]> <![CDATA[ 0.694 ]]> 0.290 SCRIBER 0.279 0.870 0.279 0.780 0.279 0.150 0.647 0.246 DELPHI <![CDATA[ 0.351 ]]> 0.884 0.351 0.803 0.351 0.235 0.710 0.319 capPPISp (ours) 0.583 0.859 0.256 0.603 <![CDATA[ 0.331 ]]> 0.131 0.637 0.323 Based on Table 4, we can see that on the Dset_186 benchmark test set, capPPISp outperforms other existing prediction tools in AUPRC, 0.4% higher than the existing optimal prediction tool, and Recall is also 23.2% higher than the existing optimal model.

[0057] Table 5 Performance comparison of capPPISp and other prediction tools on Dset_72 Prediction Tools Recall SPE PPV ACC F1 MCC AUROC AUPRC SPPIDER 0.188 0.898 0.179 0.823 0.183 0.084 0.522 0.134 PSIVER 0.152 0.899 0.152 0.820 0.152 0.052 0.604 0.141 CRFPPI 0.248 <![CDATA[ 0.911 ]]> <![CDATA[ 0.248 ]]> <![CDATA[ 0.840 ]]> 0.248 <![CDATA[ 0.158 ]]> 0.669 0.200 SSWRF 0.246 <![CDATA[ 0.911 ]]> 0.246 0.840 0.246 0.157 0.678 0.198 DLPred 0.246 0.901 0.246 0.826 0.246 0.148 <![CDATA[ 0.688 ]]> 0.215 SCRIBER 0.232 0.909 0.232 0.837 0.232 0.141 0.680 0.198 DELPHI <![CDATA[ 0.274 ]]> 0.914 0.274 0.847 <![CDATA[ 0.274 ]]> 0.189 0.711 <![CDATA[ 0.237 ]]> capPPISp (ours) 0.481 0.885 0.216 0.655 0.275 0.115 0.625 0.282 Based on Table 5, we can see that on the Dset_72 benchmark test set, capPPISp outperforms other existing prediction tools in terms of AUPRC and F1 balanced classification performance indicators by 4.5% and 0.1% respectively. In terms of Recall, capPPISp also outperforms the existing best tool by 20.7%.

[0058] The protein-protein interaction site prediction method of the present invention analyzes and obtains the biological feature matrix and the semantic feature matrix of the protein sequence to be predicted, and uses different deep learning models to extract features and then splice them; this multi-feature fusion method can make full use of the different models' ability to capture different features in the protein sequence, thereby improving the accuracy of the prediction. At the same time, by further splicing the integrated feature matrix with the attention feature matrix, the potential correlation between the features can be better captured, avoiding the problem of isolated feature processing. At the same time, combined with the capsule network, multiple parallel main capsule units are used to further extract and integrate the feature correlation of the splicing matrix from different angles, capture more diverse features, avoid feature omissions that may be caused by a single perspective, and make the extracted features more complete and representative; the classification capsule is used to more accurately match and distinguish capsule tensors with different categories through its internal dynamic routing mechanism, and the routing weight is adaptively adjusted according to the correlation and similarity of the features in the capsule tensor, and the feature information is accurately allocated to the corresponding category capsule, thereby improving the recognition accuracy of different categories and improving the accuracy of protein-protein interaction site prediction. At the same time, the present invention can better deal with the problems of data sparsity and uneven label distribution through the combination of ensemble learning and capsule network; the dynamic routing mechanism of capsule network can also adaptively adjust feature weights, thereby alleviating the impact of imbalance between positive and negative samples.

[0059] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0060] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0061] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0062] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0063] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.

Claims

1. A method for predicting protein-protein interaction sites, characterized in that: include: Obtaining biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted, and obtaining a biological feature matrix and a semantic feature matrix of the protein sequence to be predicted; After the biological feature matrix is ​​input into a variety of different deep learning models, the output feature matrix of each deep learning model is horizontally spliced ​​to obtain an integrated feature matrix; the semantic feature matrix is ​​input into a bidirectional long short-term memory network to output an attention feature matrix; The integrated feature matrix and the attention feature matrix are horizontally spliced ​​and passed through a linear layer to obtain a spliced ​​matrix of the protein sequence to be predicted; The concatenation matrix of the protein sequence to be predicted is input into the capsule network layer to obtain the classification vector of the amino acid residues at each site in the protein sequence to be predicted, including: The concatenated matrix is ​​input into multiple main capsule units for convolution operation, and the capsule matrix output by each main capsule unit is obtained for vertical stacking to obtain the capsule tensor of the protein sequence to be predicted; Input the capsule tensor into multiple classification capsules respectively; in each classification capsule, generate a real value for each amino acid residue in the protein sequence to be predicted; Concatenate the real values ​​selected by each classification capsule for each amino acid residue to obtain a classification vector for each amino acid residue; The classification vector of each amino acid residue in the protein sequence to be predicted is sent to the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

2. The method for predicting protein-protein interaction sites according to claim 1, characterized in that: A variety of deep learning models, including Transformer, Convolutional Neural Networks, and Bidirectional Long Short-Term Memory Networks.

3. The method for predicting protein-protein interaction sites according to claim 2, characterized in that: After the biological feature matrix is ​​input into a variety of different deep learning models, the output feature matrix of each deep learning model is horizontally spliced ​​to obtain an integrated feature matrix, including: The biometric matrix is ​​input into the Transformer, and after passing through the multi-head attention mechanism, residual connection, normalization layer, feedforward network, residual connection and normalization layer, the long-distance dependency between biometric features is captured, and the multi-head attention feature matrix is ​​output; The biometric matrix is ​​input into the convolutional neural network, and then concatenated after passing through multiple convolution units with different convolution kernel sizes to output a local feature matrix. Input the biological feature matrix into the bidirectional long short-term memory network to obtain the time-dependent feature matrix; The multi-head attention feature matrix, local feature matrix and time-dependent feature matrix are horizontally spliced ​​to obtain an integrated feature matrix.

4. The method for predicting protein-protein interaction sites according to claim 1, characterized in that: The concatenated matrix is ​​input into multiple main capsule units for convolution operation to obtain the capsule matrix output by each main capsule unit, including: The main capsule unit is a two-dimensional convolutional neural network, and the hyperparameters of multiple two-dimensional convolutional neural networks are the same, the convolution kernel size is 1×9, the step size is 1, the number of input channels and the number of output channels are both 1, and the padding operation is the same mode; In each main capsule unit, the splicing matrix undergoes convolution and padding operations to obtain a capsule matrix of the same size as the splicing matrix.

5. The method for predicting protein-protein interaction sites according to claim 1, characterized in that: In each classification capsule, a dynamic routing algorithm is used to generate a real value for each amino acid residue in the protein sequence to be predicted, including: Obtaining the initial weight coefficient of the capsule matrix corresponding to each amino acid residue in the capsule tensor of the protein sequence to be predicted; Use the Softmax function to normalize the initial weight coefficients to obtain the normalized weight coefficients of each capsule matrix; The capsule matrices corresponding to all amino acid residues are weighted and summed using the normalized weight coefficient to obtain a weighted capsule tensor; Use the Squashing function to activate the weighted capsule tensor and obtain the activated capsule tensor; After multiple iterations, the weight coefficient is updated, and the corresponding target capsule tensor is obtained when the weight coefficient converges; Get the target capsule vector in the capsule matrix corresponding to each amino acid residue in the target capsule tensor; The magnitude of the target capsule vector for each amino acid residue is calculated as a real value for each amino acid residue.

6. The method for predicting protein-protein interaction sites according to claim 1, characterized in that: The biological and semantic features of the amino acid residues at each site in the protein sequence to be predicted are obtained, including: Acquisition of biological features of amino acid residues at each site in the protein sequence to be predicted, including: Based on the AAindex database, the physical properties and physicochemical characteristics of the amino acid residues at each site in the protein sequence to be predicted are obtained; Psi-Blast was used to calculate the position-specific matrix of the protein sequence to be predicted; Calculate the solvent accessible surface area of ​​each amino acid residue in the protein sequence to be predicted based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of ​​each type of amino acid; Based on the length of the protein sequence to be predicted, obtaining the relative position information of each amino acid residue in the protein sequence to be predicted; Obtaining the number of amino acid species and the number of amino acid types within a window of a preset size at each site in the protein sequence to be predicted; The acquisition of semantic features of the amino acid residues at each site in the protein sequence to be predicted includes: using ProtBERT to calculate the semantic features of the amino acid residues at each site in the protein sequence to be predicted.

7. The method for predicting protein-protein interaction sites according to claim 6, characterized in that: Based on the relative solvent accessibility of each amino acid residue in the protein sequence to be predicted and the maximum accessible surface area of ​​each type of amino acid, the solvent accessible surface area of ​​each amino acid residue in the protein sequence to be predicted is calculated and expressed as: ; in, Indicates the first The solvent accessible surface area when the amino acid residue is an amino acid of the jth category; , Indicates the length of the protein sequence to be predicted; , Indicates the amino acid type Types, maximum 20; Indicates the first The relative solvent accessibility of the amino acid residues, Indicates The maximum accessible surface area of ​​an amino acid.

8. The method for predicting protein-protein interaction sites according to claim 1, characterized in that: Use Focal Loss to train a variety of deep learning models, bidirectional long short-term memory networks, and capsule network layer models, including: Obtain multiple protein sequence samples and the true labels of whether the amino acid residues at each site are protein-protein interaction sites; The protein sequence samples are input and passed through a variety of deep learning models and bidirectional long short-term memory networks to obtain the corresponding splicing matrix; The splicing matrix is ​​input into the capsule network layer to obtain the classification vector of each amino acid residue in the protein sequence. After passing through the output layer, the probability of each amino acid residue belonging to the interaction site is obtained to generate a predicted label. Based on the predicted labels and true labels of the amino acid residues at each site in each protein sequence sample, the probability of predicting a positive class is calculated; Based on the probability of predicting the positive category, the Focal Loss function is used to calculate the model prediction loss, and the model parameters of various deep learning models, bidirectional long short-term memory networks, and capsule network layers are updated until the Focal Loss function converges to obtain the trained various deep learning models, bidirectional long short-term memory networks, and capsule network layers.

9. The method for predicting protein-protein interaction sites according to claim 8, characterized in that: FocalLoss function, expressed as: ; in, represents the Focal Loss function, Indicates that the value is The weight factor of represents the probability of predicting the positive category, represents the focusing parameter, Represents the sample adjustment factor.

10. A protein-protein interaction site prediction device, characterized in that: include: A feature extraction module is used to obtain the biological features and semantic features of the amino acid residues at each site in the protein sequence to be predicted, and obtain the biological feature matrix and semantic feature matrix of the protein sequence to be predicted; The integrated learning module is used to input the biological feature matrix into multiple different deep learning models, and then horizontally splice the output feature matrix of each deep learning model to obtain an integrated feature matrix; input the semantic feature matrix into the bidirectional long short-term memory network, and output the attention feature matrix; The integrated feature matrix and the attention feature matrix are horizontally spliced ​​and passed through a linear layer to obtain a spliced ​​matrix of the protein sequence to be predicted; The capsule network module is used to input the concatenation matrix of the protein sequence to be predicted into the capsule network layer to obtain the classification vector of the amino acid residue at each site in the protein sequence to be predicted, including: inputting the concatenation matrix into multiple main capsule units for convolution operation, obtaining the capsule matrix output by each main capsule unit for vertical stacking, and obtaining the capsule tensor of the protein sequence to be predicted; inputting the capsule tensor into multiple classification capsules; generating a real value for each amino acid residue in the protein sequence to be predicted in each classification capsule; and concatenating the real values ​​selected by each classification capsule for each amino acid residue to obtain the classification vector of each amino acid residue; The prediction module is used to send the classification vector of each amino acid residue in the protein sequence to be predicted to the linear output layer to obtain the probability that each amino acid residue in the protein sequence to be predicted belongs to the interaction site.

Citation Information

Patent Citations

  • Protein-ligand binding site prediction algorithm based on deep learning

    CN110689920A

  • Protein-protein interaction site prediction method based on deep map convolutional network

    CN113192559A

  • Protein interaction site prediction method based on deep learning

    CN113643756A

  • Protein relative solvent accessibility prediction method and device based on sequence

    CN118629515A