Method and system for predicting species-specific protein post-translational modification sites
By adopting species-specific methods and semantic adversarial strategies in the prediction of post-translational modification site of proteins, the problems of inefficient prediction methods and low accuracy in the prior art are solved, and high-precision multi-species post-translational modification site prediction are achieved.
Patent Information
- Application Number
- CN202210405901.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-18
AI Technical Summary
The existing methods for predicting post-translational modification site of proteins have problems such as inefficient, expensive, time-consuming, fitting and low prediction accuracy, especially the differences in data distribution between different species, resulting in poor model prediction results.
The species-specific protein post-translation modification site prediction method is adopted. By collecting and sorting the post-translation modification sites of different species from the preset database, local sequences are extracted using sliding window technology, and data sets are constructed by single-hot encoding, and dense connections are added on the basis of the convolutional neural network, and training is combined with semantic adversarial strategies to achieve domain adaptation and knowledge transfer.
It improves the prediction accuracy of post-translational modification sites on different species, avoids the problem of model overfitting, is suitable for small data-quantity species, and expands the application scope of prediction models.
Smart Images

Figure CN114724629B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an algorithm for predicting protein post-translational modification sites, belonging to the field of bioinformatics, and particularly relates to a method and system for predicting species-specific protein post-translational modification sites. Background Art
[0002] Protein post-translational modification is one of the important fields in current proteomics research, which has a very important impact on the structure and function of proteins. So far, more than four hundred different types of post-translational modifications have been identified, and the most common modification types include phosphorylation, ubiquitination, and acetylation, etc. Post-translational modification is a key mechanism for increasing protein diversity and plays an important role in multiple processes such as cell differentiation and development, metabolism, signal transduction, muscle contraction, and gene expression. Early methods for identifying protein post-translational modification sites mainly relied on low-throughput experiments, but limited by the experimental environment and the complicated experimental process, this method is still inefficient, expensive, and time-consuming. Therefore, computational methods for predicting protein post-translational modification sites have gradually received attention.
[0003] Although there are currently various methods for predicting protein post-translational modification sites. For example, Wang et al. (WANG D, ZENG S, XU C, et al. MusiteDeep: a deep-learning framework for general and kinase-specific phosphorylation site prediction[J]. Bioinformatics, 2017, 33(24): 3909-16) constructed a phosphorylation site predictor based on a multi-layer convolutional neural network. Chen et al. (CHEN Z, LIU X, LI F, et al. Large-scale comparative assessment of computational predictors for lysine post-translational modification sites[J]. Briefings in bioinformatics, 2019, 20(6): 2267-90) designed a variety of post-translational modification site prediction tools for lysine, namely MUscADEL, based on a bidirectional long short-term memory neural network. However, these methods still have the following drawbacks: Most existing studies develop general post-translational modification site prediction models based on deep learning algorithms, ignoring the distribution differences of post-translational modification data in different species, resulting in the difficulty for the models to achieve good prediction results on different species. In addition, at present, except for several relatively common species such as humans and mice, the post-translational modification site data of most species are scarce, and the model fine-tuning strategies adopted in existing prediction methods are prone to overfitting problems, bringing challenges to the research on species-specific post-translational modification site prediction. Therefore, it is necessary to develop more effective domain adaptation methods to achieve the transfer of post-translational modification knowledge between different species, thereby improving the prediction accuracy.
[0004] The invention patent "Protein Coding Method, Protein Post-translational Modification Site Prediction Method and System" with the application number CN201910253412.9 collects modification site information, position weight training, and encoding of peptides to be encoded. The protein post-translational modification site prediction method includes collecting modification site information, feature encoding, model training, and protein post-translational modification site prediction. This invention application uses a deep neural network and penalized logistic regression to construct prediction models for digital vector features of different categories of positive and negative sites respectively, obtaining multiple prediction models; taking the prediction results of each prediction model as new features and using penalized logistic regression to construct the final model. This existing patent can capture more protein information, thus helping to improve the prediction accuracy and can quickly identify protein modification sites on a large scale. However, the logic adopted in the process of constructing models separately in the aforementioned existing patent is to use a deep neural network and penalized logistic regression. At the same time, the logic and parameters used for sample label classification in this existing patent are also different from the technical solution of this application. Therefore, the technical solution disclosed in this existing invention is significantly different from this application. This application only discloses prediction examples for humans, rats, and mice in the implementation manner and cannot reduce the distribution difference of post-translational modification data between humans and other species.
[0005] In summary, the prior art has technical problems such as inefficient, expensive, time-consuming, overfitting prediction methods and low prediction accuracy. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to solve the technical problems of inefficient, expensive, time-consuming, overfitting prediction methods and low prediction accuracy existing in the prior art.
[0007] The present invention solves the above technical problems by adopting the following technical solutions: The species-specific protein post-translational modification site prediction method includes:
[0008] S1. Collect and organize ubiquitination and acetylation post-translational modification sites of different species from a preset database, and use a sliding window technique to extract local sequences of the post-translational modification sites to construct a data set, and one-hot encode the protein local sequences in the data set to obtain post-translational modification sites and training samples of different species;
[0009] S2. Add dense connections on the basis of a convolutional neural network so that each convolutional layer receives feature maps from all the previous convolutional layers, and accordingly construct a sequence feature extraction network and set a classifier and a domain category discriminator;
[0010] S3. Use a semantic adversarial strategy to pair and process the training samples of different species to obtain a sample pairing group of human post-translational modification samples and post-translational modification samples of other species;
[0011] S4. Use the post-translation modification samples of humans to train the sequence feature extraction network and the classifier, train the domain category discriminator according to the sample pairing groups to distinguish the group information of the input sample pairs, and use an alternating method to train the sequence feature extraction network and the domain category discriminator respectively until the loss function converges, so that the sequence feature extraction network, the classifier, and the domain category discriminator learn a domain-invariant discriminative feature space;
[0012] S5. Divide an independent test set from the post-translation modification sites, and use it to evaluate the performance of the sequence feature extraction network, the classifier, and the domain category discriminator.
[0013] The sequence feature extraction network in the present invention adds dense connections on the basis of a convolutional neural network to enhance the propagation of sequence features in the model and extract more effective sequence features. The present invention realizes domain adaptation based on semantic confrontation, effectively transfers the knowledge of human post-translation modification to assist the site prediction of other species, and realizes high-precision species-specific post-translation modification site prediction.
[0014] In a more specific technical solution, the step S1 includes:
[0015] S11. Collect lysine acetylation sites from databases such as HPRD, PLMD, dbPTM, PhosphoSitePlus, and mUbiSida as positive example samples;
[0016] S12. Regard other lysine sites in the corresponding proteins that have not been reported to undergo acetylation as negative example samples;
[0017] S13. Use the CD-HIT tool to remove proteins with similarity greater than 40% in the dataset;
[0018] S14. Use a sliding window method to intercept 15 amino acids upstream and downstream of each post-translation modification site to form an amino acid sequence with a length of 31;
[0019] S15. One-hot encode the amino acid sequence into a feature matrix;
[0020] S16. Divide the dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0021] In a more specific technical solution, the step S2 includes:
[0022] S21. Set dense connections for the convolutional modules of the convolutional neural network, and use the following logic to make each convolutional layer receive the feature maps of all the previous convolutional layers:
[0023]
[0024] , where represents the splicing operation on the feature map, and W i is the weight of the convolution kernel;
[0025] S22. Add a bottleneck layer after each of the convolution modules according to the following logic to reduce the number of channels of the features:
[0026] h B = α B (W B h M + b B )
[0027] , where h M is the output of the convolution module, W B , b B and α B respectively represent the weight, bias term, and activation function of the convolution kernel in the bottleneck layer;
[0028] S23. Set the classifier and the domain category discriminator.
[0029] The sequence feature extraction network in the present invention is composed of multiple dense connection modules. Each dense connection module is connected to a bottleneck layer composed of 1×1 convolution to reduce the number of channels of the features and reduce the model burden.
[0030] In a more specific technical solution, the step S3 includes:
[0031] S31. Construct a sample pairing group p1 from two of the differential species training samples that both come from the human acetylation dataset and have the same label;
[0032] S32. Construct a sample pairing group p2 from two of the differential species training samples that respectively come from the human and other species acetylation datasets but have the same label;
[0033] S33. Construct a sample pairing group p3 from two of the differential species training samples that both come from the human acetylation dataset but have different labels;
[0034] S34. Construct a sample pairing group p4 from two of the differential species training samples that respectively come from the human and other species acetylation datasets and have different labels.
[0035] The semantic adversarial training strategy proposed during the model training stage: Pair human training samples with training samples of other species to alleviate the problem of insufficient data for other species. It can effectively reduce the distribution difference of post-translational modification data between humans and other species, enabling the model to effectively transfer human post-translational modification knowledge for assisting in site prediction of other species, thereby effectively avoiding the overfitting problem of the model, making the model more applicable to species with small data volumes, and improving the application scope of the model.
[0036] In a more specific technical solution, the step S4 includes:
[0037] S41. Freeze the parameters of the sequence feature extraction network and the classifier, and update the parameters of the domain category discriminator;
[0038] S42. Freeze the parameters of the domain category discriminator, the training feature extraction network, and the classifier, so that the domain category discriminator does not distinguish the sample pairing groups;
[0039] S43. Continuously train the sequence feature extraction network, the classifier, and the domain category discriminator with the steps S41 and S42 until the loss function converges to align the semantics of acetylation of each species.
[0040] Compared with the existing post-translational modification site prediction methods, the species-specific prediction scheme proposed by the present invention can establish prediction models for different species respectively, improving the prediction accuracy of the model on different species. The present invention provides a new idea for species-specific post-translational modification site prediction, and the prediction results can provide a basis for the research of protein molecular mechanisms, which has guiding significance for revealing the functions of proteins.
[0041] In a more specific technical solution, in the step S43, the standard cross-entropy loss function is used to optimize the domain category discriminator:
[0042]
[0043] , where E[·] represents statistical expectation, represents the label of the sample pair p i and φ represents concatenating the features of the two samples in the sample pair p i .
[0044] In a more specific technical solution, in the step S43, the following loss function is used to optimize the feature extraction network:
[0045] In a more specific technical solution, in the step S43, the following logical combination classification loss is used to ensure the PTM site prediction accuracy:
[0046]
[0047] where λ represents the weight coefficient for balancing the classification loss and the domain confusion loss, X s and X t represent human and other species samples respectively, g s and g t represent the sequence feature extraction networks for humans and other species respectively, h s and h t represent the classifiers for humans and other species respectively, and l represents the standard binary cross-entropy loss function.
[0048] In the present invention, by alternately training the sequence feature extraction network and the domain category discriminator, the model finally learns a discriminative feature space with domain invariance, in which the human acetylation knowledge can be effectively transferred, and at the same time, high-precision acetylation site prediction for multiple species can be achieved.
[0049] In a more specific technical solution, the step S5 includes:
[0050] S51. Feeding the sample pairs in the training set into the constructed deep neural network for iterative training;
[0051] S52. Looping through the training set;
[0052] S53. Every time the data in the training set is traversed, calculate the evaluation index on the validation set according to the following logic, and accordingly retain the applicable sequence feature extraction network, the classifier, and the domain category discriminator:
[0053]
[0054]
[0055]
[0056]
[0057]
[0058] , where TP represents the total number of acetylated sites correctly predicted, TN represents the total number of non-acetylated sites correctly predicted, FP represents the total number of samples mispredicted as acetylated sites, and FN represents the total number of acetylated sites not predicted.
[0059] In a more specific technical solution, the species-specific protein post-translational modification site prediction system includes:
[0060] A data preprocessing module, which is used to collect and sort out ubiquitination and acetylation post-translational modification sites of different species from a preset database, and use a sliding window technique to extract local sequences of the post-translational modification sites to construct a data set, and one-hot encode the protein local sequences in the data set to obtain post-translational modification sites and training samples of different species;
[0061] A network construction module, which is used to add dense connections on the basis of a convolutional neural network so that each convolutional layer receives the feature maps of all the previous convolutional layers, and accordingly constructs a sequence feature extraction network and sets a classifier and a domain category discriminator. The network construction module is connected to the data preprocessing module;
[0062] A training sample construction module, which is used to pair and process the training samples of different species by using a semantic adversarial strategy to obtain a sample pairing group of human post-translational modification samples and post-translational modification samples of other species. The training sample construction module is connected to the data preprocessing module;
[0063] An adversarial training module, which is used to train the sequence feature extraction network and the classifier with the human post-translational modification samples, train the domain category discriminator according to the post-translational modification samples of other species to distinguish the group information of the input sample pairs, and use an alternating method to train the sequence feature extraction network and the domain category discriminator respectively until the loss function converges, so that the sequence feature extraction network, the classifier and the domain category discriminator learn a domain-invariant discriminant feature space. The adversarial training module is connected to the training sample construction module and the network construction module;
[0064] A performance evaluation module, which is used to evaluate the performance of the sequence feature extraction network, the classifier and the domain category discriminator according to the post-translational modification sites. The performance evaluation module is connected to the adversarial training module.
[0065] The present invention has the following advantages compared with the prior art: The sequence feature extraction network in the present invention adds dense connections on the basis of a convolutional neural network, enhances the propagation of sequence features in the model, and extracts more effective sequence features. The present invention realizes domain adaptation based on semantic adversarial, effectively transfers human post-translational modification knowledge to assist in the site prediction of other species, and realizes high-precision species-specific post-translational modification site prediction.
[0066] The sequence feature extraction network in the present invention is composed of multiple dense connection modules, and each dense connection module is connected to a bottleneck layer composed of a 1*1 convolution, which is used to reduce the number of channels of the features and reduce the model burden.
[0067] The semantic adversarial training strategy proposed in the model training stage: pair human training samples with training samples of other species to alleviate the problem of insufficient data of other species. It can effectively reduce the distribution difference of post-translational modification data between humans and other species, enabling the model to effectively transfer human post-translational modification knowledge for assisting the site prediction of other species, thus effectively avoiding the overfitting problem of the model, making the model more applicable to species with small data volumes, and improving the application scope of the model.
[0068] Compared with the existing post-translational modification site prediction methods, the species-specific prediction scheme proposed by the present invention can establish prediction models for different species respectively, improving the prediction accuracy of the model on different species. The present invention provides a new idea for species-specific post-translational modification site prediction, and the prediction results can provide a basis for the research of protein molecular mechanisms and have guiding significance for revealing the functions of proteins.
[0069] By alternately training the sequence feature extraction network and the domain category discriminator, the model finally learns a discriminative feature space with domain invariance. In this feature space, it can effectively transfer human acetylation knowledge and at the same time achieve high-precision acetylation site prediction for multiple species. The present invention solves the technical problems of low efficiency, high cost, time-consuming, overfitting and low prediction accuracy existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a flowchart of the method DeepAce-Pred of the present invention;
[0071] Figure 2 It is a structural diagram of the sequence feature extraction network constructed by the method of the present invention;
[0072] Figure 3-a It is a schematic diagram of the comparison result of the first ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0073] Figure 3-b It is a schematic diagram of the comparison result of the second ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0074] Figure 3-c It is a schematic diagram of the comparison result of the third ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0075] Figure 3-d It is a schematic diagram of the comparison result of the fourth ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0076] Figure 3-e It is a schematic diagram of the comparison result of the fifth ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0077] Figure 3-f Schematic diagram of the comparison result of the sixth ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0078] Figure 3-g Schematic diagram of the comparison result of the seventh ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0079] Figure 3-h Schematic diagram of the comparison result of the eighth ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0080] Figure 3-i Schematic diagram of the comparison result of the ninth ROC curve of the method of the present invention and other existing methods in acetylation modification;
[0081] Figure 4-a Schematic diagram of the comparison result of the first performance index of the method of the present invention and other existing methods in acetylation modification;
[0082] Figure 4-b Schematic diagram of the comparison result of the second performance index of the method of the present invention and other existing methods in acetylation modification;
[0083] Figure 4-c Schematic diagram of the comparison result of the third performance index of the method of the present invention and other existing methods in acetylation modification;
[0084] Figure 4-d Schematic diagram of the comparison result of the fourth performance index of the method of the present invention and other existing methods in acetylation modification;
[0085] Figure 4-e Schematic diagram of the comparison result of the fifth performance index of the method of the present invention and other existing methods in acetylation modification;
[0086] Figure 4-f Schematic diagram of the comparison result of the sixth performance index of the method of the present invention and other existing methods in acetylation modification;
[0087] Figure 4-g Schematic diagram of the comparison result of the seventh performance index of the method of the present invention and other existing methods in acetylation modification;
[0088] Figure 4-h Schematic diagram of the comparison result of the eighth performance index of the method of the present invention and other existing methods in acetylation modification. Detailed implementation manners
[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0090] Embodiment 1
[0091] As Figure 1 shown below, the present invention will be described in detail with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention provide a method for predicting species-specific acetylation sites, which mainly includes the following steps:
[0092] Step 1: Data preprocessing.
[0093] First, the lysine acetylation sites collected from various databases are used as positive example samples, and other lysine sites in the corresponding proteins that have not been reported to be acetylated are regarded as negative example samples. Secondly, to prevent the problem of over-optimization of performance caused by protein homology, the present invention uses the CD-HIT tool to remove proteins with similarity greater than 40% in the dataset. For each site, 15 amino acids upstream and downstream of it are intercepted in a sliding window manner to form an amino acid sequence with a length of 31. Then, the amino acid sequence is encoded into a 31*21 feature matrix using a unique encoding method. Finally, the dataset is divided into three parts: a training set, a validation set, and a test set in a ratio of 8:1:1.
[0094] Step 2: Construct a deep neural network.
[0095] As Figure 2 shown, in the embodiments of the present invention, the sequence feature extraction network is constructed in the manner as Figure 2 follows: The sequence feature extraction network consists of multiple cascaded convolutional modules. This module designs dense connections on the basis of the structure of the convolutional neural network, so that each convolutional layer in the module is connected to all the previous layers. For the i-th convolutional layer in the module, the corresponding output is where represents the splicing operation on the feature map. To avoid the excessive burden on the model caused by the increase in the number of network layers, a bottleneck layer is added after each module to reduce the number of channels of the features. The bottleneck layer uses a 1*1 convolution to achieve feature compression:
[0096] h B = α B (W B h M + b B ) (1)
[0097] where h M is the output of the module, W B , b B and α B represent the weight, bias term, and activation function of the convolutional kernel in the bottleneck layer, respectively.
[0098] The sample is converted into a feature vector after passing through the sequence feature extraction network and then input into the classifier for classification. The classifier consists of two layers of fully connected neural networks. The first layer contains 6160 neurons, the second layer contains 2 neurons, and the activation function is softmax to output the prediction score.
[0099] The model also includes a domain category discriminator, whose input is composed of the concatenation of the features of the sample pair after passing through the sequence feature extraction network. The domain category discriminator consists of three layers of convolutional neural networks. The first layer contains 12320 neurons, the second layer contains 64 neurons, and the third layer contains 4 neurons, and outputs the classification result.
[0100] Step 3: Construction of training samples.
[0101] To implement adversarial training, the present invention constructs four groups of sample pairs: 1) Both samples come from the human acetylation dataset and have the same label (p1); 2) The two samples come from the human and other species acetylation datasets respectively but have the same label (p2);
[0102] 3) Both samples come from the human acetylation dataset but have different labels (p3); 4) The two samples come from the human and other species acetylation datasets respectively and have different labels (p4).
[0103] Step 4: Train the network model using the semantic adversarial training strategy.
[0104] The sequence feature extraction network and the domain category discriminator are trained separately in an alternating manner. First, use the human acetylation data to pre-train the sequence feature extraction network and the classifier. Then use the constructed four groups of sample pairs to train the domain category discriminator so that it can effectively distinguish which category the input sample pair belongs to. The domain category discriminator is optimized using the standard cross-entropy loss function:
[0105]
[0106] where E[·] represents the statistical expectation, represents the label of the sample pair p i , and φ represents the concatenation of the features of the two samples in the sample pair p i . When training the domain category discriminator, the parameters of the feature extraction network and the classifier are frozen. Next, train the feature extraction network so that the domain category discriminator cannot effectively distinguish the category of the data pair, and the corresponding loss function is as follows:
[0107]
[0108] Through such training, the domain category discriminator can be confused, thereby achieving semantic alignment of acetylation between humans and other species and effectively reducing the distribution difference between different species. In addition, to ensure the accuracy of PTM site prediction, this project further combines the classification loss based on formula (3):
[0109]
[0110] where λ represents the weight coefficient for balancing the classification loss and the domain confusion loss, X s and X t represent human and other species samples respectively, g s and g t represent the sequence feature extraction networks for humans and other species respectively, h s and h t represent the classifiers for humans and other species respectively, and l represents the standard binary cross-entropy loss function.
[0111] By alternately training the sequence feature extraction network and the domain category discriminator, the model finally learns a discriminative feature space with domain invariance. In this feature space, human acetylation knowledge can be effectively transferred, and at the same time, high-precision acetylation site prediction for multiple species can be achieved.
[0112] Step 5, Performance evaluation
[0113] In the embodiments of the present invention, the sample pairs in the training set are fed into the constructed deep neural network for iterative training. Every time the data in the training set is traversed, an evaluation index is calculated on the validation set, and the network with the best performance is retained.
[0114] Exemplarily, the evaluation index is defined as follows:
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] where TP represents the total number of acetylation sites correctly predicted, TN represents the total number of non-acetylation sites correctly predicted, FP represents the total number of samples mispredicted as acetylation sites, and FN represents the total number of acetylation sites not predicted.
[0121] Example 2
[0122] 1) Data collection and collation: First, collect and collate post-translational modification sites such as ubiquitination and acetylation of multiple species from databases such as HPRD, PLMD, dbPTM, PhosphoSitePlus, and mUbiSida, and use the sliding window technique to extract the local sequences of post-translational modification sites to construct a dataset. Then, one-hot encode the local protein sequences in the dataset. Each amino acid in the sequence is encoded into a 21-dimensional feature vector, where the position corresponding to the amino acid is 1 and the other positions are 0.
[0123] 2) Network structure design: The sequence feature extraction network adds dense connections on the basis of the convolutional neural network, enabling each convolutional layer to receive the feature maps of all previous convolutional layers, thereby enhancing the propagation of sequence features in the model and extracting more effective sequence features. The sequence feature extraction network consists of multiple dense connection modules, and each dense connection module is connected to a bottleneck layer composed of 1*1 convolutions to reduce the number of channels of the features and reduce the model burden. The classifier consists of two fully connected layers and is used to output the sample category. The domain category discriminator consists of two fully connected layers and is used to judge the domain category to which the input data belongs.
[0124] 3) Semantic adversarial training strategy: Pair human training samples with training samples of other species to alleviate the problem of insufficient data of other species. The pairing is divided into 4 cases: both samples come from humans and have the same label (p1), the two samples come from humans and other species respectively but have the same label (p2), both samples come from humans but have different labels (p3), and the two samples come from humans and other species respectively and have different labels (p4).
[0125] After pairing, in the first step, use human post-translational modification samples to pre-train the sequence feature extraction network and the classifier. In the second step, train the domain category discriminator to distinguish which of the above four groups of pairings the input sample pair belongs to. At this time, freeze the parameters of the sequence feature extraction network and the classifier, and only update the parameters of the domain category discriminator. In the third step, freeze the parameters of the domain category discriminator, and train the feature extraction network and the classifier so that the domain category discriminator cannot distinguish between group 1 and group 2 and between group 3 and group 4. Then, cycle through the training of the second and third steps until the loss function converges.
[0126] 4) Performance evaluation: Randomly divide 10% of the data from the collected post-translational modification sites as an independent test set, and evaluate the prediction results using indicators such as accuracy, recall, and precision.
[0127] As Figure 2As shown, in this embodiment, the accuracy and robustness of the present application are evaluated. The embodiments of the present invention calculate the evaluation indexes on the test set and compare them with other existing acetylation site prediction methods. It can be seen from the results that DeepAce-Pred of the present invention is superior to other existing acetylation modification site prediction methods PAIL and CapsNet-PTM. Compared with the CapsNet-PTM, the model with the highest performance among the existing methods, the AUC value of the present invention for acetylation site prediction in Mus musculus is increased from 0.695 to 0.759, with a relative increase of more than 6.4%. In addition, on yeast, the present invention also achieves an increase in the AUC value of 5.1 and 24.7 compared with PAIL and CapsNet-PTM respectively.
[0128] As shown in Figure 4, in addition to the AUC value, we also evaluate various performance indexes and results such as the ACC value, Pre value, F1 value, etc. of various methods. It can be seen from Figure 4 that the present invention achieves higher prediction performance compared with the existing acetylation site methods on multiple species. For example, on Mus musculus, the Pre value of DeepAce-Pred of the present invention is 0.795, which is increased by 13.1% and 27.0% compared with CapsNet-PTM and PAIL respectively. The above results show that the semantic adversarial strategy proposed by the present invention can effectively reduce the difference in the acetylation data distribution between humans and other species, which is beneficial to the transfer of knowledge, thus significantly improving the prediction performance of the model on other species.
[0129] In summary, the sequence feature extraction network in the present invention adds dense connections on the basis of the convolutional neural network to enhance the propagation of sequence features in the model and extract more effective sequence features. The present invention realizes domain adaptation based on semantic adversarial, effectively transfers human post-translational modification knowledge to assist the site prediction of other species, and realizes high-precision species-specific post-translational modification site prediction.
[0130] The sequence feature extraction network in the present invention is composed of multiple dense connection modules, and each dense connection module is connected to a bottleneck layer composed of 1*1 convolution to reduce the number of channels of features and reduce the model burden.
[0131] The semantic adversarial training strategy proposed in the model training stage: pair the human training samples with the training samples of other species to alleviate the problem of insufficient data of other species. It can effectively reduce the distribution difference of human and other species' post-translational modification data, enable the model to effectively transfer human post-translational modification knowledge to assist the site prediction of other species, thus effectively avoiding the overfitting problem of the model, making the model more suitable for species with small data volume, and improving the application scope of the model.
[0132] Compared with the existing methods for predicting post-translational modification sites, the species-specific prediction scheme proposed by the present invention can establish prediction models for different species respectively, improving the prediction accuracy of the models for different species. The present invention provides a new idea for the prediction of species-specific post-translational modification sites, and the prediction results can provide a basis for the study of protein molecular mechanisms, which has guiding significance for revealing the functions of proteins.
[0133] The present invention alternately trains a sequence feature extraction network and a domain category discriminator, and the model finally learns a discriminant feature space with domain invariance. In this feature space, the knowledge of human acetylation can be effectively transferred, and at the same time, high-precision acetylation site prediction for multiple species can be achieved. The present invention solves the technical problems of low efficiency, high cost, time-consuming, overfitting and low prediction accuracy existing in the prior art.
[0134] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for predicting post - translational modification sites of species - specific proteins, characterized in that, The method includes: S1. Collect and organize ubiquitination and acetylation post-translational modification sites of different species from a pre-set database, and use the sliding window technique to extract the local sequences of the post-translational modification sites to construct a data set, and one-hot encode the protein local sequences in the data set to obtain post-translational modification sites and training samples of different species; S2. Add dense connections on the basis of a convolutional neural network so that each convolutional layer receives the feature maps of all the previous convolutional layers, and accordingly construct a sequence feature extraction network and set a classifier and a domain category discriminator; wherein, step S2 includes: S21. Set dense connections for the convolutional modules of the convolutional neural network, and enable each convolutional layer to receive the feature maps of all the previous convolutional layers according to the following logic: , In the formula, represents the operation of splicing feature maps, and W i is the convolution kernel weight; S22. Add a bottleneck layer after each convolutional module according to the following logic to reduce the number of channels of the features: h B = α B (W B h M + b B ), where h M is the output of the convolution module, W B , b B and α B respectively represent the weight, bias term and activation function of the convolution kernel in the bottleneck layer; S23. Set a classifier and the domain category discriminator; S3. Use a semantic adversarial strategy to pair-process the training samples of different species to obtain a sample pairing group of human post-translational modification samples and post-translational modification samples of other species; S4. Use the human post-translational modification samples to train the sequence feature extraction network and the classifier, train the domain category discriminator according to the sample pairing group to distinguish the group information of the input sample pairs, and use an alternating method to train the sequence feature extraction network and the domain category discriminator respectively until the loss function converges, so that the sequence feature extraction network, the classifier and the domain category discriminator learn a domain-invariant discriminative feature space; S5. Divide an independent test set from the post-translational modification sites to evaluate the performance of the sequence feature extraction network, the classifier and the domain category discriminator.
2. The method for predicting post - translational modification sites of species - specific proteins according to claim 1, characterized in that, The step S1 includes: S11. Collect lysine acetylation sites from the HPRD, PLMD, dbPTM, PhosphoSitePlus and mUbiSida databases as positive example samples; S12. Regard other lysine sites in the corresponding proteins that have not been reported to undergo acetylation as negative example samples; S13. Use the CD-HIT tool to remove proteins with similarity greater than 40% in the data set; S14. Use a sliding window method to intercept 15 amino acids upstream and downstream of each post-translational modification site to construct an amino acid sequence with a length of 31; S15. One-hot encode the amino acid sequence into a feature matrix; S16. Divide the data set into a training set, a validation set and a test set according to a ratio of 8:1:
1.
3. The method for predicting post - translational modification sites of species - specific proteins according to claim 1, characterized in that, The step S3 includes: S31. Construct a sample pairing group p1 from two training samples of different species that both come from the human acetylation data set and have the same label; S32. Construct a sample pairing group p2 from two training samples of different species that come from the human and other species acetylation data sets respectively but have the same label; S33. Construct a sample pairing group p3 from two training samples of different species that both come from the human acetylation data set but have different labels; S34. Construct the sample pairing group p4 with the differential species training samples from two acetylation data sets of humans and other species with different labels.
4. The method for predicting post - translational modification sites of species - specific proteins according to claim 1, characterized in that, The step S4 includes: S41. Freeze the parameters of the sequence feature extraction network and the classifier, and update the parameters of the domain category discriminator; S42. Freeze the parameters of the domain category discriminator, train the feature extraction network and the classifier, so that the domain category discriminator does not distinguish the sample pairing group; S43. Continuously train the sequence feature extraction network, the classifier and the domain category discriminator with the steps S41 and S42 until the loss function converges to align the semantics of acetylation of each species.
5. The method for predicting post - translational modification sites of species - specific proteins according to claim 4, characterized in that, In the step S43, the standard cross-entropy loss function is used to optimize the domain category discriminator: , where E[·] represents statistical expectation, denotes the sample pair p i 's label, and φ represents concatenating the features of the two samples in the sample pair p i .
6. The method for predicting post-translational modification sites of species-specific proteins according to claim 4, wherein In the step S43, the following loss function is used to optimize the feature extraction network:
7. The method for predicting post-translational modification sites of species-specific proteins according to claim 4, wherein In the step S43, the following logical combination classification loss is used to ensure the prediction accuracy of PTM sites: where λ represents the weight coefficient for balancing the classification loss and the domain confusion loss, X s and X t represent human and other species samples respectively, g s and g t represent the sequence feature extraction networks for humans and other species respectively, h s and h t represent the classifiers for humans and other species respectively, and l represents the standard binary cross-entropy loss function.
8. The method for predicting post-translational modification sites of species-specific proteins according to claim 1, wherein The step S5 includes: S51. Send the sample pairs in the training set into the constructed deep neural network for iterative training; S52. Loop through the training set; S53. Every time the data in the training set is traversed, calculate the evaluation index on the validation set with the following logic, and accordingly retain the applicable sequence feature extraction network, the classifier and the domain category discriminator: , Among them, TP represents the total number of acetylated sites correctly predicted, TN represents the total number of non-acetylated sites correctly predicted, FP represents the total number of samples mispredicted as acetylated sites, and FN represents the total number of acetylated sites not predicted.
9. A system for predicting post-translational modification sites of species-specific proteins, which is used to execute the method for predicting post-translational modification sites of species-specific proteins according to any one of the preceding claims 1 to 8, wherein The system includes: A data preprocessing module for collecting and sorting out the ubiquitination and acetylation post-translational modification sites of different species from a preset database, using the sliding window technique to extract the local sequences of the post-translational modification sites to construct a data set, and one-hot encoding the protein local sequences in the data set to obtain post-translational modification sites and differential species training samples; A network construction module for adding dense connections on the basis of a convolutional neural network so that each convolutional layer receives the feature maps of all the previous convolutional layers, and accordingly constructs a sequence feature extraction network and sets a classifier and a domain category discriminator. The network construction module is connected to the data preprocessing module; A training sample construction module for pairing and processing the differential species training samples using a semantic adversarial strategy to obtain a sample pairing group of human post-translational modification samples and other species post-translational modification samples. The training sample construction module is connected to the data preprocessing module; An adversarial training module for training the sequence feature extraction network and the classifier using the post-modification samples translated by humans, training the domain category discriminator according to the post-modification samples translated by other species to distinguish the group information of the input sample pairs, and respectively training the sequence feature extraction network and the domain category discriminator in an alternating manner until the loss function converges, so that the sequence feature extraction network, the classifier and the domain category discriminator learn a domain-invariant discriminative feature space. The adversarial training module is connected to the training sample construction module and the network construction module; A performance evaluation module for evaluating the performance of the sequence feature extraction network, the classifier and the domain category discriminator according to the post-modification sites. The performance evaluation module is connected to the adversarial training module.
Citation Information
Patent Citations
Protein coding method and protein post-translational modification site prediction method and system
CN110033822A
Genome unit point variation pathogenicity prediction method and system and storage medium
CN110245685A