Method and system for predicting classification and cleavage sites of target peptide
By extracting evolutionary and structural information of target peptide sequences using the DeepMaT model, and combining it with multilayer perceptron and conditional random field, we have achieved efficient classification of multiple target peptides and prediction of cleavage sites. This solves the problem of insufficient accuracy of existing tools and improves the accuracy of target peptide prediction and the generalization ability of the model.
Patent Information
- Application Number
- CN202511176670.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing AI-based targeted peptide prediction tools cannot efficiently classify multiple targeted peptides and predict cleavage sites simultaneously, and their accuracy needs improvement.
The DeepMaT model is used to extract evolutionary and three-dimensional structural information of the target peptide sequence through the ISM module. The global dependencies and local details of the sequence are captured by combining the Mamba2 module, the Attention module and the feedforward layer. The MLP multilayer perceptron is used for classification and prediction, and the CRF conditional random field is used to predict the cleavage site.
It significantly improves the accuracy of target peptide classification and cleavage site prediction, especially for rare categories, and enhances the model's generalization ability and robustness.
Smart Images

Figure CN120932745A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a method and system for predicting target peptide classification and cleavage sites. Background Technology
[0002] Proteins are synthesized within cells and guided to specific regions via specific sorting signals—a process known as subcellular localization, crucial for protein function. Common specific sorting signals include N-terminal targeting peptides, such as signal peptides, mitochondrial transport peptides, chloroplast transport peptides, and thylakoid transport peptides. After guiding the synthesized protein to the specific region, the N-terminal targeting peptide is recognized and cleaved by specific enzymes. Identifying these targeting peptides and their cleavage sites is essential for understanding protein transport mechanisms. Traditional experimental methods, such as fluorescence fusion localization and Edman degradation assays, while capable of determining targeting peptides in biological experiments, have several drawbacks, including being time-consuming, costly, and requiring strict purity control of the experimental samples. To reduce costs and improve efficiency, predictive tools based on artificial intelligence are increasingly being adopted.
[0003] Currently, there are many AI-based prediction tools developed for specific target peptides, such as MitoFates, ChloroP, and SignalP, but they only predict a single category and cannot perform multi-category predictions. There are also tools that can predict multiple target peptides, such as TargetP, but the accuracy of the model in predicting classification and cleavage sites needs to be improved. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a method and system for predicting target peptide classification and cleavage sites. The method and system can simultaneously predict target peptide classification and cleavage sites with high accuracy.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] A method for predicting target peptide classification and cleavage sites includes the following steps:
[0007] Obtain a dataset that includes different types of targeted peptide sequences and non-targeted peptide sequences, wherein the targeted peptide sequences include targeted peptide sequences with cleavage site annotations and targeted peptide sequences without cleavage site annotations;
[0008] The DeepMaT model is trained by inputting sequences from the dataset. The training includes: extracting evolutionary information from the sequences through the ISM module and implicitly incorporating three-dimensional structural information to obtain peptide sequence features; inputting the peptide sequence features into the Mamba2 module, the Attention module, and the feedforward layer connected in sequence to capture and fuse long-distance global dependencies and local details in the sequences to obtain fused features; inputting the fused features into the MLP multilayer perceptron to predict the classification of target peptides, and simultaneously inputting the fused features into the CRF conditional random field to predict cleavage sites.
[0009] Model evaluation.
[0010] Preferably, the processing procedure of the Mamba2 module includes the following steps:
[0011] The peptide sequence characteristics H output by the ISM module ISM After processing by the Mamba2 layer, the first feature is obtained;
[0012] After the first feature is processed by the DROPOUT layer, it is compared with the peptide sequence feature H. ISM The residuals are connected and input into the normalization layer to obtain the feature matrix M.
[0013] Preferably, the processing procedure of the Attention module includes the following steps:
[0014] The feature matrix M output by the Mamba2 module is processed by a multi-head self-attention layer to obtain the second feature;
[0015] The second feature is processed by the DROPOUT layer, connected to the residual of the feature matrix M, and then input into the normalization layer to obtain the feature matrix A.
[0016] Preferably, the MLP multilayer perceptron employs a cross-entropy loss function.
[0017] Preferably, the loss function Loss of the CRF (Conditional Random Field) is as follows:
[0018]
[0019] Loss = w1 * Loss CRF +w2*Loss CE
[0020] Where T is the sequence length of the sample, N is the number of samples in a batch, K is the number of classes, and p i These are the true labels of the samples. It is the predicted probability value, p. i k It is the true label of the k-th class of the sample. M is the predicted probability value of the k-th class. t For a set of real labels at position t, Loss CE For classification, the cross-entropy loss and negative log-likelihood loss - log(P(y|h)) are denoted as Loss. CRF w1 and w2 are the weights of the CRF loss and the classification cross-entropy loss, respectively.
[0021] Preferably, the model evaluation includes the following steps:
[0022] Evaluation metrics were determined, and the DeepMaT model was evaluated using five-fold cross-validation and an independent test set. The independent test set included a set of target peptide sequences with cleavage site annotations and a set of target peptide sequences without cleavage site annotations.
[0023] The evaluation metrics include precision, recall, F1 score, and Matthews correlation coefficient.
[0024] Preferably, the targeting peptide sequence includes a signal peptide, a mitochondrial transport peptide, a chloroplast transport peptide, and an endoderm transport peptide.
[0025] Based on the above, the present invention also discloses a system for predicting target peptide classification and cleavage sites, comprising:
[0026] The acquisition module is used to acquire a dataset, which includes different types of targeted peptide sequences and non-targeted peptide sequences. The targeted peptide sequences include targeted peptide sequences with cleavage site annotations and targeted peptide sequences without cleavage site annotations.
[0027] The training module is used to input sequences from the dataset into the DeepMaT model for training. The training includes: extracting evolutionary information of the sequences through the ISM module and implicitly incorporating three-dimensional structural information to obtain peptide sequence features; inputting the peptide sequence features into the Mamba2 module, the Attention module, and the feedforward layer connected in sequence to capture and fuse long-distance global dependencies and local details in the sequences to obtain fused features; inputting the fused features into the MLP multilayer perceptron to predict the target peptide classification, and simultaneously inputting the fused features into the CRF conditional random field to predict cleavage sites.
[0028] The evaluation module is used for model evaluation.
[0029] Based on the above technical solution, the beneficial effects of this invention are as follows: This invention discloses a method for predicting target peptide classification and cleavage sites. The method involves acquiring a dataset, which includes different types of target peptide sequences and non-target peptide sequences. The target peptide sequences include those with cleavage site annotations and those without. The sequences in the dataset are input into a DeepMaT model for training. The training includes: extracting evolutionary information from the sequences through an ISM module and implicitly incorporating three-dimensional structural information to obtain peptide sequence features; inputting the peptide sequence features into a sequentially connected Mamba2 module, Attention module, and feedforward layer to capture and fuse long-distance global dependencies and local details in the sequences to obtain fused features; inputting the fused features into an MLP (Multilayer Perceptron) to predict target peptide classification, and simultaneously inputting the fused features into a CRF (Conditional Random Field) to predict cleavage sites; and finally, model evaluation. By fully extracting peptide sequence features using the ISM model and utilizing the global learning characteristics of Mamba2 combined with the local learning characteristics of the multi-head attention mechanism to model sequence features, the prediction accuracy of rare categories and cleavage is significantly improved. Attached Figure Description
[0030] Figure 1 This is a schematic flowchart of a method for predicting target peptide classification and cleavage sites in one embodiment;
[0031] Figure 2 This is a model framework diagram of the DeepMaT model in one embodiment;
[0032] Figure 3 This is an example of the performance of the DeepMaT model in the task of classifying and predicting cleavage sites of targeted peptides. Figure 3 A to 3E are performance comparison charts in the target peptide classification and prediction task; Figure 3 F is a comparison of performance in the cleavage site prediction task;
[0033] Figure 4 This is a comparison chart of ablation experiment results in one embodiment, wherein... Figure 4 A is a comparison of the performance of ablation experiments on signal peptides; Figure 4 B is a comparison of the performance of the ablation experiment on mitochondrial transport peptides; Figure 4 C is a comparison of the performance of ablation experiments on chloroplast transport peptides; Figure 4 D is a comparison of the performance of ablation experiments on internal vesicle transport peptides;
[0034] Figure 5 This is a schematic diagram comparing signal peptide prediction results in one embodiment, wherein, Figure 5A is a schematic diagram comparing the classification prediction results of the DeepMaT model with those of the SignalP 6.0, PEFT-SP, and USPNet target peptides. Figure 5 B is a schematic diagram comparing the prediction results of the cleavage site of the targeted peptide by the DeepMaT model with those of the SignalP 6.0, PEFT-SP, and USPNet models. Figure 5 C represents the overall performance of the DeepMaT model on the SignalP 6.0 dataset;
[0035] Figure 6 This is the test result of an independent test set in one embodiment, where, Figure 6 A compares the classification prediction results of five DeepMaT models and TargetP 2.0 on the independent test set 1; Figure 6 B is a comparison of the classification prediction results of the five DeepMaT models and TargetP2.0 on the independent test set 2; Figure 6 C is a comparison chart of the classification prediction performance of DeepMaT and TargetP 2.0 on independent test set 1; Figure 6 D is a comparison of the performance of DeepMaT and TargetP 2.0 on the independent test set 1 for the cut-off site prediction task; Figure 6 E is a comparison chart of the classification prediction performance of DeepMaT and TargetP 2.0 on independent test set 2. Detailed Implementation
[0036] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0037] like Figures 1 to 6 As shown, this embodiment provides a method for predicting target peptide classification and cleavage sites, including the following steps:
[0038] Step 1: Obtain the dataset, which includes different types of targeted peptide sequences and non-targeted peptide sequences. The targeted peptide sequences include targeted peptide sequences with cleavage site annotations and targeted peptide sequences without cleavage site annotations.
[0039] In this embodiment, publicly available datasets from TargetP 2.0 and SignalP 6.0 were used, and two independent test sets were collected. TargetP includes signal peptides (SP), mitochondrial transport peptides (MT), chloroplast transport peptides (CH), thylakoid transport peptides (TH), and non-target peptide sequences, with clearly defined cleavage site annotations; SignalP 6.0 provides multiple signal peptide subtypes. Reviewed target peptide sequences were retrieved from the UniPort database, and these sequences were categorized into those with clearly defined cleavage sites and those without. The publicly available datasets for both TargetP 2.0 and SignalP 6.0 were divided into five subsets to ensure that every sample was included in the evaluation process. In each iteration, four subsets were used for training, and the remaining one was used for testing.
[0040] Table 1 TargetP2.0 dataset
[0041]
[0042] Table 2 SignalP 6.0 dataset
[0043]
[0044] Sequences were divided into two groups based on the presence or absence of cleavage site annotations: sequences with clearly defined cleavage sites were assigned to independent test set 1, while sequences lacking such annotations were assigned to independent test set 2. Sample sizes are detailed in Table 3. The performance of the DeepMaT model was compared with TargetP 2.0 on these two independent test sets to evaluate the generalization ability and robustness of the DeepMaT model to newly collected data.
[0045] Table 3 Independent Test Set
[0046]
[0047] Step 2: Input the sequences in the dataset into the DeepMaT model for training. The training includes: extracting evolutionary information of the sequences through the ISM module and implicitly incorporating three-dimensional structural information to obtain peptide sequence features; inputting the peptide sequence features into the Mamba2 module, Attention module, and feedforward layer connected in sequence to capture and fuse long-distance global dependencies and local details in the sequence to obtain fused features; inputting the fused features into the MLP multilayer perceptron to predict the target peptide classification, and simultaneously inputting the fused features into the CRF conditional random field to predict cleavage sites.
[0048] In this embodiment, the DeepMaT model includes a feature extraction module, a feature learning module, and a prediction module, which are described in detail below:
[0049] 1) Feature Extraction Module: The ISM (Implicit Structure Model) module is used, a self-supervised protein language model that implicitly incorporates the three-dimensional structural information of proteins into the model through a self-supervised learning strategy, resulting in embeddings that conform to evolutionary laws and possess structural information. In the initial stage of feature extraction, the input sequence needs to be processed by its own specific tokenizer, which converts the amino acid sequence into a specific encoding format. The ISM model follows the Transformer architecture of ESM2, pre-trained based on a language model, and can fully understand sequence clock evolution and covariation relationships. Unlike traditional sequence models, ISM uses a "structure tuning" strategy to distill protein structural information into the model. This strategy relies on structural tokens generated by the structure encoder. Developers first need to extract the structural features of each residue, and then cluster these discrete vectors into a token set using K-means clustering. During the pre-training stage, the ISM model needs to predict the tokens corresponding to residues, enabling the model to have a further perception of protein structure without introducing explicit structural input. During the inference phase of the model, the ISM model accepts sequence data transformed by the word segmenter as input and outputs a feature matrix of shape [L×D] (L is the sequence length and D is the feature dimension).
[0050] amino acid sequence of the sample (X) seq The data is converted into ISM's input format using the ISM tokenizer. This step involves converting amino acid characters into specific numeric codes. ISM takes the encoded data as input and outputs a feature matrix of size L×1280, denoted as H. ISM , where L represents the sequence length and 1280 is the embedding dimension. This matrix contains the local semantic information of each amino acid residue and the contextual features of its three-dimensional structural environment, forming a structurally rich sequence representation.
[0051] H ISM =ISM(X) seq )
[0052] 2) Feature Learning Module: A sequential three-layer architecture is adopted, consisting of a Mamba2 module, an Attention module (MHA), and a feedforward layer. Both the Mamba2 and Attention modules employ residual structures, with each layer including LayerNorm, Dropout, and residual connections to enhance network stability and generalization ability. In the Mamba2 module, the Mamba2 layer primarily learns the feature data. As the first feature learning layer, it prioritizes learning the global dependencies of features. The Mamba2 layer is a state-space model structure that efficiently captures long-range global dependencies in sequences. The output of the Mamba2 module is related to H... ISM A feature array of uniform dimension, denoted by M:
[0053] M = LayerNorm(H) ISM +Dropout(Mamba2(H ISM The Attention module, as the second sub-module of the feature learning module, employs a multi-head attention mechanism, mapping features to multiple attention subspaces. Compared to Mamba2, the multi-head attention mechanism is more suitable for learning local detail features, thus the two complement each other in feature learning. Following this is the feedforward layer, which reduces the risk of overfitting, increases model complexity, and maps data to a more appropriate representation space.
[0054] The feature matrix M is input to the multi-head self-attention layer to further explore the details of interactions between local sequence segments. In the attention layer, the input features are processed by the corresponding weight matrix W. q W k and W v The query is projected into three distinct vectors: query (Q), key (K), and value (V). The multi-head attention mechanism calculates attention scores based on the similarity of QKT, generates a weight distribution after softmax normalization, and then sums the weights of the V vector accordingly.
[0055] Q = W q X; K = W k X; V = W v X
[0056]
[0057] Where X represents the input attention data, T represents the matrix transpose operation, and d represents the vector dimension.
[0058] The output of the Attention module is a refined representation of the input features, obtained by multiple parallel attention heads independently learning and weighting different aspects of the data. To enhance model stability and prevent overfitting, a layer normalization and Dropout layer are added after the multi-head self-attention layer. The output of the Attention module is denoted as A, maintaining the same dimension as the output of the Mamba2 module, which is L×1280.
[0059] A=LayerNorm(M+Dropout(MultiheadAttention(M)))
[0060] 3) Prediction Module: This module includes classification prediction and cleavage site prediction tasks. The classification prediction task determines whether a given sequence contains a target peptide and classifies it into a specific category (signal peptide, mitochondrial transport peptide, chloroplast transport peptide, endothelial transport peptide, or non-target peptide). In this task, a Multilayer Perceptron (MLP) is used. The output data from the feature learning module is input into this module, and the final output is the classification prediction probability of the sequence. The MLP is a multilayer classifier, including linear layers, activation layers, LayerNorm layers, and Dropout layers. The MLP operates on the first label in the sequence output by the multi-head self-attention layer (a0 ∈ R1280). Specifically, it maps the corresponding 1280-dimensional feature vector to an n-dimensional output, where n represents the number of peptide classes. Then, a softmax function is applied to the generated vector to calculate the classification probability.
[0061] The DeepMaT model performs both peptide classification and cleavage site prediction, each requiring its own objective during optimization. For the classification task, a cross-entropy loss function is used, which quantifies the difference between the predicted class probability and the ground truth label, thereby improving the model's ability to distinguish between different peptide types.
[0062] In the cleavage site prediction task, a Conditional Random Field (CRF) is introduced. This CRF can predict the label for each residue position. Here, we set the transition matrix for CRF initialization and set some unreasonable transitions as invalid. During training, the probability of the true label is optimized by maximum likelihood estimation. Then, during inference, the Viterbi algorithm is used to find the label sequence with the highest score, thus achieving global consistency modeling among labels.
[0063] First, let A(a1,a2,...,a...) T-1 The data is transformed into the dimensionality required by CRF through a linear layer. The number of labels is set to 2, so every 1280 dimensions of data is converted to 2 dimensions to adapt to CRF, but the sequence length remains L. The CRF module then converts the transformed state sequence y = y1...y2 into a 2D format.t It processes the data. It assigns a corresponding hidden state sequence h = h1...h1 to each position in the input. t This effectively simulates the structural relationships in the output sequence:
[0064]
[0065] Where Z(h) is the normalization constant for modeling; This is the learning transfer matrix of the CRF, with a shape of C*C, where C is the number of modeling labels; ψ is the learnable linear transformation from the hidden state h to the label; T is the sequence length; and t is the product index. for of y t line y t+1 Column data:
[0066] ψ(h t ) = W ψ h t +b ψ
[0067] Among them, h t W represents the hidden state sequence corresponding to the product index t. ψ Let b be the weight matrix. ψ This is the bias vector.
[0068] The negative log-likelihood loss -log(P(y|h)) was used with a Conditional Random Field (CRF), denoted as Loss. CRF It can effectively simulate label dependencies between sequences, improving the accuracy of cleavage sites. The final training objective is a weighted sum of two loss functions, in which adjustable hyperparameters are introduced to balance the performance between classification and sequence labeling tasks. The formula is shown below:
[0069]
[0070]
[0071] Loss = w1 * Loss CRF +w2*Loss CE
[0072] Where T is the sequence length of the sample, t is the summation index, N is the number of samples in a batch, K is the number of classes, and p i These are the true labels of the samples. It is the predicted probability value, p. i k It is the true label of the k-th class of the sample. M is the predicted probability value of the k-th class. t For a set of real labels at position t, LossCE is the cross-entropy loss for classification, and w1 and w2 are the weights of the CRF loss and the classification cross-entropy loss, respectively.
[0073] Step 3, Model Evaluation.
[0074] In this embodiment, the performance of the DeepMaT model in predicting signal peptides was validated through five-fold cross-training and independent test sets. Precision, recall, F1 score, and Matthews correlation coefficient (MCC) were used for quantitative analysis. Ablation experiments were also conducted to verify the contributions of the Mamba2 layer and multi-head attention layer in the feature learning module to the model's learning ability. For the evaluation on the SignalP 6.0 dataset, the Matthews correlation coefficient (MCC) was divided into MCC1 and MCC2. MCC1 treats a certain type of signal peptide as a positive sample and non-signal peptides as negative samples; MCC2 treats a certain type of signal peptide as a positive sample, but other types of signal peptides and non-signal peptides as negative samples. MCC1 reflects the model's classification ability for a single class, while MCC2 reflects the single-class performance in multi-class tasks. Due to the similarity between different classes of signal peptides, MCC2 is more difficult to calculate.
[0075] The experiment provided in this embodiment:
[0076] 1) Model Performance
[0077] To evaluate the performance of the DeepMaT model in the classification and cleavage site prediction tasks of targeted peptides, both DeepMaT and the TargetP 2.0 model were compared on the TargetP 2.0 dataset. In the classification tasks of signal peptides, mitochondrial transport peptides, chloroplast transport peptides, and endothelial transport peptides, the DeepMaT model consistently outperformed the TargetP 2.0 model in key evaluation metrics such as accuracy, F1 score, and recall. The attention mechanism used by the TargetP 2.0 model has limited ability to model long-range dependencies in sequences, while the DeepMaT model utilizes the Mamba2 architecture—a state-space model (SSM)—to enhance global sequence modeling and extract more information features from longer input sequences. Figure 3 A to Figure 3 As shown in Figure E, the DeepMaT model outperformed TargetP 2.0 on most metrics. Although chloroplast and endosome transport peptides share some common sequence features, the DeepMaT model improved the classification performance of endosome peptides without affecting the accuracy of chloroplast peptide predictions. Notably, the recall rate of mitochondrial transport peptides improved from 0.85 to 0.90, an improvement of 5 percentage points. Figure 3E further illustrates the improved ability of the DeepMaT model to distinguish between target and non-target peptides. In addition to its superior classification performance, the model also significantly improves the accuracy of predicting the cleavage sites of target peptides; see [link to relevant documentation]. Figure 3 F. Notably, the prediction accuracy for cleavage sites of endosome transport peptides improved from 0.60 in TargetP 2.0 to 0.867, representing a significant improvement. In experiments, this improved prediction accuracy did not sacrifice classification performance, highlighting the robustness and effectiveness of the DeepMaT model in multi-task learning.
[0078] 2) Ablation experiment
[0079] We conducted reduction experiments on the feature learning module of the DeepMaT model by retraining a reduced version of the model on the TargetP 2.0 dataset. Specifically, we examined the effects of simultaneously removing the Mamba2 module, the multi-head attention module, and both modules. Figure 4 As shown, removing either the Mamba2 module or the multi-head attention module alone leads to a significant performance drop, especially for thylakoid peptides, which are underrepresented in the dataset. Performance also declines to varying degrees for other peptide categories. Interestingly, removing both components simultaneously results in a partial performance recovery, but still falls short of the complete DeepMaT model. These results indicate that the Mamba2 module and the multi-head attention module have complementary advantages in capturing global and local dependencies, enabling the DeepMaT model to learn richer and more informative features.
[0080] 3) Compare with state-of-the-art (SOTA) models on the SignalP 6.0 dataset to verify the signal peptide recognition performance.
[0081] In addition to evaluating the performance on targeted peptides, the generalization ability of the DeepMaT model in signal peptide prediction was also assessed using the SignalP 6.0 dataset. To this end, the DeepMaT model was compared with three established models: SignalP 6.0, PEFT-SP, and USPNet, all of which can predict signal peptides and their cleavage sites. Results for these baseline models were derived from supplementary materials in their respective publications. For fair comparison, the DeepMaT model was retrained using the same dataset split used in SignalP 6.0. This dataset contains 20,290 samples across four species and five signal peptide types, exhibiting significant class imbalance (as shown in Table 2). PEFT-SP is an ESM2-based fine-tuned model incorporating a LoRA mechanism, while USPNet combines a BiLSTM module and a protein language model for signal peptide prediction.
[0082] like Figure 5 As shown in (A, B), the DeepMaT model exhibits improved classification performance on certain signal peptide types—for example, on the archaea Sec / SPI class, the DeepMaT model outperforms other classification models by approximately 3% to 12%. While the DeepMaT model does not dominate across all signal peptide types, it still maintains competitive performance, likely due to the extreme class imbalance in the dataset. In cleavage site prediction, the DeepMaT model also outperforms other models on Sec / SPI-labeled SPs. Figure 5 (C) shows the overall performance of the model on the entire SignalP 6.0 dataset. The vertical axis groups the metrics by species (first) and signal peptide type (second), so peptides from the same species are visualized together. The horizontal axis represents the evaluation metrics. The DeepMaT model demonstrates strong performance, particularly in predicting cleavage sites for Sec / SPIII and TAT / SPII-tagged SPs, and shows a stronger ability to identify less common SP types.
[0083] 4) Evaluate model performance on independent test sets.
[0084] To further demonstrate the generalization ability of the DeepMaT model, two independent test sets were constructed—Independent Test Set 1 and Independent Test Set 2—each of which includes signal peptides, mitochondrial transport peptides, and chloroplast transport peptides (Table 3). Figure 6 In the table, "data 1" and "data 2" represent the original data types. The prediction results of the DeepMaT model were compared with those of the TargetP 2.0 model. The results from these independent test sets preliminarily demonstrate that the DeepMaT model exhibits superior generalization performance compared to the TargetP 2.0 model. Figure 6 The classification results on these datasets are presented to visually demonstrate the model's effectiveness. The DeepMaT model achieved higher classification accuracy than the TargetP 2.0 model in predicting mitochondrial and chloroplast transport peptides and cleavage sites, indicating a significant improvement. However, it was also observed that both the DeepMaT and TargetP 2.0 models tended to misclassify a considerable number of non-target peptides, leading to classification confusion.
[0085] For the cleavage site prediction results of independent test set 2 (see Table 4), the DeepMaT model can clearly generate up to five candidate predictions, thus providing insight into the probability distribution of potential cleavage sites. Among the prediction results, only two sequences, A0FKE6 and P0DO76, were predicted as identical by both the DeepMaT model and the TargetP 2.0 model, indicating that these sequences have the highest reference value for consistent cleavage site predictions.
[0086] Table 4. Prediction results of cleavage sites on independent test set 2
[0087]
[0088] It should be understood that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart above may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0089] The above are merely preferred embodiments of the present application and are not intended to limit the embodiments of the present application. For those skilled in the art, the embodiments of the present application can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present application should be included within the protection scope of the embodiments of the present application.
Claims
1. A method for predicting target peptide classification and cleavage sites, characterized in that, Includes the following steps: Obtain a dataset that includes different types of targeted peptide sequences and non-targeted peptide sequences, wherein the targeted peptide sequences include targeted peptide sequences with cleavage site annotations and targeted peptide sequences without cleavage site annotations; The sequences in the dataset are input into the DeepMaT model for training. The training includes: extracting evolutionary information of the sequences through the ISM module and implicitly incorporating three-dimensional structural information to obtain peptide sequence features. The peptide sequence features are input into the Mamba2 module, the Attention module, and the feedforward layer, which are connected in sequence, to capture and fuse long-distance global dependencies and local details in the sequence to obtain fused features. The fused features are then input into the MLP multilayer perceptron to predict the classification of the target peptide, and simultaneously input into the CRF conditional random field to predict the cleavage site. Model evaluation.
2. The method for predicting target peptide classification and cleavage sites according to claim 1, characterized in that, The processing procedure of the Mamba2 module includes the following steps: The peptide sequence features output by the ISM module are H ISM After processing by the Mamba2 layer, the first feature is obtained; After the first feature is processed by the DROPOUT layer, it is compared with the peptide sequence feature H. ISM The residuals are connected and input into the normalization layer to obtain the feature matrix M.
3. The method for predicting target peptide classification and cleavage sites according to claim 1, characterized in that, The processing procedure of the Attention module includes the following steps: The feature matrix M output by the Mamba2 module is processed by a multi-head self-attention layer to obtain the second feature; The second feature is processed by the DROPOUT layer, connected to the residual of the feature matrix M, and then input into the normalization layer to obtain the feature matrix A.
4. The method for predicting target peptide classification and cleavage sites according to claim 1, characterized in that, The MLP (Multilayer Perceptron) employs a cross-entropy loss function.
5. The method for predicting target peptide classification and cleavage sites according to claim 1, characterized in that, The loss function Loss of the CRF (Conditional Random Field) is shown in the following formula: Loss=w1*Loss CRF +w2*Loss CE Where T is the sequence length of the sample, N is the number of samples in a batch, K is the number of classes, and p i These are the true labels of the samples. It is the predicted probability value, p. i k It is the true label of the k-th class of the sample. M is the predicted probability value of the k-th class. t For a set of real labels at position t, Loss CE For classification, the cross-entropy loss and negative log-likelihood loss - log(P(y|h)) are denoted as Loss. CRF w1 and w2 are the weights of the CRF loss and the classification cross-entropy loss, respectively.
6. The method for predicting target peptide classification and cleavage sites according to claim 1, characterized in that, The model evaluation includes the following steps: Evaluation metrics were determined, and the DeepMaT model was evaluated using five-fold cross-validation and an independent test set. The independent test set included a set of target peptide sequences with cleavage site annotations and a set of target peptide sequences without cleavage site annotations. The evaluation metrics include precision, recall, F1 score, and Matthews correlation coefficient.
7. The method for predicting target peptide classification and cleavage sites according to claim 1, characterized in that, The targeted peptide sequences include signal peptides, mitochondrial transport peptides, chloroplast transport peptides, and endothelial transport peptides.
8. A system for predicting target peptide classification and cleavage sites, characterized in that, include: The acquisition module is used to acquire a dataset, which includes different types of targeted peptide sequences and non-targeted peptide sequences. The targeted peptide sequences include targeted peptide sequences with cleavage site annotations and targeted peptide sequences without cleavage site annotations. The training module is used to input sequences from the dataset into the DeepMaT model for training. The training includes: extracting evolutionary information of the sequences through the ISM module and implicitly incorporating three-dimensional structural information to obtain peptide sequence features; inputting the peptide sequence features into the Mamba2 module, the Attention module, and the feedforward layer connected in sequence to capture and fuse long-distance global dependencies and local details in the sequences to obtain fused features; inputting the fused features into the MLP multilayer perceptron to predict the target peptide classification, and simultaneously inputting the fused features into the CRF conditional random field to predict cleavage sites. The evaluation module is used for model evaluation.