Method for predicting multifunctional peptide by using pre-trained protein language model and application

Through pre-trained protein language model and imbalanced learning strategy, combined with data sampling and feature selection, the data imbalance and computational cost problems in multifunctional peptide prediction are solved, and efficient multifunctional peptide recognition and prediction are achieved, improving accuracy and robustness.

CN120496637AInactive Publication Date: 2025-08-15GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510469843.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has the problem of small samples in the prediction of multifunctional peptides, which leads to poor overfitting and generalization capabilities, and most of them focus on the prediction of single-functional biologically active peptides. There is a lack of effective multifunctional peptide prediction algorithms and has low accuracy.

Method used

The pre-trained protein language model is adopted, combined with a large supervised model of Transformer structure, and through unbalanced learning strategies and data sampling methods, SMOTE-TOMEK data comprehensive sampling and Sharple value feature selection are used to train ESM-2 pre-trained models to generate llmpred models for prediction of multifunctional peptides.

Benefits of technology

It effectively improves the effect of biological sequence prediction, improves the recognition accuracy of multifunctional peptides, reduces calculation costs, and solves the problem of data imbalance, improving the robustness and prediction performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496637A_ABST
    Figure CN120496637A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting a multifunctional peptide by using a pre-trained protein language model, which comprises the following steps of: acquiring data, and dividing the acquired data into a training set and a test set; processing the training set to obtain an optimal feature subset; using the optimal feature subset to train an ESM-2 pre-training model to obtain an llmpred model; using the test set and the llmpred model obtained through evaluation, using each functional peptide sequence as the input of the llmpred model after evaluation, finally obtaining an embedded vector corresponding to the input peptide sequence, then using a classifier to carry out classification prediction on the embedded vector, and outputting a result. According to the method, knowledge is transferred from a large amino acid sequence set based on a Transform model so as to better encode peptide information, meanwhile, the serious class imbalance phenomenon is relieved by means of an imbalance learning strategy and a data sampling method, and the biological sequence prediction effect is effectively improved through the performance of different feature encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a method for predicting multifunctional peptides using a pre-trained protein language model. Background Art

[0002] In recent years, with the development of scientific research and in-depth exploration of the laws of life activities, more and more functional peptide molecules have been discovered, and peptides with therapeutic properties are increasingly widely used in clinical diagnosis and treatment. These properties provide new ideas for disease treatment. Therefore, identifying these therapeutic peptides is of great practical significance for discovering more effective disease treatment methods. In fact, compared with their small molecule counterparts, these multifunctional peptide molecules have more hydrogen bond donors and acceptors, which indicates that they can bind to the target with good specificity and relatively few off-target side effects. At the same time, peptides are easily degraded by enzymes, and the toxicity associated with peptide metabolism is very low, with basically no cumulative toxicity. Due to the possibility of both detecting and regulating protein-protein interactions, peptides are considered to be therapeutic drugs with great potential.

[0003] Computer-assisted methods based on machine learning strategies offer fast and efficient alternatives for proteome and synthetic sequence screening to accelerate the discovery of novel functional peptides. For example, Agrawal et al. (Agrawal, P., et al. In Silico Approach for Prediction of Antifungal Peptides, Front Microbiol 2018, 9: 323) proposed combining the compositional features of peptides and support vector machines to predict antifungal peptides; MAHTPred et al. (Manavalan, B., et al. mAHTPred: a sequence-based meta-predictor for improving the prediction of anti-hypertensive peptides using effective feature representation, Bioinformatics 2019, 35(16): 2757-2765) predicted antihypertensive peptides based on feature descriptors and ensemble learning, effectively improving the balanced prediction performance and model robustness on independent data sets; Deep-AFPpred (Sharma, R., et al. Deep-AFPpred: identifying novel antifungal peptides using pretrained embeddings from seq2vec with 1DCNN-BiLSTM, Brief Bioinform 2022,23(1)) used transfer learning and 1DCNN-BiLSTM neural network to classify antifungal peptides, improving the classification accuracy of antifungal peptides; AniAMPpred (Sharma, R., et al. AniAMPpred: artificial intelligence guided discovery of novel antimicrobial peptides in animal kingdom, BriefBioinform2021,22(6)) combined deep learning-based features with support vector machines to predict potential antimicrobial proteins; Deep-ABPpred (Sharma, R., et al.Deep-ABPpred: identifying antibacterial peptides in protein sequences using bidirectional LSTM with word2vec, BriefBioinform2021,22(5)) constructed a bidirectional long short-term memory network based on the amino acid features of word2vec to identify antimicrobial peptides; and a Transformer-based method (Pang, Y., et al. Integrating transformer and imbalanced multi-label learning to identify antimicrobial peptides and their functional activities, Bioinformatics2022,38(24):5368-5374) was also proposed, using the Transformer architecture and natural language processing knowledge to extract peptide sequence information to identify antimicrobial peptides and their functional activities.

[0004] The prediction tools, algorithms, and datasets mentioned above have promoted the rapid development of this field. However, the above methods have the following problems when used: the above machine learning-based prediction methods are affected by the small number of samples, which can easily lead to overfitting and poor generalization ability; at the same time, most of the above studies focus on the prediction of single-function bioactive peptides, and there are few algorithms for predicting multiple functions together, and the accuracy is low. Therefore, there is an urgent need to design a method for predicting multifunctional peptides using a pre-trained protein language model to solve the problems existing in the above-mentioned existing technologies. Summary of the Invention

[0005] In response to the above-mentioned problems, the present invention aims to provide a method for predicting multifunctional peptides using a pre-trained protein language model. This method transfers knowledge from a large set of amino acid sequences by utilizing a large supervised model based on the Transformer structure to better encode peptide information. Specifically, an imbalanced learning strategy is adopted and a data sampling method is used to alleviate the severe class imbalance phenomenon. Finally, the performance of different feature encodings shows that transfer learning and imbalanced learning can be integrated to effectively improve the effect of biological sequence prediction.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A method for predicting multifunctional peptides using a pre-trained protein language model includes collecting data and dividing the collected data into a training set and a test set;

[0008] Process the training set to obtain the best feature subset;

[0009] Use the best feature subset to train the ESM-2 pre-trained model to obtain the llmpred model and corresponding training parameters;

[0010] Using the test set, evaluate the obtained llmpred model:

[0011] Each functional peptide sequence is used as the input of the evaluated llmpred model, and finally the embedding vector corresponding to the input functional peptide sequence is obtained. The classifier is then used to classify and predict the embedded vector and output the result.

[0012] Preferably, the process of processing the training set to obtain the optimal feature subset includes:

[0013] The SMOTE-TOMEK data comprehensive sampling method is used to sample different classification data in the training set to obtain a balanced training set;

[0014] Calculate the Shapley value of each sample in the balanced training set to obtain the optimal feature subset;

[0015] Preferably, the process of sampling different classification data in the training set to obtain a balanced training set includes:

[0016] Based on the training set, the SMOTE method is used to generate synthetic samples;

[0017] Based on the synthetic samples, the Tomek links algorithm is used to identify the noise samples in the synthetic samples and delete them;

[0018] Preferably, the process of training the ESM-2 pre-trained model to obtain the llmpred model includes:

[0019] Load the pre-trained ESM-2 model parameters;

[0020] Input the samples in the best feature subset into the ESM-2 model, compare the model output with the true label of the corresponding sample, and calculate the loss value;

[0021] Based on the loss value, calculate its gradient with respect to the ESM-2 model parameters;

[0022] Update the trainable parameters of the ESM-2 model using gradients and an optimizer.

[0023] Repeat the steps to input the samples in the best feature subset into the ESM-2 model, compare the model output results with the true labels of the corresponding samples, and calculate the loss value. Based on the loss value, calculate its gradient with respect to the ESM-2 model parameters; until the loss value no longer decreases, the llmpred model and the corresponding training parameters are obtained.

[0024] Preferably, the formula for calculating the loss value is:

[0025]

[0026] Among them, K is the total number of categories, y k and are the predicted probability and true probability that the sample belongs to the kth class, respectively.

[0027] Preferably, the process of updating the trainable parameters of the ESM-2 model using gradients and an optimizer includes:

[0028] Using the chain rule, starting from the loss function, we calculate the gradient of each layer of the ESM-2 model layer by layer.

[0029] The calculated gradient is then used to update the parameters.

[0030] Preferably, the calculation formula for parameter update using the calculated gradient is:

[0031]

[0032] Among them, w i+1 is the updated weight, w i is the current weight, b i+1 is the updated bias term, b i is the current bias term, and γ is the learning rate.

[0033] A second object of the present invention is to provide a system for predicting functional peptides using a pre-trained protein language model, comprising:

[0034] Data collection module, used to collect data and divide the collected data into training set and test set;

[0035] Get the best dataset module, which is used to process the training set to obtain the best feature subset;

[0036] Training module, training the ESM-2 pre-trained model to obtain the llmpred model;

[0037] The evaluation module uses the test set to evaluate the obtained llmpred model:

[0038] The output module is used to take each functional peptide sequence as the input of the evaluated llmpred model, and finally obtain the embedding vector corresponding to the input functional peptide sequence. The classifier is then used to classify and predict the embedded vector and output the result.

[0039] A third object of the present invention is to provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute a method for predicting multifunctional peptides using a pre-trained protein language model.

[0040] A fourth object of the present invention is to provide an electronic device comprising:

[0041] At least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a method for predicting multifunctional peptides using a pre-trained protein language model.

[0042] The beneficial effects of the present invention are as follows: the present invention discloses a method for predicting multifunctional peptides using a pre-trained protein language model. Compared with the prior art, the present invention has the following improvements:

[0043] The present invention proposes a method for predicting multifunctional peptides using a pre-trained protein language model, including data preparation, prediction of multi-class functional peptides using a Transformer-based supervised model, and multi-label classification based on comprehensive data sampling and feature selection. The method combines sequence features and embedded features extracted by ESM-2, and uses feature selection and data sampling to solve the data imbalance problem and reduce computational costs. The model is used to identify and predict a variety of different functional peptides and toxic peptides. Experimental results show that llmpred outperforms other state-of-the-art methods in various indicators such as AUC, and calculations confirm the ability of the proposed model to solve data imbalance, reduce computational costs, and improve the prediction performance of multifunctional therapeutic peptides, and can be widely applied to other biological sequence analysis problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Flowchart for dividing the data set of the present invention.

[0045] Figure 2 This is the overall architecture diagram of the pLMFPPred architecture of the present invention.

[0046] Figure 3 This is the imbalance distribution diagram before data sampling in the present invention.

[0047] Figure 4This is a distribution diagram of the physicochemical properties of peptides classified by various sequence lengths and functional activities of different targets in Example 2 of the present invention;

[0048] exist Figure 4 In the figure, Figure (a) is the amino acid score graph of each type of sequence, Figure (b) is the global charge score graph of each type of sequence, Figure (c) is the length graph of each type of sequence, Figure (d) is the global hydrophobicity distribution graph of each type of sequence, and Figure (e) is the global hydrophobicity moment distribution graph of each type of sequence.

[0049] Figure 5 This is the ROC curve diagram of pLMFPPred in Example 2 of the present invention.

[0050] Figure 6 This is a visualization comparison diagram of uniform manifold approximation and projection (UMAP) dimensionality reduction in Example 2 of the present invention; the left side is the baseline encoding and the right side is the pLMFPPred encoding. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0052] Example 1

[0053] Refer to the attached Figure 1-6 A method for predicting multifunctional peptides using a pre-trained protein language model is shown, comprising:

[0054] Collect data and divide the collected data into training set and test set;

[0055] This example uses datasets that have been verified by previous work, as shown in Table 1. The antifungal dataset is from Antifp by Agrawal et al., the antihypertensive dataset is from MAHTPred by Manavalan et al., the antibacterial dataset is from Deep-ABPpred by Sharma et al., and the toxic peptide dataset is from ATSE by Wei et al.

[0056] We prepare the dataset in two steps:

[0057] First, the sequences under each target domain in each dataset are clustered separately, and based on the clustering results, duplicate entries are removed by comparing sequence similarities;

[0058] Then, if there are any duplicate sequences between the bioactive peptide data and the toxic peptide data, the duplicate data will be retained in the toxicity dataset, where the bioactive peptide data include three categories: antifungal, antihypertensive, and antibacterial. Finally, we obtained 10,019 sequences, including 1,921 toxic peptides, 5,619 antimicrobial peptides, 1,052 antihypertensive peptides, and 1,427 antifungal peptides.

[0059] Subsequently, using a fixed random seed, the dataset obtained in step S102 is divided into an 80% training set (8015) and a 20% test set (2004), as shown in Table 1;

[0060] Table 1: Summary of dataset categories and divisions

[0061]

[0062] Process the training set to obtain the best feature subset; the specific process is:

[0063] Using the SMOTE-TOMEK data comprehensive sampling method, different classification data in the training set are sampled to obtain a balanced training set. Specific process:

[0064] Class imbalance is the most common problem in classification tasks. Figure 3 The label ratios for different tasks in the dataset of step S103 are shown, and it is clear that there is an imbalance in the data. In order to solve the class imbalance problem of the dataset in step S1, the SMOTE-TOMEK data comprehensive sampling method is introduced. SMOTE-TOMEK is a technology that combines SMOTE (Synthetic Minority Over-sampling Technique) and Tomek Link algorithm. SMOTE is an oversampling method that balances the dataset by synthesizing new minority class samples, while Tomek Link is an undersampling method used to remove noise or duplicate samples in the dataset. The specific application process of this method in this patent includes:

[0065] Based on the training set, the SMOTE method is used to generate synthetic samples;

[0066] Specifically, we first randomly select a minority class sample from the training set, calculate the distance between it and other samples of the same class, find its k nearest neighbors, then randomly select a sample from these nearest neighbors and generate a new synthetic sample through linear interpolation to supplement the total number of minority class samples so that the total number of minority class samples is consistent with the total number of majority class samples. We repeat the above process until the total number of minority class samples is consistent with the total number of majority class samples.

[0067] (1) Based on the generated synthetic samples, the Tomek links algorithm is used to identify and delete noise samples in the synthetic samples. The specific application process of this method in this patent includes:

[0068] 1. For each sample generated in step S2011, use the Euclidean distance to find its nearest neighbor among all samples (including original data and synthesized samples);

[0069] 2. Traverse each neighboring sample and check whether it belongs to a different category. If so, the two samples form a Tomek Link;

[0070] 3. For each Tomek Link, assume that one sample has a label of 0 and the other has a label of 1, then delete both samples of the Tomek Link. The sample with label 0 is the majority class sample, and the sample with label 1 is the minority class sample.

[0071] (2) Calculate the Shapley value of each sample in the balanced training set to obtain the optimal feature subset;

[0072] The embedding vector feature dimension extracted by the ESM-2 model is 5120. Such high-dimensional data is prone to interference noise and consumes a lot of resources during calculation. To reduce the interference noise between feature redundancy, improve training speed, and reduce fitting risk during training, we introduced a feature selection method based on statistical hypothesis testing and power calculation, combined with Shapley values, to select the optimal feature subset, thereby removing features that have little contribution to the model or have a negative impact.

[0073] Specifically, the SHAP (SHapleyAdditive exPlanations) library is used to calculate the Shapley value of each sample in the balanced training set to assess sample importance and evaluate the contribution of each sample to the model prediction. The larger the sample's Shapley value, the greater its contribution to the model prediction. The samples with the highest contribution are sorted in ascending order, and the samples with the top 10% contribution are selected. These samples are then grouped together to form the optimal feature subset.

[0074] Use the best feature subset to train the ESM-2 pre-trained model to obtain the llmpred model and corresponding training parameters; the specific process is as follows:

[0075] First, load the pre-trained ESM-2 model parameters;

[0076] The pre-trained ESM-2 model is in dictionary format, i.e., key-value pairs. Loading the model involves replacing the key values in the pre-trained ESM-2 model with the corresponding values. ESM-2 is a deep learning model based on the Transformer architecture, which typically predicts the three-dimensional structure of proteins by learning evolutionary information from protein sequences. However, this patented method does not predict the three-dimensional structure of proteins, but rather predicts functional peptides. Therefore, only the feature extraction network in ESM-2 is used as the backbone network of llmpred. Fine-tuning is performed based on the parameters of the original ESM-2 model, using the optimal feature subset from the previous step.

[0077] Next, we input the samples from the best feature subset into the ESM-2 model, calculate the output through the forward propagation process of the model, compare the model output results with the true labels of the corresponding samples, and calculate the loss value. The loss value is calculated as follows:

[0078]

[0079] Among them, K is the total number of categories, y k and are the predicted probability and true probability that the sample belongs to the kth class respectively;

[0080] Again, based on the loss value, calculate its gradient with respect to the ESM-2 model parameters;

[0081] Use the automatic differentiation tool in PyTorch to calculate the gradient of the loss value with respect to the ESM-2 model parameters. Specifically, the chain rule is used to propagate the gradient from the model output layer to the input layer layer by layer, and the gradient of each layer is calculated. The cross entropy loss function is selected as the loss function to measure the difference between the model output and the true label. The ESM-2 model parameters are the parameters of its convolution kernel.

[0082] Finally, the trainable parameters of the ESM-2 model are updated using the gradient and the AdamW optimizer.

[0083] The parameter update process is:

[0084] 1. Using the chain rule, starting from the loss function, calculate the gradient of each layer of the ESM-2 model layer by layer. Assume that the weight of the i-th layer of the model is w i , the bias term is b i , the loss value is L. Then the gradient calculation method of the i-th layer is:

[0085]

[0086] Among them, z i =w i-1 δ(z i-1 )+bi-1 , δ is the activation function, Indicates L versus w i Seek derivation, Indicates L to b i Derivative, θ is the derivative symbol, w i-1 is the weight after the last update, b i-1 is the bias after the last update, w 0 and b 0 is a random initial state.

[0087] 2. Then use the calculated gradient to update the parameters. The update process is:

[0088]

[0089] Among them, γ is the learning rate, its size is 0.003, w i+1 is the updated weight, w i is the current weight, b i+1 is the updated bias term, b i is the current bias term, and γ is the learning rate.

[0090] After each update, the gradient is cleared Prepare for the next backpropagation; the AdamW optimizer updates parameters with a learning rate of 0.0003 and a block size of 16.

[0091] 3. Repeat steps S303-S304 for multiple training cycles (epochs) until the model converges (the loss value no longer decreases), and obtain the trained ESM-2 model, i.e., the llmpred model and the corresponding training parameters;

[0092] Using the test set, evaluate the obtained llmpred model:

[0093] Use the test set to evaluate the performance of the llmpred model to ensure its effectiveness on specific tasks. The specific process includes:

[0094] First, load the llmpred model parameters; specifically, copy the training parameter 0 in step S3 to the corresponding variable of the model;

[0095] Next, use the llmpred model to traverse each sample in the test set and generate a prediction result for each sample. In this example, there are a total of 2004 test samples.

[0096] Again, the generated prediction results are compared with the true labels of each sample in the test set. The comparison is to directly compare the output labels. For example, if the model outputs 1 and the true label is 0, then the prediction is wrong; if it is 1, then the prediction is correct.

[0097] Finally, using precision, accuracy, recall, F1 score, etc. as evaluation indicators, the corresponding evaluation indicators are obtained; and the performance of the model is obtained based on the evaluation indicators. The final results of accuracy, recall, and F1 score are shown in Table 2, and the final conclusion is described in Example 2.

[0098] The calculation formulas for precision, recall, and F1 score are as follows:

[0099]

[0100] Among them, true positive means that the actual result is true and the prediction is also true; true negative means that the actual result is false and the prediction is also false; false positive means that the actual result is false and the prediction is true; false negative means that the actual result is true and the prediction is false;

[0101] Each functional peptide sequence is used as the input of the evaluated llmpred model, and finally the embedding vector corresponding to the input functional peptide sequence is obtained, which is conducive to the classifier to classify and predict the embedded vector and output the result;

[0102] In order to convert functional peptides into feature vectors for classification, each functional peptide sequence is first used as the input of the evaluated llmpred model. The evaluated llmpred model encodes the input features and then outputs a 5120-dimensional embedding vector (which corresponds to its input functional peptide sequence). Finally, this embedding vector is sent to the Light Gradient Boosting classifier for functional peptide classification prediction.

[0103] Example 2

[0104] Different from the above-mentioned embodiment 1, this embodiment verifies the performance of the llmpred model evaluated in step 4 in order to verify the embodiment 1, specifically:

[0105] Performance indicators and comparative analysis

[0106] The performance of the model is evaluated using ROC-AUC, accuracy, precision, recall, and F1 score. The corresponding indicator calculation formulas are as follows:

[0107]

[0108] Among them, true positive means that the actual value is true and the prediction is also true; true negative means that the actual value is false and the prediction is also false; false positive means that the actual value is false and the prediction is true; false negative means that the actual value is true and the prediction is false.

[0109] In order to evaluate the performance of the model, four traditional sequence feature representations and two protein language models were used as baseline methods to compare with the llmpred model and the method of predicting functional peptides using the llmpred model; including composition-transition-distribution transition descriptor (CTDT), grouped amino acid composition (GAAC), pseudo amino acid composition (PAAC) and amino acid composition (AAC), the four traditional sequence representations were extracted using iFeature; the protein language model included the pre-trained BERT model named TAPE and ProtT5. For the downstream prediction model, we trained three ML-based models, including eXtreme Gradient Boosting, Random Forest and Support Vector Machine, as well as a Bi-directional Long Short-Term Memory deep learning model, in order to compare with the popular neural network-based deep learning model.

[0110] Among the various sequences collected, the length of the sequences is different. Compared with peptides without antimicrobial activity, sequences with antifungal activity tend to be longer, while most antimicrobial, antihypertensive and toxic sequences are shorter. This is consistent with the view that antimicrobial peptides maintain their membrane interaction activity through shorter positively charged amino acid chains. The distribution of physicochemical properties of other peptides classified by functional activities of different targets can be found in Figure 4 ;

[0111] Table 2: Effect comparison

[0112]

[0113] In Table 2 , the best results are in bold, and the second-best results are underlined;

[0114] Table 3: Comparison of feature extraction effects

[0115]

[0116] As shown in Table 2, llmpred achieved the best accuracy (0.9741) and the best F1 score (0.9738) among all models, followed by the extreme gradient boosting tree (accuracy = 0.9731, F1 score = 0.9729) and the bidirectional long short-term memory network model (accuracy = 0.9660, F1 score = 0.9657). Random forest and support vector machine performed poorly. The information gain obtained from the embedding calculated from the ESM-2 model enabled llmpred to achieve the best performance among all models. The model performance was far ahead of the baseline model and the llmpred model without sampling technology. The use of data sampling technology improved the performance of llmpred compared to llmpred with class imbalance. This shows that data sampling technology reduces the imbalance of different categories of data, making the model prediction results more balanced and robust. Figure 5 The AUC-ROC curves for these categories are shown. The AUC values for each category are impressive, exceeding 0.99, demonstrating the excellent and stable predictive performance of llmpred. Table 3 shows the results of feature extraction. The model's feature dimension was reduced from 5120 to 340, a 93.4% reduction in feature dimension, and training time was significantly reduced from 74 seconds to 5.5 seconds, a 92.6% reduction. llmpred only experienced a 0.2%-0.3% accuracy loss. This demonstrates that feature selection removed noisy features, avoided model overfitting, and significantly reduced the dataset dimension, resulting in a 92.6% reduction in model training time and a significant reduction in computational overhead. We consider a 0.2%-0.3% loss to be acceptable in this context. Notably, the neural network-based deep learning model did not outperform llmpred, which was built using a gradient boosted decision tree model. This suggests that the advantages of deep learning models may not be obvious in this task.

[0117] Table 4: Model comparison

[0118]

[0119] In Table 4 , the best results are in bold, and the second-best results are underlined;

[0120] Table 5: Comparison of ESM-2 models

[0121]

[0122] In Table 5, the best results are in bold and the second best results are underlined; Result 2: Performance comparison of different feature encodings

[0123] The second task of llmpred was to evaluate the performance of the upstream ESM-2 model. We conducted comparative experiments on features extracted by various methods. Table 4 summarizes the performance metrics of each model. The ESM-2 model achieved the best accuracy of 96.96% and the best f1-score of 96.94%, outperforming ProtT5 and TAPE, while traditional feature descriptors performed more modestly. The detailed performance of the different models is shown in Table X. The results show that llmpred based on ESM-2 achieved the best year-over-year performance. We also tested ESM-2 models with different parameters. As shown in the table below, the ESM-2 model with 15 billion parameters outperformed all other models. This indicates that as the number of model parameters increases, the protein language model's ability to understand the sequence improves, leading to improvements in the model's prediction accuracy and F1-score. This means that the model is better able to overcome the bias existing between imbalanced classes, enabling more accurate classification of data.

[0124] Results 3: Peptide Feature Embedding Visualization

[0125] Compared to baseline classifiers, the ESM-2-based llmpred demonstrated substantial performance improvements on all of the aforementioned tasks. These improvements can be primarily attributed to the latent features learned and extracted by ESM-2 from peptide sequences. To visually demonstrate the effectiveness of the features extracted by ESM-2, we compared the ESM-2 embeddings with the baseline dataset features using dimensionality reduction based on Uniform Manifold Approximation Projection (UMAP) to present a clear feature representation. Compared to the baseline dataset features, the ESM-2 embeddings demonstrated a significant impact in promoting the intrinsic separation of peptides from various classes. The features learned from the embedding vector space resulted in distinct clusters between peptides with different functional targets, with sharper separation boundaries compared to the baseline features. Compared to the ESM-2 embeddings, the feature encoding of the baseline method failed to achieve differentiation in the UMAP representation, resulting in many peptides with different functions being entangled with other peptides. This analysis demonstrates that the model based on the ESM-2 embeddings produces more accurate representations that help distinguish peptides with different functions.

[0126] From the above results, it can be seen that the therapeutic peptide prediction method based on the large-scale pre-trained Transformer protein language model proposed in Example 1 of the present invention combines sequence features and embedded features extracted by ESM-2, and solves the data imbalance problem and reduces computational costs through feature selection and data sampling. We use this model to identify and predict three different functional peptides and toxic peptides. The experimental results show that llmpred outperforms other state-of-the-art methods in various indicators such as AUC, and the calculations confirm the ability of the proposed model to solve data imbalance, reduce computational costs and improve therapeutic peptide prediction performance. We believe that the proposed solution, which combines large-scale language models with feature selection and data sampling techniques, can be widely applied to other biological sequence analysis problems.

[0127] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting multifunctional peptides using a pre-trained protein language model, characterized in that: include Collect data and divide the collected data into training set and test set; Process the training set to obtain the best feature subset; Use the best feature subset to train the ESM-2 pre-trained model to obtain the llmpred model; Use the test set to evaluate the obtained llmpred model; Each functional peptide sequence is used as the input of the evaluated llmpred model, and finally the embedding vector corresponding to the input peptide sequence is obtained. The classifier is then used to perform classification prediction on the embedded vector and output the result.

2. The method for predicting multifunctional peptides using a pre-trained protein language model according to claim 1, characterized in that: The process of processing the training set to obtain the best feature subset includes: The SMOTE-TOMEK data comprehensive sampling method is used to sample different classification data in the training set to obtain a balanced training set; Calculate the Shapley value of each sample in the balanced training set to obtain the optimal feature subset.

3. The method for predicting multifunctional peptides using a pre-trained protein language model according to claim 2, characterized in that: The process of sampling different classification data in the training set to obtain a balanced training set includes: Based on the training set, the SMOTE method is used to generate synthetic samples; Based on the synthetic samples, the Tomek links algorithm is used to identify the noise samples in the synthetic samples and delete them.

4. The method for predicting multifunctional peptides using a pre-trained protein language model according to claim 1, wherein: The process of training the ESM-2 pre-trained model to obtain the llmpred model includes: Input the samples in the best feature subset into the ESM-2 model, compare the output of the ESM-2 model with the true label of the corresponding sample, and calculate the loss value; Based on the loss value, calculate its gradient with respect to the ESM-2 model parameters; Using the gradient and the optimizer, the trainable parameters of the ESM-2 model are updated until the loss value no longer decreases, resulting in the llmpred model.

5. The method for predicting multifunctional peptides using a pre-trained protein language model according to claim 4, characterized in that: The formula for calculating the loss value is: Among them, K is the total number of categories, y k and are the predicted probability and true probability that the sample belongs to the kth class, respectively.

6. The method for predicting multifunctional peptides using a pre-trained protein language model according to claim 5, characterized in that: The process of updating the trainable parameters of the ESM-2 model using gradients and an optimizer involves: Using the chain rule, starting from the loss function, we calculate the gradient of each layer of the ESM-2 model layer by layer. The calculated gradient is then used to update the parameters.

7. The method for predicting multifunctional peptides using a pre-trained protein language model according to claim 6, characterized in that: The calculation formula for parameter update using the calculated gradient is: Among them, w i+1 is the updated weight, w i is the current weight, b i+1 is the updated bias term, b i is the current bias term, and γ is the learning rate.

8. A system for predicting functional peptides using a pre-trained protein language model, characterized in that: include: Data collection module, used to collect data and divide the collected data into training set and test set; Get the best dataset module, which is used to process the training set to obtain the best feature subset; Training module, training the ESM-2 pre-trained model to obtain the llmpred model; The evaluation module uses the test set to evaluate the obtained llmpred model: The output module is used to take each functional peptide sequence as the input of the evaluated llmpred model, and finally obtain the embedding vector corresponding to the input functional peptide sequence. The classifier is then used to classify and predict the embedded vector and output the result.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

10. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Landslide susceptibility evaluation method capable of explaining machine learning combined with class imbalance processing

    CN122221054A