Protein language model-based IL-4 induced peptide prediction method and system
Through a method based on protein language model, balancing data and combining deep learning models for IL-4-induced peptide prediction, the problems of data imbalance and limitations of feature extraction in existing tools are solved, and prediction accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510092820.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-02
AI Technical Summary
The existing IL-4-induced peptide prediction tools have data imbalance and feature extraction limitations, resulting in insufficient prediction accuracy and low efficiency.
Using a protein language model-based method, the positive and negative sample peptide sequences of IL-4-induced peptides were obtained and balanced, and the amino acids and protein characteristic vectors in the signal peptide sequence were extracted, and the GRU layer in the deep learning model was combined for data processing and prediction.
The accuracy and efficiency of IL-4-induced peptide prediction are improved, and the generalization ability and prediction accuracy of the model are enhanced through data balance and the use of deep learning models.
Smart Images

Figure CN119920322A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method and system for predicting IL-4 induced peptides based on a protein language model. Background Art
[0002] With the continuous progress of medicine and biotechnology, the development of antiviral drugs and vaccines has achieved remarkable results, but infectious diseases are still a major threat to global health. Interleukin-4 (IL-4) in the immune system plays a vital role in regulating immune responses, mediating allergic reactions and antibody production. IL-4 promotes the differentiation of helper T cells (Th cells) into Th2 cells through interaction with CD4+T cells, thereby activating the proliferation and differentiation of B cells, and ultimately leading to the production of immunoglobulin E (IgE). This process is crucial for regulating allergic reactions and the occurrence and development of various immune diseases. Effectively predicting the effects of IL-4-induced peptides can not only improve the accuracy of immunotherapy, but also provide new strategies for vaccine development.
[0003] Traditional methods for predicting IL-4 induced peptides include experimental verification methods, machine learning tools, and predictions using protein language models. Among them, although existing experimental verification methods (such as ELISA, flow cytometry, etc.) can provide accurate prediction results, these methods generally have problems such as long time consumption, complex operation, and high cost. Especially when facing large-scale samples, traditional experimental methods are difficult to meet the needs of high-throughput screening and are easily restricted by experimental conditions; some existing machine learning tools (such as IL-4-Pred, Meta-IL4, etc.) have improved the prediction accuracy of IL-4 induced peptides to a certain extent by adopting traditional machine learning algorithms combined with limited feature engineering, but the performance of the tools is limited by problems such as data imbalance and model complexity, resulting in certain limitations in their practical applications.
[0004] It can be seen that although the existing IL-4 induced peptide prediction tools have improved the prediction performance to a certain extent, there are still problems such as data imbalance and limitations in feature extraction, which lead to insufficient accuracy and low efficiency in IL-4 induced peptide prediction. Summary of the invention
[0005] Based on this, in order to solve the above technical problems, a method and system for predicting IL-4 induced peptides based on a protein language model are provided, which can improve the accuracy and efficiency of IL-4 induced peptide prediction.
[0006] A method for predicting IL-4 induced peptides based on a protein language model, the method comprising:
[0007] Obtaining positive and negative sample peptide sequences of IL-4-induced peptides, determining a training set from each of the positive and negative sample peptide sequences, and performing data balancing processing on the training set using edited neighbors and synthetic minority class oversampling techniques to obtain each processed IL-4-induced peptide sequence;
[0008] Extracting a signal peptide sequence from the IL-4 inducing peptide sequence, inputting the signal peptide sequence into a protein language model, extracting an amino acid feature vector and a protein feature vector from the signal peptide sequence through the protein language model, and concatenating the amino acid feature vector and the protein feature vector to obtain a joint feature vector;
[0009] Normalizing the joint feature vector, inputting the processed feature vector into a deep learning model, and processing it through a GRU layer in the deep learning model to obtain prediction features of the peptide sequence;
[0010] The predicted features are input into a classification model to obtain the category information of the predicted IL-4-induced peptide; the classification model is trained using a cross entropy loss function.
[0011] In one embodiment, obtaining positive and negative sample peptide sequences of IL-4 inducing peptides, and determining a training set from each of the positive and negative sample peptide sequences, comprises:
[0012] Collecting positive and negative sample peptide sequences of IL-4 inducing peptides; wherein the positive sample peptide sequence is a peptide sequence that can induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4 secretion;
[0013] A target division ratio is determined, and the positive and negative sample peptide sequences are divided into a training set and a test set according to the target division ratio.
[0014] In one embodiment, the training set is subjected to data balancing processing using the editing neighbor and synthetic minority class oversampling techniques to obtain processed IL-4 inducing peptide sequences, including:
[0015] Using synthetic minority class oversampling technology to oversample the positive sample peptide sequences in the training set to generate synthetic sample peptide sequences;
[0016] The synthetic sample peptide sequence was cleaned using the edited neighbor technique to obtain each processed IL-4 inducing peptide sequence.
[0017] In one embodiment, extracting the signal peptide sequence in the IL-4 inducing peptide sequence, inputting the signal peptide sequence into a protein language model, and extracting the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, including:
[0018] Determining a characteristic cutoff length of a signal peptide, and extracting a signal peptide sequence from the IL-4 inducing peptide sequence according to the characteristic cutoff length of the signal peptide;
[0019] The signal peptide sequence is input into a protein language model, and the dependency relationship between amino acid residues is captured by the protein language model, thereby extracting an amino acid feature vector of the signal peptide sequence and a context feature vector of the protein sequence.
[0020] In one embodiment, the amino acid feature vector and the protein feature vector are concatenated to obtain a joint feature vector, including:
[0021] Determine each translation unit in each of the signal peptide sequences, and search for the amino acid feature vector and protein feature vector corresponding to the translation unit;
[0022] splicing the amino acid feature vector and the protein feature vector corresponding to the same translation unit to obtain a joint feature vector of the translation unit;
[0023] The dimension of the joint feature vector is the sum of the dimensions of the amino acid feature vector and the protein feature vector.
[0024] In one embodiment, the joint feature vector is normalized, the processed feature vector is input into a deep learning model, and processed by a GRU layer in the deep learning model to obtain prediction features of the peptide sequence, including:
[0025] Using the Z-score standardization method to adjust the mean and standard deviation of each feature in the joint feature vector to obtain a processed feature vector;
[0026] The processed feature vector is input into a deep learning model, and processed through a convolutional layer, a pooling layer, and a GRU layer in the deep learning model to obtain prediction features of the peptide sequence.
[0027] In one embodiment, the method further comprises:
[0028] Selecting a hyperparameter tuning method, and using the hyperparameter tuning method to determine an optimal parameter combination;
[0029] The parameters in the deep learning model are adjusted based on the optimal parameter combination.
[0030] In one embodiment, the method further comprises:
[0031] Obtaining a peptide sequence to be predicted, inputting the peptide sequence to be predicted into a protein language model, and extracting target features through the protein language model;
[0032] Inputting the target features into the deep learning model to obtain IL-4 induced peptide prediction results;
[0033] The IL-4 induced peptide prediction results are displayed on the display interface.
[0034] A protein language model-based IL-4 induced peptide prediction system, the system comprising:
[0035] A peptide sequence processing module is used to obtain positive and negative sample peptide sequences of IL-4-induced peptides, determine a training set from each of the positive and negative sample peptide sequences, and perform data balancing processing on the training set using edited neighbors and synthetic minority class oversampling techniques to obtain each processed IL-4-induced peptide sequence;
[0036] A feature extraction module is used to extract the signal peptide sequence in the IL-4 inducing peptide sequence, input the signal peptide sequence into a protein language model, extract the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, and concatenate the amino acid feature vector and protein feature vector to obtain a joint feature vector;
[0037] A peptide sequence prediction module, used to normalize the joint feature vector, input the processed feature vector into a deep learning model, and process it through a GRU layer in the deep learning model to obtain prediction features of the peptide sequence;
[0038] The prediction result acquisition module is used to input the prediction features into the classification model to obtain the category information of the predicted IL-4 induced peptide; the classification model is trained using the cross entropy loss function.
[0039] In one embodiment, the peptide sequence processing module is also used to: collect positive and negative sample peptide sequences of IL-4 inducing peptides; wherein the positive sample peptide sequence is a peptide sequence that can induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4 secretion; determine a target division ratio, and divide the positive and negative sample peptide sequences into a training set and a test set according to the target division ratio.
[0040] The above-mentioned IL-4 induced peptide prediction method and system based on protein language model can avoid data imbalance by adopting editing neighbors and synthetic minority class oversampling technology for data processing; the use of protein language model can efficiently extract long-range dependencies and implicit patterns in protein sequences, and provide rich and high-dimensional feature representation; the GRU layer in the model has a powerful ability to process sequence data, and can capture the temporal dependencies in IL-4 induced peptide sequences. Through in-depth analysis of protein sequences and capture of long-range dependencies, the accuracy of IL-4 induced peptide prediction can be greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. 1 is an application environment diagram of an IL-4 induced peptide prediction method based on a protein language model in one embodiment;
[0042] Figure 2 FIG. 1 is a schematic diagram of a method for predicting IL-4-induced peptides based on a protein language model in one embodiment;
[0043] Figure 3 A schematic diagram of comprehensive analysis of model uncertainty, stability and subset performance in one embodiment;
[0044] Figure 4 A schematic diagram of a model framework of PLM-IL4 in one embodiment;
[0045] Figure 5 is a structural block diagram of an IL-4 induced peptide prediction system based on a protein language model in one embodiment;
[0046] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] The IL-4 induced peptide prediction method based on the protein language model provided in the embodiments of the present application can be applied to Figure 1 In the application environment shown. Figure 1As shown, the application environment includes a computer device 110. The computer device 110 can obtain positive and negative sample peptide sequences of IL-4 induced peptides, determine a training set from each positive and negative sample peptide sequence, and use the editing neighbor and synthetic minority class oversampling technology to perform data balancing on the training set to obtain each processed IL-4 induced peptide sequence; the computer device 110 can extract the signal peptide sequence in the IL-4 induced peptide sequence, input the signal peptide sequence into the protein language model, extract the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, and splice the amino acid feature vector and the protein feature vector to obtain a joint feature vector; the computer device 110 can normalize the joint feature vector, input the processed feature vector into the deep learning model, and process it through the GRU layer in the deep learning model to obtain the predicted features of the peptide sequence; the computer device 110 can input the predicted features into the classification model to obtain the predicted category information of the IL-4 induced peptide; the classification model is trained using the cross entropy loss function. Among them, the computer device 110 can be, but not limited to, various personal computers, laptops, smart phones, robots and other devices.
[0049] In one embodiment, Figure 2 As shown, a method for predicting IL-4 induced peptides based on a protein language model is provided, comprising the following steps:
[0050] Step 202, obtain positive and negative sample peptide sequences of IL-4 inducing peptides, determine training sets from each positive and negative sample peptide sequence, and use edited neighbor and synthetic minority class oversampling technology to perform data balancing processing on the training set to obtain each processed IL-4 inducing peptide sequence.
[0051] The computer equipment can collect positive and negative sample data of IL-4 induced peptides, which are from public databases such as Immune Epitope Database (IEDB) to ensure that the samples are widely representative. The data set in this embodiment is derived from Meta-IL4, downloaded from the Immune Epitope Database (IEDB), and contains 985 IL-4 induced peptides (positive samples) and 774 non-IL-4 induced peptides (negative samples), all from MHC class II molecules, and there is no restriction on the host or epitope source.
[0052] In one embodiment, a provided IL-4 induced peptide prediction method based on a protein language model may also include a process of dividing data, the specific process including: collecting positive and negative sample peptide sequences of IL-4 induced peptides; wherein the positive sample peptide sequence is a peptide sequence that can induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4 secretion; determining a target division ratio, and dividing the positive and negative sample peptide sequences into a training set and a test set according to the target division ratio.
[0053] Among them, the positive sample peptide sequence is a peptide sequence known to be able to induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4. In this embodiment, after collecting the positive and negative sample peptide sequences of IL-4 inducing peptides, the data set can be divided into a training set and a test set in a ratio of 8:2 to ensure that the training and testing of the model have independence and generalization ability. Among them, when dividing, it is ensured that the distribution of each class (positive samples and negative samples) in the training set and the test set is similar.
[0054] In one embodiment, a protein language model-based IL-4 induced peptide prediction method is provided, which may also include a data processing process, the specific process including: using a synthetic minority class oversampling technique to oversample the positive sample peptide sequences in the training set to generate synthetic sample peptide sequences; using the edit neighbor technique to clean the data of the synthetic sample peptide sequences to obtain the processed IL-4 induced peptide sequences.
[0055] Due to the inherent imbalance of the data set, the performance of the classifier tends to be biased towards the majority class samples. In view of the common data imbalance problem in biomedical data sets, in this embodiment, a dynamic hybrid data balancing strategy is constructed by innovatively combining Edited Nearest Neighbours (ENN) and Synthetic Minority Over-sampling Technique (SMOTE). The dynamic hybrid data balancing strategy not only achieves the balance of the data set, but also retains the intrinsic structural characteristics of the original data, significantly enhancing the robustness and generalization ability of the model when processing unbalanced data. The training set and test set finally generated are balanced to 773 and 145 sequences in each category, ensuring that the model can fully learn the characteristics of each category during the training process, and improving the accuracy and stability of the prediction. Through a series of data enhancement and noise reduction processing, the imbalance of the data set is optimized, and the generalization ability of the model is significantly improved.
[0056] Specifically, due to the imbalance of positive and negative samples in the IL-4 induced peptide sequence, synthetic minority oversampling technology and edited neighbor technology are used for balance. Among them, SMOTE (Synthetic Minority Oversampling Technology): oversample the minority class (positive samples) in the training set to generate synthetic samples and increase the number of positive samples; ENN (Edited Neighbors): clean the oversampled data and remove those samples that may introduce noise to improve the robustness of the model. Among them, oversampling processing (SMOTE): use the SMOTE method to oversample minority samples and generate synthetic samples to increase the number of minority samples; denoising processing (ENN): after oversampling, use the ENN method to remove samples that are different from the nearest neighbor samples, and further clean the data and eliminate noise through undersampling; balanced data construction: finally, a balanced training set and test set are constructed to ensure the balanced representativeness of positive and negative samples.
[0057] Next, the computer device can perform length normalization on all peptide sequences to ensure the consistency and standardization of the input signal peptide sequences.
[0058] Step 204, extract the signal peptide sequence in the IL-4 inducing peptide sequence, input the signal peptide sequence into the protein language model, extract the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, and concatenate the amino acid feature vector and protein feature vector to obtain a joint feature vector.
[0059] When performing feature extraction, the 30-layer ESM-2 protein language model can be used. The protein language model can automatically learn and capture complex sequence features and deep semantic information from large-scale protein sequences through its deep learning-based architecture. The ESM-2 model uses a multi-layer stacked Transformer structure to efficiently extract long-range dependencies and implicit patterns in protein sequences, providing rich and high-dimensional feature representations, laying a solid foundation for subsequent prediction tasks.
[0060] In one embodiment, a method for predicting IL-4 induced peptides based on a protein language model is provided, which may also include a process for extracting feature vectors. The specific process includes: determining a signal peptide feature cutoff length, and extracting a signal peptide sequence in an IL-4 induced peptide sequence according to the signal peptide feature cutoff length; inputting the signal peptide sequence into a protein language model, capturing the dependency between amino acid residues through the protein language model, and extracting an amino acid feature vector of the signal peptide sequence and a context feature vector of the protein sequence.
[0061] In this embodiment, for each IL-4 inducing peptide sequence, its signal peptide portion is extracted, usually the first 100 to 150 amino acid residues of the peptide sequence; the cutoff length (M) of the signal peptide feature ranges from 80 to 200 amino acids, preferably 100 to 150 amino acids.
[0062] Next, the computer device can input the extracted signal peptide sequence into a pre-trained protein language model (such as ESM-2) to obtain the amino acid feature vector of the peptide sequence and the context feature vector of the protein sequence. The protein language model can effectively capture the dependency between amino acid residues and generate a high-dimensional feature representation through self-supervised learning.
[0063] In one embodiment, a method for predicting IL-4-induced peptides based on a protein language model is provided, which may also include a feature splicing process, the specific process including: determining each translation unit in each signal peptide sequence, and searching for an amino acid feature vector and a protein feature vector corresponding to the translation unit; splicing the amino acid feature vector and the protein feature vector corresponding to the same translation unit to obtain a joint feature vector of the translation unit; wherein the dimension of the joint feature vector is the sum of the dimensions of the amino acid feature vector and the protein feature vector.
[0064] Wherein, the translation unit is the smallest unit of the translation of the mRNA sequence into the amino acid sequence. In this embodiment, for each translation unit of each peptide sequence, the corresponding amino acid residue feature vector and protein sequence feature vector are spliced to obtain the joint feature vector of the translation unit. The dimension of the spliced feature vector is the sum of the dimensions of the amino acid residue feature vector and / or protein sequence feature vector of each translation unit.
[0065] Step 206, normalize the joint feature vector, input the processed feature vector into the deep learning model, process it through the GRU layer in the deep learning model, and obtain the prediction features of the peptide sequence.
[0066] After extracting the high-dimensional joint feature vector, it can be input into the Gated Recurrent Unit (GRU) model for prediction. As an efficient variant of recurrent neural network (RNN), GRU has a strong ability to process sequence data and can capture the temporal dependencies in IL-4 induced peptide sequences. Compared with the traditional long short-term memory network (LSTM), GRU reduces computational complexity and improves training efficiency by simplifying the structure, while effectively preventing the gradient vanishing problem while maintaining high performance.
[0067] Starting from the input layer, the input dimension corresponds to the length of the feature vector extracted from the protein sequence through the ESM-2 model. The GRU layer is the core of the model and uses a gating mechanism to effectively capture long-term dependencies in sequence data. The key components of each GRU layer include the number of units (hyperparameter tuning ranges from 50 to 300) and the configuration of returning sequences (except for the last layer) to preserve the sequence information in the network. To mitigate overfitting, L2 regularization (including kernel, recursive, and bias regularization) is used. In addition, recursive Dropout is applied to improve the generalization ability of the model by randomly dropping neurons during training.
[0068] After each GRU layer, batch normalization can be used to standardize the data, reduce internal covariate shift, speed up training, and enhance the stability of the model. Batch normalization is defined as: in, and denote the mean and standard deviation of the kth feature in the mini-batch B, respectively, and ∈ is a small constant used to avoid division by zero.
[0069] A Dropout layer is added after each GRU layer and fully connected layer, and the Dropout rate is between 0.1 and 0.5, which is determined by hyperparameter tuning. The purpose of Dropout is to prevent overfitting by randomly deactivating some neurons during training.
[0070] After the GRU layer, several fully connected layers are added to extract features and perform classification. The number of neurons ranges from 32 to 256, and the specific value is determined by hyperparameter tuning. The output of the fully connected layer can be expressed as: (l) =f(W (l) x (l) )+b l ; where f is the activation function, W (l) is the weight matrix, b l is the bias vector.
[0071] The output layer consists of a fully connected layer using a sigmoid activation function, which generates a binary classification result (0 or 1). The sigmoid function maps the output value between 0 and 1, representing the probability of the prediction: Where σ is the sigmoid function, W out and b out are the weight matrix and bias vector of the output layer respectively.
[0072] During the model compilation process, Adam and RMSprop optimizers were evaluated, and the optimal optimizer was selected through hyperparameter tuning. The loss function uses binary cross entropy, and the performance indicator is accuracy. The binary cross entropy loss function is defined as: where y i represents the true label, represents the predicted probability, and N is the number of samples.
[0073] In one embodiment, a method for predicting IL-4-induced peptides based on a protein language model is provided, which may also include a process for processing to obtain prediction features, the specific process including: using the Z-score standardization method to adjust the mean and standard deviation of each feature in the joint feature vector to obtain a processed feature vector; inputting the processed feature vector into a deep learning model, processing it through the convolutional layer, pooling layer, and GRU layer in the deep learning model to obtain the prediction features of the peptide sequence.
[0074] In this embodiment, in order to avoid the influence of scale differences of different features on model training, the concatenated feature vectors are normalized. The Z-score normalization method is usually used to adjust the mean of each feature to 0 and the standard deviation to 1, thereby ensuring the balanced contribution of each feature during model training. Then, the computer device can input the normalized feature vector into the deep learning model for training. The input layer of the model receives the concatenated feature vector, which is then processed through a series of convolutional layers, pooling layers, GRU layers, etc., and finally obtains the comprehensive prediction features of the peptide sequence.
[0075] In one embodiment, a protein language model-based IL-4 induced peptide prediction method provided may also include an optimization and hyperparameter tuning process, the specific process including: selecting a hyperparameter tuning method, and using the hyperparameter tuning method to determine the optimal parameter combination; adjusting the parameters in the deep learning model based on the optimal parameter combination.
[0076] In order to further improve the performance of the deep learning model, in this embodiment, systematic hyperparameter adjustment and dynamic learning rate scheduler are used. Through advanced hyperparameter tuning methods such as grid search and Bayesian optimization, the optimal parameter combination is accurately located. At the same time, the learning rate scheduler is used to dynamically adjust the learning rate to accelerate model convergence, improve training stability, and avoid overfitting. This optimization strategy ensures efficient training and excellent performance of the model in a complex data environment.
[0077] Among them, hyperparameter tuning is performed using the keras_tuner library, and Bayesian optimization is used to determine the optimal parameter combination, including the number of GRU layers, the number of units in each layer, the Dropout rate, the number of fully connected layers, the number of neurons in each layer, and the type of optimizer used. In order to reduce overfitting and improve generalization during model training, a variety of callback functions are used. EarlyStopping monitors the validation loss and stops training when there is no further improvement. LearningRateScheduler adaptively adjusts the learning rate based on the number of training rounds. ReduceLROnPlateau reduces the learning rate when the validation loss tends to stabilize to promote better convergence.
[0078] When training a model, it can include several parts, including training sample preparation, loss function selection, and optimization algorithm, among which:
[0079] Training sample preparation: Input the feature vector containing positive and negative samples into the classification or regression model. For classification tasks, classifiers such as support vector machine (SVM) and random forest (RF) can be used; for regression tasks, deep neural network (DNN) or linear regression model can be used.
[0080] Loss function selection: If it is a classification task, the cross-entropy loss function (Cross-Entropy Loss) is used; if it is a regression task, the mean square error (MSE) or mean absolute error (MAE) is used as the loss function.
[0081] Optimization algorithm: Use the Adam optimization algorithm for back propagation to update the model parameters until the loss function converges.
[0082] Step 208, input the predicted features into the classification model to obtain the category information of the predicted IL-4 induced peptide; the classification model is trained using the cross entropy loss function.
[0083] The computer device can input the predicted features into the classification model, train it using the cross entropy loss function, and finally output the predicted positive or negative category of the IL-4 inducing peptide. Then, the computer device can evaluate the model using an independent test set, using multiple indicators such as accuracy (ACC), sensitivity (SN), specificity (SP), Matthews correlation coefficient (MCC) and AUC to ensure the comprehensiveness and accuracy of the evaluation results; in the regression task, the mean square error or mean absolute error loss function is used to predict the induction ability score of the peptide sequence.
[0084] In one embodiment, in order to comprehensively evaluate the performance and reliability of the classification model, a series of multi-level evaluation and verification methods are adopted, including model uncertainty distribution analysis, stability test, subset performance evaluation and multi-model comparative analysis. These methods not only verify the predictive ability of the classification model, but also reveal its robustness and superiority in different scenarios. Specifically, model uncertainty distribution analysis: is the model uncertainty distribution based on entropy value, indicating that most predictions have a high degree of confidence, and some samples show appropriate uncertainty; stability test: is to compare the original sample with the perturbation sample to show the prediction stability of the model after the introduction of noise and verify its robustness; subset performance evaluation: is to evaluate the performance of the GRU model on different test subsets, showing excellent performance in key indicators such as precision and recall; multi-model comparative analysis: is to compare the PLM-IL-4 model with existing tools (such as Meta-IL4) in multiple performance indicators through radar charts, highlighting the overall superiority of PLM-IL-4 in specificity (SP), Matthews Correlation Coefficient (MCC) and AUC.
[0085] In this embodiment, by introducing entropy as a quantitative indicator, a clear direction is provided for further optimizing the processing of complex samples. Figure 3 As shown in Figure A, the entropy values of most samples are close to zero, indicating that the model has a high degree of confidence in its predictions for these samples; the entropy values of a few samples are distributed between 0.1 and 0.7. GRU is able to capture key dependencies in complex sequences, allowing the model to maintain stable prediction performance even in the face of high entropy samples, indicating that these samples have complex characteristics, thus showing appropriate uncertainty in model predictions.
[0086] In this embodiment, in the model robustness analysis, the original sample and the perturbation sample are compared. Figure 3 As shown in Figure B, the perturbation sample is generated by adding random noise with a mean of 0 and a standard deviation of 0.1 to the test data. The results show that the overall trend of the original sample is consistent with that of the perturbation sample, indicating that the model can maintain the robustness and stability of prediction under slight data changes. By further introducing the GRU module in the feature processing stage, the model's anti-noise performance for perturbation data is enhanced by utilizing its modeling ability for time series data. This combination makes the model more stable in practical applications, especially for protein sequence data containing random noise.
[0087] In this embodiment, the performance evaluation of the GRU model on different test subsets is as follows: Figure 3As shown in Figure C, a variety of evaluation indicators are used, including accuracy, precision, recall, F1 score, and Matthews correlation coefficient (MCC). Figure 3 Subset 2, shown in C, performs better in most indicators (especially precision and recall); although the MCC indicator is slightly lower, it remains stable in key indicators such as accuracy and ROC-AUC. Combining the features extracted by the ESM-2 model with GRU achieves a balance of performance on different subsets and has strong adaptability to complex samples. GRU's advantage in capturing sequence context relationships enables the model to maintain high performance on subsets with different feature distributions.
[0088] Among them, the IL-4 induced peptide prediction method based on the protein language model provided in the present application can be applied to the PLM-IL4 model. In one embodiment, the framework of PLM-IL4 is as follows Figure 4 As shown, it includes A, data preprocessing; B protein language model; C, GRU network; D, IL-4 induced peptide prediction; E, computer device server. In this embodiment, combined with deep learning technology and innovative data preprocessing methods, Evolutionary Scale Modeling (ESM-2) and Gated Recurrent Unit (GRU) networks are used to perform deep feature extraction and classification prediction of protein sequences. Through in-depth analysis of protein sequences and capture of long-range dependencies, not only the accuracy of IL-4 induced peptide prediction is greatly improved, but also through innovative data processing and model optimization strategies, the data imbalance problem and overfitting risk existing in the prior art are solved, and the generalization ability and robustness of the model are significantly improved.
[0089] Specifically, ESM-2 uses the Transformer architecture to pre-train on large-scale protein sequences, extracting high-dimensional features containing evolutionary and structural information, providing strong data support for the identification of IL-4-induced peptides; the GRU network further learns the contextual information in the sequence on this basis, capturing key long-range dependencies, and ensuring the model's efficiency and accuracy in processing complex protein sequences. In addition, in order to solve the imbalance problem of positive and negative samples in the data set, the SMOTE and ENN methods were innovatively introduced to effectively optimize data quality through oversampling and noise removal, thereby ensuring the model's efficient recognition of minority samples.
[0090] Specifically, the performance of the PLM-IL-4 model, Meta-IL4 model and four other common deep learning models (LSTM+FC, GRU+FC, CNN+FC and their corresponding integrated attention mechanism variants) were evaluated through comparative experiments to verify the superiority of this application in predicting IL-4 induced peptides. Multiple performance indicators were used in the experiment to evaluate the model, including accuracy (ACC), sensitivity (SN), specificity (SP), Matthews correlation coefficient (MCC), area under the receiver operating characteristic curve (AUC), F1 score and precision.
[0091] During the model training process, the PLM-IL-4 model and Meta-IL4, LSTM+FC, GRU+FC, CNN+FC and other models were trained. The Adam optimizer and binary cross entropy loss function were used in the training process, and the optimal model parameters were selected through hyperparameter tuning, such as the number of GRU layers, the number of neurons and the learning rate.
[0092] The model evaluation process is as follows: in the test set, various indicators such as accuracy (ACC), sensitivity (SN), specificity (SP), Matthews correlation coefficient (MCC), area under the receiver operating characteristic curve (AUC), F1 score and precision are used to evaluate each model.
[0093] In this embodiment, the radar charts of the PLM-IL-4 model and the Meta-IL4 model are compared as follows: Figure 3 As shown in D, the performance comparison between the PLM-IL-4 model and the Meta-IL4 model is shown in the following table: (The values marked with * in the table represent unpublished data)
[0094] ACC SN SP MCC AUC F1 precision Meta-IL4 90.7 93.54 85.41 79.3 * * * PLM-IL-4 93.1 91.72 94.48 86.24 98.67 93.01 94.33
[0095] It can be seen that the PLM-IL-4 model is significantly better than the Meta-IL4 model in key indicators such as specificity (SP), Matthews correlation coefficient (MCC) and AUC; specificity increased from 85.41 to 94.48, MCC increased by 6.94 percentage points, and AUC increased by 0.96 percentage points. The PLM-IL-4 model also showed a significant improvement in the Matthews correlation coefficient (MCC), which increased by 6.94 percentage points, indicating that the PLM-IL-4 model has a more balanced performance between true positives, true negatives, false positives, and false negatives. This further demonstrates the ability of the PLM-IL-4 model to provide more reliable and interpretable classification results. In addition, the area under the curve (AUC) score increased by 0.96 percentage points, demonstrating the superiority of the model in distinguishing categories at various decision thresholds.
[0096] By comparing the performance of Meta-IL4 and other deep learning models, the PLM-IL-4 model performed well in multiple key indicators. The PLM-IL-4 model further improved the performance of key indicators such as specificity (SP), Matthews correlation coefficient (MCC) and AUC.
[0097] In one embodiment, a provided IL-4 induced peptide prediction method based on a protein language model may also include a process of displaying the prediction results online in real time, the specific process including: obtaining a peptide sequence to be predicted, inputting the peptide sequence to be predicted into a protein language model, and extracting target features through the protein language model; inputting the target features into a deep learning model to obtain IL-4 induced peptide prediction results; and displaying the IL-4 induced peptide prediction results on a display interface.
[0098] Specifically, in this embodiment, an online tool platform is also developed and launched for scientific researchers and scholars to predict IL-4 induced peptides. The online tool platform has several core functions such as user-friendly interface, efficient computing power, result display, multiple prediction modes, free access and use. Among them, user-friendly interface: It is a simple and intuitive interface. Users only need to upload peptide sequences to quickly obtain prediction results; efficient computing power: It is reflected in the platform's ability to quickly complete the feature extraction and prediction of protein sequences with the help of powerful server computing resources; result display: The platform provides detailed prediction results, including positive or negative classification of IL-4 induced peptides, and displays relevant evaluation indicators (such as AUC, ACC, etc.); multiple prediction modes: Provide classification and regression prediction modes to meet different scientific research needs; free access and use: Public and free access rights, any user can conduct experiments and verifications on the platform, helping scientific researchers to quickly verify experimental hypotheses and optimize prediction results.
[0099] In one embodiment, in order to verify the superiority of the GRU+FC model in the present application in the IL-4 induced peptide prediction task, the GRU+FC model in the present application is compared with existing models (such as LSTM+FC, GRU+FC, CNN+FC, etc.) for experiments. Specifically, the introduction of the attention mechanism (especially in the GRU+FC+Attention model) improves certain indicators (such as specificity (SP) and accuracy). However, the GRU+FC model still performs well in most indicators including accuracy (ACC), sensitivity (SN), Matthews correlation coefficient (MCC) and area under the receiver operating characteristic curve (AUC), proving its effectiveness in IL-4 induced peptide prediction. The comparison data of LSTM+FC, GRU+FC, CNN+FC and their corresponding integrated attention mechanism variants are shown in the following table:
[0100] LSTM+FC, GRU+FC, CNN+FC and their corresponding variants of integrated attention mechanisms
[0101] Model ACC SN SP MCC AUC F1 Precision LSTM+FC 0.8552 0.8207 0.8897 0.712 0.9452 0.85 0.8815 GRU+FC(Our) 0.931 0.9172 0.9448 0.8624 0.9867 0.9301 0.9433 CNN+FC 0.831 0.8759 0.7862 0.6647 0.9086 0.8383 0.8038 LSTM+FC+Attention 0.7966 0.8345 0.7586 0.5948 0.8728 0.804 0.7756 GRU+FC+Attention 0.9138 0.8759 0.9517 0.83 0.9815 0.9104 0.9478 CNN+FC+Attention 0.802 0.859 0.742 0.6147 0.8886 0.801 0.7658
[0102] The above table compares the performance of six model architectures in detail: LSTM+FC, GRU+FC, CNN+FC, and their respective variants with integrated attention mechanisms. By comparing the performance of different models (such as LSTM+FC, CNN+FC, and GRU+FC) on the same dataset, the significant improvement of the GRU module in indicators such as sensitivity (SN), accuracy (ACC), and Matthews correlation coefficient (MCC) is highlighted. For example, the GRU+FC model improved the accuracy (ACC) by 7.6% and the AUC (area under the curve) by 4.2% compared with LSTM+FC. The analysis results show that the GRU+FC model consistently outperforms other models in predicting IL-4-induced peptides, showing excellent prediction accuracy and robustness, and is considered to be the best model in this study. The experimental results show that the models based on ESM-2 and GRU have shown significant advantages in key indicators such as accuracy, specificity, and AUC, showing high prediction accuracy and model stability.
[0103] Among them, the variant model that introduced the attention mechanism showed improvements in specific indicators (such as accuracy and specificity), but overall, the PLM-IL-4 model still had significant advantages in overall performance.
[0104] Through detailed experimental processes and comparative analysis, it is demonstrated that the GRU module is superior to other models in different scenarios. Its effect is not just a simple performance improvement, but it can accurately capture important contextual information in IL-4-induced peptide sequences, thereby achieving higher prediction accuracy.
[0105] It should be understood that, although the various steps in the above-mentioned flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flow chart may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0106] In one embodiment, Figure 5As shown, a system for predicting IL-4-induced peptides based on a protein language model is provided, comprising: a peptide sequence processing module 510, a feature extraction module 520, a peptide sequence prediction module 530 and a prediction result acquisition module 540, wherein:
[0107] The peptide sequence processing module 510 is used to obtain positive and negative sample peptide sequences of IL-4 induced peptides, determine a training set from each positive and negative sample peptide sequence, and perform data balancing processing on the training set using edited neighbor and synthetic minority class oversampling techniques to obtain each processed IL-4 induced peptide sequence;
[0108] A feature extraction module 520 is used to extract the signal peptide sequence in the IL-4 inducing peptide sequence, input the signal peptide sequence into the protein language model, extract the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, and concatenate the amino acid feature vector and the protein feature vector to obtain a joint feature vector;
[0109] The peptide sequence prediction module 530 is used to normalize the joint feature vector, input the processed feature vector into the deep learning model, and process it through the GRU layer in the deep learning model to obtain the prediction features of the peptide sequence;
[0110] The prediction result acquisition module 540 is used to input the prediction features into the classification model to obtain the category information of the predicted IL-4 induced peptide; the classification model is trained using the cross entropy loss function.
[0111] In one embodiment, the peptide sequence processing module 510 is also used to collect positive and negative sample peptide sequences of IL-4 inducing peptides; wherein the positive sample peptide sequence is a peptide sequence that can induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4 secretion; determine the target division ratio, and divide the positive and negative sample peptide sequences into a training set and a test set according to the target division ratio.
[0112] In one embodiment, the peptide sequence processing module 510 is also used to use the synthetic minority class oversampling technology to oversample the positive sample peptide sequences in the training set to generate synthetic sample peptide sequences; and use the edit neighbor technology to clean the data of the synthetic sample peptide sequences to obtain the processed IL-4 induced peptide sequences.
[0113] In one embodiment, the feature extraction module 520 is also used to determine the signal peptide feature cutoff length, and extract the signal peptide sequence in the IL-4 induced peptide sequence according to the signal peptide feature cutoff length; input the signal peptide sequence into the protein language model, capture the dependency between amino acid residues through the protein language model, and extract the amino acid feature vector of the signal peptide sequence and the context feature vector of the protein sequence.
[0114] In one embodiment, the feature extraction module 520 is also used to determine each translation unit in each signal peptide sequence, and find the amino acid feature vector and protein feature vector corresponding to the translation unit; the amino acid feature vector and protein feature vector corresponding to the same translation unit are concatenated to obtain a joint feature vector of the translation unit; wherein the dimension of the joint feature vector is the sum of the dimensions of the amino acid feature vector and the protein feature vector.
[0115] In one embodiment, the peptide sequence prediction module 530 is also used to use the Z-score normalization method to adjust the mean and standard deviation of each feature in the joint feature vector to obtain a processed feature vector; the processed feature vector is input into the deep learning model, and processed through the convolution layer, pooling layer, and GRU layer in the deep learning model to obtain the predicted features of the peptide sequence.
[0116] In one embodiment, the peptide sequence prediction module 530 is further used to select a hyperparameter tuning method, and use the hyperparameter tuning method to determine the optimal parameter combination; and adjust the parameters in the deep learning model based on the optimal parameter combination.
[0117] In one embodiment, the prediction result acquisition module 540 is also used to obtain the peptide sequence to be predicted, input the peptide sequence to be predicted into the protein language model, and extract the target features through the protein language model; input the target features into the deep learning model to obtain the IL-4 induced peptide prediction results; and display the IL-4 induced peptide prediction results on the display interface.
[0118] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for predicting IL-4 induced peptides based on a protein language model is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0119] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0120] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of a method for predicting IL-4 induced peptides based on a protein language model are implemented.
[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for predicting IL-4-induced peptides based on a protein language model are implemented.
[0122] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0123] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0124] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for predicting IL-4-induced peptides based on a protein language model, characterized in that: The method comprises: Obtaining positive and negative sample peptide sequences of IL-4-induced peptides, determining a training set from each of the positive and negative sample peptide sequences, and performing data balancing processing on the training set using edited neighbors and synthetic minority class oversampling techniques to obtain each processed IL-4-induced peptide sequence; Extracting a signal peptide sequence from the IL-4 inducing peptide sequence, inputting the signal peptide sequence into a protein language model, extracting an amino acid feature vector and a protein feature vector from the signal peptide sequence through the protein language model, and concatenating the amino acid feature vector and the protein feature vector to obtain a joint feature vector; Normalizing the joint feature vector, inputting the processed feature vector into a deep learning model, and processing it through a GRU layer in the deep learning model to obtain prediction features of the peptide sequence; The predicted features are input into a classification model to obtain the category information of the predicted IL-4-induced peptide; the classification model is trained using a cross entropy loss function.
2. The IL-4 induced peptide prediction method based on protein language model according to claim 1, characterized in that: Obtaining positive and negative sample peptide sequences of IL-4-induced peptides, and determining a training set from each of the positive and negative sample peptide sequences, including: Collecting positive and negative sample peptide sequences of IL-4 inducing peptides; wherein the positive sample peptide sequence is a peptide sequence that can induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4 secretion; A target division ratio is determined, and the positive and negative sample peptide sequences are divided into a training set and a test set according to the target division ratio.
3. The IL-4 induced peptide prediction method based on protein language model according to claim 1, characterized in that: The training set is subjected to data balancing processing using the editing neighbor and synthetic minority class oversampling techniques to obtain processed IL-4 inducing peptide sequences, including: Using synthetic minority class oversampling technology to oversample the positive sample peptide sequences in the training set to generate synthetic sample peptide sequences; The synthetic sample peptide sequence was cleaned using the edited neighbor technique to obtain each processed IL-4 inducing peptide sequence.
4. The IL-4 induced peptide prediction method based on protein language model according to claim 1, characterized in that: Extracting the signal peptide sequence in the IL-4 inducing peptide sequence, inputting the signal peptide sequence into a protein language model, and extracting the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, including: Determining a characteristic cutoff length of a signal peptide, and extracting a signal peptide sequence from the IL-4 inducing peptide sequence according to the characteristic cutoff length of the signal peptide; The signal peptide sequence is input into a protein language model, and the dependency relationship between amino acid residues is captured by the protein language model, thereby extracting an amino acid feature vector of the signal peptide sequence and a context feature vector of the protein sequence.
5. The IL-4 induced peptide prediction method based on protein language model according to claim 1, characterized in that: The amino acid feature vector and the protein feature vector are concatenated to obtain a joint feature vector, including: Determine each translation unit in each of the signal peptide sequences, and search for the amino acid feature vector and protein feature vector corresponding to the translation unit; splicing the amino acid feature vector and the protein feature vector corresponding to the same translation unit to obtain a joint feature vector of the translation unit; The dimension of the joint feature vector is the sum of the dimensions of the amino acid feature vector and the protein feature vector.
6. The IL-4 induced peptide prediction method based on protein language model according to claim 1, characterized in that: The joint feature vector is normalized, and the processed feature vector is input into the deep learning model, and processed by the GRU layer in the deep learning model to obtain the prediction features of the peptide sequence, including: Using the Z-score standardization method to adjust the mean and standard deviation of each feature in the joint feature vector to obtain a processed feature vector; The processed feature vector is input into a deep learning model, and processed through a convolutional layer, a pooling layer, and a GRU layer in the deep learning model to obtain prediction features of the peptide sequence.
7. The IL-4 induced peptide prediction method based on protein language model according to claim 6, characterized in that: The method further comprises: Selecting a hyperparameter tuning method, and using the hyperparameter tuning method to determine an optimal parameter combination; The parameters in the deep learning model are adjusted based on the optimal parameter combination.
8. The IL-4 induced peptide prediction method based on protein language model according to claim 1, characterized in that: The method further comprises: Obtaining a peptide sequence to be predicted, inputting the peptide sequence to be predicted into a protein language model, and extracting target features through the protein language model; Inputting the target features into the deep learning model to obtain IL-4 induced peptide prediction results; The IL-4 induced peptide prediction results are displayed on the display interface.
9. A IL-4 induced peptide prediction system based on a protein language model, characterized in that: The system comprises: A peptide sequence processing module is used to obtain positive and negative sample peptide sequences of IL-4-induced peptides, determine a training set from each of the positive and negative sample peptide sequences, and perform data balancing processing on the training set using edited neighbors and synthetic minority class oversampling techniques to obtain each processed IL-4-induced peptide sequence; A feature extraction module is used to extract the signal peptide sequence in the IL-4 inducing peptide sequence, input the signal peptide sequence into a protein language model, extract the amino acid feature vector and protein feature vector in the signal peptide sequence through the protein language model, and concatenate the amino acid feature vector and protein feature vector to obtain a joint feature vector; A peptide sequence prediction module, used to normalize the joint feature vector, input the processed feature vector into a deep learning model, and process it through a GRU layer in the deep learning model to obtain prediction features of the peptide sequence; The prediction result acquisition module is used to input the prediction features into the classification model to obtain the category information of the predicted IL-4 induced peptide; the classification model is trained using the cross entropy loss function.
10. The IL-4 induced peptide prediction system based on protein language model according to claim 9, characterized in that: The peptide sequence processing module is also used to: collect positive and negative sample peptide sequences of IL-4 inducing peptides; wherein the positive sample peptide sequence is a peptide sequence that can induce IL-4 secretion, and the negative sample peptide sequence is a peptide sequence that cannot induce IL-4 secretion; determine a target division ratio, and divide the positive and negative sample peptide sequences into a training set and a test set according to the target division ratio.