General epitope polypeptide screening method and system based on deep learning

By adopting a deep learning neural network architecture in epitope prediction, combining BiLSTM, multi-head self-attention and CNN, the problem of lack of universality and reliability in the existing technology is solved, and high-precision prediction and evaluation of multiple types of epitopes is achieved, which improves the support capabilities of vaccine design.

CN120089201AActive Publication Date: 2025-06-03INST OF ANIMAL HEALTH GUANGDONG ACADEMY OF AGRI SCI +1

Patent Information

Application Number
CN202510567189.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art lacks universality in epitope prediction, cannot simultaneously capture the local characteristics and long-range dependencies of the sequence, lacks reliability evaluation of the prediction results, and does not fully consider the conservatism of the epitope and similarity to the host protein.

Method used

Using a deep learning-based neural network architecture, combining BiLSTM, multi-head self-attention mechanism and CNN, we capture the long-range dependence and local characteristics of the sequence to achieve accurate prediction of multiple types of epitopes, and improve the reliability of prediction by evaluating the conservation and similarity of the peptide.

Benefits of technology

High-precision prediction of various types of epitopes is achieved, which improves the practicality and reliability of the system, ensures that the screened epitopes have good immunogenicity and safety, and provides an important computing tool for vaccine design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089201A_ABST
    Figure CN120089201A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information, and discloses a general epitope polypeptide screening method based on deep learning, which comprises the following steps: step 1, acquiring a data set of at least one epitope polypeptide; 2, performing feature extraction on the epitope polypeptide sequence in the data set to obtain a polypeptide sequence feature vector; 3, taking the polypeptide sequence feature vectors of the epitope polypeptides in the data set as a training set to train a preset model; and 4, predicting the epitope of the polypeptide by adopting the model trained in the step 3 to obtain a scoring result of the epitope. According to the method, a mixed neural network structure is adopted, and local and global features in a polypeptide sequence can be effectively captured. The method provided by the invention not only can realize epitope screening of one type of polypeptide, but also can realize epitope screening of various polypeptides, and is a universal epitope polypeptide screening method. Meanwhile, the invention further provides a system based on the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a general epitope polypeptide screening method based on deep learning. Background Art

[0002] An epitope is a specific region in an antigen molecule that can be recognized by the immune system. Accurately predicting and screening immunogenic epitopes is crucial for vaccine design. Traditional epitope prediction methods mainly rely on experimental verification and simple sequence analysis, which have the problems of high cost and low efficiency. Although there are various computational methods for epitope prediction, these methods often only focus on a single type of epitope and fail to fully utilize the long-range dependence relationship of sequences, resulting in insufficient prediction accuracy.

[0003] Regarding polypeptide epitope prediction, the following literature can be referred to: A patent application with the publication number CN119252379A and the theme of a method, device, medium, and program product for predicting antigen epitope information discloses determining antigen surface information by combining deep learning networks based on antigen surface atom information and antibody surface atom information.

[0004] A patent application with the publication number CN119207572A and the theme of a high-throughput screening method for T cell epitopes based on a deep learning framework uses the T cell receptor CDR3β sequence and at least one antigen epitope peptide sequence as data sources to combine with a deep neural network for training to obtain a method for screening T cell epitopes.

[0005] A patent application with the publication number CN112002374A and the theme of a method for predicting MHC-I epitope affinity based on deep learning predicts epitope affinity by combining the sequence features, hydrophilicity features, polarity features, and position features of polypeptides with a CNN model.

[0006] However, the methods in the prior art mainly have the following deficiencies: 1. Most methods only predict a single type of epitope (such as B cell epitopes or MHC class I epitopes), lacking generality; 2. The model structure is relatively simple and cannot capture both local features and long-range dependence relationships of sequences simultaneously; 3. The prediction results lack reliability evaluation and are difficult to be directly applied to vaccine design; 4. The conservation of epitopes and their similarity to host proteins are not fully considered, which may lead to poor effects of the screened epitopes in practical applications; Therefore, there is an urgent need to develop a general epitope polypeptide screening method that can simultaneously predict multiple types of epitopes and has high accuracy and reliability. Summary of the Invention

[0007] The objective of the present invention is to provide a general epitope polypeptide screening method based on deep learning. This method adopts a neural network architecture, combines BiLSTM, multi-head self-attention mechanism and CNN, and can effectively capture the long-range dependence relationship and local features of sequences to achieve accurate prediction of various types of epitopes. The method of the present invention can not only screen epitopes of one type of polypeptide, but also screen epitopes of multiple polypeptides, and is a general epitope polypeptide screening method. The present invention also provides a system based on this method. An epitope is a specific region in an antigen molecule that can be recognized by the immune system and trigger an immune response. Its accurate prediction is crucial for vaccine design, immunotherapy and the development of diagnostic reagents. The present invention uses an innovative deep learning architecture to achieve high-precision prediction of various types of epitopes, providing an important computational tool for vaccine development.

[0008] The specific solution of the present invention is as follows: A general epitope polypeptide screening method based on deep learning, characterized by comprising the following steps: Step 1: Obtain a data set of at least one epitope polypeptide; Step 2: Extract features from the epitope polypeptide sequences in the data set to obtain polypeptide sequence feature vectors; Step 3: Use the polypeptide sequence feature vectors of multiple epitope polypeptides in the data set as a training set to train a preset model; the model is a deep learning prediction module, and the model has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; The bidirectional long short-term memory network layer is used to capture the long-range dependence relationship of the features within the polypeptide sequence; The multi-head self-attention layer is used to capture the attention weights of the features within the polypeptide sequence to determine the correlation between different features; The convolutional neural network layer is used to extract high-level features with biological significance; The multi-label classification layer is used to integrate the results obtained by the bidirectional long short-term memory network layer, the multi-head self-attention layer, and the convolutional neural network layer and output the epitope prediction result; Step 4: Use the model trained in Step 3 to predict the epitopes of polypeptides to obtain the scoring results of the epitopes.

[0009] In the above general epitope polypeptide screening method, in Step 4, it also includes evaluating the conservation and similarity of the polypeptide; The conservation of the polypeptide refers to the conservation of the epitopes predicted in Step 4 among different strains; The similarity of the polypeptide refers to the similarity between the epitopes predicted in Step 4 and the host protein; In step 4, different weights are assigned to the predicted probability, conservativeness, and similarity of the epitopes obtained through model presetting to obtain the scoring result of epitope prediction.

[0010] In the above general epitope polypeptide screening method, the calculation method of the conservativeness is weighted entropy conservativeness measure: ; where is the occurrence frequency of the amino acid at position , is the importance weight of position , and is the physicochemical property similarity adjustment factor.

[0011] The calculation method of the similarity is multi-level similarity analysis: ; where is the sequence-level similarity, is the structure-level similarity, is the function-level similarity, , , are the weights of each level of similarity respectively.

[0012] In the above general epitope polypeptide screening method, the bidirectional long short-term memory network layer includes: a forward LSTM layer: used to extract features from the start position to the end position of the sequence to obtain forward features; a backward LSTM layer: used to extract features from the end position to the start position of the sequence to obtain backward features; a feature fusion layer, used to fuse the forward features and the backward features to obtain the long-range dependence relationship of the features within the polypeptide sequence.

[0013] The multi-head self-attention layer includes: an attention calculation unit: used to calculate the attention weights of each position in the sequence with other positions; a multi-head parallel processing unit: used to execute multiple groups of independent attention calculations in parallel; a feature splicing unit, used to splice the results of multiple groups of attention calculations to obtain the attention weight distribution of each feature within the polypeptide sequence, thereby determining the correlation between different features.

[0014] The convolutional neural network layer includes: a plurality of convolutional kernels, used to extract local features of different scales.

[0015] In the above general epitope polypeptide screening method, step 2 is specifically to perform digital encoding and feature vector embedding on the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors; the feature vectors are the sequence information and physicochemical property information of the polypeptide sequences.

[0016] The physicochemical property information includes amino acid composition, dipeptide frequency, amino acid physicochemical properties, BLOSUM substitution matrix characteristics, PSIPRED predicted secondary structure, calculated relative surface area exposure, Kyte-Doolittle hydrophobicity index, flexibility index, position-specific scoring matrix generated using PSI-BLAST, functional domain information based on the Pfam database, BepiPred antigenicity index, and conservation analysis calculated by Shannon entropy.

[0017] In the above general epitope polypeptide screening method, the epitope polypeptides in the dataset are one or more of B-cell epitope polypeptides, MHC class I epitope polypeptides, and MHC class II epitope polypeptides.

[0018] Meanwhile, the present invention also discloses a general epitope polypeptide screening system, including the following components: Sequence data acquisition module: used to acquire a dataset of at least one epitope polypeptide; Sequence feature extraction module: used to extract features from the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors; Deep learning prediction module: used to train with the polypeptide sequence feature vectors to obtain a trained deep learning prediction module, and use the trained deep learning prediction module to predict the epitopes of polypeptides; The deep learning prediction module has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; The bidirectional long short-term memory network layer is used to capture the long-range dependence of features within the polypeptide sequence; The multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features; The convolutional neural network layer is used to extract high-level features; the multi-label classification layer is used to output the epitope prediction results; Epitope evaluation module: used to score the epitopes predicted by the deep learning prediction module to obtain the scoring results of the epitopes.

[0019] In the above system, the epitope evaluation module includes a conservation analysis sub-module, a similarity analysis sub-module, and a comprehensive scoring sub-module; the conservation analysis sub-module is used to evaluate the conservation degree of epitopes in different strains; the similarity analysis sub-module is used to evaluate the similarity between epitopes and host proteins; the comprehensive scoring sub-module is used to calculate the comprehensive score based on prediction probability, conservation, and similarity.

[0020] The beneficial effects of this application are: 1. Innovation: It adopts a hybrid neural network architecture of BiLSTM, multi-head self-attention, and CNN, which can capture both long-range dependencies and local features of sequences simultaneously, improving prediction accuracy; 2. Versatility: It supports the simultaneous prediction of B-cell epitopes, MHC class I, and MHC class II epitopes, enhancing the practicality of the system and providing support for the development of different types of vaccines; 3. Efficiency: It automatically learns sequence features through a deep learning model, reducing the workload of manual feature engineering; 4. Reliability: It introduces epitope conservation and host similarity analysis to ensure that the selected epitopes have good immunogenicity and safety. Brief Description of the Drawings

[0021] Figure 1 is the system architecture diagram of the universal epitope polypeptide screening system based on deep learning provided by the present invention; Figure 2 is the flow chart of polypeptide sequence feature extraction provided by the present invention, showing the complete process of obtaining data from a polypeptide sequence database, an experimental verification dataset, and a protein structure database, through preprocessing and feature encoding, and finally generating a feature matrix; Figure 3 is the structure diagram of the deep learning prediction module provided by the present invention, showing in detail the connection relationship and data flow between the bidirectional long short-term memory network layer, the multi-head self-attention layer, the convolutional neural network layer, and the multi-label classification layer. Detailed Embodiment

[0022] Next, the present invention will be clearly and completely described in conjunction with the embodiments of the present invention.

[0023] This embodiment provides a specific implementation of a universal epitope polypeptide screening system based on deep learning.

[0024] Refer to Figure 1 , the system includes a sequence data acquisition module, a sequence feature extraction module, a deep learning prediction module, and an epitope evaluation module. The specific implementation of each module is as follows: 1. Sequence Data Acquisition Module In this embodiment, the sequence data acquisition module is mainly responsible for obtaining training data from public databases. The implementation process of this module specifically includes the following aspects: (1) Data Source Selection This system mainly obtains training data from three professional databases: First is the IEDB (Immune Epitope Database) database, which is currently the largest immune epitope database in the world and provides a large amount of experimentally verified B-cell epitope and T-cell epitope data; Secondly, there is the UniProt database, which provides comprehensive protein sequence information and functional annotation data and can be used for sequence feature analysis and model training; Thirdly, there is the PDB database, which provides detailed protein structure information and can be used for structural feature analysis and verification.

[0025] The data in these three databases are all highly reliable and complete, and can provide high-quality data support for model training.

[0026] (2) Data type classification The system systematically classifies and organizes the acquired data, mainly divided into three types of data sets.

[0027] The first type is the B-cell epitope data set, which contains experimentally verified linear B-cell epitope sequences and their label information. Each record contains a complete amino acid sequence and the corresponding epitope / non-epitope marker. The lengths of these sequences range from 5 to 30 amino acids. To ensure the training effect, the number of positive samples in the data set is not less than 3,000.

[0028] The second type is the MHC class I epitope data set, which contains verified MHC class I epitope sequences and their binding strength data. The lengths of the peptide sequences in the records mainly concentrate on 9 to 11 amino acids. The binding strength is represented by the IC50 value, and the total amount of the data set is required to be not less than 3,000 samples.

[0029] The third type is the MHC class II epitope data set, which contains verified MHC class II epitope sequences and their binding strength data. The sequence lengths range from 15 to 25 amino acids, and at the same time contain the binding data of different alleles. The total amount of the data set is required to be not less than 4,000 samples.

[0030] (3) Data preprocessing steps To ensure data quality, the system comprehensively preprocesses the acquired raw data, specifically referring to Figure 2 .

[0031] First, perform sequence deduplication. Use professional sequence alignment tools to detect and process exactly the same sequences. During deduplication, duplicate sequences with different markers or experimental verification results are retained, and at the same time, the frequency information of each sequence is recorded, and this information can be used for subsequent weight assignment.

[0032] Secondly, perform length normalization. Different normalization strategies are adopted for different types of epitopes; Among them, B-cell epitopes are uniformly truncated or padded to a length of 30 amino acids, MHC class I epitopes are uniformly processed to a length of 9 amino acids, MHC class II epitopes are uniformly processed to a length of 15 amino acids, and all sequence padding is performed using the special character "X".

[0033] Next, format standardization is performed. All amino acid representations are unified into single-letter codes, label encodings are unified into positive sample 1 and negative sample 0. At the same time, binding strength data is converted into standard scores, and a unified data storage format is established.

[0034] Finally, in the quality control step, the system removes sequences containing non-standard amino acids, filters out data with unclear labels, checks the integrity of the sequences, and verifies the consistency of the data to ensure the quality of the dataset. The specific preprocessing algorithm is as follows: 。

[0035] Among them, the mathematical representations of each processing step are as follows: Duplicate removal: ; Among them, S represents the set of sequences in the original dataset; represents the unique sequence after duplicate removal; represents the label associated with sequence ; represents the set of labels of sequence ; represents the frequency of occurrence of sequence in the original dataset.

[0036] Length standardization: ; Among them, s represents the input amino acid sequence; represents the length of s; represents the sequence type (such as B-cell epitope, MHC class I epitope, and MHC class II epitope); represents the target length after standardization, which depends on the sequence type (such as 30 for B-cell epitopes, 9 for MHC class I epitopes, and 15 for MHC class II epitopes); represents adding ( ) "X" characters (padding characters); represents truncating the first amino acids of sequence s.

[0037] Format standardization: 。

[0038] Quality control: 。

[0039] Among them, represents the input amino acid sequence; represents the label of the sequence; represents the th amino acid in s; represents the set of 20 standard amino acids; represents the number of non-standard amino acids in sequence s.

[0040] 2. Sequence Feature Extraction Module In this embodiment, the sequence feature extraction module adopts a two-stage processing strategy to convert the amino acid sequence into numerical features that can be processed by a deep learning model. Refer to Figure 2 , and the implementation process of this module is as follows: (1) Sequence Digital Encoding In the sequence digital encoding stage, a complete amino acid dictionary is first constructed.

[0041] For the standard amino acid encoding, the system establishes a mapping dictionary containing 20 standard amino acids, and maps each amino acid to a unique integer from 0 to 19 in the order of similarity of its physicochemical properties. In terms of special character encoding, the system uses 20 to represent the padding symbol "PAD", 21 to represent the sequence start symbol "START", 22 to represent the sequence end symbol "END", and 23 to represent the unknown amino acid "X".

[0042] In the sequence encoding conversion process, the system first replaces each residue in the amino acid sequence with its corresponding digital encoding, and adds start and end symbols at the beginning and end of the sequence respectively. For non-standard amino acids appearing in the sequence, they are uniformly encoded with the special marker 23.

[0043] In the sequence length normalization step, for sequences shorter than the standard length, PAD symbols are filled at the end, and for sequences longer than the standard length, they are truncated, while preserving the original sequence length information for subsequent processing.

[0044] The specific encoding mapping is as follows: ; Among them, the map function maps each amino acid character to its corresponding digital encoding: 。

[0045] (2) Vector Embedding Processing In the vector embedding processing stage, the construction of the embedding matrix is first carried out.

[0046] During the pre-training process, the system uses millions of protein sequences in the UniProt database as training data and adopts the Word2Vec model for pre-training. The training parameter settings include a window size of 5 and a negative sampling number of 5. In terms of embedding features, the system sets a dimension configuration of n×128, mapping each amino acid into a 128-dimensional vector space. These vectors contain not only the sequence information of amino acids but also feature information such as their physicochemical properties, as described below: ① Primary sequence features: Amino acid composition (AAC), dipeptide composition (DPC), physicochemical properties of amino acids (surface accessibility, molecular weight, isoelectric point, instability coefficient), and BLOSUM substitution matrix features; ② Structural features: Predicting secondary structure using PSIPRED, calculating relative surface area exposure (RSA), Kyte-Doolittle hydrophobicity index, and flexibility index; ③ Advanced features: Position-specific scoring matrix (PSSM) generated using PSI-BLAST, functional domain information based on the Pfam database, BepiPred antigenicity index, and conservation analysis calculated using Shannon entropy.

[0047] The system constructs a comprehensive amino acid embedding representation by jointly optimizing sequence context information and physicochemical property vectors. Specifically, for each amino acid a, its embedding vector is obtained by the following model: .

[0048] where, is the context representation learned by Word2Vec, is the physicochemical property vector, and are learnable mapping matrices.

[0049] In terms of parameter optimization, the system supports fine-tuning the parameters of the embedding layer during training, using the Adam optimizer for parameter updates, and setting the initial learning rate to 0.001. To prevent overfitting, the system adopts multiple regularization measures, including using L2 regularization (coefficient of 0.01), adding a Dropout layer (ratio of 0.2), and implementing gradient clipping (threshold of 5.0). The specific optimization objective function is: .

[0050] where, is the negative sampling loss of Word2Vec, is the physicochemical property prediction loss, is the L2 regularization term of the parameter, and are the balance coefficients, set to 0.5 and 0.01 respectively.

[0051] Through the above pre-training process, a polypeptide sequence feature vector containing sequence information and physicochemical property information of the polypeptide sequence can be obtained.

[0052] 3. Deep learning prediction module In this embodiment, the deep learning prediction module adopts an innovative hybrid neural network architecture, which includes four key components. Refer to Figure 3 and its specific implementation process is as follows: (1) Design of the BiLSTM layer (bidirectional long short-term memory network layer) The network structure of the BiLSTM layer includes two main parts. In terms of basic composition, the system sets 128 hidden units for both the forward LSTM unit and the backward LSTM unit, and a total of 2 network depths are set.

[0053] The relationship between the polypeptide sequence feature vector and the BiLSTM layer is a relationship of comprehensive utilization. Specifically, the polypeptide sequence feature vector processed by the feature extraction module , where each is a 128-dimensional feature vector, which contains both sequence information and physicochemical property information and serves as the input of the BiLSTM layer.

[0054] The forward LSTM unit processes the feature vector in the order from the starting position of the sequence to the ending position to obtain the forward hidden state sequence . The calculation formula of the forward LSTM is: ; ; ; ; ; .

[0055] Among them, , , represent the forget gate, input gate and output gate respectively, represents the memory cell state, represents the hidden state, represents element-wise multiplication.

[0056] The backward LSTM unit processes the feature vector in the order from the ending position of the sequence to the starting position Process the feature vectors in order to obtain the reverse hidden state sequence . The calculation of the reverse LSTM is similar to that of the forward one, except that the processing order is reversed: ; ; ; ; ; .

[0057] The feature fusion layer concatenates the hidden states of the forward and reverse LSTMs to obtain the final BiLSTM output , where , with a dimension of 256: .

[0058] To improve the robustness and performance of the model, the following optimization techniques are also introduced: 1. Layer Normalization: Normalize the output of the BiLSTM to accelerate convergence and improve stability: ; where and are the mean and standard deviation of respectively, and and are learnable scaling and offset parameters.

[0059] 2. Residual Connection: Add residual connections to the multi-layer BiLSTM to alleviate the problem of vanishing gradients: .

[0060] 3. Gradient Clipping, to prevent gradient explosion: ; where is the clipping threshold.

[0061] This bidirectional processing method can simultaneously capture the forward and backward dependencies in the polypeptide sequence, and is particularly suitable for dealing with long-range interactions in the polypeptide sequence.

[0062] In this way, both sequence information and physicochemical property information are comprehensively considered, enabling the model to learn a more comprehensive epitope feature representation.

[0063] (2) Implementation of the multi-head self-attention layer The implementation of the multi-head self-attention layer consists of three key parts for capturing the correlations between different positions in the sequence.

[0064] After being processed by the BiLSTM layer, the polypeptide sequence feature vector enters the multi-head self-attention layer. The output features of the BiLSTM layer serve as the input to the multi-head self-attention layer. This layer implements the following computational process: First, in terms of the basic structure design, the system configures 8 attention heads, each with a dimension of 32, and the total feature dimension is 256. Each attention head contains three computational units: a query matrix (Q), a key matrix (K), and a value matrix (V).

[0065] For each attention head , three different linear transformation matrices , , are used to generate the query (Query), key (Key), and value (Value) matrices: ; ; .

[0066] Among them, the dimension of each transformation matrix is , and the generated , , all have dimensions of .

[0067] Next, the attention score matrix is calculated: ; where has a dimension of , and is the scaling factor used to stabilize the training process.

[0068] Finally, the results of the 8 attention heads are concatenated and then passed through a linear transformation: ; where is a linear transformation matrix with a dimension of .

[0069] To enhance the model performance, the following enhancement mechanisms are also introduced: 1. Positional Encoding, which provides positional information for the attention mechanism: 。

[0070] 2. Attention Dropout, which prevents overfitting: 。

[0071] 3. Residual connection and layer normalization, which improve the model stability: 。

[0072] The multi - head self - attention mechanism can simultaneously learn these different types of correlations. Each attention head focuses on capturing different types of dependencies, thus forming a comprehensive understanding of the polypeptide sequence features.

[0073] (3) Configuration of the CNN layer (Convolutional Neural Network layer) The implementation of the CNN layer mainly includes three aspects, which are used to extract high - level feature patterns with biological significance from the polypeptide sequence.

[0074] After the polypeptide sequence feature vector is processed by the BiLSTM layer and the multi - head self - attention layer, it enters the CNN layer for high - level feature extraction.

[0075] The output features of the multi - head self - attention layer are represented as , where each is a 256 - dimensional feature vector and serves as the input to the CNN layer.

[0076] First, in terms of the convolutional kernel design, the system uses three different - sized convolutional kernels ( , , ). The number of convolutional kernels of each size is 128, and the stride is uniformly set to 1. The convolutional kernels are initialized using the He initialization method (Kaiming initialization), the weight decay coefficient is set to 0.0001, and the bias term is initialized to 0.

[0077] Perform convolutional operations on using three different - sized one - dimensional convolutional kernels: ; ; 。

[0078] Among them, , , are the weights of convolutional kernels of different sizes respectively, , , is the corresponding bias term, BN represents the batch normalization operation, and * represents the convolution operation.

[0079] Next, perform max pooling operations on each convolution result respectively: ; ; .

[0080] Finally, concatenate the pooling results of the three different scales to form the final high-level features: .

[0081] To further enhance the feature representation, a channel attention mechanism (Squeeze-and-Excitation, SE) is introduced: ; ; ; Among them, is the channel descriptor obtained by global average pooling, and are the parameter matrices for dimensionality reduction and dimensionality increase, is the channel weight, represents the multiplication of the channel dimension.

[0082] The specific meaning of the high-level features and their relationship with epitope prediction:

[0083] 1. Structural pattern features: Convolution kernels of different sizes can capture the structural patterns in the polypeptide sequence. For example: Convolution kernel: Capturing short-range structural patterns such as β-turns; Convolution kernel: Capturing medium-range structural patterns such as short α-helix fragments; Convolution kernel: Capturing long-range structural patterns such as extended β-sheets.

[0084] 2. Epitope feature patterns: Different epitope types usually have specific amino acid combination patterns, and the CNN layer can learn these patterns: - B cell epitopes: Usually have specific surface exposure patterns and hydrophilicity distributions; - MHC class I epitopes: Often have specific patterns of anchor residue positions; - MHC class II epitopes: Have specific patterns of core binding regions.

[0085] 3. Physicochemical property distribution: CNN can identify the distribution patterns of physicochemical properties in polypeptide sequences, such as the spatial arrangements of hydrophobicity, charge distribution, etc.

[0086] (4) Multi-label classification layer

[0087] The multi-label classification layer is the output layer of the entire deep learning prediction module. It integrates the features extracted by the previous layers and outputs the final epitope prediction results.

[0088] The high-level features F output by the CNN layer are globally average pooled to obtain a feature vector Z of a fixed dimension: ; where Z is a 384-dimensional vector (formed by splicing 128-dimensional features of three types of convolutional kernels).

[0089] The system contains three independent classification heads, corresponding to B-cell epitopes (corresponding to the subscript B in the following formula), MHC class I epitopes (corresponding to the subscript I in the following formula), and MHC class II epitopes (corresponding to the subscript II in the following formula): ; ; ; ; ; ; where, and are the weights and biases of the first fully connected layer ( ), and are the weights and biases of the second layer ( ).

[0090] For a given polypeptide sequence, the model can simultaneously predict the probabilities of it being different types of epitopes: ; ; ; When the probability value exceeds the preset threshold of 0.5, it is determined that the polypeptide is an epitope of the corresponding type.

[0091] The model adopts a weighted binary cross-entropy loss function and simultaneously introduces label smoothing and regularization: ; ; ; ; Among them, represents the true label, represents the class weight (used to handle class imbalance problems), represents the regularization coefficient (0.01).

[0092] To handle label noise and improve the model generalization ability, label smoothing technology is introduced: ; ; where is the smoothing parameter.

[0093] 4. Epitope Evaluation Module In this embodiment, the epitope evaluation module includes three functional units to achieve comprehensive evaluation of epitopes. The specific implementation process is as follows: (1) Conservation Analysis Unit The implementation of the conservation analysis unit includes two main links. In the sequence collection stage, the system first obtains the target sequence from the NCBI database, requiring the sequence coverage to exceed 90%. During the data acquisition process, the system will perform strict quality control, including sequence integrity check, processing of repetitive sequences, and recording of mutation information.

[0094] In the conservation calculation link, the system adopts a weighted entropy conservation metric, improves the traditional Shannon entropy calculation method, and introduces position weight and physicochemical property similarity weight: .

[0095] Among them: is the occurrence frequency of amino acid a at position ; is the importance weight of position , determined based on epitope binding site analysis; is the physicochemical property similarity adjustment factor: ; where is the BLOSUM62 similarity matrix value between amino acids a and b; The conservation value ranges from 0 to 1, and the larger the value, the higher the conservation. In terms of the scoring criteria, the system sets a site conservation threshold of 0.9, allowing a maximum of 2 mutation sites, and requiring the minimum length of the conserved fragment to be 8 amino acids.

[0096] (2) Similarity Analysis Unit The implementation of the similarity analysis unit mainly includes two parts. In the database construction phase, the system integrates data from multiple sources, including host proteome data, UniProt reference protein sets, and known cross-reaction data. The data processing process includes formatting the BLAST database, establishing indexes, and implementing a regular update mechanism.

[0097] In terms of similarity calculation, the system adopts multi-level similarity analysis: ;

[0098] Among them: is the sequence-level similarity, calculated based on the BLAST algorithm: ; where S is the BLAST alignment score.

[0099] is the structure-level similarity, calculated based on predicted structural features: ; where RMSD is the root mean square deviation after structure superposition; is the function-level similarity, calculated based on binding site features: .

[0100] where J is the Jaccard similarity coefficient, measuring the overlap degree of the binding feature set; , , are the weights of each level of similarity, with default settings of 0.5, 0.3, and 0.2 respectively.

[0101] The screening criteria include setting the sequence similarity threshold to 0.5, the number of consecutive matching amino acids not exceeding 5, and requiring at least 2 mismatch sites to ensure the difference between the epitope and the host protein.

[0102] (3) Comprehensive scoring unit The implementation of the comprehensive scoring unit includes two core parts. In terms of the scoring formula design, the system adopts the method of weighted summation, and the specific formula is: .

[0103] Among them: is the epitope probability predicted by the model; is the conservation score of the epitope sequence; is the similarity score between the epitope and the host protein; is the coverage rate of the epitope for the isolate; is the accessibility of the epitope on the protein surface; to are weight parameters, optimized through grid search, with default values of 0.3, 0.25, 0.2, 0.15, and 0.1 respectively.

[0104] In terms of setting the screening criteria, the system requires that the comprehensive score is greater than 0.7, the prediction probability is greater than 0.8, the conservativeness is greater than 0.9, and the similarity is less than 0.5. The final result output includes score ranking, a detailed analysis report, and a visual display.

[0105] Through the above system, the screening of universal epitope polypeptides can be achieved.

[0106] The specific implementation process of the method for screening epitope polypeptides using the above system is as follows: (1) Data acquisition and preprocessing Obtain training data from relatively authoritative databases; this training data constitutes the training set, validation set, and test set, which are divided according to the ratio of 8:1:1; In this embodiment, the training data is divided into three types of data sets, namely the B-cell epitope data set, MHC class I epitope data set, and MHC class II epitope data set as described above; In this step, the preprocessing of the above training data is also performed, such as the deduplication processing, length normalization processing, and sequence quality control operations as described above.

[0107] Perform the following preprocessing on the obtained data: ① Sequence deduplication: Use a sequence alignment tool to detect and process exactly the same sequences, and retain the duplicate sequences with different labels or experimental verification results; ② Length normalization: Adopt different normalization strategies for different types of epitopes. B-cell epitopes: Uniformly truncate or pad to 30 amino acid lengths; MHC class I epitopes: Uniformly process to 9 amino acid lengths; MHC class II epitopes: Uniformly process to 15 amino acid lengths; ③ Format normalization: Uniformly represent all amino acids as single-letter codes, and uniformly encode the labels as positive sample 1 and negative sample 0; ④ Quality control: Remove sequences with too high a proportion of non-standard amino acids, filter data with unclear labels, and check the integrity of the sequences.

[0108] (2) Feature extraction The sequence feature extraction module is used to extract features from the epitope polypeptide sequences in the dataset, obtaining polypeptide sequence feature vectors. This step includes two main processes: Sequence digital encoding: converting the amino acid sequence into digital encoding, including standard amino acid encoding and special character encoding; Vector embedding processing: using a pre-trained embedding matrix to convert the digital encoding into a feature vector containing sequence information and physicochemical property information.

[0109] The specific implementation of feature extraction can be expressed as: ; ; where Embedding is the pre-trained embedding matrix, mapping each encoding to a 128-dimensional feature space.

[0110] (3) Model training The polypeptide sequence feature vectors of multiple epitope polypeptides in the dataset are used as the training set to train a preset model; the model is a deep learning prediction module, which has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; the bidirectional long short-term memory network layer is used to capture the long-range dependence of features within the polypeptide sequence; the multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features; the convolutional neural network layer is used to extract high-level features; the multi-label classification layer is used to output the epitope prediction results.

[0111] The following method is adopted for model training: First, it is initialized by using pre-trained weights to provide good initial values for the parameters of the model. During the optimization process, the Adam optimizer is used, and the initial learning rate is set to 0.001, combined with a weighted binary cross-entropy loss function, which takes into account label smoothing and regularization. The loss function is defined as: ;

[0112] where, , , are the loss weight coefficients of B cell epitopes, MHC class I epitopes, and MHC class II epitopes respectively, used to balance the loss contributions of different types of epitopes; , , are the weighted binary cross-entropy loss functions of B cell epitopes, MHC class I epitopes, and MHC class II epitopes respectively, used to represent the errors of the model in predicting different types of epitopes; represents the L2 regularization term, which restricts the complexity of the model by penalizing large parameter values to prevent the model from overfitting, where is the regularization coefficient (default is 0.01), are the model parameters, is the square of the L2 norm of the parameter vector.

[0113] In addition, a dynamic batch size and learning rate adjustment strategy are adopted to optimize the training process. In terms of regularization, Dropout (0.2), L2 regularization (0.01) and early stopping strategy are used to avoid overfitting.

[0114] (4) Epitope prediction and evaluation Use the model trained in step 3 to predict the epitopes of polypeptides, and at the same time evaluate the conservation and similarity of the polypeptides to obtain the scoring results of the epitopes.

[0115] In the epitope prediction stage, the system first preprocesses the target protein sequence: the sequence is segmented by a sliding window (9 - 30 amino acids) to ensure coverage of all possible epitope regions; then each peptide segment obtained by segmentation is digitally encoded and vector embedding transformation is performed. Next, use the trained deep learning model to predict the processed peptide segment sequence to obtain the prediction probabilities of three different types of epitopes. Finally, set the prediction probability threshold to 0.8, and only retain the prediction results that exceed this threshold. For the prediction results, the system will also conduct further evaluations: calculate the conservation score of each candidate epitope in different strains; calculate the similarity score between the candidate epitope and the host protein to evaluate the possible cross - reaction risk; evaluate the accessibility of the epitope on the protein surface to ensure that it can be recognized by the immune system. The comprehensive score is calculated according to the following formula: .

[0116] Finally, the system sorts all candidate epitopes according to the comprehensive score and selects the top 10 epitopes with the highest scores as the final results. The final output of the epitope prediction results includes not only the prediction probabilities, but also evaluation indicators such as conservation, similarity, accessibility, etc., as well as the comprehensive score and confidence interval, providing a comprehensive reference basis for vaccine design.

Claims

1. A method for screening universal epitope peptides based on deep learning, characterized in that: The steps include: Step 1: obtaining a dataset of at least one epitope peptide; Step 2: Extract features of epitope peptide sequences in the data set to obtain peptide sequence feature vectors; Step 3: Using the polypeptide sequence feature vectors of multiple epitope polypeptides in the data set as a training set to train a preset model; the model is a deep learning prediction module, and the model has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; The bidirectional long short-term memory network layer is used to capture the long-range dependencies of features within the polypeptide sequence; The multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features; The convolutional neural network layer is used to extract high-level features with biological significance; The multi-label classification layer is used to integrate the results obtained by the bidirectional long short-term memory network layer, the multi-head self-attention layer, and the convolutional neural network layer and output the epitope prediction results; Step 4: Use the model trained in step 3 to predict the epitope of the peptide and obtain the scoring result of the epitope.

2. The method for screening universal epitope polypeptides according to claim 1, characterized in that: The step 4 also includes evaluating the conservation and similarity of the polypeptides; The conservation of the polypeptide refers to the conservation of the epitope predicted in step 4 among different strains; The similarity of the polypeptide refers to the similarity between the epitope predicted in step 4 and the host protein; In step 4, different weights are assigned to the predicted probability, conservation, and similarity of the epitopes obtained by model preset to obtain a scoring result for epitope prediction.

3. The method for screening universal epitope polypeptides according to claim 2, characterized in that: The conservatism is calculated by weighted entropy conservatism metric: ; in, It's location Amino Acid The frequency of occurrence, It's location The importance weight of is the adjustment factor for the similarity of physical and chemical properties; The similarity calculation method is a multi-level similarity analysis: ; in, is the sequence-level similarity, is the structural similarity, is the function-level similarity, , , are the weights of similarity at each level.

4. The method for screening universal epitope polypeptides according to claim 1, characterized in that: The bidirectional long short-term memory network layer includes: a forward LSTM layer: used to extract features from the starting position of the sequence to the ending position to obtain forward features; a reverse LSTM layer: used to extract features from the ending position of the sequence to the starting position to obtain reverse features; a feature fusion layer, used to fuse the forward features and the reverse features to obtain the long-range dependency relationship of the features in the polypeptide sequence; The multi-head self-attention layer includes: an attention calculation unit: used to calculate the attention weight of each position in the sequence relative to other positions; a multi-head parallel processing unit: used to perform multiple sets of independent attention calculations in parallel; a feature splicing unit, used to splice multiple sets of attention calculation results to obtain the attention weight distribution of each feature in the polypeptide sequence, so as to determine the correlation between different features; The convolutional neural network layer includes: multiple convolution kernels for extracting local features of different scales.

5. The method for screening universal epitope polypeptides according to claim 1, characterized in that: The step 2 specifically comprises digitally encoding and embedding a feature vector of the epitope polypeptide sequence in the data set to obtain a polypeptide sequence feature vector; the feature vector is sequence information and physicochemical property information of the polypeptide sequence; The physicochemical property information includes amino acid composition, binary frequency, amino acid physicochemical properties, BLOSUM substitution matrix characteristics, PSIPRED predicted secondary structure, calculated relative surface area exposure, KyteDoolittle hydrophobicity index, flexibility index, position-specific scoring matrix generated using PSI-BLAST, functional domain information based on the Pfam database, BepiPred antigenicity index, and conservation analysis calculated by Shannon entropy.

6. The method for screening universal epitope polypeptides according to claim 1, characterized in that: The epitope polypeptides in the data set are one or more of B cell epitope polypeptides, MHC class I epitope polypeptides and MHC class II epitope polypeptides.

7. A universal epitope polypeptide screening system, characterized in that: Includes the following components: Sequence data acquisition module: used to obtain a data set of at least one epitope polypeptide; Sequence feature extraction module: used to extract features of epitope peptide sequences in the data set and obtain peptide sequence feature vectors; Deep learning prediction module: used to train with peptide sequence feature vectors to obtain a trained deep learning prediction module, and use the trained deep learning prediction module to predict the epitope of the peptide; The deep learning prediction module has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; The bidirectional long short-term memory network layer is used to capture the long-range dependencies of features within the polypeptide sequence; The multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features; The convolutional neural network layer is used to extract high-level features; the multi-label classification layer is used to output epitope prediction results; Epitope evaluation module: used to score the epitopes predicted by the deep learning prediction module and obtain the epitope scoring results.

8. The system according to claim 7, characterized in that The epitope evaluation module includes a conservation analysis submodule, a similarity analysis submodule and a comprehensive scoring submodule; the conservation analysis submodule is used to evaluate the conservation degree of epitopes in different strains; The similarity analysis submodule is used to evaluate the similarity between the epitope and the host protein; the comprehensive score submodule is used to calculate the comprehensive score based on the prediction probability, conservation and similarity.

Citation Information

Patent Citations

  • MHC-I epitope affinity prediction method based on deep learning

    CN112002374A

  • T cell epitope high-throughput screening method and device based on deep learning framework

    CN119207572A

  • Method and device for predicting epitope information, medium and program product

    CN119252379A

  • Polypeptide sequence construction method and device, equipment and storage medium

    CN117497054A

  • Method and system for predicting capacity of small open reading window coding polypeptide in non-coding RNA (Ribonucleic Acid)

    CN118038995A

Cited By

  • Deep reinforcement learning driven polypeptide drug molecule generation method

    CN120748493A