A general epitope polypeptide screening method and system based on deep learning

Through the deep learning hybrid neural network architecture, combined with BiLSTM, multi-head self-attention and CNN, the universality and accuracy of epitope prediction in the existing technology is solved, efficient and reliable prediction of multiple types of epitopes is achieved, and vaccine design is supported.

CN120089201BActive Publication Date: 2025-07-04INST OF ANIMAL HEALTH GUANGDONG ACADEMY OF AGRI SCI +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510567189.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-04
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The epitope prediction method in the prior art lacks universality, cannot capture the local characteristics and long-range dependencies of the sequence simultaneously, the prediction results lack reliability, and it is difficult to apply to vaccine design.

Method used

Using a hybrid neural network architecture based on deep learning, combining BiLSTM, multi-head self-attention mechanism and CNN, we capture the long-range dependence and local characteristics of the peptide sequence, output epitope prediction results through the multi-tagged classification layer, and introduce epitope conservatism and host similarity analysis.

Benefits of technology

High-precision prediction of multiple types of epitopes is achieved, which improves the accuracy and reliability of predictions, supports simultaneous prediction of B cells, MHC class I and MHC class II epitopes, and provides an important computing tool for vaccine design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089201B_ABST
    Figure CN120089201B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of information technology and discloses a general epitope polypeptide screening method based on deep learning, which includes the following steps: Step 1: Obtain a data set of at least one epitope polypeptide; Step 2: Extract features from the epitope polypeptide sequences in the data set to obtain polypeptide sequence feature vectors; Step 3: Use the polypeptide sequence feature vectors of multiple epitope polypeptides in the data set as a training set to train a preset model; Step 4: Use the model trained in Step 3 to predict the epitopes of polypeptides to obtain the scoring results of the epitopes. This method adopts a hybrid neural network structure and can effectively capture local and global features in polypeptide sequences. The method of the present invention can not only realize the epitope screening of one type of polypeptide, but also realize the epitope screening of multiple polypeptides, and is a general epitope polypeptide screening method. At the same time, the present invention also provides a system based on this method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a general epitope polypeptide screening method based on deep learning. Background Art

[0002] An epitope is a specific region in an antigen molecule that can be recognized by the immune system. Accurately predicting and screening immunogenic epitopes is crucial for vaccine design. Traditional epitope prediction methods mainly rely on experimental verification and simple sequence analysis, which have the problems of high cost and low efficiency. Although there are various computational methods for epitope prediction, these methods often only focus on a single type of epitope and fail to fully utilize the long-range dependence relationship of the sequence, resulting in insufficient prediction accuracy.

[0003] Regarding polypeptide epitope prediction, the following literature can be referred to:

[0004] A patent application with the publication number CN119252379A and the theme of a method, device, medium, and program product for predicting antigen epitope information discloses determining antigen surface information by combining deep learning networks based on antigen surface atom information and antibody surface atom information.

[0005] A patent application with the publication number CN119207572A and the theme of a high-throughput screening method for T cell epitopes based on a deep learning framework uses the T cell receptor CDR3β sequence and at least one antigen epitope peptide sequence as data sources to combine with a deep neural network for training to obtain a method for screening T cell epitopes.

[0006] A patent application with the publication number CN112002374A and the theme of a method for predicting MHC-I epitope affinity based on deep learning predicts epitope affinity by combining the sequence features, hydrophilicity features, polarity features, and position features of polypeptides with a CNN model.

[0007] However, the existing methods mainly have the following deficiencies:

[0008] 1. Most methods only predict a single type of epitope (such as B cell epitopes or MHC class I epitopes), lacking generality;

[0009] 2. The model structure is relatively simple and cannot capture both local features and long-range dependence relationships of the sequence simultaneously;

[0010] 3. The prediction results lack reliability evaluation and are difficult to be directly applied to vaccine design;

[0011] 4. The conservation of epitopes and the similarity with host proteins are not fully considered, which may lead to poor effects of the screened epitopes in practical applications;

[0012] Therefore, there is an urgent need to develop a general epitope polypeptide screening method that can simultaneously predict multiple types of epitopes with high accuracy and reliability. Summary of the Invention

[0013] The object of the present invention is to provide a general epitope polypeptide screening method based on deep learning. This method adopts a neural network architecture, combines BiLSTM, multi-head self-attention mechanism and CNN, can effectively capture the long-range dependence relationship and local features of the sequence, and realizes the accurate prediction of multiple types of epitopes. The method of the present invention can not only realize the epitope screening of one type of polypeptide, but also realize the epitope screening of multiple polypeptides, and is a general epitope polypeptide screening method. The present invention also provides a system based on this method. An epitope is a specific region in an antigen molecule that can be recognized by the immune system and trigger an immune response. Its accurate prediction is crucial for vaccine design, immunotherapy and the development of diagnostic reagents. The present invention uses an innovative deep learning architecture to achieve high-precision prediction of multiple types of epitopes, providing an important computational tool for vaccine development.

[0014] The specific solution of the present invention is as follows:

[0015] A general epitope polypeptide screening method based on deep learning, characterized by comprising the following steps:

[0016] Step 1: Obtain a data set of at least one epitope polypeptide;

[0017] Step 2: Extract features from the epitope polypeptide sequences in the data set to obtain polypeptide sequence feature vectors;

[0018] Step 3: Use the polypeptide sequence feature vectors of multiple epitope polypeptides in the data set as a training set to train a preset model; the model is a deep learning prediction module, and the model has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer;

[0019] The bidirectional long short-term memory network layer is used to capture the long-range dependence relationship of the features within the polypeptide sequence;

[0020] The multi-head self-attention layer is used to capture the attention weights of the features within the polypeptide sequence to determine the correlation between different features;

[0021] The convolutional neural network layer is used to extract high-level features with biological significance;

[0022] The multi-label classification layer is used to integrate the results obtained by the bidirectional long short-term memory network layer, the multi-head self-attention layer, and the convolutional neural network layer and output the epitope prediction results;

[0023] Step 4: Use the model trained in Step 3 to predict the epitopes of the polypeptide, and obtain the scoring results of the epitopes.

[0024] In the above general epitope polypeptide screening method, in Step 4, it further includes evaluating the conservation and similarity of the polypeptide;

[0025] The conservation of the polypeptide refers to the conservation of the epitopes predicted in Step 4 among different strains;

[0026] The similarity of the polypeptide refers to the similarity between the epitopes predicted in Step 4 and the host protein;

[0027] In Step 4, different weights are assigned to the prediction probability, conservation, and similarity of the epitopes preset by the model to obtain the scoring results of epitope prediction.

[0028] In the above general epitope polypeptide screening method, the calculation method of the conservation is weighted entropy conservation measure:

[0029] ;

[0030] where is the occurrence frequency of the amino acid at position , is the importance weight of position , is the physicochemical property similarity adjustment factor.

[0031] The calculation method of the similarity is multi-level similarity analysis:

[0032] ;

[0033] where is the sequence-level similarity, is the structure-level similarity, is the function-level similarity, , , are the weights of each level of similarity respectively.

[0034] In the above general epitope polypeptide screening method, the bidirectional long short-term memory network layer includes: a forward LSTM layer: used to extract features from the start position to the end position of the sequence to obtain forward features; a reverse LSTM layer: used to extract features from the end position to the start position of the sequence to obtain reverse features; a feature fusion layer, used to fuse the forward features and the reverse features to obtain the long-range dependence of the features within the polypeptide sequence.

[0035] The multi-head self-attention layer includes: an attention calculation unit for calculating the attention weights of each position in the sequence with other positions; a multi-head parallel processing unit for parallelly executing multiple groups of independent attention calculations; and a feature concatenation unit for concatenating the attention calculation results of multiple groups to obtain the attention weight distribution of each feature within the polypeptide sequence, thereby determining the correlation between different features.

[0036] The convolutional neural network layer includes: a plurality of convolutional kernels for extracting local features of different scales.

[0037] In the above general epitope polypeptide screening method, step 2 specifically involves digitally encoding and feature vector embedding of the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors; the feature vectors are the sequence information and physicochemical property information of the polypeptide sequences.

[0038] The physicochemical property information includes amino acid composition, dipeptide frequency, amino acid physicochemical properties, BLOSUM substitution matrix features, PSIPRED predicted secondary structure, calculation of relative surface area exposure, Kyte-Doolittle hydrophobicity index, flexibility index, position-specific scoring matrix generated using PSI-BLAST, functional domain information based on the Pfam database, BepiPred antigenicity index, and conservation analysis calculated using Shannon entropy.

[0039] In the above general epitope polypeptide screening method, the epitope polypeptides in the dataset are one or more of B cell epitope polypeptides, MHC class I epitope polypeptides, and MHC class II epitope polypeptides.

[0040] Meanwhile, the present invention also discloses a general epitope polypeptide screening system, including the following components:

[0041] A sequence data acquisition module for acquiring a dataset of at least one epitope polypeptide;

[0042] A sequence feature extraction module for extracting features from the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors;

[0043] A deep learning prediction module for training using the polypeptide sequence feature vectors to obtain a trained deep learning prediction module, and using the trained deep learning prediction module to predict the epitopes of polypeptides;

[0044] The deep learning prediction module has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer;

[0045] The bidirectional long short-term memory network layer is used to capture the long-range dependence relationships of the features within the polypeptide sequence;

[0046] The multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features;

[0047] The convolutional neural network layer is used to extract high-level features; the multi-label classification layer is used to output the epitope prediction result;

[0048] Epitope evaluation module: used to score the epitopes predicted by the deep learning prediction module to obtain the scoring results of the epitopes.

[0049] In the above system, the epitope evaluation module includes a conservation analysis sub-module, a similarity analysis sub-module, and a comprehensive scoring sub-module; the conservation analysis sub-module is used to evaluate the conservation degree of the epitope in different strains; the similarity analysis sub-module is used to evaluate the similarity between the epitope and the host protein; the comprehensive scoring sub-module is used to calculate the comprehensive score according to the prediction probability, conservation, and similarity.

[0050] The beneficial effects of this application are:

[0051] 1. Innovation: Adopting a hybrid neural network architecture of BiLSTM, multi-head self-attention, and CNN can capture both the long-range dependence and local features of the sequence, improving the prediction accuracy;

[0052] 2. Versatility: Supports the simultaneous prediction of B-cell epitopes, MHC class I, and MHC class II epitopes, improving the practicality of the system and providing support for the development of different types of vaccines;

[0053] 3. Efficiency: Automatically learning sequence features through the deep learning model reduces the workload of manual feature engineering;

[0054] 4. Reliability: Introducing epitope conservation and host similarity analysis ensures that the selected epitopes have good immunogenicity and safety. Description of the Drawings

[0055] Figure 1 is the system architecture diagram of the universal epitope polypeptide screening system based on deep learning provided by the present invention;

[0056] Figure 2 is the polypeptide sequence feature extraction flow chart provided by the present invention, showing the complete process of obtaining data from the polypeptide sequence database, experimental verification dataset, and protein structure database, through preprocessing and feature encoding, and finally generating a feature matrix;

[0057] Figure 3 is the deep learning prediction module structure diagram provided by the present invention, showing in detail the connection relationship and data flow between the bidirectional long short-term memory network layer, multi-head self-attention layer, convolutional neural network layer, and multi-label classification layer. Detailed implementation manners

[0058] Next, the embodiments of the present invention will be used to clearly and completely describe the present invention.

[0059] This embodiment provides a specific implementation manner of a general epitope polypeptide screening system based on deep learning.

[0060] Refer to Figure 1 , this system includes a sequence data acquisition module, a sequence feature extraction module, a deep learning prediction module, and an epitope evaluation module. The specific implementation of each module is as follows:

[0061] 1. Sequence data acquisition module

[0062] In this embodiment, the sequence data acquisition module is mainly responsible for obtaining training data from public databases. The implementation process of this module specifically includes the following aspects:

[0063] (1) Selection of data sources

[0064] This system mainly obtains training data from three professional databases:

[0065] Firstly, it is the IEDB (Immune Epitope Database) database. This database is currently the largest immune epitope database in the world and provides a large amount of experimentally verified B-cell epitope and T-cell epitope data;

[0066] Secondly, it is the UniProt database. This database provides comprehensive protein sequence information and functional annotation data, which can be used for sequence feature analysis and model training;

[0067] Thirdly, it is the PDB database. This database provides detailed protein structure information, which can be used for structural feature analysis and verification.

[0068] The data of these three databases are all highly reliable and complete, and can provide high-quality data support for model training.

[0069] (2) Classification of data types

[0070] The system systematically classifies and organizes the obtained data, mainly divided into three types of data sets.

[0071] The first type is the B-cell epitope data set. This data set contains experimentally verified linear B-cell epitope sequences and their label information. Each record contains a complete amino acid sequence and the corresponding epitope / non-epitope marker. The length range of these sequences is between 5 and 30 amino acids. To ensure the training effect, the number of positive samples in the data set is not less than 3,000.

[0072] The second category is the MHC class I epitope dataset, which contains validated MHC class I epitope sequences and their binding strength data. The recorded peptide sequence lengths mainly concentrate between 9 and 11 amino acids, and the binding strength is represented by the IC50 value. The total amount of the dataset is required to be no less than 3,000 samples.

[0073] The third category is the MHC class II epitope dataset, which contains validated MHC class II epitope sequences and their binding strength data. The sequence length ranges from 15 to 25 amino acids, and it also contains the binding data of different alleles. The total amount of the dataset is required to be no less than 4,000 samples.

[0074] (3) Data preprocessing steps

[0075] To ensure data quality, the system conducts comprehensive preprocessing on the acquired raw data, specifically referring to Figure 2 .

[0076] First, perform sequence deduplication. Use professional sequence alignment tools to detect and process exactly the same sequences. During deduplication, duplicate sequences with different labels or experimental verification results are retained, and at the same time, the occurrence frequency information of each sequence is recorded, which can be used for subsequent weight assignment.

[0077] Second, perform length normalization. Different normalization strategies are adopted for different types of epitopes;

[0078] Among them, B cell epitopes are uniformly truncated or padded to 30 amino acid lengths, MHC class I epitopes are uniformly processed to 9 amino acid lengths, and MHC class II epitopes are uniformly processed to 15 amino acid lengths. All sequence padding is performed using the special character "X".

[0079] Then perform format normalization. Uniformly represent all amino acids as single-letter codes, label encoding is unified as positive sample 1 and negative sample 0, and at the same time, convert the binding strength data into standard scores and establish a unified data storage format.

[0080] Finally, in the quality control section, the system will remove sequences containing non-standard amino acids, filter out data with unclear labels, and at the same time check the integrity of the sequences and verify the data consistency to ensure the quality of the dataset. The specific preprocessing algorithm is as follows:

[0081] .

[0082] Among them, the mathematical representations of each processing step are as follows:

[0083] Deduplication processing:

[0084] ;

[0085] Among them, S represents the sequence set in the original data set; Indicates the unique sequence after deduplication; Representation and Sequence The associated tags; Representation sequence The tag collection of Representation sequence The frequency of occurrence in the original dataset.

[0086] Length normalization:

[0087] ;

[0088] Where s represents the input amino acid sequence; Indicates the length of s; Indicates the sequence type (such as B cell epitope, MHC class I epitope, and MHC class II epitope); Indicates the standardized target length, according to the sequence type The number of epitopes depends on the number of cells (e.g., 30 for B cell epitopes, 9 for MHC class I epitopes, and 15 for MHC class II epitopes); Indicates adding ( ) "X" characters (fill characters); Indicates the first part of the sequence s to be intercepted amino acids.

[0089] Format standardization:

[0090] .

[0091] Quality Control:

[0092] .

[0093] in, represents the input amino acid sequence; The label representing the sequence; represents the first amino acids; Represents a set of 20 standard amino acids; Indicates the number of non-standard amino acids in sequence s.

[0094] 2. Sequence feature extraction module

[0095] In this embodiment, the sequence feature extraction module adopts a two-stage processing strategy to convert the amino acid sequence into numerical features that can be processed by the deep learning model. Figure 2 , the implementation process of this module is as follows:

[0096] (1) Sequence digital encoding

[0097] In the stage of sequence digital encoding, a complete amino acid dictionary is first constructed.

[0098] For the standard amino acid encoding, the system establishes a mapping dictionary containing 20 standard amino acids, and maps each amino acid to a unique integer from 0 to 19 in the order of similarity of its physicochemical properties. In terms of special character encoding, the system uses 20 to represent the padding symbol "PAD", 21 to represent the sequence start symbol "START", 22 to represent the sequence end symbol "END", and 23 to represent the unknown amino acid "X".

[0099] In the process of sequence encoding conversion, the system first replaces each residue in the amino acid sequence with the corresponding digital encoding, and adds start and end symbols at the beginning and end of the sequence respectively. For non-standard amino acids appearing in the sequence, the special marker 23 is uniformly used for encoding.

[0100] In the sequence length normalization step, for sequences shorter than the standard length, PAD symbols are filled at the end, and for sequences longer than the standard length, truncation processing is performed. At the same time, the original sequence length information is saved for subsequent processing.

[0101] The specific encoding mapping is as follows:

[0102] ;

[0103] Among them, the map function maps each amino acid character to the corresponding digital encoding:

[0104] .

[0105] (2) Vector embedding processing

[0106] In the stage of vector embedding processing, the construction of the embedding matrix is first carried out.

[0107] In the pre-training process, the system uses millions of protein sequences in the UniProt database as training data, and uses the Word2Vec model for pre-training. The training parameter settings include a window size of 5 and a negative sampling number of 5. In terms of embedding features, the system sets a dimension configuration of n×128, and maps each amino acid into a 128-dimensional vector space. These vectors not only contain the sequence information of amino acids, but also contain feature information such as their physicochemical properties, as described below:

[0108] ① Primary sequence features: Amino acid composition (AAC), dipeptide composition (DPC), physicochemical properties of amino acids (surface accessibility, molecular weight, isoelectric point, instability coefficient), and BLOSUM substitution matrix features;

[0109] ② Structural features: Predict the secondary structure using PSIPRED, calculate the relative surface area exposure (RSA), Kyte-Doolittle hydrophobicity index, and flexibility index;

[0110] ③ Advanced features: Perform conservation analysis using the position-specific scoring matrix (PSSM) generated by PSI-BLAST, functional domain information based on the Pfam database, BepiPred antigenicity index, and Shannon entropy.

[0111] The system constructs a comprehensive amino acid embedding representation by jointly optimizing sequence context information and physicochemical property vectors. Specifically, for each amino acid a, its embedding vector is obtained by the following model:

[0112] .

[0113] Among them, is the context representation learned by Word2Vec, is the physicochemical property vector, and are learnable mapping matrices.

[0114] In terms of parameter optimization, the system supports fine-tuning the parameters of the embedding layer during training, uses the Adam optimizer for parameter updates, and sets the initial learning rate to 0.001. To prevent overfitting, the system adopts multiple regularization measures, including using L2 regularization (coefficient 0.01), adding a Dropout layer (ratio 0.2), and implementing gradient clipping (threshold 5.0). The specific optimization objective function is:

[0115] .

[0116] Among them, is the negative sampling loss of Word2Vec, is the physicochemical property prediction loss, is the L2 regularization term of the parameters, and are balance coefficients, set to 0.5 and 0.01 respectively.

[0117] Through the above pre-training process, a polypeptide sequence feature vector containing sequence information and physicochemical property information of the polypeptide sequence can be obtained.

[0118] 3. Deep learning prediction module

[0119] In this embodiment, the deep learning prediction module adopts an innovative hybrid neural network architecture, which includes four key components. Refer to Figure 3 , and its specific implementation process is as follows:

[0120] (1) Design of the BiLSTM layer (Bidirectional Long Short-Term Memory Network layer)

[0121] The network structure of the BiLSTM layer consists of two main parts. In terms of basic composition, the system sets 128 hidden units for both the forward LSTM unit and the backward LSTM unit, and a total of 2 network depths are set.

[0122] The relationship between the polypeptide sequence feature vector and the BiLSTM layer is a relationship of comprehensive utilization. Specifically, the polypeptide sequence feature vector processed by the feature extraction module , where each is a 128-dimensional feature vector, which contains both sequence information and physicochemical property information, and serves as the input of the BiLSTM layer.

[0123] The forward LSTM unit processes the feature vector in the order from the starting position of the sequence to the ending position to obtain the forward hidden state sequence . The calculation formula of the forward LSTM is:

[0124] ;

[0125] ;

[0126] ;

[0127] ;

[0128] ;

[0129] .

[0130] Among them, , , represent the forget gate, input gate, and output gate respectively, represents the memory cell state, represents the hidden state, represents element-wise multiplication.

[0131] The backward LSTM unit processes the feature vector in the order from the ending position of the sequence to the starting position to obtain the backward hidden state sequence The calculation of the reverse LSTM is similar to that of the forward one, except that the processing order is reversed:

[0132] ;

[0133] ;

[0134] ;

[0135] ;

[0136] ;

[0137] .

[0138] The feature fusion layer concatenates the hidden states of the forward and reverse LSTMs to obtain the final BiLSTM output , where , with a dimension of 256:

[0139] .

[0140] To improve the robustness and performance of the model, the following optimization techniques are also introduced:

[0141] 1. Layer Normalization: Normalize the output of the BiLSTM to accelerate convergence and improve stability:

[0142] ;

[0143] where and are the mean and standard deviation of respectively, and and are learnable scaling and offset parameters.

[0144] 2. Residual Connection: Add residual connections to the multi-layer BiLSTM to alleviate the vanishing gradient problem:

[0145] .

[0146] 3. Gradient Clipping, to prevent gradient explosion:

[0147] ;

[0148] where is the clipping threshold.

[0149] This two-way processing method can capture both forward and backward dependencies in the polypeptide sequence, and is particularly suitable for dealing with long-range interactions in the polypeptide sequence.

[0150] In this way, both sequence information and physicochemical property information are comprehensively considered, enabling the model to learn a more comprehensive epitope feature representation.

[0151] (2) Implementation of the multi-head self-attention layer

[0152] The implementation of the multi-head self-attention layer consists of three key parts for capturing the correlations between different positions in the sequence.

[0153] After the polypeptide sequence feature vector is processed by the BiLSTM layer, it enters the multi-head self-attention layer. The output features of the BiLSTM layer serve as the input to the multi-head self-attention layer. This layer implements the following calculation process:

[0154] First, in terms of the basic structure design, the system configures 8 attention heads, each with a dimension of 32, and the total feature dimension is 256. Each attention head contains three computational units: a query matrix (Q), a key matrix (K), and a value matrix (V).

[0155] For each attention head , three different linear transformation matrices , , are used to generate the query (Query), key (Key), and value (Value) matrices:

[0156] ;

[0157] ;

[0158] .

[0159] Among them, the dimension of each transformation matrix is , and the generated , , all have dimensions of .

[0160] Next, the attention score matrix is calculated:

[0161] ;

[0162] Among them, has a dimension of , and is the scaling factor used to stabilize the training process.

[0163] Finally, the results of the 8 attention heads are concatenated and then passed through a linear transformation:

[0164] ;

[0165] where is a linear transformation matrix with a dimension of .

[0166] To enhance the model performance, the following enhancement mechanisms are also introduced:

[0167] 1. Positional Encoding, which provides positional information for the attention mechanism:

[0168] .

[0169] 2. Attention Dropout, which prevents overfitting:

[0170] .

[0171] 3. Residual connection and layer normalization, which improve the model stability:

[0172] .

[0173] The multi-head self-attention mechanism can learn these different types of correlations simultaneously. Each attention head focuses on capturing different types of dependencies, thus forming a comprehensive understanding of the polypeptide sequence features.

[0174] (3) Configuration of the CNN layer (Convolutional Neural Network layer)

[0175] The implementation of the CNN layer mainly includes three aspects, which are used to extract high-level feature patterns with biological significance from the polypeptide sequence.

[0176] After the polypeptide sequence feature vector is processed by the BiLSTM layer and the multi-head self-attention layer, it enters the CNN layer for high-level feature extraction.

[0177] The output feature of the multi-head self-attention layer is represented as , where each is a 256-dimensional feature vector and serves as the input to the CNN layer.

[0178] First, in terms of the convolutional kernel design, the system uses three different sizes of convolutional kernels ( , , ), and the number of convolutional kernels of each size is 128, and the stride is uniformly set to 1. The convolutional kernels are initialized using the He initialization method (Kaiming initialization), the weight decay coefficient is set to 0.0001, and the bias term is initialized to 0.

[0179] Perform convolution operations using one-dimensional convolutional kernels of three different sizes on as follows:

[0180] ;

[0181] ;

[0182] .

[0183] Among them, , , are the weights of the convolutional kernels of different sizes respectively, , , are the corresponding bias terms, BN represents the batch normalization operation, and * represents the convolution operation.

[0184] Next, perform max pooling operations on each convolution result respectively:

[0185] ;

[0186] ;

[0187] .

[0188] Finally, concatenate the pooling results of the three different scales to form the final high-level features:

[0189] .

[0190] To further enhance the feature representation, a channel attention mechanism (Squeeze-and-Excitation, SE) is introduced:

[0191] ;

[0192] ;

[0193] ;

[0194] Among them, is the channel descriptor obtained by global average pooling, and are the parameter matrices for dimensionality reduction and dimensionality increase, is the channel weight, represents the multiplication of the channel dimension.

[0195] The specific meaning of the high-level features and their relationship with epitope prediction:

[0196] 1. Structural pattern features: Convolution kernels of different sizes can capture the structural patterns in polypeptide sequences. For example:

[0197] Convolution kernel: Captures short-range structural patterns, such as β-turns;

[0198] Convolution kernel: Captures medium-range structural patterns, such as short α-helix fragments;

[0199] Convolution kernel: Captures long-range structural patterns, such as extended β-sheets.

[0200] 2. Epitope feature patterns: Different epitope types usually have specific amino acid combination patterns, and the CNN layer can learn these patterns:

[0201] - B-cell epitopes: Usually have specific surface exposure patterns and hydrophilicity distributions;

[0202] - MHC class I epitopes: Often have specific patterns of anchor residue positions;

[0203] - MHC class II epitopes: Have specific patterns of core binding regions.

[0204] 3. Physicochemical property distributions: CNN can identify the distribution laws of physicochemical properties in polypeptide sequences, such as the spatial arrangements of hydrophobicity, charge distribution, etc.

[0205] (4) Multi-label classification layer

[0206] The multi-label classification layer is the output layer of the entire deep learning prediction module. It integrates the features extracted by the previous layers and outputs the final epitope prediction results.

[0207] The high-level features F output by the CNN layer are globally average pooled to obtain a feature vector Z with a fixed dimension:

[0208] ;

[0209] where Z is a 384-dimensional vector (formed by splicing 128-dimensional features of three types of convolution kernels).

[0210] The system contains three independent classification heads, corresponding to B-cell epitopes (corresponding to the subscript B in the following formula), MHC class I epitopes (corresponding to the subscript I in the following formula), and MHC class II epitopes (corresponding to the subscript II in the following formula):

[0211] ;

[0212] ;

[0213] ;

[0214] ;

[0215] ;

[0216] ;

[0217] Among them, and are the weights and biases of the first fully connected layer( ). and are the weights and biases of the second layer( ).

[0218] For a given polypeptide sequence, the model can simultaneously predict the probabilities of it being different types of epitopes:

[0219] ;

[0220] ;

[0221] ;

[0222] When the probability value exceeds the preset threshold of 0.5, the polypeptide is determined to be an epitope of the corresponding type.

[0223] The model adopts a weighted binary cross-entropy loss function and simultaneously introduces label smoothing and regularization:

[0224] ;

[0225] ;

[0226] ;

[0227] ;

[0228] Among them, represents the true label, represents the class weight (used to handle class imbalance problems), represents the regularization coefficient (0.01).

[0229] To handle label noise and improve the generalization ability of the model, label smoothing technology is introduced:

[0230] ;

[0231] ;

[0232] Among them is the smoothing parameter.

[0233] 4. Epitope Evaluation Module

[0234] In this embodiment, the epitope evaluation module includes three functional units to achieve comprehensive evaluation of epitopes. The specific implementation process is as follows:

[0235] (1) Conservation Analysis Unit

[0236] The implementation of the conservation analysis unit includes two main links. In the sequence collection stage, the system first obtains the target sequence from the NCBI database, requiring the sequence coverage to exceed 90%. During the data acquisition process, the system will perform strict quality control, including sequence integrity check, processing of repetitive sequences, and recording of mutation information.

[0237] In the conservation calculation link, the system adopts a weighted entropy conservation metric, improves the traditional Shannon entropy calculation method, and introduces position weight and physicochemical property similarity weight:

[0238] .

[0239] Where:

[0240] is the occurrence frequency of amino acid a at position ;

[0241] is the importance weight of position , determined based on epitope binding site analysis;

[0242] is the physicochemical property similarity adjustment factor:

[0243] ;

[0244] Where is the BLOSUM62 similarity matrix value between amino acids a and b;

[0245] The conservation value is between 0 and 1, and the larger the value, the higher the conservation. In terms of the scoring standard, the system sets a site conservation threshold of 0.9, allows a maximum of 2 mutation sites, and requires the minimum length of the conserved fragment to be 8 amino acids.

[0246] (2) Similarity Analysis Unit

[0247] The implementation of the similarity analysis unit mainly includes two parts. In the database construction phase, the system integrates data from multiple sources, including host proteome data, UniProt reference protein sets, and known cross-reaction data. The data processing process includes the formatting of the BLAST database, the establishment of indexes, and the implementation of a regular update mechanism.

[0248] In terms of similarity calculation, the system adopts multi-level similarity analysis:

[0249] ;

[0250] Among them:

[0251] is the sequence-level similarity, calculated based on the BLAST algorithm:

[0252] ;

[0253] where S is the BLAST alignment score.

[0254] is the structure-level similarity, calculated based on predicted structural features:

[0255] ;

[0256] where RMSD is the root mean square deviation after structural superposition;

[0257] is the function-level similarity, calculated based on binding site features:

[0258] .

[0259] where J is the Jaccard similarity coefficient, measuring the overlap degree of the binding feature set;

[0260] , , are the weights of each level of similarity, with default settings of 0.5, 0.3, and 0.2 respectively.

[0261] The screening criteria include setting the sequence similarity threshold to 0.5, the number of consecutive matching amino acids not exceeding 5, and requiring at least 2 mismatch sites to ensure the difference between the epitope and the host protein.

[0262] (3) Comprehensive Scoring Unit

[0263] The implementation of the comprehensive scoring unit includes two core parts. In terms of the scoring formula design, the system adopts the method of weighted summation, and the specific formula is:

[0264] 。

[0265] Among them:

[0266] is the epitope probability predicted by the model;

[0267] is the conservation score of the epitope sequence;

[0268] is the similarity score between the epitope and the host protein;

[0269] is the coverage rate of the epitope for the isolate;

[0270] is the accessibility of the epitope on the protein surface;

[0271] to are weight parameters, optimized by grid search, with default values of 0.3, 0.25, 0.2, 0.15, and 0.1 respectively.

[0272] In terms of setting the screening criteria, the system requires that the comprehensive score is greater than 0.7, the prediction probability is greater than 0.8, the conservation is greater than 0.9, and the similarity is less than 0.5. The final result output includes a score ranking, a detailed analysis report, and a visual display.

[0273] Through the above system, the screening of universal epitope polypeptides can be achieved.

[0274] The specific implementation process of the method for screening epitope polypeptides using the above system is as follows:

[0275] (1) Data acquisition and preprocessing

[0276] Obtain training data from a relatively authoritative database; this training data constitutes a training set, a validation set, and a test set, which are divided according to a ratio of 8:1:1;

[0277] In this embodiment, the training data is divided into three types of data sets, namely the B-cell epitope data set, the MHC class I epitope data set, and the MHC class II epitope data set as described above;

[0278] In this step, preprocessing of the above training data is also performed, such as the deduplication process, length normalization process, and sequence quality control operations as described above.

[0279] Perform the following preprocessing on the acquired data:

[0280] ① Sequence deduplication: Use a sequence alignment tool to detect and process exactly the same sequences, and retain the repeated sequences with different labels or experimental verification results;

[0281] ② Length normalization: Different normalization strategies are adopted for different types of epitopes. B-cell epitopes: Uniformly truncated or padded to 30 amino acid lengths; MHC class I epitopes: Uniformly processed to 9 amino acid lengths; MHC class II epitopes: Uniformly processed to 15 amino acid lengths;

[0282] ③ Format normalization: All amino acid representations are unified into single-letter codes, and label encodings are unified into positive sample 1 and negative sample 0;

[0283] ④ Quality control: Remove sequences with too high a proportion of non-standard amino acids, filter data with unclear labels, and check the integrity of the sequences.

[0284] (2) Feature extraction

[0285] Use the sequence feature extraction module to extract features from the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors. This step includes two main processes:

[0286] Sequence digital encoding: Convert the amino acid sequence into digital encoding, including standard amino acid encoding and special character encoding;

[0287] Vector embedding processing: Use the pre-trained embedding matrix to convert the digital encoding into a feature vector containing sequence information and physicochemical property information.

[0288] The specific implementation of feature extraction can be expressed as:

[0289] ;

[0290] ;

[0291] Among them, Embedding is the pre-trained embedding matrix, which maps each encoding to a 128-dimensional feature space.

[0292] (3) Model training

[0293] Use the polypeptide sequence feature vectors of multiple epitope polypeptides in the dataset as the training set to train the preset model; the model is a deep learning prediction module, and the model has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; the bidirectional long short-term memory network layer is used to capture the long-range dependence of features within the polypeptide sequence; the multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features; the convolutional neural network layer is used to extract high-level features; the multi-label classification layer is used to output the epitope prediction results.

[0294] The model training adopted the following method: First, it was initialized by using pre-trained weights to provide good starting values for the model's parameters. During the optimization process, the Adam optimizer was used, with the initial learning rate set to 0.001, and combined with a weighted binary cross-entropy loss function, which considered label smoothing and regularization. The loss function was defined as:

[0295] ;

[0296] where, 、 、 are the loss weight coefficients of B-cell epitopes, MHC class I epitopes, and MHC class II epitopes respectively, used to balance the loss contributions of different types of epitopes; 、 、 are the weighted binary cross-entropy loss functions of B-cell epitopes, MHC class I epitopes, and MHC class II epitopes respectively, used to represent the errors of the model in predicting different types of epitopes; represents the L2 regularization term, which restricts the complexity of the model by penalizing large parameter values to prevent overfitting, where is the regularization coefficient (default is 0.01), are the model parameters, is the square of the L2 norm of the parameter vector.

[0297] In addition, a dynamic batch size and learning rate adjustment strategy were adopted to optimize the training process. In terms of regularization, Dropout (0.2), L2 regularization (0.01), and early stopping strategy were used to avoid overfitting.

[0298] (4) Epitope prediction and evaluation

[0299] The model trained in step 3 was used to predict the epitopes of polypeptides, and at the same time, the conservativeness and similarity of the polypeptides were evaluated to obtain the scoring results of the epitopes.

[0300] In the epitope prediction stage, the system first preprocesses the target protein sequence: the sequence is segmented through a sliding window (9 - 30 amino acids) to ensure coverage of all possible epitope regions; then, digital encoding and vector embedding transformation are performed on each segmented peptide. Next, the processed peptide sequence is predicted using a trained deep learning model to obtain the prediction probabilities of three different types of epitopes. Finally, a prediction probability threshold of 0.8 is set, and only the prediction results exceeding this threshold are retained. For the prediction results, the system also conducts further evaluations: calculates the conservation score of each candidate epitope in different strains; calculates the similarity score between the candidate epitope and the host protein to evaluate the possible cross - reaction risk; evaluates the accessibility of the epitope on the protein surface to ensure it can be recognized by the immune system. The comprehensive score is calculated according to the following formula:

[0301] .

[0302] Finally, the system sorts all candidate epitopes according to the comprehensive score and selects the top 10 epitopes with the highest scores as the final results. The final output of the epitope prediction results includes not only the prediction probabilities, but also evaluation indicators such as conservation, similarity, and accessibility, as well as the comprehensive score and confidence interval, providing a comprehensive reference basis for vaccine design.

Claims

1. A general epitope polypeptide screening method based on deep learning, characterized in that, It includes the following steps: Step 1: Obtain a dataset of at least one epitope polypeptide; Step 2: Extract features from the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors; Step 3: Use the polypeptide sequence feature vectors of multiple epitope polypeptides in the dataset as a training set to train a preset model; the model is a deep learning prediction module, and the model has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer; The bidirectional long short-term memory network layer is used to capture the long-range dependence of features within the polypeptide sequence; The multi-head self-attention layer is used to capture the attention weights of features within the polypeptide sequence to determine the correlation between different features; The convolutional neural network layer is used to extract high-level features with biological significance; The multi-label classification layer is used to integrate the results obtained by the bidirectional long short-term memory network layer, the multi-head self-attention layer, and the convolutional neural network layer and output the epitope prediction result; Step 4: Use the model trained in Step 3 to predict the epitope of the polypeptide to obtain the scoring result of the epitope; In Step 4, it also includes evaluating the conservation and similarity of the polypeptide; The conservation of the polypeptide refers to the conservation of the epitope predicted in Step 4 among different strains; The similarity of the polypeptide refers to the similarity between the epitope predicted in Step 4 and the host protein; In Step 4, different weights are assigned to the prediction probability, conservation, and similarity of the epitope preset by the model to obtain the scoring result of the epitope prediction; The calculation method of the conservation is the weighted entropy conservation metric: ; Among them, is the occurrence frequency of the amino acid at position , is the importance weight of position , is the physicochemical property similarity adjustment factor; The calculation method of the similarity is the multi-level similarity analysis: ; Among them, is the sequence-level similarity, is the structure-level similarity, is the function-level similarity, , , are the weights of the similarities at each level respectively; Step 2 is specifically: digitally encode and embed feature vectors for the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors; the feature vectors are the sequence information and physicochemical property information of the polypeptide sequence; The physicochemical property information includes amino acid composition, dipeptide frequency, amino acid physicochemical properties, BLOSUM substitution matrix features, PSIPRED predicted secondary structure, calculation of relative surface area exposure, Kyte-Doolittle hydrophobicity index, flexibility index, position-specific scoring matrix generated using PSI-BLAST, functional domain information based on the Pfam database, BepiPred antigenicity index, and conservation analysis calculated using Shannon entropy.

2. The general epitope polypeptide screening method according to claim 1, wherein The bidirectional long short-term memory network layer includes: a forward LSTM layer: used to extract features from the start position to the end position of the sequence to obtain forward features; a reverse LSTM layer: used to extract features from the end position to the start position of the sequence to obtain reverse features; a feature fusion layer, used to fuse the forward features and the reverse features to obtain the long-range dependence of features within the polypeptide sequence; The multi-head self-attention layer includes: an attention calculation unit for calculating the attention weights of each position in the sequence with other positions; a multi-head parallel processing unit for parallelly executing multiple groups of independent attention calculations; and a feature concatenation unit for concatenating the attention calculation results of multiple groups to obtain the attention weight distribution of each feature within the polypeptide sequence, thereby determining the correlation between different features. The convolutional neural network layer includes: a plurality of convolutional kernels for extracting local features of different scales.

3. The general epitope polypeptide screening method according to claim 1, characterized in that The epitope polypeptides in the dataset are one or more of B-cell epitope polypeptides, MHC class I epitope polypeptides, and MHC class II epitope polypeptides.

4. A general epitope polypeptide screening system for implementing the method according to any one of claims 1 to 3, characterized in that, It includes the following components: A sequence data acquisition module for acquiring a dataset of at least one epitope polypeptide. A sequence feature extraction module for extracting features from the epitope polypeptide sequences in the dataset to obtain polypeptide sequence feature vectors. A deep learning prediction module for training using the polypeptide sequence feature vectors to obtain a trained deep learning prediction module, and using the trained deep learning prediction module to predict the epitopes of polypeptides. The deep learning prediction module has a bidirectional long short-term memory network layer, a multi-head self-attention layer, a convolutional neural network layer, and a multi-label classification layer. The bidirectional long short-term memory network layer is used to capture the long-range dependence relationships of the features within the polypeptide sequence. The multi-head self-attention layer is used to capture the attention weights of the features within the polypeptide sequence to determine the correlation between different features. The convolutional neural network layer is used to extract high-level features; the multi-label classification layer is used to output the epitope prediction results. An epitope evaluation module for scoring the epitopes predicted by the deep learning prediction module to obtain the scoring results of the epitopes.

5. The system according to claim 4, characterized in that, The epitope evaluation module includes a conservation analysis sub-module, a similarity analysis sub-module, and a comprehensive scoring sub-module; the conservation analysis sub-module is used to evaluate the conservation degree of the epitopes in different strains. The similarity analysis sub-module is used to evaluate the similarity of the epitopes with host proteins; the comprehensive scoring sub-module is used to calculate the comprehensive score based on the prediction probability, conservation, and similarity.

Citation Information

Patent Citations

  • MHC-I epitope affinity prediction method based on deep learning

    CN112002374A

  • T cell epitope high-throughput screening method and device based on deep learning framework

    CN119207572A

  • Method and device for predicting epitope information, medium and program product

    CN119252379A

  • Polypeptide sequence construction method and device, equipment and storage medium

    CN117497054A

  • Method and system for predicting capacity of small open reading window coding polypeptide in non-coding RNA (Ribonucleic Acid)

    CN118038995A