Prediction Method, Device and Storage Medium for B-Cell Linear Epitopes
By introducing the deep learning model architecture of amino acid physicochemical feature encoding and CNN+BiGRU+Attention, the problem of failing to fully utilize amino acid position information and physicochemical characteristics in the prior art is solved, and a significant improvement in the prediction of linear epitopes in B cells has been achieved, especially in the epitope prediction of severe acute respiratory syndrome coronavirus, which has achieved high accuracy.
Patent Information
- Application Number
- CN202411297283.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-18
AI Technical Summary
The prior art fails to fully utilize amino acid position information and physicochemical characteristics in B cell linear epitope prediction, resulting in poor prediction accuracy.
采用氨基酸理化特征编码结合CNN+BiGRU+Attention的深度学习模型架构,通过特征编码、卷积模块、双向GRU模块和自注意力模块,提升预测准确性。
The prediction accuracy of B cell linear epitope was significantly improved, especially in epitope prediction of severe acute respiratory syndrome coronavirus, which achieved a 99.9% accuracy rate.
Smart Images

Figure CN119360968B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of immunology and drug development, and particularly relates to a method, device and storage medium for predicting B cell linear epitopes. Background Art
[0002] Adaptive immunity consists of B cells and T cells, which recognize antigens with different specificities. B cells mainly stimulate humoral immunity. B cells play an important immune role by secreting antibodies. B cell antigen determinants (B cell epitopes) are protein fragments present on the surface of antigens that are specifically bound by B cell receptors and cause an immune response in the body. According to the structural characteristics of epitopes, they can be divided into linear epitopes and conformational epitopes. The main difference between them is whether the amino acids constituting the epitope are continuous on the antigen protein. Experimental methods for B cell epitope identification mainly include peptide microarray, X-ray crystallography, and enzyme-linked immunosorbent assay (ELISA), etc. However, these methods are often time-consuming, costly, and inefficient. Using computational methods to predict B cell epitopes can greatly accelerate the screening of B cell epitopes, improve the efficiency of immunotherapy, and accelerate the development of related drugs.
[0003] Currently, many computational methods have been developed for B cell epitope screening. The methods for predicting linear epitopes emerged earlier and have been around since the 1980s. The earliest model based on linear B cell epitopes was the method based on propensity scales. Later, after entering the new century, many machine learning-based models began to emerge one after another, such as ABCPred, BCPreds, SVMTriP, BepiPred-2.0, epitope1D, etc. In recent years, models using deep learning to model linear B cell epitopes have emerged. For example, EpiDope uses context-sensitive embeddings and combines ELMo + LSTM for modeling; NetBCE uses the architectures of CNN and BLSTM to predict linear epitopes; CALIBER uses the ESM-2 pre-trained model and combines a BLSTM model to be able to predict both linear epitopes and conformational epitopes simultaneously. However, current models often only use the amino acid composition, k-mer features, etc. in the sequence during the encoding stage, and do not incorporate the position information of amino acids in the sequence and the physicochemical characteristics of amino acids at each position into the modeling process, which results in poor prediction effects of the models. And most current methods use machine learning methods, and there may be many hidden features that have not been fully explored. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above technical deficiencies, and provide a method, device and storage medium for predicting B cell linear epitopes, so as to solve the technical problem of how to improve the prediction accuracy of B cell linear epitopes in the prior art.
[0005] To achieve the above technical objectives, the technical solution of the present invention provides a method for predicting B-cell linear epitopes, including the following steps:
[0006] S1. Obtain the sequence to be tested, and perform preprocessing on the sequence to be tested respectively to obtain the target input sequence;
[0007] S2. Perform feature encoding on each amino acid residue in the target input sequence obtained in step S1, and output the encoding matrix of each sequence;
[0008] S3. Input the encoding matrix obtained in step S2 into the 2-layer convolutional module in the B-cell linear epitope prediction model, and output the local feature matrix of each sequence;
[0009] S4. Input the local feature matrix obtained in step S3 into the 2-layer or 5-layer BiGRU module in the trained B-cell linear epitope prediction model, and output the global feature matrix of each sequence;
[0010] S5. Input the global feature matrix of each obtained sequence into the 3-layer fully connected layer in the B-cell linear epitope prediction model, and output the probability that the sequence to be tested is a B-cell epitope.
[0011] In any implementation manner, before step S5, it further includes inputting the global feature matrix obtained in step S4 into the self-attention module in the B-cell linear epitope prediction model, and outputting the feature matrix after feature rearrangement;
[0012] Input the feature matrix after feature rearrangement of each obtained sequence into the 3-layer fully connected layer in the B-cell linear epitope prediction model, and output the probability that the sequence to be tested is a B-cell epitope.
[0013] In any implementation manner, in step S1, the preprocessing includes: for a sequence with more than 25 amino acids, intercept the first 25 amino acids as the target input sequence; for a sequence with less than 25 amino acids, append the "X" amino acid at the end of the sequence until the length is 25.
[0014] In any implementation manner, in step S2, the feature encoding includes physical and chemical feature encoding, and 28 features that have been successfully applied in the field of protein binding extracted from the AAIndex database are used as the physical and chemical feature encoding of each amino acid residue in the target input sequence.
[0015] In any embodiment, in step S2, the 28 features that have been successfully applied to the field of protein binding are: CHOP780202, CIDH920103, CIDH920105, FAUJ880109, FAUJ880111, FINA910104, GEIM800104, GEIM800106, KANM800102, KLEP840101, KRIW710101, LIFS790101, MEEJ800101, OOBM770102, PALJ810107, QLAN880123, RACS770103, RADA880108, ROSM880102, SWER830101, ZIMJ680102, ZIMJ680104, AURR980120, MUNV940103, NADH010104, NADH010106, GUYH850105, MIYS990104.
[0016] In any embodiment, in step S3, each layer in the two-layer convolutional module includes a convolutional network, a batch normalization layer, a ReLU layer, and a max pooling layer.
[0017] In any embodiment, in step S5, the self-attention model reallocates the weights of the feature vectors of each subsequence in the global feature matrix of the sequence and outputs a feature matrix with rearranged features.
[0018] In any embodiment, in step S6, the three-layer fully connected layer includes two layers of concatenated linear layers and ReLU layers, and one layer of linear layer and Sigmoid layer; the feature matrix is input into the two layers of concatenated linear layers and ReLU layers; and then through one layer of linear layer and Sigmoid layer, the probability that the input sequence is a B-cell epitope is obtained.
[0019] In addition, the present invention also proposes a prediction device for B-cell linear epitopes, including:
[0020] A data acquisition and preprocessing unit, configured to acquire a sequence to be tested and preprocess the sequence to be tested respectively to obtain a target input sequence;
[0021] A feature encoding unit, configured to perform feature encoding on each amino acid residue in the obtained target input sequence and output an encoding matrix for each sequence;
[0022] A first feature representation unit, configured to input the obtained encoding matrix into a two-layer or five-layer convolutional module in the B-cell linear epitope prediction model and output a local feature matrix for each sequence;
[0023] A second feature representation unit, configured to input the obtained local feature matrix into a two-layer bidirectional GRU module in a trained B-cell linear epitope prediction model, and output a global feature matrix for each sequence;
[0024] A prediction unit, configured to input the feature matrix of each obtained sequence into a three-layer fully-connected layer in the B-cell linear epitope prediction model, and output the probability that the to-be-detected sequence is a B-cell epitope.
[0025] In addition, the present invention also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned prediction method for B-cell linear epitopes is implemented.
[0026] Compared with the prior art, the beneficial effects of the present invention include: The prediction method for B-cell linear epitopes provided by the present invention includes the following steps: obtaining a to-be-detected sequence, and respectively preprocessing the to-be-detected sequence to obtain a target input sequence; S2, performing feature encoding on each amino acid residue in the obtained target input sequence, and outputting an encoding matrix for each sequence; inputting the obtained encoding matrix into a two-layer convolutional module in a B-cell linear epitope prediction model, and outputting a local feature matrix for each sequence; inputting the obtained local feature matrix into a two-layer or five-layer bidirectional GRU module in a trained B-cell linear epitope prediction model, and outputting a global feature matrix for each sequence; inputting the obtained global feature matrix into a self-attention module in the B-cell linear epitope prediction model, and outputting a feature matrix after feature rearrangement; inputting the feature matrix after rearrangement of each obtained sequence into a three-layer fully-connected layer in the B-cell linear epitope prediction model, and outputting the probability that the to-be-detected sequence is a B-cell epitope. The method provided by the present invention introduces amino acid physicochemical feature encoding and a deep learning model architecture of CNN+BiGRU. This method has achieved better performance than the original model in two data environments.
[0027] A further optimized solution introduces an amino acid physicochemical feature encoding and a deep learning model architecture of CNN+BiGRU+Attention. This method has achieved the best results on most data sets of six independent test sets. Finally, this method has successfully predicted two epitope data sets of severe acute respiratory syndrome coronavirus, and among the 10 epitopes of severe acute respiratory syndrome coronavirus-1, the prediction scores of 7 epitopes have reached a level exceeding 99.9%. Results in multiple aspects prove that this method is the best-performing linear B-cell epitope prediction tool at present. The prediction method provided by the present invention significantly improves the prediction accuracy of B-cell linear epitopes. Description of the Drawings
[0028] Figure 1 is a schematic flowchart of the prediction method for B-cell linear epitopes in Embodiment 1 of the present invention.
[0029] Figure 2 It is a comparison chart of the ablation experiment results of Embodiment 1 of the present invention.
[0030] Figure 3 It is the AUROC and AUPR charts of Embodiment 1 of the present invention and the method of the comparative example. Detailed implementation manners
[0031] The "range" disclosed in the present application is defined in the form of a lower limit and an upper limit. A given range is defined by selecting a lower limit and an upper limit, and the selected lower limit and upper limit define the boundary of a specific range. The range defined in this way can include the end values or not include the end values, and can be combined arbitrarily, that is, any lower limit can be combined with any upper limit to form a range. For example, if ranges of 60-120 and 80-110 are listed for a specific parameter, ranges of 60-110 and 80-120 are also contemplated. In addition, if the minimum range values 1 and 2 are listed, and if the maximum range values 3, 4, and 5 are listed, then the following ranges are all contemplated: 1-3, 1-4, 1-5, 2-3, 2-4, and 2-5. In the present application, unless otherwise specified, the numerical range "a-b" represents an abbreviated representation of any real number combination between a and b, where a and b are both real numbers. For example, the numerical range "0-5" means that all real numbers between "0-5" have been fully listed herein, and "0-5" is only an abbreviated representation of these numerical combinations. In addition, when stating that a certain parameter is an integer ≥2, it is equivalent to disclosing that the parameter is, for example, the integers 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, etc.
[0032] If there is no special description, the "including" and "comprising" mentioned in the present application mean open-ended or can also be closed-ended. For example, the "including" and "comprising" can mean that other components not listed can also be included or comprised, or can only include or comprise the listed components.
[0033] If there is no special description, in the present application, the term "or" is inclusive. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, any of the following conditions satisfies the condition "A or B": A is true (or exists) and B is false (or does not exist); A is false (or does not exist) while B is true (or exists); or both A and B are true (or exist).
[0034] This detailed implementation manner provides a method for predicting B cell linear epitopes, including the following steps:
[0035] S1. Obtain the sequence to be tested, and perform preprocessing on the sequence to be tested respectively to obtain the target input sequence; the preprocessing includes: for a sequence with more than 25 amino acids, intercept the first 25 amino acids as the target input sequence; for a sequence with less than 25 amino acids, append "X" amino acids at the end of the sequence until the length reaches 25;
[0036] S2. Perform feature encoding on each amino acid residue in the target input sequence obtained in step S1, and output the encoding matrix of each sequence; the feature encoding includes physical and chemical feature encoding, and 28 features that have been successfully applied in the field of protein binding extracted from the AAIndex database are used as the physical and chemical feature encoding of each amino acid residue in the target input sequence; the 28 features that have been successfully applied in the field of protein binding are: CHOP780202, CIDH920103, CIDH920105, FAUJ880109, FAUJ880111, FINA910104, GEIM800104, GEIM800106, KANM800102, KLEP840101, KRIW710101, LIFS790101, MEEJ800101, OOBM770102, PALJ810107, QLAN880123, RACS770103, RADA880108, ROSM880102, SWER830101, ZIMJ680102, ZIMJ680104, AURR980120, MUNV940103, NADH010104, NADH010106, GUYH850105, MIYS990104;
[0037] S3. Input the encoding matrix obtained in step S2 into the 2-layer convolution module in the B-cell linear epitope prediction model, and output the local feature matrix of each sequence; each layer in the 2-layer convolution module includes a convolutional network, a batch normalization layer, a ReLU layer, and a max pooling layer;
[0038] S4. Input the local feature matrix obtained in step S3 into the 2-layer or 5-layer bidirectional GRU module in the trained B-cell linear epitope prediction model, and output the global feature matrix of each sequence; the 2-layer bidirectional GRU module includes cascaded bidirectional GRU layers;
[0039] S5. Input the feature matrix after rearranging each obtained sequence into the three fully-connected layers in the B-cell linear epitope prediction model, and output the probability that the sequence to be detected is a B-cell epitope; the three fully-connected layers include two cascaded linear layers and ReLU layer, and one linear layer and Sigmoid layer; input the feature matrix into the two cascaded linear layers and ReLU layer; then pass through one linear layer and Sigmoid layer to obtain the probability that the input sequence is a B-cell epitope.
[0040] In some embodiments, before step S5, it further includes: inputting the obtained global feature matrix into the self-attention module in the B-cell linear epitope prediction model, and outputting the feature matrix after feature rearrangement; the self-attention model reallocates the weights of the feature vectors of each subsequence in the global feature matrix of the sequence, and outputs the feature matrix after feature rearrangement; then input the feature matrix after rearranging the features of each obtained sequence into the three fully-connected layers in the B-cell linear epitope prediction model, and output the probability that the sequence to be detected is a B-cell epitope; further, the three fully-connected layers include two cascaded linear layers and ReLU layer, and one linear layer and Sigmoid layer; input the rearranged feature matrix into the two cascaded linear layers and ReLU layer; then pass through one linear layer and Sigmoid layer to obtain the probability that the input sequence is a B-cell epitope.
[0041] In addition, this specific embodiment also proposes a prediction device for B-cell linear epitopes, including:
[0042] A data acquisition and preprocessing unit, configured to acquire a sequence to be detected, and respectively preprocess the sequence to be detected to obtain a target input sequence;
[0043] A feature encoding unit, configured to perform feature encoding on each amino acid residue in the obtained target input sequence, and output an encoding matrix for each sequence;
[0044] A first feature representation unit, configured to input the obtained encoding matrix into the two convolutional modules in the B-cell linear epitope prediction model, and output a local feature matrix for each sequence;
[0045] A second feature representation unit, configured to input the obtained local feature matrix into the two-layer or five-layer bidirectional GRU module in the trained B-cell linear epitope prediction model, and output a global feature matrix for each sequence;
[0046] A prediction unit, configured to input the feature matrix after rearranging each obtained sequence into the three fully-connected layers in the B-cell linear epitope prediction model, and output the probability that the sequence to be detected is a B-cell epitope.
[0047] In some embodiments, the above-mentioned device further includes: a feature rearrangement unit, configured to input the obtained global feature matrix into the self-attention module in the B-cell linear epitope prediction model, and output a feature matrix with rearranged features.
[0048] In addition, this specific embodiment also proposes a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned B-cell linear epitope prediction method is implemented.
[0049] It should be noted that B-cell epitopes (determinants) are composed of hydrophilic amino acids on the surface of an antigen, which are easy to approach the B-cell receptor (BCR) and antibody molecules and be recognized, and are generally composed of 3 to 5 spatially adjacent, continuous or discontinuous amino acid residues.
[0050] An antigenic epitope is a special chemical group in an antigen molecule that determines antigen specificity, including B-cell epitopes and T-cell epitopes. The fragment that can be specifically recognized and bound by the B-cell surface receptor or antibody is a B-cell epitope. According to the structural characteristics of the epitope, the B-cell epitope is a linear epitope or a conformational epitope.
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] In the present invention, terms such as "some embodiments", "this embodiment" and examples, etc. are described, which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0053] If similar descriptions such as "first / second" appear in the application documents, the following description will be added. In the following description, the terms "first\second\third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when permitted, so that the embodiments described here can be implemented in an order other than the one illustrated or described here.
[0054] In this embodiment, the term "and / or" only describes the association relationship of associated objects, indicating that there can be three relationships. For example, object A and / or object B can represent: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0055] The embodiments of the present application will be described below. The embodiments described below are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application. For those embodiments where specific technologies or conditions are not indicated, the technologies or conditions described in the literature in this field or the product instructions are followed. For reagents or instruments whose manufacturers are not indicated, they are all conventional products that can be obtained through commercial purchases.
[0056] Embodiment 1
[0057] Combined with Figure 1 , this embodiment proposes a B-cell linear epitope prediction method based on a convolutional and gated network, including the following steps:
[0058] Step 1: Obtain the sequence to be measured, and perform preprocessing on the sequence to be measured respectively to obtain the target input sequence;
[0059] For sequences with more than 25 amino acids, only the first 25 amino acids are intercepted as the target input sequence;
[0060] For sequences with less than 25 amino acids, the amino acid "X" is appended to the end of the sequence until the length reaches 25.
[0061] Step 2: Perform feature encoding on each amino acid residue in the target input sequence, and output the encoding matrix of each sequence;
[0062] Twenty-eight features that have been successfully applied to the protein binding field are extracted from the AAIndex database as the encoding of the physical and chemical characteristics of each amino acid residue in the target input sequence; the twenty-eight features that have been successfully applied to the protein binding field are: CHOP780202, CIDH920103, CIDH920105, FAUJ880109, FAUJ880111, FINA910104, GEIM800104, GEIM800106, KANM800102, KLEP840101, KRIW710101, LIFS790101, MEEJ800101, OOBM770102, PALJ810107, QLAN880123, RACS770103, RADA880108, ROSM880102, SWER830101, ZIMJ680102, ZIMJ680104, AURR980120, MUNV940103, NADH010104, NADH010106, GUYH850105, MIYS990104; specifically shown in Table 1.
[0063] Step 3: Input the encoding matrix into the 2-layer convolutional module in the trained B-cell linear epitope prediction model, and output the local feature matrix of each sequence;
[0064] Input the encoding matrix of the sequence into a two-layer convolutional layer to obtain the local feature matrix of the epitope;
[0065] Each layer in the two-layer convolutional layer contains a convolutional network, a batch normalization layer, a ReLU layer, and a max pooling layer.
[0066] Step 4: Input the local feature matrix into the two-layer bidirectional GRU module in the trained B-cell linear epitope prediction model, and output the global feature matrix of each sequence;
[0067] Input the local feature matrix of the sequence into a cascaded bidirectional GRU layer to obtain the global feature matrix.
[0068] Step 5: Input the global feature matrix into the self-attention module in the trained B-cell linear epitope prediction model, and output the feature matrix after feature rearrangement;
[0069] Use the self-attention model to reassign the weights of the feature vectors of each subsequence in the global feature matrix of the sequence.
[0070] Step 6: Input the feature matrix after rearrangement of each sequence into the three-layer fully connected layer in the B-cell linear epitope prediction model, and output the probability that the sequence to be measured is a B-cell epitope; input the feature matrix after rearrangement into two cascaded linear layers and a ReLU layer; then pass through one linear layer and a Sigmoid layer to map the value to between [0, 1], representing the probability that the input sequence is a B-cell epitope.
[0071] For example, when the sequence to be measured "PLMESELVIGAVIIRGHLRMA" is input, the model will finally output the probability that it is a B-cell linear epitope as 0.7944.
[0072] Table 1 Encoding of the physical and chemical characteristics of each amino acid
[0073]
[0074] In addition, this embodiment also proposes a prediction device for B-cell linear epitopes, including:
[0075] A data acquisition and preprocessing unit for acquiring the sequence to be measured and preprocessing the sequence to be measured respectively to obtain the target input sequence;
[0076] A feature encoding unit for performing feature encoding on each amino acid residue in the obtained target input sequence and outputting the encoding matrix of each sequence;
[0077] The first feature representation unit is configured to input the obtained encoding matrix into a 2-layer convolutional module in the B-cell linear epitope prediction model, and output a local feature matrix for each sequence;
[0078] The second feature representation unit is configured to input the obtained local feature matrix into a 2-layer bidirectional GRU module in the trained B-cell linear epitope prediction model, and output a global feature matrix for each sequence;
[0079] The feature rearrangement unit is configured to input the obtained global feature matrix into a self-attention module in the B-cell linear epitope prediction model, and output a feature matrix after feature rearrangement;
[0080] The prediction unit is configured to input the feature matrix of each sequence after rearrangement into a 3-layer fully connected layer in the B-cell linear epitope prediction model, and output the probability that the sequence to be tested is a B-cell epitope.
[0081] In addition, this embodiment also proposes a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned B-cell linear epitope prediction method is implemented.
[0082] Embodiment 2
[0083] This embodiment proposes a B-cell linear epitope prediction method based on convolution and gated networks. The difference from Embodiment 1 is that a 5-layer bidirectional GRU module is used in Step 4, and other steps are the same as those in Embodiment 1.
[0084] Embodiment 3
[0085] This embodiment proposes a B-cell linear epitope prediction method based on convolution and gated networks. The difference from Embodiment 2 is that Step 5 of Embodiment 2 is not implemented, and it includes the following steps:
[0086] Step 1, obtain a sequence to be tested, and perform preprocessing on the sequence to be tested respectively to obtain a target input sequence;
[0087] For a sequence with more than 25 amino acids, only the first 25 amino acids are intercepted as the target input sequence;
[0088] For a sequence with less than 25 amino acids, the amino acid "X" is appended to the end of the sequence until the length is 25.
[0089] Step 2, perform feature encoding on each amino acid residue in the target input sequence, and output an encoding matrix for each sequence;
[0090] Twenty-eight features that have been successfully applied in the field of protein binding and extracted from the AAIndex database are used to encode the physical and chemical characteristics of each amino acid residue in the target input sequence; the twenty-eight features that have been successfully applied in the field of protein binding are: CHOP780202, CIDH920103, CIDH920105, FAUJ880109, FAUJ880111, FINA910104, GEIM800104, GEIM800106, KANM800102, KLEP840101, KRIW710101, LIFS790101, MEEJ800101, OOBM770102, PALJ810107, QLAN880123, RACS770103, RADA880108, ROSM880102, SWER830101, ZIMJ680102, ZIMJ680104, AURR980120, MUNV940103, NADH010104, NADH010106, GUYH850105, MIYS990104; specifically shown in Table 1 as follows.
[0091] Step 3: Input the encoding matrix into the 2-layer convolutional module in the trained B-cell linear epitope prediction model, and output the local feature matrix of each sequence.
[0092] Input the encoding matrix of the sequence into the 2-layer convolutional layer to obtain the local feature matrix of the epitope.
[0093] Each layer in the 2-layer convolutional layer contains a convolutional network, a batch normalization layer, a ReLU layer, and a max pooling layer.
[0094] Step 4: Input the local feature matrix into the 5-layer bidirectional GRU module in the trained B-cell linear epitope prediction model, and output the global feature matrix of each sequence.
[0095] Input the local feature matrix of the sequence into the concatenated bidirectional GRU layer to obtain the global feature matrix.
[0096] Step 5: Input the feature matrix of each sequence into the 3-layer fully connected layer in the B-cell linear epitope prediction model, and output the probability that the sequence to be tested is a B-cell epitope; input the rearranged feature matrix into the 2-layer concatenated linear layer and ReLU layer; then pass through 1-layer linear layer and Sigmoid layer to map the value to between [0,1], indicating the probability that the input sequence is a B-cell epitope.
[0097] Experimental verification
[0098] To verify the necessity of each module of the prediction method in this embodiment, we first conducted ablation experiments. We separately selected the convolutional unit and the gating unit as the objects of ablation. Experiments were carried out using the training and test sets of epitope1D. As Figure 2 A, when the number of CNN layers was removed or reduced to 1 layer, the results of our method would decline to varying degrees. In Figure 2 B, we made a relatively wide selection of the number of BiGRU (i.e., bidirectional GRU) layers. The results showed that when the number of BiGRU layers reached 5 layers (i.e., Embodiment 2), all indicators achieved relatively optimal values. After comprehensive comparison, we selected the model with 2 layers of BiGRU retained (our Method 1, i.e., the prediction method of Embodiment 1) and the model with the self-attention module removed (our Method 2, i.e., the prediction method of Embodiment 3) as the final result comparison tools.
[0099] We trained and tested our method using the datasets of epitope1D and NetBCE respectively, which could ensure the absolute fairness of the result comparison. First, we trained and tested with the data of epitope1D. The result comparison is shown in Figure 3 the left. Our Method 1 and our Method 2 achieved 0.9022 and 0.9335 in AUC respectively, exceeding epitope1D by 4.74% and 8.38%. Further, our Method 1 and our Method 2 achieved 0.0526 and 0.0640 in AUC10% respectively, exceeding epitope1D by 56.55% and 90.40%. In AUPR, our Method 1 and our Method 2 obtained scores of 0.7023 and 0.7878 respectively, exceeding epitope1D by 26.61% and 42.02%. Then we retrained and tested the model on the data of NetBCE. The result comparison is shown in Figure 3 the right. Our Method 1 achieved scores of 0.8939, 0.0514 and 0.7224 in AUC, AUC10% and AUPR respectively, leading NetBCE by 7.14%, 46.98% and 22.67% respectively. Our Method 2 achieved scores of 0.9058, 0.0729 and 0.8458 in AUC, AUC10% and AUPR respectively, leading NetBCE by 8.57%, 108.33% and 43.63% respectively. Our method achieved better results in both ways of obtaining datasets, which indicates that it is indeed a relatively good model architecture.
[0100] We compared the results of our method on two independent test sets.
[0101] On the test set of iBCE-EL, our method achieved a significant lead. Our Method 1 (NetBCE) led EpiDope, BepiPred-2.0, CLBTope, and EpitopeVec by 114.75%, 92.01%, 86.52%, and 63.12% respectively in terms of AUC. It led EpiDope, BepiPred-2.0, CLBTope, and EpitopeVec by 3167.26%, 738.89%, 647.93%, and 621.84% respectively in terms of AUC10%. It led EpiDope, BepiPred-2.0, CLBTope, and EpitopeVec by 133.88%, 100.02%, 87.38%, and 76.63% respectively in terms of AUPR.
[0102] In the LBTope dataset, the length of all epitope sequences is 20. The LBTope dataset contains 7,824 positive samples and 7,853 negative samples in total. Our Method 1 (NetBCE) achieved the best prediction performance, with scores of 0.6158, 0.0117, and 0.6077 in terms of AUC, AUC10%, and AUPR respectively. Our Method 1 (NetBCE) led EpiDope, BepiPred-2.0, CLBTope, and EpitopeVec by 40.20%, 17.42%, 8.64%, and 12.30% respectively in terms of AUC. It led EpiDope, BepiPred-2.0, CLBTope, and EpitopeVec by 318.95%, 139.40%, 22.19%, and 60.69% respectively in terms of AUC10%. It led EpiDope, BepiPred-2.0, CLBTope, and EpitopeVec by 34.83%, 19.13%, 5.92%, and 11.58% respectively in terms of AUPR. As shown in Table 2 specifically. Moreover, both Example 2 and Example 3 of our method also achieved better results than the comparative tools.
[0103] Table 2
[0104]
[0105] We collected 10 experimentally verified severe acute respiratory syndrome coronavirus-1 epitopes, and these 10 sequences are all B-cell linear epitopes. Our Method 1 (NetBCE) achieved a probability of over 99.9% of being an epitope on 7 sequences. Our Method 1 (NetBCE) also predicted approximately 80% for the other 3 epitopes. The results are shown in Table 3.
[0106] Table 3
[0107]
[0108] The present invention discloses a method for predicting B-cell linear epitopes based on convolutional and gated networks, including: first encoding the sequence to be tested using a feature matrix based on the biochemical properties of amino acids; then these encoded matrices will be fed into a 2-layer convolutional unit. After coming out of the convolutional module, these matrices will be input into the gated module, which is composed of 2 layers of BiGRU or 5 layers of BiGRU. Subsequently, the sequence will be input into the self-attention module. Finally, it will be input into the output module, and the output module maps the result of each sequence to a value between [0,1] as the probability of whether the final output is a B-cell epitope. The present invention can be used as a method for predicting B-cell linear epitopes with relatively good prediction accuracy.
[0109] The specific embodiments of the present invention described above do not constitute a limitation to the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. A method for predicting B-cell linear epitopes, characterized in that, The following steps are involved: S1, obtaining a sequence to be tested, and preprocessing the sequence to be tested to obtain a target input sequence; S2, performing feature encoding on each amino acid residue in the target input sequence obtained in step S1, and outputting the encoding matrix of each sequence; S3, inputting the encoding matrix obtained in step S2 into a 2-layer convolution module in the B cell linear epitope prediction model, and outputting a local feature matrix of each sequence; S4, inputting the local feature matrix obtained in step S3 into the 2-layer or 5-layer BiGRU module in the trained B cell linear epitope prediction model, and outputting the global feature matrix of each sequence; S5, inputting the obtained global feature matrix of each sequence into the 3-layer fully connected layer in the B cell linear epitope prediction model, and outputting the probability that the sequence to be tested is a B cell epitope; In step S2, the feature coding includes physical and chemical feature coding, and 28 features that have been successfully applied to the protein binding field extracted from the AAIndex database are used as the physical and chemical feature coding of each amino acid residue in the target input sequence.
2. The prediction method of the B cell linear epitope according to claim 1, wherein Before step S5, the method further includes inputting the global feature matrix obtained in step S4 into a self-attention module in the B cell linear epitope prediction model, and outputting a feature matrix after feature rearrangement.
3. The prediction method of the B cell linear epitope according to claim 1, wherein In step S1, the preprocessing includes: for sequences with more than 25 amino acids, the first 25 amino acids are cut off as the target input sequence; for sequences with less than 25 amino acids, "X" is appended to the end of the sequence until the amino acid length is 25.
4. The prediction method of the B cell linear epitope according to claim 1, wherein In step S2, the 28 features that have been successfully applied to the field of protein binding are: CHOP780202, CIDH920103, CIDH920105, FAUJ880109, FAUJ880111, FINA910104, GEIM800104, GEIM800106, KANM800102, KLEP840101, KRIW710101, LIFS790101, MEEJ800 101. OOBM770102, PALJ810107, QLAN880123, RACS770103, RADA880108, ROSM880102, SWER830101, ZIM J680102, ZIMJ680104, AURR980120, MUNV940103, NADH010104, NADH010106, GUYH850105, MIYS990104.
5. The prediction method of the B cell linear epitope according to claim 1, wherein In step S3, each layer in the 2-layer convolution module includes a convolutional network, a batch normalization layer, a ReLU layer, and a maximum pooling layer.
6. The prediction method of the B cell linear epitope according to claim 2, characterized in that The self-attention module redistributes the weights of the feature vectors of each subsequence in the global feature matrix of the sequence and outputs the feature matrix after the features are rearranged.
7. The prediction method of the B-cell linear epitope according to claim 1, characterized in that In step S5, the three-layer fully connected layer includes two serially connected linear layers and ReLU layers, as well as one linear layer and Sigmoid layer; the feature matrix is input into the two serially connected linear layers and ReLU layers; and then through one linear layer and Sigmoid layer, the probability that the input sequence is a B-cell epitope is obtained.
8. A prediction device for B-cell linear epitopes, characterized in that, Comprising: A data acquisition and preprocessing unit, configured to acquire a sequence to be measured, and respectively preprocess the sequence to be measured to obtain a target input sequence; A feature encoding unit, configured to perform feature encoding on each amino acid residue in the obtained target input sequence, and output an encoding matrix for each sequence; the feature encoding includes physical and chemical feature encodings, and 28 features that have been successfully applied to the protein binding field extracted from the AAIndex database are used as the physical and chemical feature encodings for each amino acid residue in the target input sequence; A first feature representation unit, configured to input the obtained encoding matrix into two convolutional modules in the B-cell linear epitope prediction model, and output a local feature matrix for each sequence; A second feature representation unit, configured to input the obtained local feature matrix into two or five-layer bidirectional GRU modules in the trained B-cell linear epitope prediction model, and output a global feature matrix for each sequence; A prediction unit, configured to input the global feature matrix of each obtained sequence into a three-layer fully connected layer in the B-cell linear epitope prediction model, and output the probability that the sequence to be measured is a B-cell epitope.
9. A storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the prediction method of the B-cell linear epitope according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Optical spectrum measuring apparatus
EP0840101A1
Novel coding scheme for predicting compound protein affinity based on deep learning, computer equipment and storage medium
CN112562781A
Antigen presentation prediction model training method, antigen presentation prediction method, equipment and medium
CN115798592A