Protein language model-based B cell epitope prediction method, apparatus and device, and computer program product
By combining ESM-IF1 and ESM2 models, Transformer architecture and deep convolutional neural networks, efficient and accurate prediction of B cell epitopes are achieved, and the problem of insufficient prediction accuracy and generalization capabilities in the existing technology is solved, and the application value in the field of biomedical science is enhanced.
Patent Information
- Application Number
- CN202510717656.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing B-cell epitope prediction methods have shortcomings in prediction accuracy and generalization capabilities, and it is difficult to effectively capture the complex characteristics of epitopes, especially when identifying conformational epitopes, and lack of biological interpretability, which limits its application in vaccine design and experimental verification.
Using a method based on protein language model, by inputting the protein sequence to be predicted into the pre-trained ESM-IF1 model and ESM2 model, multi-dimensional structural embedding vectors and sequence feature vectors are captured, combined with the Transformer architecture and deep convolutional neural network, and using the TIM loss function for training to achieve accurate prediction of B cell epitopes.
It improves the accuracy and generalization ability of B cell epitope prediction, can understand antigen-antibody interactions more comprehensively and deeply, supports the target screening of new vaccines, the identification of therapeutic antibodies, and the development of immune diagnostic reagents, and provides key technical support for precision medicine and personalized treatment.
Smart Images

Figure CN120260680A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biomedical technology, and in particular to a method, device, equipment and computer program product for predicting B cell epitopes based on a protein language model. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the breakthroughs in deep learning and large models, the biomedical field is ushering in profound changes. B cell epitopes, as regions on antigen molecules that can be specifically recognized by B cell receptors (BCRs) or antibodies, play a key role in immune responses. Generally speaking, B cell epitopes are divided into linear (continuous) epitopes and conformational (discontinuous) epitopes. Linear epitopes are composed of continuous residue sequences, while conformational epitopes are composed of residues that are discontinuous in the primary protein sequence but aggregated together due to protein folding. It is estimated that the vast majority of B cell epitopes are conformational epitopes, and linear epitopes account for only a small part. At present, the acquisition of epitopes mainly relies on experimental means, but this process is often time-consuming and costly. Therefore, the accurate prediction of B cell epitopes with the help of computational methods is of great significance to vaccine design, antibody drug development, and immune diagnosis.
[0003] Traditional epitope prediction methods expose many limitations when processing biomedical data. On the one hand, many tools rely only on sequence features or simple artificially designed features, which makes it difficult to fully capture the complex characteristics of epitopes, especially when identifying conformational epitopes. Moreover, these methods often have high false positive and false negative rates when predicting cross-pathogens or new antigens, and their accuracy and robustness need to be improved. On the other hand, epitope recognition is not only affected by sequence information, but also closely related to three-dimensional structure and functional properties. However, existing tools have limited ability to identify complex epitopes. More importantly, the prediction results of existing tools often lack biological interpretability, and it is difficult to clearly reveal which features or regions play a key role in the prediction process, limiting their application value in vaccine design and experimental verification.
[0004] The emergence of deep learning and large models provides new ideas for solving the above problems. Deep learning can automatically learn feature representation from massive data, reduce the workload of manual feature engineering, and effectively improve the efficiency and accuracy of feature extraction. Large models have shown unique advantages in processing complex biomedical data with their powerful representation and learning capabilities. For example, AlphaFold 2 fully demonstrates that rich information can be extracted from protein sequences alone, which is of great significance to protein feature engineering.
[0005] However, current B-cell epitope prediction based on deep learning and large models still faces some problems that need to be solved urgently. Current B-cell epitope prediction methods are mostly based on sequence features, structural features, or machine learning algorithms, but there are still deficiencies in prediction accuracy and generalization ability. First, the complexity and diversity of biomedical data increase the difficulty of model training, and more effective data preprocessing and model optimization methods are urgently needed. Second, the interpretability problem of large models hinders their wide application in the biomedical field. Therefore, developing more transparent and interpretable model architectures and algorithms to comprehensively and deeply understand antigen-antibody interactions has become one of the key challenges in current research. Summary of the Invention
[0006] The purpose of this application is to provide a B-cell epitope prediction method, device, equipment, and computer program product based on a protein language model to solve the technical problem that existing B-cell epitope prediction methods in the prior art still have deficiencies in prediction accuracy and generalization ability. The many technical effects that can be produced by the preferred technical solutions provided in this application are described in detail below.
[0007] To achieve the above purpose, this application provides the following technical solutions: In the first aspect, a B-cell epitope prediction method based on a protein language model provided by this application includes: Input the protein sequence to be predicted into the pre-trained ESM-IF1 model and the pre-trained ESM2 model respectively to capture multi-dimensional structural embedding vectors and sequence feature vectors; Concatenate the sequence feature vectors and the structural embedding vectors to obtain a multi-dimensional data set; Based on the Transformer architecture, train the multi-dimensional data set through an integrated convolutional neural network and calculate the TIM loss function. When the value of the TIM loss function is within a preset range, obtain a deep convolutional neural network model; Use the deep convolutional neural network model to output the prediction probability of each amino acid position of the protein sequence; If the prediction probability is greater than the set threshold, determine that the corresponding amino acid position is an epitope, otherwise it is a non-epitope.
[0008] In some embodiments, inputting the protein sequence to be predicted into the pre-trained ESM-IF1 model to obtain multi-dimensional structural embedding vectors includes: Input the protein sequence to be predicted into the pre-trained ESM-IF1 model for the pre-trained ESM-IF1 model to learn the geometric features in the protein backbone atom coordinates and characterize each residue in the protein sequence; Based on the geometric features and the feature vectors of each residue, a multi-dimensional structural embedding vector is obtained.
[0009] In some embodiments, the protein sequence to be predicted is input into a pre-trained ESM2 model for prediction to obtain a multi-dimensional sequence feature vector, including: The protein sequence to be predicted is input into a pre-trained ESM2 model for the ESM2 model to capture the local sequence information and global sequence information of the protein sequence and generate a corresponding embedding vector for each amino acid; Based on the local sequence information, global sequence information, and the embedding vector corresponding to each amino acid, a multi-dimensional sequence feature vector is obtained.
[0010] In some embodiments, the splicing of the sequence feature vector and the structural embedding vector to obtain a multi-dimensional data set includes: The sequence feature vector and the structural embedding vector are concatenated along the last dimension to generate the multi-dimensional data set. The number of dimensions of the multi-dimensional data set is the sum of the number of dimensions of the sequence feature vector and the structural embedding vector, and the multi-dimensional data set is a three-dimensional tensor based on the batch size, sequence length, and feature dimension in shape.
[0011] In some embodiments, based on the Transformer architecture, the multi-dimensional data set is trained by integrating a convolutional neural network, including: The multi-dimensional data set is input into the Transformer architecture, and the scaled dot-product attention is adopted by the Transformer architecture to obtain the final output result of the multi-head attention as the global feature; The multi-dimensional data set is input into the convolutional neural network, and the convolutional result is output by the residual block of the convolutional neural network as the convolutional feature; Training is performed based on the global feature and the convolutional feature.
[0012] In some embodiments, the calculation of the TIM loss function includes: Based on the multi-dimensional data set and the label corresponding to the multi-dimensional data set, the empirical conditional entropy and empirical marginal entropy of the label are determined, and the cross-entropy is calculated; Based on the empirical conditional entropy, empirical marginal entropy, and the cross-entropy, the TIM loss function is calculated.
[0013] In some embodiments, the use of the deep convolutional neural network model to output the prediction probability of each amino acid position includes: The output of the fully connected layer of the deep convolutional neural network model is calculated using the following formula: ; Among them, represents the output of the fully connected layer of the deep convolutional neural network model, W represents the weight matrix, and b represents the bias vector. represents the feature vector at the i-th amino acid position in the protein sequence to be predicted; Map the output of the fully connected layer to the probability interval to obtain the prediction probability.
[0014] In a second aspect, the present application provides a B-cell epitope prediction device based on a protein language model, including: A feature extraction module for inputting the protein sequence to be predicted into a pre-trained ESM-IF1 model and a pre-trained ESM2 model respectively for prediction to obtain multi-dimensional sequence feature vectors and structure embedding vectors; A splicing module for splicing the sequence feature vectors and the structure embedding vectors to obtain a multi-dimensional data set; A training module for training the multi-dimensional data set based on the Transformer architecture through an integrated convolutional neural network and calculating the TIM loss function, and obtaining a deep convolutional neural network model when the value of the TIM loss function is within a preset range; An output module for using the deep convolutional neural network model to output the prediction probability of each amino acid position of the protein sequence; A prediction module for determining that the corresponding amino acid position is an epitope if the prediction probability is greater than a set threshold, otherwise it is a non-epitope.
[0015] In a third aspect, the present application provides a prediction device, including: One or more processors; A memory for storing one or more computer programs, and one or more of the processors are used to execute the one or more computer programs stored in the memory, so that one or more of the processors execute the method for predicting B-cell epitopes based on a protein language model according to any item in the first aspect.
[0016] In a fourth aspect, the present application provides a computer program product, which is stored on a data carrier and is designed to execute the method for predicting B-cell epitopes based on a protein language model as described above.
[0017] Implementing one of the above technical solutions of this application has the following advantages or beneficial effects: For the B-cell epitope prediction method, device, equipment, and computer program product based on the protein language model of this application, first, the protein sequences to be predicted are respectively input into the pre-trained ESM-IF1 model and the pre-trained ESM2 model, thereby capturing multi-dimensional sequence feature vectors and structural embedding vectors, which can effectively capture the structural and evolutionary features of B-cell epitopes and provide basic data for subsequent analysis; splicing the sequence feature vectors and the structural embedding vectors to achieve the effective combination of multi-modal information to construct a unified epitope feature representation, thereby improving the efficiency and reliability of prediction; then, model training is carried out based on the Transformer architecture. During the training stage, a convolutional neural network is integrated, and the TIM loss function is introduced. The convolutional neural network can effectively extract local features in multi-dimensional datasets, and the TIM loss function helps to optimize the model training process, making the model more targeted and effective during training; finally, the trained model is used to perform probability prediction on each amino acid position of the protein sequence to achieve accurate prediction of different types of B-cell epitopes.
[0018] This application effectively integrates the sequence feature vectors and structural embedding vectors output by the pre-trained ESM-IF1 model and the pre-trained ESM2 model, and combines deep learning algorithms to construct an efficient and accurate deep convolutional neural network model; moreover, the organic integration of the convolutional neural network and Transformer in the deep learning algorithm realizes the in-depth mining and efficient learning of the non-linear features of B-cell epitopes. It has the advantages of high prediction accuracy, strong generalization ability, and good robustness, and can be widely applied to multiple biomedical fields such as the screening of new vaccine targets, the identification of epitopes of therapeutic antibodies, and the development of immunodiagnostic reagents, providing key technical support for precision medicine and personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 is a flowchart of the B-cell epitope prediction method based on the protein language model of the embodiment of this application; Figure 2 is a schematic block diagram of the B-cell epitope prediction method based on the protein language model of the embodiment of this application; Figure 3It is a schematic structural diagram of the B-cell epitope prediction device based on the protein language model according to an embodiment of the present application; Figure 4 It is a schematic structural diagram of the prediction device according to an embodiment of the present application.
[0020] In the figure: 1. Prediction device; 10. Memory; 11. Processor. Detailed implementation manners
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be adopted to implement the present application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. It should be understood that they are only examples of processes, methods, devices, etc. consistent with some aspects of the present application disclosed in detail in the appended claims, and other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present application.
[0022] In the description of the present application, it should be understood that terms such as "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the orientation shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. Terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. The meaning of the term "plurality" is two or more. The terms "connected" and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and may be the internal communication of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0023] In order to illustrate the technical solutions described in the present application, the following will be described through specific examples, and only the parts related to the embodiments of the present application are shown.
[0024] As Figures 1 to 2 shown, the present application provides a B-cell epitope prediction method based on a protein language model, including the following steps S10 to step S50.
[0025] S10. Input the protein sequences to be predicted into the pre-trained ESM-IF1 model and the pre-trained ESM2 model respectively to capture multi-dimensional sequence feature vectors and structural embedding vectors.
[0026] Specifically, ESM-IF1 (Evolutionary Scale Modeling - Inverse Folding 1) is a deep learning model developed by MetaAI that focuses on the protein inverse folding task. The pre-trained ESM-IF1 model is used to solve the inverse folding problem and can recover the native sequence of a protein from the backbone atom coordinates. This model has good generalization ability and can not only handle single-chain proteins but also generalize to multi-chain protein complexes.
[0027] In some embodiments, inputting the protein sequence to be predicted into the pre-trained ESM-IF1 model to obtain multi-dimensional structural embedding vectors may include: Input the protein sequence to be predicted into the pre-trained ESM-IF1 model for the pre-trained ESM-IF1 model to learn the geometric features in the protein backbone atom coordinates and characterize each residue in the protein sequence; Based on the geometric features and the feature vectors of each residue, obtain multi-dimensional structural embedding vectors.
[0028] Specifically, input the protein sequence to be predicted into the pre-trained ESM-IF1 model. The pre-trained ESM-IF1 model is used for protein feature characterization. This model can learn the geometric features in the protein backbone atom coordinates and has rotational and translational invariance, which are crucial for understanding the three-dimensional structure and function of proteins. These features can be used for various protein-related tasks, including predicting protein binding affinity, stability, etc. The ESM-IF1 model can characterize each residue and generate informative feature vectors, with each residue's feature vector having a dimension of 512. Therefore, inputting the protein sequence to be predicted into the pre-trained ESM-IF1 model can capture multi-dimensional structural embedding vectors, and the dimension of this multi-dimensional structural embedding vector is 512.
[0029] ESM2 (Evolutionary Scale Model 2) is a large-scale protein language model developed by Meta AI and is the latest version of the "Evolutionary Scale Modeling" series. Its core goal is to capture evolutionary information and structural features from massive protein sequence data through self-supervised learning. The ESM2 model can better predict the three-dimensional structure of proteins, predict atomic-level structural information, and has a very fast prediction speed.
[0030] In some embodiments, the protein sequence to be predicted is input into a pre-trained ESM2 model for prediction, and a multi-dimensional sequence feature vector is obtained, which may include: The protein sequence to be predicted is input into a pre-trained ESM2 model for the ESM2 model to capture the local sequence information and global sequence information of the protein sequence and generate a corresponding embedding vector for each amino acid; Based on the local sequence information, global sequence information, and the embedding vector corresponding to each amino acid, a multi-dimensional sequence feature vector is obtained.
[0031] Specifically, when the protein sequence to be predicted is input into a pre-trained ESM2 model, the ESM2 model will output an embedding representation of the protein sequence. These embedding representations capture the features of the protein sequence, including local sequence information and global sequence information, and the ESM2 model will generate an embedding vector for each amino acid. These embedding vectors can also be used for various downstream tasks, such as protein structure prediction, interaction prediction, etc.
[0032] By separately inputting the protein sequence to be predicted into a pre-trained ESM-IF1 model and a pre-trained ESM2 model, the captured multi-dimensional structural embedding vector and sequence feature vector characterize the general features of the protein sequence, and at the same time contain spatial structure features and sequence features, which can be used for various downstream tasks and play a key role in the subsequent training and prediction tasks of linear epitopes and conformational epitopes. The dimension of this multi-dimensional sequence feature vector is 1280 dimensions.
[0033] S20. Concatenate the sequence feature vector and the structural embedding vector to obtain a multi-dimensional dataset.
[0034] To achieve the effective combination of multi-modal information, the sequence feature vector and the structural embedding vector can be concatenated.
[0035] In some embodiments, step S20 may include: Concatenate the sequence feature vector and the structural embedding vector along the last dimension to generate the multi-dimensional dataset. The number of dimensions of the multi-dimensional dataset is the sum of the number of dimensions of the sequence feature vector and the structural embedding vector, and the multi-dimensional dataset is a three-dimensional tensor based on the batch size, sequence length, and feature dimension in shape.
[0036] Specifically, the concatenation can be represented by the following formula: ; where X represents the multi-dimensional dataset, Concat represents concatenation along the last dimension, represents the sequence feature vector, represents the structural embedding vector.
[0037] Due to the sequence feature vector being 1280 - dimensional and the structure embedding vector being 512 - dimensional, after concatenating the two, a dataset X with 1792 dimensions is obtained.
[0038] Multidimensional dataset , where \(R\) represents the set of real numbers. The multidimensional dataset X is a three - dimensional tensor with a shape of (batch size , seqlen, dim). batch size represents the batch size, seqlen represents the sequence length, and dim represents the feature dimension.
[0039] S30. Based on the Transformer architecture, the multidimensional dataset is trained by integrating a convolutional neural network, and the TIM loss function is calculated. When the value of the TIM loss function is within a preset range, a deep convolutional neural network model is obtained.
[0040] In some embodiments, training the multidimensional dataset by integrating a convolutional neural network based on the Transformer architecture may include: Inputting the multidimensional dataset into the Transformer architecture, and using scaled dot - product attention in the Transformer architecture to obtain the final output result of multi - head attention as the global feature; Inputting the multidimensional dataset into the convolutional neural network, and outputting the convolutional result as the convolutional feature through the residual blocks of the convolutional neural network; Training based on the global feature and the convolutional feature.
[0041] Specifically, the core of the Transformer architecture is the multi - head attention mechanism, which can comprehensively extract features. Using scaled dot - product attention in the Transformer architecture to obtain the final output result of multi - head attention as the global feature may include: Based on the multidimensional dataset X, initializing the query weight matrix , key weight matrix , and value weight matrix , which are represented by the following formulas: ; ; ; Among them, Q (Query) represents the query vector, which is the vector for asking questions and finding relevant content in the information. K (Key) represents the key vector, which is the information that is matched and compared with the query vector Q. V (Value) represents the value vector, which is the vector actually containing useful information. Based on the query weight matrix the corresponding query vector Q, the key weight matrix the corresponding key vector K, and the value vector V corresponding to the value weight matrix V are used as input vectors to calculate the scaled dot-product attention, which is expressed by the following formula: ; where represents the dimension of the keyword vector. Scaling by serves as a normalization factor to ensure that the increase in dimension does not lead to a significant increase in the dot product.
[0042] The query vector, the key vector, and the value vector are linearly transformed h times, and the scaled dot-product attention is applied to each linearly transformed vector respectively to obtain the output results of each attention head. Here, h represents the number of attention heads.
[0043] The output results of each attention head are concatenated in sequence and linearly transformed to obtain the final output result of the multi-head attention, which is expressed by the following formula: ; where , represents the output weight matrix. For attention head i, , , are the weight matrices of the Q, K, and V vectors respectively, allowing the calculation of , , , represents the query of the i-th attention head, represents the key of the i-th attention head, represents the value of the i-th attention head.
[0044] In parallel with the Transformer architecture, a convolutional neural network (CNN) is also integrated into the training framework. The core component of the convolutional neural network CNN is a residual block composed of two one-dimensional convolutional layers. For the input multi-dimensional data set , the residual block is directly applied to the multi-dimensional data set X, and the calculation formula is as follows: ; ; ; where represents the input of the (l + 1)-th residual block, represents the output of the residual block as convolutional features; Conv1D represents a 1D convolutional operation, , are the weights and biases of the i-th layer within the residual block, ReLU represents the rectified linear unit activation function, and BN represents the batch normalization operation. , represents two consecutive convolutional layers. Specifically, is a one-dimensional convolutional operation that accepts the input data and processes it, is the second one-dimensional convolutional operation that accepts the output of [operation name] as input and further processes it.
[0045] The obtained global features and convolutional features are used as the input to the fully connected layer of the convolutional neural network. The fully connected layer makes classification decisions based on the features to output prediction probabilities.
[0046] In some embodiments, calculating the TIM loss function may include: Based on the multi-dimensional dataset and the labels corresponding to the multi-dimensional dataset, determining the empirical conditional entropy and empirical marginal entropy of the labels, and calculating the cross entropy; Based on the empirical conditional entropy, empirical marginal entropy, and the cross entropy, calculating the TIM loss function.
[0047] Specifically, the TIM loss function is used to calculate all the losses on the training set. The labels of the multi-dimensional dataset X are used to indicate whether the amino acid positions are epitopes. The empirical mutual information between the multi-dimensional dataset X and its corresponding labels Y is divided into two main components. The first component is the empirical conditional entropy of the labels, denoted as ; the second component is the empirical marginal entropy of the labels, denoted as .
[0048] To optimize the binary classification problem, the cross-entropy loss between the labels and the data also needs to be considered, and the cross entropy is denoted as CE.
[0049] The empirical conditional entropy , the empirical marginal entropy and the cross entropy CE are calculated using the following formulas respectively: ; ; ; where, represents the probability of belonging to the k-th class, denotes the size of the dataset X, i is used to index the dataset X, and k is used to index the label categories. denotes the probability that the i-th sequence belongs to the k-th category. denotes the indicator function of whether the sequence indexed by i belongs to the k-th category. K can be set to 2 to achieve binary classification.
[0050] Based on the empirical conditional entropy , the empirical marginal entropy and the cross entropy CE, calculate the TIM loss function, which can be calculated by the following formula: ; where and both denote hyperparameters, which are used to determine the convergence rate of each term in the loss function.
[0051] In some embodiments, can be set, and the TIM loss function at this time is the standard cross entropy theory and the standard mutual information.
[0052] If the value of the TIM loss function does not change significantly for multiple consecutive cycles, it can be considered that the model converges. For example, when the value of the TIM loss function basically does not change significantly at 100 cycles, the model converges at this time, and a deep convolutional neural network model is obtained.
[0053] S40. Use the deep convolutional neural network model to output the prediction probability of each amino acid position of the protein sequence.
[0054] S50. If the prediction probability is greater than the set threshold, determine that the corresponding amino acid position is an epitope, otherwise it is a non-epitope.
[0055] After the feature extraction and model training in the previous steps are completed, the model enters the prediction layer. The prediction layer of the deep convolutional neural network model adopts a fully connected layer architecture, and its role is to map the feature vector output by the previous layer to the final prediction result.
[0056] Specifically, for each amino acid position in the protein sequence to be predicted, the prediction layer will output a probability value, which is used to characterize the possibility that this position is an epitope.
[0057] In some embodiments, step S40 may include: Calculate the output of the fully connected layer of the deep convolutional neural network model using the following formula: ; where denotes the output of the fully connected layer of the deep convolutional neural network model, W denotes the weight matrix, and b denotes the bias vector. A feature vector representing the i-th amino acid position in the protein sequence to be predicted; Map the output of the fully connected layer to a probability interval to obtain the predicted probability.
[0058] Specifically, when calculating the output of the fully connected layer After that, map the output of the fully connected layer To a probability interval, which can be (0, 1), so as to obtain the predicted probability of each amino acid position .
[0059] The set threshold can be 0.618. When the predicted probability ≥0.618, predict that the amino acid position is an epitope; when the predicted probability <0.618, then predict that the amino acid position is a non-epitope, so as to realize the prediction of epitopes in the protein sequence.
[0060] In this application, first, the protein sequence to be predicted is respectively input into the pre-trained ESM-IF1 model and the pre-trained ESM2 model, so as to capture multi-dimensional sequence feature vectors and structural embedding vectors, which can effectively capture the structural features and evolutionary features of B cell epitopes and provide basic data for subsequent analysis; splice the sequence feature vectors and the structural embedding vectors to realize the effective combination of multi-modal information, so as to construct a unified epitope feature representation, thereby improving the efficiency and reliability of prediction; then, based on the Transformer architecture, model training is carried out. In the training stage, a convolutional neural network is integrated and the TIM loss function is introduced. The convolutional neural network can effectively extract local features in multi-dimensional datasets, and the TIM loss function helps to optimize the model training process, making the model more targeted and effective during training; finally, use the trained model to perform probability prediction on each amino acid position of the protein sequence to realize the accurate prediction of different types of B cell epitopes.
[0061] This application effectively integrates the sequence feature vectors and structural embedding vectors output by the pre-trained ESM-IF1 model and the pre-trained ESM2 model, and combines deep learning algorithms to construct an efficient and accurate deep convolutional neural network model; moreover, the organic integration of the convolutional neural network and Transformer in deep learning algorithms realizes the deep mining and efficient learning of the non-linear features of B cell epitopes. It has the advantages of high prediction accuracy, strong generalization ability and good robustness, and can be widely applied to multiple biomedical fields such as the screening of targets for new vaccines, the identification of epitopes of therapeutic antibodies, and the development of immunodiagnostic reagents, providing key technical support for precision medicine and personalized treatment.
[0062] It should be noted that there is not necessarily a certain order among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel, or they can be executed alternately, etc.
[0063] As another aspect of the embodiments of the present invention, the embodiments of the present invention provide a B-cell epitope prediction device based on a protein language model. The B-cell epitope prediction device based on the protein language model according to the embodiments of the present invention can be used as one of the software functional units. The B-cell epitope prediction device based on the protein language model includes a number of instructions, and these instructions are stored in a memory. The processor can access this memory and call the instructions for execution to complete the above-mentioned B-cell epitope prediction method based on the protein language model.
[0064] Please refer to Figure 3 , the B-cell epitope prediction device 300 based on the protein language model includes: A feature extraction module 301, configured to input the protein sequence to be predicted into a pre-trained ESM-IF1 model and a pre-trained ESM2 model respectively for prediction, so as to obtain a multi-dimensional sequence feature vector and a structural embedding vector; A splicing module 302, configured to splice the sequence feature vector and the structural embedding vector to obtain a multi-dimensional data set; A training module 303, configured to train the multi-dimensional data set based on the Transformer architecture through an integrated convolutional neural network, and calculate the TIM loss function. When the value of the TIM loss function is within a preset range, a deep convolutional neural network model is obtained; An output module 304, configured to use the deep convolutional neural network model to output the prediction probability of each amino acid position of the protein sequence; A prediction module 305, configured to determine that the corresponding amino acid position is an epitope if the prediction probability is greater than a set threshold, otherwise it is a non-epitope.
[0065] In some embodiments, the feature extraction module 301 is further configured to: Input the protein sequence to be predicted into the pre-trained ESM-IF1 model, so that the pre-trained ESM-IF1 model can learn the geometric features in the protein backbone atom coordinates and characterize each residue in the protein sequence; Based on the geometric features and the feature vector of each residue, a multi-dimensional structural embedding vector is obtained.
[0066] In some embodiments, the feature extraction module 301 is further configured to: Input the protein sequence to be predicted into the pre-trained ESM2 model, so that the ESM2 model can capture the local sequence information and global sequence information of the protein sequence and generate corresponding embedding vectors for each amino acid; Based on the local sequence information, global sequence information, and the embedding vectors corresponding to each amino acid, obtain a multi-dimensional sequence feature vector.
[0067] In some embodiments, the splicing module 302 is further configured to: Splice the sequence feature vector and the structure embedding vector along the last dimension to generate the multi-dimensional data set. The number of dimensions of the multi-dimensional data set is the sum of the number of dimensions of the sequence feature vector and the structure embedding vector, and the multi-dimensional data set is a three-dimensional tensor whose shape is based on the batch size, sequence length, and feature dimension.
[0068] In some embodiments, the training module 303 is further configured to: Input the multi-dimensional data set into the Transformer architecture, and use scaled dot-product attention through the Transformer architecture to obtain the final output result of the multi-head attention as the global feature; Input the multi-dimensional data set into the convolutional neural network, and output the convolution result as the convolution feature through the residual block of the convolutional neural network; Train based on the global feature and the convolution feature.
[0069] In some embodiments, the training module 303 is further configured to: Based on the multi-dimensional data set and the label corresponding to the multi-dimensional data set, determine the empirical conditional entropy and empirical marginal entropy of the label, and calculate the cross entropy; Calculate the TIM loss function based on the empirical conditional entropy, empirical marginal entropy, and the cross entropy.
[0070] In some embodiments, the output module 304 is further configured to: Calculate the output of the fully connected layer of the deep convolutional neural network model using the following formula: ; where, represents the output of the fully connected layer of the deep convolutional neural network model, W represents the weight matrix, b represents the bias vector, represents the feature vector at the i-th amino acid position in the protein sequence to be predicted; Map the output of the fully connected layer to the probability interval to obtain the prediction probability.
[0071] It should be noted that the above B-cell epitope prediction device based on the protein language model can execute the B-cell epitope prediction method based on the protein language model provided by the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the embodiments of the B-cell epitope prediction device based on the protein language model, reference can be made to the B-cell epitope prediction method based on the protein language model provided by the embodiments of the present invention.
[0072] Those of ordinary skill in the art can understand that all or part of the features / steps of implementing the above method embodiments can be realized by a method, a data processing system or a computer program. These features can be implemented without using hardware, entirely using software, or using a combination of hardware and software. The aforementioned computer program can be stored in one or more computer-readable storage media. When the computer program stored on the storage medium is executed (such as by a processor), it executes the steps of the above embodiments of the B-cell epitope prediction method based on the protein language model.
[0073] The aforementioned storage media that can store program codes include: a static hard disk, a solid-state drive, a random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), an optical storage device, a magnetic storage device, a flash memory, a magnetic disk or an optical disc, and / or a combination of the above devices, that is, it can be implemented by any type of volatile or non-volatile storage device or a combination thereof.
[0074] As Figure 4 shown, the present application also provides an embodiment of a prediction device 1, including one or more processors 11 and a memory 10; wherein, the memory 10 is used to store one or more computer programs, and one or more processors 11 are used to execute the one or more computer programs stored in the memory 10, so that the processor 11 executes the features / steps of the above embodiments of the B-cell epitope prediction method based on the protein language model.
[0075] The present application also provides a computer program product, which is stored on a data carrier and is designed to execute the B-cell epitope prediction method based on a protein language model as described above. Therefore, the computer program product according to the present application has the same advantages as those described in detail with reference to the device according to the present application. The computer program product can be executed as computer-readable instruction codes in each appropriate programming language such as JAVA, C++, etc. In addition, the computer program product can be provided on a network, such as the Internet, or a network user can download the computer program product from a network, such as the Internet, when needed. The computer program product can be implemented either by means of a computer program, i.e., software, or by means of one or more dedicated electronic circuits, i.e., hardware, or in any mixed form, i.e., by means of software components and hardware components, or in a form of software, hardware, or a mixture of software and hardware.
[0076] The above are only the preferred embodiments of the present application. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present application. Additionally, under the teaching of the present application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the protection scope of the present application.
Claims
1. A B-cell epitope prediction method based on a protein language model, characterized in that, Including: Input the protein sequence to be predicted into the pre-trained ESM-IF1 model and the pre-trained ESM2 model respectively to capture multi-dimensional structural embedding vectors and sequence feature vectors; Concatenate the sequence feature vector and the structural embedding vector to obtain a multi-dimensional data set; Based on the Transformer architecture, train the multi-dimensional data set through an integrated convolutional neural network, and calculate the TIM loss function. When the value of the TIM loss function is within a preset range, obtain a deep convolutional neural network model; Use the deep convolutional neural network model to output the prediction probability of each amino acid position of the protein sequence; If the prediction probability is greater than the set threshold, determine that the corresponding amino acid position is an epitope, otherwise it is a non-epitope.
2. The B cell epitope prediction method based on a protein language model according to claim 1, wherein Input the protein sequence to be predicted into the pre-trained ESM-IF1 model to obtain a multi-dimensional structural embedding vector, including: Input the protein sequence to be predicted into the pre-trained ESM-IF1 model for the pre-trained ESM-IF1 model to learn the geometric features in the protein backbone atom coordinates and characterize each residue in the protein sequence; Based on the geometric features and the feature vectors of each residue, obtain a multi-dimensional structural embedding vector.
3. The B cell epitope prediction method based on a protein language model according to claim 1, wherein Input the protein sequence to be predicted into the pre-trained ESM2 model for prediction to obtain a multi-dimensional sequence feature vector, including: Input the protein sequence to be predicted into the pre-trained ESM2 model for the ESM2 model to capture the local sequence information and global sequence information of the protein sequence and generate corresponding embedding vectors for each amino acid; Based on the local sequence information, global sequence information and the embedding vectors corresponding to each amino acid, obtain a multi-dimensional sequence feature vector.
4. The B cell epitope prediction method based on a protein language model according to claim 1, characterized in that, The concatenating the sequence feature vector and the structural embedding vector to obtain a multi-dimensional data set includes: Concatenate the sequence feature vector and the structural embedding vector along the last dimension to generate the multi-dimensional data set. The dimension number of the multi-dimensional data set is the sum of the dimension numbers of the sequence feature vector and the structural embedding vector, and the multi-dimensional data set is a three-dimensional tensor based on the batch size, sequence length and feature dimension in shape.
5. The B cell epitope prediction method based on a protein language model according to claim 1, wherein, The training the multi-dimensional data set through an integrated convolutional neural network based on the Transformer architecture includes: Input the multi-dimensional data set into the Transformer architecture, and use scaled dot-product attention in the Transformer architecture to obtain the final output result of multi-head attention as the global feature; Input the multi-dimensional data set into the convolutional neural network, and output the convolution result as the convolution feature through the residual block of the convolutional neural network; Train based on the global feature and the convolution feature.
6. The B cell epitope prediction method based on a protein language model according to claim 1, wherein The calculating the TIM loss function includes: Based on the multi-dimensional data set and the label corresponding to the multi-dimensional data set, determine the empirical conditional entropy and empirical marginal entropy of the label and calculate the cross entropy; Calculate the TIM loss function based on the empirical conditional entropy, empirical marginal entropy, and the cross entropy.
7. The B cell epitope prediction method based on a protein language model according to claim 1, wherein The utilization of the deep convolutional neural network model to output the prediction probability of each amino acid position includes: Calculate the output of the fully connected layer of the deep convolutional neural network model using the following formula: ; Among them, represents the output of the fully connected layer of the deep convolutional neural network model, W represents the weight matrix, and b represents the bias vector. represents the feature vector of the i-th amino acid position in the protein sequence to be predicted; Map the output of the fully connected layer to the probability interval to obtain the prediction probability.
8. A B-cell epitope prediction device based on a protein language model, characterized in that, Includes: A feature extraction module for inputting the protein sequence to be predicted into a pre-trained ESM-IF1 model and a pre-trained ESM2 model respectively for prediction to obtain multi-dimensional sequence feature vectors and structural embedding vectors; A splicing module for splicing the sequence feature vectors and the structural embedding vectors to obtain a multi-dimensional data set; A training module for training the multi-dimensional data set based on the Transformer architecture through an integrated convolutional neural network and calculating the TIM loss function, and obtaining a deep convolutional neural network model when the value of the TIM loss function is within a preset range; An output module for using the deep convolutional neural network model to output the prediction probability of each amino acid position of the protein sequence; A prediction module for determining that the corresponding amino acid position is an epitope if the prediction probability is greater than a set threshold, otherwise it is a non-epitope.
9. A prediction device, characterized in that, Includes: One or more processors; A memory for storing one or more computer programs, and one or more of the processors are used to execute the one or more computer programs stored in the memory, so that one or more of the processors execute the B cell epitope prediction method based on a protein language model according to any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product is stored on a data carrier and is designed to execute the B cell epitope prediction method based on a protein language model according to any one of claims 1-7.
Citation Information
Patent Citations
Method for designing mRNA (messenger ribonucleic acid) vaccine sequence based on sequence and structure information
CN116386727A
Protein complex model quality evaluation method based on Voronoi diagram
CN119132385A
DNA binding protein and RNA binding protein classification method based on ESM-2 and dual-path neural network
CN119229982A
Drug and target affinity prediction method based on parallel full-connection network
CN119763653A
Data augmentation method and method for generating epitope activity prediction model using same
WO2024162655A1
Cited By
Halophilic protein prediction method and system based on mixed deep learning architecture
CN121506251A
Protein sequence analysis method, device, medium and product
CN121905284A