BCE prediction method based on hybrid deep learning strategy
By constructing a BCE prediction model with a hybrid deep learning strategy, using the protein language big model and structural characteristics, the time-consuming and expensive BCE prediction in the existing technology is solved, and efficient and economical BCE prediction is achieved, supporting the rapid development of immune drugs and vaccines.
Patent Information
- Application Number
- CN202510055199.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The prior art is time-consuming and expensive to predict B-cell antigen epitope (BCE), making it difficult to efficiently support the development of immune drugs and vaccines.
Using a BCE prediction method based on a hybrid deep learning strategy, a prediction model including a database, feature extraction module, feature processing module and forward neural network module is constructed, and protein sequence features are extracted using a protein language model and predicted in combination with structural features.
It achieves rapid and economical prediction of BCE, reduces experimental costs and manpower investment, and improves the efficiency of vaccine and drug design.
Smart Images

Figure CN119993274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of immunological drug research and development, and in particular to a BCE prediction method based on a hybrid deep learning strategy. Background Art
[0002] B-cell epitopes (BCE) are one of the key points in the development of immune drugs. It refers to the part of the antigen recognized by the variable region of the B cell antibody. Generally speaking, BCE can be divided into two categories: linear (continuous) and conformational (discontinuous). Linear BCE exists in the primary sequence of the antigen molecule and is composed of a continuous sequence of amino acid residues. The identification of antigen epitopes is very important for vaccine design, immunodiagnostic testing, and synthetic antibody production. Although there are fewer natural linear antigen epitopes, there are still many studies on linear antigen epitopes because they are very important for the development of synthetic peptide vaccines.
[0003] Since experimental determination of BCE is time-consuming and expensive, computational predictions can play a key role in the development of new vaccines and drugs against common viral pathogens such as human immunodeficiency virus, hepatitis or influenza viruses.
[0004] At present, generative artificial intelligence technology represented by large models can learn the common features and representations of data from unlabeled data through unsupervised learning methods (autoencoders, generative adversarial networks, masked language modeling, contrastive learning, etc.) and large-scale deep learning models trained with massive data. Large protein language models based on pre-training or fine-tuning have demonstrated excellent performance in protein structure, function and other downstream tasks. The embedded representation obtained by inputting the primary sequence into the large protein language model can capture some biophysical characteristics of the protein sequence or predict the three-dimensional structure of the protein.
[0005] At present, the feature representation of proteins mostly uses features extracted from protein sequences, such as amino acid pair antigenicity scale (AAP), amino acid trimer antigenicity scale (AAT), amino acid composition (AAC), etc. to calculate BCE features, as well as physical and chemical properties, structural composition, etc. The calculation models are mainly based on random forests, support vector machines, extreme gradient boosting, etc., and the focus is more on input feature selection or feature dimensionality reduction.
[0006] Starting from the primary sequence of the protein, using generative artificial intelligence technologies such as large models, more sequence features can be obtained. Combining sequence AAP, AAT, AAC, amino acid local structure information, etc., the protein can be better characterized.
[0007] Epitopes are the basis of protein antigenicity, and determining B cell epitopes plays an important guiding role in designing vaccines and drugs. Experimental methods for determining antigen epitopes include X-ray diffraction, fluorescence polarization technology, phage display technology, peptide scanning technology, etc. These methods are relatively cumbersome and require a lot of work. With the development of computer technology and the increasing expansion of biological databases, it has become possible to summarize the sequence and structural characteristics of antigen epitopes from existing data and predict possible epitopes through computational means. Computational methods can save a lot of experimental and labor costs.
[0008] Predicting BCE using computational methods can be described as determining whether a protein sequence is an epitope, which is a binary classification problem. Summary of the invention
[0009] In order to solve the technical problems raised in the background technology, the present invention provides a BCE prediction method based on a hybrid deep learning strategy.
[0010] The present invention is implemented by the following technical solution: A BCE prediction method based on a hybrid deep learning strategy, comprising the following steps:
[0011] Step 1: Build a BCE prediction model;
[0012] Step 2, inputting the protein sequence into the BCE prediction model;
[0013] Step 3: Use the prediction model to output whether the protein sequence is BCE.
[0014] Specifically, the BCE prediction model in step 1 includes a database, a feature extraction module, a feature processing module and a forward neural network module;
[0015] in:
[0016] The database is used to obtain training data sets;
[0017] The feature extraction module is used to extract features from the data and obtain four sets of features;
[0018] The feature processing module is used to process the extracted features;
[0019] The feedforward neural network module is used to merge multiple processed features and output predicted values.
[0020] Specifically, the training data in the database comes from the IEDB database and the Bcipep database; the verification data in the database is also selected from the EDB database and the Bcipep database.
[0021] Specifically, the feature extraction module is used for ProT5 representation features, ESM-2 representation features, DSSP features and sequence features, and the sequence features are spliced by orthogonal coding of sequence residues, AAT, AAP, and AAC;
[0022] in:
[0023] AAP is an amino acid antigen pair, AAT is an amino acid trimer, and AAC is an amino acid composition;
[0024] ProtT5 indicates that the feature is extracted using the protein language large model ProtT5; the specific model is ProtT5-XL-UniRef50, the input of which is the primary sequence of the protein, and the output is the R of the ith residue in the corresponding sequence. i The feature representation vector X i , X i The dimension is 1024;
[0025] ESM-2 represents features extracted using the protein language pre-trained large model ESM-2; the specific model is Esm2_t33_650M_UR50D, the input of the model is the primary sequence of the protein, and the output is the i-th residue R in the corresponding sequence i The feature representation vector X i , X i Dimension 1280;
[0026] DSSP features are extracted using the protein structure prediction model ESMFold; ESMFold is used to predict the three-dimensional structure of the protein, and the protein file is parsed using DSSP software to obtain the secondary structure information, (φ, ψ) dihedral angle information, and solvent accessible surface area information of the eight states of the protein; the secondary structure is represented by 8-dimensional one-hot orthogonal encoding data; a dihedral angle is represented by a pair of sine-cosine functions, such as formula (1), and the (φ, ψ) dihedral angle is represented by 4-dimensional data; the solvent accessible surface area is normalized by maximum-minimum, such as formula (2), and the minimum value of the residue solvent accessible surface area is 0; the above data has a total of 13 dimensions;
[0027] y 1 = sin(θ),y 2 =cos(θ)(1)
[0028]
[0029] Sequence features were extracted using the following method;
[0030] Given that there are too few non-zero data in the above one-hot orthogonal vector encoding:
[0031] First, use formula (3) to map the sparse code to the dense code by the autoencoder; take the parameter W in the encoder as the corresponding new code vector, and the data dimension is 20;
[0032] The AminoAcid Pair (AAP) feature represents the frequency of two adjacent amino acids appearing in pairs in the epitope sequence in the dataset. The frequency of occurrence of this amino acid pair in non-epitope sequences The ratio is calculated as formula (4). Normalized to [-1,1]; amino acid antigen pair features are represented by four values: AAP feature value, maximum value in AAP, minimum value in AAP and average value of AAP, a total of 4 dimensions
[0033]
[0034] in and The calculation method is the same as formula (5) and (6).
[0035]
[0036] in, is the number of times a pair of amino acids appears in the epitope sequence, is the number of all amino acid pairs in the epitope sequence.
[0037]
[0038] in, is the number of occurrences of an amino acid pair in the non-epitope sequence, is the number of all amino acid pairs in the non-epitope sequence.
[0039] The Amino Acid Triplets (AAT) feature represents the ratio of the frequency of occurrence of three consecutive amino acid residues in the epitope sequence to the frequency of occurrence of the trimer in the non-epitope sequence in the dataset. The calculation and normalization methods are the same as the AAP feature. The AAT feature represents four sets of values including the AAT characteristic value, maximum value, minimum value and average value of the residue, with a total of 4 dimensions.
[0040] The amino acid composition (AAC) feature represents the relative proportion of each amino acid in each protein sequence in the dataset, and is calculated as shown in formula (7).
[0041]
[0042] Where R iis the number of amino acids of type i in the sequence, N is the length of the sequence, and the AAC feature is 1-dimensional;
[0043] Considering the convenience of data batch input during deep learning model training, all sequence lengths are aligned to 71 according to the longest sequence in all test sets, and zeros are added if the length is less than 71.
[0044] The residue feature ProtT5 is 71*1024, ESM-2 is 71*1280, DSSP is 71*13, and the residue code, AAT, AAP, and AAC are spliced together to form a total of 71*29;
[0045] Specifically, the feature processing module includes: module (a) structural feature processing module, module (b) ESM-2 large model feature processing module, module (c) ProtT5 large model feature processing module, and module (d) sequence feature processing module.
[0046] Specifically, in the structural feature processing module: the three-dimensional structure predicted by ESMFold is parsed into 13-dimensional structural features using DSSP software; a single sequence 71*13 feature matrix is extracted using a full-size two-dimensional convolutional neural network feature to convert the two-dimensional matrix into a one-dimensional vector; the convolution kernel size is (71, 13), the activation function is ReLU, and the output is 256 dimensions; the calculation process is as shown in formula (8).
[0047] F 1 =ReLU(Conv(W*x 1:71 +b)) (8)
[0048] Specifically, the residue representation feature 1280 dimensions output by the ESM-2 large model is input into a two-layer bidirectional LSTM network, and then a layer of forward attention network is used to map the two dimensions to a one-dimensional vector;
[0049] First, a bidirectional LSTM network is used to focus on the global information of proteins and capture the long-range dependencies of protein sequences; the unidirectional LSTM model is formulated as shown in (9).
[0050]
[0051] Among them, σ is the activation function, generally using the Sigmoid function; ⊙ represents the bitwise multiplication of the matrix; x t is the network input at time t; i t 、f t , o t 、c t and h t They represent the input gate, forget gate, output gate, internal memory unit and output at time t respectively; h t-1 is the output of the previous moment; c t-1is the output of the internal memory unit at the previous moment; the rest are learnable parameters of the neural network.
[0052] The first layer of LSTM network has an input of 1280 dimensions and a unidirectional output of 128 dimensions. When the forward LSTM and backward LSTM data converge, a merge operation is performed at the lowest dimension, as shown in formula (10), and the output is 256 dimensions;
[0053] The second layer of LSTM network inputs 256 and outputs 128 dimensions in one direction. When converging forward and backward, it also performs a merge operation at the lowest dimension and outputs 256 dimensions.
[0054] The forward attention network is mainly used to convert a two-dimensional matrix into a one-dimensional vector; h t ′ is the output of the bidirectional LSTM, representing the t-th amino acid residue in the current sequence; through a layer of forward neural network as shown in formula (11), the attention weight of the current feature representation of residue t in the sequence is obtained through formula (12); the feature representation of all residues in the sequence is weighted and summed, as shown in formula (13), to achieve the conversion of two-dimensional feature representation to one-dimensional. The input data dimension of the forward attention network is 71*256, and the output dimension is 1*256.
[0055] e t =h t 'W t (11)
[0056]
[0057]
[0058] Specifically, the ProtT5 represents a feature extraction module, which includes two layers of bidirectional LSTM networks and one layer of forward attention network. The module input is 1024 dimensions, and the module output F3 is also 256 dimensions, and the two-dimensional features are also mapped to one-dimensional vectors.
[0059] Specifically, in the sequence feature extraction module, the full-size convolutional neural network is used to extract the encoding features, AAP, AAT and AAC features of the sequence; the input size of each sequence is 71*29, the convolution kernel size is (71, 29), the activation function is ReLU, and the output feature F4 is 256 dimensions.
[0060] Specifically, the feedforward neural network module performs a data merging operation on the output one-dimensional vectors of the above four feature processing modules. The output dimension of each module is 256, and the merged dimension is 1024, as shown in formula (14);
[0061] F = concat(F 1 ,F 2 ,F 3 ,F 4) (14)
[0062] It includes two layers of feedforward neural networks, as shown in formulas (15) and (16), respectively. The first feedforward neural network uses tanh as activation function, and the output dimension is 512; the second one uses Sigmoid function as shown in formula (17), the output dimension is 1, and the output result is 0 or 1, where: 0 represents non-epitope, 1 represents antigen epitope;
[0063] F′=tanh(W*F+b) (15)
[0064]
[0065] The error loss between the model prediction value and the true value is described by the binary cross entropy loss function, as shown in formula (18);
[0066]
[0067] Compared with the prior art, the present invention has the following beneficial effects:
[0068] The present invention provides a BCE prediction method of a hybrid deep learning strategy, which uses different deep neural networks such as convolutional neural networks, bidirectional long short-term memory networks, and forward attention mechanisms to fit and extract different types of sequence features to achieve B cell linear antigen epitope prediction for classification tasks. The prediction model trained by the present invention can be run on an ordinary user-level computer.
[0069] The deep learning prediction method proposed in the present invention uses sequence features and structural features to represent a protein, uses a four-channel deep neural network to extract sequence information and structural information respectively, and can use the same model to predict linear epitopes. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 The flowchart of the BCE prediction method based on the hybrid deep learning strategy proposed in the present invention;
[0071] Figure 2 This is a structural diagram of the BCE prediction model based on the hybrid deep learning strategy proposed in the present invention. DETAILED DESCRIPTION
[0072] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.
[0073] Example:
[0074] Reference Figure 1-Figure 2 ,This scheme proposes a BCE prediction method based on a hybrid deep learning strategy,,including the following steps:
[0075] Step 1: Build a BCE prediction model;
[0076] Step 2, inputting the protein sequence into the BCE prediction model;
[0077] Step 3: Use the prediction model to output whether the protein sequence is BCE.
[0078] Specifically, the BCE prediction model in step 1 includes a database, a feature extraction module, a feature processing module and a forward neural network module;
[0079] in:
[0080] The database is used to obtain training data sets;
[0081] The feature extraction module is used to extract features from the data and obtain four sets of features;
[0082] The feature processing module is used to process the extracted features;
[0083] The feedforward neural network module is used to merge multiple processed features and output predicted values.
[0084] Specifically, the training data in the database comes from the IEDB database and the Bcipep database; the verification data in the database is also selected from the EDB database and the Bcipep database.
[0085] Specifically, the feature extraction module is used for ProtT5 representation features, ESM-2 representation features, DSSP features and sequence features, and the sequence features are spliced by orthogonal coding of sequence residues, AAT, AAP, and AAC;
[0086] in:
[0087] AAP is an amino acid antigen pair, AAT is an amino acid trimer, and AAC is an amino acid composition;
[0088] ProtT5 features are extracted using the protein language model ProtT5; the specific model is ProtT5-XL-UniRef50, the input of which is the primary sequence of the protein, and the output is the R of the ith residue in the corresponding sequence. i The feature representation vector X i , X i The dimension is 1024;
[0089] The ESM-2 features are extracted using the protein language pre-trained large model ESM-2; the specific model is esm2_t33_650M_UR50D, the input of which is the primary sequence of the protein, and the output is the i-th residue R in the corresponding sequence. iThe feature representation vector X i , X i Dimension 1280;
[0090] DSSP features are extracted using the protein structure prediction model ESMFold; ESMFold is used to predict the three-dimensional structure of the protein, and the protein file is parsed using DSSP software to obtain the secondary structure information, (φ, ψ) dihedral angle information, and solvent accessible surface area information of the eight states of the protein; the secondary structure is represented by 8-dimensional one-hot orthogonal encoding data; a dihedral angle is represented by a pair of sine-cosine functions, such as formula (1), and the (φ, ψ) dihedral angle is represented by 4-dimensional data; the solvent accessible surface area is normalized by maximum-minimum, such as formula (2), and the minimum value of the residue solvent accessible surface area is 0; the above data has a total of 13 dimensions;
[0091] y 1 = sin(θ),y 2 =cos(θ)(1)
[0092]
[0093] Sequence features were extracted using the following method;
[0094] Given that there are too few non-zero data in the above one-hot orthogonal vector encoding:
[0095] First, use formula (3) to map the sparse code to the dense code by the autoencoder; take the parameter W in the encoder as the corresponding new code vector, and the data dimension is 20;
[0096]
[0097] The AminoAcid Pair (AAP) feature represents the frequency of two adjacent amino acids appearing in pairs in the epitope sequence in the dataset. The frequency of occurrence of this amino acid pair in non-epitope sequences The ratio is calculated as formula (4). Normalized to [-1,1]; amino acid antigen pair features are represented by four values: AAP feature value, maximum value in AAP, minimum value in AAP and average value of AAP, a total of 4 dimensions
[0098]
[0099] in and The calculation method is the same as formula (5) and (6).
[0100]
[0101] in, is the number of times a pair of amino acids appears in the epitope sequence, is the number of all amino acid pairs in the epitope sequence.
[0102]
[0103] in, is the number of occurrences of an amino acid pair in the non-epitope sequence, is the number of all amino acid pairs in the non-epitope sequence.
[0104] The AminoAcid Triplets (AAT) feature represents the ratio of the frequency of occurrence of three consecutive amino acid residues in the epitope sequence to the frequency of occurrence of the trimer in the non-epitope sequence in the dataset. The calculation and normalization methods are the same as the AAP feature. The AAT feature represents four sets of values including the AAT characteristic value, maximum value, minimum value and average value of the residue, with a total of 4 dimensions.
[0105] The amino acid composition (AAC) feature represents the relative proportion of each amino acid in each protein sequence in the dataset, and is calculated as shown in formula (7).
[0106]
[0107] Where R i is the number of amino acids of type i in the sequence, N is the length of the sequence, and the AAC feature is 1-dimensional;
[0108] Considering the convenience of data batch input during deep learning model training, all sequence lengths are aligned to 71 according to the longest sequence in all test sets, and zeros are added if the length is less than 71.
[0109] The residue feature ProtT5 is 71*1024, ESM-2 is 71*1280, DSSP is 71*13, and the residue code, AAT, AAP, and AAC are spliced together to form a total of 71*29;
[0110] Specifically, the feature processing module includes: module (a) structural feature processing module, module (b) ESM-2 large model feature processing module, module (c) ProtT5 large model feature processing module, and module (d) sequence feature processing module.
[0111] Specifically, in the structural feature processing module: the three-dimensional structure predicted by ESMFold is parsed into 13-dimensional structural features using DSSP software; a single sequence 71*13 feature matrix is extracted using a full-size two-dimensional convolutional neural network feature to convert the two-dimensional matrix into a one-dimensional vector; the convolution kernel size is (71, 13), the activation function is ReLU, and the output is 256 dimensions; the calculation process is as shown in formula (8).
[0112] F 1 =ReLU(Conv(W*x 1:71 +b)) (8)
[0113] Specifically, the residue representation feature 1280 dimensions output by the ESM-2 large model is input into a two-layer bidirectional LSTM network, and then a layer of forward attention network is used to map the two dimensions to a one-dimensional vector;
[0114] First, a bidirectional LSTM network is used to focus on the global information of proteins and capture the long-range dependencies of protein sequences; the unidirectional LSTM model is formulated as shown in (9).
[0115]
[0116] Among them, σ is the activation function, generally using the Sigmoid function; ⊙ represents the bitwise multiplication of the matrix; x t is the network input at time t; i t 、f t , o t 、c t and h t They represent the input gate, forget gate, output gate, internal memory unit and output at time t respectively; h t-1 is the output of the previous moment; c t-1 is the output of the internal memory unit at the previous moment; the rest are learnable parameters of the neural network.
[0117] The first layer of LSTM network has an input of 1280 dimensions and a unidirectional output of 128 dimensions. When the forward LSTM and backward LSTM data converge, a merge operation is performed at the lowest dimension, as shown in formula (10), and the output is 256 dimensions;
[0118] The second layer of LSTM network inputs 256 and outputs 128 dimensions in one direction. When converging forward and backward, it also performs a merge operation at the lowest dimension and outputs 256 dimensions.
[0119] The forward attention network is mainly used to convert a two-dimensional matrix into a one-dimensional vector; h t ′ is the output of the bidirectional LSTM, representing the t-th amino acid residue in the current sequence; through a layer of forward neural network as shown in formula (11), the attention weight of the current feature representation of residue t in the sequence is obtained through formula (12); the feature representation of all residues in the sequence is weighted and summed, as shown in formula (13), to achieve the conversion of two-dimensional feature representation to one-dimensional. The input data dimension of the forward attention network is 71*256, and the output dimension is 1*256.
[0120] e t =h t 'W t (11)
[0121]
[0122] Specifically, the ProtT5 represents a feature extraction module, which includes two layers of bidirectional LSTM networks and one layer of forward attention network. The module input is 1024 dimensions, and the module output F3 is also 256 dimensions, and the two-dimensional features are also mapped to one-dimensional vectors.
[0123] Specifically, in the sequence feature extraction module, the full-size convolutional neural network is used to extract the encoding features, AAP, AAT and AAC features of the sequence; the input size of each sequence is 71*29, the convolution kernel size is (71, 29), the activation function is ReLU, and the output feature F4 is 256 dimensions.
[0124] It should be noted that the feedforward neural network module performs a data merging operation on the output one-dimensional vectors of the above four feature processing modules. The output dimension of each module is 256, and the merged dimension is 1024, as shown in formula (14);
[0125] F = concat(F 1 ,F 2 ,F 3 ,F 4 )(14)
[0126] It includes two layers of feedforward neural networks, as shown in formulas (15) and (16), respectively. The first feedforward neural network uses tanh as activation function, and the output dimension is 512; the second one uses Sigmoid function as shown in formula (17), the output dimension is 1, and the output result is 0 or 1, where: 0 represents non-epitope, 1 represents antigen epitope;
[0127] F′=tanh(W*F+b)(15)
[0128]
[0129] The error loss between the model prediction value and the true value is described by the binary cross entropy loss function, as shown in formula (18);
[0130]
[0131] In summary, it can be seen that predicting BCE using computational methods can be described as determining whether a protein sequence is an epitope, which is a binary classification problem.
[0132] The above-mentioned embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by technicians in this field on the basis of the present invention shall fall within the scope of protection required by the present invention.
Claims
1. A BCE prediction method based on a hybrid deep learning strategy, characterized in that: The steps include: Step 1: Build a BCE prediction model; Step 2, inputting the protein sequence into the BCE prediction model; Step 3: Use the prediction model to output whether the protein sequence is BCE.
2. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The BCE prediction model in step 1 includes a database, a feature extraction module, a feature processing module and a forward neural network module; in: The database is used to obtain training data sets; The feature extraction module is used to extract features from the data and obtain four sets of features; The feature processing module is used to process the extracted features; The feedforward neural network module is used to merge multiple processed features and output predicted values.
3. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The training data in the database are from the IEDB database and the Bcipep database; the verification data in the database are also selected from the EDB database and the Bcipep database.
4. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The feature extraction module is used for ProtT5 expression features, ESM-2 expression features, DSSP features and sequence features, and the sequence features are spliced by orthogonal coding of sequence residues, AAT, AAP, and AAC; in: AAP is an amino acid antigen pair, AAT is an amino acid trimer, and AAC is an amino acid composition; ProtT5 indicates that the feature is extracted using the protein language large model ProtT5; the specific model is ProtT5-XL-UniRef50, the input of which is the primary sequence of the protein, and the output is the i-th residue R in the corresponding sequence. i The feature representation vector X i , X i The dimension is 1024; ESM-2 represents features extracted using the protein language pre-trained large model ESM-2; the specific model is Esm2_t33_650M_UR50D, the input of the model is the primary sequence of the protein, and the output is the i-th residue R in the corresponding sequence i The feature representation vector X i , X i Dimension 1280; DSSP features are extracted using the protein structure prediction model ESMFold; ESMFold is used to predict the three-dimensional structure of the protein, and the protein file is parsed using DSSP software to obtain the secondary structure information, (φ, ψ) dihedral angle information, and solvent accessible surface area information of the eight states of the protein; the secondary structure is represented by 8-dimensional one-hot orthogonal encoding data; a dihedral angle is represented by a pair of sine-cosine functions, such as formula (1), and the (φ, ψ) dihedral angle is represented by 4-dimensional data; the solvent accessible surface area is normalized by maximum-minimum, such as formula (2), and the minimum value of the residue solvent accessible surface area is 0; the above data has a total of 13 dimensions; y1=sin(θ),y2=cos(θ)(1) Sequence features were extracted using the following method; Given that there are too few non-zero data in the above one-hot orthogonal vector encoding: First, use formula (3) to map the sparse code to the dense code by the autoencoder; take the parameter W in the encoder as the corresponding new code vector, and the data dimension is 20; The AAP feature represents the frequency of two adjacent amino acids in the AAP epitope sequence in the dataset. The frequency of occurrence of the AAP in non-epitope sequences The ratio is calculated as formula (4). Normalized to [-1,1]; amino acid antigen pair features are represented by four values: AAP feature value, maximum value in AAP, minimum value in AAP and average value of AAP, with a total of 4 dimensions; in and The calculation method is as shown in formulas (5) and (6); in, is the number of times a pair of amino acids appears in the epitope sequence, is the number of all amino acid pairs in the epitope sequence; in, is the number of occurrences of an amino acid pair in the non-epitope sequence, is the number of all amino acid pairs in the non-epitope sequence; The AAT feature represents the ratio of the frequency of occurrence of three consecutive amino acid residues in the epitope sequence in the dataset to the frequency of occurrence of the trimer in the non-epitope sequence. The calculation and normalization methods are the same as the AAP feature. The AAT feature represents four sets of values, including the AAT feature value, maximum value, minimum value and average value of the residue, with a total of 4 dimensions. AAC feature, which represents the relative proportion of each amino acid in each protein sequence in the dataset, is calculated as shown in formula (7); Where R i is the number of amino acids of type i in the sequence, N is the length of the sequence, and the AAC feature is 1-dimensional; Considering the convenience of data batch input during deep learning model training, all sequence lengths are aligned to 71 according to the longest sequence in all test sets, and zeros are added if the length is less than 71. The residue feature ProtT5 is 71*1024, ESM-2 is 71*1280, DSSP is 71*13, and the residue code, AAT, AAP, and AAC are spliced together to form a total of 71*29.
5. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The feature processing modules include: a structural feature processing module, an ESM-2 large model feature processing module, a ProtT5 large model feature processing module, and a sequence feature processing module.
6. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: In the structural feature processing module: the three-dimensional structure predicted by ESMFold is analyzed by DSSP software to obtain 13-dimensional structural features; A single sequence 71*13 feature matrix is extracted using a full-size two-dimensional convolutional neural network to convert the two-dimensional matrix into a one-dimensional vector; the convolution kernel size is (71, 13), the activation function is ReLU, and the output is 256 dimensions; the calculation process is as shown in formula (8). F1=ReLU(Conv(W*x 1:71 +b)) (8)。 7. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The residue representation feature output by the ESM-2 large model is 1280-dimensional, which is input into a two-layer bidirectional LSTM network, and then a layer of forward attention network is used to map the two dimensions to a one-dimensional vector; First, a bidirectional LSTM network is used to focus on the global information of proteins and capture the long-range dependencies of protein sequences. The unidirectional LSTM model is described in (9). Among them, σ is the activation function, generally using the Sigmoid function; ⊙ represents the bitwise multiplication of the matrix; x t is the network input at time t; i t 、f t , o t 、c t and h t They represent the input gate, forget gate, output gate, internal memory unit and output at time t respectively; h t-1 is the output of the previous moment; c t-1 is the output of the internal memory unit at the previous moment; the rest are the learnable parameters of the neural network; The first layer of LSTM network has an input of 1280 dimensions and a unidirectional output of 128 dimensions. When the forward LSTM and backward LSTM data converge, a merge operation is performed at the lowest dimension, as shown in formula (10), and the output is 256 dimensions. The second layer of LSTM network inputs 256 and outputs 128 dimensions in one direction. When converging forward and backward, it also performs a merge operation at the lowest dimension and outputs 256 dimensions. The forward attention network is mainly used to convert a two-dimensional matrix into a one-dimensional vector; h t ′ is the output of the bidirectional LSTM, representing the t-th amino acid residue in the current sequence; through a layer of forward neural network as shown in formula (11), the attention weight of the current feature representation of residue t in the sequence is obtained through formula (12); the feature representation of all residues in the sequence is weighted and summed, as shown in formula (13), to realize the conversion of two-dimensional feature representation to one-dimensional, the input data dimension of the forward attention network is 71*256, and the output dimension is 1*256; have been t =h t ′W t (11) 。 8. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The ProtT5 represents a feature extraction module, which includes two layers of bidirectional LSTM networks and one layer of forward attention network. The module input is 1024 dimensions, and the module output F3 is also 256 dimensions, and the two-dimensional features are also mapped to one-dimensional vectors.
9. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: In the sequence feature extraction module, the encoding features, AAP, AAT and AAC features of the sequence are also extracted by a full-size convolutional neural network; the input size of each sequence is 71*29, the convolution kernel size is (71, 29), the activation function is ReLU, and the output feature F4 is 256 dimensions.
10. A BCE prediction method based on a hybrid deep learning strategy as claimed in claim 1, characterized in that: The feedforward neural network module performs a data merging operation on the one-dimensional vectors output by the four feature processing modules. The output dimension of each module is 256, and the merged dimension is 1024, as shown in formula (14); F=concat(F1,F2,F3,F4) (14) It includes two layers of feedforward neural networks, as shown in formulas (15) and (16), respectively. The first feedforward neural network uses tanh as activation function, and the output dimension is 512; the second one uses Sigmoid function as shown in formula (17), the output dimension is 1, and the output result is 0 or 1, where: 0 represents non-epitope, 1 represents antigen epitope; F′=tanh(W*F+b) (15) The error loss between the model prediction value and the true value is described by the binary cross entropy loss function, as shown in formula (18);
Citation Information
Patent Citations
Protein structure prediction method based on mixed deep learning model
CN116312754A
Sequence-based antigen-antibody affinity prediction method
CN116434839A
Protein SNO site prediction method of deep learning network fusing features
CN117976035A
B cell epitope prediction method and device
CN118155723A
Protein secondary structure prediction method based on multi-task deep learning
CN118280432A