Antigen immunogenicity prediction method and system based on evolution scale model
By using the method based on evolutionary scale model to characterize antigen sequences, the problem that the existing technology is difficult to effectively characterize the deep characteristics of protein sequences is solved, and more stable and accurate antigen immunogenicity prediction is achieved.
Patent Information
- Application Number
- CN202510149879.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
Existing antigen immunogenicity prediction methods are difficult to fully characterize the deep characteristics of complex protein sequences, affecting the stability and accuracy of the prediction results.
Using an evolutionary scale model-based method, T cell antigen receptor, Class I human leukocyte antigen and antigen peptide sequences are characterized, and a joint feature representation matrix is generated through a bilinear attention network, and an antigen immunogenic prediction model is input for prediction.
By more efficiently characterizing the deep characteristics of protein sequences, the stability and accuracy of antigen immunogenicity prediction are improved, which is of great clinical significance.
Smart Images

Figure CN120072041A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computational biology, and particularly relates to antigen immunogenicity prediction. Background Art
[0002] Neoantigen peptides presented by class I human leukocyte antigen (HLA-I) molecules are recognized by T cell antigen receptors (TCRs), activating T cells to transform from naive T cells into CD8+ cytotoxic T cells, thereby stimulating the immune system and destroying target malignant cells. Complementary determining region 3 of the beta chain of the T cell antigen receptor (CDR3β) has attracted much attention due to its highly diverse antigen specificity. Since TCR must consider HLA restriction when recognizing neoantigens (because it can recognize both antigen peptides and polymorphisms), it is crucial to understand the binding mechanism of the TCR ternary complex (TCR-pHLA).
[0003] The development of artificial intelligence computing models provides technical support for accurately predicting the immunogenicity of candidate antigens. Predicting antigen immunogenicity by constructing deep learning models has been successfully applied to the selection of targets for immunotherapy, opening up new avenues for modern immunology research. In the past few years, the large output of laboratory data and the improvement of deep learning technologies have improved the prediction accuracy of antigen immunogenicity. However, existing methods still use traditional coding methods (such as Blosum, One-hot, etc.) at the coding level of amino acid sequences, making it difficult to fully represent the deep features of complex protein sequences and affecting the stability and accuracy of the final prediction results. Therefore, it is of great clinical significance to implement a coding method that can more richly represent protein sequences to predict antigen immunogenicity with higher accuracy. Summary of the Invention
[0004] The present invention aims to solve the problem that existing antigen immunogenicity prediction methods are difficult to fully represent the deep features of complex protein sequences, thereby affecting the stability and accuracy of the final prediction results. Now, an antigen immunogenicity prediction method and system based on an evolutionary scale model are provided.
[0005] An antigen immunogenicity prediction method based on an evolutionary scale model includes:
[0006] Using an evolutionary scale model to perform feature encoding on the measured immunogenicity data to generate a feature matrix of the measured immunogenicity data, where the measured immunogenicity data includes: T cell antigen receptor sequence, class I human leukocyte antigen sequence, and antigen peptide sequence;
[0007] Fusing the feature matrix into a joint feature representation matrix through a bilinear attention network and inputting it into an antigen immunogenicity prediction model to complete the prediction of antigen immunogenicity.
[0008] Further, the above-mentioned feature encoding of the measured immunogenicity data using the evolutionary scale model includes:
[0009] Using the evolutionary scale model to perform feature embedding on the complementarity-determining region 3 sequence of the T cell antigen receptor, capturing key sequence features, and generating a feature matrix of the T cell antigen receptor;
[0010] Using the evolutionary scale model to extract structural features from the class I human leukocyte antigen sequence, generating a feature matrix of the class I human leukocyte antigen;
[0011] Using the evolutionary scale model to extract context features from the antigen peptide sequence, generating a feature matrix of the antigen peptide.
[0012] Further, the above-mentioned evolutionary scale model is a Transformer-based model, and its core architecture includes: an input embedding layer, a Transformer encoder layer, and an output layer.
[0013] Further, the above-mentioned fusion of the feature matrices into a joint feature representation matrix through a bilinear attention network includes:
[0014] Generating a joint feature representation matrix O through the following formula:
[0015]
[0016] where T is the feature matrix of the T cell antigen receptor, F is the feature matrix of the class I human leukocyte antigen, A is the feature matrix of the antigen peptide, and W a is the interaction weight matrix between T and F, W b is the interaction weight matrix between F and A, and Y is the weight matrix mapped to the output feature space.
[0017] Further, the above-mentioned antigen immunogenicity prediction model includes a three-layer one-dimensional convolutional neural network and a three-layer fully connected neural network;
[0018] The three-layer one-dimensional convolutional neural network extracts features from the joint feature representation matrix, and then inputs the extracted one-dimensional feature matrix into the three-layer fully connected neural network to finally obtain the antigen immunogenicity score.
[0019] Further, the convolutional kernel sizes of each layer of the one-dimensional convolutional neural network are respectively: 1, 2, 3, and each convolutional kernel outputs 1000 feature maps, which are concatenated into a one-dimensional feature matrix with a dimension of 3000.
[0020] Further, the expression of the loss function L of the above-mentioned antigen immunogenicity prediction model is:
[0021]
[0022] Among them, p y is the predicted probability of the true class y, and z y is the true label, and z j is the logits of the j-th class output by the model. The logits are the prediction results of the model without passing through Softmax, where j = 1, 2,..., C, and C represents the total number of classes.
[0023] An antigen immunogenicity prediction system based on an evolutionary scale model, comprising:
[0024] An encoding unit: used to perform feature encoding on the measured immunogenicity data by using an evolutionary scale model to generate a feature matrix of the measured immunogenicity data. The measured immunogenicity data includes: T cell antigen receptor sequences, class I human leukocyte antigen sequences, and antigen peptide sequences;
[0025] A fusion unit: used to fuse the feature matrix into a joint feature representation matrix through a bilinear attention network,
[0026] A prediction unit: used to input the joint feature representation matrix into an antigen immunogenicity prediction model to complete the prediction of antigen immunogenicity.
[0027] An antigen immunogenicity prediction device based on an evolutionary scale model. The antigen immunogenicity prediction device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the antigen immunogenicity prediction method based on the evolutionary scale model as described above.
[0028] A computer storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by the processor to implement the antigen immunogenicity prediction method based on the evolutionary scale model as described above.
[0029] For the antigen immunogenicity prediction method and system based on the evolutionary scale model of the present invention, the original data of T cell antigen receptors, class I human leukocyte antigens, and antigen peptides are obtained from an immunogenicity database; then the data is cleaned to obtain sample data that meets the model input standard; then the evolutionary scale model is used to obtain a T cell antigen receptor feature matrix, a class I human leukocyte antigen protein feature matrix, and an antigen sequence feature matrix, and a joint feature representation matrix is obtained through a bilinear attention network; finally, the joint feature representation matrix is used as a training set in the antigen immunogenicity prediction model for training. The deep features of complex protein sequences are fully characterized, and the results are predicted stably and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flowchart of the antigen immunogenicity prediction method based on the evolutionary scale model;
[0031] Figure 2 It is a schematic diagram of module connection for an antigen immunogenicity prediction system based on an evolutionary scale model. Specific implementation manner
[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0033] Specific implementation manner 1: Refer to Figure 1 Specifically illustrate this implementation manner. The antigen immunogenicity prediction method based on the evolutionary scale model described in this implementation manner is used to predict antigen immunogenicity. The prediction model needs to be trained first, and the training data uses sample data.
[0034] The antigen immunogenicity prediction method includes:
[0035] Step S1: Obtain the original data of T cell antigen receptor sequences, class I human leukocyte antigen sequences, and antigen peptide sequences related to the candidate antigen in the immune epitope database.
[0036] Specifically, in this step, positive immunogenicity sample data is collected from the immune epitope database, and then the T cell antigen receptor, class I human leukocyte antigen, and antigen peptide sequences are recorded. At the same time, the original data samples are downloaded and stored for subsequent data processing.
[0037] Step S2: Remove duplicates and clean the original data to construct datasets of T cell antigen receptor sequences, class I human leukocyte antigen sequences, and antigen peptide sequences that meet the quality standards, including:
[0038] Step S201: Screen for duplicate samples in the collected original data, delete the duplicate samples, and then filter and remove the sample data that does not meet the following length standards:
[0039] 1. The length of the T cell antigen receptor sequence is in the range of 15 - 25 amino acids;
[0040] 2. The length of the class I human leukocyte antigen sequence is 34 amino acids;
[0041] 3. The length of the antigen peptide sequence is in the range of 8 - 15 amino acids.
[0042] Step S202: The data after cleaning is reorganized into a data set, and the processed TCR, HLA-I, and antigen peptide sequences are integrated into a training set, a validation set, and a test set to ensure the balance and diversity of the data set.
[0043] Step S203: The immunogenic data after cleaning generates negative sample data ten times the number of positive samples by randomly extracting peptide sample data to expand the training set, validation set, and test set.
[0044] Step S3: Feature encoding is performed on the T cell antigen receptor sequence, class I human leukocyte antigen sequence, and antigen peptide sequence based on the evolutionary scale model to generate feature matrices representing the above sequences. It includes:
[0045] The evolutionary scale model is a large-scale pre-trained protein language model based on the Transformer architecture, consisting of an input embedding layer, a Transformer encoder layer, and an output layer, and is used to process protein sequences and generate high-dimensional feature representations.
[0046] The input embedding layer encodes the input protein sequence, converting the amino acid sequence into a numerical vector of a fixed dimension. This process includes:
[0047] Let the input protein sequence be S=(s 1 ,s 2 ,s 3 ,...,s L ), where s i represents the i-th amino acid in the protein sequence. Each amino acid s i is mapped to the corresponding embedding vector e i ∈R d through the lookup table method, where d is the embedding dimension, set to 1280, and L is the length of the amino acid sequence.
[0048] The evolutionary scale model is implemented based on the Transformer architecture. To enhance the sequence order perception ability of the Transformer architecture, sinusoidal position encoding is specifically introduced. The calculation method of sinusoidal position encoding is:
[0049] PE (pos,2i) =sin(pos / 10000 2i / d ),
[0050] PE (pos,2i+1) =cos(pos / 10000 2i / d ),
[0051] where i represents the dimension index in the encoding vector (starting from 0); PE (pos,2i) and PE(pos,2i+1) Sine position encodings for odd and even dimensions respectively, where the odd and even dimensions enable the encoding to capture changes in different frequencies; pos is the current index position.
[0052] The position encoding is added to the amino acid embedding vector to form the final input embedding matrix X ∈ R L×d .
[0053] The input embedding matrix X undergoes feature extraction through 150 Transformer encoder layers. Each Transformer layer consists of a multi-head self-attention module, a feed-forward layer, residual connections, and normalization, as follows:
[0054] The multi-head self-attention module uses the self-attention mechanism to calculate the correlation between each amino acid within the sequence. Let the immunogenic query matrix Q, immunogenic key matrix K, and immunogenic value matrix V be linearly transformed from the input X:
[0055] Q = XW Q , K = XW K , V = XW V ,
[0056] W Q , W K , W V are the weight matrices for Q, K, and V respectively.
[0057] The attention weight calculation method of the attention mechanism is:
[0058]
[0059] d k is the dimension of each attention head
[0060] The multi-head attention mechanism concatenates the results of multiple attention heads and performs a linear transformation:
[0061] MultiHead(X) = Concat(head 1 ,..., head H )W O ,
[0062] where Concat(head 1 ,..., head H ) represents the concatenation of the results of multiple attention heads, head h is the h-th attention head, h = 1, 2,..., H, H is the total number of attention heads, and W O represents the trained projection transformation matrix used to re-project the concatenated feature dimensions back to the original dimension.
[0063] The forward feedback layer (FFN) is a two-layer fully connected layer network that uses the GELU activation function to enhance the non-linear expression ability for non-linear transformation:
[0064] FFN(X) = max(0, XW 1 + b 1 )W 2 + b 2 ,
[0065] where W 1 represents the weight matrix of the first-layer feedforward neural network, which projects the input feature matrix into a higher-dimensional space; W 2 represents the weight matrix of the second-layer feedforward neural network, which maps the high-dimensional features back to the original dimension; b 1 represents the bias vector of the linear transformation of the first-layer feedforward neural network, and b 2 represents the bias vector of the linear transformation of the second-layer feedforward neural network.
[0066] For the residual and normalization, residual connection is adopted to enhance the gradient propagation ability:
[0067] X new = X + Attention + FFN.
[0068] Layer normalization is adopted after each Transformer block to improve the training stability.
[0069] The output layer finally outputs in a residue-by-residue embedding manner, and its output dimension is [BatchSize, L, d], where BatchSize represents the number of immunogenic data samples input in one training iteration.
[0070] The evolutionary scale model is used to perform feature embedding on the complementarity-determining region 3 (CDR3) of the T cell antigen receptor to capture key sequence features and generate the T cell antigen receptor feature matrix; structural feature extraction is performed on the class I human leukocyte antigen sequence to generate the class I human leukocyte antigen feature matrix; context feature extraction is performed on the antigen peptide sequence to generate the antigen peptide sequence feature matrix. The T cell antigen receptor feature matrix, the class I human leukocyte antigen feature matrix, and the antigen peptide sequence feature matrix are regarded as the biological representation matrices of the corresponding biological protein molecules. The format of the T cell antigen receptor feature matrix is [N, 27, 1280], the format of the class I human leukocyte antigen feature matrix is [N, 35, 1280], the format of the antigen polypeptide sequence feature matrix is [N, 17, 1280], and the format of the joint feature representation matrix is [N, 79, 1280], where N represents the number of samples.
[0071] Step S4: The feature matrix is fused into a joint feature representation matrix through a bilinear attention network and input into the antigen immunogenicity prediction model to complete the prediction of the immunogenicity of the candidate antigen. It includes:
[0072] Step S401: Take the T cell antigen receptor feature matrix, the class I human leukocyte antigen protein feature matrix, and the antigen sequence feature matrix as inputs, and perform feature fusion through a bilinear attention network (BAN) to generate a comprehensive joint feature representation matrix O, regarded as the biological feature representation matrix of the interaction among the three, as shown in the following formula:
[0073]
[0074] where T is the T cell antigen receptor feature matrix,
[0075] F is the class I human leukocyte antigen feature matrix,
[0076] A is the antigen polypeptide sequence feature matrix,
[0077] W a is the interaction weight matrix of the T cell antigen receptor feature and the class I human leukocyte antigen,
[0078] W b is the interaction weight matrix of the class I human leukocyte antigen and the antigen polypeptide,
[0079] Y is the weight matrix mapped to the output feature space.
[0080] Step S402: The antigen immunogenicity prediction model includes three layers of one-dimensional convolutional neural networks and three layers of fully connected neural networks.
[0081] In this step, three convolution kernels [1, 2, 3] of the three layers of one-dimensional convolutional neural networks are set to capture features of different sizes. The output channel size of each convolution kernel is set to 1000. The three layers of one-dimensional convolutional neural networks perform feature extraction on the joint feature representation matrix, and each convolution kernel will output 1000 feature maps. The output feature maps of each layer are concatenated to finally obtain a one-dimensional feature matrix with a dimension of 3000.
[0082] The one-dimensional feature matrix with a dimension of 3000 obtained by the three layers of one-dimensional convolutional neural networks passes through 2048 and 256 neurons and then performs batch normalization. Each fully connected layer uses Dropout to prevent overfitting. The first two fully connected layers use the rectified linear unit as the activation function, indicating that the antigen immunogenicity score is generated by the last fully connected layer through the Softmax activation function.
[0083] The loss function L of the antigen immunogenicity prediction model is the cross-entropy loss function, and the maximum number of preset iterations Iterantion is setMAX is 100. The formula for the loss function L is as follows:
[0084]
[0085] where p y is the predicted probability of the true class y, representing the confidence of the model in predicting the true class; z y is the true label, representing the index of the correct class, and its value range is {1, 2,..., C}, where C represents the total number of classes; z j are the logits output by the model, that is, the scores for the j-th class, usually the output of the last layer of the network without passing through the activation function. The logits are the prediction results of the model without passing through Softmax.
[0086] Specific implementation method 2: Refer to Figure 2 to specifically describe this implementation method. The antigen immunogenicity prediction system based on the evolutionary scale model described in this implementation method includes:
[0087] Data sample extraction module: Obtain T cell antigen receptor, class I human leukocyte antigen, and antigen peptide sequence data based on unevenly distributed original data.
[0088] Processing module: Clean and deduplicate the data to meet the encoding requirements of the evolutionary scale model.
[0089] Encoding module: Encode the T cell antigen receptor, class I human leukocyte antigen, and antigen peptide sequence data into a T cell antigen receptor feature matrix, a class I human leukocyte antigen protein feature matrix, and an antigen sequence feature matrix.
[0090] Fusion module: Take the three feature matrices as inputs and use a bilinear attention network for fusion to generate a comprehensive joint feature representation matrix.
[0091] Prediction module: Input the joint feature representation matrix into the antigen immunogenicity prediction model to obtain the immunogenicity prediction result. The antigen immunogenicity includes whether the antigen can activate CD8+ T cells.
[0092] Specifically, the optimizer is selected as Adam, the initial learning rate is 0.00001, and the batch size is 128.
[0093] This implementation method can use various evaluation indicators to evaluate the performance of the model, such as AUC, AUPRC:
[0094]
[0095] Among them:
[0096]
[0097]
[0098] Among them, TP is the number of true positive samples, FP is the number of false positive samples, TN is the number of true negative samples, FN is the number of false negative samples, TPR is the true positive rate, FPR is the false positive rate, Precision is the precision rate, and Recall is the recall rate.
[0099] Specific Embodiment 3: The antigen immunogenicity prediction device based on the evolutionary scale model described in this embodiment. The antigen immunogenicity prediction device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the antigen immunogenicity prediction method based on the evolutionary scale model as described in Specific Embodiment 1.
[0100] Specific Embodiment 4: The computer storage medium described in this embodiment. At least one instruction is stored in the computer storage medium, and the at least one instruction is loaded and executed by the processor to implement the antigen immunogenicity prediction method based on the evolutionary scale model as described in Specific Embodiment 1.
[0101] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.
Claims
1. A method for predicting antigen immunogenicity based on an evolutionary scale model, characterized in that: include: The tested immunogenicity data are feature-encoded using an evolutionary scale model to generate a feature matrix of the tested immunogenicity data, wherein the tested immunogenicity data include: a T cell antigen receptor sequence, a class I human leukocyte antigen sequence, and an antigen peptide sequence; The feature matrices are fused into a joint feature representation matrix through a bilinear attention network and input into the antigen immunogenicity prediction model to complete the prediction of antigen immunogenicity.
2. The method for predicting antigen immunogenicity based on an evolutionary scale model according to claim 1, characterized in that: The method of using the evolutionary scale model to perform feature encoding on the tested immunogenicity data comprises: Using the evolutionary scale model to perform feature embedding on the T cell antigen receptor complementary determining region 3 sequence, capturing key sequence features, and generating a feature matrix of the T cell antigen receptor; Using the evolutionary scale model to extract structural features of the class I human leukocyte antigen sequence to generate a feature matrix of the class I human leukocyte antigen; The evolutionary scale model is used to extract context features from the antigen peptide sequence to generate a feature matrix of the antigen peptide.
3. The method for predicting antigen immunogenicity based on an evolutionary scale model according to claim 2, characterized in that: The evolutionary scaling model is a Transformer-based model, and its core architecture includes: an input embedding layer, a Transformer encoder layer, and an output layer.
4. The method for predicting antigen immunogenicity based on an evolutionary scale model according to claim 1, characterized in that: The step of fusing the feature matrices into a joint feature representation matrix through a bilinear attention network includes: The joint feature representation matrix O is generated by the following formula: Among them, T is the characteristic matrix of T cell antigen receptor, F is the characteristic matrix of class I human leukocyte antigen, A is the characteristic matrix of antigen peptide segment, and W a is the interaction weight matrix between T and F, W b is the interaction weight matrix between F and A, and Y is the weight matrix mapped to the output feature space.
5. The method for predicting antigen immunogenicity based on an evolutionary scale model according to claim 1, characterized in that: The antigen immunogenicity prediction model includes a three-layer one-dimensional convolutional neural network and a three-layer fully connected neural network; The three-layer one-dimensional convolutional neural network performs feature extraction on the joint feature representation matrix, and then inputs the extracted one-dimensional feature matrix into the three-layer fully connected neural network to finally obtain the antigen immunogenicity score.
6. The method for predicting antigen immunogenicity based on an evolutionary scale model according to claim 5, characterized in that: The convolution kernel sizes of each layer of the one-dimensional convolutional neural network are 1, 2, and 3 respectively. Each convolution kernel outputs 1000 feature maps, which are concatenated into a one-dimensional feature matrix with a dimension of 3000.
7. The method for predicting antigen immunogenicity based on an evolutionary scale model according to claim 5 or 6, characterized in that: The loss function L of the antigen immunogenicity prediction model is expressed as: Among them, p y is the predicted probability of the true category y, z y is the true label, z j is the j-th class logits output by the model. Logits are the prediction results of the model without Softmax. j = 1, 2, ..., C, where C represents the total number of categories. 8.Antigen immunogenicity prediction system based on evolutionary scale model, characterized in that: include: Coding unit: used for encoding the tested immunogenicity data by using the evolutionary scale model to generate a feature matrix of the tested immunogenicity data, wherein the tested immunogenicity data includes: T cell antigen receptor sequence, class I human leukocyte antigen sequence and antigen peptide sequence; Fusion unit: used to fuse the feature matrices into a joint feature representation matrix through a bilinear attention network, Prediction unit: used to input the joint feature representation matrix into the antigen immunogenicity prediction model to complete the prediction of antigen immunogenicity. 9.Antigen immunogenicity prediction device based on evolutionary scale model, characterized in that, The antigen immunogenicity prediction device comprises a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the antigen immunogenicity prediction method based on the evolutionary scale model as claimed in any one of claims 1 to 7.
10. A computer storage medium, characterized in that The computer storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the antigen immunogenicity prediction method based on the evolutionary scale model as described in any one of claims 1 to 7.