Supervised contrastive learning based interpretable mhc-ii peptide binding affinity prediction method

By combining supervised contrastive learning with a deep learning framework based on Transformer and residual modules, the binding core and anchor information of MHC-II molecules and peptides are identified, solving the prediction accuracy and interpretability problems of existing models. This achieves more efficient prediction of MHC-II peptide binding affinity, supporting vaccine design and immunotherapy.

CN117198405BActive Publication Date: 2026-04-07NANJING UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing MHC-II peptide binding affinity prediction models suffer from problems such as limited training data, low prediction accuracy, insufficient model interpretability, inadequate sequence information mining, and waste of computational resources, making it difficult to meet the needs of practical applications.

Method used

We employ a deep learning framework based on Transformer and residual modules, combined with supervised contrastive learning, to identify the binding core and anchor information between MHC-II molecules and peptides through feature encoding and pre-training, thereby improving prediction accuracy and model interpretability.

Benefits of technology

It significantly improved the prediction accuracy of MHC-II peptide binding affinity, reduced computational resource consumption, enhanced model stability and interpretability, and promoted the development of vaccine design and immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117198405B_ABST
    Figure CN117198405B_ABST
Patent Text Reader

Abstract

This invention discloses an interpretable MHC-II peptide binding affinity prediction method based on supervised contrastive learning, comprising: collecting MHC-II peptide affinity data and encoding the features of MHC-II molecules and peptide sequences into feature vectors; constructing a training set based on the data over time; constructing a pre-trained classification dataset Pre-Dataset based on the binding affinity values; constructing a deep learning framework based on Transformer and residual modules; inputting the Pre-Dataset into the constructed deep learning framework, pre-training it using supervised contrastive learning, and fine-tuning the learning; and performing prediction using the trained deep learning model. This invention assigns stable and reliable initial model weights to the model. Simultaneously, the Transformer module identifies the binding core and accurate anchor point locations based on the interaction characteristics between MHC-II molecules and peptides.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of bioinformatics predicting MHC-II peptide binding affinity, and particularly relates to an interpretable MHC-II peptide binding affinity prediction method based on supervised contrastive learning. BACKGROUND

[0002] Major histocompatibility complex (MHC) molecules, as a class of cell surface glycoproteins playing a core role in the immune system, can be divided into two categories: one is major histocompatibility complex (MHC-I) and the other is major histocompatibility complex (MHC-II). In particular, MHC-II molecules can bind to antigen peptides derived from pathogenic microorganisms or self-tissues to form MHC-antigen peptide complexes, and further present them to T cell receptors (TCR), thereby triggering and activating specific T cell immune responses to eliminate pathogens or repair damaged tissues. Therefore, accurately identifying MHC binding peptides is not only important for identifying epitopes that can effectively activate specific immune responses, but also provides key information for vaccine design, immunotherapy and research on immune-related diseases. Specifically, computational methods can help identify potential peptides that can trigger effective immune responses, enhance the safety and efficacy of vaccines, and guide the treatment of immune-related diseases such as autoimmune diseases, allergies and cancer.

[0003] Since it is time-consuming and laborious to determine the binding specificity of a large number of MHC-II molecules by experiment, it is particularly crucial to use efficient computational methods to predict the binding affinity of peptides to MHC-II molecules. Therefore, using the knowledge of bioinformatics, it is urgent to develop an intelligent prediction method for rapid and accurate prediction of MHC-II peptide binding sites from MHC-II molecules and peptide sequences, which has far-reaching significance for vaccine design and immunotherapy.

[0004] Currently, there are still few MHC-II peptide binding affinity prediction models based on MHC-II molecules and peptide sequences. By consulting relevant literature, it can be found that the current computational models designed specifically for MHC-II peptide binding affinity prediction based on MHC-II molecules and peptide sequences are: NetMHCIIpan-3.2, PUFFIN, MHCAttnNet, DeepSeqPanII and DeepMHCII, etc. Among them, NetMHCIIpan-3.2 (Jensen KK, Andreatta M, Marcatili P et al. Improved methods for predicting peptide binding affinity to MHC class II molecules, Immunology 2018; 154: 394-406.) is the first to use artificial neural networks to predict binding affinity. This method combines MHC-II pseudo sequence information and integrates 40 networks with different numbers of hidden nodes. Zeng et al. proposed a multi-model ensemble method PUFFIN (Zeng H, Gifford DK. Quantification of uncertainty in peptide-MHC binding prediction improves high-affinity peptide selection for therapeutic design, Cell systems 2019; 9: 159-166.e153.), which quantifies the prediction uncertainty and prioritizes peptides with high 'binding likelihood' to improve the accuracy of high-affinity peptide selection. MHCAttnNet (Venkatesh G, Grover A, Srinivasaraghavan G et al. MHCAttnNet: predicting MHC-peptide bindings for MHC alleles classes I and II using an attention-based deep neural model, Bioinformatics 2020; 36:i399-i406.) processes variable-length peptide sequences and MHC alleles through attention mechanism and Bi-LSTM encoder to improve MHC-I and II class peptide binding prediction.Similarly, DeepSeqPanII (Liu Z, Jin J, Cui Y et al. DeepSeqPanII: an interpretable recurrent neural network model with attention mechanism for peptide-HLA class II binding prediction, IEEE / ACM Transactions on Computational Biology and Bioinformatics 2021; 19: 2188-2196.) is a recurrent neural network model with attention mechanism that can predict peptide-HLA class II binding. Recently, You et al. proposed an interaction model DeepMHCII with binding core perception (You R, Qu W, Mamitsuka H et al. DeepMHCII: a novel binding core-aware deep interaction model for accurate MHC-II peptide binding affinity prediction, Bioinformatics 2022; 38: i220-i228.) to enhance MHC-II and peptide binding affinity prediction.

[0005] Current methods, although making some progress in prediction performance, still face four major problems: (1) the high diversity of MHC-II molecules and the scarcity of binding data lead to low prediction accuracy of MHC-II molecules with limited training data; (2) although some algorithms aim to enhance the model's interpretability by analyzing the attention module, there is still much room for improvement; (3) most methods perform single representation on MHC-II molecules and peptide sequences, which may lead to insufficient mining of sequence information; (4) most current algorithms enhance prediction performance by integrating multiple structurally similar models. However, if similar or better results can be achieved with fewer integrations or a single model, it will save computational resources for training and prediction. The current prediction accuracy is still far from large-scale practical applications, and further improvement is urgently needed. SUMMARY

[0006] The purpose of the present application is to propose an interpretable MHC-II peptide binding affinity prediction method based on supervised contrast learning, which improves the prediction accuracy of MHC-II peptide affinity.

[0007] The technical solution of the application is: an interpretable MHC-II peptide binding affinity prediction method based on supervised contrast learning, comprising the following steps:

[0008] Step 1: Collect MHC-II peptide affinity data, and encode the MHC-II molecules and peptide sequences respectively to convert them into feature vectors;

[0009] Step 2: Construct a training set according to the time of the data processed in step 1;

[0010] Step 3: Construct a pre-training classification dataset Pre-Dataset from the training set according to the binding affinity value;

[0011] Step 4: Construct a deep learning framework based on a Transformer module and a residual module;

[0012] Step 5: Input the Pre-Dataset into the constructed deep learning framework and pre-train it using supervised contrast learning;

[0013] Step 6: Optimize the deep learning framework trained in step 5;

[0014] Step 7: Input the MHC-II molecule sequence and the peptide sequence to be predicted into the deep learning model trained in step 6, and output the binding affinity prediction value of the corresponding MHC-II molecule and peptide through forward calculation of the model.

[0015] Preferably, the MHC-II molecules and peptide sequences are encoded by one-hot encoding, and the amino acids in the MHC-II molecules and peptide sequences are converted into corresponding feature vectors.

[0016] Preferably, the deep learning framework based on the Transformer module and the residual module comprises an MHC-II peptide interaction module, two groups of Transformer modules, a plurality of stacked residual modules, an average pooling layer, a linear transformation layer in the pre-training stage, and a sigmoid function for optimization learning in step 6.

[0017] Preferably, the Transformer module is:

[0018] Transformer(X)=Concat(head1,...,head h )

[0019]

[0020] In the formula, is a projection weight matrix; Represents the relative positional encoding of interactive features; d w Indicates the feature length; i represents the i-th attention head; d model d is the dimension of the input features of the Transformer; k is the dimension of the input features after projection transformation; X represents the input feature matrix of the Transformer module; Concat() is the tensor concatenation function; Attention() is the attention mechanism module.

[0021] Preferably, the residual module specifically comprises:

[0022]

[0023] Where, x l and x l+1 These represent the input and output of the l-th residual block, respectively. It is a set of weights for the l-th layer residual block. Represents the residual function. This represents the batch normalization function and the activation function.

[0024] Preferably, the deep learning framework is divided into an encoder network and a projection network, with the projection network located after the encoder network and serving as a linear transformation layer in the pre-training stage.

[0025] Preferably, the specific calculation process of the encoder network is as follows:

[0026]

[0027] In the formula, x and y represent peptide and MHC-II molecule pairs, and the encoder network maps x and y to a unit hypersphere. The representation vectors r and D on the vectors are... E The dimension of the vector output by the encoding network;

[0028] The projection network uses a multi-layer projection network to map r to vector z. The projection network is as follows:

[0029]

[0030] In the formula, Proj(·) contains two fully connected layers and an activation function, and the final output is normalized to the unit hypersphere. D P This represents the vector dimension output by the projection network.

[0031] Preferably, the D E =256, D P =16.

[0032] Preferably, the pre-training using supervised contrastive learning includes using the inner product between z samples as a distance metric between samples in the projection space, and calibrating the model parameters using supervised contrastive loss, specifically calculated as follows:

[0033]

[0034] In the formula, in a batch with a given sample size of N, {(x i y i ), label i} i=1,2,...,N This represents a sample pair containing a peptide, MHC-II, and its corresponding tag. These are the output features of the pre-trained model; i∈I≡{1, 2, ..., N} represents the index of the sample; A(i)≡I\{i} represents the set of indices not in set I; P(i)≡{p∈A(i): label p =label i} represents the set of indices of different categories from i in this batch; |P(i)| represents its quantity; . represents the inner product; This represents the temperature coefficient.

[0035] Preferably, optimizing the deep learning framework trained in step 5 specifically includes: using the MSE loss function to optimize the network weights and generate the optimal learning model.

[0036] Compared with the prior art, the significant advantages of this invention are as follows: (1) This invention designs a method to learn binding core and accurate anchor point information from multi-source interaction features (including embedding and one-hot encoding) through the Transformer module to generate attention-aware features; supervised contrastive learning pre-training is adopted to further improve the prediction accuracy of binding affinity. The combination of the two improves the prediction accuracy of the calculation model of MHC-II peptide binding affinity; (2) This invention uses supervised contrastive learning for pre-training, giving the model stable and reliable initial model weights, and accelerating the convergence speed of the model. At the same time, the Transform module identifies the binding core and accurate anchor point position based on the interaction features between MHC-II molecules and peptides. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of an interpretable MHC-II peptide binding affinity prediction method based on supervised contrastive learning.

[0038] Figure 2 This is a schematic diagram of the preprocessing process for the MHC-II peptide pre-training dataset.

[0039] Figure 3 This is a schematic diagram of the Transformer module structure. Detailed Implementation

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] like Figure 1 As shown, an interpretable method for predicting MHC-II peptide binding affinity based on supervised contrastive learning includes the following steps:

[0042] First, the latest MHC-II peptide affinity data were collected from the IEDB database, and MHC-II molecules and peptide sequences were encoded using one-hot encoding, converting the amino acids in the MHC-II molecules and peptide sequences into corresponding feature vectors. Second, the processed dataset was used to construct training, validation, and independent test sets according to time, and a pre-trained classification dataset (Pre-Dataset) was constructed based on the training set. Third, a deep learning framework based on Transformer and residual modules was constructed, and the network was iteratively trained using the Adam optimizer. Then, the Pre-Dataset was input into the constructed deep learning framework, and supervised contrastive learning was used to pre-train the Transformer and residual modules in the deep learning framework. Finally, the loss function was calculated using the Adam optimizer to update the network weights. Finally, through fine-tuning, the network weights were further fine-tuned on the affinity training dataset using MSE loss, based on the network weights trained above, to generate the corresponding learning model. The prediction process involves inputting the sequences of MHC-II molecules and peptides into the network model, and then outputting the binding affinity values ​​of the corresponding MHC-II molecules and peptides through forward computation of the network.

[0043] The aforementioned process will now be described in more detail with reference to the accompanying drawings.

[0044] Step 1: Data preprocessing. The latest MHC-II peptide affinity data were collected from the IEDB database. MHC-II molecules and peptide sequences were encoded using one-hot encoding, converting the amino acids in both molecules and sequences into corresponding feature vectors. The MHC-II molecule sequence is represented by a pseudo-sequence of 34 amino acids in contact with the peptide binding core; the first 15 come from the α chain, and the rest from the β chain. In one-hot encoding, the coding dimension is 20, representing 20 common amino acids.

[0045] Step 2: Construct the training set, validation set, and independent test set in chronological order. In this invention, data prior to April 2022 is used as the training set; the validation set consists of 1000 randomly selected samples from the training set; and the independent test set consists of publicly available data from April 2022 to April 2023, while excluding samples identical to those in the training set.

[0046] Step 3: Construct a pre-trained classification dataset (Pre-Dataset) from the training set based on the affinity values. This invention divides the training dataset into 10 categories according to normalized affinity values, such as category 0 representing affinity values ​​within the range [0.0, 0.1]. The specific pre-trained dataset processing flow is as follows: Figure 2 As shown.

[0047] Step 4: Construct a deep learning framework based on the Transformer module and the residual module. The network uses the Adam optimizer for iterative learning.

[0048] The deep learning framework takes as input the sequence features of MHC-11 molecules and peptides, including one-hot encoded features and learnable embeddings. It learns the interaction features between MHC-II molecules and peptides through interactive convolutions, and inputs these interaction features into two sets of Transformer modules, several stacked residual modules, an average pooling layer, a linear transformation layer in the pre-training stage, and a sigmoid function in the fine-tuning stage. In the prediction stage, it finally outputs the normalized binding affinity values ​​for the corresponding MHC-II molecules and peptides.

[0049] The Transformer module defined in this invention is as follows:

[0050] Transformer(X)=Concat(head1,...,head h )

[0051]

[0052] In the formula, It is the projection weight matrix; Represents the relative positional encoding of interactive features; d w Indicates the feature length; i represents the i-th attention head; d model d is the dimension of the input features of the Transformer; k is the dimension of the input features after projection transformation; X represents the input feature matrix of the Transformer module; Concat() is the tensor concatenation function; Attention() is the attention mechanism module. In this invention, d mdel =128, attention head h=4, d k =dmodel / h=32. A detailed diagram of the Transformer module is shown below. Figure 3 As shown.

[0053] The residual module in this invention is defined as follows:

[0054]

[0055] In the formula x l and x l+1 These represent the input and output of the l-th residual block, respectively. It is a set of weights for the l-th layer residual block. Represents the residual function. This represents the batch normalization function and the activation function.

[0056] Specifically, the deep learning framework is divided into encoder network and projection network. The linear transformation layer in the pre-training stage is the projection network, which follows the encoder network.

[0057] The encoder network calculation process is as follows:

[0058]

[0059] In the formula, the sample size of each batch is N. x and y represent peptide and MHC-II molecule pairs. The encoder network maps x and y to a unit hypersphere. The representation vectors r and D on the vectors are... E This represents the dimension of the vector output by the encoding network. In this invention, D is determined based on the complexity of the features. E Set to 256, the projection network uses a multi-layer projection network to map r to vector z. The projection network is represented by the following equation:

[0060]

[0061] In the formula, Proj(·) contains two fully connected layers and an activation function, and the final output is normalized to the unit hypersphere. D p =16 represents the vector dimension of the projection network output. Next, the inner product of z samples is used as the distance metric between them (between samples) in the projection space, and the model parameters are calibrated using contrastive loss. The specific calculation process is as follows:

[0062]

[0063] In the formula, in a given batch, {(x i y i ), label i} i= 1 ,2,...,NThis represents a sample pair containing a peptide, MHC-II, and its corresponding tag. These are the output features of the pre-trained model; i∈I≡{1,2,...,N} represents the index of the sample, often called the "anchor"; A(i)≡I\{i} represents the set of indices not in set i; P(i)≡{p∈A(i): label p =label i} represents the set of indices of different categories from i in this batch; |P(i)| represents its quantity; · represents the inner product; This represents the temperature coefficient. In the supervised contrastive learning pre-training stage, the network is divided into different classes based on affinity numerical information. Then, a supervised contrastive loss function is used to minimize the normalized embedding within each class, maximizing the distance between different classes. Furthermore, after completing the supervised contrastive learning pre-training, the projection network Proj(·) is removed. In the subsequent fine-tuning stage, further training is performed using the mean squared error (MSE) loss function based on the encoder Enc(·) weights.

[0064] Step 5: In the pre-training stage, the constructed Pre-Dataset dataset is input into the network structure built in Step 4. The network learns from the data through the stacked Transformer modules and residual modules. Finally, the loss function is calculated through the Adam optimizer, and the network weights are updated until the network loss no longer decreases. The best model file is then saved.

[0065] Step 6: Fine-tuning stage. Based on the network weights trained in Step 5, the network weights are further fine-tuned on the affinity training dataset using MSE loss to generate the corresponding learning model. The fine-tuning method used in this invention is to adjust the weight parameters of all layers of the network without freezing any network layer weights, and a relatively small learning rate is used during the fine-tuning stage.

[0066] Step 7: Input the MHC-II molecule sequence and peptide sequence to be predicted into the deep learning model trained in Step 6. Through the forward calculation of the model, output the predicted binding affinity values ​​of the corresponding MHC-II molecule and peptide.

[0067] This invention utilizes supervised contrastive learning for pre-training, assigning the model stable and reliable initial model weights. Simultaneously, the Transform module identifies the binding core and precise anchor sites based on the interaction characteristics between MHC-II molecules and peptides.

Claims

1. A method for predicting the interpretable binding affinity of MHC-II peptides based on supervised contrastive learning, characterized in that, Including the following steps: Step 1: Collect MHC-II peptide affinity data, and encode the features of MHC-II molecules and peptide sequences respectively, converting them into feature vectors; Step 2: Construct a training set from the data processed in Step 1 according to time; Step 3: Construct a pre-trained classification dataset (Pre-Dataset) from the training set based on the binding affinity values; Step 4: Construct a deep learning framework based on the Transformer module and the residual module; Step 5: Input the Pre-Dataset into the constructed deep learning framework and pre-train it using supervised contrastive learning; Step 6: Optimize the deep learning framework obtained in Step 5; Step 7: Input the MHC-II molecule sequence and peptide sequence to be predicted into the deep learning model trained in Step 6. Through the forward calculation of the model, output the predicted binding affinity values ​​of the corresponding MHC-II molecule and peptide. The deep learning framework based on Transformer modules and residual modules includes an MHC-II peptide interaction module, two sets of Transformer modules, several stacked residual modules, an average pooling layer, a linear transformation layer in the pre-training stage, and a sigmoid function for optimization learning in step 6. The Transformer module is: Transformer(X)=Concat(head1,…,head h ) In the formula, It is the projection weight matrix; Represents the relative positional encoding of interactive features; d w Indicates the feature length; i represents the i-th attention head; d model It is the dimension of the input features of the Transformer; d k is the dimension of the input features after projection transformation; X represents the input feature matrix of the Transformer module; Concat() is the tensor concatenation function; Attention() is the attention mechanism module.

2. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 1, characterized in that, MHC-II molecules and peptide sequences are encoded using a one-hot encoding method, converting the amino acids in the MHC-II molecules and peptide sequences into corresponding feature vectors.

3. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 2, characterized in that, The residual module is specifically: Where, x l and x l+1 These represent the input and output of the l-th residual block, respectively. It is a set of weights for the l-th layer residual block. Represents the residual function. This represents the batch normalization function and the activation function.

4. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 1, characterized in that, The deep learning framework is divided into an encoder network and a projection network. The projection network is located after the encoder network and is a linear transformation layer in the pre-training stage.

5. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 4, characterized in that, The specific calculation process of the encoder network is as follows: In the formula, x and y represent peptide and MHC-II molecule pairs, and the encoder network maps x and y to a unit hypersphere. The representation vectors r and D on the vectors are... E The dimension of the vector output by the encoding network; The projection network uses a multi-layer projection network to map r to vector z. The projection network is as follows: In the formula, Proj(·) contains two fully connected layers and an activation function, and the final output is normalized to the unit hypersphere. D P This represents the vector dimension output by the projection network.

6. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 5, characterized in that, The D E =256, D P =16.

7. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 5, characterized in that, The pre-training using supervised contrastive learning includes using the inner product between z samples as a distance metric between samples in the projection space, and calibrating the model parameters using supervised contrastive loss, specifically calculated as follows: In the formula, in a batch with a given sample size of N, {(x i ,y i ),label i } i=1,2,…,N A sample pair representing a peptide, MHC-II, and its corresponding tag; These are the output features of the pre-trained model; i∈I≡{1,2,…,N} represents the index of the sample; A(i)≡I\{i} represents the set of indices not in set I; P(i)≡{p∈A(i):label p =label i } represents the set of indices of different categories from i in this batch; |P(i)| represents its quantity; • Indicates the inner product; This represents the temperature coefficient.

8. The method for predicting interpretable MHC-II peptide binding affinity based on supervised contrastive learning according to claim 3, characterized in that, The optimization of the deep learning framework trained in step 5 specifically includes: using the MSE loss function to optimize the network weights and generate the optimal learning model.

Citation Information

Patent Citations

  • DNA-protein binding site prediction method based on self-attention residual network

    CN112382338A

  • Method and system for predicting protein-polypeptide binding site

    CN113593631A

  • Photographic image aesthetic style classification method based on improved self-supervised feature learning

    CN114140645A