A method for predicting the binding of complementarity determining region 3 to immune epitopes
Through multimodal heterogeneous graph modeling and GAT network optimization, combined with dynamic sample generation and multi-objective loss function, the data imbalance and loss function optimization problems in CDR3-immunoepitope binding prediction are solved, and high-precision prediction results are achieved, providing key residue analysis for cancer vaccine design and individualized treatment.
Patent Information
- Application Number
- CN202510807630.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The prior art has problems such as heterogeneous data modeling bias, extreme data imbalance, low negative sample quality and inefficient loss function optimization in the prediction of CDR3-immunoepitopes, resulting in insufficient prediction accuracy and difficult to meet the needs of cancer vaccine design and individualized immunotherapy.
Using multimodal heterogeneous graph modeling, graph attention network (GAT) and multi-objective loss function collaborative optimization methods, by constructing heterogeneous graphs of CDR3 and immunoepitopes, adversarial samples are generated dynamically, combining the focus loss function and AUC loss function, the classification accuracy and global sorting ability of the model are optimized, and interpretability analysis is performed.
It significantly improves the prediction accuracy of CDR3 binding to immunoepitopes, improves the accuracy of cancer vaccine design and individualized immunotherapy, and provides analytical guidance for key binding residues.
Smart Images

Figure CN120319302B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and specifically relates to a method for predicting the binding of complementarity determining region 3 (CDR3) to an immune epitope. Background Art
[0002] The specific binding of the T cell receptor (TCR) to immune epitopes through its CDR3 is a core mechanism of the adaptive immune response. The CDR3 loop, generated by V(D)J rearrangement of the TCR gene, exhibits extremely high sequence diversity, and its spatial conformation directly determines the strength of its interaction with anchor residues in the epitope-MHC complex. Traditionally, validation of CDR3-epitope binding has relied heavily on experimental methods such as surface plasmon resonance, tetramer staining, or in vitro functional assays. However, these methods have significant drawbacks, including high experimental costs and long cycle times, extremely low throughput, and differences in experimental conditions (such as pH and temperature) from the in vivo microenvironment, which can lead to false-negative results. To overcome these experimental bottlenecks, computational prediction methods have become a research hotspot. Early methods were primarily based on sequence similarity analysis, which only captured linear sequence features and ignored the three-dimensional conformational dynamics and topological dependencies of CDR3-epitope binding. For example, the spatial complementarity between the complementarity-determining region (CDR3 loop) of the TCR and the anchor residues of the epitope-MHC complex plays a decisive role in binding specificity, while sequence similarity models are unable to model such structural interactions. As the loop domain with the highest sequence diversity in the variable region of the T cell receptor, the spatial interaction between CDR3 and the epitope peptide is the molecular basis of immune recognition. The present invention breaks through the technical bottlenecks of traditional methods in heterogeneous topological characterization, extreme data imbalance, and difficult sample differentiation by integrating multimodal heterogeneous graph modeling, dynamic edge initialization, and a multi-objective collaborative optimization strategy that includes focal loss function (FocalLoss), AUC (Area Under ROC Curve) loss function, and binary cross entropy (BCE) loss function, providing high-precision computational tools for tumor neoantigen vaccine design and personalized immunotherapy.
[0003] In recent years, graph neural networks (GNNs), particularly graph convolutional networks (GCNs) and graph attention networks (GATs), have been introduced to the field of molecular interaction prediction. Existing GNNs typically model TCRs and epitopes as nodes in a homogeneous graph and extract features through neighborhood aggregation. However, these approaches still suffer from deficiencies such as insufficient modeling of heterogeneous features and limitations of attention mechanisms. At the data level, CDR3-immune epitope binding prediction faces two major challenges: extreme data imbalance and low-quality negative samples. At the algorithm optimization level, existing technologies suffer from the following deficiencies:
[0004] Single loss function: Most models use only a single loss function, such as BCE or AUC loss, which makes it difficult to optimize both classification accuracy and ranking ability. For example, BCE loss focuses on minimizing local classification errors but cannot directly optimize the AUC metric (global ranking ability). Simply maximizing AUC may lead to unstable predictions for samples near the classification threshold.
[0005] Rigid optimization strategies: Traditional optimizers (such as Adam) lack a collaborative update mechanism for multi-objective losses, leading to conflicting gradient directions for different loss functions, slow model convergence, and a tendency to fall into local optima. For example, in mixed-loss training, the gradient magnitude of the AUC loss is often much larger than the BCE loss, causing the latter to be suppressed and the model to overfit prematurely.
[0006] In summary, existing CDR3-immune epitope binding prediction technologies face multiple challenges, including heterogeneous data modeling bias, extreme data imbalance, poor negative sample quality, and inefficient loss function optimization. Developing a novel prediction framework that integrates multimodal feature representation, dynamic sample augmentation, multi-objective loss optimization, and interpretable analysis has become an urgent need to advance computational immunology. Summary of the Invention
[0007] To address the core issues identified in the background art, including heterogeneous topological representation bias, extreme data imbalance, poor negative sample quality, and inefficient loss function optimization, this paper proposes a method for predicting the binding of complementarity-determining region 3 to immune epitopes. This technical solution achieves breakthroughs through the following innovative designs:
[0008] Multimodal heterogeneous graph modeling: Constructing a heterogeneous graph with two types of nodes (CDR3 / immune epitope) and using GAT multi-head attention to explicitly encode cross-modal feature interactions;
[0009] Dynamic edge initialization: synthesize highly confusing negative samples based on feature space interpolation and dynamically optimize the training data distribution;
[0010] Multi-objective loss function collaborative optimization: Fusion of focus loss function and AUC loss function, balancing classification accuracy and global ranking ability through dynamic weighting;
[0011] Interpretability analysis: locating the interaction patterns between key CDR3 residues and immune epitope anchor sites based on GAT attention weights.
[0012] These designs improve the accuracy of predicting the binding specificity of CDR3 to immune epitopes. The present invention specifically includes the following steps:
[0013] Step 1: Data collection and multimodal feature extraction. Obtain CDR3 sequences and immune epitope peptides from databases such as pMTnet and VDJdb; process the data through denoising, length normalization, and multimodal feature extraction.
[0014] Step 2: Heterogeneous graph construction and dynamic edge initialization. Nodes are defined as CDR3 nodes and immune epitope nodes, with a feature dimension of 792 for CDR3 nodes and 1292 for immune epitope nodes. Edges are divided into two categories: positive edges (known binding pairs) and negative edges (initial random negative samples + dynamic adversarial samples).
[0015] Step 3: Design and train the graph attention network. The GAT structure is a 2-layer 8-head attention network with an output dimension of 128. Dynamic negative sample updates are then performed to generate adversarial samples and replace them with 20% of low-confidence negative samples.
[0016] Step 4: Collaborative optimization of multiple loss functions. This method invents a multi-loss function fusion strategy. The focus loss function focuses on difficult samples, improves the binary cross entropy to solve the problem of class imbalance, and maximizes the distance between positive and negative sample pairs by AUC loss. Finally, the focus loss and AUC loss are combined to perform dynamic weight fusion and dual optimizer collaborative training.
[0017] Step 5: Interpretability analysis and result output. Extract GAT attention weights and map them to CDR3 and immune epitope residues. Output residue pairs with contribution ≥ 0.15 as binding hotspots. Finally, output the results and generate a CSV file.
[0018] Step 6: Model validation and hyperparameter tuning. 5-fold cross validation was used, and AUC was used as the evaluation metric.
[0019] Step 7: Application output and experimental results. After training, the final CSV file is output to output the experimental results.
[0020] Accurate prediction of CDR3-immune epitope binding specificity is crucial for cancer immunotherapy, anti-infective vaccine development, and autoimmune disease research. Residue-level analysis based on GAT attention weights can accurately locate the key binding residue pairs between CDR3 and immune epitopes, providing direct guidance for vaccine target design. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of a method for predicting binding of complementarity determining region 3 to immune epitopes.
[0022] Figure 2 It is a schematic diagram of heterogeneous graph structure.
[0023] Figure 3 This is a diagram of the GAT multi-head attention mechanism. DETAILED DESCRIPTION
[0024] The following is a detailed description of the implementation steps of the present invention in conjunction with the accompanying drawings and examples. This method is implemented in Python 3.11.10 based on the torch 2.5.1 and torch-geometric 2.6.1 frameworks. The specific process includes seven steps: data acquisition and multimodal feature extraction, heterogeneous graph construction and dynamic edge initialization, graph attention network design and training, multi-objective loss function collaborative optimization, interpretability analysis and result output, model verification and hyperparameter tuning, and application output and experimental results. Figure 1 A flow chart of a method for predicting binding between complementarity determining region 3 and immune epitopes, with details of each part as follows:
[0025] Step 1: Data Collection and Multimodal Feature Extraction. This step begins with data collection and cleaning. Data sources include public databases such as pMTnet, VDJdb, and McPAS. TCR CDR3 β-chain sequences and their corresponding immune epitope peptide sequences are obtained to ensure coverage of human HLA class I / II-restricted epitopes. De-noising is then performed to remove sequences containing non-standard amino acid characters. Samples with abnormal lengths are also removed: TCR CDR3 sequence length is limited to 5-30 amino acids, and epitope peptide length is limited to 7-15 amino acids. Duplicate records are removed based on the MD5 hash value of the CDR3-immune epitope pair to avoid redundant training data.
[0026] The second part of this step involves sequence normalization and padding. Sequence lengths are normalized. For CDR3 sequences shorter than 30 amino acids, the special marker "[PAD]" is padded to position 30 at the C-terminus. For CDR3 sequences longer than 30 amino acids, the CDR3 core region is truncated at position 30 to preserve the CDR3 core region. Epitope sequences are maintained at their original length without padding or truncation.
[0027] The third part of this step is multimodal feature encoding. First, CDR3 feature extraction is performed: the CDR3 sequence is encoded using the pre-trained model TCRpeg to generate a 768-dimensional semantic embedding vector; additional physicochemical features are extracted, including amino acid hydrophobicity index, charge polarity, and sequence entropy, and spliced into a 24-dimensional manual feature vector. Then, epitope feature extraction is performed, and the ESM-2 pre-trained model is used to generate a 1280-dimensional structural embedding vector; combined with the three-dimensional structure of the epitope-MHC complex predicted by AlphaFold2, the solvent accessible surface area and secondary structure type are extracted to generate a 12-dimensional structural feature vector. Feature fusion is then performed to splice the semantic embedding with the manual features. The final feature dimension of the CDR3 node is 792 dimensions, and the epitope node is 1292 dimensions.
[0028] The fourth part of this step is the construction of positive and negative samples. Positive samples are directly based on verified CDR3-immune epitope binding pairs in the database and are marked as 1. Negative samples are divided into initial negative samples and dynamic negative samples. Initial negative samples are randomly selected non-repeating CDR3-immune epitope pairs to ensure no overlap with positive samples. Dynamic negative samples are adversarial samples generated dynamically during training based on the model's prediction confidence.
[0029] Step 2: Heterogeneous graph construction and dynamic edge initialization. First, define the heterogeneous graph structure, which includes node sets and edge sets, such as Figure 2 Schematic diagram of heterogeneous graph structure. The node set V contains two types of nodes:
[0030] CDR3 node (V t ): Each node corresponds to a TCR CDR3 sequence, and the feature vector dimension is 792;
[0031] Epitope node (V e ): Each node corresponds to an antigen epitope, and the feature vector dimension is 1292.
[0032] The edge set E contains positive edges and negative edges:
[0033] Positive side (E pos ): Connect known binding CDR3-immune epitope pairs, with edge weights initialized to 1;
[0034] Negative edge (E neg ): The edge corresponding to the initial negative sample has a weight of 0.
[0035] Construct a graph structure G=(V, E, X) based on the node set and edge set, where X is the node feature matrix.
[0036] Next, a dynamic negative sample pool is constructed. Initial negative samples are randomly generated negative pairs, ensuring they do not overlap with positive samples. 20% of the edge capacity is reserved for adversarial examples dynamically generated during training. The heterogeneous graph structure transforms CDR3-immune epitope binding prediction into an edge classification problem. Dynamic edge design lays the foundation for subsequent adversarial training.
[0037] Step 3: Design and train the graph attention network. This step uses the GAT architecture, which uses a 2-layer GAT, with each layer containing 8 independent attention heads. The GAT multi-head attention mechanism is shown in the figure below. Figure 3 shown.
[0038] Feature propagation formula: l Layer Node i Output features Calculated as: ;in, To normalize the attention coefficient, it is calculated by LeakyReLU. The dimensionality of the first layer output is 128, and the second layer is 128. The ELU activation function enhances nonlinear expression capabilities. The outputs of the eight attention heads in each layer are concatenated and then reduced to 128 dimensions through a fully connected layer. After each round of training, negative samples with a prediction probability between 0.4 and 0.6 are screened, that is, boundary samples that are difficult for the model to distinguish. Adversarial samples are generated according to the following formula: ;in is a randomly selected positive sample feature, is the interpolation coefficient, ~ u (0.2, 0.8); Negative sample replacement strategy: replace 20% of low-confidence negative samples in each round with a prediction probability < 0.1, adopt an incremental update strategy, and only retain the latest generated adversarial samples to avoid memory overflow.
[0039] Step 4: Collaborative optimization of multi-objective loss functions. In this step, this method invents a multi-loss function fusion strategy. Specifically, the focus loss function focuses on difficult samples to solve the problem of class imbalance, and improves the binary cross entropy loss function to dynamically reduce the loss contribution of easy-to-classify samples: ; Used to balance the weights of positive and negative samples, γ Adjust the focus strength for difficult samples, p is the predicted probability of the model;
[0040] AUC loss: maximize the distance between positive and negative sample pairs, the formula is: ;in γ is the interval threshold, is the logit value output by the model; combining the focus loss function and the AUC loss function, dynamic weight fusion is performed, and the total loss is: ; β is the weight coefficient, and the later stage of training focuses on optimizing the sorting ability. At the same time, this method performs dual optimizer collaborative training during training, using the Adam optimizer and the PESG optimizer. The Adam optimizer is responsible for optimizing the focal loss, and the PESG optimizer is dedicated to AUC loss optimization. After each round of training, the AUC is calculated on the validation set and the optimal model parameters are saved.
[0041] Step 5: Interpretability Analysis and Results Output. Extract the attention weight matrix of the last GAT layer and calculate the average cross-layer attention score for each CDR3-immune epitope pair. For CDR3 nodes, decompose the node-level attention weights to each amino acid residue based on the CDR3 sequence position. For epitope nodes, combine the ESM-2 token embedding to map the weights to each peptide residue. Set a threshold contribution score ≥ 0.15 to mark residue pairs as key interaction sites. Finally, output the results to a CSV file containing the following fields: TCR CDR3 sequence, epitope peptide sequence, predicted binding probability, key residue pairs, and contribution score.
[0042] Step 6: Model Validation and Hyperparameter Tuning. This method uses a 5-fold cross-validation strategy. The dataset is divided into five subsets, one of which is selected in turn as the test set, while the remaining subsets serve as the training set. The primary evaluation metric is the Area Under the Curve (AUC), which measures global ranking ability. After multiple training runs, the optimal focal loss function parameters determined for this method are: α∈0.5, 0.65, 0.6, and γ∈1.9, 3.2.
[0043] Step 7: Application Output and Experimental Results. After training, the final output is a CSV file containing the following fields: TCRCDR3 sequence, epitope peptide sequence, predicted binding probability, and classification label (1 / 0). Finally, after multiple experiments, on the pMTnet test set, with the focal loss function parameters α = 0.65 and γ = 1.9, the AUC reached 0.9252, a 1.16% improvement over the GTE method's AUC of 0.9136.
[0044] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for predicting the binding of complementarity determining region 3 to an immune epitope, characterized in that: Through multimodal heterogeneous graph modeling, dynamic edge initialization and multi-objective loss function collaborative optimization, the prediction accuracy of CDR3 and immune epitope binding specificity is significantly improved. The complementary determining region 3 is abbreviated as CDR3, including the following steps: Step 1: Data collection and multimodal feature extraction: CDR3 sequences and immune epitope peptides were obtained from databases such as pMTnet and VDJdb. The data were processed through denoising, length normalization, and multimodal feature extraction. Step 2: Heterogeneous graph construction and dynamic edge initialization. Nodes are defined as CDR3 nodes and immune epitope nodes. The feature dimension of CDR3 nodes is 792, and the feature dimension of immune epitope nodes is 1292. Edges are divided into two categories: positive edges and negative edges. Step 3: Design and train the graph attention network (GAT). The GAT structure is a 2-layer 8-head attention network with an output dimension of 128. Dynamic negative sample updates are then performed to generate adversarial samples and replace them with low-confidence negative samples at a ratio of 20%. Step 4: Multi-objective loss function collaborative optimization. This method invented a multi-loss function fusion strategy to solve the category imbalance problem and perform dual optimizer collaborative training. Step 5: Interpretability analysis and result output: extract GAT attention weights, map them to CDR3 and immune epitope residues, output residue pairs with contribution ≥ 0.15 as binding hotspots, and finally output the results and generate a CSV file; Step 6: Model validation and hyperparameter tuning, using 5-fold cross validation and AUC as the evaluation metric; Step 7: Apply the output and experimental results. After training, the final CSV file is output to output the experimental results.
2. The method for predicting the binding of a complementarity determining region 3 to an immune epitope according to claim 1, wherein: The process of heterogeneous graph construction and dynamic edge initialization is as follows: First, the heterogeneous graph structure is defined. The heterogeneous graph structure includes a node set and an edge set. The node set V contains two types of nodes: CDR3 Node V t : Each node corresponds to a TCR CDR3 sequence, and the feature vector dimension is 792; Epitope Node V e : Each node corresponds to an antigen epitope, and the feature vector dimension is 1292; The edge set E contains positive edges and negative edges: Positive side E pos : Connect known binding CDR3-immune epitope pairs, with edge weights initialized to 1; Negative side E neg : The edge corresponding to the initial negative sample has a weight of 0; Construct a graph structure G=(V, E, X) based on the node set and edge set, where X is the node feature matrix; Next, a dynamic negative sample pool is constructed. The initial negative samples are randomly generated negative pairs, and they are ensured not to be repeated with positive samples. 20% of the edge capacity is reserved for adversarial samples dynamically generated during training. The heterogeneous graph structure transforms CDR3-immune epitope binding prediction into an edge classification problem. Dynamic edge design lays the foundation for subsequent adversarial training.
3. The method for predicting the binding of a complementarity determining region 3 to an immune epitope according to claim 1, wherein: This method uses graph attention network design and training. In this step, the GAT architecture is used, with a 2-layer GAT, each layer containing 8 independent attention heads; Feature propagation formula: l Layer Node i Output features Calculated as: ;in, To normalize the attention coefficient, it is calculated by LeakyReLU. The dimensionality of the first layer output is 128, and the second layer is 128. The ELU activation function enhances nonlinear expression capabilities. The outputs of the eight attention heads in each layer are concatenated and then reduced to 128 dimensions through a fully connected layer. After each round of training, negative samples with a prediction probability between 0.4 and 0.6 are screened, that is, boundary samples that are difficult for the model to distinguish. Adversarial samples are generated according to the following formula: ;in is a randomly selected positive sample feature, is the interpolation coefficient, ~ u (0.2, 0.8); Negative sample replacement strategy: replace 20% of low-confidence negative samples in each round with a prediction probability < 0.1, adopt an incremental update strategy, and only retain the latest generated adversarial samples to avoid memory overflow.
4. The method for predicting the binding of a complementarity determining region 3 to an immune epitope according to claim 1, wherein: This method adopts the collaborative optimization of multi-objective loss functions. In this step, this method invented a multi-loss function fusion strategy, the specific content is: The focus loss function focuses on difficult samples, solves the problem of class imbalance, improves the binary cross entropy loss function, and dynamically reduces the loss contribution of easy-to-classify samples: ; Used to balance the weights of positive and negative samples, γ Adjust the focus strength for difficult samples, p is the predicted probability of the model; AUC loss: maximize the distance between positive and negative sample pairs, the formula is: ;in γ is the interval threshold, is the logit value output by the model; combining the focus loss function and the AUC loss function, dynamic weight fusion is performed, and the total loss is: ; β is the weight coefficient, and the later stage of training focuses on optimizing the sorting ability. At the same time, this method performs dual optimizer collaborative training during training, using the Adam optimizer and the PESG optimizer. The Adam optimizer is responsible for optimizing the focal loss, and the PESG optimizer is dedicated to AUC loss optimization. After each round of training, the AUC is calculated on the validation set and the optimal model parameters are saved.
5. The method for predicting the binding between a complementarity determining region 3 and an immune epitope according to claim 1, wherein: The process of interpretability analysis and result output is as follows: extract the attention weight matrix of the last layer of GAT, and calculate the average of its cross-layer attention scores for each CDR3-immune epitope pair; For CDR3 nodes, the node-level attention weight is decomposed into each amino acid residue according to the CDR3 sequence position; for epitope nodes, the weight is mapped to each residue in the peptide segment in combination with the token embedding of ESM-2; Residue pairs with a threshold contribution score ≥ 0.15 were marked as key interaction sites; finally, the results were output to generate a CSV file containing the following fields: TCR CDR3 sequence, epitope peptide sequence, predicted binding probability, key residue pairs, and contribution score.
Citation Information
Patent Citations
T cell receptor-epitope binding specificity prediction method based on multilevel knowledge distillation
CN119920319A
Children MPP auxiliary diagnosis system based on multi-modal time series data modeling
CN120015296A