HLA (human leukocyte antigen) and antigen peptide binding prediction method based on interactive attention
Through the HLA and antigen peptide binding prediction method based on interactive attention, integrating sequence and structural information, the accuracy and universality of HLA and antigen peptide binding prediction in the prior art are solved, and a more efficient prediction effect is achieved.
Patent Information
- Application Number
- CN202510760195.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing HLA and antigen peptide binding prediction methods have mispredictions in some cases, especially when new elution ligand data sets are validated, and existing models often ignore structural information between HLA and peptides, lacking versatility and accuracy.
Using the HLA and antigen peptide binding prediction method based on interactive attention, the HLA and antigen peptide binding data set is constructed, and the sequence embedding module, encoder, structural prediction model and normalization layer are used, combined with reverse adversarial training, sequence and structural information are integrated, multiple structural features are extracted and feature fusion is performed, to improve the generalization ability of the model.
A more comprehensive antigen immunogenicity assessment is achieved, improving prediction accuracy and model robustness, enabling high accuracy while improving coverage of positive cases and reducing sensitivity to subtle changes in inputs.
Smart Images

Figure CN120280001A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computational biology, and particularly relates to a method for predicting the binding of HLA and antigen peptides based on interactive attention. Background Art
[0002] The binding of antigen peptides to human leukocyte antigen (HLA) is an important step in antigen presentation. The polymorphism of HLA is a functional characteristic acquired during evolution, enabling the human immune system to respond to a variety of pathogens at the individual level. With the in-depth study of the immune system mechanism, experimental techniques such as mass spectrometry elution of HLA ligands and single-cell T sequencing have been developed to detect peptide-Human Leukocyte Antigen (pHLA) binding. However, these techniques are usually time-consuming, technically complex, and costly. Based on decades of laboratory research and the accumulation of HLA binding sequence data, many computational models for HLA-I alleles have been significantly improved. Existing models such as TransPHLA (Chu Y, Zhang Y, Wang Q, et al. A transformer-based model to predict peptide-HLA class I binding and optimize mutated peptides for vaccine design[J]. Nature Machine Intelligence, 2022, 4(3):300–311.), ACME (Yan H, Ziqiang W, Hailin H, et al. ACME: pan-specific peptide-MHC class I binding prediction through attention-based deep neural networks[J]. Bioinformatics, 2019, 35(23):4946–4954.), BigMHC (Alexander B A, Yunxiao Y, M.X S, et al. Deep neural networks predict class I major histocompatibility complex epitope presentation and transfer learn neoepitope immunogenicity[J]. Nature Machine Intelligence, 2023, 5(8):861–872.), NetMHCpan (Vanessa J, SinuP, Massimo A, et al. NetMHCpan-4.0: Improved Peptide-MHC Class I Interaction Predictions Integrating Eluted Ligand and Peptide Binding Affinity Data[J]. Journal of immunology, 2017, 199(9):3360–3368.) and MHCflurry (O’Donnell J T, Rubinsteyn A, Laserson U. MHCflurry 2.0: Improved pan-allele prediction of MHC I-presented peptides by incorporating antigen processing[J]. Cell Systems, 2020, 11(1):42-48.e7.) are widely used to predict the binding affinity between HLA and peptides. However, even these widely used algorithms may produce incorrect predictions in some cases (e.g., when validating on a new eluted ligand dataset). Many tools only focus on the binding between peptides and HLA, while ignoring the structural information between HLA and peptides, and still need to be experimentally verified. For the structure-based deep learning methods proposed in recent years, such as TransflGN (Hong N, Jiang D, Wang Z, et al. TransfIGN: A Structure-Based Deep Learning Method for Modeling the Interaction between HLA-A*02:01 and Antigen Peptides[J]. Journal of Chemical Information and Modeling, 2024, 64(13):5016-5027.) and NeoaPred (Jiang D, Xi B, Tan W, et al. NeoaPred: a deep-learning framework for predicting immunogenic neoantigen based on surface and structural features of peptide-human leukocyte antigen complexes[J]. Bioinformatics, 2024, 40(9):btae547.Models such as (etc.) all have their own limitations. TransflGN only focuses on the modeling of HLA-A*02:01 and antigen peptides and is not universal. NeoaPred focuses on the structural modeling of WT / Mut wild peptides and mutant peptides while ignoring the structural characteristics between peptides and HLA, and may lack accuracy at some new HLA loci. Summary of the Invention
[0003] Object of the Invention: The technical problem to be solved by the present invention is to provide a method for predicting the binding of HLA and antigen peptides based on interactive attention in view of the deficiencies of the prior art.
[0004] To solve the above technical problem, the present invention discloses a method for predicting the binding of HLA and antigen peptides based on interactive attention, including:
[0005] Step 1, collect HLA allele and unique peptide data, and construct an HLA and antigen peptide binding data set;
[0006] Step 2, construct an HLA and antigen peptide binding prediction model, and use the HLA and antigen peptide binding data set to train the prediction model to obtain a trained prediction model;
[0007] Step 3, input the HLA sequence and antigen peptide sequence into the prediction model to obtain the prediction result of the binding of HLA and antigen peptides.
[0008] Further, the HLA and antigen peptide binding prediction model in Step 2 includes a sequence embedding module, an encoder, a sequence feature fusion module, a structure prediction model, a structure feature extraction module, a structure sequence feature fusion module and a normalization layer. The sequence embedding module is used to perform amino acid embedding and position encoding on the HLA sequence and antigen peptide sequence respectively to obtain an HLA sequence embedding matrix and an antigen peptide sequence embedding matrix;
[0009] The encoder is used to extract HLA sequence features and antigen peptide sequence features from the HLA sequence embedding matrix and the antigen peptide sequence embedding matrix respectively;
[0010] The sequence feature fusion module is used to perform sequence feature fusion on the HLA sequence features and antigen peptide sequence features to obtain sequence fusion features;
[0011] The structure prediction model is used to predict the HLA structure, antigen peptide structure and spatial structure of the pHLA complex according to the HLA sequence and antigen peptide sequence;
[0012] The structure feature extraction module is used to extract structure features including interface area, number of hydrogen bonds, number of salt bridges, distance between mass centers and interaction energy from the HLA structure, antigen peptide structure and spatial structure of the pHLA complex;
[0013] The structure sequence feature fusion module is used to perform feature fusion on the sequence fusion feature and the structure feature to obtain a structure sequence fusion feature;
[0014] The normalization layer is used to obtain the HLA and antigen peptide binding prediction probability according to the structure sequence fusion feature.
[0015] Further, the training of the prediction model using the HLA and antigen peptide binding data set in step 2 includes adversarial training. Perturbations are introduced in the neighborhoods of the HLA sequence embedding and antigen peptide sequence embedding spaces respectively. The perturbations are generated along the ascending direction of the loss gradient and are constrained by the L2 norm. The perturbations require the prediction model to minimize both the empirical risk and the adversarial loss while minimizing the sensitivity to subtle changes in the input.
[0016] Further, constructing the HLA and antigen peptide binding data set in step 1 includes: to enhance the diversity of the data set and the robustness of the model, the randomly mismatched method and the unbound sequence pool method are used to perform data augmentation on the collected HLA allele and unique peptide data.
[0017] Further, the sequence embedding module performs amino acid embedding and position encoding on the HLA sequence and the antigen peptide sequence respectively to obtain the HLA sequence embedding matrix and the antigen peptide sequence embedding matrix, including:
[0018] Denote that the HLA sequence includes M1 amino acids. Each amino acid in the HLA sequence is mapped to an M-dimensional vector through a character embedding layer. ; Given that the amino acid order is crucial for the protein structure and function, sine-cosine position encoding is applied to each amino acid position; the amino acid embedding and the position encoding are added to obtain the HLA sequence embedding matrix. Finally, each HLA sequence is represented as matrix;
[0019] Denote that the antigen peptide sequence includes at most M2 amino acids. ; The antigen peptide sequence is padded to the maximum length M2 to unify the variable-length input. Each amino acid is also mapped to an M-dimensional embedding and the position encoding is added. After this processing, each antigen peptide sequence is represented as embedding matrix.
[0020] Further, the encoder extracts the HLA sequence feature and the antigen peptide sequence feature from the HLA sequence embedding matrix and the antigen peptide sequence embedding matrix respectively, including:
[0021] The encoder includes a self-attention layer, which is based on the self-attention mechanism. The self-attention mechanism learns the attention of all amino acid pairs in the input sequence. Among them, the HLA sequence embedding matrix is used as the query vector Q1, the key vector K1, and the value vector V1 respectively to extract HLA sequence features, and the antigen peptide sequence embedding matrix is used as the query vector Q1, the key vector K1, and the value vector V1 respectively to extract antigen peptide sequence features; the self-attention layer outputs the weighted sum of the value vector V1, and the weighted value is the attention score. The attention score is calculated by the normalized dot product of the query vector Q1 and the key vector K1, and then obtained through the softmax operation.
[0022] Further, the sequence feature fusion module performs sequence feature fusion on the HLA sequence feature and the antigen peptide sequence feature to obtain sequence fusion features, including:
[0023] The interactive attention mechanism is used to fuse the sequence features of the interaction between the antigen peptide segment and the HLA molecule. Among them, the HLA sequence feature matrix is used as the key vector K2 and the value vector V2, and the antigen peptide sequence feature matrix is used as the query vector Q2; in the interactive attention, the interactive vector I is used as the proxy of the query vector Q2, and the information from the key vector K2 and the value vector V2 is aggregated, and then the aggregation result is broadcast back to the query vector Q2; after being processed by the interactive attention mechanism, the sequence fusion features are obtained.
[0024] Further, the structure prediction model predicts the HLA structure, the antigen peptide structure, and the spatial structure of the pHLA complex according to the HLA sequence and the antigen peptide sequence, including: using the PepConf framework to calculate the pHLA spatial distance matrix to describe the interaction between the antigen peptide and the HLA-I molecule; using the intermolecular loss to forcibly limit the spatial distance between the antigen peptide and the HLA-I molecule.
[0025] Further, the structure sequence feature fusion module uses the cross-attention mechanism to fuse the sequence features and the structure features, including:
[0026] Construct the input features of the structure sequence feature fusion module, compress the sequence fusion features into a global sequence representation S through average pooling, and map the structure features including the interface area, the number of hydrogen bonds, the number of salt bridges, the distance between the mass centers, and the interaction energy into the structure embedding C through a multi-layer perceptron; use the sequence features as the query Q3 and the structure features as the key-value pair K3-V3 to establish a cross-modal cross-attention mechanism to achieve the adaptive alignment from sequence to structure: , Among them, Attention 3 represents the attention feature of the sequence structure, W Q is the sequence fusion feature, W K and W V are the structure features;
[0027] Through gated residual connections, the conservative binding patterns of the original sequence features are retained while injecting structural constraints to obtain the structure-sequence fusion feature z: , where LayerNorm( ) represents the layer normalization function, is a learnable gating parameter.
[0028] Furthermore, the normalization layer obtains the HLA and antigen peptide binding prediction probabilities based on the structure-sequence fusion feature, and the formula is as follows: , where p represents the prediction probability, represents the activation function, w represents the learnable weight vector, and b represents the bias term.
[0029] Beneficial effects: The prediction model of this application shows advantages over previous methods in the pHLA (peptide - human leukocyte antigen) binding prediction task:
[0030] First, by integrating sequence - and structure - based prediction tasks into a unified model, the method provided by this application can synchronously analyze and predict the potential of pHLA in terms of sequence and structure. This two - dimensional evaluation provides a more comprehensive perspective on antigen immunogenicity than existing methods and offers new insights into the quality of neoantigens that trigger immune responses.
[0031] Second, a new attention paradigm based on interactive attention is proposed. By improving the traditional query - key - value (QKV) attention structure and introducing element representation generation and interaction between elements, the sequence feature fusion module can fuse more expressive features and promote each other.
[0032] Finally, using backpropagation adversarial training effectively improves the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The following further describes the present invention in detail with reference to the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0034] Figure 1 It is a schematic diagram of the prediction model structure in a method for predicting HLA and antigen peptide binding based on interactive attention provided by an embodiment of this application.
[0035] Figure 2 It is a graph showing the performance evaluation results of a method for predicting HLA and antigen peptide binding based on interactive attention provided by an embodiment of this application and the prior art method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0037] This embodiment discloses a method for predicting the binding of HLA and antigen, based on interactive attention, which can be applied to the design of tumor antigen vaccines. The design of tumor antigen vaccines is one of the important methods of cellular immunotherapy. The design of tumor antigen vaccines refers to a treatment strategy that activates the patient's immune system against tumors by screening, synthesizing, or modifying specific tumor antigen peptides. Among them, the design of tumor antigen vaccines is mainly based on the antigen presentation process of the tumor immune cycle, and the core of this process corresponds to a binding: the binding of antigen peptide - human leukocyte antigen receptor. The binding of antigen peptide and HLA is a key step in antigen presentation and is a prerequisite for T cells to recognize foreign substances. When the antigen peptide binds to the HLA molecule to form a peptide - HLA (pHLA) complex and is presented on the cell surface, T cells can recognize the complex, thereby triggering a strong immune response. HLA is divided into two major categories: HLA class I (HLA-I) and HLA class II (HLA-II). HLA-I is encoded by three loci and is widely distributed on the surface of all nucleated cells, while HLA-II is only expressed in specific antigen-presenting cells. This embodiment focuses on HLA-I molecules (hereinafter referred to as HLA).
[0038] The first embodiment of this application discloses a method for predicting the binding of HLA and antigen peptide based on interactive attention, including:
[0039] Step 1, collect HLA allele and unique peptide data, and construct an HLA and antigen peptide binding data set;
[0040] The embodiments of this application focus on HLA class I molecules and construct a large-scale pHLA binding data set to promote the design of neoantigen vaccines and immunotherapy research. In this embodiment, four public HLA binder databases, IEDB (Immune Epitope Database), EPIMHC (Database of the relationship between MHC-binding peptides and T cell epitopes observed in real proteins), MHCBN (a curated database containing detailed information on major histocompatibility complex MHC binding, non-binding peptides, and T cell epitopes), and SYFPEITHI (Database of peptide sequences binding to class I and class II MHC molecules), are used. By removing duplicate sequences and abnormal data (such as missing values or asterisk marks), 610,654 pHLA binding pairs are finally obtained. This data set covers 142 HLA alleles and 279,924 unique peptides.
[0041] To efficiently evaluate the performance of the prediction model proposed in this embodiment, the pHLA binding dataset is divided into a training set and an independent test set at a ratio of 9:1. Finally, the training set contains 565,529 pairs of samples, covering 139 HLA alleles and 219,744 antigens; the independent test set contains 45,125 pairs of samples, covering 118 HLA alleles and 33,606 antigens. On this basis, to enhance the diversity of the dataset and the robustness of the model, this embodiment uses two methods, random mismatch and unbound sequence pool, to generate approximately twice the number of negative pHLA samples:
[0042] Random mismatch method: This method generates negative samples by shuffling the HLA and peptide sequences and randomly pairing them. Although this method may result in some negative samples having exactly the same peptide and HLA as positive samples, thus introducing false negative sample bias, since the probability of such samples occurring is low and the impact on the model is negligible, it can be ignored.
[0043] Unbound sequence pool method: By extracting long protein sequences from the Immune Epitope Database (IEDB) and randomly truncating them into short sequences, an unbound sequence pool is constructed. Subsequently, peptides are randomly drawn from the unbound sequence pool and paired with specific HLA alleles to generate negative samples.
[0044] The negative samples generated by these two methods each account for approximately 50% of the total negative sample volume. For each HLA allele, some negative peptides are derived from the truncated fragments of the IEDB HLA immunopeptidome source proteins, and the remaining negative samples are generated by random mismatch. Although there may be false negative samples, their proportion is extremely low and the impact on the results is negligible.
[0045] Step 2, construct an HLA and antigen peptide binding prediction model, and use the HLA and antigen peptide binding dataset to train the prediction model to obtain the trained prediction model;
[0046] As Figure 1 shown, the HLA and antigen peptide binding prediction model includes a sequence embedding module, an encoder, a sequence feature fusion module, a structure prediction model, a structure feature extraction module, a structure sequence feature fusion module, and a normalization layer.
[0047] The sequence embedding module is used to perform amino acid embedding and position encoding on the HLA sequence and the antigen peptide sequence respectively to obtain an HLA sequence embedding matrix and an antigen peptide sequence embedding matrix.
[0048] Denote that the HLA sequence includes M1 amino acids, and each amino acid in the HLA sequence is mapped to an M-dimensional vector through a character embedding layer. ;
[0049] The HLA sequence is fixed to a length of 34 amino acids, i.e., , and each amino acid is mapped to a 64-dimensional vector through the character embedding layer, i.e., . Given that the amino acid sequence is crucial for the structure and function of proteins, sine-cosine positional encoding is applied to each position. The amino acid embedding and positional encoding are added together to obtain the sequence embedding, and finally each HLA sequence is represented as a 34×64 matrix.
[0050] It is noted that the antigen peptide sequence includes at most M2 amino acids, ; in this embodiment , the antigen peptide sequence is padded to a maximum length of 15 to unify the variable-length input, and each amino acid is also mapped to a 64-dimensional embedding and positional encoding is added. After this processing, each antigen peptide sequence is represented as a 15×64 embedding matrix.
[0051] The encoder is used to extract HLA sequence features and antigen peptide sequence features from the HLA sequence embedding matrix and the antigen peptide sequence embedding matrix respectively;
[0052] The encoder is based on the self-attention mechanism, which has shown excellent ability in extracting global correlations and dependencies from amino acid sequences. The self-attention mechanism learns the attention of all possible amino acid pairs in the input sequence. Among them, the HLA sequence embedding matrix is used as the query vector Q1, the key vector K1, and the value vector V1 respectively to extract HLA sequence features, and the antigen peptide sequence embedding matrix is used as the query vector Q1, the key vector K1, and the value vector V1 respectively to extract antigen peptide sequence features. The attention scores are calculated by the normalized dot product of the query vector Q1 and the key vector K1, and then a softmax operation is performed. The output of the self-attention layer Attention 1 is the weighted sum of the value vector V1 weighted by the attention scores. The operation of the self-attention layer is represented in matrix form as follows: , where is the dimension of the vector (chosen as 64). The output of the self-attention block is passed through multiple fully connected layers, and the dimension of the gyro first increases and then decreases. For the antigen peptide sequence and the HLA sequence, the above encoder architecture is used for feature extraction.
[0053] In this embodiment, a masking mechanism is introduced when calculating the self-attention scores. Specifically, for peptides shorter than their maximum length, non-amino acid characters are not considered during model training. For this purpose, zero attention scores are assigned to these characters so that they do not affect the calculation of the attention scores. In the implementation of this embodiment, the encoder includes a single-layer single-head self-attention block.
[0054] The sequence feature fusion module is used to perform sequence feature fusion on HLA sequence features and antigen peptide sequence features to obtain sequence fusion features;
[0055] This embodiment proposes a new attention paradigm, namely Interaction Attention, which can effectively capture the complex correlations and global dependencies between different sequences. This mechanism is used to fuse the features of the interaction between peptide segments and HLA molecules. Different from the traditional attention mechanism, when fusing HLA-peptide segment features, the HLA sequence feature matrix serves as the key vector K2 and the value vector V2, and the antigen peptide sequence feature matrix serves as the query vector Q2. In Interaction Attention, the interaction vector I (Interaction Tokens) serves as the "agent" of the query vector Q2. By aggregating information from the key vector K2 and the value vector V2, and then broadcasting the aggregation result back to the query vector Q2. This process effectively reduces the computational complexity while retaining the global context modeling ability.
[0056] When calculating the attention scores, a masking mechanism is adopted to significantly reduce the computational overhead and accelerate the model convergence. In the specific implementation, a single-layer single-head Interaction Attention structure is adopted, and the direct similarity calculation between Q and K is reduced through proxy tokens, thereby improving the computational efficiency.
[0057] Given the input antigen peptide sequence feature matrix and HLA sequence feature matrix, the sequence feature fusion module performs feature fusion between the antigen peptide segment and the HLA molecule through the Interaction Attention mechanism.
[0058] In Interaction Attention, by introducing the interaction vector I, which serves as the "agent" of the query vector Q2, and reducing the computational complexity by aggregating information from the key vector K2 and the value vector V2. Specifically, the interaction vector I extracts global information from the key-value pair (K2, V2) for information aggregation.
[0059] The formula is as follows: , , where pooling represents the pooling operation, denotes mapping the interaction vector I to the same feature space as K2, such as a linear transformation operation, etc.; softmax() represents the normalization operation;
[0060] The interaction vector I passes the global information to the query vector Q2 to complete the broadcast of the aggregation result back to the query vector Q2.
[0061] The formula is as follows: , Among them, represents mapping the query vector Q2 to the same feature space as the interaction vector I, and W Q represents the sequence fusion feature.
[0062] The structure prediction model is used to predict the HLA structure, antigen peptide structure, and spatial structure of the pHLA complex based on the HLA sequence and antigen peptide sequence;
[0063] In this embodiment, the PepConf framework is adopted. This framework draws on the structure prediction method of AlphaFold2. According to the HLA sequence embedding matrix and antigen peptide sequence embedding matrix, the pHLA spatial distance matrix (two-dimensional matrix) is calculated to describe the interaction between the antigen peptide and the HLA-I molecule, thereby providing support for the construction of the peptide conformation. At the same time, the intermolecular loss is used to forcibly limit the spatial distance between the antigen peptide and the HLA-I molecule, and the loss function is as follows: , L represents the total loss of each example, and L pep represents the loss of the peptide itself, and L pHLA represents the loss between the peptide and the HLA-I molecule.
[0064] L pep consists of four auxiliary losses, as follows:
[0065] L FAPE is the Frame Aligned Point Error (FAPE) loss, which is used to evaluate the peptide atom coordinates relative to the coordinates of the peptide rigid group; L dist is the cross-entropy loss of the peptide internal residue distance distribution; L angle represents the side chain and backbone torsion angle loss; L viol is the structure violation loss. These auxiliary losses have been defined in AlphaFold2 and OpenFold.
[0066] L pHLA consists of two auxiliary losses, as follows:
[0067] L pHLA-FAPE is the FAPE loss for evaluating the peptide atom coordinates relative to the HLA rigid group; L pHLA-dist is the cross-entropy loss of the residue distance distribution between the peptide and the HLA.
[0068] This method has excellent performance and can successfully predict the spatial structure of the pHLA complex.
[0069] The structural feature extraction module is used to extract structural features including interface area, number of hydrogen bonds, number of salt bridges, distance between mass centers, and interaction energy from the HLA structure, antigen peptide structure, and spatial structure of the pHLA complex;
[0070] To systematically evaluate the binding characteristics between antigen peptides and HLA molecules, five key interaction features are extracted from the predicted complex structure in this embodiment: interface area, number of hydrogen bonds, number of salt bridges, distance between mass centers, and interaction energy. The following is a detailed description of each feature:
[0071] Specifically, first, the FreeSASA (where SASA stands for Solvent Accessible Surface Area, the solvent-accessible surface area of biomolecules) tool is used to quantify the reduction of the solvent-accessible area before and after binding, and the interface area InterfaceArea is used to characterize the burial degree of the peptide segment in the HLA groove; second, the number of cross-chain hydrogen bonds HBond is statistically counted through geometric thresholds (a distance of 3.5 Å and an angle of 120°) to reflect the stability of polar interactions; third, the number of salt bridges SaltBridge between positively and negatively charged residues is identified according to the 4.0 Å distance criterion to evaluate the contribution of electrostatic stability; subsequently, the distance between the mass centers of the two chains CoMDist (Euclidean distance) is calculated to describe the insertion depth and exposure degree of the peptide segment; finally, the interaction energy Contact is obtained by multiplying the number of cross-chain atomic contacts within 5.0 Å by –0.1 kcal / mol⁻¹ to provide a quick estimate of the binding strength. Without relying on expensive molecular mechanics solutions, this five-dimensional vector fully captures the structure-driven physical clues and injects fine spatial and energy information into the downstream multi-modal model.
[0072] To effectively combine the interaction information and its structural features between the antigen peptide sequence and the HLA molecule, a cross-modal feature fusion method based on the cross-attention mechanism is adopted.
[0073] This mechanism encodes the global binding pattern (such as motif conservation) for sequence features and captures local physical interactions (such as interface energy gradient) for structural features.
[0074] First, the input features of the structural sequence feature fusion module are constructed. Sequence features: The sequence fusion feature W Q is compressed into a global sequence representation through average pooling Structural features: Five physical indicators (InterfaceArea, HBond, SaltBridge, CoMDist, Contact) extracted from the pHLA complex are mapped into a structural embedding through a Multilayer Perceptron (MLP). : .
[0075] Doing so enables the non - linear transformation of the MLP to simulate the complex mapping relationship between physical indicators and binding free energy, such as the trade - off between the number of hydrogen bonds and entropy compensation.
[0076] Using the sequence feature as query Q3 and the structural feature as key - value pair K3 - V3, a cross - modal cross - attention mechanism is established to achieve the adaptive alignment from sequence to structure: , Among them, Attention 3 represents the attention feature of the sequence - structure, W Q is the sequence fusion feature, W K and W V are the structural features;
[0077] Through gated residual connection, the conserved binding pattern of the original sequence feature is retained, and at the same time, structural constraints are injected to obtain the structure - sequence fusion feature z:
[0078] , Among them, LayerNorm( ) represents the layer normalization function, is the learnable gating parameter, initialized to 0.5 to balance the contributions of the two modalities.
[0079] The normalization layer is used to obtain the prediction probability of HLA and antigen - peptide binding according to the structure - sequence fusion feature, and the formula is as follows: , Among them, p represents the prediction probability, represents the Sigmoid activation function, w represents the learnable weight vector, and b represents the bias term. In the specific implementation process, w and b can be initialized to small random values and automatically updated by the Adam optimization algorithm during training.
[0080] The training of the prediction model using the HLA and antigen peptide binding dataset in step 2 includes adversarial training. Specifically, by introducing small perturbations in the neighborhood of the sequence embedding space (instead of directly perturbing the original sequence), these perturbations are generated along the ascending direction of the loss gradient and are usually constrained by the L2 norm. The perturbations require the prediction model to minimize not only the empirical risk but also the adversarial loss, thereby reducing the sensitivity to subtle changes in the input. The adversarial loss is defined by the formula:
[0081] where D represents a function that measures the difference between two distributions, represents the probability that the model predicts the input x as label y. denotes the trainable parameters of the prediction model provided in this implementation, with random initial values, which will participate in the gradient update during the calculation of adversarial perturbations, denotes a copy of the model parameters that are "frozen" (stop-gradient) during the generation of adversarial perturbations and are used to calculate the output distribution of the original input and do not participate in the optimization of r, denotes the prediction probability distribution of the model for class y when the input is x and the parameter is r. vadv r is the adversarial perturbation in the reverse direction for the sample x, and this perturbation maximizes the difference between and by moving along the ascending direction of the gradient. denotes unlabeled samples, r represents a small perturbation vector for the input x, and the direction that maximizes the deviation of the output distribution is sought. denotes the upper limit of the L2 norm of the perturbation vector r.
[0082] Adversarial perturbations are applied to both the HLA sequence embedding and the neighborhood of the antigen peptide sequence embedding space, enabling the encoder to learn to extract discriminative features. Ablation experiments confirm that adversarial learning in the reverse direction improves the performance of the prediction model.
[0083] Step 3: Input the HLA sequence and the antigen peptide sequence into the prediction model to obtain the prediction result of the binding of HLA and the antigen peptide.
[0084] To evaluate the prediction performance of the proposed model in HLA-antigen binding, the prediction model (abbreviated as TransPed) proposed in this embodiment was compared with existing models TransPHLA, NetMHC-pan_BA, and NetMHCpan_EL (Birkir R, Bruno A, Sinu P, et al. NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data[J]. Nucleic acids research, 2020, 48(W1):W449-W454.), MHLAPre (Xu L, Yang Q, Dong W, et al. Metalearning for mutant HLA class I epitope immunogenicity prediction to accelerate cancer clinical immunotherapy[J]. Briefings in Bioinformatics, 2024, 26(1):bbae625.), and NetMHCcons (Edita K, Claus L, Ole L, et al. NetMHCcons: a consensus method for the major histocompatibility complex class I predictions[J]. Immunogenetics, 2012, 64(3):177-86.) on an independent test set. The performance metrics included accuracy, recall, area under the PR curve, and true negative rate.
[0085] Figure 2The performance evaluation results shown indicate that on the independent test set, the TransPed model provided in this embodiment significantly outperforms current mainstream methods in all key performance indicators. Specifically, the precision of TransPed reaches 0.92, a 10 percentage point increase compared to 0.82 of TransPHLA; the recall rate reaches 0.90, about 5 percentage points higher than that of the second place; the area under the PR curve (PR-AUC) is as high as 0.95, leading the range of 0.85 - 0.88 of NetMHCpanBA / EL; the true negative rate also reaches 0.89, nearly 8 percentage points higher than the average of other models. This result shows that TransPed can not only improve the coverage ability of positive examples while maintaining high accuracy, but also demonstrate stronger robustness in distinguishing difficult negative samples, providing a more effective and reliable solution for pHLA affinity prediction.
[0086] TransPed significantly improves the prediction ability of neoantigen peptides through a number of innovative designs. It organically integrates sequence information and structural information, comprehensively analyzing the peptide-MHC binding characteristics from both the sequence-structure dual perspectives. This benefits from the introduction of the interactive attention mechanism and positional encoding: the former highlights the interaction between the key sites of the peptide and the receptor in the attention framework, and the latter encodes the amino acid sequence by integrating spatial position information, enabling the model to better capture the structural environment where the residues are located. Through the cross-attention fusion of structural features and sequence features, TransPed can link sequence patterns with three-dimensional conformations, accurately identifying the key sequence-structure factors that affect peptide-MHC affinity. This is not available in other models. For example, models such as TransPHLA, ACME, and BigMHC only focus on sequence features and ignore the structural information between HLA and peptides, and still need to be verified through experiments. For recently proposed structure-based deep learning methods, such as models like TransflGN and NeoaPred, there are their respective limitations. TransflGN only focuses on the modeling of HLA-A*02:01 and antigen peptides and is not general. NeoaPred focuses on the structural modeling of WT / Mut wild peptides and mutant peptides while ignoring the structural features between peptides and HLA, and may lack accuracy at some new HLA loci.
[0087] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the inventive content of a method for predicting the binding of HLA and antigen peptides based on interactive attention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.
[0088] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium and includes several instructions for causing a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, a MUU, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present invention.
[0089] The present invention provides a method for predicting the binding of HLA and antigen peptides based on interactive attention. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.
Claims
1. A method for predicting the binding of HLA and antigen peptides based on interactive attention, characterized in that, Comprising: Step 1, collecting HLA allele and unique peptide data, and constructing an HLA and antigen peptide binding dataset; Step 2, constructing an HLA and antigen peptide binding prediction model, and training the prediction model using the HLA and antigen peptide binding dataset to obtain a trained prediction model; Step 3, inputting the HLA sequence and the antigen peptide sequence into the prediction model to obtain the prediction result of the binding of HLA and antigen peptide.
2. The HLA and antigen peptide binding prediction method based on interactive attention according to claim 1, wherein The HLA and antigen peptide binding prediction model in Step 2 includes a sequence embedding module, an encoder, a sequence feature fusion module, a structure prediction model, a structure feature extraction module, a structure sequence feature fusion module, and a normalization layer. The sequence embedding module is used to perform amino acid embedding and position encoding on the HLA sequence and the antigen peptide sequence respectively to obtain an HLA sequence embedding matrix and an antigen peptide sequence embedding matrix; The encoder is used to extract HLA sequence features and antigen peptide sequence features from the HLA sequence embedding matrix and the antigen peptide sequence embedding matrix respectively; The sequence feature fusion module is used to perform sequence feature fusion on the HLA sequence features and the antigen peptide sequence features to obtain sequence fusion features; The structure prediction model is used to predict the HLA structure, the antigen peptide structure, and the spatial structure of the pHLA complex according to the HLA sequence and the antigen peptide sequence; The structure feature extraction module is used to extract structure features including interface area, number of hydrogen bonds, number of salt bridges, distance between mass centers, and interaction energy from the HLA structure, the antigen peptide structure, and the spatial structure of the pHLA complex; The structure sequence feature fusion module is used to perform feature fusion on the sequence fusion features and the structure features to obtain structure sequence fusion features; The normalization layer is used to obtain the HLA and antigen peptide binding prediction probability according to the structure sequence fusion features.
3. The HLA and antigen peptide binding prediction method based on interactive attention according to claim 2, wherein The training of the prediction model using the HLA and antigen peptide binding dataset in Step 2 includes adversarial training. Perturbations are introduced respectively in the neighborhood of the HLA sequence embedding and antigen peptide sequence embedding spaces. The perturbations are generated along the direction of the loss gradient and are constrained by the L2 norm; the perturbations require the prediction model to minimize the adversarial loss while minimizing the empirical risk.
4. The HLA and antigen peptide binding prediction method based on interactive attention according to claim 3, wherein The construction of the HLA and antigen peptide binding dataset in Step 1 includes: performing data augmentation on the collected HLA allele and unique peptide data by means of random mismatch and unbound sequence pool.
5. A method for predicting the binding of HLA and antigen peptides based on interactive attention according to claim 4, characterized in that The sequence embedding module performing amino acid embedding and position encoding on the HLA sequence and the antigen peptide sequence respectively to obtain an HLA sequence embedding matrix and an antigen peptide sequence embedding matrix includes: It is noted that the HLA sequence consists of M1 amino acids, and each amino acid in the HLA sequence is mapped to an M-dimensional vector through a character embedding layer. ; Apply sine-cosine positional encoding to each amino acid position; Add the amino acid embedding and the positional encoding to obtain the HLA sequence embedding matrix, and finally each HLA sequence is represented as matrix; The antigen peptide sequence contains at most M2 amino acids. ; The antigen peptide sequence is padded to the maximum length M2 to unify the variable-length input. Each amino acid is also mapped to an M-dimensional embedding and a positional encoding is added. After this processing, each antigen peptide sequence is represented as an embedding matrix.
6. The HLA and antigen peptide binding prediction method based on interactive attention according to claim 5, wherein The encoder extracting HLA sequence features and antigen peptide sequence features from the HLA sequence embedding matrix and the antigen peptide sequence embedding matrix respectively includes: The encoder includes a self-attention layer, which is based on the self-attention mechanism. The self-attention mechanism learns the attention of all amino acid pairs in the input sequence. Among them, the HLA sequence embedding matrix is used as the query vector Q1, the key vector K1, and the value vector V1 respectively to extract the HLA sequence features, and the antigen peptide sequence embedding matrix is used as the query vector Q1, the key vector K1, and the value vector V1 respectively to extract the antigen peptide sequence features; the self-attention layer outputs the weighted sum of the value vector V1, and the weighted value is the attention score. The attention score is calculated by the normalized dot product of the query vector Q1 and the key vector K1, and then a softmax operation is performed to obtain it.
7. A method for predicting the binding of HLA and antigen peptides based on interactive attention according to claim 6, wherein The sequence feature fusion module performs sequence feature fusion on the HLA sequence features and the antigen peptide sequence features to obtain sequence fusion features, including: The interactive attention mechanism is used to fuse the sequence features of the interaction between the antigen peptide segment and the HLA molecule. Among them, the HLA sequence feature matrix is used as the key vector K2 and the value vector V2, and the antigen peptide sequence feature matrix is used as the query vector Q2; in the interactive attention, the interactive vector I is used as the proxy of the query vector Q2, and by aggregating the information from the key vector K2 and the value vector V2, and then broadcasting the aggregation result back to the query vector Q2; after being processed by the interactive attention mechanism, sequence fusion features are obtained.
8. A method for predicting the binding of HLA and antigen peptides based on interactive attention according to claim 7, characterized in that, The structure prediction model predicts the HLA structure, the antigen peptide structure, and the spatial structure of the pHLA complex according to the HLA sequence and the antigen peptide sequence, including: using the PepConf framework to calculate the pHLA spatial distance matrix to describe the interaction between the antigen peptide and the HLA-I molecule; using the intermolecular loss to limit the spatial distance between the antigen peptide and the HLA-I molecule.
9. A method for predicting the binding of HLA and antigen peptides based on interactive attention according to claim 8, characterized in that The structure sequence feature fusion module uses the cross-attention mechanism to fuse the sequence features and the structure features, including: Construct the input features of the structure sequence feature fusion module. The sequence fusion features are compressed into a global sequence representation S through average pooling, and the structure features including the interface area, the number of hydrogen bonds, the number of salt bridges, the distance between the mass centers, and the interaction energy are mapped into a structure embedding C through a multi-layer perceptron; using the sequence features as the query Q3 and the structure features as the key-value pair K3-V3, a cross-modal cross-attention mechanism is established to achieve the adaptive alignment from sequence to structure: , Among them, Attention 3 represents the attention feature of the sequence structure, W Q is the sequence fusion feature, W K and W V are the structural features; Through gated residual connection, the conserved binding pattern of the original sequence features is retained, and at the same time, structural constraints are injected to obtain the structure sequence fusion feature z: , Among them, LayerNorm( ) represents the layer normalization function, which is a learnable gating parameter.
10. The HLA and antigen peptide binding prediction method based on interactive attention according to claim 9, characterized in that, The normalization layer obtains the binding prediction probability of HLA and the antigen peptide according to the structure sequence fusion feature, and the formula is as follows: , where p represents the predicted probability, represents the activation function, w represents the learnable weight vector, and b represents the bias term.
Citation Information
Patent Citations
Aspect-level sentiment analysis method and system based on multi-head attention and graph convolution network
CN112633010A
Human TCR / HLA-I / Peptide ternary complex interaction identification prediction method and system
CN116597903A
Compound structure prediction method and device, computer equipment and storage medium
CN116959577A
Cited By
Multi-modal deep learning-based MHC presentation peptide fragment prediction method and system
CN120452555A
Antigen peptide immunogenicity prediction method and application thereof
CN122436004A