A peptide-HLAI allele affinity prediction system based on protein language model and dual multi-instance attention mechanism

CN119132408BActive Publication Date: 2026-09-25WUHAN KANGSHENG BEITAI BIOLOGICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411045983.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-09-25
Estimated Expiration
2044-08-01

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了基于蛋白质语言模型和双重多实例注意力机制的肽-HLA I类等位基因亲和力预测系统,具备改善HLA抗原呈递的预测效果等优点,解决了上述技术问题

Benefits of technology

[0052]1、本发明通过对于包含HLA I类分子与多肽的分子序列,利用自训练的表位肽蛋白质语言模型和HLA I类分子蛋白质语言模型来获取氨基酸序列的特征编码向量,来更好的表征分子序列;对氨基酸之间可变长度的依赖问题,利用改进的动态选择机制LSTM学习序列中不同距离氨基酸之间的依赖特征。对于质谱数据中存在的多特异性性质,采用双重多实例注意力模块学习等位基因不同粒度(氨基酸层面和等位基因层面)的重要性,达到了改善HLA抗原呈递的预测效果的有益效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132408B_ABST
    Figure CN119132408B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biological information and computational immunology, and discloses a peptide-HLA class I allele affinity prediction system based on a protein language model and a double multi-instance attention mechanism, which comprises a sequence encoding and distributed representation module of a peptide and a human leukocyte antigen HLA class I molecule, a multi-level feature extraction module of a dynamic selection mechanism long short-term memory network (LSTM), a double multi-instance attention module, and a feature fusion and affinity prediction module.The present application uses a self-training epitope peptide protein language model and an HLA class I molecule protein language model to obtain feature encoding vectors of amino acid sequences for a molecular sequence containing an HLA class I molecule and a polypeptide, so as to better represent the molecular sequence; and the improved dynamic selection mechanism LSTM is used to learn the dependent features between amino acids at different distances in the sequence, so as to improve the prediction effect of HLA antigen presentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and computational immunology, specifically to a peptide-HLA class I allele affinity prediction system based on a protein language model and a dual multi-instance attention mechanism. Background Technology

[0002] Major histocompatibility complex (MHC) molecules, as key components of the immune system, are responsible for detecting and presenting antigenic peptides to the cell surface, triggering T cell-mediated immune responses. Human leukocyte antigens (HLA) are the HLA genes and their encoded products in the human body. HLA molecules are divided into class I and class II, which interact with intracellular and extracellular antigens, respectively. A deep understanding of the interaction between HLA molecules and antigenic peptides, especially the peptide-HLA binding process, is of great significance. However, traditional experimental methods for studying peptide-HLA binding, such as affinity assays and liquid chromatography-mass spectrometry, are time-consuming, costly, and labor-intensive. Therefore, various computational prediction methods have been developed. Among them, NetMHCpan-4.1 and MHCflurry-2.0 are the latest algorithms widely used for predicting HLA class I peptide binding. NetMHCpan-4.1 consists of several shallow neural networks. Training and prediction are performed by integrating multiple shallow neural networks. However, because the data comes from a limited number of HLA molecules and peptides, it may not fully cover all HLA molecular subtypes and their interaction mechanisms with peptides, resulting in poor prediction performance for some subtypes. MHCflurry-2.0 uses in vitro binding affinity data or single HLA allele (SA) mass spectrometry data as training data, but cannot use multiple HLA allele (MA) mass spectrometry data for model training, resulting in a large amount of multi-HLA allele mass spectrometry data not being effectively utilized.

[0003] These problems lead to a high false positive rate in existing tools, necessitating the development of new methods for predicting peptide-HLA binding affinity. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] To address the shortcomings of existing technologies, this invention provides a peptide-HLA class I allele affinity prediction system based on a protein language model and a dual multi-instance attention mechanism, which has advantages such as improving the prediction effect of HLA antigen presentation and solving the above-mentioned technical problems.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the present invention provides the following technical solution: a peptide-HLA class I allele affinity prediction system based on a protein language model and a dual multi-instance attention mechanism, the prediction system comprising: a sequence encoding and distributed representation module for peptides and HLA class I molecules, a multi-level feature extraction module using a dynamic selection mechanism LSTM, a dual multi-instance attention module, and a feature fusion and affinity prediction module;

[0008] The sequence encoding and distributed representation module for peptides and HLA class I molecules includes: an epitope peptide protein language model and an HLA class I molecule protein language model;

[0009] The dynamic selection mechanism LSTM multi-level feature extraction module is: adding a dynamic selection mechanism to LSTM;

[0010] The dual multi-instance attention module comprises two sub-modules: Module 1, Seq-Attention, is used to learn the attention weights between different amino acids in a single sequence; Module 2, Bag-Attention, is used to dynamically calculate the relative weight of each allele's contribution to affinity among multiple alleles.

[0011] The feature fusion and affinity prediction module is used to fuse allele features processed by the dual multi-instance attention module and perform nonlinear transformation and feature compression on them.

[0012] Preferably, the epitope peptide protein language model and the HLA class I molecular protein language model are based on ProLLaMa and incorporate the concept of transfer learning. They are specifically optimized for epitope peptide sequence datasets and HLA class I molecular sequence datasets. Through the collected peptide-HLA I binding dataset, the peptide sequences and HLA molecular sequences are subjected to the Masked Language Modeling (MLM) task, respectively. The sequence features of peptides and HLA I molecules are learned from the data in an unsupervised manner. This not only inherits the functions of ProLLaMa, but also further refines the understanding of the unique sequence patterns and contextual relationships of epitope peptides and HLA I molecules. The epitope peptide protein language model and the HLA class I molecular protein language model convert the original amino acid sequences into high-dimensional, continuous vector representations, extract key structural information of the amino acid sequences, and also contain potential functional features.

[0013] Preferably, the dynamic selection mechanism LSTM multi-level feature extraction module is as follows: from the set of historical hidden states before the current input of the amino acid sequence, the optimal hidden state is selected and connected to the current input. The selection method of the optimal hidden state at the historical time is: add a linear neural network layer, and then use SoftMax normalization to obtain the score of each state. The one with the highest score is considered optimal. The dynamic selection mechanism replaces the original LSTM model's method of directly connecting the current input to the hidden state of the previous time step, and learns useful information at long distances in the sequence to solve the problem of variable-length dependencies between amino acids in the current affinity prediction scenario.

[0014] Preferably, the dual multi-instance attention module, upon receiving peptide sequence feature vectors related to different alleles from the dynamically selected LSTM layer, learns the importance of different features through Dual Multi-instance Attention and DMI-Attention training, at both the amino acid and allele levels of the model, adaptively highlighting allele features that have a greater impact on affinity while weakening secondary influencing factors.

[0015] Preferably, the feature fusion and affinity prediction module integrates the allele features processed by the Dual Multi-instance Attention (DMI-Attention) module into a global feature vector. This vector is then subjected to nonlinear transformation and feature compression through a series of fully connected layers (FCNs). The role of the fully connected layers is to further refine key features and map them to the actual affinity range. The final fully connected layer generates the binding probability of each HLA class I allele to a given peptide sequence.

[0016] Preferably, the sequence encoding and distributed representation module expression of the peptide and HLA class I molecules is as follows:

[0017] Let the epitope peptide sequence be p = (p1, p2, ..., p n ), where p i This represents the encoding of the i-th amino acid;

[0018] Let the HLA class I molecule sequence be h = (h1, h2, ..., h m ), where h j This represents the encoding of the j-th amino acid;

[0019] Epitope peptide protein language model LM p Mapping the epitope peptide sequence p to a high-dimensional continuous vector representation:

[0020] p = LM p (p)

[0021] HLA Class I molecular protein language model LM h Map the HLA class I molecular sequence h to a high-dimensional continuous vector representation:

[0022] h = LM h (h)

[0023] Wherein: the vectors p and h represent the sequence features of epitope peptides and HLA class I molecules, including their structural information and potential functional features.

[0024] Preferably, the expression for the multi-level feature extraction module of the dynamic selection mechanism LSTM is as follows:

[0025] Let the input amino acid sequence be x = (x1, x2, ..., x...). T ), where x t This represents the input feature vector of the sequence at time step t;

[0026] The hidden state h of LSTM t Defined by the following recursive relation:

[0027] h t =LSTM(x t ,h t-1 )

[0028] Where: LSTM is the computation function of the LSTM unit, used to update and output the hidden state h at the current time step. t Based on the current input x t The hidden state h of the previous time step t-1 ;

[0029] The dynamic selection mechanism uses a linear neural network layer—a fully connected layer and SoftMax normalization—to select the optimal state from the historical hidden states. The optimal hidden state at the current time step:

[0030] The optimal state h among the hidden states h at time t-1 k ,

[0031] h k =Max(Softmax(f(h1,…,h) t-1 ))

[0032] Where: f is used to calculate the historical hidden state h k The function of importance at the current time step t.

[0033] Preferably, the dual multi-instance attention module includes a Seq-Attention module expression:

[0034] For molecular sequence (p) i ,h i Let (p)i, i ,h i If i has n amino acids, and each amino acid vector represents a dimension of m, then Seq-Attention learns the weights α between different amino acids through an attention mechanism. i =(α i1 ,α i2 ,…,α in ):

[0035] α ik =Softmax(f att (p ik ))

[0036] Among them, f att It is a function used to calculate attention weights.

[0037] Preferably, the dual multi-instance attention module further includes a Bag-Attention module expression:

[0038] For a group of peptide-HLA class I molecules (p i ,h il )i, assuming each p i The corresponding number of alleles is l, and Bag-Attention dynamically calculates p for each peptide sequence. i The corresponding multiple alleles h il The contribution weight β of affinity i =(β) i1 ,β i2 ,…,β il ):

[0039] β ik =Softmax(g att (h ik ))

[0040] Among them, g att It is a function used to calculate attention weights.

[0041] Preferably, the expression for the feature fusion and affinity prediction module is:

[0042] Let the feature vector of the peptide-HLA class I molecule sequence after processing by the dual multi-instance attention module be (p i ,h ij ), where p i H represents the i-th peptide sequence molecule. ij This represents the feature vector of the j-th allele corresponding to the i-th peptide sequence molecule;

[0043] Feature fusion:

[0044] The allele features processed by the dual multi-instance attention module are fused using a weighted summation:

[0045]

[0046] Where: β ij The weights of multiple alleles after processing by the dual multi-instance attention module, where l represents the weights of each peptide p. i The corresponding number of multiple alleles;

[0047] Affinity prediction:

[0048] For the fused feature vector h merged Through a series of fully connected layers (FCNs) performing nonlinear transformations and feature compression, the final value is mapped to the actual affinity domain. These fully connected layers are represented as follows:

[0049]

[0050] in: It is the final affinity prediction score. FCN is a fully connected neural network containing several layers, used to convert the feature vector h merged Mapped to the output space.

[0051] Compared with existing technologies, this invention provides a peptide-HLA class I allele affinity prediction system based on protein language models and a dual multi-instance attention mechanism, which has the following beneficial effects:

[0052] 1. This invention utilizes a self-trained epitope peptide protein language model and an HLA class I molecule protein language model to obtain feature encoding vectors for amino acid sequences containing HLA class I molecules and peptides, thereby better characterizing the molecular sequences. For the variable-length dependencies between amino acids, an improved dynamic selection mechanism, LSTM, is used to learn the dependency features between amino acids at different distances in the sequence. Regarding the multi-specificity properties present in mass spectrometry data, a dual multi-instance attention module is employed to learn the importance of alleles at different granularities (amino acid level and allele level), achieving a beneficial effect in improving the prediction performance of HLA antigen presentation. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the system of the present invention;

[0054] Figure 2 This is a schematic diagram of the test results of the F1 index on the monoallelic test dataset of this invention;

[0055] Figure 3This is a schematic diagram of the test results of the F1 metric on the Anthem dataset of this invention;

[0056] Figure 4 This is a schematic diagram of the test results of the F1 value index on the multi-allelic test dataset of this invention;

[0057] Figure 5 This is a schematic diagram of the test results of the MCC index on the multi-allele test dataset of this invention;

[0058] Figure 6 This is a schematic diagram of the test results of the top 30 multi-allelic MCC indicators in this invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Please see Figure 1 A peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism. The prediction system includes: a sequence encoding and distributed representation module for peptides and HLA class I molecules, a multi-level feature extraction module using dynamic selection mechanism LSTM, a dual multi-instance attention module, and a feature fusion and affinity prediction module.

[0061] The sequence encoding and distributed representation module for peptides and HLA class I molecules includes: an epitope peptide protein language model and an HLA class I molecule protein language model;

[0062] The dynamic selection mechanism LSTM multi-level feature extraction module is: adding a dynamic selection mechanism to LSTM;

[0063] The dual multi-instance attention module comprises two sub-modules: Module 1, Seq-Attention, is used to learn the attention weights between different amino acids in a single sequence; Module 2, Bag-Attention, is used to dynamically calculate the relative weight of each allele's contribution to affinity among multiple alleles.

[0064] The feature fusion and affinity prediction module is used to fuse allele features processed by the dual multi-instance attention module and perform nonlinear transformation and feature compression on them.

[0065] Preferably, the sequence encoding and distributed representation module expression of the peptide and HLA class I molecules is as follows:

[0066] Let the epitope peptide sequence be p = (p1, p2, ..., p n ), where p i This represents the encoding of the i-th amino acid;

[0067] Let the HLA class I molecule sequence be h = (h1, h2, ..., h m ), where h j This represents the encoding of the j-th amino acid;

[0068] Epitope peptide protein language model LM p Mapping the epitope peptide sequence p to a high-dimensional continuous vector representation:

[0069] p = LM p (p)

[0070] HLA Class I molecular protein language model LM h Map the HLA class I molecular sequence h to a high-dimensional continuous vector representation:

[0071] h = LM h (h)

[0072] Wherein: the vectors p and h represent the sequence features of epitope peptides and HLA class I molecules, including their structural information and potential functional features.

[0073] The module uses a protein language model to map peptide and HLA class I molecule sequences into high-dimensional continuous vectors p and h. These vectors capture rich feature information of the sequences, preserving their structural and functional characteristics. The module can handle peptide and HLA class I molecule sequences of various lengths and features, making it suitable for diverse bioinformatics data analysis needs. Through sequence encoding and distributed representation, the module can comprehensively capture the sequence features of epitope peptides and HLA class I molecules, including but not limited to amino acid composition, structure, and functional information. This enables the module to provide more accurate and comprehensive feature descriptions when predicting peptide-HLA class I molecule binding affinity.

[0074] Preferably, the expression for the multi-level feature extraction module of the dynamic selection mechanism LSTM is as follows:

[0075] Let the input amino acid sequence be x = (x1, x2, ..., x...). T ), where x t This represents the input feature vector of the sequence at time step t;

[0076] The hidden state h of LSTM t Defined by the following recursive relation:

[0077] h t=LSTM(x t ,h t-1 )

[0078] Where: LSTM is the computation function of the LSTM unit, used to update and output the hidden state h at the current time step. t Based on the current input x t The hidden state h of the previous time step t-1 ;

[0079] The dynamic selection mechanism uses a linear neural network layer—a fully connected layer and SoftMax normalization—to select the optimal state from the historical hidden states. The optimal hidden state at the current time step:

[0080] The optimal state h among the hidden states h at time t-1 k ,

[0081] h k =Max(Softmax(f(h1,…,h) t-1 ))

[0082] Where: f is used to calculate the historical hidden state h k The function of importance at the current time step t.

[0083] The recursive relationship gives the module a significant advantage in processing time-correlated biological sequence data, such as protein or DNA sequence analysis. The module utilizes LSTM units for multi-level feature extraction, with the hidden state h at each time step... t It incorporates the accumulation and integration of past information, which can gradually improve the representational ability of features and help to more accurately capture the complex structure and patterns of sequences. It introduces a dynamic selection mechanism to weight the historical hidden state h through fully connected layers and SoftMax normalization. k This allows us to select the optimal hidden state at the current time step t. This mechanism enables modules to adaptively focus on and integrate information that is particularly important to the current task, improving the model's predictive performance and generalization ability. Due to the clear definition of the mathematical expressions, the modules have good interpretability and flexibility, and are easy to understand and adjust.

[0084] Preferably, the dual multi-instance attention module includes a Seq-Attention module expression:

[0085] For molecular sequence (p) i ,h i Let (p)i, i ,h iIf i has n amino acids, and each amino acid vector represents a dimension of m, then Seq-Attention learns the weights α between different amino acids through an attention mechanism. i =(α i1 ,α i2 ,…,α in ):

[0086] α ik =Softmax(f att (p ik ))

[0087] Among them, f att It is a function used to calculate attention weights.

[0088] Preferably, the dual multi-instance attention module further includes a Bag-Attention module expression:

[0089] For a group of peptide-HLA class I molecules (p i ,h il )i, assuming each p i The corresponding number of alleles is l, and Bag-Attention dynamically calculates p for each peptide sequence. i The corresponding multiple alleles h il The contribution weight β of affinity i =(β) i1 ,β i2 ,…,β il ):

[0090] β ik =Softmax(g att (h ik ))

[0091] Among them, g att It is a function used to calculate attention weights.

[0092] Precise attention mechanism: The Seq-Attention module utilizes the attention mechanism α i =(α i1 ,α i2 ,…,α in Learning the weights between amino acids in different peptide sequences can accurately capture important features within the sequence, helping to distinguish and enhance the contribution of key amino acids to affinity. The Bag-Attention module dynamically calculates the weights of each HLA class I molecule. j The contribution weight β of affinity j =(β) j1 ,β j2 ,…,β jmThis method effectively integrates information from multiple peptide-HLA class I molecule pairs. It considers the complex relationships between each allele and all related peptide sequences, improving the predictive model's adaptability to different pairings. The attention weights in the module are calculated using the Softmax function, offering good interpretability and flexibility. The weight distribution can be adjusted according to actual needs to optimize model performance. This flexibility makes the module widely applicable to different datasets and application scenarios. By combining Seq-Attention and Bag-Attention, the module can learn and integrate feature importance at different granularities, thereby improving the accuracy and stability of predicting peptide-HLA class I molecule affinity.

[0093] Preferably, the expression for the feature fusion and affinity prediction module is:

[0094] Let the feature vector of the peptide-HLA class I molecule sequence after processing by the dual multi-instance attention module be (p i ,h ij ), where p i H represents the i-th peptide sequence molecule. ij This represents the feature vector of the j-th allele corresponding to the i-th peptide sequence molecule;

[0095] Feature fusion:

[0096] The allele features processed by the dual multi-instance attention module are fused using a weighted summation:

[0097]

[0098] Where: β ij The weights of multiple alleles after processing by the dual multi-instance attention module, where l represents the weights of each peptide p. i The corresponding number of multiple alleles;

[0099] Affinity prediction:

[0100] For the fused feature vector h merged Through a series of fully connected layers (FCNs) performing nonlinear transformations and feature compression, the feature is ultimately mapped to the actual affinity value domain. These fully connected layers are represented as follows:

[0101]

[0102] in: It is the final affinity prediction score. FCN is a fully connected neural network containing several layers, used to convert the feature vector h merged Mapped to the output space.

[0103] Figure 2The results show the F1 score on a monoallelic test dataset, which is the evaluation dataset provided by the NetMHCpan online website and involves 15 different monoallelic genes. The results show that the F1 score of the model in this application is slightly higher than that of two widely used software programs, NetMHCpan-4.1 and MHCFlurry-2.0. Figure 2 In another publicly available dataset, Anthem, our model achieved a higher F1 score of 0.763 compared to 0.732-0.760 for peptides of 11-14 amino acids. Figure 3 The mean F1 score for each length is 0.788 vs. 0.729-0.741.

[0104] In the evaluation of multi-allelic data, the model of this application was applied to a test set containing 20,000 MA data points split from the original dataset. The F1 score of the model of this application was significantly higher than that of NetMHCpan-4.1EL for eluting ligands (0.800 vs 0.575, p = 3.992e-8), NetMHCpan-4.1BA for binding affinity (0.800 vs 0.558, p = 1.335e-8), and MHCflurry-2.0 (0.800 vs 0.586, p = 9.107e-8). Figure 4 The Matthews correlation coefficient (MCC), 0.791 vs 0.580-0.603, also showed a similar performance improvement, demonstrating that this metric can provide a reliable measure even in cases of data imbalance. Figure 5 The MCC details for the top 30 alleles in terms of data volume in the MA test set are as follows: Figure 6 As shown.

[0105] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism, characterized in that, The prediction system includes: a sequence encoding and distributed representation module for peptides and HLA class I molecules, a multi-level feature extraction module using dynamic selection mechanism LSTM, a dual multi-instance attention module, and a feature fusion and affinity prediction module. The sequence encoding and distributed representation module for peptides and HLA class I molecules includes: an epitope peptide protein language model and an HLA class I molecule protein language model; The dynamic selection mechanism LSTM multi-level feature extraction module is: adding a dynamic selection mechanism to LSTM; The dual multi-instance attention module comprises two sub-modules: Module 1, Seq-Attention, is used to learn the attention weights between different amino acids in a single sequence; Module 2, Bag-Attention, is used to dynamically calculate the relative weight of each allele's contribution to affinity among multiple alleles. The feature fusion and affinity prediction module is used to fuse allele features processed by the dual multi-instance attention module and perform nonlinear transformation and feature compression on them; The dynamic selection mechanism LSTM multi-level feature extraction module works as follows: from the set of historical hidden states before the current input of the amino acid sequence, the optimal hidden state is selected and connected to the current input. The optimal hidden state in the set of historical hidden states is selected by adding a linear neural network layer and then using SoftMax normalization to obtain the score of each state. The state with the highest score is considered optimal. This dynamic selection mechanism replaces the original LSTM model's method of directly connecting the current input to the hidden state of the previous time step, thereby learning useful information from long distances in the sequence.

2. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 1, characterized in that: The epitope peptide protein language model and HLA class I molecular protein language model are based on ProLLaMa and incorporate the concept of transfer learning. They are specifically optimized for epitope peptide sequence datasets and HLA class I molecular sequence datasets. Through the collected peptide-HLA I binding dataset, the peptide sequences and HLA I molecular sequences are subjected to the Masked Language Modeling (MLM) task, respectively. The sequence features of peptides and HLA I molecules are learned from the data in an unsupervised manner. The epitope peptide protein language model and HLA class I molecular protein language model convert the original amino acid sequences into high-dimensional, continuous vector representations, extract key structural information of the amino acid sequences, and also contain potential functional features.

3. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 1, characterized in that: The dual multi-instance attention module receives peptide sequence feature vectors related to different alleles from the dynamically selected LSTM layer. Through training, the module learns the importance of different features at both the amino acid and allele levels.

4. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 3, characterized in that: The feature fusion and affinity prediction module integrates the allele features processed by the dual multi-instance attention module into a global feature vector. Then, it performs nonlinear transformation and feature compression through a series of fully connected layers. The role of the fully connected layers is to further refine key features. The last fully connected layer generates the binding probability of each HLA class I allele with a given peptide sequence.

5. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 2, characterized in that: The sequence encoding and distributed representation module expression for the peptide and HLA class I molecules is as follows: Let the epitope peptide sequence be ,in Indicates the first The code for one amino acid; Let the HLA class I molecule sequence be ,in Indicates the first The code for one amino acid; Epitope peptide protein language model epitope peptide sequence Mapped to a high-dimensional continuous vector representation: ; HLA Class I molecular protein language model HLA class I molecular sequence Mapped to a high-dimensional continuous vector representation: ; Where: represents a vector and The sequence characteristics of epitope peptides and HLA class I molecules were captured, including their structural information and potential functional features.

6. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 3, characterized in that: The expression for the multi-level feature extraction module of the dynamic selection mechanism LSTM is as follows: Let the input amino acid sequence be ,in Indicates the sequence at time step The input feature vector; Hidden state of LSTM Defined by the following recursive relation: ; in: This is the computation function for the LSTM unit, used to update and output the hidden state at the current time step. Based on the current input The hidden state of the previous time step ; The dynamic selection mechanism uses a linear neural network layer—a fully connected layer and SoftMax normalization—to select the optimal state from the historical hidden states. As the optimal hidden state at the current time step: Hidden state at any time optimal state , ; in: It is used to calculate historical hidden states. At the current time step A function of importance.

7. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 4, characterized in that: The dual multi-instance attention module includes the Seq-Attention module expression: For molecular sequences ,set up have There are amino acids, and each amino acid vector represents a dimension of . Seq-Attention learns the weights between different amino acids through an attention mechanism. : ; in, It is a function used to calculate attention weights.

8. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 7, characterized in that: The dual multi-instance attention module also includes a Bag-Attention module expression: For a group of peptide-HLA class I molecules Assuming each The corresponding number of alleles is Bag-Attention dynamically calculates each peptide sequence Corresponding multiple alleles Weight of contribution to affinity : ; in, It is a function used to calculate attention weights.

9. The peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism according to claim 1, characterized in that: The expression for the feature fusion and affinity prediction module is as follows: Let the feature vector of the peptide-HLA class I molecule sequence after processing by the dual multi-instance attention module be: ,in Indicates the first A peptide sequence molecule, Indicates the first The corresponding peptide sequence molecule of the first Each allele feature vector; Feature fusion: The allele features processed by the dual multi-instance attention module are fused using a weighted summation: ; in: The weights of multiple alleles after processing by a dual multi-instance attention module. Indicates each peptide The corresponding number of multiple alleles; Affinity prediction: For the fused feature vector Through a series of fully connected layers, nonlinear transformations and feature compression are performed, ultimately mapping the result to the actual affinity value domain. These fully connected layers are represented as follows: ; in: It is the final affinity prediction score. It is a fully connected neural network containing several layers, used to process feature vectors Mapped to the output space.

Citation Information

Patent Citations

  • Leukocyte antigen and polypeptide binding affinity prediction method based on deep learning

    CN111951887A

  • Antigen affinity prediction method and system based on deep learning

    CN114649054A