Multi-modal deep learning-based MHC presentation peptide fragment prediction method and system
By constructing a multi-species MHC-peptide binding dataset and a multimodal deep learning model, the problems of low recall and low computational efficiency of MHC-presented peptide prediction in the existing technology are solved, and high-precision and efficient MHC-antigen peptide binding prediction are achieved, which promotes the development of applications such as individualized tumor neoantigen vaccine design.
Patent Information
- Application Number
- CN202510955009.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-11
AI Technical Summary
The existing MHC presenting peptide prediction methods have low recall in recognition of antigenic peptides, lack cross-species data support, inefficient calculation efficiency, imbalanced data labeling, and it is difficult to fully capture the high-order interactive characteristics of MHC and antigenic peptides.
A multi-species MHC-peptide binding data set was constructed, and the core site information was extracted using multi-sequence alignment and frequency-weighted amino acid feature distance algorithm, combined with multi-modal deep learning models, including convolutional neural networks, channel attention mechanisms and cross attention mechanisms, integrating MHC and peptide sequence characteristics for efficient prediction.
It significantly improves the accuracy and efficiency of MHC presentation antigen peptide prediction, covers immunopolypeptideomics data of a variety of disease types and cell lines, improves computational efficiency and prediction performance, and provides efficient and accurate tools for immunotherapy and vaccine design.
Smart Images

Figure CN120452555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the interdisciplinary field of bioinformatics and artificial intelligence, and in particular to a method and system for predicting MHC-presented peptides based on multimodal deep learning. Background Art
[0002] MHC is a core molecule in the immune system and plays an indispensable role in presenting antigenic peptides. MHC molecules can bind to antigenic peptides and present them to T cells, thereby activating specific immune responses. This process is critical for the body's immune surveillance, pathogen clearance, and maintenance of immune homeostasis. MHC is known as human leukocyte antigen (HLA), and its high degree of polymorphism provides the molecular basis for the diversity of immune responses among individuals. Accurately predicting the binding and presentation events of MHC and antigenic peptides is not only of great significance for a deeper understanding of the functional mechanisms of the immune system, but also has far-reaching application value in fields such as personalized tumor immunotherapy, vaccine design, and autoimmune disease research.
[0003] In recent years, with the rapid development of computational biology and artificial intelligence technologies, certain progress has been made in the field of MHC-presented peptide prediction. Current mainstream prediction tools, such as NetMHCpan and MHCflurry, have demonstrated certain predictive capabilities on specific data sets. However, existing technologies still face many challenges. Existing computational methods and prediction models show low recall rates in identifying antigenic peptides. This technical limitation has had a significant impact on research. Considering that antigenic peptides are already scarce in their natural state, this low recall rate problem further exacerbates the difficulty of identifying these key molecules, making them more likely to be systematically ignored during the analysis process. This not only affects the comprehensiveness and accuracy of related research, but may also have potential negative impacts on subsequent application areas such as immunology research and vaccine development.
[0004] In terms of data resources, existing research on predicting MHC-presented antigenic peptides primarily relies on data from specific populations, severely lacking HLA typing and corresponding antigenic peptide presentation data for a broader population. Due to significant differences in HLA distribution across ethnic groups, the sample size of HLA subtypes from a broader population in existing resource databases is insufficient, limiting the accuracy and applicability of prediction models across a wider range of populations. Furthermore, existing research often focuses on data from a single species, lacking systematic cross-species data support, making it difficult to meet the needs of comparative immunology and biomedical translational research.
[0005] At the methodological level, existing technologies have several limitations. First, when processing multi-allelic MHC-peptide data, it is difficult to accurately parse the correspondence between a single peptide and a specific MHC molecule, resulting in a limited amount of training data; second, when characterizing MHC sequence features, existing methods usually only select a limited number of core amino acid sites, which makes it difficult to fully reflect the diversity and functional specificity of MHC structures; third, there is a lack of effective cross-species MHC sequence feature alignment strategies, which makes the model's generalization ability in cross-species predictions significantly insufficient. In addition, traditional methods mostly rely on manual feature engineering and lack automated feature extraction and alignment strategies, making it difficult to effectively capture the high-order interaction features between MHC and antigen peptides.
[0006] In terms of computational efficiency, with the rapid growth of immunogenomics data, the demand for high-throughput MHC-presented antigen peptide prediction is increasing. However, existing systems suffer from widespread computational inefficiency when processing large amounts of data. They lack efficient distributed training mechanisms and dynamic parameter optimization strategies, making it difficult to meet the real-time requirements of clinical translational applications. This is particularly true in time-sensitive applications such as the design of personalized tumor neoantigen vaccines, where computational efficiency has become a key bottleneck for technology implementation.
[0007] Furthermore, existing prediction systems face the problem of data label imbalance during data preprocessing—an imbalance in the ratio of bound to unbound samples, which can easily lead to prediction bias. In terms of model architecture design, most methods employ a single network structure, failing to fully integrate multimodal information such as sequence information, structural features, and mass spectrometry validation data, limiting further improvements in prediction performance. Summary of the Invention
[0008] Based on this, the purpose of the present invention is to provide a method and system for predicting MHC-presented peptides based on multimodal deep learning, so as to at least address the deficiencies in the above-mentioned technologies.
[0009] The present invention proposes a method for predicting MHC-presented peptides based on multimodal deep learning, comprising: Construct a multi-species MHC-peptide binding dataset, including MHC molecule sequence information and corresponding peptide binding data of human and non-human species; performing single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence; Perform multiple sequence alignment and feature alignment on cross-species MHC sequences, and use a frequency-weighted amino acid feature distance algorithm to extract core site information characterizing MHC molecular polymorphism; A multimodal deep learning model was constructed to integrate MHC sequence features and peptide sequence features to predict the binding ability of MHC molecules and peptides.
[0010] Furthermore, the steps for constructing a multi-species MHC-peptide binding dataset include: Collect clinical tumor samples and enrich and identify immune peptides from the clinical tumor samples using a magnetic bead-based automated immune peptide mass spectrometry platform; Perform whole-exome sequencing and HLA typing on the same sample to obtain MHC genotype and its corresponding antigen peptide pairing information; Integrate immunogenicity data from public databases, including information on the pairing of MHC class I molecules and antigenic peptides.
[0011] Furthermore, the steps of performing single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence include: Construct an intersection judgment mechanism between peptide segments and their related MHC molecule sets; Based on the intersection judgment mechanism, unique MHC-peptide correspondences are screened to achieve single genotype mapping extraction of peptides.
[0012] Furthermore, the steps of performing multiple sequence alignment and feature alignment on the cross-species MHC sequences and extracting core site information characterizing MHC molecular polymorphism using a frequency-weighted amino acid feature distance algorithm include: Obtaining multiple species MHC sequences from multiple public databases, and uniformly aligning the multiple species MHC sequences using a multiple sequence alignment tool; The frequency-weighted amino acid feature distance algorithm was used to calculate the polymorphism index of each amino acid site in the aligned multi-species MHC sequences, and the sites with the longest amino acid polymorphism distance were selected as the core site information characterizing the polymorphism of MHC molecules.
[0013] Furthermore, the core site information contains 40 sites, which are embedded and encoded using the BLOSUM50 matrix to convert the biological sequence information into feature expressions processed by the machine learning model.
[0014] Furthermore, the multimodal deep learning model includes a convolutional neural network module, a channel attention mechanism module, and a cross attention mechanism module; Among them, the convolutional neural network module is used to extract local features of MHC molecules and peptide sequences, the channel attention mechanism module is used to enhance the weights of important feature channels, and the cross-attention mechanism module is used to capture the sequence interaction patterns between MHC molecules and antigen peptides.
[0015] Furthermore, the method further comprises: Balance the sampling of bound samples and unbound samples to control the ratio of positive and negative samples; Affinity values were normalized according to logarithmic transformation to improve numerical stability and model convergence speed.
[0016] The present invention also proposes a MHC-presented peptide prediction system based on multimodal deep learning, comprising: A dataset construction module is used to construct a multi-species MHC-peptide binding dataset, including MHC molecule sequence information and corresponding peptide binding data of human and non-human species; a data parsing module, configured to perform single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence; The data comparison module is used to perform multiple sequence alignment and feature alignment of MHC sequences across species, and uses a frequency-weighted amino acid feature distance algorithm to extract core site information that characterizes MHC molecular polymorphism; The data prediction module is used to build a multimodal deep learning model, integrate MHC sequence features and peptide sequence features, and predict the binding ability of MHC molecules and peptides.
[0017] Furthermore, the dataset construction module is specifically used to: Collect clinical tumor samples and enrich and identify immune peptides from the clinical tumor samples using a magnetic bead-based automated immune peptide mass spectrometry platform; Perform whole-exome sequencing and HLA typing on the same sample to obtain MHC genotype and its corresponding antigen peptide pairing information; Integrate immunogenicity data from public databases, including information on the pairing of MHC class I molecules and antigenic peptides.
[0018] Furthermore, the data analysis module is specifically used to: Construct an intersection judgment mechanism between peptide segments and their related MHC molecule sets; Based on the intersection judgment mechanism, unique MHC-peptide correspondences are screened to achieve single genotype mapping extraction of peptides.
[0019] Furthermore, the data comparison module is specifically used to: Obtaining multiple species MHC sequences from multiple public databases, and uniformly aligning the multiple species MHC sequences using a multiple sequence alignment tool; The frequency-weighted amino acid feature distance algorithm was used to calculate the polymorphism index of each amino acid site in the aligned multi-species MHC sequences, and the sites with the longest amino acid polymorphism distance were selected as the core site information characterizing the polymorphism of MHC molecules.
[0020] Furthermore, the core site information contains 40 sites, which are embedded and encoded using the BLOSUM50 matrix to convert the biological sequence information into feature expressions processed by the machine learning model.
[0021] Furthermore, the multimodal deep learning model includes a convolutional neural network module, a channel attention mechanism module, and a cross attention mechanism module; Among them, the convolutional neural network module is used to extract local features of MHC molecules and peptide sequences, the channel attention mechanism module is used to enhance the weights of important feature channels, and the cross-attention mechanism module is used to capture the sequence interaction patterns between MHC molecules and antigen peptides.
[0022] Furthermore, the system further comprises: Balanced sampling module, used to balance the sampling of bound samples and unbound samples and control the ratio of positive and negative samples; A data optimization module is used to normalize affinity values according to logarithmic transformation to improve numerical stability and model convergence speed.
[0023] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned MHC-presented peptide prediction method based on multimodal deep learning.
[0024] The present invention also proposes a computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned MHC-presented peptide prediction method based on multimodal deep learning is implemented.
[0025] The MHC-presented peptide prediction method and system based on multimodal deep learning in the present invention significantly improves the accuracy and efficiency of MHC-presented antigen peptide prediction by integrating multi-source data, innovative algorithm design and efficient computing framework; at the data level, it introduces large-scale data of HLA-presented antigen peptides from different populations for the first time, covering immune peptidomics data of various disease types and cell lines, and integrating comprehensive cross-species data resources; at the algorithm level, it develops a multi-allele MHC-peptide single genotype analysis method, a frequency-weighted amino acid feature distance algorithm and a dual-channel adaptive convolutional network architecture combined with channel attention and cross-attention mechanisms to more comprehensively and effectively capture the high-order interaction features between MHC and antigen peptides; it adopts a collaborative training framework and dynamic optimization strategy to significantly improve computing efficiency and prediction performance, providing efficient and accurate prediction tools for immunotherapy, vaccine design and autoimmune disease research, and promoting further development in the fields of precision medicine and biotechnology. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a flowchart of the method for predicting MHC-presented peptides based on multimodal deep learning in the first embodiment of the present invention; Figure 2 The overall architecture of the multimodal deep learning model in the first embodiment of the present invention; Figure 3 This is a structural block diagram of the MHC-presented peptide prediction system based on multimodal deep learning in the second embodiment of the present invention; Figure 4 FIG. 4 is a structural block diagram of a computer in a third embodiment of the present invention.
[0027] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0028] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.
[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0030] Current methods for predicting the presentation of antigenic peptides by MHC molecules face several challenges that need to be addressed, including: (1) lack of tissue source data representing a wider population in model training samples, resulting in insufficient population diversity; (2) traditional MHC sequence encoding methods fail to effectively characterize its polymorphic key residues, affecting the accuracy of modeling; (3) most methods fail to fully extract high-order interaction information between MHC and antigenic peptides, resulting in limited predictive capabilities; (4) low training efficiency, especially when faced with massive samples and complex feature structures, making it difficult to balance computing resources and model performance; and (5) severe imbalance in data labels, affecting the generalization ability of the model.
[0031] To this end, this application has achieved high-precision prediction of the binding ability of MHC class I molecules with antigenic peptides by constructing a high-quality cross-species immunogenicity dataset, developing a method for analyzing multi-allelic data, designing an MHC encoding method for expressing polymorphisms, and introducing a multimodal neural network structure. It has significant scientific innovation and application promotion value.
[0032] Example 1 See also Figure 1 , which shows a method for predicting MHC-presented peptides based on multimodal deep learning in a first embodiment of the present invention, and the method specifically includes steps S101 to S106: S101: Construct a multi-species MHC-peptide binding dataset, including MHC molecule sequence information and corresponding peptide binding data for human and non-human species; Furthermore, the step S101 specifically includes steps S1011 to S1013: S1011, collecting clinical tumor samples, and enriching and identifying immune peptides from the clinical tumor samples using a magnetic bead-based automated immune peptide mass spectrometry platform; S1012, whole exome sequencing and HLA typing were performed on the same sample to obtain MHC genotype and its corresponding antigen peptide pairing information; S1013, integrates immunogenicity data from public databases, including the pairing information of MHC class I molecules and antigenic peptides.
[0033] In practice, this example first constructed a high-quality tumor sample cohort covering multiple clinical cases. All patient samples passed ethics approval (Approval No. 2022-CDYFYYLK-06-012) from the First Affiliated Hospital of Nanchang University, and informed consent was obtained through signed informed consent forms. Colorectal and lung cancer tissues (including both cancerous and adjacent tissues) were collected from the patients. Immunopeptide mass spectrometry (MPMS) was then performed on an automated immunopeptide mass spectrometry platform integrated with magnetic beads and the KingFisher Apex system. By further optimizing the sample lysis volume and immunoprecipitation binding time, immunopeptides from the clinical samples were enriched and identified, providing mass spectrometry (MS) ligand elution data.
[0034] At the same time, whole exome sequencing was performed on the same sample, and HLA typing was performed using dedicated software, providing real human MHC genotypes and their corresponding antigen peptide pairing information for subsequent model training.
[0035] Specifically, to expand the data size and enhance the generalization capabilities of model training, this example further integrates immunogenicity data from the IEDB database and published literature, including over 800,000 MHC class I molecule-antigen peptide pairings, covering 10 species, including human, mouse, porcine, and canine, and 290 MHC alleles. Binding affinity data (BA) is presented as IC50 (in nanomolar units). A confidence level of 0.75 was set to filter MS ligand elution data, and an IC50 value of 100 nanomolar was assigned. Mass spectrometry ligand elution data was then filtered based on reliability criteria and standardized to supplement samples with missing labels.
[0036] S102, performing single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence; Furthermore, the step S102 specifically includes steps S1021 and S1022: S1021, constructing a mechanism for determining the intersection of peptide segments and their associated MHC molecule sets; S1022: Based on the intersection judgment mechanism, unique MHC-peptide correspondences are screened to achieve single genotype mapping extraction of peptides.
[0037] In specific implementation, to address the mapping uncertainty problem caused by the "multi-allele-MHC paired single peptide" data structure commonly found in public databases, this embodiment proposes a multi-allele MHC-peptide single genotype parsing method. By constructing an intersection judgment mechanism between peptides and their related MHC molecule sets, it automatically screens out unique MHC-peptide correspondences, realizes the single genotype mapping extraction of immune peptides, and significantly improves the reliability of training data and the accuracy of model training.
[0038] Specifically, for the same peptide, different patient groups may have different MHC molecule sets, and the intersection of their MHC molecule sets may contain a unique MHC molecule that can present the peptide. A peptide list is generated based on the MS ligand elution data. The peptides and all their MHC genotypes are read cyclically. The intersection of all MHC molecule sets corresponding to each peptide is taken. If the final intersection contains only one MHC genotype, it means that the peptide can be presented by at least this MHC molecule. This establishes a one-to-one correspondence between peptides and MHC genotypes, successfully achieving multi-allele MHC-peptide single genotype resolution.
[0039] S103, multiple sequence alignment and feature alignment of cross-species MHC sequences, and the frequency-weighted amino acid feature distance algorithm are used to extract core site information characterizing MHC molecular polymorphism; Furthermore, the step S103 specifically includes steps S1031 and S1032: S1031, obtaining MHC sequences of multiple species from multiple public databases, and uniformly aligning the MHC sequences of the multiple species using a multiple sequence alignment tool; S1032: A frequency-weighted amino acid feature distance algorithm is used to calculate the polymorphism index of the amino acid at each site in the aligned multi-species MHC sequences, and multiple sites with the longest amino acid polymorphism distance ranking are selected as core site information characterizing the polymorphism of the MHC molecule.
[0040] In practice, to efficiently encode MHC class I molecule sequences across species, this example incorporates multispecies MHC sequences obtained from multiple databases (such as UniProt and IMGT) and aligns them using a multiple sequence alignment tool. By calculating the amino acid polymorphism index for each site, the 40 sites with the longest amino acid polymorphism distances are selected as core "pseudo-sequences" representing MHC molecule polymorphisms. This pseudo-sequence is embedded and encoded using the BLOSUM50 matrix, converting biological sequence information into feature representations that can be processed by machine learning models, thereby achieving more efficient model input representation.
[0041] Specifically, in order to systematically evaluate the polymorphic characteristics of the amino acid sequences of MHC class I molecules in multiple species and provide structurally stable and representative MHC characteristic sites for subsequent neural network model training, a frequency-weighted amino acid characteristic distance algorithm was constructed in this example. The specific steps are as follows: 1. Data preparation and preprocessing The amino acid sequences of MHC class I molecules from multiple species were obtained from the UniProt and IMGT databases and multiple sequence alignments were performed using ClustalOmega (version 1.2.4). All sequences were normalized to a uniform length of 622 bits by inserting a gap character “-” to form a set of aligned sequences. , where n is the total number of MHC class I molecule sequences.
[0042] Define the position index k∈{1,2,...,622} to represent the amino acid position subscript in the aligned sequence. For each position k, extract the amino acid composition of all sequences at that position, recorded as vector a k =(a 1k ,a 2k ,...,a nk ), where a ik represents the amino acid at position k of the i-th sequence.
[0043] 2. Amino acid frequency statistics For each position k, count the occurrence frequencies of all amino acids:
[0044] Among them, AminoAcidSet is a set of 20 standard amino acids plus the special character "*" (indicating missing or undetermined). Is the indicator function. Construct the frequency dictionary F k ={a:f k (a)|a∈AminoAcidSet} and sort them in descending order of frequency.
[0045] 3. Conservative distance calculation 3.1 Reference amino acid selection For position k, the amino acid with the highest frequency is selected as the reference amino acid :
[0046] Where, For amino acids, Reference amino acids. It is a function that finds the input value that makes the function reach its maximum value.
[0047] 3.2 Amino acid substitution treatment Since the sequences aligned by Clustal Omega contain special characters "?" and "X", the special characters are standardized: the "?" character is replaced with "*" (indicating missing); the "X" character is replaced with the reference amino acid. .
[0048] 3.3 BLOSUM50 matrix distance calculation definition is the encoding vector of amino acid a in the BLOSUM50 matrix, Reference amino acid a ref The encoding vector in the BLOSUM50 matrix, is the number of amino acid variants at position k. Define n as the total number of MHC class I sequences. For each amino acid a at position k, calculate its weighted Euclidean distance to the reference amino acid using a frequency-weighted amino acid signature distance algorithm. The result is the amino acid polymorphism distance:
[0049] The frequency-weighted amino acid feature distance algorithm was used to calculate the polymorphism of the amino acid at each position in the aligned sequence. The longer the amino acid polymorphism distance, the higher the polymorphism of the amino acid at that position. The number of polymorphic amino acids required for modeling was weighed against the amount of calculation. Finally, the 40 amino acid positions with the highest polymorphism ranking were selected as the MHC core "pseudo-sequences", and BLOSUM50 was used to encode these 40 amino acid pseudo-sequences.
[0050] For antigen peptide sequences, the antigen peptide sequences with a length of 8 to 14 bits were screened, and the sequences with less than 14 bits were padded to 14 bits using “*”, and finally these sequences were encoded using BLOSUM50.
[0051] Each MHC-presented peptide data point is labeled with an IC50 value. IC50 values less than 5000 nM are defined as MHC-bound and presented peptide data (called positive data), while IC50 values greater than or equal to 5000 nM are defined as MHC-unbound peptide data (called negative data). Because the number of positive data points far exceeds the number of negative data points, direct model training can easily lead to bias. To avoid this bias, positive data points are randomly selected to equal the number of negative data points, and the final IC50 value is calculated as the log10.
[0052] S104, construct a multimodal deep learning model, integrate MHC sequence features and peptide sequence features, and predict the binding ability of MHC molecules and peptides.
[0053] Furthermore, the multimodal deep learning model includes a convolutional neural network module, a channel attention mechanism module, and a cross attention mechanism module; Among them, the convolutional neural network module is used to extract local features of MHC molecules and peptide sequences, the channel attention mechanism module is used to enhance the weights of important feature channels, and the cross-attention mechanism module is used to capture the sequence interaction patterns between MHC molecules and antigen peptides.
[0054] In specific implementation, the shortcomings of current MHC-presented antigen peptide prediction methods, such as the lack of samples from a wider population, the difficulty of core amino acids (or pseudo sequences) in characterizing MHC polymorphisms, the inability to effectively extract high-order interaction features between MHC and antigen peptides, low training efficiency, and data label imbalance, are addressed. This embodiment proposes a cross-species MHC-presented antigen peptide prediction method based on multimodal deep learning. This method innovatively combines convolutional neural networks (CNNs), squeeze-and-excitation (SE) and cross-attention mechanisms to accurately capture the complex interaction patterns between MHC sequences and antigen peptide sequences, thereby improving the accuracy of binding affinity prediction. This model not only improves the accuracy of antigen peptide binding ability prediction, but also exhibits good generalization performance across multiple species. The specific steps are as follows: 1. Dataset construction and preprocessing The data set used in this embodiment includes MHC class I molecule sequences, corresponding antigen peptide sequences and their experimentally determined binding affinity values, which are quantified in the form of IC50. In order to facilitate the processing of the neural network model, these amino acid sequence data need to be encoded first. In this embodiment, the BLOSUM50 substitution matrix is adopted to encode the MHC and antigen peptide sequences. The BLOSUM50 matrix is a scoring matrix widely used in bioinformatics. It is based on the statistical analysis of protein sequence alignment and converts each amino acid into a 24-dimensional feature vector, thereby capturing the evolutionary relationship and biochemical similarity between amino acids. This encoding method can convert sequence information into a numerical form that can be understood by the model.
[0055] Weighing the number of polymorphic amino acids required for neural network modeling and the amount of computational effort, we ultimately selected the 40 amino acid positions with the highest polymorphism ranking as the core pseudosequence of the MHC molecule to represent the polymorphic sites of the MHC molecule. BLOSUM50 was then used to encode these 40 amino acid pseudosequences. In this example, the fixed length of the MHC sequence was set to L. MHC =40, the fixed length of the antigen peptide sequence is L Antigen = 14. After BLOSUM50 encoding, each MHC sequence is represented as a shape of (L MHC ,24) matrix M∈R40×24, and each antigen peptide sequence is represented as a matrix with a shape of (L Antigen ,24) matrix A∈R14×24. The experimentally measured binding affinity value y is the base-10 logarithm of the IC50 value, i.e., y=log10(IC50). This logarithmic transformation is a common technique in the field to address problems with a wide range of IC50 values and make their distribution closer to a normal distribution, which facilitates model learning.
[0056] 2. Overall structure of the neural network model At the core of this example is a neural network model called MHCAntigenBindingModel, which aims to predict the binding affinity between MHC molecules and antigen peptides based on their encoded sequences. This model consists of several key modules that work together to extract relevant features from the sequences and model their interactions.
[0057] This embodiment adopts Figure 2The overall architecture shown includes a data preprocessing module (input layer), a feature extraction module (convolutional module and channel attention), a multimodal feature fusion module (cross attention and feature concatenation), and an affinity prediction module (fully connected layer and output layer). The module takes MHC sequences and antigen peptide sequences as input, and after feature encoding and deep neural network processing, it outputs binding affinity predictions (expressed as log-transformed IC50 values). In the data preprocessing module, the neural network includes an MHC input layer and a peptide input layer (i.e., the input layer). The core sequence length of the MHC input layer is 40, and the input features are 24 channels. The sequence length of the peptide input layer is padded to 14, and the input features are 24 channels. In the feature extraction module, the MHC convolution module and the peptide convolution module are both composed of a convolution layer, a batch normalization layer, and a ReLU activation function, and global average pooling is used. Both the MHC channel attention and the peptide channel attention are channel attention structures, including a linear bottleneck layer, a ReLU activation function, a linear bottleneck layer, a Sigmoid activation function, and channel reweighting. In the multimodal feature fusion module, the MHC feature is used as the query in the cross-attention module, and the antigen feature is used as the key and value. The feature concatenation layer concatenates the MHC feature with the peptide feature. In the affinity prediction module, the fully connected layer consists of three layers, and the first two layers both contain ReLU activation functions and dropout operations. Finally, the log10 (IC50) value is output through the output layer.
[0058] 2.1 Convolutional Feature Extraction Module This example uses a two-stream convolutional neural network to process MHC and antigen peptide sequence information respectively. The convolution operation can effectively capture local sequence patterns, such as binding site features and key residue interactions. (in, is the mathematical symbol for the set of real numbers, is the batch size, is the feature dimension, is the sequence length), the convolution operation is defined as:
[0059] in, and represents the convolution operation, is the transposed feature matrix, W is the convolution kernel, b is the bias term, is the ReLU activation function. Specifically, this embodiment adopts a two-layer convolution structure:
[0060]
[0061] Among them, BN represents batch normalization operation, which can accelerate network convergence and provide regularization effect. is the rectified linear unit activation function, and They are convolution kernel 1 and convolution kernel 2 respectively. and are bias term 1 and bias term 2 respectively, and They are convolution layer 1 and convolution layer 2. For MHC sequences and antigen peptide sequences, convolutional networks with the same structure but no shared parameters are used for processing respectively to fully capture their respective sequence features.
[0062] 2.2 Channel Attention Mechanism To enhance the expressive power of feature channels, this embodiment introduces a channel attention mechanism (Squeeze-and-Excitation, SE) after convolutional feature extraction. This mechanism adaptively adjusts the importance weights of different feature channels, highlighting key features related to binding affinity and suppressing redundant information. The calculation process of the SE Block is as follows: 1) Global average pooling (Squeeze operation):
[0063] in, is the global average pooling result, is the spatial size of the input features, It is the index variable for traversing spatial positions.
[0064] 2) Channel weight calculation (Excitation operation):
[0065] in, is the channel attention weight vector, is the Sigmoid activation function, is the ReLU activation function, and is the weight matrix of the fully connected layer, is the mathematical symbol for the set of real numbers, is the feature dimension, r is the dimensionality reduction ratio (4 in this embodiment).
[0066] 3) Feature recalibration:
[0067] in, is the feature vector recalibrated by the channel attention mechanism, Represents element-wise multiplication along the channel dimension.
[0068] The channel attention mechanism enables the model to identify and enhance feature channels that contribute significantly to prediction results, such as key sites in the MHC binding groove and anchor residues in antigenic peptides. This mechanism significantly improves the model's ability to extract key biological information.
[0069] 2.3 Cross-Attention Fusion Module One of the core features of this implementation is the introduction of a cross-attention mechanism to model the interaction between MHC and antigen peptides. Unlike traditional feature splicing methods, cross-attention can dynamically capture the positional correlation and complementary features between the two molecules, simulating the stereoselectivity of biomolecular binding.
[0070] Specifically, let the MHC characteristics be , the antigenic peptide features are ,in, is the mathematical symbol for the set of real numbers, is the batch size, C is the feature dimension, is the spatial dimension of MHC, is the spatial dimension of the antigen peptide segment, and the cross attention calculation process is as follows: 1) Query-key-value transformation:
[0071]
[0072]
[0073] in, is the query vector matrix in the attention mechanism, is the key vector matrix in the attention mechanism, is the value vector matrix in the attention mechanism. 、 and is the weight matrix of the 1×1 convolution operation. Specifically, is the convolution weight matrix used to generate the query vector matrix, is the convolution weight matrix used to generate the key vector matrix, is the convolution weight matrix used to generate the matrix of value vectors.
[0074] 2) Attention weight calculation:
[0075] in, is the attention weight matrix used to represent the degree of attention of each MHC position to all antigen positions, is the scaling factor for queries and keys, is the transpose operation, is the transpose of the query vector matrix, is the normalized exponential function.
[0076] 3) Attention output calculation:
[0077] in, is the output feature after attention mechanism fusion, The transpose of the matrix of value vectors.
[0078] The cross-attention mechanism enables the model to adaptively learn the interaction strength between the MHC binding groove and different positions of the antigen peptide, thereby accurately capturing the correspondence between key binding sites and improving prediction accuracy.
[0079] 2.4 Multilayer Perceptron Prediction Module After feature fusion, this embodiment uses a multi-layer perceptron to perform the final affinity prediction. The specific structure is as follows: 1) Feature flattening and splicing:
[0080] in, is the feature representation after final concatenation, represents the output of the feature tensor from MHC after the "flattening" operation, represents the output of the feature tensor from the antigen peptide after the "flattening" operation, is the batch size, is the feature dimension, is the spatial dimension of MHC, is the spatial dimension of the antigen peptide, A mathematical symbol representing the set of real numbers.
[0081] 2) Multi-layer feedforward network:
[0082]
[0083]
[0084]
[0085]
[0086] in, 、 and is the weight matrix of the fully connected layer; 、 and is the corresponding bias term. is the activation function of the corrected linear unit. Dropout is a regularization technique used in neural networks. The Dropout operation is based on the probability p Randomly set some neuron outputs to zero to effectively prevent overfitting. is the output of the first hidden layer. is the output of the second hidden layer. is the output after the first layer of Dropout. is the output after the second layer Dropout. is the predicted output of the model.
[0087] 3 Model Training and Optimization Strategy 3.1 Loss Function and Optimization Algorithm This embodiment uses mean square error (MSE) as the loss function to measure the difference between the predicted value and the true value.
[0088] The optimization algorithm uses AdamW, which implements a more reasonable weight decay strategy compared to the traditional Adam optimizer.
[0089] 3.2 Learning Rate Scheduling Strategy To improve training efficiency and prevent local optimality, this embodiment adopts the ReduceLROnPlateau learning rate scheduling strategy to adaptively adjust the learning rate according to the verification loss.
[0090] This strategy reduces the learning rate to half of its original value when the validation loss does not improve for five consecutive cycles, effectively promoting the model to converge to a better solution.
[0091] 3.3 Early Stopping Strategy To avoid overfitting and improve computational efficiency, the present invention implements an improved early stopping strategy. The early stopping judgment function is defined as:
[0092] in, is the validation loss of the t-th cycle, is the historical best validation loss, is the minimum improvement threshold (0 in this embodiment), p is the patience parameter (20 in this example). When the function returns True, training stops and resumes to the best historical model.
[0093] Furthermore, this embodiment proposes an efficient distributed parallel hyperparameter search strategy that fully utilizes multi-GPU computing resources. The specific implementation is as follows: (1) Hyperparameter space definition: predefine the hyperparameters to be optimized and their candidate value ranges; (2) Task allocation: randomly generate hyperparameter combinations and allocate them to multiple GPUs; (3) Process isolation: Use the spawn method to create a process to ensure GPU memory isolation; (4) Parallel evaluation: Each process independently performs K-fold cross validation to evaluate the performance of the corresponding hyperparameters; (5) Result summary: Collect all evaluation results and select the hyperparameter combination with the lowest validation loss.
[0094] This strategy significantly improves the efficiency of hyperparameter search and enables efficient evaluation of 400 sets of hyperparameter configurations.
[0095] 4 Cross-validation and model evaluation 4.1 K-fold Cross-Validation This example uses 5-fold cross validation to evaluate the generalization performance of the model. D Divide into 5 non-overlapping subsets , each time using 4 subsets as training sets and the remaining 1 as test set. The final model performance is based on the comprehensive evaluation of all test folds.
[0096] For each test fold, the training set and validation set are further divided (with a ratio of 9:1) for model selection and early stopping judgment:
[0097]
[0098]
[0099] in For the k Folded training data, For test data, A function to split the entire dataset into training and test sets.
[0100] 4.2 Performance Evaluation Metrics To comprehensively evaluate the model performance, this embodiment also uses the following multiple evaluation indicators: Root mean square error (RMSE), coefficient of determination (R 2 ), mean absolute error (MAE), mean absolute error (MAE), and binary classification accuracy (ACC). Except for the binary classification accuracy, the other indicators are calculated based on the recognized classic formulas.
[0101] Among them, the two-class accuracy formula is:
[0102] in, is the indicator function, is the number of samples, To distinguish the threshold value of positive data from negative data, is the index of the sample, For the The predicted value of the sample.
[0103] In some optional embodiments, the method further includes steps S1021-S1022: S1021, balance the sampling of bound samples and unbound samples to control the ratio of positive and negative samples; S1022, normalizing the affinity values according to the logarithmic transformation to improve numerical stability and model convergence speed.
[0104] In practice, to mitigate the model bias caused by label imbalance in previous algorithm models, this example proposes a balanced sampling strategy of bound and unbound samples. By controlling the ratio of positive and negative samples, the model can balance the learning needs of different data types during training. Affinity values are normalized using a logarithmic transformation to improve numerical stability and model convergence speed.
[0105] Furthermore, to demonstrate the technical advantages of this embodiment, this embodiment compares the performance of previous algorithms with this algorithm on evaluation metrics such as accuracy, precision, recall, F1 score, area under the ROC curve, and area under the precision-recall curve. All metrics are calculated using recognized classical formulas. Considering the limitations of existing algorithms in situations where data categories are unevenly distributed, all algorithm evaluations in this experiment were performed on a class-balanced dataset to ensure the fairness and reliability of the evaluation results.
[0106] This example comprehensively evaluates the current mainstream MHC-peptide binding prediction algorithms, including bigMHC, MHCflurry2.0, MHCnuggets, MixMHCpred2.2, PRIME2.0, netMHCpan4.1, and the new method proposed in this study. The performance of each algorithm was objectively compared through comprehensive analysis of multiple evaluation metrics.
[0107] 1. Accuracy analysis Accuracy is a fundamental metric for measuring the accuracy of a model's overall predictions, representing the proportion of correct predictions among all predictions. As shown in Table 1, the method proposed in this example achieved an accuracy of 86.01%, significantly higher than the other compared algorithms. MHCflurry2.0 ranked second with an accuracy of 79.89%, while MHCnuggets performed the worst with an accuracy of only 49.48%. This result demonstrates that the method proposed in this example has a clear advantage in overall predictive power.
[0108] Table 1. Performance comparison of different prediction algorithms for MHC-presented peptides
[0109] 2. Accuracy analysis Precision reflects the proportion of true positive samples among samples predicted as positive by the model and is an important metric for evaluating the reliability of model predictions. In this precision evaluation, PRIME2.0, bigmhc, and netMHCpan4.1 performed exceptionally well, reaching 98.28%, 98.19%, and 97.9%, respectively. It is worth noting that the method in this example achieved a relatively low precision of 84.37%, due to the trade-offs made in pursuing high recall.
[0110] 3. Recall analysis Recall measures a model's ability to identify true positive samples and is particularly important for biomedical applications requiring high sensitivity. This metric is also the most relevant to practical needs, as in fields such as vaccine development, the primary goal is to identify peptides that are truly presented (i.e., positive samples). During vaccine design, it is crucial to identify as comprehensively as possible peptides that can be presented to T cells by MHC molecules to stimulate an effective immune response. Missing potentially immunogenic peptides (false negatives) can reduce vaccine efficacy. A high recall rate means more potential therapeutic targets can be captured. During the experimental validation phase, predicted positive peptides are typically further screened and verified. In this metric, the method of this example significantly outperforms other algorithms with a recall rate of 91.75%, far exceeding the second-place algorithm, MHCflurry2.0, at 68.12%. PRIME2.0 and MHCnuggets achieved recall rates of only 17.92% and 18.19%, respectively, demonstrating that these algorithms have significant shortcomings in identifying true positive samples.
[0111] 4. F1 score evaluation The F1 score, as the harmonic mean of precision and recall, can comprehensively reflect the overall performance of the model and is particularly suitable for class-imbalanced datasets. The method of this embodiment achieved an F1 score of 0.88, significantly higher than other algorithms. MHCflurry2.0 ranked second with an F1 score of 0.79, while MHCnuggets and PRIME2.0 had F1 scores of only 0.29 and 0.30, respectively, indicating that these algorithms have significant problems in balancing precision and recall.
[0112] 5. Area under the ROC curve (AUROC) analysis AUROC is an important metric for evaluating the discriminative ability of binary classification models, reflecting the model's performance at different thresholds. The method in this example achieved an AUROC of 0.94, outperforming all compared algorithms. netMHCpan4.1, bigmhc, and MHCflurry2.0 all achieved AUROCs of 0.87 or 0.88, respectively, demonstrating relatively good performance. MHCnuggets achieved an AUROC of only 0.49, approaching random guessing.
[0113] 6. Area under the precision-recall curve (AUPRC) analysis In class-imbalanced datasets, the AUPRC better reflects the actual performance of a model than the AUROC. The method in this example achieved an AUPRC of 0.95, slightly higher than the 0.92 achieved by MHCflurry2.0 and netMHCpan4.1. All mainstream algorithms performed relatively similarly on this metric, with the exception of MHCnuggets, which achieved an AUPRC of only 0.57, significantly lower than the other algorithms.
[0114] A comprehensive analysis of the six key metrics above demonstrates that the proposed method in this example outperforms existing algorithms across all five metrics: precision, recall, F1 score, AUROC, and AUPRC, demonstrating excellent overall performance and balance. The most notable advantage of the method in this example lies in its recall (91.75%), which is significantly higher than other algorithms. This is of significant significance for MHC-peptide binding prediction applications requiring high sensitivity. Although its precision (84.37%) is lower than some of the compared algorithms, its significantly superior F1 score (0.88) demonstrates that it achieves a superior balance between precision and recall. MHCflurry2.0 performs well across multiple metrics and is the closest existing algorithm to this method. MHCnuggets performs the worst across all evaluated metrics, particularly with an AUROC close to random guessing, making its use in practical applications not recommended.
[0115] Therefore, the MHC-peptide binding prediction method proposed in this example significantly outperforms existing algorithms in multiple key performance indicators. In particular, it significantly improves recall while maintaining high precision, achieving a better overall performance balance. This result demonstrates that the method proposed in this example has significant application value and promotion potential in the field of MHC-peptide binding prediction.
[0116] In summary, the MHC-presented peptide prediction method based on multimodal deep learning in the above embodiments of the present invention significantly improves the accuracy and efficiency of MHC-presented antigen peptide prediction by integrating multi-source data, innovative algorithm design and efficient computing framework; at the data level, it introduces large-scale data of HLA-presented antigen peptides from different populations for the first time, covering immune peptidomics data of various disease types and cell lines, and integrating comprehensive cross-species data resources; at the algorithm level, it develops a multi-allele MHC-peptide single genotype analysis method, a frequency-weighted amino acid feature distance algorithm and a dual-channel adaptive convolutional network architecture combined with channel attention and cross-attention mechanisms to more comprehensively and effectively capture the high-order interaction features between MHC and antigen peptides; it adopts a collaborative training framework and dynamic optimization strategy to significantly improve computing efficiency and prediction performance, providing efficient and accurate prediction tools for immunotherapy, vaccine design and autoimmune disease research, and promoting further development in the fields of precision medicine and biotechnology.
[0117] Example 2 On the other hand, the present invention also proposes a MHC presented peptide prediction system based on multimodal deep learning, please refer to Figure 3 , which shows an MHC-presented peptide prediction system based on multimodal deep learning in a second embodiment of the present invention, the system includes: Dataset construction module 11, used to construct a multi-species MHC-peptide binding dataset, including MHC molecule sequence information of human and non-human species and their corresponding peptide binding data; Furthermore, the data set construction module 11 is specifically used to: Collect clinical tumor samples and enrich and identify immune peptides from the clinical tumor samples using a magnetic bead-based automated immune peptide mass spectrometry platform; Perform whole-exome sequencing and HLA typing on the same sample to obtain MHC genotype and its corresponding antigen peptide pairing information; Integrate immunogenicity data from public databases, including information on the pairing of MHC class I molecules and antigenic peptides.
[0118] A data analysis module 12 is configured to perform single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence; Furthermore, the data analysis module 12 is specifically used to: Construct an intersection judgment mechanism between peptide segments and their related MHC molecule sets; Based on the intersection judgment mechanism, unique MHC-peptide correspondences are screened to achieve single genotype mapping extraction of peptides.
[0119] The data comparison module 13 is used to perform multiple sequence alignment and feature alignment of MHC sequences across species, and to extract core site information characterizing MHC molecular polymorphism using a frequency-weighted amino acid feature distance algorithm; Furthermore, the data comparison module 13 is specifically used to: Obtaining multiple species MHC sequences from multiple public databases, and uniformly aligning the multiple species MHC sequences using a multiple sequence alignment tool; The frequency-weighted amino acid feature distance algorithm was used to calculate the polymorphism index of each amino acid site in the aligned multi-species MHC sequences, and the sites with the longest amino acid polymorphism distance were selected as the core site information characterizing the polymorphism of MHC molecules.
[0120] Furthermore, the core site information contains 40 sites, which are embedded and encoded using the BLOSUM50 matrix to convert the biological sequence information into feature expressions processed by the machine learning model.
[0121] The data prediction module 14 is used to construct a multimodal deep learning model, integrate MHC sequence features and peptide sequence features, and predict the binding ability of MHC molecules and peptides.
[0122] Furthermore, the multimodal deep learning model includes a convolutional neural network module, a channel attention mechanism module, and a cross attention mechanism module; Among them, the convolutional neural network module is used to extract local features of MHC molecules and peptide sequences, the channel attention mechanism module is used to enhance the weights of important feature channels, and the cross-attention mechanism module is used to capture the sequence interaction patterns between MHC molecules and antigen peptides.
[0123] Furthermore, the system further comprises: Balanced sampling module, used to balance the sampling of bound samples and unbound samples and control the ratio of positive and negative samples; A data optimization module is used to normalize affinity values according to logarithmic transformation to improve numerical stability and model convergence speed.
[0124] The functions or operation steps implemented when the above modules and units are executed are substantially the same as those in the above method embodiments and will not be repeated here.
[0125] The MHC-presented peptide prediction system based on multimodal deep learning provided in the embodiments of the present invention has the same implementation principles and technical effects as those in the aforementioned method embodiments. For the sake of brief description, any matters not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiments.
[0126] Example 3 The present invention also provides a computer, see Figure 4 , shown is a computer in the third embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored in the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, the above-mentioned MHC-presented peptide prediction method based on multimodal deep learning is implemented.
[0127] The memory 10 includes at least one type of storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10 may be an internal storage unit of a computer, such as the computer's hard disk. In other embodiments, the memory 10 may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 10 may include both an internal storage unit of the computer and an external storage device. The memory 10 can be used not only to store application software installed in the computer and various types of data, but also to temporarily store data that has been output or is about to be output.
[0128] Among them, in some embodiments, the processor 20 can be an electronic control unit (Electronic Control Unit, abbreviated as ECU, also known as a vehicle computer), a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run the program code stored in the memory 10 or process data, such as executing access restriction programs.
[0129] It should be pointed out that Figure 4 The structure shown does not constitute a limitation of the computer. In other embodiments, the computer may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0130] An embodiment of the present invention further provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for predicting MHC-presented peptides based on multimodal deep learning.
[0131] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device.
[0132] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in another suitable manner as necessary, and then storing it in a computer memory.
[0133] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0134] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0135] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for predicting MHC-presented peptides based on multimodal deep learning, characterized in that: include: Construct a multi-species MHC-peptide binding dataset, including MHC molecule sequence information and corresponding peptide binding data of human and non-human species; performing single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence; Perform multiple sequence alignment and feature alignment on cross-species MHC sequences, and use a frequency-weighted amino acid feature distance algorithm to extract core site information characterizing MHC molecular polymorphism; A multimodal deep learning model was constructed to integrate MHC sequence features and peptide sequence features to predict the binding ability of MHC molecules and peptides.
2. The method for predicting MHC-presented peptides based on multimodal deep learning according to claim 1, characterized in that: The steps to construct a multi-species MHC-peptide binding dataset include: Collect clinical tumor samples and enrich and identify immune peptides from the clinical tumor samples using a magnetic bead-based automated immune peptide mass spectrometry platform; Perform whole-exome sequencing and HLA typing on the same sample to obtain MHC genotype and its corresponding antigen peptide pairing information; Integrate immunogenicity data from public databases, including information on the pairing of MHC class I molecules and antigenic peptides.
3. The method for predicting MHC-presented peptides based on multimodal deep learning according to claim 1, characterized in that: The steps of performing single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence include: Construct an intersection judgment mechanism between peptide segments and their related MHC molecule sets; Based on the intersection judgment mechanism, unique MHC-peptide correspondences are screened to achieve single genotype mapping extraction of peptides.
4. The method for predicting MHC-presented peptides based on multimodal deep learning according to claim 1, characterized in that: The steps of performing multiple sequence alignment and feature alignment on cross-species MHC sequences and extracting core site information characterizing MHC molecular polymorphism using a frequency-weighted amino acid feature distance algorithm include: Obtaining multiple species MHC sequences from multiple public databases, and uniformly aligning the multiple species MHC sequences using a multiple sequence alignment tool; The frequency-weighted amino acid feature distance algorithm was used to calculate the polymorphism index of each amino acid site in the aligned multi-species MHC sequences, and the sites with the longest amino acid polymorphism distance were selected as the core site information characterizing the polymorphism of MHC molecules.
5. The method for predicting MHC-presented peptides based on multimodal deep learning according to claim 4, characterized in that: The core site information contains 40 sites, which are embedded and encoded using the BLOSUM50 matrix to convert biological sequence information into feature expressions processed by machine learning models.
6. The method for predicting MHC-presented peptides based on multimodal deep learning according to claim 1, characterized in that: The multimodal deep learning model includes a convolutional neural network module, a channel attention mechanism module and a cross attention mechanism module; Among them, the convolutional neural network module is used to extract local features of MHC molecules and peptide sequences, the channel attention mechanism module is used to enhance the weights of important feature channels, and the cross-attention mechanism module is used to capture the sequence interaction patterns between MHC molecules and antigen peptides.
7. The method for predicting MHC-presented peptides based on multimodal deep learning according to claim 1, characterized in that: The method further comprises: Balance the sampling of bound samples and unbound samples to control the ratio of positive and negative samples; Affinity values were normalized according to logarithmic transformation to improve numerical stability and model convergence speed.
8. A multimodal deep learning-based MHC-presented peptide prediction system, characterized by: include: A dataset construction module is used to construct a multi-species MHC-peptide binding dataset, including MHC molecule sequence information and corresponding peptide binding data of human and non-human species; a data parsing module, configured to perform single genotype analysis on the multi-allele MHC-peptide data in the multi-species MHC-peptide binding dataset to determine a unique MHC-peptide correspondence; The data comparison module is used to perform multiple sequence alignment and feature alignment of MHC sequences across species, and uses a frequency-weighted amino acid feature distance algorithm to extract core site information that characterizes MHC molecular polymorphism; The data prediction module is used to build a multimodal deep learning model, integrate MHC sequence features and peptide sequence features, and predict the binding ability of MHC molecules and peptides.
9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for predicting MHC-presented peptides based on multimodal deep learning as described in any one of claims 1 to 7 is implemented.
10. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the MHC-presented peptide prediction method based on multimodal deep learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
New antigen sequence generation method and system based on deep learning
CN116403639A
MHC prediction model construction method and system based on degenerate coding and deep learning
CN117457079A
MHC-II and polypeptide combination prediction method
CN117953963A
Esophageal squamous carcinoma detection method and system based on combination of multi-modal information fusion and artificial intelligence
CN119399126A
Method and system for screening and constructing small molecule polypeptide simulant based on interaction between proteins
CN120108484A
Cited By
Method for predicting immunogenicity of antigen peptides and use thereof
CN122436004A