A glycopeptide epitope prediction method, storage medium and terminal device
By constructing an antigen presentation probability prediction model and a glycosylation site prediction model, the problem of glycopeptides that cannot be effectively detected in the prior art is solved, and efficient screening of glycopeptides with high immunogenicity and glycosylation characteristics is achieved, reducing research cost and time.
Patent Information
- Application Number
- CN202111080403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-09-15
AI Technical Summary
The prior art cannot effectively detect glycopeptides with strong binding strength to MHC, resulting in high cost and low efficiency in vaccine design and immune response research.
By constructing an antigen presentation probability prediction model and a glycosylation site prediction model, combining natural MHC data and glycopeptide database information, the antigen presentation fraction and glycosylation sites of the peptide to be tested are predicted, thereby screening out glycopeptides with high immunogenicity and glycosylation characteristics.
It improves the efficiency and accuracy of glycopeptide screening, reduces the cost and time of vaccine design and immune response research, and provides higher specificity and effectiveness.
Smart Images

Figure CN113851187B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network model application, and in particular to a glycopeptide epitope prediction method, a storage medium and a terminal device. Background Art
[0002] The major histocompatibility complex (MHC) is a gene cluster in the mammalian genome that is responsible for encoding part of the immune system. It plays an important role in transplant rejection, immune response, and immune regulation in mammals. During the immune response, antigen-presenting cells engulf and ingest antigenic foreign bodies containing protein macromolecules, bind processed epitope-containing polypeptide fragments to MHC molecules, and present them to T cells on the cell surface. MHC class I molecules present intracellular antigen peptides to cytotoxic T cells, and MHC class II molecules present extracellular antigen peptides to helper T cells to stimulate cellular immunity and humoral immunity. Therefore, identifying which peptides can bind to MHC molecules is very important for understanding the interaction between hosts and pathogens, and plays a key step in vaccine design, infectious diseases, autoimmunity, and cancer research.
[0003] Protein glycosylation is a highly heterogeneous post-translational modification that participates in a variety of biological processes such as cell signal transduction, cell adhesion, and cell recognition. It mainly includes protein N-glycosylation and protein O-glycosylation. Many pathogens are covered with dense polysaccharide coatings on their surfaces, and the surface of tumor cells also has significant glycosylation differences compared with normal cells. Therefore, polysaccharides are a particularly promising new candidate vaccine. Polysaccharides are coupled to immunogenic prototype carriers (such as immunogenic proteins, lipids, nucleic acids, etc.) or incorporated into liposome nanoparticles to induce humoral responses to bacterial polysaccharides, viral spike proteins, and tumor-associated polysaccharide antigens.
[0004] Sugar-dependent epitopes can trigger neutralizing antibody responses, such as HIV virus using glycosylation to achieve immune evasion. Studies on SARS-CoV-1 have shown that the sugar coat composed of sugar chains on the surface of the virus is very similar to the glycosylation of endogenous host cells, and immune evasion can be achieved by masking virus-specific polypeptides with polysaccharides. Recent studies on SARS-CoV-2 have also shown that the S protein that mediates the binding of the SARS-CoV-2 virus to target cells carries a large number of glycosylation modifications, and the interaction between the receptor binding domain of the S glycoprotein and the glycosaminoglycans of the host cell is the basis of SARS-CoV-2 infection; in the innate immune response and the adaptive immune response, glycosylation plays a protective role and can promote immune evasion by masking viral polypeptide epitopes; SARS-CoV-2 monoclonal antibody neutralization therapy also shows interaction with polysaccharide epitopes on the SARS-CoV-2 spike protein. All of this evidence shows the importance of glycosylation in responding to viral infections, especially in the development of vaccines and infection treatments. Therefore, understanding the glycosylation of peptides is the basis for the development of effective vaccines, neutralizing antibodies, and therapeutic infection inhibitors.
[0005] The ability of exogenous antigen peptides to bind to MHC molecules is the premise for the rational design of vaccines. Accurate evaluation of the binding force between the two plays a decisive role in discovering new MHC antigen epitopes. Glycosylation plays a key role in the invasion of viruses into host cells and is an important feature of tumor cells. Glycopeptides have a greater possibility of becoming epitopes and are a very potential candidate vaccine. Therefore, finding glycopeptides with strong binding force to MHC can not only significantly reduce the experimental cost of verifying cell epitopes and save experimenters' time, but also have higher specificity.
[0006] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0007] In view of the above-mentioned deficiencies of the prior art, the object of the present invention is to provide a glycopeptide epitope prediction method, storage medium and terminal device, aiming to solve the problem that the prior art cannot effectively detect glycopeptides with strong binding force to MHC.
[0008] The technical solution of the present invention is as follows:
[0009] A method for predicting glycopeptide epitopes, comprising the steps of:
[0010] Based on natural MHC data, a prediction model for antigen presentation probability was constructed;
[0011] Collect peptide glycosylation site information and its corresponding peptide sequence from the glycopeptide database to construct a glycosylation site dataset;
[0012] Based on the glycosylation site dataset, constructing a glycosylation site prediction model;
[0013] Inputting the peptide sequence to be predicted into the antigen presentation probability prediction model to predict the antigen presentation score of the peptide sequence to be predicted;
[0014] The peptide sequence to be predicted with an antigen presentation score greater than a threshold is input into the glycosylation site prediction model to predict the glycosylation site of the peptide sequence to be predicted.
[0015] The glycopeptide epitope prediction method, wherein the step of constructing an antigen presentation probability prediction model based on natural MHC data comprises:
[0016] Measure the affinity between natural MHC binding peptides and MHC to obtain affinity measurement data;
[0017] constructing an affinity prediction model based on the affinity measurement data;
[0018] Using mass spectrometry data from monoallelic samples, a prediction model for antigen processing was constructed;
[0019] The affinity prediction model and the antigen processing prediction model are combined to construct an antigen presentation probability prediction model.
[0020] The glycopeptide epitope prediction method, wherein the glycopeptide database includes the dbPTM database, the UniProt database, the OGP database, the PhosphoSitePlus database and the UniMod database.
[0021] The glycopeptide epitope prediction method, wherein the step of constructing a glycosylation site prediction model based on the glycosylation site dataset comprises:
[0022] A glycosylation site prediction model was constructed based on the glycosylation site dataset using site alignment, sequence alignment, pattern matching and machine learning.
[0023] In the glycopeptide epitope prediction method, the peptide sequence to be predicted is subjected to peptide length screening and hydrophobicity screening.
[0024] A storage medium, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the glycopeptide epitope prediction method of the present invention.
[0025] A terminal device, comprising: a processor, a memory and a communication bus; the memory stores a computer-readable program executable by the processor;
[0026] The communication bus realizes the connection and communication between the processor and the memory;
[0027] When the processor executes the computer-readable program, the steps in the glycopeptide epitope prediction method of the present invention are implemented.
[0028] Beneficial effects: The present invention provides a method for predicting glycopeptide epitopes, which can directly predict the immunogenicity of glycopeptides. While improving the efficiency of peptide screening, it can more accurately predict immunogenic glycopeptides, providing an important tool for vaccine design, autoimmunity and cancer research. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 The first flow chart of a glycopeptide epitope prediction method of the present invention.
[0030] Figure 2 The second flow chart of a glycopeptide epitope prediction method of the present invention.
[0031] Figure 3 The present invention is a structural principle diagram of a terminal device. DETAILED DESCRIPTION
[0032] The present invention provides a glycopeptide epitope prediction method, storage medium and terminal device. In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] The ability of exogenous antigen peptides to bind to MHC molecules is the premise for the rational design of vaccines. Accurate evaluation of the binding force between the two plays a decisive role in discovering new MHC antigen epitopes. Glycosylation plays a key role in the invasion of viruses into host cells and is an important feature of tumor cells. Glycopeptides have a greater possibility of becoming epitopes and are a very potential candidate vaccine. Therefore, finding glycopeptides with strong binding force to MHC can not only significantly reduce the experimental cost of verifying cell epitopes and save experimenters' time, but also have higher specificity.
[0034] Currently, there are many softwares available for predicting MHC antigen epitopes, but there are differences in epitope prediction programs, especially the source and utility of some servers are still uncertain, so the prediction results are uneven. Moreover, most of the current prediction software only focuses on the prediction of MHC binding affinity, and does not consider the influence of post-translational modifications of peptide segments. In view of the defects of existing prediction software, the present invention provides a glycopeptide epitope prediction method, such as Figure 1 As shown, it includes the steps of:
[0035] S10. Construct an antigen presentation probability prediction model based on natural MHC data;
[0036] S20, collecting peptide glycosylation site information and its corresponding peptide sequence from the glycopeptide database to construct a glycosylation site dataset;
[0037] S30, constructing a glycosylation site prediction model based on the glycosylation site dataset;
[0038] S40, inputting the peptide sequence to be predicted into the antigen presentation probability prediction model to predict the antigen presentation score of the peptide sequence to be predicted;
[0039] S50, inputting the peptide sequence to be predicted whose antigen presentation score is greater than a threshold into the glycosylation site prediction model to predict the glycosylation site of the peptide sequence to be predicted.
[0040] The method provided in this embodiment not only predicts the ability of antigen to bind to MHC, but also adds the prediction of peptide glycosylation, which lays the foundation for subsequent epitope analysis and can effectively improve the accuracy of vaccine design.
[0041] In some embodiments, Figure 2 As shown, the glycopeptide epitope prediction method specifically includes a model building process and a glycopeptide antigen prediction process, wherein the model building process specifically includes the following steps:
[0042] Measure the affinity between natural MHC binding peptides and MHC to obtain affinity measurement data;
[0043] constructing an affinity prediction model based on the affinity measurement data;
[0044] Using mass spectrometry data from monoallelic samples, a prediction model for antigen processing was constructed;
[0045] Combining the affinity prediction model and the antigen processing prediction model to construct an antigen presentation probability prediction model;
[0046] Collect peptide glycosylation site information and its corresponding peptide sequence from the glycopeptide database to construct a glycosylation site dataset;
[0047] A glycosylation site prediction model was constructed based on the glycosylation site dataset using site alignment, sequence alignment, pattern matching and machine learning.
[0048] In this embodiment, the natural MHC data are open data sets, including Sarkizova data set (Sarkizova, S. et al. Nat Biotechnol 38, 199–209 (2020).), Abelin data set (Abelin, J Get al. Immunity 46, 315–326 (2017).), data set provided by IEDB database (Vita, R. et al. Nucleic Acids Research 47, D339–D343 (2019).), Kim data set (Kim, Y. et al. BMC Bioinformatics 15, 241 (2014).) and SysteMHC data set (Shao, W. et al. Nucleic Acids Research 46, D1237–D1247 (2018).).
[0049] In the process of constructing the glycosylation site prediction model, the glycopeptide database includes the dbPTM database, the UniProt database, the OGP database, the PhosphoSitePlus database and the UniMod database.
[0050] In this embodiment, the affinity prediction model and the antigen processing prediction model are both obtained by training with a neural network-based deep learning model as the basic model. In the process of training the affinity prediction model, the natural MHC binding peptide sequence and allele are input, and the affinity between the natural MHC binding peptide and MHC is output. As an example, the neural network-based deep learning model is a CNN model, but is not limited thereto.
[0051] In this embodiment, the monoallelic sample refers to cells expressing a single MHC-I gene using transgenic technology.
[0052] In the present embodiment, the prediction results of the affinity prediction model and the prediction results of the antigen processing model are used as input data, and the antigen presentation probability prediction model is constructed based on a support vector machine, XGBoost or a logistic regression machine learning model. The antigen presentation score is the total score value calculated by combining the predicted probability of antigen processing and the predicted affinity. As described in the embodiments, the predicted antigen processing probability and affinity are used as input, and a model can be constructed using algorithms such as logistic regression, support vector machine or XGBoost to predict the total score value, i.e., the antigen presentation score.
[0053] In this embodiment, a glycosylation site prediction model is constructed based on the glycosylation site dataset using site alignment, sequence alignment, pattern matching and machine learning, wherein site alignment refers to glycosylation site alignment, and the glycosylation site refers to the position where the glycosylation modification occurs on the protein sequence. Site alignment is to check whether there is a glycosylation site at the corresponding position of the protein in the database; sequence alignment refers to comparing the identified epitope sequence with the known glycosylated peptide sequence; pattern matching, as an example, such as N-glycosylation, only occurs on asparagine with NXS / T (X≠Pro) motif (sequence matrix), and matching this type of motif on the epitope can quickly determine whether the corresponding glycosylation type can occur; machine learning refers to using existing glycosylation sites to train a machine learning model to predict whether glycosylation can occur at the corresponding site.
[0054] In this embodiment, if Figure 2 As shown, the glycopeptide antigen prediction process specifically includes the following steps:
[0055] For the protein sequence to be predicted, it is enzymatically digested into the peptide sequence to be predicted according to the specified enzyme digestion rules. If it is a peptide sequence, no additional processing is required;
[0056] Detect the properties of the peptide sequence to be predicted, including but not limited to peptide length, motif, amino acid composition, hydrophobicity, etc., and filter some peptides as needed;
[0057] Inputting the filtered peptide sequence to be predicted into the antigen presentation probability prediction model to predict its antigen presentation score, and screening the peptide sequence to be predicted according to a threshold value;
[0058] The screened peptide sequence to be predicted is input into the glycosylation site prediction model to predict the glycosylation site of the peptide sequence to be predicted.
[0059] Specifically, the candidate polypeptide sequences predicted by different algorithms are different, and various prediction methods need to be considered comprehensively, and the final selection is made in combination with practical applications. The polypeptide sequence finally selected is generally 15-20 bases, a single epitope generally contains 5-8 bases, and a polypeptide of 15-20 bases generally contains 1 or more epitopes. Relatively long polypeptide segments can better maintain consistency with natural proteins and are more likely to produce antibodies. The glycosylation site prediction method provided by the present invention combines multiple prediction algorithms to find the most satisfactory polypeptide, improve the prediction success rate, and reduce the cost of polypeptide synthesis. With the development of glycoproteomics, there are currently a large number of open high-quality glycosylation site, glycopeptide and glycoprotein data sets, which can be used to construct models for predicting glycosylation of peptide antigens. Based on the prediction of peptide epitopes, the possibility of glycosylation of polypeptides is predicted, and peptides are further screened to reduce the number of candidate antigen peptides, thereby further reducing costs.
[0060] In summary, the method of the present invention can directly predict the immunogenicity of glycopeptides, which can improve the efficiency of peptide screening and more accurately predict immunogenic glycopeptides, providing an important tool for vaccine design, autoimmunity and cancer research.
[0061] In some embodiments, a storage medium is also provided, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the glycopeptide epitope prediction method of the present invention.
[0062] Based on the above glycopeptide epitope prediction method, the present application also provides a terminal device, such as Figure 3 As shown, it includes at least one processor (processor) 20; display screen 21; and memory (memory) 22, and may also include a communications interface (Communications Interface) 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22 and the communication interface 23 can communicate with each other through the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logic instructions in the memory 22 to execute the method in the above embodiment.
[0063] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0064] The memory 22 is a computer-readable storage medium that can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implementing the methods in the above embodiments.
[0065] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, a variety of media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, may also be a transient storage medium.
[0066] In addition, the specific process of loading and executing multiple instruction processors in the storage medium and the terminal device has been described in detail in the above method and will not be described here one by one.
[0067] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for predicting glycopeptide epitopes, characterized in that: Includes steps: Based on natural major histocompatibility complex data, a prediction model for antigen presentation probability was constructed; Collect peptide glycosylation site information and its corresponding peptide sequence from the glycopeptide database to construct a glycosylation site dataset; Based on the glycosylation site dataset, constructing a glycosylation site prediction model; Inputting the peptide sequence to be predicted into the antigen presentation probability prediction model to predict the antigen presentation score of the peptide sequence to be predicted; The peptide sequence to be predicted with an antigen presentation score greater than a threshold is input into the glycosylation site prediction model to predict the glycosylation site of the peptide sequence to be predicted.
2. The method for predicting glycopeptide epitopes according to claim 1, characterized in that: The step of constructing an antigen presentation probability prediction model based on natural major histocompatibility complex data comprises: Measuring the affinity between the natural major histocompatibility complex binding peptide and the major histocompatibility complex to obtain affinity measurement data; constructing an affinity prediction model based on the affinity measurement data; Using mass spectrometry data from monoallelic samples, a prediction model for antigen processing was constructed; The affinity prediction model and the antigen processing prediction model are combined to construct an antigen presentation probability prediction model.
3. The method for predicting glycopeptide epitopes according to claim 1, characterized in that: The glycopeptide databases include dbPTM database, UniProt database, OGP database, PhosphoSitePlus database and UniMod database.
4. The method for predicting glycopeptide epitopes according to claim 1, characterized in that: The step of constructing a glycosylation site prediction model based on the glycosylation site dataset comprises: A glycosylation site prediction model was constructed based on the glycosylation site dataset using site alignment, sequence alignment, pattern matching and machine learning.
5. The method for predicting glycopeptide epitopes according to claim 1, characterized in that: The peptide sequence to be predicted is screened for peptide length and hydrophobicity.
6. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the glycopeptide epitope prediction method according to any one of claims 1 to 5.
7. A terminal device, characterized in that: include: Processor, memory and communication bus; The memory stores a computer-readable program executable by the processor; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, the steps in the glycopeptide epitope prediction method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method and equipment for neoantigen epitope prediction
CN105524984A
Systems and methods for the preparation of peptide-MHC-i complexes with native glycan modifications
WO2021050792A2