Prediction method for tumor neoantigen of glioblastoma

By constructing a fusion gene dataset and using AGFusion, NetMHC and AlphaFold3 to predict tumor neoantigens in glioblastoma, the problem of limited tumor neoantigen database was solved, and the effectiveness of accurate tumor neoantigen screening and immunotherapy was achieved.

CN120299509APending Publication Date: 2025-07-11BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455464.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has limited tumor neoantigen databases when treating glioblastoma, resulting in unsatisfactory therapeutic effects of immunotherapy and it is difficult to effectively stimulate the patient's immune system to recognize and remove cancer cells.

Method used

By constructing a fusion gene dataset, AGFusion predicts protein domain and exon structures, combining NetMHC to predict high-affinity peptides, and using AlphaFold3 to verify the spatial structure of the peptides, screening out tumor neoantigens with strong binding.

Benefits of technology

Accurate tumor neoantigen prediction for glioblastoma has been achieved, the tumor neoantigen database has been expanded, and the targeted drug development basis is provided to stimulate the immune system to recognize and kill tumor cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299509A_ABST
    Figure CN120299509A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biology and computers, in particular to a tumor neoantigen prediction method for glioblastoma. 226 peptide fragments are screened by combining a biological technology and a computer technology, and then 20 peptide fragments with the highest affinity and 16 fusion genes corresponding to the 20 peptide fragments and three groups of tumor neoantigens of glioblastoma with the highest value are selected and used for future clinical treatment and optimization of drug research and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of biological technology and computer technology, and particularly relates to a method for predicting tumor neoantigens of glioblastoma multiforme. Background Art

[0002] Glioblastoma multiforme (GBM) is a highly malignant solid tumor in the brain. The reason is that the unique biological characteristics of GBM cells are highly lethal. Currently, the main methods for treating GBM are to remove cancer cells by surgery and supplement with chemotherapy. GBM generally escapes treatment by counteracting the effects through cell death and the rapid proliferation of cancer cells, and the curative effect is not ideal. The emerging treatment method is to use cancer immunotherapy: the immune system kills tumor cells by recognizing immune-stimulating neoantigens and T cell-mediated cytotoxicity. Targeted therapy with immune checkpoint inhibitors (such as bevacizumab) can specifically kill glioblastoma cells. Due to the particularity of its location and pathogenesis, using tumor antigen peptides to treat GBM is an excellent treatment method.

[0003] Tumor neoantigens are antigens derived from single nucleotide variations (SNVs) and insertions or deletions (Indels), and can be recognized by neoantigen-specific T cell receptors (TCRs) in the context of major histocompatibility complex (MHC) molecules. Tumor neoantigens do not exist in the normal human body. At the same time, T cells that can specifically recognize such proteins can escape negative selection in the thymus and do not produce autoimmunity toxicity. Therefore, tumor neoantigens can be used as ideal immune targets for cancer treatment. Since the SNV and Indel sources of tumor neoantigens are limited, in order to expand the database of tumor neoantigens to achieve the goal of treating more patients, we focus on another source of it - gene fusion.

[0004] A fusion gene refers to two or more originally independent gene segments that, under certain conditions (such as chromosomal structural variations, viral infections, errors during cell damage repair, etc.), undergo physical connection to form a new gene with a unique sequence and function. This fusion can be an exchange of partial gene segments or an integration of entire genes. Fusion genes have been found to be prevalent in all major types of human tumors. The identification of fusion genes plays an important role as diagnostic and prognostic markers. Gene fusions result from chromosomal structural variations, which may generate an open reading frame (ORF). The peptide in the region spanning two gene breakpoints has a different structure from that of the self-antigen and is thus recognized by the body's T cells, triggering autoimmunity. Tumorigenesis related to fusion genes accounts for approximately 20% of the global cancer incidence. Therefore, it is feasible to broaden the repertoire of tumor neoantigens by discovering fusion genes. Predicting the tumor neoantigens formed by fusion genes provides a basis for the development of targeted drugs. Researchers can develop targeted drugs for specific tumors based on these neoantigens, stimulating the patient's immune system to eliminate cancer cells by recognizing the neoantigens and T cell-mediated cytotoxicity. Summary of the Invention

[0005] In view of the above defects, the first technical solution of the present invention discloses a tumor neoantigen of glioblastoma, including at least one of the antigenic peptides shown in the amino acid sequences of SEQ ID NO.1 to SEQ ID NO.18.

[0006] In addition, the application of the above-mentioned tumor neoantigen of glioblastoma in the preparation of drugs for preventing glioblastoma and / or treating glioblastoma and / or preventing the recurrence of glioblastoma.

[0007] In addition, a glioblastoma vaccine, including the above-mentioned tumor neoantigen of glioblastoma.

[0008] The second technical solution of the present application discloses a method for predicting tumor neoantigens of glioblastoma, including the following steps:

[0009] S1. Construct an initial fusion gene dataset of CBM patients;

[0010] S2. Use AGFusion to predict the protein domains and exon structures of the initial fusion genes to obtain predicted valid fusion genes;

[0011] S3. Use NetMHC to identify the peptide segments with high affinity in the predicted valid fusion genes;

[0012] S4. Use AlphaFold3 to predict and verify the protein spatial structures of the peptide segments to obtain CBM neoantigens.

[0013] Further, S2 is specifically as follows:

[0014] Input the initial fusion gene data information;

[0015] Query the AGFusion database to annotate the functional domain information and possible protein products of the initial fusion gene;

[0016] Run AGFusion to output the detailed information of the fusion gene and obtain the predicted effective fusion gene.

[0017] Further, S3 is specifically as follows:

[0018] Upload the protein Fasta file containing somatic mutation sites of the effective fusion gene on NetMHC;

[0019] Cut the protein sequence into short peptide segments with a cutting method of 8 - 11mer for MHC molecule affinity prediction;

[0020] Screen out the peptide segments with the SB label, which are the peptide segments with high affinity.

[0021] Further, S4 is specifically as follows:

[0022] Input the sequence information of the high - affinity peptide segment and the sequence of HLA - A0201 on the AlphaFold3 server website respectively to obtain the three - dimensional structure model, quality score and detailed annotation after the fusion of the two proteins;

[0023] Visualize the fused protein structure model, view the bond positions, bond lengths and hydrogen bond lengths HB at the binding site. If HB < 3.0, it is considered that the two proteins are bound;

[0024] Verify the spatial configuration of the bound peptide segment. If it is a high - confidence antigen peptide, it is determined as a novel tumor antigen of glioblastoma.

[0025] The beneficial effects of the present invention are as follows:

[0026] 1. Efficient data integration and processing: Using the dataset in the paper DriverFusions and Their Implications in the Development and Treatment of Human Cancers published in the Cell reports journal, the fusion genes of multiple cancers were detected from the TCGA database, and the information of GBM fusion genes was efficiently screened, ensuring the reliability and accuracy of the data source.

[0027] 2. Precise fusion gene fragment analysis: Using AGFusion to predict protein domains and exon structures, visualizes the key functional units of proteins to exert biological effects, accurately locates the breakpoints of unknown gene fusions and the retained exon regions; AGFusion can display the length, molecular weight, genes, transcripts contained in the protein, and whether there is an in-frame mutation. At the same time, the chemical properties of amino acids are visualized in different colors, including hydrophobicity, polarity, charge, and acidity, so as to fully understand the physical and chemical properties of fusion genes and new peptides.

[0028] 3. Bulk screening of neoantigens: NetMHC predicts the affinity of peptides binding to MHC-I through neural networks. Its neural network is trained using 81 different human MHC alleles, which can comprehensively predict tumor neoantigens; antigens bind to MHC molecules through epitopes. The length of the epitopes that MHC class I molecules can bind is 8 to 11 amino acids, ensuring the bulk screening of antigen prediction.

[0029] 4. Reliability verification of spatial structure: The function of proteins depends to a large extent on their three-dimensional structure. AlphaFold3 can accurately predict the folded three-dimensional structure of proteins from their amino acid sequences and provide visualization results to verify the reliability of their spatial structures one by one.

[0030] 5. The neoantigens screened in this application have strong affinity for binding to MHC-I, and thus can be used for the development and optimization of future drugs. Brief Description of the Drawings

[0031] Figure 1 is a flow chart of the method for predicting tumor neoantigens in the implementation of glioblastoma of the present invention;

[0032] Figure 2 is the situation of the ABR-YWHAE fusion gene predicted by AGFusion in the embodiment of the present invention: (A) exon structure, including the breakpoint position of gene fusion and the retained exon region; (B) amino acid sequence of the fusion protein and its visualization based on the EMBOSS color scheme;

[0033] Figure 3These are four groups of strongly binding peptides obtained by screening the fusion gene data of the embodiments of the present invention through NetMHC, expressing two top 20 high-affinity peptides: (A) Peptide information and affinity of peptides after the fusion of TRIP12 gene and CPS1 gene; (B) Peptide information and affinity of peptides after the fusion of CHIC2 gene and ADGRL3 gene; (C) Peptide information and affinity of peptides after the fusion of KLHL7 gene and TMEM106B gene; (D) Peptide information and affinity of peptides after the fusion of CLEC16A gene and TXNDC11 gene.

[0034] Figure 4 This is to verify the feasibility of the predicted peptide spatial structure by the embodiments of the present invention using AlphaFold3.

[0035] Figure 5 This is the expected position error of folded residues and consensus residues when predicting the spatial structure using AlphaFold3 in the embodiments of the present invention. Detailed implementation manners

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. The components of the embodiments of the present invention usually described and shown in the accompanying drawings here can be arranged and designed in various different configurations.

[0037] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0038] Unless otherwise specified, the meanings of the scientific and technical terms in this specification are the same as those generally understood by those skilled in the art. However, if there are conflicts, the definitions in this specification shall prevail.

[0039] In the context of the present invention, the terms "comprising" or "including" do not exclude other possible elements. The compositions of the present invention (including the various embodiments described herein) may comprise, consist of, or consist essentially of the following elements; the basic elements and essential limitations of the present invention described herein, as well as any other or optional ingredients, components, or limitations described herein or as otherwise required.

[0040] Now, the present invention will be described in more detail. It should be noted that the various aspects, features, embodiments, examples, and their advantages described in the present invention can be compatible and / or combined together.

[0041] Example 1 Prediction of Glioblastoma Neoantigens

[0042] As Figure 1 shown, the flow chart of the method for predicting tumor neoantigens of glioblastoma in this example is disclosed. Specifically:

[0043] S1. Construct an initial fusion gene dataset of CBM patients;

[0044] In this example, the initial fusion gene dataset of CBM patients is derived from the CBM patient fusion gene data in the research "Gao, Q., et al., Driver Fusions and Their Implications in the Development and Treatment of Human Cancers. Cell Reports, 2018. 23(1): p. 227 - 238.e3."; This research used a variety of fusion detection tools to systematically study the fusion situations of 9,624 tumors in 33 cancer types, studied the neoantigens generated by gene fusions, and a total of 25,664 fusion events were identified, with a verification rate of 63%. Among them, 25,664 cases of CBM fusion genes were recorded, including the detailed information of 536 CBM fusion genes, including the gene names involved, breakpoint positions, fusion directions, etc., as the initial fusion gene dataset.

[0045] S2. Use AGFusion to predict the protein domains and exon structures of the initial fusion genes to obtain predicted effective fusion genes (AGfusion prediction protein domains and exon structures - 470 efficient GBM fusion genes, Figure 1 )

[0046] In this example, the purpose of this step is to output the detailed information of the fusion genes. Download the python code and environment of AGFusion on the website https: / / github.com / murphycj / AGFusion as the running basis. The specific steps are as follows:

[0047] Input the initial fusion gene data information: The data information includes the detailed information of 536 CBM fusion genes, including the gene names involved, breakpoint positions, fusion directions, etc.

[0048] Query the AGFusion database to annotate the functional domain information and possible protein products of the initial fusion genes.

[0049] Run AGFusion to output the detailed information of the fusion genes and obtain the predicted valid fusion genes: The steps first output the retained gene fragments, protein domains, etc., and finally visualize the fusion genes and protein structures. The gene isotype combinations are shown at the top of the figure, and the fusion junctions are indicated by vertical lines at amino acid position 20 (bottom), including the kinases (red regions) and GED domains (blue regions) ( Figure 2 A).

[0050] After that, AGFusion outputs the detailed information of the fusion genes, the amino acid sequences of the fusion proteins and their visualization based on the EMBOSS color scheme, the protein length, molecular weight, the genes and transcripts included, and whether there is an in-frame mutation. At the same time, the colors of the amino acids are hydrophobic (yellow), polar (blue), charged (red), and acidic (green) according to their chemical properties ( Figure 2 B).

[0051] In this example, a total of 470 predicted valid fusion genes are output, and fusion products with carcinogenic functions are initially obtained.

[0052] S3. Use NetMHC to identify the peptides with high affinity in the predicted valid fusion genes (Net MHC identified 414 high-affinity peptides, Figure 1 );

[0053] In this example, this step is to initially screen out the gene fusion peptides with antigen potential, by downloading NetMHC-4.0 to the Linux system from the website http: / / www.cbs.dtu.dk / services / NetMHCpan / as the running basis.

[0054] The specific steps are as follows:

[0055] Upload the protein Fasta file containing somatic mutation sites of the valid fusion genes on NetMHC;

[0056] Cut the protein sequence into short peptide segments, with the cutting method being 8-11mer, and perform MHC molecule affinity prediction;

[0057] Screen out the peptides with the SB label, which are the peptides with high affinity: such as Figure 3As shown, after analyzing 470 fusion genes, a total of 414 strongly affinity peptides were screened, and the peptides expressed by the fusion gene with the smallest Affinity (nM), that is, the strongest affinity ability, were selected for subsequent analysis and processing. Among them, the fusion genes TRIP12_CPS1( Figure 3 A), CHIC2_ADGRL3( Figure 3 B), KLHL7_TMEM106B( Figure 3 C), and CLEC16A_TXNDC11( Figure 3 C) express two top 20 high-affinity peptides, preliminarily proving that they may become potential tumor treatment targets.

[0058] S4. Use AlphaFold3 to predict and verify the protein spatial structure of the said peptides, and obtain the CBM neoantigen (AlphaFold3 undergoing structural validation-GBM neoantigen-Future drug research).

[0059] In this embodiment, 226 peptides with %R < 0.1 were selected from the 414 strongly affinity peptides screened by NetMHC, and AlphaFold3 was used to verify the spatial structure after each peptide was linked to the MHC-I molecule binding complex (pMHC-I).

[0060] When visualizing the spatial structure, a bond length of less than 0.30 at the selected binding site can indicate that the predicted hydrogen bond can bind. As Figure 4 shown, where 4A and 4B are the peptide sequences synthesized by the TRIP12_CPS1 fusion gene. The two ends of the hydrogen bond of the peptide sequence YLFDSFFSL are connected to A:LYS 3046:N-Z:ASP 3182:OD1, and the hydrogen bond length HB = 2.281A (4A); the two ends of the hydrogen bond of the peptide sequence YLFDSFFSL are connected to A:SER 2371:O-G:GLU 2374:OE2, and the hydrogen bond length HB = 2.822A( Figure 4 B);

[0061] 4C and 4D are the peptide sequences synthesized by the CHIC2_ADGRL3 fusion gene. The two ends of the hydrogen bond of the peptide sequence FMICGILYV are connected to A:ASP 54:OD2:ALA235:N, and the hydrogen bond length HB = 2.897A( Figure 4 C); the two ends of the hydrogen bond of the peptide sequence FMICGILYV are connected to A:SER 28:O-ASP 53:N, and the hydrogen bond length HB = 2.813A( Figure 4 D).

[0062] 4E and 4F are the peptide segment sequences FMVDILAKV synthesized by the KLHL7_TMEM106B fusion gene. The two ends of the hydrogen bond are connected to A: GLN 78: NE2-TRP 75: O, and the hydrogen bond length HB = 2.755 Å( Figure 4 E); the two ends of the hydrogen bond of the peptide segment sequence FMVDILAKV are connected to A: ARG 287: NH2-LEU 194: O, and the hydrogen bond length HB = 2.669 Å Figure 4 F.

[0063] The model validation metrics of AlphaFold3 include PAE (Predicted Aligned Error), pTM (Predicted Template Modelling), iPTM (interface predicted TM-score), etc. The Predicted Aligned Error (PAE) map is used to evaluate the accuracy of protein structure prediction, and through the correlation with the distance change matrix in molecular dynamics simulation, a structural ensemble of disordered proteins and proteins containing disordered regions is constructed( Figure 5 ).

[0064] In the examples of the present invention, it is verified that among the predicted structural models of 226 pMHC-I complexes, 220 have a PAE value lower than 1.5, indicating that the predicted confidence of the screened antigen peptides is relatively high; the PTM values of all pMHC-I complexes are above 0.5, indicating that the overall predicted folding of the complex is similar to the real structure; the iPTM values of all complexes are above 0.8, indicating that the prediction of the antigen peptide is a high-confidence and high-quality prediction.

[0065] The complete concept of MHC refers to a group of closely linked gene clusters on a certain chromosome of vertebrates that encode major histocompatibility antigens, control cell - to - cell recognition, and regulate immune responses. The stronger the affinity of a peptide segment for binding to MHC - Ⅰ, the stronger its potential to become an antigen. Twenty peptide segments with the highest affinities were screened out from the output of NetMHC (their specific information is shown in Table 1), and their amino acid sequences are shown in SEQ ID NO.1 - SEQ ID NO.18. They are expressed by a total of 16 fusion gene combinations, namely: TRIP12 - CPS1, CHIC2 - ADGRL3, NBPF3 - EPHB2, MED13 - EPHB2, LRP5 - ATG16L2, PARN - CACHD1, HARBI1 - PTPRS, FREM2 - MTRF1, KLHL7 - TMEM106B, SLC39A3 - SGTA, PXDN - TRIP12, CLEC16A - TXNDC11, STON2 - SEL1L, GPLD1 - TDP2, ZNF544 - AP2A1, HMGA2 - NUP107 fusion genes.

[0066] Table 1

[0067]

[0068] In the above table, Position indicates the position of the peptide segment in the protein, indicating the starting position of the peptide segment and the original sequence position in the protein; peptide indicates the antigen peptide sequence recognized and presented by HLA molecules; Identity indicates the percentage similarity of the peptide segment to a given HLA type, and this field evaluates the matching degree of the peptide segment at the MHC binding site; Affinity (nM) indicates the binding affinity of the peptide segment to HLA molecules, and the unit is usually nM (nanomole). The smaller the IC50 value, the stronger the binding affinity of the peptide segment to HLA; %Rank indicates the percentage ranking of the binding affinity of the peptide segment among all peptide segments; Genes indicates the two fusion genes that form the peptide segment.

[0069] In the above table, the amino acids represented by the peptide segment sequences are shown in Table 2.

[0070] Table 2

[0071] Peptide sequence Amino acid Peptide sequence Amino acid YLFDSFFSL TyrLeuPheAspSerPhePheSerLeu FMICGILYV PheMetIleCysGlyIleLeuTyrVal MMMEDILRV MetMetMetGluAspIleLeuArgVal FVMGGVYFV PheValMetGlyGlyValTyrPheVal MLLDVMHTV MetLeuLeuAspValMetHisThrVal VLLRWLPPV ValLeuLeuArgTrpLeuProProVal MMMEVDQFV MetMetMetGluValAspGlnPheVal FMVDILAKV PheMetValAspIleLeuAlaLysVal YMYDFCTLI TyrMetTyrAspPheCysThrLeuIle MLLGSLLPV MetLeuLeuGlySerLeuLeuProVal LLLDVITWV LeuLeuLeuAspValIleThrTrpVal LLLEAVPAV LeuLeuLeuGluAlaValProAlaVal YLLSQVFLI TyrLeuLeuSerGlnValPheLeuIle YLMTIIALL TyrLeuMetThrIleIleAlaLeuLeu YAVDYSWYV TyrAlaValAspTyrSerTrpTyrVal KMYNVLLFV LysMetTyrAsnValLeuLeuPheVal TLFGVLYEV ThrLeuPheGlyValLeuTyrGluVal YLLRSLDYV TyrLeuLeuArgSerLeuAspTyrVal

[0072] Among them, the two fusion genes MED13 - EPHB2 and NBPF3 - EPHB2 synthesize the same MMMEDILRV peptide segment (SEQ IDNO.3), and the two fusion genes LANCL2 - VOPP1 and VOPP1 - LANCL2 synthesize the same YLLRSLDYV peptide segment (SEQ IDNO.18).

[0073] As can be seen from the above table, there are three fusion genes expressing two of the top 20 affinity peptide segments, which are the three most valuable groups in this screening, namely the TRIP12_CPS1 fusion gene, the CHIC2_ADGRL3 fusion gene, and the KLHL7_TMEM106B fusion gene; they express the antigen peptides YLFDSFFSL and TLFGVLYEV, FMICGILYV and LLLDVITWV, and FMVDILAKV and YMYDFCTLI, respectively.

[0074] In the TRIP12_CPS1 fusion gene, the TRIP12 gene is located in the 36.3 segment of the long arm of human chromosome 2, with 45 exons (NCBI Gene database https: / / www.ncbi.nlm.nih.gov / gene / 9320), which is related to the interaction with thyroid hormone receptors and may play a role in cell cycle regulation. It is involved in multiple signal transduction pathways, especially in regulating protein degradation; the CPS1 gene is located in the 34 segment of the long arm of human chromosome 2, with 43 exons (NCBI Gene database: https: / / www.ncbi.nlm.nih.gov / gene / 1373), mainly plays a role in the liver and is involved in the urea cycle. It catalyzes the synthesis of carbamoyl phosphate from ammonia and carbon dioxide, which is an important step in ammonia metabolism.

[0075] In the CHIC2_ADGRL3 fusion gene, the CHIC2 gene is located in the 12 segment of the long arm of human chromosome 4, with 9 exons (NCBI Gene database: https: / / www.ncbi.nlm.nih.gov / gene / 26511), and its function is not fully clear; the ADGRL3 gene is located in the 13.1 segment of the long arm of human chromosome 4, with 33 exons (NCBI Gene database: https: / / www.ncbi.nlm.nih.gov / gene / 23284), belonging to the GPCR family and involved in various physiological processes such as neural development and synaptic plasticity. Its function may be related to diseases such as brain development, behavior, and autism.

[0076] In the KLHL7_TMEM106B fusion gene, KLHL7 is located in the 15.3 segment of the short arm of human chromosome 7, with 15 exons (NCBI Gene database: https: / / www.ncbi.nlm.nih.gov / gene / 55975), which is involved in protein degradation and transcriptional regulation and is associated with certain neurological diseases and cancers; TMEM106B is located in the 21.3 segment of the short arm of human chromosome 7, with 9 exons (NCBI Gene database

[0077] https: / / www.ncbi.nlm.nih.gov / gene / 54664) is related to the health of the nervous system and has a certain connection with synaptic function and neurodegenerative diseases.

[0078] In the CLEC16A_TXNDC11 fusion gene, CLEC16A is located in the 13.13 segment of the short arm of human chromosome 16, with 33 exons (NCBI Gene database:

[0079] https: / / www.ncbi.nlm.nih.gov / gene / 23274), which is related to the function of the immune system and autoimmune diseases (such as type 1 diabetes); TXNDC11 is located in the 13.13 segment of the short arm of human chromosome 16, with 15 exons (NCBI Gene database https: / / www.ncbi.nlm.nih.gov / gene / 51061), and is involved in cellular antioxidant responses and protein folding.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tumor neoantigen of glioblastoma, characterized in that, Comprising at least one of the antigenic peptides shown by the amino acid sequences as SEQ ID NO.1 to SEQ ID NO.

18.

2. Use of the neoantigen of glioblastoma according to claim 1 in the preparation of a drug for preventing glioblastoma and / or treating glioblastoma and / or preventing recurrence of glioblastoma.

3. A glioblastoma vaccine, characterized in that, Comprising the neoantigen of glioblastoma according to claim 1.

4. A method for predicting tumor neoantigens of glioblastoma, characterized in that, Comprising the following steps: S1. Construct an initial fusion gene dataset of CBM patients S2. Use AGFusion to predict the protein domains and exon structures of the initial fusion genes to obtain predicted valid fusion genes; S3. Use NetMHC to identify the peptide segments with high affinity in the predicted valid fusion genes; S4. Use AlphaFold3 to predict and verify the protein spatial structure of the peptide segments to obtain CBM neoantigens.

5. The prediction method according to claim 4, characterized in that, The specific content of S2 is as follows: Input the initial fusion gene data information; Query the AGFusion database to annotate the functional domain information and possible protein products of the initial fusion genes; Run AGFusion to output the detailed information of the fusion genes to obtain the predicted valid fusion genes.

6. The prediction method according to claim 4, wherein The specific content of S3 is as follows: Upload the protein Fasta file containing somatic mutation sites of the valid fusion genes on NetMHC; Cut the protein sequence into short peptide segments, with the cutting method being 8 - 11mer, and perform MHC molecule affinity prediction; Screen out the peptide segments with the SB label, which are the peptide segments with high affinity.

7. The prediction method according to claim 4, wherein The specific content of S4 is as follows: Input the peptide segment sequence information and the sequence of HLA - A0201 on the AlphaFold3 server website respectively to obtain the three - dimensional structure model, quality score and detailed annotation after the fusion of the two proteins; Visualize the fused protein structure model, check the bond positions, bond lengths and hydrogen bond lengths HB at the binding site. If HB < 3.0, it is considered that the two proteins are bound; Verify the spatial configuration of the bound peptide segments. If they are antigenic peptides with high confidence, they are determined to be the neoantigens of glioblastoma.