Recognition method and recognition system of tumor fusion gene neoantigen

Through the method based on nanopore sequencing technology, tumor-specific fusion gene neoantigens are identified and screened out, which solves the problems of low recognition efficiency and small number of neoantigens in the prior art, and achieves more efficient and higher quality neoantigens recognition.

CN120072064APending Publication Date: 2025-05-30BEIJING HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510147261.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing tumor neoantigens recognition methods based on second-generation sequencing technology have problems such as sequencing read length, high splicing error rate, and GC content affecting sequencing accuracy, resulting in low recognition efficiency and small number of neoantigens.

Method used

Using a nanopore sequencing method, RNA sequencing data of patients' tumor tissue and adjacent control tissue were obtained, genome replies and gene fusion event analysis were performed, and tumor-specific fusion gene neoantigens were screened out in combination with a multi-module scoring system.

Benefits of technology

The detection efficiency of tumor-specific gene fusion events has been improved, the number and quality of neoantigens produced have been significantly improved, and the immunogenic potential is also higher, making it suitable for targeted selection for tumor immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005267124020000081
    Figure BDA0005267124020000081
  • Figure BDA0005267124020000091
    Figure BDA0005267124020000091
  • Figure BDA0005267124020000101
    Figure BDA0005267124020000101
Patent Text Reader

Abstract

The invention relates to the technical field of medicine, in particular to an identification method and an identification system of a tumor fusion gene neoantigen. According to the method, a standardized gene fusion new antigen clinical screening system is established by combining specific gene fusion in tumors on the basis of the advantages of nanopore sequencing gene fusion new antigens. In practical application, the system is researched and developed to enable a user to visually screen new antigens according to scores and rankings of the new antigens, compared with a traditional fusion gene new antigen screening method, more reliable new antigen targets with higher immunogenicity can be provided for patients, tumor immunotherapy schemes suitable for the patients can be screened out, and the purposes of optimizing treatment strategies and improving the tumor immunotherapy efficiency are achieved. Clinical diagnosis and treatment are enabled, and personalized treatment decision support is provided for patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical technology, and in particular to a method and system for identifying neoantigens of tumor fusion genes. Background Art

[0002] As the most widely used biomarker in clinical tumor diagnosis, the importance of neoantigens in diagnosis and treatment is self-evident. While recent years have seen a surge in exploration of diagnostic and treatment approaches such as humoral diagnostics, targeted therapy, minimally invasive surgery, and stereotactic radiotherapy, neoantigens, once considered traditional indicators, have also garnered significant attention and progress in screening and diagnosis, postoperative monitoring, efficacy prediction for advanced patients, and as therapeutic targets.

[0003] Tumor-specific neoantigens are new oncofetal antigens produced by gene mutations, abnormal RNA splicing, or other genomic changes in tumor cells that affect the translation reading frame. Because these neoantigens differ from self-antigens in normal tissues, they are considered tumor cell-specific. Tumor neoantigens are ideal targets for cancer immunotherapy because, compared to exogenous proteins, they are truly present in human tissues. These antigens can be recognized by T cells because they have not undergone thymic screening and are not easily affected by immune tolerance mechanisms. They can be recognized as foreign bodies by the immune system and trigger an immune response, leading to the killing of tumor cells.

[0004] Gene fusion is a new gene generated by the rearrangement of fragments of two or more different genes at the chromosomal level. This gene mutation can lead to abnormal proliferation and metastasis of tumor cells. In tumors, some fusion genes have been found to be associated with the occurrence, development and treatment resistance of tumors. For example, RSPO3-related fusion genes may lead to abnormal activation of the Wnt signaling pathway, thereby promoting the proliferation, metastasis and drug resistance of tumor cells. Gene fusion neoantigens have more advantages in terms of number and immunogenic potential than neoantigens derived from single nucleotide variants (SNVs) and insertions or deletions (Indels).

[0005] Gene fusion neoantigens also play a crucial role in subsequent treatment, guiding treatment selection, monitoring treatment efficacy and prognosis, and providing a crucial reference for developing new treatment strategies. Therefore, building a gene fusion neoantigen system based on nanopore sequencing technology can help improve the personalization and precision of tumor treatment, resulting in better treatment outcomes and quality of life for patients.

[0006] At present, the research patents on tumor neoantigens at home and abroad are mainly based on the second-generation sequencing technology to detect the mutation (SNV, Indels) information of the tumor, and then predict the encoded neoantigens through the mutation information [CN118351934A]. In previous reports, researchers believe that the generation of tumor-specific neoantigens is usually related to the mutation load of tumor cells, that is, the presence of a large number of gene mutations or other genomic changes in tumor cells. These mutations may produce new protein sequences, some of which may be recognized as foreign antigens by the immune system. Therefore, tumor-specific neoantigens are of great significance in cancer immunotherapy and can be used as targets for targeted therapy.

[0007] However, there are some major drawbacks to studying neoantigens based on second-generation sequencing technology. The first is its short sequencing read length. When targeting genes encoding multiple complex transcripts, the splicing of short sequences can easily cause splicing errors, resulting in the inability to truly evaluate post-transcriptional modification events such as selective splicing and gene fusions. Secondly, its sequencing accuracy is limited by the GC content, that is, regions with a GC content of 50% on the genome are more easily detected, and this part produces more sequences and higher coverage, but regions with high and low GC content are relatively difficult to detect. Therefore, when mutations occur in these regions, the second-generation sequencing technology will affect the sequence detection of this region to a certain extent. The sequences detected by nanopore sequencing are not affected by GC content. It is a sequencing technology that relies on electrical signals rather than fluorescent signals, and its long read length sequencing advantage can make up for the short sequences and sequence splicing of second-generation sequencing. The transcript sequence information in the sample can be obtained through only one library construction and sequencing, which is more suitable for biological studies such as selective splicing and gene fusion.

[0008] The gene fusion neoantigens currently being studied also have more advantages than SNV and Indel neoantigens. First, in terms of the number of gene fusion neoantigens, compared with neoantigens generated by previous methods, gene fusion can produce more tumor-specific neoantigens, and 5.8% of fusion neoantigens are shared between patients. Secondly, in terms of immunotherapy, including MHC presentation peptides and T cell recognition of pMHC, gene fusion neoantigens often have higher immunogenic potential than SNV and Indel neoantigens. These advantages indicate that gene fusion neoantigens can serve as new therapeutic targets and provide new ideas and methods for tumor immunotherapy.

[0009] In summary, in previous technologies, there were many methods for exploring new antigens based on second-generation sequencing technology. The number of new antigens mined based on this method was small and the immunogenicity was low. This unreliable and inefficient method often led to unsatisfactory treatment effects for patients. Summary of the Invention

[0010] In light of this, the present invention provides a method and system for detecting tumor fusion gene neoantigens based on nanopore sequencing. Compared with neoantigen identification methods based on next-generation sequencing, this method increases the efficiency of tumor-specific gene fusion events by five times, and significantly improves the number and quality of neoantigens generated.

[0011] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0012] The method for identifying tumor fusion gene neoantigens comprises the following steps:

[0013] Step 1: Obtain the patient's tumor tissue and adjacent normal control tissue, extract RNA, construct nanopore sequencing libraries, and sequence them to obtain RNA sequencing data for the patient's tumor tissue and adjacent normal control tissue;

[0014] Step 2: Perform genome alignment on the RNA sequencing data to obtain alignment data;

[0015] Step 3: Analyze the comparison data to determine gene fusion events, select fusion gene sequences with ≥5 sequence supports as candidate sequences, obtain tumor-specific fusion gene sequences, and then use candidate sequence analysis to obtain neoantigens;

[0016] Each neoantigen is scored according to the scoring method shown in the following five modules, and scores S1 to S5 are obtained for each candidate sequence:

[0017] a. Tumor-specific fusion gene module: The score of the new antigen S1 = the number of candidate sequences supporting the antigen / sequencing depth;

[0018] b. Coding potential assessment module: The result score of the original candidate sequence CPPred corresponding to the new antigen is S2;

[0019] c. Protein prediction module: For the original candidate sequence corresponding to the new antigen, the Orfipy calculation result is 1 if there is a corresponding sequence, and 0 otherwise;

[0020] d. HLA typing module: For the tumor RNA sequencing data in step 1, if there is a clear typing, S4 is 1, otherwise S4 is 0;

[0021] e. Immunogenicity potential assessment module: Using protein prediction results and HLA typing as input, the new antigen score S5 is calculated by netMHCpan (v4.1):

[0022] S5=-log 10 (Rank-corrected value)*number of bound HLA;

[0023] Step 4: Based on the weight of each module, the total score S of each neoantigen is calculated according to the following formula, and the top 100 candidate sequences with the highest S scores are identified as tumor neoantigen sequences;

[0024] Total score S = a*S1+b*S2+c*S3+d*S4+e*S5;

[0025] Among them, scores a to e are the weights of the tumor-specific fusion gene module, coding potential assessment module, protein prediction module, HLA typing module, and immunogenic potential assessment module, respectively.

[0026] In step 1 of the present invention, the data volume of the RNA sequencing is not less than 15G; the adjacent cancer control tissue is a tissue that is >5 cm away from the tumor tissue.

[0027] In the present invention, step 2 includes: using minimap2 software to align the RNA sequencing data to the genome; and using samtools software to sort and index the aligned data files to obtain aligned data files.

[0028] In step 3 of the present invention, LongGF software is used to respectively identify gene fusion events in tumor tissue and adjacent control tissue in the comparison data, and determine tumor-specific fusion gene sequences; wherein, the screening conditions for the gene fusion events are: the minimum alignment length with the gene overlap is 100, the bin interval on the genome is 50, the minimum alignment length with the reference genome is 100, and the situation of pseudogenes and sequence secondary alignment is taken into account, and the results are selected using the keyword "SumGF".

[0029] In step 3 of the present invention, the coding potential assessment is performed using CPPred software, and the assessment indicators include: ORF length, ORF coverage, ORF integrity, Fickett score, Hexamer score, PI, Gravy, instability index and CTD characteristics.

[0030] In step 3 of the present invention, the HLA typing prediction specifically includes: based on the long sequence fastq data of adjacent adjacent cancer tissue, using SpecHLA software to perform fine internal correction and local assembly on it to achieve accurate HLA typing of the patient.

[0031] In step 3 of the present invention, a and e are 30% respectively, c is 20%, b and d are 10% respectively, and the sum of a to e is 100%.

[0032] The present invention also provides a tumor fusion gene neoantigen recognition system based on the recognition method, comprising: a module determination unit, a weight design unit, a score evaluation unit and a system architecture;

[0033] The module determination unit includes a gene fusion module, a coding potential assessment module, a coding protein prediction module, an HLA typing prediction module and an immunogenic potential assessment module;

[0034] The present invention also provides an electronic device, comprising a processor and a memory;

[0035] The memory stores computer program instructions, and the processor is used to execute the computer program instructions to implement the method for identifying tumor fusion gene neoantigens as described in the present invention.

[0036] The present invention also provides a computer-readable storage medium storing a computer program, which is called by a processor to implement the method for identifying tumor fusion gene neoantigens described in the present invention.

[0037] The present invention also provides new antigenic peptides for colorectal cancer, comprising at least one of the antigenic peptides having amino acid sequences as shown in SEQ ID NOs: 1 to 21.

[0038] The specific sequences are as follows:

[0039] SEQ ID NO:1: KSNDSQNRL; SEQ ID NO:2: VATSFIRTI; SEQ ID NO:3: LVDQFKVTL; SEQ ID NO:4: LSFQPSPRL; SEQ ID NO:5: RLSPEPGSAL;

[0040] SEQ ID NO: 6: WPGVDARGAAA; SEQ ID NO: 7: RLSPEPGSALG; SEQ ID NO: 8: RRLSPEPGSAL; SEQ ID NO: 9: PGAPRMLIRLG; SEQ ID NO: 10: ALPAFVALL; SEQ ID NO: 11: LLLLSPWPL; SEQ ID NO: 12: LLLSPWPLL; SEQ ID NO: 14: LLLLSPWPLL; SEQ ID NO: 15: GQFSAVHPNV; SEQ ID NO: 16: HLMDIAIIV; SEQ ID NO: 17: IMCSANWAI; SEQ ID NO: 18: IISDEVNFLV; SEQ ID NO: 19: GLVDQLVGPL; SEQ ID NO: 20: GLLNVCMNA;

[0041] The present invention also provides a biomaterial comprising any one of the following:

[0042] 1. A nucleic acid encoding the colorectal cancer neoantigen peptide of the present invention;

[0043] II. an expression cassette comprising the nucleic acid described in I;

[0044] III. A recombinant vector comprising the nucleic acid described in I or the expression cassette described in II;

[0045] IV, transfect or transform the host of the recombinant vector described in III,

[0046] The present invention also provides the use of the colorectal cancer neoantigen peptide or the biomaterial in any of the following:

[0047] i. Preparation of diagnostic reagents for colorectal cancer;

[0048] ii. Drugs for preventing and treating colorectal cancer.

[0049] In the present invention, the drugs include protein vaccines, DNA vaccines, RNA vaccines or car-T products.

[0050] The present invention provides a system and method for identifying tumor fusion gene neoantigens. This method is based on the advantages of nanopore sequencing gene fusion neoantigens and combines specific gene fusions in tumors to establish a standardized gene fusion neoantigen clinical screening system (see flowchart for details). Figure 1 , the user manual and the specific results of each part during use are shown in Figure 2 The present invention has at least one of the following advantages:

[0051] 1. The nanopore RNA sequencing technology relied upon by the present invention has the advantage of longer sequencing read lengths, and is five times more efficient than second-generation sequencing in mining specific gene fusion events in patient tumor tissues.

[0052] 2. Based on the screening system of the present invention, the resulting gene fusion can produce 6 times more neoantigens and 11 times more tumor-specific neoantigens, far exceeding the original method in terms of quantity and quality.

[0053] 3. Based on the screening system of the present invention, the immunogenicity of the new antigens screened was verified using a humanized mouse model and was superior to that of the traditional method.

[0054] 4. The recognition system provided by the present invention is a visualized candidate neoantigen system, which enables medical personnel to more intuitively and accurately screen tumor immunotherapy plans suitable for patients based on patient scores and rankings. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Shown is the experimental flow chart of the present invention;

[0056] Figure 2Show the user manual and main components of the new antigen screening system;

[0057] Figure 3 Schematic diagram showing calculation of the total score S;

[0058] Figure 4 The S score distribution results of Example 1 and Comparative Example 1 are shown;

[0059] Figure 5 The immunogenicity of Example 1 and Comparative Example 1 was verified by animal experiments.

[0060] Figure 6 The Elispot test results of the new antigen comparison of Example 1 and Comparative Example 1 are shown. DETAILED DESCRIPTION

[0061] The present invention provides a method and system for identifying neoantigens of tumor fusion genes. Those skilled in the art can refer to the contents of this article and appropriately improve the process parameters to achieve the desired results. It should be noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications herein without departing from the content, spirit, and scope of the present invention to implement and apply the technology of the present invention.

[0062] The test materials used in the present invention are all common commercial products and can be purchased in the market.

[0063] The existing methods for screening and identifying new fusion gene antigens focus on clarifying the accuracy of the fusion gene sequence, and construct a series of complex calculation formulas for scoring based on the type of fusion gene event annotated in the database, the number of fusion gene event prediction software, whether the encoding open reading frame (ORF) of the upstream and downstream genes of the fusion gene has changed, and the distance between the two genes upstream and downstream of the fusion gene breakpoint. This reduces the repeatability of the method and its effectiveness remains to be discussed.

[0064] The identification method constructed by the present invention has significant advantages in fusion gene detection. It can be intuitively verified using sequencing sequences, avoiding cumbersome calculation methods. And based on this advantage, the identification method constructed by the present invention includes modules: coding potential assessment module, protein prediction module, HLA typing module, and immunogenic potential assessment module. On the basis of screening reliable fusion sequences, it also considers the reliability of neoantigen production and function. From the perspective of identification method, it is superior to traditional methods based on second-generation sequencing to detect fusion gene neoantigens.

[0065] Specifically, the clinical prediction software (i.e., identification system) provided by the present invention mainly includes four parts, namely module determination, weight design, score evaluation and system architecture.

[0066] In terms of module selection, the software primarily includes five modules: gene fusion event mining, fusion sequence coding potential assessment, coding protein prediction, patient HLA typing, and immunogenicity assessment. Gene fusion event mining and neoantigen immunogenicity assessment are key steps in predictive software development, significantly improving model accuracy and functionality. Therefore, these two modules are given significantly higher weighting and score assessment than the potential coding protein prediction module.

[0067] In terms of weighting design, the five modules mentioned above were weighted according to their importance based on previous research experience. The immunogenicity assessment modules for exploring gene fusion events and neoantigens each accounted for 30%, the coding protein prediction module accounted for 20%, and the more mature module for assessing the coding potential of fusion sequences and the patient HLA typing module each accounted for 10%, for a total of 100%.

[0068] In terms of score evaluation, the results of clinical prediction software need to be reflected through rankings or scores, allowing users to select candidate antigens based on the quantitative results. Therefore, for the gene fusion event mining module, the number of standardized supporting sequences is used as the main measurement factor; the module for evaluating the coding potential of fusion sequences uses the coding potential of CPPred results as the main measurement indicator; the coding protein prediction module uses the cds.fa results of agfusion, with a value of 1 if there is a corresponding sequence and 0 otherwise; the patient HLA typing module is similar to the coding protein prediction module, with a value of 1 if there is a corresponding typing and 0 otherwise; and the immunogenic potential assessment module uses the bindingRank Score of netMHCpan (v4.1) as the score measurement factor.

[0069] Regarding the software system architecture, the hardware environment is primarily server-based. The web deployment server uses the CentOS Linux release 7.9.2009 operating system, equipped with JDK 9.0.4, MySQL Server, and Nginx. The business code server uses Ubuntu 20.04.4 LTS, Python 3.8 and Python 2.7, and has a hard drive capacity of >3TB. The client web server is configured to run IE 6.0 or higher, and uses the recommended hardware configuration for Windows platforms or higher. Interfaces include SSH connections between the web deployment server and the business code server; HTTP network interfaces between the web frontend and backend; and interfaces between the server and the MySQL Server database. The neoantigen prediction results of the screening and scoring system are ultimately displayed on the front-end web page.

[0070] As research on immunotherapy continues to deepen, constructing a scoring system for gene fusion neoantigens in tumors is of great significance for the subsequent targeted treatment of cancer patients. Such a scoring system can optimize treatment strategies, empower clinical diagnosis and treatment, and provide patients with personalized treatment decision support.

[0071] The present invention also provides a new colorectal cancer antigen obtained by screening by the method of the present invention, whose amino acid sequence is selected from any one of the following: KSNDSQNRL, VATSFIRTI, LVDQFKVTL, LSFQPSPRL, RLSPEPGSAL, WPGVDARGAAA, RLSPEPGSALG, RRLSPEPGSAL, PGAPRMLIRLG, ALPAFVALL, LLLLSPWPL, LLLSPWPLL, RLFFALERI, LLLLSPWPLL, GQFSAVHPNV, HLMDIAIIV, IMCSANWAI, IISDEVNFLV, GLVDQLVGPL, GLLNVCMNA, SIISDEVNFLV.

[0072] The present invention will be further described below in conjunction with the embodiments:

[0073] Example 1 Screening for Gene Fusion Neoantigens

[0074] First, a pair of tumor tissues from surgically resected colorectal cancer patients and adjacent control tissues located >5 cm away from the tumor tissues were obtained for RNA extraction and library construction using nanopore sequencing. The amount of sequenced RNA data was no less than 15G.

[0075] The data analysis is as follows:

[0076] 1. Sequencing data preprocessing: The pod5 data file set was converted into fastq sequence using MinKNOW (v23.11.8) software, and then multiple fastq files under the same sample environment were cat-integrated. The results are as follows Figure 3 As shown, each pod5 file is converted into a corresponding fastq file, and then the fastq files of the same sample are merged into an overall fastq sequence file.

[0077] 2. Genome alignment of sequencing data. The integrated fastq files were aligned to the genome using minimap2 (v2.17-r941) software. The genomic data used was hg38, and the "splice" mode was selected for alignment. The aligned data files were then sorted and indexed using samtools (v1.9) software.

[0078] 3. Mining specific fusion genes in tumors. After the early preprocessing of the data, LongGF software was used to identify gene fusion events in tumors and adjacent control tissues. The input files for both are the files after genome alignment obtained in step 2. In the process of using LongGF software, the minimum alignment length with gene overlap is 100, the bin interval on the genome is 50, and the minimum length of alignment with the reference genome is 100. In the mining process, pseudogenes and sequence secondary alignments are considered, and then the keyword "SumGF" is used to select the results. After obtaining the fusion sequences in the tumor and control tissues respectively, the specific fusion sequences that only appear in the tumor and the minimum number of sequences supporting the gene fusion sequence is 5 are picked out as candidate sequences. Table 2 shows the specific gene fusion sequence information in the colorectal cancer sample tumor, including the serial number of the fusion sequence, the constituent genes of the fusion sequence, the number of sequences supported by the fusion event, and the corresponding chromosome fusion position.

[0079] 4. Evaluate the coding potential of fusion genes and predict potential proteins. For the gene fusion sequences obtained in step 3, CPPred software was used to distinguish non-coding RNA from coding RNA based on ORF length, ORF coverage, ORF completeness, Fickett score, Hexamer score, PI, Gravy, instability index, and CTD features. The coding potential of each fusion sequence is presented in Table 3. For gene fusion sequences with coding potential, Orfipy software was then used to extract the coding reading frames within the sequences to obtain neoantigen sequences.

[0080] 5. Clarify the patient's HLA typing. SpecHLA software was used to perform fine internal correction and local assembly based on the long sequence fastq data of the control tissue to achieve accurate HLA typing of the patient. The results are shown in Table 4.

[0081] 6. Evaluate the immunogenic potential of candidate neoantigens. Using netMHCpan (v4.1), using the results from steps 4 and 5 as input, we determine the binding potential of the candidate gene fusion neoantigen to the patient's HLA, providing a quantitative score for the immunogenicity assessment. Rank is the ranking of the predicted binding score for the antigen compared to a random natural peptide, with a % rank < 0.5 indicating strong binding.

[0082] The candidate Fu1 fusion neoantigens for this colorectal cancer patient are shown in Table 5. EL refers to the natural ligand, and BA refers to the corresponding MHC ligand. The top-ranked gene fusion neoantigens are ranked according to their Rank-corrected values ​​and the number of HLA-binding antigens. These antigens have strong immunogenic potential, as shown in Table 6.

[0083] Table 1 Partial results after fastq file merging

[0084]

[0085] Table 2 Fusion sequence results

[0086]

[0087] Table 3 Coding potential of fusion sequences

[0088] Serial number ORF length ORF coverage ORF integrity Fickett Hexamer PI Gravy Instability Index Classification Coding potential Fu_1 1131 0.130 1 0.493 -0.212 7.965 -0.293 39.107 coding 1.000 Fu_2 1140 0.215 1 0.499 -0.306 6.215 -0.106 46.679 coding 1.000 Fu_3 639 0.196 1 0.485 -0.408 5.275 -0.111 37.296 coding 0.980 Fu_4 678 0.155 1 0.472 -0.156 9.746 -0.538 55.027 coding 0.965 Fu_5 153 0.075 1 0.443 -0.311 10.834 0.298 68.938 noncoding 0.027 Fu_6 564 0.379 1 0.953 -0.195 5.350 -1.097 62.036 coding 0.834

[0089] Table 4 HLA typing results of patients

[0090] HLA typing HLA_A_1 HLA_A_2 HLA_B_1 HLA_B_2 HLA_C_1 HLA_C_2 N1 A*01:03 A*02:06 B*48:01 B*46:01 C*08:01 C*01:02

[0091] Table 5 New antigen immune potential of fusion sequences

[0092]

[0093] Table 6 Top ten gene fusion neoantigens by scores

[0094]

[0095] Example 2 Operation of the tumor neoantigen screening system

[0096] Based on Example 1, a patient may have multiple gene fusion sequences with high reliability, and each fusion sequence corresponds to a different neoantigen. Therefore, constructing a tumor neoantigen screening system can help patients screen for neoantigens by ranking them.

[0097] 1. Module determination: For the patient data in Example 1, the modules are determined to be the five modules of mining gene fusion events, evaluating the coding potential of fusion sequences, predicting the encoded protein, patient HLA typing, and evaluating the immunogenicity potential.

[0098] 2. Calculate the score for each module:

[0099] a. Gene fusion module: Candidate sequence score S1 = number of supporting sequences / sequencing depth;

[0100] b. Coding potential module: CPPred result score is S2;

[0101] c. Protein prediction module: For the Orfipy results, if there is a corresponding sequence, S3 is 1, otherwise S3 is 0;

[0102] d. HLA typing module: If there is a clear typing, S4 is 1, otherwise 0;

[0103] e. Immunogenicity potential assessment module: -log calculated by new antigens 10 (Rank-corrected value)*The number of combined HLAs, i.e., S5, is used as a score measurement factor.

[0104] 3. Calculate the total score based on the weight of each module.

[0105] The immunogenicity assessment modules for mining gene fusion events and neoantigens account for 30% (a, e) respectively, the protein prediction module accounts for 20% (c), and the more mature modules for assessing the coding potential of fusion sequences and patient HLA typing account for 10% respectively (b, d), with a total weight of 100%.

[0106] S=a*S1+b*S2+c*S3+d*S4+e*S5( Figure 3 ).

[0107] Comparative Example 1

[0108] For a pair of colorectal cancer patients' tumor tissues that had been surgically resected and adjacent control tissues that were >5 cm away from the tumor tissues in Example 1, whole exome sequencing with a sequencing depth of >200X was performed using second-generation sequencing technology.

[0109] 1. Use the GATK4 analysis process to screen specific SNVs and Indels in tumor tissues.

[0110] 2. Using the method in step 4 of Example 1, obtain the amino acid sequence encoded by the mutation information.

[0111] 3. Evaluate the number and immunogenic potential of candidate antigens using the method of step 6 in Example 1. Using the results of step 2 and the patient HLA typing in step 5 of Example 1 as input, obtain candidate sequences for the traditional peptide library set.

[0112] By comparison, it was found that Example 1 can produce 6 times more neoantigens and 11 times more tumor-specific neoantigens than Comparative Example 1. Using the calculation method in Example 2, the S value corresponding to each antigen was calculated, and it was found that the score of Example 1 was better than that of Comparative Example 1, indicating that the immunogenic potential of the neoantigens screened by the present invention is better than that of the traditional method ( Figure 4 ).

[0113] The top six neoantigens from Example 1 were then screened and compared with the top six neoantigens from this comparative example using a humanized mouse model. The IFN-γ enzyme-linked immunospot assay (ELISPOT) was used to verify that the top-ranked neoantigens from the present invention were able to elicit an immune response in humanized mice and secreted more IFN-γ cytokines than those in the comparative example. The specific steps are as follows:

[0114] 1. The top six neoantigens of the present invention were synthesized and mixed in equal mass ratios to obtain peptide library 1; the top six neoantigens in comparative example 1 were synthesized and mixed in equal mass ratios to obtain peptide library 2. The peptide synthesis was entrusted to a manufacturer with GMP qualifications.

[0115] 2. Six 8-week-old humanized HLA-A2 mice were randomly divided into two groups, with 3 mice in each group. After one week of adaptation, they were divided into the peptide library group 1 of the present invention and the traditional peptide library group 2. Each time, 50 μL of the peptide was mixed with an equal volume of CTL cell adjuvant and injected subcutaneously in the left and right groin for immunization, 50 μL each, with a total dose of 100 μL / mouse, once a week for three weeks. Seven days after the third immunization, the mouse spleen was removed and the mouse lymphocyte suspension was prepared for ELISPOT detection ( Figure 5 and Figure 6 ).

[0116] In the ELISPOT test results, the peptides that showed positive results for IFN-γ were identified as positive candidate peptides. Mouse lymphocytes were diluted to a concentration of 2.5*10 6 / mL, 100μL per well was added to an ELISPOT plate pre-coated with mouse IFN-γ. Peptide library group 1 of the present invention and traditional peptide library group 2 were used as stimuli, and 100μL of the corresponding polypeptide was added respectively (the final concentration of the polypeptide was 50μg / mL). No stimulant was added to the blank control group, and the cells were incubated at 37°C for 24h. Color development was performed according to the instructions of the IFN-γ ELISPOT kit, and the number of spots produced was read using an enzyme-linked spot analyzer. A positive IFN-γ result indicates the production of antigen-specific T cells, which is considered to be the ability of the polypeptide to induce an immune response in the body. The number of spots reflects the strength of its immunity.

[0117] Comparative Example 2

[0118] Using next-generation sequencing technology, we sequenced >15GB of transcriptome data from a pair of surgically resected colorectal cancer tumor tissues and adjacent adjacent control tissues located >5cm from the tumor tissues, as described in Example 1. After aligning the data to a reference genome, we used FusionCatacher, Star_Fusion, and SOAPFuse software to predict fusion gene events. We also counted fusion gene sequences with ≥5 sequences, demonstrating a five-fold improvement in detection efficiency compared to the proposed method.

[0119] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for identifying tumor fusion gene neoantigens, characterized in that: The steps include: Step 1: Obtain the patient's tumor tissue and adjacent normal control tissue, extract RNA, construct nanopore sequencing libraries and sequence them, and obtain RNA sequencing data of the patient's tumor tissue and adjacent normal control tissue; Step 2: Perform genome sequencing on the RNA sequencing data to obtain alignment data; Step 3: Analyze the comparison data to determine the gene fusion event, take the fusion gene sequence with a sequence support number of ≥5 as the candidate sequence, obtain the tumor-specific fusion gene sequence, and then use the candidate sequence analysis to obtain the new antigen; Each new antigen is scored according to the scoring method shown in the following five modules to obtain the scores of each candidate sequence S1 to S5: a. Tumor-specific fusion gene module: The score of the new antigen S1 = the number of candidate sequences supporting the antigen / sequencing depth; b. Coding potential assessment module: The result score of the original candidate sequence CPPred corresponding to the new antigen is S2; c. Protein prediction module: For the original candidate sequence corresponding to the new antigen, the result of Orfipy calculation is used. If there is a corresponding sequence, S3 is 1, otherwise S3 is 0; d. HLA typing module: For the tumor RNA sequencing data in step 1, if there is a clear typing, S4 is 1, otherwise S4 is 0; e. Immunogenicity potential assessment module: Using protein prediction results and HLA typing as input, the new antigen score S5 is calculated by netMHCpan (v4.1): S5=-log 10 (Rank-corrected value)*Number of bound HLA; Step 4: According to the weight of each module, the total score S of each neoantigen is calculated according to the following formula, and the top 100 candidate sequences with the highest S scores are determined as tumor neoantigen sequences; Total score S = a*S1+b*S2+c*S3+d*S4+e*S5; Among them, points a to e are the weights of the tumor-specific fusion gene module, coding potential assessment module, protein prediction module, HLA typing module and immunogenicity potential assessment module, respectively.

2. The identification method according to claim 1, characterized in that: In step 1, the data volume of the RNA sequencing is not less than 15G; the adjacent cancer control tissue is a tissue that is >5cm away from the tumor tissue.

3. The identification method according to claim 1, characterized in that: The step 2 comprises: using minimap2 software to align the RNA sequencing data to the genome; using samtools software to sort and construct indexes on the aligned data files to obtain aligned data files.

4. The identification method according to claim 1, characterized in that: In step 3, LongGF software is used to identify gene fusion events in tumor tissues and adjacent control tissues in the comparison data, respectively, to determine tumor-specific fusion gene sequences; wherein, the screening conditions for the gene fusion events are: the minimum alignment length with gene overlap is 100, the bin interval on the genome is 50, the minimum alignment length with the reference genome is 100, and pseudogenes and sequence secondary alignments are considered, and the results are selected using the keyword "SumGF".

5. The identification method according to claim 1, characterized in that: The coding potential evaluation was performed using CPPred software, and the evaluation indicators included: ORF length, ORF coverage, ORF integrity, Fickett score, Hexamer score, PI, Gravy, instability index and CTD characteristics.

6. The identification method according to claim 1, characterized in that: The HLA typing prediction specifically includes: based on the long sequence fastq data of adjacent cancer control tissues, using SpecHLA software to perform fine internal correction and local assembly to achieve accurate HLA typing of patients.

7. The identification method according to claim 1, characterized in that: a and e are 30% respectively, c is 20%, b and d are 10% respectively, and the sum of a to e is 100%.

8. A tumor fusion gene neoantigen recognition system based on the recognition method according to any one of claims 1 to 7, characterized in that: include: Module determination unit, weight design unit, score evaluation unit and system architecture; The module determination unit includes a gene fusion module, a coding potential assessment module, a coding protein prediction module, an HLA typing prediction module and an immunogenicity potential assessment module.

9. An electronic device, characterized in that: including a processor and a memory; The memory stores computer program instructions, and the processor is used to execute the computer program instructions to implement the method for identifying tumor fusion gene neoantigens as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which is called by a processor to implement the method for identifying tumor fusion gene neoantigens according to any one of claims 1 to 7.

11. A new antigen peptide for colorectal cancer, characterized in that: It comprises at least one of the antigenic peptides with amino acid sequences as shown in SEQ ID NOs: 1-21.

12. Biomaterial, characterized in that Includes any of the following:

1. A nucleic acid encoding the colorectal cancer neoantigen peptide according to claim 1; II, an expression cassette comprising the nucleic acid described in I; III. A recombinant vector comprising the nucleic acid described in I or the expression cassette described in II; IV. Transfect or transform the host of the recombinant vector described in III.

13. Use of the colorectal cancer neoantigen peptide according to claim 11 or the biomaterial according to claim 12 in any of the following: i. Preparation of colorectal cancer diagnostic reagents; ii. Drugs for the prevention and treatment of colorectal cancer; The drugs include protein vaccines, DNA vaccines, RNA vaccines or car-T products.

Citation Information

Patent Citations

  • Tumor neoantigen recognition method and system based on next-generation sequencing data

    CN118351934A