Identification method and application of liver cancer neoantigen derived from transposon gene chimeric transcript

Through the joint analysis of whole-transcriptionome sequencing and mass spectrometry data, liver cancer neoantigens from transposon-gene chimeric transcripts were screened, solving the problems of large individual differences and low coverage of new antigens in the prior art, and achieving the efficiency and accuracy of personalized immunotherapy.

CN120126563APending Publication Date: 2025-06-10FUJIAN HAIXI CELL BIOENGINEERING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510075660.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing neoantigens of liver cancer mainly come from somatic mutations, and there are problems of large individual differences and low coverage, which leads to the cumbersome development process of tumor neogenic antigen vaccines and is difficult to achieve large-scale clinical promotion.

Method used

By obtaining the whole transcriptome sequencing data of tumor tissue and paired adjacent tissue of primary liver cancer patients, transposon-gene chimeric transcripts were identified and screened, their translation potential was predicted, and protein mass spectrometry data was verified to finally screen out the potential neoantigen sequence.

Benefits of technology

Personalized identification and screening of liver cancer neoantigens was achieved, and a wide coverage of neoantigens was constructed, which significantly improved the efficiency and accuracy of individualized immunotherapy, and solved the problems of large differences in individual neoantigens and low coverage in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126563A_ABST
    Figure CN120126563A_ABST
Patent Text Reader

Abstract

The invention relates to an identification method and application of a group of transposon gene chimeric transcript-derived liver cancer neoantigens, and belongs to the field of liver cancer antigen identification. The method is used for identifying and screening new antigens from transposon-gene chimeric transcripts based on transcriptome sequencing data of cancer tissues of liver cancer patients and control tissues. Through multi-center liver cancer transcriptome data set verification, mass spectrum verification and tissue toxicity filtration, 17 new antigens derived from transposon-gene chimeric transcripts are finally screened out, and the constructed new antigen pannel can quickly formulate a new antigen vaccine scheme for liver cancer patients, so that the efficiency and accuracy of individualized immunotherapy are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of identification of liver cancer antigens, and particularly relates to a method for identifying a novel liver cancer antigen derived from a group of chimeric transcripts of transposon genes and its application. Background Art

[0002] Liver cancer is one of the malignant tumors with relatively high incidence and mortality rates globally. Currently, the main clinical treatment methods for liver cancer include surgical treatment, interventional treatment, radiotherapy, chemotherapy, and targeted drug treatment. Although certain progress has been made in recent years, the treatment of liver cancer still faces bottleneck problems such as low long-term survival rate and high recurrence and metastasis rates. As the most important detoxifying organ in the human body, the liver is constantly exposed to and processes foreign antigenic substances, forming a unique immune tolerance microenvironment. This characteristic makes liver cancer cells prone to escaping the surveillance of the immune system, further exacerbating the complexity of treatment.

[0003] Immunotherapy shows important prospects in the treatment of liver cancer by reactivating the patient's own immune cells to attack tumor cells. Among them, immune checkpoint blockers represented by anti-PD-L1 and PD-1 antibodies have achieved remarkable results. However, the effectiveness of such therapies depends on the effective recognition of tumor cells by immune cells. If immune cells cannot recognize tumor cells, even if the immunosuppressive microenvironment is relieved, it is difficult to achieve effective treatment, resulting in many patients failing to respond to this therapy.

[0004] Tumor antigens are key molecules for immune cells to recognize and kill tumor cells. Among them, tumor neoantigens, as a class of specific antigens only expressed in tumor cells, have the advantages of strong targeting, low toxicity and side effects, low probability of immune tolerance, and support for individualized precision treatment. In recent years, a number of clinical trials based on tumor neoantigens have demonstrated their safety and effectiveness in the treatment and recurrence prevention of liver cancer. However, currently, the neoantigens used in clinical applications mainly originate from somatic mutations, which have limitations such as large individual differences and low coverage. This leads to a cumbersome R & D process for tumor neoantigen vaccines, including high-throughput sequencing of tumor tissues, bioinformatics analysis, and vaccine production, etc., with a long cycle and high cost, making it difficult to achieve large-scale clinical promotion. In addition, for patients with low mutation burden, the number of potential neoantigen sites derived from somatic mutations is limited, making it difficult to provide sufficient treatment targets.

[0005] Transposons account for nearly 50% of the human genome, far higher than the proportion of known protein-coding sequences. Recent studies have shown that in tumor tissues, due to abnormal regulation of epigenetic modifications, transposons may be reactivated and act as new promoters to regulate the expression of themselves and downstream genes, forming transposon-gene chimeric transcripts. The replacement of the original promoter by such transposons as new promoters may lead to changes in the open reading frame: on the one hand, the transposons themselves may be translated to generate new proteins with additional amino acids, thus producing tumor-specific neoantigens; on the other hand, the change of the promoter may trigger the frameshift of chimeric genes, thereby generating completely new protein sequences, which become potential neoantigens. Compared with point mutations, these neoantigens have stronger immunogenicity (involving more amino acid sequence changes) and a wider coverage range, so they are expected to be an ideal source of broad-spectrum neoantigens.

[0006] However, there have been no reports on the study of neoantigens derived from transposon-gene chimeric transcripts in liver cancer. Whether there are high-frequency chimeric transcripts that can serve as a source of neoantigens in the Asian population remains an unknown area. Summary of the Invention

[0007] The purpose of the present invention is to solve the limitations of liver cancer neoantigens derived from classical somatic mutations, and to provide a method for identifying and applying liver cancer neoantigens derived from a group of transposon gene chimeric transcripts.

[0008] To achieve the above object, the technical solution of the present invention is: a method for identifying liver cancer neoantigens derived from a group of transposon gene chimeric transcripts, including:

[0009] (1) Obtain the whole-transcriptome sequencing data of tumor tissues and paired adjacent tissues of patients with primary liver cancer. After the quality control of the original data, identify and quantitatively analyze the transposon-gene chimeric transcripts within the whole genome.

[0010] (2) According to the expression profiles of transposon-gene chimeric transcripts in liver cancer tissues and adjacent tissues, screen out the transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues.

[0011] (3) Predict the translation potential of transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues, and obtain the transcript sequences with translation potential and their corresponding potential protein sequences.

[0012] (4) Select the transposon-gene chimeric transcripts that are specifically highly expressed and specifically expressed in liver cancer with translation potential, and verify them based on protein mass spectrometry data to further obtain protein mass spectrometry data to support the target transcripts.

[0013] (5) Combine the transcriptome and mass spectrometry ligandomics data of 32 normal tissues in the GTEx (Genotype-Tissue Expression) database, and adopt a hierarchical filtering strategy based on tissue-specific expression. Set the tolerance level according to the importance and safety threshold of different tissues, and filter out chimeric transcripts and their derived antigens that may cause toxicity to ensure antigen safety;

[0014] (6) Obtain the original transcriptome sequencing data of liver cancer patients and control normal tissues from public datasets, and further identify and quantitatively verify the selected target transposon-gene chimeric transcripts;

[0015] (7) Compare the protein sequences generated by transposon-gene chimeric transcripts that are specific or highly expressed in liver cancer with the protein sequences corresponding to normal transcripts, and intercept the peptide segment sequences that are different from the normal gene transcripts for each transcript as the new antigen sequences of the candidate transposon-gene chimeric transcripts;

[0016] (8) Compare the protein sequences generated by the target chimeric transcripts with the protein sequences of normal gene transcripts, and intercept the peptide segment sequences in the different regions as the new antigen sequences of the candidate transposon-gene chimeric transcripts;

[0017] (9) Evaluate the immunogenicity of the new antigen sequences corresponding to the selected target transcripts, predict the number of immunogenic peptide segments in different common HLA class I genotypes, and evaluate the potential of each target transcript in immunotherapy.

[0018] In an embodiment of the present invention, in step (1), quality control processing is performed, that is, the second-generation sequencing quality control software fastp is used to perform quality control on the original transcriptome sequencing data, and a clean fastq file is obtained after removing low-quality sequences.

[0019] In an embodiment of the present invention, in step (1), the transcriptome sequencing data alignment algorithm SATR, the transcript splicing algorithm StringTie, and the TEProF2 (Transposable Element Promoter Finder 2) algorithm are used to identify and quantitatively analyze transposon-gene chimeric transcripts within the entire genome.

[0020] In an embodiment of the present invention, in step (2), the transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues are transcripts with an expression ratio lower than 1% in adjacent cancer tissues.

[0021] In one embodiment of the present invention, in step (3), the Coding Potential Calculator 2 (CPC2) algorithm is used to predict the translation potential of transposon-gene chimeric transcripts highly expressed or specifically expressed in liver cancer tissues.

[0022] In one embodiment of the present invention, in step (4), verification is performed based on the protein mass spectrometry data of two groups of liver cancer HLA ligandomics, namely PXD029882 and PXD023143.

[0023] In one embodiment of the present invention, in step (6), the original transcriptome sequencing data of 335 liver cancer patients and 152 control normal tissues from 4 public datasets, namely GSE148355, GSE214846, GSE202069, and HRA000169, is obtained, and the screened target transposon-gene chimeric transcripts are further identified and quantitatively verified.

[0024] In one embodiment of the present invention, in step (9), the new antigen immunogenicity prediction software Pvactools is used to evaluate the immunogenicity of the new antigen sequences corresponding to the screened target transcripts, predict the number of immunogenic peptide segments in different common HLA class I typings, and evaluate the potential of each target transcript in immunotherapy.

[0025] In one embodiment of the present invention, this method can be applied to the development of neoantigen-based vaccines.

[0026] The present invention also provides an application of liver cancer neoantigens derived from a group of transposon gene chimeric transcripts. Based on the above method, a neoantigen-based vaccine is developed for use in immunotherapy including liver cancer treatment and relapse prevention.

[0027] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention is used to identify and screen neoantigens derived from transposon-gene chimeric transcripts based on the transcriptome sequencing data of cancer tissues and control tissues of liver cancer patients. Through verification using multi-center liver cancer transcriptome datasets, mass spectrometry verification, and tissue toxicity filtration, 17 neoantigens derived from transposon-gene chimeric transcripts are finally screened out, and the constructed neoantigen pannel can quickly formulate a neoantigen vaccine plan for liver cancer patients, significantly improving the efficiency and accuracy of personalized immunotherapy. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Protein mass spectrometry level verification for liver cancer-specific transposon-gene chimeric transcripts.

[0029] Figure 2 Coverage of liver cancer-specific transposon-gene chimeric transcripts in each dataset.

[0030] Figure 3 The number of immune peptides corresponding to different HLA types for each liver cancer-specific transposon-gene chimeric transcript. Detailed implementation manners

[0031] The technical solution of the present invention will be specifically described below in conjunction with the accompanying drawings.

[0032] The present invention provides a method for identifying a novel antigen of liver cancer derived from a group of transposon-gene chimeric transcripts, including:

[0033] (1) Obtain the whole transcriptome sequencing data of the tumor tissues and paired adjacent tissues of patients with primary liver cancer. After the quality control of the original data, identify and quantitatively analyze the transposon-gene chimeric transcripts within the whole genome.

[0034] (2) According to the expression profiles of transposon-gene chimeric transcripts in liver cancer tissues and adjacent tissues, screen out the transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues.

[0035] (3) Predict the translation potential of the transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues, and obtain the transcript sequences with translation potential and their corresponding potential protein sequences.

[0036] (4) Select the transposon-gene chimeric transcripts that are specifically highly expressed and specifically expressed in liver cancer with translation potential, and verify them based on protein mass spectrometry data to further obtain protein mass spectrometry data to support the target transcripts.

[0037] (5) Combine the transcriptome and mass spectrometry ligandomics data of 32 normal tissues in the GTEx database (Genotype-Tissue Expression), and adopt a hierarchical filtering strategy based on tissue-specific expression. Set the tolerance level according to the importance and safety threshold of different tissues to filter out the chimeric transcripts and their derived antigens that may cause toxicity to ensure antigen safety.

[0038] (6) Obtain the original transcriptome sequencing data of liver cancer patients and control normal tissues from public datasets, and further identify and quantitatively verify the screened target transposon-gene chimeric transcripts.

[0039] (7) Compare the protein sequences generated by the transposon-gene chimeric transcripts that are specifically or highly expressed in liver cancer with the protein sequences corresponding to normal transcripts, and intercept the peptide sequences different from the normal gene transcripts for each transcript as the new antigen sequences of the candidate transposon-gene chimeric transcripts.

[0040] (8) Align the protein sequence generated from the target chimeric transcript with the protein sequence of the normal gene transcript, and intercept the peptide sequence of the differential region as the neoantigen sequence of the candidate transposon-gene chimeric transcript.

[0041] (9) Evaluate the immunogenicity of the neoantigen sequence corresponding to the selected target transcript, predict the number of immunogenic peptide segments in different common HLA class I genotypes, and evaluate the potential of each target transcript in immunotherapy.

[0042] The following is a specific application scenario of the method of the present invention, but does not constitute a limitation on the scope of the present invention.

[0043] 1. After extracting RNA from liver cancer tissue and adjacent tissue samples, perform transcriptome sequencing.

[0044] 2. Use the fastp software to perform quality control (QC) on the original transcriptome sequencing data, and obtain the clean fastq file after removing low-quality sequences.

[0045] 3. Input the clean fastq file of the normal control tissue into the HLA typing prediction software Optitype software to obtain the HLA class I typing of the patient.

[0046] 4. Use the two-step alignment method of the STAR software to align the clean fastq files of liver cancer and control tissues to the human reference genome (hg38) respectively to generate bam files.

[0047] 5. Using the aligned bam files, use the in-house python script fusion_site_detection to extract the number of reads spanning or covering the chimeric sites (junction of transposon and gene) of 17 chimeric transcripts.

[0048] 6. Using the gtf files of 17 chimeric transcripts and the gtf file of the human reference genome, use the merge result merging command cuffmerge in the transcript assembly tool cufflinks to merge the above two gtf files to generate a merged gtf reference file (merge.gtf file) for evaluating the expression level of chimeric transcripts.

[0049] 7. Based on the merge.gtf file, use the quant command in the transcript quantification algorithm Salmon to perform quantitative analysis on the expression levels of 17 chimeric transcripts in the patient's tumor tissue and adjacent tissue.

[0050] 8. Screen chimeric transcripts based on the number of reads spanning the chimeric site and the expression level. Retain chimeric transcripts with the number of reads spanning the chimeric site greater than 10, the expression level (TPM) greater than 5 in tumor tissues, and the tumor tissue expression level 10 times higher than that in adjacent tissues.

[0051] 9. Use the neoantigen immunogenicity prediction software Pvactools to predict the immunogenicity of neoantigens corresponding to the screened chimeric transcripts based on the HLA typing of the patient.

[0052] 10. Screen chimeric transcripts with immunogenicity against the patient's HLA typing, select the corresponding peptide segments from the prefabricated vaccine pannel, and prepare personalized neoantigen vaccines.

[0053] Through the combined analysis of whole-transcriptome and mass spectrometry data ( Figure 1 ), 19 tumor-specific, translationally potential and immunopeptide-generating chimeric transcript open reading frames were successfully screened out (Table 1). These 19 chimeric transcripts were verified in four external liver cancer transcriptome datasets (GSE148355, GSE212846, GSE202069, HRA000169), and finally 17 of them were confirmed (including AluSp-FIS1, HERV3-int-PRKD1, L1PA2-WDR72, MIRb-C1orf198, AluSx-DUOX1, LTR79-TMEM189-UBE2V1, L2b-KCP, MIRb-PRKCH, AluY-PRDM16, MER5A-KCNU1, L1PA2-IGSF1, AluSx-CRK, L1PA2-RABGAP1L, MIR-GIPR, Tigger3b-MAP3K9, Tigger20a-VPS8, LTR12C-TCF4, Figure 2 ), and these chimeric transcripts were replicated in at least one validation dataset. Further analysis found that 5 chimeric transcripts (AluSp-FIS1, HERV3-int-PRKD1, L1PA2-WDR72, MIRb-C1orf198, L1PA2-RABGAP1L) were identified in all datasets. These chimeric transcripts can cover approximately 50% of liver cancer patients in the exploratory dataset, while covering approximately 33.3% of patients in the validation dataset. The immunogenicity prediction results show that these chimeric transcripts can generate immunopeptides against multiple common HLA typings and have the potential to induce immune responses ( Figure 3 ). Therefore, this transposon combination has good coverage in liver cancer patients and can provide support for the development of neoantigen-based vaccines, thus realizing the application of immunotherapy such as liver cancer treatment and prevention of recurrence.

[0054] Table 1

[0055]

[0056]

[0057]

[0058] The above are the preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced shall fall within the protection scope of the present invention.

Claims

1. A method for identifying a group of liver cancer neoantigens derived from transposon gene chimeric transcripts, characterized in that: include: (1) Obtain whole transcriptome sequencing data of tumor tissues and paired adjacent paracancerous tissues from patients with primary liver cancer. After quality control of the raw data, identify and quantify transposon-gene chimeric transcripts across the genome. (2) based on the expression profiles of transposon-gene chimeric transcripts in liver cancer tissues and adjacent tissues, screening for transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues; (3) predict the translation potential of transposon-gene chimeric transcripts that are highly or specifically expressed in liver cancer tissues, and obtain transcript sequences with translation potential and their corresponding potential protein sequences; (4) selecting HCC-specific highly expressed and specifically expressed transposon-gene chimeric transcripts with translational potential, validating them based on protein mass spectrometry data, and further acquiring protein mass spectrometry data to support the target transcripts; (5) Combining the transcriptome and mass spectrometry ligandomics data of 32 normal tissues in the GTEx database, a hierarchical filtering strategy based on tissue-specific expression was adopted. According to the importance and safety threshold of different tissues, tolerance levels were set to filter out chimeric transcripts and their derived antigens that may cause toxicity, so as to ensure antigen safety; (6) Obtaining original transcriptome sequencing data from liver cancer patients and control normal tissues from public datasets, and further identifying and quantitatively verifying the selected target transposon-gene chimeric transcripts; (7) Compare the protein sequences produced by liver cancer-specific or highly expressed transposon-gene chimeric transcripts with the protein sequences corresponding to normal transcripts, and extract the peptide sequences that are different from the normal gene transcripts of each transcript as the new antigen sequences of the candidate transposon-gene chimeric transcripts; (8) comparing the protein sequence produced by the target chimeric transcript with the protein sequence of the normal gene transcript, and extracting the peptide sequence of the differential region as the new antigen sequence of the candidate transposon-gene chimeric transcript; (9) Evaluate the immunogenicity of the neoantigen sequences corresponding to the screened target transcripts, predict the number of immunogenic peptides in different common HLA class I typing, and evaluate the potential of each target transcript in immunotherapy.

2. The method for identifying a group of liver cancer neoantigens derived from transposon gene chimeric transcripts according to claim 1, characterized in that: In step (1), quality control processing is to use fastp software to perform quality control on the original transcriptome sequencing data and obtain a clean fastq file after removing low-quality sequences.

3. The method for identifying liver cancer neoantigens derived from a group of transposon gene chimeric transcripts according to claim 1, characterized in that: In step (1), TEProF2, SATR and StringTie algorithms were used to identify and quantify transposon-gene chimeric transcripts across the entire genome.

4. The method for identifying liver cancer neoantigens derived from a group of transposon gene chimeric transcripts according to claim 1, characterized in that: In step (2), the transposon-gene chimeric transcripts highly expressed or specifically expressed in liver cancer tissues are transcripts whose expression ratio in adjacent cancer tissues is less than 1%.

5. The method for identifying a group of liver cancer neoantigens derived from transposon gene chimeric transcripts according to claim 1, characterized in that: In step (3), the CPC2 algorithm is used to predict the translation potential of transposon-gene chimeric transcripts that are highly expressed or specifically expressed in liver cancer tissues.

6. The method for identifying a group of liver cancer neoantigens derived from transposon gene chimeric transcripts according to claim 1, characterized in that: In step (4), verification was performed based on protein mass spectrometry data of two groups of liver cancer HLA ligandomics, namely PXD029882 and PXD023143.

7. The method for identifying a group of liver cancer neoantigens derived from transposon gene chimeric transcripts according to claim 1, characterized in that: In step (6), the original transcriptome sequencing data of 335 liver cancer patients and 152 control normal tissues from four public datasets, namely GSE148355, GSE214846, GSE202069, and HRA000169, were obtained to further identify and quantitatively verify the selected target transposon-gene chimeric transcripts.

8. The method for identifying liver cancer neoantigens derived from a group of transposon gene chimeric transcripts according to claim 1, characterized in that: In step (9), the immunogenicity of the neoantigen sequences corresponding to the screened target transcripts is evaluated using Pvactools software, the number of immunogenic peptides in different common HLA class I typing is predicted, and the potential of each target transcript in immunotherapy is evaluated.

9. The method for identifying liver cancer neoantigens derived from a group of transposon gene chimeric transcripts according to claim 1, characterized in that: This approach can be applied to the development of vaccines based on new antigens.

10. An application of a group of liver cancer neoantigens derived from transposon gene chimeric transcripts, characterized in that: Develop a new antigen-based vaccine based on claims 1-9 for application in immunotherapy including liver cancer treatment and prevention of recurrence.