Method and gene combination for detecting tumor mutation load

By acquiring and filtering the somatic cell mutation data set and calculating the number of bases in the gene combination coding region, the problem of high cost and insufficient accuracy of detecting tumor mutations in the prior art is solved, and the detection effect of lower cost and higher accuracy is achieved.

CN120060468APending Publication Date: 2025-05-30SHENGWEI DATA INTELLIGENCE (CHENGDU) GENE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311630594.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is costly and inadequately accurate when detecting tumor mutation burden, making it difficult to effectively screen patients who respond to PD-1/PD-L1 and CTLA-4 immune checkpoint inhibitors.

Method used

By obtaining somatic mutation data sets, filtering site information, calculating the number of bases in the coding region of the gene combination, and calculating tumor mutation burden using specific formulas, reducing detection costs and improving accuracy.

Benefits of technology

It realizes lower-cost tumor mutation load detection, while improving the accuracy of the detection, making the results more in line with clinical needs and better assisting doctors in choosing treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120060468A_ABST
    Figure CN120060468A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of biological information, particularly relates to a method and a device for detecting a tumor mutation load and application thereof, and more particularly relates to a method for detecting the tumor mutation load by combining a gene combination. The invention provides a method for detecting tumor mutation load. The method comprises the following steps: S1, acquiring a somatic mutation data set; s2, filtering site information of the somatic mutation data set; s3, calculating the number of bases contained in the coding region of the related gene combination; and S4, calculating the tumor mutation load. By using the method provided by the invention, high-cost WES sequencing is not needed, and better prediction accuracy can be obtained, so that the result is more accurate and better conforms to a clinical result, and a doctor can be better assisted in medication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics. Specifically, it relates to a method, device and application for detecting tumor mutation burden, and more specifically, to a method for detecting tumor mutation burden by combining gene combinations. Background Art

[0002] Immunotherapy using immune checkpoint inhibitors (ICIs) has profoundly changed cancer treatment, enabling some cancer patients to achieve clinical efficacy far beyond expectations. ICIs enhance the activity of anti-tumor T cells by inhibiting immune checkpoint molecules such as PD-1 / PD-L1 and CTLA-4, which negatively regulate T cell activation and contribute to tumor immune response escape. In a clinical trial using nivolumab to treat non-small cell lung cancer, the 5-year survival rate reached 16%, while standard chemotherapy only reached 4%. However, the response rate of patients treated with PD-1 / PD-L1 immunotherapy is relatively low, remaining at 15 - 40%.

[0003] How to screen for populations that benefit from ICIs targeting PD-1 / PD-L1 and CTLA-4 and reduce the burden on patients is particularly important. There is an urgent need for precise biomarker molecules to screen for populations that benefit from ICIs. Detecting the expression of PD-L1 protein in tumor cells by immunohistochemistry (IHC) as a diagnostic biomarker is the first companion diagnostic test approved by the FDA for screening populations that benefit from PD-1 / PD-L1. However, the current detection of PD-L1 presents a one-drug-one-detection model, which seriously affects the application of companion diagnosis in pathological diagnosis work, is difficult to standardize, and has a high false positive rate, and is not an ideal biomarker.

[0004] Currently, biomarker molecules related to genomic stability are highly correlated with immunotherapy in multiple cancer types, including tumor mutation burden, neoantigen load, DNA mismatch repair deficiency, and microsatellite instability. Among the above biomarkers, tumor mutation burden is a robust, effective, and clinically verifiable biomarker, which has been verified in multiple cancer types. In non-small cell lung cancer, multiple studies have shown a correlation between patient efficacy and tumor mutation burden. Corresponding evaluations have also been carried out in other cancer types such as melanoma, head and neck squamous cell carcinoma, small cell lung cancer, and urothelial carcinoma. Multiple large-scale clinical studies have found that there are significant differences in the efficacy of patients with high tumor mutation burden and low tumor mutation burden when receiving ICIs treatment. Therefore, accurate calculation of tumor mutation burden is a very critical screening factor for evaluating whether tumor patients are suitable for ICIs treatment.

[0005] Typically, based on NGS high-throughput sequencing technology, the tumor mutation burden is calculated through whole-exome sequencing (WES) or related panels. The gold standard for calculating the tumor mutation burden is to statistically calculate the number of somatic mutations of base substitutions, insertions, and deletions in the protein-coding region sequence through WES technology. However:

[0006] 1) The exome detection is expensive and has a low detection depth, and there may be missed detections for low-coverage sites.

[0007] 2) Although the panel detection only includes hundreds of genes, it has a high sequencing depth compared to WES, so it can detect low-frequency somatic mutations. However, when calculating the tumor mutation burden, accuracy is a major challenge, with characteristics such as insufficient consistency with exome sequencing and poor targeting for samples with different tumor proportions.

[0008] Therefore, there is a need in the art for a method for calculating the tumor mutation burden that has a low cost and better accuracy at the same time. Summary of the Invention

[0009] In view of this, in a first aspect, the present invention provides a method for detecting the tumor mutation burden, including the following steps:

[0010] S1. Obtain a somatic mutation data set;

[0011] S2. Filter the site information of the somatic mutation data set;

[0012] S3. Calculate the number of bases contained in the coding regions of the gene combinations involved; and

[0013] S4. Calculate the tumor mutation burden.

[0014] Using the method of the present invention, there is no need to perform high-cost WES sequencing, and better prediction accuracy can be obtained, making the results more accurate, more in line with clinical results, and better assisting doctors in drug use.

[0015] Further, the somatic mutation data set includes non-synonymous and synonymous single nucleotide mutations in the coding regions of gene combinations, and insertion and deletion mutations shorter than 20 bases.

[0016] Still further, the insertion and deletion mutations include frameshift mutations and non-frameshift mutations.

[0017] Further, the filtering of the site information of the somatic mutation data set includes removing variant sites with a mutation frequency less than 5%, or a sample sequencing depth less than 20, or a variant base support number less than 4.

[0018] Furthermore, the site information of the filtered somatic mutation dataset includes filtering of driver mutations involved in different cancer types and pan-cancer driver mutations. Specifically, using these filtering methods will make the results more accurate.

[0019] Further, the gene combination can be the CDx (F1CDx) gene combination, or the gene combination shown in Table 1.

[0020] Table 1

[0021]

[0022]

[0023]

[0024] Preferably, the gene combination is the gene combination described in Table 1. Using the gene combination described in Table 1 can obtain better prediction accuracy.

[0025] Further, the formula for calculating the tumor mutation burden is: Tumor Mutation Burden (TMB) = number of somatic mutations / number of bases in the coding region of the gene combination × 1,000,000.

[0026] In a second aspect, the present invention provides a device for detecting tumor mutation burden, comprising:

[0027] S1, a module for obtaining a somatic mutation dataset;

[0028] S2, a module for filtering the site information of the somatic mutation dataset;

[0029] S3, a module for calculating the number of bases included in the coding region of the gene combination involved; and

[0030] S4, a module for calculating the tumor mutation burden.

[0031] Further, the somatic mutation dataset includes non-synonymous and synonymous single nucleotide mutations in the coding region of the gene combination, and insertion-deletion mutations shorter than 20 bases.

[0032] Furthermore, the insertion-deletion mutations include frameshift mutations and non-frameshift mutations.

[0033] Further, the site information of the filtered somatic mutation dataset includes removing mutation sites with a mutation frequency less than 5%, or a sample sequencing depth less than 20, or a variant base support number less than 4.

[0034] Furthermore, the site information of the filtered somatic mutation dataset includes filtering of driver mutations involved in different cancer types and pan-cancer driver mutations.

[0035] Further, the gene combination may be the CDx (F1CDx) gene combination or the gene combination shown in Table 1.

[0036] Preferably, the gene combination is the gene combination described in Table 1. Using the gene combination described in Table 1, better prediction accuracy can be obtained.

[0037] Further, the formula for calculating the tumor mutation burden is: Tumor Mutation Burden (TMB) = number of somatic mutations / number of bases in the coding region of the gene combination × 1,000,000.

[0038] In a third aspect, the present invention provides an application of the above-described analysis method or device in the preparation of a kit or device for detecting tumor mutation burden.

[0039] In a fourth aspect, the present invention provides a device, comprising:

[0040] at least one processor; and

[0041] a memory communicatively connected to at least one of the processors; wherein,

[0042] the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for detecting tumor mutation burden described in any one of the above.

[0043] In some embodiments, the device further comprises at least one input device and at least one output device; in the device, the processor, the memory, the input device, and the output device are connected by a bus.

[0044] In a fifth aspect, there is provided a storage medium storing computer instructions for being executed by a computer to implement the method for detecting tumor mutation burden described in any one of the above.

[0045] In some embodiments, the storage medium is a computer-readable storage medium.

[0046] In a sixth aspect, the present invention provides a kit, comprising:

[0047] sample nucleic acid extraction reagents and gene sequencing reagents; and

[0048] the above-described device or equipment or storage medium.

[0049] In a seventh aspect, the present invention provides a gene combination for assisting in calculating tumor mutation burden, comprising: the gene combination shown in Table 1. Description of the Drawings

[0050] Figure 1 Schematic example of the calculation method of the present invention;

[0051] Figure 2 Based on the F1CDx gene panel, correlation analysis of TMB and WESTMB calculated by calculation method 1;

[0052] Figure 3 Based on the F1CDx gene panel, correlation analysis of TMB and WESTMB calculated by calculation method 2;

[0053] Figure 4 Based on the F1CDx gene panel, correlation analysis of TMB and WESTMB calculated by the TMB calculation method of the present invention;

[0054] Figure 5 Based on the gene panel of the present invention, correlation analysis of TMB and WESTMB calculated by the TMB calculation method of the present invention.

[0055] Figure 6 Prognosis difference analysis for patients treated with ICIs based on the TMB grouping calculated by the gene panel and calculation method of the present invention. Detailed implementation manners

[0056] The present invention will be specifically described below in combination with specific implementation manners and examples, and the advantages and various effects of the present invention will be presented more clearly therefrom. Those skilled in the art should understand that these specific implementation manners and examples are for illustrating the present invention rather than limiting the present invention.

[0057] The schematic example of the calculation method described in the present invention is as Figure 1 shown.

[0058] Terms involved in the present invention:

[0059] Tumor mutation burden (TMB): For tumor samples, the total number of somatic mutations of base substitutions, insertions and deletions per million bases in the exon coding region of the genome.

[0060] Single nucleotide variant (SNV): Substitution of a single nucleotide base pair in a DNA fragment.

[0061] Insertion and deletion mutation (indel): Insertion / deletion of several nucleotides in a DNA fragment. If located in the coding region, it will cause an increase / decrease in the encoded amino acids, resulting in frameshift mutation or non-frameshift mutation.

[0062] Synonymous mutation: A single nucleotide mutation that occurs in the gene coding region and does not change the encoded amino acid.

[0063] Nonsynonymous mutation: A single nucleotide mutation that occurs in the gene coding region and can change the encoded amino acid.

[0064] Driver mutation: A mutation that provides a growth advantage, promotes tumor development, and includes activating mutations in proto-oncogenes and inactivating mutations in tumor suppressor genes.

[0065] Example 1. The calculation method of the present invention calculates the F1CDx panel.

[0066] In this example, the TCGA WES public dataset was selected and compared with two TMB calculation methods for its gene combination against the CDx (F1CDx) panel approved by the FDA, and with the TMB calculation method in the present invention. In this example, the TCGA WES public dataset was selected and compared with two TMB calculation methods for its gene combination against the CDx (F1CDx) panel approved by the FDA, and with the TMB calculation method in the present invention.

[0067] (1) Selection of somatic mutation dataset

[0068] The WES somatic mutation dataset of 10,967 samples from 32 cancer types of TCGA (The Cancer Genome Atlas) was selected.

[0069] For the above somatic mutation dataset, obtain The CDx (F1CDx) panel involves somatic mutations in the coding regions of 324 gene combinations, including only non-synonymous and synonymous single nucleotide mutations in the coding regions of these gene combinations, and insertion and deletion mutations shorter than 20 bases.

[0070] (2) Filtering of somatic mutation sites

[0071] Filter out mutation sites with a mutation frequency lower than 5%, or a sample sequencing depth less than 20, or a reads support number of variant bases less than 4. Examples of filtered sites are shown in Table 2.

[0072] Table 2

[0073]

[0074]

[0075] Note: HGVS represents the Human Genome Variation Society, which has established a set of gene mutation naming methods. HGVSc represents mutations in the coding region of a specific transcript; HGVSp represents amino acid changes, and amino acid names are represented by three letters; HGVSp_Short represents amino acid changes, and amino acid names are represented by a single letter.

[0076] On this basis, driver mutations involved in different cancer types and pan-cancer driver mutations are filtered. An example of the filtered sites is shown in Table 3.

[0077] Table 3

[0078]

[0079]

[0080] (3) Tumor mutation burden calculation

[0081] In this embodiment, 2 TMB calculation methods are benchmarked. Together with the TMB calculation method in the present invention, there are a total of 3 methods. Combining with the TCGA WES mutation dataset, for the 324 genes involved in the CDx (F1CDx) panel, mutations in their coding regions are selected for TMB calculation.

[0082] Benchmark TMB calculation method 1: The number of non-synonymous mutations and somatic mutations of insertions and deletions shorter than 20 bases in the coding regions of the 324 gene combinations is counted. Then, it is divided by the number of bases in the coding regions involved in the 324 gene combinations, and then multiplied by 1,000,000.

[0083] Benchmark TMB calculation method 2: The number of non-synonymous mutations and somatic mutations of frameshift mutations shorter than 20bp in the coding regions of the 324 gene combinations is counted. Then, it is divided by the number of bases in the coding regions involved in the 324 gene combinations, and then multiplied by 1,000,000. Specific example results of some samples are shown in Table 4.

[0084] Table 4

[0085]

[0086] Taking WESTMB in TCGA as the gold standard, a correlation comparison is made with benchmark method 1, benchmark method 2 and the TMB calculation method in the present invention. The specific results are shown in Figures 2 to 4 , the R-square (R2) of the TMB obtained by the TMB calculation method in the present invention and WESTMB is 0.9760, while the R-square (R2) of benchmark method 1 is 0.9608, and the R-square (R2) of benchmark method 2 is 0.9569. The TMB calculation method in the present invention shows the best performance.

[0087] Example 2: Calculating TMB Based on an Optimized Gene Panel

[0088] The WES somatic mutation dataset of 10,967 samples from 32 cancer types in TCGA (The Cancer Genome Atlas) was selected.

[0089] For the above somatic mutation dataset, somatic mutations in the coding regions of the 668 gene panels (Table 1) involved in the present invention were obtained, including only non-synonymous and synonymous single nucleotide mutations in the coding regions of these gene panels, as well as insertion and deletion mutations shorter than 20 bases.

[0090] Filtering of Somatic Mutation Sites

[0091] Mutation sites with a mutation frequency lower than 5%, or a sample sequencing depth less than 20, or a reads support number of variant bases less than 4 were filtered out. Examples of filtered sites are shown in Table 5.

[0092] Table 5

[0093]

[0094]

[0095] On this basis, driver mutations specific to different cancer types and pan-cancer driver mutations were filtered. Examples of filtered sites are shown in Table 6.

[0096] Table 6

[0097]

[0098]

[0099] Based on the tumor mutation burden calculation method in the present invention and the somatic mutations in the coding regions of the 668 gene panels used to calculate the tumor mutation burden, the tumor mutation burden was calculated. Specific example results of the tumor mutation burden in this example are shown in Table 7.

[0100] Table 7

[0101]

[0102] Taking the WESTMB in TCGA as the gold standard, a correlation analysis was performed between the TMB calculated based on the TMB calculation method of the present invention and the 668 gene panels. The specific results are shown in Figure 5 , with an R-square (R2) of 0.9864, higher than 0.9760 of the gene panel of CDx (F1CDx).

[0103] Example 3. Prognostic differences in patients treated with ICIs based on the TMB grouping of the present invention

[0104] Twenty-one patients with TCGA melanoma who received ICIs treatment and the corresponding WES mutation datasets were selected, and the clinical information is shown in Table 8.

[0105] Table 8

[0106]

[0107] Note: Ipilimumab, an immune checkpoint inhibitor, is a fully human monoclonal antibody of the IgG1 subtype that targets and downregulates the protein receptor CTLA-4 of the immune system to activate the immune system; Pembrolizumab, an immune checkpoint inhibitor, is a humanized anti-PD-1 monoclonal antibody that relieves the immune suppression mediated by the PD-1 pathway.

[0108] For the above somatic mutation datasets, somatic mutations in the coding regions of 668 gene combinations (Table 1) involved in the present invention were obtained, including only non-synonymous and synonymous single nucleotide mutations in the coding regions of these gene combinations, as well as insertion and deletion mutations shorter than 20 bases.

[0109] Filtration of somatic mutation sites

[0110] Mutation sites with a mutation frequency lower than 5%, or a sample sequencing depth less than 20, or a reads support number of variant bases less than 4 were filtered out. Examples of filtered sites are shown in Table 9.

[0111] Table 9

[0112]

[0113] On this basis, driver mutations involved in different cancer types and pan-cancer driver mutations were filtered. Examples of filtered sites are shown in Table 10.

[0114] Table 10

[0115]

[0116]

[0117] Based on the tumor mutation burden calculation method in the present invention and the somatic mutations in the coding regions of 668 gene combinations used for calculating the tumor mutation burden, the tumor mutation burden was calculated. The results of the tumor mutation burden in this example are shown in Table 11.

[0118] Table 11

[0119]

[0120] Taking the current general TMB threshold of 10 as the cut-off point, dividing TMB≥10 into TMB-H and TMB<10 into TMB-L, and combining the overall survival (OS) as the prognostic indicator, the prognostic differences between the two groups of patients were compared. The specific results are shown in Figure 6 . Compared with TMB-L, the hazard ratio for TMB-H was HR = 0.1723, the 95% confidence interval was CI = 0.0414 - 0.7164, and the significance was P = 0.0072. The median survival time for TMB-H was 152.8093 months, and the median survival time for TMB-L was 36.2626 months. The results indicate that SKCM patients in the TMB-H group had a longer overall survival after receiving ICIs treatment.

[0121] The above three examples demonstrate the device of the present invention from different perspectives, and it can objectively evaluate TMB, having certain clinical application value. Example 1 demonstrates that the TMB calculation method of the present invention can more accurately evaluate TMB; Example 2 demonstrates that based on the TMB calculation method of the present invention, using 668 gene combinations to evaluate TMB is a more preferred gene combination; Example 3 is for 21 SKCM patients treated with ICIs. Based on the TMB evaluation obtained from the TMB calculation method and 668 gene combinations in the present invention, the group with TMB-H had a longer overall survival.

[0122] The above embodiments are only the preferred embodiments of the present invention. It should be noted that: for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and equivalent replacements can be made. These technical solutions obtained by improving and equivalently replacing the claims of the present invention all fall within the protection scope of the present invention.

Claims

1. A method for detecting tumor mutation burden, comprising the following steps: S1. Obtain a somatic mutation data set; S2. Filter the locus information of the somatic mutation data set; S3. Calculate the number of bases contained in the coding regions of gene combinations; and S4. Calculate the tumor mutation burden.

2. The method according to claim 1, wherein, the somatic mutation data set includes non-synonymous and synonymous single nucleotide mutations in the coding regions of gene combinations, and insertion and deletion mutations shorter than 20 bases.

3. The method according to claim 2, wherein, the insertion and deletion mutations include frameshift mutations and non-frameshift mutations.

4. The method according to claim 1, wherein, filtering the locus information of the somatic mutation data set includes removing variant loci with a mutation frequency less than 5%, or a sample sequencing depth less than 20, or a variant base support number less than 4.

5. The method according to claim 2, wherein, The gene combination is the CDx gene combination, or the gene combination as shown in Table 1.

6. A device for detecting tumor mutation burden, comprising: S1. A module for obtaining a somatic mutation data set; S2. A module for filtering the locus information of the somatic mutation data set; S3. A module for calculating the number of bases contained in the coding regions of gene combinations; and S4. A module for calculating the tumor mutation burden.

7. Use of a method or device according to any one of claims 1 to 6 in the preparation of a kit or device for detecting tumor mutation burden.

8. A device, comprising: at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for detecting tumor mutation burden according to any one of claims 1 to 5.

9. A storage medium storing computer instructions for being executed by a computer to implement the method for detecting tumor mutation burden according to any one of claims 1 to 5.

10. A gene combination for assisting in calculating tumor mutation burden, comprising: the gene combination shown in Table 1.

Citation Information

Cited By

  • Detection reagent for tumor mutation load as well as related product and application thereof

    CN120989244A