A TMB detection method based on panel data and its application

Through a TMB detection method based on panel data, including data filtering, comparison and TMB calculation steps, the problem of statistical inaccurateness when large panels detect TMB in the prior art is solved, and the accuracy and stability of TMB values ​​are achieved, providing a basis for immunotherapy.

CN114360649BActive Publication Date: 2025-05-30星云基因科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111678171.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-30
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art has the problem of statistical inaccuracy when using large panels to detect tumor mutation load (TMB).

Method used

A TMB detection method based on panel data is proposed, including data filtering, comparison and TMB calculation steps. The specific steps include using fastp for data filtering, using GATK for variation detection and statistics, annotation and filtering through ANNOVAR, and finally using a specific formula to calculate the TMB value.

Benefits of technology

This method can improve the accuracy and stability of TMB values, providing a reliable basis for immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360649B_ABST
    Figure CN114360649B_ABST
Patent Text Reader

Abstract

A TMB detection method based on panel data and its application, belonging to the technical field of bioinformatics. To solve the problem of inaccurate statistics of TMB using large panels at present, the present invention provides a TMB detection method based on panel data, which includes: somatic cell detection, false positive site filtering, TMB calculation, and finally obtaining the TMB value. Through the above several modules, the accuracy and stability of the TMB value can be improved, providing a basis for immunotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to a TMB detection method based on panel data and its application. Background Art

[0002] At present, a number of clinical studies have confirmed that the tumor mutation burden (TMB), that is, the total number of somatic mutations in the tumor genomic region per million bases (mb), has a very high correlation with immunotherapy; the clinical benefit rate of patients with high TMB tumors after receiving immune checkpoint inhibitor treatment is relatively high. At present, whole exome sequencing (WES) is the recognized gold standard for TMB detection in the industry. However, considering the economic cost, large panels are currently often used as an alternative for TMB detection. Summary of the Invention

[0003] In view of the problem that the current statistics of TMB using large panels are inaccurate, the present invention provides a TMB detection method based on panel data and its application.

[0004] A TMB detection method based on panel data includes the following steps:

[0005] S1. Data filtering: Use fastp to filter the off-machine data of the gene detection panel.

[0006] S2. Alignment: Align the filtered data with the human reference genome hg19, perform quality control on the alignment results, then perform variant detection using GATK, and finally perform statistics on the alignment results of the data.

[0007] S3. TMB calculation: Use GATK to perform somatic SNV detection on the tumor tissue, adjacent cancer tissue or white blood cells of each sample, filter according to the information of the number of reads supporting mutations at the locus, mutation frequency, and genotype quality value, use statistical methods and models to filter out sites that do not meet strand specificity, use ANNOVAR to annotate the predicted somatic mutations, filter out mutation sites caused by the population and background noise, select non-synonymous mutations, count the length of the captured region, and calculate TMB according to the following formula

[0008]

[0009] Further defined, the filtering conditions in S1 are as follows:

[0010] 1) Remove reads containing adaptors, automatically identify the adapter sequence, and perform filtering.

[0011] 2) Remove reads with a proportion of N removed greater than 10%;

[0012] 3) Remove low-quality reads, where the low-quality reads refer to reads with a quality value Q <= 5 and the number of bases accounting for more than 50% of the entire read.

[0013] Further defined, the specific steps of S2 are as follows:

[0014] 1) Use the bwa software to align the data filtered in S1 with the human reference genome hg19;

[0015] 2) Use samtools to sort the aligned bam;

[0016] 3) Use picard to remove duplicates from the sorted bam;

[0017] 4) Statistically analyze the sequencing coverage rate, alignment rate, etc.

[0018] Further defined, the filtering conditions of S3 based on the number of reads supporting mutant sites, mutation frequency, and genotype quality value are as follows:

[0019] a. Filter reads with the number of reads supporting mutants less than 10;

[0020] b. Filter sites with a mutation frequency less than 5%;

[0021] c. Filter sites with a quality value less than 20.

[0022] Further defined, the method for filtering sites that do not meet strand specificity in S3 using statistical methods and models is as follows: Filter strand-specific sites for the P-value calculated using Fisher’s Exact Test. The calculation method of the P-value is to use Fisher’s Exact Test with the number of reads covered as the input and calculate the P-value according to the following formula:

[0023]

[0024] where R 1 and R 2 are row totals, C 1 and C 2 are column totals, N is the total number of observations in the row-column table, and nij is the value in the i-th row and j-th column of the table.

[0025] Further limit, the method for filtering mutation sites caused by background noise in S3 is to first remove germline mutation sites using annotated populations, and then remove possible germline mutation sites using control sequencing data; establish a model of the background noise mutation frequency of different mutant types at each detection site using the somatic mutation site set of the control data; calculate the mutation frequency of each candidate somatic mutation site in the candidate somatic mutation site set of the test sample, and calculate the p-value of each candidate somatic mutation site and the model, and filter the background noise sites through the P-value; filter mutation sites with a population frequency greater than 0.01 and mutation sites marked as cosmic to obtain somatic SNV sites and count the number.

[0026] The present invention also provides an application of the above TMB detection method in predicting the treatment effect of tumor patients.

[0027] Advantages of the present invention:

[0028] The present invention provides a method for detecting TMB, which includes: somatic detection, false positive site filtering, TMB calculation, and finally obtaining the TMB value. Through the above several modules, the accuracy and stability of the TMB value can be improved, providing a basis for immunotherapy. Description of the drawings

[0029] Figure 1 It is a coverage result graph; where Figure 1 A in it is the base ratio at different sequencing depths, the abscissa represents the sequencing depth, and the ordinate represents the proportion of the sequencing depth bases in all bases; Figure 1 B in it is the cumulative base ratio at different depths. Detailed implementation manners

[0030] Example 1: A TMB detection method based on panel data

[0031] A TMB detection method based on panel data, the TMB detection method includes the following steps:

[0032] S1. Data filtering: Use fastp to filter the off-machine data of the gene detection panel, and the filtering conditions are as follows:

[0033] 1) Remove reads containing adaptors, automatically identify the adaptor sequences, and perform filtering;

[0034] 2) Remove reads with a proportion of N greater than 10%;

[0035] 3) Remove low-quality reads, where the low-quality reads refer to reads with a quality value Q <= 5 and the number of bases accounting for more than 50% of the entire read.

[0036] S2. Alignment: The filtered data is aligned with the human reference genome hg19, and the alignment results are quality controlled. Then, variant detection is performed using GATK, and finally, the alignment results of the data are statistically analyzed. The specific steps are as follows:

[0037] 1) Use the bwa software to align the data filtered in S1 with the human reference genome hg19;

[0038] 2) Use samtools to sort the aligned bam;

[0039] 3) Use picard to remove duplicates from the sorted bam;

[0040] 4) Statistically analyze the sequencing coverage rate, alignment rate, etc.

[0041] S3. TMB calculation: Use GATK to perform somatic SNV detection on the tumor tissue, adjacent tissue, or white blood cells of each sample. Filter according to the information of the number of reads supporting mutations at the locus, mutation frequency, and genotype quality value. Use statistical methods and models to filter out loci that do not meet strand specificity. Use ANNOVAR to annotate the predicted somatic mutations and filter out mutation loci caused by the population and background noise. Select non-synonymous mutations, calculate the length of the captured region, and calculate TMB according to the following formula:

[0042]

[0043] The conditions for filtering according to the information of the number of reads supporting mutations at the locus, mutation frequency, and genotype quality value are as follows:

[0044] a. Filter reads with the number of reads supporting mutations less than 10;

[0045] b. Filter loci with a mutation frequency less than 5%;

[0046] c. Filter loci with a quality value less than 20.

[0047] The method of using statistical methods and models to filter out loci that do not meet strand specificity is as follows: Filter the P-value calculated by Fisher's Exact Test for strand-specific loci. The calculation method of the P-value is to use Fisher's Exact Test with the number of covered reads as the input and calculate the P-value according to the following formula:

[0048]

[0049] where R 1 and R2 is the row total, C 1 and C 2 is the column total, N is the total number of observations in the row list, and nij is the value in the i-th row and j-th column of the table.

[0050] The method for filtering mutation sites caused by background noise is to first remove germline mutation sites using annotated populations, and then remove possible germline mutation sites using control sequencing data; establish a model of the background noise mutation frequency of different mutant types at each detection site using the somatic mutation site set of the control data; calculate the mutation frequency of each candidate somatic mutation site in the candidate somatic mutation site set of the sample to be tested, and calculate the p-value of the sum of each candidate somatic mutation site and the model, and filter the background noise sites through the P-value; filter out mutation sites with a population frequency greater than 0.01 and mutation sites marked as cosmic to obtain somatic SNV sites and count the number.

Claims

1. A TMB detection method based on panel data, characterized in that, the TMB detection method comprises the following steps: S1. Data filtering: Use fastp to filter the off-machine data of the gene detection panel; S2. Alignment: Align the filtered data with the human reference genome hg19, perform quality control on the alignment results, then perform variant detection using GATK, and finally perform statistics on the alignment results of the data; S3. TMB calculation: Use GATK to perform somatic SNV detection on the tumor tissue, adjacent tissue, or white blood cells of each sample. Filter according to the information of the number of reads supporting mutations at the locus, mutation frequency, and genotype quality value. Use statistical methods and models to filter out loci that do not conform to strand specificity. Use ANNOVAR to annotate the predicted somatic mutations, filter out mutation loci caused by the population and background noise, select non-synonymous mutations, count the length of the captured region, and calculate TMB according to the following formula: ; The method for filtering mutation sites caused by background noise is to first remove germline mutation sites using annotated populations, and then remove germline mutation sites using the sequencing data of the control; Use the set of somatic mutation sites in the control data to establish a model of the background noise mutation frequency of different mutant types at each detection site; Calculate the mutation frequency of each candidate somatic mutation site in the candidate somatic mutation site set of the sample to be tested, and calculate the p-value of each candidate somatic mutation site and the model, and filter the background noise sites through the P-value; Filter mutation sites with a population frequency greater than 0.01 and mutation sites marked as cosmic to obtain somatic SNV sites and count the number; The method for filtering sites that do not conform to strand specificity using statistical methods and models is as follows: Filter the P-value calculated using Fisher’s Exact Test for strand-specific sites. The calculation method of the P-value is to use Fisher’s Exact Test with the number of covered reads as the input and calculate the P-value according to the following formula: , where R 1 and R 2 are row totals, C 1 and C 2 are column totals, N is the total number of observations in the row list, and nij is the value in the ith row and jth column of the table.

2. The TMB detection method according to claim 1, characterized in that, the filtering conditions in S1 are as follows: 1) Remove reads containing adaptors, automatically identify the adapter sequence, and perform filtering; 2) Remove reads with a proportion of N greater than 10%; 3) Remove low-quality reads, where low-quality reads refer to reads with a quality value Q <= 5 and the number of bases accounting for more than 50% of the entire read.

3. The TMB detection method according to claim 1, characterized in that, the specific steps of S2 are as follows: 1) Use the bwa software to align the data filtered in S1 with the human reference genome hg19; 2) Use samtools to sort the aligned bam; 3) Use picard to remove duplicates from the sorted bam; 4) Perform statistics on the sequencing coverage rate and alignment rate.

4. The TMB detection method according to claim 1, characterized in that, the filtering conditions for filtering according to the number of reads supporting mutation at the site, mutation frequency, and genotype quality value in S3 are as follows: a. Filter reads with the number of reads supporting mutation less than 10; b. Filter sites with a mutation frequency less than 5%; c. Filter sites with a quality value less than 20.

Citation Information

Patent Citations

  • Method for analyzing tumor mutation load on the basis of next-generation sequencing data of single sample

    CN108470114A

  • Method for accurately predicting BSA-seq candidate gene in low-density SNP genomic region

    CN109360606A

  • Tumor mutation analysis method and system, terminal and readable storage medium

    CN110570904A