Filtering methods and applications of DNA damage false positive mutations in NGS sequencing

A method for filtering DNA damage false positives in FFPE samples constructs a baseline and calculates dynamic error rates to identify and filter mutations across different sequencing types, addressing limitations of current methods and improving sensitivity and accuracy in FFPE sample analysis.

JP7825248B1Active Publication Date: 2026-03-06GENECAST (BEIJING) BIOTECHNOLOGY CO LTD +3
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Current methods for filtering DNA damage false-positive mutations in FFPE samples are limited by mutant allele frequency (VAF) thresholds, primarily affecting low-frequency mutations and single-end sequencing data, and do not effectively accommodate different sequencing platforms and methods.

Method used

A method involving constructing a baseline from normal samples, performing binomial or negative binomial tests to identify DNA damage, calculating dynamic error rates based on sequencing depth, and determining significance probabilities to filter false-positive mutations without VAF limitations, applicable to both single-end and paired-end sequencing.

Benefits of technology

The method accurately filters DNA damage false positives across various sequencing depths and platforms, enhancing sensitivity and reducing false positives in FFPE samples, especially at low and ultra-high depths, while maintaining the detection of clinically relevant mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007825248000001_ABST
    Figure 0007825248000001_ABST
Patent Text Reader

Abstract

A method for filtering DNA damage false-positive mutations in NGS sequencing is provided. The filtering method involves constructing a baseline based on the mutation subtype distribution of normal samples, performing mutation testing on test samples, and using a binomial or negative binomial test to determine the difference between the mutation subtype corresponding to the test sample and the baseline based on the test results. Based on the difference, the test sample is then determined to have DNA damage. SNV mutations with a VAF of 5% or less are extracted from the DNA-damaged samples, a reference error rate is determined for the occurrence of the corresponding mutation subtype at the reference sequencing depth, and a significance p-value for DNA damage at the site is calculated based on the reference error rate. Based on the significance p-value, the corresponding site is determined to be a true mutation site or a false-positive mutation site. The present invention does not limit the frequency of mutant alleles of DNA damage, making it applicable to a wider range of technology platforms.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of genetic testing technology, and in particular to a method and application for filtering DNA damage and false positive mutations in NGS (Next-Generation Sequencing, also known as second-generation sequencing). [Background technology]

[0002] Formalin-fixed, paraffin-embedded (FFPE) samples can be stored for long periods of time, making them a popular method for storing patient tissue samples for clinical pathology, oncology genetic testing, and medical science research. These tissue samples are essential for precision medicine using next-generation sequencing technology. However, some sequence mutations detected in FFPE DNA samples may be due to damage during sample processing, primarily in the following circumstances:

[0003] 1) Oxidative damage

[0004] As shown in Figure 1, reactive oxygen species (ROS) can oxidize guanine (G) to 8-oxoguanine (8-oxoG), which pairs with A instead of C, resulting in a false-positive mutation (G:C / T:A G->T / C->A mutation) after two copies.

[0005] 2) DNA damage caused by FFPE processing

[0006] a. Cross-linking: That is, cross-linking occurs between macromolecules such as proteins and nucleic acids. DNA damage caused by formaldehyde is mainly due to its carbonyl group, which is positively charged and less susceptible to steric hindrance, making cross-linking between nucleic acids and proteins more likely. In vitro, formaldehyde first reacts with free amino groups on proteins or nucleic acids to form unstable hydroxymethyl adducts, which then react with other nucleic acids or proteins to form stable cross-links. This causes local denaturation of DNA, further affecting the quality of nucleic acid extraction.

[0007] b. Base conversion: FFPE DNA samples contain a large number of C>U or C>T mutations. In vivo, as shown in Figure 2, cytosine is converted to uracil by ammonia self-hydrolysis, occurring approximately 190 times per cell per day. Base excision repair can restore cytosine, but this cannot be repaired in vitro, resulting in a large number of false-positive C:G / T:A mutations (C->T / G->A) in the sequencing data.

[0008] c. Base loss: Abasic sites are formed after DNA molecules lose a purine or pyrimidine base. When tissues are fixed using formalin, the formaldehyde in formalin is oxidized to formic acid in the air, and the presence of formic acid lowers the pH value of formalin. On the other hand, the N-glycosidic bond between the purine base and the sugar backbone is easily hydrolyzed at low pH, causing DNA base loss. In addition, abasic sites can self-cleave by B-elimination reaction, further leading to breakage of the DNA chain.

[0009] 3) DNA fragmentation

[0010] The degree of FFPE DNA fragmentation increases with prolonged storage time and with a decrease in the pH value of the formalin used for fixation. PCR success rates are lower when using FFPE DNA that has been stored for a long time compared to fresh FFPE DNA. Therefore, DNA fragmentation is likely to occur throughout the sample storage period.

[0011] Based on the cause of DNA damage and the principles of NGS sequencing, it can be seen that such false positive sites are strand-preferential, i.e., mutations exist only on the positive or negative strand of the DNA molecule, and the reads obtained by NGS sequencing show the characteristics F1R21 or F2R1.

[0012] Currently, commonly used filtering methods include filtering initial candidate mutations based on statistical rules. For example, Yost et al. used a binomial test to compare the allele frequency of each mutation with the mismatch rate of its corresponding mutation type, controlling for formalin-induced deamination, and then removing less significant allele frequencies. Similarly, Kerick et al. used a cutoff value for the sequencing depth of the mutation site for filtering. The FIREVAT algorithm uses a series of filtering parameters, including the VAF, mapping results, and the respective depths of the reference and allele sequences.

[0013] Additionally, filtering algorithms exist to identify false-positive mutations in low-quality or FFPE sequencing data. For example, the cisCall algorithm employs sequential filters for clustered low-VAF mutations. If the mutation error rate within the clustered interval exceeds the expected value, it is filtered out by internal statistical calculations. Bayesian mutation detection algorithms are also used, which incorporate prior information about specific cancer-associated mutations in addition to general quality filtering. Another algorithm, LoLoPicker, uses mutations within a panel of FFPE samples to assess the background error rate of specific sites and incorporates it into hypothesis testing to identify true or false mutations. Meanwhile, mutation detection software developed by Pisces includes a model that recalibrates mutation quality based on the deviation of the average mutation rate for each possible mutation type, minimizing thermal damage and FFPE deamination.

[0014] FFPE-related deamination, like false positives such as oxidative DNA damage, typically occurs only on a single strand of the original DNA template, resulting in directional deviation between reads 1 and 2 during paired-end sequencing. This strand preference allows for effective quantification of DNA damage. This method of calculating linkage disequilibrium is adopted by several software programs, such as the LearnReadOrientationModel module in GASK4, which filters false positive mutations that occur only on a single strand during sequencing. This tool primarily uses a Bayesian statistical model to calculate mononucleotides, instead of strand preference generated per three bases. SOBDetector software also incorporates a strand preference parameter, but instead of directly filtering, it calculates a strand deviation value for each mutation and adds the value to the mutation result file (vcf) for manual screening.

[0015] However, some of the above algorithms discard mutations with a VAF <5% or only evaluate their performance for VAF >5%. Because FFPE-related false positives typically occur at low frequencies, C:G>T:A mutations are discarded when the variant allele frequency (VAF) is <5% or <10%. However, this restriction is undesirable because it prevents the detection of potentially clinically relevant low-frequency mutations and limits the use of low-tumor content samples. FFPE-related false positives have also been observed at frequencies of 10% or higher.

[0016] Because DNA damage mutations exist on a single strand of the original DNA template, current software or algorithms incorporate DNA template preference into calculations or feature training, but do not apply to single-end sequencing NGS data. Based on current literature, most test data are from capture methods, with relatively little amplicon data available, and single-end sequencing amplicon data has yet to be confirmed for DNA damage samples. Summary of the Invention [Problem to be solved by the invention]

[0017] SUMMARY OF THE INVENTION An object of the present invention is to solve at least one of the technical problems in the related art mentioned above to some extent.

[0018] To this end, the objective of the present invention is to provide a method and application for filtering DNA damage false-positive mutations in NGS sequencing, which does not limit the mutant allele frequency (VAF) of DNA damage, can be applied to more technology platforms, and can accommodate both single-end and paired-end sequencing.

[0019] In order to solve the above technical problems, the present invention is realized as follows.

[0020] An embodiment of the present invention provides a method for filtering DNA damage false positive mutations in NGS sequencing, the method comprising:

[0021] S1, constructing a baseline based on the mutation subtype distribution of normal samples;

[0022] S2, performing a mutation test on the test sample, and based on the test results, determining the difference between the mutation subtype corresponding to the test sample and the baseline using a binomial test or a negative binomial test, and determining whether or not the test sample has DNA damage based on the difference;

[0023] S3. Extract SNV mutations with VAF ≤ 5% for DNA damage samples and determine the baseline error rate at which the corresponding mutation subtypes occur at the baseline sequencing depth.

[0024] S4: Calculate the significance probability p-value of DNA damage at the site based on the reference error rate, and determine whether the corresponding site is a true mutation site or a false positive mutation site based on the significance probability p-value of DNA damage, thereby realizing filtering.

[0025] In addition, the method for filtering DNA damage false-positive mutations in NGS sequencing of the present invention has the following additional technical features:

[0026] In some embodiments, constructing a baseline based on the mutation subtype distribution of S1 normal samples comprises:

[0027] A number of normal samples without DNA damage are selected and subjected to mutation testing, the original SNV results of the mutation testing are statistically analyzed, the distribution of the measured mutation subtypes in the normal samples is obtained, and a baseline is constructed based on this distribution.

[0028] In some embodiments, step S2 includes:

[0029] The frequencies of different measured mutation subtypes in the mutation test results of the measured sample are statistically analyzed, and a one-sided binomial test or a negative binomial test is performed based on the baseline to calculate a p-value. The obtained p-value is compared with a preset threshold value. If the p-value is less than or equal to the preset threshold value and the number of corresponding mutation subtypes in the measured sample is greater than or equal to the preset reference value, it is determined that DNA damage exists in the sample; conversely, it is determined that DNA damage does not exist.

[0030] In some embodiments, the preset threshold corresponding to the p-value is 0.01.

[0031] In some embodiments, the method for obtaining the reference sequencing depth is as follows, where the median of the sequencing depths of all SNVs in the corresponding samples is taken as the reference sequencing depth.

[0032] In some embodiments, in step S3, the sequence depth, supporting sequence base sequence, and VAF of all measured mutation subtypes are statistically analyzed, and SNV mutations with a VAF of 5% or less are extracted based on the statistical information, and the 95% quantile is used as the reference error rate occurring at the reference sequence depth of the measured mutation subtype in the sample.

[0033] In some embodiments, the baseline error rate is negatively correlated with a power of log(sequence depth).

[0034] In some embodiments, the relationship between the reference error rate and the sequence depth is specifically as follows (the following formula):

[0035]

number

[0036] In some embodiments, the ultra-high depth cutoff value needs to be preset,

[0037] The sequencing depth is compared with the ultra-high depth cutoff value. If the sequencing depth is greater than the ultra-high depth cutoff value, the error rate will no longer decrease with increasing sequencing depth, and all error rates at sites higher than the depth cutoff value will remain unchanged, using the error rate corresponding to the cutoff value depth.

[0038] The embodiments of the present invention further provide an application of the method for filtering DNA damage false-positive mutations in NGS sequencing described in any one of the above, wherein the method is applied to single-end sequencing SE and paired-end sequencing PE. [Effects of the Invention]

[0039] Compared with the prior art, the present invention has at least the following advantageous effects:

[0040] The method for filtering DNA damage false-positive mutations in NGS sequencing provided by the embodiments of the present invention can solve the current problem of strand preference in identifying and filtering DNA damage false-positive mutations, and can be applied to mutation sites with different VAFs at different depths, and to data from different sequencing platforms and sequencing methods, which can be useful for accurate testing of oncogenes.

[0041] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. [Brief explanation of the drawings]

[0042] [Figure 1] Figure 1 shows a schematic diagram of the oxidative DNA damage 8-oxoG as used in the prior art. [Figure 2] Figure 2 shows a schematic diagram of DNA damage caused by conventional FFPE processing, specifically, C:G / T:A (G->A / C->T mutations) due to the conversion of cytosine to uracil caused by DNA cytosine deamination. [Figure 3] FIG. 3 is a flow chart of a method for filtering DNA damage false positive mutations in NGS sequencing, in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0043] The technical solutions of the present invention will be described below clearly and completely with reference to the drawings of the embodiments of the present invention. Obviously, the embodiments described herein are not all the embodiments but only some of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of the present invention.

[0044] Hereinafter, the embodiments of the present invention will be described in detail with reference to the accompanying drawings through specific examples and application scenarios.

[0045] As shown in Figure 1, some embodiments of the present invention provide a method for filtering DNA damage false-positive mutations in NGS sequencing, which overcomes the shortcomings of current technology, does not limit the mutant allele frequency (VAF) of DNA damage, is applicable to more technology platforms, and is compatible with both single-end and paired-end sequencing, making it highly likely to be applied to tumor screening in the future.

[0046] In some embodiments of the present invention, the method for filtering DNA damage false positive mutations in NGS sequencing specifically includes:

[0047] First, a baseline is constructed based on the distribution of 12 subtype mutations in normal samples. A binomial test or negative binomial test is used to determine whether there is a significant difference between the subtype corresponding to the test sample and the baseline. The cutoff for DNA damage samples is determined by a binomial p-value of ≤ 0.01. Typical DNA damage assessment or filtering software typically uses the GATK CollectOxoGMetrics or FilterByOrientationBias modules to calculate the TOTAL_QSCORE score for the 12 subtypes of a sample; lower scores indicate a higher likelihood of artifactual alteration. At the same time, the F1R2 and F2R1 values ​​are used to determine whether the mutation is a false positive due to DNA damage. Therefore, this method does not directly assess DNA damage in the sample; subsequent site assessment also requires filtering based on strand preference (i.e., whether there is variability in the F1R2 or F2R1 values ​​of the mutation's supporting reads). Here, GATK is the Genome Analysis Toolkit, i.e., a genome analysis kit, CollectOxoGMetrics is a module for collecting quantitative indicators of oxidized guanine (OxoG), and FilterByOrientationBias is a module for filtering orientation bias.

[0048] Next, after confirming that the test sample is a DNA-damaged sample, the median depth and 95% quantile allele frequency (Vaf) for each of the 12 subtypes of the test sample are statistically calculated. These depths and frequencies are used as the reference depth and reference error rate for the corresponding mutation type. For each potential DNA damage false-positive mutation type (G → A / C → T, G → T / CA) site in each test sample, the dynamic error rate ε (see Equation 2) is used as the background error probability for that site based on the actual sequencing depth of each site, and the authenticity of the next site is evaluated. As the depth increases, a lower error rate is used to avoid filtering out high-frequency true mutations. To prevent the error rate from dropping to an unrealistically low value at very high depths, the median depth of the test sample is used as a cutoff value. This ensures that the error rate does not decrease with increasing depth when the depth is greater than this value. The dynamic error rate continues to increase as the depth decreases, which is consistent with the fact that the background Vaf of DNA damage sites is higher at low depths. The dynamic error rate can effectively evaluate the actual situation of DNA damage sites at ultra-low depths (<100X) or ultra-high depths (≧5000X). Current software or algorithms evaluate DNA damage sites using binomial distribution or other algorithms, but the measurement depth of tissue samples is relatively low, and the test effect on DNA damage at ultra-high depths is unknown. The dynamic error rate of the present invention can compensate for this technical shortcoming.

[0049] Finally, for each target site in the sample, a binomial test or negative binomial test is performed using the current dynamic error rate ε and sequencing depth to obtain a p-value (see Equation 2), which represents the significance probability of DNA damage at that site. If p-damage > 0.001, there is no significant difference between the background error rate of that site and the corresponding mutation type in the sample, and the site is filtered. Conventional software or algorithms rely on strand preference as a key condition for determining whether a mutation is due to DNA damage, and therefore are unable to calculate single-end sequencing data, which has drawbacks that affect the filtering effect of the algorithm. However, the present invention can calculate both single-end and paired-end sequencing data. Furthermore, while some software or algorithms limit DNA damage sites to a 5% range, the present invention can eliminate potential false-positive DNA damage sites even at low frequencies.

[0050] In actual tumor clinical samples, the present invention tested FFPE samples with a storage period of more than two or more than five years. The simultaneously calculated p-sample and the experimentally detected DNA degradation showed a high degree of linear correlation; the larger the p-sample value, the higher the degradation level and the more severe the DNA damage in the sample. The p-sample value of the present invention effectively reduces experimental procedures and provides higher sensitivity than experimental results. The mutations filtered by the present invention significantly reduce false positive mutation sites, making it suitable for low-depth or ultra-high-depth sample testing. Furthermore, the present invention is easy to operate and can be applied to multiple platforms and sequencing methods (single-end sequencing SE or paired-end sequencing PE).

[0051] Example 1

[0052] To accurately filter base mutations caused by DNA damage, the present invention has designed a method for filtering false-positive mutations caused by DNA damage. This method first identifies DNA-damaged samples, and in this step, distinguishes between DNA-damaged samples and normal samples, thereby avoiding false-positive mutations in normal samples during filtering. Next, the identified DNA-damaged samples are filtered, the p-value of the measured site is calculated, and based on the p-value, it is determined whether the site is caused by DNA damage. Specifically,

[0053] a. Statistical analysis is performed using the original SNV (mononucleotide variation) results after mutation testing of multiple normal DNA samples (read as having no DNA damage), to obtain the distribution of the measured mutation subtype in normal samples, and to construct a corresponding baseline.

[0054] b. The original SNV results after mutation testing of the test sample are processed in the same way, and the frequencies of different test mutation subtypes are statistically calculated. A one-sided binomial test or negative binomial test is performed on the baseline to calculate the p-value. If the p-value is less than or equal to 0.01, it is determined that there is DNA damage within the sample. If the number of corresponding mutation subtypes in the test sample is less than the reference value, it is not considered to be DNA damage even if there is a significant difference in the p-value.

[0055] c. For samples determined to have DNA damage, the median of all SNV sequence depths for that sample is calculated as the reference sequence depth, and information such as the sequence depth, support reads (support sequence base sequences), and VAF for all measured mutation subtypes is calculated. SNV mutations with a VAF ≤ 5% are extracted, and the 95% quantile is used as the reference error rate ∈ for the reference sequence depth for the measured mutation subtype in that sample.

[0056] Dynamic error rate of site depth ∈, the error rate of DNA damage type sites is different from the constant error rate of other error types, and its error rate increases with decreasing depth,

[0057]

number

[0058] where e is the vaf at the 90% quantile of the measured sample for the measured mutation type, e.g., CT>T / G->A, and is used as the reference error rate;

[0059] Median Depth is the median depth of the sample,

[0060] Depth is the depth of the area to be measured.

[0061] Formula (1) has a relatively good effect in some examples. In theory, other formulas can be used, but comparative verification through a large number of tests has shown that formula (1) has the best filtering effect on samples with low DNA damage depth.

[0062] d. Each sequencing depth corresponds to an error rate, specifically shown in Equation (1). The error rate calculated using this formula is different from a fixed error rate; its exponent decreases with increasing sequencing depth. Therefore, a large error rate can be obtained at low sequencing depths, which can be used to filter mutations caused by DNA damage. A low error rate can be obtained at high sequencing depths, avoiding the filtering of frequent true mutations. To prevent the error rate from dropping to an extremely low level at ultra-high sequencing depths, the present invention sets a cutoff value. When the sequencing depth is greater than this cutoff value, the error rate does not decrease with increasing sequencing depth, and the error rate corresponding to the depth of the cutoff value is used as the error rate for the current sequencing depth. The error rate corresponding to this cutoff value depth can also be calculated using Equation (1).

[0063] e. Using the above error rate, calculate the significance probability p-value of the DNA damage site (damage_p-value), specifically refer to Equation (2). Filter the observed mutation subtypes based on the damage_p-value. If the damage_p-value is ≦0.001, this site is considered to be a true mutation, or conversely, a false positive site due to DNA damage.

[0064] Calculate the significance probability p-value of the DNA damage site, i.e., perform a binomial test (or negative binomial test) at the current error rate and depth of each C->T / G->A mutation site to obtain the p-value.

[0065]

number

[0066] where SR is the number of support reads for the mutation,

[0067] D is the depth, the number of all reads at the site (including unsupported mutations, i.e., reads of the reference allele),

[0068] ∈ is the error rate.

[0069] The method for filtering DNA damage false positive mutations in the NGS sequencing of the present invention first clearly determines potential DNA damage samples based on baseline samples and calculated p-samples, and can determine the severity of DNA damage of the sample based on the magnitude of the p-sample value (shown in Table 1). DNA samples generally use the N / Q value as the sample quality determination criterion in the experimental process. When 0 < N / Q value ≤ 3.5, it is determined that the integrity of the DNA sample is good. When 3.5 < N / Q value ≤ 6.5, the integrity of the DNA sample is slightly good. When the N / Q value > 6.5, it is determined that the integrity of the DNA sample is poor. As can be seen from the table, as the ratio of N to Q increases, the value of p-sample becomes smaller and smaller, and it can be seen that the quality of the sample becomes worse and worse. Therefore, the present invention can assist the experiment in determining the quality of the sample. As an important parameter for sample quality control, it can reduce the experimental steps and save the experimental cost. Secondly, it can solve the problem of relying on strand preference for the identification and filtering of current DNA damage false positive mutations. Some FFPE samples continuously accumulate mutations of C>T and G>A damage in some fragments due to long-term storage. Due to the existence of these mutations, such fragments have low PCR amplification efficiency in library preparation before sequencing, resulting in low measured depth. Furthermore, the sequence data of these regions shows the characteristics of low depth (Depth) and high mutation frequency (VAF). In similar software such as GATK, because the software relies on strand preference, it cannot effectively filter single-end sequence data from DNA damage samples. However, through comparative tests, the method of the present invention can effectively filter such errors. Furthermore, the present invention can also be applied to mutation sites with different VAFs at different depths. For example, in low-quality samples, DNA damage is very serious at the same time as severe degradation, so the amplification efficiency is relatively low. When the depth of the sample is insufficient, it is difficult to distinguish the authenticity of mutations. However, the present invention can perform false positive filtering on such samples and increase the probability of patient examination and medication.Finally, NGS testing has been widely applied in precision medicine. Apart from the general capture method, amplicon amplification is widely used due to its fast testing speed, usually producing results within 24 hours. However, this testing method differs between the MGI platform and multiple Illumina platforms due to differences in sequencing error rates. The present invention can simultaneously accommodate sequencing data from multiple platforms and different sequencing methods, providing powerful support for accurate cancer gene testing.

[0070] [Table 1]

[0071] For the parts of the present invention that are not described in detail, reference should be made to existing techniques in the art or techniques known to those skilled in the art.

[0072] Although the embodiments of the present invention have been described above with reference to the drawings, the present invention is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are merely examples and do not limit the present invention. Those skilled in the art may make other modifications based on the present invention without departing from the spirit of the present invention and the scope protected by the claims, and all of them are within the protection scope of the present invention.

Claims

1. 1. A method for filtering DNA damaging false positive mutations in NGS sequencing, the method comprising: S1, constructing a baseline based on the mutation subtype distribution of normal samples; S2: Performing a mutation test on the test sample, and based on the test results, determining the difference between the mutation subtype corresponding to the test sample and the baseline using a binomial test or a negative binomial test, and determining the presence or absence of DNA damage in the test sample based on the difference; S3: After confirming that the test sample is a DNA-damaged sample, the median depth and 95% quantile allele frequency (Vaf) are calculated for each of the 12 subtypes of the test sample, and the depth and frequency are used as the reference sequence depth and reference error rate corresponding to the mutation type; S4: For potential DNA damage false-positive mutation type sites in the measured sample, a dynamic error rate ε is calculated based on the actual sequencing depth of each site, and a significance probability p-value of DNA damage at each site is calculated using the dynamic error rate ε and the sequencing depth. Based on the significance probability p-value of DNA damage, it is determined whether the corresponding site is a true mutation site or a false-positive mutation site, thereby realizing filtering.

2. Constructing a baseline based on the mutation subtype distribution of normal samples of S1 includes: The method for filtering DNA damage false positive mutations in NGS sequencing described in claim 1, characterized in that multiple normal samples without DNA damage are selected and subjected to mutation testing, the original SNV results of the mutation testing are statistically analyzed, a distribution of the measured mutation subtypes in the normal samples is obtained, and a baseline is constructed based on this distribution.

3. Step S2 includes: The method for filtering DNA damage false-positive mutations in NGS sequencing according to claim 1, characterized in that the frequency of different measured mutation subtypes in the mutation test results of the measured sample is statistically calculated, a one-sided binomial test or a negative binomial test is performed based on the baseline, and a p-value is calculated by comparing the obtained p-value with a preset threshold value. If the p-value is less than or equal to the preset threshold value and the number of corresponding mutation subtypes in the measured sample is greater than or equal to the preset reference value, it is determined that DNA damage is present in the sample, and conversely, it is determined that DNA damage is not present.

4. The method for filtering DNA damage false positive mutations in NGS sequences according to claim 3, characterized in that the preset threshold corresponding to p-value is 0.

01.

5. The method for obtaining the reference sequencing depth is as follows: the median of the sequencing depths of all SNVs in the corresponding sample is used as the reference sequencing depth. The method for filtering DNA damaging false positive mutations in NGS sequencing described in claim 1 is characterized in that:

6. The method for filtering DNA damaging false positive mutations in NGS sequencing described in claim 1, characterized in that in step S3, the sequence depth, supporting sequence base sequence and VAF of all measured mutation subtypes are statistically analyzed, SNV mutations with VAF≦5% are extracted based on the statistical information, and the 95% quantile is used as the reference error rate occurring at the reference sequence depth of the measured mutation subtype in the sample.

7. The method for filtering DNA damaging false positive mutations in NGS sequences according to claim 1, characterized in that the reference error rate shows a negative correlation with the power of log(sequencing depth).

8. The relationship between the reference error rate and the sequence depth is specifically as follows: [Equation 4] The method for filtering DNA damage false positive mutations in NGS sequences according to claim 1.

9. Ultra-high depth cutoff value preset The method for filtering DNA damaging false positive mutations in NGS sequencing described in claim 1, characterized in that the sequencing depth is compared with the ultra-high depth cutoff value, and if the sequencing depth is greater than the ultra-high depth cutoff value, the error rate corresponding to the depth cutoff value is used as the error rate of the current sequencing depth.

10. A method for applying the method for filtering DNA damage false positive mutations in NGS sequencing described in any one of claims 1 to 9, applied to single-end sequencing SE or paired-end sequencing PE.

Citation Information

Patent Citations

  • Method for detecting somatic cell mutation of paraffin section samples based on next-generation sequencing and device thereof

    CN110729025A

  • Mutation analysis method and device for cell-free DNA (cfDNA) sequencing data

    CN112111565A

  • False positive nucleotide variation site filtering method and computing device

    CN114613430A

  • Filtering method for breaking false positive mutation generated by artificial fragments in library building through enzyme digestion method

    CN116895332A

  • Method for increasing ratio of intrinsic fragment used in NGS analysis for detecting low-frequency mutation of cfdna

    WO2022050654A1