How next-generation sequencing panels are validated

By employing homozygous DNA and control genomic DNA at varying dilutions, the method addresses the limitations of existing NGS panel validation methods, providing precise detection limits and reducing errors in mutation frequency analysis, thus improving the reliability of NGS panel validation.

JP2025526421AActive Publication Date: 2025-08-13NATIONAL CANCER CENTER(JP)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025504509
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-28
Filing Date
2023-07-27
Publication Date
2025-08-13
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

Current methods for validating next-generation sequencing (NGS) panels, such as the Frampton method, struggle to accurately assess the sensitivity and specificity of NGS panels, particularly for alleles with a variant allelic fraction (VAF) less than 5%, and require extensive use of HapMap cell lines, leading to inefficiencies and inconsistencies in evaluating detection limits and mutation frequencies.

Method used

A composition and method using homozygous DNA, such as from a hydatidiform mole cell line, and control genomic DNA at varying dilutions to analyze false-negative and false-positive mutations, and derive detection limits and stringency cutoffs for NGS panels, enhancing sensitivity and specificity analysis.

Benefits of technology

The method provides objective validation of NGS panels by accurately determining false-negative and false-positive mutation frequencies, allowing for improved detection limits and reduced errors in NGS results, thereby enhancing the reliability of NGS panel performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025526421000001_ABST
    Figure 2025526421000001_ABST
Patent Text Reader

Abstract

The present invention relates to a composition for validating a next-generation sequencing (NGS) panel containing homozygous DNA and control genomic DNA; a kit for validating a next-generation sequencing (NGS) panel containing the composition; a method for validating a next-generation sequencing (NGS) panel through analysis of false-negative mutations, analysis of detection limit, and analysis of false-positive mutations; and a method for providing information to increase the specificity of a next-generation sequencing (NGS) panel. In particular, the validation method according to the present invention can analyze the frequency of false-negative mutations, the frequency of false-positive mutations, and the detection limit for a next-generation sequencing (NGS) panel as objective results, and can be useful for validating a next-generation sequencing (NGS) panel.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a composition for validating a next-generation base sequence analysis panel; a kit for validating a next-generation base sequence analysis panel comprising the same; and a method for validating a next-generation base sequence analysis panel. [Background technology]

[0002] Next-generation sequencing (NGS), also known as massively parallel sequencing or high-throughput sequencing, is a technology capable of analyzing millions of base sequences at once. Previously, Sanger sequencing required several years and billions of dollars to analyze one genome. However, NGS allows analysis to be completed in a few days for around $1,000. NGS technology is currently being applied to the screening, diagnosis, prognosis, and therapeutic selection of various diseases, including cancer, genetic disorders, and infectious diseases. More than 55,000 NGS diagnostic panels are currently in clinical use for more than 11,000 diseases in the United States alone.

[0003] To apply developed NGS panels to clinical practice, the panels must be evaluated to ensure they accurately produce the desired results. For this reason, the US FDA has published guidelines for the validation of NGS panels, and the Centers for Medicare & Medicaid Services (CMS) monitors panels, with the guidelines and monitoring methods being continuously updated.

[0004] Currently, the most commonly used method for validating NGS panels is the one introduced by Frampton et al. in Nature Biotechnology in 2013 (hereafter referred to as the "Frampton method"), and subsequent studies have primarily used the Frampton method (Shin et al., Nature Comm 2017; Froyen et al., Cancers 2022). The Frampton method validates NGS panels using a DNA pool containing DNA from HapMap cell lines or cancer cell lines with well-known sequences. Specifically, the Frampton method involves mixing equal amounts of DNA from 10 HapMap cell lines to create a DNA pool (called pool XX) with 5% DNA, and another DNA pool (called pool YY) with 10 other cell lines. These were then analyzed with an NGS panel to evaluate sensitivity and specificity (Frampton et al., Nature Biotechnology 2013). The proportion of each single nucleotide variant (SNV) allele in a DNA mixture (hereafter referred to as "variant allelic fraction" or VAF) can be calculated. Haplotype alleles that appear only in one cell line account for 5% of the total mixture, while the remaining SNVs that appear multiple times have a VAF of 10% or more. When Frampton et al. analyzed the panel using a mixture of selected HapMap cell lines, 20% of the alleles contained in the panel genes had a VAF of approximately 5%, and approximately 30% had a VAF of 10%. Therefore, 20% of the alleles could be used to measure the detection rate of 5% VAF, and 30% of the alleles could be used to measure the detection rate of 10% VAF. However, in Frampton et al.'s paper, the HapMap cell line mixture they created did not contain any alleles with a VAF of less than 5%, so they were unable to measure the detection rate of alleles with a VAF of less than 5%.Furthermore, to measure the detection rate of alleles with a 1% VAF using the Frampton method, at least 50 HapMap cell lines must be used, and mutations with a mutation detection frequency of approximately 1% must be used. In addition, in the results of Frampton's paper, approximately 200 alleles with a VAF of approximately 5% were obtained using 10 HapMap cell lines. However, based on this estimate, only 40 or fewer alleles with a 1% VAF could be obtained, which creates the disadvantage of having to evaluate the sensitivity of an NGS panel with a small number of alleles.

[0005] In addition, when evaluating sensitivity using the Frampton method, sensitivity was closely related to (1) the down-sampled coverage of the in silico model (corresponding to read counts) and (2) the VAF of the allele. Specifically, SNVs with a 5% VAF were detected at 96.1% of alleles at 500 coverage, 83.5% at 300 coverage, and 44.2% at 150 coverage. SNVs with a 10% VAF were detected at 99.3% of alleles at 500 coverage, 98.9% at 300 coverage, and 82.3% at 150 coverage. SNVs with a VAF greater than 10% were detected at 99.5% of alleles even at 150 coverage. In addition, specificity was evaluated using the Framton method. According to the results of the published paper, only when the downsampling coverage was greater than 500, two SNVs with a fraction of less than 5% were detected. In the remaining cases, no false-positive mutations were detected, resulting in a positive predictive value of over 99.9%, indicating that the specificity of the NGS panel used was very high.

[0006] In addition, to detect indels in cancer cell lines using the Frampton method, we created 41 pools of DNA from cancer cell lines, mixing DNA from 2 to 10 copies, and measured the sensitivity and specificity of indel detection. We also confirmed that indel detection sensitivity varied depending on (1) downsampling coverage and (2) the VAF of indel mutations within the mixture. Specifically, detection of indels with a 10% VAF was 65% at 190 coverage and 88.3% at 667 coverage. Detection of indels with a 20% VAF was 86.3% at 190 coverage and 97.3% at 667 coverage. Regarding false-positive indels, no false-positive indels with a VAF of 20% or more were detected, but 1-2% of indels with a VAF of 20% or less were detected. The same method was used to measure the sensitivity and specificity of gene amplification / deletion detection using DNA from a mixture of various cancer cell lines. No false-positive mutations were found, but false-negative mutations were detected in approximately 10-20% of cases.

[0007] As mentioned above, the Frampton method accurately mixes approximately 10 DNA samples in equal amounts during the dilution process, but frequent errors make it difficult to evaluate the sensitivity of NGS panels and do not provide accurate information on how many diluted mutations can be detected at a particular dilution ratio. In addition to these limitations of the Frampton method, the detection of indels and amplifications / deletions is less sensitive than that of SNVs. Variables such as allelic bias, duplicate reads, and the power of the analytical pipeline can affect sensitivity and specificity, making it difficult to establish guidelines for the sensitivity and specificity of NGS panels (Shin et al., Nature Comm 2017). Therefore, the current guidelines of the American College of Medical Genetics and Genomics for next-generation sequencing-based assays do not set specific coverage thresholds or minimum sensitivity or specificity (US Food and Drug Administration (FDA). Considerations for Design, Development, and Analytical Validation of Next-Generation Sequencing-Based In Vitro Diagnostics Intended to Aim in the Diagnosis of Suspected Germline Diseases. Updated 13 April 2018).

[0008] However, information on VAF, which is the minimum read count detectable at the target site, and limit of detection (LoD) information are provided through standard dilution experiments (Rehm et al., Genet Med, 2013; ACMG clinical laboratory standards for next-generation sequencing). While regulations for NGS methods are still unclear, the American Clinical Laboratory Associations (ACLA) oversees NGS panels currently used for genetic testing through the Clinical Laboratory Improvement Amendments (CLIA), and the FDA has established guidelines for the design, development, and analytical validation of such panels, thereby enabling NGS to be applied clinically. Therefore, as new reference materials and methods for measuring the sensitivity and specificity of NGS panels are developed, researchers will be able to easily compare and evaluate NGS panels, and the minimum regulations required for clinical application may also be revised. Summary of the Invention [Problem to be solved by the invention]

[0009] An object of the present invention is to provide a composition for validation of next-generation sequencing (NGS) panels comprising homozygote DNA and control genomic DNA.

[0010] Another object of the present invention is to provide a kit for verifying a next-generation base sequence analysis (NGS) panel, comprising the composition for verifying a next-generation base sequence analysis panel (NGS).

[0011] It is yet another object of the present invention to provide a method for validating a next-generation sequencing (NGS) panel through analysis of false-negative mutations, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting a Ho-N pair allele (homozygote-null pair allele) or a He-N pair allele (heterozygote-null pair allele) from the results of the next-generation sequencing (NGS) performed in step (a); and (c) analyzing the frequency of false-negative mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected Ho-N pair allele or He-N pair allele.

[0012] Another object of the present invention is to provide a method for validating an NGS panel through a detection limit analysis, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) selecting a Ho-N pair allele (homozygote-null pair allele) or a He-N pair allele (heterozygote-null pair allele) from the results of the NGS obtained in step (a); (c) analyzing a dilution mutation detection rate, which indicates the number of diluted mutations found in the Ho-N pair allele or He-N pair allele as a percentage, in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios based on the selected Ho-N pair allele or He-N pair allele; and (d) analyzing the detection limit.

[0013] Another object of the present invention is to provide a method for validating a next-generation sequencing (NGS) panel through analysis of false-positive mutations, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using a next-generation sequencing (NGS) panel; (b) selecting an NN pair allele (Null-Null pair allele) containing no mutation in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) performed in step (a); and (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected NN pair allele.

[0014] It is yet another object of the present invention to provide a method for providing information for increasing the specificity of a next-generation sequencing (NGS) panel, comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting NN pair alleles (Null-Null pair alleles) containing no mutations in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) performed in step (a); (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected NN pair alleles; and (d) analyzing variant allelic fraction (VAF) for removing false-positive mutations and deriving a stringency cutoff value.

[0015] Yet another object of the present invention is to provide a method for evaluating errors in the process of generating raw data or errors in the bioinformatics analysis process in a company's NGS panel analysis process, the method comprising: (a) collecting analysis results for false negative and false positive mutations from the company's raw data using bioinformatics analysis results held by the company; and (b) comparing and analyzing the bioinformatics analysis results from the company's raw data using commercially available software with the analysis results of step (a). [Means for solving the problem]

[0016] To achieve the above object, the present invention provides a composition for validation of a next-generation sequencing (NGS) panel comprising homozygote DNA and control genomic DNA.

[0017] In one embodiment of the present invention, the homozygous DNA may be DNA isolated from a hydatidiform mole cell line or a parthenogenetic cell line.

[0018] In one embodiment of the present invention, the control genomic DNA may be genomic DNA isolated from normal human blood, genomic DNA isolated from patient blood or cancer tissue, circulating tumor DNA (ctDNA), or synthetic or cloned DNA.

[0019] In one embodiment of the present invention, the dilution ratio of the homozygous DNA and the control genomic DNA may be selected from the group consisting of 1:9999 to 9999: 1. Preferably, the dilution ratio of the homozygous DNA and the control genomic DNA may be selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5 and 9999:1, but is not limited thereto.

[0020] The present invention also provides a kit for verifying a next-generation base sequence analysis (NGS) panel, comprising the composition for verifying a next-generation base sequence analysis panel (NGS).

[0021] The present invention also provides a method for validating a next-generation sequencing (NGS) panel through analysis of false-negative mutations, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting a Ho-N pair allele (homozygote-null pair allele) or a He-N pair allele (heterozygote-null pair allele) from the results of the next-generation sequencing (NGS) obtained in step (a); and (c) analyzing the frequency of false-negative mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected Ho-N pair allele or He-N pair allele.

[0022] The present invention also provides a method for validating an NGS panel through a detection limit analysis, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) selecting a Ho-N pair allele (homozygote-null pair allele) or a He-N pair allele (heterozygote-null pair allele) from the NGS results obtained in step (a); (c) analyzing a dilution mutation detection rate, which indicates the number of diluted mutations found in the Ho-N pair allele or He-N pair allele as a percentage, in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios, based on the selected Ho-N pair allele or He-N pair allele; and (d) analyzing the detection limit.

[0023] In one embodiment of the present invention, the method may validate a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNVs), insertions / deletions (Indels), and chromosomal amplifications / deletions.

[0024] In one embodiment of the present invention, the detection limit in step (d) may be a percentage of the VAF (variant allelic fraction) corresponding to the maximum dilution ratio that shows a 90% or 95% dilution mutation detection rate in a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios.

[0025] The present invention also provides a method for validating a next-generation sequencing (NGS) panel through analysis of false-positive mutations, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting an NN pair allele (Null-Null pair allele) containing no mutation in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) obtained in step (a); and (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected NN pair allele.

[0026] In one embodiment of the present invention, the method may validate a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNVs), insertions / deletions (Indels), and chromosomal amplifications / deletions.

[0027] The present invention also provides a method for providing information for increasing the specificity of a next-generation sequencing (NGS) panel, comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting an NN pair allele (Null-Null pair allele) containing no mutation in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) obtained in step (a); (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected NN pair allele; and (d) analyzing a variant allelic fraction (VAF) for removing false-positive mutations and deriving a stringency cutoff value.

[0028] In one embodiment of the present invention, the stringency cutoff value may be a 95% stringency cutoff value that removes 95% of false positive mutations or a 99% stringency cutoff value that removes 99% of false positive mutations.

[0029] In one embodiment of the invention, the method may provide information on stringency cutoff values for eliminating false positive mutations to increase the specificity of next generation sequencing (NGS) panels.

[0030] The present invention also provides a method for evaluating errors in the process of generating raw data or errors in the bioinformatics analysis process in a company's NGS panel analysis process, including: (a) collecting analysis results for false negative and false positive mutations from the company's raw data using bioinformatics analysis results held by the company; and (b) comparing and analyzing the bioinformatics analysis results from the company's raw data using commercially available software with the analysis results of step (a). [Effects of the Invention]

[0031] The compositions according to the present invention can be useful for validating NGS panels. In particular, the validation method according to the present invention (i) can analyze the frequency of false-negative mutations, the frequency of false-positive mutations, and the detection limit for an NGS panel as objective results; (ii) can pre-determine the presence or absence of specific chromosomal sites or specific gene sites where false-negatives are particularly frequent during validation of an NGS panel, thereby removing sites with a high number of false-negatives during analysis and reducing errors in the analysis of NGS results associated with false-negatives occurring in NGS panels; and (iii) can analyze the number of false-positive mutations and the distribution of variant allelic fraction (VAF) to provide an objective stringency cutoff value, thereby reducing errors in the analysis of NGS results associated with false-positives occurring in NGS panels. [Brief explanation of the drawings]

[0032] [Figure 1] FIG. 1 shows that the relative amounts of DNA can be measured in a DNA mixture in which DNA1, a homozygous variant, and DNA2, a null, are mixed at various dilution ratios. [Figure 2] Figure 2A shows the results of the analysis of the detection limit for the NGS panel of AA company using the Ho-N pair allele. Figures 2B and 2C show the results of the analysis of the detection limit for the NGS panel of BB company. Figures 2D and 2E show the results of the analysis of the detection limit for the NGS panel of CC company. (X-axis: alleles listed by chromosome number and position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); and FN: false-negative allele; each color indicates a false-negative result at the corresponding dilution ratio; X-axis below each graph: alleles listed in order by chromosome position; and Y-axis below each graph: chromosome number). The dilution factors of control genomic DNA and H mole DNA are as follows: CH100, 0:100; CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; CH1, 99:1; and CH0, 100:0. [Figure 3]Figure 3A shows the results of an analysis of the detection limit for the NGS panel of AA company using the He-N pair allele. Figure 3B shows the results of an analysis of the detection limit for the NGS panel of BB company. Figure 3C shows the results of an analysis of the detection limit for the NGS panel of CC company. (X-axis: alleles listed in order by chromosome number and position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false-negative allele; each color indicates a false-negative result at the corresponding dilution; X-axis at the bottom of each graph: alleles listed in order by chromosome position; and Y-axis at the bottom of each graph: chromosome number). The dilution factors of control genomic DNA and H mole DNA are as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; CH1, 99:1; and CH0, 100:0. [Figure 4] Figure 4A shows the results of indel detection analysis by dilution factor of the NGS panel from AA using Ho-N pair alleles or He-N pair alleles. Figure 4B shows the results of indel detection analysis by dilution factor of the NGS panel from BB (X axis: alleles sorted by chromosomal location; Y axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false-negative allele; and each color represents a false-negative result at that dilution factor). The dilutions of control genomic DNA and H mole DNA were as follows: CH100 (0:100); CH95 (5:95); CH90 (10:90); CH80 (20:80); CH50 (50:50); CH20 (80:20); CH10 (90:10); CH5 (95:5); and CH0 (100:0). [Figure 5]Figure 5A shows the results of an analysis of the detection limit for the NGS panel of AA company using the Ho-He pair allele. Figure 5B shows the results of an analysis of the detection limit for the NGS panel of BB company. Figure 5C shows the results of an analysis of the detection limit for the NGS panel of CC company. (X-axis: alleles listed by chromosomal position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false-negative allele; each color indicates a false-negative result at a corresponding dilution ratio; X-axis at the bottom of the graph in 5C: alleles listed in order by chromosomal position; and Y-axis at the bottom of the graph in 5C: chromosome number). The dilution factors of control genomic DNA and H mole DNA are as follows: CH100, 0:100; CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; CH1, 99:1; and CH0, 100:0. [Figure 6]Figure 6A shows the results of an analysis of the detection limit for chromosome 6 mutations using the Ho-N pair allele or He-N pair allele for CC's NGS panel. Figure 6B shows the results of an analysis of the detection limit for chromosome 6 mutations using the null Ho-N pair allele in control DNA. Figure 6C shows the results of an analysis of the detection limit for chromosome 6 mutations using the He-N pair allele. Figure 6D shows the results of an analysis of the detection limit for chromosome 6 mutations using the homozygous Ho-N pair allele in control DNA. (X-axis: alleles listed by chromosomal position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false negative allele; and each color: false negative result at the corresponding dilution ratio.) The dilution factors of control genomic DNA and H mole DNA are as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH50, 50:50; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; and CH1, 99:1. [Figure 7]Figure 7A shows the results of analyzing the frequency of false-positive mutations using an NGS panel from company AA using an NN pair allele in which both H mole DNA and control genomic DNA were null. Figure 7B shows the results of analyzing the frequency of false-positive mutations using an NGS panel from company BB. Figure 7C shows the results of analyzing the frequency of false-positive mutations using an NGS panel from company CC. (X-axis: alleles sorted by chromosomal position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count).) The dilution factors of control genomic DNA and H mole DNA were as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH50, 50:50; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; and CH1, 99:1. [Figure 8] Figure 8 shows the stringency cutoff values derived from the CC NGS panel results after removing false-positive mutations on chromosome 6 and using only the remaining false-positive mutations. (X-axis: alleles listed in order by chromosomal position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count).) The dilution factors of the control genomic DNA and H mole DNA were as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH50, 50:50; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; and CH1, 99:1. [Figure 9]Figure 9A shows the results of analyzing false-positive mutations derived from each company (in-house) and Dragen software (default, solid, and liquid). The graph shows the results (y-axis: arithmetic mean of false-positive mutations found across multiple dilution samples). Figure 9B shows the results of analyzing false-negative mutations detected by the NGS panels from AA, BB, and CC companies to measure sensitivity. The graph shows the results of analyzing false-negative mutations derived from each company's bioinformatic method (in-house) and Dragen software (default, solid, and liquid) under three conditions. The graph also shows the dilution mutation detection rate (y-axis). The x-axis shows the percentage of samples with mutations. Therefore, the percentage of mutations in each sample's total DNA is half the percentage of the mutant sample. For false positive error detection, if a variant or reference base that was not detected in the analysis of all alleles in the human mole DNA and control genomic DNA was found, it was determined to be a false positive error. DETAILED DESCRIPTION OF THE INVENTION

[0033] The present invention will be described in detail below.

[0034] The terms used in the present invention are currently commonly used and general terms that have been selected as much as possible in consideration of the functions of the present invention, but these terms may change depending on the intentions of engineers in the relevant technical field or the emergence of new technologies. In addition, in certain cases, arbitrarily selected terms may be used, and in such cases, their meanings will be described in detail in the description of the embodiments. Therefore, the terms used in the present invention should be defined based on the meanings of the terms and the overall content of the present invention, rather than simply by the names of the terms.

[0035] In the present invention, when an optional component or optional step is referred to as "comprising," this does not mean excluding other components or other steps, but means that other components or other steps may be further included, unless otherwise specified.

[0036] The present invention provides compositions for validation of next-generation sequencing (NGS) panels comprising homozygote DNA and control genomic DNA.

[0037] The term "next generation sequencing" used in this invention refers to a technology that can analyze millions of base sequences at once, and is also called massively parallel sequencing or high-throughput sequencing. In this invention, next generation sequencing and NGS are used interchangeably.

[0038] As used herein, the term "allele" refers to a gene that forms a pair of homologous chromosomes and has different traits, and refers to the DNA sequence of an allele. Homologous chromosomes are classified into homozygotes and heterozygotes, with "homozygotes" being those in which alleles with the same properties are combined, and "heterozygotes" being those in which alleles with different properties are combined. In the present invention, the terms "allele" and "allele" are used interchangeably.

[0039] In the present invention, the composition contains two different genomic DNAs for validating an NGS panel, one of which is homozygous DNA.

[0040] In the present invention, the homozygous DNA may be, but is not limited to, DNA isolated from a hydatidiform mole cell line or a parthenogenetic cell line.

[0041] As used herein, the term "control genomic DNA" refers to the other genomic DNA among the two different genomic DNAs contained in the composition, and may be either homozygous or heterozygous genomic DNA.

[0042] In the present invention, the control genomic DNA may be genomic DNA isolated from normal human blood. Alternatively, the control genomic DNA may be genomic DNA isolated from patient blood or cancer tissue. Alternatively, the control genomic DNA may be circulating tumor DNA (ctDNA). Alternatively, the control genomic DNA may be, but is not limited to, synthetic DNA or cloned DNA.

[0043] In the present invention, the composition may contain homozygous DNA and control genomic DNA at a dilution ratio selected from the group consisting of 1:9999 to 9999:1, and preferably, the dilution ratios of homozygous DNA and control genomic DNA may be selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5 and 9999:1, but are not limited thereto.

[0044] The present invention also provides a kit for verifying a next-generation base sequence analysis (NGS) panel, comprising the composition for verifying the next-generation base sequence analysis panel.

[0045] The present invention also provides a method for validating a next-generation sequencing (NGS) panel through analysis of false-negative mutations, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting a Ho-N pair allele (homozygote-null pair allele) or a He-N pair allele (heterozygote-null pair allele) from the results of the next-generation sequencing (NGS) obtained in step (a); and (c) analyzing the frequency of false-negative mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected Ho-N pair allele or He-N pair allele.

[0046] As used herein, the term "false negative" refers to a case where a test result that should actually be positive is mistakenly negative.

[0047] The present invention also provides a method for validating an NGS panel through a detection limit analysis, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) selecting a Ho-N pair allele (homozygote-null pair allele) or a He-N pair allele (heterozygote-null pair allele) from the NGS results obtained in step (a); (c) analyzing a dilution mutation detection rate, which indicates the number of diluted mutations found in the Ho-N pair allele or He-N pair allele as a percentage, in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios, based on the selected Ho-N pair allele or He-N pair allele; and (d) analyzing the detection limit.

[0048] As used herein, the term "Ho-N pair allele (Homozygote-Null pair allele)" refers to a mixed allele in which one allele site is a homozygous variant and a null with no mutation.

[0049] As used herein, the term "He-N pair allele (Heterozygote-Null pair allele)" refers to a mixed allele in which one allele site is a heterozygous variant and a null with no mutation.

[0050] The term "dilution mutation detection rate" as used herein means the percentage of diluted mutations found in Ho-N pair alleles or He-N pair alleles in a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios.

[0051] In the present invention, the method may be for validating a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variations (SNVs), insertions / deletions (Indels), and chromosomal amplifications / deletions.

[0052] As used herein, the term "variant" refers to a modification of a chromosome, gene, or base sequence that is genetically distinct from the wild type. In the present invention, the terms "mutation" and "variant" are used interchangeably.

[0053] In the present invention, the detection limit of step (d) may be expressed as a percentage of the VAF (variant allelic fraction) corresponding to the maximum dilution ratio that shows a 90% or 95% dilution mutation detection rate in a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios.

[0054] In the present invention, to validate the NGS panel, a composition for validating the NGS panel was used, which included a DNA mixture sample containing DNA isolated from a hydatidiform mole cell line (hereinafter referred to as "H mole DNA"), which is homozygous DNA, and control genomic DNA at various dilution ratios.

[0055] For example, if DNA1 contains a homozygous variant at one allele site and DNA2 contains a null with no mutation, DNA1 and DNA2 are mixed and the VAF (variant allelic fraction) is calculated, the dilution factors of DNA1 and DNA2 can be determined. The limit of detection can then be calculated by expressing the VAF (variant allelic fraction) corresponding to the maximum dilution factor that shows a 90% or 95% dilution mutation detection rate for each allele as a percentage.

[0056] More specifically, Figure 1 shows how the relative amounts of DNA in a DNA mixture containing a homozygous variant, DNA1, and a null, DNA2, can be measured at various dilutions. In this case, an allele containing a homozygous variant and a null at one allele site is defined as a "Ho-N pair allele (homozygote-null pair allele)." The probability of an Ho-N pair allele appearing among alleles with different VAFs (variant allelic fractions) is shown in Table 1 below.

[0057] [Table 1] *p = VAF (variant allelic fraction); q = 1-p; P1: The probability of an Ho-N pair allele appearing when both DNAs are common genomic DNA (P1 = p 2 q2 +q 2 p 2 =2p 2 q 2 ); and P2: When one of the two DNAs is H mole DNA, the probability of the Ho-N pair allele appearing (P2 = p 2 q+q 2 p=pq(p+q)=pq).

[0058] As shown in Table 1, when calculating and comparing the number of Ho-N pair alleles from alleles with nine different VAFs (variant allelic fraction, p) in a mixture of two DNAs, one of which is homozygous H mole DNA, and the other of which is general genomic DNA, the sum of the probability of an Ho-N pair allele appearing in the nine different VAF alleles is calculated to be 0.67 when both DNAs are general genomic DNA, while it is calculated to be 1.65 when one DNA is homozygous H mole DNA. In other words, when one DNA is homozygous H mole DNA, the number of Ho-N pair alleles can be increased by about 2.45 times (= P2 / P1 ratio) compared to when both DNAs are general genomic DNA.

[0059] In particular, the P2 / P1 ratio did not change when the probability of a mutation in each allele was calculated separately for when it was analyzed as a real mutation and when it was analyzed as a reference sequence. Furthermore, even when the frequency of the allele determined to be the reference allele was lower than the frequency of the variant allele, the P2 / P1 ratio was the same. Even when calculating the P2 / P1 ratio by taking into account the frequency at which a mutation appears in each allele distribution and integrating the p-value between 0 and 1, the P2 / P1 ratio was 2.50, which was almost unchanged (for example, assuming the frequency of allele 1 is 0.1 and the frequency of allele 9 is 0.9).

[0060] Therefore, in the present invention, when one of the two DNAs in a mixture is homozygous H mole DNA, it is possible to obtain approximately 2.5 times more Ho-N pair alleles than when one of the two DNAs is not homozygous DNA. Therefore, it was confirmed that it may be more useful to analyze the detection limit of an NGS panel using homozygous H mole DNA.

[0061] The present invention also provides a method for validating a next-generation sequencing (NGS) panel through analysis of false-positive mutations, the method comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting an NN pair allele (Null-Null pair allele) containing no mutation in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) obtained in step (a); and (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected NN pair allele.

[0062] As used herein, the term "NN pair allele (Null-Null pair allele)" refers to an allele in which all allele sites are null with no mutation.

[0063] As used herein, the term "false positive" refers to a case where a test result that should actually be negative is mistakenly shown as positive.

[0064] In the present invention, based on the NN pair allele where both the homozygous DNA and the control genomic DNA are null, if a mutation was detected in each DNA mixture sample where the homozygous DNA and the control genomic DNA were mixed at different dilutions, it was determined to be a false positive mutation.

[0065] In the present invention, the method may be for validating a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variations (SNVs), insertions / deletions (Indels), and chromosomal amplifications / deletions.

[0066] The present invention also provides a method for providing information for increasing the specificity of a next-generation sequencing (NGS) panel, comprising: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions using an NGS panel; (b) selecting an NN pair allele (Null-Null pair allele) containing no mutation in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) obtained in step (a); (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected NN pair allele; and (d) analyzing a variant allelic fraction (VAF) for removing false-positive mutations and deriving a stringency cutoff value.

[0067] In the present invention, the stringency cutoff value may be a 95% stringency cutoff value that removes 95% of false-positive mutations or a 99% stringency cutoff value that removes 99% of false-positive mutations.

[0068] In the present invention, the method may provide information on stringency cutoff values for eliminating false positive mutations to increase the specificity of next generation sequencing (NGS) panels.

[0069] The present invention will be described in more detail below with reference to examples. These examples are intended to explain the present invention more specifically, but the scope of the present invention is not limited to these examples.

[0070] Example 1. DNA mixture preparation and NGS analysis In this study, we compared the sensitivity and specificity of three companies' NGS panels using homozygous H mole DNA for NGS panel validation. H mole DNA derived from human hydatidiform moles was purchased from Coriell (NA07489, Camden, NJ). Control genomic DNA was extracted from the blood of a collaborator with approval from the Institutional Review Board of the National Cancer Center.

[0071] The prepared H mole DNA, control genomic DNA, and a mixture of H mole DNA and control genomic DNA were sent to companies AA, BB, and CC, respectively, and experiments were performed using each company's NGS panel on Illumina's Novaseq 6000 NGS equipment. The resulting FASTQ files were then analyzed for mutations according to each company's method, and the NGS panels were validated. However, the number of mutations obtained from each company's NS panel varied depending on the analysis method used by each company, in addition to the number of gene mutations. All companies commonly used the NCBI human genome assembly hg19 as a reference sequence, while company CC used LoFreq in addition to MuTect as a variant caller to analyze more mutations.

[0072] Example 2. NGS validation method 2.1. Selection of Ho-N pair alleles and He-N pair alleles The results for the mutations analyzed by each company were aligned based on chromosome number and position, along with sample information such as the reference base, variant allele information, VAF, read count, and mutation information. For the analysis results of AA and BB companies, samples with a read count of less than 100 were excluded from the analysis, while for CC company, samples with a read count of less than 300 were excluded from the analysis. Furthermore, for AA and BB companies, intron site SNVs and indel alleles were included in addition to exon site SNVs, while CC company included exon site SNVs and splice site SNVs in its analysis, but excluded intron site SNVs and indel mutations.

[0073] Among the selected alleles, pairs in which the alleles in the H mole DNA and the control genomic DNA consist of a homozygous variant and a null were defined as "Ho-N pair alleles (homozygote-null pair alleles)," and pairs in which the H mole DNA was null and the control genomic DNA consisted of a heterozygous variant were defined as "He-N pair alleles (heterozygote-null pair alleles)."

[0074] 2.2. Determination of false-negative mutations, dilution mutation detection rates, and detection limits using Ho-N pair alleles and He-N pair alleles Alleles in which mutations are present in homozygous H mole DNA or control genomic DNA, but no mutations are detected in the NGS panel analysis of the DNA mixture sample of H mole DNA and control genomic DNA, were defined as false negative alleles. The frequencies of false negative alleles were analyzed in the DNA mixture sample of H mole DNA and control genomic DNA, divided into Ho-N pair alleles and He-N pair alleles.

[0075] The dilution mutation detection rate was defined as the percentage of alleles in which mutations were found in Ho-N pair alleles or He-N pair alleles in DNA mixtures in which homozygous H mole DNA and control genomic DNA were mixed at different dilutions.

[0076] The limit of detection (LoD) was defined as the percentage of the variant allelic fraction (VAF) corresponding to the maximum dilution factor that resulted in 90% or 95% dilution mutation detection in a DNA mixture containing homozygous H mole DNA and control genomic DNA at different dilutions. Here, VAF is the read count of the sequence detected as a variant at a specific allele position divided by the total read count of the sequence detected as a variant and the reference sequence at that specific allele position.

[0077] 2.3. Determination of false positive mutations and stringency cutoff values Among null alleles in which no mutations were detected in either the homozygous H mole DNA or the control genomic DNA, a false positive was defined as a mutation detected in the NGS analysis of a DNA mixture sample in which H mole DNA and control genomic DNA were mixed at different dilutions. The VAF (variant allelic fraction) values that can remove 95% or 99% of these false positive mutations were defined as the 95% and 99% stringency cutoff values, respectively.

[0078] Example 3. Sensitivity analysis results for NGS panel using Ho-N pair allele Among all mutation sites identified from the results of AA Company, BB Company, and CC Company, the median read counts of the mutant alleles were confirmed to be 197 (Q1, Q3; 51, 464), 813 (433, 1189), and 679 (577, 717), respectively. Meanwhile, the median read counts in the samples for the Ho-N pair allele and He-N pair allele of the present invention were confirmed to be 380 (Q1, Q2; 221, 633) for AA Company, 901 (Q1, Q; 503, 1234) for BB Company, and 703 (Q1, Q2; 679, 727) for CC Company, respectively.

[0079] First, we analyzed Ho-N pair alleles to evaluate the sensitivity of the three companies' NGS panels. Because Company AA's mutation analysis results showed few alleles in exon sites, we analyzed sensitivity and specificity including intron site mutations. After selecting Ho-N pair alleles from Company AA's analysis results, 15 (10.2%, 15 / 147) were confirmed to be exon site mutations, and the remaining were confirmed to be intron site mutations or UTR site mutations.

[0080] In addition, the mutation analysis results from BB Company included the sent SNVs and indels, but they also analyzed sensitivity and specificity for all exon and intron mutations, selecting Ho-N pair alleles, and confirming that 134 (89.9%, 134 / 149) were exon site mutations or splice site mutations. In the mutation analysis results from CC Company, the number of analyzed alleles was relatively high, so they selected only exon site mutations and splice site mutations, removed indels, and analyzed only SNVs, confirming that there were a total of 306 Ho-N pair alleles.

[0081] We analyzed the sensitivity of each company's NGS panel using Ho-N pair alleles. First, the AA company's NGS panel detected all SNVs in a DNA mixture containing 50% H-mol DNA and 50% control genomic DNA (CH50). It also detected 134 of 147 Ho-N pair alleles in a DNA mixture containing 80% H-mol DNA and 20% control genomic DNA (CH80), confirming a dilution mutation detection rate of 91.2% (Figure 2A). Furthermore, it detected only 16 SNVs in a DNA mixture containing 90% H-mol DNA and 10% control genomic DNA (CH90), confirming a dilution mutation detection rate of 10.9% (Figure 2A). This indicates that the maximum dilution factor required for 90% dilution mutation detection rate for the AA company's NGS panel is approximately 20%, which means that the detection limit for 90% dilution mutation detection rate is approximately 20%.

[0082] The BB company's NGS panel detected all 149 Ho-N pair alleles in 10% DNA mixture samples (i.e., DNA mixtures containing 90% H-mol DNA and 10% control genomic DNA (CH90) and 10% H-mol DNA and 90% control genomic DNA (CH10)), and 137 of 149 Ho-N pair alleles in 5% DNA mixture samples (i.e., DNA mixtures containing 95% H-mol DNA and 5% control genomic DNA (CH95) and 5% H-mol DNA and 95% control genomic DNA (CH5)), confirming a dilution mutation detection rate of approximately 88.6% (Figures 2B and 2C). This indicates that the maximum dilution factor for 90% dilution mutation detection rate for the BB company's NGS panel is between 5% and 10%, and even closer to 5%, meaning that the detection limit for 90% dilution mutation detection rate is approximately 5%.

[0083] According to CC's results, some genes on chromosome 6 had many false negatives and were therefore excluded from the detection limit analysis. Chromosome 6 contains 91 Ho-N pair alleles, and excluding these, 215 Ho-N pair alleles were analyzed. The CC company's NGS panel detected all 215 Ho-N pair alleles in the 2.5% DNA mixture samples (i.e., DNA mixtures containing 97.5% H-mol DNA and 2.5% control genomic DNA (CH97.5) and 2.5% H-mol DNA and 97.5% control genomic DNA (CH2.5)), and 197 of the 215 Ho-N pair alleles in the 1% DNA mixture samples (i.e., DNA mixtures containing 99% H-mol DNA and 1% control genomic DNA (CH99) and 1% H-mol DNA and 99% control genomic DNA (CH1)), confirming a dilution mutation detection rate of approximately 91.6% (Figures 2D and 2E). This confirms that the maximum dilution factor required for a 90% dilution mutation detection rate for the CC company's NGS panel is 1%, which means that the detection limit for a 90% dilution mutation detection rate is approximately 1%.

[0084] From the above results, it was confirmed that by selecting Ho-N pair alleles using a DNA mixture containing the H mole DNA of the present invention, the detection limit, which is expressed as a percentage of the VAF (variant allelic fraction) corresponding to the maximum dilution ratio showing a 90% dilution mutation detection rate, can be analyzed, and that each company's NGS panel can be verified through analysis of the frequency of false-negative mutations and analysis of the detection limit.

[0085] Example 4. Sensitivity analysis results for NGS panel using He-N pair alleles We investigated the sensitivity of the He-N pair alleles, where the two DNA mixtures consisted of null and heterozygous variants. AA's analysis yielded 269 He-N pair alleles (53 exonic variants, 19.7% of the total). AA's NGS panel detected 266 of the 269 He-N pair alleles in a 50% DNA mixture (i.e., a DNA mixture containing 50% H mole DNA and 50% control genomic DNA (CH50)), confirming a dilution mutation detection rate of 98.9% (Figure 3A). This means that the mutant allele with a VAF (variant allelic fraction) of 0.25 was detected 98.9%, with a detection limit of 25% or less. These results were similar to those obtained using the Ho-N pair alleles described above.

[0086] The BB Company analysis yielded 172 He-N pair alleles (147 exonic or splicing variants, 85.5% of the total). The BB Company NGS panel detected 166 of the 172 He-N pair alleles in 10% DNA mixture samples (i.e., a DNA mixture containing 90% H mole DNA and 10% control genomic DNA (CH90) and a DNA mixture containing 10% H mole DNA and 90% control genomic DNA (CH10)), confirming a dilution mutation detection rate of 96.5% (Figure 3B). This means that 96.5% of mutant alleles with a VAF (variant allelic fraction) of 0.05 were detected, with a detection limit of less than 5%. These results were similar to those obtained using the Ho-N pair alleles described above.

[0087] The CC analysis yielded 276 He-N pair alleles (all exon site SNVs), excluding 54 He-N pair alleles on chromosome 6. The CC NGS panel detected 246 of the 276 He-N pair alleles in 2.5% DNA mixture samples (i.e., a DNA mixture containing 97.5% H-mol DNA and 2.5% control genomic DNA (CH97.5) and a DNA mixture containing 2.5% H-mol DNA and 97.5% control genomic DNA (CH2.5)), confirming a dilution mutation detection rate of 89.1% (Figure 3C). This means that 89.1% of mutant alleles with a VAF (variant allelic fraction) of 0.0125 were detected, with a detection limit of 1.25%. These results were similar to those obtained using the Ho-N pair alleles described above.

[0088] For reference, in addition to SNVs, AA and BB companies contained 38 and 8 indels in the Ho-N pair alleles or He-N pair alleles, respectively, which had little effect on the calculation of the detection limit (Figure 4).

[0089] From the above results, the present invention confirmed that by analyzing the Ho-N pair allele and the He-N pair allele and determining the maximum dilution ratio showing a 90% dilution mutation detection rate, the detection limit for each NGS panel can be analyzed as an objective result, and can be useful for validating NGS panels. Furthermore, it was confirmed that using both the results of the Ho-N pair allele and the He-N pair allele has the advantage of being able to obtain sensitivity information for more alleles than using only the Ho-N pair allele.

[0090] Example 5. Relationship between detection limit and read count or VAF (variant allelic fraction) Previous literature has reported that when detecting SNVs, a read count (or local sequence coverage) of 350 can detect alleles with a VAF (variant allelic fraction) of less than 5% and 5-10%, respectively, at 89.4% and 99.2%, and when a read count of 738 can detect alleles with a VAF (variant allelic fraction) of less than 5% and 5-10%, respectively, at 98.5% and 99.7%, and that read count and VAF (variant allelic fraction) are related to the sensitivity or limit of detection of mutations (Frampton et al., Nature Biotechnology 2013).

[0091] However, in the actual analysis results, the median read count of AA company was about 380, but it detected only 10.9% of mutations with a VAF (variant allelic fraction) of 10% and 91.2% of mutations with a VAF of 20%, indicating significantly low sensitivity for AA company's NGS panel. Furthermore, the read count of BB company was about 900, and it detected 88.6% of mutations with a VAF (variant allelic fraction) of 5%, showing similar detection limits to those of the previous literature.

[0092] However, CC's NGS panel demonstrated a 100% detection rate for mutations with a read count of approximately 700 and a variant allelic fraction (VAF) of 2.5%, and a 91.6% detection rate for mutations with a VAF of 1%, demonstrating that the detection limit for CC's NGS panel was superior to the results of the previous literature. These results also suggest that there are various variables that determine the sensitivity of an NGS panel, in addition to the relationship between read count and variant allelic fraction (VAF). Therefore, it is not possible to evaluate the sensitivity or detection limit of an NGS panel solely based on read count and variant allelic fraction (VAF). Researchers must verify the sensitivity of a particular NGS panel using their own standard mixtures to assess whether the panel has the sensitivity required to achieve their research objectives.

[0093] Example 6. Sensitivity analysis results for NGS panel using Ho-He pair allele Next, we defined a pair consisting of a homozygous variant in H mole DNA and a heterozygous variant in the control genomic DNA as a "Ho-He pair allele (homozygote and heterozygous pair allele)" and obtained the reference allele fraction (RAF) value, which is the fraction occupied by the diluted reference allele compared to the variant allele. We confirmed whether this value could be used to analyze the detection limit.

[0094] In the RAF analysis, the AA NGS panel detected 94.9% (131 / 138) Ho-He pair alleles with a 2.5% RAF (reference allelic fraction) in a 5% DNA mixture (i.e., a DNA mixture containing 95% H-mol DNA and 5% control genomic DNA (CH95)) (Figure 5A). The BB NGS panel detected 96.7% (116 / 120) Ho-He pair alleles with a 2.5% RAF (reference allelic fraction) in a 5% DNA mixture (i.e., a DNA mixture containing 95% H-mol DNA and 5% control genomic DNA (CH95) and a DNA mixture containing 5% H-mol DNA and 95% control genomic DNA (CH5)) (Figure 5B). In addition, CC's NGS panel detected 97.8% (221 / 226) Ho-He pair alleles with 0.5% RAF in 1% DNA mixture samples (i.e., a DNA mixture containing 99% H mole DNA and 1% control genomic DNA (CH99) and a DNA mixture containing 1% H mole DNA and 99% control genomic DNA (CH1)) (Figure 5C).

[0095] As mentioned above, the results of the detection limit analysis using the Ho-He pair allele confirmed that the detection rate of the reference sequence base diluted at the maximum dilution factor requested by each company was 90% or higher. In addition, no false-negative alleles on chromosome 6 were clearly detected in the CC company's NGS panel.

[0096] As described above, the results of the analysis using the Ho-He pair allele were completely different from the results of the dilution mutation detection analysis using the Ho-N pair allele or He-N pair allele. Therefore, we confirmed that the method of analyzing the detection limit of the reference sequence base using the reference allelic fraction (RAF) value of the Ho-He pair allele is not appropriate for evaluating NGS panels.

[0097] Example 7. Additional analysis results of detection limits for chromosome 6 CC company's NGS panel was found to have a serious problem with a high number of false negatives on chromosome 6, and analysis using the He-N pair allele was relatively less than analysis using the Ho-N pair allele. Therefore, the chromosome 6 sites with a high number of false negatives were matched with the chromosomal location of the Ho-N pair allele or He-N pair allele, and then it was confirmed whether there was a difference in the analysis results using the Ho-N pair allele and the He-N pair allele.

[0098] As a result, we confirmed that there are specific gene sites within chromosome 6 where false negatives are more common, and that the He-N pair allele contains fewer of these gene sites, resulting in fewer false negatives than the Ho-N pair allele (Figure 6). The existence of specific chromosomal sites associated with false negatives, as in CC company NGS panels, means that we must consider the possibility of false negatives even for mutations with a high VAF (variant allelic fraction).

[0099] From the above results, the method of the present invention has the advantage that it can determine the final results by taking into account the possibility of false negatives at specific chromosomal sites or specific gene sites that are particularly prone to false negatives during NGS panel validation in advance, and further confirmed that by removing such sites from the design of new NGS panels, it is possible to reduce errors in NGS panel analysis related to false negatives that occur in NGS panels.

[0100] Example 8. Analysis of false positive mutations and stringency cutoff values for NGS panels As mentioned above, in the null allele where no mutations were found in either H mole DNA or control genomic DNA (defined as the "NN pair allele"), if a mutation was found in the NGS panel analysis of a DNA mixture sample where H mole DNA and control genomic DNA were mixed at different dilutions, this was defined as a false positive mutation, and the detection frequency of false positive mutations in each company's NGS panel was analyzed.

[0101] As a result, the NGS panels from AA and BB companies yielded relatively fewer false-positive mutations than those from CC company (Figure 7). This is likely due to the VAF values, which indicate the detection limit of the AA and BB company NGS panels, being relatively higher than that of CC company, which means that the sensitivity at which mutations can be detected is reduced. However, it was confirmed that the NGS panel from CC company yielded relatively fewer false-positive mutations than the NGS panels from AA and BB companies when the VAF (variant allelic fraction) was 0.1 or higher, except for chromosome 6 (Figure 7).

[0102] For reference, two issues were found in CC's NGS panel analysis. First, a large number of false-positive mutations were found in the 10% DNA mixture sample (i.e., a DNA mixture (CH10) containing 10% H mole DNA and 90% control genomic DNA). However, if the experiment had been conducted under similar conditions as the other samples, a similar number of false-positive mutations should have been found. However, because CH10 alone showed approximately 10 times more false-positive mutations, we concluded that this was due to an error in the experimental process using CC's NGS equipment. Therefore, the false-positive mutations found only in CH10 were excluded from Figure 7C. Second, the VAF of false-positive mutations for chromosome 6 alleles in CC's NGS panel was significantly lower or higher than that of false-positive mutations at the remaining chromosome sites, as can be seen in Figure 7C. Therefore, the alleles on chromosome 6 were excluded from the false-positive mutation detection analysis.

[0103] Subsequently, as mentioned above, the 95% or 99% stringency cutoff value was derived, defined as the VAF value that can eliminate 95% or 99% of false-positive mutations. To date, most researchers have determined that only mutations with a certain VAF value or higher are "true mutations" after NGS panel analysis in order to eliminate false-positive mutations, and have referred to this VAF (variant allelic fraction) value as the "stringency cutoff value." However, no clear criteria for scientifically determining this have been presented.

[0104] The present invention clearly presents a method for deriving a VAF value that can remove 95% or 95% of false-positive mutations as a "stringency cutoff value." To confirm this, the distribution of VAF values of false-positive mutations was analyzed for the NN pair alleles of CC companies. In this case, the alleles corresponding to the two problems with the NGS panel of the CC companies mentioned above (i.e., 1. large-scale false-positive mutations appear in the CH10 sample; and 2. severe changes in VAF due to false-positive mutations on chromosome 6) were excluded from the analysis, and then the stringency cutoff value was derived.

[0105] As a result, the number of false-positive mutations obtained from the CC company's NGS panel was ultimately 1,977, and the VAF (variant allelic fraction) value required to remove 95% of false-positive mutations was 0.067, while the VAF (variant allelic fraction) value required to remove 99% of false-positive mutations was 0.0875 (Figure 8). In other words, when analyzing the CC company's NGS panel to remove 95% of false-positive mutations, the stringency cutoff value was calculated as 0.067, and when analyzing the CC company's NGS panel to remove 99% of false-positive mutations, the stringency cutoff value was calculated as 0.087.

[0106] From the above results, it was confirmed that the present invention can verify an NGS panel by analyzing the frequency of false-positive mutations, provide information on objective stringency cutoff values for eliminating false-positive mutations to increase the specificity of an NGS panel, and reduce errors in analyzing NGS results associated with false-positives occurring in NGS panels.

[0107] Example 9. Evaluation of errors contained in FASTQ files and bioinformatics analysis errors in each company's NGS panel analysis results The errors found in the analysis of the three companies' NGS panels can be divided into (1) errors inherent in the raw data (i.e., FASTQ files) generated using the NGS panels and (2) errors inherent in the bioinformatic analysis process used to identify mutations in each of the three companies' raw data. To confirm the contribution of these two errors, we used commercial software to identify mutations in the raw data and then used this to confirm the contribution of the two errors. For bioinformatic analysis, we used Illumina's Dragen software. While this process requires a BED file from each company, companies AA and CC did not provide BED files. Therefore, we analyzed only the mutations in the exons of the genes provided by those two companies. For company BB, we analyzed all mutations in the entire captured region using the BED file provided by the company.

[0108] False-negative errors in the mutation region were assessed as described above. However, for false-positive error assessment, any variant or reference base not detected in the analysis of H mole DNA and control genomic DNA in all alleles, including the Ho-N pair, He-N pair, and Ho-He pair, as well as the NN pair allele, was assessed as a false-positive error. The frequency of false-positive mutation detection was then analyzed using each company's NGS panel.

[0109] Figure 9 shows the false negatives and false positives obtained by analyzing mutation results provided by the three companies' bioinformatics methods, as well as the results of analyzing the false negatives and false positives obtained by analyzing mutations from each company's raw data (FASTQ files) using Dragen software. Mutation analysis was performed using Dragen software under three conditions: default, solid, and liquid. Default is a commonly used condition, while solid and liquid are conditions with increased sensitivity compared to the default condition.

[0110] Figure 9A shows the results of analyzing the false-positive mutations detected by each company (in-house) and Dragen software (under three conditions: default, solid, and liquid). The y-axis in Figure 9A shows the arithmetic mean of false-positive mutations detected across multiple dilutions. Of the three Dragen software analysis conditions, the default condition yielded fewer false-positives than the solid or liquid conditions (companies BB and CC had fewer than 30 false-positives, indicated by the blue line). Considering that the default Dragen software analysis conditions have lower sensitivity than the other two conditions, these results are consistent with the general phenomenon of increasing specificity as sensitivity decreases.

[0111] However, mutation analysis using AA Company's in-house method yielded few false positive mutations (average 0.5), whereas analysis using Dragen software's default conditions resulted in many false positives (Figure 9A, square; average false positives: 818). For this reason, as discussed below, we suspect that AA Company's NGS panel analysis generated many FP errors. This suggests that the conditions were changed to increase specificity at the expense of sensitivity. Furthermore, CC Company's in-house method was found to generate many false positives (Figure 9A, arrow; average false positives: 1680). This is believed to be related to CC Company's intentionally increased sensitivity in its mutation analysis, as discussed below.

[0112] Figure 9B shows the sensitivity of the mutation detection results of the NGS panels from companies AA, BB, and CC, analyzed for false negatives. The results of the analysis of false negatives for mutations derived from each company's bioinformatics method (in-house) and the analysis of false negatives for mutations derived from Dragen software's bioinformatics method under three conditions (default, solid, and liquid) are shown in a graph of the dilution mutation detection rate (y-axis). The percentage on the x-axis in Figure 9B represents the percentage of samples with heterozygote mutations. Therefore, the percentage of mutations in the total DNA of each sample is half the percentage of mutant samples displayed on the x-axis (e.g., for a sample displayed as 5% on the x-axis in Figure 9B, the actual percentage of mutations is 2.5%). Figure 9B shows only the results using the Dragen software's bioinformatics method under the default conditions. However, to analyze the mutation detection rate using Dragen software in Figure 9B, for companies AA and CC, only mutations present in the exons of the target genes provided by each company were analyzed. For company BB, mutations present in the sites captured by company BB's NGS panel were analyzed using the bed file. All sensitivity analysis results in Figure 9B were obtained using the He-N pair allele.

[0113] The sensitivity analysis results in Figure 9B show that AA's bioinformatics method detected almost no mutations in samples containing 20% heterozygote variant DNA. This indicates that the mutation detection limit of AA's NGS panel is greater than 10%. Looking at the Ho-N pair allele analysis results, the final results of AA's raw data and their bioinformatics method analysis indicate a mutation detection limit of only about 20%. However, analyzing mutations obtained using Dragen software's default conditions revealed that the sensitivity mutation detection limit was approximately 5%, as most mutations present in samples containing 10% heterozygote variant DNA were detected. Therefore, the bioinformatics method used by AA was considered to have selected conditions with low sensitivity. This lowering of the sensitivity is thought to be related to the discovery of a very large number of false positive mutations (818) when analyzing mutations obtained under the default conditions of the Dragen software in Figure 9A. Ultimately, it can be understood that AA Company changed the analysis conditions to sacrifice sensitivity and increase specificity (changing the conditions so that the sensitivity was around 20%).

[0114] The sensitivity analysis results for CC Company (Figure 9B) show that the analysis results using their bioinformatics conditions detected almost all dilution mutations in the 2.5% heterozygote variant DNA sample, indicating a sensitivity of approximately 1.25% mutations. Analysis of mutations obtained using Dragen software's default conditions detected only approximately 50% of dilution mutations in the 2.5% heterozygote variant DNA sample, and almost all dilution mutations in the 5% heterozygote variant DNA sample, indicating a mutation detection limit of approximately 2.5%. Therefore, CC Company's bioinformatics method appears to have a higher sensitivity than the Dragen software analysis. However, the specificity analysis results (Figure 9A) show that CC Company's in-house bioinformatics conditions detected 1,680 false positive mutations, indicating that their bioinformatics conditions intentionally increased sensitivity, resulting in a significant decrease in specificity.

[0115] Summarizing the above results, it can be concluded that AA Company had many errors in the process of generating raw data, but that this company adjusted its analytical conditions to sacrifice sensitivity in order to solve the problem of frequent false positive mutations during the bioinformatics analysis. BB Company generated more false negative dilution mutations during its bioinformatics analysis compared to mutation analysis using Dragen software, and as a result, the sensitivity of the NGS panel analyzed by the company was slightly lower than that of Dragen software analysis. CC Company adjusted its analytical conditions to increase sensitivity during the bioinformatics analysis, resulting in frequent false positives.

[0116] As described above, analyzing the results of the present invention using commercial software has the advantage of being able to identify errors in the NGS panel analysis process of a specific company separately into (1) errors in the process of generating raw data and (2) errors in the bioinformatics analysis process.

Claims

1. A composition for validation of a next-generation sequencing (NGS) panel comprising homozygous DNA and control genomic DNA.

2. 2. The composition of claim 1, wherein the homozygous DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic cell line.

3. The composition of claim 1 , wherein the control genomic DNA is genomic DNA isolated from normal human blood.

4. The composition of claim 1 , wherein the control genomic DNA is genomic DNA isolated from the patient's blood or cancer tissue.

5. The composition of claim 1 , wherein the control genomic DNA is circulating tumor DNA (ctDNA).

6. The composition of claim 1 , wherein the control genomic DNA is synthetic or cloned DNA.

7. 2. The composition of claim 1, wherein the dilution ratio of the homozygous DNA and the control genomic DNA is selected from the group consisting of 1:99 to 99:

1.

8. 2. The composition of claim 1, wherein the dilution factors of the homozygous DNA and the control genomic DNA are selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5 and 9999:

1.

9. A kit for verifying a next-generation base sequence analysis (NGS) panel, comprising a composition for verifying a next-generation base sequence analysis (NGS) panel according to any one of claims 1 to 8.

10. (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using a next-generation sequencing (NGS) panel; (b) selecting Ho-N pair alleles (homozygote-null pair alleles) or He-N pair alleles (heterozygote-null pair alleles) from the results of the next-generation sequencing (NGS) performed in step (a); and (c) A method for validating a next-generation sequencing (NGS) panel through analysis of false negative mutations, the method comprising: analyzing the frequency of false negative mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected Ho-N pair allele or He-N pair allele.

11. (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using a next-generation sequencing (NGS) panel; (b) selecting Ho-N pair alleles (homozygote-null pair alleles) or He-N pair alleles (heterozygote-null pair alleles) from the results of the next-generation sequencing (NGS) performed in step (a); (c) analyzing the dilution mutation detection rate, which indicates the percentage of diluted mutations found in the Ho-N pair allele or He-N pair allele, in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios based on the selected Ho-N pair allele or He-N pair allele; and (d) A method for validating a next-generation sequencing (NGS) panel through analysis of the limit of detection, including a step of analyzing the limit of detection.

12. The validation method according to claim 11, wherein the method validates a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variations (SNVs), insertions / deletions (Indels), and chromosomal amplifications / deletions (amplifications / deletions).

13. 12. The method of claim 11, wherein the detection limit in step (d) is a percentage of a variant allelic fraction (VAF) corresponding to a maximum dilution ratio that shows a 90% or 95% dilution mutation detection rate in a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios.

14. (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using a next-generation sequencing (NGS) panel; (b) selecting N-N pair alleles (Null-Null pair alleles) that do not contain mutations in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) performed in step (a); and (c) A method for validating a next-generation sequencing (NGS) panel through analysis of false positive mutations, the method comprising: analyzing the frequency of false positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected N-N pair allele.

15. The validation method according to claim 14, wherein the method validates a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variations (SNVs), insertions / deletions (Indels), and chromosomal amplifications / deletions (amplifications / deletions).

16. (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios using a next-generation sequencing (NGS) panel; (b) selecting N-N pair alleles (Null-Null pair alleles) that do not contain mutations in the homozygous DNA and the control genomic DNA from the results of the next-generation sequencing (NGS) performed in (a); (c) analyzing the frequency of false positive mutations in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilutions based on the selected N-N pair allele; and (d) A method for providing information to increase the specificity of a next-generation sequencing (NGS) panel, comprising the step of analyzing a variant allelic fraction (VAF) that removes false positive mutations and deriving a stringency cutoff value.

17. 17. The information providing method according to claim 16, wherein the stringency cutoff value is a 90% stringency cutoff value that removes 90% of false-positive mutations, a 95% stringency cutoff value that removes 95% of false-positive mutations, or a 99% stringency cutoff value that removes 99% of false-positive mutations.

18. 17. The information providing method of claim 16, wherein the method provides information on stringency cutoff values for removing false positive mutations to increase the specificity of a next generation sequencing (NGS) panel.

19. (a) collecting analysis results for false negative and false positive mutations from the company's raw data using the company's bioinformatics analysis results; and (b) a step of comparing and analyzing the bioinformatic analysis results from the company's raw data using commercially available software with the analysis results of step (a); A method for evaluating errors in the process of generating raw data or errors in the bioinformatic analysis process in a company's NGS panel analysis process, the method comprising the steps of:

Citation Information

Patent Citations

  • Method for measuring the copy number of chromosomes, genes, or specific nucleotide sequences using SNP arrays

    JP2011510626A

  • Linear DNA assembly for nanopore sequencing

    WO2021108532A2

  • Synthetic polynucleotides and method of use thereof in genetic analysis

    WO2022235315A1