Validation method for next-generation DNA sequence analysis panels
Patent Information
- Application Number
- JP2025504509
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-07-28
- Filing Date
- 2023-07-27
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-07-27
AI Technical Summary
【0031】 本発明による組成物は、NGSパネル検証に有用に使用され得、特に、本発明による検証方法は、(i)NGSパネルに対する偽陰性変異の頻度、偽陽性変異の頻度および検出限界を客観的結果として分析することができ、(ii)NGSパネルの検証時に偽陰性が特に多く出る特定の染色体部位または特定の遺伝子部位の有無をあらかじめ判別することによって、偽陰性が多い部位を分析時に除去することができ、NGSパネルに発生する偽陰性に関連したNGS結果の解析に対するエラーを減らすことができ、(iii)偽陽性変異の数とVAF(variant allelic fraction)の分布を分析して客観的な厳密性カットオフ値(stringency cutoff value)を提供することができ、NGSパネルに発生する偽陽性と関連したNGS結果の解析に対するエラーを減らすことができる。
Smart Images

Figure 0007914330000002 
Figure 0007914330000003 
Figure 0007914330000004
Abstract
Description
Technical Field
[0001] The present invention relates to a composition for validation of a next-generation sequencing panel; a kit for validation of a next-generation sequencing panel comprising the same; and a method for validation of a next-generation sequencing panel.
Background Art
[0002] Next generation sequencing (hereinafter referred to as "NGS") is a technology that can analyze millions of base sequences at once, and is also called massively parallel sequencing or high-throughput sequencing. The previously used Sanger sequencing method required several years and billions of dollars to analyze a single human genome, but with the NGS method, analysis can be completed within several days at a cost of about 1,000 US dollars. Currently, NGS technology is applied to the screening, diagnosis, prognosis prediction, and therapeutic selection of various diseases including not only cancer but also genetic diseases and infectious diseases. In the United States alone, more than 55,000 NGS diagnostic panels for more than 11,000 diseases using NGS technology are used clinically.
[0003] In order to apply a developed NGS panel to clinical use, it is necessary to evaluate whether the panel can accurately produce desired results. Accordingly, the US Food and Drug Administration (FDA) has issued guidelines for the validation of NGS panels, and the Centers for Medicare & Medicaid Services (CMS) monitors such panels, and the guidelines and monitoring methods are continuously updated.
[0004] Currently, the most widely used method for validating NGS panels is the method introduced by Frampton et al. in Nature Biotechnology in 2013 (hereinafter referred to as the "Frampton method"), and the Frampton method has been the primary method used in subsequent studies (Shin et al., Nature Comm 2017; and Froyen et al., Cancers 2022). The Frampton method is a method for validating NGS panels using a DNA pool containing DNA from HapMap cell lines with well-known base sequences, or a DNA pool containing DNA from cancer cell lines. Specifically, in the Frampton method, the DNA from 10 HapMap cell lines is mixed in equal amounts to create a DNA mixture with a DNA content of 5%, which is called pool XX. Then, a DNA mixture is created using 10 other cell lines, which is called pool YY, and this is analyzed with an NGS panel to evaluate sensitivity and specificity (Frampton et al., Nature Biotechnology 2013). While it is possible to calculate the proportion of each single nucleotide variant (SNV) allele (hereinafter referred to as "variant allelic fraction" or VAF) in a DNA mixture, haplotype alleles that appear only in one cell line account for 5% of the total mixture, while the remaining SNVs that appear several times have a VAF of 10% or more. When Frampton et al. analyzed a panel using a mixture of HapMap cell lines selected by them, 20% of the alleles in the panel's genes had approximately 5% VAF, and approximately 30% of the alleles had 10% VAF. Therefore, 20% of the alleles could be used to measure the detection rate of 5% VAF, and 30% of the alleles could be used to measure the detection rate of 10% VAF. However, in Frampton et al.'s paper, the HapMap cell line mixture created by the researchers did not contain any alleles with less than 5% VAF, so it was not possible to measure the detection rate of alleles with less than 5% VAF.Furthermore, in order to measure the detection rate of alleles with 1% VAF using the Frampton method, at least 50 HapMap cell lines must be used, and mutations with a detection frequency of about 1% must be used. In addition, in Frampton's paper, about 200 alleles with about 5% VAF were obtained using 10 HapMap cell lines. However, it can be estimated from this that only 40 or fewer alleles with 1% VAF can be obtained, resulting in the disadvantage that the sensitivity of the NGS panel must be evaluated with a small number of alleles.
[0005] Furthermore, when sensitivity was evaluated using the Frampton method, sensitivity was closely related to (1) the down-sampled coverage of the in silico model corresponding to the read count, and (2) the VAF in the allele. More specifically, SNVs with 5% VAF could be detected at 96.1% with 500 coverage, 83.5% with 300 coverage, and 44.2% with 150 coverage. SNVs with 10% VAF could be detected at 99.3% with 500 coverage, 98.9% with 300 coverage, and 82.3% with 150 coverage. SNVs with VAF greater than 10% could be detected at 99.5% even with 150 coverage. Furthermore, specificity was evaluated using the Framton method. According to the results of the published paper, two SNVs with a fraction of less than 5% were detected only when the downsampling coverage was greater than 500. In the remaining cases, no false positive mutations were detected at all, showing a positive predictive value of over 99.9%, and the specificity of the NGS panel used was reported to be very high.
[0006] Furthermore, to detect indels appearing in cancer cell lines using the Frampton method, 41 different pools were created by mixing 2 to 10 DNA samples from cancer cell lines, and the sensitivity and specificity of indel detection were measured. It was confirmed that the indel detection sensitivity also changes depending on (1) downsampling coverage and (2) the VAF of indel mutations within the mixture. Specifically, indels with 10% VAF were detected at 65% with 190 coverage and 88.3% with 667 coverage. Indels with 20% VAF were detected at 86.3% with 190 coverage and 97.3% with 667 coverage. In the case of false-positive indels, no false-positive indels with more than 20% VAF were detected, but indels with less than 20% VAF, approximately 1-2% were detected as false-positive indels. The sensitivity and specificity of gene amplification / deletion detection were measured using the same method with DNA from various cancer cell lines. No false-positive mutations were found, but false-negative mutations were detected in approximately 10-20% of cases.
[0007] As described above, the Frampton method mixes approximately 10 DNA molecules in equal amounts during the dilution process, but errors occur frequently, making it difficult to evaluate the sensitivity of NGS panels and failing to provide accurate information on how many diluted mutations can be detected at a particular dilution ratio. In addition to these limitations of the Frampton method, the detection of indels and amplifications / deletions is less sensitive than SNV detection, and sensitivity and specificity vary depending on variables such as allelic bias, duplicate reads, and the power of the analytical pipeline, making it difficult to specify and set the sensitivity and specificity of NGS panels in guidelines (Shin et al., Nature Comm 2017). Therefore, the current guidelines from the American College of Medical Genetics and Genomics for next-generation sequencing-based assays do not specify a particular coverage threshold or minimum sensitivity or specificity (US Food and Drug Administration (FDA). Considerations for Design, Development, and Analytical Validation of Next Generation Sequencing-Based In Vitro Diagnostics Intended to Aim in the Diagnosis of Suspected Germline Diseases. Updated 13 April 2018).
[0008] However, VAF information, which can be detected as a preferred minimum read count at the target site, and Limit of Detection (LoD) information, obtained through dilution experiments of standard samples, are provided (Rehm et al., Genet Med, 2013; ACMG clinical laboratory standards for next-generation sequencing). In this situation, regulations for NGS methods are still not clear, but the ACLA (American Clinical Laboratory Associations) oversees NGS panels currently used in genetic testing through the CLIA (Clinical Laboratory Improvement Amendments) regulations, and the FDA has established guidelines for the design, development, and analytical validation of such panels, enabling NGS to be applied clinically. Therefore, if new standards and methods are developed that can measure the sensitivity and specificity of NGS panels, researchers will be able to easily compare and evaluate NGS panels, and the minimum requirements for clinical application may also be revised. [Overview of the project] [Problems that the invention aims to solve]
[0009] The object of the present invention is to provide a composition for verifying next-generation sequencing (NGS) panels, which include homozygote DNA and control genomic DNA.
[0010] Another object of the present invention is to provide a validation kit for a next-generation sequencing (NGS) panel, comprising a validation composition for the next-generation sequencing (NGS) panel.
[0011] A further object of the present invention is to provide a method for validating a next-generation sequencing (NGS) panel through the analysis of false negative mutations, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the NGS results obtained in step (a); and (c) analyzing the frequency of false negative mutations in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios, using the selected Ho-N pair alleles or He-N pair alleles as a reference.
[0012] A further object of the present invention is to provide a method for validating a next-generation sequencing (NGS) panel, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the NGS results obtained in step (a); (c) analyzing the dilution mutation detection rate, which is expressed as a percentage, of the number of diluted mutations found in the Ho-N pair alleles or He-N pair alleles in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios, based on the selected Ho-N pair alleles or He-N pair alleles; and (d) analyzing the limit of detection.
[0013] A further object of the present invention is to provide a method for validating a next-generation sequencing (NGS) panel through the analysis of false positive mutations, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting NN pair alleles (Null-Null pair alleles) that do not contain mutations in the isozygous DNA and control genomic DNA from the next-generation sequencing (NGS) results obtained in step (a); and (c) analyzing the frequency of false positive mutations in each DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios, using the selected NN pair alleles as a reference.
[0014] A further object of the present invention is to provide an information-providing method for increasing the specificity of a next-generation sequencing (NGS) panel, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting NN pair alleles (Null-Null pair alleles) that do not contain mutations in the isozygous DNA and control genomic DNA from the NGS results obtained in step (a); (c) analyzing the frequency of false positive mutations in each DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios, using the selected NN pair alleles as a reference; and (d) analyzing the variant allelic fraction (VAF) that removes false positive mutations to derive a stringency cutoff value.
[0015] A further object of the present invention is to provide a method for evaluating errors in the raw data generation process or errors in the bioinformatics process in a company's NGS panel analysis process, which includes the steps of (a) collecting analysis results for false negative and false positive variants using bioinformatics analysis results held by the company from the company's raw data; and (b) comparing and analyzing the bioinformatics analysis results with the analysis results from step (a) using commercially available software from the company's raw data. [Means for solving the problem]
[0016] To achieve the above objectives, the present invention provides a validation composition for next-generation sequencing (NGS) panels comprising homozygote DNA and control genomic DNA.
[0017] In one embodiment of the present invention, the isozygote DNA may be DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
[0018] In one embodiment of the present invention, the control genomic DNA may be genomic DNA isolated from the blood of a normal human. Alternatively, the control genomic DNA may be genomic DNA isolated from the blood or cancerous tissue of a patient. Furthermore, the control genomic DNA may be circulating tumor DNA (ctDNA). Additionally, the control genomic DNA may be synthetic DNA or cloned DNA.
[0019] In one embodiment of the present invention, the dilution ratio of isozygote DNA and control genomic DNA may be selected from the group consisting of 1:9999 to 9999:1. Preferably, the dilution ratio of isozygote DNA and control genomic DNA may be selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5, and 9999:1, but is not limited thereto.
[0020] Furthermore, the present invention provides a validation kit for a next-generation sequencing (NGS) panel, which includes a validation composition for the next-generation sequencing (NGS) panel.
[0021] Furthermore, the present invention provides a method for validating a next-generation sequencing (NGS) panel through the analysis of false negative mutations, which includes the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the NGS results obtained in step (a); and (c) analyzing the frequency of false negative mutations in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios, using the selected Ho-N pair alleles or He-N pair alleles as a reference.
[0022] Also, the present invention provides a method for validating a next-generation sequencing (NGS) panel through limit of detection analysis, comprising the steps of: (a) performing next-generation sequencing (NGS) using a next-generation sequencing (NGS) panel on a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios from each other; (b) selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the results of the next-generation sequencing (NGS) obtained in step (a); (c) analyzing the diluted mutation detection rate, which is expressed as a percentage of the number of diluted mutations found in the Ho-N pair alleles or He-N pair alleles in each DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios, with the selected Ho-N pair alleles or He-N pair alleles as a reference; and (d) analyzing the limit of detection.
[0023] In one embodiment of the present invention, the method may be for validating a next-generation sequencing (NGS) panel capable of detecting a mutation selected from the group consisting of single nucleotide variation (SNV), insertion / deletion (Indel), and chromosomal amplification / deletion.
[0024] In one embodiment of the present invention, the limit of detection in step (d) may be expressed as a percentage of VAF (variant allelic fraction) corresponding to the maximum dilution ratio that exhibits a diluted mutation detection rate of 90% or 95% in a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed at different dilution ratios from each other.
[0025] The present invention also provides a method for verifying a next-generation sequencing (NGS) panel through false-positive mutation analysis, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample, in which homozygous DNA and control genomic DNA are mixed at different dilution ratios from each other, using the next-generation sequencing (NGS) panel; (b) selecting N-N pair alleles (Null-Null pair alleles) that do not contain mutations in both the homozygous DNA and the control genomic DNA from the next-generation sequencing (NGS) results obtained in step (a); and (c) analyzing the frequency of false-positive mutations in each DNA mixture sample, in which homozygous DNA and control genomic DNA are mixed at different dilution ratios from each other, with the selected N-N pair alleles as a reference.
[0026] In one embodiment of the present invention, the method may be used to verify a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNV), insertions / deletions (Indel), and chromosomal amplification / deletion.
[0027] The present invention also provides a method for providing information for increasing the specificity of a next-generation sequencing (NGS) panel, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample, in which homozygous DNA and control genomic DNA are mixed at different dilution ratios from each other, using the next-generation sequencing (NGS) panel; (b) selecting N-N pair alleles (Null-Null pair alleles) that do not contain mutations in both the homozygous DNA and the control genomic DNA from the next-generation sequencing (NGS) results obtained in step (a); (c) analyzing the frequency of false-positive mutations in each DNA mixture sample, in which homozygous DNA and control genomic DNA are mixed at different dilution ratios from each other, with the selected N-N pair alleles as a reference; and (d) analyzing VAF (variant allelic fraction) for removing false-positive mutations to derive a stringency cutoff value.
[0028] In one embodiment of the present invention, the stringency cutoff value may be a 95% stringency cutoff value that removes 95% of false positive mutations, or a 99% stringency cutoff value that removes 99% of false positive mutations.
[0029] In one embodiment of the present invention, the method may provide information on a strict cutoff value for removing false-positive mutations in order to increase the specificity of a next-generation sequencing (NGS) panel.
[0030] Furthermore, the present invention provides a method for evaluating errors in the raw data generation process or errors in the bioinformatics process in a company's NGS panel analysis process, which includes the steps of (a) collecting analysis results for false negative and false positive variations from the company's raw data using bioinformatics analysis results held by the company; and (b) comparing and analyzing the bioinformatics analysis results from the company's raw data using commercially available software with the analysis results from step (a). [Effects of the Invention]
[0031] The compositions according to the present invention can be usefully used for NGS panel validation. In particular, the validation method according to the present invention can (i) objectively analyze the frequency of false-negative mutations, the frequency of false-positive mutations, and the detection limit for an NGS panel; (ii) remove areas with a high frequency of false negatives during analysis by pre-determining the presence or absence of specific chromosomal regions or specific gene regions that tend to produce a particularly high number of false negatives during NGS panel validation, thereby reducing errors in the analysis of NGS results related to false negatives occurring in the NGS panel; and (iii) provide an objective stringency cutoff value by analyzing the distribution of the number of false-positive mutations and VAF (variant allelic fraction), thereby reducing errors in the analysis of NGS results related to false positives occurring in the NGS panel. [Brief explanation of the drawing]
[0032] [Figure 1] Figure 1 shows that the relative amount of DNA can be measured in a DNA mixture in which the homozygous variant DNA1 and the null DNA2 are mixed at various dilution ratios. [Figure 2] Figure 2A shows the results of analyzing the detection limits for AA Company's NGS panel using Ho-N pair alleles, Figures 2B and 2C show the results of analyzing the detection limits for BB Company's NGS panel, and Figures 2D and 2E show the results of analyzing the detection limits for CC Company's NGS panel (X-axis: Listed by allele chromosome number and position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); and FN: false negative allele; each color: false negative result at the corresponding dilution; X-axis at the bottom of each graph: Listed in order by chromosome position; and Y-axis at the bottom of each graph: chromosome number). The dilution ratios for control genomic DNA and H mole DNA are as follows: CH100, 0:100; CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; CH1, 99:1; and CH0, 100:0. [Figure 3]Figure 3A shows the results of analyzing the detection limit of AA Company's NGS panel using He-N pair alleles, Figure 3B shows the results of analyzing the detection limit of BB Company's NGS panel, and Figure 3C shows the results of analyzing the detection limit of CC Company's NGS panel (X-axis: alleles listed in order by chromosome number and position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false negative allele; each color: false negative result at the corresponding dilution ratio; X-axis at the bottom of each graph: listed in order by chromosome position; and Y-axis at the bottom of each graph: chromosome number). The dilution ratios for control genomic DNA and H mole DNA are as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; CH1, 99:1; and CH0, 100:0. [Figure 4] Figure 4A shows the results of analyzing indel detection using Ho-N pair alleles or He-N pair alleles with dilution ratios from AA Company's NGS panel, and Figure 4B shows the results of analyzing indel detection using dilution ratios from BB Company's NGS panel (X axis: alleles listed in order according to chromosomal position; Y axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false negative allele; and each color: false negative result at the corresponding dilution ratio). The dilution ratios for control genomic DNA and H mole DNA are as follows: CH100, 0:100; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; and CH0, 100:0. [Figure 5]Figure 5A shows the results of analyzing the detection limit of AA Company's NGS panel using Ho-He pair alleles, Figure 5B shows the results of analyzing the detection limit of BB Company's NGS panel, and Figure 5C shows the results of analyzing the detection limit of CC Company's NGS panel (X-axis: Listed by chromosomal position of alleles; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false negative allele; each color: false negative result at the corresponding dilution ratio; X-axis of the lower part of the 5C graph: Listed in order by chromosomal position; and Y-axis of the lower part of the 5C graph: chromosome number). The dilution ratios for control genomic DNA and H mole DNA are as follows: CH100, 0:100; CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH80, 20:80; CH50, 50:50; CH20, 80:20; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; CH1, 99:1; and CH0, 100:0. [Figure 6]Figure 6A shows the results of analyzing the detection limit of chromosome 6 mutations for Ho-N pair alleles or He-N pair alleles against CC Company's NGS panel; Figure 6B shows the results of analyzing the detection limit of chromosome 6 mutations for Ho-N pair alleles that are null in the control DNA; Figure 6C shows the results of analyzing the detection limit of chromosome 6 mutations for He-N pair alleles; and Figure 6D shows the results of analyzing the detection limit of chromosome 6 mutations for Ho-N pair alleles that are homozygote in the control DNA (X axis: Listed by chromosomal position of alleles; Y axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count); FN: false negative allele; and each color: false negative result at the corresponding dilution factor). The dilution ratios for control genomic DNA and H molecular DNA are as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH50, 50:50; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; and CH1, 99:1. [Figure 7]Figure 7A shows the results of analyzing the frequency of false positive mutations using AA Company's NGS panel with NN pair alleles where both H mole DNA and control genomic DNA are null; Figure 7B shows the results of analyzing the frequency of false positive mutations using BB Company's NGS panel; and Figure 7C shows the results of analyzing the frequency of false positive mutations using CC Company's NGS panel (X-axis: alleles listed in order by chromosomal position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count). The dilution ratios of control genomic DNA and H mole DNA are as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH50, 50:50; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; and CH1, 99:1. [Figure 8] Figure 8 shows the results obtained by removing false-positive mutations on chromosome 6 from CC Company's NGS panel results and deriving the stringency cutoff value from the remaining false-positive mutations (X-axis: alleles listed in order by chromosomal position; Y-axis: Log(VAF); VAF (variant allelic fraction) = (variant allele read count) / (total read count). The dilution ratios of control genomic DNA and H mole DNA are as follows: CH99, 1:99; CH97.5, 2.5:97.5; CH95, 5:95; CH90, 10:90; CH50, 50:50; CH10, 90:10; CH5, 95:5; CH2.5, 97.5:2.5; and CH1, 99:1. [Figure 9]Figure 9A shows the results of analyzing false-positive mutations derived from the identified mutations. The results of analyzing the presence or absence of false-positive mutations derived from mutations by each company (Inhouse) and the results of analyzing the presence or absence of false-positive mutations derived from mutations by Dragen software (under three conditions: default, solid, and liquid) are displayed graphically (y-axis: arithmetic mean of false-positive mutations found in a large number of dilution samples). Figure 9B shows the results of measuring sensitivity by analyzing false negatives from the mutation detection results of AA, BB, and CC NGS panels. The dilution mutation detection rate is displayed graphically (y-axis) from the results of analyzing the presence or absence of false negative mutations derived from mutations derived by each company using their respective bioinformatics methods (Inhouse) and the results of analyzing the presence or absence of false negative mutations derived from mutations derived by Dragen software using three conditions (default, solid, and liquid). (x-axis: % represents the percentage of samples containing mutations, and thus the percentage of mutations in the total DNA of each sample is half the percentage of mutated samples). For the purpose of determining false positive errors, if a variant or reference base that did not appear when analyzed in H mole DNA and control genomic DNA was found in all alleles, it was determined to be a false positive error. [Modes for carrying out the invention]
[0033] The present invention will be described in detail below.
[0034] The terminology used in this invention has been selected to the greatest extent possible from currently widely used general terms, taking into account the function of the invention, but this may change depending on the intent of the technicians engaged in the relevant art or the emergence of new technologies. In certain cases, some terms have been arbitrarily selected, in which case their meaning will be described in detail in the description of the embodiment. Therefore, the terminology used in this invention must be defined not merely as a name of a term, but based on the meaning of the term and the overall content of the invention.
[0035] In the present invention, when we say that any component or any step "includes," this means that, unless otherwise stated, it does not exclude other components or steps, but rather that other components or steps may be further included.
[0036] The present invention provides a composition for verifying next-generation sequencing (NGS) panels, which include homozygote DNA and control genomic DNA.
[0037] In this invention, the term "next-generation sequencing" refers to a technique capable of analyzing millions of base sequences at once, and is also known as massively parallel sequencing or high-throughput sequencing. In this invention, the terms next-generation sequencing and NGS are used interchangeably.
[0038] In this invention, the term "allele" refers to a pair of homologous chromosomes, each containing a different trait, and signifies the DNA sequence of the allele. Homologous chromosomes are classified into isozygotes and heterozygotes. An "isozygote" is formed when alleles with the same properties are joined, while a "heterozygote" is formed when alleles with different properties are joined. In this invention, the terms "allele" and "allele" are used interchangeably.
[0039] In the present invention, the composition contains two different genomic DNAs for verifying an NGS panel, one of which is an isozygote DNA.
[0040] In the present invention, the isozygote DNA may be, but is not limited to, DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
[0041] In this invention, the term "control genomic DNA" means one of the two distinct genomic DNAs contained in the composition, other than the isozygous DNA, and may be isozygous or heterozygous genomic DNA.
[0042] In the present invention, the control genomic DNA may be genomic DNA isolated from the blood of a normal human. Alternatively, the control genomic DNA may be genomic DNA isolated from the blood or cancerous tissue of a patient. Furthermore, the control genomic DNA may be circulating tumor DNA (ctDNA). Additionally, the control genomic DNA may be, but is not limited to, synthetic DNA or cloned DNA.
[0043] In the present invention, the composition may be selected from the group consisting of dilution ratios of isozygote DNA and control genomic DNA from 1:9999 to 9999:1, and preferably the dilution ratios of isozygote DNA and control genomic DNA may be selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5, and 9999:1, but is not limited thereto.
[0044] Furthermore, the present invention provides a verification kit for a next-generation sequencing (NGS) panel, which includes a verification composition for the next-generation sequencing panel.
[0045] Furthermore, the present invention provides a method for validating a next-generation sequencing (NGS) panel through the analysis of false negative mutations, which includes the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the NGS results obtained in step (a); and (c) analyzing the frequency of false negative mutations in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios, using the selected Ho-N pair alleles or He-N pair alleles as a reference.
[0046] In this invention, the term "false negative" refers to a case where a test result that should be positive is mistakenly shown as negative.
[0047] Furthermore, the present invention provides a method for validating a next-generation sequencing (NGS) panel, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the NGS results obtained in step (a); (c) analyzing the dilution mutation detection rate, which is expressed as a percentage, of the number of diluted mutations found in the Ho-N pair alleles or He-N pair alleles in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios, based on the selected Ho-N pair alleles or He-N pair alleles; and (d) analyzing the limit of detection.
[0048] In this invention, the term "Ho-N pair allele (Homozygote-Null pair allele)" refers to an allele in which one allele region is mixed with a homozygous variant and a null allele that does not undergo mutation.
[0049] In this invention, the term "He-N pair allele (Heterozygote-Null pair allele)" refers to an allele in which one allele region is a mixture of a heterozygous variant and a null allele without mutation.
[0050] In this invention, the term "diluted mutation detection rate" refers to the number of diluted mutations found in Ho-N pair alleles or He-N pair alleles in a DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios, expressed as a percentage.
[0051] In the present invention, the method may involve verifying a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNVs), insertions / deletions (Indels), and chromosomal amplification / deletions.
[0052] In this invention, the term "variant" refers to a modification of a chromosome, gene, or base sequence that is genetically distinct from the wild type. In this invention, the terms "variant" and "mutation" are used interchangeably.
[0053] In the present invention, the detection limit in step (d) may be expressed as a percentage of the VAF (variant allelic fraction) corresponding to the maximum dilution ratio that shows a 90% or 95% detection rate of dilution mutations in a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios.
[0054] In this invention, for the validation of the NGS panel, an NGS panel validation composition was used that included a DNA mixture sample containing DNA isolated from a hydatidiform mole cell line, which is isozygote DNA (hereinafter referred to as "H mole DNA"), and control genomic DNA at various different dilution ratios.
[0055] For example, if DNA1 contains a homozygous variant in one allele site and DNA2 contains a null allele, mixing DNA1 and DNA2 and calculating the variant allelic fraction (VAF) allows us to determine the dilution ratio of DNA1 and DNA2. By expressing the VAF corresponding to the maximum dilution ratio that shows a 90% or 95% dilution variant detection rate for each allele as a percentage, we can derive the limit of detection.
[0056] More specifically, Figure 1 shows that the relative amounts of mixed DNA can be measured in each DNA mixture in which DNA1, a homozygous variant, and DNA2, a null, are mixed at various different dilution ratios. In this case, an allele in which one allele site is mixed with a homozygous variant and a null is defined as an "Ho-N pair allele (Homozygote-Null pair allele)," and the probability of an Ho-N pair allele appearing among alleles with different VAFs (variant allelic fractions) is shown in Table 1 below.
[0057] [Table 1] *p=VAF(variant allelic fraction);q=1-p;P1: The probability of an Ho-N pair allele occurring when both DNAs are general genomic DNA (P1=p 2 q2 +q 2 p 2 =2p 2 q 2 ); and P2: The probability of getting an Ho-N pair allele when one of the two DNAs is an H mole DNA (P2=p 2 q+q 2 p = pq (p + q) = pq).
[0058] As shown in Table 1 above, when comparing the number of Ho-N pair alleles from alleles with nine different VAFs (variant allelic fraction, p) in two cases: when one of the two DNA mixtures is isozygous H mole DNA, and when both DNAs are general genomic DNA, the sum of the probabilities of Ho-N pair alleles appearing from alleles with nine different VAFs is calculated to be 0.67 when both DNAs are general genomic DNA, whereas it is calculated to be 1.65 when one DNA is isozygous H mole DNA. In other words, when one DNA is isozygous H mole DNA, the number of Ho-N pair alleles can be obtained approximately 2.45 times higher (= P2 / P1 ratio) compared to when both DNAs are general genomic DNA.
[0059] In particular, the P2 / P1 ratio remained unchanged even when calculating the probability separately for each allele, depending on whether the mutation was analyzed as an actual mutation or as a reference sequence. Furthermore, the P2 / P1 ratio remained the same even when the frequency of the allele determined to be the reference sequence allele was lower than the frequency of the variant allele. When calculating the p-value considering the frequency of mutations appearing in the distribution of each allele and integrating it over a range of 0 to 1, the P2 / P1 ratio remained almost unchanged at 2.50 (for example, considering that the frequency of allele 1 is 0.1 and the frequency of allele 9 is 0.9).
[0060] From this, we confirmed that in this invention, when one of the two DNA mixtures is an isozygous H mole DNA, it is possible to determine approximately 2.5 times more Ho-N pair alleles compared to the case where one is not an isozygous H mole DNA, and that analyzing the detection limit of an NGS panel using isozygous H mole DNA may be even more useful.
[0061] Furthermore, the present invention provides a method for verifying a next-generation sequencing (NGS) panel through false-positive mutation analysis, which includes the steps of: (a) performing next-generation sequencing (NGS) analysis on a DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) selecting NN pair alleles (Null-Null pair alleles) that do not contain mutations in the isozygous DNA and control genomic DNA from the NGS results obtained in step (a); and (c) analyzing the frequency of false-positive mutations in each DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios, using the selected NN pair alleles as a reference.
[0062] In this invention, the term "NN pair allele (Null-Null pair allele)" refers to an allele in which all allele regions are null, meaning that no mutations are present.
[0063] In this invention, the term "false positive" refers to a case where a test result that should be negative is mistakenly shown as positive.
[0064] In this invention, a false positive mutation was determined when a mutation was detected in each DNA mixture sample in which isozygous DNA and control genomic DNA were mixed at different dilution ratios, using an NN pair allele (where both isozygous DNA and control genomic DNA are null) as the standard.
[0065] In the present invention, the method may involve verifying a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNVs), insertions / deletions (Indels), and chromosomal amplification / deletions.
[0066] Furthermore, the present invention provides an information-providing method for increasing the specificity of a next-generation sequencing (NGS) panel, comprising the steps of: (a) performing next-generation sequencing (NGS) on a DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios; (b) selecting NN pair alleles (Null-Null pair alleles) that do not contain mutations in the isozygous DNA and control genomic DNA from the NGS results obtained in step (a); (c) analyzing the frequency of false positive mutations in each DNA mixture sample in which isozygous DNA and control genomic DNA are mixed at different dilution ratios, using the selected NN pair alleles as a reference; and (d) analyzing the variant allelic fraction (VAF) that removes false positive mutations to derive a stringency cutoff value.
[0067] In the present invention, the stringency cutoff value may be a 95% stringency cutoff value that removes 95% of false positive mutations, or a 99% stringency cutoff value that removes 99% of false positive mutations.
[0068] In the present invention, the method may also provide information on a strict cutoff value for removing false-positive mutations in order to increase the specificity of a next-generation sequencing (NGS) panel.
[0069] The present invention will be described in more detail below based on examples. These examples are provided to illustrate the present invention more concretely, and the scope of the present invention is not limited to these examples.
[0070] Example 1. DNA mixture preparation and NGS analysis In this invention, for NGS panel validation, the sensitivity and specificity of H mole DNA, which is isozygote DNA, against NGS panels from three different companies were compared and analyzed. The H mole DNA derived from human hydatidiform mole used in this invention was purchased from Coriell (NA07489, Camden, NJ), and the control genomic DNA was extracted from the blood of a collaborator with the approval of the Institutional Review Board of the National Cancer Center.
[0071] The prepared H mole DNA, control genomic DNA, and DNA mixtures containing both H mole DNA and control genomic DNA were sent to companies AA, BB, and CC, respectively. Experiments were then conducted using Illumina's Novaseq 6000 NGS equipment and each company's NGS panel. Subsequently, files were received from the FASTQ files obtained from the experiments, analyzed for mutations using methods established by each company, and the NGS panels were validated. However, the number of mutations obtained from each company's NS panel differed not only in the number of gene mutations but also in the analytical methods used by each company. All companies commonly used hg19 from the NCBI human genome assembly as the reference sequence, and company CC further analyzed more mutations by using LoFreq in addition to MuTect as a variant caller.
[0072] Example 2. NGS Verification Method 2.1 Selection of Ho-N pair alleles and He-N pair alleles The results for the mutations analyzed by each company were sorted by chromosome number and position along with the sample information, including reference base, variant allele information, VAF, read count, and mutation information. In the analysis results from companies AA and BB, read counts less than 100 were excluded from the analysis, and in the analysis results from company CC, read counts less than 300 were excluded. Furthermore, in the cases of companies AA and BB, all intron region SNVs and indel alleles were included in the analysis in addition to exon region SNVs, while in the case of company CC, exon region SNVs and splice region SNVs were included in the analysis, but intron region SNVs and indel mutations were excluded.
[0073] Among the selected alleles, pairs in which the H mole DNA and the control genomic DNA alleles consist of a homozygous variant and a null were defined as "Ho-N pair allele (Homozygote-Null pair allele)," and pairs in which the H mole DNA is null and the control genomic DNA consists of a heterozygous variant were defined as "He-N pair allele (Heterozygote-Null pair allele)."
[0074] 2.2. Determination of false-negative mutation detection rates, dilution mutation detection rates, and detection limits using Ho-N pair alleles and He-N pair alleles. Alleles in which mutations are present in the isozygous DNA (H mole DNA) or control genomic DNA, but which were not detected by NGS panel analysis of a DNA mixture sample of H mole DNA and control genomic DNA, were defined as false negative alleles. The frequency of false negative alleles was analyzed by separating the DNA mixture sample of H mole DNA and control genomic DNA into Ho-N pair alleles and He-N pair alleles.
[0075] The dilution mutation detection rate was defined as the percentage of alleles in which mutations were found in Ho-N pair alleles or He-N pair alleles in a DNA mixture in which isozygous DNA (H mole DNA) and control genomic DNA were mixed at different dilution ratios.
[0076] The Limit of Detection (LoD) was defined as the variant allelic fraction (VAF), expressed as a percentage, of the maximum dilution ratio at which a 90% or 95% variant detection rate is achieved in a DNA mixture in which isozygous H mole DNA and control genomic DNA are mixed at different dilution ratios. Here, VAF is the value obtained by dividing the read count of sequences detected as variants at a specific allele site by the total read count, which includes sequences detected as both variants and reference sequence bases at that specific allele site.
[0077] 2.3. Determination of false-positive mutations and stringency cutoff values Among the null alleles in which no mutations were found in either the isozygous H mole DNA or the control genomic DNA, a false positive mutation was defined as a mutation found in NGS analysis of a DNA mixture sample in which the H mole DNA and the control genomic DNA were mixed at different dilution ratios. The VAF (variant allelic fraction) value that can remove 95% or 99% of such false positive mutations was defined as the 95% stringency cutoff value or the 99% stringency cutoff value, respectively.
[0078] Example 3. Sensitivity analysis results for an NGS panel using Ho-N pair alleles. From the results obtained from companies AA, BB, and CC, the median read counts of the mutant alleles among all the mutant sites were confirmed to be 197 (Q1, Q3; 51, 464), 813 (433, 1189), and 679 (577, 717), respectively. On the other hand, the median read counts in the samples for the Ho-N pair allele and He-N pair allele of the present invention were confirmed to be 380 (Q1, Q2; 221, 633) from company AA, 901 (Q1, Q; 503, 1234) from company BB, and 703 (Q1, Q2; 679, 727) from company CC, respectively.
[0079] First, we analyzed Ho-N pair alleles to evaluate the sensitivity to NGS panels from three companies. Since the mutation analysis results from company AA showed very few alleles in the exon region, we included intron region mutations in the sensitivity and specificity analysis. After selecting Ho-N pair alleles from the analysis results from company AA, 15 (10.2%, 15 / 147) were confirmed to be exon region mutations, and the remainder were confirmed to be intron region mutations or UTR region mutations.
[0080] Furthermore, while BB Company's mutation analysis results included the submitted SNVs and indels, sensitivity and specificity were analyzed including all exon and intron mutations, and after selecting Ho-N pair alleles, 134 (89.9%, 134 / 149) were confirmed to be exon or splice site mutations. CC Company's mutation analysis results had a relatively large number of alleles analyzed, so only exon and splice site mutations were selected, indels were removed, and only SNVs were analyzed, but a total of 306 Ho-N pair alleles were confirmed.
[0081] Analysis of the sensitivity of each company's NGS panel using Ho-N pair alleles revealed that, firstly, AA Company's NGS panel detected all SNVs in a DNA mixture (CH50) containing 50% H mole DNA and 50% control genomic DNA, and detected 134 SNVs out of 147 Ho-N pair alleles in a DNA mixture (CH80) containing 80% H mole DNA and 20% control genomic DNA, confirming a dilution mutation detection rate of 91.2% (Figure 2A). Furthermore, in a DNA mixture (CH90) containing 90% H mole DNA and 10% control genomic DNA, only 16 SNVs were detected, confirming a dilution mutation detection rate of 10.9% (Figure 2A). From this, it was confirmed that the maximum dilution factor showing a 90% dilution mutation detection rate for AA Company's NGS panel is approximately 20%, which means that the detection limit showing a 90% dilution mutation detection rate is approximately 20%.
[0082] BB Company's NGS panel detected all 149 Ho-N pair alleles in 10% DNA mixture samples (i.e., DNA mixture containing 90% H mole DNA and 10% control genomic DNA (CH90) and DNA mixture containing 10% H mole DNA and 90% control genomic DNA (CH10)), and detected 137 out of 149 Ho-N pair alleles in 5% DNA mixture samples (i.e., DNA mixture containing 95% H mole DNA and 5% control genomic DNA (CH95) and DNA mixture containing 5% H mole DNA and 95% control genomic DNA (CH5)), confirming a dilution mutation detection rate of approximately 88.6% (Figures 2B and 2C). From this, it was confirmed that the maximum dilution ratio showing a 90% dilution mutation detection rate for BB Company's NGS panel is between 5 and 10%, and moreover, around 5%, which means that the detection limit showing a 90% dilution mutation detection rate is around 5%.
[0083] CC Company's results showed a high number of false negatives in some genes on chromosome 6, and these were excluded from the limit-of-detection analysis. Chromosome 6 contains 91 Ho-N pair alleles, and the remaining 215 Ho-N pair alleles were analyzed. CC's NGS panel detected all 215 Ho-N pair alleles in 2.5% DNA mixture samples (i.e., a DNA mixture containing 97.5% H mole DNA and 2.5% control genomic DNA (CH97.5) and a DNA mixture containing 2.5% H mole DNA and 97.5% control genomic DNA (CH2.5)), and detected 197 out of 215 Ho-N pair alleles in 1% DNA mixture samples (i.e., a DNA mixture containing 99% H mole DNA and 1% control genomic DNA (CH99) and a DNA mixture containing 1% H mole DNA and 99% control genomic DNA (CH1)), confirming a dilution mutation detection rate of approximately 91.6% (Figures 2D and 2E). From this, it was confirmed that the maximum dilution factor showing a 90% dilution mutation detection rate for CC's NGS panel is 1%, which means that the detection limit showing a 90% dilution mutation detection rate is approximately 1%.
[0084] From the above results, it was confirmed that Ho-N pair alleles can be selected using a DNA mixture containing the H mole DNA of the present invention, and the detection limit, expressed as a percentage of the VAF (variant allelic fraction) corresponding to the maximum dilution ratio showing a 90% dilution mutation detection rate, can be analyzed. This allows for the verification of each company's NGS panel through the analysis of false-negative mutation frequency and detection limit.
[0085] Example 4. Sensitivity analysis results for an NGS panel using He-N pair alleles. We investigated whether sensitivity could be analyzed using He-N pair alleles consisting of null and heterozygous variants from two mixed DNAs. First, AA's analysis yielded 269 He-N pair alleles (53 exonic variants, 19.7% of the total). AA's NGS panel detected 266 out of 269 He-N pair alleles in a 50% DNA mixture sample (i.e., a DNA mixture containing 50% H mole DNA and 50% control genomic DNA (CH50)), confirming a dilution variant detection rate of 98.9% (Figure 3A). This means that 98.9% of variant alleles with a VAF (variant allelic fraction) of 0.25 were detected, and the detection limit was less than 25%. This result was confirmed to be similar to the analysis results using the aforementioned He-N pair alleles.
[0086] BB Company's analysis yielded 172 He-N pair alleles (147 exonic or splicing variants, representing 85.5% of the total). BB Company's NGS panel detected 166 out of 172 He-N pair alleles in a 10% DNA mixture sample (i.e., a DNA mixture containing 90% H mole DNA and 10% control genomic DNA (CH90) and a DNA mixture containing 10% H mole DNA and 90% control genomic DNA (CH10)), confirming a dilution variant detection rate of 96.5% (Figure 3B). This means that 96.5% of variant alleles with a VAF (variant allelic fraction) of 0.05 were detected, with a detection limit of 5% or less. This result was confirmed to be similar to the analysis results using the aforementioned He-N pair alleles.
[0087] CC's analysis yielded 276 He-N pair alleles (all exon-site SNVs), excluding 54 He-N pair alleles located on chromosome 6. CC's NGS panel detected 246 out of 276 He-N pair alleles in a 2.5% DNA mixture sample (i.e., a DNA mixture containing 97.5% H mole DNA and 2.5% control genomic DNA (CH97.5) and a DNA mixture containing 2.5% H mole DNA and 97.5% control genomic DNA (CH2.5)), confirming a dilution variant detection rate of 89.1% (Figure 3C). This means that 89.1% of the variant alleles with a VAF (variant allelic fraction) of 0.0125 were detected, with a detection limit of 1.25%. This result was confirmed to be similar to the analysis results using the aforementioned He-N pair alleles.
[0088] For reference, the Ho-N pair alleles and He-N pair alleles of company AA and company BB contained 38 and 8 indels respectively, in addition to SNVs, and these had little effect on the calculation of the detection limit (Figure 4).
[0089] From the above results, it was confirmed that in this invention, by analyzing Ho-N pair alleles and He-N pair alleles and confirming the maximum dilution ratio that shows a 90% dilution mutation detection rate, the detection limit for each NGS panel can be analyzed as an objective result, and that it can be usefully used for validating NGS panels. Furthermore, it was confirmed that using all the results from both Ho-N pair alleles and He-N pair alleles has the advantage of providing sensitivity information for more alleles compared to using only Ho-N pair alleles.
[0090] Example 5. Relationship between detection limit and read count or VAF (variant allelic fraction) Previous literature has reported that when the read count (or local sequence coverage) is 350, alleles with a VAF (variant allelic fraction) of less than 5% and alleles with a VAF of 5-10% can be detected with 89.4% and 99.2% accuracy, respectively, and when the read count is 738, alleles with a VAF (variant allelic fraction) of less than 5% and alleles with a VAF (variant allelic fraction) of 5-10% can be detected with 98.5% and 99.7% accuracy, respectively, indicating that the read count and VAF (variant allelic fraction) are related to the sensitivity or detection limit of mutations (Frampton et al., Nature Biotechnology 2013).
[0091] Incidentally, the actual analysis results showed that while Company AA had a median read count of approximately 380, it detected only 10.9% of mutations with a VAF (variant allelic fraction) of 10%, and 91.2% of mutations with a VAF of 20%, indicating that Company AA's sensitivity to the NGS panel was remarkably low. Furthermore, Company BB had a read count of approximately 900, and detected 88.6% of mutations with a VAF (variant allelic fraction) of 5%, showing that the detection limit for Company BB's NGS panel was similar to the results in the aforementioned prior literature.
[0092] However, CC Company's read count was approximately 700, and it detected 100% of mutations with a VAF (variant allelic fraction) of 2.5% and 91.6% of mutations with a VAF of 1%, indicating that the detection limit for CC Company's NGS panel was superior to the results of the aforementioned prior literature. Furthermore, these results suggest that there may be various variables other than the relationship between read count and VAF that determine the sensitivity of an NGS panel. Therefore, it is clear that the sensitivity or detection limit of an NGS panel cannot be evaluated solely by read count and VAF, and that researchers need to verify the sensitivity of a particular NGS panel using standard mixtures in order to evaluate whether that panel has the sensitivity to achieve the research objectives.
[0093] Example 6. Sensitivity analysis results for an NGS panel using Ho-He pair alleles. Next, we defined a pair consisting of a homozygous variant in H mole DNA and a heterozygous variant in control genomic DNA as a "Ho-He pair allele (Homozygote and Heterozygous pair allele)." We then calculated the RAF (reference allelic fraction) value by determining the fraction occupied by the diluted reference allele compared to the variant allele, and confirmed whether the detection limit could be analyzed based on this value.
[0094] In analysis using RAF, AA's NGS panel detected 94.9% (131 / 138) Ho-He pair alleles with a 2.5% RAF (reference allelic fraction) in a 5% DNA mixture sample (i.e., a DNA mixture containing 95% H mole DNA and 5% control genomic DNA (CH95)) (Figure 5A). Similarly, BB's NGS panel detected 96.7% (116 / 120) Ho-He pair alleles with a 2.5% RAF (reference allelic fraction) in a 5% DNA mixture sample (i.e., a DNA mixture containing 95% H mole DNA and 5% control genomic DNA (CH95) and a DNA mixture containing 5% H mole DNA and 95% control genomic DNA (CH5)) (Figure 5B). Furthermore, CC's NGS panel detected 97.8% (221 / 226) of Ho-He pair alleles with 0.5% RAF in 1% DNA mixture samples (i.e., a DNA mixture containing 99% H mole DNA and 1% control genomic DNA (CH99) and a DNA mixture containing 1% H mole DNA and 99% control genomic DNA (CH1)) (Figure 5C).
[0095] As described above, analysis of the detection limit using Ho-He pair alleles confirmed that the detection rate of the reference sequence bases diluted at the maximum dilution ratio requested from each company was 90% or higher. Furthermore, no false negative alleles found on chromosome 6 in the CC company's NGS panel were clearly detected.
[0096] As described above, the analysis results using Ho-He pair alleles were found to be completely different from the results of dilution mutation detection analysis using Ho-N pair alleles or He-N pair alleles. Therefore, it was confirmed that the method of analyzing the detection limit of reference sequence bases using the RAF (reference allelic fraction) value of Ho-He pair alleles is not suitable for evaluating NGS panels.
[0097] Example 7. Additional analysis results for the detection limit of chromosome 6 CC Company's NGS panel was found to have a serious problem with a high number of false negatives on chromosome 6, and analyses using He-N pair alleles were relatively less frequent compared to analyses using Ho-N pair alleles. Therefore, after matching the chromosome 6 region with the chromosomal location of Ho-N pair alleles or He-N pair alleles, we checked for differences in the analysis results between Ho-N pair alleles and He-N pair alleles.
[0098] As a result, we found that there are specific gene regions within chromosome 6 that produce more false negatives, and that the He-N pair allele contains fewer of these gene regions, resulting in fewer false negatives compared to the Ho-N pair allele (Figure 6). The existence of specific chromosomal regions associated with false negatives, as seen in CC's NGS panel, means that even mutations with large VAFs (variant allelic fractions) may result in false negatives, and this must be taken into consideration.
[0099] From the results described above, it was confirmed that the method of the present invention has the advantage of being able to determine in advance the presence or absence of specific chromosomal regions or specific gene regions that tend to produce a particularly high number of false negatives during NGS panel validation, thereby allowing the final result to be judged while considering the possibility of false negatives in those regions during analysis. Furthermore, it was confirmed that by removing such regions from the design of a new NGS panel, errors in NGS panel analysis related to false negatives occurring in the NGS panel can be reduced.
[0100] Example 8. Analysis results of false-positive mutations and stringency cutoff values for the NGS panel. As mentioned above, in null alleles where no mutations were found in either the H mole DNA or the control genomic DNA (defined as "NN pair allele"), if mutations were detected in DNA mixture samples where the H mole DNA and the control genomic DNA were mixed at different dilution ratios, these were defined as false positive mutations, and the detection frequency of false positive mutations was analyzed using each company's NGS panel.
[0101] As a result, the NGS panels of companies AA and BB detected relatively fewer false-positive mutations compared to those of company CC (Figure 7). This is thought to be due to the fact that the VAF values, which indicate the detection limit of the NGS panels of companies AA and BB, were relatively higher than the VAF value of company CC, in other words, the sensitivity for detecting mutations was reduced. However, it was confirmed that company CC's NGS panel had relatively fewer false-positive mutations compared to the NGS panels of companies AA and BB when the VAF (variant allelic fraction) was 0.1 or higher, except for chromosome 6 (Figure 7).
[0102] For reference, two problems were discovered in CC Company's NGS panel analysis. First, a large number of false positive mutations were found in the 10% DNA mixture sample (i.e., a DNA mixture containing 10% H mole DNA and 90% control genomic DNA (CH10)). However, if the experiment was conducted under similar conditions to other samples, a similar number of false positive mutations should have been found. Since approximately 10 times more false positive mutations were found in CH10 alone, this was judged to be an error in the experimental process using CC Company's NGS instrument, and in Figure 7C, the false positive mutations that appeared only in CH10 were excluded. Second, the VAF of false positive mutations in the chromosome 6 allele of CC Company's NGS panel was significantly lower or significantly higher than the VAF of false positive mutations in the remaining chromosomal regions, as can be seen in Figure 7C. Therefore, the chromosome 6 allele was excluded from the false positive mutation detection analysis.
[0103] Subsequently, as mentioned above, a 95% stringency cutoff value or a 99% stringency cutoff value was derived, defined as a VAF value that can remove 95% or 99% of false-positive mutations. To date, most researchers, after NGS panel analysis, have determined only those above a certain VAF value to be "true mutations" in order to remove false-positive mutations, and have referred to this VAF (variant allelic fraction) value as the "stringency cutoff value," but no clear scientific criteria for determining this value have been presented.
[0104] The present invention aims to clearly present a method for deriving a VAF value that can remove 95% or 95% of false-positive mutations as a "stringency cutoff value." To confirm this, the distribution of VAF values of false-positive mutations was analyzed using CC Company's NN pair allele. In this case, the alleles corresponding to the two issues that arose in the aforementioned CC Company NGS panel (i.e., 1. a large number of false-positive mutations appearing in CH10 samples; and 2. severe VAF changes in false-positive mutations on chromosome 6) were excluded from the analysis before deriving the stringency cutoff value.
[0105] As a result, the number of false positive mutations obtained from CC Company's NGS panel was ultimately 1,977. It was confirmed that the VAF (variant allelic fraction) value that can remove 95% of false positive mutations is 0.067, and the VAF (variant allelic fraction) value that can remove 99% of false positive mutations is 0.0875 (Figure 8). In other words, when attempting to analyze CC Company's NGS panel to remove 95% of false positive mutations, a stringency cutoff value of 0.067 was derived, and when attempting to analyze CC Company's NGS panel to remove 99% of false positive mutations, a stringency cutoff value of 0.087 was derived.
[0106] From the above results, it was confirmed that the present invention can validate an NGS panel by analyzing the frequency of false positive mutations, provide information on an objective strict cutoff value for removing false positive mutations to increase the specificity of the NGS panel, and reduce errors in the analysis of NGS results associated with false positives occurring in the NGS panel.
[0107] Example 9. Evaluation of errors and bioinformatics analysis errors included in FASTQ files from each company's NGS panel analysis results. The errors found in the analysis of the NGS panels from three companies can be divided into (1) errors inherent in the raw data (FASTQ file) created using the NGS panel, and (2) errors in the bioinformatics analysis process used to find mutations in the raw data from each of the three companies. To determine the contribution of these two errors, we used commercially available software to find mutations in the raw data and used this to determine the contribution of the two errors. For the bioinformatics analysis, we used Illumina's Dragen software. This process requires bed files from each company, but companies AA and CC did not provide bed files, so we only analyzed mutations in the exon portion of the genes provided by these two companies. In the case of company BB, we used the bed file provided by the company to analyze all mutations in the entire captured region.
[0108] The determination of false negative errors in the mutated region was performed as described above. However, for the purpose of determining false positive errors, if a variant or reference base that did not appear when analyzed in H mole DNA and control genomic DNA was found in all alleles, including not only NN pair alleles but also Ho-N pair, He-N pair, Ho-He pair, etc., it was all determined to be a false positive error, and the detection frequency of false positive mutations was then analyzed using each company's NGS panel.
[0109] Figure 9 shows the results of analyzing false negatives and false positives from mutation results provided by three companies using bioinformatics methods, and the results of analyzing false negatives and false positives from mutations in each company's raw data (FASTQ files) using Dragen software. Mutation analysis was performed using Dragen software, with three conditions: default, solid, and liquid. Default is the commonly used condition, while solid and liquid are conditions with increased sensitivity compared to the default condition.
[0110] Figure 9A shows the results of the analysis of false positive mutations derived from the mutations. The graph displays the results of the analysis of the presence or absence of false positives of mutations derived from each company, as well as the results of the analysis of the presence or absence of false positives of mutations derived from Dragen software (under three conditions: default, solid, and liquid). The y-axis in Figure 9A shows the arithmetic mean of false positive mutations found in a large number of diluted samples. It can be seen that the results of analysis under the default condition of Dragen software had fewer false positives compared to the results of analysis under the solid or liquid conditions (the number of false positives for companies BB and CC was 30 or less, shown by the blue line). Considering that the default condition of Dragen software has lower sensitivity than the other two conditions, these results are consistent with the general phenomenon that specificity increases when sensitivity is reduced.
[0111] Incidentally, analysis of mutations using AA Company's in-house method yielded a low number of false positives (average 0.5), while analysis using the default conditions of the Dragen software resulted in a high number of false positives (squares in Figure 9A, average 818 false positives). For this reason, as will be explained later, it is suspected that a large number of FP errors occurred in the NGS panel result analysis at AA Company. This suggests that the conditions were changed to increase specificity at the expense of sensitivity. Furthermore, it can be confirmed that CC Company's in-house method also generates a high number of false positives (arrows in Figure 9A, average 1680 false positives), which, as will be explained later, is considered to be related to CC Company intentionally increasing the sensitivity when analyzing mutations.
[0112] Figure 9B shows the results of measuring sensitivity by analyzing false negatives from mutation detection results of NGS panels from companies AA, BB, and CC. The dilution mutation detection rate is shown on the (y-axis) graph from the results of analyzing the presence or absence of false negatives of mutations derived by each company's bioinformatics method (Inhouse) and the results of analyzing the presence or absence of false negatives of mutations derived by the bioinformatics method of Dragen software under three conditions (default, solid, and liquid). The % shown on the x-axis of Figure 9B represents the percentage of samples containing heterozygote mutations, and therefore the percentage of mutations in the total DNA of each sample is half of the mutation sample percentage shown on the x-axis (for example, for a sample shown as 5% on the x-axis in Figure 9B, the actual percentage of mutations is 2.5%). In Figure 9B, only the results using the default condition are shown for the results using the bioinformatics method with Dragen software. However, in the case of companies AA and CC, when analyzing the mutation detection rate using the Dragen software in Figure 9B, only mutations present in the exons of the target genes provided by each company were analyzed. In the case of company BB, mutations present in the regions captured by company BB's NGS panel using a bed file were analyzed. Furthermore, all sensitivity analysis results in Figure 9B are based on results using He-N pair alleles.
[0113] As seen in the sensitivity analysis results in Figure 9B, the AA Company's analysis using these bioinformatics methods showed almost no detection of mutations in samples containing 20% heterozygote variant DNA. This means that the mutation detection limit of the AA Company's NGS panel is greater than 10%. Referring to the results of the analysis of Ho-N pair alleles, the AA Company's raw data and the final results analyzed using these bioinformatics methods indicate that the mutation detection limit is only about 20%. On the other hand, analyzing the mutations obtained under the default conditions of the Dragen software shows that it detects almost no mutations present in samples containing 10% heterozygote variant DNA, indicating that the sensitivity mutation detection limit is about 5%. Therefore, it can be considered that the bioinformatics method used by the AA Company selected conditions with low sensitivity. This adjustment to lower sensitivity appears to be related to the fact that a very large number of false-positive mutations (818) were found when analyzing the mutations obtained under the default conditions of the Dragen software in Figure 9A. In short, it can be understood that AA changed the analysis conditions (changing the conditions so that the sensitivity was around 20%) in order to increase specificity at the expense of sensitivity.
[0114] Figure 9B shows the sensitivity analysis results. In the case of CC Company, analysis under those bioinformatics conditions revealed that almost all diluted mutations in the 2.5% heterozygote variant DNA sample were detected, indicating a sensitivity of approximately 1.25%, capable of detecting almost all mutations. Analysis of mutations obtained under the default conditions of Dragen software revealed that only about 50% of the diluted mutations in the 2.5% heterozygote variant DNA sample were detected, while almost all diluted mutations in the 5% heterozygote variant DNA sample were detected, indicating a mutation detection limit of approximately 2.5%. Therefore, CC Company's bioinformatics method appeared to have higher sensitivity than the analysis using Dragen software. However, Figure 9A shows the specificity analysis results. CC Company's in-house bioinformatics analysis conditions detected 1680 false-positive mutations, indicating that these bioinformatics analysis conditions intentionally use conditions that increase sensitivity, resulting in a significant decrease in specificity.
[0115] In summary, while Company AA can be acknowledged for having many errors in the process of generating raw data, it is recognized that the company adjusted its analytical conditions at the expense of sensitivity in order to solve the problem of many false-positive mutations occurring during the bioinformatics analysis process. Company BB experienced more false-negative dilution mutations during the bioinformatics analysis process compared to mutation analysis using Dragen software, and consequently, the sensitivity of the NGS panel analyzed by the company is recognized as being slightly lower than that of analysis using Dragen software. Company CC adjusted its analytical conditions to increase sensitivity during the bioinformatics analysis process, and it is recognized that this resulted in many false positives.
[0116] As described above, analyzing the results of the present invention with commercially available software has the advantage of allowing for the identification of errors in (1) the raw data generation process and (2) the bioinformatics analysis process during the NGS panel analysis process of a specific company.
Claims
1. A validation composition for a next-generation sequencing (NGS) panel containing homozygote DNA and control genomic DNA, The composition wherein the isozygote DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
2. The composition according to claim 1, wherein the control genomic DNA is genomic DNA isolated from the blood of a normal human.
3. The composition according to claim 1, wherein the control genomic DNA is genomic DNA isolated from the patient's blood or cancer tissue.
4. The composition according to claim 1, wherein the control genomic DNA is circulating tumor DNA (ctDNA).
5. The composition according to claim 1, wherein the control genomic DNA is synthetic DNA or cloned DNA.
6. The composition according to claim 1, wherein the dilution ratio of isozygote DNA and control genomic DNA is selected from the group consisting of 1:99 to 99:
1.
7. The composition according to claim 1, wherein the dilution ratios of isozygote DNA and control genomic DNA are selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5, and 9999:
1.
8. A kit for verifying a next-generation sequencing (NGS) panel, comprising a verification composition for a next-generation sequencing (NGS) panel according to any one of claims 1 to 7.
9. (a) The step of performing next-generation sequencing (NGS) analysis on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) A step of selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the next-generation sequencing (NGS) results obtained in step (a); and (c) A method for verifying a next-generation sequencing (NGS) panel through the analysis of false negative mutations, comprising the step of analyzing the frequency of false negative mutations in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios based on the selected Ho-N pair allele or He-N pair allele, The method wherein the isozygote DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
10. (a) The step of performing next-generation sequencing (NGS) analysis on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) A step of selecting Ho-N pair alleles (Homozygote-Null pair alleles) or He-N pair alleles (Heterozygote-Null pair alleles) from the next-generation sequencing (NGS) results obtained in step (a); (c) The step of analyzing the dilution mutation detection rate, expressed as a percentage, in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios based on the selected Ho-N pair allele or He-N pair allele; and (d) A method for validating a next-generation sequencing (NGS) panel through an analysis of the limit of detection, which includes a step of analyzing the limit of detection, The method wherein the isozygote DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
11. The verification method according to claim 10, wherein the method verifies a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNVs), insertions / deletions (Indels), and chromosome amplification / deletions (amplification / deletions).
12. The verification method according to claim 10, wherein the detection limit in step (d) is expressed as a percentage of the VAF (variant allelic fraction) corresponding to the maximum dilution ratio that shows a 90% or 95% detection rate of dilution mutations in a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios.
13. (a) The step of performing next-generation sequencing (NGS) analysis on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) A step of selecting N-N pair alleles (Null-Null pair alleles) that do not contain mutations in the isozygote DNA and control genomic DNA based on the next-generation sequencing (NGS) results obtained in step (a); and (c) A method for verifying a next-generation sequencing (NGS) panel through the analysis of false positive mutations, which includes the step of analyzing the frequency of false positive mutations in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios based on the selected N-N pair allele, The method wherein the isozygote DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
14. The verification method according to claim 13, wherein the method verifies a next-generation sequencing (NGS) panel capable of detecting mutations selected from the group consisting of single nucleotide variants (SNVs), insertions / deletions (Indels), and chromosomal amplification / deletions.
15. (a) The step of performing next-generation sequencing (NGS) analysis on a DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios using an NGS panel; (b) A step of selecting N-N pair alleles (Null-Null pair alleles) that do not contain mutations in the isozygote DNA and control genomic DNA based on the next-generation sequencing (NGS) results obtained in step (a); (c) The step of analyzing the frequency of false positive mutations in each DNA mixture sample in which isozygote DNA and control genomic DNA are mixed at different dilution ratios based on the selected N-N pair allele; and (d) A method for providing information to increase the specificity of a next-generation sequencing (NGS) panel, which includes a step of analyzing variant allelic fractions (VAFs) to remove false-positive mutations and deriving a strength cutoff value, The method wherein the isozygote DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic germ cell line.
16. The information provision method according to claim 15, wherein the stringency cutoff value is a 90% stringency cutoff value that removes 90% of false positive mutations, a 95% stringency cutoff value that removes 95% of false positive mutations, or a 99% stringency cutoff value that removes 99% of false positive mutations.
17. The information providing method according to claim 15, wherein the method provides information on a strict cutoff value for removing false-positive mutations in order to increase the specificity of a next-generation sequencing (NGS) panel.
18. (a) A step of collecting analysis results for false negative variants calculated according to the method of claim 9 and analysis results for false positive variants calculated according to the method of claim 15 from the raw data of a company; and (b) A method for evaluating errors in the raw data generation process or the bioinformatics analysis process in a company's NGS panel analysis process, including a step of comparing the results of a bioinformatics analysis obtained using commercialized software from the company's raw data with the results of the analysis in step (a).
Citation Information
Patent Citations
Method for measuring the copy number of chromosomes, genes, or specific nucleotide sequences using SNP arrays
JP2011510626A
Linear DNA assembly for nanopore sequencing
WO2021108532A2
Synthetic polynucleotides and method of use thereof in genetic analysis
WO2022235315A1