Method for validation and false positive error analysis of next-generation sequencing panel by using DNA mixture

A DNA mixture-based method for NGS panels addresses the lack of standard tools for sensitivity and specificity measurement, enhancing accuracy by quantifying false positives and improving validation of NGS panels.

WO2025159599A1PCT designated stage Publication Date: 2025-07-31NATIONAL CANCER CENTER(JP)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/001552
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current methods for validating next-generation sequencing (NGS) panels lack standardized tools for measuring sensitivity and specificity, particularly in assessing false positive errors, and existing reference samples are limited in their applicability and cannot universally evaluate different NGS panels.

Method used

A method involving the use of a DNA mixture containing homozygous and control genomic DNA to perform NGS, select Ho-N and He-N pair alleles, analyze dilution mutation detection rates, and verify NGS panels by assessing false positive errors and detection limits, along with a composition for verifying NGS panels using such DNA mixtures.

Benefits of technology

Enhances the evaluation of NGS panel accuracy by quantifying false positive errors and improving sensitivity and specificity through the use of a DNA mixture standard sample, allowing for comprehensive validation of NGS panels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025001552_31072025_PF_FP_ABST
    Figure KR2025001552_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for validation and false positive error analysis of a next-generation sequencing (NGS) panel by using mixed DNA, and to a composition comprising mixed DNA for validation of a next-generation sequencing (NGS) panel. The method for validation and false positive error analysis of a next-generation sequencing (NGS) panel by using mixed DNA according to the present invention has an effect of evaluating the appropriateness of an NGS analysis result by analyzing false positive errors from a result of analyzing variations, and thus such a method for verification and false positive error analysis of a next-generation sequencing (NGS) panel by using mixed DNA and a composition containing mixed DNA for validation of a next-generation sequencing (NGS) panel can be helpfully used for validation of the next-generation sequencing (NGS) panel.
Need to check novelty before this filing date? Find Prior Art

Description

Method for validation and false positive error analysis of next-generation sequencing panels using DNA mixtures

[0001] The present invention relates to a method for verifying and analyzing false positive errors of a next-generation sequencing (NGS) panel using mixed DNA, and a composition for verifying a next-generation sequencing (NGS) panel including mixed DNA.

[0002]

[0003] To develop and apply NGS panels to clinical practice, the panel's accuracy and sensitivity must be evaluated to ensure they meet the desired standards. To this end, the US FDA has published guidelines for the validation of NGS panels, and the Center for Medicare & Medicaid Services (CMS) monitors these panels, with ongoing updates to these guidelines and monitoring methods. Furthermore, the College of American Pathologists certifies NGS tests through the CAP survey, and in Korea, the Ministry of Food and Drug Safety (MFD) certifies NGS tests. However, these certification bodies do not provide tools for measuring the sensitivity and specificity of NGS panels and for comparative evaluations.

[0004] To measure the sensitivity and specificity of NGS panels, there is a method introduced by Frampton et al. in Nature Biotechnology in 2013 (hereinafter referred to as the "Frampton method"), and the Frampton method has been primarily used in subsequent studies (Shin et al., Nature Comm 2017; and Froyen et al., Cancers 2022). The Frampton method is a method for verifying NGS panels using a DNA pool mixed with DNA from HapMap cell lines with well-known base sequences or a DNA pool mixed with DNA from cancer cell lines. However, the Frampton method also had difficulties in assessing the sensitivity of NGS panels due to errors that occurred when the expected allele frequency (allelic fraction) and the actual value were different in the NGS results measured by mixing exactly 10 or so DNAs in equal amounts during the dilution process. In particular, it has limitations in that it cannot measure specificity, so it is not used to establish guidelines.

[0005] In addition to the Frampton method, methods have been introduced to measure the sensitivity of mutation detection by mixing multiple sequences with mutations, diluting them with normal genomic DNA at a certain ratio, and testing them with an NGS panel to determine up to a dilution ratio of the mixed mutations. A method that can detect mutations of less than 1% by providing a sequence mixture containing 40 variants in 28 genes (Seraseq® Tumor Mutation DNA Mix v2 AF 10 HC, SeraCase, Mildford) and mixing it with wild-type DNA (GM24385) from a person known to have no variants in the mixture at a desired ratio and using it for NGS panel analysis has been commercialized (Vicente-Garces et al., Frontiers in Mol Biosciences, 2022). Additionally, the company provides a DNA mixture containing wild-type DNA and various concentrations of each DNA to measure the sensitivity of 25 variants in 16 genes, enabling monitoring of ctDNA to assess the sensitivity of NGS panels (Seraseq® ctDNA Complete™ Reference Material, SeraCase, Mildford). However, these reference samples can only be used to test specific NGS panels and cannot be applied to other NGS panels. Crucially, they are limited in that they cannot measure specificity.

[0006] In this situation where there is no standard material that can universally measure sensitivity and specificity, the researchers conducted a study to obtain the sensitivity and specificity of NGS using a DNA mixture standard sample containing homozygous DNA, heterozygous DNA, or control genomic DNA, and to improve the analysis parameters for sensitivity and specificity. In particular, they developed a groundbreaking method to improve false positive and false negative errors of NGS, and completed the present invention.

[0007]

[0008] The present invention aims to solve the above-mentioned problems and other problems related thereto.

[0009] One exemplary purpose of the present invention is

[0010] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0011] (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results according to step (a);

[0012] (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and

[0013] (d) A method for verifying a next generation sequencing (NGS) panel through analysis of the limit of detection, including a step of analyzing the limit of detection, is provided.

[0014] Another exemplary purpose of the present invention is

[0015] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0016] (b) a step of selecting false positive mutations from the next generation sequencing (NGS) results according to step (a); and

[0017] (c) A method for verifying a next-generation sequencing (NGS) panel through analysis of a false positive error rate is provided, including a step of analyzing a false positive error rate in the DNA mixture sample based on the selected false positive mutations.

[0018] Another exemplary purpose of the present invention is

[0019] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0020] (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results according to step (a);

[0021] (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and

[0022] (d) It provides a method for verifying a next-generation sequencing (NGS) panel, including a step of comparing and analyzing the expected VAF and the actual VAF according to the above dilution ratio.

[0023] Another exemplary purpose of the present invention is

[0024] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0025] (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results obtained by step (a);

[0026] (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and

[0027] (d) Provided is a method for verifying a next-generation sequencing (NGS) panel, including a step of analyzing the deviation of the number of variant reads (read count) per chromosomal region used for detecting dilution mutations in step (c).

[0028] Another exemplary object of the present invention is to provide a composition for validation of a next generation sequencing (NGS) panel comprising a DNA mixture comprising two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof.

[0029] Another exemplary object of the present invention is to provide a composition for validation of a next generation sequencing (NGS) panel comprising a mixture of two or more cells selected from homozygous cells, control cells, or a combination thereof.

[0030]

[0031] The technical problem to be achieved according to the technical idea of ​​the invention disclosed in this specification is not limited to the problem to solve the above-mentioned problem, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0032]

[0033] This is explained in detail as follows. Meanwhile, each description and embodiment disclosed in this application can also be applied to each other description and embodiment. In other words, all combinations of the various elements disclosed in this application fall within the scope of this application. Furthermore, the scope of this application is not limited by the specific descriptions described below.

[0034]

[0035] In order to solve the problems presented in the above background technology, the inventors of the present invention have conducted research to obtain the sensitivity and specificity of NGS using a DNA mixture standard sample mixed with homozygous DNA or control genomic DNA, and since there was a limit to the information that could be obtained from the analysis results of sensitivity and specificity during the research, further research was conducted to improve the analysis parameters.

[0036] For example, it was difficult to directly compare the number of false-positive variants according to each NGS method because the size of each target sequence was different, and it was easy to compare the specificity of various NGS methods by calculating the false-positive error rate by dividing the number of false-positive errors by the number of target sequence bases of a specific NGS method. It is more efficient and intuitive to calculate the limit of detection by integrating them rather than calculating the limit of detection separately from the Ho-N pair allele (Homozygote-Null pair allele) or He-N pair allele (Heterozygote-Null pair allele).

[0037] In addition, when analyzing NGS, we find various cases where false-positive mutations occur in alleles other than NN (Null-Null) pairs or reference base-reference base pairs (RR pairs). Specifically, false-positives occur in alleles where only mutations are found (variant-variant pairs, VV pairs) or alleles where the reference base and the mutation are mixed (reference base-variant mixed pairs, RV mixed pairs or RV pairs). However, during the analysis process, we found that the case where the reference base was detected as a false positive in the VV pair was much higher than the detection rate of other false-positive mutations. Since such errors are caused by bias in bioinformatics analysis, we judged that removing them and obtaining the entire false-positive mutation would be advantageous for comparing and analyzing the false-positive errors of various NGSs.

[0038] In addition to parameters such as sensitivity and specificity, when analyzing the linear relationship between the expected variant allelic fraction (VAF) and the VAF measured in NGS, we found that there were cases where the linear relationship was not found and the line was curved, which newly derived that this analysis will be helpful in measuring NGS errors. In addition, if the read depth of the variant for each chromosome is shown, we can see how the capture or amplification efficiency of the NGS kit is distributed for each chromosome, and from these results, if the deviation for each chromosome location is large, it can be determined that the false negative error is high and improvement is needed. In another analysis, when showing the regions of chromosome where false positive errors occurred, there are alleles with a particularly high number of false positives, and since the mutation results of these alleles are the result of errors, it is desirable to display the false positive error rate for specific alleles in the NGS result report. In the present invention, such false positive errors can be discovered, and it was newly derived that these alleles with a high number of false positives can be used as parameters to improve the specificity of NGS.

[0039]

[0040] As one aspect for achieving the above object, the present invention provides a method for producing a DNA mixture sample comprising two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof, comprising the steps of: (a) performing next-generation sequencing (NGS) using a next-generation sequencing (NGS) panel;

[0041] (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results according to step (a);

[0042] (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and

[0043] (d) A method for verifying a next generation sequencing (NGS) panel through analysis of the limit of detection, including a step of analyzing the limit of detection, is provided.

[0044] The term "allele" of the present invention refers to a gene that forms a pair of homologous chromosomes and has different characteristics, and refers to the DNA sequence of the allele. Each allele exists in a homozygous and heterozygous state, and a "homozygous" allele refers to a case where two genes of one individual have alleles with the same base, and a "heterozygous" allele refers to a case where one individual or cell has alleles with different bases.

[0045] The term "homozygous cell" in the present invention means a cell in which almost all alleles are composed of homozygous alleles, except for a few alleles with exceptions, and "homozygous DNA" means DNA isolated from a homozygous cell.

[0046] Additionally, a "heterozygous cell" refers to a cell that is composed of heterozygous alleles from many alleles because the traits inherited from each parent are different, and "heterozygous DNA" refers to DNA isolated from a heterozygous cell, and is a concept distinct from heterozygous alleles.

[0047] In one embodiment of the present invention, the DNA mixture sample, which is a mixture of two or more cell-derived DNAs selected from the homozygous DNA, the control genomic DNA, or a combination thereof, may be a mixture of two or more homozygous DNAs having different mutations; a mixture of the homozygous DNA and the control genomic DNA; or a combination thereof.

[0048] An example of two or more homozygous DNAs having different mutations is that each homozygous DNA is isolated from H. mole cells derived from two different patients, and each homozygous DNA may have different mutations and insertions and deletions (indels) that make up the homozygous DNA.

[0049] In one embodiment of the present invention, the homozygous DNA may be DNA isolated from a hydatidiform mole (H. mole) cell line, a parthenogenetic cell line, or a cell in which most alleles are homozygous.

[0050] In one embodiment of the present invention, the "control genomic DNA" may be homozygous DNA; heterozygous DNA; or genomic DNA isolated from normal human blood cells or tissue cells, cultured blood cells, tissue cells, or cancer cells, or from blood cells or cancer tissue cells of a patient. Furthermore, the control genomic DNA may be circulating tumor DNA (ctDNA). Furthermore, the control genomic DNA may be synthetic DNA or cloned DNA.

[0051] In one embodiment of the present invention, the dilution ratio of the homozygous DNA and the control genomic DNA may be selected from the group consisting of 1:99 to 99:1. Preferably, the dilution ratio of the homozygous DNA and the control genomic DNA may be selected from the group consisting of 1:9999, 2.5:9997.5, 5:9995, 1:999, 2.5:997.5, 5:995, 1:99, 2.5:97.5, 5:95, 10:90, 20:80, 50:50, 80:20, 90:10, 95:5, 97.5:2.5, 99:1, 995:5, 997.5:2.5, 999:1, 9995:5, 9997.5:2.5 and 9999:1, but is not limited thereto.

[0052] In one embodiment of the present invention, the homozygous DNA or control genomic DNA may be isolated from fresh cells, fresh tissues, or formalin fixed paraffin embedded (FFPE) cell tissues.

[0053] The term 'fresh cell' in the present invention refers to a living cell obtained from culture and may also include frozen cells that have not been chemically treated, and 'fresh tissue' refers to a tissue directly separated from a living body and may also include frozen tissue that has not been chemically treated.

[0054] The 'formalin fixed paraffin embedded (FFPE) cell tissue' of the present invention can be obtained without limitation by methods known in the art, and can be obtained through, for example, tissue fixation and embedding processes.

[0055] The term "next-generation sequencing (NGS)" used in the present invention refers to a technology capable of analyzing millions of base sequences simultaneously, also known as massively parallel sequencing or high-throughput sequencing. In the present invention, the terms "next-generation sequencing" and "NGS" are used interchangeably.

[0056] The term "Ho-N pair allele (Homozygote-Null pair allele)" of the present invention means a pair in which the alleles of the homozygous DNA (or homozygous cell) including H. mole DNA among the selected alleles and the control genomic DNA (or control genomic cell) are each composed of a homozygous variant and a null without a mutation, and the term "He-N pair allele (Heterozygote-Null pair allele)" of the present invention means a pair in which the homozygous DNA among the selected alleles is null and the control genomic DNA is composed of a heterozygous variant.

[0057] More specifically, the criteria for selecting He-N pair alleles may vary depending on the type and number of DNAs included in the DNA mixture constituting the sample. That is, in a sample where heterozygous DNA and homozygous DNA are mixed in a 50:50 ratio, if the reference base (or null allele) is in the homozygous DNA and the heterozygous variant is in the heterozygous DNA, the He-N pair is selected, and at this time, the ratio of the reference base and the variant allele is 3:1. However, when mixing multiple mixed samples, even if each single DNA is mixed in equal amounts, the allele ratio of the reference base and the variant may not be 3:1. For example, when a reference standard material is made by mixing equal amounts of DNA from two homozygous cells and one type of control genomic DNA in a DNA mixture, if the reference base appears in both homozygous DNAs for a specific allele, and a heterozygous variant in which the reference base and the variant are mixed appears in the control genomic DNA, the ratio of null alleles to variant alleles becomes 5:1.

[0058] In the present invention, "informative allele" refers to alleles that have certain conditions for analyzing sensitivity or specificity. For example, Ho-N pair and He-N pair alleles can be said to be informative alleles used to measure sensitivity. In these pairs, the mutant base or reference base exists only in either homozygous DNA (DNA1) or control genomic DNA (DNA2), and the sensitivity can be measured by determining whether a diluted base among the informative alleles is detected in the analysis result of a mixed DNA sample. In addition, informative alleles necessary for measuring specificity are distinguished into RR pair, VV pair, VR mixed pair, etc., and when a base that does not exist in DNA1 and DNA2 appears in the analysis result of a mixed DNA sample, it is detected as a false-positive allele, thereby measuring specificity. These informative alleles vary depending on whether the alleles that make up the homozygous cells and the control genome cells are variant, null, or heterozygous when the reference standard material is made by mixing homozygous cells and control genome cells.

[0059] In the present invention, the "dilution mutation detection rate" is defined as the number of diluted mutations found in a dilution ratio of the expected mutation among the Ho-N pair allele or He-N pair allele in a DNA mixture sample in which homozygous DNA and control genomic DNA are mixed, expressed as a percentage.

[0060] In the present invention, "detection limit" refers to the minimum VAF (variant allelic fraction) value at which the detection rate is 90% or higher. Normally, the detection rate, which serves as the standard for the detection limit, can be lowered or increased. Here, the condition of a detection rate of 90% or higher, which defines the detection limit, can use various values, such as 95% or 99%, depending on the conditions.

[0061] The term 'VAF (variant allelic fraction)' of the present invention means a value obtained by dividing the number of reads (read count) of a sequence detected as a variant at a specific allelic position by the total number of reads (read count) obtained by adding the sequence detected as a variant and a reference sequence base at a specific allelic position.

[0062] As another aspect for achieving the above purpose, the present invention

[0063] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0064] (b) a step of selecting false positive mutations from the next generation sequencing (NGS) results according to step (a); and

[0065] (c) A method for verifying a next-generation sequencing (NGS) panel through analysis of a false positive error rate is provided, including a step of analyzing a false positive error rate in the DNA mixture sample based on the selected false positive mutations.

[0066] The above ‘homozygous DNA’, ‘control genomic DNA’ and ‘next-generation sequencing (NGS)’ are as described above.

[0067] In one embodiment of the present invention, the DNA mixture sample containing two or more cell-derived DNAs selected from the homozygous DNA, the control genomic DNA, or a combination thereof may be a mixture of two or more homozygous DNAs having different mutations; a mixture of the homozygous DNA and the control genomic DNA; or a combination thereof.

[0068] The term "false positive mutation" in the present invention refers to a mutation in which a test result that should have been originally negative is erroneously positive, and is defined as a case in which a base that was not detected in another analysis of a homozygous DNA or control genomic DNA sample is detected as a result of NGS analysis of the sample.

[0069] In one embodiment of the present invention, the false positive mutation may be a case where a mutation is found to exist among null alleles in which no mutation is found in two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof, as a result of analyzing a mixed DNA sample by NGS.

[0070] In one embodiment of the present invention, the false positive mutation may be a case where, among alleles (VV pairs) in which only mutations are found and no reference bases are found in two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof, a mutation or reference base different from the VV pair is found in the results of NGS analysis of a DNA mixture sample.

[0071] In one embodiment of the present invention, the false positive mutation may be a case where, among alleles (RV pair alleles) in which a reference base and a variant are mixed in two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof, a new mutation other than a base constituting the RV pair is discovered as a result of analyzing a DNA mixture sample by NGS.

[0072] In the present invention, the step (c) may additionally include a process of eliminating VVR false positive error results in which a reference base is found in a VV pair. The 'VVR false positive error' refers to a false positive error that occurs when a VV pair is judged to be a reference base.

[0073] In the present invention, the step (c) may additionally include a process of excluding erroneous results of alleles in which false positive errors appear at a high frequency. The alleles in which false positive errors appear at a high frequency refer to alleles in which false positive errors appear at a high concentration, and these are caused by errors that occur in the process of producing raw data for NGS, such as differences in experimental materials used for producing raw data for NGS, differences in minute sample concentrations during the experimental process using these materials, or differences in PCR devices, and are unrelated to the performance of the NGS panel. Therefore, in order to improve the accuracy of the NGS panel verification process, erroneous results of alleles in which false positive errors appear at a high frequency can be excluded.

[0074] In the present invention, the step (c) may additionally include a process for examining the batch effect.

[0075] The term "batch effect" in the present invention refers to a change in data caused by conditions other than biological factors that may occur during an experiment. In particular, when NGS samples are sent to a company at different times and the results are analyzed, it refers to an error in the form of finding differences in results depending on the time of sending, showing a different pattern from differences found in samples sent at the same time. A specific example of batch effect testing includes examining whether reference standard material was sent simultaneously with clinical or research samples to produce NGS raw data.

[0076]

[0077] As another aspect for achieving the above purpose, the present invention

[0078] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0079] (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results according to step (a);

[0080] (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and

[0081] (d) A method for verifying a next-generation sequencing (NGS) panel is provided, including a step of comparing and analyzing the expected VAF and the actual VAF according to the above dilution ratio.

[0082] The above 'homozygous DNA', 'control genomic DNA', 'next-generation sequencing (NGS)', 'Ho-N pair allele', 'He-N pair allele', 'dilution mutation detection rate' and 'VAF (variant allelic frequency)' are as described above.

[0083] In one embodiment of the present invention, the DNA mixture sample containing two or more cell-derived DNAs selected from the homozygous DNA, the control genomic DNA, or a combination thereof may be a mixture of two or more homozygous DNAs having different mutations; a mixture of the homozygous DNA and the control genomic DNA; or a combination thereof.

[0084] As another aspect for achieving the above purpose, the present invention

[0085] (a) a step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof;

[0086] (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results obtained by step (a);

[0087] (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and

[0088] (d) A method for verifying a next-generation sequencing (NGS) panel is provided, including a step of analyzing the deviation of the number of variant reads (read count) per chromosomal region used for detecting dilution mutations in step (c).

[0089] The above 'homozygous DNA', 'control genomic DNA', 'next-generation sequencing (NGS)', 'Ho-N pair allele', 'He-N pair allele' and 'dilution mutation detection rate' are as described above.

[0090] In one embodiment of the present invention, the DNA mixture sample containing two or more cell-derived DNAs selected from the homozygous DNA, the control genomic DNA, or a combination thereof may be a mixture of two or more homozygous DNAs having different mutations; a mixture of the homozygous DNA and the control genomic DNA; or a combination thereof.

[0091] The above "variant read count" refers to the number of times each mutation is read for a chromosomal region. If there is a deviation in the read count for each chromosomal region, the sensitivity of the NGS panel is relatively low and the reliability of the NGS results is reduced. Therefore, it is necessary to analyze the deviation in the variant read count for each chromosomal region used for dilution mutation detection to verify an appropriate next-generation sequencing (NGS) panel.

[0092] In another aspect for achieving the above object, the present invention provides a composition for verifying a next-generation sequencing (NGS) panel, comprising a DNA mixture in which two or more cell-derived DNAs are mixed, selected from homozygous DNA, control genomic DNA, or a combination thereof.

[0093] The above 'next-generation sequencing (NGS)' is as described above.

[0094] In one embodiment of the present invention, the DNA mixture comprising two or more cell-derived DNAs selected from the homozygous DNA, the control genomic DNA, or a combination thereof may be a mixture of two or more homozygous DNAs having different mutations; a mixture of the homozygous DNA and the control genomic DNA; or a combination thereof.

[0095] In one embodiment of the present invention, the homozygous DNA may be DNA isolated from a hydatidiform mole cell line or a parthenogenetic cell line.

[0096] In one embodiment of the present invention, the control genomic DNA may be homozygous DNA; heterozygous DNA; or genomic DNA isolated from blood cells or tissue cells of a normal person, cultured blood cells, tissue cells, or cancer cells, or blood cells or cancer tissue cells of a patient.

[0097] In one embodiment of the present invention, the DNA mixture may be isolated from fresh cells, fresh tissues, or formalin fixed paraffin embedded (FFPE) cell tissues.

[0098] In one embodiment of the present invention, the composition may be a reference standard material for verification of a next-generation base sequence analysis panel.

[0099] As another aspect for achieving the above object, the present invention provides a composition for verifying a next-generation sequencing (NGS) panel, comprising a mixture of two or more cells selected from homozygous cells, control cells, or a combination thereof.

[0100] The above ‘homozygosity’ and ‘next-generation sequencing (NGS)’ are as described above.

[0101] In one embodiment of the present invention, the mixture of two or more cells selected from the homozygous cells, control cells, or a combination thereof may be a mixture of two or more homozygous cells having different mutations; a mixture of two or more homozygous cells having different mutations and control cells; a mixture of homozygous cells and control cells; or a combination thereof.

[0102] In one embodiment of the present invention, the homozygous cell or control cell may be a fresh cell or a formalin fixed paraffin embedded (FFPE) cell.

[0103] In one embodiment of the present invention, the composition may include a mixture of two or more cells selected from homozygous cells, control cells, or a combination thereof, in the form of a formalin fixed paraffin embedded (FFPE) cell block.

[0104] The 'formalin-fixed paraffin cell block' of the present invention can be obtained through a process of collecting, fixing, and embedding single cells or complex cells. In general, when formalin-fixed paraffin embedded (FFPE) cell tissues are used, there is a difference in the sensitivity and specificity results of T-NGS using reference standard material using DNA isolated from fresh cells or fresh tissues due to DNA fragmentation. Therefore, formalin-fixed paraffin embedded (FFPE) cells can be produced in the form of a cell block to eliminate or reduce the occurrence of such differences in sensitivity and specificity.

[0105] In one embodiment of the present invention, the homozygous cell may be a hydatidiform mole cell or a parthenogenetic cell.

[0106] In one embodiment of the present invention, the control cell may be a homozygous cell; a heterozygous cell; a blood cell or tissue cell of a normal person; a cultured blood cell, tissue cell, or cancer cell; or a blood cell or cancer tissue cell of a patient. The blood cell or tissue cell of a normal person; the cultured blood cell, tissue cell, or cancer cell; or the blood cell or cancer tissue cell of a patient may be a cell in which alleles are mixed as heterozygotes and homozygotes.

[0107] In one embodiment of the present invention, the composition may be a reference standard material for verification of a next-generation base sequence analysis panel.

[0108]

[0109] The method for verifying and analyzing false positive errors of a next-generation sequencing (NGS) panel using mixed DNA according to the present invention has the effect of evaluating the appropriateness of NGS analysis results by analyzing false positive errors from the results of analyzing variants, and thus the method for verifying and analyzing false positive errors of a next-generation sequencing (NGS) panel using mixed DNA and the composition for verifying a next-generation sequencing (NGS) panel including mixed DNA can be usefully used for verifying a next-generation sequencing (NGS) panel.

[0110]

[0111] Figure 1 shows the variant detection rate in the theoretical VAF (variant allele fraction) of the Ho-He-N pair allele (X-axis: theoretical VAF (%) of the Ho-He-N pair allele, Y-axis: variant detection rate in which a variant is detected among the Ho-He-N pair alleles having the theoretical VAF value).

[0112] Figure 2 shows the types of errors judged as false positives (DNA1 and DNA2 are homozygous DNA and control genomic DNA, respectively, and the alphabets indicated in red lowercase letters represent bases not found in DNA1 and DNA2).

[0113] Figure 3 shows the analysis of false positive error rates according to the analysis method by company and NGS (X-axis: analysis method by company and NGS, Y-axis: false positive error rate (number of false positive errors / number of target bases)). A. False positive error rate in RR pairs. B. False positive error rate in RV pairs. C. False positive error rate in VV pairs. D. False positive error rate in VV pairs after removing VVR errors.

[0114] Figure 4 compares the difference in the overall false positive error rate between the results that do not include (VVR-) and the results that include (VVR+) VVR errors in which the reference base appears in the VV pair in the in-house analysis results of BB Company.

[0115] Figure 5 shows the results of sensitivity measurements based on the dilution ratio of the reference base (X-axis: Allele fraction percentage (AF percentage) of the reference base, Y-axis: reference base detection rate).

[0116] Figure 6 shows the analysis of the overall false positive error rate after removing reference base detection errors from the VV pair (X-axis: analysis method according to company and NGS, Y-axis: overall false positive error rate).

[0117] Figure 7 shows the correlation between the expected and actual values ​​of the dilution ratio of the variant allele of the Ho-N pair and He-N pair alleles according to each company and analysis method (X-axis: actual VAF value of the variant allele (actually observed VAF), Y-axis: expected VAF value according to the dilution ratio).

[0118] Figure 8 shows the results of read count analysis of variants analyzed from the results of several NGS service providers (AA to DD) (X-axis: chromosome position number, Y-axis: read count, median value is indicated in the center of the box, and the top and bottom of the box indicate 25% and 75% of the read count value).

[0119] Figure 9 shows the distribution of alleles that show false positive errors in the CC company's NGS results (X-axis, alleles that show false positive errors are listed in order according to the chromosomal region. Y-axis, allele frequency (allele fraction) of the erroneous base in the alleles that show errors).

[0120] Figure 10 illustrates a method for mixing two different cells and producing a formalin-fixed paraffin-embedded (FFPE) tissue block to provide a reference standard material.

[0121]

[0122] Hereinafter, the present invention will be described in more detail through the following examples. However, these examples are intended to exemplify the present invention and the scope of the present invention is not limited to these examples.

[0123]

[0124] Example 1. Sensitivity analysis integrating variant alleles in Ho-N pairs and He-N pairs

[0125] 1.1. Preparation of DNA mixture

[0126] In the present invention, for the validation of the NGS panel, the sensitivity and specificity of the NGS panels of three companies were compared and analyzed using the human hydatidiform mole derived H. mole DNA (hydatidiform mole, DNA1), which is homozygous DNA, the control genomic DNA (DNA2), which is heterozygous DNA derived from the patient's blood cells, and a mixture sample of these two DNAs as reference standard materials. In order to use the separated DNAs as a DNA mixture, several mixture complex samples were prepared by mixing H. mole DNA (DNA1) and the control genomic DNA (DNA2) at various ratios. The human hydatidiform mole derived H. mole DNA used in the present invention was purchased from Coriell (NA07489, Camden, NJ), and the control genomic DNA was extracted from the blood of a collaborator with the approval of the Institutional Review Board of the National Cancer Center.

[0127]

[0128] 1.2. NGS analysis

[0129] The prepared H. mole DNA (DNA1), control genomic DNA (DNA2), and a DNA mixture containing H. mole DNA and control genomic DNA were sent to AA Company, BB Company, and CC Company, respectively, and experiments were performed using Illumina's Novaseq 6000 NGS equipment using each company's NGS panel. Afterwards, a file analyzing the mutations using the method provided by each company was received from the FASTQ file obtained in the experiment, and verification of the NGS panel was performed.

[0130]

[0131] 1.3. NGS panel sensitivity analysis through detection limit analysis

[0132] In order to measure the detection limit of variants of each company's NGS using a DNA mixture containing H. mole DNA and control genomic DNA as a reference standard material, all alleles corresponding to the Ho-N pair or He-N pair were first selected and defined as Ho-He-N pair alleles. In order to integrate the detection limit results of each selected Ho-He-N pair, the theoretical VAF (variant allele fraction) value of the He-N pair was calculated as half of that of Ho-N, thereby allowing the results of each Ho-N pair or He-N pair to be integrated into a single result to calculate the sensitivity.

[0133] Specifically, among the Ho-He-N pair alleles, in the case of the Ho-N pair, the theoretical VAF (or VAF expectation value) that the variant occupies among the total alleles is 50% in the case of CH50 (a sample in which H. mole and control DNA are mixed 50:50), and in the case of the He-N pair, the VAF expectation value of the CH50 sample is 25%. In the same way, all VAF expectation values ​​of all Ho-He-N pair alleles in samples with other mixing ratios were obtained, and the number of variants actually detected at a specific VAF expectation value among the Ho-He-N pair alleles was defined as the variant detection rate at a specific VAF expectation value and calculated as the value divided by the number of Ho-He-N pair alleles having a specific VAF expectation value.

[0134] These variant detection rates were displayed graphically, and the minimum VAF value at which the variant detection rate was 90% or higher at each VAF expectation value was determined as the detection limit. Figure 1 shows the variant detection rates at theoretical VAF values ​​for Ho-He-N pair alleles, where AA to DD represent T-NGS kits from different companies, inhouse represents the variant calls analyzed by the in-house bioinformatics analysis method used by each company, and dragen represents the analysis using the DRAGEN system from Illumina using the fastq files provided by each company to relatively compare the results of the bioinformatics methods. For example, AA-dragen represents the variant analysis using the DRAGEN system from the T-NGS raw data obtained from company AA, and AA-inhouse represents the variant analysis using the bioinformatics analysis method used by company AA from the T-NGS raw data obtained from company AA. In addition, DD-default represents the variant call result analyzed by the bioinformatics analysis method used by DD company from the T-NGS raw data obtained from DD company, and DD AF>0.005 represents that the variants were analyzed by the bioinformatics analysis method used by DD company from the T-NGS raw data obtained from DD and only the variants with AF>0.005 were selected. The X-axis represents the theoretical VAF of the Ho-He-N pair allele, and the Y-axis represents the variant detection rate among the Ho-He-N pair alleles with the theoretical VAF value. When estimating the detection limit, it means that the detection limit of VAF is 20% for AA-inhouse and about 1% for DD-default.These results indicate that the results of analyzing the AA company kit in-house can detect about 90% of variants when the variant accounts for 20%, so the sensitivity for detecting variants is relatively low. On the other hand, the results of analyzing the DD company kit in-house can detect about 90% of variants when the variant accounts for 1%, so the sensitivity is relatively high.

[0135]

[0136] Example 2. Method for comparing the degree of false positives by calculating the false positive error rate.

[0137] When analyzing the number of false positive errors (FP errors) in a mixture of homozygous DNA (DNA1) and control genomic DNA (DNA2), or in a mixture of homozygous DNAs with different mutations, it is difficult to directly compare the number of false positives between companies' NGS because the number of targets and the number of bases of the targets being analyzed are different. To address this, we calculated and compared the false positive error rate (FP error) in the NGS results of different companies. The FP error rate is the number of false positives found in a DNA mixture sample divided by the number of bases of the NGS target gene. Even when the number of bases targeted by NGS differs between companies, the FP error rate generated from the NGS results can be easily used to directly compare the relative degree of false positive errors. For reference, in cases where the number of target genes is the same, such as whole exome sequencing, the number of false positive errors can also be used to compare the relative degree of error.

[0138] False positive errors were defined as bases that were not detected in the analysis results of DNA1 and DNA2, but were newly detected in the analysis results of the mixed sample of DNA1 and DNA2. Depending on whether the bases of DNA1 and DNA2 are reference bases or variant bases, the status of alleles can be divided into three pairs: 1) RR pair, when only the reference base was analyzed in DNA1 and DNA2; 2) VV pair, when only the variant base was analyzed in both DNA1 and DNA2; and 3) VR mixed, when both the reference base and variant base(s) were analyzed in DNA1 and DNA2. Since the frequency and error rate of false positive errors differed in each of these three types of pairs, each error was measured separately (Fig. 2).

[0139] In order to obtain the false positive error rate using samples mixed with DNA1 and DNA2, false positives were determined in the following three cases: 1) If a mutation is found in the NGS analysis results of a specific DNA mixture sample among the RR pairs, it is determined that the specific DNA mixture has a false positive for the pair allele (if an error mutation such as a, not the reference base, appears in the RR pair of Figure 2). 2) If another mutation or reference base that was not found in DNA1 and DNA2 is found in the NGS analysis result of a specific DNA mixture sample among the VV pairs, it is determined that the specific DNA mixture has a false positive for the VV pair allele (if an error mutation or an error reference base such as g or t appears in the VV pair of Fig. 2), 3) If a new mutation that was not found in DNA1 and DNA2 is found in the NGS analysis result of a specific DNA mixture sample among the RV mixed pairs, it is determined that the specific DNA mixture has a false positive for the RV pair mixed allele (if an error mutation such as c appears in the RV mixed pair of Fig. 2).

[0140] Here, the reference base refers to the base corresponding to a specific chromosomal location in the reference sequence used to compare the DNA sequence of the raw data analyzed by NGS. Various reference sequences can be used, but here, it refers specifically to the reference sequence used to determine chromosomal location and variant calling during bioinformatics analysis of raw data to determine sensitivity and specificity. Examples of such reference sequences include the GRCh37 Genome Reference Consortium Human Build 37 (GRCh37) or hg19.

[0141] To measure the false positive error rate for the present invention, for Company BB, errors occurring in all sequences in the BED file were detected. For the remaining companies, namely, Company AA, Company CC, and Company DD, only errors found in the exon region of the target gene were detected, and from the in-house analysis results conducted by each company, only errors in the exon region were selected to calculate the false positive error rate.

[0142] To obtain RR pairs for measuring the false positive error rate, cases in which no variants were detected at all in H. mole DNA and control genomic DNA were included in the RR pair. In the case of VV pairs, all cases in which the two DNAs, DNA1 and DNA2, contained even a small amount of variants were included in the VV pair. And in the case of RV mixed pairs, all cases in which the two DNAs, DNA1 and DNA2, contained both the reference base and the variant were included.

[0143] To calculate the false positive error rate, the number of alleles corresponding to each VV pair and VR pair was calculated, and the number of false positive errors occurring in each pair was divided by the calculated number of alleles corresponding to each VV pair and VR pair, and the error rate in each VV pair and RV pair was calculated respectively. To calculate the false positive error rate of the RR pair, in the case of BB company, the number of alleles of the VV pair and RV pair was subtracted from the number of bases of the entire sequence in the BED file, and the remaining number was calculated as the total number of alleles of the RR pair. In the case of the other companies, the number of alleles of the RR pair was calculated in the same way, considering only the number of bases of the target gene exon. The number of false positive errors occurring in each RR pair site was divided by the number of all alleles of the RR pair, and the false positive error rate of the RR pair was calculated.

[0144] For example, calculating the false positive error rate, in the results of Figure 2, the number of RR pairs is 10, and 1 false positive error occurred in the mixed sample, so the false positive error rate of the RR pair is 0.1 or 10%. Similarly, calculating, the error rate of the VV pair is 0.2, and the error rate of the VR mixed pair is 0.1. In addition, the total number of alleles is 30, and 4 errors were found in the mixed sample, so the false positive error rate is 4 / 30 or 13.3%.

[0145]

[0146] Example 3. Method for removing false positive error results in which a reference base appears in a VV pair and analyzing the false positive error rate

[0147] In fact, the false positive error rates found in the three pairs by analyzing the NGS results of several companies were as shown in AC in Fig. 3. A is the false positive error rate in the RR pair, B is the false positive error rate in the RV pair, C is the false positive error rate in the VV pair, and D is the final version of the analysis result of the false positive error rate after performing a transformation to exclude the VVR errors that result in false positives as reference bases in the VV pair.

[0148] We confirmed that the analysis results for the VV pair had an unusually high error rate compared to the analysis results for other pairs. This may be due to a high rate of PCR errors at bases with a high incidence of errors. However, a detailed analysis of the BB-in-house sample analysis data revealed that false positive errors occurred in most VV pairs (error rates approaching 1). Based on these results, we believe that this is likely a systematic error occurring in the bioinformatics analysis process rather than an error occurring in a specific company's NGS experiment.

[0149] Further analysis of the reason why false positive errors occur particularly frequently in VV pairs revealed that there were particularly many cases in which the appearance of the reference base in the VV pair was judged as an error, and when these errors were removed, the false positive error rate was significantly reduced, as shown in Figure 3D. It can be thought that the appearance of the reference base in this way is due to the fact that the criteria for judging a specific allele as a mutation and the criteria for judging it as a reference base are different in the bioinformatics analysis process in NGS results where mutations and reference bases are mixed.

[0150] To explain a little more, if the raw data has a specific allele site where the reference base is mostly mixed with a small amount of variant bases, the conditions for calling that there is a variant base at this allele site are strict; however, if the raw data has a majority of variant bases and a small amount of reference bases, the conditions for calling that there is a reference base at this allele site are more permissive. As a result, if the reference base exists in the same ratio in a specific VV pair of the raw data, the reference base will be called much more often than the variant base, and therefore, the false positive error that occurs when the reference base is called in the VV pair will increase sharply. Therefore, the false positive error that occurs when the VV pair is called as the reference base is defined as the VVR error.

[0151] To determine how much these VVR errors affect the overall false-positive error rate, we evaluated the overall false-positive error rate calculated by removing VVR errors and the overall false-positive error rate calculated without removing them. The median false-positive error rate differed by more than twice depending on whether VVR errors were included or not (Fig. 4). This means that if VVR errors are not errors in the actual NGS results but errors in the bioinformatics analysis process that makes variant calls, removing VVR errors in which the reference base appears in the VV pair can more accurately determine the overall false-positive error rate and allow for more accurate comparison and analysis of the error rates of various other NGS panels.

[0152] This study aimed to determine whether the reference base bias phenomenon, in which the criteria for calling a reference base as true differs when a reference base appears in a specific sequence and when a variant base appears, is also observed in other analyses using the data analyzed in this study. Sensitivity can be measured from NGS data in two ways: 1) sensitivity when measured based on the detection rate of the variant, and 2) sensitivity when measured based on the detection rate of the reference base. If there is a difference in the sensitivity values ​​according to these two methods, it means that this reference base bias appears repeatedly, and this can be judged as an error that occurs during the bioinformatics analysis process. In this analysis process, to confirm the variant detection rate, the detection rate of the diluted variant was calculated from the expected dilution rate of the variant. In contrast, in order to confirm the detection rate of the reference base, the expected value of the reference base dilution ratio was calculated to determine how diluted the reference base was in the alleles of each mixed sample, and the detection rate of the reference base in the alleles corresponding to the expected value of the reference base dilution ratio of the DNA mixed sample was calculated and defined as the reference base detection rate, and analyzed.

[0153] When we measured the sensitivity by detecting the reference base based on the dilution ratio of the reference base, we obtained results that detected more than 80% up to 0.5% AF (allele fraction) under all conditions such as multiple companies, in-house, and DRAGEN analysis, and we could also see that there was no difference between the analysis methods of each company (Fig. 5). These results confirmed that the sensitivity was significantly increased compared to the sensitivity measured based on the actual mutation (1-25%), confirming the bias (reference base bias), and these results suggest that measuring the sensitivity based on the detection of the reference base may lead to erroneous conclusions.

[0154] To summarize the results above, we simply converted the detection rate of variants into the detection rate of the reference base for sensitivity measurement, but we obtained results where the T-NGS sensitivity of all companies increased. This result suggests that reference base bias may exist in the bioinformatics analysis process, and that most VVR errors may be simple errors that occur repeatedly in the bioinformatics analysis process rather than actual errors. Therefore, removing VVR errors related to reference base bias that occurs in the bioinformatics analysis process and determining the overall false positive error rate would be a way to determine the accurate false positive error rate.

[0155] As described above, it was confirmed that reference base bias can cause errors in determining sensitivity and specificity during the bioinformatics analysis process. Based on this, after removing the VVR error that detects the reference base in the VV pair, the overall false positive error rate of each T-NGS was calculated, and the result was as shown in Figure 6.

[0156] The results of the overall false positive error rate revealed several other findings. Specifically, we observed that the false positive error rate analyzed from a company's raw data production varied within a batch, and also showed a batch effect, where results could differ between batches.

[0157] The red-filled boxes in CC-dragen and CC-inhouse in Fig. 6 (boxes marked A in the figure) are CH10 samples, which were sent to CC company at a different time from the other samples and contained a 90:10 mixture of control genomic DNA and H. mole DNA. Compared to the samples sent at the same time, the false positive error rate was 100-1000 times higher for only the CH10 sample. This suggests that a specific error occurred for this sample during the raw data generation process. These results suggest that some of the samples sent in a batch may have a particularly high false positive error rate compared to the results of other samples, and using the results of these samples in the final analysis may lead to significant errors in the research conclusions.

[0158] Also, the two samples marked with green filled boxes (boxes marked B in the figure) in CC-dragen in Fig. 6 are the results of samples that were independently sent for sensitivity testing and produced raw data in different batches from other samples analyzed by CC company. The results of variant analysis with DRAGEN (CC-dragen false positive error rate in Fig. 6) showed that the false positive error rate was about 100 times higher than the results of analyzing other samples. This result suggests that the false positive error rate shows a batch effect when NGS samples are sent at different times, which means that the results from one company do not always produce NGS results of consistent quality. This means that, unlike the error in a single sample, the false positive error rate of all samples sent at once can be significantly different from the expected value, which means that it is necessary to measure the false positive error rate when sending samples to be analyzed to examine the batch effect.

[0159] The above results mean that in order to ensure the reliability of NGS raw data, reference standard materials that can test sensitivity and specificity must be sent simultaneously with clinical or research samples to produce NGS raw data, and only when the sensitivity and specificity of the test results of these standard materials meet a certain quality can the NGS results of samples sent simultaneously in batches be trusted.

[0160]

[0161] Example 4. Comparative analysis of expected VAF and actual VAF according to dilution ratio

[0162] We analyzed the correlation between the expected VAF and the actual VAF according to the dilution ratio in the Ho-N pair and He-N pair alleles used to measure sensitivity. As shown in Figure 7, two out of three companies showed a linear correlation in the results of calling variants in-house and with DRAGEN.

[0163] However, in the case of Company AA, while the correlation was linear when analyzed with DRAGEN, the correlation when analyzed with the in-house method was not linear, as shown in Figure 7F. In particular, cases where the expected VAF was low were rarely observed in the actual VAF, and the actual VAF values ​​appeared to be higher than the expected VAF values. This result suggests that Company AA's bioinformatics method encountered a problem in detecting variants with low VAF values, and it appears that Company AA's bioinformatics method needs improvement.

[0164]

[0165] Example 5. Analysis of variant read counts by chromosomal location

[0166] We requested four NGS service providers to analyze with a read depth of approximately 1000x, and when the actual read count of each variant is displayed by chromosomal region, the read count of each variant in the NGS results of three companies was less than 1000x, and there was a large difference in the read count by chromosomal location in the results of each company (Fig. 8). In the case of company CC, the read count was higher than that of companies AA and BB, and the difference by chromosome was also small. However, in the case of company AA, the total read count was significantly lower, and the deviation in the read count by location was also large. In the case of company BB, the median read count was lower than that of company CC, but the deviation in the read count was much larger. The median read depth of the DD company was the only one that exceeded 1000x, and the read count for each chromosome was also generally constant, but the deviation in the read count was relatively large compared to the CC company.

[0167] If we analyze the results of the read count by chromosome region and the results of each company's sensitivity, we can see that although the NGS results of CC company had a larger read count than those of AA and BB companies, the deviation by region was small, so it can be said that it has conditions where the FN error for mutations in the target gene can be low. In addition, it is possible that this result is related to the fact that CC company has a relatively high sensitivity in the sensitivity results of each company. Therefore, analyzing the number of variant reads by chromosome region can be one way to analyze the reliability of NGS results.

[0168]

[0169] Example 6. Selection of chromosomal regions with high error rates and improvement of NGS analysis methods.

[0170] To determine whether false positive errors are concentrated in a specific region by displaying the AF (allele fraction) and location of false positive bases in a graph, the distribution of alleles that produce false positive errors in the CC company NGS results was analyzed (Fig. 9).

[0171] As a result, it shows a pattern that is evenly distributed throughout the entire chromosome region rather than appearing in a specific part. A is a false positive error in the entire mixed sample, B is a false positive error in three mixed samples, 99.5:0.5 (CH0.5), 5:95 (CH95), 0.5:99.5 (CH99.5), C is a false positive error in the remaining mixed samples excluding the three mixed samples above, and D is a false positive error. Alleles were divided into three types, and the left side shows the distribution of alleles with 3 or fewer errors, the middle side shows the chromosomal distribution of alleles with 4-6 errors, and the right side shows the distribution of alleles with 7 or more errors.

[0172] As shown in Figure 9A, there are alleles where errors are concentrated, and there were sites where more than 4 false positive errors appeared in the analysis results of 10 mixed samples, and in particular, when 7 or more false positive errors appeared, the size of the AF of the false positive base was found to be relatively large. By excluding alleles where such false positive errors appear at a high frequency from the final mutation analysis results, it can be a way to prevent the reporting of false positive errors, and by reporting such sites to researchers or patients in advance, it will be possible to prevent errors arising in the interpretation of the research results.

[0173] Also, as shown in Fig. 9B, in the three mixed samples where the control genomic DNA and H. mole DNA were mixed in ratios of 99.5:0.5 (CH0.5), 5:95 (CH95), and 0.5:99.5 (CH99.5), it can be confirmed that the false positive errors of alleles with a high frequency of false positive errors were not high, unlike other mixed samples. Considering that the reference standard material was constant in this study and the mixed samples were simply two samples mixed in different ratios, the cause of this batch effect can be said to occur during the production of the NSG raw data. In other words, it can be thought that it is not a difference in the sample DNA or an error in the bioinformatics analysis, but an error that occurred due to a difference in the experimental materials used to produce the NGS raw data, a slight difference in the sample concentration during the experimental process using them, or an error that occurred due to the use of a different PCR device. Therefore, if each laboratory selects conditions that minimize error rates at these allele sites during its experimental process, overall NGS errors are likely to be reduced. Furthermore, alleles with particularly high error rates can be used as indicators for reducing errors in specific NGS kits.

[0174]

[0175] Example 7. Sensitivity and specificity analysis of a next-generation sequencing (NGS) panel using formalin-fixed paraffin-embedded (FFPE) cell blocks from mixed cell samples.

[0176] Typically, the reference standard material (RSM) is provided in the form of formalin-fixed paraffin-embedded (FFPE) tissues using cancer tissue or fresh cancer cells, which are then used to measure the sensitivity and specificity of T-NGS for mutation detection. This is because, compared to using DNA isolated from fresh cancer tissue or fresh cancer cells as the RSM, there are many false-negative and false-positive errors in the measured data when formalin-fixed paraffin-embedded (FFPE) cancer tissue or cancer cells are provided as the RSM and DNA isolated from them is used.

[0177] Here, "fresh cells" or "fresh tissue" refers to living cells obtained from culture or tissues directly isolated from a living body. It may also include frozen tissues or cells that have not been chemically treated. Currently, for mutation analysis in cancer tissue using the NGS method, FFPE tissue slides made from tissues containing at least 50%, preferably at least 75%, cancer tissue are used as reference standard material.

[0178] Therefore, in order to provide a composition for NGS panel verification for utilizing the verification method of the present invention in the form of a reference standard material, two types of cells, such as H. mole cells and control cells, were mixed, and an FFPE cell block was produced, and then a cell block slide was made to provide a reference standard material, or DNA was isolated from the cell block slide to provide a reference standard material. That is, when providing a reference standard material, it can be provided in the form of a paraffin-embedded tissue slice on a slide, or DNA extracted from a paraffin-embedded tissue slice can be provided (Fig. 10).

[0179] At this time, the two types of cells can be composed of homozygous cells like H. mole cells and cells with heterozygous alleles like normal blood cells, or they can both be composed of H. mole cells composed of different variants.

[0180] When using DNA from FFPE cell blocks of mixed cell samples, 1) when creating mixed cell samples using homozygous cells and control cells, the sensitivity and specificity of NGS are determined in the same way as when using mixed DNA samples. However, for NGS analysis of homozygous cells and control cells, either fresh cell DNA or FFPE cell block DNA can be used. However, since NGS results using FFPE cell block DNA have lower sensitivity and specificity, it would be advantageous to use fresh cell DNA. Here, control cells can be normal 2n cells or cancer cells, and can be cells in which alleles are mixed as heterozygotes and homozygotes, unlike homozygous cells in which all alleles are composed of homozygotes.

[0181] 2) When determining the sensitivity of NGS by making a mixed cell sample FFPE cell block using two types of homozygous cells with different mutations, since there are no heterozygous alleles present in the control DNA (He-N pair, Ho-He-N pair rarely exist) and homozygous alleles exist, the sensitivity is analyzed using only the Ho-N pair. When determining the specificity of NGS by mixing two homozygous cells, the specificity can be determined by analyzing the RR pair, RV pair, and VV pair in the same way as in the analysis of the mixture of homozygous and control DNA.

[0182] 3) When using two types of homozygous cells with different mutations and control cells as a mixture or using their DNA, the expected VAF can be determined according to Ho-N pair, He-N pair, and Ho-He-N-pair, and the sensitivity can be determined in the same way as in the analysis of the mixture of homozygous and control DNA. In addition, when determining specificity, the specificity can be determined in the same way as in the analysis of the mixture of homozygous and control DNA by classifying them into RR pair, RV pair, VV pair, etc.

[0183]

[0184] From the above description, those skilled in the art will understand that the present invention can be implemented in other specific forms without altering its technical spirit or essential characteristics. In this regard, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. The scope of the present invention should be interpreted as encompassing all changes or modifications derived from the meaning and scope of the following claims and their equivalent concepts, rather than the detailed description above.

Claims

1. (a) A step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof; (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results according to step (a); (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and (d) A method for verifying a next generation sequencing (NGS) panel through analysis of the limit of detection, comprising a step of analyzing the limit of detection.

2. In paragraph 1, A method for verifying a next-generation sequencing (NGS) panel through analysis of the detection limit, wherein the DNA mixture sample, which is a mixture of two or more cell-derived DNAs selected from the homozygous DNA, control genomic DNA, or a combination thereof, is a mixture of two or more homozygous DNAs having different mutations; a mixture of homozygous DNA and control genomic DNA; or a combination thereof.

3. In paragraph 1, A method for verifying a next-generation sequencing (NGS) panel through analysis of the detection limit, wherein the homozygous DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic cell line.

4. In paragraph 1, A method for verifying a next-generation sequencing (NGS) panel through analysis of the detection limit, wherein the control genomic DNA is homozygous DNA; heterozygous DNA; or genomic DNA isolated from normal human blood cells or tissue cells, cultured blood cells, tissue cells, or cancer cells, or blood cells or cancer tissue cells of a patient.

5. In paragraph 1, A method for validating a next-generation sequencing (NGS) panel through analysis of the limit of detection, wherein the homozygous DNA or control genomic DNA is isolated from fresh cells, fresh tissues, or formalin-fixed paraffin embedded (FFPE) cell tissues. 6.(a) A step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof; (b) a step of selecting false positive mutations from the next generation sequencing (NGS) results according to step (a); and (c) A method for verifying a next-generation sequencing (NGS) panel through analysis of a false positive error rate, comprising a step of analyzing a false positive error rate in the DNA mixture sample based on the selected false positive mutations.

7. In paragraph 6, A method for verifying a next-generation sequencing (NGS) panel through analysis of false positive error rate, wherein the DNA mixture sample, which is a mixture of two or more cell-derived DNAs selected from the homozygous DNA, control genomic DNA, or a combination thereof, is a mixture of two or more homozygous DNAs having different mutations; a mixture of homozygous DNA and control genomic DNA; or a combination thereof.

8. In paragraph 6, The above false positive mutation is a case where a mutation is found to exist in the result of analyzing a DNA mixture sample by NGS among null alleles in which no mutation is found in two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof, and a verification method of a next-generation sequencing (NGS) panel through analysis of a false positive error rate.

9. In paragraph 6, The above false positive mutation is a case where a mutation or reference base different from the VV pair is found in the result of analyzing a DNA mixture sample by NGS among alleles (VV pairs) in which only mutations are found and no reference base is found in two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof. A verification method of a next-generation sequencing (NGS) panel through analysis of a false positive error rate.

10. In paragraph 6, The above false positive mutation is a case where a new mutation other than a base constituting an RV pair is discovered among alleles (RV pair alleles) in which a reference base and a variant are mixed in two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof, when analyzing a DNA mixture sample by NGS, a verification method of a next-generation sequencing (NGS) panel through analysis of a false positive error rate.

11. In paragraph 9, The above step (c) is a verification method of a next-generation sequencing (NGS) panel through analysis of a false positive error rate, which additionally includes a process of removing VVR false positive error results in which a reference base is found in a VV pair.

12. In paragraph 6, A method for verifying a next-generation sequencing (NGS) panel through analysis of a false positive error rate, which additionally includes a process of excluding error results of alleles in which false positive errors occur at a high frequency in the above step (c).

13. In paragraph 6, A method for verifying a next-generation sequencing (NGS) panel through analysis of false positive error rate, further comprising a process of examining batch effect in step (c) above. 14.(a) A step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof; (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results according to step (a); (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and (d) A method for verifying a next-generation sequencing (NGS) panel, comprising a step of comparing and analyzing the expected VAF and the actual VAF according to the above dilution ratio.

15. In paragraph 14, A method for verifying a next-generation sequencing (NGS) panel, wherein the DNA mixture sample, which is a mixture of two or more cell-derived DNAs selected from the homozygous DNA, control genomic DNA, or a combination thereof, is a mixture of two or more homozygous DNAs having different mutations; a mixture of homozygous DNA and control genomic DNA; or a combination thereof.

16. In paragraph 14, A method for validating a next-generation sequencing (NGS) panel, wherein the homozygous DNA or control genomic DNA is isolated from fresh cells, fresh tissues, or formalin-fixed paraffin embedded (FFPE) cell tissues. 17.(a) A step of performing next generation sequencing (NGS) using a next generation sequencing (NGS) panel on a DNA mixture sample containing two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof; (b) a step of selecting a Ho-N pair allele (Homozygote-Null pair allele) and / or a He-N pair allele (Heterozygote-Null pair allele) from the next generation sequencing (NGS) results obtained by step (a); (c) a step of analyzing the dilution mutation detection rate, which represents the number of diluted mutations found among the Ho-N pair allele and He-N pair allele in the DNA mixture sample based on the selected Ho-N pair allele and / or He-N pair allele; and (d) A method for verifying a next-generation sequencing (NGS) panel, comprising a step of analyzing the deviation of the number of variant reads (read count) per chromosomal region used for detecting dilution mutations in step (c).

18. In paragraph 17, A method for verifying a next-generation sequencing (NGS) panel, wherein the DNA mixture sample, which is a mixture of two or more cell-derived DNAs selected from the homozygous DNA, control genomic DNA, or a combination thereof, is a mixture of two or more homozygous DNAs having different mutations; a mixture of homozygous DNA and control genomic DNA; or a combination thereof.

19. In paragraph 17, A method for validating a next-generation sequencing (NGS) panel, wherein the homozygous DNA or control genomic DNA is isolated from fresh cells, fresh tissues, or formalin-fixed paraffin embedded (FFPE) cell tissues.

20. A composition for validation of a next generation sequencing (NGS) panel, comprising a DNA mixture comprising two or more cell-derived DNAs selected from homozygous DNA, control genomic DNA, or a combination thereof.

21. In paragraph 20, A composition wherein the DNA mixture comprising two or more cell-derived DNAs selected from the homozygous DNA, the control genomic DNA, or a combination thereof is a mixture of two or more homozygous DNAs having different mutations; a mixture of the homozygous DNA and the control genomic DNA; or a combination thereof.

22. In paragraph 20, A composition wherein the homozygous DNA is DNA isolated from a hydatidiform mole cell line or a parthenogenetic cell line.

23. In paragraph 20, A composition wherein the control genomic DNA is homozygous DNA; heterozygous DNA; or genomic DNA isolated from a normal person's blood cell or tissue cell, cultured blood cell, tissue cell, or cancer cell, or a patient's blood cell or cancer tissue cell.

24. A composition for validation of a next-generation sequencing (NGS) panel, comprising a mixture of two or more cells selected from homozygous cells, control cells, or a combination thereof.

25. In paragraph 24, A composition wherein the mixture of two or more cells selected from the homozygous cells, control cells, or a combination thereof is a mixture of two or more homozygous cells having different mutations; a mixture of two or more homozygous cells having different mutations and control cells; a mixture of homozygous cells and control cells; or a combination thereof.

26. In paragraph 24, A composition wherein the homozygous cells or control cells are fresh cells or formalin fixed paraffin embedded (FFPE) cells.

27. In paragraph 24, The composition comprises a mixture of two or more cells selected from homozygous cells, control cells, or a combination thereof, in the form of a formalin fixed paraffin embedded (FFPE) cell block.

28. In paragraph 24, A composition wherein the homozygous cell is a hydatidiform mole cell or a parthenogenetic cell.

29. In paragraph 24, A composition wherein the control cells are homozygous cells; heterozygous cells; blood cells or tissue cells of a normal person; cultured blood cells, tissue cells or cancer cells; or blood cells or cancer tissue cells of a patient.

Citation Information

Patent Citations

  • A method for measuring the chromosome, gene or nucleotide sequence copy number using SNP array

    KR1020090097792A

  • Semiconductor package

    KR1020250044565A

  • Method to confirm variants in NGS panel testing by SNP genotyping

    US20210180128A1

  • Systems and methods for contamination detection in next generation sequencing samples

    US20220392572A1