Methods for detecting cancer DNA in a sample
By enriching and combining measurements from multiple genetic variations in a test sample, the method addresses the sensitivity and false-positive issues of current MRD detection, enabling early and accurate identification of cancer recurrence.
Patent Information
- Application Number
- JP2025508497
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-19
- Filing Date
- 2023-08-17
- Publication Date
- 2025-08-15
AI Technical Summary
Current methods for detecting minimal residual disease (MRD) in cancer patients are limited by low sensitivity due to low tumor DNA levels and high false-positive rates, particularly in solid tumors, making it difficult to accurately predict cancer recurrence.
A method that enriches a test sample for multiple target regions with different genetic variations, including single nucleotide variants, multi-nucleotide variants, copy number variants, structural variants, and phase variants, and combines measurements from these regions using an error model to identify cancer DNA.
Enhances the detection of cancer DNA by combining evidence from multiple target regions, improving sensitivity and reducing false positives, allowing for early detection of tumor recurrence.
Smart Images

Figure 2025526854000013 
Figure 2025526854000014 
Figure 2025526854000015
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of UK Patent Application No. GB2212094.3, filed August 19, 2022, which is incorporated herein by reference for all purposes.
[0002] Field The present disclosure relates generally to the field of liquid biopsies, such as diagnosing the presence of cancer in blood or other liquid samples obtained from a patient. [Background technology]
[0003] Detection and monitoring of circulating tumor DNA (ctDNA) is rapidly becoming a diagnostic, prognostic, and predictive tool in the care of cancer patients. Following cancer treatment, a small number of cancer cells often remain in the body, even in patients who appear to be in remission. These remaining cells are often referred to as "minimal residual disease" (MRD) or residual disease. These residual cells ultimately cause recurrence in many cancers. It is important to determine the likelihood of a patient experiencing disease recurrence and relapse after initial treatment, so that those most likely to require additional treatment can receive that treatment. Meanwhile, patients who do not require additional treatment can be spared, thereby reducing patient harm and treatment costs. Thus, an effective method for detecting minimal residual disease is highly desirable. It is also important to provide a highly sensitive method for detecting the risk of cancer recurrence more quickly than current methods (e.g., those typically performed by imaging or clinical analysis).
[0004] MRD has been successfully detected in some hematological malignancies because relatively large amounts of DNA can be analyzed and the frequency of common tumor-specific fusion genes can be measured in a straightforward manner. There is now strong evidence that MRD can be detected in many solid tumors by assessing cell-free DNA (cfDNA) in relation to circulating tumor DNA (ctDNA). However, a challenge with detecting minimal residual disease in cfDNA is the insufficient sensitivity of many of the tests used to detect sequence variations in samples. Most molecular tests today are performed by sequencing cfDNA for panels of known genes. A challenge with detecting minimal residual disease by sequencing cfDNA is that the amount of tumor DNA in cell-free DNA is often well below the detection limit of such methods. Specifically, the frequency with which individual tumor sequence variations are expected to occur in the cfDNA of patients with minimal residual disease is typically well below the frequency with which sequencing artifacts caused by PCR errors, base miscalls, and / or DNA damage would occur. This challenge is compounded by the fact that in some cases, tumor DNA levels may be low, resulting in less than one copy of each mutation being evaluated in the cfDNA sample being analyzed on average.In addition, the relatively small amount of mutant DNA derived from leukocytes dissolved in the bloodstream may lead to erroneous results.Therefore, the detection of minimal residual disease by sequencing-based approaches remains difficult.
[0005] Assays for detecting MRD can utilize a variety of approaches, including sequencing patient tumor tissue to identify tumor-specific genetic variants (mutants). These variants can include single-nucleotide variants, small insertions and deletions, doublet base substitutions, and larger structural changes. Identifying these tumor-specific variants in a patient's cfDNA sample is theoretically indicative of MRD. However, as the number of tumor-specific variants in an assay increases, the likelihood of a false-positive result may increase. Furthermore, different types of variants increase or decrease the likelihood of a false-positive result. Therefore, there is a need for improved detection of ctDNA. Summary of the Invention
[0006] Disclosed herein is a method for detecting cancer DNA in a test sample of DNA obtained from a patient. In some embodiments, the method includes enriching the test sample for multiple target regions, the multiple target regions including a first target region having a first class and a second target region having a second class. The multiple target regions are measured, and for each of the first target region and the second target region, measurements supporting the target region class are compared with an error model that models the probability that the target region class is observed in DNA that does not have the target region class. These comparisons can then be combined for at least the first target region and the second target region. Cancer DNA is then identified in the test sample based on the combined comparisons.
[0007] The methods described herein, in one example, recognize that the problem of low sensitivity in determining the presence of cancer DNA in a patient's test sample can be solved by combining evidence from multiple target regions with different types and numbers of genetic variations. Measurements of each target region provide several pieces of evidence, which can be combined to support a reliable conclusion that the test sample contains cancer DNA and therefore the patient has cancer or residual disease. Furthermore, the methods described herein can combine evidence from multiple target region classes, such as single nucleotide variations, multi-nucleotide variants (such as doublet and triplet base substitutions), short insertions or deletions, copy number variants, structural variants (SVs), polygenic variants, and multiphase variants (i.e., a target region has two or more variants, all located on the same chromosome within the target region). Each of these classes can provide different levels of support or confidence for the presence (or absence) of cancer, and must be combined in a rational manner, as further described herein.
[0008] These and other advantages may become apparent in light of the following discussion.
[0009] Those skilled in the art will understand that the drawings, described below, are for illustrative purposes only and are not intended to limit the scope of the present teachings in any way. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a flow chart illustrating one embodiment of a method for detecting cancer DNA in a test sample of DNA obtained from a patient. [Figure 2A] FIG. 1 shows one embodiment of the highly sensitive tagged amplicon sequencing (eTAm-Seq™) approach, in which target regions are amplified by polymerase chain reaction (PCR). [Figure 2B] FIG. 1 illustrates some typical classes of target regions. [Figure 3] FIG. 1 shows an exemplary assay of a target region with genetic variation, according to one embodiment of the present disclosure. [Figure 4] 4A shows an example of an error probability distribution according to one embodiment of the present disclosure. In the model shown in FIG. 4A, data corresponding to infrequent high-signal events are shaded. Two models are shown in FIG. 4B, one for background noise and one for DNA damage. "VAF" refers to variant allele fraction. These models are obtained from DNA that does not have genetic variations, and these models show the probability of different variant allele fractions (or the number of variant reads in total reads) in this non-cancerous DNA. [Figure 5] FIG. 1 is a block diagram of an exemplary computer system that may be used in implementing some embodiments of the techniques described herein. [Figure 6] FIG. 1 shows one embodiment of the assay described herein and illustrates some of the difficulties of detecting cancer DNA by methods that score individual target regions as to whether they harbor a particular genetic variant. [Figure 7] 1 shows a schematic representation of some of the principles of an embodiment of the method; [Figure 8] FIG. 1 illustrates how cancer DNA fraction is calculated by comparing actual dilution data with a mathematical model. [Figure 9] Diagram adapted from Kurtz et al. (Nat Biotechnol 2021, 39:1-11), where the authors show that the genome has a small number of phase variants (part b). [Figure 10]Figure adapted from Li et al. (Nature 2020 578, 112-121) showing the range of structural variants and types in different types of cancer. Some types of cancer, such as breast cancer and certain sarcomas, often have a large number of structural variants, while other cancers, such as CLL, typically have a smaller number. DETAILED DESCRIPTION OF THE INVENTION
[0011] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. However, certain elements are defined for clarity and ease of reference.
[0012] The terms and symbols of nucleic acid chemistry, biochemistry, genetics, and molecular biology used herein follow standard treatises and texts in the field, such as Kornberg and Baker, DNA Replication, 2nd Edition (WH Freeman, New York, 1992); Lehninger, Biochemistry, 8th Edition (Worth Publishers, New York, 2021); Strachan and Read, Human Molecular Genetics, 5th Edition (Wiley-Liss, New York, 2018); Eckstein, ed., Oligonucleotides and Analogs: A Practical Approach (Oxford University Press, New York, 1992); Gait, ed., Oligonucleotide Synthesis: A Practical Approach (IRL Press, Oxford, 1984), the contents of which are incorporated by reference in their entireties, and the like.
[0013] As used herein, depending on the context, the term "calling" can mean indicating whether a particular genetic variation is present in a sequence, whether a sample has a genetic variation, or whether a sample contains cancer DNA.
[0014] If two nucleic acids are " complementary ", they will hybridize with each other under high stringency conditions.The term " completely complementary " refers to a double strand in which each base of one nucleic acid is base-paired with the complementary nucleotide of the other nucleic acid.In many cases, two complementary sequences have at least 10, for example, at least 12 or 15 complementary nucleotides.
[0015] As used herein, the term " detecting recurrence " refers to detecting tumor recurrence through the identification of cancer DNA.In this context, the term " early detection " refers to the detection of mutant DNA before cancer recurrence can be reliably detected through conventional standard treatment / surveillance monitoring methods, such as radiological imaging.This can be carried out, for example, by monitoring the presence of ctDNA in cfDNA in blood samples that are collected continuously at multiple time points, as described below.
[0016] The terms "determining," "measuring," "evaluating," "assessing," and "assaying" are used interchangeably and include quantitative and qualitative determinations. Assessments can be relative or absolute.
[0017] The term "genetic variation," as used herein, refers to a variation (e.g., a nucleotide substitution, indel, or rearrangement) that is present or likely to be present in a test sample. Genetic variations can be from any cause. For example, genetic variations can be caused by mutations (e.g., somatic mutations), or genetic variations can be germline, e.g., germline mutations that become incorporated into the DNA of all cells in the body. When a sequence variation is called as a genetic variation, the call indicates that the sample is likely to have the variation, but the "call" may be inaccurate. In many cases, the term "genetic variation" can be replaced with the term "mutation." For example, when the method is used to detect sequence variations associated with cancer or other diseases caused by mutations, the term "genetic variation" can be replaced with the term "mutation."
[0018] As used herein, the term "minimal residual disease" (MRD) refers to the presence of cancer cells after treatment with curative intent. MRD is also referred to in some publications as "molecular residual disease" or "residual disease."
[0019] The terms "nucleic acid," "oligonucleotide," and "polynucleotide" refer to any nucleic acid composed of nucleotides, e.g., deoxyribonucleotides or ribonucleotides, of any length, e.g., greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1,000 bases, greater than 10,000 bases, greater than 100,000 bases, greater than about 1 million bases, up to about 10 10 It is used interchangeably herein to describe a base or a larger polymer.
[0020] The terms "plurarity," "population," and "collection" are used interchangeably to refer to something that has at least two members. In certain cases, a plurality, population, or collection may have at least 5, at least 10, at least 100, at least 1,000, at least 10,000, at least 100,000, at least 100,000, or at least 10 6 , at least 10 7 , at least 10 8 , or at least 10 9 , or more members.
[0021] The term "reference sequence" as used herein refers to the reference sequence of a reference genome or a sequence derived from a patient sample that is expected to have no somatic variants, such as a buccal swab.The reference sequence corresponds to a sequence that has or is suspected to have a sequence variation (e.g., a target sequence).Therefore, the presence (or absence) of a sequence variation can be determined by comparing the sequence that has or is suspected to have a sequence variation (e.g., a target sequence) with the reference sequence.Since the reference sequence and the sequence that has or is suspected to have a sequence variation (e.g., a target sequence) originate from the same genome location, the reference sequence differs from the sequence that has or is suspected to have a sequence variation (e.g., a target sequence) only in the sequence variation itself.
[0022] The term "reference genome" as used herein can refer to a single genome, a collection of genomes, or a consensus genome. A reference genome can be obtained from one or more public databases. A reference genome is used to determine the location of the sequence to be analyzed in the genome of an organism. As will be recognized by those skilled in the art, a consensus genome is a genome constructed from multiple genomes of the same species.
[0023] The term "sequence variation" as used herein refers to a variant that differs from an expected sequence or reference sequence, for example, a reference genome or reference sequence derived from a patient sample that is expected to have no somatic variants, such as a buccal swab. Sequence variation can refer to a combination of the position and type of sequence modification. For example, sequence variation is represented by the position of the variation and the type of substitution (e.g., G to A, G to T, G to C, A to G, etc., or insertion / deletion of G, A, T, or C, etc.) present at that position. Sequence variation can be a rearrangement due to substitution, deletion, or insertion of one or more nucleotides. In the context of this method, sequence variation can be caused by, for example, PCR error, sequencing error, or genetic variation. In many cases, sequence variation is a variation that occurs at a frequency of less than 50% relative to other molecules in a sample. Many sequence variations, such as indels and nucleotide substitutions, are substantially identical to molecules that do not have sequence variation. In some cases, a particular sequence variation may be present in a sample at a frequency of less than 20%, less than 10%, less than 5%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.001%, or less than 0.0001%.
[0024] The term "substantially" refers to sequences that are nearly identical as measured by similarity functions, including but not limited to Hamming distance, Levenshtein distance, Jaccard distance, cosine distance, etc. (See Kemena et al., Bioinformatics 2009 25:2455-65, the entire contents of which are incorporated herein by reference). The exact threshold depends on the sample preparation and sequencing error rate used to perform the analysis; the higher the error rate, the lower the similarity threshold required. In certain cases, substantially identical sequences have at least 98% or at least 99% sequence identity.
[0025] As used herein, the term "threshold" refers to the level of evidence (e.g., ratio or set amount) required to make a call.
[0026] As used herein, the term "value" refers to a number, letter, word (e.g., "high," "medium," or "low"), or descriptor (e.g., "+++" or "++") that may indicate the strength of evidence. A value may have one component (e.g., a single number) or two or more components, depending on how the value is analyzed.
[0027] Definitions of other terms can be found throughout the specification. It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a prelude to the use of exclusive terminology such as "solely," "only," and the like, or the use of a "negative" limitation in connection with the recitation of claim elements.
[0028] Detailed Description The method described herein is based on the recognition that cancer DNA can be easily detected by incorporating and combining evidence from various target regions with one or more genetic variations, including single nucleotide variants, multi-nucleotide variants, copy number variants, short insertions and deletions, epigenetic variants, structural variants and phase variants into the assay of target regions.This method is useful in the enrichment-based method of examining target regions of genome.
[0029] FIG. 1 illustrates one embodiment of a method 100 for detecting cancer DNA in a test sample, such as a blood sample collected from a cancer patient. In this embodiment, method 100 may include (a) enriching a plurality of target regions from the test sample (step 102). The plurality of target regions may include a first target region comprising a first class and a second target region comprising a second class. Method 100 may then continue with (b) measuring the plurality of target regions in the test sample (step 104). For each of the first and second target regions, method 100 may then continue with (c) comparing measurements supporting the target region class to an error model that models the probability that the target region class is observed in DNA that does not have the target region class (step 106). Method 100 may then continue with (d) combining the comparisons for at least the first and second target regions (step 108). Method 100 may continue with (e) identifying cancer DNA in the test sample based on the combined comparison of the first target region and the second target region (step 110).
[0030] The methods disclosed herein may include obtaining a test sample from a patient. Alternatively, the test sample may be previously obtained from the patient. The test sample may include any nucleic acid sample or body fluid containing DNA, RNA, or cDNA. A genomic DNA sample obtained from a mammal (e.g., a mouse or a human) is one type of test sample. The test sample may be approximately 10 4 , 10 5 , 10 6 or 10 7 , 10 8 , 10 9 or 10 10The sample may contain more than 100 different nucleic acid molecules. Any sample containing nucleic acid, such as genomic DNA or RNA, derived from tissue culture cells, or a tissue sample can be used herein. Furthermore, while many embodiments describe the detection or use of cancer DNA, the methods according to the present disclosure are applicable to all forms of nucleic acid, and thus the use of "cancer DNA" can also refer to RNA or other detectable nucleic acids associated with cancer. Test samples can include plasma, serum, cerebrospinal fluid, urine, saliva, stool, amniotic fluid, aqueous humor, bile, breast milk, earwax, chyle, exudate, gastric juice, lymph, mucus, pericardial effusion, peritoneal fluid, pleural effusion, pus, sebum, semen, sputum, synovial fluid, sweat, tears, vomit, or whole blood.
[0031] In some embodiments, the test sample contains cell-free DNA (cfDNA), i.e., DNA that is free in bodily fluids and not contained within cells. cfDNA can be obtained by centrifuging the test sample to remove all cells and then isolating DNA from the remaining fluid (e.g., plasma or serum). Such methods are well known (see, e.g., Lo et al., Am J Hum Genet 1998;62:768-75). Circulating cell-free DNA can be double-stranded or single-stranded. The term cfDNA encompasses free DNA molecules circulating in the bloodstream and DNA molecules present in extracellular vesicles (e.g., exosomes) circulating in the bloodstream. Cell-free DNA can contain cancer DNA, i.e., DNA derived from cancerous cells. Cancer DNA derived from solid tumors can be found in cfDNA, which in this case can be described as tumor DNA (tDNA) or circulating tumor DNA (ctDNA). Cancer DNA is identifiable because it contains mutations. In a preferred embodiment, the test sample is bloodstream-derived cell-free DNA (circulating cell-free DNA), which is DNA circulating in the peripheral blood of a patient.
[0032] In some embodiments, the test sample contains cancer DNA isolated directly from a tissue biopsy, from circulating tumor cells (CTCs), or from other cells that are no longer part of tumor tissue but are not circulating, such as cells in a urine or stool sample. In some embodiments, the test sample may contain DNA isolated from cells, such as bone marrow cells, lymph node-derived cells, or circulating white blood cells in the case of blood cancer, or from other sample types such as lymph node-derived cells, tumor margin-derived cells, or cerebrospinal fluid (CSF) and whole blood that are currently screened by other means for the presence of cancer cells from solid tumors. Cells can be obtained from a tissue sample (e.g., a cancer tissue sample, or a sample suspected to be a cancer tissue sample, or a tissue sample containing or suspected to contain cancer cells) or a body fluid sample (e.g., any of the body fluids listed above) obtained from a patient.
[0033] DNA molecules in cell-free DNA can be highly fragmented and have a median size below 1 kb (e.g., 50 bp to 500 bp, 80 bp to 400 bp, or 100-1000 bp), although fragments with median sizes outside this range may also exist. Typically, cfDNA has an average fragment size of approximately 100-250 bp, e.g., 150-200 bp in length, or approximately 160 bp. ctDNA is tumor-derived and originates directly from the tumor or from circulating tumor cells (CTCs), which are viable, intact tumor cells that are released from the primary tumor and can enter the bloodstream or lymphatic system. The exact mechanism by which cancer DNA is released is unclear, but it is presumed to involve apoptosis and necrosis from dying cells or the active release of viable tumor cells. The amount of ctDNA in circulating cell-free DNA samples isolated from cancer patients varies greatly: typical samples contain less than 10% ctDNA, but many samples from patients being evaluated for MRD may have less than 0.01% ctDNA, and some samples have more than 10% ctDNA. Molecules of cancer DNA can often be identified because they harbor tumorigenic mutations.
[0034] In one embodiment, the test sample is a plasma sample, and cell-free DNA (cfDNA) is isolated from the plasma sample. The cancer DNA ratio (compared to non-cancerous DNA) in the test sample can be 0.01% or less, 0.002% or less, 0.005% or less, or 0.001% or less. In some embodiments, the detectable cancer DNA ratio in the DNA test sample can be from about 0.0001%, although the actual detection limit can vary. In some embodiments, the test sample contains less than 25,000 genome equivalents of DNA (e.g., cfDNA), for example, less than 20,000 genome equivalents, less than 10,000 genome equivalents, less than 5,000 genome equivalents, or less than 1,000 genome equivalents of DNA. In some embodiments, the test sample contains about 100 to about 25,000 genome equivalents (i.e., enrichable or amplifiable copies) of DNA. In some embodiments, the test sample contains about 10 ng to about 100 ng of DNA. In some embodiments, the test sample contains at least 10 ng, at least 20 ng, at least 30 ng, at least 40 ng, at least 50 ng, at least 60 ng, at least 70 ng, at least 80 ng, at least 90 ng, or at least 100 ng of DNA. In some embodiments, the test sample contains 66 ng of DNA.
[0035] Furthermore, the methods described herein can be used to detect cancer DNA from both solid tumors and hematological (blood) cancers. Thus, the term "cancer" can refer to any disease characterized by uncontrolled cell division, and can be, for example, a blood cancer such as leukemia, lymphoma, or multiple myeloma, or a neoplastic cancer involving an abnormal amount of tissue in which cells grow and divide more than they should or do not die when they should. Neoplastic cancers, such as lung cancer, breast cancer, or liver cancer, are associated with solid tumors. In solid tumor embodiments, the method can identify cancer DNA (here, tumor DNA) in cfDNA (e.g., circulating cfDNA). In blood cancer embodiments, the method can identify cancer DNA in DNA extracted from cells collected from bone marrow, lymph nodes, or circulating white blood cells, or in cfDNA. For example, in blood cancer embodiments, bone marrow aspirate can be collected from an AML patient (before treatment), and the variants in the patient's AML can be determined (e.g., by sequencing DNA from AML cells), and the patient can be treated. Sometime after treatment, bone marrow aspirate, cell-free DNA, or urine can be examined for evidence of these variants to determine whether the patient still has cancer. In some embodiments, the method can identify cancer DNA in tissue samples such as surgical margins or lymph nodes.
[0036] A "target region" or "region" refers to a DNA region that contains or is suspected of containing one or more genetic variations. When referring to a genome or target polynucleotide, such a region refers to a continuous subregion or segment of the genome or target polynucleotide. This term refers to any continuous portion of a genomic sequence, which may be within or associated with a gene, e.g., a coding sequence. A target region can be from one nucleotide to a segment of hundreds or thousands of nucleotides or more in length. Typically, the length of a target region is about the average length of the nucleic acids present in the test sample, or less. For example, in cfDNA embodiments, the target region is typically about 160 bp. However, in other embodiments, the length of a target region can be about 50 bp, 100 bp, about 200 bp, about 300 bp, about 400 bp, and about 500 bp. For example, in some embodiments where the test sample is a tissue sample (e.g., derived from a lymph node or surgical margin), the length of the target region can be equivalent to the desired sequencing length or average fragment length. In practice, the target region can be any region targeted by a PCR primer pair, and its length is therefore the length of the resulting amplicon.
[0037] Enrichment of multiple target regions (step 102) can be accomplished by a variety of methods, including, but not limited to, hybridization to nucleic acid probes, polymerase chain reaction (PCR), ligation target capture, molecular inversion probes, ligation, and ATOM-Seq. In some embodiments, enrichment involves capturing multiple target regions from a test sample by contacting the test sample with a pool of oligonucleotides. For example, the oligonucleotide pool can have oligonucleotides that contain reverse complements (or substantially reverse complements) of multiple target regions. When the test sample is (for example) heated and the nucleic acids are denatured into single strands, the oligonucleotides can bind to any target region and can be selected (e.g., by a probe).
[0038] In some embodiments, enrichment involves amplifying multiple target regions by polymerase chain reaction (PCR), an enzymatic reaction that uses one or more sequence-specific primer pairs to amplify specific template DNA. As shown in FIG. 2, forward primer 202a and reverse primer 204a can be designed to contain sequences complementary to the beginning and end of target region 206a. The forward and reverse primers are then added to a test sample containing cancer DNA and subjected to PCR conditions, including one or more rounds of thermal cycling appropriate for denaturation, renaturation, and extension with appropriate reagents (e.g., nucleotides, buffers, polymerase, etc.) known in the art, to generate multiple PCR products, e.g., amplicons 208a. The term "amplicon" as used herein refers to the product (or "band") amplified by a specific primer pair in a PCR reaction. The amplicons 208a are then sequenced, and the number of reads with sequence variation 210a can then be counted for that target region 206a.
[0039] PCR can be multiplex PCR, which utilizes two or more primer pairs for different targets. When two or more targets are present in a reaction, multiplex PCR generates two or more amplified DNA products, which are co-amplified in one reaction using a corresponding number of sequence-specific primer pairs. As shown in Figure 2, multiplex PCR can include three forward primer pairs 202a, 202b, 202c and a reverse primer pair 204a, 204, 204c, each individually designed for multiple target regions 206a, 206b, 206c, generating amplicons 208a, 208b, 208c, which can then be sequenced. Observation of sequence reads with sequence variations 210a, 210b, 210c provides support that the observed sequence variations are true genetic variations present in the sample, indicating the presence of cancer DNA in the test sample.
[0040] In some embodiments, the test sample can first be pre-amplified, for example, by whole genome amplification. Pre-amplification can be performed, for example, by ligating adapters and performing PCR targeting the ligated adapters. In these embodiments, sequencing adapters can be added during amplification or ligated after amplification. In other embodiments, the target region can be enriched using a "target enrichment-based" approach, in which adapters are ligated to the test sample, fragments containing the target region are enriched by hybridization to nucleic acid probes, and then amplified using primers that hybridize to the adapters. In such embodiments, during any ligation reaction, adapters with multiple barcodes can be ligated to the DNA, effectively separating the molecular population into individual barcode populations or replicates. In this way, the sequence of the target region can be enriched from the sample by PCR or by hybridization to nucleic acid probes. Other enrichment methods can also be used. In other embodiments, any other method involving physical replication or the use of molecular barcodes, such as molecular inversion probes (MIPs) or anchored multiplex PCR (AMP), can also be utilized. In some embodiments, the target region can be enriched during the targeting step using methods including COLD-PCR, allele-specific PCR, digestion of wild-type sequences utilizing adjacent germline variations, or other methods known to those skilled in the art. In a preferred embodiment, the pre-amplification step is performed using multiplex PCR, and the sample is then divided into two or more samples for further PCR analysis (either singleplex or multiplex). In this embodiment, the samples may be pooled and an additional barcoding step may be performed to enable sequencing of the amplicons.
[0041] While the remainder of this disclosure describes in detail the use of PCR and "amplicon" sequencing, embodiments of the present disclosure may also be applied to other methods, including methods that use pre-amplified samples or molecular barcodes or molecular beacons, e.g., random sequences, that are added to nucleic acids prior to amplification. In such embodiments, comparing measurements supporting the presence of a target region class to one or more error models may include estimating the probability that a sequence variation is present within the target region by (for example) measuring or counting the number of index sequences for that target region.
[0042] Target region "class" can refer to target regions that have one or more types of genetic variations.For example, target region class can include target regions that have: single nucleotide variants (SNVs), such as A>T or C>G single base changes; multi-nucleotide variants (MNVs), such as CA>TG doublet base substitution or AAA>TTT triplet base substitution; short insertion or deletion (INDEL) of one or more nucleotides, such as TTTT insertion or CG deletion; copy number variants (CNVs), including gene amplification, chromosomal aneuploidy, or tandem repeats, which can often be detected as target regions with significantly increased sequence coverage; structural variants (SVs), which reflect relatively large genetic changes, such as gene fusions or large insertions or deletions of, for example, 1000s, 10000s, 100,000s, or 1 million nucleotides; and epigenetic variants (EVs), such as changes in DNA methylation, DNA-protein modifications, chromatin accessibility, histone modifications, and the like. In addition to the type of change, the target region class may also refer to specific changes. For example, a target region class may include an A to T SNV change at a particular position or in a particular sequence context (e.g., a trinucleotide context, i.e., a particular nucleotide immediately adjacent to the genetic variation, or a pentanucleotide context, i.e., two adjacent bases on either side of the change), an AAAA to AA INDEL change, an A to T SNV change at a first position and a C to G SNV change at a second position, and the like.
[0043] A target region class may also include multiple (i.e., two or more) genetic variations. The two or more genetic variations may be of the same type (e.g., two or more SNVs, INDELs, SVs, and EVs) or two or more different types (e.g., one SNV and one INDEL; one SNV, one INDEL, and one EV; etc.). The two or more genetic variations may be separated by at least one nucleotide. Two or more genetic variations present in the same DNA molecule may be described as phase variants (PVs). The term "phase" refers to determining whether the genetic variations are located in either the maternal or paternal copy of the chromosome, e.g., chromosome 1. Two or more genetic variations are considered to be PVs relative to each other if they are both present in the same chromosome (i.e., the maternal or paternal copy), and therefore present in the same DNA molecule in the test sample. If two PVs are sufficiently close together (e.g., within the same target region), they can be amplified and sequenced together and therefore can be seen in the same sequence read. As further illustrated in Figure 2A, amplicon 208c resulting from target region 206c may contain two sequence variations 210c, 210d (each individually indicated by an "X") present in the same amplicon 208c, while amplicons 208a, 208b have only one sequence variation 210a, 210b. Thus, the two sequence variations 210c, 210d (if true genetic variations) are PVs. While target region classes with PVs can include any combination of phase genetic variations, in the context of cfDNA, target region classes with PVs often include two or more SNVs present in the same DNA molecule.
[0044] In some embodiments, the genetic variations are somatic variations, i.e., they are non-germline genetic variations that may be associated with diseases such as cancer. In some embodiments, the genetic variations may include germline genetic variations, i.e., genetic variations that constitute the patient's (non-tumor) genome. Germline genetic variations may be useful in target region classes that have two or more genetic variations. For example, target regions that have both germline SNVs and somatic tumor SNVs can be enriched and sequenced. Preferably, the germline SNVs and tumor SNVs are phase variants. In such cases, the observation of a tumor SNV in combination with a germline SNV in one sequence read provides unique identifying information that increases the probability that the tumor SNV is true.
[0045] The various target region classes are further illustrated in Figure 2B, where for each exemplary class, two DNA molecules are shown as lines, representing the two copies of each chromosome (paternal and maternal) that are amplified by PCR primer pairs targeting specific regions. As shown in Figure 2B, target region classes can include an SNV (250); an MNV (252), here a doublet base substitution; two SNVs (254), located on opposite chromosomes and therefore on different DNA molecules; two PVs (256), which are two SNVs located on the same chromosome and therefore on the same DNA molecule; a germline SNV and a tumor SNV (258), located on the same chromosome and therefore on the same DNA molecule, and therefore a PV as shown; one INDEL (260), representing a deletion of one base; one INDEL (262), representing an insertion of one base; an SNV and a one-base deletion (264), located on the same chromosome and therefore on the same DNA molecule, and therefore a PV as shown; and an EV (266), representing a methylated cytosine that has not been converted to uracil by bisulfite treatment.
[0046] In some embodiments, enrichment of multiple target regions may involve performing a multiplex PCR assay in which multiple target regions are simultaneously amplified in a test sample. Figure 3 illustrates one embodiment of a genetic variation profiling assay 300 using the methods described herein. The assay 300 may involve measuring multiple target regions, each containing a single class. As shown in Figure 3, each box represents a different target region measured by the assay, preferably in the same reaction volume. The assay detects many different target region classes, which may include single nucleotide variations (SNVs) 302, multi-nucleotide variations (MNVs) 304, copy number variations (CNVs) 306, short insertions / deletions (INDELs) 308, structural variations (SVs) 310, epigenetic variations (EVs) 312, and phase variants (PVs) 314. As shown, some target regions may contain more than one genetic variation, including two SNVs (316) (not on the same chromosome) and two PVs (314). A target region with two SNVs on separate chromosomes has the advantage of doubling the utility of a particular target region by simultaneously profiling two distinct variations, for example, by using the same PCR primer pair in a multiplex reaction. In some embodiments, either of the two SNVs or the PV may be a germline SNV. As previously described, a PV may contain any type or combination of genetic variations. For example, as further shown in Figure 3, a two-PV region (314) may contain a region with both an SNV (318) and a deletion of a single nucleotide INDEL (320).
[0047] The assays described herein may include any number of target region classes, and may include multiple target regions having the same class. In a preferred embodiment, the assay includes at least two different target region classes. The assay 300 can be applied to a test sample to determine the state of each target region profiled by the assay in the sample. In some embodiments, each target region can be enriched and measured (e.g., sequenced). For any given target region, measurements can be determined that support the presence of a target region class (e.g., a specific SNV, CNV, INDEL, SV, EV, and / or two or more PVs located within a target region).
[0048] As described in further detail below, the methods described herein can combine comparisons from multiple target region classes, such as multiple target regions with genetic variations in assay 300. One advantage of combining comparisons from various target region classes is that evidence from different variant types, when combined, can be considered to support a reliable conclusion that cancer DNA (or RNA) is present in the test sample. For example, in some embodiments, the first target region comprises a first class, where the first class comprises SNVs. In these embodiments, the second target region may comprise a second class that includes CNVs, INDELs, SVs, EVs, or, particularly, two or more PVs. In other embodiments, the first target region comprises a first class, where the first class comprises CNVs. In these embodiments, the second target region may comprise a second class, where the second class includes SNVs, INDELs, SVs, EVs, or, particularly, two or more PVs. In some embodiments, the third target region comprises a third class, and the third class comprises SNV, CNV, INDEL, SV, EV, or two or more PV, particularly two or more PV. Various combinations of target region classes are considered herein.
[0049] Target region can be selected by first identifying multiple gene variations of interest, for example, the gene variations associated with patient's cancer.Genetic variations can include previously identified sequence variations, for example, the variations that are known to be associated with patient's cancer or suspected to be associated with patient's cancer.Variations can also be identified from patient's cancer, for example, somatic mutations that exist in the genome of patient's cancer cells or exist in the genome of patient's cancer cells before any cancer treatment. For example, genetic variations can include variations present or previously identified in various cancer-associated genes, including, but not limited to, TP53, EGFR, BRAF, and KRAS, as well as other genes frequently mutated in cancer (e.g., those in the COSMIC Cancer Gene Census, available at cancer.sanger.ac.uk / census; see also Sondka et al., The COSMIC Cancer Gene Census: Describing Genetic Dysfunction Across All Human Cancers, Nature Reviews Cancer 18, 696-705 (2018), the contents of which are incorporated herein by reference); regions of common structural rearrangements (e.g., common gene fusions or edges of common amplifications, e.g., MYC), and regions of common amplifications, rearrangements (e.g., chromosomal thrombosis), commonly localized hypermutations (e.g., kataegis), epigenetic alterations, and the like.
[0050] In some embodiments, genetic variations may include cancer-specific genetic variations identified by sequencing DNA isolated from a patient's cancer cells. For example, cancer-specific variations can be identified by sequencing DNA or RNA isolated from a biological sample containing cancer cells obtained from a cancer patient. Tumor-specific variations can be identified by sequencing DNA or RNA isolated from a tissue sample obtained from a cancer patient's tumor biopsy. Alternatively, tumor-specific variants can be identified by sequencing cell-free DNA or RNA or DNA or RNA isolated from a patient's circulating cancer cells. In hematological cancers, genetic variations can be identified by sequencing DNA or RNA samples obtained, for example, from bone marrow, circulating blood cells, or lymph nodes. In such embodiments, the assays described herein may be "personalized" in that the genetic variations are obtained from the same patient.
[0051] In some embodiments, cancer-specific variants are identified using targeted sequencing methods such as hybrid capture sequencing.In another embodiment, cancer-specific variants are identified using pull-down or non-pull-down techniques to enrich selected sequences.These methods can sequence different genomic regions, such as exomes, that is, whole exome sequencing (WES), which can include genomic regions with common mutations in cancer genes or regions with frequent mutations outside genes.In a preferred embodiment, cancer-specific variants are identified through WES of tumor tissue.In other embodiments, tumor-specific variants can be identified using whole genome sequencing (WGS), where samples are sequenced without any specific enrichment.
[0052] WES and similar targeted sequencing methods effectively limit the genomic search space by selecting certain pre-identified sequences, thereby increasing coverage and confidence in calling somatic variations. Such methods also allow for the identification of genetic variations that are more likely to have functional effects. However, limiting the search space can reduce the total genetic variations obtained, which can also affect the types of variations identified. For example, in hematological cancers such as lymphoma, phase variants (PVs) tend to cluster in known "hotspot" regions, allowing them to be identified using either targeted techniques (such as WES) or WGS. However, in solid tumors, PVs tend to be randomly distributed throughout the genome, resulting in fewer PVs that are close enough together to reside within a single target region. Thus, prior art methods focused on identifying a single variant type, such as a PV, for use in cancer diagnostic assays rely on WGS, and in hematological cancers, on identifying a large number of PVs in "hotspot" regions (see, e.g., Kurtz, DM et al., Enhanced detection of minimal residual disease by targeted sequencing of phased variants in circulating tumor DNA, Nat Biotech 1-11 (2021), incorporated herein by reference in its entirety). The inventors recognized and appreciated that, for example, in methods for identifying genetic variations in target regions of solid tumors, PVs that are sufficiently close to each other can provide a significant increase in specificity, whereas (e.g.) WES or targeted techniques may discover only a few PVs within a sufficient distance. The methods described herein solve this problem by combining (e.g.) measurement of target regions with two or more PVs with measurement of target regions with other types of variants, enabling a "hybrid" approach that can interrogate an economical number of genetic variations in a test sample and utilize all available evidence.For example, in assay 300, evidence from two PV target regions 314 can be combined with evidence from one SNV target region 302. Conversely, prior art WGS-based methods rely on the identification of multiple PVs. Furthermore, focusing solely on phase variants misses important information. Therefore, methods that use only phase variants as target region classes are less sensitive than methods that utilize multiple classes.
[0053] In some embodiments, cancer-specific genetic variations are compared to genetic variations obtained from a matched normal sample. The sample is "normal" because it is derived from non-cancerous biological material, and "matched" because it is derived from the same patient. For example, a matched normal sample of non-cancerous DNA obtained from the same patient, such as buccal swab DNA, whole blood DNA, or adjacent non-cancerous DNA (i.e., derived from tissue adjacent to the tumor that appears normal), can be sequenced and compared to the patient's cancer-specific genetic variations. These matched normal samples can be sequenced simultaneously with, before, or after sequencing the patient's cancer cells. Genetic variations detected in cancer cells (cancer DNA) but not in the matched normal sample (non-cancerous DNA) are more likely to be cancer-specific and can therefore be selected for inclusion in the assays described herein. Variations detected in the matched normal sample (non-cancerous DNA) are likely not cancer-specific and can therefore be excluded. Those skilled in the art will be aware of various software packages for calling tumor-specific variants, such as MuTect2 (Cibulskis et al., Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol. 2013;31:213-9) and VarScan2 (Koboldt et al., VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res. 2012;22:568-76).
[0054] In some embodiments, the genetic variation is a clonal genetic variation. Cancer DNA contains both clonal and subclonal mutations. During tumor evolution, there is a transition between clonal and subclonal mutations. Subclonal mutations are present only in a subset of cells within a tumor. They arise after the most recent common ancestor of all cancer cells in a tumor sample. Conversely, clonal mutations arise before the most recent common ancestor of all cancer cells. Therefore, clonal mutations are present in all cells within a tumor unless some mechanism eliminates the mutation, for example, unless it is a structural variation, in which case the entire locus is lost in a subset of cells. Clonal changes typically occur early in cancer evolution and are present in all cancer cells. Genetic variations can be considered clonal if they are present in multiple biological samples or can be inferred from sequence reads obtained from tumor bulk tissue. Clonality can be difficult to determine because tumors are often heterogeneous, the entire tumor cannot be sequenced, and the heterogeneity of bulk sequencing data is difficult to quantify. Various approaches have been proposed to determine clonality, including Bayesian mixture models, clustering probability distributions of cancer cell fractions, and phylogenetic methods. Software tools for determining clonality include PyClone-VI, EXPANDS, QuantumClone, and PhyloWGS.See also Gillis, S., Roth, A., PyClone-VI: scalable inference of clonal population structures using whole genome data. BMC Bioinformatics 21, 571 (2020); Andor et al., EXPANDS: expanding ploidy and allele frequencies on nested subpopulations, Bioinformatics 30(1):50-60 (2013); Deveau et al., QuantumClone: clonal assessment of functional mutations in cancer based on a genotype-aware method for clonal reconstruction, Bioinformatics 34(11):1808-1816 (2018); Deshwar et al., PhyloWGS: reconstructing subclonal composition and evolution from whole-genome sequencing of tumors, Genome Biology 16(35) (2015), the entire contents of each of which are incorporated by reference. Samples can be sequenced by whole genome sequencing, whole exome sequencing, or targeted sequencing (e.g., by sequencing a panel of cancer genes or by sequencing a panel of sequences that are mutational hotspots), etc.
[0055] The target region with genetic variation can be ranked or filtered based on the type of genetic variation that exists.For example, the target region can be ranked based on one or more of the following: clonality, i.e., the allele proportion in cancer samples; the likelihood of unique alignment; the estimated background error rate, where the genetic variation that shows sequence evidence or PCR polymerase error rate is penalized or filtered; the high signal background event, where the genetic variation that shows DNA damage or PCR error in early cycle is penalized or filtered; the target region class, for example, the prioritization of PV pairs in consideration of its predictive utility; the proximity of any germline (not somatic) variant that can be useful for enrichment; the likelihood of somatic change; and the like.
[0056] Once a genetic variation is selected, the corresponding target region can be determined, for example, by selecting the upstream and downstream positions of the genetic variation.For example, the target region can include a genome section that starts (for example) 75 bp before the genetic variation and ends (for example) 75 bp after the genetic variation.In some embodiments, PCR primers are designed to amplify the target region.In some embodiments, oligonucleotide probes are designed to enrich the target region.In the embodiment where the test sample contains cfDNA, the target region is preferably designed to be about 150 bp long, reflecting the average fragment length of cfDNA molecules.
[0057] Measuring the plurality of target regions (step 104) can be done in a variety of ways. In some embodiments, the measuring is done by digital PCR (dPCR) or droplet digital PCR (ddPCR). The measuring can also be done by quantitative PCR or other fluorescence-based assays. In some embodiments utilizing molecular barcoding, the measuring can include generating a consensus sequence for the target regions and determining whether the consensus supports the target region class.
[0058] In some embodiments, measuring the plurality of target regions in the enriched sample comprises sequencing the plurality of target regions of step (a) to obtain a plurality of sequence reads corresponding to the first target region and the second target region. In such embodiments, comparing the measurements (e.g., step 106 of method 100) comprises comparing the amount of sequence reads supporting the presence of a target region class to one or more error models that model the probability that the target region class is observed in DNA or RNA that does not have the target region class.
[0059] Sequencing generally refers to the method of obtaining the identity of at least 10 consecutive nucleotides of a polynucleotide (for example, the identity of at least 20, at least 50, at least 100, or at least 200 or more consecutive nucleotides).In a preferred embodiment, sequencing is carried out using next-generation sequencing, i.e., a so-called highly parallelized nucleic acid sequencing method, including sequencing-by-synthesis, sequencing-by-ligation, and sequencing by combination platform currently utilized by Illumina, Life Technologies, Pacific Biosciences, Element Biosciences, Singular Genomics, Omniome, Genapsys, Ultima Genomics, and Roche.Next-generation sequencing methods can also include, but are not limited to, nanopore sequencing methods such as those provided by Oxford Nanopore, or electronic detection-based methods, such as the ion torrent technology commercially available from Life Technologies. In some embodiments, sequencing is performed using an Illumina NextSeq system or a NovaSeq system. In some embodiments, sequencing is performed using pyrosequencing, such as a Roche 454 GS FLX system.
[0060] The output of the sequencing process is multiple sequence reads, i.e., a string of letters indicating the order in which certain nucleotides (e.g., A, C, G, T) occur within the sequenced DNA molecule or amplicon. Sequence reads can vary in length from 25 to 1000 bp or more, and in many cases, each base in a sequence read may be accompanied by a score indicating the quality of the base call. As previously described, cfDNA in blood is typically highly fragmented, with an average length of approximately 160 bp. Thus, in some embodiments, the target region may comprise a length of approximately 160 bp, and the amplicon may comprise a length of approximately 160 bp or less. In these embodiments, the sequence read is preferably at least 160 bp in length to sequence the entire amplicon, and thus the entire target region.
[0061] In some embodiments, the sequence read corresponds to a first target region and a second target region. In some embodiments, the sequencing adaptor can be directly ligated to the amplicon having the sequences of the first target region and the second target region. In other embodiments, the sequencing adaptor can be incorporated into the amplicon during amplification, i.e., PCR. Various embodiments and modifications are considered within the scope of the present disclosure.
[0062] In some embodiments, the target region comprises a class that includes two or more phase variants (PVs). In such embodiments, two or more PVs are present (or are expected to be present) within the same target region. Because the PVs are present on the same DNA molecule, the sequence reads of the corresponding amplicon can contain both PVs, providing highly specific evidence of cancer DNA present in the sample.
[0063] In some embodiments, due to the low proportion of cancer DNA, high-depth sequencing may be required to identify genetic variations.In some embodiments, sequencing of multiple target regions comprises sequencing to a minimum read depth of at least 10,000, at least 25,000, at least 50,000, or at least 100,000, at least 200,000, or at least 500,000.In some embodiments, sequencing of multiple target regions comprises sequencing to a maximum read depth of at least 25,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, or at least 1,000,000.In any embodiment, the read depth of each step can be from about 10,000 to about 500,000.In any embodiment, the read depth can be from about 10,000 to about 200,000.
[0064] In some embodiments, sequence reads are processed by computer, for example, by trimming, demultiplexing, aligning, matching, collapsing, and / or filtering.Typically, processing assigns each sequence read to one of the target regions that have or are suspected to have one or more genetic variations associated with the patient's cancer.For example, sequence reads can be analyzed to identify which reads correspond to multiple target regions.As will be appreciated by those skilled in the art, sequence reads that are identical or nearly identical to the target region can be analyzed to determine whether potential genetic variations exist within the target sequence.Sequences can be aligned with reference sequences, such as genome sequences, or can be matched with a database of predicted sequences to determine their most likely location on the reference sequence.
[0065] After the sequence reads are processed, the amount (e.g., number) of sequence reads (k) having the gene variation or multiple gene variations and the total amount (e.g., total number) of sequence reads (n) can then be determined for each target region. Methods for quantifying reads can be adapted from those described, for example, by Forshew et al. (Sci.Transl.Med.2012 4:136ra68), Gale et al. (PLoS One 2018 13:e0194630), and Weaver et al. (Nat.Genet.2014 46:837-843), all of which are incorporated herein by reference in their entirety. Similar results can be obtained using approaches that utilize molecular indexes. In these methods, the total number of sequenced molecules and the number of variant molecules can be estimated using the index. Such molecular identifier sequences can be used in combination with other features of the fragments (e.g., the final sequence of the fragment that defines the breakpoint) to distinguish between fragments. Molecular identifier sequences are described in (Casbon Nucl. Acids Res. 2011, 22e81), which is incorporated herein by reference in its entirety.
[0066] Comparing the measurements supporting the presence of a target region class to one or more error models that model the probability that the target region class is observed in DNA that does not have the target region class (step 106) can be done in a variety of ways. In some embodiments, the comparison includes comparing, for each target region, the measurements supporting the presence of a class (k) to a binomial distribution model, a variance binomial distribution model, a beta-binomial distribution model, a multinomial distribution model, a normal distribution model, an exponential distribution model, or a gamma error probability distribution model. For example, in one embodiment, the error probability distribution model for a first target region class is a beta-binomial error probability distribution model, and the error probability distribution model for a second target region class is a multinomial error probability distribution model. In some embodiments, the number of sequence reads supporting the presence of a class (k) includes the number of sequence reads that have a genetic variation. In some embodiments, the comparison further includes obtaining a statistical evaluation or score indicating the degree of evidence supporting a conclusion that a given target region has one or more genetic variations in a sample. In some embodiments, the statistical evaluation can be, for example, a p-value, a likelihood, a likelihood ratio, or a probability distribution. The statistical evaluation also preferably includes a likelihood ratio approach, in which the likelihood of observing n sequence reads with one or more genetic variations in a test sample is determined when cancer DNA is present in the sample and when cancer DNA is not present in the sample. These values can then be used (for example) to calculate a likelihood ratio to determine whether one or more genetic variations in the target region are present in the sample.
[0067] As previously mentioned, cancer DNA, if present, is often a small percentage of cell-free DNA. For example, in MRD, the cancer fraction can be as low as 0.01 ppm. At this level, the inventors have recognized and understood several issues that can result in false-positive results, for example. First, sequencing is not perfect, and background errors can lead to misleading bases, potentially resulting in false-positive ctDNA calls. Second, errors can also be introduced during PCR. For example, bases can "switch" due to DNA damage (e.g., oxidation, deamination) before amplification, and subsequent amplification by PCR can result in many sequence reads that support an inaccurate conclusion. Furthermore, as the number of genetic variants included in an assay increases, the probability of a false-positive call increases. Some assays require (for example) at least two individual genetic variations to be called positive for accurate diagnosis. However, this approach can be flawed by limitations in the number of variants tested by the assay, resulting in a low amount of available ctDNA "signal" and low sensitivity.
[0068] One way to account for background error is to model the error as a probability distribution and determine whether the observed genetic variation is likely to arise from background error. For example, the probability of observing k sequence reads with genetic variation in a target region with a background error rate p can be determined using the following binomial probability distribution:
[0069]
number
[0070] Probability (e.g., P(X=k)) refers to the likelihood of a particular outcome occurring or the likelihood of that outcome occurring. Probability may be based on the values of parameters in a model. Probability refers to an unknown event and is associated with possible outcomes. Because possible outcomes are mutually exclusive and exhaustive, probability can be expressed on a linear scale. For example, probability may be expressed as a value between 0 (impossible) and 1 (certain), or alternatively, as a percentage or proportion. For example, in the context of the present disclosure, probability can be used as a measure to determine whether cancer DNA is present in a sample.
[0071] An error probability distribution (which may also be referred to as an "error model," "error distribution," or "error probability distribution model") refers to a distribution that estimates or models the probability that an observation (such as a variant allele proportion) is due to an error. These terms can refer to any type of error, including errors due to DNA damage or early-cycle PCR errors, and sequencing errors. A hypothetical error model is shown as a frequency distribution in Figures 4A-B. In these examples, multiple samples (e.g., hundreds of samples) not known to have somatic genetic variations (i.e., healthy control samples) are sequenced, and the proportion of sequence reads with a particular type of sequence variation is calculated for each sample. All sequence variations within a sequence read are primarily due to errors that occur during PCR, base miscalls, and pre-PCR events such as DNA damage (e.g., oxidation of guanine to 8-oxoguanine, which base pairs with A and causes a G to T variation in the sequence read). These proportions can be plotted as a frequency distribution, which can be used to calculate the probability that the sequence variation observed in a sequence read is truly a genetic variation.
[0072] In some embodiments, likelihood ratios (LRs) can be used to estimate the degree of evidence supporting the existence of a target region class.Likelihood ratios refer to the ratio of at least two likelihoods, each associated with a different hypothesis, and can be used to determine which hypothesis is more likely to be the experimental outcome.Each likelihood refers to the hypothesized probability that a particular outcome will be obtained by an event that has already occurred.Likelihoods can be used to evaluate how well a sample supports a particular value of a parameter in a model.Likelihoods therefore refer to past events with known outcomes and are associated with hypotheses.
[0073] The likelihood ratio can be used to determine the potential utility of a particular diagnostic test and to determine the degree to which a patient is likely to have a disease or condition, and can therefore be used as a measure of diagnostic accuracy. When applied to diagnostic tests, the likelihood ratio is the likelihood of expecting a given result in a sample without any cancer DNA compared to the likelihood of expecting the same result in a sample containing cancer DNA. As shown in equations (2), (3), and (4) below, two hypotheses can be determined: H0: the likelihood of observing k reads with genetic variations, given the absence of cancer in the sample (the null hypothesis); and H1: the likelihood of observing k reads with genetic variations, given the presence of at least one cancer molecule (thus supporting the target region class), where z is the total number of input DNA molecules, which can be estimated, for example, by light diffraction or digital PCR. Each hypothesis incorporates a background error rate (p), which indicates the frequency with which sequence reads with genetic variations in the target region class may be due to errors. The background error rate (p) can be selected individually for each target region. The ratio of these two values (H1 / H0) can then be determined and compared to a threshold. A value above 1 suggests that H1 is a more likely hypothesis, while a value between 0 and 1 suggests that H0 is more likely. The threshold required to call a variant associated with the current cancer can vary depending on the desired sensitivity and specificity, which can be determined (for example) from a set of known samples.
[0074]
number
[0075] Thus, in any embodiment, comparing the number of sequence reads (k) supporting the presence of a target region class to one or more error models (step 106) may include calculating a likelihood ratio between the likelihood of observing a number of sequence reads with a genetic variation (i) when cancer DNA is present and (ii) when cancer DNA is not present. Similarly, in any embodiment, this may be calculated as a likelihood ratio (LR) between the likelihood of observing k reads in each target region (i) when cancer DNA is present and (ii) when cancer DNA is not present. i ) As described in more detail below, in these embodiments, the individual likelihood ratios LR i are combined to produce a cumulative LR score (e.g., LR = ∑ i ... i (product of
[0076] In some embodiments, the target region class includes two or more phase variants (PVs). The two or more PVs may include a first genetic variation and a second genetic variation located on the same DNA molecule. Therefore, each variation is sequenced together on one sequence read. In these embodiments, comparing the amount of sequence reads supporting the existence of a target region class containing two or more phase variants with one or more error models that model the observation of the target region class in DNA that does not have the target region class (step 106) can be performed by comparing the number of sequence reads that have both the first genetic variation and the second genetic variation, the number of reads that have only the first genetic variation, the number of reads that have only the second genetic variation, and the number of reads that have neither genetic variation with a multinomial distribution. For example, the probability of observing two phase variants in a target region can be modeled as follows:
[0077]
number
[0078] A variety of probability density functions can be used for X. One candidate for this role is the standard Dirichlet distribution:
[0079]
number
[0080]
number
[0081] Individual probabilities can also be calculated for X. For example, X i =P(k i ), i.e., k i is the probability that a lead of θ is observed. Considering the tumor fraction θ, i The probability of observing is:
[0082]
number
number
[0083] While the above embodiment describes two phase variants, the methods according to the present disclosure can be further modified by one skilled in the art to have target regions with additional PVs, for example, three or more, four or more, five or more, or ten or more PVs. However, in preferred embodiments, two PVs per target region class is typically sufficient, since two PVs are likely to be found in close proximity to each other and present on a single DNA fragment.
[0084] Various error models can be used to model the probability that a target region class is observed in DNA that does not have that target region class. In some embodiments, an error model corresponding to the target region class is selected. In some embodiments, the target region class is SNV, and the error model is a binomial distribution. In some embodiments, the target region class is two or more PVs, and the error model is a multinomial distribution. In some embodiments, the comparison includes comparing two or more error models. In such embodiments, the two or more error models can model various types of errors, including but not limited to sequencing errors, PCR errors, DNA damage, polymerase errors, and the like. For example, a first error model (e.g., a binomial probability distribution) can be used to estimate the background error rate from sequencing, and a second error model (e.g., a Poisson distribution) can be used to estimate the background error rate from DNA damage. In some embodiments, a single distribution can account for one or more types of errors. For example, the two shape parameters (α, β) of the beta-binomial distribution can be tuned to accommodate the estimated background error rate and DNA damage.
[0085] In some embodiments, the background error rate is estimated using a probability distribution. In some embodiments, there may be two distributions of the same family or type (e.g., two binomial distributions), or, if two different families or types of distributions are used, one distribution may be for the background error rate and another distribution may be for PCR errors.
[0086] In some embodiments, H1, the likelihood of observing k reads with a genetic variation (and thus supporting the target region class) given the presence of at least one cancer molecule, is multiplied by an additional probability, e.g., the probability that the genetic variation is cancer-specific (V C For example, equation (3) above can be modified as follows:
number
[0087] Furthermore, the inventors have recognized and appreciated that short insertions and deletions (INDELs) can be particularly useful when included as target region classes according to one embodiment of the present disclosure. For example, a target region class may include, for example, INDELs 1 to 5, 1 to 10, 1 to 15, or 1 to 20 nucleotides in length. In such embodiments, INDELs 1 to 5 nucleotides in length are preferred because they are more likely to be observed in test samples. In other embodiments, a target region class may include INDELs 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, or 20 or more nucleotides in length. In some embodiments, the maximum INDEL length is 20 nucleotides because longer variations can affect the accuracy of sequence alignment. Depending on the type of sequencing technology used, the background error rate associated with INDELs may be very low. Thus, observing an INDEL in a target region can provide a relatively large amount of evidence that an INDEL (and thus cancer DNA) is present in a test sample.
[0088] The error model can be trained using a control sample, for example, a DNA sample known to lack any sequence variations or collected from a healthy patient (e.g., a patient without cancer). By sequencing a control sample known to lack any sequence variations, all observed sequence variations in the control sample are attributed to errors. Such observations can be used to set the parameters of an error model according to the present disclosure. Preferably, the control sample is processed under conditions similar to those of the test sample. For example, the primers can amplify the same or similar target region, and the sequencing technique can be identical. Many control samples, for example, at least about 50 samples, can be used to build the error model. The error model can be recorded in a computer database and accessed as needed. Thus, in one embodiment, the error model is trained according to a set of control samples. In one embodiment, the set of control samples is derived from healthy donors. In such an embodiment, training the error model based on the set of control samples establishes the background error (p) for the target region class in the absence of cancer.
[0089] In some embodiments, a multiple comparison correction is applied to the comparison to prevent false positives, for example, by setting a more stringent threshold or applying a Bonferroni correction. In any embodiment, the threshold can be determined using a binomial distribution model, an overdispersed binomial distribution model, a beta distribution model, a normal distribution model, an exponential distribution model, or a gamma probability distribution model of the background error rate of sequence variation, where the frequency is selected so that signals are observed above the threshold less than 0.1% of the time, less than 0.01% of the time, or less than 0.001% of the time, preferably less than 0.1% of the time, depending on the desired pre-defined specificity per variant in the absence of mutant molecules.
[0090] In some embodiments, CNVs are identified using a read depth approach, which uses non-overlapping sliding windows to count the number of sequence reads that are mapped to genome regions that overlap the windows.Regions with significant increases in read depth (exceeding what is expected according to typical background errors associated with sequencing) can be further analyzed to identify copy number.Alternatively, a paired-end approach can be used, in which copy number variations are detected based on the distance between paired mapped sequence reads.Sequence reads can also be assembled de novo, and the resulting assembled continuous sequence can be aligned to a reference genome to identify copy number variations.
[0091] In some embodiments, epigenetic variants (EVs) are identified by processing test samples and then sequencing.For example, methylated nucleotides can be identified by treating them with sodium bisulfite, which converts unmethylated cytosine to uracil.By amplifying and sequencing the sample, the uracil base is then converted to thymine (T).Therefore, the presence of unmodified cytosine bases in the target region can be supported by the presence of the target region with EVs.The comparison of EVs can be performed using an error model that also measures the background level of sequencing errors, which can sometimes also reveal all errors associated with the bisulfite conversion process.
[0092] The step of combining comparisons of at least the first and second target regions (step 108) can be performed in a variety of ways. In some prior art methods, each genetic variation is called individually rather than as a whole set. As the number of variants increases, statistical corrections can be applied to individual variant calls, but the higher stringency required to eliminate false positives can also have a negative impact on sensitivity, causing most variant calls to be ignored. The inventors recognize and understand that independent analysis of each target region can provide some level of evidence for a cumulative statistical evaluation. Rather than considering each target region individually, combining the scores of two or more target regions can result in a reliable call for the test sample without a corresponding decrease in sensitivity or increase in false positives. Thus, in some embodiments, comparisons of each of the first and second target regions (and any other target regions considered) can be combined into a score or statistical evaluation that measures the overall degree of evidence supporting the conclusion that cancer DNA is present in the test sample. In some embodiments, the comparisons can be combined to provide a cumulative statistical evaluation that corresponds to the probability or likelihood that cancer DNA is present in the test sample. A variety of methods can be used to obtain a cumulative statistical assessment to identify whether cancer DNA is present in the test sample, including joint statistical measures (such as joint probability, joint likelihood, or joint likelihood ratio) or otherwise combining the results for each target region (e.g., summing, averaging).
[0093] In some embodiments, the combining includes calculating an average of each comparison. In one embodiment, the average is a weighted average. For example, a comparison from a first target region having a class of two or more PVs can be weighted 1.0, while a comparison from a second target region having a class of single SNVs can be weighted 0.5. In this way, the class of two or more PVs provides additional weight because the probability of two genetic variations being observed together in one sequence read is less likely to be the result of an error. Thus, in one embodiment, the step of combining the comparisons of at least the first and second target regions (step 108) can further include adjusting each comparison by a weight. In some embodiments, the weight of the class of two or more PVs is 1.0. In some embodiments, the weight of the class of single SNVs is 0.5. In some embodiments, each comparison may be performed using a different calculation, for example, using a different error model or statistical technique for evaluating each target region. For example, in some embodiments, a first target region having a class of single SNVs uses an error model derived from a binomial distribution, and a second target region having a class of two or more PVs uses an error model derived from a multinomial distribution. In such embodiments, different weights can be applied to each type of comparison, such that the evidence supporting the presence of a cancer call depends on the statistical evaluation being made.
[0094] In some embodiments, the comparison involves a statistical evaluation, such as obtaining a p-value that indicates the probability or likelihood that a genetic variant is present in the test sample. In these embodiments, the p-values for each target region can be combined, for example, using Fisher's method. If U is distributed as Uniform (0, 1), -2logU is calculated as a chi-squared (X2) with 2 degrees of freedom. 2 ) is distributed as χ1, ,χ k But, X v1 2 ,···,X vk 2 If they are independently distributed as Χ1+···+Χk ,X 2 v1+···+vk p1, ,p k are independently distributed as Uniform(0,1), the combined p-value is:
number
number
[0095] In some embodiments, the statistical evaluation may include a likelihood ratio, which indicates the likelihood that a genetic variant is present in the test sample. Individual likelihood ratios are combined to produce a cumulative LR score (LR, which is equal to the sum of the log-likelihoods) across all target regions of the sample. i The likelihood or likelihood ratio can be calculated as a product of (a product of) likelihoods. In these embodiments, the likelihoods or likelihood ratios can be combined by finding the log-likelihood of each individual event, i.e., the sum of the likelihoods calculated for each target region. Thus, when estimating parameters using log-likelihood in maximum likelihood estimation, each data point is used by adding it to the sum of the log-likelihoods. Thus, each data point is evidence supporting the estimated parameter, where each data point adds independent evidence to the identification of whether cancer DNA is present in the sample.
[0096] In one embodiment, a log-likelihood or log-likelihood ratio can be determined for each genetic variation and then combined, for example, by summing. The likelihood ratio for the entire sample can be determined by summing the log-likelihood ratios for each of the target regions (1 ..V), as shown in equation (6) below. This means that
number
number
[0097] By combining and considering evidence from multiple target regions, embodiments described herein can be scaled to any practical number of target regions. For example, while assay 300 of FIG. 3 includes 30 target regions, embodiments of the present disclosure can be scaled to 1-100 target regions, 100-200 target regions, 200-300 target regions, 300-400 target regions, or 400-500 target regions. In some embodiments, the number of target regions is 500-1000 target regions, 1000-2000 target regions, 2000-3000 target regions, 3000-4000 target regions, or 4000-5000 target regions. In some embodiments, the number of target regions is 5,000-10,000 target regions, 10,000-20,000 target regions, 20,000-30,000 target regions, 30,000-40,000 target regions, or 40,000-50,000 target regions. In some embodiments, the number of target regions depends on the type of cancer. For example, melanoma can have up to 1 million SNVs, while most cancers have approximately 100,000 SNVs. Thus, in these embodiments, the number of target regions can include 50,000-100,000 target regions, 100,000-200,000 target regions, 200,000-300,000 target regions, 300,000-400,000 target regions, 400,000-500,000 target regions, or 500,000-1 million target regions. In any embodiment, the number of target regions is at least 2, at least 4, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, or at least 5000 target regions. In many embodiments, between 2 and 200, for example, between 6 and 100 target regions are interrogated.
[0098] The step of identifying cancer DNA in the test sample based on a combined comparison of the first and second target regions (step 110) can be performed in various ways. In some embodiments, the step of identifying cancer DNA in the test sample can include comparing a cumulative score of multiple target regions to a threshold. In any embodiment, cancer DNA can be identified or otherwise considered to be present in the test sample based on the cumulative score. For example, if the cumulative score exceeds a threshold, cancer DNA can be identified in the test sample. An appropriate threshold can be determined empirically, for example, by sequencing a quantity of a test sample in which it is known in advance whether cancer DNA is present, and then selecting a threshold that has both high sensitivity (i.e., the ability to detect) and high specificity (i.e., the ability to distinguish). In such embodiments, the threshold can be determined by running at least 10, at least 100, at least 1,000, or at least 10,000 samples containing non-cancerous DNA (or at least not known to have cancer DNA) through the assay and selecting a threshold that exceeds the signal identified in the control samples or that estimates a false positive rate of 1%, 0.1%, or 0.01% or less as determined using the control samples. The samples run can be from the same patient or from different patients. For example, a run of 200 samples could involve obtaining samples from 20 healthy donors (suspected not to have cancer) and performing 10 assays per patient until 200 samples are reached. A likelihood ratio analysis can be performed on each control sample to obtain an overall likelihood ratio for healthy patients. Calculating likelihood ratios for all samples run would provide a range of likelihood ratios for healthy patients, and the threshold can be set somewhere above the maximum likelihood ratio. This threshold can be pre-calculated from a pool of healthy donors and therefore will not vary from patient to patient. As will become apparent, methods according to the present disclosure may further include identifying a patient with cancer if the result is above a threshold, and, for example, administering a treatment to the patient.In these embodiments, the patient may have previously received a first therapy, in which case the method includes administering to the patient a second therapy that is different from the first therapy.
[0099] As described above, the inventors recognize and understand that different target region classes containing different types of variants can provide different levels of evidence to support the identification of cancer DNA in test samples.For example, the observation of phase variants and INDELs (especially INDELs longer than 1 nucleotide, for example, 2, 3, 4, 5 or more nucleotides, preferably 2 or 3 nucleotides) within target regions is highly unlikely to be caused by background error.On the other hand, single-base variations provide relatively little evidence for the conclusion that cancer DNA exists, because these types of variations are more likely to be the result of error.Therefore, the assay according to the present disclosure can combine the results from multiple target region classes together to utilize all tumor-related variants that may be available, and obtain a highly sensitive and specific cancer DNA detection assay.
[0100] In some embodiments, the test sample can be divided into two or more aliquots, which can be processed in the same manner according to embodiments of the present disclosure. For example, each aliquot can be similarly enriched, measured, and compared for a specific target region, and this comparison can be combined to obtain a reliable identification of cancer DNA present in the test sample. Further information about aliquot generation and the use of replicate samples can be found in jointly filed International Patent Application No. PCT / IB2022 / 051195, filed February 10, 2022, which is incorporated herein by reference in its entirety.
[0101] In any embodiment, variant allele fraction (VAF) can be determined for test sample. VAF can be determined by, for example, the amount of sequence reads that support the existence of target region class. The term "variant allele fraction", "estimated variant allele fraction", "VAF" or "eVAF" refers to the estimated allele fraction of variant cancer DNA in test sample.
[0102] In any embodiment, the amount of cancer DNA in the test sample may be quantified. Quantification may include estimated variant allele proportions. In some embodiments, the estimated allele proportions may include the mean or median variant allele proportions for each target region determined to have a target region class present. In some embodiments, the estimated variant allele proportions may include the mean (k / n) of the variant allele proportions for each variant. This may be preferable in situations where variant levels are low and results are probabilistic; therefore, including evidence from all variants can provide a more realistic measurement. The quantified cancer DNA can be compared to cancer DNA quantified from one or more additional samples, for example, cancer DNA quantified from samples obtained from a patient at at least a first time point and a second time point, where the first time point is before treatment and the second time point is after treatment. Similarly, individual variants or groups of variants can be tracked across samples at different time points.
[0103] In any embodiment, the methods described herein can be performed on test samples obtained from a patient at at least a first time point and a second time point, where the first time point is before treatment and the second time point is after treatment, and the method includes determining whether there is a change in the amount of cancer DNA or the range of possible amounts of cancer DNA between the first and second time points. In any embodiment, additional samples can be obtained at additional time points, for example, additional samples are taken after the second time point on a monthly, twice-monthly, quarterly, or annual schedule. This embodiment can be used to monitor whether a treatment being administered to a patient continues to be effective. The change in cancer DNA over time can be determined using point estimates, confidence intervals, or both, where a significant (e.g., statistically significant) decrease indicates that the treatment is effective, and the lack of a significant change or increase indicates that the treatment is not effective. This embodiment can also be used to monitor whether cancer is recurring after curative surgery. Changes in cancer DNA over time can be determined using point estimates, confidence intervals, or both, with the absence of detectable cancer DNA indicating that cancer has not recurred, and a significant change or increase indicating a high likelihood of cancer progression. In these cases, a change of at least 2-fold, at least 4-fold, at least 6-fold, at least 8-fold, or at least 10-fold can be considered significant. In these cases, a change of at least 20%, at least 30%, at least 50%, at least 70%, or at least 90% can be considered significant. In some embodiments, a change is considered significant if it exceeds a threshold, such as 50%, and if the confidence intervals do not overlap when quantifying cancer DNA at the first and second time points. In these embodiments, a significant decrease indicates that the treatment is effective, and the absence of a significant change or increase indicates that the treatment is ineffective.
[0104] In some embodiments, the disclosed method may further include providing a report indicating whether cancer DNA is present in the sample. In some embodiments, the report may include a likelihood ratio or score (or another number indicating the same) as described above, and a threshold value to which the likelihood ratio can be compared to determine whether the test sample contains cancer DNA. If the report indicates that cancer DNA is not present in the sample, but the likelihood ratio or score or another number indicating the same is close to the threshold value, the report may advise scheduling a follow-up test soon (e.g., within one month, two months, or three months) to reevaluate whether the value now exceeds the threshold value for determining whether the sample contains cancer DNA. In some embodiments, the report may additionally list approved (e.g., FDA- or EMA-approved) treatments for treating disease or residual disease, such as chemotherapy or immunotherapy. This information may aid in disease diagnosis (e.g., whether the patient has MRD) and / or treatment decisions made by a physician.
[0105] In some embodiments, a sample can be collected from a patient at a first location, e.g., a clinical setting such as a hospital or doctor's office, and the sample can be sent to a second location, e.g., a laboratory, where the sample is processed, the above-described methods are performed, and a report is generated. A "report," as used herein, is an electronic or physical document that includes a report element that provides test results that may indicate the presence and / or amount of cancer DNA in the sample. Once generated, the report can be sent to another location (which may be the same location as the first location), where it can be interpreted by a medical professional (e.g., a clinician, laboratory technician, or physician, e.g., an oncologist, surgeon, pathologist, or virologist) as part of clinical judgment.
[0106] The patient whose sample is analyzed by this method may have any type of cancer or may have previously undergone treatment for any type of cancer. For example, the patient may have or have previously had melanoma, carcinoma, lymphoma, sarcoma, or glioma. For example, the cancer may be, among others, melanoma, lung cancer (e.g., non-small cell lung cancer), breast cancer, head and neck cancer, bladder cancer, Merkel cell carcinoma, cervical cancer, hepatocellular carcinoma, gastric cancer, cutaneous squamous cell carcinoma, classical Hodgkin's lymphoma, B-cell lymphoma, colorectal cancer, pancreatic cancer, gastric cancer, or breast cancer, including other solid tumors and blood cancers. In some embodiments, the cancer is a type of cancer that exhibits an average mutation rate of at least 0.1 mutations per megabase, or at least 0.2 mutations per megabase, or at least 0.5 mutations per megabase, or at least 1 mutation per megabase, or at least 10 mutations per megabase. In some embodiments, the cancer is a cancer that exhibits an average mutation rate of at least 0.5 mutations per megabase. Methods for calculating mutation rates are known in the art (e.g., Schumacher TN, Schreiber RD. Neoantigens in cancer immunotherapy. Science. 2015;348(6230):69-74, incorporated herein by reference in its entirety).
[0107] In some embodiments, the method can be used to guide treatment decisions. In some embodiments, the method can be used to determine whether a patient should be treated again, for example, with the same or a second therapy. For example, if a patient has previously been treated with a first cancer therapy and the patient is identified as having MRD using the method, the patient can be treated with a second cancer therapy that is the same or different from the first cancer therapy. For example, if a patient has previously been treated with surgery or an immune checkpoint inhibitor and the patient is identified as having MRD, the patient can be treated with additional surgery, the same or a different immune checkpoint inhibitor, or another type of therapy, where immune checkpoint therapy includes administration of a CTLA-4, PD1, PD-L1, TIM-3, VISTA, LAG-3, IDO, or KIR checkpoint inhibitor; other types of therapy include, for example, (a) anthracycline therapy (e.g., by administration of daunomycin, doxorubicin, or mitoxantrone); (b) alkylating agent therapy (e.g., mechlorethamine, cyclophosphamide, (c) topoisomerase II inhibitor therapy (e.g., by administration of etoposide or teniposide), (d) bleomycin therapy, (e) antimetabolite therapy (e.g., by administration of methotrexate, 5-fluorocysteine, cytarabine, 6-mercaptopurine, or 6-thioguanine), (f) vinca alkaloid therapy (e.g., by administration of vincristine or vinblastine), (g) steroid therapy (e.g., by administration of prednisone or dexamethasone), and (h) radiation treatment.Alternative therapies include targeted therapy and non-targeted chemotherapy. Targeted therapy includes erlotinib (Tarceva), afatinib (Giotrif), gefitinib (Iressa), or osimertinib (Tagrisso), which may be administered to patients with activating mutations in EGFR; crizotinib (Xalkori), ceritinib (Zykadia), alectinib (Alecensa), or brigatinib (Alumbrig), which may be administered to patients with ALK fusion genes; and crizotinib (Crimalib), which may be administered to patients with ROS1 fusion genes. Treatment with zotinib (Xalkori), entrectinib (RXDX-101), lorlatinib (PF-06463922), crizotinib (Xalkori), entrectinib (RXDX-101), lorlatinib (PF-06463922), repotrectinib (TPX-0005), DS-6051b, ceritinib, ensartinib, or cabozantinib; or dabrafenib (Tafinlar) or trametinib (Mekinist), which may be administered to patients with activating mutations in BRAF. Many other actionable mutations are known. If the patient intends to switch to non-targeted chemotherapy, the treatment may be, for example, a platinum-based doublet chemotherapy (the platinum-based doublet chemotherapy may include a platinum-based agent selected from cisplatin (CDDP), carboplatin (CBDCA), and nedaplatin (CDGP)), and one third-generation agent (selected from docetaxel (DTX), paclitaxel (PTX), vinorelbine (VNR), gemcitabine (GEM), irinotecan (CPT-11), pemetrexed (PEM), and tegafur-gimeracil-oteracil (S1)).
[0108] Methods for diagnosing cancer are described herein and include subjecting a test sample obtained from a patient to a method for detecting cancer DNA in the test sample according to any of the methods disclosed herein.
[0109] The method for treating cancer in patients is described herein, and includes determining the presence or absence of cancer DNA detected in test samples obtained from patients according to any method described herein, and administering cancer treatment or treatment to patients or recommending the administration of cancer treatment or treatment to patients.Administering or recommending is based on the identification of cancer DNA in test samples.For example, if cancer DNA is detected, treatment or treatment can be administered or recommended.
[0110] The method for treating cancer in a patient is described herein, wherein the patient is diagnosed as having or suspected of having cancer based on the presence or absence of cancer DNA detected in the test sample obtained from the patient, which is determined according to any method disclosed herein.The method comprises administering cancer treatment or treatment to the patient based on the identification of cancer DNA detected in the test sample obtained from the patient.In some embodiments, the method comprises recommending cancer treatment or treatment to the patient based on the identification of cancer DNA present in the sample obtained from the patient.
[0111] Methods for determining the effectiveness of a cancer treatment or therapy are described herein and include administering a cancer treatment or therapy to a patient, obtaining a test sample from the patient, and determining the presence, absence, or amount of cancer DNA in the test sample according to any method disclosed herein. In some embodiments, the method may include obtaining a test sample from the patient before administration of the cancer treatment or therapy, and comparing the presence, absence, or amount of cancer DNA in the test sample obtained before administration of the cancer treatment or therapy with the presence, absence, or amount of cancer DNA in the test sample obtained after administration of the cancer treatment or therapy. The difference may be an indicator of the effectiveness of the cancer treatment or treatment. For example, an increase in the amount of cancer DNA may indicate that the cancer treatment or treatment is ineffective. Thus, the method may include administering an alternative and / or additional cancer treatment or treatment to the patient, or recommending an alternative and / or additional cancer treatment or treatment to the patient. Conversely, a reduction or disappearance of cancer DNA in the test sample (which is an apparent disappearance, i.e., below the detection level of the method) may indicate that the cancer treatment or treatment is effective. Therefore, the method can comprise continuing or stopping the administration of cancer treatment or treatment to patient, or continuing or stopping the recommendation of cancer treatment or treatment.In some embodiments, the method can comprise carrying out cancer DNA detection method to monitor the effect of cancer treatment or treatment by using the patient's test sample collected at least two time points during the administration of cancer treatment or treatment, for example, the test sample obtained over time for one or several days, one or several months, one or several years, or other time points disclosed herein.
[0112] The present disclosure also provides a method for detecting or monitoring minimal residual disease (MRD), comprising obtaining or having obtained a test sample from a patient who has been treated or treated for cancer, and performing a method for detecting cancer DNA in the test sample according to the methods disclosed herein.
[0113] Treatment or therapy recommendations may be made in any suitable manner, for example, by providing a report containing the recommendations.
[0114] The cancer treatment or therapy can be any suitable treatment.For example, the cancer treatment or therapy can be tumor resection.The cancer treatment or therapy can be the administration of a pharmacological treatment for cancer.In some embodiments, the method of the present disclosure can be performed on a patient undergoing surgery to remove a tumor.In some embodiments, the cancer treatment or therapy that is administered or recommended after detecting the presence or amount of cancer DNA in a test sample obtained from a patient can be a pharmacological cancer treatment or therapy.
[0115] The methods described herein can be used to monitor treatment. For example, in some embodiments, the method can include analyzing a sample obtained at a first time point using the method, and analyzing a sample obtained at a second time point using the method, and comparing the results, i.e., determining whether cancer DNA is present in the sample, or determining whether there is a change in the amount of cancer DNA or the range of possible amounts of cancer DNA between the first and second time points. In some embodiments, such changes can be determined using point estimates or confidence intervals, and a significant decrease can indicate that the treatment is effective, while the absence of a significant decrease or increase can indicate that the treatment is not effective. The first and second time points can be before and after treatment, or two or more time points after treatment. For example, by comparing the results obtained at one time point with the results obtained at another time point, the method can be used to determine whether previously identified variations no longer exist, are reduced, or are increased in a subject during the course of treatment. The period between the first and second time points can be at least one month, at least six months, or at least one year, and in some cases, patients can be examined periodically, for example, every three months, every six months, or every year for several years, for example, for more than five years. In another embodiment, the method can be used to evaluate the effectiveness of a treatment by monitoring the patient's ctDNA level at certain time intervals after administration of the treatment. For example, if the treatment is effective, the ctDNA level should rise shortly after administration due to apoptosis of cancer cells, and then decrease significantly as ctDNA degrades. In such embodiments, the time between administration of the treatment and the first time point can be, for example, at least 15 minutes, at least 30 minutes, at least 45 minutes, and at least 1 hour. In such embodiments, the time between the first and second time points can be, for example, every 15 minutes, every 30 minutes, every 45 minutes, every hour, every 2 hours, or every hour for several hours, for example, for 8 hours or more.
[0116] The method according to the present disclosure can also be used to determine whether the subject is disease-free or whether the disease recurs.As mentioned above, this method can be used to analyze minimal residual disease and detect recurrence.In these embodiments, the primer pair used in this method can be designed to amplify the sequence with gene variation that has been previously identified in patient's cancer through sequencing cancer material, cfDNA at early time point or other suitable sample.
[0117] In some embodiments, when testing for minimal residual disease or recurrence, the DNA sample obtained from the patient is cell-free DNA. This cell-free DNA can be collected from the patient at any time after treatment. In some embodiments, this cell-free DNA can be collected at the point when any remaining ctDNA from cancer has disappeared if cancer treatment is successful. This time point can depend on factors such as the initial amount of ctDNA and the treatment method. In methods such as surgery in which all tumors are removed at once, the time point can be 1 week, 2 weeks, 3 weeks, or 4 weeks after treatment with curative intent. If treatment gradually removes cancer, these time points can be longer, such as 1 month or 2 months.
[0118] In some embodiments, the method can be used in clinical trials. For example, the methods described herein can be used to identify specific patient groups for clinical enrollment or to evaluate the effectiveness of new drugs (e.g., neoadjuvant or adjuvant therapy, or any combination therapy, which may be non-specific or targeted to the patient's cancer). In some embodiments, the amount of ctDNA in a patient's bloodstream can be estimated at multiple time points, which allows, for example, varying the dose of a drug administered during the patient's trial. In some embodiments, the amount of ctDNA in a patient's bloodstream can be estimated at multiple time points during a clinical trial and used to determine whether a particular treatment, level of treatment, duration of treatment, or type of treatment is working in combination with the patient.
[0119] As can be easily understood, many steps described herein, such as processing sequence and generating a report indicating the presence of cancer DNA in the test sample of DNA, can be implemented by computer.Therefore, in some embodiments, the method can include: based on the analysis of sequence reads, execute an algorithm to calculate the likelihood that a patient has cancer DNA present in the test sample of DNA collected from the patient, and output the likelihood.In some embodiments, the method can include: inputting sequence into computer, and executing an algorithm that can use the input calculation to calculate the likelihood.
[0120] As will be apparent, the described computational steps may be computer-implemented, and thus the instructions for performing the steps may be represented as programming that may be recorded on a suitable physical computer-readable recording medium. The sequencing reads may be analyzed by a computer.
[0121] The methods disclosed herein may be computer-implemented methods, i.e., methods performed by or executed on a computer. The present disclosure also provides one or more computer-readable recording media having recorded thereon instructions for performing the methods disclosed herein. The one or more computer-readable recording media may implement the above methods when executed on a computing device. The present disclosure also provides a system including one or more computer-readable media, i.e., a memory that records instructions and data units (the data units optionally including one or more error probability distribution models) for performing the methods, and a processor for executing the instructions.
[0122] An exemplary implementation of a computer system 500 that may be used in combination with any of the embodiments of the disclosure provided herein is shown in FIG. 5. The computer system 500 may include one or more processors 510 and one or more articles of manufacture including non-transitory computer-readable storage media (e.g., memory 520 and one or more non-volatile storage media 530). The processor 510 may control the writing of data to and reading of data from the memory 520 and the non-volatile storage device 530 in any suitable manner, and the aspects of the disclosure provided herein are not limited in this respect. To perform all of the functionality described herein, the processor 510 may execute one or more processor-executable instructions stored on one or more non-transitory computer-readable storage media (e.g., memory 520), which may function as non-transitory computer-readable storage media storing the processor-executable instructions for execution by the processor 510.
[0123] The terms "program" or "software" are used herein in a generic sense to refer to any type of computer code or set of processor-executable instructions that may be used to program a computer or other processor to implement various aspects of the embodiments described above. Furthermore, it is understood that, according to one aspect, one or more computer programs that, when executed, perform the methods provided herein need not reside on one computer or processor, but may be modularly distributed among different computers or processors to implement various aspects provided herein.
[0124] Processor-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0125] Additionally, the data structure may be recorded on one or more non-transitory computer-readable recording media in any suitable form. For simplicity of explanation, the data structure may be depicted as having multiple fields, with the fields related by location within the data structure. Such association may similarly be achieved by assigning records to the fields with locations within the non-transitory computer-readable media that convey the relationship between the fields. However, any suitable mechanism may be used to establish relationships between information within fields of the data structure, for example, using pointers, tags, or other mechanisms that establish relationships between data elements.
[0126] Also, various inventive concepts may be embodied as one or more processes, examples of which are provided, for example, with reference to Figure 1. The acts performed as part of each process may be ordered in any suitable manner. Thus, embodiments may be constructed in which acts are performed in an order different from that described, and such embodiments may include performing some acts simultaneously even though shown as sequential acts in the exemplary embodiment. [Example]
[0127] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use embodiments of the present disclosure, and are not intended to limit the scope of what the inventors consider to be their disclosure.
[0128] Example 1 Tumor-informed assays to best call whether cancer DNA is present should maximize all available evidence. Typical cancers have many single-nucleotide variants present throughout the genome. In many cases, there are over 10,000 (see Figure 9). Individually, these SNVs, if targeted and observed in test samples during sequencing, provide some evidence of ctDNA, but there is a risk that any signal is actually an error, such as DNA damage, polymerase error, or sequencer error.
[0129] If two variants (e.g., two SNVs) are close together and on the same DNA strand (e.g., a cell-free DNA molecule), and both variants are observed in one sequencing read or multiple sequencing reads, this provides much more evidence that the sequencing signal is not a false positive (such as DNA damage, polymerase error, or sequencer error) and that cancer is in fact present. This is because the chance of getting two errors that match two true variants on one DNA molecule or sequencing read is very small.
[0130] These variants on the same DNA molecule are far more informative than any individual SNV, and there are far fewer of them: typically, in many solid cancers, there are 10-1000 of these phase variants (see Figure 9).
[0131] A typical cancer genome is populated with additional alterations, including, for example, germline alterations, epigenetic alterations, and structural variants. Again, for example, some cancers have many structural variants, while others have very few (see Figure 10).
[0132] Therefore, an assay that can utilize all of this information will always have better performance than, for example, an assay that only looks at individual SNVs or only looks at phase variants.The challenge is how to combine this information.In target region-based approaches, the genome is cut into sections, then target regions are selected (for example, the target region is the sequence between two primers in PCR products), and then each target region is evaluated based on all the genetic variants within it, and when sequenced, all of this information is combined together to determine whether cancer DNA is present.
[0133] Such assays are also much more universal. For example, assays relying solely on phase variants may provide very little information for targeting in some cases. For example, some osteosarcoma patients in Figure 9 have <10 phase variants. While the same patient will always have multiple SNVs, osteosarcoma patients also typically have multiple structural variants (Figure 10). A targeted region-based calling approach that combines information allows for consistent and sensitive cancer DNA detection across patients.
[0134] Example 2 To design an optimal MRD assay, the system is designed to examine data from as many high-quality regions as possible. Regions are considered higher quality if they are easier to distinguish from noise, for example, in the presence of cancer (e.g., by having two-phase SNVs), and are easier to amplify and sequence than other regions. To do this, tumor biopsies are first obtained, macrodissected to a target of 50% tumor content, exome capture is performed, and the sample is then sequenced using an Illumina sequencer. All possible variants are identified using the standard Illumina pipeline, and each variant is then assigned a combined score based on: 1) likelihood of being true, 2) likelihood of being somatic, 3) variant background error rate, 4) high-signal background error rate, 5) probability of being clonal, and 6) level of variant amplification or copy number gain. The genome is divided into 50-bp windows with a 25-bp overlap. Each window is assigned a combined score that includes: 1) the scores of all variants present within the window; 2) a score for the ability to uniquely align the region (where regions that are not uniquely aligned are penalized, with larger penalties corresponding to more misalignments); and 3) a score for the ability to amplify and sequence the region (where features known to make sequencing suspicious, including repeats, are penalized). The regions are then sorted by score, and the top 100 are selected for PCR primer design. If two overlapping regions are in the top 100 list, the region with the highest score is kept, and the region with the smaller score is discarded. The 101st region is then added to the list, and this process is repeated. A multiplex PCR is designed for the top 48 regions. In silico PCR is performed using all primer pairs. If a primer combination that produces two or more non-specific regions is identified, the primer for the lowest-scoring region that produces this non-specific product is discarded, and alternative primers are designed.If the problem of non-specific PCR is not overcome, the region is discarded and the next region is added to the primer design.
[0135] One of the challenges in this tumor information-based method for detecting cancer DNA in test samples is the number of regions that can be robustly and economically targeted.This strategy of ranking regions can maximize the number and quality of regions that can be successfully interrogated in test DNA samples.If variants are phase variants (PV), that is, if variants are cis, adjacent to each other, and on the same chromosome, they can be read together, which increases the ability to distinguish signal from noise.If variants are trans but can still be read with the same primer pair (or other targeting reagent such as bait), the amount of information obtained from one target region can be doubled.This approach also limits the number of reads that are discarded due to non-specific products.
[0136] Example 3 In order to detect cancer DNA in test samples with high sensitivity, it is advantageous to target multiple regions and multiple region classes. In some cancer types, targeting only one region class is sufficient. However, the inventors recognize and understand that it is better to target multiple target region classes with different types of genetic variations. In this example, it is identified that certain breast cancer patients have a large number of structural variants (SVs), while other patients have more SNVs and INDELs. Furthermore, many patients have regions with multiple somatic genetic and epigenetic changes. Furthermore, these patients also have many regions with both somatic and germline alterations. A larger panel is designed to sequence breast cancer tumor DNA and evaluate somatic SNVs, INDELs, and SVs, as well as germline alterations. Optimal target regions are identified according to the methods disclosed herein. Primers are designed to target these regions. If the target region has one or more SNVs or INDELs, primers are designed to flank all of the SNVs and INDELs. If the target region is identified as having a rearrangement (e.g., SV), two different parts of the same chromosome or two different chromosomes are combined. The rearrangement sequence is used to design primers, one 3' and one 5' from the rearrangement. If an SNV, INDEL, or other genetic variant (e.g., EV) is in cis with the rearrangement, primers are designed to flank both the rearrangement and the other variant using the rearrangement sequence obtained from the tumor. If the target region is identified as having a pair of phase variants (PV), primers are designed to flank the 5' and 3' PV. The advantage of this approach is the ability to consistently obtain multiple regions for evaluation of cancer DNA in test samples.
[0137] Because each of these different classes is incorporated, a different error model may be required for each type. The inventors have realized that the results from such different error models can be combined in a reasonable manner, for example by summing the log-likelihoods of each target region class, thereby providing a reliable conclusion about the cancer DNA in the test sample.
[0138] Example 4 Figures 6A-B illustrate why it can be difficult to call a sample containing cancer DNA, especially in test samples with low tumor fraction. As shown in Figure 6A (top), in test samples with high relative tumor fraction (TF), most, if not all, target regions have multiple cancer DNA molecules, resulting in high signals (e.g., likelihood) across multiple target regions, so cancer DNA can be easily called, thus eliminating most false positives and false negatives. As shown in Figure 6B (bottom), samples with low tumor fraction are even more difficult to call because the data for individual regions may not be sufficiently distinguishable from the background error rate. Furthermore, input DNA at such low levels may not have cancer DNA in some target regions, and therefore, in many target regions, once amplified, will not produce a true signal and will result in a true negative (see the white boxes in Figure 6B). For example, if an assay tests for multiple SNVs, and some SNVs are present in only one cancer DNA molecule but not others, it is not possible to call any one region positive, but by combining information across different target regions (i.e., SNVs), the likelihood of making a confident call increases. In Figure 6B, the region with two SNVs each has one mutant molecule (see the light gray boxes in Figure 6B). In isolation, none of these SNVs, which show a small signal, provide sufficient evidence to make a positive call. Combined, they provide more evidence, but in this example, they still do not provide sufficient support to make a confident call on the overall cancer DNA. By including not only target regions with a single SNV, but also target regions with multiple phase variants and other variants such as MNVs or INDELs, the likelihood of cancer DNA being detected, if present, is much higher. In Figure 6B, three regions with two PVs are present. In two of these regions, no cancer DNA is present and there is no signal (see white boxes in Figure 6B for two PVs).However, in the third target region with two PVs, one cancer DNA molecule is present (see the dark gray box in Figure 6B for two PVs). In this case, even a single cancer molecule read can provide a large amount of information, showing that two phase variants are found in the same sequencing read, providing much more evidence (dark gray box in Figure 6B). Combining information from this two-PV target region with information from two target regions with individual SNVs (light gray box) allows for much more sensitive calls. In this example, the region with indels provides further supporting information (2 / 3 indels are present—dark gray and white boxes). Still, to utilize this information and confidently determine whether cancer DNA is present, an approach that combines information from different regions, as described in this invention, is required.
[0139] Example 5 Figure 7 shows one embodiment of how evidence can be combined across multiple regions. In diluted samples (<<0.1% tumor fraction), the proportion of mutant reads in each individual target region for each sample is not expected to approximate the overall tumor fraction due to the effects of dropout. For example, many target regions exhibit zero variant molecules. Conversely, the effect of viewing n / input reads as a discrete distribution is modeled. In this example, the tumor fraction is not measured directly. Rather, the tumor fraction is unweighted across all possible inputs, providing an accurate estimate of the tumor fraction of the sample. Specifically, instead of estimating the number of cancer DNA molecules with variants in all target regions, the likelihood of all possible values is calculated based on (i) the number of sequence reads with genetic or epigenetic variations or a combination of multiple expected genetic variations in the target region (which varies for each target region), (ii) the total number of sequencing reads, (iii) the number of input DNA molecules, and (iv) the estimated background error rate for each target region class, from which the most probable value is identified. This eliminates the need for assumptions. As shown in Figure 6, the inclusion of several PV regions (including three PV regions) and several INDEL regions, when considered together according to any method of the present disclosure, provides a large amount of evidence that can support a reliable conclusion that cancer DNA is present in a test sample, for example, by comparing the amount of sequence reads in the target regions with different error models for each target region class. The accuracy of the mathematical model can be verified by comparing it with the true dilution data in the ground truth line plot (Figure 8).
[0140] Although the present method has been described in more detail, it is understood that the present disclosure is not limited to the specific embodiments described, and therefore may, of course, vary. It is also understood that the terminology used herein is for the purpose of describing specific embodiments only, and is not intended to be limiting. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, preferred methods and materials are described.
[0141] All publications and patents cited in this specification are incorporated herein by reference to disclose and describe the methods and / or materials for which the publications are cited, as if each individual publication or patent was specifically and individually indicated to be incorporated by reference.
[0142] As used herein, and in the appended claims, the singular forms "a," "an," and "the" should be understood to include plural references unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement serves as a predicate for the use of exclusive terminology such as "solely," "only," and the like, or the use of a "negative" limitation in connection with the recitation of claim elements.
[0143] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features which may be readily separated from, or combined with, the features of all other embodiments without departing from the scope or spirit of the disclosure. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.
[0144] Although several embodiments of the technology described herein have been described in detail, various modifications and improvements will be readily apparent to those skilled in the art. Such modifications and improvements are within the spirit and scope of the present disclosure. Accordingly, the foregoing description is illustrative only and is not intended to be limiting. The technology is limited only as defined by the following claims and equivalents thereof.
Claims
1. 1. A method for detecting cancer DNA in a test sample obtained from a patient, comprising: (a) enriching or having enriched a test sample for a plurality of target regions, the plurality of target regions comprising a first target region having a first class and a second target region having a second class; (b) measuring or measuring the plurality of target regions of step (a) in the enriched test sample; (c) for each of the first and second target regions, comparing or comparing the measurements of step (b) that support the presence of a target region class to one or more error models that model the probability that the target region class is observed in DNA that does not have that target region class; (d) combining or combining the comparisons of step (c) for at least the first target region and the second target region; and (e) identifying or having identified cancer DNA in the test sample based on the combined comparison of step (d).
2. 2. The method of claim 1, wherein the enriching in step (a) comprises amplifying a plurality of target regions by polymerase chain reaction (PCR) to obtain PCR products, and the measuring in step (b) comprises sequencing the PCR products or their progeny to obtain a plurality of sequence reads.
3. 2. The method of claim 1, wherein the enriching in step (a) comprises contacting the test sample with a pool of oligonucleotides, the pool of oligonucleotides comprising oligonucleotides substantially complementary to a plurality of target regions.
4. 4. The method of any one of claims 1 to 3, wherein the cancer is a solid tumor and the multiple target regions are identified by sequencing the solid tumor.
5. 5. The method of claim 4, wherein the sequencing of the solid tumor is performed by targeted sequencing or whole exome sequencing (WES).
6. (i) in step (b), measuring includes sequencing the plurality of target regions of step (a) to obtain a plurality of sequence reads corresponding to the first target region and the second target region; and (ii) In step (c), the comparison comprises comparing the amount of sequence reads supporting the presence of a target region class to one or more error models that model the probability that the target region class is observed in DNA or RNA that does not have the target region class.
7. 7. The method of claim 1, wherein in step (c), the comparison comprises comparing the amount of sequence reads that do not support the presence of the target region with one or more error models that model the probability that the target region class is observed in DNA or RNA that does not have the target region class.
8. 8. The method of claim 1, wherein the one or more error models are based on a background error rate in each of the first target area class and the second target area class.
9. 9. The method of claim 1, further comprising training one or more error models based on a set of control samples.
10. 10. The method of claim 1, wherein the one or more error models in step (c) comprise a first error model for the first target region and a second error model for the second target region.
11. 11. The method of claim 1, wherein the first error model for the first target region comprises a beta-binomial model and the second error model for the second target region comprises a multivariate beta-binomial distribution.
12. 12. The method of claim 11, wherein the multivariate beta-binomial distribution is a standard Dirichlet distribution or a generalized Dirichlet distribution.
13. 13. The method of claim 1, wherein the target region class is related to the type of genetic variation within the target region.
14. 14. The method of any one of claims 1 to 13, wherein a first class of first target regions are regions with a single genetic variation and a second class of second target regions are regions with two or more genetic variations.
15. 15. The method of claim 14, wherein the first class of the first target region is a single nucleotide variant (SNV) and the second class of the second target region comprises a first phase variant (PV) and a second PV.
16. 15. The method of claim 14, wherein the single genetic variation is a single nucleotide variant (SNV) and the two or more genetic variations comprise a tumor SNV and a germline SNV.
17. (i) the comparison in step (c) for the first target region includes comparing the amount of sequence reads having a single gene variation and the total amount of sequence reads for the first target region to a first error model; and 17. The method of any one of claims 14 to 16, wherein (ii) the comparison in step (c) for the second target region comprises comparing the amount of sequence reads having two or more genetic variations and the total amount of sequence reads for the second target region to a second error model.
18. 18. The method of any one of claims 14 to 17, wherein the first error model comprises an error probability distribution that models the probability of a single genetic variation being observed in DNA that does not have a single genetic variation, and the second error model comprises an error probability distribution that models the probability of two or more genetic variations being observed in DNA that does not have two or more genetic variations.
19. 19. The method of any one of claims 14 to 18, wherein the two or more genetic variations are located within 160 bp of each other.
20. 20. The method of any one of claims 14 to 19, wherein the two or more genetic variations are separated by at least one nucleotide.
21. 21. The method of any one of claims 14 to 20, wherein the one or more error models take into account the distance between two or more genetic variations.
22. The comparison in step (c) for the second target region determines the amount of sequence reads that have both the first PV and the second PV (k 1 ), the amount of sequence reads with only the first PV (k 2 ), the amount of sequence reads with only the second PV (k 3 ), and the amount of sequence reads that have neither the first nor the second PV (k 4 22. The method of claim 15, comprising comparing the signal strength of the signal to one or more error models.
23. 23. The method of any one of claims 1 to 22, wherein the comparison of step (c) comprises likelihood or log-likelihood, and the combination of step (d) comprises combining the comparison of step (c) for the first target region and the comparison of step (c) for the second target region.
24. 24. The method of any one of claims 1 to 23, further comprising calculating a variant allele fraction (VAF) at each of the first and second target regions based on the measurements in step (b).
25. 25. The method of any one of claims 1 to 24, further comprising the step (f) of determining whether cancer DNA is present in the test sample.
26. 26. The method of any one of claims 1 to 25, further comprising providing a report.
27. 27. The method of any one of claims 1 to 26, further comprising treating the patient based on the identification of cancer DNA in the test sample of step (e) or the determination of step (f).
28. A. Obtaining a second test sample from the patient at a second time point; B. Enriching the second test sample for multiple target regions; C. Measuring the plurality of target regions of step B from the enriched second test sample; D. For each of the first and second target regions, comparing the measurements of step C that support the presence of a target region class to one or more error models that model the probability that the target region class is observed in DNA that does not have that target region class; E. Combining the comparisons of step D for at least the first target region and the second target region; 27. The method of any one of claims 1-26, further comprising the step of: F. identifying cancer DNA in the second test sample based on the combined comparison of step E.
29. 29. The method of claim 28, wherein the method of steps A to F comprises the additional features of any one of claims 2 to 27 as applied to steps A to F.
30. 30. The method of claim 28 or claim 29, further comprising administering a cancer treatment or therapy to the patient prior to obtaining the test sample, and determining the effectiveness of the cancer treatment or therapy based on determining whether cancer DNA is present in the second test sample and / or whether the level of cancer DNA is altered.