Method of detecting cancer DNA in sample

By enriching test samples for multiple target areas and combining different types and amounts of genetic variation evidence, using error models for comparison, the problem of insufficient sensitivity to detecting micro-residual diseases in the prior art is solved, and earlier detection of cancer recurrence risk is achieved.

CN119948176APending Publication Date: 2025-05-06INIVATA LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069281.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-19
Filing Date
2023-08-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is inadequate in detecting circulating tumor DNA (ctDNA) in micro-residual diseases (MRDs), making it difficult to detect the risk of cancer recurrence early.

Method used

By enriching the test samples against multiple target regions and combining evidence from different types and quantities of genetic variation, an error model was used to identify cancer DNA present in the test samples.

Benefits of technology

It improves the detection sensitivity of micro-residual diseases, can detect the risk of cancer recurrence earlier, and reduces the possibility of false positive results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005331300500000231
    Figure BDA0005331300500000231
  • Figure BDA0005331300500000241
    Figure BDA0005331300500000241
  • Figure BDA0005331300500000251
    Figure BDA0005331300500000251
Patent Text Reader

Abstract

In one embodiment, a method may include enriching a test sample for a plurality of target regions, where the plurality of target regions includes a first target region having a first category and a second target region having a second category. A plurality of target regions may be measured, and for each of a first target region and a second target region, the measurements supporting a category of the target region may be compared to an error model that models a probability that the category of the target region is observed in DNA of the category that does not contain the target region. These comparison results may then be combined for at least a first target region and a second target region. The cancer DNA may then be identified in the test sample based on the combined comparison results.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims the benefit of United Kingdom patent application serial number GB2212094.3 filed on August 19, 2022, which is incorporated herein by reference for all purposes.

[0003] field

[0004] The present disclosure relates generally to the field of liquid biopsy, e.g., diagnosing the presence or absence of cancer from a blood or other liquid sample from a patient.

[0005] background

[0006] Detecting and monitoring circulating tumor DNA (ctDNA) is rapidly becoming an established diagnostic, prognostic, and predictive tool in the care of cancer patients. Often, after cancer treatment, a small number of cancer cells may remain in patients who appear to be in remission. These residual cells are often referred to as "minimal residual disease" (MRD) or residual disease. These residual cells will ultimately be the cause of many cancer relapses. It is critical to determine a patient's likelihood of disease recurrence and relapse after initial treatment so that those most likely to require additional treatment can receive it and those who do not can be spared, thereby reducing harm to patients and lowering treatment costs. Therefore, effective methods for detecting minimal residual disease are highly desirable. It is also critical to have sensitive methods to detect the risk of cancer recurrence earlier than current methods (e.g., which are often done through imaging or clinical analysis).

[0007] MRD has been successfully detected in some hematological malignancies because relatively large amounts of DNA can be analyzed and the frequency of common tumor-specific fusions can be measured in a straightforward manner. There is now strong evidence that MRD can be detected in many solid tumors by evaluating cell-free DNA (cfDNA) against circulating tumor DNA (ctDNA). However, the problem with detecting minimal residual disease in cfDNA is that many of the tests used to detect sequence variants in the sample are not sensitive enough. Many of today's molecular tests are done by sequencing cfDNA against a panel of known genes. The problem with detecting minimal residual disease by sequencing cfDNA is that the amount of tumor DNA in cell-free DNA is often well below the detection limit of such methods. Specifically, the frequency of individual tumor sequence variants expected to occur in cfDNA of patients with minimal residual disease is often much lower than the frequency of sequencing artifacts generated by PCR errors, base miscalling, and / or DNA damage. This problem is compounded by the fact that in some cases, the level of tumor DNA may be so low that, on average, less than one copy of each mutation being evaluated is present in the cfDNA sample analyzed. In addition, the relatively small amount of mutant DNA derived from lysed white blood cells in the blood may lead to erroneous results. Therefore, detection of minimal residual disease by sequencing-based approaches remains challenging.

[0008] Assays for detecting MRD can employ a variety of approaches, including sequencing a patient’s tumor tissue to identify tumor-specific genetic variants. These variants can include single nucleotide variants, small insertions and deletions, double-base substitutions, and larger structural changes. In theory, identifying these tumor-specific variants in a patient’s cfDNA sample should be indicative of MRD. However, as the number of tumor-specific variants in an assay increases, the likelihood of false-positive results may increase. In addition, different types of variants may produce more or less false-positive results. Therefore, improved ctDNA testing is needed.

[0009] Overview

[0010] This paper describes a method for detecting cancer DNA in a DNA test sample from a patient. In some embodiments, the method may include enriching a test sample for multiple target areas, wherein the multiple target areas include a first target area with a first category and a second target area with a second category. Multiple target areas can be measured, and for each of the first target area and the second target area, the measured value of the category supporting the target area can be compared with an error model, which models the probability of observing the category of the target area in the DNA of the category that does not include the target area. These comparisons can then be combined for at least the first target area and the second target area. Cancer DNA can then be identified in a test sample based on the comparison of the combination.

[0011] In one example, the methods described herein are derived from such recognition: by combining evidence from multiple target regions with different kinds of genetic variations and different numbers of genetic variations, the problem of low sensitivity when determining the presence of cancer DNA in a patient's test sample can be solved. The observations from each target region provide some evidence that can be combined to support a high-confidence conclusion that the test sample contains cancer DNA, so the patient suffers from cancer or residual disease. In addition, the methods described herein can combine evidence from multiple categories of target regions, such as single nucleotide variations, polynucleotide variations (such as duplex and triplex base substitutions), short insertions or deletions, copy number variations, structural variations (SVs), multiple genetic variations, and multiple phase variations (i.e., wherein the target region has two or more variations all on the same chromosome in the target region). Each of these categories can provide different degrees of support or confidence about the presence (or absence) of cancer, and therefore must be combined in a principled manner, as further described herein.

[0012] These and other advantages may become apparent in view of the following discussion.

[0013] BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Those skilled in the art will appreciate that the drawings described below are for illustration purposes only and are not intended to limit the scope of the present teachings in any way.

[0015] Figure 1 is a flow chart depicting an embodiment of a method of detecting cancer DNA in a DNA test sample from a patient.

[0016] Figure 2A Enhanced Tag-Amplicon Sequencing (eTAm-Seq TM ) is a diagram of an embodiment of the method, wherein the target region is amplified by polymerase chain reaction (PCR).

[0017] Figure 2B Several exemplary categories of target regions are depicted.

[0018] Figure 3 is a diagram depicting an exemplary assay for a target region comprising a genetic variation according to an embodiment of the present disclosure.

[0019] Figure 4A -B shows an example of error probability distribution according to an embodiment of the present disclosure. Figure 4A In the model shown, data corresponding to low-frequency high-signal events are shaded. Figure 4B Two models are shown, one for background noise and the other for DNA damage. "VAF" refers to variant allele fraction. Such models can be obtained from DNA that does not contain genetic variation, and they indicate the probability of different variant allele fractions in this non-cancerous DNA (or the number of variant reads in the total reads).

[0020] Figure 5 is a block diagram of an illustrative computer system that can be used to implement some embodiments of the techniques described herein.

[0021] Figure 6A-6B Embodiments of the assays described herein are depicted and illustrate some of the difficulties of detecting cancer DNA by scoring each target region to determine whether it contains a particular genetic variation.

[0022] Figure 7 Some principles of embodiments of the present method are schematically illustrated.

[0023] Figure 8 shows how the fraction of cancer DNA can be calculated by comparing the true dilution data with a mathematical model.

[0024] Fig. 9 This is an adaptation of a figure from Kurtz et al. (Nat Biotechnol 2021 39:111), in which the authors show that the genome has a small amount of phase variation (part b).

[0025] Fig.10 This is a figure adapted from Li et al. (Nature 2020 578, 112-121), which shows the number of structural variants and the range of structural variant types in different types of cancer. Some types of cancer typically have a large number of structural variants, such as breast cancer and certain sarcomas, while other types of cancer, such as CLL, typically have a smaller number.

[0026] definition

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. However, definitions of certain elements are provided for clarity and ease of reference.

[0028] The terms and symbols of nucleic acid chemistry, biochemistry, genetics, and molecular biology used in this article follow standard treatises and textbooks in the field, such as Kornberg and Baker, DNA Replication, Second Edition (WH Freeman, New York, 1992); Lehninger, Biochemistry, Eighth Edition (Worth Publishers, New York, 2021); Strachan and Read, Human Molecular Genetics, Fifth Edition (Wiley-Liss, New York, 2018); Eckstein, editor, Oligonucleotides and Analogs: A Practical Approach (Oxford University Press, New York, 1992); Gait, editor, Oligonucleotide Synthesis: A Practical Approach (IRL Press, Oxford, 1984); (the contents of which are incorporated herein by reference in their entirety), etc.

[0029] As used herein, the term "determining" can mean indicating whether a particular genetic variation is present in a sequence, whether a sample contains a genetic variation, or whether a sample contains cancer DNA, depending on the context.

[0030] If two nucleic acids are "complementary", they will hybridize to each other under high stringency conditions. The term "fully complementary" describes a duplex in which every base of one nucleic acid is base paired with a complementary nucleotide in the other nucleic acid. In many cases, two complementary sequences have at least 10, e.g., at least 12 or 15 complementary nucleotides.

[0031] As used herein, the term "detecting recurrence" refers to detecting the recurrence of a tumor by identifying cancer DNA. In this context, the term "early detection" refers to reliably detecting mutant DNA by conventional standard care / monitoring monitoring methods (e.g., radiographic imaging, etc.) before cancer relapses. This can be achieved, for example, by continuously monitoring blood samples collected for the presence of ctDNA in cfDNA at multiple time points, as described below.

[0032] The terms "determine", "measure", "assess", "evaluate" and "measure" are used interchangeably and include both quantitative and qualitative measurements. Assessments can be relative or absolute.

[0033] The term "genetic variation" used herein refers to a variation (such as nucleotide substitution, insertion / deletion or rearrangement) that exists or is thought to be present in a test sample. Genetic variation can come from any source. For example, genetic variation can be caused by mutation (such as somatic mutation), or it can be a germline, such as a mutation derived from a germ cell, which is integrated into the DNA of each cell of the body. If the sequence variation is referred to as a genetic variation, the determination indicates that the sample may contain the variation; but in some cases, the "determination" may not be correct. In many cases, the term "genetic variation" can be replaced by the term "mutation". For example, if a method is used to detect sequence variation associated with cancer or other diseases caused by mutations, then "genetic variation" can be replaced by the term "mutation".

[0034] As used herein, the term "minimal residual disease" (MRD) refers to the presence of cancer cells after treatment with curative intent. In some publications, MRD may also be referred to as "molecular residual disease" or "residual disease."

[0035] The terms "nucleic acid," "oligonucleotide," and "polynucleotide" are used interchangeably herein to describe a polymer of any length composed of nucleotides (e.g., deoxyribonucleotides or ribonucleotides), for example, a polymer of greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, greater than 10,000 bases, greater than 100,000 bases, greater than about 1,000,000 bases, up to about 10 10 or more bases.

[0036] The terms "plurality," "group," and "collection" are used interchangeably to refer to something that contains at least 2 members. In some cases, a plurality, group, or collection may have at least 5, at least 10, at least 100, at least 1,000, at least 10,000, at least 100,000, at least 10 6 At least 10 7 At least 10 8 or at least 10 9 or more members.

[0037] The term "reference sequence" as used herein is a reference sequence from a reference genome or a sequence from a patient sample (e.g., a buccal swab) that is not expected to contain somatic variation. The reference sequence corresponds to a sequence that contains or may be suspected of containing a sequence variation (e.g., a target sequence), thereby enabling the presence (or absence) of the sequence variation to be determined by comparing the sequence that contains or may be suspected of containing a sequence variation (e.g., a target sequence) to the reference sequence. The reference sequence differs from the sequence that contains or may be suspected of containing a sequence variation (e.g., a target sequence) only in the sequence variation itself, because the reference sequence and the sequence that contains or may be suspected of containing a sequence variation (e.g., a target sequence) are derived from the same genomic location.

[0038] The term "reference genome" as used herein may refer to a single genome, a collection of genomes, or a consensus genome. A reference genome may be from one or more publicly available databases. A reference genome is used to determine the position of the sequence being analyzed in the genome of an organism. Those skilled in the art will appreciate that a consensus genome is a genome constructed from multiple genomes from the same species.

[0039] The term "sequence variation" as used herein is a variation different from an expected sequence or reference sequence, such as a reference genome or a sequence from a patient sample (such as a buccal swab) that is not expected to contain somatic variation. Sequence variation can refer to a combination of position and sequence change type. For example, sequence variation can be referred to by the position of the variation and which type of substitution (for example, G to A, G to T, G to C, A to G, etc., or insertion / deletion of G, A, T or C, etc.) exists at that position. Sequence variation can be a substitution, deletion, insertion rearrangement of one or more nucleotides. In the context of the present method, sequence variation can be produced by, for example, PCR errors, sequencing errors or genetic variation. In many cases, sequence variation is a variation that exists at a frequency of less than 50% relative to other molecules in the sample. Many sequence variations (such as insertions and deletions and nucleotide substitutions) are substantially the same as molecules that do not contain sequence variation. In some cases, a particular sequence variation may be present in a sample at a frequency of less than 20%, less than 10%, less than 5%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.001%, or less than 0.0001%.

[0040] The term "substantially" refers to sequences that are approximately duplicates as measured by a similarity function (including, but not limited to, Hamming distance, Levenstein distance, Jaccard distance, cosine distance, etc. (see generally Kemena et al., Bioinformatics 2009 25:2455-65, the contents of which are hereby incorporated by reference in their entirety). The exact threshold depends on the error rate of sample preparation and sequencing used to perform the analysis, with the higher the error rate, the lower the similarity threshold required. In some cases, substantially identical sequences have at least 98% or at least 99% sequence identity.

[0041] As used herein, the term "threshold" refers to the level of evidence (eg, a ratio or a set amount) required to make a determination.

[0042] As used herein, the term "value" refers to a number, letter, word (e.g., "high," "medium," or "low"), or descriptor (e.g., "+++" or "++") that can indicate the strength of the evidence. A value can contain one component (e.g., a single number) or multiple components, depending on how the value is analyzed.

[0043] Other definitions of terms may appear throughout this specification. It should also be noted that claims may be drafted to exclude any optional elements. Therefore, this statement is intended to serve as a prerequisite basis for use of the terminology "solely," "only," and the like when reciting claim elements or using a "negative" limitation.

[0044] Detailed description

[0045] The methods described herein are derived from the recognition that cancer DNA can be reliably detected by incorporating and combining evidence from various target regions containing one or more genetic variations (including single nucleotide variations, multi-nucleotide variations, copy number variations, short insertions and deletions, epigenetic variations, structural variations, and phase variations) into an assay for the target region. These methods can be used for enrichment-based approaches to examine target regions of the genome.

[0046] Figure 1An embodiment of a method 100 for detecting cancer DNA in a test sample (e.g., a blood sample) collected from a cancer patient is depicted. In this embodiment, the method 100 may include (a) enriching a plurality of target regions from the test sample (step 102). The plurality of target regions may include a first target region comprising a first category and a second target region comprising a second category. The method 100 may continue by (b) measuring the plurality of target regions in the test sample (step 104). For each of the first target region and the second target region, the method 100 may continue by (c) comparing the measurement results supporting the category of the target region with an error model that models the probability of observing the category of the target region in DNA that does not contain the category of the target region (step 106). The method 100 may continue by (d) combining comparisons for at least the first target region and the second target region (step 108). The method 100 may continue by (e) identifying cancer DNA in the test sample based on a comparison for the combination of the first target region and the second target region (step 110).

[0047] The methods disclosed herein may include the step of obtaining a test sample from a patient. Alternatively, the test sample may be previously obtained from a patient. The test sample may include any nucleic acid sample or fluid containing DNA, RNA, or cDNA. A genomic DNA sample from a mammal (e.g., mouse or human) is a type of test sample. The test sample may have more than about 10 4 , 10 5 , 10 6 or 10 7 , 10 8 , 10 9 or 10 10 Different nucleic acid molecules. Any sample containing nucleic acid can be used herein, such as genomic DNA or RNA from tissue culture cells or tissue samples. In addition, although many embodiments describe the detection or use of cancer DNA, the method according to the present disclosure can be applied to any form of nucleic acid, so the use of "cancer DNA" can also refer to RNA or other detectable nucleic acids associated with cancer. The test sample can include plasma, serum, cerebrospinal fluid, urine, saliva, feces, amniotic fluid, aqueous humor, bile, milk, cerumen, chyle, exudate, gastric juice, lymph, mucus, pericardial fluid, peritoneal fluid, pleural fluid, pus, sebum, serous fluid, semen, sputum, synovial fluid, sweat, tears, vomitus or whole blood.

[0048] In some embodiments, the test sample includes cell-free DNA (cfDNA), i.e., DNA that is free in body fluids and not contained in cells. cfDNA can be obtained by centrifuging the test sample to remove all cells and then separating DNA from the remaining liquid (e.g., plasma or serum). Such methods are well known (see, e.g., Lo et al., Am J Hum Genet 1998; 62: 768-75). Circulating cell-free DNA can be double-stranded or single-stranded. The term cfDNA is intended to cover free DNA molecules circulating in the bloodstream and DNA molecules present in extracellular vesicles (e.g., exosomes) circulating in the bloodstream. Cell-free DNA may contain cancer DNA, i.e., DNA from cancer cells. Cancer DNA from solid tumors can be found in cfDNA, in which case it can be referred to as tumor DNA (tDNA) or circulating tumor DNA (ctDNA). Cancer DNA can be identified because it contains mutations. In a preferred embodiment, the test sample is cell-free DNA (circulating cell-free DNA) from the bloodstream, which is DNA circulating in the peripheral blood of the patient.

[0049] In some embodiments, the test sample includes cancer DNA isolated directly from a tissue biopsy, circulating tumor cells (CTCs), or other cells that are no longer part of the tumor tissue but are not circulating (e.g., cells in urine or fecal samples). In some embodiments, the test sample may include DNA isolated from cells such as bone marrow cells, cells from lymph nodes, or circulating leukocytes (in the case of blood cancer or cells from lymph nodes), from tumor margins, or other sample types (e.g., cerebrospinal fluid (CSF) and whole blood) that are currently screened for the presence of cancer cells from solid tumors by other means. These cells can be obtained from a patient's tissue sample (e.g., a cancer tissue sample or a suspected cancer tissue sample or a tissue sample containing or suspected of containing cancer cells) or a liquid sample (e.g., any of the above-mentioned liquids).

[0050] DNA molecules in cell-free DNA may be highly fragmented and have a median size of less than 1 kb (e.g., in the range of 50 bp to 500 bp, 80 bp to 400 bp, or 100-1,000 bp), although fragments with median sizes outside this range may exist. Typically, the average fragment size of cfDNA is about 100-250 bp, for example, 150 to 200 bp long, or about 160 bp. ctDNA is tumor-derived and originates directly from tumors or circulating tumor cells (CTCs), which are living, intact tumor cells that are shed from primary tumors and can enter the blood or lymphatic system. The exact mechanism of how cancer DNA is released is unclear, although it is speculated that it involves apoptosis and necrosis from dying cells, or active release from living tumor cells. The amount of ctDNA in samples of circulating cell-free DNA isolated from cancer patients varies greatly: a typical sample contains less than 10% ctDNA, although many samples from patients undergoing MRD evaluation may have less than 0.01% ctDNA, while some samples have more than 10% ctDNA. Molecules of cancer DNA can often be identified because they contain tumorigenic mutations.

[0051] In one embodiment, the test sample is a plasma sample, and cell-free DNA (cfDNA) is separated from the plasma sample. The fraction of cancer DNA in the test sample (compared with non-cancer DNA) can be equal to or less than 0.01%, equal to or less than 0.002%, equal to or less than 0.005%, or equal to or less than 0.001%. In some embodiments, the detectable cancer DNA fraction in the test sample of DNA can be about 0.0001%, but the actual detection limit may be different. In some embodiments, the test sample contains less than 25,000 genome equivalents of DNA (e.g., cfDNA), such as less than 20,000, less than 10,000, less than 5,000, or less than 1,000 genome equivalents of DNA. In some embodiments, the test sample contains about 100 to about 25,000 genome equivalents (i.e., enriched or amplifiable copies) of DNA. In some embodiments, the test sample contains about 10ng to about 100ng of DNA. In some embodiments, the test sample comprises at least 10 ng, at least 20 ng, at least 30 ng, at least 40 ng, at least 50 ng, at least 60 ng, at least 70 ng, at least 80 ng, at least 90 ng, or at least 100 ng of DNA. In some embodiments, the test sample comprises 66 ng of DNA.

[0052] In addition, the methods described herein can be used to detect cancer DNA from solid tumors and hematological (blood) cancers. Therefore, the term "cancer" can refer to any disease characterized by uncontrolled cell division, and can be blood cancer, such as leukemia, lymphoma or multiple myeloma, or tumor cancer, such as related to abnormal tissue mass, in which cell growth and division exceed the degree that should be or do not die when they should die. Tumor cancer, such as lung cancer, breast cancer or liver cancer, is related to solid tumors. For solid tumor embodiments, the method can identify cancer DNA (here tumor DNA) in cfDNA (e.g., circulating cfDNA). For blood cancer embodiments, the method can identify cancer DNA in DNA extracted from cells obtained from bone marrow, lymph nodes or circulating white blood cells, or identify cancer DNA in cfDNA. For example, in blood cancer embodiments, bone marrow aspirates (pretreatment) can be taken out from AML patients, the variation in their AML can be determined (e.g., by sequencing DNA from AML cells), and the patient can be treated. At a certain time after treatment, bone marrow aspirates, cell-free DNA or urine can be checked for evidence of these variations to determine whether the patient still has cancer. In some embodiments, the method can identify cancer DNA in a tissue sample (eg, surgical margins or lymph nodes).

[0053] "Target region" or "region" refers to a DNA region containing or suspected of containing one or more genetic variations. For a genome or a target polynucleotide, such a region refers to a continuous subregion or segment of a genome or a target polynucleotide. The term refers to any continuous portion of a genomic sequence, whether it is within a gene (e.g., a coding sequence) or associated therewith. The target region can be a fragment from a single nucleotide to a length of hundreds or thousands of nucleotides or longer. Typically, the length of the target region is about or less than the average length of the nucleic acid present in the test sample. For example, in cfDNA embodiments, the target region is typically about 160bp. However, in other embodiments, the length of the target region can be about 50bp, 100bp, about 200bp, about 300bp, about 400bp, and about 500bp. For example, in some embodiments where the test sample is a tissue sample (e.g., from a lymph node or surgical margin), the length of the target region can be commensurate with the desired sequencing length or average fragment length. In practice, the target region can be any region targeted by a pair of PCR primers, so its length will be the length of the resulting amplicon.

[0054] Enriching multiple target regions (step 102) can be performed in a variety of ways, including but not limited to hybridization with nucleic acid probes, polymerase chain reaction (PCR), linked target capture, molecular inversion probes, ligation, and ATOM-Seq. In some embodiments, enrichment includes capturing multiple target regions from a test sample by contacting the test sample with a pool of oligonucleotides. For example, the oligonucleotide pool may contain oligonucleotides that include reverse complements (or substantially reverse complements) of multiple target regions. When the test sample (for example) is heated and the nucleic acid is denatured into a single strand, the oligonucleotide can bind to any target region and then be selected (for example, by a probe).

[0055] In some embodiments, enrichment includes amplifying multiple target regions by polymerase chain reaction (PCR) (i.e., an enzymatic reaction using one or more pairs of sequence-specific primers to amplify specific template DNA). As shown in Figure 2, forward primer 202a and reverse primer 204a can be designed to include sequences complementary to the starting portion and the termination portion of target region 206a. The forward and reverse primers are then added to the test sample including cancer DNA, and subjected to PCR conditions, including one or more rounds of suitable thermal cycles for denaturation, renaturation and extension using appropriate reagents known in the art (e.g., nucleotides, buffers, polymerases, etc.), to produce multiple PCR products, such as amplicon 208a. The term "amplicon" used herein refers to a product (or "band") amplified by a specific primer pair in a PCR reaction. Amplicon 208a can then be sequenced, and the number of reads containing sequence variation 210a can then be counted for the target region 206a.

[0056] PCR can be a multiplex PCR, which uses two or more primer pairs for different targets. If there are two or more targets in the reaction, the multiplex PCR will produce two or more amplified DNA products, which are amplified together using a corresponding number of sequence-specific primer pairs in a single reaction. As shown in Figure 2, the multiplex PCR can include three pairs of forward primers 202a, 202b, 202c and reverse primers 204a, 204, 204c, which are individually designed for multiple target regions 206a, 206b, 206c, thereby producing amplicons 208a, 208b, 208c, which can then be sequenced. The observation of sequence reads containing sequence variations 210a, 210b, 210c supports that the observed sequence variations are true genetic variations present in the sample, indicating the presence of cancer DNA in the test sample.

[0057] In some embodiments, the test sample can first be pre-amplified, such as by whole genome amplification. Pre-amplification can be achieved by, for example, connecting adapters and performing PCR for the connected adapters. In these embodiments, sequencing adapters can be added during amplification, or can be connected after amplification. In other embodiments, the target region can be enriched using a "target-based enrichment" method, wherein the adapter is connected to the test sample, and before amplification, a primer hybridized with the adapter is used to enrich the fragment containing the target region by hybridizing with a nucleic acid probe. In such an embodiment, a ligation reaction can be performed, or an adapter with multiple barcodes can be connected to DNA, so that the molecular group can be effectively separated into a separate barcode group or repeat. Therefore, the sequence of the target region can be enriched from the sample by PCR or hybridization with a nucleic acid probe. Other enrichment methods can be used. In other embodiments, any other method with physical replication or the use of molecular barcodes, such as molecular inversion probes (MIP) or anchored multiple PCR (AMP) can be used. In some embodiments, COLD-PCR, allele-specific PCR, methods for digesting wild-type sequences by utilizing adjacent germline variations, or other methods known to those skilled in the art can be used to enrich the target region during the targeting step. In a preferred embodiment, the pre-amplification step is performed using multiplex PCR, and the sample is then divided into two or more samples for further PCR analysis (single or multiple). In this embodiment, the samples can be pooled and further barcoding steps can be performed to sequence the amplicons.

[0058] Although the remainder of this disclosure describes in detail the use of PCR and "amplicon" sequencing, embodiments of the present disclosure may also be applicable to other methods, including methods that pre-amplify samples or pre-amplify using molecular barcodes or indexes (e.g., random sequences attached to nucleic acids). In such embodiments, comparing the measurement results supporting the presence of the class of the target region to one or more error models can include estimating the probability that the sequence variation is present in the target region by, for example, measuring or counting the number of index sequences for the target region.

[0059] The "category" of a target region can refer to a target region with one or more types of genetic variation. For example, the category of a target region can include a target region comprising: a single nucleotide variation (SNV), such as a single base change of A>T or C>G; a multiple nucleotide variation (MNV), such as a doublet base substitution of CA>TG or a triplet base substitution of AAA>TTT; a short insertion or deletion (INDEL) of one or more nucleotides, such as an insertion of TTTT or a deletion of CG; a copy number variation (CNV), including cases of gene amplification, chromosomal aneuploidy, or tandem duplication, which can usually be detected as a target region with a significant increase in sequencing coverage; a structural variation (SV) reflecting relatively large genetic changes, such as a gene fusion or a large fragment insertion or deletion of, for example, 1,000, 10,000, 100,000, or 1,000,000 nucleotides; and an epigenetic variation (EV), such as changes in DNA methylation, DNA protein modification, chromatin accessibility, histone modification, etc. In addition to the type of change, the category of a target region can also refer to a specific change. For example, the categories of target regions may include SNV changes of A to T at specific positions or in specific sequence contexts (e.g., a trinucleotide context, i.e., a specific nucleotide immediately adjacent to a genetic variation, or a pentanucleotide context, i.e., two adjacent bases on either side of the variation), INDEL changes of AAAA to AA, SNV changes of A to T at the first position and C to G at the second position, etc.

[0060] The category of the target region may also include multiple (ie, two or more) genetic variations. The two or more genetic variations may be of the same type (eg, two or more SNVs, INDELs, SVs, and EVs) or two or more different types (eg, 1 SNV and 1 INDEL; 1 SNV and 1 INDEL and 1 EV; etc.). The two or more genetic variations may be separated by at least one nucleotide. Two or more genetic variations present on the same DNA molecule may be referred to as phase variations (PVs). The term "phase" refers to determining whether a genetic variation is located on the maternal or paternal copy of the chromosome (eg, chromosome 1). When two or more genetic variations are both present on the same chromosome (ie, the maternal or paternal copy), they can be considered as PVs in the context of each other, so they will be present on the same DNA molecule in the test sample. If two PVs are close enough (eg, within the same target region), they may be amplified and sequenced, thereby being observed in the same sequence read length. As Figure 2AAs shown, amplicon 208c generated from target region 206c may contain two sequence variations 210c, 210d (each variation is individually indicated by an "X") present on the same amplicon 208c, while amplicons 208a, 208b contain only a single sequence variation 210a, 210b; therefore, both sequence variations 210c, 210d (if true genetic variations) are PVs. Although the class of target regions containing PVs may contain any combination of phased genetic variations, in the context of cfDNA, the class of target regions containing PVs will typically contain two or more SNVs present on the same DNA molecule.

[0061] In some embodiments, genetic variation is somatic variation, that is, they are non-germline genetic variation that may be associated with a disease (e.g., cancer). In some embodiments, genetic variation may include germline genetic variation, i.e., genetic variation constituting the patient's (non-tumor) genome. Germline genetic variation may be useful in a target region category with two or more genetic variations. For example, target regions containing germline SNVs and somatic tumor SNVs simultaneously can be enriched and sequenced. Preferably, germline SNVs and tumor SNVs are phase variations. In this case, observing the tumor SNVs combined with germline SNVs in a single sequence read provides unique identification information, which increases the possibility that the tumor SNVs are true.

[0062] Figure 2B The various types of target regions are further described. Figure 2B In each example category, two DNA molecules are depicted as lines, illustrating the two copies of each chromosome (paternal and maternal) that will be amplified by a pair of PCR primers targeting a specific region. Figure 2B As shown, the categories of target regions may include: SNV (250); MNV (252), which is a double base substitution here; two SNVs (254), which are located on opposite chromosomes and therefore on different DNA molecules; two PVs (256), which are two SNVs located on the same chromosome and therefore on the same DNA molecule; germline SNV and tumor SNV (258), which are PVs as shown in the figure because they are located on the same chromosome and therefore on the same DNA molecule; a single INDEL, showing a single base deletion (260); a single INDEL, showing a single base insertion (262); SNV and a single base deletion (264), which are PVs as shown in the figure because they are located on the same chromosome and therefore on the same DNA molecule; and EV, which shows methylated cytosine that has not yet been converted to uracil by bisulfite treatment (266).

[0063] In some embodiments, enriching for multiple target regions can include performing a multiplex PCR assay in which multiple target regions are amplified simultaneously in a test sample. Figure 3 An embodiment of a genetic variation analysis assay using the method 300 described herein is depicted. The assay 300 may include measuring a plurality of target regions, each target region comprising a category. Figure 3 As shown, each box represents a different target region measured by the assay, preferably in the same reaction volume. The assay detects multiple different categories of target regions, including: single nucleotide variation (SNV) 302, multinucleotide variation (MNV) 304, copy number variation (CNV) 306, short insertion / deletion (INDEL) 308, structural variation (SV) 310, epigenetic variation (EV) 312, and phase variation (PV) 314. As shown, some target regions may contain two or more genetic variations, including 2 SNVs (316) (not on the same chromosome) and 2 PVs (314). Target regions with 2 SNVs on different chromosomes have advantages because they double the utility of a particular target region by analyzing two different variations simultaneously (e.g., by using the same pair of PCR primers in a multiplex reaction). In some embodiments, either of the two SNVs or PVs can be a germline SNV. As previously described, PVs can include any type of genetic variation or combination of genetic variations. For example, as Figure 3 As further shown in FIG. 2 , two PV regions ( 314 ) may include regions having both SNVs ( 318 ) and single nucleotide INDEL ( 320 ) deletions.

[0064] The assay described herein can include any number of target region categories, and can include multiple target regions with the same category. In a preferred embodiment, the assay includes at least two different categories of target regions. Assay 300 can be applied to a test sample to determine the sample state of each target region analyzed by the assay. In some embodiments, each target region can be enriched and measured (e.g., sequencing). For any given target region, it is possible to determine the measurement results (e.g., specific SNV, CNV, INDEL, SV, EV and / or two or more PVs located in a single target region) supporting the presence of the category of the target region.

[0065] As will be described in further detail below, the methods described herein can combine comparisons of target regions from multiple categories (e.g., multiple target regions comprising genetic variations determined 300). One benefit of combining comparisons of target regions from various categories is that evidence from different types of variations can be considered jointly to support a high-confidence conclusion that cancer DNA (or RNA) is present in the test sample. For example, in some embodiments, the first target region includes a first category, wherein the first category includes SNVs. In these embodiments, the second target region may include a second category, which includes CNVs, INDELs, SVs, EVs, or in particular two or more PVs. In another embodiment, the first target region includes a first category, wherein the first category includes CNVs. In these embodiments, the second target region may include a second category, wherein the second category includes SNVs, INDELs, SVs, EVs, or two or more PVs, in particular two or more PVs. In some embodiments, the third target region may include a third category, wherein the third category includes SNVs, CNVs, INDELs, SVs, EVs, or two or more PVs, in particular two or more PVs. Various combinations of target region categories are contemplated herein.

[0066] The target region can be selected by first identifying a plurality of genetic variations of interest, such as genetic variations associated with the patient's cancer. The genetic variations can include previously identified sequence variations, such as variations known or suspected to be associated with the patient's cancer. These variations can also be identified from the patient's cancer, such as somatic mutations in the genome of the patient's cancer cells or somatic mutations in the genome of the patient's cancer cells prior to any cancer treatment. For example, genetic variations can include variations present or previously identified in various cancer-related genes, including but not limited to TP53, EGFR, BRAF, and KRAS, as well as other genes frequently mutated in cancer (e.g., genes in the COSMIC Cancer Gene Census, available at cancer.sanger.ac.uk / census; see also Sondka et al., The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers, Nature Reviews Cancer 18, 696-705 (2018), the contents of which are incorporated herein by reference); regions of common structural rearrangements (e.g., common gene fusions or common amplification edges, such as MYC), and regions of common amplification, rearrangement (e.g., chromothripsis), common local hypermutation (e.g., Kataegis), epigenetic changes, etc.

[0067] In some embodiments, genetic variation can include cancer-specific genetic variation identified by sequencing DNA isolated from cancer cells of the patient. For example, cancer-specific variation can be identified by sequencing DNA or RNA isolated from a biological sample containing cancer cells obtained from a cancer patient. Tumor-specific variation can be identified by sequencing DNA or RNA isolated from a tissue sample obtained from a tumor biopsy of a cancer patient. Alternatively, tumor-specific variation can be identified by sequencing DNA or RNA isolated from cell-free DNA or RNA, or from the circulating cancer cells of the patient. For blood cancer, genetic variation can be identified by sequencing DNA or RNA samples from bone marrow, circulating blood cells or lymph nodes. In such embodiments, the assay described herein can be a "personalized" assay because the genetic variation is obtained from the same patient.

[0068] In some embodiments, targeted sequencing methods (e.g., hybrid capture sequencing) are used to identify cancer-specific variations. In another embodiment, pull-down or non-pull-down techniques designed to enrich selected sequences are used to identify cancer-specific variations. These methods can sequence different regions of the genome, such as exomes, i.e., whole exome sequencing (WES), which can include genomic regions containing common mutations in cancer genes or regions containing frequent mutations that are not in genes. In a preferred embodiment, cancer-specific variations are identified by WES of tumor tissue. In other embodiments, whole genome sequencing (WGS) can be used to identify tumor-specific variations, wherein samples are sequenced without any specific enrichment.

[0069] WES and similar targeted sequencing methods effectively limit the search space across the genome by selecting certain predetermined sequences, resulting in higher coverage and increased confidence in somatic variant calls. Such methods may also produce genetic variants that are more likely to have functional effects. However, limiting the search space may produce fewer total genetic variants, which can also affect the types of variants identified. For example, in blood cancers such as lymphomas, phase variants (PVs) tend to cluster in known "hotspot" regions and can therefore be identified using targeted techniques such as WES or WGS. However, in solid cancers, PVs tend to be randomly scattered throughout the genome, so fewer PVs that are close enough to each other to be located within a single target region will be identified. Therefore, prior art methods focus on identifying a single type of variation (e.g., PV) for cancer diagnostic testing, which rely on WGS and identify a large number of PVs in "hotspot" regions in blood cancers (see, e.g., Kurtz, DM et al., Enhanced detection of minimal residual disease by targeted sequencing of phased variants in circulating tumor DNA, Nat Biotech 1-11 (2021), which is incorporated herein by reference in its entirety). The inventors have recognized and appreciated that PVs that are sufficiently close to each other can significantly improve specificity, but for example, using (e.g.) WES or targeted technology in methods for identifying genetic variations in target regions of solid tumors can only find a few PVs within a sufficient distance. The method described herein solves this problem by combining the measurement results of a target region containing (e.g.) two or more PVs and a target region containing other types of variations, thereby achieving a "hybrid" approach that can study an economical number of genetic variations in a test sample and can utilize any available evidence. For example, in assay 300, evidence from two PV target regions 314 can be combined with evidence from a single SNV target region 302. In contrast, state-of-the-art WGS-based methods rely on identifying a large number of PVs. Furthermore, by focusing only on phase variation, important information is missed; therefore, methods that use only phase variation as a target region category will have lower sensitivity than methods that use multiple categories.

[0070] In some embodiments, cancer-specific genetic variation is compared with the genetic variation obtained from the normal sample of matching. The sample is "normal" because it is derived from non-cancerous biological material, and is "matched" because it is from the same patient. For example, the normal sample (such as oral swab DNA, whole blood DNA or adjacent non-cancerous DNA (i.e., from tissue adjacent to a seemingly normal tumor)) of the matching of the non-cancerous DNA from the same patient can be sequenced, and compared with the cancer-specific genetic variation of the patient. The sequencing of these matched normal samples can be carried out simultaneously with the sequencing of the cancer cells from the patient, or it can be carried out before or after the sequencing of the cancer cells from the patient. The genetic variation detected in the cancer cell (cancer DNA) can be selected, rather than the genetic variation detected in the normal sample (non-cancerous DNA) of the matching, to be included in the mensuration described herein, because these variations are more likely to be cancer-specific. The variations detected in the normal sample (non-cancerous DNA) of the matching can be excluded, because they may not be cancer-specific. Those skilled in the art will be aware of various software packages for determining tumor-specific mutations, such as MuTect2 (Cibulskis et al., Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol. 2013; 31: 213-9) and VarScan2 (Koboldt et al., VarScan 2: somatic mutation and copy number alteration discovery in cancer by exomesequencing, Genome Res. 2012; 22: 568-76).

[0071] In some embodiments, the genetic variation is a clonal genetic variation. Cancer DNA includes clonal and subclonal mutations. During the evolution of a tumor, there is a transition between clonal mutations and subclonal mutations. Subclonal mutations are present only in a group of cells in the tumor: these mutations occur after the most recent common ancestor of all cancer cells in the tumor sample. In contrast, clonal mutations occur before the most recent common ancestor of all cancer cells. Therefore, clonal mutations are present in all cells in the tumor unless there is some mechanism to eliminate mutations, such as structural variation, in which case the entire locus will be lost in a group of cells. Clonal changes usually appear early in cancer evolution and are present in all cancer cells. When a genetic variation is present in multiple biological samples or can be inferred from sequence reads generated from a large number of tumor tissues, it can be considered a clone. Clonality can be difficult to determine because tumors are often heterogeneous, the entire tumor cannot be sequenced, and it is challenging to quantify heterogeneity from a large amount of sequencing data. Various methods have been proposed to determine clonality, including Bayesian mixed models, clustered probability distributions of cancer cell parts, and phylogenetic methods. Software tools for determining clonality include PyClone-VI, EXPANDS, QuantumClone, and PhyloWGS.See also Gillis, S., Roth, A. PyClone-VI: scalable inference of clonal population structures using whole genome data. BMC Bioinformatics 21, 571 (2020); Andor et al., EXPANDS: expanding ploidy and allelefrequencies on nested subpopulations, Bioinformatics 30(1): 50-60 (2013); Deveau et al., QuantumClone: ​​clonal assessment of functional mutations in cancer based on a genotype-aware method for clonal reconstruction, Bioinformatics 34(11): 1808-1816 (2018); Deshwar et al., PhyloWGS: reconstructing subclonal composition and evolution from whole-genome sequencing of tumors, Genome Biology 16(35) (2015); the contents of which are incorporated herein by reference in their entirety. The sample can be sequenced by whole genome sequencing, whole exome sequencing, or targeted sequencing (e.g., by sequencing a panel of cancer genes or sequencing a panel of sequences that are mutation hotspots), etc.

[0072] Target regions containing genetic variations can be ranked or filtered based on the type of genetic variation present. For example, target regions can be ranked based on one or more of the following: clonality or allele fraction within a cancer sample; likelihood of unique alignment; estimated background error rate, where genetic variations showing evidence of sequence or PCR polymerase error rate are penalized or filtered; high signal background events, where genetic variations showing DNA damage or early cycle PCR errors are penalized or filtered; the category of the target region, such as prioritizing a pair of PVs based on their predictive utility; proximity of any germline (non-somatic) variation that may contribute to enrichment; likelihood of being a somatic change; etc.

[0073] Once a genetic variation is selected, the corresponding target region can be determined by, for example, selecting the upstream and downstream positions of the genetic variation. For example, the target region can include a portion of the genome that begins at (for example) 75bp before the genetic variation and ends at (for example) 75bp after the genetic variation. In some embodiments, PCR primers are designed to amplify the target region. In some embodiments, oligonucleotide probes are designed to enrich the target region. In embodiments where the test sample includes cfDNA, the length of the target region is preferably designed to be about 150bp, reflecting the average fragment length of the cfDNA molecule.

[0074] Measuring multiple target regions (step 104) can be performed in a variety of ways. In some embodiments, the measurement is performed by digital PCR (dPCR) or droplet digital PCR (ddPCR). The measurement can also be performed by quantitative PCR or other fluorescence-based assays. In some embodiments using molecular barcodes, the measurement can include generating a consensus sequence for the target region and determining whether the consensus supports the category of the target region.

[0075] In some embodiments, measuring multiple target regions in the enriched sample includes sequencing the multiple target regions of step (a) to generate multiple sequence reads corresponding to the first target region and the second target region. In such embodiments, comparing the measurement results (e.g., step 106 of method 100) includes comparing the number of sequence reads supporting the presence of the category of the target region with one or more error models, which model the probability of observing the category of the target region in DNA or RNA that does not contain the target region.

[0076] Sequencing generally refers to a method of obtaining the identity of at least 10 consecutive nucleotides of a polynucleotide (e.g., the identity of at least 20, at least 50, at least 100, or at least 200 or more consecutive nucleotides). In a preferred embodiment, sequencing is performed using next generation sequencing, a so-called highly parallelized nucleic acid sequencing method, and includes synthetic sequencing, ligation sequencing, and combined sequencing platforms currently used by companies such as Illumina, Life Technologies, Pacific Biosciences, Element Biosciences, Singular Genomics, Omniome, Genapsys, Ultima Genomics, and Roche. Next generation sequencing methods may also include, but are not limited to, nanopore sequencing methods or electronic detection-based methods such as those commercialized by Life Technologies, such as the Ion Torrent technology. In some embodiments, sequencing is performed using an Illumina NextSeq or NovaSeq system. In some embodiments, sequencing is performed using pyrophosphate sequencing such as the Roche 454GS FLX system.

[0077] The output of the sequencing process is a plurality of sequence reads, i.e., a string of letters indicating the order in which certain nucleotides (e.g., A, C, G, T) exist in the sequenced DNA molecule or amplicon. The length of the sequence read may vary from 25-1000bp or longer, and in many cases, each base of the sequence read may be associated with a score indicating the quality of the base determination. As previously mentioned, cfDNA in blood is usually highly fragmented, with an average length of about 160bp. Therefore, in some embodiments, the length of the target region may be about 160bp, and the length of the amplicon may be about 160bp or less. In these embodiments, the length of the sequence read is preferably at least 160bp, so that the entire amplicon is sequenced, thereby sequencing the entire target region.

[0078] In some embodiments, the sequence read length corresponds to the first target region and the second target region. In some embodiments, the sequencing adapter can be directly connected to the amplicon having the sequence of the first target region and the second target region. In other embodiments, the sequencing adapter can be incorporated into the amplicon during amplification (i.e., during PCR). Various embodiments and modifications are considered within the scope of the present disclosure.

[0079] In some embodiments, the target region includes a class that includes two or more phase variations (PVs). In such embodiments, two or more PVs are present (or expected to be present) in the same target region. Since the PVs are present on the same DNA molecule, the sequence read of the corresponding amplicon may include both PVs, thereby providing highly specific evidence for the presence of cancer DNA in the sample.

[0080] In some embodiments, due to the low proportion of cancer DNA, high depth sequencing may be required to identify genetic variation. In some embodiments, sequencing multiple target regions includes sequencing to a minimum read depth of at least 10,000, at least 25,000, at least 50,000 or at least 100,000, at least 200,000 or at least 500,000. In some embodiments, sequencing multiple target regions includes sequencing to a maximum read depth of at least 25,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000 or at least 1,000,000. In any embodiment, the read depth of the step may be about 10,000 to about 500,000. In any embodiment, the read depth may be about 10,000 to about 200,000.

[0081] In some embodiments, the sequence reads are processed by computation, for example, by pruning, demultiplexing, alignment, matching, folding and / or filtering. Typically, each sequence read is assigned to one of the target regions containing or suspected of containing one or more genetic variations associated with the patient's cancer. For example, the sequence reads can be analyzed to determine which reads correspond to multiple target regions. As will be appreciated by those of ordinary skill in the art, the sequence reads identical or nearly identical to the target region can be analyzed to determine whether there is a potential genetic variation in the target sequence. The sequence can be aligned with a reference sequence (e.g., a genomic sequence), or matched with a database of expected sequences to determine their most likely positions on the reference sequence.

[0082] After processing the sequence reads, the number (e.g., number) (k) of sequence reads containing genetic variation or multiple genetic variations and the total amount (e.g., number) (n) of sequence reads (e.g., number) can be determined for each target region. The method of quantitative read length can be adapted from, for example, the methods described by Forshew et al. (Sci. Transl. Med. 2012 4: 136ra68), Gale et al. (PLoS One 2018 13: e0194630) and Weaver et al. (Nat. Genet. 2014 46: 837-843), all of which are incorporated herein by reference in their entirety. Similar results can be obtained using methods employing molecular indexes. In these methods, an index can be used to estimate the total number of molecules sequenced and the number of variant molecules. Such molecular identifier sequences can be used in combination with other features of the fragment (e.g., the end sequence of the fragment defining the breakpoint) to distinguish fragments. Molecular identifier sequences are described in (Casbon Nucl. Acids Res. 2011, 22e81), which is incorporated herein by reference in its entirety.

[0083] The measurement results supporting the existence of the category of the target region can be compared with one or more error models in a variety of ways, and the error model simulates the probability of observing the category of the target region in the DNA of the category that does not contain the target region (step 106). In some embodiments, the comparison includes comparing the measurement results (k) supporting the existence of the category with a binomial, dispersed binomial, β-binomial, polynomial, normal, exponential or gamma error probability distribution model for each target region. For example, in one embodiment, the error probability distribution model of the first category of the target region is a β-binomial error probability distribution model, and the error probability distribution model of the second category of the target region is a polynomial error probability distribution model. In some embodiments, the sequence read length number (k) supporting the existence of the category includes the sequence read length number containing genetic variation. In some embodiments, the comparison also includes generating a statistical evaluation or score, which describes the degree of evidence supporting the conclusion that a given target region contains one or more genetic variations in a sample. In some embodiments, the statistical evaluation can be, for example, a p-value, likelihood, likelihood ratio or probability distribution. The statistical evaluation may also preferably include a likelihood ratio method, in which the likelihood of observing n sequence reads containing one or more genetic variations in the test sample is determined if i) cancer DNA is present in the sample, and ii) cancer DNA is not present in the sample. These values ​​can then be used to calculate (for example) likelihood ratios to determine whether one or more genetic variations in the target region are present in the sample.

[0084] As mentioned above, if present, cancer DNA typically represents a small fraction of cell-free DNA. For example, in MRD, the cancer fraction may be as low as 0.01ppm. At this level, the inventors have recognized and realized several problems that may produce, for example, false positive results. First, sequencing is not perfect, and background errors may cause misreading of bases, which may lead to false positive ctDNA determinations. Secondly, errors may also be introduced during PCR. For example, bases may be "switched" due to DNA damage (e.g., oxidation, deamination) before amplification, and subsequent PCR amplification may ultimately lead to many sequence reads that support erroneous conclusions. In addition, as the number of genetic variants included in the assay increases, the possibility of false positive determinations also increases accordingly. Some assays require (for example) to determine that at least two separate genetic variants are positive in order to make an accurate diagnosis. However, this approach is flawed because it limits the number of variants detected by the assay, thereby reducing the amount of available ctDNA "signals" and reducing sensitivity.

[0085] One way to account for background error is to model the error as a probability distribution and then determine whether an observed genetic variation is unlikely to come from background error. For example, given a background error rate p, the probability of observing k sequence reads containing a genetic variation in a target region can be determined using a binomial probability distribution:

[0086]

[0087] Where n is the total number of sequence reads. The background error rate can be estimated from a set of control samples that do not contain any cancer-associated variants. If the determined probability is less than a threshold level (e.g., 0.05, 0.01, 0.001, 0.0001), the genetic variant can be said to be present in the sample.

[0088] Probability (e.g., P(X=k)) refers to the chance of a particular outcome occurring, or the likelihood of that outcome occurring. Probability may be based on the values ​​of parameters in a model. Probability refers to an unknown event and is attached to possible outcomes. Since possible outcomes are mutually exclusive and exhaustive, probabilities can be expressed on a linear scale. For example, probability can be expressed as a value between 0 (impossible) and 1 (definite), or can equally be expressed as a percentage or fraction. For example, in the context of the present disclosure, probability can be used as a measure to determine whether cancer DNA is present in a sample.

[0089] An error probability distribution (also referred to as an "error model", "error distribution", or "error probability distribution model") is a distribution that estimates or models the probability that an observation (e.g., variant allele fraction) is attributable to error. These terms can refer to any type of error, including errors due to DNA damage or early cycle PCR errors, as well as sequencing errors. The error model assumed in Figure 4A -B is shown as a frequency distribution. In these examples, multiple samples (e.g., hundreds of samples) that are not known to contain somatic genetic variation (i.e., healthy control samples) are sequenced, and the fraction of sequence reads with a particular type of sequence variation is calculated for each sample. Any sequence variation in the sequence read is mainly caused by errors that occur during PCR, base miscalls, and pre-PCR events (e.g., DNA damage, for example, guanine is oxidized to 8-oxoguanine, which pairs with the A base, resulting in a G to T variation that appears in the sequence read). These scores can be plotted as a frequency distribution, which can then be used to calculate the probability of whether the sequence variation observed in the sequence read is a true genetic variation.

[0090] In some embodiments, likelihood ratio (LR) can be used to estimate the degree of evidence supporting the existence of the category of the target area. Likelihood ratio refers to the ratio of at least two possibilities (each possibility is associated with a different hypothesis), which can be used to determine which hypothesis is more likely in the case of a given experimental result. Each possibility refers to the hypothesis probability that an event that has occurred produces a specific result. Likelihood is used to assess the support of the sample for a specific value of a parameter in the model. Therefore, likelihood refers to a past event with a known result and is associated with a hypothesis.

[0091] Likelihood ratios can be used as a measure of diagnostic accuracy because they can be used to determine the potential utility of a particular diagnostic test, as well as the likelihood that a patient has a disease or condition. When applied to diagnostic testing, the likelihood ratio is the ratio of the likelihood that a given result would be expected to occur in a sample that does not contain any cancer DNA to the likelihood that the same result would be expected to occur in a sample that contains cancer DNA. As shown in equations (2), (3), and (4) below, two hypotheses can be determined: H0, the likelihood of observing k reads containing genetic variation assuming that there is no cancer in the sample (null hypothesis); H1, the likelihood of observing k reads containing genetic variation (thus supporting the class of the target region) assuming that there is at least one cancer molecule (where z is the total number of input DNA molecules, which can be estimated, for example, by optical diffraction or digital PCR). Each hypothesis contains a background error rate (p) indicating the frequency with which sequence reads containing genetic variation in the class of the target region may be due to error. The background error rate (p) can be selected separately for each target region. The ratio of these two values ​​(H1 / H0) can then be determined and compared to a threshold. Values ​​greater than 1 indicate that H1 is a more likely hypothesis, while values ​​between 0 and 1 indicate that H0 is more likely. The threshold required to call a cancer-associated variant present can vary depending on the desired sensitivity and specificity, which can be determined from, for example, a set of known samples.

[0092] H0=Binom(k,n,p) (2)

[0093]

[0094] Thus, in any embodiment, comparing the number of sequence reads (k) supporting the presence of the target region class to one or more error models (step 106) can include calculating a likelihood ratio between the likelihood of observing the number of sequence reads containing the genetic variation: (i) if the cancer DNA is present, and (ii) if the cancer DNA is not present. Similarly, in any embodiment, this can be done by calculating a likelihood ratio (LR) between the likelihood of observing the k reads for each target region. i ) to complete: (i) if cancer DNA is present, and (ii) if cancer DNA is not present. As described in more detail below, in these embodiments, the individual likelihood ratio LR i The cumulative LR score for all target regions of the test sample can be combined (e.g., LR equal to the sum of the log-likelihoods i product of ).

[0095] In some embodiments, the target region category contains two or more phase variations (PVs). The two or more PVs may include a first genetic variation and a second genetic variation located on the same DNA molecule. Therefore, each variation will be sequenced together on a single sequence read. In these embodiments, by comparing the number of sequence reads containing the first genetic variation and the second genetic variation, the number of reads containing only the first genetic variation, the number of reads containing only the second genetic variation, and the number of reads containing no genetic variation with a multinomial distribution, the number of sequence reads supporting the existence of the category of the target region (wherein the category contains two or more phase variations) can be compared with one or more error models (step 106), and the error model models the category of the target region observed in the DNA that does not contain the category of the target region. For example, the probability of observing two phase variations in the target region can be modeled as:

[0096]

[0097] Where k1 is the number of sequence reads where two genetic variations are observed (Alt, Alt), k2 is the number of sequence reads where only the first genetic variation is observed (Alt, Ref), k3 is the number of reads where only the second genetic variation is observed (Ref, Alt), k4 is the number of reads where no variation is observed (Ref, Ref), and X = (X1, X2, X3, X4) is a random vector whose components are not independent but satisfy the condition that the sum is 1.

[0098] Various probability density functions can be used for X. One candidate for this role is the standard Dirichlet distribution:

[0099]

[0100] Where δ(-) is the Dirac delta distribution, α1,α2,α3,α4≥0 are parameters, and 0≤x1,x2,x3,x4≤1. This distribution has an important property, that is, the marginal distribution is a beta distribution:

[0101] X i ~Beta(α i ,α0-α i ) (7)

[0102] Among them, for i = 1, 2, 3, 4, α0 = α1 + α2 + α3 + α4. It is also possible to consider using more advanced models (such as generalized Dirichlet distribution) to capture more complex correlations between the components of X.

[0103] It is also possible to calculate the individual probabilities of X. For example, X i =P(k i ), that is, observe k i Given a tumor fraction θ, we observe that each k i The probability is:

[0104] Alt,Alt:P(k1)=(1-θ*e1*e2)+(θ*1-e′1*1-e′2) (8)

[0105] Alt,Ref:P(k2)=(1-θ*e1*1-e2)+(θ*1-e′1*e′2) (9)

[0106] Ref,Alt:P(k3)=(1-θ*1-e1*e2)+(θ*e′1*1-e′2) (10)

[0107] Ref,Ref:P(k4)=(1-θ*1-e1*1-e2)+(θ*e′1*e′2) (11)

[0108] where e1 and e2 are the error rates of the first and second phase variants, respectively, where the first and second PV observations are from non-cancerous DNA and are due to sequencing errors, and e1 ’ and e2 ’ is the error rate at which the corresponding observations of non-cancerous DNA at the first and second phase variation positions are from tumor DNA and are attributed to sequencing errors. In such embodiments, the error rate can be replaced by a random variable (e.g., a binomial distribution or a beta-binomial distribution), resulting in a random variable P(θ):

[0109] P(θ)=mean(P(k1,k2,k3,k4)), (12)

[0110] It can be incorporated into the hypothesis (H1) described herein that cancer DNA is present in the sample.

[0111] Although the above embodiments describe two phase variations, the methods according to the present disclosure can be further modified by those skilled in the art to accommodate target regions with additional PVs (e.g., three or more, four or more, five or more, or ten or more PVs). However, in preferred embodiments, two PVs are generally sufficient for the target region category because these PVs are more likely to be found close to each other and present on a single DNA fragment.

[0112] Various error models can be used to model the probability of observing the category of the target region in the DNA of the category that does not contain the target region. In some embodiments, an error model corresponding to the category of the target region is selected. In some embodiments, the category of the target region is SNV, and the error model is a binomial distribution. In some embodiments, the category of the target region is two or more PVs, and the error model is a multinomial distribution. In some embodiments, the comparison includes comparison with two or more error models. In such embodiments, two or more error models can model various types of errors, including but not limited to sequencing errors, PCR errors, DNA damage, polymerase errors, etc. For example, a first error model (e.g., a binomial probability distribution) can be used to estimate the background error rate from sequencing, and a second error model (e.g., a Poisson distribution) can be used to estimate the background error rate from DNA damage. In some embodiments, a single distribution can explain one or more types of errors. For example, the two shape parameters (α, β) in the β-binomial distribution can be adjusted to accommodate the estimated background error rate and DNA damage.

[0113] In some embodiments, a probability distribution is used to estimate the background error rate. In some embodiments, there may be two distributions of the same family or type (e.g., 2 binomial distributions), or, if two different families or types of distributions are used, there may be one distribution for the background error rate and another for PCR errors.

[0114] In some embodiments, H1, the likelihood of observing k reads containing a genetic variant (thus supporting the class of the target region) assuming the presence of at least one cancer molecule, is adjusted by additional probabilities, such as the probability that the genetic variant is cancer-specific (V C ). For example, the above equation (3) can be modified as follows:

[0115]

[0116] This approach is useful in embodiments that combine, for example, structural variants, INDELs, and phase variants, because any background errors are unlikely to result in the observation of such genetic variants. In this case, a single sequence read can provide substantial evidence for the presence of cancer DNA in the test sample. Thus, in such embodiments, it may be desirable to account for the probability (V) that a genetic variant is not tumor-specific. C ) to modify H1, which may be due to clonal hematopoietic mutations of uncertain likelihood (CHIP) or potential contamination. In a preferred embodiment, V C is set to a value such that a single observed sequence read that alone supports the presence of a class is insufficient to provide substantial evidence to support the conclusion that cancer DNA is present in the sample.

[0117] In addition, the inventors have recognized and appreciated that, according to the embodiments of the present disclosure, when short insertions and deletions (INDELs) are included as categories of target regions, they may be particularly useful. For example, the category of the target region may include INDELs of, for example, 1-5, 1-10, 1-15, or 1-20 nucleotides in length. In such embodiments, INDELs of 1-5 nucleotides in length are preferred, because these INDELs are more likely to be observed in the test sample. In other embodiments, the category of the target region may include INDELs of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, or 20 or more nucleotides in length. In some embodiments, the maximum INDEL length is 20 nucleotides, because longer variations may affect the accuracy of sequence alignment. Depending on the type of sequencing technology used, the background error rate associated with INDELs may be very low. Therefore, observing INDELs in the target region can provide relatively large amounts of evidence that INDELs (and therefore cancer DNA) are present in the test sample.

[0118] The error model can be trained using a control sample (e.g., a DNA sample known not to contain any sequence variation or collected from a healthy patient (e.g., a patient not suffering from cancer). By sequencing a control sample known not to contain any sequence variation, any sequence variation observed in the control sample must be due to error. Such observations can be used to set the parameters of the error model according to the present disclosure. Preferably, the control sample is treated under conditions similar to the test sample. For example, primers can amplify the same or similar target region, and the sequencing technology can be the same. An error model can be constructed using many control samples, such as at least about 50 samples. The error model can be stored in a computer database and accessed as needed. Therefore, in one embodiment, the error model is trained based on a group of control samples. In one embodiment, the group of control samples is from a healthy donor. In such embodiments, a background error (p) for a target region category in the absence of cancer is established based on a group of control sample training error models.

[0119] In some embodiments, multiple comparison correction is applied to the comparison to prevent false positives, such as setting a more stringent threshold or applying a Bonferroni correction. In any embodiment, the threshold can be determined using a binomial, overdispersed binomial, Beta, normal, exponential or Gamma probability distribution model of the background error rate of sequence variation, wherein the frequency is selected so that when there are no mutant molecules, depending on the desired predefined specificity per variant, a signal above this frequency is observed in less than 0.1%, 0.01% or 0.001% (preferably 0.1%) of the time.

[0120] In some embodiments, CNV is identified using a read depth method, wherein a non-overlapping sliding window is used to count the number of sequence reads mapped to genomic regions overlapping with the window. Regions where the read depth is significantly increased (exceeding that expected based on typical background errors associated with the sequence) can be further analyzed to identify copy number. Alternatively, a double-end method can be used, wherein copy number variation is detected based on the distance between mapped paired sequence reads. Sequence reads can also be assembled from scratch, and the resulting assembled continuous sequence is aligned with a reference genome to identify copy number variation.

[0121] In some embodiments, epigenetic variations (EVs) are identified by treating a test sample and then sequencing. For example, methylated nucleotides can be identified by treating with sodium bisulfite, which converts unmethylated cytosine to uracil. Amplification and sequencing of the sample then converts uracil bases to thymine (T). Therefore, the presence of unmodified cytosine bases in the target region can support the presence of a target region containing EVs. Comparison of EVs can be performed using an error model that similarly measures the background level of sequencing errors, optionally taking into account any errors associated with the bisulfite conversion process.

[0122] The comparison of at least the first target region and the second target region can be combined in a variety of ways (step 108). In some prior art methods, each genetic variation is determined separately, rather than as a whole set. Although statistical corrections can be applied to single variation determinations as the number of variants increases, the higher stringency required to eliminate false positives may also have an adverse effect on sensitivity and discount most variation determinations. The inventor has recognized and understood that each independent analysis of each target region can contribute a certain degree of evidence to the cumulative statistical evaluation. Compared with considering each target region alone, combining the scores of two or more target regions can produce high confidence determinations for the test sample without correspondingly reducing sensitivity or increasing false positives. Therefore, in some embodiments, the comparison of each of the first target region and the second target region (and any other target region considered) can be accumulated as a component number or statistical evaluation, thereby measuring the overall degree of evidence supporting the conclusion that cancer DNA is present in the test sample. In some embodiments, the comparisons are combined to produce a cumulative statistical evaluation, indicating the probability or possibility that cancer DNA is present in the test sample. Various methods can be used to create a cumulative statistical estimate, including combining statistical measures (e.g., joint probability, joint likelihood, or joint likelihood ratio) or otherwise combining (e.g., summing, averaging) the results for each target region to identify the presence or absence of cancer DNA in the test sample.

[0123] In some embodiments, the combination includes calculating the average of each comparison. In one embodiment, the average is a weighted average. For example, a comparison of a first target region from a category with two or more PVs may be assigned a weight of 1.0, while a comparison of a second target region from a category with a single SNV may be assigned a weight of 0.5. In this way, two or more PV categories provide additional weights because the probability of observing two genetic variations simultaneously on a single sequence read is less likely to be the result of an error. Therefore, in one embodiment, combining the comparison of at least the first target region and the second target region (step 108) may further include adjusting each comparison by a weight. In some embodiments, the weight of the category of two or more PVs is 1.0. In some embodiments, the weight of the category of a single SNV is 0.5. In some embodiments, each comparison may be performed using a different calculation, such as by using a different error model or statistical technique to evaluate each target region. For example, in some embodiments, the first target region of the category with a single SNV uses an error model derived from a binomial distribution, while the second target region of the category with two or more PVs uses an error model derived from a multinomial distribution. In such an embodiment, different weights may be applied to each type of comparison so that the evidence supporting a call for the presence of cancer is proportional to the statistical evaluation made.

[0124] In some embodiments, the comparison includes generating a statistical evaluation, such as a p-value, describing the probability or likelihood of the presence of a genetic variation in the test sample. In these embodiments, the p-values ​​for each target region can be combined using, for example, Fisher's method. If U is distributed as Uniform(0,1), then -2logU is distributed as a chi-square with 2 degrees of freedom. If X1,…,X k Independent distribution is Then X1+…+X k Distribution Since p1,…,p k The independent distribution is Uniform(0,1), so the p-value of the combination is:

[0125]

[0126] For example, if there are three target regions with independent statistical estimates (eg, p-values) of 0.145, 0.263, and 0.081, the combined statistical estimate may be:

[0127]

[0128] As shown, this provides moderate evidence against the null hypothesis.Other methods for combining statistical estimates from multiple independent statistical tests can be found at least in Jiang and Wong, Open Journal of Statistics Vol. 5 No. 01 (2015), which is hereby incorporated by reference in its entirety.

[0129] In some embodiments, the statistical evaluation may include a likelihood ratio, which indicates the likelihood that a genetic variation is present in a test sample. Individual likelihood ratios may be combined into a cumulative LR score for all target regions of a sample (LR equal to the sum of the log-likelihoods). i In these embodiments, the likelihoods or likelihood ratios can be combined by finding the sum of the log-likelihoods for each individual event (i.e., the likelihoods calculated for each target region). In this way, when the log-likelihood is used to estimate the parameters of the maximum likelihood estimate, each data point is used by adding to the total log-likelihood. Therefore, each data point is evidence in support of the estimated parameter, where the addition of each data point increases independent evidence to identify the presence or absence of cancer DNA in the sample.

[0130] In one embodiment, the log-likelihood or log-likelihood ratio can be determined for each genetic variation and then combined, for example, by summing. As shown in the following formula (6), the likelihood ratio for the entire sample can be determined by summing the log-likelihood ratios for each target region (1..V) considered:

[0131]

[0132] Similarly, we can determine the sum of the log-likelihoods for H0 and H1 for each variant and then find the likelihood ratio:

[0133]

[0134] In this way, the score for each individual target region is not simply compared to a threshold individually. Instead, the score for each target region in the assay is considered jointly. This provides a further benefit because evidence from target regions of different categories can also be considered jointly.

[0135] The embodiments described herein can be extended to any practical number of target regions by jointly considering evidence from multiple target regions. Figure 3The assay 300 of the present disclosure includes 30 target regions, but embodiments of the present disclosure can be extended to 1-100 target regions, 100-200 target regions, 200-300 target regions, 300-400 target regions, or 400-500 target regions. In some embodiments, the number of target regions is 500-1000 target regions, 1000-2000 target regions, 2000-3000 target regions, 3000-4000 target regions, or 4000-5000 target regions. In some embodiments, the number of target regions is 5000-10,000 target regions, 10,000-20,000 target regions, 20,000-30,000 target regions, 30,000-40,000 target regions, or 40,000-50,000 target regions. In some embodiments, the number of target regions is proportional to the cancer type. For example, melanoma may have up to one million SNVs, while most cancers have about 100,000 SNVs. Thus, in these embodiments, the number of target regions can include 50,000-100,000 target regions, 100,000-200,000 target regions, 200,000-300,000 target regions, 300,000-400,000 target regions, 400,000-500,000 target regions, or 500,000-1,000,000 target regions. In any embodiment, the number of target regions is at least 2, at least 4, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, or at least 5,000 target regions. In many embodiments, 2-200 (e.g., 6-100) target regions can be examined.

[0136] Identifying cancer DNA in a test sample based on a comparison of a combination of a first target region and a second target region (step 110) can be performed in a variety of ways. In some embodiments, identifying cancer DNA in a test sample may include comparing the cumulative evaluation of multiple target regions with a threshold value. In any embodiment, it is possible to identify or otherwise consider that there is cancer DNA in a test sample based on a cumulative evaluation. For example, if the cumulative evaluation exceeds a threshold value, cancer DNA can be identified in a test sample. An appropriate threshold value can be determined empirically, such as by sequencing a test sample previously known to have cancer DNA, and then selecting a threshold value with high sensitivity (i.e., detection capability) and high specificity (i.e., ability to distinguish). In such embodiments, a threshold value can be determined by running at least 10, at least 100, at least 1000, or at least 10,000 samples containing non-cancerous DNA (or at least unknown to have cancerous DNA) in a determination and selecting a threshold value higher than the signal identified in a control sample or making the false positive rate determined using a control sample estimated to be 1% or less, 0.1% or less, or 0.01% or less. The sample being run may be from the same patient or from different patients. For example, running 200 samples may involve sampling from 20 healthy donors (assuming no cancer) and running 10 assays for each patient to reach 200 samples. For each control sample, likelihood ratio analysis can be applied to give the overall likelihood ratio of healthy patients. Calculating the likelihood ratios of all samples that have been run will result in a series of likelihood ratios for healthy patients, and a threshold value can be set above the highest likelihood ratio. The threshold value can be calculated in advance from a healthy donor pool, so it will not vary from patient to patient. Obviously, the method according to the present disclosure may also include identifying the patient as having cancer if the result is equal to or higher than the threshold value, and, for example, treating the patient. In these embodiments, the patient may have previously received a first therapy. In these cases, the method includes performing a second therapy different from the first therapy on the patient.

[0137] As previously described, the inventors have recognized and appreciated that different categories of target regions, including different kinds of variations, can provide various levels of evidence to support the identification of cancer DNA in a test sample. For example, phase variations and INDELs (particularly INDELs longer than 1 nucleotide, such as 2, 3, 4, 5 nucleotides or more, preferably 2 nucleotides or 3 nucleotides) observed within a target region are highly unlikely to be caused by background errors. In contrast, single nucleotide variations may provide relatively little evidence for concluding that cancer DNA is present because these types of variations are more likely to be the result of an error. Therefore, assays according to the present disclosure can combine results from multiple categories of target regions to take advantage of all tumor-associated variants that may be present, thereby producing a highly sensitive and highly specific cancer DNA detection assay.

[0138] In some embodiments, the test sample can be divided into two or more aliquots, which can be processed in the same manner according to the embodiments of the present disclosure. For example, each aliquot can be similarly enriched, measured, and compared for certain target areas, and these comparisons can be combined to produce a high-confidence identification of cancer DNA present in the test sample. More information about packaging and using duplicate samples can be found in co-owned International Patent Application No. PCT / IB2022 / 051195, which was filed on February 10, 2022, and is hereby incorporated by reference in its entirety.

[0139] In any embodiment, the variant allele fraction (VAF) of the test sample can be determined. VAF can be determined, for example, by the number of sequence reads that support the presence of the target region category. The terms "variant allele fraction," "estimated variant allele fraction," "VAF," or "eVAF" refer to the estimated allele fraction of variant cancer DNA in the test sample.

[0140] In any embodiment, the amount of cancer DNA in the test sample can be quantified. Quantification can include an estimated variant allele fraction. In some embodiments, the estimated allele fraction can include the mean or median of the variant allele fraction of each target region in which the category of the target region is determined to exist. In some embodiments, the estimated variant allele fraction can include the mean (k / n) of the variant allele fraction of each variation. In the case where the variation level is low and the result is random, this may be preferred, so including evidence from all variations can produce more realistic measurements. The quantified cancer DNA can be compared with the quantified cancer DNA from one or more additional samples, such as comparing the quantified cancer DNA from the sample obtained from the patient during at least the first time point and the second time point, wherein the first time point is before treatment, and the second time point is after treatment. Similarly, individual variations or variation groups in samples from different time points can be tracked.

[0141] In any embodiment, the method described herein may be performed on a test sample obtained from a patient during at least a first time point and a second time point, wherein the first time point is before treatment, the second time point is after treatment, and the method includes determining whether the amount of cancer DNA or the range of possible cancer DNA amount changes between the first and second time points. In any embodiment, more samples may be obtained at additional time points, for example, additional samples are collected monthly, every two months, every quarter or every year after the second time point. This embodiment can be used to monitor whether the treatment implemented on the patient is effective. Point estimates, confidence intervals or both can be used to determine the changes in cancer DNA over time, wherein a significant (e.g., statistically significant) reduction indicates that the treatment is effective, and no significant change or increase indicates that the treatment is ineffective. This embodiment can also be used to monitor whether cancer recurs after surgery for the purpose of cure. Point estimates, confidence intervals or both can be used to determine the changes in cancer DNA over time, wherein no detectable cancer DNA indicates that the cancer has not recurred, and a significant change or increase indicates that the cancer may increase. In these cases, changes of at least two times, at least four times, at least six times, at least eight times or at least ten times can be considered significant. In these cases, a change of at least 20%, at least 30%, at least 50%, at least 70%, or at least 90% can be considered significant. In some embodiments, a change is considered significant if the change is greater than a threshold value (e.g., 50%) and the confidence intervals when quantifying the cancer DNA at the first and second time points do not overlap. In these embodiments, a significant decrease indicates that the treatment is effective, while no significant change or increase indicates that the treatment is ineffective.

[0142] In some embodiments, the method of the present disclosure may further include providing a report indicating whether cancer DNA is present in the sample. In some embodiments, the report may include a likelihood ratio or score (or another number representing it) as described above, and a threshold value that can be compared with the likelihood ratio to determine whether the test sample contains cancer DNA. If the report indicates that there is no cancer DNA in the sample, but the likelihood ratio or score or another number representing it is close to the threshold, the report may suggest arranging a follow-up test as soon as possible (e.g., within one month, two months, or three months) to reassess whether the value now exceeds the threshold value for determining whether the sample contains cancer DNA. In some embodiments, the report may additionally list approved (e.g., FDA or EMA approved) therapies for treating the disease or residual disease, such as chemotherapy or immunotherapy. This information can help diagnose the disease (e.g., whether the patient suffers from MRD) and / or the treatment decision made by the doctor.

[0143] In some embodiments, a sample can be collected from a patient at a first location (e.g., in a clinical setting such as a hospital or doctor's office), and the sample can be forwarded to a second location (e.g., a laboratory) where the sample is processed and the above method is performed to generate a report. A "report" as described herein is an electronic or tangible document that includes report elements that provide test results that may indicate the presence and / or amount of cancer DNA in the sample. Once generated, the report can be forwarded to another location (which can be the same location as the first location) where it can be interpreted by a health care professional (e.g., a clinician, laboratory technician, or physician such as an oncologist, surgeon, pathologist, or virologist) as part of a clinical decision.

[0144] The patient whose sample is analyzed in this method may suffer from any type of cancer, or may have previously received treatment for any type of cancer. For example, the patient may suffer from or may have suffered from melanoma, carcinoma, lymphoma, sarcoma or glioma. For example, cancer may be melanoma, lung cancer (such as non-small cell lung cancer), breast cancer, head and neck cancer, bladder cancer, Merkel cell carcinoma, cervical cancer, hepatocellular carcinoma, gastric cancer, squamous cell carcinoma of the skin, classical Hodgkin lymphoma, B cell lymphoma, colorectal cancer, pancreatic cancer, gastric cancer or breast cancer, including other solid tumors and blood cancers. In some embodiments, cancer is a type of cancer, which on average shows at least 0.1 mutations per megabase, at least 0.2 mutations per megabase, at least 0.5 mutations per megabase, at least 1 mutation per megabase, and an average mutation rate of at least 10 mutations per megabase. In some embodiments, cancer is a cancer showing an average mutation rate of at least 0.5 mutations per megabase. Methods for calculating mutation rates are known in the art (eg, Schumacher TN, Schreiber RD. Neoantigens in cancer immunotherapy. Science. 2015; 348(6230): 69-74, which is hereby incorporated by reference in its entirety).

[0145] In some embodiments, the method can be used to guide treatment decisions. In some embodiments, the method can be used to identify whether a patient should be treated again, for example, with the same therapy or a second therapy. For example, if a patient has previously been treated with a first cancer therapy and the patient is determined to have MRD using the method of the present invention, the patient can be treated with a second cancer therapy that is the same or different from the first cancer therapy. For example, if a patient has previously been treated with surgery or an immune checkpoint inhibitor and the patient is identified as having MRD, the patient can be treated with further surgery, the same or different immune checkpoint inhibitor, or other types of therapy, wherein immune checkpoint therapy includes administration of CTLA-4, PD1, PD-L1, TIM-3, VISTA, LAG-3, IDO or KIR checkpoint inhibitors, and other types of therapy include, for example, (a) anthracycline therapy (e.g., by administration of daunomycin, doxorubicin or mitoxantrone), (b) alkylating agent therapy (e.g., by administration of nitrogen mustard, cyclophosphamide, Ifosfamide, melphalan, cisplatin, carboplatin, nitrosoureas, dacarbazine and procaine or busulfan), (c) topoisomerase II inhibitor therapy (e.g., by administration of etoposide or teniposide), (d) bleomycin therapy, (e) antimetabolite therapy (e.g., by administration of methotrexate, 5-fluoropyrimidine (5-fluorocil), cytarabine, 6-mercaptopurine or 6-thioguanine), (f) vinca alkane therapy (e.g., by administration of vincrisene or vinblastine), (g) steroid therapy (e.g., by administration of prednisone or dexamethasone and (h) radiation therapy, etc. Alternative therapies include targeted therapies and non-targeted chemotherapy, where targeted therapies include treatment with erlotinib (Tarceva), afatinib (Gilotrif), gefitinib (Iressa), or osimertinib (Tagrisso), which can be administered to patients with activating mutations in EGFR, crizotinib (Xalkori), ceritinib (Zykadia), alectinib (Alecensa), or brigatinib (Alunbig), which can be administered to patients with ALK fusions, crizotinib (Xalkori), entrectinib (RXDX-101), lovatinib (Vegas), or seleniformis (Salutamide), which can be administered to patients with ALK fusions. Latinib (PF-06463922), crizotinib (Xalkori), entrectinib (RXDX-101), lorlatinib (PF-06463922), ripretinib (TPX-0005), DS-6051b, ceritinib, ensartinib, or cabozantinib, which can be administered to patients with ROS1 fusions, or dabrafenib (Tafinar) or trametinib (Mekinist), which can be administered to patients with activating mutations in BRAF. Many other actionable mutations are known.If the patient is to be switched to non-targeted chemotherapy, the therapy may be, for example, platinum-based double chemotherapy (wherein the platinum-based double chemotherapy may include a platinum-based agent selected from cisplatin (CDDP), carboplatin (CBDCA) and nedaplatin (CDGP)) and a third generation agent (selected from docetaxel (DTX), paclitaxel (PTX), vinorelbine (VNR), gemcitabine (GEM), irinotecan (CPT-11), pemetrexed (PEM) and tegafur gimeracil oteracil (S1)).

[0146] Described herein are methods of diagnosing cancer comprising performing a method of detecting cancer DNA in a test sample obtained from a patient according to any of the methods disclosed herein.

[0147] Described herein are methods of treating cancer in a patient, comprising determining the presence or absence of cancer DNA detected in a test sample obtained from the patient according to any of the methods disclosed herein, and administering a cancer therapy or treatment to the patient, or recommending administration of a cancer therapy or treatment to the patient. Administration or recommendation is based on identification of cancer DNA in the test sample. For example, if cancer DNA is detected, a therapy or treatment may be administered or recommended.

[0148] Methods of treating cancer in a patient are described herein, wherein the patient has been diagnosed as having cancer or suspected of having cancer based on the presence or absence of cancer DNA detected in a test sample obtained from the patient as determined according to any of the methods disclosed herein. The method includes administering cancer therapy or treatment to the patient based on the identification of cancer DNA detected in a test sample obtained from the patient. In some embodiments, the method may alternatively include recommending cancer therapy or treatment to the patient based on the identification of the presence of cancer DNA in a sample obtained from the patient.

[0149] Methods for determining the effectiveness of cancer treatment or therapy are described herein, and include administering cancer treatment or therapy to a patient, obtaining a test sample from the patient, and determining the presence, absence or amount of cancer DNA in the test sample according to any method disclosed herein. In some embodiments, the method may include the following steps: obtaining a test sample from the patient before administering cancer treatment or therapy, and comparing the presence, absence or amount of cancer DNA in the test sample obtained before administering cancer treatment or therapy with the presence, absence or amount of cancer DNA in the test sample obtained after administering cancer treatment or therapy. The difference may indicate the effectiveness of cancer therapy or therapy. For example, an increase in the amount of cancer DNA may indicate that cancer therapy or therapy is ineffective. Therefore, the method may include administering an alternative and / or additional cancer therapy or therapy to the patient or recommending an alternative and / or additional cancer therapy or therapy to the patient. On the contrary, a reduction or disappearance of cancer DNA in the test sample (i.e., a significant disappearance, i.e., below the detection level of the method) may indicate that cancer therapy or therapy is effective. Therefore, the method may include continuing or stopping administering cancer therapy or therapy to the patient, or suggesting continuing or stopping cancer therapy or therapy. In some embodiments, the method can include monitoring the effectiveness of a cancer therapy or treatment by performing a cancer DNA detection method using test samples from the patient taken at least two time points during the administration of the cancer therapy or treatment (e.g., test samples obtained during one or more days, months or years, or other time points disclosed herein).

[0150] The present disclosure also provides a method for detecting or monitoring minimal residual disease (MRD), comprising obtaining or having obtained a test sample from a patient who has received cancer therapy or treatment, and performing a method for detecting cancer DNA in the test sample according to the methods disclosed herein.

[0151] Recommendation regarding treatment or therapy may be implemented in any suitable manner, such as by providing a report containing the recommendation.

[0152] Cancer therapy or treatment can be any suitable therapy. For example, cancer therapy or treatment can be tumor resection. Cancer therapy or treatment can be drug treatment of cancer. In some embodiments, the methods of the present disclosure can be performed on patients who have undergone surgery to remove tumors. In some embodiments, the cancer therapy or treatment administered or recommended after detecting the presence or amount of cancer DNA in a test sample obtained from a patient can be a drug cancer therapy or treatment.

[0153] The methods described herein can be used to monitor treatment. For example, in some embodiments, the method can include using the method to analyze a sample obtained at a first time point, and analyzing a sample obtained at a second time point by the method, and comparing the results, i.e., determining whether there is cancer DNA in the sample or determining whether the amount of cancer DNA or the range of possible cancer DNA amounts between the first and second time points has changed. In some embodiments, point estimates or confidence intervals can be used to determine this change, and a significant reduction may indicate that the treatment is effective, while no significant reduction or increase may indicate that the treatment is ineffective. The first and second time points can be before and after treatment, or two or more time points after treatment. For example, by comparing the results obtained from one time point with the results of another time point, the method can be used to determine whether the previously identified changes no longer exist in the subject during the treatment process, whether they have been reduced or increased. The time period between the first and second time points can be at least one month, at least 6 months, or at least one year, and in some cases, the patient can be tested regularly, such as every three months, every six months, or every year, for several years, such as 5 years or more. In another embodiment, the method can be used to assess the effectiveness of treatment by monitoring the patient's ctDNA levels at several time intervals after treatment administration. For example, if treatment is effective, ctDNA levels should be increased due to apoptosis of cancer cells shortly after administration, and then significantly reduced as ctDNA degrades. In such an embodiment, the time period between treatment administration and the first time point can be, for example, at least 15 minutes, at least 30 minutes, at least 45 minutes and at least one hour. In such an embodiment, the time period between the first and second time points can be, for example, every 15 minutes, every 30 minutes, every 45 minutes, every hour, every two hours or several hours (e.g., 8 hours or longer) per hour.

[0154] The method according to the present disclosure can also be used to determine whether the subject is disease-free or whether the disease has recurred. As described above, the method can be used to analyze minimal residual disease and recurrence detection. In these embodiments, the primer pairs used in the method can be designed to amplify sequences containing genetic variations previously identified in the patient's cancer by sequencing cancer material, cfDNA at an earlier time point, or sequencing another suitable sample.

[0155] In some embodiments, when testing minimal residual disease or recurrence detection, the DNA test sample from the patient will be cell-free DNA. This cell-free DNA can be collected from the patient at any time point after treatment. In some embodiments, if the cancer is successfully treated, this cell-free DNA can be collected at any time point when the remaining ctDNA from the cancer has been cleared. This time point may depend on factors such as the initial amount of ctDNA and the treatment method. For methods such as surgery that remove all tumors at once, the time point may be 1 week, 2 weeks, 3 weeks, or 4 weeks after treatment for the purpose of cure. When treatment can remove cancer more gradually, these time points may be longer, such as 1 month or 2 months.

[0156] In some embodiments, the method can be used for clinical trials. For example, the methods described herein may potentially be used to identify a specific patient group for clinical registration or to evaluate the efficacy of a new drug (e.g., a neoadjuvant therapy or adjuvant therapy that may be non-specific or targeted to a patient's cancer, or any combination therapy). In some embodiments, the ctDNA amount in the patient's blood can be estimated at multiple time points, thereby allowing changes in the drug dose administered to the patient in the mid-term of the trial. In some embodiments, the ctDNA amount in the patient's blood can be estimated at multiple time points during clinical trials, and is used to determine whether a specific therapy, treatment level, treatment duration, or treatment type and patient's combination is effective.

[0157] As will be readily appreciated, many of the steps described herein (e.g., sequence processing and generating a report indicating the presence of cancer DNA in a DNA test sample) can be implemented on a computer. Thus, in some embodiments, the method may include executing an algorithm that calculates the likelihood of whether a patient has cancer DNA in a test sample of DNA collected from the patient based on an analysis of the sequence reads, and outputs the likelihood. In some embodiments, the method may include inputting a sequence into a computer and executing an algorithm that can use the input measurements to calculate the likelihood.

[0158] It is obvious that the described computational steps can be computer-implemented, and thus the instructions for executing these steps can be set forth as a program that can be recorded in a suitable physical computer-readable storage medium. Computational analysis can be performed on sequencing reads.

[0159] The method disclosed herein can be a computer-implemented method, i.e., a method performed by a computer or performed on a computer. The present disclosure also provides a computer-readable storage medium storing instructions for performing the method disclosed herein. The computer-readable storage medium can make it possible to implement the method as described above when executed on a computing device. The present disclosure also provides a system comprising one or more computer-readable media, a memory for storing instructions and data units (the data units optionally include one or more error probability distribution models) for performing the method, and a processor for executing instructions.

[0160] Figure 5 An illustrative implementation of a computer system 500 that can be used in conjunction with any embodiment of the present disclosure provided herein is shown. The computer system 500 may include one or more processors 510 and one or more manufactured products, the manufactured products including non-temporary computer-readable storage media (e.g., memory 520 and one or more non-volatile storage media 530). The processor 510 can control writing data to the memory 520 and the non-volatile storage device 530 and reading data from the memory 520 and the non-volatile storage device 530 in any suitable manner, because the aspects of the present disclosure provided herein are not limited in this regard. In order to perform any function described herein, the processor 510 can execute one or more processor-executable instructions stored in one or more non-temporary computer-readable storage media (e.g., memory 520), which can be used as non-temporary computer-readable storage media storing processor-executable instructions for execution by the processor 510.

[0161] The term "program" or "software" is used in this article in a general sense to refer to any type of computer code or processor executable instruction set that can be used to program a computer or other processor to implement various aspects of the embodiments as described above. In addition, it should be understood that according to one aspect, one or more computer programs that perform the methods provided herein when executed do not have to reside on a single computer or processor, but can be distributed in a modular manner between different computers or processors to implement various aspects provided herein.

[0162] Processor executable instructions can have many forms, such as program modules that are executed by one or more computers or other devices. Generally speaking, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In general, the functions of program modules can be combined or distributed as needed in various implementations.

[0163] In addition, the data structure may be stored in one or more non-transitory computer-readable storage media in any suitable form. For ease of illustration, the data structure may be shown as having fields that are related by location in the data structure. This relationship may also be achieved by allocating storage for the fields using locations in a non-transitory computer-readable medium that conveys the relationship between the fields. However, any suitable mechanism may be used to establish the relationship between information in the fields of the data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.

[0164] Furthermore, various inventive concepts may be embodied as one or more processes, examples of which have been provided, including reference to Figure 1 The actions performed as part of each process may be ordered in any suitable manner. Thus, embodiments may be constructed in which the actions are performed in an order different from that shown, which may include performing some actions simultaneously even though shown as sequential actions in illustrative embodiments. Example

[0165] The following examples are intended to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use embodiments of the disclosure and are not intended to limit the scope of what the inventors regard as the disclosure.

[0166] Example 1

[0167] In order for a tumor-informed assay to best determine whether cancer DNA is present, it should take advantage of all the evidence that is practically available. A typical cancer has many single nucleotide variants throughout the genome. There are often more than 10,000 (see Fig. 9 ). Taken individually, if these SNVs are located and observed in a test sample when sequenced, they can provide some evidence of ctDNA, although there is a risk that any signal is actually an error, such as DNA damage, polymerase errors, or sequencer errors.

[0168] When two variants (e.g., 2 SNVs) are close to each other and on the same DNA strand (e.g., a cell-free DNA molecule), then if both variants are seen in one or more sequencing reads, this provides more evidence that the sequencing signal is not a false positive (e.g., DNA damage, polymerase errors, or sequencer errors) and that cancer is actually present. This is because the probability of having two errors matching two exact variants on a single DNA molecule or sequencing read is extremely low.

[0169] Although these variations on the same DNA molecule provide much more information than any single SNV, they are much fewer in number. Typically, there are 10-1000 such phase variants in many solid cancers (see Fig. 9 ).

[0170] The typical cancer genome is full of additional changes. This includes germline changes, epigenetic changes, and structural variations. Again, as an example, some cancers have a large number of structural variations, while others have very few (see Fig.10 ).

[0171] So an assay that can take advantage of all of this information always has the potential to be better than one that only looks at individual SNVs or only looks at phase variations. The challenge is how to combine this information. Targeted region-based approaches, where the genome is broken down into multiple parts, then target regions are selected (a target region is, for example, the sequence between two primers in a PCR product), and then each target region is evaluated based on all of the genetic variation within it, and all of this information is combined when sequencing to determine if cancer DNA is present.

[0172] Such assays are also more versatile. Assays that rely only on, for example, phase variation may in some cases have very little targeting information. For example, Fig. 9 Some osteosarcoma patients in the study had <10 phase variants. However, the same patient always had a large number of SNVs, and osteosarcoma patients also often had a large number of structural variants ( Fig.10 ). The combined information of the target region-based decision method can achieve consistent and highly sensitive cancer DNA detection in patients.

[0173] Example 2

[0174] In order to design the best MRD assay, the system is designed to query as many high-quality regions as possible. If, for example, a region is more easily distinguished from noise when cancer is present (e.g., by having 2 phase SNVs), and is easier to amplify and sequence than other regions, then the region may be considered to be of higher quality. To this end, a tumor biopsy sample is first obtained, macroscopically dissected, 50% of the tumor content is targeted, exome capture is performed, and then the sample is sequenced using an Illumina sequencer. All potential variations are identified using the standard Illumina process, and then a combined score is given based on 1) the possibility of being true, 2) the possibility of somatic variation, 3) the background error rate of variation, 4) the high signal background error rate, 5) the probability of clonality, 6) the amplification level of variation or the copy number gain. The genome is divided into 50bp windows, which overlap 25bp. Each window is given a combined score that includes 1) the scores of all variants present within the window, 2) a score for the ability to uniquely align regions (regions that cannot be uniquely aligned are penalized, and the greater the number of misalignments, the higher the penalty), and 3) a score for the ability to amplify and sequence regions (penalties are given for features known to pose challenges for sequencing, including repeats). The regions are then sorted by score, and the top 100 are selected to design PCR primers for them. If two overlapping regions are in the top 100 list, the region with the highest score is retained and the region with the lower score is discarded. The 101st region is then added to the list, and so on. Multiplex PCR is designed for the first 48 regions. Computer simulated PCR is performed using all primer pairs. When a primer combination is identified that produces ≥2 non-specific regions, the primer for the region with the lowest score that results in that non-specific product is discarded, and an alternative primer is designed. If the non-specific PCR problem cannot be overcome, the region is discarded and the next region is added to the primer design.

[0175] One challenge facing this tumor-informed approach to cancer DNA in test samples is the number of regions that can be robustly and cost-effectively targeted. This strategy for sorting regions can maximize the number and quality of regions successfully detected in test DNA samples. When the variants are phase variants (PVs), that is, when the variants are cis, adjacent to each other and located on the same chromosome, they can be read together, which increases the ability to separate signal from noise. When the variants are trans but can still be read using the same primer pair (or other targeting reagents, such as baits), the amount of information from a single target region can be doubled. The method should also limit the number of reads wasted on nonspecific products.

[0176] Example 3

[0177] In order to detect cancer DNA in test samples with high sensitivity, it is advantageous to target multiple regions and multiple regional categories. For some cancer types, only one category of the targeted region is sufficient. However, the inventors have recognized and realized that it is better to target multiple target region categories containing different kinds of genetic variations. In this example, a large number of structural variations (SVs) are identified for some breast cancer patients, while more SNVs and INDELs are present in other patients. In addition, the regions of many patients have multiple somatic genetic and epigenetic changes. In addition, they also have many regions with both somatic and germline changes. A large panel is designed to sequence breast cancer tumor DNA to evaluate somatic SNVs, INDELs and SVs and germline changes. The best target region is identified according to the method disclosed herein. Primers are designed to target these regions. If the target region contains 1 or more SNVs or INDELs, the primers are designed to be located on both sides of all SNVs and insertions / deletions. When identifying that the target region contains rearrangements (such as SVs), two different parts of the same chromosome or two different chromosomes will be introduced together. The rearrangement sequence is used for primer design, with one primer located 3' and the other 5' to the rearrangement. In cases where the SNV, INDEL, or other genetic variation (e.g., EV) is in cis with the rearrangement, primers are designed to flank the rearrangement and other variation using the rearranged sequence obtained from the tumor. In cases where the target region for identification contains a pair of phase variations (PVs), primers are designed to flank the 5' PV and 3' PV. The advantage of this approach is the ability to consistently obtain a large number of regions to assess cancer DNA in a test sample.

[0178] As each of these different categories are incorporated, a different error model may be required for each type. The inventors have realized that the results from these different error models can be combined in a principled way, such as by summing the log-likelihoods for each target region category, to reach a high confidence conclusion that cancer DNA is in the test sample.

[0179] Example 4

[0180] Fig. 6A -B illustrates why it can be challenging to judge a sample as containing cancer DNA, especially for test samples with low tumor fractions. Fig. 6A As shown in (top), in test samples with high relative tumor fractions (TFs), cancer DNA can be easily called because most, if not all, target regions will contain multiple cancer DNA molecules, resulting in high signals (e.g., likelihood) across multiple target regions, thereby eliminating most false positives and false negatives. Figure 6B(Bottom) shows that samples with low tumor fractions are more difficult to judge because the data for individual regions may not be sufficiently distinguishable from the background error rate. In addition, such low levels of input DNA may not contain cancer DNA in some target regions, so that for many target regions, once amplified, they will not produce true signals but true negatives (see Figure 6B For example, if the assay tests for multiple SNVs and only one cancer DNA molecule is present for some SNVs but not for others, it may not be possible to call any one region positive, but by combining information from different target regions (i.e., SNVs), a confident call is more likely. Figure 6B In the 2 regions containing SNVs, each has a mutant molecule (see Figure 6B ). Individually, none of these SNVs showing a small amount of signal would provide enough evidence to make a positive call. Together they provide more evidence, but in this example they still do not provide enough support to make a confident call for cancer DNA as a whole. By including not only target regions with a single SNV, but also target regions with multiple phase variants and target regions with other variants such as MNVs or INDELs, cancer DNA is more likely to be detected when it is present. Figure 6B In 3 of the 10 regions, there were 2 PVs. For two of these regions, there was no cancer DNA and no signal (see Figure 6B However, for the third target region containing 2PV, there is a single cancer DNA molecule (see Figure 6B For this reason, even reading a single cancer molecule can be very informative, as 2 phase variants can be seen on the same sequencing read, providing much more evidence ( Figure 6B ). When information from this 2PV target region is combined with information from the two target regions where individual SNVs are present (light grey squares), it can make the call more sensitive. In this example, the region with the indel provides further supporting information (presence of 2 / 3 INDELs - dark grey squares vs white squares). However, in order to utilize this information and confidently determine if cancer DNA is present, a method is needed to combine information from different regions as described in the present invention.

[0181] Example 5

[0182] Figure 7An implementation scheme for combining evidence across multiple regions is shown. For dilute samples (<<0.1% tumor fraction), the fraction of mutant reads for individual target regions of each sample is not expected to be close to the overall tumor fraction due to dropout effects. For example, many target regions will show zero variant molecules. Instead, the effect of n / input reads is modeled as a discrete distribution. In this example, the tumor fraction is not measured directly. Instead, it is marginalized over all possible inputs, which provides an accurate estimate of the tumor fraction of the sample. Specifically, instead of guessing the number of cancer DNA molecules containing variants in the target region, the probabilities of all possible values ​​are calculated based on the following factors: (i) the number of sequence reads with genetic or epigenetic variants or a combination of multiple expected genetic variants in the target region (which will vary by target region); (ii) the total number of sequencing reads; (iii) the input number of DNA molecules; and (iv) the estimated background error rate for each target region category, and the value with the highest probability is identified from them. This avoids making assumptions. As shown in FIG6 , the inclusion of several PV regions (including 3PV regions) and several INDEL regions provides a large amount of evidence that, when considered together according to any of the methods of the present disclosure, can support a high confidence conclusion that cancer DNA is present in the test sample, such as by comparing the number of sequence reads for the target region with different error models for each target region category. The accuracy of the mathematical model can be verified by comparing with actual dilution data in the ground truth line map ( Figure 8 ).

[0183] Now that the present method has been described in more detail, it should be understood that the present disclosure is not limited to the specific embodiments described, as these embodiments may of course vary. It should also be understood that the terms used herein are only used to describe specific embodiments and are not intended to be limiting. Although any methods and materials similar or equivalent to those described herein may also be used in the practice or testing of the present disclosure, preferred methods and materials have been described.

[0184] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent was specifically and individually indicated to be incorporated by reference and are incorporated herein to disclose and describe the methods and / or materials in connection with the cited publication.

[0185] It must be noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should also be noted that the claims can be drafted to exclude any optional elements. As such, this statement is intended to serve as antecedent basis for using such exclusive terminology as "only," "solely," and the like when reciting claim elements, or for using a "negative" limitation.

[0186] It will be clear to those skilled in the art after reading this disclosure that each individual embodiment described and illustrated herein has discrete components and features that can be easily separated or combined with the features of any of the other several embodiments without departing from the scope or spirit of the disclosure. Any described method can be performed in the order of events described or in any other order that is logically possible.

[0187] Having described in detail several embodiments of the technology described herein, various modifications and improvements will readily occur to those skilled in the art. Such modifications and improvements are intended to be within the spirit and scope of the present disclosure. Therefore, the foregoing description is intended only as an example and is not intended to be limiting. These technologies are limited only by the definitions of the following claims and their equivalents.

Claims

1. A method for detecting cancer DNA in a test sample from a patient, the method comprising: (a) enriching or enriching a test sample for a plurality of target regions, the plurality of target regions comprising a first target region having a first category and a second target region having a second category; (b) measuring or having measured the multiple target regions of step (a) in the enriched test sample; (c) for each of the first target region and the second target region, comparing or having compared the measurements of step (b) supporting the presence of the class of the target region to one or more error models that model the probability of observing the class of the target region in DNA that does not contain the class of the target region; (d) combining or having combined the comparison of step (c) for at least the first target area and the second target area; and (e) identifying or having identified cancer DNA in the test sample based on the combined comparison of step (d).

2. The method of claim 1, wherein the enrichment of step (a) comprises amplifying multiple target regions by polymerase chain reaction (PCR) to produce PCR products, and wherein the measurement of step (b) comprises sequencing the PCR products or their progeny to produce multiple sequence reads.

3. The method of claim 1, wherein the enrichment of step (a) comprises contacting the test sample with an oligonucleotide pool, wherein the oligonucleotide pool comprises oligonucleotides that are substantially complementary to a plurality of target regions.

4. The method of any one of claims 1-3, wherein the cancer is a solid tumor, and the plurality of target regions are identified by sequencing the solid tumor.

5. The method of claim 4, wherein sequencing of solid tumors is performed by targeted sequencing or whole exome sequencing (WES).

6. The method of any one of claims 1 to 5, wherein: (i) in step (b), the measuring comprises sequencing the plurality of target regions of step (a) to generate a plurality of sequence reads corresponding to the first target region and the second target region; as well as (ii) in step (c), the comparison comprises comparing the number of sequence reads supporting the presence of a class of the target region to one or more error models that model the probability of observing the class of the target region in DNA or RNA that does not contain the target region.

7. The method of any one of claims 1-6, wherein in step (c), the comparison comprises comparing the number of sequence reads that do not support the presence of the target region to one or more error models that model the probability of observing the class of target region in DNA or RNA that does not contain the target region.

8. The method of any one of claims 1-7, wherein the one or more error models are based on a background error rate for each of the first category of target regions and the second category of target regions.

9. The method of any one of claims 1-8, further comprising training one or more error models based on a set of control samples.

10. The method of any one of claims 1-9, wherein the one or more error models in step (c) include a first error model for a first target region and a second error model for a second target region.

11. The method of any one of claims 1-10, wherein the first error model for the first target region comprises a β-binomial model and the second error model for the second target region comprises a multivariate β-binomial distribution.

12. The method of claim 11, wherein the multivariate beta-binomial distribution is a standard Dirichlet distribution or a generalized Dirichlet distribution.

13. The method of any one of claims 1 to 12, wherein the category of the target region is associated with the type of genetic variation within the target region.

14. The method of any one of claims 1-13, wherein the first category of the first target regions are regions comprising a single genetic variation, and the second category of the second target regions are regions comprising two or more genetic variations.

15. The method of claim 14, wherein the first category of the first target region is a single nucleotide variation (SNV), and the second category of the second target region comprises a first phase variation (PV) and a second PV.

16. The method of claim 14, wherein the single genetic variation is a single nucleotide variation (SNV), and the two or more genetic variations include a tumor SNV and a germline SNV.

17. The method of any one of claims 14 to 16, wherein: (i) the comparing of step (c) for the first target region comprises comparing the number of sequence reads having the single genetic variation and the total amount of sequence reads for the first target region to a first error model, and (ii) The comparison of step (c) for the second target region comprises comparing the number of sequence reads having two or more genetic variations and the total amount of sequence reads for the second target region to a second error model.

18. The method of any one of claims 14-17, wherein the first error model comprises an error probability distribution that models the probability of observing a single genetic variation in DNA that does not contain the single genetic variation, and the second error model comprises an error probability distribution that models the probability of observing two or more genetic variations in DNA that does not contain the two or more genetic variations.

19. The method of any one of claims 14-18, wherein two or more genetic variations are located within 160 bp of each other.

20. The method of any one of claims 14-19, wherein the two or more genetic variations are separated by at least 1 nucleotide.

21. The method of any one of claims 14-20, wherein one or more error models take into account the distance between two or more genetic variations.

22. The method of any one of claims 15-21, wherein the comparison for the second target region in step (c) comprises comparing the number of sequence reads having both the first PV and the second PV (k1), the number of sequence reads having only the first PV (k2), the number of sequence reads having only the second PV (k3), and the number of sequence reads having neither the first PV nor the second PV (k4) to one or more error models.

23. The method of any one of claims 1-22, wherein the comparison of step (c) comprises likelihood or log likelihood, and wherein the combining of step (d) comprises adding the comparison of step (c) for the first target region to the comparison of step (c) for the second target region.

24. The method of any one of claims 1-23, further comprising calculating a variant allele fraction (VAF) for each of the first and second target regions based on the measurement results in step (b).

25. The method of any one of claims 1-24, further comprising the step (f) of determining whether cancer DNA is present in the test sample.

26. The method of any one of claims 1-25, further comprising the step of providing a report.

27. The method of any one of claims 1-26, further comprising treating the patient based on the identification of the cancer DNA in the test sample of step (e) or the determination of step (f).

28. The method of any one of claims 1 to 26, further comprising: A. obtaining a second test sample from the patient at a second time point after the test sample; B. enriching the second test sample for multiple target regions; C. measuring the plurality of target regions from the enriched second test sample of step B; D. for each of the first target region and the second target region, comparing the measurement results of step C supporting the presence of the class of the target region to one or more error models that model the probability of observing the class of the target region in DNA that does not contain the class of the target region; E. combining the comparison of step D for at least the first target region and the second target region; and F. Identifying cancer DNA in the second test sample based on the combined comparison of step E.

29. The method of claim 28, wherein the method of steps A to F includes the additional features of any one of claims 2 to 27 as applied to steps A to F.

30. The method of claim 28 or claim 29, further comprising administering a cancer treatment or therapy to the patient prior to obtaining the test sample, and determining the effectiveness of the cancer treatment or therapy based on determining whether cancer DNA is present and / or whether the level of cancer DNA is altered in the second test sample.