Compositions and Methods for Separating Cell-Free DNA

By capturing and sequencing the cfDNA of the variable target group and epigenetic target group, the low concentration and heterogeneity problems in cell-free DNA analysis are solved, and efficient and accurate cfDNA isolation and sequencing are achieved.

CN113661249BActive Publication Date: 2025-07-01GUARDANT HEALTH INC
View PDF 36 Cites 0 Cited by

Patent Information

Application Number
CN202080026244.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-31
Filing Date
2020-01-31
Publication Date
2025-07-01
Estimated Expiration
2040-01-31

AI Technical Summary

Technical Problem

The prior art is difficult to develop accurate and sensitive methods for analyzing cell-free DNA (cfDNA), especially in liquid biopsies, due to the low concentration and heterogeneity of cell-free DNA.

Method used

By capturing the cfDNA of the sequence variable target group and the epigenetic target group, the cfDNA was separated into different target group using the target-specific probe group, and during the sequencing process, the cfDNA of the sequence variable target group was captured with a higher capture yield than the cfDNA of the epigenetic target group to achieve deeper sequencing coverage.

Benefits of technology

Efficient isolation and sequencing of cfDNA is achieved, improving the accuracy and sensitivity of the analysis, and enabling more efficient capture and detection of DNA produced by tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113661249B_ABST
    Figure CN113661249B_ABST
Patent Text Reader

Abstract

The present disclosure relates to compositions and methods for isolating DNA such as cell-free DNA (cfDNA). In some embodiments, the cell-free DNA is from a subject having or suspected of having cancer, and / or the cell-free DNA includes DNA produced by a tumor. In some embodiments, the DNA isolated by the method is captured using a sequence-variable target region set and an epigenetic target region set, wherein the sequence-variable target region set is captured with a higher capture yield than the epigenetic target region set. In some embodiments, the captured cfDNA of the sequence-variable target region set is sequenced to a deeper sequencing depth than the captured cfDNA of the epigenetic target region set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 799,637, filed on January 31, 2019, which is incorporated herein by reference for all purposes. Background of the Invention

[0003] Cancer causes millions of deaths worldwide each year. Early detection of cancer may lead to improved outcomes, as early cancers tend to be more sensitive to treatment.

[0004] Improperly controlled cell growth is a hallmark of cancer, which is typically caused by the accumulation of genetic and epigenetic alterations such as copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions, and / or deletions (indels). Epigenetic variations include 5 - methylation of cytosine (5 - methylcytosine) and the association of DNA with chromatin proteins and transcription factors.

[0005] Biopsy represents a traditional method for detecting or diagnosing cancer, in which cells or tissues are extracted from a potential cancer site and the associated phenotypic and / or genotypic characteristics are analyzed. Biopsy has the disadvantage of being invasive.

[0006] Cancer detection based on the analysis of body fluids such as blood (“liquid biopsy”) is an interesting alternative based on the observation that DNA from cancer cells is released into body fluids. Liquid biopsy is non - invasive (possibly only requiring a blood draw). However, given the low concentration and heterogeneity of cell - free DNA, it is challenging to develop accurate and sensitive methods for analyzing liquid biopsy materials. Isolating the cell - free DNA fraction that can be used for further analysis in liquid biopsy procedures is an important part of the process. Accordingly, there is a need for improved methods and compositions for isolating cell - free DNA (e.g., for liquid biopsy). Summary of the Invention

[0007] The present disclosure provides compositions and methods for isolating DNA, such as cell-free DNA. The present disclosure is based in part on the following recognition. It may be beneficial to isolate cell-free DNA in order to capture two target sets - a sequence variable target set and an epigenetic target set (where the capture yield of the sequence variable target set is greater than the capture yield of the epigenetic target set). In all embodiments described herein involving a sequence variable target set and an epigenetic target set, the sequence variable target set includes regions not present in the epigenetic target set, and vice versa, although in some cases, fractions of the regions may overlap (e.g., a fraction of genomic positions may be present in both target sets). The difference in capture yield can allow, for example, deep and thus more accurate sequence determination in the sequence variable target set during simultaneous sequencing, such as in the same sequencing cell or in the same pool of material to be sequenced, and shallow and broader coverage in the epigenetic target set.

[0008] The epigenetic target set can be analyzed in various ways, including methods that do not rely on a high degree of accuracy in determining the specific nucleotide sequence within the target. Examples include determining methylation and / or the distribution and size of the fragments, which can indicate normal or abnormal chromatin structure in the cells from which the fragments were obtained. Such analysis can be performed by sequencing and requires less data (e.g., the number of sequence reads or the depth of sequencing coverage) compared to determining the presence or absence of sequence mutations, such as base substitutions, insertions, or deletions.

[0009] Compared to the methods described herein, isolating the epigenetic target set and the sequence variable target set with the same capture yield would result in the generation of unnecessary redundant data for the epigenetic target set and / or provide a lower accuracy than desired for determining the genotypes of the members of the sequence variable target set.

[0010] The present disclosure is directed to meeting the need for improved isolation of cell-free DNA and / or providing other benefits. Accordingly, the following exemplary embodiments are provided.

[0011] In one aspect, the present disclosure provides a method for isolating cell-free DNA (cfDNA), the method comprising: capturing more than one target set of cfDNA obtained from a test subject, wherein the more than one target set includes a sequence variable target set and an epigenetic target set, thereby generating a group of captured cfDNA molecules; wherein in the group of captured cfDNA molecules, the cfDNA molecules corresponding to the sequence variable target set are captured with a higher capture yield than the cfDNA molecules corresponding to the epigenetic target set.

[0012] In another aspect, the present disclosure provides a method for isolating cell-free DNA (cfDNA), the method comprising: contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes comprises target-binding probes specific for a set of sequence-variable target regions and target-binding probes specific for a set of epigenetic target regions, and the set of target-specific probes is configured to capture cfDNA corresponding to the set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to the set of epigenetic target regions, thereby forming a complex of the target-specific probes and cfDNA; and separating the complex from cfDNA not bound to the target-specific probes, thereby providing a set of captured cfDNA molecules. In some embodiments, the method further comprises sequencing the set of captured cfDNA molecules. In some embodiments, the method further comprises sequencing cfDNA molecules corresponding to the set of sequence-variable target regions to a deeper sequencing depth than cfDNA molecules corresponding to the set of epigenetic target regions.

[0013] In another aspect, the present disclosure provides a method for identifying the presence of DNA produced by a tumor, the method comprising: collecting cfDNA from a test subject, capturing more than one set of target regions from the cfDNA, wherein the more than one set of target regions comprises a set of sequence-variable target regions and a set of epigenetic target regions, thereby producing a set of captured cfDNA molecules, sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the set of sequence-variable target regions are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the set of epigenetic target regions.

[0014] In another aspect, the present disclosure provides a method for determining the likelihood that a subject has cancer, comprising: a) collecting cfDNA from a test subject; b) capturing more than one set of target regions from the cfDNA, wherein the more than one set of target regions comprises a set of sequence-variable target regions and a set of epigenetic target regions, thereby producing a set of captured cfDNA molecules; c) sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the set of sequence-variable target regions are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the set of epigenetic target regions; d) obtaining more than one sequence read generated by a nucleic acid sequencer by sequencing the captured cfDNA molecules; e) mapping the more than one sequence read to one or more reference sequences to produce mapped sequence reads; and f) processing the mapped sequence reads corresponding to the set of sequence-variable target regions and the set of epigenetic target regions to determine the likelihood that the subject has cancer.

[0015] In some embodiments, the captured cfDNA molecules of the sequence variable target region set are sequenced to a sequencing depth that is at least 2-fold deeper than the captured cfDNA molecules of the epigenetic target region set. In some embodiments, the captured cfDNA molecules of the sequence variable target region set are sequenced to a sequencing depth that is at least 3-fold deeper than the captured cfDNA molecules of the epigenetic target region set. In some embodiments, the captured cfDNA molecules of the sequence variable target region set are sequenced to a sequencing depth that is 4- to 10-fold deeper than the captured cfDNA molecules of the epigenetic target region set. In some embodiments, the captured cfDNA molecules of the sequence variable target region set are sequenced to a sequencing depth that is 4- to 100-fold deeper than the captured cfDNA molecules of the epigenetic target region set.

[0016] In some embodiments, cfDNA amplification includes the step of ligating an adaptor containing a barcode to the cfDNA. In some embodiments, cfDNA amplification includes the step of ligating an adaptor containing a barcode to the cfDNA.

[0017] In some embodiments, capturing more than one target region set of cfDNA includes contacting the cfDNA with a target binding probe specific for the sequence variable target region set and a target binding probe specific for the epigenetic target region set. In some embodiments, the target binding probe specific for the sequence variable target region set is present at a higher concentration than the target binding probe specific for the epigenetic target region set. In some embodiments, the target binding probe specific for the sequence variable target region set is present at a concentration that is at least 2-fold higher than the target binding probe specific for the epigenetic target region set. In some embodiments, the target binding probe specific for the sequence variable target region set is present at a concentration that is at least 4-fold or 5-fold higher than the target binding probe specific for the epigenetic target region set. In some embodiments, the target binding probe specific for the sequence variable target region set has a higher target binding affinity than the target binding probe specific for the epigenetic target region set.

[0018] In some embodiments, cfDNA obtained from a test subject is partitioned into at least 2 fractions based on methylation level, and the subsequent steps of the method are performed on each fraction.

[0019] In some embodiments, the partitioning step includes contacting the collected cfDNA with a methyl-binding reagent immobilized on a solid support.

[0020] In another aspect, the present disclosure provides a collection of target-specific probes for capturing cfDNA produced by tumor cells, the collection comprising target-binding probes specific for a set of sequence-variable target regions and target-binding probes specific for an epigenetic target region set, wherein the capture yield of the target-binding probes specific for the set of sequence-variable target regions is at least 2-fold higher than the capture yield of the target-binding probes specific for the epigenetic target region set. In some embodiments, the capture yield of the target-binding probes specific for the set of sequence-variable target regions is at least 4-fold or 5-fold higher than the capture yield of the target-binding probes specific for the epigenetic target region set.

[0021] In some embodiments, there are at least 10 regions in the set of sequence-variable target regions and at least 100 regions in the epigenetic target region set.

[0022] In some embodiments, the probes are present in a single solution. In some embodiments, the probes comprise a capture moiety.

[0023] In another aspect, the present disclosure provides a system that includes a communication interface that receives, via a communication network, more than one sequence read produced by a nucleic acid sequencer by sequencing a captured set of cfDNA molecules, wherein the captured set of cfDNA molecules is obtained by capturing more than one target region set from a cfDNA sample, wherein the more than one target region set includes a sequence-variable target region set and an epigenetic target region set, wherein the captured cfDNA molecules corresponding to the sequence-variable target regions are sequenced to a deeper sequencing depth than the captured cfDNA molecules corresponding to the epigenetic target region set; and a controller that includes or has access to a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (i) receiving, via the communication network, the sequence reads produced by the nucleic acid sequencer; (ii) mapping the more than one sequence read to one or more reference sequences to produce mapped sequence reads; (iii) processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.

[0024] In some embodiments, the sequencing depth corresponding to the sequence-variable target region set is at least 2-fold deeper than the sequencing depth corresponding to the epigenetic target region set. In some embodiments, the sequencing depth corresponding to the sequence-variable target region set is at least 3-fold deeper than the sequencing depth corresponding to the epigenetic target region set. In some embodiments, the sequencing depth corresponding to the sequence-variable target region set is 4-10-fold deeper than the sequencing depth corresponding to the epigenetic target region set. In some embodiments, the sequencing depth corresponding to the sequence-variable target region set is at least 4-100-fold deeper than the sequencing depth corresponding to the epigenetic target region set.

[0025] In some embodiments, prior to sequencing, the captured cfDNA molecules of the sequence variable target regions are pooled with the captured cfDNA molecules of the epigenetic target regions. In some embodiments, the captured cfDNA molecules of the sequence variable target regions and the captured cfDNA molecules of the epigenetic target regions are sequenced in the same sequencing pool.

[0026] In some embodiments, the epigenetic target regions include hypermethylated variable target regions. In some embodiments, the epigenetic target regions include hypomethylated variable target regions. In some embodiments, the epigenetic target regions include methylation control target regions.

[0027] In some embodiments, the epigenetic target regions include fragmented variable target regions. In some embodiments, the fragmented variable target regions include transcription start site regions. In some embodiments, the fragmented variable target regions include CTCF binding regions.

[0028] In some embodiments, the footprint of the epigenetic target regions is at least 2-fold larger than the size of the sequence variable target regions. In some embodiments, the footprint of the epigenetic target regions is at least 10-fold larger than the size of the sequence variable target regions.

[0029] In some embodiments, the footprint of the sequence variable target regions is at least 25 kB or 50 kB.

[0030] In yet another aspect, the present disclosure provides a composition comprising captured cfDNA, wherein the captured cfDNA comprises captured sequence variable targets and captured epigenetic targets, and the concentration of the sequence variable targets is greater than the concentration of the epigenetic targets, wherein the concentration is normalized for the footprint sizes of the sequence variable targets and the epigenetic targets.

[0031] In some embodiments, the captured cfDNA comprises sequence tags. In some embodiments, the sequence tags include barcodes. In some embodiments, the concentration of the sequence variable targets is at least 2-fold greater than the concentration of the epigenetic targets. In some embodiments, the concentration of the sequence variable targets is at least 4-fold or 5-fold greater than the concentration of the epigenetic targets. In some embodiments, the concentration is a mass / volume concentration normalized for the footprint size of the target regions.

[0032] In some embodiments, the epigenetic target regions include one, two, three, or four of hypermethylated variable target regions; hypomethylated variable target regions; transcription start site regions; and CTCF binding regions; optionally, wherein the epigenetic target regions further include methylation control target regions.

[0033] In some embodiments, the composition is produced according to the methods disclosed elsewhere herein. In some embodiments, the capture is performed in a single container.

[0034] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report can be in paper format or electronic format. For example, information about sequence information as determined by the methods or systems disclosed herein and / or information derived from sequence information can be presented in such a report. In some embodiments, the information is the cancer status of a subject as determined by the methods or systems disclosed herein. The methods or systems disclosed herein can also include the step of transmitting the report to a third party, such as the subject from whom the sample originated or a healthcare practitioner.

[0035] In another aspect, the present disclosure provides a method for determining the risk of cancer recurrence in a test subject, the method comprising: at one or more preselected time points after one or more previous cancer treatments of the test subject, collecting DNA originating from or derived from tumor cells from the test subject diagnosed with cancer; capturing more than one set of target regions from the DNA, wherein the more than one set of target regions includes a sequence variable target region set and an epigenetic target region set, thereby generating a set of captured DNA molecules; sequencing the captured DNA molecules, wherein the captured DNA molecules of the sequence variable target region set are sequenced to a deeper sequencing depth than the captured DNA molecules of the epigenetic target region set, thereby generating a set of sequence information; using the set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at the preselected time point; and determining a cancer recurrence score, the cancer recurrence score indicating the presence or absence of DNA originating from or derived from the tumor cells of the test subject, wherein when the cancer recurrence score is determined to be at or above a predetermined threshold, the cancer recurrence status of the test subject is determined to be at risk of cancer recurrence, or when the cancer recurrence score is below the predetermined threshold, the cancer recurrence status of the test subject is determined to be at a lower risk of cancer recurrence.

[0036] In another aspect, the present disclosure provides a method of classifying a test subject as a candidate for subsequent cancer treatment, the method comprising: at one or more preselected time points after one or more prior cancer treatments of the test subject, collecting DNA originating from or derived from tumor cells from the test subject diagnosed with cancer; capturing more than one set of target regions from the DNA, wherein the more than one set of target regions includes a sequence variable target region set and an epigenetic target region set, thereby producing a set of captured DNA molecules; sequencing more than one captured DNA molecule from the set of DNA molecules, wherein the captured DNA molecules of the sequence variable target region set are sequenced to a deeper sequencing depth than the captured DNA molecules of the epigenetic target region set, thereby producing a set of sequence information; using the set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at one or more preselected time points; determining a cancer recurrence score that indicates the presence or absence of DNA originating from or derived from tumor cells; and comparing the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold, thereby classifying the test subject as a candidate for subsequent cancer treatment when the cancer recurrence score is higher than the cancer recurrence threshold, or not classifying the test subject as a candidate for therapy when the cancer recurrence score is lower than the cancer recurrence threshold.

[0037] The following is an exemplary list of embodiments according to the present disclosure.

[0038] Embodiment 1 is a method of isolating cell-free DNA (cfDNA), the method comprising:

[0039] capturing more than one set of target regions of cfDNA obtained from a test subject,

[0040] wherein the more than one set of target regions includes a sequence variable target region set and an epigenetic target region set,

[0041] thereby producing a set of captured cfDNA molecules;

[0042] wherein in the set of captured cfDNA molecules, the cfDNA molecules corresponding to the sequence variable target region set are captured with a higher capture yield than the cfDNA molecules corresponding to the epigenetic target region set.

[0043] Embodiment 2 is a method of isolating cell-free DNA (cfDNA), the method comprising:

[0044] contacting cfDNA obtained from a test subject with a set of target-specific probes,

[0045] Wherein the target-specific probe set includes target-binding probes specific for a group of sequence-variable target regions and target-binding probes specific for a group of epigenetic target regions, and the target-specific probe set is configured to capture cfDNA corresponding to the group of sequence-variable target regions with a higher capture yield than cfDNA corresponding to the group of epigenetic target regions,

[0046] thereby forming a complex of the target-specific probe and cfDNA; and

[0047] separating the complex from cfDNA that has not bound to the target-specific probe, thereby providing a set of captured cfDNA molecules.

[0048] Embodiment 3 is the method according to Embodiment 1 or 2, further comprising sequencing the set of captured cfDNA molecules.

[0049] Embodiment 4 is the method according to Embodiment 3, the method comprising sequencing cfDNA molecules corresponding to the group of sequence-variable target regions to a deeper sequencing depth than cfDNA molecules corresponding to the group of epigenetic target regions.

[0050] Embodiment 5 is a method for identifying the presence of DNA produced by a tumor, the method comprising:

[0051] collecting cfDNA from a test subject,

[0052] capturing more than one target region group from the cfDNA,

[0053] wherein the more than one target region group includes a group of sequence-variable target regions and a group of epigenetic target regions, thereby producing a set of captured cfDNA molecules,

[0054] sequencing the captured cfDNA molecules,

[0055] wherein the captured cfDNA molecules of the group of sequence-variable target regions are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the group of epigenetic target regions.

[0056] Embodiment 6 is the method according to any one of Embodiments 3-5, wherein the captured cfDNA molecules of the group of sequence-variable target regions are sequenced to a sequencing depth at least 2-fold deeper than the captured cfDNA molecules of the group of epigenetic target regions.

[0057] Embodiment 7 is the method according to any one of Embodiments 3-5, wherein the captured cfDNA molecules of the group of sequence-variable target regions are sequenced to a sequencing depth at least 3-fold deeper than the captured cfDNA molecules of the group of epigenetic target regions.

[0058] Embodiment 8 is the method according to any one of Embodiments 3-5, wherein the captured cfDNA molecules of the sequence variable target region set are sequenced to a sequencing depth 4-10 times deeper than that of the captured cfDNA molecules of the epigenetic target region set.

[0059] Embodiment 9 is the method according to any one of Embodiments 3-5, wherein the captured cfDNA molecules of the sequence variable target region set are sequenced to a sequencing depth 4-100 times deeper than that of the captured cfDNA molecules of the epigenetic target region set.

[0060] Embodiment 10 is the method according to any one of Embodiments 3-9, wherein before sequencing, the captured cfDNA molecules of the sequence variable target region set are pooled with the captured cfDNA molecules of the epigenetic target region set.

[0061] Embodiment 11 is the method according to any one of Embodiments 3-10, wherein the captured cfDNA molecules of the sequence variable target region set and the captured cfDNA molecules of the epigenetic target region set are sequenced in the same sequencing pool.

[0062] Embodiment 12 is the method according to any one of the foregoing embodiments, wherein the cfDNA is amplified before capture.

[0063] Embodiment 13 is the method according to Embodiment 12, wherein the cfDNA amplification includes the step of ligating an adaptor containing a barcode to the cfDNA.

[0064] Embodiment 14 is the method according to any one of the foregoing embodiments, wherein the epigenetic target region set includes a hypermethylated variable target region set.

[0065] Embodiment 15 is the method according to any one of the foregoing embodiments, wherein the epigenetic target region set includes a hypomethylated variable target region set.

[0066] Embodiment 16 is the method according to Embodiment 14 or 15, wherein the epigenetic target region set includes a methylation control target region set.

[0067] Embodiment 17 is the method according to any one of the foregoing embodiments, wherein the epigenetic target region set includes a fragmented variable target region set.

[0068] Embodiment 18 is the method according to Embodiment 17, wherein the fragmented variable target region set includes a transcription start site region.

[0069] Embodiment 19 is the method according to Embodiment 17 or 18, wherein the fragmented variable target region set includes a CTCF binding region.

[0070] Embodiment 20 is the method according to any one of the preceding embodiments, wherein capturing more than one target region set of the cfDNA comprises contacting the cfDNA with target binding probes specific for the sequence variable target region set and target binding probes specific for the epigenetic target region set.

[0071] Embodiment 21 is the method according to embodiment 20, wherein the target binding probes specific for the sequence variable target region set are present at a higher concentration than the target binding probes specific for the epigenetic target region set.

[0072] Embodiment 22 is the method according to embodiment 20, wherein the target binding probes specific for the sequence variable target region set are present at a concentration at least 2-fold higher than the target binding probes specific for the epigenetic target region set.

[0073] Embodiment 23 is the method according to embodiment 20, wherein the target binding probes specific for the sequence variable target region set are present at a concentration at least 4-fold or 5-fold higher than the target binding probes specific for the epigenetic target region set.

[0074] Embodiment 24 is the method according to any one of embodiments 20-23, wherein the target binding probes specific for the sequence variable target region set have a higher target binding affinity than the target binding probes specific for the epigenetic target region set.

[0075] Embodiment 25 is the method according to any one of the preceding embodiments, wherein the footprint of the epigenetic target region set is at least 2-fold larger than the size of the sequence variable target region set.

[0076] Embodiment 26 is the method according to embodiment 25, wherein the footprint of the epigenetic target region set is at least 10-fold larger than the size of the sequence variable target region set.

[0077] Embodiment 27 is the method according to any one of the preceding embodiments, wherein the footprint of the sequence variable target region set is at least 25 kB or 50 kB.

[0078] Embodiment 28 is the method according to any one of the preceding embodiments, wherein the cfDNA obtained from the test subject is partitioned into at least 2 fractions based on the methylation level, and the subsequent steps of the method are performed on each fraction.

[0079] Embodiment 29 is the method according to embodiment 28, wherein the partitioning step comprises contacting the collected cfDNA with a methyl-binding reagent immobilized on a solid support.

[0080] Embodiment 30 is the method according to embodiment 28 or 29, wherein the at least two fractions include a hypermethylated fraction and a hypomethylated fraction, and the method further comprises differentially tagging the hypermethylated fraction and the hypomethylated fraction or sequencing the hypermethylated fraction and the hypomethylated fraction separately.

[0081] Embodiment 31 is the method according to embodiment 30, wherein the hypermethylated fraction and the hypomethylated fraction are differentially tagged, and the method further comprises pooling the differentially tagged hypermethylated fraction and hypomethylated fraction prior to the sequencing step.

[0082] Embodiment 32 is the method according to any one of the foregoing embodiments, the method further comprising determining whether a cfDNA molecule corresponding to the set of sequence variable target regions contains a cancer-related mutation.

[0083] Embodiment 33 is the method according to any one of the foregoing embodiments, the method further comprising determining whether a cfDNA molecule corresponding to the set of epigenetic target regions contains or indicates a cancer-related epigenetic modification or copy number variation (e.g., focal amplification), optionally, wherein the method comprises determining whether a cfDNA molecule corresponding to the set of epigenetic target regions contains or indicates a cancer-related epigenetic modification and copy number variation (e.g., focal amplification).

[0084] Embodiment 34 is the method according to embodiment 33, wherein the cancer-related epigenetic modification comprises hypermethylation in one or more hypermethylated variable target regions.

[0085] Embodiment 35 is the method according to embodiment 33 or 34, wherein the cancer-related epigenetic modification comprises one or more perturbations of CTCF binding.

[0086] Embodiment 36 is the method according to any one of embodiments 33-35, wherein the cancer-related epigenetic modification comprises one or more perturbations of transcription start sites.

[0087] Embodiment 37 is the method according to any one of the foregoing embodiments, wherein the captured set of cfDNA molecules is sequenced using: high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing (NGS), single molecule synthesis sequencing (SMSS) (Helicos), massively parallel sequencing, clonal single molecule arrays (Solexa), shotgun sequencing, Ion Torrent, Oxford nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent or nanopore platforms.

[0088] Embodiment 38 is a set of target-specific probes for capturing cfDNA produced by tumor cells, the set comprising target-binding probes specific for a group of sequence-variable target regions and target-binding probes specific for a group of epigenetic target regions, wherein the capture yield of the target-binding probes specific for the group of sequence-variable target regions is at least 2-fold higher than the capture yield of the target-binding probes specific for the group of epigenetic target regions.

[0089] Embodiment 39 is the set of target-specific probes according to Embodiment 38, wherein the capture yield of the target-binding probes specific for the group of sequence-variable target regions is at least 4-fold or 5-fold higher than the capture yield of the target-binding probes specific for the group of epigenetic target regions.

[0090] Embodiment 40 is the set of target-specific probes according to Embodiment 38 or 39, wherein the group of epigenetic target regions includes a group of hypermethylation variable target region probes.

[0091] Embodiment 41 is the set of target-specific probes according to any one of Embodiments 38-40, wherein the group of epigenetic target regions includes a group of hypomethylation variable target region probes.

[0092] Embodiment 42 is the set of target-specific probes according to Embodiment 40 or 41, wherein the group of epigenetic target region probes includes a group of methylation control target region probes.

[0093] Embodiment 43 is the set of target-specific probes according to any one of Embodiments 38-42, wherein the group of epigenetic target region probes includes a group of fragmented variable target region probes.

[0094] Embodiment 44 is a set of target-specific probes as described in Embodiment 43, wherein the fragmented variable target region probe set includes transcription start site region probes.

[0095] Embodiment 45 is a set of target-specific probes as described in Embodiment 43 or 44, wherein the fragmented variable target region probe set includes CTCF binding region probes.

[0096] Embodiment 46 is a set of target-specific probes as described in any one of Embodiments 38-45, wherein there are at least 10 regions in the sequence variable target region set and at least 100 regions in the epigenetic target region set.

[0097] Embodiment 47 is a set of target-specific probes as described in any one of Embodiments 38-46, wherein the footprint of the epigenetic target region set is at least 2 times larger than the size of the sequence variable target region set.

[0098] Embodiment 48 is a set of target-specific probes as described in Embodiment 47, wherein the footprint of the epigenetic target region set is at least 10 times larger than the size of the sequence variable target region set.

[0099] Embodiment 49 is a set of target-specific probes as described in any one of Embodiments 38-48, wherein the footprint of the sequence variable target region set is at least 25 kB or 50 kB.

[0100] Embodiment 50 is a set of target-specific probes as described in any one of Embodiments 38-49, wherein the probes are present in a single solution.

[0101] Embodiment 51 is a set of target-specific probes as described in any one of Embodiments 38-50, wherein the probes include a capture moiety.

[0102] Embodiment 52 is a composition comprising captured cfDNA, wherein the captured cfDNA includes a captured sequence variable target region and a captured epigenetic target region, and the concentration of the sequence variable target region is greater than the concentration of the epigenetic target region, wherein the concentration is normalized to the footprint size of the sequence variable target region and the epigenetic target region.

[0103] Embodiment 53 is the composition as described in Embodiment 52, wherein the captured cfDNA contains sequence tags.

[0104] Embodiment 54 is the composition as described in Embodiment 53, wherein the sequence tags include barcodes.

[0105] Embodiment 55 is the composition according to any one of Embodiments 52-54, wherein the concentration of the sequence-variable target region is at least 2-fold greater than the concentration of the epigenetic target region.

[0106] Embodiment 56 is the composition according to any one of Embodiments 52-54, wherein the concentration of the sequence-variable target region is at least 4-fold or 5-fold greater than the concentration of the epigenetic target region.

[0107] Embodiment 57 is the composition according to any one of Embodiments 52-56, wherein the concentration is a mass / volume concentration normalized to the footprint size of the target region.

[0108] Embodiment 58 is the composition according to any one of Embodiments 52-57, wherein the epigenetic target region includes one, two, three, or four of a hypermethylated variable target region; a hypomethylated variable target region; a transcription start site region; and a CTCF binding region; optionally, wherein the epigenetic target region further includes a methylation control target region.

[0109] Embodiment 59 is the composition according to any one of Embodiments 52-58, the composition produced according to the method according to any one of Embodiments 1-37.

[0110] Embodiment 60 is a method for determining the likelihood that a subject has cancer, the method comprising:

[0111] Collecting cfDNA from a test subject;

[0112] Capturing more than one set of target regions from the cfDNA;

[0113] Wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby producing a set of captured cfDNA molecules;

[0114] Sequencing the captured cfDNA molecules,

[0115] Wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set;

[0116] Obtaining more than one sequence read generated by a nucleic acid sequencer by sequencing the captured cfDNA molecules;

[0117] Mapping the more than one sequence read to one or more reference sequences to produce mapped sequence reads;

[0118] Processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.

[0119] Embodiment 61 is the method according to embodiment 60, having the features of any one of embodiments 21-38.

[0120] Embodiment 62 is the method according to embodiment 60 or 61, wherein the captured cfDNA molecules of the set of sequence variable target regions are pooled with the captured cfDNA molecules of the set of epigenetic target regions before obtaining the plurality of sequence reads and / or before sequencing in the same sequencing pool.

[0121] Embodiment 63 is the method according to any one of embodiments 60-62, wherein the cfDNA is amplified before capture, optionally, wherein the cfDNA amplification comprises the step of ligating an adaptor comprising a barcode to the cfDNA.

[0122] Embodiment 64 is the method according to any one of embodiments 60-63, wherein the set of epigenetic target regions is as described in any one of embodiments 15-19.

[0123] Embodiment 65 is the method according to any one of embodiments 60-64, wherein capturing the plurality of target regions of the cfDNA comprises contacting the cfDNA with target binding probes specific for the set of sequence variable target regions and target binding probes specific for the set of epigenetic target regions.

[0124] Embodiment 66 is a system, the system comprising:

[0125] A communication interface that receives, via a communication network, a plurality of sequence reads generated by a nucleic acid sequencer by sequencing a set of captured cfDNA molecules, wherein the set of captured cfDNA molecules is obtained by capturing a plurality of target regions from a cfDNA sample, wherein the plurality of target regions comprises a set of sequence variable target regions and a set of epigenetic target regions, and wherein the captured cfDNA molecules corresponding to the sequence variable target regions are sequenced to a deeper sequencing depth than the captured cfDNA molecules corresponding to the set of epigenetic target regions; and

[0126] A controller that comprises or is able to access a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that comprises:

[0127] (i) receiving, via the communication network, the sequence reads generated by the nucleic acid sequencer;

[0128] (ii) mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads;

[0129] (iii) Process sequence reads corresponding to the mapping of the sequence variable target region set and the epigenetic target region set to determine the likelihood that a subject has cancer.

[0130] Embodiment 67 is the system according to embodiment 66, wherein the sequencing depth corresponding to the sequence variable target region set is at least 2 times deeper than the sequencing depth corresponding to the epigenetic target region set.

[0131] Embodiment 68 is the system according to embodiment 66, wherein the sequencing depth corresponding to the sequence variable target region set is at least 3 times deeper than the sequencing depth corresponding to the epigenetic target region set.

[0132] Embodiment 69 is the system according to embodiment 66, wherein the sequencing depth corresponding to the sequence variable target region set is 4 - 10 times deeper than the sequencing depth corresponding to the epigenetic target region set.

[0133] Embodiment 70 is the system according to any one of embodiments 66, wherein the sequencing depth corresponding to the sequence variable target region set is 4 - 100 times deeper than the sequencing depth corresponding to the epigenetic target region set.

[0134] Embodiment 71 is the system according to any one of embodiments 66 - 70, wherein before sequencing, the captured cfDNA molecules of the sequence variable target region set are pooled with the captured cfDNA molecules of the epigenetic target region set.

[0135] Embodiment 72 is the system according to any one of embodiments 66 - 71, wherein the captured cfDNA molecules of the sequence variable target region set and the captured cfDNA molecules of the epigenetic target region set are sequenced in the same sequencing pool.

[0136] Embodiment 73 is the system according to any one of embodiments 66 - 72, wherein the epigenetic target region set includes a hypermethylation variable target region set.

[0137] Embodiment 74 is the system according to any one of embodiments 66 - 73, wherein the epigenetic target region set includes a hypomethylation variable target region set.

[0138] Embodiment 75 is the system according to embodiment 72 or 73, wherein the epigenetic target region set includes a methylation control target region set.

[0139] Embodiment 76 is the system according to any one of embodiments 66 - 75, wherein the epigenetic target region set includes a fragmentation variable target region set.

[0140] Embodiment 77 is the system according to any one of embodiments 66 - 76, wherein the fragmentation variable target region set includes a transcription start site region.

[0141] Embodiment 78 is the system according to embodiment 76 or 77, wherein the fragmented variable target region set includes CTCF binding regions.

[0142] Embodiment 79 is the system according to any one of embodiments 66 - 78, wherein the footprint of the epigenetic target region set is at least 2 times larger than the size of the sequence variable target region set.

[0143] Embodiment 80 is the system according to embodiment 79, wherein the footprint of the epigenetic target region set is at least 10 times larger than the size of the sequence variable target region set.

[0144] Embodiment 81 is the system according to any one of embodiments 66 - 80, wherein the footprint of the sequence variable target region set is at least 25 kB or 50 kB.

[0145] Embodiment 82 is the method or system according to any one of the above embodiments, wherein the capture is performed in a single container.

[0146] Embodiment 83 is the method according to any one of embodiments 1 - 37, wherein the test subject was previously diagnosed with cancer and received one or more previous cancer treatments, optionally, wherein the cfDNA is obtained at one or more preselected time points after the one or more previous cancer treatments.

[0147] Embodiment 84 is the method according to the previous embodiment, the method further comprising sequencing the captured set of cfDNA molecules to generate a set of sequence information.

[0148] Embodiment 85 is the method according to the previous embodiment, wherein the captured DNA molecules of the sequence variable target region set are sequenced to a deeper sequencing depth than the captured DNA molecules of the epigenetic target region set.

[0149] Embodiment 86 is the method according to embodiment 84 or 85, the method further comprising detecting the presence or absence of DNA originating from or derived from tumor cells at a preselected time point using the set of sequence information.

[0150] Embodiment 87 is the method according to the previous embodiment, the method further comprising determining a cancer recurrence score, the cancer recurrence score indicating the presence or absence of DNA originating from or derived from the tumor cells of the test subject.

[0151] Embodiment 88 is the method described in the previous embodiment, the method further comprising determining a cancer recurrence status based on the cancer recurrence score, wherein when the cancer recurrence score is determined to be at or above a predetermined threshold, the cancer recurrence status of the test subject is determined to be at risk of cancer recurrence, or when the cancer recurrence score is below the predetermined threshold, the cancer recurrence status of the test subject is determined to be at a lower risk of cancer recurrence.

[0152] Embodiment 89 is the method described in Embodiment 87 or 88, the method further comprising comparing the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold, and when the cancer recurrence score is higher than the cancer recurrence threshold, the test subject is classified as a candidate for subsequent cancer treatment, or when the cancer recurrence score is lower than the cancer recurrence threshold, the test subject is not classified as a candidate for subsequent cancer treatment.

[0153] Embodiment 90 is a method for determining the risk of cancer recurrence in a test subject, the method comprising:

[0154] (a) Collecting DNA originating from or derived from tumor cells from a test subject diagnosed with cancer at one or more preselected time points after one or more previous cancer treatments of the test subject;

[0155] (b) Capturing more than one set of target regions from the DNA, wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby generating a set of captured DNA molecules;

[0156] (c) Sequencing the captured DNA molecules, wherein the captured DNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured DNA molecules of the epigenetic target region set, thereby generating a set of sequence information;

[0157] (d) Detecting the presence or absence of DNA originating from or derived from tumor cells at the preselected time point using the set of sequence information; and

[0158] (e) Determining a cancer recurrence score, the cancer recurrence score indicating the presence or absence of DNA originating from or derived from the tumor cells of the test subject, wherein when the cancer recurrence score is determined to be at or above a predetermined threshold, the cancer recurrence status of the test subject is determined to be at risk of cancer recurrence, or when the cancer recurrence score is below the predetermined threshold, the cancer recurrence status of the test subject is determined to be at a lower risk of cancer recurrence.

[0159] Embodiment 91 is a method for classifying a test subject as a candidate for subsequent cancer treatment, the method comprising:

[0160] (a) At one or more preselected time points after one or more previous cancer treatments of the test subject, collect DNA originating from or derived from tumor cells from a test subject diagnosed with cancer;

[0161] (b) Capture more than one set of target regions from the DNA, wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby generating a set of captured DNA molecules;

[0162] (c) Sequence more than one captured DNA molecule from the set of DNA molecules, wherein the captured DNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured DNA molecules of the epigenetic target region set, thereby generating a set of sequence information;

[0163] (d) Use the set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at one or more preselected time points,

[0164] (e) Determine a cancer recurrence score that indicates the presence or absence of DNA originating from or derived from the tumor cells; and

[0165] (f) Compare the cancer recurrence score of the test subject with a predetermined cancer recurrence threshold, thereby classifying the test subject as a candidate for subsequent cancer treatment when the cancer recurrence score is higher than the cancer recurrence threshold, or not classifying the test subject as a candidate for therapy when the cancer recurrence score is lower than the cancer recurrence threshold.

[0166] Embodiment 92 is the method according to embodiments 88-90, wherein the test subject is at risk of cancer recurrence and is classified as a candidate for subsequent cancer treatment.

[0167] Embodiment 93 is the method according to any one of embodiments 89, 91 or 92, wherein the subsequent cancer treatment includes chemotherapy or administration of a therapeutic composition.

[0168] Embodiment 94 is the method according to any one of embodiments 90-93, wherein the DNA originating from or derived from tumor cells is cell-free DNA.

[0169] Embodiment 95 is the method according to any one of embodiments 90-93, wherein the DNA originating from or derived from tumor cells is obtained from a tissue sample.

[0170] Embodiment 96 is the method according to any one of embodiments 87-95, the method further comprising determining a disease-free survival (DFS) period for the test subject based on the cancer recurrence score.

[0171] Embodiment 97 is the method according to embodiment 96, wherein the DFS period is 1 year, 2 years, 3 years, 4 years, 5 years or 10 years.

[0172] Embodiment 98 is the method according to any one of embodiments 84-97, wherein the sequence information group includes sequence variable target region sequences, and determining the cancer recurrence score includes determining at least a first sub-score indicative of the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequences.

[0173] Embodiment 99 is the method according to embodiment 98, wherein the number of mutations in the sequence variable target regions selected from 1, 2, 3, 4 or 5 is sufficient to result in a cancer recurrence score classified as positive for cancer recurrence, optionally wherein the number of mutations is selected from 1, 2 or 3.

[0174] Embodiment 100 is the method according to any one of embodiments 84-99, wherein the sequence information group includes epigenetic target region sequences, and determining the cancer recurrence score includes determining a second sub-score indicative of the amount of abnormal sequence reads in the epigenetic target region sequences.

[0175] Embodiment 101 is the method according to embodiment 100, wherein the abnormal sequence reads include reads indicative of methylation of hypermethylated variable target sequences and / or reads indicative of abnormal fragmentation in fragmented variable target regions.

[0176] Embodiment 102 is the method according to embodiment 101, wherein a proportion of reads corresponding to the hypermethylated variable target region group and / or the fragmented variable target region group indicative of hypermethylation in the hypermethylated variable target region group and / or abnormal fragmentation in the fragmented variable target region group that is greater than or equal to a value in the range of 0.001%-10% is sufficient to classify the second sub-score as positive for cancer recurrence.

[0177] Embodiment 103 is the method according to embodiment 102, wherein the range is 0.001%-1% or 0.005%-1%.

[0178] Embodiment 104 is the method according to embodiment 102, wherein the range is 0.01%-5% or 0.01%-2%.

[0179] Embodiment 105 is the method according to embodiment 102, wherein the range is 0.01%-1%.

[0180] Embodiment 106 is the method according to any one of Embodiments 84 - 105, the method further comprising determining a fraction of tumor DNA from one or more read fractions in the sequence information set that indicate features originating from tumor cells.

[0181] Embodiment 107 is the method according to Embodiment 106, wherein one or more features indicating origin from tumor cells include one or more of alterations in sequence variable target regions, hypermethylation of hypermethylated variable target regions, and abnormal fragmentation of fragmented variable target regions.

[0182] Embodiment 108 is the method according to Embodiment 106 or 107, the method further comprising determining a cancer recurrence score based at least in part on the fraction of tumor DNA, wherein a fraction of tumor DNA greater than or equal to a predetermined value in the range of 10 -11 to 1 or 10 -10 to 1 is sufficient to classify the cancer recurrence score as positive for cancer recurrence.

[0183] Embodiment 109 is the method according to Embodiment 108, wherein a fraction of tumor DNA greater than or equal to a predetermined value in the range of 10 –10 to 10 –9 、10 –9 to 10 –8 、10 –8 to 10 –7 、10 –7 to 10 –6 、10 –6 to 10 –5 、10 –5 to 10 –4 、10 –4 to 10 –3 、10 –3 to 10 –2 or 10 –2 to 10 –1 is sufficient to classify the cancer recurrence score as positive for cancer recurrence.

[0184] Embodiment 110 is the method according to Embodiment 108 or 109, wherein the predetermined value is in the range of 10 –8 to 10 –6 or is 10 -7 。

[0185] Embodiment 111 is the method according to any one of Embodiments 107-110, wherein if the cumulative probability that the fraction of the tumor DNA is greater than or equal to a predetermined value is at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995 or 0.999, then the fraction of the tumor DNA is determined to be greater than or equal to the predetermined value.

[0186] Embodiment 112 is the method according to Embodiment 111, wherein the cumulative probability is at least 0.95.

[0187] Embodiment 113 is the method according to Embodiment 111, wherein the cumulative probability is in the range of 0.98-0.995 or is 0.99.

[0188] Embodiment 114 is the method according to any one of Embodiments 84-113, wherein the sequence information set includes a sequence variable target region sequence and an epigenetic target region sequence, and determining the cancer recurrence score includes determining a first sub-score indicating the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequence and a second sub-score indicating the amount of abnormal sequence reads in the epigenetic target region sequence, and combining the first sub-score and the second sub-score to provide the cancer recurrence score.

[0189] Embodiment 115 is the method according to Embodiment 114, wherein combining the first sub-score and the second sub-score includes independently applying a threshold to each sub-score (e.g., greater than a predetermined number of mutations (e.g., >1) in the sequence variable target region and greater than a predetermined fraction of abnormal (e.g., tumor) reads in the epigenetic target region), or training a machine learning classifier to determine the status based on more than one positive and negative training sample.

[0190] Embodiment 116 is the method according to Embodiment 115, wherein a combined score value in the range of -4 to 2 or -3 to 1 is sufficient to classify the cancer recurrence score as cancer recurrence positive.

[0191] Embodiment 117 is the method according to any one of Embodiments 83-116, wherein one or more preselected time points are selected from the group consisting of 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 1.5 years, 2 years, 3 years, 4 years and 5 years after administration of one or more previous cancer treatments.

[0192] Embodiment 118 is the method according to any one of Embodiments 83-117, wherein the cancer is colorectal cancer.

[0193] Embodiment 119 is the method according to any one of Embodiments 83-118, wherein one or more prior cancer treatments include surgery.

[0194] Embodiment 120 is the method according to any one of Embodiments 83-119, wherein one or more prior cancer treatments include administering a therapeutic composition.

[0195] Embodiment 121 is the method according to any one of Embodiments 83-120, wherein one or more prior cancer treatments include chemotherapy.

[0196] Each step of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same time or at different times and / or at the same geographical location or at different geographical locations such as countries. Each step of the methods disclosed herein may be performed by the same person or different persons. BRIEF DESCRIPTION OF THE DRAWINGS

[0197] The drawings incorporated in and forming a part of this specification illustrate certain embodiments and, together with the written description, are used to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the drawings, which are included by way of example and not of limitation. It should be understood that, unless the context otherwise requires, like reference numerals denote like elements in all the drawings. It should also be understood that some or all of the drawings may be schematic for illustrative purposes and do not necessarily depict the actual relative sizes or positions of the elements shown.

[0198] Figure 1 An overview of the partitioning method is shown.

[0199] Figure 2 is a schematic diagram of an example of a system suitable for some embodiments of the present disclosure.

[0200] Figure 3 Shows the sensitivity of detecting cancer at different stages using one or both of epigenetic target regions and sequence variable target regions in a liquid biopsy test as described in Example II.

[0201] Figure 4 Shows the recurrence-free survival of subjects with or without detectable ctDNA over time as described in Example III. DETAILED DESCRIPTION

[0202] Certain embodiments of the present invention will now be described in detail. Although the invention will be described in connection with these embodiments, it should be understood that they are not intended to limit the invention thereto. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents that may be included within the invention as defined by the appended claims.

[0203] Before describing the teachings in detail, it should be understood that the present disclosure is not limited to specific compositions or process steps, as these may vary. It should be noted that, unless the context clearly dictates otherwise, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents. Thus, for example, reference to "a nucleic acid" includes more than one nucleic acid, reference to "a cell" includes more than one cell, and so forth.

[0204] Numeric ranges include the numbers defining the range. Measured and measurable values should be understood as approximate values, taking into account significant figures and errors associated with the measurements. In addition, the use of "comprise", "comprises", "comprising", "contain", "contains", "containing", "include", "includes", and "including" is not intended to be limiting. It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only, and not restrictive of the teachings.

[0205] Unless specifically stated otherwise in the above specification, embodiments described in the specification as "comprising" various components are also considered to be "consisting of" or "consisting essentially of" the recited components; embodiments described in the specification as "consisting of" various components are also considered to be "comprising" or "consisting essentially of" the recited components; and embodiments described in the specification as "consisting essentially of" various components are also considered to be "consisting of" or "comprising" the recited components (this interchangeability does not apply to the use of these terms in the claims).

[0206] The section headings used herein are for organizational purposes and are not to be construed as limiting the disclosed subject matter in any way. If any document or other material incorporated by reference conflicts with any explicit content (including definitions) of this specification, then this specification shall govern.

[0207] I. Definitions

[0208] "Cell-free DNA", "cfDNA molecule", or simply "cfDNA" includes DNA molecules that exist in a subject in an extracellular form (e.g., in blood, serum, plasma, or other body fluids such as lymph, cerebrospinal fluid, urine, or sputum), and includes DNA that is not contained within a cell or otherwise bound to a cell. Although the DNA initially exists in one or more cells of an organism (e.g., a mammal) of a large complex organism, the DNA has undergone release from the one or more cells into the fluid present in the organism. Typically, cfDNA can be obtained by obtaining a fluid sample without performing an in vitro cell lysis step, and also includes removing cells present in the fluid (e.g., centrifuging blood to remove cells).

[0209] The "capture yield" of a probe set for a given set of target regions refers to the amount of nucleic acid corresponding to the set of target regions captured by the set under typical conditions (e.g., relative to the amount of another set of target regions or an absolute amount). Exemplary typical capture conditions are incubation of the sample nucleic acid and the probe in a small reaction volume (about 20 μL) containing a stringent hybridization buffer at 65 °C for 10 - 18 hours. The capture yield can be expressed in absolute terms, or for more than one probe set, in relative terms. When comparing the capture yields for more than one set of target regions, they are normalized against the footprint size of the set of target regions (e.g., per kilobase). Thus, for example, if the footprint sizes of a first target region and a second target region are 50 kb and 500 kb, respectively (giving a normalization factor of 0.1), then the DNA corresponding to the first set of target regions is captured at a higher yield than the DNA corresponding to the second set of target regions when the mass / volume concentration of the DNA captured corresponding to the first set of target regions is more than 0.1 times the mass / volume concentration of the DNA captured corresponding to the second set of target regions. As another example, using the same footprint size, if the mass / volume concentration of the DNA captured corresponding to the first set of target regions is 0.2 times the mass / volume concentration of the DNA captured corresponding to the second set of target regions, then the DNA corresponding to the first set of target regions is captured at a capture yield twice as high as the DNA corresponding to the second set of target regions.

[0210] "Capturing" or "enriching" one or more target nucleic acids refers to preferentially isolating or separating one or more target nucleic acids from non-target nucleic acids.

[0211] A "captured" nucleic acid set refers to nucleic acids that have undergone capture.

[0212] "Target-region set" or "set of target regions" or "target region" refers to more than one genomic locus or more than one genomic region that is targeted for capture and / or targeted by a set of probes (e.g., by sequence complementarity).

[0213] "Corresponding to a target-region set" means that a nucleic acid, such as cfDNA, originates from a locus in the target-region set or that one or more probes specifically bind to the target-region set.

[0214] In the context of a probe or other oligonucleotide and a target sequence, "specifically binds" means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a replica thereof to form a stable probe:target hybrid, while the formation of stable probe:non-target hybrids is minimized. Thus, the probe hybridizes to the target sequence or a replica thereof to a much greater extent than to non-target sequences, enabling capture or detection of the target sequence. Appropriate hybridization conditions are well known in the art, can be predicted based on sequence composition, or can be determined by using conventional testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) at §§1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57, particularly §§9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57, which are incorporated herein by reference).

[0215] "Sequentially variable target-region set" refers to a target-region set that may exhibit sequence alterations, such as nucleotide substitutions, insertions, deletions, or gene fusions or transpositions, in neoplastic cells (e.g., tumor cells and cancer cells).

[0216] "Epigenetic targetome" refers to a targetome that may exhibit non-sequence modifications in neoplastic cells (e.g., tumor cells and cancer cells) and non-tumor cells (e.g., immune cells, cells from the tumor microenvironment). These modifications do not change the sequence of DNA. Examples of non-sequence modifications that are altered include, but are not limited to, methylation, nucleosome distribution, CTCF binding, transcription start sites, regulatory protein binding regions, and alterations (increases or decreases) in any other proteins that may bind to DNA. For the purposes of the present invention, loci that are sensitive to focal amplifications and / or gene fusions associated with a neoplasm, tumor, or cancer may also be included in the epigenetic targetome, because the detection of copy number alterations by sequencing or mapping fusion sequences to more than one locus in a reference genome tends to be more similar to the detection of the exemplary epigenetic alterations discussed above than to the detection of nucleotide substitutions, insertions, or deletions. For example, because focal amplifications and / or gene fusions can be detected at relatively shallow sequencing depths, since their detection does not rely on the accuracy of base calling at one or several individual positions. For example, the epigenetic targetome may include a targetome for analyzing fragment length or fragment end position distribution. The terms "epigenetics" and "epigenomics" are used interchangeably herein.

[0217] Circulating tumor DNA or ctDNA is a cfDNA component that originates from tumor cells or cancer cells. In some embodiments, cfDNA includes DNA that originates from normal cells and DNA that originates from tumor cells (i.e., ctDNA). Tumor cells are neoplastic cells that originate from a tumor, whether they remain within the tumor or become separate from the tumor (as, for example, in the case of metastatic cancer cells and circulating tumor cells).

[0218] The term "hypermethylation" refers to an increase in the level or degree of methylation of one or more nucleic acid molecules relative to other nucleic acid molecules within a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypermethylated DNA may include DNA molecules that contain at least 1 methylated residue, at least 2 methylated residues, at least 3 methylated residues, at least 5 methylated residues, at least 10 methylated residues, at least 20 methylated residues, at least 25 methylated residues, or at least 30 methylated residues.

[0219] The term "hypomethylation" refers to a decrease in the level or degree of methylation of one or more nucleic acid molecules relative to other nucleic acid molecules within a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypomethylated DNA includes non-methylated DNA molecules. In some embodiments, hypomethylated DNA may include DNA molecules that contain 0 methylated residues, up to 1 methylated residue, up to 2 methylated residues, up to 3 methylated residues, up to 4 methylated residues, or up to 5 methylated residues.

[0220] As used herein, the terms "a combination thereof" and "combinations thereof" refer to any and all permutations and combinations of the terms listed before that term. For example, "A, B, C, or a combination thereof" is intended to include at least one of the following: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing with this example, combinations that include repeats of one or more items or terms are explicitly included, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc. Unless otherwise apparent from the context, those skilled in the art will understand that there is generally no limit to the number of items or terms in any combination.

[0221] Unless the context otherwise requires, "or" is used in an inclusive sense, i.e., equivalent to "and / or".

[0222] II. Exemplary Methods

[0223] Methods are provided herein for isolating cell-free DNA (cfDNA) and / or identifying the presence of DNA produced by a tumor (or neoplastic or cancerous cells).

[0224] In some embodiments, the method includes capturing cfDNA obtained from a test subject for more than one set of target regions. The target regions include epigenetic target regions, which may exhibit differences in methylation levels and / or fragmentation patterns depending on whether the epigenetic target regions are of tumor origin or from healthy cells. The target regions also include sequence variable target regions, which may exhibit sequence differences depending on whether the sequence variable target regions are of tumor origin or from healthy cells. The capturing step produces a set of captured cfDNA molecules, and the cfDNA molecules corresponding to the set of sequence variable target regions in the set of captured cfDNA molecules are captured with a higher capture yield than the cfDNA molecules corresponding to the set of epigenetic target regions.

[0225] In some embodiments, the method includes contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to a set of sequence variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions.

[0226] Capturing cell-free DNA (cfDNA) corresponding to a set of sequence variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions can be beneficial because analyzing sequence variable target regions may require a deeper sequencing depth than analyzing epigenetic target regions to a sufficient confidence or accuracy. The deeper sequencing depth can result in more reads per DNA molecule and can facilitate sequencing by capturing more unique molecules per region. The amount of data required to determine fragmentation patterns (e.g., testing for perturbations of transcription start sites or CTCF binding sites) or fragment abundances (e.g., in hypermethylated and hypomethylated compartments) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations. Capturing target regions at different yields can facilitate sequencing the target regions to different sequencing depths in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing pool).

[0227] In various embodiments, the method further includes, for example, sequencing the captured cfDNA to different degrees of sequencing depth for the set of epigenetic target regions and the set of sequence variable target regions, consistent with the discussion above.

[0228] 1. Capture step; amplification; adaptors; barcodes

[0229] In some embodiments, the methods disclosed herein include the step of capturing one or more sets of target regions of DNA such as cfDNA. Any suitable method known in the art can be used for the capture.

[0230] In some embodiments, the capture includes contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes can have any of the characteristics described herein for sets of target-specific probes, including but not limited to the embodiments illustrated above and the portions associated with the probes below.

[0231] The capture step can be performed using conditions suitable for specific nucleic acid hybridization, which generally depends to some extent on the characteristics of the probes, such as length, base composition, etc. Those skilled in the art will be familiar with the appropriate conditions of the general knowledge of nucleic acid hybridization known in the art. In some embodiments, a complex of target-specific probes and DNA is formed.

[0232] In some embodiments, the complex of target-specific probes and DNA is separated from the DNA that has not bound to the target-specific probes. For example, in the case where the target-specific probes are covalently or non-covalently bound to a solid support, washing or aspiration steps can be used to separate the unbound material. Optionally, chromatography can be used in cases where the complex has chromatographic properties different from the unbound material (e.g., where the probe contains a ligand that binds to a chromatography resin).

[0233] As discussed in detail elsewhere herein, the set of target-specific probes can include more than one set, such as probes for a set of sequence-variable target regions and probes for a set of epigenetic target regions. In some such embodiments, the capture step is performed simultaneously in the same container with the probes for the set of sequence-variable target regions and the probes for the set of epigenetic target regions, e.g., the probes for the set of sequence-variable target regions and the probes for the set of epigenetic target regions are in the same composition. This method provides a relatively simplified workflow. In some embodiments, the concentration of the probes for the set of sequence-variable target regions is greater than the concentration of the probes for the set of epigenetic target regions.

[0234] Optionally, the capture step is performed with the set of sequence-variable target probes in a first container and the set of epigenetic target probes in a second container, or the contacting step is performed with the set of sequence-variable target probes in a first container at a first time and the set of epigenetic target probes at a second time before or after the first time. This method allows for the preparation of separate first and second compositions, the first and second compositions comprising captured DNA corresponding to the set of sequence-variable target regions and captured DNA corresponding to the set of epigenetic target regions. The compositions can be processed separately as needed (e.g., fractionated based on methylation as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis such as sequencing.

[0235] In some embodiments, the DNA is amplified. In some embodiments, the amplification is performed before the capture step. In some embodiments, the amplification is performed after the capture step. Methods for non-specific amplification of DNA are known in the art. See, e.g., Smallwood et al., Nat. Methods 11:817-820 (2014). For example, random primers having an adaptor sequence at their 5' end and random bases at their 3' end can be used. Typically, there are 6 random bases, but the length can be between 4 and 9 bases. This method is suitable for low-input / single-cell amplification and / or bisulfite sequencing.

[0236] In some embodiments, adaptors are included in the DNA. This can be done, e.g., as described above, e.g., by providing the adaptor in the 5' portion of the primer and performing the amplification procedure simultaneously. Optionally, the adaptor can be added by other methods such as ligation.

[0237] In some embodiments, a tag (which can be or include a barcode) is included in the DNA. The tag can facilitate identification of the source of the nucleic acid. For example, the barcode can be used to allow identification of the source (e.g., the subject) from which the DNA originated after pooling more than one sample for parallel sequencing. This can be done, for example, as described above, e.g., by providing the barcode in the 5' portion of the primer simultaneously with the amplification procedure. In some embodiments, the adaptor and the tag / barcode are provided by the same primer or primer set. For example, the barcode can be located on the 3' side of the adaptor and the 5' side of the target hybridization portion of the primer. Optionally, the barcode can be added by other methods, such as ligation, optionally together with the adaptor in the same ligation substrate.

[0238] Additional details regarding amplification, tagging, and barcoding are discussed in the "General Features of the Methods" section below, which can be combined, to the extent practicable, with any of the foregoing embodiments and the embodiments set forth in the Introduction and Overview sections.

[0239] 2. Captured Sets

[0240] In some embodiments, a set of captured DNA (e.g., cfDNA) is provided. With respect to the disclosed methods, for example, after the capture and / or isolation steps as described herein, a set of captured DNA can be provided. The captured set can include DNA corresponding to a set of sequence variable target regions and a set of epigenetic target regions. In some embodiments, when normalizing for differences in the size (footprint size) of the targeted regions, the amount of captured sequence variable target region DNA is greater than the amount of captured epigenetic target region DNA.

[0241] Optionally, a first captured set and a second captured set can be provided that respectively include DNA corresponding to the set of sequence variable target regions and DNA corresponding to the set of epigenetic target regions. The first captured set and the second captured set can be combined to provide a combined captured set.

[0242] In a group that includes capture of DNA corresponding to a set of sequence variable target regions and a set of epigenetic target regions (including a combined capture group as discussed above), the DNA corresponding to the set of sequence variable target regions can be present at a higher concentration than the DNA corresponding to the set of epigenetic target regions, e.g., at a concentration 1.1-fold to 1.2-fold higher, 1.2-fold to 1.4-fold higher, 1.4-fold to 1.6-fold higher, 1.6-fold to 1.8-fold higher, 1.8-fold to 2.0-fold higher, 2.0-fold to 2.2-fold higher, 2.2-fold to 2.4-fold higher, 2.4-fold to 2.6-fold higher, 2.6-fold to 2.8-fold higher, 2.8-fold to 3.0-fold higher, 3.0-fold to 3.5-fold higher, 3.5-fold to 4.0-fold higher, 4.0-fold to 4.5-fold higher, 4.5-fold to 5.0-fold higher, 5.0-fold to 5.5-fold higher, 5.5-fold to 6.0-fold higher, 6.0-fold to 6.5-fold higher, 6.5-fold to 7.0-fold higher, 7.0-fold to 7.5-fold higher, 7.5-fold to 8.0-fold higher, 8.0-fold to 8.5-fold higher, 8.5-fold to 9.0-fold higher, 9.0-fold to 9.5-fold higher, 9.5-fold to 10.0-fold higher, 10-fold to 11-fold higher, 11-fold to 12-fold higher, 12-fold to 13-fold higher, 13-fold to 14-fold higher, 14-fold to 15-fold higher, 15-fold to 16-fold higher, 16-fold to 17-fold higher, 17-fold to 18-fold higher, 18-fold to 19-fold higher, or 19-fold to 20-fold higher. The degree of difference in concentration accounts for normalization to the target footprint size, as discussed in the definition section.

[0243] a. Set of epigenetic target regions

[0244] The set of epigenetic target regions can include one or more types of target regions that may distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells (e.g., non-neoplastic circulating cells). Exemplary types of such regions are discussed in detail herein. In some embodiments, the methods according to the present disclosure include determining whether a cfDNA molecule corresponding to the set of epigenetic target regions contains or indicates a cancer-related epigenetic modification (e.g., hypermethylation in one or more hypermethylated variable target regions; one or more perturbations in CTCF binding; and / or one or more perturbations in transcription start sites) and / or copy number variations (e.g., focal amplification). The set of epigenetic target regions can also include, for example, one or more control regions as described herein.

[0245] In some embodiments, an epigenetic target set has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, an epigenetic target set has a footprint in the range of: 100 - 1000 kb, e.g., 100 - 200 kb, 200 - 300 kb, 300 - 400 kb, 400 - 500 kb, 500 - 600 kb, 600 - 700 kb, 700 - 800 kb, 800 - 900 kb, and 900 - 1,000 kb.

[0246] i. Hypermethylated variable targets

[0247] In some embodiments, an epigenetic target set includes one or more hypermethylated variable targets. Generally, a hypermethylated variable target refers to a region where an increase in the observed methylation level indicates an increased likelihood that a sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells such as tumors or cancer cells. For example, hypermethylation of tumor suppressor gene promoters has been repeatedly observed. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and the references cited therein.

[0248] A comprehensive discussion of methylation variable targets in colorectal cancer is provided in: Lam et al., Biochim Biophys Acta. 1866:106 - 20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylated variable targets containing genes or portions thereof based on colorectal cancer (CRC) studies is provided in Table 1. Many of these genes may be associated with cancers other than colorectal cancer; for example, TP53 is widely regarded as a crucial tumor suppressor, and inactivation of this gene based on hypermethylation may be a common carcinogenic mechanism.

[0249] Table 1. Exemplary hypermethylated targets (genes or portions thereof) based on CRC studies.

[0250]

[0251] In some embodiments, the hypermethylated variable target regions comprise more than one gene or portions thereof listed in Table 1, such as at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the genes or portions thereof listed in Table 1. For example, for each locus included as a target region, there may be one or more probes that have a hybridization site that binds between the transcription start site and the stop codon (the last stop codon of an alternatively spliced gene) of the gene. In some embodiments, the one or more probes bind within 300 bp upstream and / or downstream of the genes or portions thereof listed in Table 1, for example, within 200 bp or 100 bp.

[0252] Methylation variable target regions in various types of lung cancer are discussed in detail in the following: for example, Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinsky, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11:102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10):1492–1495 (2006); Kim et al., Cancer Res. 61:3419–3424 (2001); Furonaka et al., Pathology International 55:303-309 (2005); Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014); Kim et al., Oncogene. 20:1765-70 (2001); Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003); Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005); Heller et al., Oncogene 25:959–968 (2006); Licchesi et al., Carcinogenesis. 29:895–904 (2008); Guo et al., Clin. Cancer Res. 10:7917-24 (2004); Palmisano et al., Cancer Res. 63:4620–4625 (2003); and Toyooka et al., Cancer Res. 61:4556–4560, (2001).

[0253] An exemplary group of hypermethylation variable target regions containing genes or portions thereof based on lung cancer studies is provided in Table 2. Many of these genes may be associated with cancers other than lung cancer; for example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and inactivation of this gene based on hypermethylation may be a common carcinogenic mechanism not limited to lung cancer. Additionally, many genes appear in both Table 1 and Table 2, indicating universality.

[0254] Table 2. Exemplary hypermethylated target regions (genes or portions thereof) based on lung cancer studies

[0255] Gene Name Chromosome MARCH11 chr5 TAC1 chr7 TCF21 chr6 SHOX2 chr3 p16 chr3 Casp8 chr2 CDH13 chr16 MGMT chr10 MLH1 chr3 MSH2 chr2 TSLC1 chr11 APC chr5 DKK1 chr10 DKK3 chr11 LKB1 chr11 WIF1 chr12 RUNX3 chr1 GATA4 chr8 GATA5 chr20 PAX5 chr9 E-cadherin chr16 H-cadherin chr16

[0256] Any of the foregoing embodiments regarding the target regions identified in Table 2 can be combined with any of the above-described embodiments regarding the target regions identified in Table 1. In some embodiments, the hypermethylated variable target regions comprise more than one gene or portion thereof listed in Table 1 or Table 2, such as at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the genes or portions thereof listed in Table 1 or Table 2.

[0257] Additional hypermethylated target regions can be obtained from, for example, The Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017), describe the construction of a probabilistic method called Cancer Locator using hypermethylated target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions can be specific for one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subgroups of hypermethylated target regions that together exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.

[0258] ii. Hypomethylated variable target regions

[0259] Global hypomethylation is a phenomenon commonly observed in various cancers. See, e.g., Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (review article noting hypomethylation observations in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions that are normally methylated in healthy cells, such as repetitive elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, as well as intergenic regions, may show decreased methylation in tumor cells. Thus, in some embodiments, the epigenetic target region set includes hypomethylated variable target regions, where a decrease in the observed methylation level indicates an increased likelihood that the sample (e.g., a cfDNA sample) contains DNA produced by neoplastic cells (e.g., a tumor or cancer cells).

[0260] In some embodiments, the hypomethylated variable target regions include repetitive elements and / or intergenic regions. In some embodiments, the repetitive elements include one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.

[0261] For example, according to the hg19 or hg38 human genome constructs, exemplary specific genomic regions showing cancer-associated hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1. In some embodiments, the hypomethylated variable target regions overlap with or include one or both of these regions.

[0262] iii. CTCF binding regions

[0263] CTCF is a DNA-binding protein that contributes to chromatin organization and is typically colocalized with cohesin. Perturbations of CTCF binding sites have been reported in a variety of different cancers. See, for example, Katainen et al., Nature Genetics, doi:10.1038 / ng.3335, published online on June 8, 2015; Guo et al., Nat. Commun. 9:1520 (2018). CTCF binding generates recognizable patterns in cfDNA that can be detected by sequencing, such as by fragment length analysis. For example, details regarding sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016); WO 2018 / 009723; and US20170211143A1, each of which is incorporated herein by reference.

[0264] Thus, perturbations of CTCF binding result in variations in the cfDNA fragmentation pattern. Thus, CTCF binding sites represent one type of fragmentation variable target region.

[0265] There are many known CTCF binding sites. See, e.g., the CTCFBSDB (CTCF Binding Site Database) available on the Internet at insulatordb.uthsc.edu / ; Cuddapah et al., Genome Res. 19:24 - 32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18:708 - 14 (2011); Rhee et al., Cell. 147:1408 - 19 (2011), each of which is incorporated by reference. Exemplary CTCF binding sites are located at nucleotides 56014955 - 56016161 on chromosome 8 and nucleotides 95359169 - 95360473 on chromosome 13 according to, e.g., the hg19 or hg38 human genome constructs.

[0266] Thus, in some embodiments, the epigenetic target set includes CTCF binding regions. In some embodiments, the CTCF binding regions include at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 CTCF binding regions, e.g., such as one or more of those described above or in the Cuddapah et al., Martin et al., or Rhee et al. articles cited above or in the CTCFBSDB.

[0267] In some embodiments, at least some of the CTCF sites can be methylated or non - methylated, where the methylation status is associated with whether the cell is a cancer cell. In some embodiments, the epigenetic target set includes at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp in the upstream and / or downstream regions of the CTCF binding sites.

[0268] iv. Transcription start sites

[0269] Transcription start sites may also show perturbations in neoplastic cells. For example, the nucleosome organization at different transcription start sites in healthy cells of the hematopoietic lineage (which substantially contributes to cfDNA in healthy individuals) may be different from that at those transcription start sites in neoplastic cells. This results in different cfDNA patterns that can be detected by sequencing, e.g., as commonly discussed in: Snyder et al., Cell 164:57 - 68 (2016); WO 2018 / 009723; and US20170211143A1.

[0270] Therefore, perturbations in the transcription start site can also lead to variations in the fragmentation pattern of cfDNA. Therefore, the transcription start site also represents a type of variable target region for fragmentation.

[0271] Human transcription start sites are available from DBTSS (Database of Human Transcription Start Sites) available on the internet at dbtss.hgc.jp and described in Yamashita et al., Nucleic Acids Res. 34(Database Issue):D86-D89 (2006), which is incorporated herein by reference.

[0272] Thus, in some embodiments, the epigenetic target region group includes a transcription start site. In some embodiments, the transcription start site includes at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, for example, such as the transcription start sites listed in DBTSS. In some embodiments, at least some of the transcription start sites may be methylated or unmethylated, wherein the methylation state is associated with whether the cell is a cancer cell. In some embodiments, the epigenetic target region group includes at least 100bp, at least 200bp, at least 300bp, at least 400bp, at least 500bp, at least 750bp, at least 1000bp in the upstream and / or downstream regions of the transcription start site.

[0273] v. Copy number variation; focused amplification

[0274] Although copy number variations such as focused amplification are somatic mutations, they can be detected by sequencing based on read frequency in a manner similar to methods for detecting certain epigenetic changes such as methylation changes. Therefore, regions that can show copy number variations such as focused amplification in cancer can be included in the epigenetic target region group and can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the epigenetic target region group includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the aforementioned targets.

[0275] vi. Methylation control region

[0276] It can be useful to include control regions to facilitate data validation. In some embodiments, the epigenetic target set includes control regions that are expected to be methylated or unmethylated in substantially all samples, regardless of whether the DNA is derived from cancer cells or normal cells. In some embodiments, the epigenetic target set includes control hypomethylated regions that are expected to be hypomethylated in substantially all samples. In some embodiments, the epigenetic target set includes control hypermethylated regions that are expected to be hypermethylated in substantially all samples.

[0277] b. Sequence variable target set

[0278] In some embodiments, the sequence variable target set includes more than one region known to undergo somatic mutations in cancer (referred to herein as cancer-related mutations). Thus, the method can include determining whether cfDNA molecules corresponding to the sequence variable target set contain cancer-related mutations.

[0279] In some embodiments, the sequence variable target set targets more than one different gene or genomic region (“panel”), which is selected such that a defined proportion of subjects with cancer exhibit genetic variants or tumor markers in one or more of the different genes or genomic regions in the panel. The panel can be selected to limit the sequencing region to a fixed number of base pairs. The panel can be selected, for example, to sequence a desired amount of DNA by adjusting the affinity and / or amount of the probes as described elsewhere herein. The panel can be further selected to achieve a desired sequence read depth. The panel can be selected to achieve a desired sequence read depth or sequence read coverage for a certain number of sequencing base pairs. The panel can be selected to achieve the theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample.

[0280] Probes for detecting a region panel can include probes for detecting genomic regions of interest (hotspot regions) and nucleosome-sensing probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation affected by nucleosome binding patterns and GC sequence composition. Regions used herein can also include non-hotspot regions optimized based on nucleosome position and GC models.

[0281] Examples of lists of genomic positions of interest can be found in Tables 3 and 4. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 genes from Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 SNVs from Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 fusions from Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least a portion of at least 1, at least 2, or 3 indels from Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 genes from Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 SNVs from Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 fusions from Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 indels from Table 4. Each of these genomic positions of interest can be identified as a framework region or a hotspot region of a given panel. Examples of lists of hotspot genomic positions of interest can be found in Table 5. The coordinates in Table 5 are based on the hg19 assembly of the human genome, but those skilled in the art will be familiar with other assemblies and can identify the set of coordinates corresponding to exons, introns, codons, etc. indicated in the assembly of their choice.In some embodiments, the set of sequence-variable target regions used in the methods of the present disclosure includes at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 genes of Table 5. Each hot-spot genomic region lists several characteristics, including the associated gene, the chromosome on which it resides, the start and end positions of the genome representing the locus of the gene, the length of the locus of the gene in base pairs, the exons covered by the gene, and key characteristics (e.g., mutation type) that a given genomic region of interest may attempt to capture.

[0282] Table 3

[0283]

[0284] Table 4

[0285]

[0286] Table 5

[0287]

[0288]

[0289]

[0290] Additionally, or alternatively, suitable target region sets are available from the literature. For example, Gale et al., PLoS One 13:e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of a set of sequence-variable target regions. These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.

[0291] In some embodiments, the set of sequence-variable target regions includes target regions from at least 10, 20, 30, or 35 cancer-related genes such as the cancer-related genes listed above.

[0292] 3. Partitioning; Analysis of Epigenetic Features

[0293] In certain embodiments described herein, populations of different forms of nucleic acids (e.g., hypermethylated and hypomethylated DNA in a sample, such as a captured cfDNA set as described herein) can be physically partitioned based on one or more characteristics of the nucleic acids prior to analysis such as sequencing or tagging and sequencing. This approach can be used to determine, for example, whether hypermethylated variable epigenetic target regions exhibit hypermethylation characteristics of tumor cells, or whether hypomethylated variable epigenetic target regions exhibit hypomethylation characteristics of tumor cells. Additionally, by partitioning a heterogeneous nucleic acid population, one can increase rare signals, for example, by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, genetic variants that are present in hypermethylated DNA but less (or not at all) in hypomethylated DNA can be more readily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing more than one fraction of a sample, a multi-dimensional analysis of a single locus of the genome or nucleic acid species can be performed and thus greater sensitivity can be achieved.

[0294] In some cases, a heterogeneous nucleic acid sample is partitioned into two or more partitions (e.g., at least 3, 4, 5, 6, or 7 partitions). In some embodiments, each partition is differentially tagged. The tagged partitions can then be combined together for common sample preparation and / or sequencing. The partition-tag-combine steps can occur more than once, where each round of partitioning occurs based on a different characteristic (examples provided herein) and is tagged using differential tagging that is distinct from other partitions and partitioning means.

[0295] Examples of features that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins that bind DNA. The resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids not having one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; methylation level; type of methylation (e.g., 5-methylcytosine relative to other types of methylation such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins such as histones. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules that associate with nucleosomes and nucleic acid molecules that do not contain nucleosomes. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned based on nucleic acid length (e.g., molecules up to 160 bp and molecules having a length greater than 160 bp).

[0296] In some cases, each partition (representing a different nucleic acid form) is differentially labeled and the partitions are combined together prior to sequencing. In other cases, the different forms are sequenced separately.

[0297] Figure 1 The figure illustrates one embodiment of the present disclosure. A population 101 of different nucleic acids is partitioned 102 into two or more different partitions 103a, 103b. Each partition 103a, 103b represents a different nucleic acid form. Each partition is differentially tagged 104. Prior to sequencing 108, the tagged nucleic acids are combined together 107. The reads are analyzed computationally. The tags are used to sort the reads from different partitions. Analysis for detecting genetic variants can be performed at the partition-partition level as well as at the level of the entire nucleic acid population. For example, the analysis can include computational analysis to determine genetic variants in the nucleic acids in each partition, such as CNVs, SNVs, indels, fusions. In some cases, the computational analysis can include determining chromatin structure. For example, the coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage may be associated with higher nucleosome occupancy in a genomic region, while lower coverage may be associated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0298] A sample can include nucleic acids that are different with respect to modifications including post-replicative modifications of nucleotides and non-covalent binding to one or more proteins.

[0299] In one embodiment, the nucleic acid population is a nucleic acid population obtained from a serum, plasma, or blood sample of a subject suspected of having a neoplasm, tumor, or cancer or previously diagnosed with a neoplasm, tumor, or cancer. The nucleic acid population comprises nucleic acids having different methylation levels. Methylation can occur by any one or more post-replicative modifications or transcriptional modifications. Post-replicative modifications include modifications to nucleotide cytosine, particularly at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.

[0300] In some embodiments, the nucleic acids in the original population can be single-stranded and / or double-stranded. Partitioning based on the single-stranded versus double-stranded form of the nucleic acid can be achieved, for example, by using capture probes labeled for ssDNA partitioning and duplex adaptors for dsDNA partitioning.

[0301] The affinity agent can be an antibody, a natural binding partner, or a variant thereof having a desired specificity (Bock et al., Nat Biotech 28:1106-1114 (2010); Song et al., Nat Biotech 29:68-72 (2011)), or an artificial peptide specific for a given target selected, for example, by phage display.

[0302] Examples of capture moieties contemplated herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) as described herein.

[0303] Similarly, histone-binding proteins can be used for partitioning different forms of nucleic acids, which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4 (RbAp48) and SANT domain peptides.

[0304] For some affinity agents and modifications, although binding to the agent can occur in a substantially all-or-none manner depending on whether the nucleic acid bears the modification, the separation can be to some degree. In such cases, nucleic acids overrepresented in a modification bind to the agent to a greater extent than nucleic acids underrepresented in the modification. Optionally, nucleic acids having the modification can bind in an all-or-none manner. However, then, various levels of the modification can be eluted sequentially from the binding agent.

[0305] For example, in some embodiments, the partitioning can be binary or based on the degree / level of modification. For example, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMiner Methylated DNA Enrichment Kit (Thermo Fisher Scientific)). Subsequently, additional partitioning can include eluting fragments with different methylation levels by adjusting the salt concentration of the solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with a greater methylation level are eluted.

[0306] In some cases, the final partitioning represents nucleic acids with different degrees of modification (overrepresentative or under representative modifications). Overrepresentation and underrepresentation can be defined by the number of modifications carried by a nucleic acid relative to the median of the modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids of a sample is 2, then that modification of a nucleic acid containing more than two 5-methylcytosine residues is overrepresented, while a nucleic acid with 1 or 0 5-methylcytosine residues is underrepresented. The role of affinity separation is to enrich nucleic acids with overrepresented modifications in the binding phase from nucleic acids with underrepresented modifications in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted prior to subsequent processing.

[0307] When using the MethylMiner Methylated DNA Enrichment Kit (Thermo Fisher Scientific), various levels of methylation can be partitioned using sequential elution. For example, hypomethylated partitioning (e.g., no methylation) can be separated from methylated partitioning by contacting a nucleic acid population with the MBD attached to magnetic beads from the kit. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different methylation levels. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is again used to separate nucleic acids with a higher level of methylation from nucleic acids with a lower level of methylation. The elution and magnetic separation steps themselves can be repeated to generate various partitions, such as hypomethylated partitioning (e.g., representing no methylation), methylated partitioning (representing a low methylation level), and hypermethylated partitioning (representing a high methylation level).

[0308] In some methods, the nucleic acids that bind to the agent for affinity separation undergo a washing step. The washing step washes away the nucleic acids that bind weakly to the affinity agent. Such nucleic acids can enrich for nucleic acids having a degree of modification that is close to the mean or median (i.e., intermediate between the nucleic acids that remain bound to the solid phase and those that do not remain bound to the solid phase upon initial contact of the sample with the agent).

[0309] Affinity separation results in at least two and sometimes three or more partitions of nucleic acids having different degrees of modification. Although the partitions remain separate, the nucleic acids of at least one partition and usually two or three (or more) partitions are linked to nucleic acid tags, which are typically provided as part of an adaptor, wherein the nucleic acids in different partitions receive different tags that distinguish the members of one partition from the members of another partition. The tags of the nucleic acid molecules linked to the same partition can be the same or different from each other. However, if different from each other, the tags can have a common encoded portion so as to identify the molecules to which they are attached as a particular partition.

[0310] For more details regarding partitioning nucleic acid samples based on features such as methylation, see WO2018 / 119452, which is incorporated herein by reference.

[0311] In some embodiments, nucleic acid molecules can be fractionated into different partitions based on nucleic acid molecules that bind to a particular protein or a fragment thereof and nucleic acid molecules that do not bind to the particular protein or a fragment thereof.

[0312] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (such as histone methylation or acetylation), or enzymatic activities. Examples of proteins that can bind DNA and serve as a basis for fractionation can include, but are not limited to, Protein A and Protein G. Any suitable method can be used to fractionate nucleic acid molecules based on protein binding regions. Examples of methods for fractionating nucleic acid molecules based on protein binding regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow field-flow fractionation (AF4).

[0313] In some embodiments, partitioning of the nucleic acids is performed by contacting the nucleic acids with the methylation binding domain (“MBD”) of a methylation binding protein (“MBP”). The MBD binds 5-methylcytosine (5mC). The MBD is conjugated to paramagnetic beads such as M-280 streptavidin via a biotin linker. Partitioning into fractions having different degrees of methylation can be performed by eluting the fractions by increasing the NaCl concentration.

[0314] Examples of MBP contemplated herein include, but are not limited to:

[0315] (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine compared to unmodified cytosine.

[0316] (b) RPL26, PRP8, and DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine compared to unmodified cytosine.

[0317] (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind 5-formyl-cytosine compared to unmodified cytosine (Iurlaro et al., Genome Biol. 14:R119 (2013)).

[0318] (d) Antibodies specific for one or more methylated nucleobases.

[0319] Generally, elution varies with the number of methylation sites per molecule, and in the case of increasing salt concentration, the molecule has more methylation elution. To elute DNA into different populations based on the degree of methylation, one can use a series of elution buffers with increasing NaCl concentrations. The salt concentration can range from about 100 mM to about 2500 mM NaCl. In one embodiment, the process results in three (3) partitions. The molecule is contacted with a solution at a first salt concentration and including a molecule containing a methyl-binding domain, which molecule can be attached to a capture moiety such as streptavidin. At the first salt concentration, a population of molecules will bind to the MBD, while a population will remain unbound. The unbound population can be separated as the "low-methylated" population. For example, the first partition representing the low-methylated DNA form is the partition that remains unbound at a low salt concentration, such as 100 mM or 160 mM. The second partition representing moderately methylated DNA is eluted using an intermediate salt concentration, such as a concentration between 100 mM and 2000 mM. This is also separated from the sample. The third partition representing the hypermethylated DNA form is eluted using a high salt concentration, such as at least about 2000 mM.

[0320] a. Label the partitions

[0321] In some embodiments, two or more partitions, such as each partition, are differentially labeled. The label can be a molecule that contains information indicating the characteristics of the molecule associated with the label, such as a nucleic acid. For example, the molecule can carry a sample label (which distinguishes molecules in one sample from molecules in different samples), a partition label (which distinguishes molecules in one partition from molecules in different partitions), or a molecule label (which distinguishes different molecules from each other in both the case of adding unique and non-unique labels). In certain embodiments, the label can include a barcode or a combination of barcodes. As used herein, the term "barcode" refers to a nucleic acid molecule having a specific nucleotide sequence, or refers to the nucleotide sequence itself, depending on the context. The barcode can have, for example, between 10 and 100 nucleotides. Depending on the needs of a particular purpose, the set of barcodes can have degenerate sequences, or can have sequences with a specific Hamming distance. Thus, for example, a sample index, a partition index, or a molecule index can include a barcode or a combination of two barcodes, each barcode attached to a different end of the molecule.

[0322] Labels can be used to label individual polynucleotide population partitions so as to associate one label (or more than one label) with a specific partition. Optionally, labels can be used in embodiments of the present invention that do not employ a partitioning step. In some embodiments, a single label can be used to label a specific partition. In some embodiments, more than one different label can be used to label a specific partition. In embodiments where more than one different label is used to label a specific partition, the set of labels used to label one partition can be readily distinguished from the set of labels used to label other partitions. In some embodiments, the label can have additional functions. For example, the label can be used to index the sample source or to serve as a unique molecular identifier (which can be used to improve the quality of sequencing data by distinguishing sequencing errors and mutations, such as as described in Kinde et al., Proc Nat'l Acad Sci USA 108:9530-9535 (2011), Kou et al., PLoS ONE, 11:e0146638 (2016)) or to serve as a non-unique molecular identifier, such as as described in U.S. Patent No. 9,598,731. Similarly, in some embodiments, the label can have additional functions. For example, the label can be used to index the sample source or to serve as a non-unique molecular identifier (which can be used to improve the quality of sequencing data by distinguishing sequencing errors and mutations).

[0323] In one embodiment, partitioning and tagging includes tagging the molecules in each partition with a partition tag. After recombining the partitions and sequencing the molecules, the partition tag identifies the source partition. In another embodiment, different partitions are tagged with different sets of molecular tags, for example, including a pair of barcodes. In this way, each molecular barcode indicates the source partition and can also be used to distinguish molecules within the partition. For example, 35 barcodes of a first set can be used to tag the molecules in a first partition, and 35 barcodes of a second set can be used to tag the molecules in a second partition.

[0324] In some embodiments, after partitioning and tagging with partition tags, the molecules can be pooled for sequencing in a single run. In some embodiments, for example, in a step after adding the partition tags and pooling, a sample tag is added to the molecules. The sample tag can facilitate pooling of material from more than one sample for sequencing in a single sequencing run.

[0325] Optionally, in some embodiments, the partition tag can be associated with the sample as well as the partition. As a simple example, a first tag can indicate a first partition of a first sample; a second tag can indicate a second partition of the first sample; a third tag can indicate a first partition of a second sample; and a fourth tag can indicate a second partition of the second sample.

[0326] Although the tags can be attached to molecules that have been partitioned based on one or more characteristics, the ultimately tagged molecules in the library may no longer have that characteristic. For example, although single-stranded DNA molecules can be partitioned and tagged, the ultimately tagged molecules in the library may be double-stranded. Similarly, although DNA may be partitioned based on different methylation levels, in the final library, the tagged molecules derived from these molecules may be unmethylated. Thus, the tags attached to the molecules in the library generally indicate the characteristics of the "parent molecule" from which the ultimately tagged molecule originated, rather than necessarily indicating the characteristics of the tagged molecule itself.

[0327] For example, barcodes 1, 2, 3, 4, etc. are used to tag and label the molecules in a first partition; barcodes A, B, C, D, etc. are used to tag and label the molecules in a second partition; and barcodes a, b, c, d, etc. are used to tag and label the molecules in a third partition. The differentially tagged partitions can be combined prior to sequencing. The differentially tagged partitions can be sequenced separately or simultaneously together, for example, in the same flow cell of an Illumina sequencer.

[0328] After sequencing, the analysis of reads that detect genetic variants can be performed at the partition-partition level and at the level of the entire nucleic acid population. Reads from different partitions are screened using tags. The analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinate length, coverage, and / or copy number. In some embodiments, higher coverage may be associated with higher nucleosome occupancy in a genomic region, while lower coverage may be associated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0329] b. Determination of the 5-methylcytosine pattern of nucleic acids; bisulfite sequencing

[0330] Bisulfite-based sequencing and its variants provide another means of determining the methylation pattern of nucleic acids without relying on pre-sequencing partitioning based on methylation levels. In some embodiments, determining the methylation pattern includes differentiating 5-methylcytosine (5mC) from non-methylated cytosine. In some embodiments, determining the methylation pattern includes differentiating N-methyladenine from non-methylated adenine. In some embodiments, determining the methylation pattern includes differentiating 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC) from non-methylated cytosine. Examples of bisulfite sequencing include, but are not limited to, oxidative bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reduced bisulfite sequencing (redBS-seq). In some embodiments, determining the methylation pattern includes whole-genome bisulfite sequencing, e.g., as in MethylC-seq (Urich et al., Nature Protocols 10:475-483 (2015)). In some embodiments, determining the methylation pattern includes array-based methylation pattern determination, e.g., as in the methylation EPIC Beadchip or using an Illumina Infinium array (e.g., the HumanMethylation450 array) (see The Cancer Genome Atlas Research Network, Nature 507:315-322 (2014)). In some embodiments, determining the methylation pattern includes bisulfite PCR. In some embodiments, determining the methylation pattern includes EM-Seq (US 2013 / 0244237 A1). In some embodiments, determining the methylation pattern includes TAPS (WO2019 / 136413 A1).

[0331] Oxidative bisulfite sequencing (OX-BS-seq) is used to distinguish 5mC and 5hmC by first converting 5hmC to 5fC and then performing bisulfite sequencing. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish 5mc and 5hmC. In TAB-seq, 5hmC is protected by glycosylation. Then, before performing bisulfite sequencing, 5mC is converted to 5caC using Tet enzymes. Reductive bisulfite sequencing is used to distinguish 5fC from modified cytosines.

[0332] Typically, in bisulfite sequencing, a nucleic acid sample is split into two aliquots and one aliquot is treated with bisulfite. Bisulfite converts native cytosine and certain modified cytosine nucleotides (such as 5-formylcytosine or 5-carboxylcytosine) to uracil, while other modified cytosines (such as 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Comparison of the nucleic acid sequences of molecules from the two aliquots indicates which cytosines were converted to uracil and which were not. Thus, modified and unmodified cytosines can be determined. Initially splitting the sample into two aliquots is disadvantageous for samples containing only small amounts of nucleic acids and / or including heterogeneous cell / tissue sources such as body fluids containing cell-free DNA.

[0333] Thus, in some embodiments, bisulfite sequencing is performed, for example, as follows without initially dividing the sample into two aliquots. In some embodiments, the nucleic acids in a population are ligated to a capture moiety, such as any of the capture moieties described herein, i.e., a label that can be captured or immobilized. After the capture moiety is ligated to the sample nucleic acid, the sample nucleic acid serves as an amplification template. After amplification, the original template remains ligated to the capture moiety, but the amplicons are not ligated to the capture moiety.

[0334] The capture moiety can be ligated to the sample nucleic acid as a component of an adaptor, which can also provide amplification and / or sequencing primer binding sites. In some methods, the sample nucleic acid is ligated to adaptors at both ends, where the two adaptors carry the capture moiety. Preferably, any cytosine residues in the adaptor are modified, such as being modified with 5-methylcytosine, to protect against the action of bisulfite. In some cases, the capture moiety is linked to the original template via a cleavable linker (e.g., a photocleavable desthiobiotin-TEG or a uracil residue cleavable by USER TM enzyme, Chem.Commun.(Camb).51:3266-3269(2015)), in which case the capture moiety can be removed if desired.

[0335] The amplicons are denatured and contacted with an affinity reagent for capturing the tags. The original template binds to the affinity reagent, while the nucleic acid molecules generated by amplification do not bind. Thus, the original template can be separated from the nucleic acid molecules generated by amplification.

[0336] After the original template is separated from the nucleic acid molecules generated by amplification, the original template can be subjected to bisulfite treatment. Optionally, the amplification product can be subjected to bisulfite treatment while the original template population is not. After such treatment, the corresponding populations can be amplified (in the case of the original template population, this converts uracil to thymine). The populations can also be subjected to biotin probe hybridization for capture. Then the corresponding populations are analyzed and the sequences are compared to determine which cytosines were 5-methylated (or 5-hydroxymethylated) in the original sample. Detection of a T nucleotide in the template population (corresponding to an unmethylated cytosine that was converted to uracil) and a C nucleotide at the corresponding position in the amplified population indicates an unmodified C. The presence of a C at the corresponding position in the original template and the amplified population indicates a modified C in the original sample.

[0337] In some embodiments, a method uses sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation to add molecular tags to a DNA library (see WO 2018 / 119452, for example in Figure 4)。This process is carried out by labeling adapters (e.g., biotin), DNA-seq amplification of the full library, parental molecule recovery (e.g., streptavidin bead pull-down), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosine at single-base resolution by sequential NGS preparative amplification of parental library molecules with and without bisulfite treatment. This can be achieved by modifying the 5-methylated NGS adapter (directional adapter; Y-shaped / fork-shaped, with 5-methylcytosine substitution) used in BIS-seq with a marker (e.g., biotin) on one of the two adapter strands. Sample DNA molecules are ligated adapters and amplified (e.g., by PCR). Since only parental molecules will have labeled adapter ends, they can be selectively recovered from their amplified progeny by a label-specific capture method (e.g., streptavidin magnetic beads). Since parental molecules retain the 5-methylation marker, bisulfite conversion on the captured library will result in a single-base resolution 5-methylation status after BIS-seq, retaining the molecular information into the corresponding DNA-seq. In some embodiments, the bisulfite-treated library can be combined with the untreated library before capture / NGS by adding sample tag DNA sequences in a standard multiplex NGS workflow. As with the BIS-seq workflow, bioinformatics analysis can be performed for genome alignment and 5-methylated base identification. In summary, the method provides the ability to selectively recover parental, ligated molecules carrying the 5-methylcytosine marker after library amplification, allowing for parallel processing of bisulfite-converted DNA. This overcomes the damaging nature of bisulfite treatment to the quality / sensitivity of DNA-seq information extracted from the workflow. With this method, the recovered ligated, parental DNA molecules (via labeled adapters) allow for amplification of the complete DNA library and parallel application of treatments that cause epigenetic DNA modifications. The present disclosure discusses the use of the BIS-seq method to identify cytosine-5-methylation (5-methylcytosine), but the BIS-seq method is not required in many embodiments. Variations of BIS-seq have been developed to identify hydroxymethylcytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxylcytosine. These methods can be implemented with the sequential / parallel library preparation described herein.

[0338] c. Alternative methods for analyzing modified nucleic acids

[0339] In some such methods, depending on the degree of modification, the population of nucleic acids with different degrees of modification (e.g., 0, 1, 2, 3, 4, 5 or more methyl groups per nucleic acid molecule) is contacted with an adaptor before fractionation. The adaptor is attached to one or both ends of the nucleic acid molecules in the population. Preferably, the adaptor includes a sufficient number of different tags such that the number of tag combinations results in a low probability, e.g., 95%, 99% or 99.9%, that two nucleic acids with the same start and end points receive the same tag combination. After attaching the adaptor, the nucleic acids are amplified from primers that bind to primer binding sites within the adaptor. The adaptors, whether with the same or different tags, may include the same or different primer binding sites, but preferably the adaptors include the same primer binding sites. After amplification, the nucleic acids are contacted with an agent that preferably binds to the modified nucleic acids (such as those previously described). The nucleic acids are separated into at least two fractions that differ in the degree of binding of the modified nucleic acids to the agent. For example, if the agent has an affinity for the modified nucleic acids, the nucleic acids with overrepresented modification (compared to the median representation in the population) preferentially bind the agent, while the nucleic acids with underrepresented modification do not bind the agent or are more readily eluted from the agent. After separation, the different fractions can then undergo additional processing steps, which typically include additional amplification and sequence analysis in parallel but separately. The sequence data from different fractions can then be compared.

[0340] Such a separation scheme can be carried out using the following exemplary procedure. The nucleic acids are ligated at both ends to a Y-shaped adaptor that includes a primer binding site and a tag. The molecules are amplified. The amplified molecules are then fractionated by contacting them with an antibody that preferentially binds 5-methylcytosine to produce two fractions. One fraction includes the original molecules lacking methylation and the amplified copies with lost methylation. The other fraction includes the original DNA molecules with methylation. The two fractions are then processed and sequenced separately, with further amplification of the methylated fraction. The sequence data from the two fractions can then be compared. In this example, the tags are not used to distinguish methylated DNA from unmethylated DNA, but rather to distinguish different molecules within these fractions so that one can determine whether reads with the same start and end points are based on the same or different molecules.

[0341] The present disclosure provides additional methods for analyzing a nucleic acid population, wherein at least some of the nucleic acids include one or more modified cytosine residues, such as 5-methylcytosine and any other modification previously described. In these methods, the nucleic acid population is contacted with an adaptor that includes one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosines in the primer binding region of the adaptor are modified. The adaptor is attached to both ends of the nucleic acid molecules in the population. Preferably, the adaptor includes a sufficient number of different tags such that the number of tag combinations results in a low probability, e.g., 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive the same tag combination. The primer binding sites in such an adaptor can be the same or different, but are preferably the same. After attaching the adaptor, the nucleic acids are amplified from a primer that binds to the primer binding site of the adaptor. The amplified nucleic acids are divided into a first aliquot and a second aliquot. The sequence data of the first aliquot is determined with or without additional processing. Thus, the sequence data of the molecules in the first aliquot is determined regardless of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosine to uracil. The bisulfite-treated nucleic acids then undergo amplification that is primed by a primer directed to the original primer binding site of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (as distinct from their amplification products) are amplifiable because these nucleic acids retain cytosine at the primer binding site of the adaptor, while the amplification products have lost the methylation of these cytosine residues that have undergone conversion to uracil during the bisulfite treatment. Thus, only the original molecules in the population are amplified, at least some of which are methylated. After amplification, these nucleic acids undergo sequence analysis. Comparing the sequences determined from the first and second aliquots can indicate, among other things, which cytosines in the nucleic acid population are methylated.

[0342] Such an analysis can be performed using the following exemplary procedure. Methylated DNA is ligated at both ends to a Y-shaped adaptor that includes a primer binding site and a tag. The cytosines in the adaptor are 5-methylated. Methylation of the primer is used to protect the primer binding site in subsequent bisulfite steps. After attaching the adaptor, the DNA molecule is amplified. The amplification product is split into two aliquots for sequencing with and without bisulfite treatment. The aliquot not undergoing bisulfite sequencing can undergo sequence analysis with or without additional treatment. The other aliquot is treated with bisulfite, which converts unmethylated cytosines to uracil. Only the primer binding site protected by cytosine methylation can support amplification when contacted with a primer specific for the original primer binding site. Thus, only the original molecule and not the copies from the first amplification undergo additional amplification. The molecules from the additional amplification then undergo sequence analysis. The sequences from the two aliquots can then be compared. As in the separation schemes discussed above, the nucleic acid tag in the adaptor is not used to distinguish methylated DNA from unmethylated DNA, but rather to distinguish nucleic acid molecules within the same partition.

[0343] d. Methylation-sensitive PCR

[0344] In some embodiments, methylation in hypermethylated variable regions and / or hypomethylated variable regions is evaluated using methylation-sensitive amplification. By adapting known methods to the methods described herein, various steps can be made methylation-sensitive.

[0345] For example, a sample can be split into aliquots, e.g., before or after a capture step as described herein, and one aliquot can be digested with a methylation-sensitive restriction enzyme, e.g., as described in Moore et al., Methods MolBiol. 325:239-49 (2006), which is incorporated herein by reference. Unmethylated sequences are digested in this aliquot. The digested and undigested aliquots can then be processed through appropriate steps (amplification, optionally tagging, sequencing, etc.) as described herein, and the sequences are analyzed to determine the degree of digestion in the treated sample, which reflects the presence of unmethylated cytosines. Optionally, splitting into aliquots can be avoided by: amplifying the sample, separating the amplified material from the original template, and then digesting the original material with a methylation-sensitive restriction enzyme followed by further amplification, e.g., as discussed above with respect to bisulfite sequencing.

[0346] In another example, a sample can be divided into aliquots and one aliquot can be processed prior to capture to convert non-methylated cytosine to uracil, e.g., as described in US 2003 / 0082600, which is incorporated herein by reference. Conversion of non-methylated cytosine to uracil will reduce the capture efficiency of target regions with low methylation by altering the sequence of the region. The processed and unprocessed aliquots can then be subjected to the appropriate steps (capture, amplification, optionally tagging, sequencing, etc.) as described herein, and the sequences are analyzed to determine the extent of depletion of the target region in the processed sample, which reflects the presence of non-methylated cytosine.

[0347] 4. Subject

[0348] In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject having neoplasia. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject suspected of having neoplasia. In some embodiments, DNA (e.g., cfDNA) is obtained from a subject in remission from a tumor, cancer, or neoplasia (e.g., after chemotherapy, surgical resection, radiotherapy, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia can be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the lung. In some embodiments, the cancer, tumor, or neoplasm or suspected cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the breast. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the prostate. In any of the foregoing embodiments, the subject can be a human subject.

[0349] In some embodiments, the subject has been previously diagnosed with cancer, e.g., any cancer mentioned above or elsewhere herein. Such a subject may have previously received one or more prior cancer treatments, e.g., surgery, chemotherapy, radiation, and / or immunotherapy. In some embodiments, samples (e.g., cfDNA) are obtained from a subject previously diagnosed and treated at one or more preselected time points after one or more prior cancer treatments.

[0350] Samples obtained from a subject (e.g., cfDNA) can be sequenced to provide a set of sequence information, and the sequencing can include sequencing captured DNA molecules of a set of sequence variable target regions to a deeper sequencing depth than captured DNA molecules of an epigenomic target region, as described in detail elsewhere herein.

[0351] 5. Exemplary methods for molecular tagging of libraries of MBD beads partitions

[0352] Exemplary methods for molecular tagging of libraries of MBD beads partitions by NGS are as follows:

[0353] i) Physically partition the extracted DNA sample (e.g., plasma DNA extracted from a human sample, which has optionally undergone target capture as described herein) using a methyl-binding domain protein-bead purification kit, saving all elutions for downstream processing.

[0354] ii) Apply differential molecular tags and adapter sequences that initiate NGS in parallel to each partition. For example, hypermethylated, residual methylated ("washed"), and hypomethylated partitions are ligated to NGS adapters with molecular tags.

[0355] iii) Recombine all tagged partitions and then amplify using adapter-specific DNA primer sequences.

[0356] iv) Capture / hybridize the recombined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).

[0357] v) Re-amplify the captured DNA library, adding sample tags. Different samples are pooled and multiplexed on an NGS instrument.

[0358] vi) Bioinformatics analysis of the NGS data, where the molecular tags are used to identify unique molecules and to deconvolute the samples into differentially MBD-partitioned molecules. This analysis can generate information about relative 5-methylcytosine of genomic regions simultaneously with standard genetic sequencing / variant detection.

[0359] The exemplary methods set forth above can also include any compatible features of the methods according to the present disclosure set forth elsewhere herein.

[0360] 6. Exemplary workflow

[0361] An exemplary workflow for partitioning and library preparation is provided here. In some embodiments, some or all features of the partitioning and library preparation workflow can be used in combination. The exemplary workflow set forth above can also include any compatible features of the methods according to the present disclosure set forth elsewhere herein.

[0362] a. Partitioning

[0363] In some embodiments, sample DNA (e.g., between 1 ng and 300 ng) is mixed with an appropriate amount of methyl-binding domain (MBD) buffer (the amount of MBD buffer depending on the amount of DNA used) and magnetic beads conjugated to MBD protein, and incubated overnight. Methylated DNA (hypermethylated DNA) binds to the MBD protein on the magnetic beads during this incubation. Unmethylated (hypomethylated DNA) or less methylated (intermediately methylated) DNA is washed off the beads with a buffer containing increasing concentrations of salt. For example, one, two, or more fractions containing unmethylated DNA, hypomethylated DNA, and / or intermediately methylated DNA can be obtained from such washes. Finally, highly methylated DNA (hypermethylated DNA) is eluted from the MBD protein using a high-salt buffer. In some embodiments, these washes result in three partitions of DNA with increasing levels of methylation (a hypomethylated partition, an intermediately methylated partition, and a hypermethylated partition).

[0364] In some embodiments, the three partitions of DNA are desalted and concentrated in preparation for the enzymatic steps of library preparation.

[0365] b. Library Preparation

[0366] In some embodiments (e.g., after concentrating the DNA in the partitions), the partitioned DNA is made ligatable, e.g., by extending the end overhangs of the extended DNA molecules, adding adenosine residues to the 3' ends of the fragments, and phosphorylating the 5' ends of each DNA fragment. DNA ligase and adapters are added to ligate each partitioned DNA molecule to an adapter at each end. These adapters contain partition tags (e.g., non-random, non-unique barcodes) distinguishable from the partition tags in the adapters used in the other partitions. After ligation, the three partitions are pooled together and amplified (e.g., by PCR, such as with primers specific to the adapters).

[0367] After PCR, the amplified DNA can be washed and concentrated before capture. The amplified DNA is contacted with a collection of probes that target specific regions of interest described herein, which can be, for example, biotinylated RNA probes. The mixture is incubated overnight, for example, in a salt buffer. The probes are captured (e.g., using streptavidin magnetic beads) and separated from the un-captured amplified DNA, such as by a series of salt washes, to provide a captured DNA set. After capture, the DNA of the captured set is amplified by PCR. In some embodiments, the PCR primers contain sample tags, such that the sample tags are incorporated into the DNA molecules. In some embodiments, DNA from different samples is pooled together and then multiplex sequenced, for example, using an Illumina NovaSeq sequencer.

[0368] III. General Features of the Method

[0369] 1. Sample

[0370] The sample can be any biological sample isolated from a subject. The sample can be a bodily sample. The sample can include bodily tissues such as a known or suspected solid tumor, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsy, cerebrospinal fluid, synovial fluid, lymph fluid, ascites, interstitial fluid or extracellular fluid, fluid in the spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. The sample is preferably a bodily fluid, particularly blood and its fractions, and urine. The sample can be in the form initially isolated from the subject or can have undergone additional processing to remove or add components, such as cells, or to enrich one component relative to another. Thus, the preferred bodily fluids for analysis are plasma or serum containing cell-free nucleic acids. The sample can be isolated or obtained from the subject and transported to a sample analysis site. The sample can be stored and transported at a desired temperature, such as room temperature, 4 °C, -20 °C, and / or -80 °C. The sample can be isolated or obtained from the subject at the sample analysis site. The subject can be a human, mammal, animal, companion animal, service animal, or pet. The subject can have cancer. The subject can not have cancer or detectable cancer symptoms. The subject may have been treated with one or more cancer therapies, such as any one or more of chemotherapy, antibodies, vaccines, or biologics. The subject may be in remission. The subject may or may not be diagnosed as being predisposed to cancer or any cancer-related genetic mutations / disorders.

[0371] The volume of plasma can depend on the read depth required for the sequencing region. Exemplary volumes are 0.4 ml - 40 ml, 5 ml - 20 ml, 10 ml - 20 ml. For example, the volume can be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled can be 5 mL to 20 mL.

[0372] The sample can contain different amounts of nucleic acid, and the amount includes genomic equivalents. For example, a sample of about 30 ng of DNA can contain about 10,000 (10 4 ) haploid human genome equivalents, and in the case of cfDNA, contains about 200 billion (2x10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, contains about 600 billion individual molecules.

[0373] The sample can contain nucleic acids from different sources, such as cellular and cell-free nucleic acids from the same subject, cellular and cell-free nucleic acids from different subjects. The sample can contain nucleic acids carrying mutations. For example, the sample can contain DNA carrying germline mutations and / or somatic mutations. A germline mutation refers to a mutation present in the germline DNA of a subject. A somatic mutation refers to a mutation originating from a somatic cell of a subject, such as a cancer cell. The sample can contain DNA carrying cancer-related mutations (e.g., cancer-related somatic mutations). The sample can contain epigenetic variants (i.e., chemical modifications or protein modifications), where the epigenetic variant is associated with the presence of a genetic variant such as a cancer-related mutation. In some embodiments, the sample contains an epigenetic variant associated with the presence of a genetic variant, where the sample does not contain the genetic variant.

[0374] Exemplary amounts of cell-free nucleic acids in a sample prior to amplification range from about 1 fg to about 1 μg, such as from 1 pg to 200 ng, from 1 ng to 100 ng, from 10 ng to 1000 ng. For example, the amount of cell-free nucleic acid molecules can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng. The amount of cell-free nucleic acid molecules can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng. The amount of cell-free nucleic acid molecules can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, 200 ng, 250 ng, or 300 ng. The method can include obtaining from 1 femtogram (fg) to 200 ng.

[0375] Cell-free nucleic acids are nucleic acids that are not contained within cells or otherwise bound to cells, or in other words, are nucleic acids that remain in a sample after removal of intact cells. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death processes, such as necrosis and apoptosis. Some cell-free nucleic acids are released from cancer cells, such as circulating tumor DNA (ctDNA), into body fluids. Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.

[0376] Cell-free nucleic acids have an exemplary size distribution of about 100 - 500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of the molecules, a mode of about 168 nucleotides, and a second minor peak in the range of 240 to 440 nucleotides.

[0377] Cell-free nucleic acids can be isolated from body fluids by a fractionation or partitioning step in which, as found in solution, the cell-free nucleic acids are separated from intact cells and other insoluble components of the body fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid can be lysed and the cell-free nucleic acids and cellular nucleic acids can be processed together. Typically, after addition of buffer and washing steps, the nucleic acids can be precipitated with alcohol. Further cleaning steps such as silica-based columns can be used to remove contaminants or salts. Nonspecific bulk carrier nucleic acids, such as C1DNA, or DNA or proteins for bisulfite sequencing, hybridization, and / or ligation can be added throughout the reaction to optimize certain aspects of the procedure such as yield.

[0378] After such processing, the sample can contain nucleic acids in various forms, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, the single-stranded DNA and RNA can be converted to double-stranded form so that they are included in subsequent processing and analysis steps.

[0379] The double-stranded DNA molecules in the sample and the single-stranded nucleic acid molecules that have been converted to double-stranded DNA molecules can be ligated to adapters at one or both ends. Typically, the double-stranded molecules are blunt-ended by treatment with a polymerase having 5'-3' polymerase and 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotides. Klenow fragment and T4 polymerase are examples of suitable polymerases. The blunt-ended DNA molecules can be ligated to at least partially double-stranded adapters (e.g., Y-shaped adapters or bell-shaped adapters). Alternatively, complementary nucleotides can be added to the blunt ends of the sample nucleic acids and adapters to facilitate ligation. Both blunt-end ligation and sticky-end ligation are envisioned herein. In blunt-end ligation, both the nucleic acid molecule and the adapter tag have blunt ends. In sticky-end ligation, typically, the nucleic acid molecule has an "A" overhang and the adapter has a "T" overhang.

[0380] 2. Tags

[0381] Tags containing barcodes can be incorporated into the adapters or otherwise linked to the adapters. The tags can be incorporated by ligation, overlap extension PCR, and other methods.

[0382] a. Molecular tagging strategies

[0383] Molecular tagging refers to a tagging practice that allows one to distinguish the molecules from which sequence reads are derived. Tagging strategies can be classified into unique tagging strategies and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample carry different tags such that reads can be assigned to the original molecule based on the tag information alone. Tags used in such an approach are sometimes referred to as "unique tags". In non-unique tagging, different molecules in the same sample can carry the same tag such that information other than the tag information is used to assign sequence reads to the original molecule. Such information can include start and stop coordinates, coordinates to which the molecule maps, individual start or stop coordinates, etc. Tags used in such an approach are sometimes referred to as "non-unique tags". Thus, it is not necessary to tag each molecule in the sample uniquely. It is sufficient to tag uniquely the molecules in the sample that fall into recognizable categories. Thus, molecules in different recognizable families can carry the same tag without loss of information about the identity of the tagged molecule.

[0384] In certain embodiments of non-unique tagging, the number of different tags used can be sufficient such that the likelihood that all molecules in a particular group carry different tags is very high (e.g., at least 99%, at least 99.9%, at least 99.99%, or at least 99.999%). It should be noted that when barcodes are used as tags and when barcodes are attached to the ends of molecules, for example, randomly, the combinations of barcodes can together constitute a tag. In terms of this number, it is a function of the number of molecules called. For example, a category can be all molecules that map to the same start-stop position on a reference genome. A category can be all molecules that span a particular genetic locus, e.g., a particular base or a particular region (e.g., up to 100 bases or a gene or gene exon). In certain embodiments, the number of different tags used to uniquely identify multiple molecules z in a class can be any one of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11*z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z, or 100*z (e.g., the lower limit) and any one of 100,000*z, 10,000*z, 1000*z, or 100*z (e.g., the upper limit).

[0385] For example, in a sample of human cell-free DNA of about 3 ng to 30 ng, one would expect about 10 3 -10 4Individual molecules are mapped to specific nucleotide coordinates, and molecules with any starting coordinate between about 3 and 10 share the same ending coordinate. Thus, about 50 to about 50,000 different tags (e.g., barcode combinations between about 6 and 220) are sufficient to uniquely tag all such molecules. To uniquely tag all 10 3 -10 4 molecules mapped across one nucleotide coordinate would require about 1 million to about 20 million different tags.

[0386] Generally, the designation of unique tag barcodes or non-unique tag barcodes in a reaction follows the methods and systems described in U.S. Patent Applications 20010053519, 20030152490, 20110160078 and U.S. Patents No. 6,582,908, No. 7,537,898 and No. 9,598,731. Tags can be ligated to the sample nucleic acid randomly or non-randomly.

[0387] In some embodiments, the tagged nucleic acid is sequenced after being loaded into a microplate. The microplate can have 96, 384 or 1536 microwells. In some cases, they are introduced at an expected ratio of unique tags to microwells. For example, unique tags can be loaded such that more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, unique tags can be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the average number of unique tags loaded per sample genome is less than or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags per genomic sample.

[0388] A preferred format uses 20 - 50 different tags (e.g., barcodes) attached to both ends of the target nucleic acid. For example, 35 different tags (e.g., barcodes) are attached to both ends of the target molecule, creating a 35×35 arrangement, which equals 1225 tag combinations for 35 tags. The number of such tags is sufficient such that different molecules with the same start and end points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different tag combinations. Other barcode combinations include any number between 10 and 500, e.g., approximately 15x15, approximately 35x35, approximately 75x75, approximately 100x100, approximately 250x250, approximately 500x500.

[0389] In some cases, the unique tag can be a predetermined sequence or an oligonucleotide of a random or semi - random sequence. In other cases, more than one barcode can be used such that the barcodes do not have to be unique relative to each other among the more than one barcode. In this instance, the barcodes can be linked to individual molecules such that the combination of the barcode and the sequence to which it can be linked produces a unique sequence that can be traced individually. As described herein, the detection of non - unique barcodes in combination with sequence data at the beginning (start) and end (termination) portions of the sequence read can allow the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequence read can also be used to assign a unique identity to such a molecule. As described herein, fragments from a single - stranded nucleic acid to which a unique identity has been assigned can thereby allow subsequent identification of fragments from the parental strand.

[0390] 3. Amplification

[0391] Sample nucleic acids flanked by adapters can be amplified by PCR and other amplification methods. Amplification is typically initiated by the binding of primers to primer - binding sites in the adapters flanking the DNA molecule to be amplified. The amplification method can involve cycles of denaturation, annealing, and extension caused by thermal cycling, or can be isothermal, such as in transcription - mediated amplification. Other amplification methods include ligase - chain reaction, strand - displacement amplification, nucleic acid sequence - based amplification, and self - sustaining sequence replication.

[0392] Preferably, the method performs dsDNA "T / A ligation" with T - tail and C - tail adapters, which results in amplification of at least 50%, 60%, 70%, or 80% of the double - stranded nucleic acids before ligation to the adapters. Preferably, the method increases the amount or number of amplified molecules by at least 10%, 15%, or 20% relative to a control method performed with only T - tail adapters.

[0393] 4. Bait set; capture moiety; enrichment

[0394] As discussed above, nucleic acids in a sample can undergo a capture step where molecules having a target sequence are captured for subsequent analysis. Target capture can include using a bait set that includes oligonucleotide baits labeled with a capture moiety such as biotin or other examples mentioned below. The probes can have sequences selected to tile across a set of regions, such as a gene. In some embodiments, as discussed elsewhere herein, for those of target sets such as sequence variable target sets and epigenetic target sets, the bait sets can have relatively high and low capture yields, respectively. Such bait sets are combined with the sample under conditions that allow hybridization of the target molecules to the baits. Then, the captured molecules are separated using the capture moiety, e.g., a bead-based streptavidin biotin capture moiety. Such methods are further described, for example, in U.S. Patent No. 9,850,523, published on December 26, 2017, which is incorporated herein by reference.

[0395] Capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids comprising a specific nucleotide sequence, haptens recognized by an antibody, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety attached to an analyte is captured by its binding pair attached to a separable moiety, such as a magnetically attractable particle or a large particle that can be sedimented by centrifugation. A capture moiety can be any type of molecule that allows affinity separation of a nucleic acid bearing the capture moiety from a nucleic acid lacking the capture moiety. Exemplary capture moieties are biotin, which allows affinity separation by binding to streptavidin attached or attachable to a solid phase; or an oligonucleotide, which allows affinity separation by binding to a complementary oligonucleotide attached or attachable to a solid phase.

[0396] 5. Sequencing

[0397] Optionally flanked by adaptors, sample nucleic acids, with or without pre-amplification, are typically sequenced. Optionally utilized sequencing methods or commercially available formats include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing (NGS), single molecule synthesis sequencing (SMSS) (Helicos), massively parallel sequencing, clonal single molecule arrays (Solexa), shotgun sequencing, Ion Torrent, Oxford nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms. The sequencing reaction can be carried out in a variety of sample processing units, which can include multiple lane, multi-channel, multi-well, or other devices that can process multiple sets of samples substantially simultaneously. The sample processing unit can also include multiple sample chambers to enable simultaneous processing of multiple runs.

[0398] The sequencing reaction can be performed on one or more nucleic acid fragment types or regions containing markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragments present in the sample. The sequencing reaction can be performed on at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome. In other cases, the sequencing reaction can be performed on less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome.

[0399] Multiplex sequencing techniques can be used to perform simultaneous sequencing reactions. In some embodiments, cell-free polynucleotides are sequenced with at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, cell-free polynucleotides are sequenced with fewer than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. The sequencing reactions are typically performed sequentially or simultaneously. Subsequent data analysis is typically performed on all or a portion of the sequencing reactions. In some embodiments, data analysis is performed on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, data analysis is performed on fewer than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An example of read depth is from about 1000 to about 50000 reads per locus (e.g., base position).

[0400] a. Differential depth of sequencing

[0401] In some embodiments, nucleic acids corresponding to the sequence variable targetome are sequenced to a deeper sequencing depth than nucleic acids corresponding to the epigenetic targetome. For example, the sequencing depth of nucleic acids corresponding to the sequence variant targetome can be at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold deeper than the sequencing depth of nucleic acids corresponding to the epigenetic targetome, or 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, 14-fold to 15-fold, or 15-fold to 100-fold deeper. In some embodiments, the sequencing depth is at least 2-fold deeper. In some embodiments, the sequencing depth is at least 5-fold deeper. In some embodiments, the sequencing depth is at least 10-fold deeper. In some embodiments, the sequencing depth is 4-fold to 10-fold deeper. In some embodiments, the sequencing depth is 4-fold to 100-fold deeper. Each of these embodiments relates to the degree to which nucleic acids corresponding to the sequence variable targetome are sequenced to a deeper sequencing depth than nucleic acids corresponding to the epigenetic targetome.

[0402] In some embodiments, the captured cfDNA corresponding to the sequence variable targetome and the captured cfDNA corresponding to the epigenetic targetome are sequenced simultaneously, e.g., in the same sequencing flow cell (such as a flow cell of an Illumina sequencer) and / or in the same composition, which can be a pooled composition obtained by recombining separately captured groups, or a composition obtained by capturing cfDNA corresponding to the sequence variable targetome and captured cfDNA corresponding to the epigenetic targetome in the same container.

[0403] 6. Analysis

[0404] Sequencing can generate more than one sequence read or read. Sequence reads or reads can include data of nucleotide sequences having a length less than about 150 bases or a length less than about 90 bases. In some embodiments, the length of the read is between about 80 bases and about 90 bases, for example, about 85 bases. In some embodiments, the methods of the present disclosure are applied to very short reads, for example, having a length less than about 50 bases or about 30 bases. Sequence read data can include sequence data as well as meta-information. Sequence read data can be stored in any suitable file format, including, for example, VCF files, FASTA files, or FASTQ files.

[0405] FASTA can refer to a computer program for retrieving sequence databases, and the name FASTA can also refer to a standard file format. FASTA is described, for example, by Pearson & Lipman, 1988, Improved tools for biological sequence comparison, PNAS 85:2444 - 2448, which is hereby incorporated by reference in its entirety. Sequences in FASTA format begin with a single-line description, followed by lines of sequence data. The description line is distinguished from the sequence data by a greater-than (“>”) symbol in the first column. The word following the “>” symbol is the identifier of the sequence, and the remainder of the line is a description (both are optional). There must be no space between the “>” and the first letter of the identifier. It is recommended that all lines of text be less than 80 characters. If another line starting with “>” appears, the sequence ends; this indicates the start of another sequence.

[0406] The FASTQ format is a text-based format for storing biological sequences (usually nucleotide sequences) and their corresponding quality scores. It is similar to the FASTA format but has quality scores after the sequence data. For brevity, both the sequence letters and the quality scores are encoded using a single ASCII character. The FASTQ format is the de facto standard for storing the output of high-throughput sequencing instruments such as the Illumina Genome Analyzer, as described, for example, by Cock et al. (“The Sanger FASTQ file format for sequences with quality scores, and the Solexa / Illumina FASTQ variants,” Nucleic Acids Res 38(6):1767 - 1771, 2009), which is hereby incorporated by reference in its entirety.

[0407] For FASTA and FASTQ files, the meta-information includes the description line but not the sequence data lines. In some embodiments, for FASTQ files, the meta-information includes quality scores. For FASTA and FASTQ files, the sequence data begins after the description line and is typically presented using a subset of the IUPAC ambiguity codes, optionally with "-". In one embodiment, the sequence data can use the A, T, C, G, and N characters, optionally including "-" as needed or including U (e.g., to represent a gap or uracil).

[0408] In some embodiments, at least one primary sequence read file and the output file are stored as plain text files (e.g., using an encoding such as ASCII, ISO / IEC 646, EBCDIC, UTF-8, or UTF-16). The computer systems provided by the present disclosure can include a text editor program capable of opening plain text files. A text editor program can refer to a computer program capable of presenting the content of a text file (such as a plain text file) on a computer screen and allowing a person to edit the text (e.g., using a monitor, keyboard, and mouse). Examples of text editors include, but are not limited to, Microsoft Word, emacs, pico, vi, BBEdit, and TextWrangler. The text editor program can be capable of displaying the plain text file in a human-readable format on the computer screen, showing the meta-information and sequence reads (e.g., not binary-encoded but using alphanumeric characters as they can be used for printing or human writing).

[0409] Although the methods have been discussed with reference to FASTA or FASTQ files, the methods and systems of the present disclosure can be used to compress any suitable sequence file format, including, for example, files in Variant Call Format (VCF) format. A typical VCF file can include a header section and a data section. The header contains any number of meta-information lines, each starting with the character '##', and a TAB-delimited field definition line starting with a single '#' character. The field definition line names eight required columns, and the body section contains data lines that populate the columns defined by the field definition line. The VCF format is described, for example, by Danecek et al. ("The variant call format and VCFtools," Bioinformatics 27(15):2156-2158, 2011), which is hereby incorporated by reference in its entirety. The header section can be regarded as the meta-information to be written to the compressed file, and the data section can be regarded as lines, each of which can be stored in the main file only if it is unique.

[0410] Some embodiments provide for the assembly of sequence reads. For example, in assembly by alignment, sequence reads are aligned to each other or to a reference sequence. By aligning each read, and then to a reference genome, all reads are positioned relative to each other to create an assembly. Additionally, aligning or mapping sequence reads to a reference sequence can also be used to identify variant sequences in the sequence reads. Identifying variant sequences can be used in combination with the methods and systems described herein to further aid in the diagnosis or prognosis of a disease or condition or for guiding treatment decisions.

[0411] In some embodiments, any or all of the steps are automated. Optionally, the methods of the present disclosure may be implemented in whole or in part in one or more dedicated programs, for example each optionally written in a compiled language such as C++, and then compiled and distributed in binary. The methods of the present disclosure may be implemented in whole or in part as a module within an existing sequence analysis platform or by calling functions within an existing sequence analysis platform. In some embodiments, the methods of the present disclosure include multiple steps that are all automatically invoked in response to a single initiation queue (e.g., one event or combination of events from a triggering event such as human activity, another computer program, or a machine). Thus, the present disclosure provides methods in which any step or any combination of steps can occur automatically in response to a queue. "Automatically" generally means without human input, influence, or interaction intervening (e.g., only in response to original or pre-queued human activity).

[0412] The methods of the present disclosure can also include various forms of output, the various forms of output including accurate and sensitive interpretation of a nucleic acid sample of a subject. The output of the retrieval can be provided in the format of a computer file. In some embodiments, the output is a FASTA file, a FASTQ file, or a VCF file. The output can be processed to generate a text file or an XML file containing sequence data such as a nucleic acid sequence aligned with a reference genome. In other embodiments, the processing generates an output containing coordinates or strings describing one or more mutations in the subject's nucleic acid relative to the reference genome. Alignment strings can include Simple UnGapped Alignment Report (SUGAR), Verbose UsefulLabeled Gapped Alignment Report (VALGAR), and Compact Idiosyncratic GappedAlignment Report (CIGAR) (e.g., as described in Ning et al., Genome Research 11(10):1725-9, 2001, which is hereby incorporated by reference in its entirety). These strings can be implemented, for example, in the Exonerate sequence alignment software from the EuropeanBioinformatics Institute (Hinxton, UK).

[0413] In some embodiments, a sequence alignment containing a CIGAR string is generated—such as, for example, a Sequence Alignment / Map (SAM) or Binary Alignment / Map (BAM) file (the SAM format is described, for example, in Li et al., “The Sequence Alignment / Mapformat and SAMtools,” Bioinformatics, 25(16):2078-9, 2009, which is hereby incorporated by reference in its entirety). In some embodiments, the CIGAR shows or includes an alignment with one gap per line. CIGAR is a compressed pairwise alignment format reported as a CIGAR string. The CIGAR string can be used to present long (e.g., genomic) pairwise alignments. The CIGAR string can be used in the SAM format to represent the alignment of reads to a reference genome sequence.

[0414] CIGAR strings can follow established motifs. Each character is preceded by a number giving the base count for the event. The characters used can include M, I, D, N, and S (M = match; I = insertion; D = deletion; N = gap; S = substitution). A CIGAR string defines a sequence of matches and / or mismatches and deletions (or gaps). For example, the CIGAR string 2MD3M2D2M can indicate that the alignment contains 2 matches, 1 deletion (the number 1 is omitted to save some space), 3 matches, 2 deletions, and 2 matches.

[0415] In some embodiments, nucleic acid populations for sequencing are prepared by enzymatically forming blunt ends on double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Examples of enzymes or catalytic fragments thereof that can be optionally used include the Klenow fragment and T4 polymerase. At a 5' overhang, the enzyme typically extends the recessed 3' end on the opposite strand until it is flush with the 5' end to produce a blunt end. At a 3' overhang, the enzyme typically digests from the 3' end up to and sometimes beyond the 5' end of the opposite strand. If the digestion proceeds beyond the 5' end of the opposite strand, the nick can be filled by an enzyme having the same polymerase activity as that used for the 5' overhang. The formation of blunt ends on double-stranded nucleic acids facilitates, for example, the attachment of adapters and subsequent amplification.

[0416] In some embodiments, nucleic acid populations are subjected to additional processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA (e.g., complementary DNA or cDNA). These forms of nucleic acids are also optionally ligated to adapters and amplified.

[0417] With or without prior amplification, nucleic acids that have been subjected to the blunt-end formation process described above, and optionally other nucleic acids in the sample, can be sequenced to produce sequenced nucleic acids. Sequenced nucleic acids can refer to the sequence of a nucleic acid (e.g., sequence information) or a nucleic acid whose sequence has been determined. Sequencing can be performed in order to directly or indirectly provide sequence data for individual nucleic acid molecules in a sample from the consensus sequence of the amplification products of the individual nucleic acid molecules in the sample.

[0418] In some embodiments, double-stranded nucleic acids with single-stranded overhangs in a sample are ligated at both ends to adapters containing barcodes after blunt ends are formed, and sequencing determines the nucleic acid sequence and the in-line barcodes introduced by the adapters. The blunt-ended DNA molecules are optionally ligated to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Optionally, the blunt ends of the sample nucleic acids and the adapters can be tailed with complementary nucleotides to facilitate ligation (e.g., sticky-end ligation).

[0419] The nucleic acid sample is typically contacted with a sufficient number of adapters such that the probability that any two copies of the same nucleic acid receive the same adapter barcode combination from the adapters ligated to both ends is low (e.g., less than about 1% or 0.1%). Using adapters in this way can allow for the identification of families of nucleic acid sequences that have the same start and end points on a reference nucleic acid and are ligated to the same barcode combination. Such families can represent the amplified product sequences of the nucleic acids in the sample prior to amplification. The sequences of the family members can be assembled to obtain the consensus nucleotides or the complete consensus sequence of the nucleic acid molecules in the original sample, which are modified by blunt-end formation and adapter attachment. In other words, the nucleotides occupying a particular position in the nucleic acids in the sample can be determined as the consensus nucleotides of the nucleotides occupying the corresponding positions in the family member sequences. A family can include the sequences of one or both strands of a double-stranded nucleic acid. If the members of a family include the sequences of both strands from a double-stranded nucleic acid, then for the purpose of assembling the sequences to obtain the consensus nucleotides or sequence, the sequences of one strand can be converted to their complementary sequences. Some families contain only a single member sequence. In this case, the sequence can be considered the sequence of the nucleic acid in the sample prior to amplification. Optionally, families with only a single member sequence can be excluded from subsequent analysis.

[0420] By comparing the sequenced nucleic acid to a reference sequence, nucleotide variations (e.g., SNVs or indels) in the sequenced nucleic acid can be determined. The reference sequence is typically a known sequence, e.g., a known whole or partial genomic sequence from a subject (e.g., a whole genome sequence of a human subject). The reference sequence can be, for example, hG19 or hG38. As described above, the sequenced nucleic acid can represent the sequence of the nucleic acid in a directly determined sample or the consensus sequence of an amplification product of such nucleic acid. The comparison can be made at one or more specified positions on the reference sequence. When the corresponding sequences are maximally aligned, a subset of the sequenced nucleic acid can be identified that includes positions corresponding to the specified positions on the reference sequence. In such a subset, it can be determined which (if any) of the sequenced nucleic acids contain nucleotide variations at the specified positions and optionally which (if any) contain reference nucleotides (e.g., the same as in the reference sequence). If the number of sequenced nucleic acids in the subset that contain nucleotide variations exceeds a selected threshold, the variant nucleotide can be called at the specified position. The threshold can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids in the subset that contain nucleotide variations, or the threshold can be a ratio of the sequenced nucleic acids in the subset that contain nucleotide variations, such as at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20, and other possibilities. The comparison can be repeated for any specified position of interest in the reference sequence. Sometimes the comparison is made for specified positions that occupy at least about 20, 100, 200, or 300 consecutive positions on the reference sequence, e.g., about 20 - 500 or about 50 - 300 consecutive positions.

[0421] Additional details regarding nucleic acid sequencing, including the forms and applications described herein, are also provided in the following references: e.g., Levy et al., Annual Review of Genomics and Human Genetics, 17:95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55:641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7:287-296 (2009), Astier et al., J Am Chem Soc., 128(5):1705-10 (2006); U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503, each of which is hereby incorporated by reference in its entirety.

[0422] IV. Sets of Target-Specific Probes; Compositions

[0423] 1. Sets of Target-Specific Probes

[0424] In some embodiments, sets of target-specific probes are provided, the sets comprising target-binding probes specific for a group of sequence-variable target regions and target-binding probes specific for a group of epigenetic target regions. In some embodiments, the capture yield of the target-binding probes specific for the group of sequence-variable target regions is higher (e.g., at least 2-fold higher) than the capture yield of the target-binding probes specific for the group of epigenetic target regions. In some embodiments, the sets of target-specific probes are configured to have a higher capture yield for the group of sequence-variable target regions than they have for the group of epigenetic target regions (e.g., at least 2-fold higher).

[0425] In some embodiments, the capture yield of target binding probes specific for a set of sequence-variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the capture yield of target binding probes specific for an epigenetic target region set. In some embodiments, the capture yield of target binding probes specific for a set of sequence-variable target regions is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, or 14-fold to 15-fold higher than the capture yield of target binding probes specific for an epigenetic target region set. In some embodiments, the capture yield of target binding probes specific for a set of sequence-variable target regions is at least 10-fold higher than the capture yield of target binding probes specific for an epigenetic target region set, e.g., 10-fold to 20-fold higher than the capture yield of target binding probes specific for an epigenetic target region set.

[0426] In some embodiments, a set of target-specific probes is configured to have a capture yield for a set of sequence-variable target regions that is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than its capture yield for an epigenetic target region set. In some embodiments, a set of target-specific probes is configured to have a capture yield for a set of sequence-variable target regions that is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, or 14-fold to 15-fold higher than its capture yield for an epigenetic target region set. In some embodiments, a set of target-specific probes is configured to have a capture yield for a set of sequence-variable target regions that is at least 10-fold higher than its capture yield for an epigenetic target region set, e.g., 10-fold to 20-fold higher than its capture yield for an epigenetic target region set.

[0427] The set of probes can be configured to provide higher capture yields for the set of sequence-variable target regions in various ways, including concentration, different lengths, and / or chemical properties (e.g., affecting affinity), and combinations thereof. The affinity can be adjusted by adjusting the probe length and / or including nucleotide modifications as discussed below.

[0428] In some embodiments, the target-specific probes specific for the set of sequence-variable target regions are present at a higher concentration than the target-specific probes specific for the epigenetic target region set. In some embodiments, the concentration of the target-binding probes specific for the set of sequence-variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the concentration of the target-binding probes specific for the epigenetic target region set. In some embodiments, the concentration of the target-binding probes specific for the set of sequence-variable target regions is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, or 14-fold to 15-fold higher than the concentration of the target-binding probes specific for the epigenetic target region set. In some embodiments, the concentration of the target-binding probes specific for the set of sequence-variable target regions is at least 2-fold higher than the concentration of the target-binding probes specific for the epigenetic target region set. In some embodiments, the concentration of the target-binding probes specific for the set of sequence-variable target regions is at least 10-fold higher, e.g., 10-fold to 20-fold higher than the concentration of the target-binding probes specific for the epigenetic target region set. In such embodiments, the concentration can refer to the average mass / volume concentration of the individual probes in each set.

[0429] In some embodiments, target-specific probes specific for a sequence-variable targetome have a higher affinity for their targets than target-specific probes specific for an epigenetic targetome. The affinity can be modulated in any manner known to those skilled in the art, including by using different probe chemistries. For example, certain nucleotide modifications, such as cytosine 5-methylation (in certain sequence contexts), modifications that provide a heteroatom at the 2'-sugar position, and LNA nucleotides, can increase the stability of double-stranded nucleic acids, indicating that oligonucleotides with such modifications have a relatively high affinity for their complementary sequences. See, e.g., Severin et al., Nucleic Acids Res. 39:8740–8751 (2011); Freier et al., Nucleic Acids Res. 25:4429–4443 (1997); U.S. Patent No. 9,738,894. Additionally, longer sequence lengths will generally provide increased affinity. Other nucleotide modifications, such as substituting guanine with the nucleobase hypoxanthine, decrease affinity by reducing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, target-specific probes specific for a sequence-variable targetome have modifications that increase their affinity for their targets. In some embodiments, optionally or additionally, target-specific probes specific for an epigenetic targetome have modifications that decrease their affinity for their targets. In some embodiments, target-specific probes specific for a sequence-variable targetome have a longer average length and / or a higher average melting temperature than target-specific probes specific for an epigenetic targetome. As discussed above, these embodiments can be combined with each other and / or in concentration differences to achieve a desired fold difference in capture yield, such as any of the fold differences or ranges described above.

[0430] In some embodiments, the target-specific probe comprises a capture moiety. The capture moiety can be any capture moiety described herein, e.g., biotin. In some embodiments, the target-specific probe is linked to a solid support, e.g., covalently or non-covalently, such as by interaction of a binding pair of the capture moiety. In some embodiments, the solid support is a bead, such as a magnetic bead.

[0431] In some embodiments, the target-specific probe specific for a sequence-variable targetome and / or the target-specific probe specific for an epigenetic targetome is a bait set as discussed above, e.g., a probe comprising a capture moiety and a sequence selected to tile across a panel of regions (such as genes).

[0432] In some embodiments, the target-specific probe is provided in a single composition. The single composition can be a solution (liquid or frozen). Optionally, it can be a lyophilized product.

[0433] Optionally, the target-specific probes can be provided as more than one composition. For example, a first composition including probes specific for an epigenetic target set and a second composition including probes specific for a sequence-variable target set. These probes can be mixed in appropriate proportions to provide a combined probe composition having any of the foregoing fold differences in concentration and / or capture yield. Optionally, they can be used in separate capture procedures (e.g., for aliquots of a sample or sequentially for the same sample) to provide a first composition and a second composition respectively containing captured epigenetic target regions and sequence-variable target regions.

[0434] a. Probes specific for epigenetic target regions

[0435] Probes for an epigenetic target set can include probes specific for one or more types of target regions that may distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells (e.g., non-neoplastic circulating cells). For example, exemplary types of such regions are discussed in detail herein in the section above regarding capture sets. Probes for an epigenetic target set can also include probes for one or more control regions, for example, as described herein.

[0436] In some embodiments, the probes for an epigenetic target probe set have a footprint of at least 100 kb, for example, at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the probes for an epigenetic target set have a footprint in the range of: 100 - 1000 kb, for example, 100 - 200 kb, 200 - 300 kb, 300 - 400 kb, 400 - 500 kb, 500 - 600 kb, 600 - 700 kb, 700 - 800 kb, 800 - 900 kb, and 900 - 1,000 kb.

[0437] i. Hypermethylation variable target regions

[0438] In some embodiments, the probes for an epigenetic target panel include probes specific for one or more hypermethylated variable targets. The hypermethylated variable targets can be any of those described above. For example, in some embodiments, the probes specific for hypermethylated variable targets include probes specific for more than one locus listed in Table 1, e.g., probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, the probes specific for hypermethylated variable targets include probes specific for more than one locus listed in Table 2, e.g., probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. In some embodiments, the probes specific for hypermethylated variable targets include probes specific for more than one locus listed in Table 1 or Table 2, e.g., probes specific for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target, there may be one or more probes that have a hybridization site that binds between the transcriptional start site of the gene and the stop codon (the last stop codon for alternatively spliced genes). In some embodiments, the one or more probes bind within 300 bp of the listed position, e.g., within 200 bp or 100 bp. In some embodiments, the probes have a hybridization site that overlaps with the position listed above. In some embodiments, the probes specific for hypermethylated targets include probes specific for one, two, three, four, or five subgroups of hypermethylated targets that together show hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.

[0439] ii. Hypomethylated variable targets

[0440] In some embodiments, the probes for an epigenetic target panel include probes specific for one or more hypomethylated variable targets. The hypomethylated variable targets can be any of those described above. For example, the probes specific for one or more hypomethylated variable targets can include probes for regions such as repetitive elements (e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA), and intergenic regions that are normally methylated in healthy cells can show decreased methylation in tumor cells.

[0441] In some embodiments, probes specific for hypomethylated variable target regions include probes specific for repetitive elements and / or intergenic regions. In some embodiments, probes specific for repetitive elements include probes specific for one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.

[0442] Exemplary probes specific for genomic regions showing cancer-related hypomethylation include probes specific for nucleotides 8403565 - 8953708 and / or 151104701 - 151106035 of human chromosome 1. In some embodiments, probes specific for hypomethylated variable target regions include probes specific for regions overlapping with or containing nucleotides 8403565 - 8953708 and / or 151104701 - 151106035 of human chromosome 1.

[0443] iii. CTCF binding regions

[0444] In some embodiments, probes for an epigenomic target set include probes specific for CTCF binding regions. In some embodiments, probes specific for CTCF binding regions include probes specific for at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 CTCF binding regions, e.g., CTCF binding regions described in one or more of Cuddapah et al., Martin et al., or Rhee et al. as above or in CTCFBSDB or the articles cited above. In some embodiments, probes for an epigenomic target set include at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp in the upstream and downstream regions of CTCF binding sites.

[0445] iv. Transcription start sites

[0446] In some embodiments, probes for an epigenomic target panel include probes specific for transcription start sites. In some embodiments, probes specific for transcription start sites include probes specific for at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 transcription start sites, such as, for example, probes specific for transcription start sites listed in DBTSS. In some embodiments, probes for an epigenomic target panel include probes for sequences of at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of a transcription start site.

[0447] v. Focused amplification

[0448] As mentioned above, although focused amplifications are somatic mutations, they can be detected by sequencing based on read frequencies in a manner similar to methods for detecting certain epigenetic alterations such as methylation alterations. Thus, as discussed above, regions showing focused amplifications in cancer can be included in an epigenomic target panel. In some embodiments, probes specific for an epigenomic target panel include probes specific for focused amplifications. In some embodiments, probes specific for focused amplifications include probes specific for one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, probes specific for focused amplifications include probes specific for at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the foregoing targets.

[0449] vi. Control region

[0450] Including control regions to facilitate data validation can be useful. In some embodiments, probes specific for an epigenomic target panel include probes specific for control methylated regions expected to be methylated in substantially all samples. In some embodiments, probes specific for an epigenomic target panel include probes specific for control hypomethylated regions expected to be hypomethylated in substantially all samples.

[0451] b. Probes specific for sequence-variable target regions

[0452] Probes for a sequence variable target group can include probes specific for more than one region known to undergo somatic mutations in cancer. The probes can be specific for any sequence variable target group described herein. For example, in the section above regarding capture groups, exemplary sequence variable target groups are discussed in detail herein.

[0453] In some embodiments, the sequence variable target probe set has a footprint of at least 10 kb, e.g., at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target probe set has a footprint in the range of: 10 - 100 kb, e.g., 10 - 20 kb, 20 - 30 kb, 30 - 40 kb, 40 - 50 kb, 50 - 60 kb, 60 - 70 kb, 70 - 80 kb, 80 - 90 kb, and 90 - 100 kb.

[0454] In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 genes in Table 3. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 SNVs in Table 3. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least 1, at least 2, at least 3, at least 4, at least 5, or 6 fusions in Table 3. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 1, at least 2, or 3 indels in Table 3. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 genes in Table 4. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 SNVs in Table 4. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least 1, at least 2, at least 3, at least 4, at least 5, or 6 fusions in Table 4. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 indels in Table 4. In some embodiments, probes specific for a set of sequence-variable target regions include probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 genes in Table 5.

[0455] In some embodiments, probes specific for a set of sequence variable target regions include probes specific for target regions from at least 10, 20, 30, or 35 cancer-related genes such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.

[0456] c. Probe compositions

[0457] In some embodiments, a single composition is provided that contains probes for a set of sequence variable target regions and probes for an epigenetic target region. The probes can be provided in such a composition at any concentration ratio described herein.

[0458] In some embodiments, a first composition containing probes for an epigenetic target region and a second composition containing probes for a set of sequence variable target regions are provided. The ratio of the concentration of the probes in the first composition to the concentration of the probes in the second composition can be any ratio described herein.

[0459] 2. Compositions containing captured cfDNA

[0460] In some embodiments, a composition containing captured cfDNA is provided. The captured cfDNA can have any of the characteristics of the capture set described herein, including, for example, a DNA concentration corresponding to a set of sequence variable target regions (such as normalized for footprint size as discussed above) greater than the DNA concentration corresponding to an epigenetic target region. In some embodiments, the cfDNA of the capture set includes sequence tags that can be added to the cfDNA as described herein. Typically, the inclusion of sequence tags results in cfDNA molecules being different from their naturally occurring untagged form.

[0461] Such a composition can also contain the probe sets or sequencing primers described herein, each of which can be different from naturally occurring nucleic acid molecules. For example, the probe sets described herein can contain capture moieties, and the sequencing primers can contain non-naturally occurring markers.

[0462] V. Computer systems

[0463] The methods of the present disclosure can be implemented using or by means of a computer system. For example, such methods can include: collecting cfDNA from a test subject; capturing more than one set of target regions from the cfDNA, where the more than one set of target regions includes a sequence variable target region set and an epigenetic target region set, thereby generating a set of captured cfDNA molecules; sequencing the captured cfDNA molecules, where the captured cfDNA molecules of the sequence variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining more than one sequence read generated by a nucleic acid sequencer by sequencing the captured cfDNA molecules; mapping the more than one sequence read to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.

[0464] Figure 2 FIG. 201 shows a computer system 201 that is programmed or otherwise configured to implement the methods of the present disclosure. The computer system 201 can control aspects of sample preparation, sequencing, and / or analysis. In some instances, the computer system 201 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing.

[0465] The computer system 201 includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”) 205, which can be a single-core or multi-core processor or more than one processor for parallel processing. The computer system 201 also includes a memory or memory location 210 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 215 (e.g., hard disk), a communication interface 220 for communicating with one or more other systems (e.g., network adapter), and peripheral devices 225, such as a cache, other memories, data storage, and / or an electronic display adapter. The memory 210, storage unit 215, interface 220, and peripheral devices 225 communicate with the CPU 205 via a communication network or bus (solid lines), such as a motherboard. The storage unit 215 can be a data storage unit (or data repository) for storing data. The computer system 201 can be operatively coupled to a computer network 230 via the communication interface 220. The computer network 230 can be the Internet, an intranet, and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the computer network 230 is a telecommunications and / or data network. The computer network 230 can include one or more computer servers, which can initiate distributed computing, such as cloud computing. In some cases, via the computer system 201, the computer network 230 can implement a peer-to-peer network, which can initiate devices coupled to the computer system 201 to operate as clients or servers.

[0466] The CPU 205 can execute a series of machine-readable instructions, which can be embodied as a program or software. The instructions can be stored in a memory location, such as the memory 210. Examples of operations performed by the CPU 205 can include reading, decoding, executing, and writing back.

[0467] The storage unit 215 can store files, such as drivers, libraries, and saved programs. The storage unit 215 can store user-generated programs and recorded sessions, as well as outputs related to programs. The storage unit 215 can store user data, e.g., user preferences and user programs. In some cases, the computer system 201 can include one or more additional data storage units, which are external to the computer system 201, such as on a remote server that communicates with the computer system 201 via an intranet or the Internet. Data can be transferred from one location to another using, for example, a communication network or a physical data transporter (e.g., using a hard disk drive, a thumb drive, or other data storage mechanisms).

[0468] The computer system 201 can communicate with one or more remote computer systems via a network 230. For an implementation, the computer system 201 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., iPad, Galaxy Tab), telephones, smart phones (e.g., iPhone, Android-supported devices, ) or personal digital assistants. The user can access the computer system 201 via the network 230.

[0469] The methods described herein can be implemented in the form of machine (e.g., computer processor) executable code that is stored at an electronic storage location of the computer system 201, such as, for example, the memory 210 or the electronic storage unit 215. The machine executable code or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 205. In some cases, the code can be retrieved from the storage unit 215 and stored in the memory 210 for ready access by the processor 205. In some cases, the electronic storage unit 215 may not be included and the machine executable instructions may be stored in the memory 210.

[0470] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing more than one set of target regions from the cfDNA, wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby generating a set of captured cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining more than one sequence read generated by a nucleic acid sequencer by sequencing the captured cfDNA molecules; mapping the more than one sequence read to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.

[0471] The code can be pre-compiled and configured to be used with a machine having a processor suitable for executing the code or can be compiled during run-time. The code can be provided in the form of a programming language that can be selected such that the code can be executed in a pre-compiled or as-compiled manner.

[0472] Aspects of the systems and methods provided herein, such as computer system 201, can be embodied in programming. Aspects of the technology can be regarded as a “product” or “articles of manufacture” in the form of machine (or processor) executable code and / or associated data that are typically carried on or embodied in a type of machine-readable medium. The machine executable code can be stored in an electronic storage unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A “storage” type medium can include any or all tangible memories of a computer, a processor, etc. or associated modules such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time.

[0473] All or part of the software can sometimes be communicated via the Internet or various other communication networks. For example, such communication can cause the software to be loaded from one computer or processor to another, e.g., from an administrative server or host to the computer platform of an application server. Thus, another type of medium that can carry software elements includes optical, electrical, and electromagnetic waves such as those used across physical interfaces between local devices, via wired and fiber optic landline networks, and over various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be regarded as media that carry software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0474] Thus, a machine-readable medium, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks, such as any storage device in any computer, as shown in the accompanying figures, such as may be used to implement a database, etc. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wire and optical fiber, including the wires that make up a bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, DVD or DVD-ROM, any other optical medium, punched cards, paper tape, any other physical storage medium with hole patterns, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in transporting one or more strings of one or more instructions to a processor for execution.

[0475] Computer system 201 may include or communicate with an electronic display that includes a user interface (UI) to provide, for example, one or more results of a sample analysis. Examples of UIs include but are not limited to graphical user interfaces (GUIs) and web-based user interfaces.

[0476] Additional details regarding computer systems and networks, databases, and computer program products are provided in the following literature: e.g., Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Edition (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Edition (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Edition (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Edition (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Edition (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety.

[0477] VI. Applications

[0478] 1. Cancer and Other Diseases

[0479] The method can be used to diagnose the presence of a condition, particularly cancer, in a subject, to characterize the condition (e.g., stage cancer or determine cancer heterogeneity), to monitor the response to treatment of the condition, and to obtain a prognostic risk of the developing or subsequent course of the condition. The present disclosure can also be used to determine the efficacy of a particular treatment option. If the treatment is successful, the successful treatment option can increase the amount of copy number variations or rare mutations detected in the blood of the subject because more cancer may die and shed DNA. In other instances, this may not occur. In another instance, perhaps certain treatment options may be associated with the genetic signature profile of the cancer over time. This correlation can be used to select a therapy.

[0480] Additionally, if cancer is observed to remit after treatment, the method can be used to monitor the remaining disease or recurrence of the disease.

[0481] In some embodiments, the methods and systems disclosed herein can be used to identify customized or targeted therapies to treat a patient's specific disease or condition based on classifying nucleic acid variants as being of somatic or germline origin. Generally, the disease considered is a type of cancer. Non-limiting examples of such cancers include biliary cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colon adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, solid pseudopapillary tumor, acinar cell carcinoma. Prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer or uterine sarcoma. The type and / or stage of cancer can be detected from genetic variants, including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal alterations in nucleic acid chemical modifications, abnormal alterations in epigenetic patterns, and abnormal alterations in nucleic acid 5-methylcytosine.

[0482] Genetic data can also be used to characterize specific forms of cancer. Cancer is often heterogeneous both in composition and staging. Genetic signature data can permit the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that specific subtype. This information can also provide clues to a subject or practitioner regarding the prognosis of the specific type of cancer and permit the subject or practitioner to adjust treatment options based on the progression of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers can remain benign, inactive, or dormant. The systems and methods of the present disclosure can be used to determine disease progression.

[0483] In addition, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, for example, generating a genetic signature from extracellular polynucleotides of the subject, wherein the genetic signature comprises more than one data obtained by copy number variation and rare mutation analysis. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition can be a condition that results in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of cancer. In other examples, the heterogeneity can include more than one disease focus. Again, in the example of cancer, there can be more than one tumor focus, perhaps where one or more of the foci are the result of metastases that have spread from a primary site.

[0484] The methods can be used to generate or characterize a fingerprint or data set that is the sum of genetic information obtained from different cells in a heterogeneous disease. The data set can include copy number variation, epigenetic variation, and mutation analysis, either alone or in combination.

[0485] The methods can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods herein do not involve diagnosing, prognosing, or monitoring a fetus and thus do not involve non-invasive prenatal testing. In other embodiments, these methods can be used for a subject who is pregnant to diagnose, prognose, monitor, or observe cancer or other diseases of the unborn subject, whose DNA and other polynucleotides can co-circulate with maternal molecules.

[0486] Non-limiting examples of other genetic-based diseases, disorders or conditions optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth disease (CMT), cri du chat syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency disease (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, and the like.

[0487] In some embodiments, the methods described herein include using a set of sequence information obtained as described herein to detect the presence or absence of DNA originating from or derived from tumor cells at a preselected time point after a previous cancer treatment in a subject previously diagnosed with cancer. The method can also include determining a cancer recurrence score that indicates the presence or absence of DNA originating from or derived from tumor cells of the test subject.

[0488] In the case where a cancer recurrence score is determined, it can also be used to determine the cancer recurrence status. For example, when the cancer recurrence score is higher than a predetermined threshold, the cancer recurrence status may be at risk of cancer recurrence. For example, when the cancer recurrence score is higher than a predetermined threshold, the cancer recurrence status may be at low risk or lower risk of cancer recurrence. In certain embodiments, a cancer recurrence score equal to the predetermined threshold can result in a cancer recurrence status that is at risk of cancer recurrence or at low risk or lower risk of cancer recurrence.

[0489] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold, and when the cancer recurrence score is higher than the cancer recurrence threshold, the test subject is classified as a candidate for subsequent cancer treatment, or when the cancer recurrence score is lower than the cancer recurrence threshold, the test subject is not classified as a candidate for therapy. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold can result in being classified as a candidate for subsequent cancer treatment or not being classified as a candidate for therapy.

[0490] The methods discussed above may also include any one or more compatible features described elsewhere herein (including in the sections regarding methods for determining the risk of cancer recurrence in a test subject and / or classifying a test subject as a candidate for subsequent cancer treatment).

[0491] 2. Methods for Determining the Risk of Cancer Recurrence in a Test Subject and / or Classifying a Test Subject as a Candidate for Subsequent Cancer Treatment

[0492] In some embodiments, the methods provided herein are methods for determining the risk of cancer recurrence in a test subject. In some embodiments, the methods provided herein are methods for classifying a test subject as a candidate for subsequent cancer treatment.

[0493] Any such method may include collecting DNA (e.g., originating from or derived from tumor cells) from a test subject diagnosed with cancer at one or more preselected time points after one or more prior cancer treatments of the test subject. The subject can be any subject described herein. The DNA can be cfDNA. The DNA can be obtained from a tissue sample.

[0494] Any such method may include capturing more than one set of target regions from the DNA from the subject, wherein the more than one set of target regions includes a sequence-variable set of target regions and an epigenetic set of target regions, thereby producing a set of captured DNA molecules. The capturing step can be performed according to any of the embodiments described elsewhere herein.

[0495] In any such method, the prior cancer treatment may include surgery, administration of a therapeutic composition, and / or chemotherapy.

[0496] Any such method may include sequencing the captured DNA molecules, thereby producing a set of sequence information. The captured DNA molecules of the sequence-variable set of target regions may be sequenced to a deeper sequencing depth than the captured DNA molecules of the epigenetic set of target regions.

[0497] Any such method may include detecting the presence or absence of DNA originating from or derived from tumor cells at the preselected time point using the set of sequence information. The detection of the presence or absence of DNA originating from or derived from tumor cells can be performed according to any of the embodiments described elsewhere herein.

[0498] A method for determining the risk of cancer recurrence in a test subject can include determining a cancer recurrence score that indicates the presence, absence, or amount of DNA originating from or derived from tumor cells in the test subject. The cancer recurrence score can also be used to determine the cancer recurrence status. For example, when the cancer recurrence score is higher than a predetermined threshold, the cancer recurrence status may be at risk of cancer recurrence. For example, when the cancer recurrence score is higher than a predetermined threshold, the cancer recurrence status may be at low risk or lower risk of cancer recurrence. In certain embodiments, a cancer recurrence score equal to the predetermined threshold can result in a cancer recurrence status that is at risk of cancer recurrence or at low risk or lower risk of cancer recurrence.

[0499] A method for classifying a test subject as a candidate for subsequent cancer treatment can include comparing the cancer recurrence score of the test subject to a predetermined cancer recurrence threshold, such that when the cancer recurrence score is higher than the cancer recurrence threshold, the test subject is classified as a candidate for subsequent cancer treatment, or when the cancer recurrence score is lower than the cancer recurrence threshold, the test subject is not classified as a candidate for therapy. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold can result in being classified as a candidate for subsequent cancer treatment or not being classified as a candidate for therapy. In some embodiments, subsequent cancer treatment includes chemotherapy or administration of a therapeutic composition.

[0500] Any such method can include determining the disease-free survival (DFS) period of the test subject based on the cancer recurrence score; for example, the DFS period can be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.

[0501] In some embodiments, the sequence information set includes sequence variable target region sequences, and determining the cancer recurrence score can include determining at least a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence variable target region sequences.

[0502] In some embodiments, the number of mutations in the sequence variable target region selected from 1, 2, 3, 4, or 5 is sufficient to result in a cancer recurrence score that classifies the first subscore as positive for cancer recurrence. In some embodiments, the number of mutations is selected from 1, 2, or 3.

[0503] In some embodiments, the sequence information set includes epigenetic target region sequences, and determining the cancer recurrence score includes determining a second subscore indicative of the amount of abnormal sequence reads in the epigenetic target region sequences. The abnormal sequence reads can be reads indicative of an epigenetic state different from the DNA present in a corresponding sample from a healthy subject (e.g., cfDNA present in a blood sample from a healthy subject, or DNA present in a tissue sample from a healthy subject, where the tissue sample is of the same tissue type as that obtained from the test subject). The abnormal reads can be consistent with cancer-related epigenetic alterations, e.g., methylation of hypermethylated variable target regions and / or fragmentation of fragmented variable target regions that is perturbed, where "perturbed" means different from the DNA present in a corresponding sample from a healthy subject).

[0504] In some embodiments, a proportion of reads corresponding to the hypermethylated variable target region group and / or the fragmented variable target region group that indicate hypermethylation in the hypermethylated variable target region group and / or abnormal fragmentation in the fragmented variable target region group that is greater than or equal to a value in the range of 0.001%-10% is sufficient to classify the second subscore as positive for cancer recurrence. The range can be 0.001%-1%, 0.005%-1%, 0.01%-5%, 0.01%-2%, or 0.01%-1%.

[0505] In some embodiments, any such method can include determining a fraction of tumor DNA from the fraction of reads in the sequence information set that indicate one or more features originating from tumor cells. This can be done for reads corresponding to some or all of the epigenetic target regions, e.g., including one or both of the hypermethylated variable target region and the fragmented variable target region (hypermethylation of the hypermethylated variable target region and / or abnormal fragmentation of the fragmented variable target region can be considered an indication of originating from tumor cells). This can be done for reads corresponding to sequence variable target regions, e.g., reads containing alterations consistent with cancer such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA can be determined based on a combination of reads corresponding to epigenetic target regions and reads corresponding to sequence variable target regions.

[0506] Determination of the cancer recurrence score can be at least partially based on the fraction of tumor DNA, where a fraction of tumor DNA greater than 10 -11 to 1 or 10 -10 to 1 within a threshold range is sufficient to classify the cancer recurrence score as positive for cancer recurrence. In some embodiments, a fraction of tumor DNA greater than or equal to a threshold within the following ranges is sufficient to classify the cancer recurrence score as positive for cancer recurrence: 10 –10 to 10 –9 、10 –9 to 10 –8 、10–8 from 1 to 10 –7 and 10 –7 from 1 to 10 –6 and 10 –6 from 1 to 10 –5 and 10 –5 from 1 to 10 –4 and 10 –4 from 1 to 10 –3 and 10 –3 from 1 to 10 –2 or 10 –2 from 1 to 10 –1 。In some embodiments, a fraction of tumor DNA greater than a threshold of at least 10 -7 is sufficient to classify a cancer recurrence score as cancer recurrence positive. The fraction of tumor DNA greater than the threshold can be determined based on cumulative probability, such as a threshold corresponding to any of the foregoing embodiments. For example, if the cumulative probability that the tumor fraction is greater than the threshold within any of the foregoing ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999, the sample is considered positive. In some embodiments, the probability threshold is at least 0.95, such as 0.99.

[0507] In some embodiments, the sequence information set includes a sequence-variable target region sequence and an epigenetic target region sequence, and determining the cancer recurrence score includes determining a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the sequence-variable target region sequence and determining a second subscore indicative of the amount of abnormal sequence reads in the epigenetic target region sequence, and combining the first subscore and the second subscore to provide the cancer recurrence score. In the case of combining the first subscore and the second subscore, they can be combined by: independently applying a threshold to each subscore (e.g., greater than a predetermined number of mutations (e.g., >1) in the sequence-variable target region and greater than a predetermined fraction of abnormal (e.g., tumor) reads in the epigenetic target region), or training a machine learning classifier to determine the status based on more than one positive and negative training sample.

[0508] In some embodiments, a value of the combined score in the range of -4 to 2 or -3 to 1 is sufficient to classify the cancer recurrence score as cancer recurrence positive.

[0509] In any embodiment in which the cancer recurrence score is classified as cancer recurrence positive, the subject's cancer recurrence status may be at risk of cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment.

[0510] In some embodiments, the cancer is any type of cancer described elsewhere herein, e.g., colorectal cancer.

[0511] 3. Treatment and related management

[0512] In certain embodiments, the methods disclosed herein involve identifying a customized therapy and administering the customized therapy to a patient in view of the status of a nucleic acid variant as being of somatic origin or germline origin. In some embodiments, substantially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy therapy, and / or the like) can be included as part of these methods. Generally, the customized therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to methods of enhancing an immune response against a particular cancer type. In certain embodiments, immunotherapy refers to methods of enhancing a T cell response against a tumor or cancer.

[0513] In certain embodiments, the status of a nucleic acid variant as being of somatic origin or germline origin from a sample of a subject can be compared to a database of comparator results from a reference population to identify a customized or targeted therapy for the subject. Generally, the reference population includes patients having the same cancer or disease type as the tested subject and / or patients who are receiving or have received the same therapy as the tested subject. When the nucleic acid variant and the comparison results meet certain classification criteria (e.g., a basic or approximate match), a customized or targeted therapy (or more than one therapy) can be identified.

[0514] In certain embodiments, the customized therapies described herein are generally administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are generally administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) can also be administered by methods such as buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intratympanic, and the administration can include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.

[0515] While the preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited to the specific examples provided in this specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Many variations, changes, and substitutions will now occur to those skilled in the art without departing from the present invention. In addition, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein in accordance with various conditions and variables. It should be understood that various alternatives of the embodiments of the present disclosure described herein may be employed in practicing the present invention. Accordingly, it is contemplated that the present disclosure should also cover any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the present invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

[0516] While, for purposes of clarity and understanding, some of the foregoing disclosure has been described in detail by way of illustration and example, it will be apparent to those of ordinary skill in the art upon reading the present disclosure that various changes in form and detail may be made without departing from the true scope of the present disclosure and may be practiced within the scope of the appended claims. For example, all method, system, computer-readable medium, and / or component features, steps, elements, or other aspects may be used in various combinations.

[0517] All patents, patent applications, websites, other publications or documents, accession numbers, etc., cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference in this manner. If different versions of a sequence are associated with an accession number at different times, the version associated with the accession number on the actual filing date of the present application is meant. If applicable, the effective filing date means the earlier of the actual filing date or the filing date of the priority application that refers to the accession number. Similarly, if different versions of a publication, website, etc., are published at different times, the version published most recently on the actual filing date of the present application is meant, unless otherwise indicated.

[0518] VII. EXAMPLES

[0519] i) Characterization of a panel of target probes with probes for sequence variable target panels and probes for epigenetic target panels at different concentrations

[0520] This example describes the performance evaluation of a panel of probes comprising probes for sequence variable target panels and probes for epigenetic target panels as part of an effort to combine epigenetic and genotype analysis of liquid biopsy cfDNA.

[0521] Prior to contacting the target region probe set, the cfDNA sample is processed by partitioning based on methylation status, end repair, ligation to adapters, and amplification by PCR (e.g., using primers targeting the adapters).

[0522] The processed sample is contacted with a target region probe set that includes probes for a set of sequence variable target regions and probes for an epigenetic target region set. The target region probes are in the form of biotinylated oligonucleotides designed to tile the region of interest. The probes for the set of sequence variable target regions have a footprint of approximately 50 kb, and the probes for the epigenetic target region set have a target footprint of approximately 500 kb. The probes for the set of sequence variable target regions include oligonucleotides targeting a series of regions identified in Tables 3 - 5, and the probes for the epigenetic target region set include oligonucleotides targeting a series of hypermethylated variable regions, hypomethylated variable regions, CTCF binding target regions, transcription start site target regions, focused amplification target regions, and methylation control regions.

[0523] The captured cfDNA isolated in this way is then prepared for sequencing and sequenced using an Illumina HiSeq or NovaSeq sequencer. The diversity (number of unique families of sequence reads) and read family size (number of individual reads in each family) of the sequence reads corresponding to the probes for the set of sequence variable target regions and the probes for the epigenetic target region set are analyzed. The values reported below were obtained using 70 ng of input DNA. 70 ng input is considered a relatively high amount and represents challenging conditions for maintaining the desired diversity level and family size.

[0524] Probe ratios of 2:1 and 5:1 (epigenetic probe set: mass / volume concentration ratio of sequence variable probe set) gave a reduction in the diversity of the sequence variable target regions, indicating that the amount of the epigenetic target region causes interference in generating the expected number of different read families from the sequence variable target regions.

[0525] Probe ratios of 1:2 or 1:5 (epigenetic probe set: sequence variable probe set) gave a higher level of diversity of the sequence variable target regions, which was generally close to the expected number of different read families, indicating that at these ratios, the epigenetic target region is not present in an amount that significantly interferes with generating the expected number of different read families from the sequence variable target regions.

[0526] For the epigenetic target regions, all ratios gave diversity levels significantly lower than the expected number of different read families. However, this is not considered problematic because the analysis of methylation, copy number, etc. of the epigenetic target regions does not require the same degree of dense and deep sequencing coverage as determining the presence or absence of nucleotide substitutions or indels as intended for the sequence variable regions.

[0527] ii) Detection of cancer using a combined set of epigenetic target regions and sequence variable target regions

[0528] As described above, cfDNA sample cohorts from cancer patients with different cancer stages from I to IVA (7 stages in total) were analyzed using probes in a 1:5 (epigenetic probe set: sequence variable probe set) ratio. The sequence variable target sequences were analyzed by detecting genomic alterations such as SNVs, insertions, deletions, and fusions, which could be adjudged with sufficient support to distinguish true tumor variants from technical errors. The epigenetic target sequences were analyzed independently to detect methylated fragments in regions that have been shown to be differentially methylated in cancer compared to blood cells. Finally, the results of the two analyses were combined to produce a final tumor presence / absence determination to see if they showed characteristics consistent with cancer with 95% specificity.

[0529] Figure 3 The sensitivity of cancer detection based on either the sequence variable target sequences alone or in combination with the epigenetic target sequences is shown. For the IIIA and IIIC stage cohorts, the detection of cancer was 100% sensitive to either individual method. For all other cohorts except one, including the analysis of the epigenetic target sequences, the sensitivity was increased by approximately 10% - 30%. One exception was the IIB stage cohort, where each sample was either a true positive or a false negative according to both methods.

[0530] Thus, the disclosed methods and compositions can provide captured cfDNA that can be used to simultaneously sequence epigenetic target regions and sequence variable target regions to different sequencing depths for sensitive, combined sequence - and epigenetic - based detection of cancer.

[0531] iii) Identifying the risk level of colorectal cancer recurrence

[0532] Assays were developed and performed to identify whether patients undergoing treatment for colorectal cancer (CRC) were at high risk of recurrence. Plasma samples (3 mL to 4 mL) were taken from 72 patients undergoing standard of care treatment for CRC (surgery + / - neoadjuvant therapy in 42 cases, adjuvant therapy + / - neoadjuvant therapy in 30 cases).

[0533] cfDNA (median amount 27 ng) was extracted from the samples and analyzed using a method substantially as described herein, which was validated in early CRC and integrated the assessment of genomic alterations and epigenomic features indicative of cancer, including hypermethylated variable target regions. The method differentiates tumor-derived alterations from non-tumor-derived alterations (such as germline or clonal hematopoiesis of indeterminate potential (CHIP) alterations) in a tumor tissue-uninformative assay (LUNAR assay, Guardant Health, CA). The assay uses a single input sample and integrates the detection of genomic alterations with the quantification of cancer-associated epigenomic signals, and was validated using 80 plasma samples from presumably cancer-free donors aged 50 - 75 years, and resulted in a single false positive (99% specificity). Analytical sensitivity (limit of detection) was established using dilution series from 4 different patients with advanced CRC, tested in triplicate across multiple batches with a clinically relevant DNA input (30 ng). 100% sensitivity was maintained even at the lowest test level (estimated at 0.1% tumor level).

[0534] Plasma samples were collected a median of 31 days (N = 42) after surgical resection or a median of 37 days (N = 27) after completion of adjuvant therapy, after completion of SOC therapy. The median follow-up time was 515 days (33 - 938 days). A sample was considered positive for ctDNA if genomic alterations or epigenomic alterations indicative of cancer were detected. Genomic alterations were detected using Guardant Health's digital sequencing platform to differentiate true mutations from sequencing errors. Variant filters were applied to distinguish tumor mutations from non-tumor mutations (such as CHIP). Epigenomic determination was based on measuring whether the methylation rate observed in tumor hypermethylated regions was higher than expected based on methylation levels in blood. In particular, in this embodiment, if the number of genomic alterations indicative of cancer detected exceeded a threshold, where the threshold was 1, 2, or 3 alterations, the genomic result was considered positive. Epigenomic results included methylation analysis to determine the proportion of reads indicative of hypermethylation in a set of hypermethylated variable target regions. The overall "tumor fraction" was also calculated based on the overall proportion of reads with methylation-based tumor-like features, and a sample was considered positive if the cumulative probability of the tumor fraction being greater than or equal to -7 10 exceeded a probability threshold of 0.99. A total of 14 samples were positive, of which 10 samples were positive for both the epigenomic and genomic prongs, only 3 samples were positive for the epigenomic prong, and only 1 sample was positive for the genomic prong.

[0535] ctDNA was positive in 7 / 11 patients who relapsed 1 year after surgery and underwent CRC resection. ctDNA was negative in 30 / 31 patients who did not relapse 1 year after surgery and underwent CRC resection. ctDNA was negative in 20 / 22 patients who did not relapse 1 year after adjuvant therapy and completed SOC adjuvant therapy. ctDNA was positive in 4 / 5 patients who relapsed 1 year after adjuvant therapy and completed SOC adjuvant therapy. Overall, ctDNA testing after completion of standard-of-care therapy had a 100% positive predictive value (PPV) for relapse, a 76% negative predictive value (NPV), and a relapse risk ratio of 9.22 (p < 0.0001)( Figure 4 ).

[0536] The assay performance statistics for only the genomic arm and for the combined analysis using genomic and epigenomic arms are summarized in the table below.

[0537] Table 6

[0538]

[0539] Group results of genomic sequencing and epigenomic analysis. Among 14 patients with positive ctDNA after completion of SOC therapy, 10 were positive by both genomic and epigenomic assessments.

[0540] In the surgery group, ctDNA testing had a 100% PPV for relapse, a 76% NPV, and a relapse risk ratio of 8.7 (p < 0.0001). In the adjuvant therapy group, ctDNA testing had a 100% PPV for relapse, a 76% NPV, and a relapse risk ratio of 9.3 (p < 0.0001).

[0541] Patients with negative ctDNA after therapy completion were further classified according to whether their ctDNA was positive or negative before therapy. Patients who were positive before therapy and negative after therapy were called "cleared" patients, while patients who were negative before and after therapy were called "negative" patients. The cleared group included 6 individuals, 3 of whom relapsed and 3 of whom did not relapse. The negative group included 26 individuals, 7 of whom relapsed and 19 of whom did not relapse.

[0542] Thus, in resected CRC, ctDNA testing of plasma alone, tumor-uninformed integrated genomic and epigenomic assays have high recurrence PPV and NPV after completion of standard-of-care therapy. In the post-resection setting, ctDNA testing identifies patients who may benefit from adjuvant therapy. After completion of adjuvant therapy, ctDNA testing identifies patients who may benefit from additional or modified therapy. These findings suggest that ctDNA from a single post-resection or post-adjuvant blood draw can identify high-risk patients and inform treatment decisions. In contrast, current ctDNA minimal residual disease assays only assess genomic alterations, are limited by low levels of ctDNA, and rely on tumor tissue sequencing to distinguish tumor-derived alterations from confounding non-tumor-derived alterations (e.g., clonal hematopoiesis of indeterminate potential; CHIP).

Claims

1. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform a method for identifying the presence of DNA produced by a tumor, the method comprising: Collecting cfDNA from a test subject; Using a methyl-binding domain to partition the cfDNA obtained from the test subject into two or more fractions based on methylation levels and performing subsequent steps of the method on each fraction; Capturing more than one set of target regions from the cfDNA using target-specific probes; Wherein (i) the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby producing a set of captured cfDNA molecules, (ii) among the set of captured cfDNA molecules, the cfDNA molecules corresponding to the sequence-variable target region set are captured with a capture yield higher than that of the cfDNA molecules corresponding to the epigenetic target region set; and (iii) the epigenetic target region set is selected from one or more of the following: (a) a hypermethylated variable target region set; (b) a hypomethylated variable target region set; (c) a methylation control target region set; and (d) a fragmented variable target region set, wherein the fragmented variable target region set includes a transcription start site region and / or a CTCF binding region; Sequencing the captured cfDNA molecules; Wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a sequencing depth at least 2-fold deeper than that of the captured cfDNA molecules of the epigenetic target region set; Obtaining more than one sequence read generated by a nucleic acid sequencer by sequencing the captured cfDNA molecules; Mapping the more than one sequence read to one or more reference sequences to produce mapped sequence reads; And Processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.

2. The non-transitory computer-readable medium according to claim 1, wherein prior to sequencing, the captured cfDNA molecules of the sequence-variable target region set are pooled with the captured cfDNA molecules of the epigenetic target region set.

3. The non-transitory computer-readable medium according to claim 2, wherein the captured cfDNA molecules of the sequence-variable target region set and the captured cfDNA molecules of the epigenetic target region set are sequenced in the same sequencing pool.

4. The non-transitory computer-readable medium according to claim 1, wherein the cfDNA is amplified prior to capture.

5. The non-transitory computer-readable medium according to claim 4, wherein cfDNA amplification includes the step of ligating an adaptor containing a barcode to the cfDNA.

6. The non-transitory computer-readable medium according to claim 1, wherein capturing the more than one set of target regions of the cfDNA includes contacting the cfDNA with a target-binding probe specific for the sequence-variable target region set and a target-binding probe specific for the epigenetic target region set.

7. The non-transitory computer-readable medium according to claim 6, wherein the target-binding probes specific to the set of sequence-variable target regions are present at a concentration that is at least 4-fold or 5-fold higher than the target-binding probes specific to the epigenetic target region set.

8. The non-transitory computer-readable medium according to claim 1, wherein the footprint of the epigenetic target region set is at least 2-fold larger than the size of the set of sequence-variable target regions.

9. The non-transitory computer-readable medium according to claim 8, wherein the footprint of the epigenetic target region set is at least 10-fold larger than the size of the set of sequence-variable target regions.

10. The non-transitory computer-readable medium according to claim 1, wherein the two or more fractions include a hypermethylated fraction and a hypomethylated fraction, and the method further comprises differentially tagging the hypermethylated fraction and the hypomethylated fraction or sequencing the hypermethylated fraction and the hypomethylated fraction separately.

11. The non-transitory computer-readable medium according to claim 10, wherein the hypermethylated fraction and the hypomethylated fraction are differentially tagged, and the method further comprises pooling the differentially tagged hypermethylated and hypomethylated fractions prior to the sequencing step.

12. The non-transitory computer-readable medium according to claim 1, the method further comprising determining whether the cfDNA molecules corresponding to the epigenetic target region set contain or indicate cancer-related epigenetic modifications.

13. The non-transitory computer-readable medium according to claim 12, wherein the cancer-related epigenetic modifications include hypermethylation in one or more hypermethylated variable targets.

14. The non-transitory computer-readable medium according to claim 12, wherein the cancer-related epigenetic modifications include one or more perturbations in CTCF binding.

15. The non-transitory computer-readable medium according to claim 12, wherein the cancer-related epigenetic modifications include one or more perturbations in transcription start sites.

Citation Information

Patent Citations

  • Improvement in lanterns

    US170510A

  • Oligonucleotides

    US20010053519A1

  • Highly sensitive method for the detection of cytosine methylation patters

    US20030082600A1

  • Method and apparatus for imaging a sample on a device

    US20030152490A1

  • Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels

    US20110160078A1